You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
As a larger experiment in AI-assisted coding, I have a pretty nice implementation of file compression for FORM files, including
sort files
scratch files, including with bracket indexing
store and save files
spectator files
The changes are, roughly,
remove all gzip-related code from the sorting routines
remove compress.c, unixfile.c, all U* functions/macros
add io.c, which provides an interface between the rest of FORM and the physical files
The IO layer in io.c provides transparent compression using the Zstd seekable API. The rest of FORM only needs to know about uncompressed sizes and offsets, never compressed ones. This simplifies things a lot in, eg, sort.c, and made implementing compression for sort, scratch, spectator files rather straightforward. Store files were a bit more complicated and need some extra routines compared to the other files, because FORM goes back and over-writes metadata earlier in the file, which doesn't mix so well with compression.
Currently, unrelated to compression specifically, the new IO layer improves performance over the current code in some situations. This is due to fewer mutexes and more private caches and file positions, to not disrupt the global cache and file position. For example, each thread has a cache specifically for reading RHS expressions from the scratch file, without disrupting the master thread's cache.
I think this can be a nice FORM 5.1 feature. Before cleaning things up and running detailed benchmarks, I would like to open a few discussion topics:
Enabling scratch-file compression by default incurs a ~5-10% performance penalty, but only on rather fast storage. On slow storage, it is a good performance improvement. Should we enable by default? Note, that the situation is the same for the current default-enabled sort file compression.
I would like to make Zlib and Zstd mandatory build requirements (Zlib is still used for Tablebase compression)
Compressed save files imply a new incompatible format. Currently one needs to give Save +compress filename.sav;, otherwise the old format is used. Load transparently loads either format. I have left 32 bytes of "reserved" space in the file header for future use. What else do we need to ensure about a new save format?
I need to benchmark possible regressions more, but it might be good to default-enable bracket indexing. Omitting a bracket index and making heavy use of bracket content on RHSs is an enormous performance penalty and easy to do.
On compress; has new sub-keys to independently control compression for each file type (sort, scratch, store, spectator). The old On compress,gzip,6; is parsed without error, but ignored.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
As a larger experiment in AI-assisted coding, I have a pretty nice implementation of file compression for FORM files, including
The changes are, roughly,
The IO layer in io.c provides transparent compression using the Zstd seekable API. The rest of FORM only needs to know about uncompressed sizes and offsets, never compressed ones. This simplifies things a lot in, eg, sort.c, and made implementing compression for sort, scratch, spectator files rather straightforward. Store files were a bit more complicated and need some extra routines compared to the other files, because FORM goes back and over-writes metadata earlier in the file, which doesn't mix so well with compression.
Currently, unrelated to compression specifically, the new IO layer improves performance over the current code in some situations. This is due to fewer mutexes and more private caches and file positions, to not disrupt the global cache and file position. For example, each thread has a cache specifically for reading RHS expressions from the scratch file, without disrupting the master thread's cache.
I think this can be a nice FORM 5.1 feature. Before cleaning things up and running detailed benchmarks, I would like to open a few discussion topics:
Save +compress filename.sav;, otherwise the old format is used.Loadtransparently loads either format. I have left 32 bytes of "reserved" space in the file header for future use. What else do we need to ensure about a new save format?On compress;has new sub-keys to independently control compression for each file type (sort, scratch, store, spectator). The oldOn compress,gzip,6;is parsed without error, but ignored.Please let me know any thoughts!
All reactions