Skip to content

Dictionary support #16

Description

@hoelzro

Hello there!

I'm interested in using this module, but I'd really like support for zstd dictionaries. I'm willing to add this myself, but I thought I'd open an issue here first to see if a) it's a feature you'd be interested in merging, and b) what you think the interface should look like. Let me know what you think!

-Rob

Activity

  1. spiritloose commented on Oct 25, 2018

    @spiritloose
    Owner

    @hoelzro

    Hi Rob,

    I agree with you. I'd like to support streaming compression/decompression.

    I have tried to implement like compress_using_dict() function but advanced streaming functions are needed for it.

    Advanced streaming functions are experimental and are documented to "Use them only in association with static linking".

    Compress::Zstd uses static linking now but I'm planning to use dynamic linking because libzstd is already popular library and already included in any major package systems like Homebrew, Ubuntu, and so on.

    I have never used the dictionary compression with production use so far, so I have not decided which way to choose.

    Do you use the dictionary compression?
    Please tell me your usecase.

    Thanks,

  2. hoelzro commented on Oct 25, 2018

    @hoelzro
    Author

    @spiritloose I've only used dictionary compression in experiments to see how much space it would save me; I haven't used it in production yet because this module doesn't support it.

  3. plambert commented on Mar 27, 2019

    @plambert

    Dictionary support would be very useful to me as well. Specifically, I'd like to be able to feed data into an object or function, and then extract the dictionary. Later, I'd like to provide the dictionary and some data to the compress function, and get a result which can be decompressed with the same dictionary.

    This would allow me to compress a lot of small pieces of data while still being able to address them individually, and without the huge overhead of having a new dictionary for each.

    Thanks,

  4. plambert commented on Apr 5, 2019

    @plambert

    Thanks for the update!

    I'd like to test this; it's not clear how I'd go about accomplishing my use case:

    1. Create a dictionary from 1,000-10,000 in-memory strings (typically around 80-1000 bytes each).
    2. Write the dictionary to a database.
    3. Compress a series of small strings (also about 80-1000 bytes each) and write them to the database.

    Then, to decompress:

    1. Read the dictionary from the database.
    2. Read each compressed string from the database and decompress with the dictionary.

    I apologize if this should be obvious to me; I've looked at the source and while I see where it's now possible to pass a dictionary to compression and decompression routines, it's not clear how to train a dictionary.

    Thanks for your help!

  5. epa commented on Jan 8, 2024

    @epa

    May I suggest that decompression using an existing dictionary might be easier to implement than compression or dictionary-building, and so that might be the thing to add first?

  6. plambert commented on Jan 23, 2024

    @plambert

    I don't have an existing dictionary and compressed strings to use. I suppose I could dump the strings to use to build the dictionary to a lot of temp files and use the zstd command line tool to create the dictionary. Then compress the strings with the zstd command line tool, and finally decompress them with Compress::Zstd. Would that be something to characterize as "easier?" Maybe using zstd to generate the dictionary would be. Maybe I could use Compress::Zstd to compress the strings, then decompress them, to prove it works.

    Obviously this hasn't been a high priority. I'm still interested in it though.

  7. epa commented on Jan 23, 2024

    @epa

    Hi @plambert, sorry I wasn't clear. Adding decompression support only, without implementing support for compression using a dictionary, wouldn't help your use case. I only meant it might be easier to implement. And it would help my use case, where I have a fixed dictionary I prepared as a one-off; so apologies for squatting on your feature request.

  8. scottchiefbaker commented on Apr 11, 2025

    @scottchiefbaker

    I could definitely use dictionary support for both compression and decompression. I have an existing dictionary that I'd like to feed this module. I'm seeing MASSIVE savings on cpantesters text output. Compressing 32K down to 2.1K with a dictionary is awesome.

  9. scottchiefbaker commented on Apr 12, 2025

    @scottchiefbaker

    Since dictionary support isn't available I'm shelling out... Kinda ghetto but it works.

    use IPC::Open3;
    use Symbol qw(gensym);
    
    sub zstd_comp_with_dict {
        my ($str, $dict_file) = @_;
    
        my $cmd = "/usr/bin/zstd -q -D $dict_file -o /dev/stdout";
        my @cmd = split(' ', $cmd);
    
        # Open the command with various file handles
        my $pid = open3(my $chld_in, my $chld_out, my $chld_err = gensym, @cmd);
    
        # Write to the STDIN of the process
        print $chld_in $str;
        close($chld_in);
    
        # Read the STDOUT from the process
        local $/ = undef; # Input rec separator (slurp)
        my $ret  = readline($chld_out);
    
        waitpid($pid, 0);
    
        return $ret;
    }
  10. scottchiefbaker commented on May 20, 2025

    @scottchiefbaker

    FWIW this library does support compression with dictionaries:

    use Compress::Zstd::CompressionContext;
    use Compress::Zstd::CompressionDictionary;
    
    my $cdict          = Compress::Zstd::CompressionDictionary->new_from_file($filename, $level);
    my $cctx           = Compress::Zstd::CompressionContext->new;
    my $compressed_str = $cctx->compress_using_dict($raw_str, $cdict);
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions