1069

you are viewing a single comment's thread
view the rest of the comments
[-] tal@lemmy.today 4 points 4 hours ago

It's humorous, but last I checked, the best general-purpose compressors with the highest levels of compression---even lossless, which is probably not what most people think of when they think of neural nets---are neural net based.

Neural net-based compressors are computationally expensive, which is why we don't normally use them for most day-to-day tasks, but they really can produce really small outputs.

I'm going to take the text of the US Constitution and stick it in a text file.

$ wget https://www.gutenberg.org/cache/epub/5/pg5.txt
$ stat -c %s pg5.txt 
48326

Okay, so 48326 bytes.

Let's do lzo. You'd expect a limited amount of compression


LZO is "fast" compression, usually only used where compression speed is really important, like where you want to be compressing stuff that's going to be decompressed once and your bottleneck is throughput to disk:

$ lzop <pg5.txt >pg5.txt.lzo 
$ stat -c %s pg5.txt.lzo
24843

Okay, how about gzip? That's Deflate, an older, but pretty-widely-used general-purpose compression algorithm.

$ gzip <pg5.txt >pg5.txt.gz
$ stat -c %s pg5.txt.gz
16660

Okay, what about LZMA? That's a newer, more-CPU-intensive thing that's probably a good general-purpose choice that'll generally give better compression ratios. It's the kind of thing that I'd probably use in a lot of cases. (Personally, these days, I tend to use pixz, which provides both indexed access for tarballs and parallel compression and decompression, which is important for modern processors.)

$ xz <pg5.txt >pg5.txt.xz
$ stat -c %s pg5.txt.xz 
15488

Okay, now PAQ, a neural-net-based compressor:

$ zpaq a pg5.txt.zpaq a pg5.txt -method 5
$ stat -c %s pg5.txt.zpaq 
13063
[-] BradleyUffner@lemmy.world 4 points 4 hours ago* (last edited 4 hours ago)

"Neutral net based" compression isn't even in the same universe as "compressed to a prompt" via LLM

[-] tal@lemmy.today 1 points 3 hours ago* (last edited 3 hours ago)

It actually is. I mean, it's building a dictionary off of a variety of content ahead-of-time, rather than training it on the specific item in question, but that's not uncommon for non-general-purpose compressors.

I mean, doing so to a (probably short) prompt is (a) lossy (and I gave a lossless example) and (b) lossy to an extreme degree, to where it's probably not incredibly useful option for the kinds of systems that exist today.

But...existing diffusion models aren't actually intended for this, either. I'd bet that you could train a model to do image compression along these lines, with a large dictionary, that could do usable compression along the lines of what is (jokingly) described in the article. Probably have a larger compressed form than what they're thinking of.

EDIT: At one point in time, about over a quarter-century ago now, I went out and banged on a neural net post-processor for JPEG artifacts. The idea here is that JPEG very probably isn't optimally representing the final image, as a human, using their knowledge of what the world looks like, can manually (if time-consumingly) clean these up. I didn't meet with a lot of success; I only wanted to put a small amount of time into it, and I was working with much weaker hardware than people are running neural nets on today. But that generated a pre-existing dictionary, a pre-trained neural net, off a training corpus of uncompressed images. It didn't try to reconstruct the image from scratch, the way something like this would, just clean up artifacts, but it has that same pre-generated neural net approach.

[-] mrmanager@lemmy.today 2 points 4 hours ago

I did not know about PAQ actually, thats cool. :)

this post was submitted on 02 Sep 2026
1069 points (98.5% liked)

Programmer Humor

33083 readers
2047 users here now

Welcome to Programmer Humor!

This is a place where you can post jokes, memes, humor, etc. related to programming!

For sharing awful code theres also Programming Horror.

Rules

founded 3 years ago
MODERATORS