1064

top 50 comments
sorted by: hot top new old
[-] Snapz@lemmy.world 1 points 2 hours ago

♪ Hey, hey, billy, can you

compress the Buick ? ♪

Well, all right, but

he'll probably Pu-ick.

[-] tal@lemmy.today 4 points 3 hours ago

It's humorous, but last I checked, the best general-purpose compressors with the highest levels of compression---even lossless, which is probably not what most people think of when they think of neural nets---are neural net based.

Neural net-based compressors are computationally expensive, which is why we don't normally use them for most day-to-day tasks, but they really can produce really small outputs.

I'm going to take the text of the US Constitution and stick it in a text file.

$ wget https://www.gutenberg.org/cache/epub/5/pg5.txt
$ stat -c %s pg5.txt 
48326

Okay, so 48326 bytes.

Let's do lzo. You'd expect a limited amount of compression


LZO is "fast" compression, usually only used where compression speed is really important, like where you want to be compressing stuff that's going to be decompressed once and your bottleneck is throughput to disk:

$ lzop <pg5.txt >pg5.txt.lzo 
$ stat -c %s pg5.txt.lzo
24843

Okay, how about gzip? That's Deflate, an older, but pretty-widely-used general-purpose compression algorithm.

$ gzip <pg5.txt >pg5.txt.gz
$ stat -c %s pg5.txt.gz
16660

Okay, what about LZMA? That's a newer, more-CPU-intensive thing that's probably a good general-purpose choice that'll generally give better compression ratios. It's the kind of thing that I'd probably use in a lot of cases. (Personally, these days, I tend to use pixz, which provides both indexed access for tarballs and parallel compression and decompression, which is important for modern processors.)

$ xz <pg5.txt >pg5.txt.xz
$ stat -c %s pg5.txt.xz 
15488

Okay, now PAQ, a neural-net-based compressor:

$ zpaq a pg5.txt.zpaq a pg5.txt -method 5
$ stat -c %s pg5.txt.zpaq 
13063
[-] BradleyUffner@lemmy.world 2 points 3 hours ago* (last edited 3 hours ago)

"Neutral net based" compression isn't even in the same universe as "compressed to a prompt" via LLM

[-] tal@lemmy.today 0 points 2 hours ago* (last edited 2 hours ago)

It actually is. I mean, it's building a dictionary off of a variety of content ahead-of-time, rather than training it on the specific item in question, but that's not uncommon for non-general-purpose compressors.

I mean, doing so to a (probably short) prompt is (a) lossy (and I gave a lossless example) and (b) lossy to an extreme degree, to where it's probably not incredibly useful option for the kinds of systems that exist today.

But...existing diffusion models aren't actually intended for this, either. I'd bet that you could train a model to do image compression along these lines, with a large dictionary, that could do usable compression along the lines of what is (jokingly) described in the article. Probably have a larger compressed form than what they're thinking of.

EDIT: At one point in time, about over a quarter-century ago now, I went out and banged on a neural net post-processor for JPEG artifacts. The idea here is that JPEG very probably isn't optimally representing the final image, as a human, using their knowledge of what the world looks like, can manually (if time-consumingly) clean these up. I didn't meet with a lot of success; I only wanted to put a small amount of time into it, and I was working with much weaker hardware than people are running neural nets on today. But that generated a pre-existing dictionary, a pre-trained neural net, off a training corpus of uncompressed images. It didn't try to reconstruct the image from scratch, the way something like this would, just clean up artifacts, but it has that same pre-generated neural net approach.

[-] mrmanager@lemmy.today 2 points 3 hours ago

I did not know about PAQ actually, thats cool. :)

[-] whereitsat@lemmy.zip 13 points 5 hours ago

brilliant satire that critiques all of magazine journalism.

i love going to [insert publication here] and reading another article about 'so and so is ready for their next chapter in life.' the so and so always an uninteresting, overly wealthy fuckwit that hasn't accomplished anything other than going to college and having a wealthy parent.

[-] Jankatarch@lemmy.world 2 points 3 hours ago

Lossy compression, don't mind.

[-] M1k3y@discuss.tchncs.de 11 points 5 hours ago

The sad thing is that this has been possible for decades using convolutional autoencoders, but with LLMs we forgot that AI architectures other than transformers still exist.

[-] CanadaPlus@lemmy.sdf.org 5 points 4 hours ago* (last edited 4 hours ago)

Yeah, it's really awful. With any luck, AI winter will follow AI summer, like usual, and the serious people can come out again. Although, aren't CNNs more of a this century thing? I guess two decades is still decades...

IIRC autoencoders actually produce the same image to within our ability to notice, as well.

[-] gera@feddit.nu 20 points 7 hours ago
[-] richardisaguy@lemmy.world 4 points 5 hours ago

This is not funny, you don't understand how long and how many nights i spend overengineering compression pipelines for family photos and videos...

[-] grrgyle@slrpnk.net 5 points 4 hours ago

I've got mine down to 32 byte string to describe the location, datetime, and people in the photograph, with flags for who's smiling and/or blinking.

[-] Knock_Knock_Lemmy_In@lemmy.world 5 points 4 hours ago

I want to see your porn metadata.

[-] qarbone@lemmy.world 3 points 4 hours ago

Absolutely debauched. Especially when you haven't even bought them coffee yet.

[-] CanadaPlus@lemmy.sdf.org 1 points 4 hours ago

Then I bet your stuff is actually good.

[-] Phoenix3875@lemmy.world 3 points 5 hours ago* (last edited 5 hours ago)

laughs when reading this

cries when paying thousands of dollars for DLSS

[-] Bluewing@lemmy.world 5 points 7 hours ago

New Math and vibe coding. Am I right? (insert canned laughter here).

[-] wallwood@lemmy.ml 1 points 5 hours ago

I suspected everything the moment I read 13-year-old-boy

[-] bitjunkie@lemmy.world 4 points 7 hours ago
[-] dylanTheDeveloper@lemmy.world 44 points 13 hours ago* (last edited 13 hours ago)

This is so sad, Chat GPT generate me an email responding to this article

[-] kamen@lemmy.world 31 points 12 hours ago

Lossless -> lossy -> mindless.

[-] anon_8675309@lemmy.world 5 points 9 hours ago

That’s … redefining the word I would think.

[-] pressanykeynow@lemmy.world 27 points 14 hours ago
[-] sukhmel@programming.dev 15 points 12 hours ago* (last edited 12 hours ago)

That's nice, albeit I want to point out for anyone wondering that this is only conjectured and not guaranteed:

One of the properties that π is conjectured to have is that it is normal, which is to say that its digits are all distributed evenly, with the implication that it is a disjunctive sequence, meaning that all possible finite sequences of digits will be present somewhere in it.

There is no guarantee for any specific sequence to appear in π, but for short chunks chances are better (it's not really a probability, but it's simpler to say and I can't explain in details anyway). That's because (from wiki):

It is widely believed that the (computable) numbers √2, π, and e are normal, but a proof remains elusive.

load more comments (4 replies)
load more comments
view more: next ›
this post was submitted on 02 Sep 2026
1064 points (98.5% liked)

Programmer Humor

33083 readers
1984 users here now

Welcome to Programmer Humor!

This is a place where you can post jokes, memes, humor, etc. related to programming!

For sharing awful code theres also Programming Horror.

Rules

founded 3 years ago
MODERATORS