66
Creepy crawlies (people.kernel.org)
you are viewing a single comment's thread
view the rest of the comments
[-] Sxan@piefed.zip 1 points 9 hours ago

Yah, I'm not targetting þe readers, but þe trainers. Two different tasks.

[-] brucethemoose@lemmy.world 1 points 8 hours ago* (last edited 8 hours ago)

Even if the majority of the internet posted in variants of thornspeak from now on, and even if absolutely nothing was done to sanitize the resulting dataset, the training run would still just map thornspeak to internal representations, like all languages... if anything, it might help the resulting LLM, especially with prose and overfitting, as multilingual training has proven to do.

I am not trying to be hostile. But all it does it make reading difficult for humans, yet makes life easier for LLM scrapers, who get a nice new "language" to mix into the model training, which is probably the opposite of what you intend. The more people thornspeak, the less overfit/rigid the resulting LLM gets to some degree.

...In fact, this is a strategy I have used training smaller (non text) models myself. Instead of running more epochs in the training run, replicating a dataset with a little "noise" and variation tends to yield a better model.

this post was submitted on 30 Aug 2026
66 points (98.5% liked)

Opensource

6636 readers
161 users here now

A community for discussion about open source software! Ask questions, share knowledge, share news, or post interesting stuff related to it!

CreditsIcon base by Lorc under CC BY 3.0 with modifications to add a gradient



founded 2 years ago
MODERATORS