[-] LatheOperator@leminal.space 3 points 2 months ago

That's expensive and a bit impractical to access, isn't it?

[-] LatheOperator@leminal.space 4 points 2 months ago* (last edited 2 months ago)

I want the poisoning to work on file level because it's inevitable and welcome for the documents to be shared between teachers and students for free on all kinds of existing platforms. I don't own any domains, anyway, and it might be best to scatter the documents around to make them harder to blacklist.

Edit: Google has both search crawlers and AI scraping bots. Even if both are separate, easy to filter or even abiding to robots.txt, the company has indicated that opting out of or hindering scraping will impact search ranking. Of course I'd use a throwaway gibberish $1 domain and couldn't care less about search ranking, but the power of their opaque, corporate algorithm is immense and maybe would spread to DNS blocking (they control 8.8.8.8). I don't want to play a cat-and-mouse game (and expect users of the docs to play along).

[-] LatheOperator@leminal.space 8 points 2 months ago* (last edited 2 months ago)

You don't understand just how shit AI is when asked about school topics in Czech. For example, here is a bit of Czech language litany every third grader must know or they will embarrass themselves with awful spelling mistakes. (Skip the bullet points if you just want to hear about the AI)

  • The vowels I and Y (and long versions Í/Ý) sound the same [ɪ] ([ɪː]) unless preceded by D, T, or N but using the wrong one is a big no-no. (Y is never a consonant in Czech)
  • Luckily, in pretty much every native Czech word, I (Í) follows C, J, Č, Ř, Š and Ž, while Y (Ý) follows H, K, R. Consonants Q, W and X basically don't occur and G, Ď, Ť, and Ň are never followed by I or Y. Foreign words are a huge mess of course, as evident by the existence of the Spelling Bee (we don't have that, Czech is phonetic with just a few difficult bits like I/Y).
  • The most difficult are remaining consonants B, F, L, M, P, S, V, Z. They are mostly followed by I (Í) but there is a list of about 15 common exceptions on each (vyjmenovaná slova or BY-FY-LY-MY-PY-SY-VY-ZY words), plus their relative words, where Y (Ý) is written instead. For example, there are just 4 ZY-words so I'll just post the list so you'll get an idea:
    • brzy - early
      • you love exceptions so I put an exception in your exception: brzičko - diminutive of early - is spelled with an I
    • jazyk - tongue/language
      • ... and relative words like jazykolam - tongue twister
        • a well-known one is Strč prst skrz krk, I swear this language is normal
    • nazývat se - be called
      • nazívat se - yawn a lot - also exists for a goddamn reason
        • we have a lot of homonyms for a fully phonetic language, the most common are být - (to) be / bít - (to) beat; my - we / mi - (to) me
    • Ruzyně - Prague quarter where the international airport, until 2012 also called Ruzyně, is located
      • like another part of Prague Výtoň, which has been removed from the lists earlier, nobody cares what the quarter is called now that the airport bears our first president's name instead (he hated flying but it's for the better: the same year, there were efforts to name it after fucking Reagan similar to the former Prague W. Wilson (now Main) train station; RR only got a street), but a set of 4 makes for a nice cadence in reciting the ZY-words so it stays
  • The ends of most words are not governed by spelling but the grammar of declination and conjugation. That's another chapter.

Well, you'd expect AI to know all cca 100 exception words by heart because they're public domain and the most famous piece of third grade teaching material (like times tables in second grade) that barely changed in 100+ years so almost every Czech could recite them as a kid? Hell no. There's dozens of screenshots where Gemini or ChatGPT spewed utter nonsense instead. (DuckDuckGo does not appear to search corporate social media for images because they're not providing direct links to the files). Granted, some are from users asking for nonexistent XY and HY words but so many are unforced errors. I can't find my favorite, a Reddit post where Gemini listed dozens of variants of babička with all kinds of endings like Italian "babičetto" before just adding "etc." but a close second are ones where it adds non-Latin scripts:

Does the apparent incompetence stop Czech students from cheating with AI? Nope. But the longer the AI stays obviously terrible, the better.

[-] LatheOperator@leminal.space 3 points 2 months ago* (last edited 2 months ago)

I know how to edit fonts and replace characters. That would ruin searchability (and screen readers), making the PDF as good as a picture scan without OCR, which I hate (and someone would OCR it sooner or later if they realized the text content is useless). However, common words carry meaning (for example "are" is very different from "are not" etc.) and could be replaced with gibberish without most people searching for them. This also forces plagiators to take more steps.

Anyway, how do I easily add to/edit the PostScript layer in bulk, which consists of a list of individual characters and their positions? As I said, most PDF tools for adding text just add another layer, and that can be easily removed.

[-] LatheOperator@leminal.space 4 points 2 months ago* (last edited 2 months ago)

Yeah, I think that if I strategically replaced "cells" with "little gnomes" or every third "are" with "are not" in the OCR layer, nobody would notice because they're reading the graphical layer and the text is only for searching within the document (they wouldn't be searching for "cells" or "are" in a biology text because it occurs so often). Yes, that would make it hard to plagiarize or listen to the documents but I can live with that.

And the text is in Czech, whose document corpus is not nearly as big as English, a few thousand pages of mild nonsense could make a dent in basic biology knowledge.

The question remains: how? The OCR layer is basically invisible individual characters and coordinates for each, I can't write a PostScript parser from scratch to surgically remove some at the right place and add a few more there, that's outside my scope for the project.

[-] LatheOperator@leminal.space 6 points 2 months ago* (last edited 2 months ago)

Nah, I'm not hosting an entire procedurally generated site that will get blacklisted from search results once Google realizes what I did. I just want PDFs. People will share them around anyway.

119
submitted 2 months ago* (last edited 2 months ago) by LatheOperator@leminal.space to c/fuck_ai@lemmy.world

TL;DR: Please help me fuck (with) AI. See bold sections

Hi,
I haven't been keeping up with anti-AI combat so I'm asking for help. I inherited thousands of pages of materials my late grandpa made or used for his grammar school teaching job in the 1990s-2000s. They are A4 pages of documents made using what seems to be a typewriter, Text602 (DOS rich text editor) and Word. They were most likely not all made by him but he treasured them in nice binding and they have sources (mostly books and journals, almost no webpages, and absolutely no AI) and a cursory look shows meticulous compilation of every important fact on each subject (frankly, the level of detail is excruciating and I'm glad I went to a different grammar school). There's obviously no original scientific research but the materials can still be useful to someone, I bet. They were almost thrown away by the widowed grandma (she already removed and disposed of the plastic bindings and front covers so I'll have to guess document titles) but I think grandpa would prefer them to be shared. With an ADF scanner and OCR software (I have no chance of accessing the work computers he used so I'll have to scan), I can quickly make searchable PDFs of each document, and share them via torrent and DDL sites (there are Czech sites dedicated to sharing teaching materials but they have paywalls or an upload-credit system so best avoid them, not to mention some materials contain newspaper clippings and textbook photocopies for images so best stay anonymous and not try to assert copyright).

I'm afraid these texts could become a major part of some commercial LLM's Czech-language biology/social sciences knowledge corpus unless poisoned. How to best reduce the value of the documents when people try to feed them to AI (training/rewriting) with them while keeping their value for most legitimate users? (Sorry, people with screen readers, there may need to be extra steps for you.) I'm thinking about adding a huge volume of thesaurized or otherwise fuzzed public domain text like f4mi did with .ass subtitles (a technique that would probably still work if she didn't get 1M views detailing it, making YouTube reduce subtitle formatting support). Prompt injection or replacements (cell→gnome) might be interesting too. However, tools I know add an extra PDF layer, which is too obvious. I'm thinking about adding tiny text in the header and footer or between paragraphs in the OCR layer (not overlaid to reduce interference when selecting/searching), but how? I need an automated way to do this with such a huge page count. I can use both Linux and Windows machines for the job. None of them are very powerful but speed is not a concern, it's summer break and nobody will need school materials until September. I'll be happy to include multiple layers and techniques to make them too frustrating to remove.

The paper smells musty but does not seem to be moldy. It's all blank on the other side so I'll interleave it with recent newspaper to allow for the odor-neutralizing chemicals to seep into the sheets so I can eventually reuse them.

Illustration pic is an actual sheet from the collection, to make the post more engaging. Of course I won't be adding watermarks like that, that would just aggrevate people and make them try extra hard to extract the actual content. (And this one is easy to remove with color channel mixing.)

9
submitted 5 months ago* (last edited 5 months ago) by LatheOperator@leminal.space to c/onehundredninetysix@lemmy.blahaj.zone

Real news story. Pic unrelated.

[-] LatheOperator@leminal.space 3 points 1 year ago

The prompt (which I now found and pasted in the body text) includes the entire script so any coherency in the story is not by the image generator. Yes, the script is so stupid it was probably by an LLM.

[-] LatheOperator@leminal.space 4 points 1 year ago

Yeah, this is the stupidest one.

I found the prompt and looks like somebody shat out lots of comics with it, I embedded them in the post

[-] LatheOperator@leminal.space 3 points 1 year ago

The RFK one with autistic people and taxes? How do you "know" it's AI? There are no giveaways either way.

[-] LatheOperator@leminal.space 4 points 1 year ago

It's spelled "Decitions, erctiʋns..." Remember, THE COMPUTER is infallible!

149
submitted 1 year ago* (last edited 1 year ago) by LatheOperator@leminal.space to c/fuck_ai@lemmy.world

TensorArt bragged how their FluxAI tool can generate comics... The results attached to the post speak for themselves. It's just 8 months old, so not an early model.

Transcript
AI-generated comic, mostly black and white with a few shades of grey.

Big panel at top: Fashion store with T-shirts and an undistinct item on display. Two very similar women walking into a wall next to its entrance, as well as a girl with impossible legs walking either left or right. Store sign: "Size leye light Is got neaıl boutique"

Panel 2: Woman in a shirt, cross-eyed: "Decitions, erctiʋns..."

Panel 3: Same girl, looking as if around a corner, with a disfigured disembodied hand holding a bag-like item: "We.let tame it caƨt harrng sishil.. ଚaraa I tine ջur ડeas tmall shopping ppree..

Panel 4: Same girl, looking at one or more large items held in her 4- and 11-fingered hands. Concerningly blank stare, circles around eyes, eyeballs lined with red: "ooh! the bres or ኗ5бo season.

Panel 5: Different girl with no arms and disfigured legs next to clothes racks. Confused look, indistinct symbols above her head: "What!' "thought I has cute shop, top.?"

Wide panel 6 at bottom: Both characters from above looking at a group of schoolgirls in front of badly drawn lockers; one has question marks above her head. The girl from panels 2-4: "Will a hebes. Red!" in a two-tailed speech bubble.

Indistinct signature and initial in the bottom right corner.

[-] LatheOperator@leminal.space 5 points 1 year ago

Yes.

It wouldn't be the first community defined by antagonism to something, there is !fuckcars@lemmy.world, !fucksubscriptions@lemmy.world, !thepoliceproblem@lemmy.world and even !fuck_ai@lemmy.world despite a non-significant number of AI proponents around. The corresponding subreddit r/fuckalegriaart has 57K members. So it's not just me, and hating something is a valid basis for a community. I technically could create it but there are better hands it could be in. So I guess I just should, based on the reactions I see I don't expect much traffic anyway.

126

FYI: Alegria "Art", or the larger "Corporate Memphis" style is flat-color, anatomically-distorted slop associated with corporations and used by Facebook, Google and others.

Yes, I could make a community but I'm not here often enough to moderate one.

view more: next ›

LatheOperator

0 post score
0 comment score
joined 2 years ago