I thought of sharing some issues that preoccupy me here, because of how active piefed development is. Btw thank you all for your hard work!
I have noticed that more and more articles are created with LLMs without disclosing it. So, it seems to me that if someone wants to avoid posting this sort of content, one needs to at least:
- check how many articles the author posts per day in the specific site,
- then if the author really exist and
- finally copy-paste part of the text in a couple of ai-detector sites.
Initially, I thought of making a post for a feature request like the one that detects AI generated images, but for text. But I can't because if I got this right, the ai-detectors may flag an article as ai generated when the author is not a native english speaker [1].
Of course the feature that allows us to label ai-generated content ourselves before posting it is very important. In a way my questions are related to something else: what to do before that. As briefly as I can, here they are:
- does the process mentioned above seem adequate?
- if yes, what else can someone do to check an article before posting, and are there any alternatives/variations to this process?
- if no, what would you suggest?
Sadly, these AI detectors are notoriously unreliable. To the point they're completely useless for automatic content filtering in bulk. Also rarely Free and Open-Source. Even the "ai-detector sites" I tried aren't really all that great. Though I welcome suggestions in case someone knows a good (free) one.
Some AI companies said they're doing watermarking. And I think that'd solve the issue. But then they probably won't disclose how the watermarking works, so we're not going to be able to use that one, either.
I don't see any substantially better solution, than what Rimu already implemented...
It'd be great if we had that, though. And I guess we'd find some people to implement it if someone comes up with a feasible solution.
(Preferably a library(?) or maybe a scientific paper or an algorithm / NLP / machine learning approach we can implement ourselves. Main point: It has to have a low false positive rate, ideally a good detection rate. Has to run on whatever we can afford to run it on. And if you ask me, we better not use any big tech cloud services, like Microsoft Azure AI detector ๐ )
Fun fact: I used Libre Office to write an essay for school. I then saved as docx and uploaded it to my class page, the teacher had formatting weirdness so she asked me to resubmit. So I copy pasted the text into word for web, resubmitted... And promptly got called in for submitting AI. Her rationale? "No edits, no version history..." I had the email proof. What a moron.
That's messed up. I find people staring at my version history a major blocker from writing anything. Unless it's the git history that I explicitly publish, or the classes that I teach where every step is designed to be followed through, my writing or coding process has always been a highly private, personal thing.
Lol. Yeah. All that spying on students isn't a great solution anyway. I'm glad no one knew how many assignments I typed in on deadline day at 1:30am. And in university I tried to typeset the larger homework assignments in LaTeX. With minimal metadata attached to the PDFs.