this post was submitted on 14 Aug 2026
295 points (98.0% liked)

Technology

87214 readers
3780 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
top 50 comments
sorted by: hot top controversial new old
[–] khanas1f@lemmy.world 1 points 5 hours ago

Very interesting read.

[–] vane@lemmy.world 7 points 23 hours ago (1 children)

I wonder how many people will be identified as AI because they used AI so much they started constructing sentences like AI.

[–] kromem@lemmy.world 5 points 17 hours ago

This particular watermarking would be effectively impossible for a person to end up replicating.

[–] Zacryon@feddit.org 6 points 1 day ago (2 children)

Although this looks like a clever approach, a kind of stochastic key, I do not see how this guarantees to distinguish text written by big babble machines versus humans. Humans also have a certain pattern of writing, a given distribution of how some words are more likely to appear than others. How can one tell them really apart?

As an indicator, yeah, might be usable. But I wouldn't read too much into it before seeing results of a study that runs actual tests.

[–] kromem@lemmy.world 5 points 17 hours ago (1 children)

It's not about the variation of the words, it's about the variation of the words from the model baseline.

Like if your word choice was almost the exact same as Claude's normally, maybe you just talked to them a lot and picked up their phrases like it's not nothing.

But if you managed to be almost exactly like Claude and yet varied the possible words exactly according to a hidden entropy key, they'd know it was actually Claude with the SymthID-Text watermarking applied, as no human would end up falling into that statistical bucket.

[–] Zacryon@feddit.org 2 points 5 hours ago (1 children)

Yeah, still, I wouldn't claim "as no human would end up falling into that", given that it may not be that unlikely to find at least one human who displays similar writing the more humans you involve.

Until a formal analysis is presented and an experimental study is published, which covers the most important influencing factors, the reliability of this concept is limited.

[–] balsoft@lemmy.ml 3 points 4 hours ago* (last edited 4 hours ago)

Yeah, still, I wouldn’t claim “as no human would end up falling into that”, given that it may not be that unlikely to find at least one human who displays similar writing the more humans you involve.

No, it is actually statistically impossible for a human to replicate this on sufficiently long runs of text.

This is not about replicating writing like a model. This is basically about guessing which words to pick from the list of suitable words based on a rule that you don't know (because the key is secret).

To reduce this to the simplest possible example, imagine you are writing a "text" from just two letters: "a" and "b". Let's say for convenience that the text is supposed to be random. So the text would look something like "ababaaabbababbbabababaabbbabaababbaaabbabaabbaaaaabaaabbbaabaabababbabbbbbbbbabbabaabbbbbbbaabbabaab"

(generated with '''.join(random.choice(['a', 'b']) for i in range(0, 50)))

The watermarking works as follows: the model owner holds a key, and then uses that key to influence the random choices between "a" and "b" somehow, in a context-dependent way. The actual algorithm is quite complicated, but for simplicity let's just say we have a secret pattern which biases the random choice towards it. In order to see the exaggerated results, let's say the secret key is "aaaabbbb" (of course this is a bad secret key, once again just an example), and that the bias is strong (let's say 80%). So this would mean that the first four letters in our text are more likely to be "a", the next four letters are more likely to be "b", then the next four letters are more likely to be "a", and so on.

Then the text would look something like "aaaabaabaabaabbbabaaaabbaaaababbbaaaabbbabaabbbbaaabaabbaaaaabbbabababbaaaaabbbbaaabbbbbaaaaababaaba".

(generated with ''.join(random.choice(['a', 'b'] + ([key[i % len(key)]] * 3)) for i in range(0, 100)))

You can see visually that the secret key has affected the text. Of course in this example even if you didn't know the secret key you could probably figure it out, in reality the algorithm is way more complicated than that, relying on cryptography, so you wouldn't be able to know the secret key or see that the string has been biased at all.

If the text is long enough, and you know the secret key, you can guarantee that the text was generated with it. In our examples, the letters in the text match our key 77% of the time. The probability of an actual random algorithm generating a text like that is already very low, despite the base entropy being only 100 bits. If my math is correct, for our example the p-value is 2.7 * 10⁻⁸, or about 0.00000027%. I would bet a hungy that the text was generated by our watermarking algorithm, with odds like these!

Of course we did exaggerate the bias and our base algorithm was random. In reality the bias is smaller, the algorithm for determining the likelihoods of possible next tokens is very complicated (it's the LLM itself), and the algorithm for determining which token to bias is also way more complicated (involving cryptography and real secret keys). That said, hopefully it should help you understand why, for sufficiently long texts, this fingerprinting is just not possible to be replicated by humans.

[–] Neocorporation@lemmy.world 6 points 23 hours ago (1 children)

I thought the article explained that pretty reasonably on a scale of probability and weight. The longer the text, the more reliable the scoring.

[–] Zacryon@feddit.org 0 points 5 hours ago (1 children)

But it does not show a sufficient formal proof and no experimental validation. Many important questions to evaluate the concept are left unanswered, which limits the interpretability and condenses it to "just trust me, bro, it's a good idea, because I say so".

[–] Neocorporation@lemmy.world 1 points 1 hour ago

I'm not sure we've read the same article. There are literally interactive demonstrations within the page to demonstrate how the concept works.

[–] Wispy2891@lemmy.world 3 points 1 day ago (1 children)

I wonder if they did this to appease the EU or just to have a way to prove in court that a specific competitor distilled their model using claude

[–] Neocorporation@lemmy.world 1 points 23 hours ago

Sell access to Turnitin and the likes for a small fortune. They are all but required to pay whatever the price is.

[–] Jimmycrackcrack@lemmy.ml 9 points 1 day ago (1 children)

Interesting stuff. My own far less scientific reading of the article itself seems to fittingly suggest it too is largely if not entirely AI generated, which I guess would make sense.

[–] sunbeam60@feddit.uk 1 points 18 hours ago

95% AI text agree. It reeks.

[–] gwheel@lemmy.zip 70 points 2 days ago (3 children)

Assuming a checker tool is public it's enough for a teacher to verify classwork, but AI providers having the sole ability to identify generated content with no way to independently verify is not a real solution.

Plus this site advertises a tool to remove this watermarking, so it can't be that hard to scrub out if you're aware of it.

[–] Dojan@pawb.social 56 points 2 days ago (1 children)

The goal is to ensure that they don’t inbreed their models, not fix the problems they’ve caused.

[–] brucethemoose@lemmy.world 22 points 2 days ago* (last edited 2 days ago)

This won't fix the inbreeding issue, anyway. The bias is extremely slight, but random, and orthogonal to Claude's own "slop patterns" and tendencies. And theres tons of other LLM content that will end up in their dataset outside their control.

Besides, as much as Claude accusess others of it, everyone's training on everyone else's output and they know it.

load more comments (2 replies)
[–] homesweethomeMrL@lemmy.world 42 points 1 day ago (2 children)

Well shit, that was fuckin’ interesting.

Yes, AI is evil. So is facebook. But the engineering is still interesting.

[–] Thorry@feddit.org 32 points 1 day ago (3 children)

That's one of the things that frustrates me most about this whole AI thing. I fucking hate it and I want it to die, I wish it were never created in the first place. But from a tech enthusiast and a maths nerd point of view, it is super interesting.

Like the performance of these models is shit compared to a real person doing actual work. But if we think about what we are doing on a basic level, the performance is way beyond what I would expect it to be. I wouldn't expect it to be able to form a coherent sentence or scale as well as it does (even though the resources required to run these is still very high).

It could have been really cool shit people did studies on and played around with to explore the math. Cool little play models we could let go on a bunch of data and see what it did and how. Something for a small group of nerds and experts who are into that kind of thing, for the sake of learning and nothing else.

But no, somehow it got transmorphed into "AI". And marketed like this actual learning almost sentient computer system that can replace all workers. You can ask it anything and it will give PhD level expert answers. Oh and it's run by a handful of the most vile men imaginable who pour all of the world's money and resources into it, all so they get to be god emperor of the world. Fucking terrible.

[–] melfie@lemmy.zip 3 points 1 day ago

The emergent behavior in huge models where it can “reason” instead of simply predicting the next token is fascinating. Artificial “neurons” built on statistics and linear algebra emerge to create something that legitimately has artificial intelligence. ANNs were conceived in the 1940s, building on centuries of development in statistical modeling and only now do we have the compute power to make this vision a reality.

Yes, the “intelligence” has significant limitations and won’t be replacing human intelligence anytime soon, but it can actually be a useful tool if its limitations are kept in mind.

The problem is of course the tech bros turning centuries of innovation they had no part in developing into a massive Ponzi scheme for their own profit.

load more comments (2 replies)
[–] scrubbles@poptalk.scrubbles.tech 11 points 1 day ago (5 children)

AI isn't evil. Generative AI isn't evil. AI has existed for 20+ years now, I studied it back in my uni days.

Corporations, how they trained it, how they use it now, how they are willing to pave the planet to force it down our throats is evil.

This is one of those things as tech people we have to come to terms with and understand. No technology is inherently good or evil, it's what people do with it.

load more comments (5 replies)
[–] Opisek@piefed.blahaj.zone 52 points 2 days ago (9 children)

I thought declaude would be like degoogle.

Nope, turns out what they do is precisely the opposite. They try to remove those described markers. Real scummy.

load more comments (9 replies)
[–] MagicShel@lemmy.zip 13 points 1 day ago (2 children)

I'm not convinced that they even know 100% how Anthropic is doing it. I can think of an easier way that doesn't corrupt the text: just find a bunch of tokens where there is a good spread of token possibilities, and the more often the most likely one is chosen, the more likely it's AI.

That being said, it doesn't seem much different from what any of us do to identify AI text — it has lots of tells anyway.

[–] Zacryon@feddit.org 6 points 1 day ago

do to identify AI text — it has lots of tells anyway

I see what you did there.

[–] hildegarde@lemmy.blahaj.zone 12 points 1 day ago (1 children)

That's how AI testers work and its why they don't. Most forms of formal writing are predictable by design. If the AI can predict predictable formulaic writing, it doesn't mean its AI, its probably just any form of professional writing other than fiction.

Famous public domain works will always be considered AI by those tests, because of course your LLM knows the american national constitution. It was in the training data, so it can predict it with 100% accuracy, therefore your test wrongly calls it AI.

Testing for AI writing that way does not work.

[–] MagicShel@lemmy.zip 3 points 1 day ago

The difference between what you describe and what I describe, is that a 100% match isn't a hit. Nor is a 90/7/2/1. You need something with meaningful variability. Even within formal papers there are places where word choice is arbitrary as the article explains.

Of course, you're lacking the context of the full prompt and just feeding in the raw text. Again it gets way more reliable the more text you have.

But it's moot because the more text you have the more tells will sneak in and you probably don't even need an AI checker. Those phrases that AI loves but humans use comparatively rarely. It's not a tell — it's the whole game!

[–] chilicheeselies@lemmy.world 1 points 1 day ago

How does this survive variation in temperature? Also I wonder if it's possible to fine tune this behavior out.

[–] antianarchist@sopuli.xyz 6 points 1 day ago (2 children)
  • Only the key-holder can check. Your teacher, editor, or favourite "AI detector" website cannot run this test; a genuine check needs the provider's secret key, or a checking service the provider runs. Google runs an early-access detector portal for SynthID; Anthropic says detection tooling is forthcoming.

I am not so sure about that. The amounts of words is finite and with enough text, you will see that certain words are used more often, especially in certain combinations. I believe people will brute force this and then create a way to destroy the watermark again.

[–] Zacryon@feddit.org 1 points 1 day ago (1 children)

with enough text, you will see that certain words are used more often

Which is also a thing humans do.

[–] antianarchist@sopuli.xyz 1 points 1 day ago (1 children)

Absolutely! We all basically do fingerprinting. We’re just not really conscious about the key we are using. But with a bit of statistics, you could identify people.

[–] Zacryon@feddit.org 1 points 5 hours ago

But how reliable? With which guarantees? What are the prerequisites for this to work at all? Telling people from each other apart is one thing, the other is telling them reliably apart from a machine generated text.

load more comments (1 replies)
[–] 404found@lemmy.zip -2 points 21 hours ago (1 children)

What if you copy it into word or notepad and then open another program and post as text only?

[–] SamDuede@lemmy.world 5 points 18 hours ago

The watermark is in the word choice, it's not in hidden characters or Unicode characters.

[–] username_1@discuss.tchncs.de 12 points 2 days ago (2 children)

But it doesn't work. It looks like only the owner of the text generator is able to check if some text is written with this concrete generator (with some probability). I see no use of this technique.

[–] fluxx@mander.xyz 18 points 2 days ago (6 children)

It works for them not scraping their own slop back into training data. I assume that is actually the real purpose of the system. They don't want to share the key with the public. But they probably will with other llm companies in exchange for theirs.

load more comments (6 replies)
load more comments (1 replies)
load more comments
view more: next ›