36
you are viewing a single comment's thread
view the rest of the comments
[-] mindbleach@sh.itjust.works 1 points 2 months ago

There is no difference. A chatbot that's read every book in the library is the same as a denoiser that's seen Zootopia. Published contents are public... that's what the word means. It's not theft for the robot to know what Superman looks like, and the only way it'd know that is by lookin'.

None of these models are supposed to reproduce a complete work. That's a failure to generalize. It means they could be smaller, and still change new footage to fit a novel description. If you film yourself in a bathrobe and motorcycle helmet, and say "Darth Vader," you won't get a single from of A New Hope. It only removes the fine details that don't match the concept. If you're skateboarding, so is Vader.

[-] omarthemediocre@lemmy.zip 1 points 2 months ago

Yeah, but the robot can be set to recognize only based on the training data (majority of people would be ok if their work is used when training only for recognition purposes) the problem is the production when the same data set is used to produce images/video. Here it matters what the robot has been trained on. Model that never "seen" picture of pikachu will not be able to reproduce pikachu (For example Firefly by Adobe) Model that has pikachu in its training data will reproduce pikachu even if you don't use the name pikachu but just describe what pikachu looks like. And that is only part of the problem with original characters (many being done by small authors who don't have the means to sue companies behind those models) Second issue I see is the AI companies using the training data as a free buffet, even though out of all the sources - electricity, compute, training data the training data are the most valuable - no matter how much electricity and compute power you have, your model cannot exist without training data. Yet no one is paying for those training data anything (again, unless you are big publisher that is able to threat the companies and make a licence deal). All the companies behind those models are for profit companies yet they are using any loophole in the system to pay as little as possible and exploit work of others.

[-] mindbleach@sh.itjust.works 1 points 1 month ago

Generators are recognizers. Like how speakers and microphones are the same mechanism, backwards. ImageNet was only a classifier for a couple thousand single-word labels. Dall-E used the same approach to walk an image toward those labels. That's how you get "avocado chair."

Yet no one is paying for those training data

They don't have to, because training is transformative. It's fair use. If permission is not a factor, why would money be required? Your rights are not a loophole. This is fine for the same reason you're free to reference, parody, or quote commercial works guarded by flesh-eating lawyers.

Even a vegan model trained on bespoke data could reproduce trademark-infringing characters. If you can describe what Naruto looks like, the model only needs to know what anime means.

this post was submitted on 05 Jul 2026
36 points (97.4% liked)

Europe

11936 readers
752 users here now

News and information from Europe ๐Ÿ‡ช๐Ÿ‡บ

(Current banner: La Mancha, Spain. Feel free to post submissions for banner images.)

Rules

  1. This is an English-language community. Comments should be in English. Posts can link to non-English news sources when providing a full-text translation in the post description. Automated translations are fine, as long as they don't overly distort the content.
  2. No links to misinformation or commercial advertising. When you post outdated/historic articles, add the year of publication to the post title. Infographics must include a source and a year of creation; if possible, also provide a link to the source.
  3. Be kind to each other, and argue in good faith. Don't post direct insults nor disrespectful and condescending comments. Don't troll nor incite hatred. Don't look for novel argumentation strategies at Wikipedia's List of fallacies.
  4. No bigotry, sexism, racism, antisemitism, islamophobia, dehumanization of minorities, or glorification of National Socialism. We follow German law; don't question the statehood of Israel.
  5. Be the signal, not the noise: Strive to post insightful comments. Add "/s" when you're being sarcastic (and don't use it to break rule no. 3).
  6. If you link to paywalled information, please provide also a link to a freely available archived version. Alternatively, try to find a different source.
  7. Light-hearted content, memes, and posts about your European everyday belong in other communities.
  8. Don't evade bans. If we notice ban evasion, that will result in a permanent ban for all the accounts we can associate with you.
  9. No posts linking to speculative reporting about ongoing events with unclear backgrounds. Please wait at least 12 hours. (E.g., do not post breathless reporting on an ongoing terror attack.)
  10. Always provide context with posts: Don't post uncontextualized images or videos, and don't start discussions without giving some context first.

(This list may get expanded as necessary.)

Posts that link to the following sources will be removed

Unless they're the only sources, please also avoid The Sun, Daily Mail, any "thinktank" type organization, and non-Lemmy social media (incl. Substack). Don't link to Twitter directly, instead use xcancel.com. For Reddit, use old:reddit:com

(Lists may get expanded as necessary.)

Ban lengths, etc.

We will use some leeway to decide whether to remove a comment.

If need be, there are also bans: 3 days for lighter offenses, 7 or 14 days for bigger offenses, and permanent bans for people who don't show any willingness to participate productively. If we think the ban reason is obvious, we may not specifically write to you.

If you want to protest a removal or ban, feel free to write privately to the admin that applied the rule (check modlog first to find who was it.)

founded 2 years ago
MODERATORS