view the rest of the comments
Ask Lemmy
A Fediverse community for open-ended, thought provoking questions
Rules: (interactive)
1) Be nice and; have fun
Doxxing, trolling, sealioning, racism, toxicity and dog-whistling are not welcomed in AskLemmy. Remember what your mother said: if you can't say something nice, don't say anything at all. In addition, the site-wide Lemmy.world terms of service also apply here. Please familiarize yourself with them
2) All posts must end with a '?'
This is sort of like Jeopardy. Please phrase all post titles in the form of a proper question ending with ?
3) No spam
Please do not flood the community with nonsense. Actual suspected spammers will be banned on site. No astroturfing.
4) NSFW is okay, within reason
Just remember to tag posts with either a content warning or a [NSFW] tag. Overtly sexual posts are not allowed, please direct them to either !asklemmyafterdark@lemmy.world or !asklemmynsfw@lemmynsfw.com.
NSFW comments should be restricted to posts tagged [NSFW].
5) This is not a support community.
It is not a place for 'how do I?', type questions.
If you have any questions regarding the site itself or would like to report a community, please direct them to Lemmy.world Support or email info@lemmy.world. For other questions check our partnered communities list, or use the search function.
6) No US Politics.
Please don't post about current US Politics. If you need to do this, try !politicaldiscussion@lemmy.world or !uspolitics@lemmy.world
7) No Hit-and-Run questions.
Please don't delete your post for no apparent reason. If you plan on deleting a question later, say so in the post, or if you feel that you have a good reason to remove it, message a mod beforehand. It's not fair to the ones who took their time to answer, and it's not in the spirit of the community.
8) No Bots.
Posts or comments from bots, LLM's, AIs, Neural Networks, Transformers, or Marvin the Paranoid Android are not welcome in AskLemmy. Real humans only please.
Reminder: The terms of service apply here too.
Partnered Communities:
Logo design credit goes to: tubbadu
Very useful 35B models need (more or less) 8GB of VRAM and 32GB+ CPU RAM to be usable locally. 16GB RAM might work on a lean system with an RTX Nvidia card and an exl3 quantized model.
I know we are in a RAM apocalypse, but pre-apocalypse, that’s a quite reasonable requirement, IMO.
Personally, I run Deepseek V4 07-31 Flash at 19 tokens/second on a desktop with a single RTX 3090 and 128GB CPU RAM, and that’s an extremely capable model. Again, that’s expensive these days, but pre-ram apocalypse, that is not an unreasonable workstation/homelab.
Most open weights LLMs are Chinese. And they are:
Trained on pennies. They have to be, as they simply do not have a sea of GPUs like Big Tech. Training costs for their large models are in the millions or tens of millions; a single steel forge has used more energy than all of those training runs combined.
China relies more on renewables, and I believe the datacenters aren’t so hastily constructed with gas turbines in the middle of cities.
And as of now, they are transitioning away from Nvidia GPUs. Some labs already use Huawei accelerators.
Yep.
This is a huge caveat.
You can avoid this. Nvidia Nemotron models, for example, are trained on completely open datasets you can download and inspect yourself: https://huggingface.co/nvidia
They are very good, but just behind state of the art.
But in practice, the SOTA models most run use private datasets. Lord knows where the Chinese get it from, but given some common quirks between models, at least some data sources are shared (and possibly government provided?)
…However.
I would argue providing the result of the training as Apache licensed weights counts as “fair use,” in the same way non commercial fan works do.
They aren’t making a dime off releasing those weights. I’m not trying to sell anyone anything when I use them. Where is the IP theft if money isn’t changing hands?
Now, the Chinese LLM services they charge for? I have no excuse for that. Once money is on the table, it is definitely IP theft.
Not completely -- it is mostly open, but they use a dozen or so private datasets for things like training on global regulations, minesweeper (for some reason), etc. To their credit, they do indicate this on the model cards, but it's not entirely clear what is in those datasets either.
(I fell for that bit of marketing myself awhile back.)
Money not changing hands is pretty much the definition of theft.
Unlike the definition of heft which is a photo of OP’s mom.
What I'm saying is it's akin to writing a fanfic or making fanart of your favorite franchise. Or getting inspired by a painting you see, and making something similar yourself.
Do that for your personal enjoyment? That's fair use, under the law.
But the moment you start trying to sell it is when you get in legal hot water, and when it's indeed morally problematic.
The scale is different, but I'd argue a similar principle applies: if you use some model trained on public works from a protected IP, but the model and its outputs are not resold, nor profited from, it's not theft. The point is beyond money not changing hands; there's no profit being made from the original author's stuff. They aren't being taken advantage of any more than someone viewing their public stuff for free, or someone creating derivatives from private work.
But all that is off the table the moment profit and distribution is involved.