63
submitted 16 hours ago* (last edited 8 hours ago) by A_norny_mousse@piefed.social to c/asklemmy@lemmy.world

I usually don't even understand the lingo they use. "Open-weighted" is the most recent one, then it usually goes down to specific "models" that everybody is supposed to know about.

These are my thoughts (I will stick to the vague "it" for now, but of course therein lies another question: "and how does all this apply to various specialised AIs"):

  • Is it really feasible to run it 100% locally? I know there's plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying "it would, in theory, be possible to run that locally, therefore your concerns are invalid"?
  • If yes to the previous: the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

If what I wrote above is true, what exactly are people arguing when they say it's still possible to use LLMs ethically or true to FOSS philosophy, because ... ???


edit

Thanks to all who answered.

I guess it's my fault for asking several questions in one, but this thread has attracted exactly the type of people I'm writing about; several even used the term "open-weighted models" without explaining it.

Asking to get arguments explained, I got more arguments instead.

you are viewing a single comment's thread
view the rest of the comments
[-] NoLemurs@lemmy.world 10 points 14 hours ago

It is absolutely possible to run 100% locally, but in practice, at the high end, only quantized models. The full size top end models require hundreds of GB of video ram, and while you can buy that, it's stupidly expensive. Quantized models can often perform nearly as well with a small fraction of the ram, but they do sacrifice a little in precision.

These models (well the good ones) ultimately all trace their origins to what you'd likely consider "stolen" data. Whether that's ethical is debatable. If you're in the "information should be free" camp, there may not be an issue here.

As for the power/environmental impact, for what they do LLMs are actually very low impact per-request. If you're concerned about your personal AI power use, then I hope you never fly in an airplane, and minimize your driving because those are much bigger issues.

It's the scale of use that makes AI an environmental problem, and that's a question about corporate use of AI, not personal use of AI.

[-] pemptago@lemmy.ml 1 points 6 hours ago

As for the power/environmental impact, for what they do LLMs are actually very low impact per-request.

Worth noting that a request is often dozens of requests now that there's "reasoning," even for search. As I understand it, a model will take a question , figure out the context (one request), reform the question so it yields better results (another request), if it's doing a web search there's requests for each result, another to compare, another to check if it answers the original request, if not it loops and does it all over again. So one request is easily, and often, dozens of requests. This is one way Ai companies can say to investors, "see, look at how much usage has increased."

Also, we need to factor in the power to scrape and train each of those models, build the datacenters which is near impossible as these companies are not transparent about it and actively try to obstruct investigations into it. Then there's the redundancy of all these different companies competing and doing roughly the same thing at the same time, as fast as they can, so it's orders of magnitude inefficient energy consuming before it gets its first user prompt.

Comparing it to other assaults on the environment is not only hard to do, but a case of "the worse negates the bad" fallacy.

[-] NoLemurs@lemmy.world 1 points 6 hours ago

I do see what you're saying. You can account for all of these factors and it still turns out that, largely, individual LLM use just doesn't use that much power compared to most things people do day to day. Inference is so cheap that even dozens of requests don't amount to much. I could look up and give you a bunch of numbers, but I don't think that's likely to convince anyone who doesn't do the research themselves. It's so easy to come up with sources that say what you want. I'd encourage you to actually look into this yourself.

Training costs are higher, but you train once and use repeatedly. Right now, total training costs are stupidly high, but that's because we've got an arms race between the frontier labs to spend as much money and compute as they can for truly marginal gains in quality. The solution to that problem isn't for individuals to stop using AI, it's to stop those assholes from wasting so much power.

Individual LLM use is so cheap, that it really isn't worth wasting people's energies thinking about limiting that. Instead of being distracted by attempts to make this an issue of personal responsibility, we should be focusing on what will actually make a difference. We should be focused on supporting policies that lead to systemic change. A carbon tax would change corporate behavior right quick, and not just for AI companies.

[-] A_norny_mousse@piefed.social 1 points 8 hours ago

Thanks. What are quantized models?

[-] NoLemurs@lemmy.world 1 points 7 hours ago* (last edited 6 hours ago)

I'm going to simplify a little here, so don't take this completely at face value.

Models are, quite literally, long series of numbers (called weights). A model might store the weights in 16 bit numbers (that is, 16 binary digits). The size of the model (and how much memory it needs) is determined by how many weights there are, and how many bits each weight takes. You can take a 16 bit model and rework it to use 8 bit, or even 4 bit numbers. The result intuitively behaves a lot like the same model, but with less precision to the weights. That makes the model take way less space in ram, but also makes it more likely for concepts (encoded in the weights) to overlap, which impacts model quality. Often the effect is that fine distinctions get lost.

[-] partial_accumen@lemmy.world 1 points 7 hours ago* (last edited 7 hours ago)

Kind of like "compressed". It takes longer/more effort to run them, to produce the same result as a the same model that has not been quantized, where that non-quantized version would consume significantly more RAM but produce the result faster. You would typically only run the quantized model when you're starved for RAM, which most of us are running LLMs locally.

Think like zipping a file with file compression. It takes less space, but has to be unzipped for you to have usable files again.

[-] HubertManne@piefed.social 2 points 11 hours ago

The power impact was something that in the early days worried me. Looking into it I agree with you to some degree. Like using it instead of a search engine I think is by and large a wash. One prompt will likely take more energy but will give you information that likely would have required searching several times modifying the words and jumping between sites which are rendering all sorts of things. Heck If I booted into a command line and connected to an llm Im almost sure it would be significantly less energy. If you chat for entertainment instead of streaming vidoe also lower energy use. Now I think one thing is in making things. It lets people who otherwise couldn't make pictures and videos and code. In the large majority of cases what is made is going to be disposed even for folks that eventually make something they care to keep around or use. While using software to do the same uses a lot of energy the only people doing it generally where making long lasting things for projects or such. So that is where I question it. Still I will have it make a picture to use in an rpg or such.

this post was submitted on 06 Oct 2026
63 points (87.1% liked)

Ask Lemmy

41734 readers
2410 users here now

A Fediverse community for open-ended, thought provoking questions


Rules: (interactive)


1) Be nice and; have funDoxxing, trolling, sealioning, racism, toxicity and dog-whistling are not welcomed in AskLemmy. Remember what your mother said: if you can't say something nice, don't say anything at all. In addition, the site-wide Lemmy.world terms of service also apply here. Please familiarize yourself with them


2) All posts must end with a '?'This is sort of like Jeopardy. Please phrase all post titles in the form of a proper question ending with ?


3) No spamPlease do not flood the community with nonsense. Actual suspected spammers will be banned on site. No astroturfing.


4) NSFW is okay, within reasonJust remember to tag posts with either a content warning or a [NSFW] tag. Overtly sexual posts are not allowed, please direct them to either !asklemmyafterdark@lemmy.world or !asklemmynsfw@lemmynsfw.com. NSFW comments should be restricted to posts tagged [NSFW].


5) This is not a support community.
It is not a place for 'how do I?', type questions. If you have any questions regarding the site itself or would like to report a community, please direct them to Lemmy.world Support or email info@lemmy.world. For other questions check our partnered communities list, or use the search function.


6) No US Politics.
Please don't post about current US Politics. If you need to do this, try !politicaldiscussion@lemmy.world or !uspolitics@lemmy.world


7) No Hit-and-Run questions.
Please don't delete your post for no apparent reason. If you plan on deleting a question later, say so in the post, or if you feel that you have a good reason to remove it, message a mod beforehand. It's not fair to the ones who took their time to answer, and it's not in the spirit of the community.


8) No Bots.
Posts or comments from bots, LLM's, AIs, Neural Networks, Transformers, or Marvin the Paranoid Android are not welcome in AskLemmy. Real humans only please.


Reminder: The terms of service apply here too.

Partnered Communities:

Tech Support

No Stupid Questions

You Should Know

Reddit

Jokes

Ask Ouija


Logo design credit goes to: tubbadu


founded 3 years ago
MODERATORS