57
submitted 12 hours ago* (last edited 4 hours ago) by A_norny_mousse@piefed.social to c/asklemmy@lemmy.world

I usually don't even understand the lingo they use. "Open-weighted" is the most recent one, then it usually goes down to specific "models" that everybody is supposed to know about.

These are my thoughts (I will stick to the vague "it" for now, but of course therein lies another question: "and how does all this apply to various specialised AIs"):

  • Is it really feasible to run it 100% locally? I know there's plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying "it would, in theory, be possible to run that locally, therefore your concerns are invalid"?
  • If yes to the previous: the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

If what I wrote above is true, what exactly are people arguing when they say it's still possible to use LLMs ethically or true to FOSS philosophy, because ... ???


edit

Thanks to all who answered.

I guess it's my fault for asking several questions in one, but this thread has attracted exactly the type of people I'm writing about; several even used the term "open-weighted models" without explaining it.

Asking to get arguments explained, I got more arguments instead.

you are viewing a single comment's thread
view the rest of the comments
[-] e0qdk@reddthat.com 2 points 1 hour ago

There's a lot of info that you need to know to explore this space, so I'll take my own shot at answering. Let me know if anything needs further explaining!

An "open weight" model is an AI model where you can download the data needed to run the model on your own hardware for free. Contrast this with proprietary models-as-a-service like ChatGPT and Claude where you have no access to the data needed to run the model yourself -- you can only use it through the services provided, usually for a fee, and which can be taken away from you or changed at any time with no recourse.

The mapping to traditional open source terms does not work well since what you get is a binary artifact.

Those artifacts are released with a license -- and many of the models are licensed permissively (e.g. MIT or Apache license terms). You can take those weights, modify them, and then release them as new models -- and people do actually do this in practice!

Is it really feasible to run it 100% locally?

Yes. I run models on my own computers and have tried a number of configurations to figure out what works well. The Qwen family of open weight models (from Alibaba) are the ones I've found most useful so far. Gemma4 models (from Google) are also useful.

I prefer models that have been modified by the community to remove corporate censorship -- i.e. stripping that "As a large language model..." cover-your-ass crap and evasiveness on topics like Tiananmen Square. If that means the model is technically capable of telling me to go kill myself too, so be it; I've spent 25+ years dealing with assholes on the internet and can handle abuse from a stupid robot if I have to. (In practice though, they're usually pretty nice still unless I deliberately tell them to act like an asshole -- and then Qwen, at least, starts to sound like a snarky redditor; it's quite funny most of the time, actually.)

If yes to the previous: the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

Models do require training to create, yes. It's generally not clear where the training physically happened IRL -- so, yes, some of them probably used power from gas turbines, but others may be drawing power from the Three Gorges Dam in China or solar plants or nuclear plants or whatever else is hooked up to the electric grid where the training happened. Most of them are also not very open about the data sets they were trained on. (There are exceptions to this though!) The Chinese models in particular are almost certainly trained heavily on logs extracted from Western models in addition to using whatever other data they could get ahold of. Whether you think that's ethical or not is a matter of perspective; how do you feel about Robin Hood?

Once a model has been trained though, it can run on a normal GPU. The power requirements to run an LLM are basically the same as running a video game, or, equivalently, about the same as turning on a few incandescent lightbulbs. (The iGPU in one of my systems uses 100W; the discrete GPU in another system I've tried uses 215W under load with appropriate tuning -- or 300W if you run it naively.)

If you want to run a model yourself, I recommend using llama.cpp -- there are instructions on how to get started with it here: https://llama.app/

These are the models I've found most useful:

If you have an iGPU only, I recommend using one of the so called "Mixture of Experts" (MoE) releases. These are typically named like 35B-A3B or similar; the first number indicates the total number of weights (35 billion) and the second indicates how many are "active" (i.e. actually used during computation) at one time while the model is running (3 billion in the example). These models need less computation to run and stay fast on weaker GPUs. Qwen3.6-35B-A3B is very good in this space and was my go-to model for a long time.

If you have a discrete GPU and enough VRAM, I recommend using a dense model (i.e. one that activates all its weights while answering) like Qwen3.8-27B.

It's worth noting that people don't usually use the full quality weights (which are typically ~2 bytes per weight); they use a "quantized" version -- compressed in a lossy fashion like a JPEG. Going down to 4-bits (half a byte) on average per weight is about as low as most people like to go -- you will see this indicated in names like Q4_K_M. (Quantized to ~4 bits with the K quantizaation scheme, medium variant.) Usually a bigger number is better in the sense of "closer to the original quality" -- at the cost of needing more RAM.

Full quality weights are often found as safetensor files on HuggingFace. Quantized weights intended for use with llama.cpp are usually in GGUF file format.

Does that help?

this post was submitted on 06 Oct 2026
57 points (86.1% liked)

Ask Lemmy

41734 readers
1651 users here now

A Fediverse community for open-ended, thought provoking questions


Rules: (interactive)


1) Be nice and; have funDoxxing, trolling, sealioning, racism, toxicity and dog-whistling are not welcomed in AskLemmy. Remember what your mother said: if you can't say something nice, don't say anything at all. In addition, the site-wide Lemmy.world terms of service also apply here. Please familiarize yourself with them


2) All posts must end with a '?'This is sort of like Jeopardy. Please phrase all post titles in the form of a proper question ending with ?


3) No spamPlease do not flood the community with nonsense. Actual suspected spammers will be banned on site. No astroturfing.


4) NSFW is okay, within reasonJust remember to tag posts with either a content warning or a [NSFW] tag. Overtly sexual posts are not allowed, please direct them to either !asklemmyafterdark@lemmy.world or !asklemmynsfw@lemmynsfw.com. NSFW comments should be restricted to posts tagged [NSFW].


5) This is not a support community.
It is not a place for 'how do I?', type questions. If you have any questions regarding the site itself or would like to report a community, please direct them to Lemmy.world Support or email info@lemmy.world. For other questions check our partnered communities list, or use the search function.


6) No US Politics.
Please don't post about current US Politics. If you need to do this, try !politicaldiscussion@lemmy.world or !uspolitics@lemmy.world


7) No Hit-and-Run questions.
Please don't delete your post for no apparent reason. If you plan on deleting a question later, say so in the post, or if you feel that you have a good reason to remove it, message a mod beforehand. It's not fair to the ones who took their time to answer, and it's not in the spirit of the community.


8) No Bots.
Posts or comments from bots, LLM's, AIs, Neural Networks, Transformers, or Marvin the Paranoid Android are not welcome in AskLemmy. Real humans only please.


Reminder: The terms of service apply here too.

Partnered Communities:

Tech Support

No Stupid Questions

You Should Know

Reddit

Jokes

Ask Ouija


Logo design credit goes to: tubbadu


founded 3 years ago
MODERATORS