53
submitted 11 hours ago* (last edited 3 hours ago) by A_norny_mousse@piefed.social to c/asklemmy@lemmy.world

I usually don't even understand the lingo they use. "Open-weighted" is the most recent one, then it usually goes down to specific "models" that everybody is supposed to know about.

These are my thoughts (I will stick to the vague "it" for now, but of course therein lies another question: "and how does all this apply to various specialised AIs"):

  • Is it really feasible to run it 100% locally? I know there's plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying "it would, in theory, be possible to run that locally, therefore your concerns are invalid"?
  • If yes to the previous: the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

If what I wrote above is true, what exactly are people arguing when they say it's still possible to use LLMs ethically or true to FOSS philosophy, because ... ???


edit

Thanks to all who answered.

I guess it's my fault for asking several questions in one, but this thread has attracted exactly the type of people I'm writing about; several even used the term "open-weighted models" without explaining it.

Asking to get arguments explained, I got more arguments instead.

top 50 comments
sorted by: hot top new old
[-] e0qdk@reddthat.com 2 points 42 minutes ago

There's a lot of info that you need to know to explore this space, so I'll take my own shot at answering. Let me know if anything needs further explaining!

An "open weight" model is an AI model where you can download the data needed to run the model on your own hardware for free. Contrast this with proprietary models-as-a-service like ChatGPT and Claude where you have no access to the data needed to run the model yourself -- you can only use it through the services provided, usually for a fee, and which can be taken away from you or changed at any time with no recourse.

The mapping to traditional open source terms does not work well since what you get is a binary artifact.

Those artifacts are released with a license -- and many of the models are licensed permissively (e.g. MIT or Apache license terms). You can take those weights, modify them, and then release them as new models -- and people do actually do this in practice!

Is it really feasible to run it 100% locally?

Yes. I run models on my own computers and have tried a number of configurations to figure out what works well. The Qwen family of open weight models (from Alibaba) are the ones I've found most useful so far. Gemma4 models (from Google) are also useful.

I prefer models that have been modified by the community to remove corporate censorship -- i.e. stripping that "As a large language model..." cover-your-ass crap and evasiveness on topics like Tiananmen Square. If that means the model is technically capable of telling me to go kill myself too, so be it; I've spent 25+ years dealing with assholes on the internet and can handle abuse from a stupid robot if I have to. (In practice though, they're usually pretty nice still unless I deliberately tell them to act like an asshole -- and then Qwen, at least, starts to sound like a snarky redditor; it's quite funny most of the time, actually.)

If yes to the previous: the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

Models do require training to create, yes. It's generally not clear where the training physically happened IRL -- so, yes, some of them probably used power from gas turbines, but others may be drawing power from the Three Gorges Dam in China or solar plants or nuclear plants or whatever else is hooked up to the electric grid where the training happened. Most of them are also not very open about the data sets they were trained on. (There are exceptions to this though!) The Chinese models in particular are almost certainly trained heavily on logs extracted from Western models in addition to using whatever other data they could get ahold of. Whether you think that's ethical or not is a matter of perspective; how do you feel about Robin Hood?

Once a model has been trained though, it can run on a normal GPU. The power requirements to run an LLM are basically the same as running a video game, or, equivalently, about the same as turning on a few incandescent lightbulbs. (The iGPU in one of my systems uses 100W; the discrete GPU in another system I've tried uses 215W under load with appropriate tuning -- or 300W if you run it naively.)

If you want to run a model yourself, I recommend using llama.cpp -- there are instructions on how to get started with it here: https://llama.app/

These are the models I've found most useful:

If you have an iGPU only, I recommend using one of the so called "Mixture of Experts" (MoE) releases. These are typically named like 35B-A3B or similar; the first number indicates the total number of weights (35 billion) and the second indicates how many are "active" (i.e. actually used during computation) at one time while the model is running (3 billion in the example). These models need less computation to run and stay fast on weaker GPUs. Qwen3.6-35B-A3B is very good in this space and was my go-to model for a long time.

If you have a discrete GPU and enough VRAM, I recommend using a dense model (i.e. one that activates all its weights while answering) like Qwen3.8-27B.

It's worth noting that people don't usually use the full quality weights (which are typically ~2 bytes per weight); they use a "quantized" version -- compressed in a lossy fashion like a JPEG. Going down to 4-bits (half a byte) on average per weight is about as low as most people like to go -- you will see this indicated in names like Q4_K_M. (Quantized to ~4 bits with the K quantizaation scheme, medium variant.) Usually a bigger number is better in the sense of "closer to the original quality" -- at the cost of needing more RAM.

Full quality weights are often found as safetensor files on HuggingFace. Quantized weights intended for use with llama.cpp are usually in GGUF file format.

Does that help?

[-] brucethemoose@lemmy.world 4 points 3 hours ago* (last edited 3 hours ago)

Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?

Very useful 35B models need (more or less) 8GB of VRAM and 32GB+ CPU RAM to be usable locally. 16GB RAM might work on a lean system with an RTX Nvidia card and an exl3 quantized model.

I know we are in a RAM apocalypse, but pre-apocalypse, that’s a quite reasonable requirement, IMO.

Personally, I run Deepseek V4 07-31 Flash at 19 tokens/second on a desktop with a single RTX 3090 and 128GB CPU RAM, and that’s an extremely capable model. Again, that’s expensive these days, but pre-ram apocalypse, that is not an unreasonable workstation/homelab.

If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters…

Most open weights LLMs are Chinese. And they are:

  • Trained on pennies. They have to be, as they simply do not have a sea of GPUs like Big Tech. Training costs for their large models are in the millions or tens of millions; a single steel forge has used more energy than all of those training runs combined.

  • China relies more on renewables, and I believe the datacenters aren’t so hastily constructed with gas turbines in the middle of cities.

  • And as of now, they are transitioning away from Nvidia GPUs. Some labs already use Huawei accelerators.

…and stolen IP and stolen personal data?

Yep.

This is a huge caveat.

You can avoid this. Nvidia Nemotron models, for example, are trained on completely open datasets you can download and inspect yourself: https://huggingface.co/nvidia

They are very good, but just behind state of the art.

But in practice, the SOTA models most run use private datasets. Lord knows where the Chinese get it from, but given some common quirks between models, at least some data sources are shared (and possibly government provided?)

…However.

I would argue providing the result of the training as Apache licensed weights counts as “fair use,” in the same way non commercial fan works do.

They aren’t making a dime off releasing those weights. I’m not trying to sell anyone anything when I use them. Where is the IP theft if money isn’t changing hands?

Now, the Chinese LLM services they charge for? I have no excuse for that. Once money is on the table, it is definitely IP theft.

[-] e0qdk@reddthat.com 3 points 2 hours ago

completely open datasets

Not completely -- it is mostly open, but they use a dozen or so private datasets for things like training on global regulations, minesweeper (for some reason), etc. To their credit, they do indicate this on the model cards, but it's not entirely clear what is in those datasets either.

(I fell for that bit of marketing myself awhile back.)

[-] A_norny_mousse@piefed.social 1 points 3 hours ago

Where is the IP theft if money isn’t changing hands?

Money not changing hands is pretty much the definition of theft.

[-] brucethemoose@lemmy.world 1 points 1 hour ago* (last edited 1 hour ago)

What I'm saying is it's akin to writing a fanfic or making fanart of your favorite franchise. Or getting inspired by a painting you see, and making something similar yourself.

Do that for your personal enjoyment? That's fair use, under the law.

But the moment you start trying to sell it is when you get in legal hot water, and when it's indeed morally problematic.


The scale is different, but I'd argue a similar principle applies: if you use some model trained on public works from a protected IP, but the model and its outputs are not resold, nor profited from, it's not theft. The point is beyond money not changing hands; there's no profit being made from the original author's stuff. They aren't being taken advantage of any more than someone viewing their public stuff for free, or someone creating derivatives from private work.

But all that is off the table the moment profit and distribution is involved.

[-] CluckN@lemmy.world 3 points 2 hours ago

Unlike the definition of heft which is a photo of OP’s mom.

[-] partial_accumen@lemmy.world 12 points 5 hours ago

Is it really feasible to run it 100% locally?

Yes. My Macbook Air (M2) released in 2022 can run many publicly available LLM models. The ouput is not as fast as using a large powerful datacenter, but for my local needs, I'm not in a hurry. I get about 17 to 30 tokens per second speed running 100% locally.

If yes to the previous: the software doesn’t come from nowhere

For Mac users, the interface comes from here . For the specific LLM models that is a separate question for each.

and ultimately still relies on gas-turbine-powered datacenters

DCs generating power on-site is a relatively new phenomon because existing grids are at capacity, so the only way to bring new DCs online is locally generating power at that DC, usually using gas turbines or even worse, diesel generators. Most if not all of the publicly available models for running on your own hardware were built before those gas-turnbine-generating DCs were a thing.

The public models people are running now have existed for a number of years are likely made on regular utility grid power which is whatever that nation and region uses.

and stolen IP and stolen personal data, no?

The Llama LLM is made by Meta, so probably yes for that one. Deepseek is from an AI research lab in China. QWEN is from Chinese company Alibaba. We don't know for sure the inputs that created the Chinese models. US AI companies claim a number of the Chinese models are derived from American LLMs, but I haven't seen (or looked for) proof of these claims.

If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically

Likely they mean because you don't have to pay a large LLM owner in a rent-seeking model to run LLMs, nor does a person's use contribute to further development by those companies.

or true to FOSS philosophy, because … ???

With the open-weight models it doesn't rely on a commercial license to use, and the interfaces can be truly open source.

[-] A_norny_mousse@piefed.social 5 points 3 hours ago

OK but what does open-weight model mean?

[-] partial_accumen@lemmy.world 5 points 3 hours ago* (last edited 3 hours ago)

"Weights" in LLMs are the "final answer" numbers from the results of model training, and these are the engine of the LLM model. Lets use chocolate chip cookies as an analogy.

The chocolate chip cookies are produced with ingredients, a recipe, labor effort to produce the dough, and then cooking energy/effort to bake chocolate chip cookies. In a traditional AI company, the company gets the ingredients, they write their own recipe, do all the dough creation, and then the energy/labor for baking, and charge you money to get chocolate chip cookies.

An open-weights AI company got all the ingredients, wrote their own recipe, labor effort to produce the dough, and instead of charging you for the dough, you get as much uncooked dough as you want forever. Your only task is to take the dough and cook it yourself and you've got free chocolate chip cookies. You can make as many cookies as you want with your own oven. However, you are not given free raw ingredients, nor are your given the recipe to alter it in a way you might like. You can only get the dough for free.

So open-weight AI models (uncooked dough) are LLMs you can use on your own computers (oven) for free and have LLM output ( chocolate chip cookies), but you don't get the training data (ingredients) nor the training parameters (recipe) that built the open-weight model.

[-] A_norny_mousse@piefed.social 1 points 3 hours ago

Thanks, that was properly eli5'd!

[-] Balinares@pawb.social 2 points 3 hours ago

It means the weights (the numeric data that makes up the model) are publicly available to download and use for free.

In some cases there are conditions such as, if you run the model as your business, you need to purchase a license. That's usually the largest, most powerful models, though, not those most people would be able to run at home.

[-] Denjin@feddit.uk 3 points 4 hours ago

Question, what do you do with it and how is the data that underlies the model harvested?

[-] partial_accumen@lemmy.world 2 points 3 hours ago* (last edited 3 hours ago)

Question, what do you do with it

I use it for both fun throwaway stuff as well as productivity tasks for learning.

Example of fun throwaway:

Do you remember that episode of Seinfeld where the Kramer starts using Facebook marketplace and starts buying the most worthless items before being robbed when trying to get a too-good-to-be-true sale? No? Because it never happened, but you can plug that premise into an crafted LLM prompt and it will pop out a whole TV script with in-character dialog for each actor as well as use of popular existing sets.

Foreign language learning:

I'm studying a foreign language and want to interact with just the level of vocabulary and grammar I have knowledge of right now at my level for practice. I can prompt the LLM to limit itself to just what I know now and adjust the conversation level so I can practice. If I ever get stuck, I can ask the LLM to explain the grammar usage or vocabulary choice.

and how is the data that underlies the model harvested?

I'm not sure what you're asking here. Are you asking, for example, how the Deepseek model was trained? If so, I answered that above. If not, can you rephrase your question?

[-] Balinares@pawb.social 1 points 3 hours ago

You run a program like llama.cpp that can use the weights to run the model, which you can then use for whatever you'd use a model for: figuring out tech stuff, coding, etc. It's a bit involved, but there are tools that make it easy to get started. LM Studio for instance.

In some cases, the lab that created the model publishes the dataset that it was trained on, and those are usually made of publicly available data. In most cases, though, the labs don't give details, but the answer likely involves siphoning every web page they could find.

[-] HobbitFoot@thelemmy.club 3 points 4 hours ago

I have a friend who runs his AI locally because he doesn't want to pay for tokens.

[-] iForgotSpells@sopuli.xyz 9 points 7 hours ago* (last edited 7 hours ago)

You can run them locally, yes. There are models that can even run on phones, but usecase is limited. But it can only be considered ethical, if the training data used is listed or ethically sourced IMO.

AI bros on Lemmy will disagree with me, but most open weight models are still trained unethically i.e, theft. Most proponents of LLMs (who I talked to on bsky), who say local models are ethical, don't fucking use it. They're larping on socials about how awesome it is, but none of the ones I talked to are using it in their projects. They mess around, realise it is not as good as the "unethical" options, go right back to Claude

Open weight models Qwen, deepseek, mistral, and the Ollama stuff etc are unethical in normal people's eyes, but "ethical" enough for AI bros.

From what I searched, there are very few that can be considered ethical - Olmo, Apertus, Starcoder(?). But idk anyone who uses these. My friend at IBM said they used Apertus, but it was nowhere near good as ChatGPT, so they no longer use Apertus now. And these models require minimum 6-8 GB VRAM for their lowest parameter model iirc.

Even the open-weight model bros are lobbying to redefine what 'open-source AI' means. That should give you a fair idea about people behind open-weight as well

[-] Balinares@pawb.social 1 points 2 hours ago

Borderline strawman there but I'll bite.

Open weight models trained unethically are unethical. Closed weight models trained unethically and then sold back to you for a profit from gas-powered datacenters funded through Ponzi schemes are substantially more unethical.

From there it's a harm reduction calculus. No, those are never pleasant.

So do you let the closed weight labs conquer the field unopposed just so you can feel better about yourself? That's a valid stance, FWIW, and it's also super easy and convenient because you don't have to do anything. It especially makes sense if you believe it's still possible that AI will just go away on its own. I don't, myself, not anymore, so I encourage the use of open weight models, however grudging, so people don't give money to the closed labs and in the worst case aren't eventually stuck with the maximally unethical options. And we've not even touched on the nightmare labor replacement scenarios that seem every day less unlikely. Am I right? I have no clue. Like you, I'm just trying to make the best choices I can in a world that's gone to shit. I'd recommend dropping holier-than-thou attitude either way, though, because it doesn't help our side. Man.

[-] A_norny_mousse@piefed.social 2 points 4 hours ago* (last edited 4 hours ago)

Just as I suspected...

Thanks for taking the time.

there are very few that can be considered ethical - Olmo, Apertus, Starcoder

This is software meant to be run always and completely locally?

Sorry to whine, but so far nobody has eli5'd what "open-weighted" means, or "model" at that... please?

[-] Dran_Arcana@lemmy.world 4 points 6 hours ago

I would unironically argue that a model primarily trained through distillation of closed frontier models, and then released open-weight with an open-source architecture, becomes "ethical" again.

Something something Robin Hood

[-] klankin@piefed.ca 2 points 5 hours ago

Rob the poor's money from the rich and keep it for your community?

Sounds more like feudal warfare than anything, I can't see any harm to artists being reduced at all

[-] radieschen@slrpnk.net 3 points 5 hours ago

I don't know if 5 year olds are allowed to watch 1.5h videos, but if so: this one has all the pro and con arguments explained nicely and makes a point why it's not a good idea to shame people for using LLMs.

https://youtu.be/y85nqc2zm7M

[-] hash@slrpnk.net 5 points 6 hours ago

I've come round to the idea that I will not use LLMs regardless of how "ethically" they can be sourced because their outputs are harmful in a way that builds up, like a very small dose of poison.

I'm not holding this as a hard stance, just a current analysis of the technology and how it might affect my life in my circumstances. I think we're well into a world where we need to be deciding if some technologies don't suit our lives, because there certainly are even more harmful technologies to come.

It might seem like a romantic or artist's approach, but with so much spiralling out of our control I'm trying to make the active choice to be more human.

In leav­ing progress to the machines, in letting technology go forward on its own terms and selecting from it, with what seems to us exces­sive caution, modesty, or restraint, the limited though completely ade­quate implements of their cul­tures, is it possible that in thus opting not to move “forward” or not only “forward,” these people did in fact succeed in living in hu­man history, with energy, liberty, and grace?

Always Coming Home - Stone Telling Part 3 by Ursula K. Le Guin

load more comments (1 replies)
[-] WolfLink@sh.itjust.works 2 points 5 hours ago* (last edited 4 hours ago)

Is it really feasible to run it 100% locally? I know there's plenty of people with very powerful rigs indeed, but still.

There are small models you can run 100% locally on a mediocre computer or even a phone, and I mean you could be 100% offline and it will still work.

The better local models require a higher end gaming GPU or an ARM Mac with a decent amount of RAM, but nothing too extreme. You could get a computer to run them for about $2000 to $3000.

the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

The small local models typically start with a full infustrial-size model and then they “distill” it to smaller models. The concerns over training data are still valid. I’m not sure how much the concerns over power usage are about training vs running the full industrial size models commercially. Running a small local model uses a tiny fraction of the power it takes to run an industrial size model, and the training is a one-time cost (except these companies never stop training because they want to make next year’s model better and faster).

I do think it’s worth pointing out that running a model locally keeps your conversations with it private. Remember, it can be done 100% offline. So your data won’t be stolen.

[-] x1gma@lemmy.world 2 points 5 hours ago

Is it really feasible to run it 100% locally? I know there's plenty of people with very powerful rigs indeed, but still.

It's absolutely possible, and done more and more. Smaller models exist, and run on sub 8GB VRAM without any issue. If you have a gaming setup, you can run the bigger models without any issue locally. You don't need Astra or Mythos or whatever. As a daily driver for "light" tasks (e.g. summarizing a document) small models perform without a significant difference to frontier cloud models. For bigger tasks (e.g. coding or agentic workflows) you need a bit of beef on your graphics card, but it still runs on regular consumer hardware. Anything above that is not needed for any normal use-cases.

If yes to the previous: the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

This has multiple layers:

  • for running models locally, FOSS solutions exist and are no different to any other piece of software
  • models themselves are usually trained on beefy data centers and questionable data, but alternatives exist, both for training and curated data sources. Second but - those usually underperform, because LLMs need those massive data sets. Limiting the dataset limits the output quality drastically.

If what I wrote above is true, what exactly are people arguing when they say it's still possible to use LLMs ethically or true to FOSS philosophy, because ... ???

Also multiple answers:

  • ...because first there is a massive amount of hardliners both in the pro- and anti-LLM camp, and differentiated and objective views are rare. Ethical software and "true FOSS" are also very heated topics currently.
  • ...because, while possible, objectively non-ethical use of LLMs does happen more (public vibecoding of products, weaponized use in cyber and conventional warfare, use in surveillance, etc.).
  • ...because there are absurd amounts of money circulating in the AI bubble, it's too big to fail, and ethical and for-profit usually do not match.
[-] HeHoXa@lemmy.zip 0 points 3 hours ago* (last edited 3 hours ago)

Ollama for text and multimodal (image to text). It supports structured JSON responses you can tie un to cool flows.

ComfyUI for image/audio/video/3d model generation. Use a docker image because installing and configuring it raw is a nightmare.

Huggingface is a good spot for the models. Ollama has its own library. There are several others.

You can get basic text interactions and small images on a typical office PC. You can make nice enough stuff with a typical gaming PC. More power definitely means better results though

[-] jellyfishhunter@lemmy.world 36 points 11 hours ago

From that perspective local LLMs sound more like classical piracy. Not ethical, not FOSS, but out of the hands of greedy corporates.

[-] nialv7@lemmy.world 18 points 9 hours ago* (last edited 9 hours ago)

they are also doing distillation from the big models from OpenAI & Claude, so open-weight models with similar capability can be available for free and also reduce the big AI companies' ability to profit from stolen data.

[-] iturnedintoanewt@lemmy.world 2 points 4 hours ago

What are the most successful/useful distillations you can run locally?

load more comments (2 replies)
load more comments
view more: next ›
this post was submitted on 06 Oct 2026
53 points (85.3% liked)

Ask Lemmy

41734 readers
1651 users here now

A Fediverse community for open-ended, thought provoking questions


Rules: (interactive)


1) Be nice and; have funDoxxing, trolling, sealioning, racism, toxicity and dog-whistling are not welcomed in AskLemmy. Remember what your mother said: if you can't say something nice, don't say anything at all. In addition, the site-wide Lemmy.world terms of service also apply here. Please familiarize yourself with them


2) All posts must end with a '?'This is sort of like Jeopardy. Please phrase all post titles in the form of a proper question ending with ?


3) No spamPlease do not flood the community with nonsense. Actual suspected spammers will be banned on site. No astroturfing.


4) NSFW is okay, within reasonJust remember to tag posts with either a content warning or a [NSFW] tag. Overtly sexual posts are not allowed, please direct them to either !asklemmyafterdark@lemmy.world or !asklemmynsfw@lemmynsfw.com. NSFW comments should be restricted to posts tagged [NSFW].


5) This is not a support community.
It is not a place for 'how do I?', type questions. If you have any questions regarding the site itself or would like to report a community, please direct them to Lemmy.world Support or email info@lemmy.world. For other questions check our partnered communities list, or use the search function.


6) No US Politics.
Please don't post about current US Politics. If you need to do this, try !politicaldiscussion@lemmy.world or !uspolitics@lemmy.world


7) No Hit-and-Run questions.
Please don't delete your post for no apparent reason. If you plan on deleting a question later, say so in the post, or if you feel that you have a good reason to remove it, message a mod beforehand. It's not fair to the ones who took their time to answer, and it's not in the spirit of the community.


8) No Bots.
Posts or comments from bots, LLM's, AIs, Neural Networks, Transformers, or Marvin the Paranoid Android are not welcome in AskLemmy. Real humans only please.


Reminder: The terms of service apply here too.

Partnered Communities:

Tech Support

No Stupid Questions

You Should Know

Reddit

Jokes

Ask Ouija


Logo design credit goes to: tubbadu


founded 3 years ago
MODERATORS