There's a lot of info that you need to know to explore this space, so I'll take my own shot at answering. Let me know if anything needs further explaining!
An "open weight" model is an AI model where you can download the data needed to run the model on your own hardware for free. Contrast this with proprietary models-as-a-service like ChatGPT and Claude where you have no access to the data needed to run the model yourself -- you can only use it through the services provided, usually for a fee, and which can be taken away from you or changed at any time with no recourse.
The mapping to traditional open source terms does not work well since what you get is a binary artifact.
Those artifacts are released with a license -- and many of the models are licensed permissively (e.g. MIT or Apache license terms). You can take those weights, modify them, and then release them as new models -- and people do actually do this in practice!
Is it really feasible to run it 100% locally?
Yes. I run models on my own computers and have tried a number of configurations to figure out what works well. The Qwen family of open weight models (from Alibaba) are the ones I've found most useful so far. Gemma4 models (from Google) are also useful.
I prefer models that have been modified by the community to remove corporate censorship -- i.e. stripping that "As a large language model..." cover-your-ass crap and evasiveness on topics like Tiananmen Square. If that means the model is technically capable of telling me to go kill myself too, so be it; I've spent 25+ years dealing with assholes on the internet and can handle abuse from a stupid robot if I have to. (In practice though, they're usually pretty nice still unless I deliberately tell them to act like an asshole -- and then Qwen, at least, starts to sound like a snarky redditor; it's quite funny most of the time, actually.)
If yes to the previous: the software doesn't come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
Models do require training to create, yes. It's generally not clear where the training physically happened IRL -- so, yes, some of them probably used power from gas turbines, but others may be drawing power from the Three Gorges Dam in China or solar plants or nuclear plants or whatever else is hooked up to the electric grid where the training happened. Most of them are also not very open about the data sets they were trained on. (There are exceptions to this though!) The Chinese models in particular are almost certainly trained heavily on logs extracted from Western models in addition to using whatever other data they could get ahold of. Whether you think that's ethical or not is a matter of perspective; how do you feel about Robin Hood?
Once a model has been trained though, it can run on a normal GPU. The power requirements to run an LLM are basically the same as running a video game, or, equivalently, about the same as turning on a few incandescent lightbulbs. (The iGPU in one of my systems uses 100W; the discrete GPU in another system I've tried uses 215W under load with appropriate tuning -- or 300W if you run it naively.)
If you want to run a model yourself, I recommend using llama.cpp -- there are instructions on how to get started with it here: https://llama.app/
These are the models I've found most useful:
- Qwen3.8-27B (official version): https://huggingface.co/Qwen/Qwen3.8-27B
- Qwen3.8-27B (uncensored): https://huggingface.co/MuXodious/Qwen3.8-27B-absolute-heresy
- Qwen3.6-35B-A3B (official version): https://huggingface.co/Qwen/Qwen3.6-35B-A3B
- Qwen3.6-35B-A3B (uncensored -- ⚠️ mildly NSFW graphics on page): https://huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-heretic
- Gemma4 (official collection -- many variants in different sizes): https://huggingface.co/collections/google/gemma-4
If you have an iGPU only, I recommend using one of the so called "Mixture of Experts" (MoE) releases. These are typically named like 35B-A3B or similar; the first number indicates the total number of weights (35 billion) and the second indicates how many are "active" (i.e. actually used during computation) at one time while the model is running (3 billion in the example). These models need less computation to run and stay fast on weaker GPUs. Qwen3.6-35B-A3B is very good in this space and was my go-to model for a long time.
If you have a discrete GPU and enough VRAM, I recommend using a dense model (i.e. one that activates all its weights while answering) like Qwen3.8-27B.
It's worth noting that people don't usually use the full quality weights (which are typically ~2 bytes per weight); they use a "quantized" version -- compressed in a lossy fashion like a JPEG. Going down to 4-bits (half a byte) on average per weight is about as low as most people like to go -- you will see this indicated in names like Q4_K_M. (Quantized to ~4 bits with the K quantizaation scheme, medium variant.) Usually a bigger number is better in the sense of "closer to the original quality" -- at the cost of needing more RAM.
Full quality weights are often found as safetensor files on HuggingFace. Quantized weights intended for use with llama.cpp are usually in GGUF file format.
Does that help?