view the rest of the comments
LocalLLaMA
Welcome to LocalLLaMA! Here we discuss running and developing machine learning models at home. Lets explore cutting edge open source neural network technology together.
Get support from the community! Ask questions, share prompts, discuss benchmarks, get hyped at the latest and greatest model releases! Enjoy talking about our awesome hobby.
As ambassadors of the self-hosting machine learning community, we strive to support each other and share our enthusiasm in a positive constructive way.
Rules:
Rule 1 - No harassment or personal character attacks of community members. I.E no namecalling, no generalizing entire groups of people that make up our community, no baseless personal insults.
Rule 2 - No comparing artificial intelligence/machine learning models to cryptocurrency. I.E no comparing the usefulness of models to that of NFTs, no comparing the resource usage required to train a model is anything close to maintaining a blockchain/ mining for crypto, no implying its just a fad/bubble that will leave people with nothing of value when it burst.
Rule 3 - No comparing artificial intelligence/machine learning to simple text prediction algorithms. I.E statements such as "llms are basically just simple text predictions like what your phone keyboard autocorrect uses, and they're still using the same algorithms since <over 10 years ago>.
Rule 4 - No implying that models are devoid of purpose or potential for enriching peoples lives.
Qwen 3.8 27b is a beast. Better quality than all the free tier commercial AIs. If you can afford a big enough context window, it can very well assist you as good as paid tier commercial ai.
It's the first model I feel confident to assign a complex coding task and have it actually deliver a final product.
On my hardware it's pretty slow (15t/s) but also quite usable provided you are not in a hurry.
I'm about 3h ingested a complex software and added a full ffmpeg pipeline to deliver video transcoding to the browser with backend side caching and more.
It would have taken me days just to dig into how to use ffmpeg for that taks, and more just to debug it.
Thanks. Looks I need a Strix Halo 128 Gbyte box https://www.compute-market.com/blog/qwen-3-8-27b-local-hardware-guide-2026 but it's hard to justify at current memory prices. Speculating on AI bubble crash, which will bring prices down and/or bring used AMD enterprise DC gear on the used hardware market. Meanwhile, will start small with a 8 GB HBM card and 64 GB RAM box that I already own.
Yeah i managed to acquire two 16gb Nvidia cards and built a Ryzen box around it. With 32gb vram and llamacpp you can do wonders . Maybe slowly.
Anyway the cards alone are over 2k€ nowadays, and they are OLD, so really crazy.
And the power consumption.... My rig is round 500w when operating with full GPUs... So can get quite expensive quickly. Easy to get to 5kw per day with only a few hours of llm running.
I don't know if there's a good way to do it with NVIDIA cards, but with AMD, I can set a power cap. I cap my R9700 to 210W (normally it uses 300W) and still get almost all the decode speed (prefill takes a bit of a hit, but still decent enough) -- plus it's quieter that way.
It is possible with nvidia-smi, I recently found out about. I need to esperiment ..
5 kWh/day is what my current DIY solar PV system produces, as an annual average. I have about enough free roof space to double it. So this sounds good. Presumably, a Halo Strix like system would burn less, at still sufficient tokens/s.
No, that's the wrong hardware for Qwen 3.8-27B. Strix Halo is good for MoE models like Qwen3.6-35B-A3B and some of the Gemma4 variants, but it's slow for dense models like Qwen3.8-27B.
You want a discrete card (or more likely, cards) to run that model. I run it primarily on an AMD R9700 (32GB) card, and spill over to an AMD 7900 XTX (24GB) when I need more context space.
If you're hunting for a GPU right now and want to keep costs down, keep your eyes open for deals (I snagged the 24GB card for $800 US a few months ago when I saw an offer for a refurb. card) and look at unusual combos like 3x 16GB cards if you've got a motherboard that can handle it. If you're seriously considering a 128GB Strix Halo box at current prices, I'd recommend just getting two R9700s instead unless you actually have a non-LLM use for 128GB of RAM.
Thanks, good info. In that case I'll speculate on snagging some used AMD HBM enterprise gear with 64 GB or more.
I recently got a 7900 XTX. It’s about the cheapest 24GB card still available at a reasonably sane price. I’m running Qwen 3.8 27B on llama.cpp with 200k context at 4 bit quant with 700 t/s prefill and 30 t/s decode.
Edit:
I just read through the linked article and the one thing it doesn’t mention is quantized KV cache. I’m running 4 bit on both the model and the cache so I can fit 200k context in 24GB. It works really well and I don’t really notice a difference between my setup and Claude Sonnet 5 that I use at work.
128GB sounds nice for running even larger models, but probably not something I can justify in my budget at today’s prices. My XTX was like $700 after trading in my 3070, so I was able to do that.
Edit 2:
Damn, the cheapest XTX I see new now is $1400. Mine was in the $900s a couple months ago. They’re still going in the $800s used on eBay, which is still better than $1500 for a used 3090.
Absolutely insane prices. This bubble got to go.