view the rest of the comments
LocalLLaMA
Welcome to LocalLLaMA! Here we discuss running and developing machine learning models at home. Lets explore cutting edge open source neural network technology together.
Get support from the community! Ask questions, share prompts, discuss benchmarks, get hyped at the latest and greatest model releases! Enjoy talking about our awesome hobby.
As ambassadors of the self-hosting machine learning community, we strive to support each other and share our enthusiasm in a positive constructive way.
Rules:
Rule 1 - No harassment or personal character attacks of community members. I.E no namecalling, no generalizing entire groups of people that make up our community, no baseless personal insults.
Rule 2 - No comparing artificial intelligence/machine learning models to cryptocurrency. I.E no comparing the usefulness of models to that of NFTs, no comparing the resource usage required to train a model is anything close to maintaining a blockchain/ mining for crypto, no implying its just a fad/bubble that will leave people with nothing of value when it burst.
Rule 3 - No comparing artificial intelligence/machine learning to simple text prediction algorithms. I.E statements such as "llms are basically just simple text predictions like what your phone keyboard autocorrect uses, and they're still using the same algorithms since <over 10 years ago>.
Rule 4 - No implying that models are devoid of purpose or potential for enriching peoples lives.
Yeah i managed to acquire two 16gb Nvidia cards and built a Ryzen box around it. With 32gb vram and llamacpp you can do wonders . Maybe slowly.
Anyway the cards alone are over 2k€ nowadays, and they are OLD, so really crazy.
And the power consumption.... My rig is round 500w when operating with full GPUs... So can get quite expensive quickly. Easy to get to 5kw per day with only a few hours of llm running.
I don't know if there's a good way to do it with NVIDIA cards, but with AMD, I can set a power cap. I cap my R9700 to 210W (normally it uses 300W) and still get almost all the decode speed (prefill takes a bit of a hit, but still decent enough) -- plus it's quieter that way.
It is possible with nvidia-smi, I recently found out about. I need to esperiment ..
5 kWh/day is what my current DIY solar PV system produces, as an annual average. I have about enough free roof space to double it. So this sounds good. Presumably, a Halo Strix like system would burn less, at still sufficient tokens/s.