48

I currently run the top 8 or so per day from r/LocalLLama through my RSS reader. Actually going to reddit is eww, so I want it to be worthwhile. I usually get enough to keep roughly up to date just with title and text / graphs and only go if it looks really interesting.

With the upcoming demise of reddit RSS, I'm wondering if people have recommended (hopefully RSS friendly) places roughly equivalent. I guess I could filter HN which may be useful for many tech interests, but I don't know what I don't know, so I thought I'd pose the question to the crowd.

you are viewing a single comment's thread
view the rest of the comments
[-] EnsignWashout@startrek.website 4 points 1 day ago* (last edited 1 day ago)

I'm not the OP, but I'm trying to figure out if local models are always painfully slow or if I'm missing something obvious in my tuning.

Stable Diffusion can whip up a picture in less time on the same hardware, than lama.cpp takes to decide to call an MCP function.

It seems like I must be missing something in my lama.cpp setup, but none of the guides I've read have clued me in to what I've done wrong.

Ollama performs similarly poorly on the same harsware, so I've probably managed to make the se mistake(s) at least twice.

Anyway, that's the main thing I'm reading along for. Trying to increase my understanding until I catch my own mistakes.

[-] Dran_Arcana@lemmy.world 1 points 7 hours ago

what's your hardware, and what's your launch command?

[-] BeefAndPoultry@lemmus.org 2 points 1 day ago* (last edited 1 day ago)

I have a guide for performance tuning that should be a pretty good start

https://lemmus.org/post/24235317

Let me know if you have questions, or maybe just make a post asking how to optimize for your hardware and I'll try to answer

I would suggest you don't use Ollama https://sleepingrobots.com/dreams/stop-using-ollama/ If you want a GUI, Unsloth Studio is probably best and open source. LM Studio is good too but closed source.

I just use llama.cpp llama-server with the built-in Web UI

Also check the llama.cpp docs

https://github.com/ggml-org/llama.cpp/blob/master/tools/cli/README.md

https://github.com/ggml-org/llama.cpp/blob/master/docs/development/token_generation_performance_tips.md

[-] EnsignWashout@startrek.website 2 points 21 hours ago

I will study these. Thank you!

this post was submitted on 04 Oct 2026
48 points (86.4% liked)

LocalLLaMA

5207 readers
39 users here now

Welcome to LocalLLaMA! Here we discuss running and developing machine learning models at home. Lets explore cutting edge open source neural network technology together.

Get support from the community! Ask questions, share prompts, discuss benchmarks, get hyped at the latest and greatest model releases! Enjoy talking about our awesome hobby.

As ambassadors of the self-hosting machine learning community, we strive to support each other and share our enthusiasm in a positive constructive way.

Rules:

Rule 1 - No harassment or personal character attacks of community members. I.E no namecalling, no generalizing entire groups of people that make up our community, no baseless personal insults.

Rule 2 - No comparing artificial intelligence/machine learning models to cryptocurrency. I.E no comparing the usefulness of models to that of NFTs, no comparing the resource usage required to train a model is anything close to maintaining a blockchain/ mining for crypto, no implying its just a fad/bubble that will leave people with nothing of value when it burst.

Rule 3 - No comparing artificial intelligence/machine learning to simple text prediction algorithms. I.E statements such as "llms are basically just simple text predictions like what your phone keyboard autocorrect uses, and they're still using the same algorithms since <over 10 years ago>.

Rule 4 - No implying that models are devoid of purpose or potential for enriching peoples lives.

founded 3 years ago
MODERATORS