view the rest of the comments
LocalLLaMA
Welcome to LocalLLaMA! Here we discuss running and developing machine learning models at home. Lets explore cutting edge open source neural network technology together.
Get support from the community! Ask questions, share prompts, discuss benchmarks, get hyped at the latest and greatest model releases! Enjoy talking about our awesome hobby.
As ambassadors of the self-hosting machine learning community, we strive to support each other and share our enthusiasm in a positive constructive way.
Rules:
Rule 1 - No harassment or personal character attacks of community members. I.E no namecalling, no generalizing entire groups of people that make up our community, no baseless personal insults.
Rule 2 - No comparing artificial intelligence/machine learning models to cryptocurrency. I.E no comparing the usefulness of models to that of NFTs, no comparing the resource usage required to train a model is anything close to maintaining a blockchain/ mining for crypto, no implying its just a fad/bubble that will leave people with nothing of value when it burst.
Rule 3 - No comparing artificial intelligence/machine learning to simple text prediction algorithms. I.E statements such as "llms are basically just simple text predictions like what your phone keyboard autocorrect uses, and they're still using the same algorithms since <over 10 years ago>.
Rule 4 - No implying that models are devoid of purpose or potential for enriching peoples lives.
what are you looking for? I could start posting more
also not OP but I am interested in local models usable on low-end hardware, models for tasks like document summarizing and text rewording, fuzzy search of various types of local data, and search agents such that I can ask it a vague natural-language question and it runs 17 web searches and comes back with a summary of what it found.
I'm not the OP, but I'm trying to figure out if local models are always painfully slow or if I'm missing something obvious in my tuning.
Stable Diffusion can whip up a picture in less time on the same hardware, than lama.cpp takes to decide to call an MCP function.
It seems like I must be missing something in my lama.cpp setup, but none of the guides I've read have clued me in to what I've done wrong.
Ollama performs similarly poorly on the same harsware, so I've probably managed to make the se mistake(s) at least twice.
Anyway, that's the main thing I'm reading along for. Trying to increase my understanding until I catch my own mistakes.
what's your hardware, and what's your launch command?
I have a guide for performance tuning that should be a pretty good start
https://lemmus.org/post/24235317
Let me know if you have questions, or maybe just make a post asking how to optimize for your hardware and I'll try to answer
I would suggest you don't use Ollama https://sleepingrobots.com/dreams/stop-using-ollama/ If you want a GUI, Unsloth Studio is probably best and open source. LM Studio is good too but closed source.
I just use llama.cpp llama-server with the built-in Web UI
Also check the llama.cpp docs
https://github.com/ggml-org/llama.cpp/blob/master/tools/cli/README.md
https://github.com/ggml-org/llama.cpp/blob/master/docs/development/token_generation_performance_tips.md
I will study these. Thank you!