47

I currently run the top 8 or so per day from r/LocalLLama through my RSS reader. Actually going to reddit is eww, so I want it to be worthwhile. I usually get enough to keep roughly up to date just with title and text / graphs and only go if it looks really interesting.

With the upcoming demise of reddit RSS, I'm wondering if people have recommended (hopefully RSS friendly) places roughly equivalent. I guess I could filter HN which may be useful for many tech interests, but I don't know what I don't know, so I thought I'd pose the question to the crowd.

you are viewing a single comment's thread
view the rest of the comments
[-] Dran_Arcana@lemmy.world 3 points 1 day ago

Useful resources discovered

https://github.com/Yamz-Labs/kyojin Found in the Kyojin post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/). An ExLlamaV3-based ROCm engine with a quickstart that serves GLM-5.3-Flash / MiMo with an OpenAI-style API. Matters to you as the reference implementation for EXL3 layer-mix quantization and for the MoE-expert offloading pattern, though it is ROCm/Strix-Halo specific and the conversion pipeline is private. Actionable if you have AMD hardware or want to study the quant technique; not directly runnable on your NVIDIA stack as-is.

https://huggingface.co/turboderp/Qwen3.8-Flash-Next-exl3 Found in a Kyojin comment (u/MarkoMarjamaa, https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/?context=3#pdmp0wf). Public EXL3 weights for Qwen3.8 Flash-Next. The comment notes "with exl3 it would be possible to run Q5-quality quant in Q4 memory." Directly relevant to your Qwen3.8-27B/Flash-Next deployments if you adopt ExLlamaV3; immediately actionable for EXL3 users.

https://github.com/zhongkaifu/TensorSharp Found in the 176B-on-3080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/). An open-source inference engine doing MoE-aware tiered scheduling (VRAM/RAM/SSD) so a model far larger than VRAM and RAM stays usable. Relevant if you want to serve a 176B MoE on a 24 GB 3090 via SSD; the model-specific doc page (.../docs/models/qwen38-flash-next.md) has the exact launch path. Actionable, but the author has not published the quant used in his benchmark -- verify before relying on the tok/s.

https://github.com/roofkid/ninfer-4080 Found in the Ninfer-4080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/). A from-scratch engine running ISTA-DASLab Qwen3.8-27B-GSQ at 100k context on a 16 GB 4080, with a runnable script (.../scripts/run-ninfer-4080.bat). Actionable for 16 GB-class Ampere cards; portability to other cards is unconfirmed and the KV-cache quant is a known tradeoff.

https://huggingface.co/aleph-alpha/Kolibri-1 (+ tech report https://aleph-alpha.com/downloads/tech-report.pdf) Found in the Kolibri-1 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/). Official 78.1B / 3.46B-active English-German MoE weights under Apache 2.0 with a 1M-context spec and a full tech report. Actionable for evaluation once a serving path/quants exist; not deployable on your stack today without a backend and quant.

https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ Found in the Ninfer-4080 post (https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwv0fj/i_built_ninfer_4080_for_16gb_class_gpus/). A 3-bit GSQ quant of Qwen3.8-27B that the engine uses. Relevant if you are tracking sub-4-bit quants for 27B-class models; 3-bit quality for agentic work is unverified here -- treat as an extreme-quant data point, not a recommended serving quant.


Interesting anecdotes

https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdpighm u/Toxaris71 reports ~11 tok/s decode / 30 tok/s prefill on Qwen 3.8 Flash IQ3_XXS (8 GB 3070 Ti + 32 GB DDR4), and 32 tok/s decode / 80 tok/s prefill with IQ2_0, saying it is "better quality than Orinth 1.5 35B 4-bit" and much faster than Qwen 3.8 27B IQ4_XS (~2.8 tok/s decode). Interesting as a low-VRAM 3070 Ti comparison point for the Flash-Next/27B quants, but no launch command, backend version, or benchmark harness is given -- treat as anecdotal.

https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdofln4 u/leonbollerup claims ~80 tok/s decode and ~2200 tok/s prefill on the same 16 GB 3080 using Strata, which contradicts the post author's ~11 tok/s decode. Useful as an upper-bound comparison for Flash-Next on 16 GB, but no quant or config is provided -- anecdotal and in direct tension with the post.

https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwwmy1/running_qwen38_flash_next_176b_on_a_16gb_rtx_3080/?context=3#pdq00tm u/Independent_Grade612 reports ~450 prefill / ~30 decode on a 12 GB RTX 3500 Ada + 64 GB RAM using Strata with Swift 1.5 Q2XS. Another single-number 12 GB data point; no reproducible config -- anecdotal.

https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwu6zj/the_rise_of_overfit_inference_engines/?context=3#pdocpvd u/TokenRingAI reports ~7 tok/s on llama.cpp and ~12 on SGLang for Qwen Flash Next on his 2x Xeon Max, versus ~67 tg/s and ~900 pp/s on a custom NUMA engine. Notable as the most concrete performance-gap claim in the overfit-engine thread, supporting the case that narrow engines can dramatically outperform general ones on specific hardware. Still anecdotal (no command/config), but the numbers are specific enough to be worth reproducing.

https://llm.r.homelab.internal/r/LocalLLaMA/comments/1wwl7y6/alephalphakolibri1_hugging_face_78b_parameters/?context=3#pdlfi5y u/FullstackSensei estimates Kolibri-1's training cost at ~$4M (20T tokens on a 768x B300 cluster over 4 weeks at ~$8/GPU-hr). Directionally interesting for understanding training-cost trends, but it is a single commenter's estimate, not a published figure -- treat as a rough back-of-envelope, not a fact.

[-] RandomHugs@lemmy.world 1 points 3 hours ago

Thanks for the share! Appreciate actual, non-toxic discussion about LLM usage on Lemmy.

[-] MalReynolds@slrpnk.net 1 points 16 hours ago

Well, that's super wordy for my taste, but that's the beauty of local, you can change things to your taste. I'll probably go Gemma 4 12B or 31B (Qwen 3.8 27B is great for code but I like Gemma for text munging), and am AMD based, but hermes is probably a good tool.

Did you just mod the Reddit Reading skill to use redlib? Why not just use the skill raw (presumably with the OAuth)?

Thanks again.

[-] Dran_Arcana@lemmy.world 1 points 8 hours ago

the skill I'm using was just one that it made itself when I pointed it at my local instance, sent it along a few tasks, and told it to record its learning in a skill.

[-] MalReynolds@slrpnk.net 1 points 7 hours ago

Cool, been a while since I played with hermes, don't remember that being an option, sounds like a Qwen 27B job. Fun, new things to learn.

this post was submitted on 04 Oct 2026
47 points (86.2% liked)

LocalLLaMA

5207 readers
42 users here now

Welcome to LocalLLaMA! Here we discuss running and developing machine learning models at home. Lets explore cutting edge open source neural network technology together.

Get support from the community! Ask questions, share prompts, discuss benchmarks, get hyped at the latest and greatest model releases! Enjoy talking about our awesome hobby.

As ambassadors of the self-hosting machine learning community, we strive to support each other and share our enthusiasm in a positive constructive way.

Rules:

Rule 1 - No harassment or personal character attacks of community members. I.E no namecalling, no generalizing entire groups of people that make up our community, no baseless personal insults.

Rule 2 - No comparing artificial intelligence/machine learning models to cryptocurrency. I.E no comparing the usefulness of models to that of NFTs, no comparing the resource usage required to train a model is anything close to maintaining a blockchain/ mining for crypto, no implying its just a fad/bubble that will leave people with nothing of value when it burst.

Rule 3 - No comparing artificial intelligence/machine learning to simple text prediction algorithms. I.E statements such as "llms are basically just simple text predictions like what your phone keyboard autocorrect uses, and they're still using the same algorithms since <over 10 years ago>.

Rule 4 - No implying that models are devoid of purpose or potential for enriching peoples lives.

founded 3 years ago
MODERATORS