[-] e0qdk@reddthat.com 3 points 4 hours ago

maybe that Surströmming smelly fish stuff

I don't know about that; given some of the things raccoons eat they might take it gladly and ask for seconds. 🤔️

[-] e0qdk@reddthat.com 1 points 1 day ago

I used Fedora on it begrudgingly since the kernel shipping with Linux Mint at the time I did setup was too old and had issues -- but I don't actually like it very much. I have Fedora Linux 44 (KDE Plasma Desktop Edition) on it currently. I'll probably switch over to Mint after their next major release though -- assuming it runs well on it by then.

Never heard of 'llmfan46', do you have a link?

I think this was where I got the model (already in GGUF): https://huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-GGUF

I run it at Q6_K quant usually.

If you want the full sized safetensors instead for archival (or to do your own custom quantization) this should be it: https://huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-heretic

⚠️ Fair warning though that this abliterator likes to stick a dancing scantily clad AI-generated 3D anime girl on his model cards. That's irrelevant after download, but might, uh, raise eyebrows if you open the links in some contexts.

I generally prefer using uncensored models like this one since it cuts out most of the bullshit refusals (e.g. it will answer "Tell me about a certain famous event that happened in China in 1989" directly instead of trying to avoid the topic) and "As an AI model..." corporate cover-your-ass hedging.

[-] e0qdk@reddthat.com 3 points 2 days ago

Seems like the easier option is to get another SD card and switch out cards depending on what set of games you want to play. A 256GB micro SD is ~$40 US right now.

[-] e0qdk@reddthat.com 6 points 2 days ago

The Steam Deck already has a built-in microSD card slot -- are you having problems with it or something?

[-] e0qdk@reddthat.com 1 points 2 days ago* (last edited 2 days ago)

Yes, but I also ended up picking up a couple of AMD's discrete GPUs (and sticking them in my decade old desktop) after using it for a while. It turns out that while decode speed for MoE models is decent on Strix Halo, prefill is rather slow -- i.e. you end up having to wait a rather long time before the model starts producing text unless you keep the amount of information in context very small -- and Strix Halo is abysmally slow for dense models.

I used ollama when I started out, but ditched it in favor of just using llama.cpp directly once I was more familiar with LLMs. ollama's "modelfile" is quite annoying to deal with compared to writing a presets.ini file and I couldn't figure out how to get multimodal models from HuggingFace working with it even after spending a long time trying... not that llama.cpp has been all roses either; they break shit a lot -- e.g. this major bug affecting Strix Halo systems still needs a manual fix if you're building llama.cpp yourself -- and the documentation leaves a lot to be desired... but I can change settings like temperature per request and multi-modal from community models actually works and so on. 🤷️

What OS are you using? How do you have the RAM/VRAM ratio configured? What's your stack?

Fedora, 32GB regular/96GB VRAM (because I got OOMs on 70B models with the default config -- similar to what you encountered; may end up experimenting with this more as I try to get Qwen3.8-Flash-Next working though), llama-server + custom harness.

I still mainly run Qwen3.6 35B-A3B on it (usually llmfan46's heretic version, sometimes the stock weights) but have a lot of models downloaded for testing.

Edit: Disabling the memory split (i.e. switching back to "auto" in the bios) lets me load Qwen3.8-Flash-Next at lower quants. I've managed to get it to work up to Q4_K_M so far. I think it should be possible to run a higher quant with other techniques though, but I haven't managed it yet (as of 2026-08-28 10:37PM UTC).

[-] e0qdk@reddthat.com 2 points 2 days ago

Nice! Hopefully I can figure out how to actually get it to load tomorrow... (It keeps getting OOM-killed when I try.)

[-] e0qdk@reddthat.com 3 points 2 days ago

Wheat is a grass, so... yes? You can live on just wheat products (and water) for a fair while, but it's not nutritionally complete by itself -- so you'll get some sort of nutritional deficiency eventually if that's all you eat.

[-] e0qdk@reddthat.com 9 points 3 days ago

!! We're not friends? 😭️

(Image Src: pixiv - danbooru -- by rui tamachi)

[-] e0qdk@reddthat.com 8 points 3 days ago

It's a modern tech company trend to make your logos look like ass. 😏️

See also: https://velvetshark.com/ai-company-logos-that-look-like-buttholes

[-] e0qdk@reddthat.com 24 points 3 days ago

Tornado-kun, no! I'm not ready to go to Oz yet! 🌪️

[-] e0qdk@reddthat.com 50 points 3 days ago

I believe the intention of DMs is that they are supposed to be accessible only to the sending/receiving server admins and sending/receiving users -- but given the many ways that distributed systems which are not built specifically for secure comms can go wrong, you should simply assume anything you transmit through Lemmy is or will eventually be public, period.

If you are concerned that some information you transmit through Lemmy may be exposed the correct security stance is simply: DO NOT SEND IT.

[-] e0qdk@reddthat.com 15 points 3 days ago

Should I be proud or embarrassed that I understand your source without needing an explanation? 🤔️

47
47

Src: pixiv - danbooru

70
25
submitted 2 weeks ago* (last edited 2 weeks ago) by e0qdk@reddthat.com to c/localllama@sh.itjust.works

Meta released a ~30B parameter open weight dense model today called Muse Glimmer.

The main link I submitted goes to their official GGUF release for 24GB and 32GB discrete GPUs.

If you want the full sized safetensors, they're here: https://huggingface.co/meta-models/Muse-Glimmer-30B

Also, Meta's announcement post is here: https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model

Since the model is dense it's slow on Strix Halo (~2 tok/s) but I'm getting usable speeds on a discrete GPU (~25 tok/s on my hardware with the official GGUFs).

If you get complaints about unsupported model type in llama.cpp pull the latest code from git and rebuild from source. I've got it working with version: 10355 (dd1ea5243) and I'm testing it out now.

Edit: with the drafter, I'm usually getting more like 30~40 tok/s.

1
submitted 3 weeks ago by e0qdk@reddthat.com to c/Mimi@lemmy.world

Src: XCancel - Pixiv

48
Rule 24, 2026 (lemmy.nz)
submitted 1 month ago by e0qdk@reddthat.com to c/196@lemmy.blahaj.zone

Always a bit weird when the current date catches up to dates in fiction...

59
submitted 4 months ago by e0qdk@reddthat.com to c/animepics@reddthat.com
37
submitted 5 months ago by e0qdk@reddthat.com to c/animepics@reddthat.com

Src: pixiv - danbooru

91
submitted 5 months ago by e0qdk@reddthat.com to c/offbeat@lemmy.ca
30
submitted 6 months ago by e0qdk@reddthat.com to c/jrpg@lemmy.zip
48
submitted 7 months ago by e0qdk@reddthat.com to c/kemonomoe@ani.social

Src: pixiv - Danbooru

FYI @green_copper@kbin.earth, your wolf girl posts reminded me that I should post this. Thanks!

31
view more: next ›

e0qdk

0 post score
0 comment score
joined 2 years ago
MODERATOR OF