167
submitted 8 months ago* (last edited 8 months ago) by mr_MADAFAKA@lemmy.ml to c/linux_gaming@lemmy.world
you are viewing a single comment's thread
view the rest of the comments
[-] brucethemoose@lemmy.world 5 points 8 months ago* (last edited 8 months ago)

Kinda already done:

https://arxiv.org/abs/2504.07866

Huawei's model splits the experts into 8 groups, routed so that each group always has the same number of experts active. This means that (on an 8 NPU server) intercommunication is minimized and the load is balanced.


There's another big MoE (ERNIE? Don't quote me) that ships with native 2-bit QAT, too. It's basically explicitly made to cram into 8 gaming GPUs.


If you can get good results on gaming cards, then suddenly ordinary gaming hardware, run in parallel, may be quite capable of running the important models

I mean. I can run GLM 4.6 350B at 7 tokens/sec on a single 3090 + Ryzen CPU. With modest token convergence compared to the full model. Most can run GLM air and replace base tier ChatGPT.

Some businesses are already serving models split across cheap GPUs. It can be done, but its not turnkey like it is for NVLink HBM cards.

Honestly the only thing keeping OpenAI in place is name recognition, a timing lead, SEO/convenience and... hype. Basically inertia + anticompetitiveness. The tech to displace them is there, it's just inaccessible and unknown.

this post was submitted on 04 Jan 2026
167 points (95.1% liked)

Linux Gaming

27444 readers
453 users here now

Discussions and news about gaming on the GNU/Linux family of operating systems (including the Steam Deck). Potentially a $HOME away from home for disgruntled /r/linux_gaming denizens of the redditarian demesne.

This page can be subscribed to via RSS.

Original /r/linux_gaming pengwing by uoou.

No memes/shitposts/low-effort posts, please.

Resources

Help:

Launchers/Game Library Managers:

General:

Discord:

IRC:

Matrix:

Telegram:

founded 3 years ago
MODERATORS