38
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
this post was submitted on 16 Sep 2026
38 points (86.5% liked)
Technology
88080 readers
5259 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
Trust me, consumers aren't the ones buying 5k€ displays, nor the nearly 20k€ maxed out Mac Studios.
We don't know what the pricing will be.
Something to consider: All those nVidia GPU servers have discrete GPUs, meaning they need to have both RAM and VRAM, and any time there's need to transfer something between RAM and VRAM, that's overhead. Apple runs everything in a shared pool of memory - which at present is significantly slower than the VRAM on those nVidia GPUs since it's DDR rather than HBM, but if they do what they did with the Ultra line of chips and go even further, e.g glue together 8 chips instead of 2, they might make up a bit of the difference in memory bandwidth. Or they could add HBM to these chips I guess.
And a single nVidia DGX B300 with 2.1 TB VRAM is several hundred thousand. And a single one of those is not enough to run the biggest models. By offering less powerful chips and cheaper memory, they might theoretically be able to offer tons of memory and acceptable, though lower, performance for much less money. Allows security-conscious companies to run high-end LLMs locally. And since they deal in shared memory, you might be able to use CPU instead of GPU if that's more efficient in some specific part of inference, without transferring things between different memory spaces.