this post was submitted on 28 Jul 2026
116 points (89.7% liked)

Technology

43007 readers
138 users here now

This is the official technology community of Lemmy.ml for all news related to creation and use of technology, and to facilitate civil, meaningful discussion around it.


Ask in DM before posting product reviews or ads. All such posts otherwise are subject to removal.


Rules:

1: All Lemmy rules apply

2: Do not post low effort posts

3: NEVER post naziped*gore stuff

4: Always post article URLs or their archived version URLs as sources, NOT screenshots. Help the blind users.

5: personal rants of Big Tech CEOs like Elon Musk are unwelcome (does not include posts about their companies affecting wide range of people)

6: no advertisement posts unless verified as legitimate and non-exploitative/non-consumerist

7: crypto related posts, unless essential, are disallowed

founded 7 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] yogthos@lemmy.ml 3 points 1 day ago (1 children)

Sure, a benchmark doesn't capture all the subtleties and different use cases, but it does give a general idea of the capabilities of a model. Obviously, you have to run the model and see if it does what you need. But the chart isn't really about the nuance, it's showing how drastically the efficiency of the models has improved in just a year. The fact that we can even reasonably compare a model you can run on a desktop to one that needed a data center just a year ago is phenomenal.

[–] leanleft@lemmy.ml 1 points 15 hours ago (1 children)

i see what your saying. i didnt mean to discredit standard benchmarks entirely.
i guess its obvious that it measures capability regardless of imprecision.
2 major proposed changes:
**first, i dont really know. aside from saying "benchmark your own prompt+usecase"
a proposed plan:

  • approach one: pay attention and credit new or improved architecture designs and research.
  • approach two: spend more attention on benchmarks. especially specific benchmarks ( that are not focused with industrial domain tasks.) **domain task pursuit, is useful!.. but it depends on if your interest align to popular domains.
  • approach three: if willing to utilize remotely hosted models. rating should also take in consideration.. tools and everything else: websearch performance, RAG performance, smooth interface, pref/balance between speed vs comprehensiveness, cost (if relevant), etc.. .
[–] yogthos@lemmy.ml 1 points 15 hours ago

Honestly, I think the most reasonable approach is just to see what other people's experience is like and which models are well regarded, then try them out and see which one is the best fit for what you're doing. You might not even need the top performing one necessarily, and speed or lower resource usage might be a bigger factor.