this post was submitted on 28 Jul 2026
100 points (90.3% liked)

Technology

43001 readers
199 users here now

This is the official technology community of Lemmy.ml for all news related to creation and use of technology, and to facilitate civil, meaningful discussion around it.


Ask in DM before posting product reviews or ads. All such posts otherwise are subject to removal.


Rules:

1: All Lemmy rules apply

2: Do not post low effort posts

3: NEVER post naziped*gore stuff

4: Always post article URLs or their archived version URLs as sources, NOT screenshots. Help the blind users.

5: personal rants of Big Tech CEOs like Elon Musk are unwelcome (does not include posts about their companies affecting wide range of people)

6: no advertisement posts unless verified as legitimate and non-exploitative/non-consumerist

7: crypto related posts, unless essential, are disallowed

founded 7 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[โ€“] leanleft@lemmy.ml 4 points 16 hours ago* (last edited 16 hours ago) (1 children)

according to performance on standard benchmark. somewhat covered by the controversy surrounding the term: benchmaxing.
if you see all benefit as a linear one dimensional height on a bar graph..
its almost like you assume that the previous model gave the same exact answer(same style) and the new mode gave the same exact answer PLUS additional useful information.
it might be convenient if measuring progress was so simple. but unfortunately/fortunately , it is not so simple . the most important benchmark are the comparison of outcomes on the problems that YOU have & prompts that YOU can(will) write. nothing else matters for YOU.

  • i admit benchmarks are well designed to objectively measure competence on challenging problems that require skill and really need only ONE correct answer.
[โ€“] yogthos@lemmy.ml 2 points 15 hours ago

Sure, a benchmark doesn't capture all the subtleties and different use cases, but it does give a general idea of the capabilities of a model. Obviously, you have to run the model and see if it does what you need. But the chart isn't really about the nuance, it's showing how drastically the efficiency of the models has improved in just a year. The fact that we can even reasonably compare a model you can run on a desktop to one that needed a data center just a year ago is phenomenal.