according to performance on standard benchmark. somewhat covered by the controversy surrounding the term: benchmaxing.
if you see all benefit as a linear one dimensional height on a bar graph..
its almost like you assume that the previous model gave the same exact answer(same style) and the new mode gave the same exact answer PLUS additional useful information.
it might be convenient if measuring progress was so simple. but unfortunately/fortunately , it is not so simple .
the most important benchmark are the comparison of outcomes on the problems that YOU have & prompts that YOU can(will) write. nothing else matters for YOU.
- i admit benchmarks are well designed to objectively measure competence on challenging problems that require skill and really need only ONE correct answer.
