top of page

Benchmarks and beyond: measuring progress in LLMs.

4 days ago
2 min read

Updated: 2 days ago

Original paper 1: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation (ICML 2026)

Authors: Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vilem Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Šuppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Pawar, Robert Scholz, Arjun Subramonian, Jingwei Ni, Mykel J. Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, Irene Solaiman



Original paper 2: The Benchmark Trap: Structures of Power and Injustice in AI Evaluations (AIES 2026)

Authors: Jason Branford, Angelie Kraft



These papers were recommended by the researcher

Dr. Vinutha Magal Shreenath, Women in AI (WAI). Find out more here: https://www.linkedin.com/in/vinuthams/


What is noteworthy about this research?

Paper 1 defines the problems of benchmark saturation as the loss of reliable discriminative power among top-performing models under comparison and identifies benchmark properties usually associated with saturation after an analysis of 60 benchmarks, and makes concrete anecdotally known characteristics of benchmarks, with implications for benchmark lifecycle management.


Paper 2 questions the benchmark culture itself, and how this is perpetuating structural harms towards less resourced stakeholders. They indicate that benchmarks are not used by a single community for a single purpose, and serve to identify researchers, engineers, decision-makers, operators, and regulators as relevant stakeholders who rely on evaluations in different ways and contexts. Additionally,  with various actors involved in the chain of construction of benchmarks, and how gamified benchmarks are, they cannot be treated as a narrow methodological concern internal to machine learning.


Comments

Benchmarks have been incredibly important to the progress of machine learning. The paradigm of LLMs however is distinct from that of classical models, which adhered to theory, explanation and extent of model behavior that could therefore be fully validated. In the current AI landscape, the separation between model behavior, the environment the model operates in, and simply bad practices are not well-known/opaque, giving plenty of space for conjectures and claims.


Right now progress is measured in beating benchmarks. In a fragmented evaluation landscape itself, marginalization of lesser resourced interests are deeply problematic. The incentives for orgs that build the models are to productise, to maximize token usage, and claim as much capability as possible. But the trajectory towards true progress needs to bend towards baseline of predictable, verifiable behavior, where concerns of responsibility, safety and trust are addressed foremost, by truly independent evaluators.


Read More

Access the full paper 2: https://arxiv.org/pdf/2608.15326


Have research to share or recommend? https://forms.gle/kaTJEd3gfJV8vtDx8 






Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.
bottom of page