# Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
> Ficha editorial pública de Research IA. Estado: Lectura primaria completa. La interpretación editorial no sustituye la fuente primaria.

- Página canónica: https://luiseduardodemiguel.com/research-ia/papers/beyond-static-leaderboards-predictive-validity-for-the-evaluation-of-llm
- Fuente primaria: https://arxiv.org/abs/2606.19704
- Versión leída: v1
- Fuente comprobada: 2026-08-19 · lectura primaria completa; extracción editorial automatizada, revisión humana pendiente
- Autores: Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin, Yusheng Li, Tianjun Feng, Chun-Yi Tsai, Yihan Sun, Wei Alexander Xin, Akshat Bhandari, Tanisha Rathod, Aaron Fan, Sanskruti Vijay Shejwal, Tomas Pasiecznik, Sagar Chethan Kumar, Tanmay Agarwal, Rohith Kanathur, Sam Colman, Amaan Sheikh, Dev Bahl, Ann Li, Krish Veera, Alimurtaza Mustafa Merchant, Shambhawi Baswaraj Bhure, Sajal Kumar Goyla, Chengrui Li, Kirthana Natarajan, Rui Li, Thomas Ajai, Rujing Li, Vivek G. Iyer, Sanjaii Vijayakumar, Yitong Bai, Ayal Yakobe, Darief Maes, Yassine Jebbouri, Tianyang Xu, Thai Quoc On, Vera Mazeeva, Winston Li, Yuval Shemla, Yeshitha Bhuvanesh, Rushin Bhatt, Siddharth Chethan Gowda, Alisha Vinod, Caroline Cahill, Shriya Aishani Rachakonda, Yunfeng Chen, Aryaman Agrawal, Aman Upganlawar, Mao Le Jonathan Ang, Yubin Sally Go, Madhav Rajkondawar, Yang-Jung Chen, Trisha Maturi, Ananya Kapoor, Andrew Li, Shrey Arora, Mana Abbaszadeh, Shen Li, Charles Xu, Byeolah Kwon
- Fecha del corte: 18 JUNIO 2026.
- Área: AGENTES · EVALUACIÓN

## Tesis y contexto

Evalúa si los benchmarks de agentes predicen desempeño fuera de la muestra, no solo rankings estáticos.

- Problema: Los leaderboards pueden ordenar modelos sin decir si generalizan a tareas futuras o dominios reales.
- Por qué importa: Para elegir agentes en producción, la validez predictiva vale más que una puntuación vistosa.

## Evidencia reportada

- **reported-result**: Figure 3 shows the headline improvement each team reports against its own baseline. [localizador](https://arxiv.org/html/2606.19704#S4)
- **reported-result**: Annotations state what was held constant or co-improved (“quality preserved”, accuracy gain, etc.). [localizador](https://arxiv.org/html/2606.19704#S4)

## Lectura y límite

- Método: La lectura de 2 The Argument describe la intervención y su construcción: The position rests on three structural critiques of aggregate-score leaderboards; the transferability assumption, judge reflexivity, and the circularity through which evaluation constitutes the capability it measures ( 13 ) ; developed below. The core argument: a Pass 1 score of 0.75 can be achieved by many qualitatively different configurations; one that is reasoning-heavy and cost-expensive, one that is retrieval-rich and latency-bound, one that is tool-hygiene-fragile but artifact-reuse-efficient. Aggregate scores treat these as equivalent; deployment treats them as not. Three concrete cases of aggregation…
- Límite: La lectura primaria permite comprobar método y resultados en el HTML, pero no convierte sus conclusiones en validación independiente. La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 4 Predictive Validity as Ranking Criterion.
- Confianza editorial: Media
- Limitación: El cierre de la fuente señala: The fourteen implementation studies are unpublished implementation reports, not peer-reviewed publications. Their value here is convergence under architectural diversity, not the independent peer review each would warrant as a standalone empirical contribution.
- Limitación: La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 4 Predictive Validity as Ranking Criterion.

## Localizadores de evidencia
- [Fuente primaria · canonical](https://arxiv.org/abs/2606.19704): tipo abstract
- [HTML · lectura completa](https://arxiv.org/html/2606.19704): tipo abstract
- [Método · 2 The Argument](https://arxiv.org/html/2606.19704#S2): tipo section
- [Evaluación · 4 Predictive Validity as Ranking Criterion](https://arxiv.org/html/2606.19704#S4): tipo section
- [Cierre · Limitations](https://arxiv.org/html/2606.19704#Sx1): tipo section

## Próxima prueba

- ¿La propuesta mejora evaluación interna frente a la línea base actual?
- Métrica: Comparar la métrica principal de la fuente junto con calidad, coste, latencia y tasa de errores.
- Regla de parada: Parar si no aparece una mejora reproducible o si aumenta el riesgo, la complejidad o el coste sin compensación.

## Recursos reproducibles
- [Link](https://github.com/YuvalShemla/hpml-2026-project.git)
- [Link](https://github.com/siddharthgowda/AssetOpsBench)
- [Link](https://github.com/Rohith-Kanathur/AssetOpsBench)
- [Link](https://github.com/kmn01/AssetOpsBench/tree/dev)

## Enlaces relacionados

- [VAKRA](https://luiseduardodemiguel.com/research-ia/markdown/papers/vakra)
- [The Devil Is in the Interface](https://luiseduardodemiguel.com/research-ia/markdown/papers/devil-interface)
- [SkillSentry](https://luiseduardodemiguel.com/research-ia/markdown/papers/skillsentry)