# Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies
> Ficha editorial pública de Research IA. Estado: Lectura primaria completa. La interpretación editorial no sustituye la fuente primaria.

- Página canónica: https://luiseduardodemiguel.com/research-ia/papers/can-ai-agents-conduct-open-ended-ai-research-early-evidence-from-two-cas
- Fuente primaria: https://arxiv.org/abs/2607.27191
- Versión leída: v2
- Fuente comprobada: 2026-08-19 · lectura primaria completa; extracción editorial automatizada, revisión humana pendiente
- Autores: Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
- Fecha del corte: 29 JULIO 2026.
- Área: AGENTES · RAG · EVALUACIÓN

## Tesis y contexto

Introduce las shadow evaluations: se entrega a un agente la pregunta central de un paper de alta calidad todavía no publicado y los autores originales evalúan el trabajo resultante. Los agentes recibieron seis días y miles de dólares de cómputo. Completaron la ingeniería experimental, pero no produjeron avances científicos sustanciales.

- Problema: Los benchmarks verificables miden tareas estrechas, mientras que enviar papers generados a revisión académica introduce mucho ruido. Por qué puede ser importante: Es una corrección necesaria frente a predicciones de automatización total de la ciencia. Identifica fallos recurrentes en criterio científico, creatividad, backtracking, uso de recursos y mantenimiento del objetivo.
- Por qué importa: La relevancia práctica todavía necesita contraste editorial.

## Evidencia reportada

- **reported-result**: Those runs suffered from the same failure modes, but also suffered from poor literature review and much worse writing quality. [localizador](https://arxiv.org/html/2607.27191#S4)
- **reported-result**: This shows that despite the weak performance observed in our final experiments, more reasoning helped improve performance. [localizador](https://arxiv.org/html/2607.27191#S4)
- **reported-result**: We tentatively think that more reasoning effort within model calls might improve performance, but more wall-clock time or resources would not significantly change the results. [localizador](https://arxiv.org/html/2607.27191#S4)
- **reported-result**: After our initial pilots, in addition to increasing reasoning effort, we conducted a comprehensive manual log analysis ourselves, and then tasked Claude Fable 5 with improving our scaffold based on the failures observed in the pilots. [localizador](https://arxiv.org/html/2607.27191#S4)

## Lectura y límite

- Método: La lectura de 2 Shadow evaluations: A new method for measuring progress towards automating AI research describe la intervención y su construcción: Evaluating automated AI research requires a way to measure the quality of the research that agents produce. Most existing evaluations use one of two approaches: evaluations on verifiable tasks or blind review. In evaluations on verifiable tasks, the agent improves a fixed metric, and an automatic verifier scores the result. There are many examples of such evaluations: RE-Bench ( 44 ) asks agents to solve research-engineering problems, MLE-Bench ( 4 ) asks them to compete in Kaggle competitions, and PostTrainBench asks them to post-train small language models against Q&A benchmarks ( 35 ) . MLS-Bench ( 24 )…
- Límite: La lectura primaria permite comprobar método y resultados en el HTML, pero no convierte sus conclusiones en validación independiente. La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 4 Results.
- Confianza editorial: Media
- Limitación: El cierre de la fuente señala: Nevertheless, there are good reasons to view the distinction between AI capabilities on verifiable tasks and open-ended research problems as germane to the question of accelerating AI R&D. Anthropic’s post on self-improvement explicitly cites Claude’s rising success rate on LLM-judged "open-ended" Claude Code sessions as evidence of self-improvement ( 10 ) . A recent report from the Elasticity Institute on the economics of recursive self-improvement…
- Limitación: La ficha no demuestra transferencia fuera de los datasets, modelos, herramientas y condiciones descritos en 4 Results.

## Localizadores de evidencia
- [Fuente primaria · canonical](https://arxiv.org/abs/2607.27191): tipo abstract
- [HTML · lectura completa](https://arxiv.org/html/2607.27191): tipo abstract
- [Método · 2 Shadow evaluations: A new method for measuring progress towards automating AI research](https://arxiv.org/html/2607.27191#S2): tipo section
- [Evaluación · 4 Results](https://arxiv.org/html/2607.27191#S4): tipo section
- [Cierre · 7 Limitations](https://arxiv.org/html/2607.27191#S7): tipo section

## Próxima prueba

- ¿La propuesta mejora evaluación de AI scientists frente a la línea base actual?
- Métrica: Comparar la métrica principal de la fuente junto con calidad, coste, latencia y tasa de errores.
- Regla de parada: Parar si no aparece una mejora reproducible o si aumenta el riesgo, la complejidad o el coste sin compensación.

## Recursos reproducibles
- [Link](https://github.com/karpathy/autoresearch)
- [the following issues](https://github.com/arXiv/html_feedback/issues)
- [list of packages that need conversion](https://github.com/brucemiller/LaTeXML/wiki/Porting-LaTeX-packages-for-LaTeXML)
- [developer contributions](https://github.com/brucemiller/LaTeXML/issues)

## Enlaces relacionados

- [VAKRA](https://luiseduardodemiguel.com/research-ia/markdown/papers/vakra)
- [The Devil Is in the Interface](https://luiseduardodemiguel.com/research-ia/markdown/papers/devil-interface)
- [SkillSentry](https://luiseduardodemiguel.com/research-ia/markdown/papers/skillsentry)