AI4AI-Bench
A benchmark for evaluating LLM agents on algorithmic design for recursive self-improvement.
AI4AI-Bench evaluates whether LLM agents can improve the algorithms that train AI systems, using frozen research repositories and reproducible, hidden evaluations.
- Project page: https://lab.einsia.ai/ai4ai/
- Paper: https://arxiv.org/abs/2608.20318
- Code: https://github.com/Einsia/AI4AI-Bench