If AI is going to improve itself, who writes the test?
This research explores whether AI can create its own tests, and if humans still have a role in that process.
As AI systems get better at improving themselves, a key question arises: can they also design the benchmarks—the tests and challenges—used to measure their own progress? Creating good benchmarks is hard, creative work currently done by human scientists. It involves deciding what to measure, gathering source material, and designing tasks that are difficult but fair.
The researchers built a system calledmark to see if an AI agent could handle this entire process on its own, and to find out where humans might still be needed.
The AI agent runs in a loop. It proposes a benchmark, which includes creating tasks and a reference solution. Other AI “solver” agents then attempt these tasks, and their scores tell the creator if the benchmark is hard enough. A separate AI “judge” reviews the benchmark for quality, checking things like whether it’s valid, solvable, and actually measures what it claims to. The agent uses all this feedback to revise its benchmark over several iterations.
The main finding is that humans still matter, but only when they give concrete, detailed guidance.
- No Human Help: When the AI worked completely alone, it created benchmarks that were too easy—solver AIs scored above 80 out of 100, meaning the tests were “saturated” and didn’t reveal much about model capabilities.
- Vague Human Help: A one-sentence suggestion about what to build helped only marginally.
- Detailed Human Help: When humans provided a detailed specification and curated the source material the AI could use, the benchmarks became much harder. Solver scores dropped dramatically, in some cases by nearly half. This shows that human input at the planning stage is crucial for creating a meaningful challenge.
- Real-Time Rescue: In one case, the AI’s benchmark creation process stalled. Human experts intervened mid-way with specific advice on *how* to build the tasks (not *what* to build), and this successfully got the process back on track, resulting in a much harder benchmark.
The AI judge was also critical. It caught flawed benchmarks that a low score alone couldn’t reveal—for example, cases where answers accidentally leaked to the solvers, making a broken task look deceptively difficult.
Current AI agents can manage the entire benchmark creation loop, but they tend to produce tests that are too easy. To create genuinely challenging and high-quality benchmarks, specific and detailed human direction is still essential, especially in defining the problem and providing the right raw materials. The most important and difficult future question isn’t technical, but philosophical: deciding which problems are meaningful and worth making difficult is a judgment we don’t yet know how to delegate to an AI.
#ArtificialIntelligence #benchmarking #RecursiveSelfImprovement #AIAgents #HumanAICollaboration
A framework to study AI models in Reasoning, Alignment, and use of Memory (RAM).
