{"id":244772,"date":"2026-10-01T21:04:36","date_gmt":"2026-10-02T02:04:36","guid":{"rendered":"https:\/\/lifeboat.com\/blog\/2026\/10\/autobenchmark-benchmark-creation-the-role-of-humans"},"modified":"2026-10-01T21:04:36","modified_gmt":"2026-10-02T02:04:36","slug":"autobenchmark-benchmark-creation-the-role-of-humans","status":"publish","type":"post","link":"https:\/\/lifeboat.com\/blog\/2026\/10\/autobenchmark-benchmark-creation-the-role-of-humans","title":{"rendered":"AutoBenchmark: benchmark creation &amp; the role of humans"},"content":{"rendered":"<p><a class=\"aligncenter blog-photo\" href=\"https:\/\/lifeboat.com\/blog.images\/autobenchmark-benchmark-creation-the-role-of-humans2.jpg\"><\/a><\/p>\n<p>If AI is going to improve itself, who writes the test?<\/p>\n<p>This research explores whether AI can create its own tests, and if humans still have a role in that process.<\/p>\n<p>As AI systems get better at improving themselves, a key question arises: can they also design the benchmarks\u2014the tests and challenges\u2014used to measure their own progress? Creating good benchmarks is hard, creative work currently done by human scientists. It involves deciding what to measure, gathering source material, and designing tasks that are difficult but fair.<\/p>\n<p>The researchers built a system calledmark to see if an AI agent could handle this entire process on its own, and to find out where humans might still be needed.<\/p>\n<p>The AI agent runs in a loop. It proposes a benchmark, which includes creating tasks and a reference solution. Other AI \u201csolver\u201d agents then attempt these tasks, and their scores tell the creator if the benchmark is hard enough. A separate AI \u201cjudge\u201d reviews the benchmark for quality, checking things like whether it\u2019s valid, solvable, and actually measures what it claims to. The agent uses all this feedback to revise its benchmark over several iterations.<\/p>\n<p>The main finding is that humans still matter, but only when they give concrete, detailed guidance.<\/p>\n<p>- No Human Help: When the AI worked completely alone, it created benchmarks that were too easy\u2014solver AIs scored above 80 out of 100, meaning the tests were \u201csaturated\u201d and didn\u2019t reveal much about model capabilities.<\/p>\n<div class=\"more-link-wrapper\"> <a class=\"more-link\" href=\"https:\/\/lifeboat.com\/blog\/2026\/10\/autobenchmark-benchmark-creation-the-role-of-humans\">Continue reading \u201cAutoBenchmark: benchmark creation &amp; the role of humans\u201d | &gt;<\/a><\/div>\n","protected":false},"excerpt":{"rendered":"<p>If AI is going to improve itself, who writes the test? This research explores whether AI can create its own tests, and if humans still have a role in that process. As AI systems get better at improving themselves, a key question arises: can they also design the benchmarks\u2014the tests and challenges\u2014used to measure their [\u2026]<\/p>\n","protected":false},"author":709,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1635,6],"tags":[],"class_list":["post-244772","post","type-post","status-publish","format-standard","hentry","category-materials","category-robotics-ai"],"_links":{"self":[{"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/posts\/244772","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/users\/709"}],"replies":[{"embeddable":true,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/comments?post=244772"}],"version-history":[{"count":0,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/posts\/244772\/revisions"}],"wp:attachment":[{"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/media?parent=244772"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/categories?post=244772"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/tags?post=244772"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}