Toggle light / dark theme

Open-source benchmark tests whether AI agents can engineer working robots

With the rapid rise of artificial intelligence in daily life, software coding has become increasingly automated, with powerful AI systems known as coding agents able to write and revise computer programs almost autonomously. But what happens when an AI agent must contend not just with digital command lines, but with the physical world of robotics?

Researchers at the Harvard John A. Paulson School of Engineering and Applied Sciences (SEAS) and the Georgia Tech School of Computational Science and Engineering are taking a systematic approach to finding out.

A team led by Na Li, the Winokur Family Professor of Electrical Engineering and Applied Mathematics at Harvard SEAS, and Bo Dai, assistant professor at Georgia Tech, have developed an evaluation tool known as a benchmark that tests how well AI coding agents can perform the challenging task of engineering an actual, physical robot. The team includes Harvard graduate student Haitong Ma and Chenxiao Gao and Rushi Qiang at Georgia Tech.

Leave a Comment

Lifeboat Foundation respects your privacy! Your email address will not be published.

/* */