A new benchmark from Dalian University of Technology (VA-Bench) tested 12 multimodal AI models on robot-arm manipulation tasks.
The results are consistent and uncomfortable: • Object location accuracy: ~100% • Task understanding: ~99% • Whole-task success (best model): only 53.93%
The more important gap is between detecting an error (73.6%) and correcting it in real time (46.7%). Dual-arm tasks collapsed further — single-arm success around 65%, dual-arm only 11%.
Simulation is teaching models what to see and what needs to be done. It is still failing to teach them how to reliably complete the action when conditions change.
For buyers evaluating robotics vendors: treat simulation demo success rates as an upper bound, not a production prediction. Ask for dual-arm success rates on held-out tasks, error-correction rates, and performance when object geometry varies before approving any pilot.
Full analysis:
#Robotics #Simulation #IndustrialAI #Procurement






