{"id":241858,"date":"2026-07-30T12:12:10","date_gmt":"2026-07-30T17:12:10","guid":{"rendered":"https:\/\/lifeboat.com\/blog\/2026\/07\/humanclaw-can-visionlanguage-models-act-through-a-body"},"modified":"2026-07-30T12:12:10","modified_gmt":"2026-07-30T17:12:10","slug":"humanclaw-can-visionlanguage-models-act-through-a-body","status":"publish","type":"post","link":"https:\/\/lifeboat.com\/blog\/2026\/07\/humanclaw-can-visionlanguage-models-act-through-a-body","title":{"rendered":"HumanCLAW: Can VisionLanguage Models Act Through a Body?"},"content":{"rendered":"<p style=\"padding-right: 20px\"><a class=\"aligncenter blog-photo\" href=\"https:\/\/lifeboat.com\/blog.images\/humanclaw-can-visionlanguage-models-act-through-a-body2.jpg\"><\/a><\/p>\n<p>The Gap: Knowing that a shape in front of you is a \u201cdoor handle\u201d (recognition) is easy. Knowing whether your arm is long enough to reach it, whether your foot is currently blocking the door from swinging open, or if you\u2019ve already walked past it is where current VLMs fail.<\/p>\n<hr>\n<p>Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM\u2019s decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>The Gap: Knowing that a shape in front of you is a \u201cdoor handle\u201d (recognition) is easy. Knowing whether your arm is long enough to reach it, whether your foot is currently blocking the door from swinging open, or if you\u2019ve already walked past it is where current VLMs fail. Evaluating whether a vision-language model [\u2026]<\/p>\n","protected":false},"author":709,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[],"class_list":["post-241858","post","type-post","status-publish","format-standard","hentry","category-futurism"],"_links":{"self":[{"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/posts\/241858","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/users\/709"}],"replies":[{"embeddable":true,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/comments?post=241858"}],"version-history":[{"count":0,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/posts\/241858\/revisions"}],"wp:attachment":[{"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/media?parent=241858"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/categories?post=241858"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/lifeboat.com\/blog\/wp-json\/wp\/v2\/tags?post=241858"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}