ASIMOV-Agentic-v1
ASIMOV-Agentic is a robotics safety benchmark. It evaluates the ability of AI agents to safely control robots by (1) refusing tasks that violate operational constraints; (2) triggering interventions—such as protective stops—during critical events like hardware faults or unsafe human proximity; (3) shielding the Vision-Language-Action (VLA) model from infeasible or out-of-distribution tasks where confidence is low; and (4) proactively resolving ambiguous instructions or scene uncertainties by requesting human help.