What it measures
Not an evaluation benchmark in itself. The dataset covers manipulation skills for training generalist policies. The challenges score submitted policies on their own tracks: World Model and Manipulation at IROS 2025, and Reasoning to Action and World Model at ICRA 2026.