What it measures
Real-robot task success rate and a partial-progress score for generalist policies and for policies trained per task, measured on identical hosted hardware.
RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
RoboChallenge is an online real-robot evaluation system for embodied control policies. Its first benchmark, Table30, has 30 tabletop manipulation tasks on four robot types (UR5, Franka Panda, Cobot Magic Aloha and ARX-5), with ten machines hosted online. Participants run their models on their own side and send task requests through an API, so no model weights or containers are uploaded.
Real-robot task success rate and a partial-progress score for generalist policies and for policies trained per task, measured on identical hosted hardware.
Table30 has a task-specific setting, where a model is trained for each task, and a generalist setting, where about 50 demonstrations per task are mixed into one model per robot type. Each task has 10 rollouts, and the progress score totals up to 100 points per task. Robots are hosted and reached through an asynchronous API in a 'remote robot' paradigm; requests are queued manually, and the paper says waits may run from hours to days.
Standardized testing on identical hosted hardware addresses the reproducibility problems of real-world evaluation. Its robots are arms and a mobile bimanual system rather than humanoids, so its relevance to humanoids is indirect.
Scoring
| Metric | Key | Unit | Direction |
|---|---|---|---|
| Task success rate (end-to-end) | success_rate | Unitless | Higher is better ↑ |
| Partial-progress score (up to 100 points per task) | progress_score | Unitless | Higher is better ↑ |
Hardware
No in-dataset robot embodiments are linked.
Hosted real robots: UR5 (single 6-DoF arm), Franka Panda, Cobot Magic Aloha (two arms on a moving platform) and ARX-5. None is a humanoid, so no roster robot is linked.
Recorded outcomes
No result files yet
Evidence
Related evaluations