What it measures
Relative real-world performance of generalist policies across many scenes and task instructions, ranked by pairwise preferences rather than by a fixed task list.
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
RoboArena is a distributed framework for real-world evaluation of generalist robot policies. Evaluators at seven academic institutions run blinded, pairwise comparisons of policies on the DROID platform in their own scenes and with their own tasks, and the preferences are aggregated into a global ranking. The paper reports 612 pairwise comparisons across seven generalist policies.
Relative real-world performance of generalist policies across many scenes and task instructions, ranked by pairwise preferences rather than by a fixed task list.
Evaluators choose tasks and scenes, run double-blind comparisons between two policies, and record a preference with a short written explanation. A task-aware Bradley-Terry model, fit with an EM algorithm, turns these preferences into a global ranking and per-policy strengths and weaknesses. The paper reports 4,284 evaluation episodes in total, including exhaustive evaluations used as an oracle, and states that about 100 pairwise comparisons are enough for high-quality rankings.
Real-world testing is the ground truth for policy quality, but single-lab and fixed-task evaluations are narrow and costly. RoboArena shows a distributed protocol that could extend to other platforms, including humanoids, although the paper's evaluations use a single Franka arm.
Scoring
| Metric | Key | Unit | Direction |
|---|---|---|---|
| Policy ranking score (task-aware Bradley-Terry strength) | ranking_score | Unitless | Higher is better ↑ |
Hardware
No in-dataset robot embodiments are linked.
Real-robot evaluations on the DROID platform: a 7-DoF Franka Panda with a Robotiq 2F-85 gripper and ZED stereo cameras, run at seven academic institutions. No humanoid is used, so no roster robot is linked.
Recorded outcomes
No result files yet
Evidence
Related evaluations