Evaluation landscape
Benchmarks and competitions
A structured guide to the tests, events, datasets, and standards that make humanoid capabilities measurable.
26 of 26 records
Simulation
Repeatable virtual tasks for policies, control systems, and whole-body behavior.
BEHAVIOR-1K
Stanford University
BEHAVIOR-1K is a simulation benchmark of 1,000 everyday household activities, instantiated in 50 interactive scenes with 9,000+ annotated objects in the OmniGibson simulator (the project site says 10,000+ objects). The BEHAVIOR Challenge runs competitions on subsets of these activities: 50 tasks at NeurIPS 2025 and 100 tasks in the 2026 edition, which is open for submissions until 2026-10-16. A preliminary version of the benchmark appeared at CoRL 2022.
Humanoid relevance: Medium
GR-1 Tabletop Tasks
NVIDIA
GR-1 Tabletop Tasks is a simulated benchmark of 24 tabletop manipulation tasks for the Fourier GR-1 humanoid with dexterous hands. It was introduced in the GR00T N1 paper and focuses on dexterous hand control: 18 rearrangement tasks and 6 tasks that place objects into articulated containers such as cabinets, drawers and microwaves. The GR00T repository lists it among its simulation benchmarks.
Humanoid relevance: High
HumanoidBench
Carmelo Sferrazza et al.
HumanoidBench is a simulated benchmark of 27 whole-body tasks for a humanoid robot with dexterous hands: 12 locomotion tasks and 15 whole-body manipulation tasks. It runs in MuJoCo, with a Unitree H1 fitted with two Shadow Dexterous Hands as the primary robot. The authors report that off-the-shelf reinforcement learning methods fall below the success threshold on most tasks.
Humanoid relevance: High
LIBERO
Bo Liu et al.
LIBERO is a simulated tabletop manipulation benchmark for lifelong robot learning. It provides 130 language-conditioned tasks in four suites (LIBERO-Spatial, LIBERO-Object, LIBERO-Goal and LIBERO-100, which is split into LIBERO-90 for pretraining and LIBERO-10 for testing), with 50 human-teleoperated demonstrations per task. Its focus is how knowledge transfers between tasks rather than the difficulty of any single task.
Humanoid relevance: Low
LocoMuJoCo
Firas Al-Hafez et al.
LocoMuJoCo is an imitation-learning benchmark for locomotion in MuJoCo, presented at a NeurIPS 2023 robot learning workshop. It contains 12 humanoid and 4 quadruped environments, plus biomechanical human models, each with noisy motion-capture reference data mapped to the embodiment, expert demonstrations and sub-optimal demonstrations. The repository adds MJX support for parallel simulation.
Humanoid relevance: High
ManiSkill3
Stone Tao et al.
ManiSkill3 is an open-source, GPU-parallelized robotics simulator and benchmark built on the SAPIEN engine for generalizable embodied AI. The paper describes environments spanning 12 domains, including mobile manipulation and humanoids, and reports throughput of up to 30,000+ FPS in benchmarked environments. The robot library includes Unitree G1 and H1 models.
Humanoid relevance: Medium
RoboCasa
The University of Texas at Austin · NVIDIA Research
RoboCasa is a large-scale simulation framework for training and evaluating generalist robots on everyday tasks, centered on kitchens. RoboCasa365 (v1.0, released 2026-02-18) expands it to 365 everyday tasks across 2,500 kitchen environments, with 3,200+ objects in more than 150 categories and 2,200+ hours of robot demonstration data. A public leaderboard ranks multi-task policies on a 50-task benchmark.
Humanoid relevance: Medium
SIMPLER
Xuanlin Li et al.
SIMPLER is a simulation-based evaluation suite for generalist manipulation policies. It recreates common real-robot setups in simulation (Google Robot and WidowX with the Bridge setup), and its authors report that simulated success correlates with real-world performance and reflects sensitivity to distribution shifts. The environments and the workflow for creating new ones are open-sourced.
Humanoid relevance: Low
SkillBench
Yuxuan Kuang et al.
SkillBench is a simulated, cross-embodiment benchmark of eight humanoid loco-manipulation tasks, introduced with the SkillBlender framework. It covers three Unitree humanoids (H1, G1 and H1-2) and four primitive skills. It tests whether pre-trained, goal-conditioned skills can be blended into complex whole-body tasks with little task-specific reward engineering.
Humanoid relevance: High
Real-world evaluation
Physical protocols that measure capability under controlled, reproducible conditions.
RoboArena
Pranav Atreya et al.
RoboArena is a distributed framework for real-world evaluation of generalist robot policies. Evaluators at seven academic institutions run blinded, pairwise comparisons of policies on the DROID platform in their own scenes and with their own tasks, and the preferences are aggregated into a global ranking. The paper reports 612 pairwise comparisons across seven generalist policies.
Humanoid relevance: Low
RoboChallenge
Adina Yakefu et al.
RoboChallenge is an online real-robot evaluation system for embodied control policies. Its first benchmark, Table30, has 30 tabletop manipulation tasks on four robot types (UR5, Franka Panda, Cobot Magic Aloha and ARX-5), with ten machines hosted online. Participants run their models on their own side and send task requests through an API, so no model weights or containers are uploaded.
Humanoid relevance: Low
Competitions
Organized events where teams and robots perform under a shared rulebook.
Beijing Humanoid Robot Half Marathon
Beijing E-Town
A 21.0975 km half marathon in Beijing's Yizhuang (E-Town) area in which humanoid robots run on a parallel course beside human runners. It is organized by Beijing E-Town and held annually. The first edition took place on 19 April 2025 and the second on 19 April 2026.
Humanoid relevance: High
CMG Mecha Fighting Series
China Media Group
The CMG World Robot Competition – Mecha Fighting Series is a humanoid robot combat tournament organized by China Media Group, with Unitree Robotics as partner. The first edition was held in Hangzhou on 25 May 2025, with four operator teams controlling Unitree G1 humanoids in boxing-style matches.
Humanoid relevance: High
DARPA Robotics Challenge
DARPA
The DARPA Robotics Challenge was a U.S. Defense Advanced Research Projects Agency prize competition (2012–2015) for semi-autonomous ground robots performing disaster-response tasks in dangerous, degraded, human-engineered environments. The Trials were held in Florida in December 2013 and the Finals at the Fairplex in Pomona, California, on 5–6 June 2015. It is a historical reference.
Humanoid relevance: Medium
RoBoLeague
Shangyicheng Group
RoBoLeague is a Beijing robot football tournament in which teams field humanoid robots that play fully autonomous, AI-driven 3-on-3 matches. The inaugural 2025 final was won by a Tsinghua University team, which beat China Agricultural University's Mountain-Sea 5–3. No later edition had been found as of 2026-10-08.
Humanoid relevance: High
ROBO-ONE
Biped Robot Association (general incorporated association, Japan)
ROBO-ONE is a Japanese tournament for bipedal combat robots, organized by the Biped Robot Association and sponsored by MISUMI. The 44th edition was held on 20–21 September 2025 at the Kanagawa Prefectural Youth Center, and the defending champion KOBIS retained its title.
Humanoid relevance: Medium
RoboCup Humanoid League
RoboCup Federation
The RoboCup Humanoid League runs autonomous humanoid robot soccer as part of the annual RoboCup championship. Teams field their own humanoid platforms, and matches are played without human intervention. Historically the league was organized into KidSize and AdultSize classes. For 2026 it was reorganized as the Humanoid Soccer League, formed by merging the Humanoid League and the Standard Platform League, with Small, Middle and Large divisions.
Humanoid relevance: High
URKL
Shenzhen EngineAI Robotics Technology Co., Ltd. · Shenzhen Quanmingxing Robotics Technology Co., Ltd.
URKL (Ultimate Robot Knock-out Legend) is a humanoid robot fighting league organized by Shenzhen EngineAI Robotics. Its inaugural season began on 9 February 2026 in Shenzhen and runs through December 2026. The championship belt is described as a 10 kg gold belt valued at about RMB 10 million (about USD 1.44 million).
Humanoid relevance: High
World Humanoid Robot Games
People's Government of Beijing Municipality · China Media Group (CMG) · World Robot Cooperation Organization · RoboCup Asia-Pacific Confederation Board of Trustees
The World Humanoid Robot Games is a recurring multi-event competition in Beijing spanning robot sports and real-world application scenarios. The inaugural 2025 edition had 280 teams and 26 events. The second edition ran 22–26 August 2026 with 666 teams, 2,056 robots and 51 events, including 21 scenario-based events.
Humanoid relevance: High
Datasets
Curated observations, demonstrations, or trajectories used to train and evaluate robot systems.
AgiBot World
AgiBot · OpenDriveLab
AgiBot World is a large real-robot manipulation dataset and platform from AgiBot, released alongside the GO-1 generalist policy. The Beta release contains 1,003,672 trajectories across five target domains, and the arXiv paper describes over 1 million trajectories across 217 tasks. Public challenges run on top of the dataset: the AgiBot World Challenge @ IROS 2025 and @ ICRA 2026 (the latter uses the AGIBOT G2 robot).
Humanoid relevance: Medium
Humanoid Everyday
Zhenyu Zhao et al.
Humanoid Everyday is a real-world humanoid manipulation dataset with 10.3k trajectories and over 3 million frames across 260 tasks in 7 broad categories. It was collected on Unitree G1 and H1 humanoids and records RGB, depth, LiDAR and tactile streams with language annotations. The authors also benchmark policy-learning methods and describe a cloud evaluation platform for external policies, which the project page lists as coming soon.
Humanoid relevance: High
Open X-Embodiment
Open X-Embodiment Collaboration
Open X-Embodiment is a large, cross-institution collection of real robot manipulation data: over 1 million trajectories from 22 robot embodiments, contributed by 21 institutions. The project also released RT-X models trained on the pooled data. It is a training resource rather than an evaluation benchmark, and its embodiments are single arms, bimanual robots and quadrupeds.
Humanoid relevance: Low
RoboMIND
Kun Wu et al. (X-Humanoid)
RoboMIND is a teleoperated, multi-embodiment manipulation dataset covering four robot types: a Franka Emika Panda, a UR5e, an AgileX dual-arm robot and a humanoid with dual dexterous hands. The project page reports about 107k real-world demonstration trajectories across 479 tasks and 96 object classes, including 5k real-world failure demonstrations, plus an Isaac Sim digital twin.
Humanoid relevance: Medium
Standards
Formal test methods, terminology, and safety requirements from standards bodies.
Standards System for Humanoid Robots and Embodied Intelligence (2026 Edition)
Ministry of Industry and Information Technology (MIIT) · MIIT/TC08 National Technical Committee on Humanoid Robots and Embodied Intelligence · China Electronics Standardization Institute (CESI)
China's first national standards system for humanoid robots and embodied intelligence, released on 28 February 2026 in Beijing. It was developed by the MIIT technical committee TC08 with more than 120 research institutions, enterprises and industry users. It is a top-level framework of six pillars, including safety and ethics, that will guide individual standards across the industry's lifecycle.
Humanoid relevance: High
IEEE RAS Humanoid Study Group
IEEE Robotics and Automation Society
The IEEE Robotics and Automation Society (RAS) Study Group – Humanoid Robots was launched in June 2024 to speed up humanoid standards work. It was led by Aaron Prather of ASTM International. Its final report, "A Pathway Study for Future Humanoid Standards" (September 2025), reviews existing standards, identifies gaps and proposes a roadmap for standards development organizations.
Humanoid relevance: High
ISO 25785-1
ISO/TC 299 (Robotics)
ISO 25785-1 is a planned ISO safety standard for industrial mobile robots whose stability depends on active control, including legged, wheeled and other balancing robots. It is at the Committee Draft (CD) stage under ISO/TC 299. The comment and voting period closed on 8 July 2026, and no publication date is shown in the project record.
Humanoid relevance: Medium