BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
BEHAVIOR-1K is a simulation benchmark of 1,000 everyday household activities, instantiated in 50 interactive scenes with 9,000+ annotated objects in the OmniGibson simulator (the project site says 10,000+ objects). The BEHAVIOR Challenge runs competitions on subsets of these activities: 50 tasks at NeurIPS 2025 and 100 tasks in the 2026 edition, which is open for submissions until 2026-10-16. A preliminary version of the benchmark appeared at CoRL 2022.
First published
2024
Organizers
Stanford University
Tasks
1000
Simulator
OmniGibson
License
MIT (code)
What it measures
Whether a robot can complete long-horizon, full-length household activities in realistic simulation. Challenge entries are scored with partial credit: the fraction of BDDL goal predicates satisfied at the end of an episode, averaged over tasks.
How it works
Activities are specified as BDDL goal conditions and executed in OmniGibson, which simulates rigid, deformable and liquid objects. The 2025 challenge provided about 10,000 teleoperated expert demonstrations (1,200+ hours) for 50 tasks; the 2026 edition's demonstrations were collected with JoyLo, a whole-body teleoperation interface. The standard track allows RGB, depth, segmentation and proprioception, and the privileged track allows simulator queries; the Q-score, the success score averaged across tasks, ranks submissions.
Why it matters
Household chores are a primary target for mobile manipulators and humanoids. The challenge robot is the R1 Pro, a mobile base with a torso and two arms, so its tasks test whole-body bimanual household work, although the benchmark itself is not humanoid-specific.
Scoring
Metrics
Metric
Key
Unit
Direction
Q-score (partial-credit task success, averaged across tasks)
Challenge robot is the simulated Galaxea R1 Pro ('r1pro' config in OmniGibson): a wheeled base with torso and two arms, with a 23-dimensional joint action space by default in the 2025 evaluation. The 2026 default is also the R1 Pro, but participants may supply other OmniGibson-supported robots. The BEHAVIOR-1K paper's sim-to-real study uses a mobile manipulator that is not named.
Timeline
Editions and next edition
2025
Dec 7, 2025 · NeurIPS 2025, Convention Center (in person) and remote
Foundation Models Meet Embodied Agents Challenge @ NeurIPS 2025, co-hosted with the Embodied Agent Interface Competition. 50 tasks, Standard and Privileged tracks, Q-score ranking. The archived leaderboard remains headed 'Provisional', although the archive index says the leaderboard was officially online and congratulates the winners; no replacement final table was published as of 2026-10-08.
Next: 2026
Oct 16, 2026
Open now: launched 2026-07-02; submission deadline 2026-10-16; winners announced 2026-11-04. 100 full-length household tasks, ranked by average task success score with BDDL partial credit. Prize pool USD 11,000. Leaderboard on Hugging Face.
Recorded outcomes
Results
Foundation Models Meet Embodied Agents Challenge @ NeurIPS 2025: organizer-published leaderboard (Standard and Privileged tracks)
Edition 2025 · Dec 7, 2025
Values are the test-set scores shown in bold on the archived organizer leaderboard. The page still carries the heading 'Provisional 2025 Challenge Leaderboard', but the archive index says the leaderboard was officially online and congratulates the winners; no replacement final table was published as of 2026-10-08. Only the five rows with complete score sets are included. Rows 6 to 18 show ambiguous two-value listings and are omitted. Method names are not listed, so policy is null. The 2025 evaluation page refers to the 'r1pro' configuration without naming a manufacturer; the R1 Pro is mapped to the roster slug galaxea-r1.