RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
RoboCasa is a large-scale simulation framework for training and evaluating generalist robots on everyday tasks, centered on kitchens. RoboCasa365 (v1.0, released 2026-02-18) expands it to 365 everyday tasks across 2,500 kitchen environments, with 3,200+ objects in more than 150 categories and 2,200+ hours of robot demonstration data. A public leaderboard ranks multi-task policies on a 50-task benchmark.
First published
2024
Organizers
The University of Texas at Austin, NVIDIA Research
Tasks
365
Simulator
robosuite (MuJoCo-based)
License
MIT (code); CC BY 4.0 (assets and datasets)
What it measures
How well a single multi-task policy performs everyday kitchen tasks in new layouts and on composite tasks it was not trained on. The leaderboard reports success rates on Atomic-Seen, Composite-Seen and Composite-Unseen splits, where Composite-Unseen tasks are held out from pretraining and evaluated zero-shot.
How it works
Scenes use diverse 3D assets and generative-AI-enriched content, and the data mix includes human demonstrations and automatically generated trajectories. The original paper evaluates atomic tasks with 50 trials across five fixed evaluation scenes and reports success rates. The RoboCasa365 leaderboard's Overall is the average success rate across the 50-task multi-task benchmark; submissions are JSON files added to the repository by pull request.
Why it matters
Kitchens are a central target for household robots, and the benchmark measures transfer to unseen layouts and composite tasks, which is the core generalization problem for humanoid household work. The framework states support for humanoid robots, but the published evaluations use a mobile Franka arm.
Scoring
Metrics
Metric
Key
Unit
Direction
Overall success rate (average across the 50-task multi-task benchmark)
overall
%
Higher is better ↑
Atomic-Seen success rate
atomic_seen
%
Higher is better ↑
Composite-Seen success rate
composite_seen
%
Higher is better ↑
Composite-Unseen success rate (zero-shot)
composite_unseen
%
Higher is better ↑
Hardware
Robots involved
No in-dataset robot embodiments are linked.
The main experiments use a Franka Panda arm on an Omron mobile base (PandaOmron). The paper and project site say the framework supports humanoid robots and quadrupeds with arms, but no specific humanoid is named in the sources opened, so no roster robot is linked. The GR00T repository lists a separate 'RoboCasa GR1 tabletop tasks' benchmark; see gr1-tabletop-tasks.
The leaderboard lists 17 models ranked by Overall, with success rates in percent. Overall is the average success rate across the 50-task multi-task benchmark (18 Atomic-Seen, 16 Composite-Seen and 16 Composite-Unseen tasks); Composite-Unseen tasks are held out from pretraining and evaluated zero-shot. The RoboCasa365 paper defines the benchmark embodiment as the simulated PandaOmron (Franka Panda arm on an Omron mobile base), which is not in the roster. The 'Open Source' column carries an asterisk that the page does not define; the notes reflect the checkmark only. The page is dated 'Updated 10/06/2026'.
Rank
Robot / policy
Team
Overall success rate (average across the three splits) (%)