What it measures
Knowledge transfer in lifelong learning: whether a policy can learn new tasks while keeping earlier ones, and how that depends on task order, architecture, visual encoder and pretraining. Performance is measured as success rate.
LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
LIBERO is a simulated tabletop manipulation benchmark for lifelong robot learning. It provides 130 language-conditioned tasks in four suites (LIBERO-Spatial, LIBERO-Object, LIBERO-Goal and LIBERO-100, which is split into LIBERO-90 for pretraining and LIBERO-10 for testing), with 50 human-teleoperated demonstrations per task. Its focus is how knowledge transfers between tasks rather than the difficulty of any single task.
Knowledge transfer in lifelong learning: whether a policy can learn new tasks while keeping earlier ones, and how that depends on task order, architecture, visual encoder and pretraining. Performance is measured as success rate.
Tasks are built with BDDL on top of robosuite and each has a natural-language instruction. The suites vary spatial layout, object types and goals to isolate different transfer types, and LIBERO-Long contains longer-horizon tasks. The paper studies knowledge-transfer types, policy architectures, lifelong learning algorithms, robustness to task ordering and the effect of pretraining. The repository notes that only a sparse reward (+1 on task completion) is supported.
Its language-conditioned, demonstration-based format matches many generalist-policy pipelines. The robot is a single tabletop arm and no humanoid is used, so results transfer to humanoids only indirectly.
Scoring
| Metric | Key | Unit | Direction |
|---|---|---|---|
| Success rate | success_rate | Unitless | Higher is better ↑ |
Hardware
No in-dataset robot embodiments are linked.
Single-arm tabletop manipulation in robosuite. The arXiv text opened does not name the arm model; secondary summaries describe a Franka Emika Panda, which is not verified here. No humanoid is used, so no roster robot is linked.
Recorded outcomes
No result files yet
Evidence
Related evaluations