Humanoid robotics glossary
Vision-language-action (VLA) model
Also known as: VLA, VLA model, Vision-language-action model
A robot model that takes camera images and language instructions and outputs robot actions, rather than only text.
Last verified Oct 10, 2026
What it means
A vision-language-action (VLA) model takes camera images and a language instruction as input and outputs robot actions. The term was introduced in the 2023 RT-2 paper, whose authors represented robot actions as text tokens so that a vision-language model could output them directly.
Many VLA models start from a vision-language model pretrained on web data and then train on robot demonstrations, which is the idea behind the RT-2 title, 'Transfer Web Knowledge to Robotic Control'. OpenVLA, an open-source example, is described as a 7-billion-parameter model trained on 970,000 real-world robot demonstrations.
How it applies to humanoid robots
Figure describes its Helix model as a generalist VLA that produces high-rate continuous control of the entire humanoid upper body, including the wrists, torso, head and individual fingers. The announcement is dated 20 February 2025.
The π₀ paper describes a flow-matching architecture built on top of a pre-trained vision-language model, a different way to attach action generation to a vision-language backbone.
On humanoidversus.com, the ai_stack field records a named model only when a maker or a cited source names it. The field is text, so compare robots by the model name and its source rather than by a score.
What makes a model a VLA
The label is used loosely. Under the RT-2 definition, a model needs vision and language input and robot actions as output, so a vision-language model that outputs only text, or a planner that does not control motion, does not meet it. Companies sometimes coin similar labels, so the site keeps each maker's wording.
Action dimensions are not joint counts. Helix's 35-DoF action space describes what the model outputs, not how many mechanical joints the robot has (see the degrees-of-freedom entry).
Figure's Helix post is a company announcement, so its capability claims are the company's own until independent results are cited.
How to read the ai_stack field
Each ai_stack value carries a source and a confidence level. 'Official' means the maker stated it, and 'reported' means a credible third party did. A null means no named model was found, which is not evidence that the robot has no AI stack.
A named model does not mean the robot is on sale or deployed. Check the status field, and the price and availability sections, before reading a model name as a product feature.
Further reading
RT-2 (Brohan et al., 2023), OpenVLA (Kim et al., 2024), π₀ (Black et al., 2024) and Figure's Helix announcement (2025).
Live dataset
Named AI stacks used by humanoid robots
Top 20 records with a published value; unknown fields are excluded rather than scored as zero.
| Robot | AI stack | Source |
|---|---|---|
| 1X NEO | Redwood AI (1X generalist vision-language model) | Source ↗ |
| Aeolus Robotics aeo | Cross-Sensory Integration combining visual, audio and tactile data in a unified AI model | Source ↗ |
| AgiBot Genie G2 | Genie RL (real-robot RL toolchain) | Source ↗ |
| AgiBot RAISE A1 | EI-Brain embodied intelligence framework | Source ↗ |
| Agile Robots Agile ONE | Layered AI architecture described as a Robot Foundation Model, running on the AgileCore platform | Source ↗ |
| Agility Digit 5 | Agility proprietary human-detection AI (Cooperative Safety) with NVIDIA Halos Core | Source ↗ |
| AI2 Robotics AlphaBot | AI2R Brain embodied intelligence system (renamed Alpha Brain in 2025) | Source ↗ |
| AI2 Robotics AlphaBot 1S | AI2R Brain embodied intelligence system (renamed Alpha Brain in 2025) | Source ↗ |
| AI2 Robotics AlphaBot 2 | GOVLA (Global & Omni-body Vision-Language-Action model), with on-device inference | Source ↗ |
| AiMOGA (Chery) Mornine M1 | MoNet 6B and a customizable large-language-model API with retrieval-augmented generation | Source ↗ |
| Andromeda Abi | Hybrid cloud and local AI architecture; GPT-4 was named for the 2023 prototype | Source ↗ |
| Apptronik Apollo 2 | Apptronik Artemis software platform; Gemini Robotics used in the Google DeepMind research partnership | Source ↗ |
| Astribot S1 | Design for AI (DFAI) architecture | Source ↗ |
| Astribot T1 | Astribot Lumo-2 foundation model | Source ↗ |
| Booster K1 | Doubao LLM (bundled; free for six months) | Source ↗ |
| BOSHIAC DexBot | Huozi-Rixin (活字-日新) large model from HIT, used as the intelligent core | Source ↗ |
| Britannia University Mia-1 | Python 3 application with computer vision and internet information retrieval | Source ↗ |
| Cartwheel Robotics Yogi | Motion Language Model for generative motion from speech or text prompts | Source ↗ |
| CASBOT 01 | Adversarial motion prior with whole-body control (motion control framework) | Source ↗ |
| CASBOT W1 | CASBOT Embodied Brain 2.0 | Source ↗ |
Related data
Explore connected records
References
Sources
- [1]RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control ↗arXiv · paper · accessed Oct 10, 2026
- [2]OpenVLA: An Open-Source Vision-Language-Action Model ↗arXiv · paper · accessed Oct 10, 2026
- [3]π₀: A Vision-Language-Action Flow Model for General Robot Control ↗arXiv · paper · accessed Oct 10, 2026
- [4]Helix: Figure AI announcement ↗Figure AI · official · accessed Oct 10, 2026