Humanoid robotics glossary

Vision-language-action (VLA) model

Also known as: VLA, VLA model, Vision-language-action model

A robot model that takes camera images and language instructions and outputs robot actions, rather than only text.

Last verified Oct 10, 2026

What it means

A vision-language-action (VLA) model takes camera images and a language instruction as input and outputs robot actions. The term was introduced in the 2023 RT-2 paper, whose authors represented robot actions as text tokens so that a vision-language model could output them directly.

Many VLA models start from a vision-language model pretrained on web data and then train on robot demonstrations, which is the idea behind the RT-2 title, 'Transfer Web Knowledge to Robotic Control'. OpenVLA, an open-source example, is described as a 7-billion-parameter model trained on 970,000 real-world robot demonstrations.

How it applies to humanoid robots

Figure describes its Helix model as a generalist VLA that produces high-rate continuous control of the entire humanoid upper body, including the wrists, torso, head and individual fingers. The announcement is dated 20 February 2025.

The π₀ paper describes a flow-matching architecture built on top of a pre-trained vision-language model, a different way to attach action generation to a vision-language backbone.

On humanoidversus.com, the ai_stack field records a named model only when a maker or a cited source names it. The field is text, so compare robots by the model name and its source rather than by a score.

What makes a model a VLA

The label is used loosely. Under the RT-2 definition, a model needs vision and language input and robot actions as output, so a vision-language model that outputs only text, or a planner that does not control motion, does not meet it. Companies sometimes coin similar labels, so the site keeps each maker's wording.

Action dimensions are not joint counts. Helix's 35-DoF action space describes what the model outputs, not how many mechanical joints the robot has (see the degrees-of-freedom entry).

Figure's Helix post is a company announcement, so its capability claims are the company's own until independent results are cited.

How to read the ai_stack field

Each ai_stack value carries a source and a confidence level. 'Official' means the maker stated it, and 'reported' means a credible third party did. A null means no named model was found, which is not evidence that the robot has no AI stack.

A named model does not mean the robot is on sale or deployed. Check the status field, and the price and availability sections, before reading a model name as a product feature.

Further reading

RT-2 (Brohan et al., 2023), OpenVLA (Kim et al., 2024), π₀ (Black et al., 2024) and Figure's Helix announcement (2025).

Live dataset

Named AI stacks used by humanoid robots

Top 20 records with a published value; unknown fields are excluded rather than scored as zero.

RobotAI stackSource
1X NEORedwood AI (1X generalist vision-language model)Source ↗
Aeolus Robotics aeoCross-Sensory Integration combining visual, audio and tactile data in a unified AI modelSource ↗
AgiBot Genie G2Genie RL (real-robot RL toolchain)Source ↗
AgiBot RAISE A1EI-Brain embodied intelligence frameworkSource ↗
Agile Robots Agile ONELayered AI architecture described as a Robot Foundation Model, running on the AgileCore platformSource ↗
Agility Digit 5Agility proprietary human-detection AI (Cooperative Safety) with NVIDIA Halos CoreSource ↗
AI2 Robotics AlphaBotAI2R Brain embodied intelligence system (renamed Alpha Brain in 2025)Source ↗
AI2 Robotics AlphaBot 1SAI2R Brain embodied intelligence system (renamed Alpha Brain in 2025)Source ↗
AI2 Robotics AlphaBot 2GOVLA (Global & Omni-body Vision-Language-Action model), with on-device inferenceSource ↗
AiMOGA (Chery) Mornine M1MoNet 6B and a customizable large-language-model API with retrieval-augmented generationSource ↗
Andromeda AbiHybrid cloud and local AI architecture; GPT-4 was named for the 2023 prototypeSource ↗
Apptronik Apollo 2Apptronik Artemis software platform; Gemini Robotics used in the Google DeepMind research partnershipSource ↗
Astribot S1Design for AI (DFAI) architectureSource ↗
Astribot T1Astribot Lumo-2 foundation modelSource ↗
Booster K1Doubao LLM (bundled; free for six months)Source ↗
BOSHIAC DexBotHuozi-Rixin (活字-日新) large model from HIT, used as the intelligent coreSource ↗
Britannia University Mia-1Python 3 application with computer vision and internet information retrievalSource ↗
Cartwheel Robotics YogiMotion Language Model for generative motion from speech or text promptsSource ↗
CASBOT 01Adversarial motion prior with whole-body control (motion control framework)Source ↗
CASBOT W1CASBOT Embodied Brain 2.0Source ↗

Related data

Explore connected records

References

Sources

  1. [1]RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control ↗arXiv · paper · accessed Oct 10, 2026
  2. [2]OpenVLA: An Open-Source Vision-Language-Action Model ↗arXiv · paper · accessed Oct 10, 2026
  3. [3]π₀: A Vision-Language-Action Flow Model for General Robot Control ↗arXiv · paper · accessed Oct 10, 2026
  4. [4]Helix: Figure AI announcement ↗Figure AI · official · accessed Oct 10, 2026