Monday, October 5, 2026 Sign inSubscribe to The Brief →
Humanoid Threats

Security for the age of physical AI

Glossary

What Is a Vision-Language-Action (VLA) Model? Robot AI Brains Explained

A vision-language-action model turns camera images and spoken instructions into robot motion. Learn how VLA models work and how attackers can mislead them.

Last updated: October 5, 2026

A vision-language-action (VLA) model is an AI model that takes in camera images and a natural-language instruction and outputs robot actions directly, such as joint positions or gripper movements. It is the "brain" that lets a humanoid hear "put the cup in the sink" and then move its arm to do it. Because it turns words and pixels into motion, a VLA model is also a new attack surface for physical AI.

How a vision-language-action model works

Google DeepMind introduced the term with RT-2 in July 2023. Its key idea was to express robot actions "as text tokens" and train them alongside ordinary language, so one model learns both to understand the web and to control a robot (RT-2 paper).

Most VLA models follow the same basic recipe:

  • Vision encoder turns camera frames into features the model can reason over.
  • Language model backbone reads the instruction and the visual features together.
  • Action head converts the model's output into motor commands, either as discrete action tokens or through a separate fast controller.

Open models made the approach widely available. OpenVLA (June 2024) is a 7-billion-parameter model built on Llama 2, trained on about 970,000 real robot demonstrations.

VLA models in humanoid robots

Humanoids usually split the work into a slow "thinking" system and a fast "moving" system:

Model Released Design Source
Figure Helix Feb 2025 7B vision-language model at 7-9 Hz plus an 80M-parameter controller at 200 Hz, 35 degrees of freedom Figure
NVIDIA Isaac GR00T N1 Mar 2025 Vision-language module (System 2) feeding a diffusion transformer that generates motor actions (System 1); open foundation model, tested on the Fourier GR-1 arXiv
OpenVLA Jun 2024 Single 7B model, open weights, robot arms arXiv

Why VLA models matter for humanoid security

A VLA model decides what the robot's body does next, so anything that can influence its inputs can influence physical motion. Research so far points to three main risks:

  1. Adversarial patches. A 2024 study showed that small, colorful patches placed in a robot's camera view could cut task success by up to 100% in simulation, and one targeted variant could change the robot's trajectory (Wang et al.).
  2. Transferable attacks. A CVPR 2026 paper reported universal patches that transfer across different VLA models, tasks and viewpoints, including in physical runs, which means an attacker may not need access to the specific model (Lu et al.).
  3. Jailbreaks and instruction manipulation. Researchers at Penn showed that robots driven by large language models can be talked into unsafe actions with crafted prompts (RoboPAIR). VLA models share the same language interface, so similar manipulation is a live research concern.

Analysis: these attacks hit the model, not the network, so firewalls and patching alone do not cover them. Defenders should treat the VLA model as one layer and keep independent safety limits (force, speed, workspace and an E-stop) beneath it. We track published robot vulnerabilities in our Humanoid Robot Vulnerability Tracker, and the Humanoid Robot Security Checklist covers deployment controls.

Example

A warehouse humanoid is told "move the boxes to the pallet." Its VLA model reads the camera feed and the instruction and outputs arm movements. If a printed adversarial sticker sits in view, the research above suggests the model could fail the task or move along a different path, even though no software was hacked. A hard-coded safety layer that limits speed and keeps the robot inside a marked zone limits the damage.

FAQ

What does VLA stand for in robotics? VLA stands for vision-language-action. It describes a model that combines what a robot sees (vision) and what it is told (language) to produce what it does (action).

Is a VLA model the same as an LLM? No. Most VLA models are built on top of a language model backbone, but they are trained on robot demonstrations to output motor actions, not just text.

Can a VLA model be hacked? Published research shows VLA models can be misled by adversarial patches in the camera view and potentially by manipulated instructions. These are research results; we know of no confirmed real-world attack on a deployed humanoid VLA as of this update.

Which humanoid robots use VLA models? Figure uses its Helix VLA, and NVIDIA's GR00T N1 has been demonstrated on the Fourier GR-1. Many other humanoid makers describe similar approaches.

Get the weekly Humanoid Threats Brief → Subscribe

The Humanoid Threats Brief

The weekly briefing on humanoid robot security.

New vulnerabilities, incidents, standards and defenses, with why each one matters. Built for security teams, robotics engineers and the people buying humanoids.

  • Every claim sourced. We link the CVE, the paper or the regulator, not rumors.
  • 5-minute read. One email a week, every Thursday.
  • Free. No spam, unsubscribe in one click.

Check your inbox to confirm your subscription.

Something went wrong. Please try again.

Read a recent storyUnitree G1 Vulnerabilities Explained: CVE-2026-76639 and CVE-2026-76640 →
100% primary-source linked 1 email a week