Skip to main content
OpenVLA is an open-source Vision-Language-Action model for general robot manipulation. Its initial presentation prompted a “record number of questions by far for any of the robot paper discussions” hosted by Hugging Face. This response indicated strong interest among researchers.
OpenVLA addressed a gap in the available models. Earlier VLA models such as Google’s RT-2 were closed-source. Other open models were trained entirely on simulated data and did not transfer directly to real robots. OpenVLA was the first fully open-source generalist manipulation policy of comparable scope.

Architecture and training

Model and source release

OpenVLA is a 7-billion parameter model built on a Prismatic VLM backbone. It combines LLaMA 2 with the DINOv2 and SigLIP vision encoders. The researchers released the following materials:
  • All pre-training and fine-tuning code
  • Model weights
  • Data mixtures used for training
The released code, weights, and data allow researchers to reproduce and extend the model.

Real-world pre-training

OpenVLA is trained on nearly 1 million real-world robot episodes from the Open X-Embodiment dataset. The training data span 27 robotic datasets. The resulting model can control several robots without additional training, including:
  • WidowX
  • Google RTX
  • Franka Panda
Training on real demonstrations helps the model transfer to physical hardware.

Performance

Without additional training, OpenVLA outperforms earlier open-source models such as Octo and RT-1X. It performs as well as or better than the 55-billion parameter closed-source RT-2X in most task categories. Its average absolute success rate is 20% higher. OpenVLA performs well on:
  • Language grounding, mapping instructions to the correct visual referents
  • Multi-instruction tasks with distractor objects, staying on-task in cluttered scenes

Next-token prediction

OpenVLA treats robotic control as a classification problem in the same way as a text-based LLM:
  1. The robot’s 7-dimensional continuous action space (position, rotation, gripper state) is discretized into 255 uniform bins
  2. The model predicts physical actions as standard text tokens using cross-entropy loss
  3. No architectural modification of the underlying VLM is required
This is the language-model training paradigm, with some vocabulary tokens assigned to robot actions.

Compute requirements

You do not need a server cluster to use OpenVLA:
  • Parameter-Efficient Fine-Tuning (PEFT) with LoRA matches full fine-tuning performance while training only 1.4% of the model’s parameters
  • 4-bit quantization reduces the GPU memory requirement from 16 GB to 7 GB of GPU VRAM, with no observed performance degradation
These methods allow OpenVLA to run on consumer-grade GPUs. This is unusual for a 7B-parameter foundation model.

Current limitations

As a large autoregressive model, OpenVLA has several constraints:
  • Single-frame inputs only, no temporal context across multiple frames
  • Single-step action prediction, predicts one action at a time, no action chunking
  • Inference speed, caps at roughly 3-9 Hz depending on hardware
These constraints make OpenVLA unsuitable for high-frequency control tasks or complex bimanual manipulation without further optimization. Possible changes include action chunking, distillation, and more efficient inference backends.

Further reading