OpenVLA addressed a gap in the available models. Earlier VLA models such as Google’s RT-2 were closed-source. Other open models were trained entirely on simulated data and did not transfer directly to real robots. OpenVLA was the first fully open-source generalist manipulation policy of comparable scope.
Architecture and training
Model and source release
OpenVLA is a 7-billion parameter model built on a Prismatic VLM backbone. It combines LLaMA 2 with the DINOv2 and SigLIP vision encoders. The researchers released the following materials:- All pre-training and fine-tuning code
- Model weights
- Data mixtures used for training
Real-world pre-training
OpenVLA is trained on nearly 1 million real-world robot episodes from the Open X-Embodiment dataset. The training data span 27 robotic datasets. The resulting model can control several robots without additional training, including:- WidowX
- Google RTX
- Franka Panda
Performance
Without additional training, OpenVLA outperforms earlier open-source models such as Octo and RT-1X. It performs as well as or better than the 55-billion parameter closed-source RT-2X in most task categories. Its average absolute success rate is 20% higher. OpenVLA performs well on:- Language grounding, mapping instructions to the correct visual referents
- Multi-instruction tasks with distractor objects, staying on-task in cluttered scenes
Next-token prediction
OpenVLA treats robotic control as a classification problem in the same way as a text-based LLM:- The robot’s 7-dimensional continuous action space (position, rotation, gripper state) is discretized into 255 uniform bins
- The model predicts physical actions as standard text tokens using cross-entropy loss
- No architectural modification of the underlying VLM is required
Compute requirements
You do not need a server cluster to use OpenVLA:- Parameter-Efficient Fine-Tuning (PEFT) with LoRA matches full fine-tuning performance while training only 1.4% of the model’s parameters
- 4-bit quantization reduces the GPU memory requirement from 16 GB to 7 GB of GPU VRAM, with no observed performance degradation
Current limitations
As a large autoregressive model, OpenVLA has several constraints:- Single-frame inputs only, no temporal context across multiple frames
- Single-step action prediction, predicts one action at a time, no action chunking
- Inference speed, caps at roughly 3-9 Hz depending on hardware
Further reading
- OpenVLA project page, paper, code, models, demos
- Kim et al. (2024). OpenVLA: An Open-Source Vision-Language-Action Model
- Open X-Embodiment dataset, the training corpus
- Prismatic VLMs, the backbone architecture

