Skip to main content
SmolVLA is Hugging Face’s compact Vision-Language-Action model based on the LeRobot library. You can fine-tune and deploy it on a single consumer GPU. It remains competitive with larger open VLAs on standard manipulation benchmarks.

Resources