Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Awesome Vision-Language-Action

🦾 Awesome Vision-Language-Action (VLA)

A curated list of Vision-Language-Action (VLA) models for robotics and embodied AI: foundation models, architectures, benchmarks, and datasets, with a focus on 2025–2026.

Awesome License: MIT GitHub Stars GitHub Forks Last Commit


Vision-Language-Action (VLA) models map pixels and language instructions directly to robot actions, bringing internet-scale knowledge to embodied control. This list tracks the foundation models, action representations, benchmarks, and datasets driving the field — with a focus on 2025–2026 and open releases.

Found a great model or paper we missed? Contributions welcome.


📚 Contents


📖 Surveys & Overviews

🤖 Generalist & Foundation VLA Models

🏗️ Architectures & Action Representations

Diffusion-action

Flow-matching & latent action

Autoregressive & spatial/3D action

⚡ Efficient & Real-time VLA

🧠 Reasoning & World-model-augmented VLA

🚗 Autonomous-Driving VLA

VLA applied to end-to-end driving — a large, fast-moving subfield distinct from tabletop manipulation.

🧪 Data, Simulation & Benchmarks

🗂️ Datasets

💻 Open-source Models & Code

  • OpenVLA (code) — Reference PyTorch implementation, weights, and fine-tuning scripts for the 7B OpenVLA model (2024).
  • Octo (code) — Jax/Flax training and inference for the Octo generalist policy with the diffusion action head (2024).
  • openpi (π0 / π0.5, code) — Physical Intelligence's official open release of π0 and π0.5 weights and inference code (2024-2025).
  • LeRobot — Hugging Face's end-to-end robot-learning library and the home of SmolVLA, plus datasets, training, and real-robot deployment (2024).
  • RoboVLMs (code) — The unified framework from the "What Matters in Building VLAs" study, letting you swap VLM backbones and action heads (2024).
  • Isaac GR00T (code) — NVIDIA's official repo for the GR00T humanoid foundation models (through N1.7), including fine-tuning recipes (2026).

🔗 Related Awesome Lists


🤝 Contributing

Contributions are welcome! See CONTRIBUTING.md. In short: add your entry to the right section using - [Title](URL) — note (year)., verify the link resolves, and open a PR.

📄 License

MIT for the curation itself. Linked resources remain under their own licenses.


Star ⭐ this repo if it helps you keep up with vision-language-action models.

About

A curated list of Vision-Language-Action (VLA) models for robotics and embodied AI — foundation models, architectures, benchmarks, datasets, focused on 2025–2026.

Topics

Resources

Contributing

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors