Action Models for Robot Learning

by Pavan Kumar Kandapagari · open-access draft, updated continuously

A textbook on the policies that turn perception and language into robot behavior. The book traces the lineage behind today's vision-language-action systems: RT-1, RT-2, OpenVLA, π0, Helix, and GR00T N1. It organizes the field around four families of action models — symbolic planners, geometric and inverse-dynamics controllers, value-based policies learned from reward, and learned policies trained from demonstration — and builds the modern recipe from first principles, one section at a time, ending with what it actually takes to fine-tune and deploy a VLA on a real robot.

New to the underlying math or machine learning? Start with the six appendices: linear algebra, probability, optimization, and the rest of the toolkit.

Start reading → Chapter 1 Browse the table of contents →
Download

Recently drafted

The book at a glance

About this book

This is a textbook for engineers and researchers entering the vision-language-action and robot-learning field in 2026. It's written from first principles and built up, chapter by chapter, to the frontier systems shipping today: RT-2, OpenVLA, π0, Helix, and GR00T N1. I'm drafting it openly, one section at a time; every new section appears on this site the day it's written, and the table of contents shows exactly what's finished and what isn't. The goal is one coherent path from the basics of robot control to fine-tuning your own VLA, with no paywalls and no waiting for the manuscript to be done.