About

An open-access textbook on action models: the policies that turn perception and language into robot behavior, drafted in public, one section at a time.

The book organizes the field around four families of action models: symbolic planners, geometric and inverse-dynamics controllers, value-based policies learned from reward, and policies learned from demonstration. From there it builds the modern recipe behind RT-2, OpenVLA, π₀, Helix, and GR00T N1 from first principles. It doesn't survey the literature, and it doesn't just retrace one lab's recipe. The table of contents shows what's finished and what's still pending.

I am Pavan Kumar Kandapagari. I lead a foundation-models-for-robotics team in Munich. For the past three years I've been working toward production vision-language-action policies, moving through LLM-as-planner systems, imitation-learning baselines from RT-1, RTX, and Octo, and diffusion- and flow-matching action decoders, before designing and pretraining the architecture our team now ships. Along the way I built and hired the fifteen-person research, infrastructure, and evaluation team behind that work, plus the hardware benchmarking suite we use to test our policies on real robots against open-source baselines like π₀ and GR00T. This book is the long-form version of the notes that work produced.

I write this in the open because the field moves fast enough that drafts read by strangers teach me more than a manuscript polished in private ever would, and because every section that survives a public read is one less section I have to revise later. No paywall exists, and none is planned.

The book lives on GitHub at github.com/kandapagari/in-action-book — issues and pull requests are welcome. You can also reach me at pavan.kandapagari@gmail.com.