After a few long years of finding time to document my lessons from training open models, my post-training book is done! It’s published by Manning, under the title Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs.

Telling the story of the book is a useful way to explain why you may want a copy.

The book started as a website where I wanted to document key methods of post-training that had potentially no online material explaining them. If there was something, I couldn’t find it. This existed for more topics than you would expect, given post-training was already popular in 2024 (when I bought the domain), and continues to this day. Topics like rejection sampling, outcome reward models, and character training are prime examples. This has helped make the website fairly popular, as it’s still one of the few places discussing these topics at a foundational, intuitive way.

Otherwise, most of the book is about communicating intuitions and history. Much of the LLM industry is defined by core techniques that haven’t changed much in the last few years. This book was my attempt to explain in simple terms why post-training works, what trade-offs people need to make to get it right, and what misconceptions people often get stuck on. To do this, some of the older, foundational blog posts on Interconnects were reworked to stitch together the story behind key mathematical topics. For this reason, a lot of the explanatory text is likely higher voice than your average textbook.