Hello, I'm Shrijith Venkatramana, and I'm building LiveReview — a blast-radius aware AI code review built for your business-critical systems. Star us to help devs discover the project, give it a try, and share your feedback to help improve the product.

Imagine you are training an LLM on 4,096 GPUs.

At the end of a training step, 4,095 GPUs have finished.

One GPU is still working.

So what happens?