Visualizing AI evals with Inspect Viz
Welcome back to our blog series on running, analyzing and visualizing AI evals. Last time we discussed how to design and run evals for analysis and visualization using inspect eval and inspect view. Many of these methods, such as rearranging the dashboard columns and sorting on them to determine model vs model and skill vs no-skill differences in metrics, can give you a broad overview of the field. While this can inspire further and deeper inquiry, it runs into issues with how predictive it is and how to communicate findings to other people. Imagine the following:
The 40-log gridlock
Imagine a data science lead needing to present a performance breakdown of four LLM candidates across ten internal tools during a high-stakes, time-crunched live meeting. Reordering and filtering the evals across multiple dimensions is hard to read on a presentation screen and requires doing mental math with an audience: riveting stuff that they definitely won’t fall asleep during of course.
What if instead you had a tool that allowed you to easily automate rendering a comparison of the different configurations? What if lining them up by model or by skill showed clear and legible patterns?






