We've been building a new system for running evals against different models, prompts, and harnesses, with the goal of being able to identify the most appropriate small and inexpensive models for different categories of task.

I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models.…

We've been building a new system for running evals against different models, prompts, and harnesses, with the goal of being able to identify the most appropriate small and…