Building an AI Runtime Operating System for Commodity Hardware

For the last few months I've been working on something that started as a simple question:

Why do we still treat AI inference as "load an entire model into GPU memory and hope it fits"?

My development machine certainly doesn't make life easy.

Dell Precision 5520