The first AI model in a product rarely creates an architecture problem.

You build a form, call an API, and display a result.

The second model adds a dropdown and a few conditions. The third adds another input mode. Before long, the product supports text-to-image, image-to-image, text-to-video, image-to-video, reference-based generation, video editing, background removal, and upscaling.

At that point, calling the provider is no longer the hardest part.

The difficult questions are product questions: