Ask your AI assistant to "make an OG image for this post" and you get one of two failures. If it reaches for a diffusion model you get soft gradients, warped typography and a title that looks photographed through water. If it does not, it writes you a perfectly good block of HTML and CSS and then apologises, because it has nowhere to render it.

The second failure is the interesting one. The model already did the design work. Layout, spacing, type scale, brand colour, all expressed precisely in the one design language every LLM has read millions of examples of. It was not missing ability. It was missing a render step.

This post wires that step in over MCP, so your agent can design, render and refine real image assets inside the conversation. It is a condensed version of our full guide, How to generate images from an AI agent with MCP, which also covers VS Code, Windsurf and the other clients.

Why diffusion is the wrong tool

Diffusion models are remarkable for photographic concepts and a lottery for layout. They cannot reliably render text, cannot hit exact pixel dimensions, cannot reuse your brand tokens and cannot produce the same output twice. A pricing card, an invoice, a certificate or an OG image needs all four.