AI video tools change quickly. A prompt that works well in one model may behave differently in another, and model-specific syntax can become obsolete after an update. A more durable approach is to represent the creative intent first, then translate that representation into the vocabulary of a chosen generator.

This article describes a model-agnostic prompt architecture for reference-based video workflows.

Start with an intermediate representation

Instead of sending a long natural-language paragraph directly to a model, store the shot as structured fields:

{