Meta released its new Muse Spark 1.2 model along with its first dedicated coding agent. The cheapest tier runs 20 cents per million output tokens, but users pay for it with their data. And the benchmarks have a glaring gap.
Muse Spark 1.2 is primarily a coding upgrade to Muse Spark 1.1, which shipped earlier this year, according to Meta. The company claims improvements in code generation, debugging, and the ability to reason over large codebases. Meta put more compute into training on programming tasks and scaled up the number of training environments.
The model was trained mainly on long-running tasks like generating entire repositories or conducting independent research. To stay on track during these extended sessions, it plans steps ahead, works toward a fixed goal, and compresses its prior context instead of cutting it off.
Some of the training data came from the predecessor model itself. Muse Spark 1.1 generated programming tasks and instruction templates, then scored how well attempted solutions met the requirements. Meta says this process helps 1.2 follow complex instructions more accurately than its predecessor.
Meta cites Terminal-Bench 2.1, DeepSWE v1.1, and 440 tasks from its own codebase as evidence, comparing Spark 1.2 against Grok 4.5, Claude Opus 5, GPT-5.6 Terra (not OpenAI's stronger Sol model), and Gemini 3.6 Flash. Spark 1.2 shows a clear step up from Spark 1.1 but doesn't always close the gap to the top performers.














