The software wrapper around an AI model has a major impact on what you pay. AI tooling company Composio tested DeepSeek V4 Flash across four agent frameworks (Claude Code, Codex, OpenCode, and Oh My Pi) on 30 tasks using real-world tools like Gmail, GitHub, Slack, and Notion. No single framework won across all categories. Oh My Pi had the highest success rate (17/30) but was the slowest at 272 seconds per task. OpenCode was cheapest at $0.073 per successful task, while Claude Code was fastest at 122 seconds but most expensive at $0.195, despite using the fewest tool calls and generating the least output tokens.
Deepseek V4 Flash tested across four agent frameworks on 30 real-world tasks. Success rates were similar, but cost and speed varied widely. | Image: Composio via X
While seven tasks passed or failed based solely on which framework ran them, overall success rates stayed close. Only OpenCode trailed slightly at 14/30. The real gaps were in cost and speed, with nearly a 3x price difference and a 2.2x speed difference depending on the framework.
AI News Without the Hype – Curated by Humans
Subscribe now










