In the last post, on 2026-08-18, I published what 14 MCP servers cost a context window before an agent does any work, and said the next thing was Tier 2: real clients, real tasks, every frame logged. That has now run. Ninety trials, three servers, two clients, fifteen scripted tasks, three trials each, through a proxy on the MCP stdio pipe.
The finding worth the post is not in the token table. It came out of the shakedown hours earlier that same day, on the client version the run then replaced: one of the clients was failing calls inside itself, before a byte reached the server, and on the wire that looked exactly like a model that tried very little and answered wrong.
87 of 90 trials completed, and the three that did not failed three different ways
The matrix is filesystem, playwright and github, five scripted tasks each, three trials per task per client, suite version 1.0.1. Ninety trials, 87 successes. Servers pinned at @modelcontextprotocol/server-filesystem@2026.7.10, @playwright/mcp@0.0.79, and the ghcr.io/github/github-mcp-server container, launched untagged, which reported itself as v1.9.0. Clients: Claude Code 2.1.235 on claude-sonnet-5, Gemini CLI 0.55.1 on gemini-2.5-flash, both model ids read back out of each trial's own client JSON rather than assumed from the flag. Claude Code also invokes claude-haiku-4-5 on every trial for its own bookkeeping, under a thousand input tokens a time. That never touches the MCP pipe and is in none of the figures below, but it is in the manifest, so it is worth knowing it is there.






