So, you installed an AI skill. Did the work get better?It sounds like too obvious a question to ask. That’s exactly why almost nobody asks it. Installing feels like the accomplishment. A folder arrives, a name shows up in Claude or Codex, and the whole thing has the shape of adding an app to your phone.What moved into your setup was somebody else’s set of decisions about the job: which tools to use, which shortcuts were fine, what counted as a good result, and when the work was done. Those decisions arrived the moment I copied the folder, and I never looked at a single one of them.I found that out with a design skill. I was tired of the same design. You know the look. Terracotta, maroon, a tasteful rounded rectangle, a landing page that’s technically fine and somehow looks like the last six landing pages your AI made. The skill came recommended, and what I wanted was concrete: my front-end work should stop coming out of that same narrow color range. It didn’t. The skill ran exactly as written. It was written by somebody whose idea of good front-end wasn’t mine.An agent skill is closer to a note you leave for a worker who may not ask a follow-up question before getting started. A narrow, mechanical task from a source you trust travels well. Once the note carries taste, business rules, approval boundaries, or your own definition of done, installing it is only the first test.And the collecting has a price you can measure. Both Codex and Claude Code put a hard cap on how much of your skill list the model ever sees, and they start trimming it the moment you cross the line. Twenty-five skills in, your agent is averaging out their conflicts and handing you duller work than it did at five. Nobody tells you when that starts.Here’s what’s inside:The skill builder you already have. The exact command in Codex, ChatGPT Work, and Claude Code, plus the prompt that rebuilds a skill you already installed.Why installing proves nothing. What a skill is under the hood, and how the front door degrades as your library grows.The one-job test. Seven steps from naming the job to rerunning it, ending in one of three decisions: keep it, fork it, or delete it.A test record you can copy. The nine lines that turn “this feels better” into evidence you can still check six months from now.What a crowded library actually costs. Why adding a skill to fix bad output is the loop that made it bad.By the end you’ll be able to take any skill you’ve installed, put one real job through it, and know whether it earns its place.
Twenty-five skills in, your agent is averaging out their conflicts and handing you duller work than it did at five.
A practical test for deciding when to keep a shared skill, rebuild it around your judgment, or remove it.









