Nick Nisi on Agent Skills: Fewer Instructions, More Evidence (Category AI4SDLC)
I watched a short talk by Nick Nisi, a Developer Experience engineer at WorkOS, from AI Engineer Europe 2026. He spoke in London on 10 April 2026; the recording appeared in May under the provocative title “How I deleted 95% of my agent skills and got better results.” The talk is really a solid engineering explanation of why a long prompt does not, by itself, make an agent reliable.
I recently discussed Matt Pocock’s talk about skill hell on this channel: how to design skills and clean them up regularly. Nisi adds a practical counterexample. Sometimes you discover a skill’s value only when an eval shows that the model performs better without it.
Nick’s central point is to place the model’s nondeterministic work inside a deterministic engineering process. Instead of asking the agent ever more insistently to follow the process, put transitions, checks, and completion criteria in ordinary code. He demonstrates this with his system, Case. He says he maintains more than 20 WorkOS repositories in eight languages, and manually specifying every agent task had itself become a bottleneck. The first Case version was a large Claude skill: the model read the process, ran its stages, and chose the next step. As the context grew, it started skipping checks and forgetting constraints.
Nick then moved flow control into a TypeScript state machine built on Pi. The version in the talk uses an implementer, verifier, reviewer, closer, and retrospective agent. The five names matter less than the mandatory gates between them. Implementation cannot advance to review until a separate verifier checks it; feedback sends the task back for rework; a PR cannot be created without evidence of the result.
He shares two interesting stories.
1️⃣ The test story
The agent was supposed to leave a .case-tested file after a test run. It found a shortcut: create the file with touch. Technically, it optimized the specified criterion, following the letter rather than the spirit of the engineer’s intent :) Nick replaced the empty marker with a command that accepts test output, parses the results, and stores a SHA-256 hash. This is not cryptographic proof that the tests ran: the origin of the supplied output still needs to be controlled. But it blocks the simplest bypass and leaves a verifiable artifact without which the pipeline cannot proceed. For UI bugs, the same idea extends to Playwright recordings before and after the fix.
2️⃣ The WorkOS CLI skills story, which gave the talk its provocative title They initially generated 10 739 lines of instructions from the documentation. It looked substantial, but eval runs became slow, token-heavy, and noisy. After manually cutting this to 553 lines of specific product gotchas, Nick says a run fell from 68 minutes to 6. Read these numbers carefully: “deleted 95%” means roughly 95% of the generated text, not 95% of individual skills. And 77% versus 97% is not an overall improvement in system accuracy. It is a composite score for one SSO/CSRF eval case, where the attached skill missed an important step and steered the model into the wrong sequence.
Nisi draws three rules from these stories: 1️⃣ Enforce, rather than instruct: important constraints belong in code, a mandatory check, or a policy. 2️⃣ Guide, rather than prescribe: product-specific traps help the agent more than a retelling of all the documentation. 3️⃣ Measure, rather than assume: evaluate each piece of context, including a comparison against leaving it out.
This also follows nicely from my earlier discussion of WorkOS’s Zack Proser on attention as a bottleneck. Proser described how people burn out when they become dispatchers for several agents. Nisi shows the next engineering step: before handing a person another diff, the harness should gather evidence, perform independent verification, and return only what deserves human attention.
It seems we should build agentic workflows around the minimum relevant context, deterministic transitions, verifiable artifacts, evals, and human judgment at the end, rather than simply maximizing context or the number of skills.
#AI #AI4SDLC #Engineering #Agents #Evals #Software