Corrections Are the Moat

Most people use Claude like a vending machine. Question in, answer out, walk away. Then they complain the answers feel generic.

They feel generic because a blank chat knows nothing about you, so it gives the median answer for the median person. Every new thread, you pay the same teaching tax.

I stopped paying it this year. My Claude now runs on about 150 skills, and a weekly loop proposes new ones from my own chat history for roughly ten minutes of my time on Monday. Below: how I write a skill, how the loop works, the five setup choices that mattered most, and the trap I walked into anyway.

Let’s go.

🧠 The model is the cheap part

In 2024 everyone chased the perfect prompt: a 300-word incantation that would make the model understand your business forever.

It never existed. A prompt gets read once and thrown away, while your business changes every week and includes a CFO who hates tables.

What compounds is the context you keep. The formula I work from:

Output = model × context × corrections kept

Everyone with a subscription gets roughly the same model. Context is table stakes; anyone can paste a company overview. Corrections kept is the only term nobody can copy, because it’s built from your own mistakes.

🛠 How I write a skill

A skill is an instruction file Claude loads when a trigger phrase matches. “Review the GST sheet” fires one, “council this” fires another. Most of mine started the same way: I corrected Claude on the same thing twice, got annoyed, and wrote it down.

Every skill I write has four parts:

  1. Triggers. The phrases I actually type, pulled from real chats rather than guessed.
  2. A do-not-use line. It names the neighbouring skill that owns the job instead. Without it, two skills fight over one request and the wrong one wins.
  3. A pre-output gate. A short checklist Claude runs before answering: nothing invented, every number sourced, money in a stated currency, drafts never sent.
  4. Rules with their origin attached. “Confirm extraction before any delete” sits next to the week an agent trashed three docs without saving what was in them.

That fourth part does most of the work. A rule with a scar attached gets followed by models and people alike; a bare rule gets skimmed.

The instinct is to keep adding. Resist it. Every few weeks I merge two skills that overlap and delete one nobody triggers, because a small library with clean boundaries beats a big one with turf wars.

🔄 The Correction Loop

Every correction you make is a rule the model should never need again. In a default setup, all of them evaporate when you close the tab.

So I mine them. One scheduled task runs weekly and makes three passes over recent chats:

  • Skill miner. Flags any workflow that showed up twice or took several rounds of correction, ranks candidates by frequency times time saved, checks the existing library, then proposes build, patch, merge, or skip.
  • Instructions miner. Diffs each project’s custom instructions against what actually happened in that project. Stale rules get cut. Repeated patterns with no covering rule get added.
  • Memory pass. New facts in, dead facts out, and the more recent source wins any tie.

Monday morning I get one report and spend about ten minutes approving or killing.

But here’s the kicker: the miner finds workflows I didn’t know I had. I never decided to build a process for auditing meeting notes. I corrected Claude the same way four times, and the miner handed back the process I’d been running on autopilot.

It gets worse:

Self-improving systems have a nasty failure mode. Bad instructions produce bad chats, and mining bad chats produces worse instructions with total confidence.

The fix is boring. Anything backed by a single ambiguous data point gets flagged and waits for my eyes. From the inside, “self-improving” and “self-reinforcing” look identical, so a human checks the direction.

⚙️ Five setup choices that mattered

1. Give everything one home. Messy setups share one bug: the same rule lives in memory and in two different skills, slowly drifting apart. My split is simple. Memory holds facts (client names, standing decisions, the tool I killed in April). Project instructions hold context. Skills hold anything with steps. Preferences hold voice. When I correct Claude now, the fix goes into the most specific place that always loads, and if a rule on that point already exists, I sharpen it instead of adding a second one. Once each rule had one home, Claude stopped contradicting itself between projects.

2. One owner per job. My preferences carry a routing table: rewriting prose goes to one skill, cutting scope to another, Slack drafts to a third, marketing copy to a fourth. When two skills could plausibly take a request, one gets it and the other gets a redirect. Ambiguity is where generic answers come from.

3. Ban the tells. My style doc is mostly a list of things I never want to read, from overused words to sentence shapes that scream “a model wrote this.” It sits at the preference layer, so every skill inherits it. This summer I rebuilt it from about 2,500 words down to roughly 950, using 200 of my own email threads as the evidence for how I actually write. The shorter version gets followed; the long one got skimmed. Editing drafts used to mean rewriting them. Now it’s mostly cutting.

4. Patch before you build. Before creating a new skill, doc, or automation, check whether an existing one can absorb the job. Usually one can.

5. Verify every agent run. A delegated job isn’t done because the agent says so. Every prompt I hand an agent ships with the read-back check I’ll run once it reports, because agents report success on work that half-happened. An unverified run doesn’t count.

🪤 The trap nobody warns you about

Building the system can quietly become the work.

A new skill feels like progress. So does a new agent, a new integration, a new dashboard. None of them ship anything a client pays for, and I caught myself stacking capability instead of using it. In my own notes I call this the optionality prison: building more ways to do the work as a substitute for doing it.

So in August I froze my agent setup. No new runtimes and no new integrations until a fixed review date; anything that touches the freeze gets flagged before it gets built.

Skill count is a vanity metric. The number that matters is how many fired this month and saved an hour doing it.

📌 If I were starting today

What’s the bottom line?

Skip the prompt phase. Pick the task you do most often, write it as a skill the way you’d brief a sharp new hire, and correct the output until the corrections harden into rules. Do that five times.

Then run the loop, even by hand. Scroll last month’s chats and write down everything you had to explain twice. That list is your roadmap, and once a month, delete the skills that never fired.

The models will keep getting better. So will everyone else’s.

Everyone has the same Claude. Nobody has your corrections.

Discover more from Aditya Sheth

Subscribe now to keep reading and get access to the full archive.

Continue reading