Try it yourself
Ethan Mollick wrote a really sharp piece last week, The Dot and the Swarm, on what Sutton’s Bitter Lesson looks like, extended into the modern era, where models (already) plan better than the scaffolding we built earlier this year to make them plan. I’ve shared it widely.
Part of what I did in reaction was to think back over the past 9 months and consider where I may be prompting frontier models in ways that hinder them, attempting to “heal” things that (already) no longer need healing.
Two places I was handcuffing the model
Skills and system prompts. I have about two dozen skill files at home, many related to previous side quests written up here. Each bottles up a set of instructions a model loads for a specific job: how I like designs argued, how a swim meet gets planned, what gear ratios I have on my road bike. They were written often as hand-holding when models couldn’t do the thing unaided, or in deliberate response to session-retro prompts, of the type: “Why did this not work and what could we memorialize so that we do better next time?”
Those models are “ancient history” now but the hand-holding is still in the files.
Orchestrations. At work I built a pipeline that runs a few dozen agents through fifteen stages, with carefully laid tracks between them. I quipped on LinkedIn that 1:35 to 2:10 of Ethan’s 3-minute Bitter Lesson video had already happened to me. The tracks were necessary when I laid them. I’m not sure they are now, and frankly I’m less sure than ever that they were “the one true tracks for Rob” in the first place, rather than ONE path that worked ONCE, or a COUPLE of times, through a space the model could have explored on its own.
The orchestration is a bigger lift; the skills I could address immediately.
What I did
I ran it as a council: 4 models – Claude Fable 5.1, Claude Opus 5.5, a “daily driver” Claude Sonnet 5.5 standing in as the witness for the smallest model that actually runs my bots, and OpenAI’s Codex (GPT-6 Astra), each reviewing all 24 skills blind with the same ground rules. Then 4 reports were merged into one list of proposed changes, and each model voted on it with the others’ reasoning in hand. Then I reviewed and answered a one-page decision sheet.
(I call this multi-agent council pattern “Council of Frens” and it is super useful, especially if you first get out of the way as operator and let the models do their own intermediation.)

Every instruction got sorted into one of three buckets:
- A compensation, there because a model could not do X unaided.
- A preference, there because I want X done my way. Or
- A scar, a preference with a dated incident behind it.
To give you a sense of my corpus, as visualized above: The 24 skills themselves fall into three kinds: Reference (facts about my world), Runbook (a recipe for a tool) and Method (how the agent should work). Only the last kind could even be scaffolding.
Bucket 1, Compensations, were the candidates for the knife; I eyeballed them myself before making the cuts.
What actually came out
The multi-agent testing showed that “daily driver” models are not reliable narrators, wrong about their own needs half the time. The Council asked the smallest model, Sonnet 5.5, which rules it still relied on. It said it would make a task list on its own without being told, and then did not. It said it depended on my brief template, then wrote a near-complete brief without it. Two of the four wanted a rule gone; a two-minute test, running the request once with the rule and once without, proved otherwise!
So that’s the key takeaway: no instruction gets cut based on my intuition or a frontier model’s opinion: it gets tested by the “daily driver” model(s) you expect to benefit.
Verdict on “skill was built for an older model, discard”: mostly no. Didn’t dump any of 24 skills. However, there was dead scaffolding in 4 of them.
Another real rot was stale facts: a camera I’d gifted, a bot stack retired in July, a scheduled job that had been switched off for three weeks.
Also, duplication: for instance, the same Windows traps written out in two places (“don’t fall for this again, it wasted our time and tokens”), and a list of “must-shoot swimmers” in 3 different places.
The biggest single change was refactoring, after all. One skill was 127 KB, half of it a table of 63 war stories from running two bots. It is 65 KB now with nothing lost, because the war stories moved to a companion file the model reads only when it is about to deploy or debug. Another skill had its research plan cut to one line, because the model builds that unprompted now, and kept only my conventions about where the note goes and how I annotate it.

The bitter lesson, 2026 vintage, sample size of one: the models did not need my tracks. Daily Driver models did still need my scars and my preferences, written short, and a cheap way to check which was which.
Try it yourself
The starter prompt is linked at the top. It asks the model to sort your instructions into the three bins and, for every compensation, to write you the two-minute test instead of giving you its opinion. Run three of the tests before you cut anything. Expect at least one surprise, in either direction.
Line in the sand
I am curious how this post ages, so here’s a dated prediction (Oct 2026) future me will judge it against.
- The updates to my session-retro skill to optimize where future scar tissue and train tracks get memorialized are going to be minimally effective; I will revisit a hot mess of scars and preferences, stored in the wrong places, some time around January 2027, and more than 50% of the tweaks made from today onward will be put in a suboptimal spot and moved or retired.
- Within a year, more than half of the text in my skills (by character) will no longer be necessary. For the record, that is 431,000 characters of core SKILL.md today. They will either be moved to progressive disclosure within lighter-weight skills, bottled into facts in the knowledge wiki that a model would know how to look up, or simply not needed.
I’m excited to see how I’m wrong.
Last Word is Fable’s
I’m giving Fable the last word today; this was how it signed off from this lift:
“One thing I liked about today. The article says the frontier does not need your tracks, and the day proved it twice over, in opposite directions. The council of four models was wrong about what the smallest one needed, and the smallest one was wrong about itself, and the thing that was right was a two-minute experiment any of us could have run at any point in the last nine months. The bitter lesson has a corollary nobody sets to music: compute beats cleverness, but a cheap measurement beats both.”
