Most frameworks for AI-assisted development try to own the whole process — hand over a spec, get back an app, trust the pipeline in between. Matt Pocock's skills repo, built by the well-known TypeScript educator behind Total TypeScript and AI Hero, takes the opposite position: small, composable Agent Skills that fix specific, named failure modes in AI-assisted engineering, without taking control away from the person actually building the thing.
If you use Claude Code or Codex every day, you have probably seen the same problems repeat: the agent builds something slightly different from what you meant, explains itself at length, ships code that looks right but fails, and leaves the codebase a little more tangled each session. Matt Pocock's skills are aimed squarely at those problems. The bigger question for most teams is whether a set of small, inspectable skills is a better fit than a framework that promises to run the whole process for them.
This article walks through the argument against end-to-end frameworks, the four failure modes the repo targets and the skills that address each, why the shared-vocabulary technique is worth adopting on its own, how the TDD and architecture skills work, the difference between user-invoked and model-invoked skills, first-time setup, the two installation routes, and what all of this means in practice for a team adopting it.
The Argument Against End-to-End Frameworks
The README makes its positioning explicit: approaches that try to own the whole development process end up making it hard to actually debug what goes wrong, because you've traded visibility for automation. Pocock's skills are built to be small, easy to adapt, and composable instead — designed to be read, understood, and modified rather than trusted as a black box. That's a meaningfully different bet than a framework promising to handle requirements-to-deployment in one pass; it's closer to a toolbox of specific fixes for specific problems than a replacement for engineering judgment.
Four Named Failure Modes, Four Fixes
Rather than a generic "best practices" list, the repo is organized around four concrete problems the author has actually observed, each with a specific skill built to address it:
| Problem | Fix | Skill |
|---|---|---|
| The agent builds the wrong thing | A structured "grilling session" — the agent asks detailed clarifying questions before writing code | /grill-me, /grill-with-docs |
| The agent is unnecessarily verbose | A shared project vocabulary that replaces long descriptive phrases with concise, agreed-on terms | Built into /grill-with-docs |
| The code doesn't actually work | Enforced feedback loops — static types, browser access, and a red-green-refactor TDD discipline | /tdd, /diagnosing-bugs |
| The codebase turns into unmanageable complexity | Periodic architecture surveys that surface concrete refactoring candidates | /to-spec, /improve-codebase-architecture |
The alignment skills are described as the most-used in the set, and the reasoning behind them is straightforward: most software misalignment between a person and whoever's building for them — human or agent — comes from an unstated gap in understanding, not a technical failure. A structured round of clarifying questions before code gets written is a cheap fix for an expensive problem, and it's notable that the fix here is a conversation discipline, not a new tool.
The Shared-Vocabulary Trick Is Worth Understanding on Its Own
One detail is worth calling out specifically because it's a genuinely transferable idea outside this specific toolset: building a project-specific glossary document that replaces long, jargon-laden descriptions with short, agreed-on terms doesn't just make output more concise — it makes an agent's variable, function, and file names more consistent (since they draw from the same shared vocabulary), makes the codebase easier for the agent to navigate on later sessions, and reduces the token cost of the agent's own reasoning because it has a more compact language to think in. That's a nice example of a very simple intervention — write down a glossary — compounding into several unrelated benefits.
The Feedback-Loop and Architecture Discipline
The /tdd skill enforces a red-green-refactor loop specifically because agents given no feedback signal on whether their code actually runs tend to produce code that looks plausible and doesn't work — a failing test written first gives the agent a concrete, checkable target rather than a vibe to match. The /diagnosing-bugs skill wraps debugging into a disciplined, phase-gated loop rather than letting an agent guess at fixes; a recent release specifically hardened this skill to redact secrets from captured command output and artifacts by default, which is a small but meaningful detail for a skill whose whole job is showing terminal output back to a human.
On the architecture side, /improve-codebase-architecture is explicitly described as a survey tool, not a rescue tool — it's built to periodically scan a codebase for concrete "deepening" opportunities (per the module-design principle that the best modules expose a lot of functionality through a simple interface) and hand back real candidates, but the README is direct that on a genuinely tangled legacy codebase, it will find problems without untangling them for you. That's an honest scoping of what an automated architecture review can actually do.
The Full Reference: User-Invoked vs. Model-Invoked
The repo's skill list splits on one axis worth understanding before adopting any of it: who's allowed to invoke a given skill. User-invoked skills only run when you type them yourself — /grill-me, /tdd, /to-spec — and their job is to orchestrate a workflow. Model-invoked skills can be triggered the same way, but the agent can also reach for them automatically mid-task when the situation calls for it — they hold the reusable discipline that a user-invoked skill leans on. The rule that keeps this from turning into a tangle: a user-invoked skill can call a model-invoked one, but never another user-invoked skill directly.
Beyond the four flagship skills already covered, the engineering set includes several more worth knowing about: /ask-matt routes you to the right user-invoked skill when it's not obvious which one fits; /triage moves issues through a defined state machine of roles; /to-tickets breaks a plan into tracer-bullet tickets with explicit blocking dependencies; /implement builds out a spec, driving /tdd at pre-agreed seams and closing with /code-review before anything gets committed; and /wayfinder is built for work too large for a single agent session — it lays out a shared map of decision tickets and resolves them one at a time until the path to the destination is clear.
On the model-invoked side, a few are notable beyond /tdd and /diagnosing-bugs: /prototype builds a throwaway artifact to answer a design question before real code gets written, /research investigates a question against primary sources and writes the findings to a cited file in the repo, /resolving-merge-conflicts works an in-progress git merge or rebase hunk by hunk rather than reaching for --abort, and /wizard generates an interactive bash script that walks a human through steps only a person can do — provisioning infrastructure or running a one-off migration. A separate, non-code-specific productivity set includes /handoff (compact a conversation into a document another agent can pick up cold), /teach, /to-questionnaire, and /wait-what — all four sit on top of /grilling, the same interview primitive behind /grill-me.
Setting Up a Repo for the First Time
Installing the skills is only step one — the README calls for a required second step before the engineering skills are usable in a given repository: running /setup-matt-pocock-skills once. It asks three concrete questions — which issue tracker to use (GitHub, Linear, or local files), what labels get applied during /triage, and where generated docs should be saved — and wires those answers into every other skill that depends on them. Skipping it is the likely reason a skill like /triage would behave unexpectedly on a fresh install.
Two Installation Philosophies
The repo ships two distinct ways to get the skills, and they represent genuinely different philosophies rather than just two download options:
- Claude Code plugin (
claude plugins install mattpocock-skills) — installs the whole set as a managed, read-only bundle from Claude Code's official marketplace that updates automatically when the author ships changes. You subscribe rather than fork. - skills.sh (
npx skills@latest add mattpocock/skills) — copies editable skill files directly into your project, so you own and can modify them, withnpx skills updateto pull upstream changes on your own schedule.
The README is direct that installing both leaves you with every skill twice — pick the philosophy that matches whether you want to hack on the skills or just consume updates passively.
Benefits of Matt Pocock's Skills
The individual skills are small, but together they change how an AI-assisted session runs. These are the benefits teams are most likely to notice.
Fewer rebuilds of the wrong thing
The grilling skills move the hard conversation to the start of a task. Instead of discovering halfway through a pull request that the agent misunderstood the requirement, you answer pointed questions before any code exists. Each clarifying question costs a few seconds; each wrong implementation costs a session of rework. For teams whose main frustration is the agent confidently building something adjacent to what they meant, this is the most immediate gain.
Code that has been shown to work
The /tdd skill gives the agent a failing test to turn green rather than a description to approximate. That changes the definition of done from "looks plausible" to "passes a check written before the code." Combined with type checks and browser access, it closes the gap between code that reads well and code that runs, which is where a lot of hidden AI-generated bugs live.
Shorter, more consistent output
A shared glossary gives the agent and the team the same compact terms. Explanations get shorter, names in the code become consistent across sessions, and the agent spends fewer tokens describing concepts it could simply name. Those effects build on each other: consistent names make the codebase easier for the next session to navigate, which makes the next session more efficient again.
Visibility instead of a black box
Because each skill is a small, readable file, you can see exactly what it tells the agent to do. When something goes wrong, you can inspect the skill, change it, or stop using it, rather than debugging an opaque pipeline. That makes the toolkit easier to trust and easier to adapt to house conventions, particularly with the editable install.
Gradual adoption without changing your workflow
The skills sit on top of the agent and process you already use. A team can start with one alignment skill on one project and add others when they prove useful, without migrating to a new framework or retraining everyone. That keeps the cost of trying it low and makes it easy to drop anything that doesn't earn its place.
Matt Pocock Skills Use Cases
The repo's skills map onto everyday situations in AI-assisted development. These are the most common places they fit.
Starting a feature with an unclear brief
When a request is vague, such as a ticket that says "add export" with no detail, /grill-me or /grill-with-docs has the agent question you about formats, edge cases, permissions, and scope before writing code. The answers can be captured in project docs, so the next session starts from the agreed understanding. The outcome is a first implementation much closer to what was actually needed.
Fixing a bug without guesswork
Agents asked to fix a bug often change several things at once and hope one of them works. /diagnosing-bugs structures the work into phases: reproduce, isolate, explain, then fix. Captured command output is redacted for secrets by default, which matters when the debugging session is shared with a teammate. The result is a fix with a clear explanation of the cause, rather than a patch nobody fully understands.
Building new code test-first
For new functionality with clear inputs and outputs, /tdd runs a red-green-refactor loop: write a failing test, make it pass, then clean up. /implement uses the same discipline across a whole spec, driving tests at agreed points and ending with a code review before anything is committed. Teams that already value tests get an agent that follows the same habits instead of skipping them.
Reviewing a codebase that has grown messy
After months of fast, AI-assisted changes, a codebase can drift toward shallow modules and tangled dependencies. /improve-codebase-architecture surveys the code and returns concrete candidates for deepening modules. It doesn't do the refactoring for you, but it turns a vague sense that the code is getting worse into a list of specific places to invest engineering time.
Work too large for one session
Large changes, such as a migration or a multi-part feature, exceed what a single agent session can hold. /to-tickets breaks a plan into tickets with explicit dependencies, /wayfinder keeps a shared map of decisions to resolve one at a time, and /handoff compacts a conversation so another agent or person can continue cold. Together they make long-running work manageable without losing context between sessions.
Matt Pocock Skills Best Practices
These practices help a team get value from the skills without turning them into another layer of process.
- These are additive to whatever agent and workflow you already use, not a replacement framework — worth adopting incrementally, starting with the alignment skills, rather than all at once.
- The shared-vocabulary technique is worth stealing even without adopting the whole skill set — writing a project glossary is a low-cost habit any team doing AI-assisted development can adopt directly.
- The
/improve-codebase-architecturesurvey is scoped honestly — treat it as a candidate-finder for a legacy codebase, not an automated cleanup, and budget real engineering time for what it surfaces. - Choose the plugin or the editable-copy install deliberately based on whether your team wants to track upstream changes passively or actively customize the skills for house conventions. Installing both leaves every skill duplicated, so agree on one route across the team.
- Run the setup step on every repository before judging the skills.
/setup-matt-pocock-skillsrecords your issue tracker, triage labels, and docs location. Skipping it is the most likely reason a skill behaves strangely on a fresh install, and it takes only a few minutes. - Keep the glossary in the repository and review it like code. Treat the shared vocabulary as a living document: add terms when new concepts appear, remove ones that fall out of use, and review changes in pull requests so the whole team stays aligned with what the agent is using.
- Read each skill before relying on it. The files are short. Knowing exactly what a skill instructs the agent to do makes it much easier to spot when it doesn't fit your codebase, and to adapt it if you use the editable install.
- Measure rework before and after. Note how often agent output had to be redone or heavily corrected on a project before adopting the skills, then track the same thing afterwards. A simple before-and-after comparison is enough to decide which skills are worth rolling out more widely.
Common Mistakes When Adopting Matt Pocock's Skills
The skills are simple, which makes it easy to misuse them in ways that cancel out their benefits.
Installing everything at once
Adding the full set on day one gives the team dozens of new commands to learn and no clear sense of which ones are helping. It also makes it hard to tell whether any improvement came from the skills or from something else. Start with the alignment skills on one project, observe the effect on rework, and add others deliberately.
Rushing through the grilling session
The value of /grill-me comes from answering its questions carefully. Developers in a hurry give one-word answers or tell the agent to make assumptions, which brings back exactly the misalignment the skill is meant to prevent. If a task isn't worth a few minutes of clarification, it probably doesn't need the skill; if it is, give the questions real answers.
Expecting the architecture survey to clean up the code
/improve-codebase-architecture finds candidates. It does not untangle a legacy system, and the README says so plainly. Teams that run it expecting an automated refactor end up disappointed, or worse, let the agent make sweeping changes without a plan. Treat its output as a prioritised list for engineers to act on with normal review.
Letting the glossary go stale
A shared vocabulary only helps if it matches the code. When the glossary describes concepts that have been renamed or removed, the agent uses outdated terms and the consistency benefit turns into confusion. Make updating the glossary part of the change that introduces or renames a concept.
Installing both the plugin and the editable copy
Using the marketplace plugin and the skills.sh copy together gives you every skill twice, possibly at different versions. The agent may pick either one, and changes you make to the editable copy may not be the ones that run. Pick one installation route per team and remove the other.
Practical Takeaway
This skill set is a useful counterpoint to any spec-driven or end-to-end AI development framework: instead of trying to own the whole pipeline, it targets four specific, well-understood failure modes with small, inspectable fixes — and it's built by someone with a long, public track record in software engineering education, which shows in how precisely each problem is named before a fix is offered. For teams comparing it against Diagram Design, Hallmark, or other Agent Skill collections, the differentiator here is scope: engineering process discipline rather than design or diagram generation.
Teams building house-specific Claude Code or Codex skills, or evaluating engineering-discipline tooling for AI-assisted development, can get hands-on help from Woyce Technologies.
FAQ
What is Matt Pocock's skills repo?
It's an open-source collection of Claude Code and Codex Agent Skills built by TypeScript educator Matt Pocock, targeting specific engineering failure modes — misalignment with requirements, agent verbosity, unreliable code, and codebase complexity — rather than trying to automate the whole development process end-to-end. Each skill is small enough to read and adapt, and the alignment skills are a sensible place to start.
How is this different from frameworks like GSD, BMAD, or Spec-Kit?
Those frameworks try to own the entire development process from spec to deployment. Pocock's skills are deliberately small, composable, and designed to be read and modified — they fix specific problems without taking control of the overall workflow away from the engineer. The practical difference shows up when something goes wrong. With an end-to-end pipeline, it can be hard to see which stage caused a bad result. With small skills, each one has a narrow job you can read, test, and change, so debugging the process stays manageable.
What does the "grilling session" skill do?
/grill-me and /grill-with-docs have the agent ask detailed clarifying questions about what you're building before it writes any code, closing the alignment gap that causes an agent to build something different from what you actually wanted. The /grill-with-docs variant also records the agreed terms in a shared vocabulary document, so later sessions start from the same understanding. The idea is that a few minutes of structured questions costs far less than rewriting code built on a wrong assumption.
Is Matt Pocock's skills repo free to use?
Yes, it's MIT-licensed and open source, installable as a managed Claude Code plugin from the official marketplace or as editable files via the skills.sh installer. The skills themselves cost nothing, but running them still uses your normal Claude Code or Codex plan, so the real cost is the model usage of the sessions in which you invoke them, plus the time spent adapting them to your team's conventions.
What does the TDD skill enforce?
A red-green-refactor loop, where the agent writes a failing test first and then makes it pass — giving it a concrete, checkable feedback signal instead of producing code that looks plausible without verification that it actually works. The value is less about testing as a ritual and more about giving the agent a clear definition of done. Static types and browser access play a similar role in the repo, adding further ways for the agent to check its own output.
Can I edit these skills to fit my own team's conventions?
Yes, if installed via the skills.sh method rather than the Claude Code plugin — that route copies the skill files directly into your project as ordinary, editable files you own, rather than a managed read-only bundle. You can then pull upstream changes with npx skills update when it suits you and resolve any conflicts with your edits. The plugin route suits teams that want automatic updates and no maintenance; the editable route suits teams with house conventions they want built into each skill. Installing both duplicates every skill, so choose one.
What's the difference between a "user-invoked" and a "model-invoked" skill in this repo?
User-invoked skills, like /grill-me or /tdd, only run when you type them and are meant to orchestrate a workflow. Model-invoked skills can also be reached for automatically by the agent when a task calls for them, and hold the reusable discipline that user-invoked skills lean on. A user-invoked skill can call a model-invoked one, but never call another user-invoked skill directly.
Do I need to configure anything before using the engineering skills?
Yes — run /setup-matt-pocock-skills once per repository first. It asks which issue tracker you use, what labels your team applies during triage, and where generated documentation should live, and every other engineering skill in the set depends on those answers being set. Skipping this step is the likely reason a skill like /triage behaves unexpectedly on a fresh install.
What does /wayfinder do?
It's built for work too large for one agent session to hold: it lays out a shared map of decision tickets on your issue tracker and resolves them one at a time, so a large piece of work stays navigable across many sessions instead of getting lost. Use it when a task is too big to finish in one sitting.
Conclusion
AI coding agents fail in recognisable ways: they build the wrong thing, say too much, ship code that has never run, and add complexity faster than anyone removes it. Matt Pocock's skills repo treats each of those as a separate, named problem and answers it with a small, readable skill rather than a framework that takes over the whole development process.
The most useful parts are also the simplest. A structured grilling session before any code is written closes most alignment gaps. A shared project vocabulary makes agent output shorter and naming more consistent. A red-green-refactor loop gives the agent a concrete definition of done. Even if you never install the repo, those three habits transfer to any agent workflow.
Keep the scope in mind. The architecture survey finds candidates but does not untangle a legacy codebase for you, the setup step is required before the engineering skills behave properly, and the plugin and editable installs suit different teams. Start with the alignment skills on one project, measure how much rework they save, and add the rest gradually. If you want help designing house-specific skills or an AI-assisted engineering workflow for your team, talk to our LLM integration team.
