Reading a good business or strategy book and actually applying what's in it months later is a genuinely hard problem — most of what gets read ends up compressed into a highlight or a note that never gets revisited. Cangjie Skill is built around a specific bet on how to fix that: instead of producing another summary for a human to read, it runs long-form content — books, podcast transcripts, interviews, long videos — through a structured pipeline that outputs installable Agent Skills, so the methodology inside a source becomes something an AI agent can actually invoke in a real task.
That matters for two groups. Individuals who read a lot want the frameworks they read to show up when they're working, not sit in a notes app. Teams have the same problem at larger scale: playbooks, onboarding guides, and recorded expert talks that nobody rereads. Cangjie Skill treats both as a distillation problem with a quality bar, which is a different idea from summarization, and close in spirit to related projects like Book-to-Skill.
This explainer covers the seven-stage RIA-TV++ pipeline, why its aggressive filtering is the most interesting design choice, what a finished run produces, how it handles long and interrupted jobs, how it compares with single-prompt summarization, how to get started safely, and the copyright question it raises, which deserves more attention than most coverage gives it.
A Seven-Stage Extraction Pipeline
The project calls its method RIA-TV++, and it's a genuinely structured process rather than a single prompt asking a model to "summarize the key points." Seven stages run in sequence: whole-content comprehension using a structural/interpretive/critical/applicability framework adapted from Mortimer Adler's analytical reading method; parallel extraction by five specialized extractors pulling candidate frameworks, principles, cases, counter-examples, and terminology; a triple-verification filter that only keeps a candidate methodology if it has independent supporting evidence in at least two places in the source, can answer a genuinely new question the source doesn't explicitly address, and isn't just common knowledge restated; a structuring step organizing surviving candidates into six dimensions (a quoted excerpt, a rewritten explanation, an example from the source, a projected future use case, executable steps, and known boundaries); a linking stage that maps relationships between the resulting skills; a pressure-testing stage that runs each skill against trick questions and cross-skill confusion tests before it ships; and a final delivery stage that installs anything that passed into a Claude Code or Cursor skills directory.
The Filtering Is the Actual Point
The detail worth taking seriously here is how aggressive the rejection rate is: the project states that only 25-50% of candidate methodologies survive the triple-verification stage. That's a deliberate design choice against the more common failure mode of AI summarization tools — producing a comprehensive-looking output that includes everything, most of which turns out to be restated common sense or context-dependent detail that doesn't generalize. A pipeline that throws away half or more of what it extracts, specifically because it can't prove the surviving content is both independently supported and non-obvious, is making a real trade-off in favor of precision over coverage.
What the Output Actually Looks Like
The delivery isn't a single file — it's a small repository. A completed run produces BOOK_OVERVIEW.md as the whole-content understanding from stage one, INDEX.md as a skill map showing how the extracted skills relate to and depend on each other, individual SKILL.md files for each surviving methodology, a GLOSSARY.md of source-specific terminology, and test-prompts.json recording the trigger scenarios each skill was validated against. Version 2.0.0 added a fifth delivery stage on top of that: a reader-facing DIGEST.md — a long-form summary aimed at a human who wants the substance without reading the entire source — generated alongside the machine-facing skill files, plus the actual installation step that copies validated skills into a Claude Code or Cursor skills directory so they're immediately callable rather than sitting in a repo waiting to be wired up manually.
Built for Long Runs That Might Get Interrupted
Distilling an entire book or a multi-hour video transcript is not a quick job, and version 2.0.0 added infrastructure for that reality directly: a PIPELINE_STATE.md file that tracks which stage is active, what's been produced so far, and the status of each individual skill candidate, so a run can resume where it left off rather than restarting from scratch after an interruption. The same release added chunking strategies for unusually long source material — splitting along natural boundaries like chapters, volumes, or episode breaks — with a serial fallback for cases where parallel extraction across chunks isn't practical. The quality gates got stricter at the same time: a lightweight user confirmation step, blind testing by an independent agent that didn't see the extraction process, cross-skill confusion tests specifically designed to catch two similar skills stepping on each other, and a length cap on quoted excerpts in translated output — a direct, practical response to the copyright question the extraction format raises in the first place.
A Growing Library of Published Skill Packs
The project's own repository lists more than twenty published skill packs built with the pipeline, spanning Warren Buffett's shareholder letters, Charlie Munger's Poor Charlie's Almanack, Influence, Contagious, business and investment texts, a 165-prompt system-prompt design collection, and classical Chinese texts including the Huangdi Neijing and the Sunzi Bingfa — each contributing anywhere from six to twenty-five individual skills depending on how much extractable methodology the source actually contained. A handful of these are external contributions brought in with the original authors' consent, worth noting both as evidence the pipeline generalizes beyond its own author's use and as a reminder that "consent-based" isn't the same as "rights-cleared" when the source material is itself under commercial copyright.
Part of a Named "Distillation" Ecosystem
Cangjie Skill positions itself as one piece of a small, explicitly connected ecosystem: a companion project distills specific people's thinking styles and communication patterns into "human skills," and a separate evolution tool is meant to keep any of these skills current over time as they get used. Cangjie's own scope is content people have written or spoken systematically — the idea being that a book or a long, deliberate interview represents a different kind of distillable value than imitating how a specific individual talks.
Benefits of Cangjie Skill
The pipeline is more expensive than a summary, so the benefits need to be concrete. These are the ones that follow directly from its design, rather than from any claims about results.
Methods Show Up During the Work
A summary waits to be reread; a skill gets invoked. When a framework from a book or an internal playbook is installed as a skill, an agent working in Claude Code or Cursor can apply it in the task where it is relevant, without anyone remembering where the note was saved. That moves knowledge from a reference shelf into the workflow, which is the problem the project set out to solve.
Less Noise in What Gets Kept
The triple-verification filter discards most candidates, keeping only those with evidence in more than one place, transfer to new questions, and genuine non-obviousness. The result is a smaller set of skills that each carry weight, rather than a long list mixing real methods with restated common sense. For an agent choosing which skill to use, fewer and sharper options also reduce the chance of picking the wrong one.
Skills That Know Their Limits
Each skill carries six dimensions, including executable steps and known boundaries. The boundaries matter as much as the steps: they tell the agent, and the person reviewing it, where the method stops applying. Free-form summaries rarely state when advice does not hold, which is how good frameworks get misapplied.
Testing Before Anything Ships
Trick questions, blind testing by an independent agent, and cross-skill confusion tests catch problems that would otherwise surface in real use. A skill that fires on the wrong prompt, or two skills that overlap, gets caught before installation. The generated test prompts also give teams a starting regression suite when they regenerate skills later.
Long Sources Become Practical
Chunking by chapter or episode and a resumable state file make it realistic to distill a full book or a long lecture series. An interrupted run picks up where it stopped instead of repeating expensive model calls, which matters for the dense, long sources where the pipeline is most worth using.
Cangjie Skill Use Cases
The pipeline suits any source that teaches a repeatable method. These are the situations where it fits best, starting with the clearly safe ones where you own the content outright.
Internal Playbooks and Process Documents
Sales playbooks, incident runbooks, and onboarding guides often sit in a wiki nobody rereads. Running them through the pipeline turns the methods inside into skills an agent can apply while people work, with steps and boundaries spelled out. Because the organisation owns the content, the copyright question does not arise, and skills can be regenerated whenever the source document is updated.
Recorded Expert Talks and Interviews
Teams often have hours of recorded internal talks, customer interviews, or workshops where an experienced person explains how they approach a problem. Once transcribed, those recordings can be distilled into skills that capture the method rather than the anecdotes. The filter is useful here, because long conversational transcripts contain a lot of material that does not generalise.
Personal Reading on Licensed or Public-Domain Material
Individuals who read heavily can distill public-domain texts, or works they have clear rights to, into a private skill library that keeps frameworks available in their daily tools. Keeping the output private and limited to permitted sources is the line that matters; publishing packs built from commercial books is the case the copyright section below addresses.
A Reference Design for Skill-Building Pipelines
Even teams that never run the tool can borrow its structure. The verification filter, six-dimension format, and pressure-testing stage are a blueprint for any process that turns institutional knowledge into agent skills, whether the source is documentation, code review guidelines, or support macros.
Onboarding New Team Members
New hires usually learn a team's methods by reading documents once and then asking colleagues. Distilling onboarding material and recorded walkthroughs into skills means the methods are available inside the coding agent a new engineer is already using, at the moment they need them. The executable steps and stated boundaries also give reviewers something concrete to check when a newcomer's work drifts from team practice, and the reader-facing digest doubles as a structured onboarding read.
Common Cangjie Skill Mistakes
The pipeline is disciplined, but it cannot stop users from making these errors around it. Most of them come from treating the output as finished rather than as a draft that still needs a human owner.
Running It on Content You Do Not Hold Rights To
The output includes direct quotes, and the project has been applied to commercially sold books. Running it on such material, and especially publishing the resulting pack, raises real derivative-work questions. Treating a personal or non-commercial framing as automatic clearance is the most consequential mistake, and the one most worth getting legal advice on before acting.
Installing Skills Without Reading Them
It is tempting to trust the pressure-testing and install everything that passed. But tests check behaviour against prompts the pipeline wrote; they do not confirm the steps faithfully reflect the source. Read each SKILL.md, compare the steps and boundaries with the original, and remove anything that overreaches.
Using It Where a Summary Would Do
A full run makes many model calls. For a short article, a news piece, or narrative content with little transferable method, the cost buys little, and most candidates will be filtered out anyway. Match the tool to dense, method-rich sources that will be used repeatedly.
Letting Skills Drift From Their Source
When the underlying playbook changes, hand-editing the skill files feels quicker than regenerating. Over time the skills and the source diverge, and nobody can tell which is authoritative. Keep both under version control and regenerate on change, then diff the results.
Installing Too Many Overlapping Packs
Several packs covering similar ground give an agent near-duplicate skills to choose between. The cross-skill tests run within a pack, not across every pack you install, so check for overlap yourself before adding another library, and remove packs nobody uses.
Cangjie Skill vs. Single-Prompt Summarization
The easiest way to see what the pipeline adds is to compare it with what most people do today: paste a chapter or transcript into a chat model and ask for the key points.
| Dimension | Single-prompt summary | Cangjie Skill pipeline |
|---|---|---|
| Output | Prose summary for a human | Installable SKILL.md files plus index, glossary, and digest |
| Selection rule | Whatever the model judges important | Must pass triple verification (multi-point evidence, transfer to new questions, non-obviousness) |
| Coverage vs. precision | Broad, includes restated common sense | Narrow, keeps roughly a quarter to a half of candidates |
| Structure | Free-form | Six fixed dimensions including steps and known boundaries |
| Testing | None | Trick questions, blind testing, cross-skill confusion tests |
| Long sources | Context-window limited | Chunking by chapter or episode, resumable state file |
| Where it's used | Read once | Invoked by an agent during real tasks |
The trade-off is cost and time. A full run uses far more model calls than a one-shot summary, and for a short article that's rarely worth it. For a dense book or a set of internal playbooks that a team will rely on repeatedly, the extra verification is the point.
Cangjie Skill Best Practices
If you want to try the pipeline, a low-risk path looks like this. The same habits apply if you build your own distillation pipeline modeled on it, and they matter more as the number of installed skill packs grows.
- Start with content you own. An internal playbook, your own conference talk transcript, or a set of team process documents removes the copyright question entirely.
- Follow the repository's README for setup. Installation details and supported environments change between versions, so treat the project README as the source of truth rather than a third-party guide.
- Review the rejected candidates, not just the survivors. The filtered-out material tells you whether the verification thresholds suit your content.
- Read every generated
SKILL.mdbefore installing it. Check the executable steps and stated boundaries against what the source actually says. - Test the skills in real tasks. Use the generated
test-prompts.jsonas a starting point, then add prompts from your own work. Our guide to what agent skills are explains how skills get selected and invoked at run time. - Keep the source and skills in version control. When the source document changes, you'll want to regenerate and diff rather than edit skills by hand.
- Run a small source first. A single chapter or one long talk shows how the pipeline behaves, how many model calls it makes, and what the output looks like before you commit to a full book and its usage costs.
- Assign an owner to each skill pack. Someone should be responsible for keeping the pack in sync with its source and for retiring skills that no longer apply. Unowned skills drift, and an agent will keep invoking outdated methods with full confidence. A short review each quarter is usually enough to catch skills that have quietly gone stale.
The Real Question This Raises: What You're Actually Allowed to Distill
This is the part worth being direct about rather than treating as a footnote. The extraction pipeline's own output structure includes a direct quoted excerpt from the source as one of its six required dimensions, and the tool has already been used, by its own documentation, against dozens of specific, actively-copyrighted, commercially-sold books — not just older public-domain texts. Turning a copyrighted book's frameworks, cases, and direct quotes into a structured, redistributable skill package and publishing that package publicly is a materially different act than writing a personal reading note for your own use, and it sits in territory that copyright law — particularly around derivative works and the scope of fair use for commentary versus wholesale extraction — hasn't cleanly settled for this specific kind of AI-generated output. Before running this against a book you don't hold rights to, and especially before publishing the resulting skill pack anywhere public, it's worth getting real legal guidance specific to your jurisdiction and the specific source material, rather than assuming a personal, non-commercial framing automatically clears it. The broader debate about who owns AI training data and how AI content licensing deals are being structured is useful background for that conversation.
Practical Implications
- The extraction methodology is worth studying on its own terms, independent of what source material you'd actually run it against — the triple-verification filter and the six-dimension structuring format are a genuinely more disciplined approach to knowledge distillation than a single-pass summarization prompt.
- Applying it to your own team's internal documents, meeting transcripts, or proprietary playbooks is the clearly safe use case — content you actually hold rights to, being turned into agent-callable structure rather than left in a wiki nobody rereads.
- Applying it to commercially copyrighted books you don't hold rights to, and then distributing the result, is the use case that needs real legal scrutiny before you do it, not after.
- The pressure-testing stage is a genuinely transferable idea for anyone building Agent Skills generally — designing trick questions and cross-skill confusion tests before a skill ships is a discipline most skill-building processes skip entirely.
Practical Takeaway
Cangjie Skill's extraction pipeline — parallel candidate generation, triple verification, structured six-dimension output, and pressure-testing before delivery — is a legitimately well-thought-out approach to a real problem: most consumed content never gets structured into something reusable. The methodology is worth learning from directly, whether or not you use the tool itself, and it's worth applying first to content you actually own or have clear rights to before extending it to material you don't.
Teams building internal knowledge-distillation pipelines or evaluating Agent Skills for institutional knowledge capture can get hands-on architecture help from Woyce Technologies.
FAQ
What is Cangjie Skill?
Cangjie Skill is an open-source pipeline that extracts methodologies from long-form content — books, podcasts, interviews, long videos — and structures them into installable Claude Code and Cursor Agent Skills, using a seven-stage process it calls RIA-TV++. Rather than producing a summary for a person to read, it outputs skill files an AI coding agent can load and apply during real tasks, along with an index, a glossary, test prompts, and a reader-facing digest.
How does Cangjie Skill decide what's worth extracting?
Candidate methodologies must pass a triple-verification filter: independent supporting evidence in at least two places in the source, the ability to answer a question the source doesn't explicitly address, and genuine non-obviousness rather than restated common knowledge. Only 25-50% of candidates typically survive. That filter is what separates a usable skill from a restated summary, and it means a long book may yield fewer skills than you might expect.
Is it legal to run Cangjie Skill against a copyrighted book?
This is a genuine open question rather than a settled one. The tool's output structure includes direct quoted excerpts from the source, and turning a copyrighted book into a redistributable skill package — especially one published publicly — raises real derivative-work and fair-use questions that depend on the specific source, jurisdiction, and how the result is used or distributed. Get actual legal advice before doing this with material you don't hold rights to.
What content types does Cangjie Skill support?
Books, long-video transcripts, podcast transcripts, interviews, lectures, courses, long-form articles, and general document collections — anything with extractable, verifiable, transferable methodology in text or transcript form. Audio and video sources need a transcript first, because the pipeline works on text. Content that teaches a repeatable method suits it best, since narrative or news-style material rarely survives the verification filter.
Is Cangjie Skill free to use?
It's open source under the GNU AGPL v3.0 license, which has real implications if you modify and redistribute it or offer it as a network service — read the license terms directly before building a service on top of it. The software itself costs nothing, but each run makes many language-model calls across its extraction, verification, and testing stages, so expect model usage costs that scale with the length of the source.
What's the safest way to use Cangjie Skill?
Running it against content you actually own or have clear rights to — internal documentation, your own team's playbooks, transcripts of your own meetings or talks — sidesteps the copyright question entirely and is the most straightforward, defensible use case. Keep the generated skill packs private to your organization, review each skill before installing it, and regenerate when the source changes rather than letting skills drift from the documents they came from.
What files does a Cangjie Skill run actually produce?
A structured mini-repository: BOOK_OVERVIEW.md for whole-source understanding, an INDEX.md skill map, individual SKILL.md files per surviving methodology, a GLOSSARY.md, test-prompts.json recording validation scenarios, and, as of version 2.0.0, a reader-facing DIGEST.md summary generated as a separate deliverable alongside the machine-facing skill files. Together these give an AI agent loadable skills and give people a readable record of what was extracted and how it was tested.
Can a Cangjie Skill run resume if it's interrupted partway through?
Yes, as of version 2.0.0. A PIPELINE_STATE.md file tracks the active stage, what's been produced so far, and each candidate skill's status, so a long extraction job on a full book or lengthy video transcript can pick back up instead of restarting. That matters because a full run makes many model calls, so restarting from scratch wastes both time and usage costs.
How does Cangjie Skill handle content too long to process in one pass?
Version 2.0.0 added chunking along natural boundaries — chapters, volumes, or episode breaks in a long video series — with a serial-processing fallback for cases where the parallel five-extractor stage isn't practical across every chunk at once. Splitting on natural boundaries keeps each methodology in context, and combined with the resumable pipeline state, it lets a very long source be worked through in stages rather than in one fragile pass.
Is Cangjie Skill only useful for books written in Chinese?
No — the pipeline and its README are available in English, Japanese, and Chinese, and the published skill packs include English-language sources, though the project's own examples so far lean heavily toward Chinese-language business, investment, and classical texts given the maintainer's background as a Chinese-language AI educator. Any source with a clear, transferable methodology can work, regardless of its language.
Conclusion
Most of what people read, watch, and record never becomes something they can use in the moment of work. Cangjie Skill attacks that gap by converting long-form sources into tested, installable agent skills rather than another summary, and its real contribution is the discipline around that conversion: multi-point evidence, a test for transfer to new questions, a non-obviousness filter, and pressure-testing before anything ships.
Two caveats matter. First, the pipeline is expensive relative to a quick summary and only pays off for dense sources that will be used repeatedly. Second, the copyright question is unresolved. The output format includes direct quotes, and the project has already been applied to commercially sold books. Using it on material you don't hold rights to, and especially publishing the results, needs legal advice first.
The most useful way to adopt the idea is internally: run it, or a pipeline modeled on it, against your own playbooks, process documents, and expert recordings, then review and test the resulting skills like any other code. If you're planning that kind of knowledge-capture system for your team, our LLM integration services can help you design it.
