Ask a coding agent to review a pull request and, without help, it often ends up reading far more of the codebase than the change actually touches — every file gets pulled into context just so the agent can reason about what might be affected. code-review-graph fixes that by building a persistent, structural map of a codebase ahead of time, so an AI coding tool can query precisely what's relevant through MCP instead of re-reading files it doesn't need.
The problem matters more than it first appears. Every token an agent spends re-reading unrelated files costs money, adds latency, and pushes useful information out of a finite context window. Worse, an agent that reads everything still doesn't know which downstream callers a change might break, so reviews end up both expensive and shallow. As teams hand more review and refactoring work to background coding agents, that waste compounds across every pull request.
This explainer walks through how code-review-graph works: the Tree-sitter parsing and incremental update model, what the graph actually stores, and why blast-radius analysis is the real value rather than raw token savings. It also looks closely at the project's benchmark numbers and its unusually candid list of known weaknesses, how language support can be extended without forking, and how the same analysis runs as a risk-scored review inside CI. It closes with practical implications for teams deciding whether to adopt it.

Built on Tree-sitter, Updated Incrementally
The core mechanism is a structural graph parsed with Tree-sitter across a genuinely broad set of languages and frameworks — Python, JavaScript, Rust, Go, PHP/Laravel, Java, C#, Terraform, Ansible, and more — plus Jupyter notebook support. The graph updates incrementally rather than requiring a full re-parse on every change, which is what makes it practical to keep current on an actively-developed repository rather than something you rebuild from scratch each session. Once built, a single install command auto-detects which AI coding tools are present — Claude Code, Cursor, Codex, Windsurf, Zed, GitHub Copilot, and over a dozen others — and writes the correct MCP configuration for each one automatically.
Blast-Radius Analysis Is the Actual Value Proposition
Beyond just answering "what does this function do," the graph supports blast-radius analysis — tracing which files and functions are affected by a given change through call and import edges, so a review tool can flag downstream impact instead of only looking at the diff in isolation. That's the harder, more valuable problem token-efficient retrieval is actually solving for: knowing what to check, not just answering questions faster about what's already open.
The Benchmark Numbers Are Reported With Real Honesty
This is the part worth taking seriously as a model for how technical claims should be published. The headline number — a ~65x median per-question token reduction across six real repositories — comes with a documented, reproducible methodology (pinned commit SHAs, a fixed random seed, deterministic CPU embeddings) rather than an unverifiable marketing figure, and the project is explicit that the 376x maximum is a single best-case result on the largest test corpus, not the typical outcome. More notably, the "Limitations and known weaknesses" section is genuinely substantive rather than boilerplate: it states directly that the impact-analysis "recall 1.0" figure is circular by construction (the ground truth is derived from the same graph the predictor traverses), that a more honest co-change evaluation mode currently returns zero predicted files on every graded commit and isn't yet a usable measurement, and that search ranking quality (MRR 0.35) and flow detection accuracy (33% recall) both need work in specific named areas. A project willing to publish "this evaluation mode doesn't work yet" in its own README is a stronger signal of engineering integrity than a clean benchmark table with no caveats would be.
What's Actually in the Graph
The graph isn't a vague "index" — it's a concrete structure with typed nodes and edges. Nodes represent functions, classes, and imports; edges represent calls, inheritance, and test coverage. That's parsed straight out of the Tree-sitter AST and stored in SQLite, which is what makes both the blast-radius traversal and the incremental updates tractable: querying "what calls this function" or "what tests cover this class" is a graph query against structured data, not a fresh text search over source files every time.
Incremental updates lean on the same structure. When a hook or watch mode fires, the tool diffs changed files against Git, walks the graph's own import and call edges to find dependents, and re-parses only the files whose SHA-256 hash actually changed — untouched files are skipped entirely rather than re-hashed and re-walked for no reason. On django, a roughly 3,000-file repository, a two-file edit re-indexes in about 2.5 seconds on the hook path, and close to 1.4 seconds of that is process start-up rather than parsing work; a no-op update (nothing actually changed) costs only that start-up overhead. That's a meaningfully different cost profile than a full re-parse on every save, and it's the detail that makes "keep the graph current on an actively-developed repo" a realistic claim rather than an aspiration.
Build performance scales with repo size in a predictable way: express (141 files) builds its graph — 1,910 nodes, 17,553 edges — in around 106ms with sub-millisecond search latency, while fastapi (1,122 files) takes about 128ms to produce 6,285 nodes and 27,117 edges. Even at the larger end, initial build stays fast enough to run as a normal part of getting started rather than something you kick off and walk away from.
Extending Language Support Without Forking
Tree-sitter coverage is already broad — Python, JavaScript/TypeScript, Go, Rust, Java, C/C++, C#, VB.NET, Ruby, Kotlin, Swift, PHP, Scala, Solidity, Dart, R, Perl, Lua, Objective-C, shell scripts, Elixir, Zig, PowerShell, Julia, Terraform/OpenTofu structure, Ansible playbooks, Vue/Svelte SFCs, and more — but the project doesn't require a fork when a codebase uses something the built-in parsers don't cover. Dropping a languages.toml file into .code-review-graph/ that maps file extensions to a grammar bundled in tree_sitter_language_pack, plus the specific Tree-sitter node types for functions, classes, imports, and calls, is enough for the generic tree-sitter walker to start extracting structure from that language — no code changes required, and the built-in language definitions can't be overridden by a misconfigured custom entry. PHP gets extra treatment beyond generic Tree-sitter parsing: repository-bounded Composer PSR-4 resolution, Blade template references, and Laravel-specific Route and Eloquent semantic edges when the source shows explicit framework imports and model inheritance.
Risk-Scored Reviews Inside CI, Not Just Locally
Beyond the MCP integration for interactive coding sessions, the same analysis ships as a composite GitHub Action that runs entirely on the CI runner — the knowledge graph is built and queried locally in that job, with no source code sent to an external service. On each pull request it posts a single sticky comment listing risk-scored functions, affected execution flows, and test gaps, and that comment updates in place on every subsequent push rather than accumulating a new comment each time. An optional fail-on-risk input turns the review from an informational comment into an actual merge gate. The project dogfoods this on itself, running the same action against its own pull requests rather than only recommending it for other repos.
code-review-graph vs Full-File Context Reading
The alternative the tool replaces is the default behaviour of most coding agents: open the changed files, then keep opening related files until the agent feels it has enough context. The table compares the two approaches on the dimensions that matter for review.
| Dimension | Full-file context reading | code-review-graph |
|---|---|---|
| How context is gathered | Agent reads whole files, often many unrelated ones | Agent queries a prebuilt graph for specific callers, imports, and tests |
| Token cost per question | Grows with the size of the files read | Much lower on multi-file questions; can be higher on tiny single-file changes |
| Downstream impact | Only found if the agent happens to open the right files | Traced through call and import edges as blast-radius analysis |
| Setup effort | None | Initial graph build, plus hooks or watch mode to keep it current |
| Freshness | Always reads the current file contents | Depends on incremental updates running after each change |
| Where it runs | Inside the agent session | Locally via MCP, or on the CI runner as a GitHub Action |
The core difference is how the agent decides what to look at. With full-file reading, relevance is a guess: the agent opens files that look related by name or by search hits, and anything it doesn't open is invisible to it. That approach needs no setup and always sees the latest code, which is why it remains the default, but it becomes expensive and unreliable as repositories grow and changes touch shared functions.
The graph turns relevance into a structural question. Instead of reading a module to discover who calls a function, the agent asks the graph and gets a compact list of callers, importing modules, and covering tests. That is where the token savings come from, and more importantly it is how the review catches a broken caller three directories away that a file-reading agent would never have opened.
The trade-off is maintenance and edge cases. The graph has to be built and kept current, and for small single-file edits its structural metadata can cost more than simply reading the file. Many teams will end up using both: the graph for review, refactoring, and impact questions, and plain file reads for quick local edits.
Benefits of code-review-graph
Reviews that look beyond the diff
The most valuable thing the graph adds is awareness of downstream impact. When a function's signature or behaviour changes, blast-radius analysis lists the callers, importing modules, and tests that might be affected. A reviewer, human or AI, gets a checklist of places to verify rather than relying on memory or luck. That shifts reviews from checking whether the changed lines look right to checking whether the change is safe for everything that depends on it.
Lower token spend on multi-file questions
On questions that involve dependency reasoning across files, the project's benchmarks report large median reductions in tokens per question. Fewer tokens mean lower model bills and faster responses, and they leave more of the context window free for the reasoning that matters. The saving compounds for teams running many automated reviews a day, though it shrinks or reverses on tiny single-file changes.
Context that stays current cheaply
Incremental updates re-parse only files whose hash has changed and follow graph edges to find dependents. On a large repository a small edit re-indexes in seconds, much of it start-up time. That makes it realistic to keep the graph fresh on an active codebase, so agents aren't reasoning about structure that changed last week.
One index for many tools
Because it speaks MCP and auto-configures over a dozen coding tools, one graph can serve a team using different editors and agents. Teams avoid maintaining separate context setups per tool, and everyone's agent sees the same structural picture of the codebase. Switching or adding a tool later doesn't mean rebuilding context infrastructure from scratch.
Local-first processing
Both the MCP workflow and the GitHub Action build and query the graph locally or on the CI runner, with no source code sent to an external service by the tool itself. For teams with strict policies about where code can go, that removes one of the usual objections to adding AI review tooling. The retrieved context still reaches whichever model provider the coding tool uses, so the overall data flow should be reviewed as a whole.
code-review-graph Use Cases
AI-assisted pull request review
The headline use is giving a coding agent precise context when it reviews a pull request. Instead of reading every file that might be relevant, the agent asks the graph what the change touches and which tests cover it. The review becomes cheaper and more focused on likely breakage, and the reviewer gets a clearer explanation of why particular files deserve attention. On large pull requests, it also helps the agent prioritise which parts of the diff carry the most risk.
Risk gating in CI
Teams can run the GitHub Action on every pull request and get a single updating comment with risk-scored functions, affected flows, and test gaps. With the optional fail-on-risk input, high-risk changes can be blocked until someone looks. That suits teams that want an automated second opinion without sending source code to an external review service. Because the comment updates in place, the pull request thread stays readable across many pushes.
Planning refactors on shared code
Before renaming or changing a widely used function, engineers can query its blast radius to see every caller and test involved. That makes it easier to size the work, split it into safe steps, or decide the refactor isn't worth the disruption. Background coding agents can use the same information to keep their changes within scope.
Onboarding into an unfamiliar codebase
New team members, and agents dropped into a repository for the first time, can ask structural questions such as what calls this, what this class inherits from, and where it is tested, without reading large swathes of code. Answers come back as concrete file and function references they can open directly. It doesn't replace documentation, but it shortens the time it takes to find the parts of a codebase that matter for a task.
Supporting custom or less common languages
Teams with a language the built-in parsers don't cover can add it through a configuration file mapping extensions to an existing grammar. That brings structural context to internal tools or niche stacks without forking the project, which is often the blocker for adopting code-intelligence tooling in mixed-language repositories.
Common code-review-graph Mistakes
Expecting the headline multiplier on every change
The token-reduction claim is real but scenario-dependent. For small, single-file changes, the project's own data shows graph context can actually exceed a naive file read — the overhead is the structural metadata that only pays off once multiple files or dependency reasoning are involved. Teams that judge the tool on a week of tiny commits will conclude it doesn't work, when they simply measured the case it isn't built for.
Treating blast-radius output as a definitive list
Blast-radius analysis trades precision for recall on purpose. It's designed to over-flag potentially-affected files rather than risk missing a broken dependency. Reviewers who treat every flagged item as a confirmed problem waste time, while those who treat the list as complete miss dynamic dispatch and other links the graph can't see. Use it as a starting point for review, not a verdict.
Ignoring how naming conventions affect search
Search quality varies significantly by construct naming convention — the project specifically notes that Express queries return zero hits due to how that framework's modules are typically named. Any code-graph tool's quality depends partly on how idiomatically a codebase is structured. Assuming search works equally well across every repository leads to silent gaps in the context an agent receives.
Trusting published numbers without rerunning them
Documented, pinned-SHA benchmarks you can rerun yourself are a real asset, but they describe six open-source repositories, not yours. Skipping the reproduction recipe means adopting the tool on the strength of numbers measured on different languages, sizes, and structures. A short run on your own codebase tells you which figures hold.
Forgetting the rest of the data flow
The graph itself stays local, which is easy to read as "no code leaves the building". The context an AI coding tool retrieves from the graph still goes to that tool's model provider. Teams with strict data policies need to evaluate the whole pipeline, not just the indexing step.
code-review-graph Best Practices
- Pilot on one active repository first. Pick a codebase with regular multi-file changes and a team willing to compare reviews with and without the graph. Measure tokens per review, review time, and missed-impact incidents over a few weeks before rolling it out widely.
- Rerun the benchmark recipe on your own code. Use the project's pinned methodology against your repository to see which token and search figures hold for your languages and structure. Discount the numbers the project itself flags as circular or unreliable.
- Keep the graph fresh automatically. Install the hooks or watch mode so incremental updates run after every change. A stale graph produces confident answers about structure that no longer exists, which is worse than no graph at all.
- Start the CI action as informational, then gate. Run the GitHub Action with comments only until the team trusts its risk scores. Turn on fail-on-risk once false-alarm rates are understood and a clear override process exists.
- Use blast radius as a review checklist. Ask reviewers and agents to work through flagged callers and tests explicitly, noting which were checked and which were irrelevant. Feed recurring false positives back into how you interpret the output.
- Add custom languages through configuration. If part of the codebase uses a language not covered by default, map it in the configuration file rather than forking. Validate extraction on a few files before relying on it in reviews.
- Track whether reviews actually improve. Count regressions caught in review versus those found after merge, before and after adoption. Token savings are pleasant, but fewer escaped breakages is the outcome that justifies keeping the tool.
- Review the end-to-end data flow. Confirm what context reaches your model provider and whether that matches your policies. Local graph building is one part of the picture, not the whole answer.
Practical Takeaway
code-review-graph is a well-engineered answer to a real, specific cost problem in AI-assisted code review: agents re-reading far more of a codebase than a change actually requires. What sets it apart isn't just the token-efficiency numbers — plenty of tools claim efficiency gains — it's the unusually transparent disclosure of exactly where the measurements are weak, circular, or not yet trustworthy. For teams evaluating AI code review tooling or building GraphRAG-style context systems of their own, that honesty is worth as much as the benchmark table itself when deciding whether to trust the claims.
Teams building code-intelligence infrastructure or evaluating MCP-based context tooling for AI coding agents can get hands-on architecture help from Woyce Technologies.
FAQ
What is code-review-graph?
code-review-graph is an open-source tool that builds a persistent, Tree-sitter-based structural graph of a codebase and serves precise, relevant context to AI coding tools via MCP, reducing how much of the codebase an agent needs to read for reviews and impact analysis. Instead of pasting whole files into a prompt, the agent asks targeted questions, such as which functions call this one or which tests cover this class, and gets back a compact answer drawn from the graph.
How much does code-review-graph actually reduce token usage?
Its published benchmarks show a median ~65x per-question token reduction across six real open-source repositories, with a documented, reproducible methodology. The maximum measured reduction (376x) is a single best-case result, not the typical outcome — for small single-file changes, the reduction can be smaller or even negative. The savings show up once multiple files or dependency reasoning are involved, so rerun the pinned-SHA reproduction recipe on your own codebase before trusting the numbers.
Which AI coding tools does code-review-graph support?
It auto-detects and configures MCP integration for a wide range of tools including Claude Code, Cursor, Codex, Windsurf, Zed, GitHub Copilot, and over a dozen others, writing the correct configuration for each automatically. Because the integration runs over the Model Context Protocol, any tool that speaks MCP can in principle query the same graph, which means a team using several different editors and agents can share one structural index rather than maintaining separate context setups for each tool.
Is code-review-graph free to use?
Yes, it's MIT-licensed and open source, installable via pip or pipx and distributed on PyPI. The core workflow builds and stores the graph locally, so the main costs are the compute to build it and the tokens your AI coding tool spends on the context it retrieves. For commercial teams, the MIT license permits use inside private repositories, though it's still worth having someone review any dependency before rolling it out across an organisation.
What is blast-radius analysis?
It's the graph's ability to trace which files and functions are likely affected by a given code change through call and import relationships, so a review tool can flag downstream impact rather than only examining the diff in isolation. In practice it answers the reviewer's real question: if this function changes, what else could break? The tool deliberately over-flags rather than under-flags, so treat its list as a checklist to work through, not a verdict.
What are code-review-graph's known limitations?
The project's own documentation discloses several: its headline impact-analysis recall figure is circular by construction, a more honest co-change evaluation mode isn't yet a usable measurement, search ranking quality and flow detection accuracy both need improvement in specific languages, and token savings can be negative for small single-file changes. That candour is useful when evaluating it, because you know exactly which numbers to discount and which to verify yourself.
How does code-review-graph handle a language it doesn't support out of the box?
You can add one without forking the project: a languages.toml file in .code-review-graph/ maps file extensions to a grammar already bundled in tree_sitter_language_pack, along with the Tree-sitter node types for functions, classes, imports, and calls. The generic parser handles extraction from there. No code changes are required, and a misconfigured custom entry can't override the built-in language definitions, so experimenting with a new grammar won't break support for the languages the project already covers.
Does code-review-graph work with GitHub Actions?
Yes — it ships as a composite GitHub Action that builds and queries the graph locally on the CI runner and posts a single sticky risk-review comment on each pull request, updated in place on every push. The comment lists risk-scored functions, affected execution flows, and test gaps. An optional fail-on-risk input can turn it into a merge gate.
How fast are incremental updates on a large repository?
On django, a roughly 3,000-file repo, a two-file edit re-indexes in about 2.5 seconds through the hook path, of which roughly 1.4 seconds is process start-up rather than parsing. A no-op update costs only that start-up overhead, since only files with a changed SHA-256 hash get re-parsed. That cost profile is what makes keeping the graph current on an actively developed repository realistic.
Is code-review-graph's analysis sent to an external service?
No — both the CLI/MCP workflow and the GitHub Action build and query the graph locally. The GitHub Action explicitly runs the knowledge graph construction and querying on the CI runner itself, with no source code sent externally. The graph itself is stored locally in SQLite. Keep in mind that the context your AI coding tool retrieves from the graph still goes to whichever model provider that tool uses.
Conclusion
AI coding tools waste a surprising amount of effort re-reading code a change never touches, and they still miss the downstream effects that matter most in a review. code-review-graph tackles both by building a persistent structural map of a repository and letting agents query it over MCP for exactly the context they need.
The most useful idea here isn't the headline token multiplier. It's blast-radius analysis: knowing which callers, flows, and tests a change might affect before a reviewer signs off. The incremental update path, extensible language support, and local-only CI action make that practical on real, actively developed codebases without sending source code anywhere.
The caveats come straight from the project's own documentation. Token savings depend on the scenario and can be negative on small single-file changes, the impact-analysis recall figure is circular, and search and flow detection still need work in some areas. Rerun the published benchmarks on your own repository before relying on the numbers.
If your team is evaluating MCP-based context tooling or building code-intelligence infrastructure for AI-assisted review, a short pilot on one active repository will tell you more than any benchmark. Our LLM integration team can help you design and measure that pilot.
