Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

Pi Explained: A Coding Agent Harness Built as Composable Packages

Pi is an open-source AI agent toolkit split into a unified multi-provider LLM API, an agent runtime, a terminal UI library, and an interactive coding agent CLI — usable together or independently.

Pi Explained: A Coding Agent Harness Built as Composable Packages — Woyce Technologies

Loading repository details…

——

Most coding agent CLIs ship as one monolithic tool — the model API, the agent loop, and the terminal interface are bundled together, take it or leave it. Pi, from earendil-works, is built the opposite way: five separate, independently usable packages — a unified multi-provider LLM API, an agent runtime, a terminal UI library, a telemetry layer, and the coding agent CLI that ties them together. It's already established enough in the ecosystem that it shows up as a comparison point in other agent tooling's own benchmarks and supported-agent lists — jcode measures its RAM usage directly against Pi, and Orca lists it as one of its natively supported agents.

The problem Pi addresses is familiar to any team that has tried to build on a coding agent rather than just use one. Most harnesses tie you to their interface, their provider choices, and their idea of how the agent loop should run. If you want the same tool-calling runtime inside a Slack bot, a CI job, or an internal product, you end up forking a CLI or rebuilding the plumbing yourself. That matters more as background coding agents move from novelty to everyday infrastructure.

This explainer covers the Pi agent harness's package architecture, how its extension system works, why it deliberately ships without a sandbox, its multi-provider design, its unusually strict supply-chain practices, reproducible releases, the project's request for real session data, and how to decide whether it fits your team.

Five Packages, Not One Monolith

The architecture is the most distinctive thing about Pi, and it's worth understanding on its own terms:

PackageWhat it does
@earendil-works/pi-aiUnified API across OpenAI, Anthropic, Google, and other providers
@earendil-works/pi-agent-coreThe agent runtime — tool calling and state management
@earendil-works/pi-coding-agentThe interactive coding agent CLI built on top of the other packages
@earendil-works/pi-tuiA terminal UI library with differential rendering
@earendil-works/pi-telemetryVendor-neutral telemetry contracts, a reference adapter, and conformance tests

Splitting the unified LLM API and the agent runtime out as their own packages means a team can build a completely different interface — a chat automation tool, a custom agent product, a Slack integration — on the same underlying provider abstraction and tool-calling runtime that powers Pi's own CLI, without inheriting the CLI itself. That's exactly what earendil-works has done with a companion project, pi-chat, for Slack and chat-based automation built on the same core packages.

Pi package architecture: the coding agent CLI, pi-chat and a custom product all sit on pi-agent-core, which sits on the unified pi-ai provider API, alongside pi-telemetry.

Self-Extensible by Design

The project describes itself as a "self extensible coding agent" — the CLI isn't just configurable, it's built to be extended, including through a tool_call event system that lets an extension terminate an all-terminating batch of tool calls without triggering another model call, a detail that matters for anyone building extensions that need tight control over when the agent loop actually continues versus stops. Recent releases show this extensibility being actively used and refined, not just documented as a theoretical capability.

No Built-In Sandbox — By Explicit Design Choice

Pi's documentation is direct about something a lot of agent CLIs leave implicit: it does not include a built-in permission system restricting filesystem, process, network, or credential access, and runs with the full permissions of whatever launched it by default. Rather than bolting on a partial, easy-to-misconfigure sandbox, the project points to three explicit containerization patterns instead — a "Gondolin" extension that routes tool calls and shell commands into a local Linux micro-VM while keeping pi and provider auth on the host, a plain Docker container for simpler isolation, and a policy-controlled sandbox option called OpenShell. That's a meaningfully more honest stance than a tool that implies safety it doesn't actually enforce — the security boundary is explicitly your responsibility to add, with concrete, documented options for doing it rather than a false sense of built-in protection.

Three isolation options for Pi, which has no built-in sandbox: a Gondolin micro-VM for tool calls, a plain Docker container, or the policy-controlled OpenShell sandbox.

Multi-Provider by Default, Not by Afterthought

Because the LLM API layer is unified across providers from the ground up rather than added on top of an Anthropic-specific or OpenAI-specific core, Pi supports a genuinely broad provider list, including subscription-model providers like Qwen's Token Plan alongside the usual API-key-based options. A recent addition — pi auth check — lets you verify provider or model credentials are actually valid before you're mid-task and hit an auth failure, a small but practical piece of operational polish that comes from treating multi-provider support as core functionality rather than a bolted-on feature.

Supply-Chain Hardening Most Agent CLIs Don't Bother With

Pi's dependency management is unusually strict for an open-source CLI, and the project documents it in enough detail that it reads as a genuine practice rather than a compliance checkbox. Direct external dependencies are pinned to exact versions rather than ranges; internal workspace packages stay version-ranged, since those are reviewed within the same repo. .npmrc sets save-exact=true and min-release-age=2, so a same-day npm release can't get pulled into a build during dependency resolution — a real defense against a specific class of supply-chain attack where a compromised package version ships and gets adopted within hours. package-lock.json is treated as ground truth, and a pre-commit hook blocks accidental lockfile changes unless you explicitly set PI_ALLOW_LOCKFILE_CHANGE=1. The published CLI package ships its own shrinkwrap file, generated from the root lockfile, to pin transitive dependencies for npm users specifically — and new dependencies that introduce lifecycle scripts fail checks until a maintainer explicitly reviews and allowlists them. CI installs with --ignore-scripts, and a scheduled workflow runs npm audit plus npm audit signatures against production dependencies. None of this is exotic tooling — it's disciplined use of features npm already provides, applied consistently rather than left as defaults.

Five supply-chain gates in Pi: exact version pins, a minimum release age, a protected lockfile, reviewed lifecycle scripts and CI audits, plus checksum-verified reproducible releases.

Releases Ship as Reproducible, Checksum-Verified Source

Standalone binaries aren't just built and uploaded on faith. GitHub releases include a versioned source archive covered by a SHA256SUMS file, and the exact script used to build the official binaries — scripts/build-binaries.sh — is included and runnable against that archive, with an --offline-model-data flag to rebuild using the release's pinned provider-model snapshot instead of refreshing it live. A package maintainer or a security-conscious team can independently rebuild what Pi shipped and confirm it matches, rather than trusting the binary a maintainer uploaded is what the source actually produces.

Benefits of the Pi Agent Harness

Build your own interface on a tested runtime

Because the LLM API and the agent runtime are separate packages, a team can build a Slack bot, a CI job, or an internal product on the same provider abstraction and tool-calling loop that powers Pi's CLI. That avoids forking a terminal tool or rewriting agent plumbing from scratch. The pi-chat companion project is a working example of this pattern, which gives some confidence that the split is practical rather than theoretical.

Freedom to switch model providers

The unified API was designed across providers from the start, so moving between OpenAI, Anthropic, Google, and subscription-based options does not require rewriting integrations. Teams can follow price, quality, or availability changes without being locked into one vendor's tooling, and pi auth check helps catch credential problems before a long task fails partway through.

Extension points that control the agent loop

The extension system, including the tool_call event that can end a batch of tool calls without triggering another model call, gives developers precise control over when the agent continues and when it stops. That matters for teams building guardrails, custom workflows, or cost controls on top of the agent, where an extra model call can mean wasted spend or an unwanted action.

A stronger supply-chain posture

Exact version pins, a minimum release age for new npm packages, a protected lockfile, reviewed lifecycle scripts, and scheduled audits reduce exposure to compromised dependencies. Reproducible, checksum-verified releases let a security team rebuild a binary and confirm it matches the source. For organisations that treat developer tools as part of their attack surface, these practices remove a lot of review work.

Honest security boundaries

Pi states plainly that it has no built-in sandbox and documents concrete isolation options instead. That honesty helps teams make an informed decision about where the security boundary sits, instead of relying on a partial permission system that implies protection it does not provide. The responsibility is clear, and so are the tools to meet it, which makes security reviews shorter because nobody has to reverse-engineer what a half-built permission layer really enforces.

Pi Agent Harness Use Cases

Interactive coding in the terminal

The most direct use is the coding agent CLI itself: a developer working in a repository asks the agent to implement, refactor, or debug, and the agent calls tools and edits files. Teams that switch models often can keep the same workflow while changing providers. Running it inside a container or micro-VM keeps the fully permissioned process away from credentials and systems it should not touch.

Chat and Slack automation

Using the core packages without the CLI, teams can build agents that respond in chat channels, run routine tasks, or answer questions about a codebase. This is the pattern pi-chat demonstrates. The benefit is reuse: the same runtime and provider layer power both the terminal tool developers use and the automation the wider team interacts with. Fixes and provider changes made in the core packages benefit both at once, instead of being duplicated across two separate agent implementations.

Agents in CI pipelines

A CI job can invoke an agent built on the runtime to triage failures, propose fixes, or update documentation. Here the lack of a built-in sandbox matters most, so teams typically run the job in an isolated container with only the access the task requires. The typed telemetry contracts help track what the agent did across many automated runs, so reviewers can audit a proposed fix before it is merged.

Custom internal agent products

Companies building their own agent-based product can adopt the provider API and agent core as a foundation, then design their own interface and permissions model. This saves building provider integrations and a tool-calling loop from scratch while keeping full control of the user experience and security design. Because the packages are MIT-licensed, the product team can build on them commercially without negotiating a separate licence.

Reproducible tooling for security-conscious teams

Organisations that must verify what runs on developer machines can rebuild Pi releases from the published source archive and checksums, using the pinned model-data snapshot for an offline build. That makes it feasible to approve an agent CLI through a stricter internal review than most tools would pass.

The Project Wants Your Agent Sessions, Not Just Your Code

One detail that says something about how earendil-works thinks about improving coding agents: the README actively asks users to publish their real Pi (and other agent) sessions publicly, arguing that real-world session data — actual tasks, tool calls, failures, and fixes — improves agent tooling more than toy benchmarks do. The maintainer publishes their own sessions to a public Hugging Face dataset and points to a companion tool, pi-share-hf, for anyone who wants to do the same with just a Hugging Face account and its CLI. It's a small ask, but it's a concrete bet that the field's benchmarks are a poor substitute for messy, real usage data — worth noting if you're weighing whether to contribute your own session logs back.

When Pi Is the Right Fit, and When It Isn't

Pi's design choices make it a strong match for some teams and a poor one for others. Based on what the project documents, the split looks roughly like this.

SituationPi is a good fitLook elsewhere or add tooling
You want to build a custom agent productThe LLM API and agent-core packages can be adopted without the CLI
You switch between model providers oftenMulti-provider support is part of the core design
You need built-in permission prompts and sandboxingPi has no built-in permission system; plan on Gondolin, Docker, or OpenShell
Supply-chain risk is a board-level concernPinned dependencies, release-age delays, and reproducible builds
You want a polished, opinionated tool with no setup decisionsA monolithic CLI may be simpler to roll out
You plan to contribute code upstreamNew contributor PRs are auto-closed by default and triaged later

The common thread is control. Pi rewards teams that want to own decisions about isolation, providers, and interfaces, and it asks more of teams that would rather a vendor made those choices for them. If you are combining it with agent skills or external tool servers, apply the same scrutiny you would to any tool that runs with full permissions, including the risks covered in our piece on MCP tool poisoning.

Common Pi Agent Harness Mistakes

Running it unisolated on a workstation full of credentials

Pi runs with the full permissions of whatever launched it. Starting it directly on a laptop with cloud credentials, SSH keys, and production access in reach gives the agent, and anything that manipulates it, the same reach. The documented Gondolin, Docker, and OpenShell options exist precisely to prevent this, and choosing one should come before the first real task.

Adopting the whole CLI when you only need the runtime

Teams building a custom product sometimes wrap the CLI instead of using the LLM API and agent-core packages directly. That inherits a terminal interface and assumptions they do not need, and makes upgrades harder. If the goal is your own interface, start from the packages and treat the CLI as a reference for how they fit together.

Treating third-party extensions and tool servers as trusted

Because the agent runs with broad permissions, any extension, skill, or external tool server it loads runs with that reach too. Installing them without review undoes the care the project takes with its own dependencies. Apply the same scrutiny to these additions as to any code that executes on your machines.

Misreading the contribution model

New contributors sometimes see their first issue or pull request auto-closed and assume the project is unmaintained or hostile. The auto-close is a deliberate triage mechanism, with maintainers reviewing closed items daily. Reading the contribution guidelines first avoids frustration and wasted effort, and it helps a first contribution land in a form the maintainers can actually review.

Floating on the latest version

The project releases steadily, and behaviour around extensions and providers evolves. Building internal tooling on an unpinned version can lead to surprises when a release changes something your workflow relied on. Pin a version and upgrade deliberately, ideally after rebuilding from the checksum-verified source.

Pi Agent Harness Best Practices

  • The package split is the real reason to look at Pi even if you don't use its CLI. A team building a custom agent product can adopt the LLM API and agent-core packages directly, skipping the CLI and TUI layers entirely — a more modular starting point than most agent frameworks offer.
  • Containerize deliberately, not as an afterthought. Given there's no built-in permission system, decide on Gondolin, Docker, or OpenShell — or an equivalent of your own — before running Pi against anything you wouldn't want a fully-permissioned process touching.
  • Keep provider credentials on the host and tools in the sandbox. Patterns like the Gondolin extension route tool calls and shell commands into a micro-VM while provider authentication stays outside it. Aim for the same separation in any isolation setup you build.
  • Verify credentials before long tasks. Run pi auth check when switching providers or rotating keys, so an authentication failure does not interrupt an agent halfway through a multi-step change.
  • The contribution model is worth knowing before opening a PR. New issues and PRs from new contributors are auto-closed by default, with maintainers reviewing auto-closed items daily — a deliberate triage mechanism, not a sign the project is unmaintained.
  • Telemetry is vendor-neutral and typed, which matters if you're instrumenting agent behavior across a fleet — worth checking the pi-telemetry package's contracts before building custom monitoring on top of Pi.
  • Mirror the project's dependency discipline in your own extensions. Pin exact versions, review lifecycle scripts, and keep lockfiles protected for anything you build on Pi, so your additions do not become the weakest link in an otherwise hardened toolchain.
  • Review session data before sharing it. If you take up the project's invitation to publish agent sessions, check them for secrets, internal code, and customer data first.

Practical Takeaway

Pi's real differentiator isn't a single flashy feature — it's the decision to build a coding agent CLI as a byproduct of a genuinely modular agent toolkit, rather than the toolkit as a byproduct of the CLI. For teams that need multi-provider LLM access and a real agent runtime as building blocks for something custom, not just an interactive terminal tool, that architecture is worth evaluating directly — independent of whether Pi's own CLI ends up being the interface you actually use day to day.

Teams building custom agent products, evaluating coding agent CLIs, or designing containerized agent deployments can get hands-on architecture help from Woyce Technologies.

FAQ

What is Pi?

Pi is an open-source AI agent toolkit from earendil-works, split into five independently usable packages: a unified multi-provider LLM API, an agent runtime, a terminal UI library, a telemetry layer, and an interactive coding agent CLI built on top of them. Because the packages are separate, a team can adopt the provider abstraction and tool-calling runtime for its own product, such as a Slack bot or CI job, without inheriting the CLI.

Is Pi free to use?

Yes, it's MIT-licensed and open source, distributed as npm packages and standalone binaries built from versioned, checksum-verified release source. The software itself costs nothing, but you still pay your chosen model providers for API usage or subscriptions, and running it safely may involve container or micro-VM infrastructure you maintain yourself.

Does Pi include a sandbox or permission system?

No — by design, Pi runs with the full permissions of the process that launched it and doesn't restrict filesystem, network, or credential access on its own. The project documents three containerization patterns (a Linux micro-VM extension, plain Docker, or a policy-controlled sandbox) for teams that need stronger isolation. Decide on one before running Pi against anything you wouldn't want a fully permissioned process touching.

Which AI providers does Pi support?

A broad, unified set including OpenAI, Anthropic, Google, and subscription-model providers like Qwen's Token Plan, all accessed through one consistent API layer rather than provider-specific integrations bolted on separately. The pi auth check command lets you confirm credentials work before a long task starts. Because the API layer was unified from the ground up rather than added on top of one provider's core, switching between models is part of the core design, which suits teams that change providers often.

Can I use Pi's packages without its CLI?

Yes — the LLM API and agent runtime packages are usable independently. A companion project, pi-chat, uses the same core packages to build Slack and chat-based automation rather than the interactive terminal CLI. That makes Pi a reasonable foundation for internal tools or products that need a tested tool-calling runtime and provider abstraction without a terminal interface.

How does Pi compare to jcode or Orca?

Pi is referenced directly by both — jcode benchmarks its own RAM usage against Pi's, and Orca lists Pi as one of its natively supported coding agents — reflecting that Pi is an established, independently-usable agent harness other tools already build around or compare themselves to. The difference is mainly one of focus: Pi's distinguishing feature is its modular package split rather than any single headline capability.

How seriously does Pi take dependency and supply-chain security?

Seriously enough to document it as its own practice: pinned exact versions for direct dependencies, a minimum release-age delay before new npm packages can be pulled in, a pre-commit hook that blocks accidental lockfile changes, a dedicated shrinkwrap file for npm users, and a scheduled npm audit / npm audit signatures workflow in CI.

Can I verify that a Pi release binary matches its published source?

Yes — GitHub releases include a versioned, checksum-verified source archive along with the actual build script used for the official binaries, so you can rebuild a release yourself and confirm it matches what was published. A flag lets the rebuild use the release's pinned provider-model snapshot, so the result does not depend on live data.

Does Pi expect users to share their agent session data?

It asks, but doesn't require it. The README encourages publishing real Pi and other agent sessions to help improve agent tooling with real-world usage rather than benchmarks, and points to a companion tool, pi-share-hf, for publishing sessions to Hugging Face. The argument is that real tasks, tool calls, failures, and fixes improve agent tooling more than toy benchmarks do. Sharing is opt-in, so review what a session contains before publishing it.

Conclusion

Most coding agent CLIs bundle the model API, agent loop, and interface into one tool. Pi takes the opposite approach: a set of independently usable packages, with the CLI as one consumer among several. That makes it most interesting for teams that want to build their own agent products or automations on a shared provider abstraction and tool-calling runtime.

The practical lessons: the package split is the main reason to evaluate Pi, even if you never use its CLI. Its lack of a built-in sandbox is a deliberate, documented choice, so isolation is your job, and the documented container and micro-VM patterns are the place to start. Its supply-chain discipline and reproducible releases are a useful model for any team shipping developer tools.

The caveats are about fit. Teams that want a fully opinionated tool with permission prompts out of the box will find more setup work here, and the contribution model is strict for newcomers. Pin a version if you build on it.

If you are designing a custom agent product or a containerised agent deployment and want help with the architecture, our AI agent development team can help.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.