You have a budget approved, a developer or agency shortlisted, and a rough idea of what the agent should do: answer support emails, qualify leads, process documents. What you usually don't have is a written AI agent scope of work that both sides would read the same way. That gap is where most AI projects lose money. The client assumes the agent will handle refunds; the developer quoted for order-status lookups only. Nobody agreed what "65% automation" means or how it gets measured. Three months later, both parties are arguing over an email thread instead of shipping.
This matters more for AI agents than for ordinary software because the behaviour is probabilistic. An agent will sometimes answer things it was never meant to touch, and it will hit integrations that looked trivial in a sales call. A good scope of work turns those surprises into decisions made in advance: what is in scope, what is explicitly out, when the agent escalates, which systems it reads and writes, and what counts as acceptance.
This guide gives you a complete, section-by-section SOW template you can copy, using a realistic e-commerce support agent as the worked example. It covers project goals and success metrics, in-scope and out-of-scope interactions, escalation rules, technical and non-functional requirements, deliverables, ownership, and acceptance criteria. It finishes with the mistakes we see most often in AI agent scopes and answers to common questions about writing and signing one.
Why Most AI Projects Go Wrong Before They Start
Most AI agent projects that fail or disappoint do so for reasons that were visible before a line of code was written. A scope that was too vague. Integrations that were assumed but not specified. Success metrics that were never agreed. Edge cases nobody surfaced in the kickoff.
A good scope of work prevents most of these failures. It forces clarity before money changes hands. It creates a shared understanding between client and developer. And — maybe most importantly — it gives both parties a reference point when something gets disputed three months in.
What follows is a complete template for an AI agent SOW. Use it to brief a developer, to evaluate a proposal you've received, or to structure the pre-project conversation that determines whether a build is going to succeed.
AI Agent Scope of Work Template
Section 1: Project Overview
1.1 Background
Describe the business context in two to three paragraphs:
- What does the business do?
- What problem is this agent solving?
- How is this handled today?
- What's prompting this project now?
Example: "Acme Retail is a UK-based online clothing retailer processing 2,000 orders per month. Customer support is currently handled by two part-time staff who spend approximately 35 hours per week responding to support emails. The majority of queries are order status checks, return requests, and product questions. Response times average 8–12 hours. The business is growing and cannot scale support headcount proportionally."
1.2 Project Goal
One sentence. What's the agent supposed to achieve?
Example: "Build and deploy an AI customer support agent that handles at least 65% of inbound support emails automatically, reducing average response time to under 5 minutes."
1.3 Success Criteria
List three to five measurable outcomes that define project success. These become the basis for acceptance at delivery.
| Metric | Target | Measurement method | Measurement window |
|---|---|---|---|
| Deflection rate | ≥65% | System logs | 90 days post-launch |
| Average first response time | <5 minutes | Email timestamp log | 90 days post-launch |
| CSAT score | ≥4.0 / 5 | Post-conversation survey | 90 days post-launch |
| Weekly human hours on in-scope queries | <10 hours | Team time tracking | 90 days post-launch |
Section 2: Agent Scope
2.1 In Scope — What the Agent Handles
List every interaction type the agent is expected to handle. Be specific.
| Interaction type | Description | Resolution path |
|---|---|---|
| Order status query | Customer asks where their order is | Agent retrieves from Shopify, provides tracking link |
| Return request | Customer wants to return an item | Agent checks eligibility, generates return label if eligible |
| Product question | Customer asks about product details | Agent answers from product catalogue |
| Delivery issue | Customer reports non-arrival or damage | Agent logs report, escalates to human team |
2.2 Out of Scope — What the Agent Does NOT Handle
Be equally specific about what the agent won't do. Ambiguity here is what causes scope creep — usually three months in, when someone says "but I assumed the agent would also handle…"
- Complaints involving legal claims or formal dispute processes
- Requests for partial refunds outside the defined return policy
- Custom orders or bespoke product requests
- Any query not listed in the In Scope table above
2.3 Escalation Definition
Define exactly when and how the agent escalates to a human:
| Trigger | Escalation action | Escalation destination |
|---|---|---|
| Query type not in scope | Agent acknowledges, forwards to human queue | Support inbox |
| Customer expresses significant frustration | Agent acknowledges, offers human callback | Priority queue |
| Return outside policy window | Agent explains, offers exception request form | Exceptions inbox |
| Classification confidence below threshold | Agent drafts response for human review | Draft queue |
Section 3: Technical Specification
3.1 Communication Channels
List every channel the agent will operate on:
- Website chat widget (specify platform: Intercom / Crisp / custom)
- Email (specify: inbound address, send-from address)
- WhatsApp (specify: WhatsApp Business account required)
- SMS
- Other: ___
3.2 Integrations Required
List every system the agent needs to connect to, with read/write requirements:
| System | Purpose | Access required | Read / Write |
|---|---|---|---|
| Shopify | Order data, customer data | Admin API key | Read + limited write |
| Royal Mail API | Tracking information | API key | Read only |
| Return portal | Return eligibility and label generation | API key | Read + write |
| Gmail | Inbound email reading and outbound sending | OAuth | Read + write |
3.3 Data and Authentication
How will the agent verify customer identity before accessing account data?
Example: "Customer identity verified by matching email address in the incoming email against the Shopify customer record. For sensitive actions (refund initiation), additional verification via order number required."
3.4 Knowledge Base
What information will the agent use to answer questions?
| Content type | Source | Format | Update frequency |
|---|---|---|---|
| Product information | Shopify product catalogue | Live API | Real-time |
| Return policy | Policy document | Static document | Updated as needed |
| Shipping information | Internal FAQ document | Static document | Monthly |
| Brand FAQs | Existing help centre | Static documents | As needed |
3.5 LLM and Infrastructure
Specify (or ask the developer to specify):
- Which LLM provider and model (see Anthropic's documentation or an equivalent provider's docs for current model specifications)
- Hosting infrastructure (cloud provider, region)
- Data residency requirements
- Expected query volume (queries per day/month)
- Expected response time requirement (p95 target)
Section 4: Non-Functional Requirements
4.1 Security
These should mirror the practices covered in AI agent security:
- Data encryption requirements (in transit, at rest)
- Authentication and authorisation requirements
- Audit logging requirements
- Third-party data processing agreements
- Penetration testing requirements (if applicable)
4.2 Compliance
List applicable regulatory requirements:
- GDPR / UK GDPR data handling requirements — see the ICO for current guidance
- Industry-specific requirements (FCA, HIPAA, etc.)
- Consumer protection requirements
4.3 Performance
| Requirement | Target |
|---|---|
| Response time (p50) | <2 seconds |
| Response time (p95) | <5 seconds |
| Uptime | ≥99.5% |
| Error rate | <1% |
4.4 Monitoring and Alerting
Define what monitoring is required and who receives alerts:
- Real-time error rate monitoring — alert if error rate exceeds 5% in any 30-minute window
- Daily performance dashboard accessible to client
- Weekly automated performance report delivered to [email]
- Monitoring access: client has read-only access to all monitoring dashboards
Section 5: Deliverables
Every deliverable the developer is responsible for at the end of the project:
- Deployed AI agent on [channels specified in 3.1]
- All integrations listed in 3.2 connected and tested
- Knowledge base loaded with content listed in 3.4
- Monitoring dashboards configured and accessible to client
- Handover documentation (architecture, how to update knowledge base, how to adjust escalation rules)
- Source code transferred to client repository
- Post-launch support period: [X weeks] of monitoring and rapid response included
Section 6: Timeline and Milestones
| Milestone | Description | Target date |
|---|---|---|
| Kick-off | Scope confirmed, access granted | Week 1 |
| Integration complete | All API connections live in staging | Week ___ |
| Knowledge base complete | All content loaded and reviewed | Week ___ |
| Internal testing complete | Developer testing against test set | Week ___ |
| Client UAT | Client testing period | Week ___ |
| Launch | Agent live in production | Week ___ |
| Post-launch review | 30-day performance review | Week ___ |
Section 7: Ownership and IP
Explicit statements on:
7.1 Code ownership: All code, prompts, configuration, and documentation created for this project is the sole property of [Client name] upon final payment.
7.2 Data ownership: All customer data, conversation data, and knowledge base content is the sole property of [Client name]. The developer has no rights to use this data for any purpose other than delivering the project.
7.3 Third-party licences: List any third-party tools, APIs, or services used, and confirm the client has or will have their own accounts for these.
Section 8: Acceptance Criteria
Define how delivery will be accepted:
Technical acceptance: All integrations pass automated test suite. Error rate below 1% over 48-hour monitoring period.
Functional acceptance: Agent correctly handles [X] representative test scenarios — built the way we outline in AI agent testing and QA — covering all in-scope interaction types.
Performance acceptance: Response time p95 below [X] seconds under [Y] concurrent users.
Success criteria acceptance: Measured against success metrics in Section 1.3 at [30/60/90] days post-launch.
Benefits of a Written AI Agent Scope of Work
Filling in eight sections before a build starts feels like overhead. Here is what that effort buys you, on both sides of the contract.
Quotes You Can Actually Compare
Two vendors quoting from a one-paragraph brief are pricing two different projects. One assumes read-only Shopify access; the other assumes refunds. Once both quote against the same in-scope table, integration list, and acceptance criteria, the gap between their numbers reflects real differences in approach, not guesses about what you meant. That makes the shortlist conversation about method and risk rather than about whose assumptions were more generous.
Fewer Change Requests and Disputes
An explicit out-of-scope list settles most "I assumed it would also..." arguments before they start. When a new requirement comes up in week five, both sides can point to the document and agree it is new work. That conversation is short and polite when the scope is written, and long and expensive when it lives in memory and email threads that each side reads differently.
Objective Acceptance
Success metrics with a baseline, a target, and a measurement window turn acceptance into a check rather than a negotiation. The developer knows what they are being judged on, and the client knows what they are paying for. Tying part of acceptance to a 30 to 90 day window also rewards building something that holds up in production, not just in the demo.
Risks Surface Early
Writing out integrations with read and write access, identity verification, and data residency forces the questions that usually appear mid-build. A legacy system with no API, a compliance requirement nobody mentioned, or an unclear owner for credentials is cheaper to discover in a scoping session than after the first sprint. Some projects get reshaped or split into phases at this stage, which is a good outcome.
Clean Handover and Ownership
Sections 5 and 7 make sure you end up owning the code, prompts, configuration, and data, with documentation good enough for someone else to maintain it. If the relationship with the developer ends, the agent does not have to end with it.
AI Agent Scope of Work Use Cases
The template uses a support agent as its example, but the same structure does different jobs depending on where you are in a project. Sales qualification agents, scheduling agents, and document-processing agents all fit it; only the example content in each section changes.
Briefing a Developer Before Proposals
The most common use: a client fills in Sections 1 and 2 (context, goal, success metrics, in-scope and out-of-scope interactions) and sends them to two or three developers. Each developer completes the technical sections as part of their proposal. The result is a set of proposals built on the same foundation, and an early read on which teams can answer specific technical questions and which fall back on generic language.
Reviewing a Proposal You Already Received
Sometimes the vendor writes the first draft. Mapping their proposal onto the eight sections shows what is missing: no escalation rules, integrations named without access levels, acceptance tied only to launch. Those gaps become your list of questions before signing, and the answers often change the price or the timeline. A vendor who fills them in quickly is usually a safer bet than one who resists putting specifics on paper.
Internal Agents Built by Your Own Team
Internal builds skip the contract, but they still suffer from unclear scope. A knowledge assistant for staff or a document-processing agent for finance benefits from the same in-scope list, escalation rules, and success metrics. Here the SOW acts as an agreement between the business team that requested the agent and the engineering team building it, and it keeps the project from quietly growing into a general-purpose assistant nobody can test.
Scoping Phase Two of an Existing Agent
After launch, teams often want to add interaction types: refunds after order status, or a new channel such as WhatsApp. Running the expansion through the template, with its own success criteria and acceptance, keeps the new work from degrading what already works. It also gives you a clean record of what changed, which matters when you review performance later and need to know whether a dip came from the new scope or from something else.
Where Even a Good SOW Fails You
Two honest caveats before you treat this as a complete safety net. First, an SOW only protects you if both sides actually read it after signing. We've watched projects where the SOW was beautifully detailed and then immediately ignored, because the day-to-day conversation drifted into Slack and decisions got made that contradicted what was on paper. The SOW is a reference point, not a contract that runs itself.
Second, the level of specificity this template asks for can paralyse a kickoff if you treat it as a checklist to fill in. The point isn't to answer every question perfectly on day one — it's to surface the questions you can't answer yet, because those are the things that will bite you in week six if they aren't surfaced. A "we don't know yet, we'll figure out by week 2" answer in Section 3.5 is more useful than a confident wrong answer.
Common AI Agent Scope Mistakes to Avoid
The template above is only as good as what goes into it. These are the gaps that show up most often when we review scopes written by clients or other vendors.
Success Metrics Without a Baseline
"Reduce response time by 80%" means nothing if nobody measured today's response time. Capture the baseline in Section 1.1 (current volume, current handling time, current CSAT) before the build starts, or the acceptance conversation turns into an argument about what the starting point was.
Integrations Listed by Name Only
"Connects to Shopify" hides a lot of decisions: which API, which scopes, whether the agent can write, who owns the API key, and what happens when the API rate-limits. Every row in Section 3.2 should state read or write access and who provisions credentials. Integration work is usually the largest variable in AI agent development cost, so vagueness here becomes a budget problem.
No Owner for the Knowledge Base
Agents answer from whatever content they are given. If the SOW doesn't say who keeps the return policy and FAQs current after launch, the agent will confidently quote last year's policy. Name the owner and the update cadence in Section 3.4.
Acceptance Tied Only to Launch
A demo that works on launch day says little about week six. Tie at least part of acceptance to measured performance over a 30–90 day window, and agree upfront how post-launch maintenance is handled.
Escalation Rules That Stop at "Hand Off to a Human"
Saying the agent escalates is not the same as saying where the conversation goes, who picks it up, and how fast. If the destination queue has no owner or no response target, escalated customers wait longer than they did before the agent existed. Each row in Section 2.3 should name a queue, an owner, and what the customer is told while they wait.
AI Agent Scope of Work Best Practices
The mistakes above have a mirror image: habits that make a scope useful from kickoff through the 90-day review.
- Split ownership by section. The client writes the business context, goals, and scope; the developer writes the technical and non-functional requirements. Both sign the whole document. Each side writes what it knows best, and neither can later claim the other's section was a surprise.
- Write the out-of-scope list first. It is easier to agree on what the agent will never do than to enumerate everything it will. Start with refunds outside policy, legal complaints, account changes, and anything regulated, then fill in the in-scope table around those boundaries.
- Capture baselines before the build. Record current volume, handling time, response time, and satisfaction scores for the in-scope queries. Without them, a success metric is a number with nothing to compare against.
- Mark unknowns with an owner and a date. "Model choice to be confirmed by week 2, owner: developer lead" is a valid entry. A blank or a confident guess is not. Unknowns tracked this way get resolved; unknowns hidden get discovered late.
- Include a test set in the acceptance criteria. Agree on a set of representative conversations, including awkward edge cases and out-of-scope requests the agent should refuse or escalate. Run it before launch and again after any prompt or model change.
- Keep the SOW where decisions are made. Link it from the project channel and update it through a lightweight change log whenever scope shifts. If a decision in chat contradicts the document, either the document changes or the decision does.
- Schedule the 30-day review at signing. Put the post-launch review on the calendar when the contract is signed, with the success metrics from Section 1.3 as the agenda. Reviews booked later tend not to happen.
Related guides
Using This Template
Share this template with any AI developer you're evaluating. Ask them to complete Sections 3 and 4 as part of their proposal. The quality and specificity of their answers tells you more about their capability than their portfolio does.
For developers who can't fill in Section 3 specifically — who can't tell you which LLM they plan to use and why, who can't describe the integration architecture — that's important information before you sign anything.
If you want help drafting a scope of work before you decide who to work with — even if you ultimately don't work with us — we're happy to do that.
Talk to us about your project — no commitment, just a conversation.
Frequently Asked Questions
What is an AI agent scope of work and why do I need one?
An AI agent scope of work (SOW) is a written document that defines exactly what an AI system will do, what it will not do, how success is measured, and what both parties are responsible for delivering. You need one because AI projects are particularly prone to misalignment — the technology is new, expectations are often inflated, and integration complexity is easy to underestimate. A clearly written SOW prevents the "but I assumed it would also…" conversations that derail projects and erode client-developer trust.
How detailed does an AI agent SOW need to be?
It needs to be detailed enough that both parties can independently read it and reach the same understanding of what is being built. In practice, that means every integration listed by name with its read/write requirements, every interaction type the agent will handle listed explicitly, success metrics that are measurable and time-bound, and a clear escalation definition. Vague language like "the agent will handle customer queries" is not a scope — it is the beginning of a dispute.
Who should write the AI agent scope of work — the client or the developer?
The best SOWs are written collaboratively. The client should own Sections 1 and 2 — the business context, what the agent will and will not handle, and what success looks like. The developer should own Sections 3 and 4 — the technical architecture, integration design, and non-functional requirements. Both parties should review and sign off on the complete document. If a developer hands you a proposal without completing the technical sections, treat that as a red flag.
What happens if something is missing from the scope of work after the project starts?
Anything not in the scope of work is out of scope by default. If a new requirement emerges mid-project, it should go through a formal change request process — the new work is scoped, priced, and agreed before it begins. This is not a technicality; it is what keeps AI projects financially predictable. The SOW template in this article includes an explicit out-of-scope list precisely to prevent the "but I assumed" problem.
How long does it take to write an AI agent scope of work?
For a well-understood use case — a customer support agent for a Shopify store, for example — a thorough SOW can be drafted in two to four hours if the right people are in the room. More complex projects with multiple integrations, compliance requirements, or novel agent behaviour may take several sessions across a week. The time spent on the SOW is almost always repaid by the time it saves in disputed work, rework, and misaligned expectations later in the project.
What should I check before signing an AI agent SOW from a developer?
Check that Sections 3 and 4 are completed specifically — not with placeholders like "TBD" or "as agreed." Confirm the success criteria in Section 1.3 are measurable and tied to a measurement window, not vague outcomes. Verify the out-of-scope list is explicit, not implied. Make sure the acceptance criteria define how delivery will actually be tested. And confirm that IP and data ownership are stated clearly: you should own the code and all customer data with no exceptions.
Can I use this SOW template for AI agents beyond customer support?
Yes. The structure applies to any AI agent project — sales qualification agents, internal knowledge assistants, document processing agents, scheduling agents, or anything else. Swap out the example content in each section for your use case. The principles are the same regardless of the agent type: define what it handles, what it does not handle, how it escalates, what systems it connects to, and how success is measured.
Does a scope of work change how much an AI agent costs?
Usually it makes the price more accurate rather than higher. Without a scope, developers pad estimates to cover unknowns, or quote low and recover the difference through change requests. A specific SOW lets a vendor price integrations, knowledge-base work, and testing directly, and makes quotes from different vendors comparable. It also exposes expensive requirements early, such as write access to a legacy system, so you can decide whether they belong in phase one or a later release.
Conclusion
AI agent projects rarely fail because the model was too weak. They fail because the client and the developer had different pictures of what was being built, and nothing on paper settled the difference. A scope of work is the cheapest fix for that: a few hours of structured conversation that turns assumptions about channels, integrations, escalation, and success into written decisions.
The parts that matter most are the ones people skip. An explicit out-of-scope list prevents the "I assumed it would also..." dispute. Measurable success criteria with a baseline and a time window make acceptance objective. Named integrations with read and write access expose the real cost. Clear ownership of code, prompts, and data protects you if the relationship ends.
Keep the caveats in view too. A detailed SOW does nothing if decisions drift into chat threads and contradict it, and trying to answer every question perfectly on day one can stall a kickoff. Mark unknowns honestly and give each one an owner and a date.
Your next step: take this template, fill in Sections 1 and 2 yourself, and ask any developer you're evaluating to complete Sections 3 and 4. If you'd like a second pair of eyes on the draft, book a call with our team.
