An AI agent discovery workshop is the single day of planning that decides whether your build succeeds. Most AI agent projects that fail don't fail because of the technology. They fail because the wrong problem was chosen, the scope was never quite clear, or three stakeholders walked into the project with three different mental pictures of what they were getting — and nobody noticed until the demo.
A discovery workshop is how you stop all of that before it gets expensive. One day, the right people in the room, real workflows on the wall, real numbers attached to them. By the end, you've decided what the agent will and won't do — and just as importantly, everyone in your business has decided the same thing.
This guide walks you through a full-day discovery workshop: who to invite, what to cover in each session, the specific questions to ask, and what outputs you should leave with. It works whether your team is running it internally or whether you've engaged an AI development partner to facilitate. (We run a lot of these. The structure below is the one that consistently produces a clean scope and a project we can confidently price.)
Quick answer: Run a single full day with 4–8 people who actually do the work, not observers. Structure it as five sessions — problem definition, automation opportunity assessment, scope definition (including what the agent explicitly won't do), integration/data requirements, and success metrics — and leave with five concrete documents: a workflow map, a prioritised opportunity list, a scope document, integration requirements, and agreed success metrics. That's what turns a vague brief into a proposal you can actually trust.
Before the Workshop
Who to invite (4–8 people is the sweet spot):
- The business owner or decision-maker who'll approve the spend.
- The operational manager most affected by the workflow being automated.
- Two or three team members who actually do the work that's being automated. These are your most valuable participants — not because they're senior, but because they know how the work really happens, not how the org chart says it happens.
- Your IT or technical lead if you have one.
- Your AI development partner, if you've engaged one.
A note worth taking seriously: do not invite people who are only going to observe. Every person in the room should have something to contribute. We've sat in workshops with twelve people, half of whom said nothing for six hours, and they were worse workshops than the four-person ones.
What to prepare:
- Access to your current workflow data — volume, timing, error rates, anything quantitative.
- Examples of the interactions you handle (anonymised if needed).
- A list of your existing tools and systems.
- Any previous attempts at automating this area, including the ones that didn't work — those are often more informative than the wins.
Pre-read for participants:
Send a one-pager 48 hours before: the purpose of the workshop, the specific area of the business you'll examine, and three to five questions they should think about before walking in. Primed participants produce dramatically better workshops than cold ones. This step costs you twenty minutes of prep time and saves you two hours of warm-up.
Session 1: Problem Definition (90 minutes)
Goal: Agree on the problem you're solving and why it actually matters.
Opening (15 minutes)
Quick round of introductions, then have each person complete this sentence out loud:
"The thing that frustrates me most about how we currently handle [this workflow] is…"
Listen closely. Take notes. These frustrations are your requirements — they're more useful than any formal requirements document, because they're real and they're concrete.
Workflow Mapping (45 minutes)
Draw the current workflow on a whiteboard or shared document. Every step, every decision point, every handoff. Ask the people who do the work to correct you as you go — and they will, often loudly. That's the point.
For each step, capture: who does it, how long it takes, how often it happens, how often it goes wrong, what goes wrong, what information it needs, what it produces.
This map will surprise you. Things you assumed were simple turn out to have three exceptions. Things you assumed were complex turn out to be five lines of logic. The people in the room will disagree about how the workflow actually works — and that disagreement is some of the most useful output of the whole day. You're not learning the process. You're learning the gap between how leadership thinks it works and how it actually works.
Impact Sizing (30 minutes)
Quantify the problem. Get numbers on the wall:
- How many times per week or month does this run?
- How much staff time does it consume in total?
- What does that cost?
- What goes wrong, how often, and what does that cost?
- What would the team do with the time saved?
These numbers become the basis for your ROI calculation and your success metrics later. They also tend to be the first time the team has actually counted, which can be a quiet shock.
With the problem sized, the next question is which parts of it are actually worth handing to an agent.
Session 2: Automation Opportunity Assessment (60 minutes)
Goal: Figure out which parts of the workflow are actually worth automating.
Automation Spectrum Exercise (30 minutes)
Take each step from your workflow map and place it on a spectrum:
- Fully automatable: Rule-based, structured, predictable. Same input always produces the same output. No judgment.
- Partially automatable: Mostly predictable, with exceptions. AI handles the common cases, humans handle the exceptions.
- Not automatable: Requires judgment, relationship, creativity, or authority that an agent shouldn't be making.
This exercise consistently produces the same insight in workshop after workshop: most people overestimate how much judgment their routine work requires, and underestimate how repetitive their day actually is. This isn't a criticism of the team — it's how memory works. The five interesting decisions a week loom larger than the three hundred boring ones.
Prioritisation Matrix (30 minutes)
Plot each automatable step on two axes:
- Impact: How much time or pain does automating this remove? How much does it improve life for customers or staff?
- Feasibility: How well-defined is the step? How clean are the inputs? How available is the data?
High impact + high feasibility = your first automation. Start there. Not with the most impressive-sounding automation. We've watched plenty of teams pick the second-most-valuable, hardest-to-build option first because it photographed better. Don't.
| Workflow Step | Automation Category | Impact (1–5) | Feasibility (1–5) | Recommended Priority |
|---|---|---|---|---|
| Order status lookup | Fully automatable | 5 | 5 | Build first |
| Invoice data extraction from PDFs | Fully automatable | 4 | 4 | Build first |
| FAQ and policy responses | Fully automatable | 4 | 5 | Build first |
| Appointment scheduling | Partially automatable | 4 | 3 | Build second |
| Lead qualification triage | Partially automatable | 5 | 3 | Build second |
| Refund decision approval | Partially automatable | 3 | 2 | Build second |
| Contract negotiation | Not automatable | — | — | Human only |
| Complex complaint resolution | Not automatable | — | — | Human only |
Use this table as a starting template during the prioritisation matrix exercise. Your actual scores will differ — the point is to get concrete numbers on the wall so the room is debating real trade-offs rather than abstract preferences.
Session 3: Scope Definition (90 minutes)
Goal: Define exactly what the agent will do — and exactly what it won't.
In-Scope Definition (30 minutes)
List every specific interaction type the agent will handle. Be precise. Vague scope is the single biggest cause of overrun projects.
For each, write:
- What triggers it (the input)
- What information the agent needs
- What the agent does (the process)
- What it produces (the output)
Example:
- Trigger: Customer asks for order status
- Needs: Customer email or order number
- Process: Look up the order in Shopify, retrieve tracking status
- Output: A specific update with tracking link
The more concrete this list, the easier the build.
Out-of-Scope Definition (20 minutes)
Equally explicit about what the agent will not handle. This is the step most teams skip, and it's the step that quietly determines whether your project goes well.
Ask the room: "What are five things a customer might ask that we definitely do NOT want the agent to try to answer?" Write them on the wall and put a big red line through them. They go to humans, full stop.
Escalation Design (40 minutes)
Map every escalation path. For each out-of-scope situation or edge case:
- What does the agent say to the customer?
- Where does the conversation go?
- Who picks it up?
- What does the agent hand over with the conversation?
Draw the escalation flows on the whiteboard. They are exactly as important as the main flows — arguably more so, because they're what users remember when the agent reaches its limits. A graceful escalation can leave a customer happier than a perfect resolution. An ungraceful one can lose the relationship.
Session 4: Integration and Data Requirements (60 minutes)
Goal: Surface what the agent has to connect to, and whether your data can actually support what you want it to do.
Systems Inventory (30 minutes)
List every system the agent needs to read from or write to. For each:
- What data does it need from here?
- Does it need to write back, or just read?
- Who can give us API access? (Have we checked that this is possible?)
- Are there compliance restrictions on this data?
The "write or just read" distinction matters more than people realise. Read-only integrations are dramatically simpler, faster to build, and safer. Write access is where the genuine risk lives.
Data Quality Assessment (30 minutes)
The agent is only as good as the data it can see. For each data source:
- Is the data complete? Are there gaps that would cause the agent to fail?
- Is the data accurate? When was it last audited?
- Is the data consistent? Does the same fact live in multiple places with different answers?
Data quality problems found in discovery are cheap to fix. Data quality problems found after the build is live are very, very expensive — usually a combination of emergency fixes, rebuilt prompts, and a stretch of low trust in the agent that takes months to recover.
Session 5: Success Definition (45 minutes)
Goal: Agree on what success looks like in numbers, not adjectives.
Metric Agreement (30 minutes)
Write three to five metrics on the board. For each: a target number and a measurement method.
Worth considering:
- Automation rate (% of in-scope interactions handled without a human)
- Response time (average time from contact to first response)
- Customer satisfaction (CSAT)
- Human hours saved (per week)
- Error rate (% of responses requiring correction)
Get everyone in the room to agree on the targets. Disagreements here are important to surface and resolve now — stakeholders with different definitions of success will be disappointed by the same outcome. We've seen the same project be celebrated by the COO and called a failure by the head of customer service, purely because they had different metrics in their heads at the start.
Failure Definition (15 minutes)
Ask the room: "What does this look like if it's failing? At what point would we stop the project?"
It's an uncomfortable question. Ask it anyway. Defining failure clarifies the project as much as defining success. It also gives you a clean off-ramp if the project genuinely needs one — something most teams don't have, which is why they end up keeping projects alive long past their useful life out of pure sunk-cost momentum.
Workshop Outputs
By the end of the day you should walk out with:
- Workflow map — current state, annotated with volume, time, and pain points.
- Automation opportunity list — prioritised by impact and feasibility.
- Scope document — in scope, out of scope, and escalation flows.
- Integration requirements — systems, data, and access needs.
- Success metrics — agreed targets with measurement methods.
These five documents are the foundation for a development brief, a vendor evaluation, and a project tracking framework. Any half-decent AI development partner should be able to turn them into a specific, accurate proposal. Any partner who can't is a partner you've already learned something useful about.
Benefits of an AI Agent Discovery Workshop
A day away from normal work is a real cost, so it's fair to ask what you get back. These are the returns we see most consistently.
A Scope Everyone Has Actually Agreed To
The biggest payoff is alignment, not documentation. When the decision-maker, the operational manager and the people doing the work all watch the same workflow get drawn and argued over, they leave with one picture of the project instead of three. That shared picture is what keeps the demo from becoming the moment someone says "that's not what I thought we were getting." Written sign-off later formalises it, but the agreement itself happens in the room.
Estimates You Can Trust
Developers price uncertainty. A vague brief gets a padded quote or a low quote with change requests waiting behind it. A scope document with triggers, inputs, outputs and escalation paths lets a partner price the actual work. You also get a fairer comparison between vendors, because each one is quoting against the same written definition rather than their own interpretation of a conversation.
Problems Found While They're Still Cheap
Missing API access, inconsistent customer records, a compliance restriction on a data source nobody mentioned: every one of these turns up eventually. Discovery moves them to the point where fixing them means a phone call or a data clean-up task, not a rebuilt integration and a delayed launch. Session 4 exists almost entirely for this reason.
The Right First Automation
The impact-feasibility exercise stops teams from starting with the impressive option. Building the high-impact, high-feasibility step first gets something live quickly, produces real numbers, and earns the trust you need for the harder second phase. Picking the photogenic, difficult option first tends to stall the whole programme.
A Clear Definition of Failure
Agreeing in advance on what failure looks like gives you an honest off-ramp. Without it, struggling projects drift on because nobody wants to be the one to call it. With it, stopping becomes a decision the group already made, which is much easier to act on. It also protects the people who run the project, because a pause against agreed criteria reads as discipline rather than as a personal failure.
AI Agent Discovery Workshop Use Cases
The five-session structure works across very different workflows. These are the situations where we see it used most often.
Customer Support Deflection
A support team is buried in repetitive tickets and leadership wants an agent to "handle support." The workshop breaks that ambition into specific interaction types: order status lookups, FAQ and policy answers, and the complaints that must stay human. The mapping usually shows that a handful of request types make up most of the volume. The outcome is a narrow, buildable first scope and an escalation design the support lead has signed off, rather than a vague mandate to automate everything.
Finance Back-Office Processing
Accounts teams spend hours pulling data from supplier invoices and keying it into the accounting system. In the workshop, the people doing that work explain the exceptions: missing PO numbers, mismatched totals, suppliers who send scans instead of PDFs. Those exceptions become the out-of-scope and escalation list. The team leaves knowing which invoices the agent will process end to end and which it will route to a person with the extracted fields pre-filled.
Sales Lead Qualification
Sales leaders often want an agent that qualifies inbound leads, but each rep qualifies differently. The workshop forces the room to write down the actual criteria and decide which questions the agent may ask. It also surfaces the CRM write access the agent would need, which is frequently the slowest thing to arrange. The outcome is an agreed qualification rubric, a hand-off point to a human rep, and a realistic integration timeline.
Booking and Scheduling
Appointment scheduling looks simple until the team explains cancellations, double bookings, staff preferences and the customers who always call instead of booking online. Mapping those cases lets the group decide whether the agent books directly or proposes slots for a human to confirm. That single decision changes the build effort and the risk profile, and it is far cheaper to make on a whiteboard than mid-project.
Choosing Between Vendors
Some teams run the workshop before they have chosen a partner at all. The five outputs become the brief they send to several vendors, which makes the proposals directly comparable. A vendor who cannot respond specifically to a defined scope tells you something important before you have committed any budget.
Common Discovery Workshop Mistakes
Even teams that follow the structure above trip over the same handful of problems. Watch for these.
Arriving With the Answer Already Chosen
If leadership arrives having already decided "we're building a sales agent", the workflow mapping turns into a justification exercise. Participants sense the conclusion and stop raising the awkward details that would challenge it. Bring a business area and a problem, not a solution. If a particular agent idea is already popular, put it on the wall as one candidate and let the impact-feasibility matrix judge it alongside everything else.
Mapping the Documented Process Instead of the Real One
The SOP binder is not the workflow. Written procedures describe how the work was meant to happen when someone last updated them, not the workarounds staff use every day. If nobody who does the work daily is in the room, your map will be wrong in ways you only find during testing, when the agent meets an input nobody mentioned and fails in front of a customer.
Skipping the Out-of-Scope List
Teams love listing what the agent will do. The list of what it won't do is what protects your budget and your customers. Without it, every edge case that appears during the build becomes a debate about whether it was "obviously" included, and the developer either absorbs the work or raises a change request. Both outcomes damage the relationship.
Leaving Integrations and Ownership as "TBC"
"We'll sort out CRM access later" is how a four-week build becomes a ten-week build. Confirm API access, or at least who owns it, before you leave the room. The same applies to the outputs: if nobody is named to write up the five documents within 48 hours, the momentum from the day disappears within a week.
Agreeing on Adjectives Instead of Numbers
"Faster responses" and "happier customers" can't be measured, so they can't settle an argument later. "First response under 60 seconds for 80% of in-scope tickets" can. Vague success criteria mean different stakeholders will judge the same result differently, which is exactly the disagreement the workshop was meant to prevent.
If you've already run a workshop and the outputs feel vague, it is usually one of these five. Fixing it means a focused half-day follow-up on the weak area, not a full re-run. For the commercial side of that follow-up, our guide to AI agent pricing models explains how scope clarity changes the type of contract you can sensibly sign.
AI Agent Discovery Workshop Best Practices
The sessions above give you the structure. These habits are what make a given day productive rather than merely busy.
- Send the pre-read and insist people read it. A one-pager with the purpose, the business area and three to five questions turns the first hour from orientation into real work. Follow up the day before with anyone who hasn't opened it.
- Keep the facilitator neutral. Whoever runs the room should not be the most senior person or the one with the strongest view on the outcome. Their job is to draw out the quiet operators and challenge assumptions, which is hard to do while defending a preferred answer.
- Capture disagreement on the wall, not in side conversations. When two people describe the same step differently, write both versions down and resolve them before moving on. Those gaps between how leadership thinks the work happens and how it really happens are the most useful material you'll get all day.
- Put a number on every step you can. Volume, time per occurrence and error rate turn opinions into trade-offs. Rough estimates are fine; a figure on the board beats a confident adjective.
- Design escalation with the same care as the main flow. For every out-of-scope case, decide what the agent says, where the conversation goes, who picks it up and what context travels with it. Users remember the hand-off far more than the routine answers.
- Separate read access from write access. Mark every integration as read-only or read-write. Read-only connections are faster to build and safer to run, so a first phase that only reads is often the smarter starting point.
- Write the failure condition down. Agree on the point at which you would pause or stop the project and record it next to the success metrics. It turns a painful future argument into a check against something already agreed.
- Name an owner and a deadline for the write-up. One person documents the five outputs within 48 hours and circulates them for corrections. Questions that surface in that review are cheaper to answer now than after development starts.
Related guides
- How to write an AI agent scope of work
- How to write a good AI agent brief
- How to evaluate AI agent vendors
- How much it costs to build an AI agent
- AI agent development services
After the Workshop
Within 48 hours: Document the outputs cleanly and share them with everyone who attended. Ask for corrections. Confusion that surfaces in review is gold — it means something wasn't clear in the room and needs resolving before development starts, not after.
Within one week: Share the scope document with your AI development partner for a project estimate. A well-run discovery workshop dramatically reduces estimate uncertainty. Developers can price against a defined scope rather than a vague brief, which means the proposal you get back is one you can actually trust.
Before development starts: Have the team review and sign off on the scope document. Scope changes after sign-off are manageable. Scope disagreements after development starts are expensive — and almost always trace back to something that was assumed rather than written down.
Talk to us about running a discovery workshop — we facilitate them as part of every client engagement, and we're happy to run one for your team before you commit to a build, even if you ultimately go elsewhere.
Frequently Asked Questions
How long does an AI agent discovery workshop typically take?
A full discovery workshop runs one day — roughly six to seven hours of structured sessions with breaks. Some teams compress it into a half-day by pre-completing the workflow mapping exercise before the session, but the full day consistently produces better outputs because it allows enough time for disagreements to surface and resolve. Rushing the discovery almost always costs you more time later in the build.
Who should facilitate the AI discovery workshop?
An experienced facilitator — ideally someone from your AI development partner — produces the best results because they know which questions to push on and can spot scope risk in real time. If you're running it internally, appoint a facilitator who is not the most senior person in the room: seniority tends to suppress the honest answers from the people who actually do the work, which is the opposite of what you need.
What if we don't have workflow data ready before the workshop?
You can still run a useful workshop without clean data, but you'll spend more time estimating and less time deciding. Encourage participants to pull whatever numbers they can in the days before — even rough figures for how often a task runs and how long it takes are better than nothing. The impact sizing session is significantly more valuable when the team has had a chance to look at actual volume beforehand.
How is a discovery workshop different from just writing a requirements document?
A requirements document captures what someone already decided they want. A discovery workshop surfaces what the team actually needs — which is often different, and sometimes very different. The workshop format brings the people who do the work into the same room as the people who approve the spend, and the gap between those two perspectives is frequently the most valuable output of the whole day. Requirements documents written without that conversation tend to describe the problem as leadership understands it, not as it exists.
Can we run the discovery workshop ourselves, or do we need an external facilitator?
You can run it internally, and the structure in this guide is designed to work without external help. The honest caveat is that an internal facilitator who is also a stakeholder in the outcome will find it harder to push back on scope creep or challenge assumptions from senior colleagues. External facilitation pays for itself quickly in cleaner scope and fewer mid-project surprises, but a well-prepared internal facilitator who commits to staying neutral can run an effective session.
How do we know if the scope we defined in the workshop is actually buildable?
Share your scope document with your AI development partner within a week of the workshop and ask them to flag anything that looks ambiguous, technically risky, or underspecified before they price the work. The five outputs from the workshop — workflow map, opportunity list, scope document, integration requirements, and success metrics — give a competent development team enough to give you a reliable estimate. If they can't give you a specific answer based on those documents, that tells you something about them.
What happens if stakeholders disagree during the workshop?
Disagreements during the workshop are valuable, not a problem. They mean you're surfacing something that would otherwise have been an assumption. The facilitator's job is to get the disagreement on the wall, understand the source of it, and work toward a written decision before the day ends. The only dangerous disagreement is one that stays hidden — those resurface during development as scope changes, which is the most expensive time to resolve them.
Conclusion
AI agent projects rarely fail on model quality. They fail because the team chose the wrong workflow, never wrote down what the agent shouldn't do, or discovered halfway through the build that a key system had no usable API. A one-day discovery workshop moves all of those discoveries to the cheapest possible moment: before anyone has written code or signed a contract.
The structure matters less than the people and the outputs. Put the people who actually do the work in the room, quantify the problem with real volumes, place each step on the automation spectrum, and draw the escalation paths as carefully as the happy paths. Leave with five written documents, not a shared feeling.
Two caveats. First, a workshop can't fix bad data; it can only expose it, and the data clean-up may need its own budget. Second, the scope you agree on is a starting point. Live usage will show you edge cases nobody predicted, so plan for an iteration phase after launch.
If you want an outside facilitator who has run these sessions across industries, book a call with our AI agent team and we'll help you plan the day.
