Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

How We Built a Support Agent for an E-commerce Brand: A Case Study

An AI support agent case study: a UK fashion retailer spent 40 hours a week on support emails. We built one that handled 71% — the build and 90-day numbers.

How We Built a Support Agent for an E-commerce Brand: A Case Study — Woyce Technologies

Growing e-commerce brands hit the same wall: order volume climbs, and the support inbox climbs with it. "Where is my order?", "Can I return this?", "Does this run small?" Each email is easy to answer, but together they swallow dozens of staff hours a week, push response times past half a day, and start showing up in reviews. Hiring more people fixes the backlog for a while and then the volume catches up again.

This AI support agent case study shows what happened when a UK fashion retailer tried a different approach. The brand was shipping roughly 2,000 orders a month with two part-time support staff and a founder answering overflow emails on Sunday nights. The brief was to cut the manual workload without making the customer experience worse, and the second half of that brief shaped every design decision.

Below you'll find the full story: how the support volume was categorised before any code was written, the integration layer with Shopify, courier APIs, the returns portal and Gmail, how the confidence-based classifier decides what to auto-send and what to draft for review, and how the agent was tuned to the brand's voice. It also covers the three things that broke in the first three weeks, the 90-day results, the full cost and payback calculation, and what we'd do differently next time.

If you run support for an online store and want a realistic picture of what an AI agent can take off your team's plate, this is a grounded reference point.

The Client

A UK-based fashion retailer selling through their own website and two marketplaces. Monthly order volume: 1,800–2,400 orders. Support team: two part-time staff plus the founder handling overflow.

They came to us in January with a clear problem: support was consuming 35–40 hours per week across the team, and the volume was growing faster than they could manage. Their average email response time was 11 hours. Reviews were starting to mention slow support. The founder was spending Sunday evenings answering order queries instead of working on the business.

Their ask: build an AI support agent that reduces the manual support load without making the customer experience worse. That second clause mattered a lot to them — and to us.

The Scope We Agreed

Before writing a line of code, we spent a week with the client mapping their actual support volume. They gave us access to their support inbox and we categorised every email received over the previous 30 days.

The results:

CategoryVolume% of total
Order status / tracking31234%
Return requests18721%
Product questions (size, material, care)14316%
Delivery issues (damaged, missing)9611%
Account queries (login, address change)748%
Other / miscellaneous9110%

The first four categories — 82% of volume — had clear, automatable resolution paths. We scoped the agent to handle those. We left "other / miscellaneous" and any delivery issues requiring compensation to the human team.

Agreed success metric before build: 65% deflection rate within 90 days of launch.

The AI Support Agent We Built: A Case Study

Integration Layer

The agent integrates with:

  • Shopify — for order data, tracking information, customer details, and order status
  • Royal Mail and DPD APIs — for real-time tracking data
  • Return portal — their existing returns management system (a third-party tool) via API
  • Gmail — reading inbound support emails and sending replies via their support address

The email integration was the most complex piece. We built a system that reads new emails in the support inbox, classifies the query type, retrieves relevant context from Shopify and courier APIs, generates a response, and either sends it automatically (for high-confidence resolutions) or drafts it for human review (for lower-confidence or policy-edge cases).

The Classification System

Every incoming email is classified before any response is attempted. The classifier uses the email subject, body, and any Shopify order data linked to the customer's email address to determine:

  1. Query category (from our taxonomy above)
  2. Confidence score (how certain we are about the classification)
  3. Recommended action (auto-respond, draft for review, escalate immediately)

Auto-respond threshold: 85%+ confidence on categories we know the agent handles well (order status, return eligibility). Below that, it drafts for human review. The threshold was deliberately conservative — we'd rather have a human approve a perfect response than send a confident wrong one.

Response Generation

For order status queries: the agent retrieves the order from Shopify, gets the current tracking status from the relevant courier API, and generates a response that includes the specific tracking link, last scan location, and estimated delivery date. The response is personalised — it uses the customer's name and references the specific items ordered.

For return requests: the agent checks the order date against the return policy (28 days from delivery), determines eligibility, and if eligible, generates a return label from the returns portal and sends it with instructions. If outside the window, it explains clearly and provides the contact for exception requests.

For product questions: the agent searches the product catalogue for the relevant item and answers from the product specifications. Size guide queries reference the actual measurement tables. Care instructions come from the product metadata.

Tone and Style Calibration

We spent more time on this than clients typically expect. The client had a distinct brand voice — warm, direct, slightly irreverent. Generic AI responses would have felt off-brand and undermined the customer experience they'd worked hard to build.

We gave the classifier 50 examples of good and bad responses from the existing support inbox, with annotations explaining what made each one good or bad. That informed the prompt design and gave us a reference set for tone evaluation during testing.

The Human Review Queue

Not everything is auto-sent. The agent drafts responses for:

  • Classifications below the confidence threshold
  • Return requests outside standard policy (partial returns, condition disputes)
  • Delivery issues involving potential compensation
  • Any email containing words indicating distress or dissatisfaction with the brand

The human team sees a queue in their inbox tool with draft responses pre-populated. For most drafts, they read, approve, and send in under 30 seconds. The workload shifts from reading-researching-writing to reviewing-and-approving.

What Broke in the First Three Weeks

Problem 1: Marketplace order numbers. Customers who ordered through the marketplaces (not the direct site) used marketplace order numbers in their emails. Our Shopify integration looked up orders by Shopify order ID or customer email. Marketplace order IDs didn't match.

Fix: Added a mapping layer that extracts marketplace order numbers from email body text and looks them up via the marketplace APIs.

Problem 2: Bundle product descriptions. Some orders included bundle items — three products listed as one SKU. The agent was describing the bundle SKU number when asked about specific items within the bundle, which was confusing for customers.

Fix: Updated the product catalogue mapping to expand bundle SKUs into their component products before the response generation step.

Problem 3: Overly formal tone on escalations. When the agent escalated a query to the human queue, it sent the customer a holding message. The initial version sounded corporate and cold — "Your enquiry has been received and will be addressed by a member of our team."

The founder flagged this immediately: "That sounds like it's from a bank, not us."

Fix: Rewrote the holding messages in the client's brand voice. Small change, big difference in how escalations felt to customers.

None of these were catastrophic, but all three reinforced something we already believed: shadow mode and tight monitoring in the first month catch the problems that demos never will.

The Results at 90 Days

MetricBeforeAfter (90 days)Change
Auto-resolved without human~5%71%+66pp
Average first response time11.2 hours4.1 minutes-98%
Weekly support hours (human)37–40 hours9–12 hours-74%
Customer satisfaction (CSAT)3.9 / 54.5 / 5+15%
Negative reviews mentioning support3–4/month0–1/month-75%

We exceeded the agreed 65% deflection target by week eight (68%) and reached 71% by week twelve as we tuned the confidence thresholds and expanded product question coverage.

The CSAT improvement was the most surprising result. We'd expected deflection to improve satisfaction (faster responses) but not to that degree. Post-survey comments consistently mentioned speed — "got an answer in minutes, amazing" — and personalisation — "the reply actually referenced my specific order."

The founder's Sunday evenings are no longer spent in the support inbox.

Benefits of an AI Support Agent for E-commerce

Replies in minutes instead of hours

Response time moved from most of a working day to a few minutes. For an online shopper, that gap is the difference between a reassured customer and one who opens a second ticket, files a chargeback, or writes a review about being ignored. Fast answers also stop the follow-up emails that pile on when someone hasn't heard back, so speed reduces total volume as well as improving the experience for each individual customer.

Staff time moves to the cases that need people

The two part-time staff went from researching and writing every reply to reviewing drafts and handling the genuinely tricky emails: compensation, condition disputes, unhappy customers. Those are the conversations where a person changes the outcome. The routine lookups that used to fill their week now take seconds of approval, which makes the role more interesting and more valuable to the business. It also makes experienced staff easier to retain over time.

Answers that reference the actual order

Because the agent reads live Shopify and courier data, replies include the specific tracking link, last scan location, and items ordered. Customers noticed that in their survey comments. Personalised, accurate answers build more trust than a template, and they remove the back-and-forth that happens when a generic reply forces the customer to ask again.

Capacity that grows with order volume

Support load used to rise in step with orders, which meant hiring every time the business grew. With the agent taking the routine categories, a rise in orders mostly raises API usage rather than headcount. Peak periods such as sales and holidays become manageable without temporary staff who need training on the brand voice.

Fewer negative reviews about support

Slow replies had started appearing in reviews. Once response time dropped, mentions of support in negative reviews fell sharply. For a direct-to-consumer brand, review sentiment feeds directly into conversion, so this effect is worth counting even though it is harder to price than staff hours. It also means the founder no longer has to treat every review notification as a potential support crisis.

E-commerce AI Support Agent Use Cases

Order status and tracking

This was the largest category, at around a third of all email. The agent matches the sender to an order, queries the courier for the latest scan, and replies with the tracking link and expected delivery date. Previously each query meant opening Shopify, finding the order, copying a tracking number, and checking the courier site. Automating that path removed the biggest single block of manual work. When the tracking data shows a parcel has stalled, the agent flags the email for review instead of sending an optimistic estimate.

Return eligibility and labels

Return requests follow a clear policy: within 28 days of delivery, item condition permitting. The agent checks the delivery date, confirms eligibility, and generates a label from the returns portal. Requests outside the window get a clear explanation and a route for exceptions. Customers receive their label immediately instead of waiting a day, and staff only see the edge cases.

Product, sizing, and care questions

Questions about fit, materials, and washing instructions are answered from the product catalogue and size tables. The agent quotes the measurements for the specific item rather than a generic size chart. This category matters before purchase as well as after, since a quick, accurate sizing answer can save a sale and prevent a return.

Triage for delivery problems

Damaged or missing parcels often involve compensation, so the agent doesn't resolve them alone. It gathers the order and tracking details, classifies the issue, and drafts a response for a person to approve. The human makes the judgement call with all the context already assembled, which cuts handling time without handing compensation decisions to software.

Account and address changes

Login problems and address updates are routine but time-sensitive, especially when an order hasn't shipped yet. The agent can explain how to reset access and flag address changes against open orders for quick action. Fewer orders go to the wrong address, and customers get a clear next step right away. Anything involving account security, such as a request to change the email on an account, still goes to a person to verify.

The Cost and Payback

Build cost: £9,500

Monthly running cost: £220 (hosting, API costs, email processing)

Monthly maintenance retainer: £650 (weekly review, prompt updates as the product catalogue changes, integration maintenance)

Total first-year cost: £9,500 + (12 × £870) = £19,940

Value from time saved: 25–28 hours per week recovered at an effective rate of £18/hour = approximately £24,000 per year in recovered productive time.

Additional value: Reduction in negative reviews and associated brand damage. Not easily quantifiable but clearly meaningful for a direct-to-consumer brand.

Payback: approximately 5 months.

Common AI Support Agent Mistakes

Some of these we made ourselves on this project; others we see regularly when reviewing support agents built elsewhere.

Calibrating tone only on the main reply

We got to the right place on tone, but the holding-message issue should have been caught in testing, not after launch. Teams tend to polish the primary answers and forget acknowledgements, escalation notices, and out-of-policy explanations. Those are often the messages customers read when they're already frustrated. Tone review across every message type belongs on the pre-launch checklist.

Ignoring every sales channel but the main one

Marketplace order lookup was a predictable requirement given the retailer's sales channels. We should have scoped it into the build rather than retrofitting it in week two. Any brand selling on marketplaces will receive emails quoting order numbers its storefront doesn't recognise, and an agent that can't find the order can only reply with a generic answer.

Leaving monitoring until later

We had logging from launch, but the client-facing dashboard took three weeks to build. The first three weeks of operational data were available but not easily visible to the client. Early visibility accelerates the tuning cycle, because the people who know the business best can spot wrong answers and gaps sooner.

Setting the auto-send threshold too low

It is tempting to lower the confidence threshold to push the deflection rate up quickly. Every confident wrong answer sent to a customer costs more trust than a draft that waits thirty seconds for approval. Start conservative, then lower thresholds category by category once the review queue shows the agent is consistently right.

Building before categorising the inbox

Teams that skip the volume analysis end up automating the categories that are interesting rather than the ones that are frequent. A month of tickets sorted by type shows where the hours actually go and gives you a baseline to measure against later.

AI Support Agent Best Practices

  • Categorise a month of tickets first. Sort real emails by type and volume before writing code. Automate the high-volume categories with clear resolution paths, and leave the miscellaneous and judgement-heavy ones with people.
  • Agree a success metric before build. A target such as a deflection rate within 90 days gives the project a finish line and a basis for tuning decisions. Record the baseline response time and staff hours at the same time.
  • Connect to live order, courier, and returns data. Specific answers depend on real data. Map every channel the brand sells through, including marketplaces, so the agent can find any order a customer mentions.
  • Use confidence-based routing with a fast review queue. Auto-send only high-confidence answers in well-understood categories, and draft everything else for approval inside the tool the team already uses. Make approval a one-click action.
  • Escalate distress, compensation, and policy exceptions. Define the triggers explicitly and send those emails straight to a person with context attached. Don't leave the agent to decide when a refund or goodwill gesture is appropriate.
  • Calibrate tone on real examples across every message type. Use annotated good and bad replies from the existing inbox, and review holding messages and refusals as carefully as the main answers.
  • Run shadow mode and watch a dashboard from day one. Let the agent draft without sending for the first period, track accuracy by category, and keep weekly reviews going after launch as the catalogue and policies change. Feed every wrong answer back into the prompts and test set.

A fair caveat: this kind of result is realistic for businesses with clean order data and well-defined policies. If the underlying systems are messy — duplicate customers, inconsistent product data, undocumented exceptions — you'll spend half the project cleaning that up, and the deflection numbers come more slowly.

If This Sounds Like Your Business

The pattern we built — email triage, order lookup, return processing, auto-response with human review queue — is repeatable across e-commerce businesses at this scale. The specific integrations vary; the architecture is consistent.

Talk to us about your support volume — we'll map your current inbox against what the agent can handle and give you a realistic projection of what deflection rate you could expect. If the numbers don't work, we'll tell you.

Frequently Asked Questions

How much does it cost to build an AI support agent for an e-commerce store?

Build costs typically range from £6,000 to £15,000 depending on the number of integrations, complexity of your return and fulfilment policies, and how much custom tone calibration is needed. Ongoing costs — hosting, API usage, and maintenance — usually run £500–£1,000 per month. For a retailer processing 1,500+ orders per month, payback is commonly within 4–6 months through recovered staff time alone.

What percentage of support tickets can an AI agent actually handle without human involvement?

For e-commerce businesses with clean order data and well-defined policies, a realistic deflection rate is 60–75% within the first 90 days. The largest categories — order status, tracking, and return requests — are the most automatable. Queries involving compensation decisions, nuanced complaints, or undocumented policy exceptions almost always stay with the human team.

Will an AI support agent make my customer experience feel impersonal or robotic?

Only if tone calibration is skipped. In this case study, CSAT improved from 3.9 to 4.5 out of 5 after the agent launched — customers rated the experience higher, not lower. The key is training the agent on real examples of your brand voice and reviewing every response type before launch, including holding messages and escalation notifications.

What systems does an AI support agent need to integrate with?

At minimum: your order management or e-commerce platform (Shopify, WooCommerce, etc.) and your email inbox. For most retailers, you also need courier API access for real-time tracking and a returns portal integration. If you sell on marketplaces like Amazon or eBay, those order IDs need a separate lookup layer — a common oversight that adds build time if not scoped upfront.

How long does it take to build and launch an AI support agent?

For a scope similar to this case study — Shopify integration, two courier APIs, returns processing, and email triage — build time is typically 6–10 weeks. That includes a shadow mode period where the agent drafts responses without sending them, which is where most edge cases surface before they reach customers.

What happens when the AI agent gets it wrong?

The system is designed so wrong answers are caught before they reach customers. Anything below an 85% confidence threshold is drafted for human review rather than auto-sent. The human team approves in under 30 seconds for most drafts. Agents don't eliminate human judgement — they shift it from high-volume routine queries to the small percentage that genuinely needs it.

Is an AI support agent worth building if I already use a helpdesk tool like Zendesk or Gorgias?

Yes — and they work alongside each other. The agent handles the automated triage and response layer; your helpdesk tool remains the workspace for the human team. Most integrations we build connect directly to the helpdesk so the human review queue appears inside the tool the team already uses, rather than requiring them to switch to a new interface.

Conclusion

The retailer's problem was a familiar one: a support inbox growing faster than the team, slow replies, and a founder losing evenings to order-status emails. The solution wasn't a generic chatbot. It was an agent that could look up real orders, real tracking scans, and real return eligibility, then decide whether it was confident enough to reply on its own or should hand a draft to a person.

A few lessons carry over to almost any online store. Categorise your inbox before you build, because the volume breakdown tells you what's worth automating. Live data integration is what turns generic answers into useful ones. Conservative confidence thresholds and a fast human review queue protect the customer experience while the system earns trust. And tone matters across every message type, including holding replies.

The caveat is that these numbers depend on clean order data and clearly documented policies. Messy catalogues, marketplace order IDs, and undocumented exceptions slow things down, and some queries, especially compensation decisions, should stay with people.

If your team is spending more than 20 hours a week on repetitive order and returns emails, start by exporting a month of tickets and categorising them. Our AI agent development team can help you turn that breakdown into a realistic deflection estimate.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.