Customer service
The Steward
Resolves, remembers, and never loses patience.
Motif · A steady flame, an open hand

Figure — Reported time saved
9–14%
less handling time per issue and more issues resolved per hour; up to 34% for newer staff
Handles password resets, order-status checks and FAQ triage, so people are freed to handle the emotional, complex, relationship-defining conversations.
Less handling time per issue and more issues resolved per hour; up to 34% for newer staff. Every figure on this page is traced to a named source below, with its method and its limits.
Read the evidence01
What these agents do
Service agents handle inbound customer conversations: answering from a knowledge base, looking up orders and accounts, resolving routine requests, and escalating what they cannot finish. Some draft replies for a human to send; others reply directly.
02
Genuinely good at
- High-volume, well-documented, repetitive requests
- Being available at three in the morning without degradation
- Retrieving the right article or record faster than a person can search
- Keeping a consistent tone across thousands of conversations
03
Genuinely bad at
- Long threads that contain several unrelated problems
- Situations where the customer is distressed and the policy answer is the wrong answer
- Anything where the knowledge base is stale, contradictory, or absent
- Recognising the limits of its own confidence without an explicit threshold
04
The risks that matter
- Reliability: a confident wrong answer to a customer is worse than no answer.
- Transparency: in the EU, people must be told they are dealing with an AI system.
- Data: conversation logs are personal data, frequently including special categories customers volunteer.
- Displacement: deflection targets set without a quality floor tend to produce both.
05
How to evaluate one responsibly
- Ask how escalation is triggered, and what confidence threshold sits behind it.
- Ask for quality figures alongside deflection figures. One without the other is not evidence.
- Read a real transcript sample, not a demo.
- Confirm where conversation data is stored and for how long.
Now compare what is actually declared.
The registry holds each agent’s stated facts — autonomy, oversight, compliance, residency, sustainability disclosure — with provenance on every field.
Compare customer service agents06
The evidence
- Independent research
- No commercial interest in the outcome. Method published.
- Analyst research
- Structured method, but the publisher sells advice in this market.
- Practitioner survey
- Self-reported. Good for direction, unreliable for magnitude.
- Vendor or customer material
- A claim about the publisher's own product. Not a finding.
Figures below are reproduced as published. We label who paid for the work and what the method can and cannot show. Where a number is self-reported or vendor-published, we say so next to the number rather than in a footnote.
4 of 4 sources
- 01Independent research
Access to a generative AI assistant raised issues resolved per hour by 14%, cut handling time 9%, and lifted the least experienced agents by 34%. Requests to speak to a manager fell 25%.
- Source
- National Bureau of Economic Research — Brynjolfsson, Li & Raymond, “Generative AI at Work”, NBER working paper #31161April 2023, revised 2025
- Method
- Staggered rollout across 5,179 support agents at a Fortune 500 software firm; difference-in-differences on 3m+ conversations.
- What it does not show
- One firm, one language, text chat only. Gains concentrated in less experienced staff; the most experienced saw little or no benefit.
- 02Independent research
Agents using the tool handled 13.8% more inquiries per hour (p<0.01); the lowest-performing quintile improved throughput by 35%.
- Method
- Secondary analysis of three controlled studies, including the NBER dataset above.
- What it does not show
- Re-analysis rather than new data. The headline 66% figure aggregates across unrelated task types and should not be read as a service benchmark.
- 03Vendor or customer material
Around 65% of incoming queries were resolved without human intervention in 2025, up from about 52% in 2023.
- Source
- Service-platform industry reporting — Annual state-of-service reporting2025
- Method
- Platform telemetry from customers of the reporting vendor, plus a user survey.
- What it does not show
- Resolution is defined by the platform, usually as “conversation closed without escalation”. It is not a measure of whether the customer's problem was solved.
- 04Vendor or customer material
Headline case-study figures — “32 hours to 32 minutes”, or one assistant doing the work of 700 agents.
- Source
- Vendor case studies — Assorted supplier case studies2023–2025
- Method
- Single-customer accounts, published with the supplier's approval, no baseline disclosed.
- What it does not show
- Selection bias by construction: failed deployments are not written up. Treat as an existence proof, not a rate.
We take no payment from any organisation named on this page, and no source is listed or omitted on commercial grounds. If a figure here is wrong or out of date, tell us and we will correct it with the date of the change.
Ideas in this domain
Continue