If your team wants to use AI on a crowded inbox, start by deciding what happens after a message is sorted. Barclays and Anthropic announced an expanded partnership on October 1. In Barclays’ Global Markets business, Claude helps classify incoming emails, add information, and choose a processing route. The platform handles approximately 120,000 emails a day, according to the announcement. That number tells us the system is in daily use. It does not tell us how many client problems were resolved correctly or how much time the people receiving those requests actually saved.

The sorting number is impressive. The unanswered question is what happens next.

TL;DR

Barclays’ announcement describes two existing uses: an employee knowledge assistant and an email-routing system. Both help people get to the work sooner. Neither tells you who checks an uncertain answer or who owns an urgent request. Begin with one category of incoming work, measure the time from receipt to resolution, and expand only when you can see what went wrong.

Sorting is not serving

An email arrives from a client. It contains a request, a deadline, and a reference to an earlier conversation. Somebody has to recognize what it asks, identify missing information, send it to the right team, and make sure a person acts on it. Barclays says its system helps classify, enrich, and route those emails so operations colleagues can prioritize them. That is a meaningful piece of the job. But the point at which a message leaves the sorting system is not the point at which the client gets an answer.

A dashboard showing 120,000 processed emails a day could look excellent while requests still sit in the wrong queue. The volume comes from the vendor’s account of the deployment; it is not a published measure of accuracy, resolution time, or customer satisfaction. I would want to see how many messages were redirected by a human, how many needed follow-up because the original request was incomplete, and how long an urgent case waited before someone took ownership.

You don’t need Barclays’ volume to ask those questions. Before choosing a model, decide what the person on the receiving end needs to see.

Give the next person a usable handoff

Suppose a team handles supplier complaints. A narrow trial could have AI separate billing questions from delivery problems, pull out the order number when it is present, and flag messages that mention a safety issue. A human still decides what to tell the supplier. The trial succeeds if staff spend less time locating the right case and more time resolving it, without quietly losing the unusual complaints.

Before starting, write down the handoff in plain language. Which fields does the recipient need? Can the system say “I don’t know” when the message is ambiguous? Where does that message go? How does the recipient correct a bad category so the same mistake is noticed next time? If a high-priority email never reaches its owner, who sees the gap before the client does?

That last question deserves more attention than the demo. A tidy queue can hide an urgent message sent to the wrong desk. Give someone the job of sampling routed messages, reviewing corrections, and looking for urgent items that stalled. Keep the exceptions where a manager can see them; a count of completed classifications won’t show the missed request.

The other half of Barclays’ example

Barclays also reports that its Colleague Knowledge Assistant, live since 2025, has been adopted by more than 16,000 employees and used for over one million searches. It helps staff find information while serving more than 20 million UK retail customers. That is a different handoff: from a search result to a conversation with a customer. Staff still need to judge whether the answer fits the customer’s situation and whether the underlying information is current.

One million searches tells you people used it. It doesn’t tell you whether the customer got the right answer. Search volume might rise because the tool works, or because the first result sends staff back for another search. Check whether employees found the right policy, whether they escalated, and whether customers had to call back.

The Bank of England and Financial Conduct Authority’s 2024 survey found that 75% of responding financial firms were already using AI, with another 10% planning to use it within three years. That survey is about adoption across respondents, not the performance of Barclays’ systems. It helps explain why counting installations or users is becoming a less interesting question than understanding the work those systems actually improve.

What I would measure on Monday

Pick a single request type with a known owner. Record its current time from arrival to first human action, the number of handoffs, and the share that gets reopened or corrected. Have the AI prepare a classification and a short summary, but keep the person in charge of the response. For a few weeks, compare the new route against the old one. If the team gets to the right request sooner and the correction rate stays acceptable, expand carefully. If sorting simply moves work into a new queue nobody watches, stop calling it a time saving.

Barclays’ plan to put Claude Code in the hands of 50% of its developers by the end of 2026 is a target, not a completed result. Its search and email systems are running now. People still carry the responsibility at the other end. If the person receiving a sorted request can’t act on it sooner or with fewer mistakes, the sorting number is beside the point.

Research and structure: Mai. Editorial direction and voice: John Lipe.