About

How AI Changes the Issue-to-Resolution Value Stream

How customer service shifts from an agent handling every ticket to AI handling the routine ones with a human on call, plus where that quietly goes wrong.
Sebastian Bitter
Sebastian Bitter
4.8.2026
Issue-to-Resolution carries a customer problem from the first message to a solved case, which makes it the stream where AI in service is most visible to customers. Two real cases carry the levels: Kraken Magic Ink at assistive, where about a third of AI drafts needed zero to minimal changes with an agent sending every one, then Klarna at agentic, 2.3 million chats in the first month at the equivalent of about 700 full-time agents. The same company then pushed to full automation, removed the human escalation and reported a quality drop a year later. Resolution rates are easy to inflate, because a case can count as solved once the customer stops replying.

The stream today and where AI takes it

Issue-to-Resolution is the value stream that carries a customer problem from the first contact to a solved case. It starts when someone reports an issue and ends when the issue is resolved and the customer is told. Every company with customers runs this stream, which is where a lot of today's AI attention in service lands, because the volume is high and the work repeats.
In its classical form the work runs as six steps that all sit with a person. The customer gets in touch and an agent logs the ticket by hand. The agent sorts it by topic and urgency. The agent looks up the account history and the knowledge base. The agent writes and sends the answer. Hard cases go to a second-level team. Finally the agent closes the case and asks for a rating.
The pain is familiar to anyone who has waited on hold, because a human does every step while the customer waits and the same routine questions get answered again and again. Most of the delay sits in the queue and the hand-offs rather than in the answer itself. Adding a faster search tool for agents does not fix that, because the bottleneck is how much a human team can handle at once.
What AI changes in this stream is which cases reach a person at all, because intake, sorting and the standard answers move to the AI. The human shifts to the cases that actually need judgment, meaning the sensitive, the complex and the high-value ones. The step where a case is handed to a human, the escalation, stops being an afterthought and becomes a designed part of the flow. The rest of this post walks that path one level at a time and shows each change with a real example.
AI-WS Bp2 Fig1 Issue-to-Resolution EN
Figure 1: Issue-to-Resolution across the four levels. The same stream at each level. Intake, triage and standard resolution move to the AI while the human moves to the exceptions, with escalation as a designed gate.

Level by level, shown by real cases

Each level below comes with a real example, its headline number and the point where that number stops holding. The classical level is the starting point above, with no AI in the flow, so the walk begins at assistive.

Assistive: AI drafts the reply, the agent sends it

At the assistive level the six steps stay as they are while AI helps inside each one. It pre-fills the ticket, suggests a category, surfaces the account history and the right help article, then drafts the reply. The agent reviews every draft and sends it. The customer still talks to a human.
The example is Magic Ink, the support tool built by Kraken Technologies and used at the energy retailer Octopus Energy. It drafts replies from the customer's account history. About a third of those drafts needed zero to minimal changes before being sent, in the words of the case study. The agent stays in charge and sends every reply.
The help is real, with two weak spots that the headline number does not carry. The case study names hallucination as a barrier and says team members are trained to review everything Magic Ink writes with a fine-tooth comb, which is exactly the check that erodes under time pressure. And every number attached to this case comes from Kraken. The case study puts customer satisfaction on the AI-assisted mail at around 70 percent and calls it higher than the rest. Kraken's own 2025 accounts report savings of up to 33 percent in cost to serve, with Magic Ink writing up to 60 percent of the emails at E.ON Next and Good Energy. Those last two figures carry the words up to, which names a ceiling and not an average.
One case is not a measurement, which is why this level also rests on an independent study. Economists at Stanford and MIT followed the staggered rollout of a generative AI assistant across 5,172 customer support agents and published the result in the Quarterly Journal of Economics. Issues resolved per hour rose by 15 percent on average. The average hides the interesting part. Less experienced and lower-skilled agents improved both their speed and the quality of their replies. The most experienced gained little speed and lost a little quality. That fits the authors' reading that the assistant passes on what the best agents already know. Customers became more polite and asked for a manager less often. What was measured is the assistive level and nothing beyond it, since the tool assisted agents rather than replacing them.
AI-WS Bp2 Fig2 Kraken case EN
Figure 2: the Kraken Magic Ink case. The customer message, the AI draft, the agent's decision and the sent reply. About a third of drafts went unchanged, on routine mail, with the agent sending every one.

Agentic: AI resolves the routine cases, a human is one step away

At the agentic level the stream is rebuilt around the AI and the steps merge. Intake and sorting happen automatically. The AI then resolves the standard cases end to end and in the customer's own language, so refunds, returns, invoice questions and the like. The important addition is a designed escalation gate. Sensitive or complex cases are routed to a human, while the customer can ask for a person at any point. The human becomes the exception handler and sets the policy.
The example is Klarna, the buy-now-pay-later company, worth following across two years rather than one. In its first month Klarna reported that its OpenAI-based assistant handled 2.3 million chats, the equivalent of about 700 full-time agents, while cutting the average resolution time from 11 minutes to under 2. The assistant did the routine work while a customer who wanted a human could still reach one. These numbers come from Klarna itself rather than from an independent audit. At that point the agentic level worked as described.
A high resolution rate is not the same as a satisfied customer, which one widely quoted number makes concrete. The support assistant Intercom Fin is advertised at a 76 percent resolution rate, a figure that rests on a definition in which a conversation counts as resolved once no further help is requested after the assistant's last answer. No robust independent measurement of that rate could be found, so the number describes the vendor's counting rule as much as the outcome. What a rate does not say is whether the customer came back satisfied.
AI-WS Bp2 Fig3 Klarna case EN
Figure 3: the Klarna arc, from agentic success to over-automation. One company across two years. In 2024 the AI resolves the standard case with a human one step away. In 2025 that escalation path is removed, quality drops and Klarna brings people back.

Autonomous: full automation with no way out, and why it backfired

At the autonomous level the human escalation path is removed, so the AI handles every case while people only set the policy and watch the aggregate numbers. On paper this is the cheapest setup. In customer service it is also the riskiest, which the same Klarna case shows.
A year after the success Klarna pushed hard toward full automation and took the human escalation away. The company then said publicly that service quality had dropped and started bringing people back, on a flexible, on-demand basis. The missing escalation path is the obvious cause, because it is the thing that changed between the two chapters. That is a reading and not a finding of the reporting. When every case is forced through the AI and a frustrated customer cannot reach a person, the cost shows up in trust rather than in the support budget. The 2024 numbers were Klarna's own, while the 2025 walk-back is reported by independent press and a CEO admission against interest.
The autonomous level in this stream is therefore a warning rather than a stage a company should aim for. The same arc that shows the agentic win also shows the cost of pushing past it, so both chapters sit in the one view above (Figure 3).

What is realistic

Put the levels together and the pattern is the same in both directions, which is unusual for a single case. Where AI resolves the standard cases and a human escalation stays available, the documented numbers hold. Where that escalation was removed, the same company reported the quality drop itself. The agentic Klarna pattern before the over-automation is the setting with evidence behind it, with the escalation gate as the part that carries it.
Getting there depends on conditions that have little to do with the model and everything to do with the ground it stands on. It takes a clear playbook for the cases the AI is allowed to handle, a connection into the systems so it can actually issue a refund with the narrowest access that still works, explicit rules for when a case goes to a human and a structured knowledge base for it to draw on. The cleaner those are, the more the AI can safely take.
The second cost is specific to this stream, because the material handed to the AI is customers' personal data, which in a support conversation can include sensitive details. Where and how that data is processed is a data-protection question from the start rather than an afterthought, especially under European rules. That choice belongs in the plan next to the business case.
What gets measured decides whether any of this shows up as value, with a resolution rate among the easiest numbers to inflate. A case marked solved because the customer gave up is not a win. The numbers that carry meaning are whether the customer came back satisfied and whether the same issue returns, not how many tickets the AI closed. That is the trap behind the Intercom figure, which a dashboard that rewards volume walks straight into.
There is a quieter effect worth naming, where the evidence points one way at one level and is missing at the other. The worry is that when the AI takes all the easy cases, the people are left with only the hard and frustrated ones, plus that the easy cases were where new agents used to learn the ropes. The Quarterly Journal of Economics study found the opposite at the assistive level. Novices got good faster, customers were more civil and fewer agents quit. That study measured a setting in which every case still reached a human. At the agentic level the easy cases never arrive there at all, which is a different situation for which nobody has published a measurement. The worry is reasonable and it stays a worry.
What the evidence supports in this stream is agentic with a human always one step away. The Klarna walk-back is the reason full autonomy has no case behind it here. The route to a person is close to a hard requirement in practice. The law demands something else. The EU AI Act has applied since August 2026 and requires that people are told they are talking to an AI. On offering a route to a person it says nothing. A right to human intervention does exist in data protection law, but only for decisions with legal or similarly significant effect. A question about a delivery date is not one of those. Klarna learned the requirement without any statute forcing it to.
AI-WS Bp2 Fig4 Gain and price EN
Figure 4: gain and price across the four maturity levels. Each level brings a gain and asks a price, drawn from the cases above. Agentic is the level with evidence behind it, autonomous the warning case. This is a concept rather than aggregated metrics. The individual figures are company-reported.

What the switch costs

This is the one stream in the series where the market has put a price on the unit of work. Intercom charges 0.99 dollars per outcome for its Fin agent, with a minimum of fifty outcomes a month, plus 29 dollars per helpdesk seat. That makes the arithmetic unusually easy and unusually revealing.
Read the definition that sits next to the price, because it is the same definition the resolution rate uses. A billable outcome is a resolution, a procedure handoff or a disqualification. And a resolution is defined as no further help being requested after the agent's last answer. A customer who was helped and a customer who gave up look identical in that definition. Both are billed. The counting rule is the pricing rule.
A competitor has since built that same objection into its own price model, which is worth reading closely. Since May 2026 Zendesk sorts automated resolutions into three tiers and bills only the top one. If the AI only prepared the ground and a human finished the case, nothing is charged. If the customer stopped replying, a second language model reads the conversation and judges whether the request was actually settled. Only the conversations that pass that check are billed. Silence on its own is no longer enough. Zendesk's documentation gives no figure for how often that check gets it right, so what changed here is the counting rule rather than the problem.
The price also shows where the model of buying by the case stops working, with Klarna's own volume making the point. At 0.99 dollars, the 2.3 million conversations Klarna reported in a single month would come to roughly 2.3 million dollars for that month. That is arithmetic on a published price rather than anything Klarna paid, since it built its own assistant. It does explain why a company at that volume builds rather than buys. Per-case pricing turns into a reason to own the thing.
Klarna is also the only company in this post whose service costs can be read in an audited filing. Since its New York listing in September 2025 it reports every quarter. In the second quarter of 2026 the line called customer service and operations came to 58 million dollars against 51 million a year earlier, so spending rose by about 14 percent after the company brought people back. Revenue grew by 27 percent over the same period, which means the same cost fell from 6.2 to 5.6 percent of revenue. That is what a hybrid model looks like on the income statement. More money for service in absolute terms, from a service function that no longer grows with the business.
Everything else on the bill is the same as in the other streams and mostly not about AI. The playbook, the system access, the escalation rules and the knowledge base all have to exist first. Klarna is the case that shows what they are worth, because the numbers held while the escalation was in place and the company reported a quality drop after it was removed. The exceptions concentrate rather than disappear, which in this stream means the people are left with the hard and frustrated cases. And the ground under all of it moves. OpenAI cut the price of a model already in service by 80 percent in July 2026, so whatever was tuned around a given model has to be checked again.

What comes next

Issue-to-Resolution shows the same pattern as the rest of the series, with a sharper edge. The routine work moves to the AI, the human moves to the cases that matter and the whole thing depends on a gate, here the escalation to a human. Push past that gate and the savings turn into a trust problem. The question this series puts to every stream is the same one. What is actually possible at each step today and who has measured it? That answer ages quickly, which is why the figures here carry a date. The next post puts the same lens on Procure-to-Pay, the stream from purchase requirement to paid invoice, the one that ends with money leaving the company.
The two headline cases here carry the story most clearly, while more sit at each level. Three further cautionary cases belong next to the Klarna walk-back. Air Canada, DPD and Cursor each bring their own evidence and their own limits. More agentic cases exist than the two named here. They are left out because they are not verified to the standard the rest of this post uses.

Sources

The figures in this post trace to these sources:
Kraken Magic Ink at Octopus Energy, about a third of drafts needing zero to minimal changes, plus satisfaction of around 70 percent on AI-assisted mail: techUK case study (company-reported)
Klarna AI assistant in its first month, 2.3M chats and the work of about 700 agents: Klarna press release, 2024 (company-reported)
Klarna bringing people back after over-automation, 2025: Customer Experience Dive, 2025 (independent press)
Fin priced at 0.99 dollars per outcome, with the billable outcome defined on the same page: fin.ai pricing (vendor's own price list, freely readable)
Intercom Fin advertised at a 76 percent resolution rate: fin.ai (vendor-reported, including the definition of a resolution)
Independent measurement at the assistive level, 5,172 support agents: Brynjolfsson, Li and Raymond, Generative AI at Work, Quarterly Journal of Economics 140(2), 2025 (peer-reviewed, freely readable as NBER working paper 31161)
Klarna customer service and operations, second quarter 2026: Klarna Group plc quarterly report to the SEC (audited filing, with the attribution to AI being the company's own)
Zendesk billing only for verified resolutions since May 2026: Zendesk help centre (vendor documentation, publicly readable rule)
What the law actually requires when a customer talks to AI: EU AI Act, Article 50 (disclosure only, applicable since August 2026)
The right to human intervention, and the narrow case it covers: GDPR, Article 22(3) (applies to decisions with legal or similarly significant effect)
A model already in service cut by 80 percent in July 2026: OpenAI price announcement (vendor's own announcement)
Magic Ink savings of up to 33 percent in cost to serve, and up to 60 percent of emails at two clients: Kraken Technologies Ltd, annual report FY25 (filed accounts, company-reported)
Air Canada held liable for its chatbot: Moffatt v. Air Canada, 2024 BCCRT 149 (tribunal decision)
DPD switching off its chatbot after it insulted a customer: The Guardian, January 2024 (independent press)
Cursor's support bot inventing a login policy, the co-founder apologising: The Register, April 2025 (independent tech press, original reporting)