Unintended AI: What Your Customer Facing AI Actually Does at Scale
Unintended AI
Unintended AI is the gap between what an organisation intends its customer facing AI to do and what actually happens when that AI interacts with real customers at scale. It is neither misuse nor malfunction. It is the set of outcomes that nobody designed, nobody approved, and in most cases nobody has examined. Isara closes that gap by reading what actually happened in customer conversations rather than what the deployment plan said should happen.
Key points in this article:
- Unintended AI describes outcomes that emerge in live customer conversations, not defects in a model.
- Governance describes intent. Technical monitoring confirms function. Neither reports outcomes.
- Recent research indicates that organisations with the most mature guardrails report the highest rollback rates, which is a visibility finding rather than a technology finding.
- Regulation now asks for evidence of outcomes, not evidence of intent.
- The intention gap can be assessed on four questions: divergence, detection, direction and durability.
What is Unintended AI?
Unintended AI is the difference between intended AI behaviour and actual AI outcomes in live customer interactions.
The concept has three parts. What the organisation intended its AI to do. What actually happened when that AI met real customers. And the space in between, which is where Unintended AI sits.
Examples of Unintended AI in a customer service context include the following:
- Issuing refunds that policy did not require.
- Failing to recognise a complaint as a complaint.
- Giving inconsistent information to customers asking the same question.
- Processing a customer showing signs of vulnerability rather than escalating them.
- Creating a repeat contact from a conversation that closed successfully.
- Discouraging a customer who was ready to buy.
- Missing a buying signal that recurred across hundreds of conversations.
Any single instance may be trivial. At AI scale, the aggregate may not be. The operative question is not whether these things happen. It is whether the organisation would know.
Why does Unintended AI matter now?
It matters now because three pressures arrived in the same twelve months: measurable rollback rates, visible customer dissatisfaction, and enforceable regulation.
Rollback rates rise with governance maturity. Sinch surveyed more than 2,500 AI decision makers across regions and industries for its AI Production Paradox study, published in May 2026. Seventy four per cent of enterprises that deployed AI customer communications agents later rolled them back or shut them down, and that figure rose to 81 per cent among organisations with what the study describes as fully mature guardrails. Sinch's chief product officer put the implication plainly: if governance were the fix, the most mature teams would roll back less rather than more, and higher rollback rates reflect better monitoring rather than worse performance.
Read that as a statement about visibility. Mature teams are not deploying worse AI. They can see what it did.
Effort is concentrated before deployment, not after. The same study found that 84 per cent of AI engineering teams spend at least half their time on safety infrastructure, and that most firms rank spending on AI trust, security and compliance above spending on AI development itself. Almost all of that investment constrains behaviour up front. Very little examines behaviour afterwards.
Analysts already name unintended results as a cause of abandonment. Brian Weber, VP analyst in Gartner's Customer Service and Support practice, told The Register in May 2026 that vendor evaluations show a fully agentless contact centre is neither technically feasible nor operationally desirable, and that unexpected costs and unintended results were contributing to firms abandoning their plans.
Customers experience it specifically in service. The Qualtrics 2026 Customer Experience Trends Report, reported in April 2026, found that nearly one in five consumers who used AI for customer service saw no benefit at all, a failure rate close to four times higher than for AI use generally. Consumers ranked customer service among the worst AI applications for convenience, time savings and usefulness.
Which AI failures are invisible on a dashboard?
The dangerous failures are the ones that register as success in reporting.
A study of live AI customer experience deployments published in 2026 found that the most common pattern was answer only fallback, in which the agent accurately explains a task it cannot actually perform. It appeared in ten of the fifteen deployments observed. The rarest but most dangerous pattern was unsafe confidence.
The reporting consequence is the important part. The conversation registers as contained. The ticket arrives later through a different channel. The outcome was poor, the metric was healthy, and the two were never connected.
Diagnosis is the recurring weak point. Post mortems on underperforming deployments tend to surface the same causes: siloed data, poorly scoped automation, weak handoff design, metrics that measure the wrong things, and deflection pressure applied too early in a rollout. None of those are AI problems, yet all of them produce outcomes that look like AI problems from the outside. Wrong diagnosis produces the wrong fix, and the wrong fix produces the same failure faster.
This is why Isara reads 100 per cent of customer conversations rather than a quality assurance sample. Failure modes that pass as success in aggregate reporting are not detectable by sampling for them.
What do regulators now require?
Regulators are moving from asking what AI was designed to do towards asking what happened to customers.
European Union. From 2 August 2026 the European Commission's AI Office, alongside national authorities, began enforcing the AI Act, and new transparency rules require certain AI systems to inform users that they are interacting with AI and that content has been generated or altered by it. Article 50 binds both providers and deployers, applies globally to anyone whose AI outputs are used within the European Union, and carries fines of up to 15 million euros or 3 per cent of worldwide annual turnover, whichever is higher. Scope explicitly includes chatbots, AI voice assistants and agentic systems that autonomously contact individuals. Note for planning: the transparency duties took effect on schedule, and only the high risk tier was deferred by the Digital Omnibus.
United Kingdom. The approach is outcomes based rather than technology based, which is harder to satisfy. The FCA's 2026 regulatory priorities keep Consumer Duty as the foundation of supervision, expect boards to monitor outcomes rather than processes, and require outcome monitoring frameworks that are robust and evidence led, including granular management information on outcomes delivered for vulnerable customers. They also call for clear accountability, governance and testing of AI models and data use, with greater scrutiny where firms deploy third party AI tools.
Evidence of outcomes, for every customer, including the vulnerable ones. That is a conversation level obligation.
How do you measure the AI intention gap?
Most organisations can report a containment rate. Very few can report an intention gap.
A containment rate confirms that a conversation ended without a human. It says nothing about whether the customer got what they needed, whether a complaint was recognised, whether a refund was warranted, or whether a vulnerable customer was escalated.
Unintended AI can be assessed on four questions, applied to conversations rather than to policy documents.
- Divergence. What proportion of AI handled conversations produced an outcome outside stated policy or intent? Not clumsy phrasing, but outcomes the business would not have authorised.
- Detection. Of those divergent outcomes, what proportion were flagged by an existing system before a human noticed by accident? This is usually the number that surprises people.
- Direction. Do divergences cluster in customer harm, avoidable cost, or missed growth? The same blindness produces all three, which is why Isara organises findings as Protect, Save and Grow rather than treating risk and revenue as separate disciplines.
- Durability. Does the divergence rate fall after intervention, or does it reappear in a different conversation type? Suppressed failure modes tend to move rather than disappear.
Two consequences follow from running this honestly.
Detection will be the weak score. On the evidence above, the failure modes that look like success on a dashboard are the common ones. An organisation with strong governance and no conversation level examination will score well on controls and badly on detection. That is precisely the pattern visible in the Sinch rollback data.
The second consequence is a prediction rather than a finding. Within eighteen months, evidencing what customer facing AI actually did will be a routine requirement in regulated sectors rather than a competitive differentiator. The Article 50 transparency duties already in force, combined with the FCA's insistence on evidence led outcome monitoring, point one way. Organisations that can produce that evidence from their own conversation data will treat it as reporting. Organisations that cannot will treat it as an incident.
One caution against overclaiming. The scale of Unintended AI is not independently quantified in any published body of work, because the category is new. The adjacent evidence is strong and consistent, but any precise global figure offered today is an estimate. Isara's position is that the number should come from conversations rather than surveys, and that it should be measured per organisation before it is measured per market.
Frequently asked questions about Unintended AI
What is Unintended AI in simple terms?
Unintended AI is what your customer facing AI does that you did not intend, did not expect, and may not know about. As set out above, it covers unnecessary refunds, unrecognised complaints, mishandled vulnerable customers, repeat contacts from apparently contained conversations, and missed buying signals. Isara categorises these as Protect, Save and Grow so that one analysis serves risk, cost and revenue owners.
Is Unintended AI the same as AI hallucination?
No. Hallucination is one contributor among several. Most of the failure modes described in this article sit downstream of the model itself, in execution, context transfer, escalation and policy.
Our AI vendor already provides analytics. Why would we need Isara?
Vendor analytics report how the vendor's system performed against its own design, which is the intent side of the gap described above. Isara reads the conversations independently of the system that produced them, which is what makes the result usable as evidence rather than as self assessment. It also produces one view where AI and human agents handle the same customers.
We run several helpdesks across markets. Does that break the analysis?
That is the normal case. Isara reads from Zendesk, Freshdesk, Intercom, HubSpot, Front and Gorgias, so the intention gap can be measured across the whole estate rather than in whichever tool has the best native reporting.
How does this help with the regulatory position described above?
Both the Article 50 transparency duties and the FCA's outcome monitoring expectations require evidence about what happened to customers. Isara covers 100 per cent of conversations rather than a sample, which is the difference between reporting outcomes and extrapolating from a subset. Coverage does not by itself constitute compliance, and this article is not legal advice on your obligations.
Can Isara run the four question assessment in this article?
Divergence and direction map onto Isara's current conversation analysis. Detection and durability require comparison against what your existing tooling flagged, which Isara works through with customers directly. The wider Unintended AI Index referenced in our category thinking is a planned research initiative rather than a shipped feature, and the methodology will be published before any numbers are.