← Blog
Also in:
Ryanair's 120,000 Daily AI Chats: Containment Is Not Resolution
2026-09-08

Ryanair's 120,000 Daily AI Chats: Containment Is Not Resolution

LEMAIT decision: PILOT, not copy-and-paste adoption. Ryanair's published case is credible enough to justify a bounded pilot for repetitive, high-volume customer-service questions. It does not disclose total cost, causal measurement, repeat-contact rates, complaint severity, or audited savings, so it cannot support a transferable ROI claim.

What the public case reports

An AWS customer story reports that Ryanair's assistant handles 120,000 daily chat interactions across seven languages at an 80% containment rate. It says five foundation models were evaluated against 12,000 real production questions. The selected model reportedly reduced response time from 18 seconds to 2.9 seconds and improved accuracy by 25% in that evaluation.

AWS also reports 10 million answers since launch, 94% accuracy, a 70% reduction in customer-service contacts per passenger carried, and more than 500,000 agent hours saved across the combined chat and voice transformation. These are E2 claims in LEMAIT's evidence scale: a named vendor-hosted customer account, not an independent audit. The source does not provide the underlying logs, definitions, counterfactual, cost base, or reconciliation between its different success metrics.

The business case: true resolution is the unit that matters

The stated baseline was a largely manual chat channel and a static chatbot that caused abandonment and downstream contact. The value tree is plausible: automate routine questions → reduce human contacts → avoid linear headcount growth; lower latency → reduce abandonment; multilingual reuse → avoid separate systems by language. Future ancillary sales are an option, not a realised benefit, and should not enter today's ROI.

The downside is false containment: a conversation ends in the AI channel but the customer retries by phone, email, complaint, or chargeback. The base case counts only verified avoided contacts multiplied by loaded handling cost, less inference, translation, observability, engineering, human review, and incident costs. The upside can include proven headcount avoidance or incremental margin, but only after controlled measurement.

The non-AI comparator is deterministic intent classification, approved FAQ retrieval, and human escalation. It is cheaper and more predictable for narrow intents. Generative AI earns its place only where language variation and long-tail questions deliver incremental true resolution without increasing policy errors.

The pilot scorecard

  • Eligible contacts, verified resolution, seven-day repeat contact, escalation, and abandonment.
  • CSAT, complaint severity, policy-error severity, latency, and language-level variance.
  • Loaded cost per resolved contact, including inference, translation, monitoring, engineering, review, and incidents.

Adopt only if the pilot beats both the current process and deterministic retrieval on risk-adjusted cost per resolved contact. Containment alone is not an acceptance criterion.

The implementation pattern

The reported architecture uses a unified agent, Amazon Nova 2 Lite, Amazon Translate, Amazon Connect, and a multi-dimensional, multi-model guardrail that scores every answer before delivery. It replaced an earlier architecture of 12 domain-specific agents. A reusable implementation must additionally define the approved retrieval sources, tool permissions, identity and booking-context boundaries, audit logs, retention, incident response, and a reachable human escalation queue.

The 12,000-question production set is a strong starting pattern, but it needs time-split holdouts, per-language scoring, disruption and refund red-team suites, a precise definition of 'accuracy', and continuous drift evaluation. Human agents must be trained, reachable, and authorised to correct the record rather than serving as a nominal fallback.

Failure modes that matter

  • Incorrect or obsolete baggage, refund, booking, and flight-disruption guidance.
  • Mistranslation, cross-language quality drift, booking-data leakage, and prompt injection.
  • Unsafe actions through connected systems and false containment that moves cost to another channel.

AI Act and GDPR

A customer-facing deployment should assess AI Act transparency for the interaction, staff AI literacy, provider/deployer allocation, logging, and effective human escalation. The GDPR analysis must cover purpose, legal basis, minimisation, retention, processor terms, international transfers, security, and data-subject rights. If connected tools begin making consequential decisions rather than providing service information, reclassify the use case before release. This is operational analysis, not legal advice.

What remains unknown

The public case does not disclose total cost of ownership, loaded labour baseline, the denominators and test methods behind containment and accuracy, repeat-contact rates, complaint outcomes, language-level quality, human escalation capacity, or independent verification. Those gaps are the reason for the PILOT decision rather than an ROI claim.

Sources and disclosure

Primary case source: AWS customer story, https://aws.amazon.com/solutions/case-studies/innovators/ryanair-agentic-ai/ . Company context: Ryanair corporate site, https://corporate.ryanair.com/about-us/ . The first source is a vendor-hosted customer story, not an independent study. This review was prepared with AI assistance and editorially checked by LEMAIT against the cited public sources on 8 September 2026.

Have a task like this in mind?

Try the free estimate on the homepage — no account needed.

Get a price in seconds →