← Blog
Also in:
Epilot's 87% Faster Email Summaries: Time Saved Is Not ROI
2026-09-08

Epilot's 87% Faster Email Summaries: Time Saved Is Not ROI

LEMAIT decision: PILOT. Epilot's published case supports testing email-thread summarisation in a bounded workflow. It does not support a purchase decision based on ROI because the absolute before-and-after time, total cost, failure methodology, and cash effect are not public.

What the public evidence supports

An AWS customer story reports that Cologne-based Epilot launched a minimum viable product in two months and now generates 55,000 email summaries per month. It reports an 87% reduction in handling time and says about 80% of users find the feature simplifies their work. Before production, product and engineering teams used human evaluation to compare prompts and models across different email threads.

The reported architecture sends a summarisation request through an API to Amazon SQS, invokes AWS Lambda, and orchestrates the task in Amazon Bedrock. A related suggested-actions feature can extract changes such as a customer address and call Epilot's Entity API, while retaining human verification before the action.

These are E2 claims in LEMAIT's evidence scale: a named supplier-hosted customer story, useful for a pilot hypothesis but not independently verified. The source describes the failure rate as negligible without publishing a threshold, denominator, severity distribution, or audit method.

The business case

The value tree is shorter reading time → faster recognition of required action → lower handling effort and fewer transcription errors. Released time becomes financial value only if it removes cost, avoids hiring, increases measurable throughput, or is redeployed to contribution-producing work.

The non-AI comparators are structured email templates, deterministic thread extraction, and rules-based field parsing. They are less flexible but may be cheaper and more predictable for stable message types. AI earns its place only when it improves risk-adjusted cost and action recall on real mail, not merely summary fluency.

  • Downside: a concise but wrong summary hides a deadline, obligation, complaint, or safety issue and increases rework.
  • Base case: verified time saved after review and correction, minus model, queue, hosting, evaluation, maintenance, security, and compliance costs.
  • Upside: measurable throughput or avoided hiring with no increase in missed actions, complaints, or data errors.

The 30-day pilot

Start with one mailbox class and read-only summaries. Use a held-out evaluation set, named reviewers, and no automatic changes to customer data. Log the source-message reference, model and prompt version, output, reviewer decision, correction, and processing cost. Define go/no-go thresholds for factual accuracy, action recall, severe omissions, review time, and cost per accepted summary before launch.

Only after the read-only pilot passes should write capability be tested separately. Suggested actions need field-level validation, least-privilege permissions, an explicit confirmation screen, idempotency, rollback, and an audit trail tying each change to its source message and human approver.

Failure modes that matter

  • Incorrect summary, missed action, wrong customer or entity, and loss of qualification or negation in a long thread.
  • Prompt injection inside inbound email, malicious attachments, personal-data exposure, and stale context.
  • Automation bias: the reviewer clicks approve because the generated output looks routine, not because it was checked.

AI Act and GDPR

Determine provider and deployer roles for the actual configuration. Document purpose, prohibited uses, staff AI literacy, human oversight, logging, and incident handling. Under GDPR, verify data categories, legal basis, processor terms, subprocessors, access, retention, deletion, and international-transfer conditions. EU-region processing and Bedrock's stated data controls are useful inputs, not a complete compliance conclusion. This is operational analysis, not legal advice.

What remains unknown

The public sources do not disclose the absolute minutes saved, sample design, error distribution, correction time, customer-level economics, total run cost, realised capacity change, or independent audit. The 87% figure is a promising operational KPI, not proven cash ROI.

Sources and disclosure

Primary case source: AWS customer story, https://aws.amazon.com/solutions/case-studies/epilot-genai-case-study/ . Company context: Epilot, https://www.epilot.cloud/en . The first source is a vendor-hosted customer story, not an independent study. This review was prepared with AI assistance and editorially checked by LEMAIT against the cited public sources on 8 September 2026.

Have a task like this in mind?

Try the free estimate on the homepage — no account needed.

Get a price in seconds →