Three translucent speech bubbles suspended by fine threads above a large circular metal plate engraved with measurement markings, with a small seal at its centreThree translucent speech bubbles suspended by fine threads above a large circular metal plate engraved with measurement markings, with a small seal at its centre
Three translucent speech bubbles suspended by fine threads above a large circular metal plate engraved with measurement markings, with a small seal at its centre

What It Takes to Let AI Near Your Customers

8
5 minutes
29.08.2026

Does AI reduce customer service costs? Yes, and the numbers are on the record. In February 2026 the largest online travel group reported customer service cost per booking down about 10% year over year while it handled roughly 10% more bookings. That is a genuine win. It is also the easy half. The harder half is everything that has to be true before an AI is allowed to say anything at all to a customer: every figure traceable to a source, every internal metric walled off, every claim carrying its own confidence. At Gimmonix we built that layer — we call it the Gimmonix Trust Architecture — for our own customer success operation, before we connected any of it to a customer touchpoint. This is what it took.

TL;DR

  • The P&L proof exists. Large travel platforms showed in 2026 that generative AI moves a real line — customer service cost per booking, down roughly 10% on rising volume.
  • Deflection is the visible win. It lowers the cost of a ticket. It does not lower the number of tickets.
  • Ticket volume is created at commit. Wrong room delivered, stale price at confirmation, ambiguous cancellation terms, failed supplier confirmation — these arrive as support contacts hours later.
  • Six capabilities are non-negotiable — together, the Gimmonix Trust Architecture. Canonical sourcing, prose-level reconciliation, permanent human corrections, an architectural internal/client boundary, published confidence states, and deterministic scoring with a human on send.
  • There is a six-point audit at the end. It takes about an afternoon.

What Did AI Actually Change in OTA Customer Service in 2026?

It changed one number, and that number turned out to matter more than the demos did.

On the 18 February 2026 earnings call, Booking Holdings CFO Ewout Steenbergen reported that customer service cost per booking had fallen 10% year over year while the group handled about 10% more bookings. He presented it as a rare case of generative AI showing up in the P&L rather than in a demo.

The pattern held through the year. In the first quarter of 2026 Agoda posted a double-digit year-over-year reduction in customer service cost per booking. By the second quarter, voice AI had been scaled across the majority of eligible inbound traveller calls.

For the first time, generative AI in travel stopped being a demo and became a line item.

That is why every post-sale team in this industry is now being asked the same question by their CFO. It is a fair question. It is just not the whole one.

Why Is Deflection the Easy Half of AI Customer Service ROI?

Deflection is resolving a customer contact without a human — a bot answers, an automated flow completes the change, the ticket closes. It reduces the unit cost of a contact.

It does not reduce how many contacts arrive. And in this industry, a lot of them arrive for reasons that have nothing to do with the support stack.

Start with cancellations. Cloudbeds' 2026 State of Independent Hotels report, drawn from 90 million bookings across 180 countries, found OTA bookings cancelling at 21.8% against 10.6% for direct. Roughly one in five bookings you win comes back. Each one is a servicing event, a refund, a reconciliation entry and a possible dispute.

Then look at where the rest originates. A guest who was sold a sea view and given a courtyard. A price that moved between display and confirmation. A cancellation policy the traveller could not parse. A booking that never confirmed against the supplier. None of those are support failures. They are booking-layer failures that surface as support tickets six hours later.

PhocusWire made the same point about AI agents in July 2026: completing a real transaction depends on live inventory, accurate pricing, supplier connectivity, payment handling and the ability to move across a fragmented stack. The booking step is the hard test — for agents, and for your cost-to-serve.

You can automate the answer, or you can stop generating the question. Only one of those compounds.You can automate the answer, or you can stop generating the question. Only one of those compounds.

This is the unglamorous half of the work: room and rate data that means the same thing across every supplier, a price that is still true at the moment of commit, and terms a machine can read without guessing. Fix that layer and ticket volume falls without anyone deflecting anything.This is the unglamorous half of the work, and it is the half Gimmonix works on: room and rate data that means the same thing across every supplier, a price that is still true at the moment of commit, and terms a machine can read without guessing. Mapping.Works standardises room and hotel data across suppliers so the guest receives what was actually sold. RateFox operates at the moment of booking, catching the better rate that was already available through the customer's own supply. Fix that layer and ticket volume falls without anyone deflecting anything.

What Does It Take to Let AI Near a Customer?

We put an AI layer into our own customer success operation at Gimmonix — account reviews, pre-meeting briefs, weekly client reporting, follow-up drafts. It has been rebuilt repeatedly on the strength of what the people using it told us. Six capabilities turned out to be structural rather than optional. Together they are what we call the Gimmonix Trust Architecture.

1. Ground every figure in one canonical source

Any number that reaches a customer resolves to a single agreed source. Not a re-derivation, not a convenient nearby table — one source, every time.

This is more than retrieval-augmented generation (RAG), the technique of having a model pull from your own documents before it answers. RAG governs where the model looks. Canonical sourcing governs which answer wins when two sources disagree — and in any real business, eventually they will.

2. Audit the prose, not just the dashboard

We run an automated reconciliation that reads the AI's own generated narrative and flags any figure it cannot trace back to a canonical source.

This is the check almost nobody builds, and it is the one we would build first if we started over. Dashboards get reviewed by people who know what the numbers should look like. Sentences get skimmed.

A number that exists only as generated prose has nothing auditing it. Dashboards get reviewed. Sentences get skimmed.

3. Make human corrections permanent

Every serious AI programme now says it keeps a human in the loop. Far fewer can say what happens to the human's correction after they make it.

Ours are generated as drafts two business days ahead and posted to the team. A reply overrides the generated content outright. The correction is then stored and treated as authoritative on every subsequent run.

That last clause is the whole thing. If a colleague has to make the same correction twice, they stop correcting. An AI system that forgets feedback trains its own users into silence.

4. Put the internal/client boundary in the architecture, not the prompt

Entire categories of internal metric are blocked from customer-facing output at the system level. Where an internal figure has to sit alongside a client-safe one, the client-facing report carries it in an explicitly badged, collapsed internal-only block, so there is no ambiguity about what may be forwarded.

Asking a model nicely not to mention something is not a control. The first time an internal metric lands in a customer's inbox, the programme is over.

5. Publish confidence, not just answers

Meeting outcomes are scored against a fixed rubric with fixed weights, not a model's impression of how things went. The same inputs produce the same score in March and in August, which is the only way a trend line means anything.

Follow-up emails are drafted in the actual sender's voice, modelled on their own previous correspondence. They are drafted. A human reads them and a human sends them. Nothing reaches a customer unattended.Every output declares its own state: fully sourced, partial, or operating on fallback. Each carries a provenance line naming what it was built from.

We also run language gates. The system cannot describe an account as a renewal risk unless an actual renewal is open. Loaded commercial language has to be earned by a fact in the record, not inferred from tone.

6. Deterministic scoring, human on send

Meeting outcomes are scored against a fixed rubric with fixed weights, not a model's impression of how things went. The same inputs produce the same score in March and in August, which is the only way a trend line means anything.

Follow-up emails are drafted in the actual sender's voice, modelled on their own previous correspondence. They are drafted. A human reads them and a human sends them. Nothing reaches a customer unattended.

What Is the Difference Between Deflection-First and Trust-First AI?

Deflection-first vs trust-first AI in customer service
Dimension Deflection-first Trust-first
Optimises for Cost per contact Contacts that never happen, and output people will act on
Wins early on Handle time, containment rate, headcount Accuracy, forwardability, adoption by your own team
Typical failure Volume stays flat; customers escalate around the bot Slower to show a P&L number
Breaks when An automated reply is confidently wrong Nothing much — but it takes longer to build
You know you have it if Your metric is containment rate You can trace any customer-facing number to its source in under a minute

How Do You Tell Whether Your AI Customer Success Layer Is Trustworthy?

Six checks. Run them this week.

  1. Pick any number your AI put in front of a customer in the last month and trace it to a single canonical source. Time yourself. Over a minute is a finding.
  2. Find a correction a colleague made to AI output six weeks ago. Confirm the system still applies it. If it does not, your feedback loop is decorative.
  3. List every internal-only metric in your business. Check whether the block on each one is enforced in code or merely requested in a prompt.
  4. Open the last three customer-facing outputs. Can a reader tell which parts were fully sourced and which were partial? If not, everything is being read at the same confidence, including the weak parts.
  5. Take last month's top five support contact reasons and mark each one as a support failure or a booking-layer failure. Most teams are surprised by the split.
  6. Confirm nothing reaches a customer unattended. Then confirm it by checking the send logs rather than the documentation

Key takeaways

  • Cost per booking fell about 10% at the largest travel platform in 2026 while volume rose — generative AI in customer service is real and measurable.
  • Deflection lowers the cost of a ticket. It does not lower ticket volume, and volume is where the compounding is.
  • A large share of post-booking contact originates at commit: wrong room, stale price, unreadable terms, failed confirmation.
  • Ground every figure in one canonical source, then audit the generated prose for figures it cannot ground.
  • Corrections that do not persist train your team to stop correcting.
  • The internal/client boundary belongs in the architecture. A prompt is not a control.
  • Publish confidence states. Score deterministically. Keep a human on send.Cost per booking fell about 10% at the largest travel platform in 2026 while volume rose — generative AI in customer service is real and measurable.
  • Deflection lowers the cost of a ticket. It does not lower ticket volume, and volume is where the compounding is.
  • A large share of post-booking contact originates at commit: wrong room, stale price, unreadable terms, failed confirmation.
  • Ground every figure in one canonical source, then audit the generated prose for figures it cannot ground.
  • Corrections that do not persist train your team to stop correcting.
  • The internal/client boundary belongs in the architecture. A prompt is not a control.
  • Publish confidence states. Score deterministically. Keep a human on send.
  • Together these six make up the Gimmonix Trust Architecture — the standard anything customer-facing has to clear before it ships.

FAQ

Does AI actually reduce customer service costs in travel?

Yes, and it is on the record. Booking Holdings reported customer service cost per booking down roughly 10% year over year in its February 2026 results while handling about 10% more bookings, and Agoda posted a double-digit reduction in the first quarter of 2026. The gains come from automating high-volume, low-complexity contacts.

What is the difference between deflection-first and trust-first AI in customer service?

Deflection-first optimises for cost per contact and wins early on handle time and containment rate. Trust-first optimises for output a customer can act on — every figure traceable, no internal metric leaking, nothing sent unattended. Deflection-first breaks the moment an automated reply is confidently wrong. Trust-first takes longer to show a P&L number, then compounds.

What causes high post-booking support ticket volume for OTAs?

Most of it originates at the moment of commit rather than in the inbox. Cloudbeds' 2026 report, covering 90 million bookings across 180 countries, found OTA bookings cancelling at 21.8% against 10.6% for direct — and every cancellation is a servicing event. Beyond cancellations, the recurring causes are a wrong room delivered, a stale price at confirmation, cancellation terms the traveller cannot parse, and bookings that never confirmed against the supplier. Each surfaces as a support ticket hours later.

What should AI never do in a customer success workflow?

Three things: state a figure it cannot trace to a canonical source, surface an internal-only metric in customer-facing output, and send anything to a customer without a human reviewing it. Each of these should be enforced in the system rather than requested in a prompt.

How do I reduce support volume rather than just support cost?

Work upstream of the ticket. Standardise room and rate data across suppliers so the guest receives what was sold, validate the price at the moment of commit rather than at display, and make cancellation terms machine-readable. Contacts that never happen cost nothing to deflect.

Related reading

For the other side of this argument — why most of the AI panic aimed at OTAs does not survive contact with a real distribution stack — see Will Agentic AI Replace OTAs? The 2026 Reality Check.

Article Author
Maryna Gaidak
Head of Marketing
Close-up of textured white frosted glass with an abstract pattern.Close-up of a white pigeon imprint on a textured white surface.White square with a subtle diamond pattern and a faint diamond shape in the center on a light background.Rounded square icon with a white flame symbol on a textured white background.Close-up of a white textured powder against a plain background.White square textured surface with subtle patterns and rounded corners.Dark green square with rounded corners featuring the brand name GIMMONIX in the center.

Become a Leader in Travel Tech with Gimmonix

Have questions or want to partner? Fill out the form, and our team will get back to you promptly.

By submitting, you agree to the Privacy Policy.
Thank You!
Your message has been sent. Our team will get back to you shortly.
Oops! Something went wrong while submitting the form.