

Does AI reduce customer service costs? Yes, and the numbers are on the record. In February 2026 the largest online travel group reported customer service cost per booking down about 10% year over year while it handled roughly 10% more bookings. That is a genuine win. It is also the easy half. The harder half is everything that has to be true before an AI is allowed to say anything at all to a customer: every figure traceable to a source, every internal metric walled off, every claim carrying its own confidence. At Gimmonix we built that layer — we call it the Gimmonix Trust Architecture — for our own customer success operation, before we connected any of it to a customer touchpoint. This is what it took.
It changed one number, and that number turned out to matter more than the demos did.
On the 18 February 2026 earnings call, Booking Holdings CFO Ewout Steenbergen reported that customer service cost per booking had fallen 10% year over year while the group handled about 10% more bookings. He presented it as a rare case of generative AI showing up in the P&L rather than in a demo.
The pattern held through the year. In the first quarter of 2026 Agoda posted a double-digit year-over-year reduction in customer service cost per booking. By the second quarter, voice AI had been scaled across the majority of eligible inbound traveller calls.
For the first time, generative AI in travel stopped being a demo and became a line item.
That is why every post-sale team in this industry is now being asked the same question by their CFO. It is a fair question. It is just not the whole one.
Deflection is resolving a customer contact without a human — a bot answers, an automated flow completes the change, the ticket closes. It reduces the unit cost of a contact.
It does not reduce how many contacts arrive. And in this industry, a lot of them arrive for reasons that have nothing to do with the support stack.
Start with cancellations. Cloudbeds' 2026 State of Independent Hotels report, drawn from 90 million bookings across 180 countries, found OTA bookings cancelling at 21.8% against 10.6% for direct. Roughly one in five bookings you win comes back. Each one is a servicing event, a refund, a reconciliation entry and a possible dispute.
Then look at where the rest originates. A guest who was sold a sea view and given a courtyard. A price that moved between display and confirmation. A cancellation policy the traveller could not parse. A booking that never confirmed against the supplier. None of those are support failures. They are booking-layer failures that surface as support tickets six hours later.
PhocusWire made the same point about AI agents in July 2026: completing a real transaction depends on live inventory, accurate pricing, supplier connectivity, payment handling and the ability to move across a fragmented stack. The booking step is the hard test — for agents, and for your cost-to-serve.
You can automate the answer, or you can stop generating the question. Only one of those compounds.You can automate the answer, or you can stop generating the question. Only one of those compounds.
This is the unglamorous half of the work: room and rate data that means the same thing across every supplier, a price that is still true at the moment of commit, and terms a machine can read without guessing. Fix that layer and ticket volume falls without anyone deflecting anything.This is the unglamorous half of the work, and it is the half Gimmonix works on: room and rate data that means the same thing across every supplier, a price that is still true at the moment of commit, and terms a machine can read without guessing. Mapping.Works standardises room and hotel data across suppliers so the guest receives what was actually sold. RateFox operates at the moment of booking, catching the better rate that was already available through the customer's own supply. Fix that layer and ticket volume falls without anyone deflecting anything.
We put an AI layer into our own customer success operation at Gimmonix — account reviews, pre-meeting briefs, weekly client reporting, follow-up drafts. It has been rebuilt repeatedly on the strength of what the people using it told us. Six capabilities turned out to be structural rather than optional. Together they are what we call the Gimmonix Trust Architecture.
Any number that reaches a customer resolves to a single agreed source. Not a re-derivation, not a convenient nearby table — one source, every time.
This is more than retrieval-augmented generation (RAG), the technique of having a model pull from your own documents before it answers. RAG governs where the model looks. Canonical sourcing governs which answer wins when two sources disagree — and in any real business, eventually they will.
We run an automated reconciliation that reads the AI's own generated narrative and flags any figure it cannot trace back to a canonical source.
This is the check almost nobody builds, and it is the one we would build first if we started over. Dashboards get reviewed by people who know what the numbers should look like. Sentences get skimmed.
A number that exists only as generated prose has nothing auditing it. Dashboards get reviewed. Sentences get skimmed.
Every serious AI programme now says it keeps a human in the loop. Far fewer can say what happens to the human's correction after they make it.
Ours are generated as drafts two business days ahead and posted to the team. A reply overrides the generated content outright. The correction is then stored and treated as authoritative on every subsequent run.
That last clause is the whole thing. If a colleague has to make the same correction twice, they stop correcting. An AI system that forgets feedback trains its own users into silence.
Entire categories of internal metric are blocked from customer-facing output at the system level. Where an internal figure has to sit alongside a client-safe one, the client-facing report carries it in an explicitly badged, collapsed internal-only block, so there is no ambiguity about what may be forwarded.
Asking a model nicely not to mention something is not a control. The first time an internal metric lands in a customer's inbox, the programme is over.
Meeting outcomes are scored against a fixed rubric with fixed weights, not a model's impression of how things went. The same inputs produce the same score in March and in August, which is the only way a trend line means anything.
Follow-up emails are drafted in the actual sender's voice, modelled on their own previous correspondence. They are drafted. A human reads them and a human sends them. Nothing reaches a customer unattended.Every output declares its own state: fully sourced, partial, or operating on fallback. Each carries a provenance line naming what it was built from.
We also run language gates. The system cannot describe an account as a renewal risk unless an actual renewal is open. Loaded commercial language has to be earned by a fact in the record, not inferred from tone.
Meeting outcomes are scored against a fixed rubric with fixed weights, not a model's impression of how things went. The same inputs produce the same score in March and in August, which is the only way a trend line means anything.
Follow-up emails are drafted in the actual sender's voice, modelled on their own previous correspondence. They are drafted. A human reads them and a human sends them. Nothing reaches a customer unattended.
Six checks. Run them this week.
For the other side of this argument — why most of the AI panic aimed at OTAs does not survive contact with a real distribution stack — see Will Agentic AI Replace OTAs? The 2026 Reality Check.



.avif)
.avif)
.avif)
.avif)
%20(1).avif)
Have questions or want to partner? Fill out the form, and our team will get back to you promptly.