A study of an AI review-summary rollout on Ctrip found that hotel ratings declined and became more varied after the feature appeared. A separate randomized experiment found that positively slanted summaries raised travellers’ expectations before their trip, indicating that a tool designed to simplify booking can also make the eventual stay feel disappointing.
The research appears in the August 2026 issue of Tourism Management. It was conducted by Yujie Wang, Ziqiong Zhang, Chuxin Wang, Zongyi Li, Rob Law and Zili Zhang.
The study does not show that all AI summaries are harmful or that the quality of the hotels deteriorated. It identifies a narrower risk: when a summary selects too many positive comments and leaves out qualifications, it can raise expectations beyond what the stay delivers.
What is an AI-generated review summary?
Hotel pages can contain hundreds or thousands of customer reviews. An AI-generated review summary uses software to condense recurring comments into a short overview, usually displayed above the individual reviews.
For a traveller, the attraction is obvious. Reading a few sentences takes less time than searching through pages of comments about cleanliness, service, noise, breakfast or location.
Ctrip, an online travel agency that connects travellers with hotels and other travel services, began introducing such summaries for selected hotels in late 2023. The feature was rolled out at different times in different cities, giving the researchers an opportunity to compare what happened before and after its arrival.
The staggered rollout provided a natural comparison
The researchers analysed review data from five cities. Hotels in places that received the summary feature earlier formed one group, while hotels still waiting for it provided a comparison group.
They used a method called difference-in-differences. In plain English, it compares the change over time in the group exposed to a new policy or feature with the change over the same period in a group that has not yet been exposed. This helps separate the effect of the feature from broader movements affecting all hotels, such as seasonal travel patterns.
After the AI summaries appeared, average post-stay ratings were lower and the ratings were more dispersed. Greater dispersion means that scores were spread further apart rather than clustering around a similar view of the hotel.
The authors interpret the combination of lower ratings and wider disagreement as reduced satisfaction with the booking decision. That wording matters. A rating records a guest’s assessment after the stay, not a direct measure of whether the hotel itself became better or worse.
Positive summaries raised expectations
To examine why ratings might change, the researchers also ran a randomized experiment. Participants shown positively biased AI summaries formed higher expectations about their forthcoming trip.
The explanation comes from expectation disconfirmation theory. People compare what they receive with what they expected. A service can therefore disappoint when it falls short of an elevated expectation, even if its underlying quality has not changed.
This is particularly relevant to hotels because much of the product cannot be checked properly before purchase. A map can show the distance to an airport, but it cannot fully convey how noisy a room will feel, whether service will be attentive or how comfortable a bed will be.
The study found that the summaries had a limited effect on expectations about objective hotel attributes. The researchers used geographical distance as one such attribute. The larger concern was the selective presentation of experience-based information, where potential guests must rely more heavily on other people’s accounts.
The business risk is an expectation gap
A booking platform may be tempted to assess an AI summary through immediate measures such as reading time, clicks and completed reservations. The study points to a second set of measures that arrives later: post-stay ratings, complaints, repeat use and trust.
The research did not measure bookings, refunds, customer retention or financial returns. Even so, it identifies a practical question for any business summarizing customer feedback. Does the summary help people make a suitable choice, or does it simply make them decide more quickly?
Those outcomes can move in different directions. A polished summary might remove friction from the booking process while also hiding the minority experiences that matter most to a particular customer. A light sleeper needs to see recurring noise complaints even if most visitors liked the hotel’s location and breakfast.
For platforms, this means summary accuracy cannot be judged only by checking whether every sentence can be traced to a real review. Selection matters as well. A digest made entirely from genuine positive comments can still give an unbalanced account if credible warnings are consistently omitted.
AI summaries may also change who writes reviews
There may be another consequence. A separate 2026 study in Decision Support Systems examined TripAdvisor’s introduction of AI-generated hotel summaries in Singapore. It estimated that the number of new user reviews fell by 27.5% after the feature was introduced.
The decline was greater among highly rated reviews and lower-tier hotels. The people who continued contributing wrote longer reviews, gave slightly lower ratings and moved from general remarks towards more specific topics.
Taken together, the two studies point to a possible feedback problem. A summary may affect what customers expect, while its presence may also alter which customers bother to write the next set of reviews. Those changed reviews then become the source material for future summaries.
Neither study followed that full loop over the long term, so it remains a risk to test rather than an established chain of events. However, it gives platforms a reason to monitor the quantity and mix of new reviews, not just how readers interact with the summary box.
The findings should not be applied to every platform
The main field evidence comes from one travel platform and five cities. A difference-in-differences design is stronger than a simple before-and-after comparison, but it still depends on the assumption that the early and later rollout groups would otherwise have followed comparable trends.
The randomized experiment provides separate evidence for the expectation mechanism, but a controlled task cannot reproduce every part of a real booking and hotel stay. The results also concern accommodation, an experience that customers can judge fully only after buying it. They may not transfer directly to standardized goods with easily checked specifications.
Nor did the researchers test every possible summary design. A balanced digest that gives negative themes suitable prominence may produce a different result from a strongly positive one.
The authors argue that travel platforms should make AI summaries more balanced and transparent. A useful next test would compare different ways of displaying positive and negative themes while keeping the original reviews easy to reach. At the moment, the evidence supports caution about overly positive compression, not the removal of AI summaries altogether.