Tourism Research.
2026, 18(5):
83-98.
In the context of image-text communication, traditional text-based analysis of online tourism word-of-mouth tends to encounter problems such as insufficient evidence, compressed sentiment intensity, and ambiguous attribution of emotional evaluations when identifying highly contextualized tourist experiences. Taking tourist reviews of the Nalati Tourist Scenic Area on Ctrip, Meituan, and Douyin as research samples, this study develops two comparative analytical paths, namely a “ text-only” path and a “ text-imagejoint” path, based on the complementarity between linguistic and visual coding. Sentiment identification is conducted on the textual side, while Qwen3-VL is employed on the visual side to extract cues such as landscape features, crowding degree, weather conditions, on-site atmosphere and facility status. The robustness and credibility of the multimodal evidence chain are further examined through consistency rates, divergence rates, and manual review. The results show that the text-only path remains applicable in identifying explicit attitudes, and the text-image joint path has only a limited impact on the overall sentiment distribution. However, the joint path can reduce neutral classifications, supplement sentiment intensity, and further clarify the attribution objects of experiential evaluations. The study suggests that a multimodal evidence chain can provide a transferable analytical approach for methodological expansion, validity testing, and credibility control in online tourism word-of-mouth research.