Related Experiment Video
Updated: Sep 25, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
A comparative study of DistilBERT and Perceiver for deceptive review detection in hospitality review data
Ioana Ximena Remeş1, Felicia Mirabela Costea1, Cornelia Aurora Gyorödi1
1Department of Computers and Information Technology, Faculty of Electrical Engineering and Information Technology, University of Oradea, Oradea, Romania.
Abstract:
Deceptive reviews on hospitality platforms can undermine consumer trust and affect the reliability of online reputation systems. This study presents a comparative empirical evaluation of two deep learning pipelines for deceptive review detection: a DistilBERT-based pipeline fine-tuned end-to-end on review text and complemented with rating and sentiment features, and a Perceiver-based pipeline operating on fixed sentence embeddings generated by a frozen encoder. The experiments were conducted on a combined English-language corpus of 3,322 hotel reviews assembled from the Myleott benchmark corpus and a Kaggle dataset of authentic Ritz-Carlton New York reviews, comprising approximately 2,520 genuine and 800 deceptive reviews. The corpus was evaluated using stratified 80/20 train-test splits, with 2,657 reviews used for training and 665 for testing in each split, and the deep learning models were trained for 10 epochs. To assess the robustness of the findings, the evaluation was repeated across three independent random seeds and compared with majority-class and TF-IDF plus logistic regression baselines. DistilBERT achieved the most stable performance, reaching an average accuracy of 0.9348 ± 0.0014, an F1-score of 0.9569 ± 0.0013, and a Matthews correlation coefficient of 0.8240 ± 0.0012. In contrast, the original Perceiver pipeline showed a strong tendency to collapse toward the majority class, which limited its ability to identify deceptive reviews despite apparently high recall. Although class-weighted training improved Perceiver's behavior, its results remained less reliable and more variable than those obtained with DistilBERT. The balanced TF-IDF plus logistic regression baseline performed close to DistilBERT, indicating that well-configured traditional natural language processing methods remain competitive on datasets of this size. The ablation analysis further indicates that the rating and sentiment features contribute primarily to the stability of the model across different train-test splits, rather than providing a substantial independent gain in discriminative performance. The results indicate that the fine-tuned DistilBERT pipeline provides the most robust option in the evaluated setting, while also emphasizing the importance of reporting corpus composition, test-set size, class imbalance, and imbalance-aware metrics when developing deceptive review detection systems for hospitality review data.
Related Concept Videos
Understanding Deception
Dark Triad and Person Perception
Self-Discrepancy Theory
Social Proof
Deindividuation
Stereotype Content Model
