Related Experiment Video
Updated: May 25, 2026

Design and Analysis for Fall Detection System Simplification
Published on: April 6, 2020
From Generation to Detection: A Multimodal LLM-Based Framework for Realistic Vishing Attacks in Turkish
1Department of Computer Engineering, Engineering Faculty, Düzce University, Düzce, Turkey.
Abstract:
This study proposes a multimodal and AI-supported approach for both generation and detection of vishing threats, also known as audio social engineering attacks. By considering text and audio modalities together, fake contents specific to Turkish language were generated with GPT-2 and Tacotron-2 models; detection processes were carried out with discriminator structures such as BERT, CNN, LSTM, and Whisper. The developed system was designed with an integrated analysis (multimodal fusion) approach at feature and decision levels and was evaluated at both technical and psychosocial levels. In user tests conducted with 394 participants, the rate of perception of fake contents as real was measured in the range of 92.8-95.3%; it was observed that contents containing official institution imitations were the most convincing class. Experimental results show that the F1 scores obtained with single models remained in the range of 0.86-0.91%; however, 96.8% accuracy and 0.947 F1 score were achieved with the BERT + CNN + Whisper architecture. It has also been determined that GPT-4-supported content generation increases the plausibility of attack scenarios and challenges the performance of classifier models. The study is a pioneer in the field of fake audio content generation and detection in Turkish and reveals how LLM-supported multimodal systems can be used in both individual and institutional defense applications.
