Attention consistency for robust deepfake detection : An informationtheoretic perspective
Swati Jadhavswati.jadhav@mitwpu.edu.inDepartment of Computer Engineering & TechnologyDr. Vishwanath Karad MIT World Peace UniversityPune, Maharashtra, IndiaView full profile → , *Sagar MohiteCorresponding authorsgmohite@bvucoep.edu.inDepartment of Computer Science and Engineering,Bharati Vidyapeeth (Deemed to be University), College of EngineeringPune , Maharashtra, 411043, IndiaView full profile → , Vaishali Rajputvaishali.rajput@vit.eduDepartment of Artificial Intelligence and Data ScienceVishwakarma Institute of TechnologyMaharashtra, IndiaView full profile → , Rashmi Ashtagirashmi.ashtagi@mitwpu.edu.inDepartment of Computer Engineering & TechnologyDr. Vishwanath Karad MIT World Peace UniversityPune, Maharashtra, IndiaView full profile → , Anusha Paianusha.pai@mitwpu.edu.inDepartment of Computer Engineering & TechnologyDr. Vishwanath Karad MIT World Peace UniversityPune, Maharashtra, IndiaView full profile → , Santushti Betgerisantushti.betgeri@mitwpu.edu.inDepartment of Computer Engineering & TechnologyDr. Vishwanath Karad MIT World Peace UniversityPune, Maharashtra, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 01 Jan 2026
- Published Online:
- 30 Jul 2026
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JIOS-2374
- Pages:
- 1–20
Abstract
The increasing realism of Deepfake media generated by modern models poses significant risks to digital trust and forensics. Although Vision Transformer-based detectors outperform Convolutional Methods, they remain vulnerable to unstable attention under common perturbations and often operate in closed-set settings, leading to overconfidence on unseen forgeries. This work analyzes attention consistency as an information-theoretic approach to enhance detection reliability. We propose an Attention-Consistent Vision Transformer (ViT-ACR) that employs Kullback–Leibler divergence to stabilize self-attention distributions. An entropy-based decision rule enables uncertainty-aware open-set inference. Experiments under identity-disjoint protocols across multiple public benchmarks demonstrate strong cross-dataset generalization, achieving 99.28% accuracy on FaceForensics++ and 89.57% mean AUC on six unseen datasets, with reduced predictive uncertainty under compression.
Keywords
Subject Classifications
References
[1] D. Afchar, V. Nozick, J. Yamagishi, and I. Echizen, “MesoNet: A Compact Facial Video Forgery Detection Network,” in Proc. IEEE Int. Workshop Inf. Forensics Secur. (WIFS), pp. 1–7 (2018).
[2] S. Agarwal, H. Farid, Y. Gu, M. He, K. Nagano, and H. Li, “Protecting World Leaders Against Deepfakes,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW) (2019).
[3] S. Agarwal, H. Farid, O. Fried, and M. Agrawala, “Detecting Deep-Fake Videos from Phoneme-Viseme Mismatches,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), pp. 660–661 (2020).
[4] A. Aghasanli, D. Kangin, and P. Angelov, “Interpretable-Through-Prototypes Deepfake Detection for Diffusion Models,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 467–474 (2023).
[5] Z. Akhtar, “Deepfakes Generation and Detection: A Short Survey,” Journal of Imaging, vol. 9, no. 1, pp. 18 (2023).
[6] I. Amerini, L. Galteri, R. Caldelli, and A. Del Bimbo, “Deepfake Video Detection through Optical Flow-Based CNN,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. Workshops (ICCVW) (2019).
[7] N. Bansal, T. Aljrees, D. P. Yadav, K. U. Singh, A. Kumar, G. K. Verma, and T. Singh, “Real-Time Advanced Computational Intelligence for Deep Fake Video Detection,” Applied Sciences, vol. 13, no. 5, p. 3095 (2023). (note: the manuscript’s title also drops “Real-Time” — worth restoring for accuracy)
[8] N. Carlini and H. Farid, “Evading Deepfake-Image Detectors with White- and Black-Box Attacks,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), pp. 658–659 (2020).
[9] Y. Hou, Q. Guo, Y. Huang, X. Xie, L. Ma, and J. Zhao, “Evading DeepFake Detectors via Adversarial Statistical Consistency,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 12271–12280 (2023).
[10] P. Korshunov and S. Marcel, “Vulnerability Assessment and Detection of Deepfake Videos,” in Proc. IEEE Int. Conf. Biometrics (ICB) (2019).
[11] H. Dang, F. Liu, J. Stehouwer, X. Liu, and A. K. Jain, “On the Detection of Digital Face Manipulation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 5781–5790 (2020).
[12] S. Kingra, N. Aggarwal, and N. Kaur, “Emergence of Deepfakes and Video Tampering Detection Approaches: A Survey,” Multimedia Tools and Applications, vol. 82, no. 7, pp. 10165–10209 (2023).
[13] A. Kohli and A. Gupta, “Detecting DeepFake, FaceSwap, and Face2Face Facial Forgeries Using Frequency CNN,” Multimedia Tools and Applications, vol. 80, pp. 18461–18478 (2021).
[14] X. Yang, Y. Li, and S. Lyu, “Exposing Deepfakes Using Inconsistent Head Poses,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), pp. 8261–8265 (2019).
[15] L. Gong and X. Li, “A Contemporary Survey on Deepfake Detection: Datasets, Algorithms, and Challenges,” Electronics, vol. 13, no. 3, pp. 585 (2024).
[16] W. Zhuang, Q. Chu, Z. Tan, Q. Liu, H. Yuan, C. Miao, Z. Luo, and N. Yu, “UIA-ViT: Unsupervised Inconsistency-Aware Method Based on Vision Transformer for Face Forgery Detection,” in Proc. Eur. Conf. Comput. Vis. (ECCV), pp. 391–407 (2022).
[17] Y. Heo, W. Yeo, and B. Kim, “Deepfake Detection Based on Improved Vision Transformer,” Applied Intelligence, vol. 53, no. 7, pp. 7512–7527 (2023).
[18] A. Khormali and X. Yuan, “DFDT: An End-to-End Deepfake Detection Framework Using Vision Transformer,” Applied Sciences, vol. 12, no. 6, pp. 2953 (2022).
[19] D. Wodajo, P. Lambert, G. Van Wallendael, S. Atnafu, and H. Mareen, “Improved Deepfake Video Detection Using Convolutional Vision Transformer,” in Proc. IEEE Gaming, Entertainment, Media Conf. (GEM), pp. 1–6 (2024).
[20] J. Wang, Z. Wu, W. Ouyang, X. Han, J. Chen, Y.-G. Jiang, and S.-N. Li, “M2TR: Multi-Modal Multi-Scale Transformers for Deepfake Detection,” in Proc. ACM Int. Conf. Multimedia Retrieval (ICMR), pp. 615–623 (2022).
[21] A. Haliassos, K. Vougioukas, S. Petridis, and M. Pantic, “LipForensics: Using Lip Movements for Deepfake Detection,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 5039–5049 (2021).
[22] X. Dong, J. Bao, D. Chen, T. Zhang, W. Zhang, N. Yu, D. Chen, F. Wen, and B. Guo, “Protecting Celebrities from Deepfake with Identity Consistency Transformer,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 9468–9478 (2022).
[23] B. Liu, B. Liu, M. Ding, and T. Zhu, “MeST-Former: Motion-Enhanced Spatiotemporal Transformer for Generalizable Deepfake Detection,” Neurocomputing, vol. 610, p. 128588 (2024).
[24] A. Rössler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner, “FaceForensics++: Learning to Detect Manipulated Facial Images,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV) (2019).
[25] B. Dolhansky, J. Bitton, B. Pflaum, J. Lu, R. Howes, M. Wang, and C. Canton Ferrer, “The Deepfake Detection Challenge (DFDC) Dataset,” arXiv:2006.07397 (2020).
[26] J. Wang, X. Jin, H. Wang, and L. Jiang, “VAD-Lip: Visual and Audio Deepfake Detection via Lip Features,” in Proc. ACM Workshop Deepfake Forensics (DFF), pp. 110–117 (2025).
[27] K. Zhang, Z. Zhang, Z. Li, and Y. Qiao, “Joint Face Detection and Alignment Using Multitask Cascaded Convolutional Networks,” IEEE Signal Processing Letters, vol. 23, no. 10, pp. 1499–1503 (2016).
[28] H. Zhao, W. Zhou, D. Chen, T. Wei, W. Zhang, and N. Yu, “Multi-Attentional Deepfake Detection,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 2185–2194 (2021).
[29] P. Narooka, D. Sharma, M. T. Basu, V. Kumar, S. C. Addimulam, N. Kapila, and N. K. Verma, “Deepfake detection: A new frontier in cybersecurity for protecting digital identities, preventing misinformation, and ensuring data integrity,” Journal of Discrete Mathematical Sciences and Cryptography, vol. 28, no. 8, pp. 3013–3023 (2025), doi: 10.47974/JDMSC-2445.
[30] S. Sharma, P. Sharma, S. Shanker, P. Veluvali, S. Tanwar, and P. Vats, “A PKI-integrated cryptographic framework for Deepfake detection via facial micro-expression analysis,” Journal of Discrete Mathematical Sciences and Cryptography, vol. 28, no. 8, pp. 3101–3109 (2025), doi: 10.47974/JDMSC-2586.


