CFP last date
20 August 2026
Reseach Article

From Spatial CNNs to Multimodal Fusion: A Quantitative Survey of Cross-Dataset Generalization in Deepfake Detection

by Neha Gupta, Rahul Kumar
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Number 122
Year of Publication: 2026
Authors: Neha Gupta, Rahul Kumar
10.5120/ijca00d0437ece9c

Neha Gupta, Rahul Kumar . From Spatial CNNs to Multimodal Fusion: A Quantitative Survey of Cross-Dataset Generalization in Deepfake Detection. International Journal of Computer Applications. 187, 122 ( Jul 2026), 55-62. DOI=10.5120/ijca00d0437ece9c

@article{ 10.5120/ijca00d0437ece9c,
author = { Neha Gupta, Rahul Kumar },
title = { From Spatial CNNs to Multimodal Fusion: A Quantitative Survey of Cross-Dataset Generalization in Deepfake Detection },
journal = { International Journal of Computer Applications },
issue_date = { Jul 2026 },
volume = { 187 },
number = { 122 },
month = { Jul },
year = { 2026 },
issn = { 0975-8887 },
pages = { 55-62 },
numpages = {9},
url = { https://ijcaonline.org/archives/volume187/number122/from-spatial-cnns-to-multimodal-fusion-a-quantitative-survey-of-cross-dataset-generalization-in-deepfake-detection/ },
doi = { 10.5120/ijca00d0437ece9c },
publisher = {Foundation of Computer Science (FCS), NY, USA},
address = {New York, USA}
}
%0 Journal Article
%1 2026-07-25T01:15:13.857511+05:30
%A Neha Gupta
%A Rahul Kumar
%T From Spatial CNNs to Multimodal Fusion: A Quantitative Survey of Cross-Dataset Generalization in Deepfake Detection
%J International Journal of Computer Applications
%@ 0975-8887
%V 187
%N 122
%P 55-62
%D 2026
%I Foundation of Computer Science (FCS), NY, USA
Abstract

With the advent of AI-generated synthetic media, including that produced by generative adversarial networks (GANs), variational autoencoders, and diffusion models, deepfake detection has become a formidable challenge in digital forensics and media integrity. The survey examines 30 representative papers published from 2020 to 2025, and includes every type of spatial detector: CNN-based, frequency-domain analysis, temporal and recurrent networks, vision transformers, contrastive and self-supervised learning and multimodal audio-visual fusion. This survey provides mathematical expressions of several important loss functions, comparison tables of performance across three benchmark datasets (FaceForensics++, Celeb-DF, DFDC), and cross-dataset generalization analysis. The analysis reveals that the gap between the transformer-based and spatial-frequency hybrid detectors is ~24% (relative to the CNN baseline) and highlights open challenges in adversarial robustness, diffusion-model deepfakes, and fairness.

References
  1. Altuncu, E., Franqueira, V. N. L., & Li, S. (2024). Deepfake: Definitions, performance metrics and standards, datasets, and a meta-review. Frontiers in Big Data, 7, 1400024. https://doi.org/10.3389/fdata.2024.1400024
  2. Amerini, I., Galteri, L., Caldelli, R., and Del Bimbo, A. 2019. Deepfake video detection through optical flow based CNN. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). IEEE, 1205–1207. https://doi.org/10.1109/ICCVW.2019.00152
  3. Aqiao, M., Tian, R., and Wang, Y. 2025. Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion. arXiv preprint arXiv:2504.17223.
  4. Chai, L., Bau, D., Lim, S. N., and Isola, P. 2020. What makes fake images detectable? Understanding properties that generalize. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, Cham, 103–120.
  5. Chen, L., Zhang, Y., Song, Y., Liu, L., and Wang, J. 2022. Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18710–18719.
  6. Coccomini, D. A., Messina, N., Gennaro, C., and Falchi, F. 2022. Combining EfficientNet and vision transformers for video deepfake detection. In Proceedings of the International Conference on Image Analysis and Processing (ICIAP). Springer, Cham, 219–229.
  7. Fang, Z., Zhao, H., Wei, T., Zhou, W., Wan, M., Wang, Z., and Yu, N. 2025. UniForensics: Face forgery detection via general facial representation. IEEE Transactions on Dependable and Secure Computing (2025).
  8. Haliassos, A., Vougioukas, K., Petridis, S., and Pantic, M. 2021. Lips don't lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5039–5049.
  9. Huang, B., Wang, Z., Yang, J., Ai, J., Zou, Q., Wang, Q., and Ye, D. 2023. Implicit identity driven deepfake face swapping detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4490–4499.
  10. Kaur, A., Noori Hoshyar, A., Saikrishna, V., Firmin, S., and Xia, F. 2024. Deepfake video detection: Challenges and opportunities. Artificial Intelligence Review 57, 6 (2024), 159.
  11. Khan, I., Khan, K., and Ahmad, A. 2025. A comprehensive survey of DeepFake generation and detection techniques in audio-visual media. ICCK Journal of Image Analysis and Processing 1, 2 (2025), 73–95.
  12. Deepa, K. P., Lokesh, C. K., Umamaheswari, D., Ayshwarya, B., Yethish, P. V., and Suhaas, B. 2026. An enhanced deep learning framework for DeepFake detection using EfficientNet-B3: Comparative evaluation of deep and machine learning techniques. Discover Computing 29, 1 (2026), 18.
  13. Li, Y., Chang, M. C., and Lyu, S. 2018. In ictu oculi: Exposing AI created fake videos by detecting eye blinking. In Proceedings of the 2018 IEEE International Workshop on Information Forensics and Security (WIFS). IEEE, 1–7.
  14. Li, L. and Lyu, S. 2020. Face X-ray for more general face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5001–5010.
  15. Li, Z., Tang, W., Gao, S., Wang, Y., and Wang, S. 2026. Multiple-context and frequency aggregation network for deepfake detection. PLoS ONE 21, 1 (2026), e0337409.
  16. Liu, H., Li, X., Zhou, W., Chen, Y., He, Y., Xue, H., and Yu, N. 2021. Spatial-phase shallow learning: Rethinking face forgery detection in frequency domain. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 772–781.
  17. Masi, I., Killekar, A., Mascarenhas, R. M., Gurudatt, S. P., and AbdAlmageed, W. 2020. Two-branch recurrent network for isolating deepfakes in videos. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, Cham, 667–684.
  18. Nguyen, H. H., Yamagishi, J., and Echizen, I. 2021. Capsule-network-based method for detecting deepfake videos and images. arXiv preprint arXiv:1910.12467.
  19. Ni, Y., Meng, D., Yu, C., Quan, C., Ren, D., and Zhao, Y. 2022. CORE: Consistent representation learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 12–21.
  20. Raza, A., Basit, A., Amin, A., Arfeen, Z. A., Masud, M. I., Fayyaz, U., and Jumani, T. A. 2026. A comprehensive review of deepfake detection techniques: From traditional machine learning to advanced deep learning architectures. AI 7, 2 (2026), 68.
  21. Rössler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., and Nießner, M. 2019. FaceForensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 1–11.
  22. Shiohara, K. and Yamasaki, T. 2022. Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18720–18729.
  23. Sunil, R., Mer, P., Diwan, A., Mahadeva, R., and Sharma, A. 2025. Exploring autonomous methods for deepfake detection: A detailed survey on techniques and evaluation.
  24. Tan, C., Zhao, Y., Wei, S., Gu, G., Liu, P., & Wei, Y. (2024, March). Frequency-aware deepfake detection: Improving generalizability through frequency space domain learning. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, No. 5, pp. 5052-5060).
  25. Tian, C., Hua, G., and Tang, H. 2023. Deepfake detection with masked attention. In Proceedings of the ACM International Conference on Multimedia (MM '23).
  26. Wang, Z., Bao, J., Zhou, W., Wang, W., and Li, H. 2023. AltFreezing for more general video face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4129–4138.
  27. Wang, J., Wu, Z., Ouyang, W., Han, X., Chen, J., Jiang, Y. G., and Li, S. N. 2022. M2TR: Multi-modal multi-scale transformers for deepfake detection. In Proceedings of the 2022 International Conference on Multimedia Retrieval (ICMR '22). 615–623.
  28. Wodajo, D., Atnafu, S., and Akhtar, Z. 2023. Deepfake video detection using generative ConvViT. arXiv preprint.
  29. Xu, J., et al. 2023. Leveraging real-world data for deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW).
  30. Yang, Z., Liang, J., Xu, Y., Zhang, X. Y., and He, R. 2023. Masked relation learning for deepfake detection. IEEE Transactions on Information Forensics and Security 18 (2023), 1696–1708.
  31. Yan, Z., et al. 2023. LSDA: Large scale deepfake detection algorithm. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  32. Zsai, C. C., Wu, T. H., and Lai, S. H. 2022. Multi-scale patch-based representation learning for image anomaly detection and segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). 3992–4000.
  33. Zhan, N., Nguyen, T., Bermak, A., and Khalil, I. 2025. CAMME: Adaptive deepfake image detection with multi-modal cross-attention. arXiv preprint arXiv:2505.18035.
  34. Zheng, Y., Bao, J., Chen, D., Zeng, M., and Wen, F. 2021. Exploring temporal coherence for more general video face forgery detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 15044–15054.
Index Terms

Computer Science
Information Sciences

Keywords

Deepfake Detection GAN Vision Transformer FaceForensics++ Frequency Domain Contrastive Learning Multimodal Fusion Digital Forensics