Deep Learning-Based Pre-Processing Pipeline and Hybrid Vision Transformer Model for Deepfake Video Detection

Main Article Content

Harith A. Hussein
Khalid Shaker
Salwani Abdullah

Abstract

Deepfake technology has increased the generation of hyper-realistic synthetic images and videos using AI that pose a threat to personal privacy, integrity in business and national security. Machine Learning (ML) and Deep Learning (DL) have been extensively used in Deepfake generation and detection due their feature learning abilities and precise predictions. While DL models overtake ML in Deepfake detection, model complexity, and high resource requirements are the major limitations. Considering such limitations, this paper developed an advanced pre-processing pipeline and the detection model for improving the Deepfake detection in videos. The pre-processing pipeline model includes Adaptive Cascaded Temporal Convolutional Networks (ACTCN) for face detection and alignment, Multi-scale ResNet-based feature fusion technique for up-scaling detection face region, CNN-based Auto-Encoder for color space transformation, Deep CNN with Fast Fourier Transformation for frequency and spatial feature enhancement, 3D-CNN for temporal consistency and ResNet for noise and texture anomaly detection. Then, the Hybrid Localized Artifact Adaptively Weighted Multi-scale Attention Network-Spatio-Temporal Hypergraph Optimized Convolutional Vision Transformer Networks (HLAAWMSAN-STHGOCVT) is developed for Deepfake video Detection. This method combined the HLAAWMSAN for feature extraction from Deepfake frames and the STHGOCVT module for Deepfake detection . This framework ensures high robustness against evolving Deepfake manipulation techniques while maintaining computational efficiency. Evaluated using benchmark and real-time video data, this proposed model has achieved 99.15% accuracy and 98.21% AUC.

Article Details

Section

Articles

How to Cite

Deep Learning-Based Pre-Processing Pipeline and Hybrid Vision Transformer Model for Deepfake Video Detection (Harith A. Hussein, Khalid Shaker, & Salwani Abdullah, Trans.). (2026). Babylonian Journal of Machine Learning, 2026, 110–142. https://doi.org/10.58496/BJML/2026/012