A CNN-transformer multimodal architecture for weakly-supervised audio-visual violence detection
1 Arizona State University.
2 Ohio University.
3 Florida State University.
4 Trine University.
5 Tribhuvan University.
Research Article
Open Access Research Journal of Engineering and Technology, 2026, 11(01), 042–051.
Article DOI: 10.53022/oarjet.2026.11.1.0056
Publication history:
Received on 11 June 2026; revised on 22 July 2026; accepted on 24 July 2026
Abstract:
Automated detection of violent events in surveillance-scale video is important for timely intervention, but frame-accurate labels are too costly to collect at scale. This has made weakly supervised learning from video-level labels the dominant paradigm. Most existing methods score short video snippets from a single modality and with limited temporal context. We propose a CNN-Transformer multimodal architecture that extracts per-snippet features from frozen, pretrained CNN backbones (I3D for video, VGGish for audio), projects each modality through a lightweight temporal 1D-CNN, and fuses the two modalities across the full temporal extent of a video with a cross-modal Transformer encoder. Snippets are then scored under a multiple-instance-learning (MIL) ranking objective. On XD-Violence, the largest public audio-visual violence detection benchmark, our 4.24M-parameter fusion network reaches 80.14% frame-level average precision (AP) and 93.29% ROC-AUC, outperforming several established baselines while staying much smaller than the CNN backbones it builds on. Ablations show that the audio modality and the Transformer fusion stage each add several points of AP over a visual-only, non-attentive baseline. We also test the model qualitatively on independently sourced, real-world YouTube footage drawn from outside the training distribution. That test exposes a limitation we believe is underappreciated: replacing the benchmark’s undisclosed official visual feature extractor with an independently implemented one collapses the visual predictions, whereas our faithfully reproduced audio pipeline still transfers. Feature-extractor fidelity, and not model architecture alone, is therefore critical for real-world deployment of weakly supervised video anomaly detection systems.
Keywords:
Anomaly Detection; Multimodal Learning; Convolutional Neural Networks; Transformers; Multiple Instance Learning; Situational Awareness
Full text article in PDF:
Copyright information:
Copyright © 2026 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution Liscense 4.0
