Multi-modal human aggression detection

作者：

Highlights：

•

摘要

This paper presents a smart surveillance system named CASSANDRA, aimed at detecting instances of aggressive human behavior in public environments. A distinguishing aspect of CASSANDRA is the exploitation of complementary audio and video cues to disambiguate scene activity in real-life environments. From the video side, the system uses overlapping cameras to track persons in 3D and to extract features regarding the limb motion relative to the torso. From the audio side, it classifies instances of speech, screaming, singing, and kicking-object. The audio and video cues are fused with contextual cues (interaction, auxiliary objects); a Dynamic Bayesian Network (DBN) produces an estimate of the ambient aggression level.Our prototype system is validated on a realistic set of scenarios performed by professional actors at an actual train station to ensure a realistic audio and video noise setting.

论文关键词：

论文评审过程：Received 19 December 2014, Revised 16 June 2015, Accepted 18 June 2015, Available online 1 April 2016, Version of Record 1 April 2016.

论文官网地址：https://doi.org/10.1016/j.cviu.2015.06.009