Multimodal Attentive Fusion Network for audio-visual event recognition

扫码查看

原文链接

NSTL
Elsevier

外文摘要：Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multimodal Attentive Fusion Network (MAFnet), an architecture that can dynamically fuse visual and audio information for event recognition. Inspired by prior studies in neuroscience, we couple both modalities at different levels of visual and audio paths. Furthermore, the network dynamically highlights a modality at a given time window relevant to classify events. Experimental results in AVE (Audio-Visual Event), UCF51, and Kinetics-Sounds datasets show that the approach can effectively improve the accuracy in audio-visual event classification. Code is available at: https://github.com/numediart/MAFnet

外文关键词：

Audio-visual fusionModality conditioningAttentionMultimodal deep learningEvent recognition

作者：

Brousmiche, Mathilde、Rouat, Jean、Dupont, Stephane

展开 >

作者单位：

Univ Sherbrooke

Univ Mons

出版年：

2022

DOI：

10.1016/j.inffus.2022.03.001

Information Fusion

EISCI

ISSN：1566-2535

年,卷(期)：2022.85

被引量6
参考文献量54