Research on cross-domain multimodal feature fusion method for ground moving target recognition
Target recognition of unattended system plays an essential role in guarding key areas. Most of the existing studies based on acoustic and seismic fusion only conduct single-modal analysis on the sensing signals, which leads to the problem of single and one-sided feature extraction, seriously affecting the accuracy of classification and recognition. In this paper, an UGS system specifically designed for collecting seismic and acoustic signals is constructed, and a cross-domain multimodal feature fusion network architecture is proposed. Firstly, based on the different sensing characteristics of acoustic signals and vibration signals, time–frequency graph data are obtained respectively through MEL spectrum calculation and Continuous Wavelet Transform, and the information expression of sensing signals is enriched through time–frequency conversion. Secondly, the time series is input into the Long Short-Term Memory network to extract the time-domain features, and the time–frequency graph is input into the VGG-19 network to extract deep image features. In this way, rich feature information can be extracted and the differences in multimodal data features can be alleviated. The accuracy of the proposed network reached 98.17% on the self-built acoustic-seismic data set. Compared with the single-modal network, the performance of the multi-modal fusion network is significantly improved.