AMDM-SE: Attention-based Multichannel Diffusion Model for Speech Enhancement

Block diagram of AMDM-SE
Network architecture, AMDM-SE in the green frame and Cross Channel TF-Attention in the blue frame.

AMDM-SE is an attention-based multichannel diffusion model for speech enhancement, built upon the single-channel SGMSE framework. The diffusion process is defined on a single-channel clean speech representation, while multichannel noisy observations are incorporated as conditioning through a novel cross-channel time–frequency attention mechanism. Experiments on the CHiME-3 benchmark demonstrate consistent improvements over the single-channel diffusion baseline, a multichannel variant without attention, and strong DNN-based predictive methods.


Noisy SGMSE AMDM-SE (our)