AMDM-SE is an attention-based multichannel diffusion model for speech enhancement, built upon the single-channel SGMSE framework. The diffusion process is defined on a single-channel clean speech representation, while multichannel noisy observations are incorporated as conditioning through a novel cross-channel time–frequency attention mechanism. Experiments on the CHiME-3 benchmark demonstrate consistent improvements over the single-channel diffusion baseline, a multichannel variant without attention, and strong DNN-based predictive methods.
| Noisy | SGMSE | AMDM-SE (our) |
|---|---|---|