Audio-noise Power Spectral Density Estimation Using Long Short-term Memory

Xiaofei Li 1 Simon Leglaive 1 Laurent Girin 2, 1 Radu Horaud 1
1 PERCEPTION - Interpretation and Modelling of Images and Videos
Inria Grenoble - Rhône-Alpes, LJK - Laboratoire Jean Kuntzmann, INPG - Institut National Polytechnique de Grenoble
2 GIPSA-CRISSP - CRISSP
GIPSA-DPC - Département Parole et Cognition
Abstract : We propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short time Fourier transform (STFT) domain. An LSTM network common to all frequency bands is trained, which processes each frequency band individually by mapping the noisy STFT magnitude sequence to its corresponding noise PSD sequence. Unlike deep-learning-based speech enhancement methods that learn the full-band spectral structure of speech segments, the proposed method exploits the sub-band STFT magnitude evolution of noise with a long time dependency, in the spirit of the unsupervised noise estimators described in the literature. Speaker-and speech-independent experiments with different types of noise show that the proposed method outperforms the unsupervised estimators, and generalizes well to noise types that are not present in the training set.
Liste complète des métadonnées

https://hal.inria.fr/hal-02100059
Contributor : Team Perception <>
Submitted on : Monday, April 15, 2019 - 3:05:45 PM
Last modification on : Friday, April 19, 2019 - 11:47:25 AM

File

noise_psd.pdf
Files produced by the author(s)

Identifiers

Citation

Xiaofei Li, Simon Leglaive, Laurent Girin, Radu Horaud. Audio-noise Power Spectral Density Estimation Using Long Short-term Memory. IEEE Signal Processing Letters, Institute of Electrical and Electronics Engineers, 2019, pp.1-5. ⟨10.1109/LSP.2019.2911879⟩. ⟨hal-02100059⟩

Share

Metrics

Record views

59

Files downloads

15