Log-Mel Spectrogram

Turkish equivalent: Log-Mel spektrogramDomain: Speech Processing

A time-frequency representation formed by applying a Mel-scaled filter bank to a spectrum and taking logarithmic energies, widely used as speech-model input.

Speech Processing Context

A log-Mel spectrogram maps short-time spectral power into perceptually motivated Mel bands and then applies logarithmic compression. It is a common acoustic front end for modern ASR systems, where sample rate, window, hop, FFT size, Mel-band count, and normalization form part of the model input contract.

Representation Boundary

A log-Mel spectrogram and MFCC are not identical. MFCC pipelines normally apply an additional DCT to log-Mel energies and retain a selected number of cepstral coefficients.

Related technical article: Audio Activity and Anomaly Detection.

Related technical publications

Publications whose title or summary directly references this concept.

Whisper Architecture in Speech Recognition Systems

Whisper converts 16 kHz audio to a log-Mel spectrogram and generates tokens autoregressively with an encoder-decoder Transformer; language, task, and timestamp information share the same token sequence.