STFT-VAE — neural audio codec comparison

Speech reconstruction vs. SNAC, Mimi, and MioCodec over 8 LibriTTS-R clips (24 kHz mono, distinct speakers). Each codec encodes a clip and decodes it back to a waveform. Listen below; the table averages the objective metrics.

Codecs

CodecTypeFrame/token rateBitrateTrained on
STFT-VAE (ours)continuous latent (STFT/ISTFT)3.125 Hzcontinuous, 128-dgeneral audio
MimiRVQ, acoustic+semantic12.5 Hz1.1 kbpsspeech
SNACmulti-scale RVQ (3 levels)12 / 23 / 47 Hz~0.98 kbpsspeech
MioCodecsemantic tokens + global embedding25 Hz0.34 kbpsspeech

Reading the numbers. STFT-VAE is a continuous codec at 3.125 Hz — 4–15× lower frame rate than the others — so its competitiveness is the point of interest, not a like-for-like bitrate match. MioCodec is a semantic codec: it keeps linguistic content plus a single global speaker embedding and discards fine acoustic detail, so its waveform-fidelity scores (SI-SDR) are low by design. Scores use ±50 ms cross-correlation alignment.

Objective metrics (mean over 8 clips)

CodecSI-SDR↑ (dB)Mel-L1↓ (dB)PESQ↑STOI↑
STFT-VAE (ours)9.393.512.560.932
Mimi12.683.323.560.969
SNAC4.353.462.280.921
MioCodec-15.807.981.560.770

Listen

Speech 1 (LibriTTS-R)

Original
STFT-VAE (ours)
SI-SDR 10.62 · PESQ 2.33
Mimi
SI-SDR 12.98 · PESQ 3.50
SNAC
SI-SDR 5.16 · PESQ 1.80
MioCodec
SI-SDR -17.41 · PESQ 1.45

Speech 2 (LibriTTS-R)

Original
STFT-VAE (ours)
SI-SDR 8.78 · PESQ 2.73
Mimi
SI-SDR 13.27 · PESQ 3.61
SNAC
SI-SDR 3.68 · PESQ 2.53
MioCodec
SI-SDR -12.52 · PESQ 1.79

Speech 3 (LibriTTS-R)

Original
STFT-VAE (ours)
SI-SDR 7.99 · PESQ 2.50
Mimi
SI-SDR 10.74 · PESQ 3.48
SNAC
SI-SDR 3.13 · PESQ 2.11
MioCodec
SI-SDR -11.34 · PESQ 1.61

Speech 4 (LibriTTS-R)

Original
STFT-VAE (ours)
SI-SDR 8.89 · PESQ 2.54
Mimi
SI-SDR 11.58 · PESQ 3.64
SNAC
SI-SDR 4.45 · PESQ 2.30
MioCodec
SI-SDR -15.53 · PESQ 1.53

Speech 5 (LibriTTS-R)

Original
STFT-VAE (ours)
SI-SDR 9.45 · PESQ 2.99
Mimi
SI-SDR 13.60 · PESQ 3.73
SNAC
SI-SDR 3.45 · PESQ 2.66
MioCodec
SI-SDR -16.18 · PESQ 1.79

Speech 6 (LibriTTS-R)

Original
STFT-VAE (ours)
SI-SDR 9.24 · PESQ 2.49
Mimi
SI-SDR 12.94 · PESQ 3.55
SNAC
SI-SDR 4.73 · PESQ 2.34
MioCodec
SI-SDR -18.74 · PESQ 1.44

Speech 7 (LibriTTS-R)

Original
STFT-VAE (ours)
SI-SDR 10.62 · PESQ 2.33
Mimi
SI-SDR 13.12 · PESQ 3.39
SNAC
SI-SDR 6.45 · PESQ 2.28
MioCodec
SI-SDR -17.13 · PESQ 1.40

Speech 8 (LibriTTS-R)

Original
STFT-VAE (ours)
SI-SDR 9.50 · PESQ 2.59
Mimi
SI-SDR 13.21 · PESQ 3.57
SNAC
SI-SDR 3.73 · PESQ 2.22
MioCodec
SI-SDR -17.52 · PESQ 1.47

Test clips: LibriTTS-R (public-domain LibriVox speech). Reproduce with comparisons/ in the repo. Audio is lossless WAV.