
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER).
The encoder, decoder, and training recipe are held fixed, and only the bottleneck changes: the Gaussian decomposition against residual vector quantization (RVQ) and finite scalar quantization (FSQ) trained in the same framework at the same bitrate, shown at both 6 kbps and 3 kbps. DAC and EnCodec are the published models. Every system in a row is rendered from the same source utterance; the same four ground-truth clips head both tables.
| Utterance | Ground truth | GS-Codec5.83 kbpslearned pos., 100 prim., 5-bit | GS-Codec6.05 kbpsfixed grid, 110 prim., 5-bit | GS-Codec6.00 kbpspredictor, 109 prim., 5-bit | RVQ6.00 kbps | FSQ6.00 kbps | DAC6.0 kbps | EnCodec6.0 kbps |
|---|---|---|---|---|---|---|---|---|
| 4446-2275-0001 | ||||||||
| 5142-33396-0015 | ||||||||
| 7021-79759-0000 | ||||||||
| 7729-102255-0023 |
| Utterance | Ground truth | GS-Codec2.98 kbpslearned pos., 63 prim., 4-bit | GS-Codec2.99 kbpsfixed grid, 68 prim., 4-bit | GS-Codec2.91 kbpspredictor, 80 prim., 3-bit | RVQ3.00 kbps | FSQ3.00 kbps |
|---|---|---|---|---|---|---|
| 4446-2275-0001 | ||||||
| 5142-33396-0015 | ||||||
| 7021-79759-0000 | ||||||
| 7729-102255-0023 |