GS-Codec: A Gaussian-Splatting Bottleneck for Neural Audio Coding

Ron Aluf    Alon Canfi    Eliya Nachmani

School of Electrical and Computer Engineering
Ben-Gurion University of the Negev

alufr@post.bgu.ac.il, canfia@post.bgu.ac.il, eliyanac@bgu.ac.il

Animation of 80 Gaussians fitting the encoder latent over the inner-loop optimization steps

Abstract

Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER).

Audio samples

The encoder, decoder, and training recipe are held fixed, and only the bottleneck changes: the Gaussian decomposition against residual vector quantization (RVQ) and finite scalar quantization (FSQ) trained in the same framework at the same bitrate, shown at both 6 kbps and 3 kbps. DAC and EnCodec are the published models. Every system in a row is rendered from the same source utterance; the same four ground-truth clips head both tables.

≈6 kbps

UtteranceGround truthGS-Codec5.83 kbpslearned pos., 100 prim., 5-bitGS-Codec6.05 kbpsfixed grid, 110 prim., 5-bitGS-Codec6.00 kbpspredictor, 109 prim., 5-bitRVQ6.00 kbpsFSQ6.00 kbpsDAC6.0 kbpsEnCodec6.0 kbps
4446-2275-0001
5142-33396-0015
7021-79759-0000
7729-102255-0023

≈3 kbps

UtteranceGround truthGS-Codec2.98 kbpslearned pos., 63 prim., 4-bitGS-Codec2.99 kbpsfixed grid, 68 prim., 4-bitGS-Codec2.91 kbpspredictor, 80 prim., 3-bitRVQ3.00 kbpsFSQ3.00 kbps
4446-2275-0001
5142-33396-0015
7021-79759-0000
7729-102255-0023

Method

Figure 1
Figure 1: GS-Codec pipeline. Training (top): the encoder latent is approximated by NG Gaussian primitives via an inner optimization loop. Inference (bottom): a GS Predictor Net regresses primitive parameters in a single pass, which are then quantized, transmitted, and decoded.
Figure 2
Figure 2: Visualization of the GS inner loop on a single latent channel across optimization steps 0, 50, and 100. Ground-truth is shown in blue, the rendered approximation in red, and the individual weighted Gaussian primitives underneath.