Total: 1
Real-time speech applications demand codecs that stream with low latency while preserving both acoustic quality and linguistic content. Waveform-oriented codecs suppress fine-grained phonetic cues at low bitrates, while semantic-aware codecs rely on non-causal encoders or dual-branch architectures incompatible with streaming. We propose LitCodec, a semantic-aware neural speech codec that injects ASR-guided supervision directly into a fully causal encoder before quantization, embedding linguistic structure into a single-codebook representation via Finite Scalar Quantization. This avoids separate semantic branches, enabling single-pass streaming without sacrificing semantic fidelity. On LibriSpeech, LitCodec achieves the highest PESQ and STOI among streaming codecs (PESQ 2.56, STOI 0.925) with strong semantic preservation (WER 2.8%) at 800 bps and 50 tokens per second. At 640 bps, LitCodec maintains 3.1% WER while EnCodec degrades to 29.0% at 750 bps.