Total: 1
Self-supervised learning (SSL) has become popular in speech processing because it generalizes well across downstream tasks. However, many SSL methods focus on capturing long-term contextual and speaker information, often making their representations invariant to background acoustics, which is an issue for speech quality assessment as it depends heavily on non-speech and fine-grained acoustic cues. These models also tend to be parameter heavy, limiting their use on resource constrained devices. In this work, we introduce an encoder that incorporates acoustic detail by combining local spectral-temporal modeling, a frame wise spectral relationship aggregator, and explicit noise and reverberation information alongside speech content. Using this pre-training framework for speech quality assessment across multiple datasets, we show that the encoder effectively extracts fine-grained acoustic features and achieves performance comparable to much larger models.