Total: 1
Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features they encode, nor which are useful for a given task. This limits transparency and extension to rare species or data-scarce domains. We ask which speech-like features bioacoustic embeddings encode, framing evaluation as interpretability rather than benchmarking. Using the 88 eGeMAPS features across six taxonomic groups, we apply linear and nonlinear regression probes to quantify which acoustic properties each model captures. Results confirm a "no free lunch" pattern: no single model captures the full feature space, while a concatenated embedding performs best, suggesting complementary coverage. Loudness is best encoded (R² = 0.76) while F0 is hardest to recover (R² = 0.33). Cross-referencing recoverability with per-species feature salience (NMI) offers data-driven hypotheses for model selection in bioacoustics.