Total: 1
Self-supervised learning (SSL) yields powerful, context-rich representations for speech emotion recognition (SER), yet the aggregation of these representations into holistic descriptors remains a bottleneck. Conventional first-order aggregation implicitly assumes feature independence, violating the latent Riemannian geometry and discarding higher-order relationships essential to the backbone's representational power. To address this problem, this paper proposes a novel Second-Order Correlation (SOC) layer. Instead of treating features in isolation, our SOC models correlations among features as covariance descriptors to capture synergistic co-occurrence patterns, which act as discriminative signatures for robust emotion recognition. By mapping these descriptors from the Riemannian manifold to a Euclidean tangent space through Log-Euclidean mapping (LEM), our method preserves geometric integrity while enabling direct linear discriminative learning. Extensive experiments on ESD and RAVDESS datasets demonstrate that SOC recovers discriminative information lost in first-order pooling, effectively aggregating high-dimensional SSL features.