Total: 1
This paper investigates audio-to-audio retrieval using self-supervised learning (SSL) models to generate audio representations without labeled data. To enhance retrieval accuracy, we explore the use of SSL embeddings with sequence matching techniques, including Dynamic Time Warping (DTW), and clustering methods, such as K-Means combined with TF-IDF and BM25. We evaluate our framework on two distinct tasks: music retrieval via query-by-humming and spoken content retrieval via query-by-example. Experimental evidence shows that clustering-based methods, which reduce SSL embeddings into discrete hidden units, are particularly effective for speech retrieval. Conversely, DTW applied directly on SSL embeddings, which preserves full sequence information, excels in music retrieval. Extensive experiments demonstrate that combining SSL representations with appropriate sequence matching improves retrieval accuracy across different audio domains.