VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

Maris Basha, Anja T Zai, Sabine Stoll, Richard Hahnloser
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:6799-6861, 2026.

Abstract

General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting. Unlike supervised classification benchmarks that measure adaptability via parameter updates, we introduce VocSim, a training-free benchmark probing the intrinsic geometric alignment of frozen embeddings, with no parameters updated and no labels used (a label-free PCA whitening is fit per subset to correct anisotropy). VocSim aggregates 125k single-source clips from 19 corpora spanning human speech, animal vocalizations, and environmental sounds, isolating content representation from source separation (polyphonic mixtures are out of scope). We evaluate embeddings with Precision@k for local purity and the Global Separation Rate (GSR) for point-wise class separation, calibrated by lift over an empirical permutation baseline. A simple pipeline of frozen Whisper features, time–frequency pooling, and label-free PCA yields strong zero-shot performance with stable GSR rankings across domains (Kendall’s $\tau$ = 0.60). However, on blind low-resource speech (Shipibo-Conibo, Chintang), local retrieval collapses while remaining above chance, exposing a cross-lingual speech generalization gap. As external validation, our top embeddings predict avian perceptual similarity, improve bioacoustic classification, and achieve state-of-the-art on the HEAR benchmark. We release data, code, and a public leaderboard.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-basha26a, title = {{V}oc{S}im A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio}, author = {Basha, Maris and Zai, Anja T and Stoll, Sabine and Hahnloser, Richard}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {6799--6861}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/basha26a/basha26a.pdf}, url = {https://proceedings.mlr.press/v306/basha26a.html}, abstract = {General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting. Unlike supervised classification benchmarks that measure adaptability via parameter updates, we introduce VocSim, a training-free benchmark probing the intrinsic geometric alignment of frozen embeddings, with no parameters updated and no labels used (a label-free PCA whitening is fit per subset to correct anisotropy). VocSim aggregates 125k single-source clips from 19 corpora spanning human speech, animal vocalizations, and environmental sounds, isolating content representation from source separation (polyphonic mixtures are out of scope). We evaluate embeddings with Precision@k for local purity and the Global Separation Rate (GSR) for point-wise class separation, calibrated by lift over an empirical permutation baseline. A simple pipeline of frozen Whisper features, time–frequency pooling, and label-free PCA yields strong zero-shot performance with stable GSR rankings across domains (Kendall’s $\tau$ = 0.60). However, on blind low-resource speech (Shipibo-Conibo, Chintang), local retrieval collapses while remaining above chance, exposing a cross-lingual speech generalization gap. As external validation, our top embeddings predict avian perceptual similarity, improve bioacoustic classification, and achieve state-of-the-art on the HEAR benchmark. We release data, code, and a public leaderboard.} }
Endnote
%0 Conference Paper %T VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio %A Maris Basha %A Anja T Zai %A Sabine Stoll %A Richard Hahnloser %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-basha26a %I PMLR %P 6799--6861 %U https://proceedings.mlr.press/v306/basha26a.html %V 306 %X General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting. Unlike supervised classification benchmarks that measure adaptability via parameter updates, we introduce VocSim, a training-free benchmark probing the intrinsic geometric alignment of frozen embeddings, with no parameters updated and no labels used (a label-free PCA whitening is fit per subset to correct anisotropy). VocSim aggregates 125k single-source clips from 19 corpora spanning human speech, animal vocalizations, and environmental sounds, isolating content representation from source separation (polyphonic mixtures are out of scope). We evaluate embeddings with Precision@k for local purity and the Global Separation Rate (GSR) for point-wise class separation, calibrated by lift over an empirical permutation baseline. A simple pipeline of frozen Whisper features, time–frequency pooling, and label-free PCA yields strong zero-shot performance with stable GSR rankings across domains (Kendall’s $\tau$ = 0.60). However, on blind low-resource speech (Shipibo-Conibo, Chintang), local retrieval collapses while remaining above chance, exposing a cross-lingual speech generalization gap. As external validation, our top embeddings predict avian perceptual similarity, improve bioacoustic classification, and achieve state-of-the-art on the HEAR benchmark. We release data, code, and a public leaderboard.
APA
Basha, M., Zai, A.T., Stoll, S. & Hahnloser, R.. (2026). VocSim A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:6799-6861 Available from https://proceedings.mlr.press/v306/basha26a.html.

Related Material