Finetuning Strategies for Querying Sounds by Vocal Imitation

arXiv:2608.19174v1 Announce Type: cross
Abstract: This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: