Abstract: Speaker recognition neural networks recognise speaker identities from input utterances by learning latent representations (i.e. speaker embeddings). However, these networks' internal mechanisms remain largely opaque, motivating research in explainable artificial intelligence (XAI) to understand them. Existing studies have analysed how speaker embeddings are organised, but rarely frame these analyses within XAI. This work proposes to explain and interpret the organisation of speaker embeddings from an XAI perspective.
To this end, we apply a hierarchical clustering algorithm, Single-Linkage Clustering (SLINK), to analyse whether our prepared speaker embeddings naturally form clusters with hierarchical relationships. The resulting hierarchical organisation (i.e. hierarchical clusters) is evaluated using the Cluster-Class Matching (CCM) method. Moreover, we propose a new method, termed Hierarchical Cluster-Class Matching (HCCM), to identify which hierarchical clusters best match individual semantic classes like male and conjunctive semantic classes like UK & male, thereby interpreting the clusters using their matched classes. We quantify the matching degree with a new metric called the L-score, which makes imperfect matches diagnosable. HCCM's results reveal that the hierarchical clusters analysed by SLINK are well interpreted using classes related to speaker identity, gender, and nationality, showing the semantics inside the hierarchical organisation of the examined speaker embeddings.
Read the original article:
