Abstract: Conditional generation and disentangled representation learning are central to controlled generation across audio, vision, and multimodal domains. However, despite strong empirical progress, particularly in speech style transfer, most existing approaches rely on heuristic objectives and architectural choices, offering limited theoretical understanding of when and why reliable attribute control is achievable. In this work, we develop a formal framework for speech attribute conversion and provide a theoretical analysis of sufficient conditions for exact and consistent transfer. Our analysis focuses on a deterministic autoencoder setting augmented with an independence constraint between the learned latent representation and the controllable attribute. Under explicit population-level assumptions about the data-generating process, we establish guarantees linking reconstruction, independence, and the feasibility of attribute manipulation while preserving task-relevant content. We further show how the theoretical framework translates into practice by proposing a practical voice conversion method that directly implements its core principles. Experimental evaluations on voice and pitch conversion tasks demonstrate the applicability of the theoretical analysis to real-world speech conversion settings and show that the resulting method achieves competitive performance against existing approaches.
Read the original article:
