Abstract: Generating human-centric videos that preserve both visual identity and person-specific expressive behavior remains a fundamental challenge. In addition to reproducing appearance, a model must replicate the facial behaviors that characterize how a subject expresses emotion over time. However, most state-of-the-art methods condition generation on a single reference image, which contains no information about these temporal dynamics. As a result, they tend to preserve the subject's visual identity but often produce expressions with limited variation and weak subject specificity. To mitigate this issue, we introduce BEACON, a lightweight framework for person-specific video generation that produces more expressive videos by disentangling visual identity from expressive behavior. BEACON conditions generation on two complementary signals: a reference image encoding the identity and a reference video capturing subject-specific facial dynamics. By conditioning on these complementary signals, BEACON generates videos that better preserve both the subject's appearance and characteristic facial dynamics, while also supporting identity-expression transfer. Our experiments on the MEAD and RAVDESS datasets show that by fine-tuning on approximately 2,000 pairs and updating about 1% of the pretrained Wan video diffusion model, BEACON improves facial expressivity over state-of-the-art video generation methods while maintaining competitive identity preservation.
Read the original article:
