SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis
Researchers have introduced SDTalk, a novel framework for high-quality, real-time talking head synthesis that addresses the limitations of existing identity-specific models. Published on arXiv in May 2026, this one-shot 3D Gaussian Splatting (3DGS) based system generalizes to unseen identities without requiring personalized training or fine-tuning. The framework utilizes a two-stage training strategy. First, it incorporates structured facial priors into a reconstruction module, separately predicting 3DGS parameters for visible and occluded regions to enable complete head reconstruction from a single image. Second, it employs a dual-branch motion field to model both coarse and fine facial dynamics, significantly improving detail fidelity and lip synchronization. Experimental results indicate that SDTalk outperforms current methods in both visual quality and inference efficiency. This advancement represents a significant step forward in computer vision, offering a more scalable and efficient solution for digital human animation and virtual communication technologies by eliminating the need for extensive per-user calibration.
Wire timeline
SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis
Researchers have introduced SDTalk, a novel framework for high-quality, real-time talking head synthesis that addresses the limitations of existing identity-specific models. Published on arXiv in May 2026, this one-shot 3D Gaussian Splatting (3DGS) based system generalizes to unseen identities without requiring personalized training or fine-tuning. The framework utilizes a two-stage training strategy. First, it incorporates structured facial priors into a reconstruction module, separately predicting 3DGS parameters for visible and occluded regions to enable complete head reconstruction from a single image. Second, it employs a dual-branch motion field to model both coarse and fine facial dynamics, significantly improving detail fidelity and lip synchronization. Experimental results indicate that SDTalk outperforms current methods in both visual quality and inference efficiency. This advancement represents a significant step forward in computer vision, offering a more scalable and efficient solution for digital human animation and virtual communication technologies by eliminating the need for extensive per-user calibration.
cs.AI updates on arXiv.org