Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Parametric implicit face representation for audio-driven facial reenactment

Huang, Ricong, Lai, Peiwen, Qin, Yipeng ORCID: https://orcid.org/0000-0002-1551-9126 and Li, Guanbin 2023. Parametric implicit face representation for audio-driven facial reenactment. Presented at: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2023, Vancouver, Canada, 18 - 22 June 2023. Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp. 12759-12768. 10.1109/CVPR52729.2023.01227

[thumbnail of 733_parametric_implicit_face_repre-Camera-ready PDF.pdf]
Preview
PDF - Presentation
Download (2MB) | Preview

Abstract

Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial landmarks or 3D face models) or implicit ones (e.g., Neural Radiance Fields), thus suffering from the trade-offs between interpretability and expressive power, hence between controllability and quality of the results. In this work, we break these trade-offs with our novel parametric implicit face representation and propose a novel audio-driven facial reenactment framework that is both controllable and can generate high-quality talking heads. Specifically, our parametric implicit representation parameterizes the implicit representation with interpretable parameters of 3D face models, thereby taking the best of both explicit and implicit methods. In addition, we propose several new techniques to improve the three components of our framework, including i) incorporating contextual information into the audio-to-expression parameters encoding; ii) using conditional image synthesis to parameterize the implicit representation and implementing it with an innovative tri-plane structure for efficient learning; iii) formulating facial reenactment as a conditional image inpainting problem and proposing a novel data augmentation technique to improve model generalizability. Extensive experiments demonstrate that our method can generate more realistic results than previous methods with greater fidelity to the identities and talking styles of speakers.

Item Type: Conference or Workshop Item - published (Paper)
Date Type: Published Online
Status: Published
Schools: Schools > Computer Science & Informatics
Publisher: IEEE
ISBN: 9798350301304
ISSN: 1063-6919
Date of Acceptance: 27 February 2023
Last Modified: 26 Mar 2026 15:15
URI: https://orca.cardiff.ac.uk/id/eprint/158085

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics