Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

IM-Animation: an implicit motion representation for identity-decoupled character animation

Xu, Zhufeng, Gau, Xuan, Liu, Feng-Lin, Zhang, Haoxian, Fang, Zhixue, Lai, Yu-Kun ORCID: https://orcid.org/0000-0002-2094-5680, Liu, Xiaoqiang, Wan, Pengfei and Gao, Lin 2026. IM-Animation: an implicit motion representation for identity-decoupled character animation. Presented at: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Denver, CO, USA, 3-7 June 2026. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 4635-4646.
Item availability restricted.

[thumbnail of Provisional file] PDF (Provisional file) - Accepted Post-Print Version
Download (17kB)
[thumbnail of IM-Animation-CVPRFindings.pdf] PDF - Accepted Post-Print Version
Restricted to Repository staff only

Download (3MB)

Abstract

Recent progress in video diffusion models has markedly advanced character animation, which synthesizes motioned videos by animating a static identity image according to a driving video. Despite these advances, existing methods still face challenges in robustness, particularly when the source and driving identities exhibit substantial differences in body shape or spatial layout in the video frame. Explicit methods represent motion using skeleton, DWPose or other explicit structured signals, but struggle to handle spatial mismatches and varying body scales. Implicit methods, on the other hand, capture high-level implicit motion semantics directly from the driving video, but suffer from identity leakage and entanglement between motion and appearance. To address the above challenges, we propose a novel implicit motion representation that compresses per-frame motion into compact 1D motion tokens. This design relaxes strict spatial constraints inherent in 2D representations and effectively prevents identity information leakage from the motion video. Furthermore, we design a temporally consistent mask token-based retargeting module that enforces a temporal training bottleneck, mitigating interference from the source image's motion and improving retargeting consistency. Our methodology employs a three-stage training strategy to enhance the training efficiency and ensure high fidelity. Extensive experiments demonstrate that our implicit motion representation and the proposed IM-Animation achieves superior or competitive performance compared with state-of-the-art methods.

Item Type: Conference or Workshop Item - published (Paper)
Status: In Press
Schools: Schools > Computer Science & Informatics
Publisher: IEEE
Related URLs:
Date of First Compliant Deposit: 26 June 2026
Last Modified: 26 Jun 2026 13:30
URI: https://orca.cardiff.ac.uk/id/eprint/187761

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics