Publication Details
Overview
 
 
Shuo Yang, Bart Jansen, Hichem Sahli, Xuan-son Nguyen, Aymeric Histace
 

Chapter in Book/ Report/ Conference proceeding

Abstract 

Monocular 3D Human Pose Estimation (3DHPE) remains challenging due to inherent depth ambiguities and occlusions. Existing video-based approaches often incur high computational costs by alternating spatiotemporal processing, while current diffusion-based methods typically rely on generic denoisers with naive input concatenation. To address these limitations, we propose a specialized diffusion framework incorporating Pose Patchification (PoPatch) and Adaptive Pose Modulation (AdaPoseMod). PoPatch extracts spatiotemporal features simultaneously, reducing computational complexity; our method reduces Multiply-Accumulate Operations (MACs) by over 200 times compared to state-of-the-art diffusion baselines. AdaPoseMod facilitates effective interaction between 2D observations and contaminated 3D poses through a dedicated modulation mechanism. Our approach achieves state-of-the-art performance on Human3.6M, MPI-INF-3DHP, and HumanEva datasets. Extensive ablation studies further validate the efficacy of our design choices in balancing efficiency and robustness.

Reference