Monocular 3D Human Pose Estimation (3DHPE) remains challenging due to inherent depth ambiguities and occlusions. Existing video-based approaches often incur high computational costs by alternating spatiotemporal processing, while current diffusion-based methods typically rely on generic denoisers with naive input concatenation. To address these limitations, we propose a specialized diffusion framework incorporating Pose Patchification (PoPatch) and Adaptive Pose Modulation (AdaPoseMod). PoPatch extracts spatiotemporal features simultaneously, reducing computational complexity; our method reduces Multiply-Accumulate Operations (MACs) by over 200 times compared to state-of-the-art diffusion baselines. AdaPoseMod facilitates effective interaction between 2D observations and contaminated 3D poses through a dedicated modulation mechanism. Our approach achieves state-of-the-art performance on Human3.6M, MPI-INF-3DHP, and HumanEva datasets. Extensive ablation studies further validate the efficacy of our design choices in balancing efficiency and robustness.
Yang, S, Jansen, B, Sahli, H, Nguyen, X & Histace, A 2026, RETHINKING DIFFUSION FOR 3D HUMAN POSE ESTIMATION: SPATIOTEMPORAL PATCHIFICATION AND ADAPTIVE MODULATION. in 2026 IEEE International Conference on Image Processing (ICIP). IEEE, pp. 1-6. https://doi.org/10.1109/ICIP61757.2026.11630088
Yang, S., Jansen, B., Sahli, H., Nguyen, X., & Histace, A. (2026). RETHINKING DIFFUSION FOR 3D HUMAN POSE ESTIMATION: SPATIOTEMPORAL PATCHIFICATION AND ADAPTIVE MODULATION. In 2026 IEEE International Conference on Image Processing (ICIP) (pp. 1-6). IEEE. https://doi.org/10.1109/ICIP61757.2026.11630088
@inproceedings{e1c1d763fd994da98ce4395880b7c291,
title = "RETHINKING DIFFUSION FOR 3D HUMAN POSE ESTIMATION: SPATIOTEMPORAL PATCHIFICATION AND ADAPTIVE MODULATION",
abstract = "Monocular 3D Human Pose Estimation (3DHPE) remains challenging due to inherent depth ambiguities and occlusions. Existing video-based approaches often incur high computational costs by alternating spatiotemporal processing, while current diffusion-based methods typically rely on generic denoisers with naive input concatenation. To address these limitations, we propose a specialized diffusion framework incorporating Pose Patchification (PoPatch) and Adaptive Pose Modulation (AdaPoseMod). PoPatch extracts spatiotemporal features simultaneously, reducing computational complexity; our method reduces Multiply-Accumulate Operations (MACs) by over 200 times compared to state-of-the-art diffusion baselines. AdaPoseMod facilitates effective interaction between 2D observations and contaminated 3D poses through a dedicated modulation mechanism. Our approach achieves state-of-the-art performance on Human3.6M, MPI-INF-3DHP, and HumanEva datasets. Extensive ablation studies further validate the efficacy of our design choices in balancing efficiency and robustness.",
author = "Shuo Yang and Bart Jansen and Hichem Sahli and Xuan-son Nguyen and Aymeric Histace",
year = "2026",
doi = "10.1109/ICIP61757.2026.11630088",
language = "English",
pages = "1--6",
booktitle = "2026 IEEE International Conference on Image Processing (ICIP)",
publisher = "IEEE",
}