We propose a novel framework, Guided Sparse Point-Voxel Diffusion (G-SPVD), for Point Cloud generation guided from a single visual input - either an image or a rough hand-drawn sketch, both from an unknown viewing angle. G-SPVD combines a Vision Transformer with a Diffusion Model that iteratively forms a noisy set of points to match the requested input. Our quantitative evaluation demonstrates that our framework achieves state-of-the-art results compared to other methods in single-image reconstruction on the ShapeNet dataset. Moreover, despite the reduced information available in sketch-based inputs, our sketch-guided model still attains competitive reconstruction metrics. We present several qualitative results for both tasks to further illustrate the effectiveness of our method. Finally, we evaluate our method on unconditional generation, demonstrating that our model can generate shapes with quality and diversity on par with the current state-of-the-art. Our code will be released upon publication.
Romanelis, I, Fotis, V, Munteanu, A & Moustakas, K 2026, G-spvd: Image and Sketch Guided Point Cloud Generation with Sparse Point-Voxel Diffusion Models. in 2025 International Conference On Visual Communications And Image Processing, Vcip. Ieee International Conference On Visual Communications And Image Processing, IEEE, 2025 Conference on Visual Communications and Image Processing-VCIP-Annual, Klagenfurt am Woerthersee, Austria, 1/12/25. https://doi.org/10.1109/VCIP67698.2025.11396860
Romanelis, I., Fotis, V., Munteanu, A., & Moustakas, K. (2026). G-spvd: Image and Sketch Guided Point Cloud Generation with Sparse Point-Voxel Diffusion Models. In 2025 International Conference On Visual Communications And Image Processing, Vcip (Ieee International Conference On Visual Communications And Image Processing). IEEE. https://doi.org/10.1109/VCIP67698.2025.11396860
@inproceedings{7972626f61c1407288291ed7c5dd3025,
title = "G-spvd: Image and Sketch Guided Point Cloud Generation with Sparse Point-Voxel Diffusion Models",
abstract = "We propose a novel framework, Guided Sparse Point-Voxel Diffusion (G-SPVD), for Point Cloud generation guided from a single visual input - either an image or a rough hand-drawn sketch, both from an unknown viewing angle. G-SPVD combines a Vision Transformer with a Diffusion Model that iteratively forms a noisy set of points to match the requested input. Our quantitative evaluation demonstrates that our framework achieves state-of-the-art results compared to other methods in single-image reconstruction on the ShapeNet dataset. Moreover, despite the reduced information available in sketch-based inputs, our sketch-guided model still attains competitive reconstruction metrics. We present several qualitative results for both tasks to further illustrate the effectiveness of our method. Finally, we evaluate our method on unconditional generation, demonstrating that our model can generate shapes with quality and diversity on par with the current state-of-the-art. Our code will be released upon publication.",
keywords = "Deep Learning, Diffusion, Generation, Image-to-PointCloud, Point Cloud",
author = "Ioannis Romanelis and Vlassis Fotis and Adrian Munteanu and Konstantinos Moustakas",
year = "2026",
month = feb,
doi = "10.1109/VCIP67698.2025.11396860",
language = "English",
isbn = "979-8-3315-6868-9",
series = "Ieee International Conference On Visual Communications And Image Processing",
publisher = "IEEE",
booktitle = "2025 International Conference On Visual Communications And Image Processing, Vcip",
note = "2025 Conference on Visual Communications and Image Processing-VCIP-Annual ; Conference date: 01-12-2025 Through 04-12-2025",
}