Focusing: View-Consistent Sparse Voxels for Efficient 3D VAE Training

Xuhui Chen, Chao Long, Fei Hou, Dongbo Zhang, Shaohui Jiao, Wencheng Wang, Ying He
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:15775-15789, 2026.

Abstract

and reconstruct complex geometry at high resolution. Recent methods often convert raw meshes into signed distance fields (SDFs), but this preprocessing can be lossy, especially for open or non-watertight assets. Render-supervised VAEs such as TripoSF avoid this conversion by matching rendered depth and normal maps, yet render losses supervise only the visible geometry for each view, leaving many latent voxels weakly constrained while still incurring unnecessary decoding and attention costs. We introduce Focusing, a view-consistent sparse-voxel training scheme for efficient 3D VAEs. Given a training view, Focusing performs depth-driven voxel carving directly in the structured latent space: voxels inconsistent with the rendered depth are removed before decoding, allowing the decoder and attention layers to operate only on locally relevant geometry. This view-dependent sparsity reduces memory and computation while concentrating learning on surface regions that contribute to the render. To stabilize training across shapes and viewpoints, we further propose adaptive zooming, which adjusts camera intrinsics to keep the number of active voxels within a target range and strengths supervision for fine details. The VAE is trained with render-based depth, normal, mask, and perceptual losses, together with sparse-voxel total variation and a brief TSDF warm-up to improve convergence and suppress holes. Across standard reconstruction benchmarks, Focusing improves Chamfer Distance and F-score over strong baselines while substantially reducing video random access memory (VRAM) consumption, enabling $1024^3$-resolution VAE training with $512^3$ trunks using as little as 50 GB of VRAM. These results demonstrate that local, view-consistent sparsity is an effective path toward higher-resolution and more efficient 3D VAE training.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26cn, title = {Focusing: View-Consistent Sparse Voxels for Efficient 3{D} {VAE} Training}, author = {Chen, Xuhui and Long, Chao and Hou, Fei and Zhang, Dongbo and Jiao, Shaohui and Wang, Wencheng and He, Ying}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {15775--15789}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26cn/chen26cn.pdf}, url = {https://proceedings.mlr.press/v306/chen26cn.html}, abstract = {and reconstruct complex geometry at high resolution. Recent methods often convert raw meshes into signed distance fields (SDFs), but this preprocessing can be lossy, especially for open or non-watertight assets. Render-supervised VAEs such as TripoSF avoid this conversion by matching rendered depth and normal maps, yet render losses supervise only the visible geometry for each view, leaving many latent voxels weakly constrained while still incurring unnecessary decoding and attention costs. We introduce Focusing, a view-consistent sparse-voxel training scheme for efficient 3D VAEs. Given a training view, Focusing performs depth-driven voxel carving directly in the structured latent space: voxels inconsistent with the rendered depth are removed before decoding, allowing the decoder and attention layers to operate only on locally relevant geometry. This view-dependent sparsity reduces memory and computation while concentrating learning on surface regions that contribute to the render. To stabilize training across shapes and viewpoints, we further propose adaptive zooming, which adjusts camera intrinsics to keep the number of active voxels within a target range and strengths supervision for fine details. The VAE is trained with render-based depth, normal, mask, and perceptual losses, together with sparse-voxel total variation and a brief TSDF warm-up to improve convergence and suppress holes. Across standard reconstruction benchmarks, Focusing improves Chamfer Distance and F-score over strong baselines while substantially reducing video random access memory (VRAM) consumption, enabling $1024^3$-resolution VAE training with $512^3$ trunks using as little as 50 GB of VRAM. These results demonstrate that local, view-consistent sparsity is an effective path toward higher-resolution and more efficient 3D VAE training.} }
Endnote
%0 Conference Paper %T Focusing: View-Consistent Sparse Voxels for Efficient 3D VAE Training %A Xuhui Chen %A Chao Long %A Fei Hou %A Dongbo Zhang %A Shaohui Jiao %A Wencheng Wang %A Ying He %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26cn %I PMLR %P 15775--15789 %U https://proceedings.mlr.press/v306/chen26cn.html %V 306 %X and reconstruct complex geometry at high resolution. Recent methods often convert raw meshes into signed distance fields (SDFs), but this preprocessing can be lossy, especially for open or non-watertight assets. Render-supervised VAEs such as TripoSF avoid this conversion by matching rendered depth and normal maps, yet render losses supervise only the visible geometry for each view, leaving many latent voxels weakly constrained while still incurring unnecessary decoding and attention costs. We introduce Focusing, a view-consistent sparse-voxel training scheme for efficient 3D VAEs. Given a training view, Focusing performs depth-driven voxel carving directly in the structured latent space: voxels inconsistent with the rendered depth are removed before decoding, allowing the decoder and attention layers to operate only on locally relevant geometry. This view-dependent sparsity reduces memory and computation while concentrating learning on surface regions that contribute to the render. To stabilize training across shapes and viewpoints, we further propose adaptive zooming, which adjusts camera intrinsics to keep the number of active voxels within a target range and strengths supervision for fine details. The VAE is trained with render-based depth, normal, mask, and perceptual losses, together with sparse-voxel total variation and a brief TSDF warm-up to improve convergence and suppress holes. Across standard reconstruction benchmarks, Focusing improves Chamfer Distance and F-score over strong baselines while substantially reducing video random access memory (VRAM) consumption, enabling $1024^3$-resolution VAE training with $512^3$ trunks using as little as 50 GB of VRAM. These results demonstrate that local, view-consistent sparsity is an effective path toward higher-resolution and more efficient 3D VAE training.
APA
Chen, X., Long, C., Hou, F., Zhang, D., Jiao, S., Wang, W. & He, Y.. (2026). Focusing: View-Consistent Sparse Voxels for Efficient 3D VAE Training. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:15775-15789 Available from https://proceedings.mlr.press/v306/chen26cn.html.

Related Material