Overview: In stage I, we generate voxels with multi-view cycle conditioning and point cloud guidance. The LiDAR point cloud is first preprocessed and rotated by a learnable parameter $\hat{R}$ to a aligned orientation. Then the voxel guidance is applied to optimize the sampled latent during the denoising process. Stage II perform a 3DGS generation with multi-view conditioning, ensuring the texture fidelity of the generation. Finally, opacity-based mesh refinement is performed in Stage III: a voxel mask is obtained by thresholding Gaussian opacity and filtering SLAT features, and decoding the filtered features produces the final clean mesh.
Qualitative comparison with baseline methods in novel view synthesis. The four images in Input column represent the multi-view inputs for multi-view methods, while the image in the black box is used for the single-view method.
Qualitative comparison with baseline methods in 3D geometry. Rightmost column shows the reference LiDAR point cloud.
Qualitative comparison with baseline methods. MM-TRELLIS generates more accurate vehicle shapes and cleaner meshes by combining multi-view cycle-conditioning with LiDAR-guided optimization.
Extension to image-only settings. MM-TRELLIS can generate accurate geometry using point clouds estimated from multi-view images via VGGT, without the need for a LiDAR sensor.