Multi-Agent View Input Tokenization
A VGGT-based geometric Transformer embeds ego and collaborator images with camera, time and agent tokens. Cross-agent calibration is not required as a model input.
NeurIPS 2026
Hong Kong JC STEM Lab of Smart City, City University of Hong Kong
Feedforward dynamic scene reconstruction from uncalibrated collaborative driving views. Recorded 3D Gaussian visualization on V2X-Real.
We present FRUC, a feedforward 3D Gaussian Splatting framework for dynamic scene reconstruction from uncalibrated collaborative driving views. Existing multi-agent reconstruction frameworks are often hindered by rigid prerequisites, demanding precise spatial calibration and slow per-scene optimization. In this paper, we rethink this task by conceptualizing a distributed multi-vehicle network as a spatio-temporally unstructured ego-centric multi-camera system, where the core challenge lies in enhancing ego-centric occluded geometry through collaboration without degrading the ego’s accurately observed visible geometry, while preserving reconstruction efficiency. For efficient reconstruction, FRUC is built upon a visual grounded geometric Transformer backbone to enable one-shot, calibration-free inference from a flexible number of multi-vehicle views. To achieve non-destructive geometric supplementation under uncalibrated cross-agent misalignment, FRUC first introduces an ego-centric causal occlusion field that explicitly derives occlusion evolution as latent priors by modeling agent-wise spatio-temporal correlations. Guided by these occlusion priors, it further formulates cross-agent integration as a deterministic residual denoising process via zero-initialized injection, turning challenging cross-agent fusion into bounded residual learning for robust collaborative blind-spot completion. Through extensive evaluations on real-world V2X-Real and UrbanIng-V2X datasets, FRUC is shown to be a new state-of-the-art for the scene reconstruction of dynamic collaborative driving environments, significantly outperforming existing methods in both rendering quality and efficiency. Code is available at https://github.com/yihangtao/FRUC.
V2X-Real input views and recorded reconstruction outputs. Use the shared timeline to compare scene geometry, occlusion priors and what-if analysis.
The Gaussian representation separates static background from dynamic foreground, making both layers available for rendering and editing.
Within-agent causal attention tracks latent motion. A learned residual expands the dynamic footprint into a structured occlusion prior.
Reconstructed geometry supports camera shifts and foreground editing. These clips illustrate the recorded reconstruction and rendering outputs.
FRUC uses three components to reconstruct a 4D scene in a single forward pass: multi-agent view input tokenization, an ego-centric causal occlusion field, and cross-agent latent residual denoising.
A VGGT-based geometric Transformer embeds ego and collaborator images with camera, time and agent tokens. Cross-agent calibration is not required as a model input.
Agent-wise masked attention links a view to its own history. Motion evolution tokens guide a residual expansion of the dynamic footprint into an occlusion uncertainty map.
CALRD conditions mixed features on the occlusion prior and clean ego-only features. Zero-initialized convolutions learn a controlled correction before Gaussian decoding.
Stage I: Single-Agent Pre-training. Establish ego-centric reconstruction and dynamic priors with a frozen VGGT backbone. Stage II: Cross-Agent Adaptation. Train the new FRUC modules and rendering decoders on multi-agent sequences, guided by photometric reconstruction and auxiliary losses.
FRUC separates static and dynamic Gaussians. Removing foreground Gaussians reveals the static background completed from collaborative context, supporting scene editing and blind-spot recovery.
Original paper example: ego-only and cooperative background rendering, dynamic probability and inferred occlusion uncertainty.
On V2X-Real, FRUC improves rendering quality and blind-spot completion over the compared methods. The paper also studies generalization to UrbanIng-V2X.
V2X-Real · multi-frame
V2X-Real · multi-frame
V2X-Real · multi-frame
Reported in the paper
Paper Table 1. Two timestamps of ego/collaborator views reconstruct an intermediate ego view. Timing is the reported experimental value; hardware and protocol affect runtime. See the full benchmark.
The paper reports a blind-spot NIQE of 3.949 and dynamic-only PSNR of 25.93 dB for FRUC on V2X-Real, complementing the full-image benchmark.
UrbanIng-V2X experiments compare direct transfer from V2X-Real with target-dataset training. Both settings test collaborative reconstruction beyond the primary benchmark.
Training, inference and V2X-Real preparation code.
Pretrained FRUC weights are available on Hugging Face.
@article{tao2026fruc,
title={{FRUC}: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views},
author={Yihang Tao and Yu Guo and Zhengru Fang and Haonan An and Yuguang Fang},
journal={arXiv preprint arXiv:2605.29997},
year={2026},
url={https://arxiv.org/abs/2605.29997},
}