Preferred Networks' T3lescope Rebuilds 3D Scenes From Rooms to City Blocks

Preferred Networks' T3lescope reconstructs 3D meshes from multi-view images in a single pass, scaling from desks to city blocks without per-scene fitting.

·
·
·
Preferred Networks' T3lescope Rebuilds 3D Scenes From Rooms to City BlocksPRO
  • T3lescope reconstructs 3D meshes from multi-view images with no per-scene optimization, from objects to city blocks.
  • A single fixed-resolution generator is applied in an inference-time coarse-to-fine cascade with shared weights.
  • Finer cascade levels refine parent geometry via SDEdit-style noised initialization instead of starting from pure noise.
  • Image conditioning uses a tile pyramid so each cell attends to features matching its voxel spacing.
  • Matches or beats per-scene optimizers on ScanNet++, Tanks and Temples, Mip-NeRF 360, and city blocks.
  • Handles glossy and transparent surfaces; code is announced but not yet released.

T3lescope reconstructs 3D scenes at selectable scales

Reconstructing a clean 3D mesh from photographs usually requires either fitting a representation to each capture or running a learned model within a fixed spatial volume. Preferred Networks’ T3lescope uses one trained generator to process posed multi-view images, whose camera positions and orientations are known, and produce meshes ranging from individual objects to city blocks. Developers can choose the reconstruction depth and spatial resolution during inference without fitting a new model for every scene.

The paper reports stronger results than existing feed-forward and generative baselines on its evaluation suite. T3lescope also matches or exceeds several per-scene optimization methods while recovering fine structures, reflective materials, and transparent surfaces across indoor and outdoor captures.

Fixed grids cap the field of view

Methods such as NeRF and Gaussian Splatting optimize a separate scene representation from each set of images. Their photometric objectives depend on consistent observations across views, which can make sparse coverage, reflections, and transparency difficult to reconstruct. Per-scene fitting also adds training time before the result can be rendered or converted into geometry.

Feed-forward reconstructors use patterns learned from training data to infer missing geometry in a single inference process. Most operate at one resolution inside a limited volume. Increasing that volume reduces spatial detail, while processing a large scene as independent tiles can create gaps and misaligned surfaces along tile boundaries.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads