SILSA replaces voxel tokens with slice latents for 3D generation

SILSA replaces voxel tokens with slice latents for 3D generation

A paper on Hugging Face Papers introduces SILSA, a topology-aware framework for high-resolution 3D generation. The authors start from a problem with current methods. High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict the active structure and then synthesize local geometry. That design, they say, fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes.

SILSA represents a shape with compact sliding-window slice latents instead of voxel tokens. It uses a fixed set of overlapping slices placed along the three canonical axes. Each token summarizes a local depth window, which is meant to preserve cross-sectional continuity and support single-stage rectified-flow generation.

The system has three main parts. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder. A Volumetric Anchor Lattice coordinates the directional slice streams through a shared 3D workspace. To keep shapes structurally correct, the authors add slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices.

The abstract reports that SILSA improves PSNR by 8.7%, coverage by 5.96 absolute points, and Betti error by 9.2% over the strongest baseline. It uses 70.0% fewer tokens than the next-most compact baseline and over 98% fewer tokens than sparse or hierarchical tokenizers. The authors say this effectively reduces training memory by 40.4% and inference time by 58.5%. Qualitative results are said to show better preservation of thin structures, repeated components, and long-range connectivity.

Key facts

  • SILSA swaps voxel tokens for compact sliding-window slice latents: a fixed set of overlapping slices along the three canonical axes, each token summarizing a local depth window.
  • Generation is single-stage with rectified flow, built on a Slice VAE with a sparse volumetric decoder and a Volumetric Anchor Lattice that coordinates the slice streams.
  • Slice-level topology supervision matches persistence diagrams and aligns Betti transitions across neighboring slices to protect thin and highly connected shapes.
  • Reported against the strongest baseline: PSNR up 8.7%, coverage up 5.96 absolute points, Betti error improved by 9.2%.
  • Token use is 70.0% lower than the next-most compact baseline and over 98% lower than sparse or hierarchical tokenizers; training memory is cut by 40.4% and inference time by 58.5%.

Why it matters

Voxel latents and multi-stage pipelines are a common route to high-resolution 3D generation, but the authors argue they chop continuous surfaces into many local tokens, raise cost, and weaken topology on thin or highly connected shapes. SILSA tries to address both cost and structure at once: a much smaller token set, plus a training signal built directly on topology (persistence diagrams and Betti transitions). If the reported gains hold, shape quality and efficiency would not have to be traded against each other.

Who it affects

The work is aimed at researchers building 3D generative models, especially anyone working with thin structures, repeated components, or shapes with long-range connectivity, which the qualitative results highlight. The token and memory savings would matter to teams limited by training memory or inference time.

How to use it

No code, model release, or project page is mentioned in the source, so there is nothing to run yet. For now it is a method to read and, if the full paper gives enough detail, to reimplement. The core ideas to borrow are the three-axis overlapping slice tokens, single-stage rectified-flow generation, and slice-level topology supervision.

How solid is it

The material is an abstract, so the figures are the authors' own claims. The abstract says experiments show improved structural fidelity and substantially lower generation cost. It does not name the baselines, datasets, or resolutions, and it gives only improvements, not absolute PSNR, coverage, or Betti error values. The baseline for the 40.4% training memory and 58.5% inference time reductions is not stated. No authors or institutions are named, and whether the results are peer-reviewed is not stated.

Risks and caveats

The headline numbers cannot be judged without knowing which baselines and datasets they refer to. The 8.7% PSNR gain is a relative improvement, while the coverage gain is 5.96 absolute points, so the two should not be compared directly. The abstract does not say whether the 9.2% Betti error figure is relative or in points. The claims about thin structures, repeated components, and long-range connectivity rest on qualitative results. No hardware, model size, or token counts are given.

“Experiments show that SILSA improves structural fidelity while substantially reducing generation cost.”

— SILSA paper abstract