Tex-Zero trains 3D texture generation without any 3D assets

Native 3D texture generation paints colors directly in 3D space for a given geometry, guided by multi-view reference images. The common belief, as the paper describes it, is that training such models requires large-scale, high-quality real 3D asset data, and getting that data has long been a hard problem.
The paper proposes Tex-Zero, a framework meant to show that a high-fidelity native 3D texture generator can be trained without 3D assets at all. Its key observation is that only high-quality, fine-grained color information is essential for 3D texture training. The geometric information matters less and can be constructed by hand rather than taken from real 3D assets. That makes it possible to turn abundant, high-quality 2D images into training samples.
The conversion works in two steps. Each 2D image is represented as a plane in 3D space. Then patch-wise random rotations and aggregation are applied to build complex geometric structures out of those planes.
On this constructed image data the authors train the Tex-Zero VAE. They say it can reconstruct real 3D assets with high quality even though it never saw one during training. On top of that VAE they train the Tex-Zero DiT, again exclusively on the constructed image data. For the DiT, the conditioning 2D multi-view images are also transformed into planes in 3D space and encoded by the Tex-Zero VAE, which the authors say reduces the representation gap and improves generation quality.
The authors report that extensive experiments show Tex-Zero generates high-fidelity 3D textures with fine-grained details using only images as training data. They describe this as a promising perspective on the data paradigm for scaling 3D texture generation.
Key facts
- Tex-Zero is a native 3D texture generation framework trained without any real 3D assets.
- Core idea: only fine-grained color information is essential; geometry is less critical and can be constructed manually.
- Each 2D image becomes a plane in 3D space, and patch-wise random rotations and aggregation build complex geometric structures.
- The Tex-Zero VAE reconstructs real 3D assets with high quality despite never observing them in training.
- The Tex-Zero DiT is also trained only on the constructed image data, with conditioning multi-view images encoded by the same VAE.
Why it matters
The usual assumption is that native 3D texture models need large collections of high-quality real 3D assets, which are hard to get. This paper argues the bottleneck can be sidestepped: abundant 2D images, converted into synthetic 3D samples, can serve as the only training data. If that holds, the data paradigm for scaling 3D texture generation shifts from scarce 3D assets to plentiful images.
Who it affects
Researchers and teams working on 3D texture generation are the direct audience, particularly anyone limited by access to real 3D asset data. The paper frames the result as relevant to how such models are scaled.
How to use it
The text describes the recipe rather than a product. Represent each high-quality 2D image as a plane in 3D space, apply patch-wise random rotations and aggregation to build complex geometry, train a VAE on the result, then train a DiT on the same constructed data, encoding the conditioning multi-view images as planes through that VAE. No code or model release is mentioned in the source.
How solid is it
The claims come from the authors' own description of their work. They say extensive experiments show high-fidelity textures with fine-grained details from image-only training, and that the VAE reconstructs real 3D assets it never saw. The source text gives no quantitative results, metrics, benchmark names or baseline comparisons, so the size of the result cannot be judged from it. It also gives no quantified comparison against models trained on real 3D assets.
Risks and caveats
The headline claim rests on qualitative statements; no numbers, model sizes, dataset sizes or training compute are stated. The key observation that geometry is less critical for texture training is the authors' own finding, and how far it generalizes is not addressed in the source. Treat it as a promising direction, which is how the authors themselves put it, rather than a settled replacement for real 3D data.