Pixal3D: 3D Models That Follow the Input Image

Keep the silhouette, proportions and visible details of your reference image closer to the generated 3D asset. Pixal3D links image pixels directly to the geometry it builds.

Pixal3D reference images and generated geometry for characters, creatures and architecture
The official gallery pairs reference images with generated geometry, revealing fine shapes beyond the color texture. Pixal3D

Pixal3D is a SIGGRAPH 2026 project from Tsinghua University, Tencent ARC and Victoria University of Wellington. Its current implementation uses TRELLIS.2, with single-image inference, PBR textures and a multi-view path released in September 2026.

Why pixel alignment matters

Many image-to-3D models generate in a canonical object space and use attention to inject image features. Pixal3D instead back-projects features into 3D so that a surface location has an explicit connection to pixels in the input view.

That connection targets faithful silhouettes and local details, not just a similar-looking object. The official examples include intricate creatures, mechanical parts and architecture; inspect geometry as well as textured views when judging a result.

Back-projected image features condition spatially aligned generation, linking the visible reference to the reconstructed shape.
Back-projected image features condition spatially aligned generation, linking the visible reference to the reconstructed shape. Pixal3D

Single-image and multi-view inputs

Single-image inference takes an object image and exports a GLB. Multi-view inference combines several views of the same object using separate multi-view weights. It needs a directory of images plus transforms.json describing their cameras.

Keep image framing consistent with the camera data: the multi-view path does not crop or rescale the views. The first frame should be the canonical front view. The supplied four-view example assumes 90° intervals, zero elevation and a 20° field of view, not arbitrary photographs.

Generate and export a GLB

Set up the TRELLIS.2 environment, then install the extra Pixal3D dependencies, including NATTEN and the specified utils3d wheel. Use the official Hugging Face demo if you want to try an image without local installation.

The commands below run the single-image and multi-view paths. Standard inference defaults to resolution 1536; --low_vram loads models on demand and defaults to 1024. --resolution overrides that choice. ATTN_BACKEND=sdpa provides a fallback when FlashAttention is unavailable.

python inference.py \
  --image assets/images/0_img.png \
  --output output.glb --low_vram

python inference_mv.py \
  --views_dir assets/mv_images/example \
  --output output_mv.glb --low_vram

Use Pixal3D in ComfyUI

ComfyUI natively supports Pixal3D and TRELLIS.2 in a shared image-to-model template. Pixal3D is the default; the switch selects TRELLIS.2. The linked workflow guide covers model folders, mesh cleanup, UV unwrapping, PBR baking and export.

For reproducible paper experiments, use the original paper version based on Direct3D-S2. For the newer generation workflow, use the TRELLIS.2-based implementation. They are different versions of Pixal3D, so avoid mixing their environments and weights.

Which Pixal3D workflow should you use?

WorkflowInput or backboneChoose it for
Current single-imageOne image; TRELLIS.2A detailed asset with geometry and PBR textures
Current multi-viewImages plus camera transforms; multi-view weightsCombining known viewpoints of the same object
Paper versionDirect3D-S2Reproducing the published paper experiments
ComfyUI templateNode-based single-image workflowMesh processing, material baking and export

Frequently asked questions

Can I use multiple images without camera parameters?

The local multi-view path expects transforms.json. Reuse the supplied camera file only when your views match its four-view orbit; otherwise provide matching camera transforms and field of view.

What does low-VRAM mode change?

It loads models as needed and changes the default inference resolution from 1536 to 1024. You can set --resolution explicitly, but higher resolutions still require more memory.

Is a generated GLB ready for a game engine?

GLB is a convenient asset format, but check polygon count, UVs, material appearance and geometry before integration. The ComfyUI workflow offers mesh simplification and material-baking steps.

Official sources

Generate a model without a local setup

Use the browser-based image-to-3D generator when you want a single asset without managing a local workflow.

Generate a 3D Asset