main branch uses a TRELLIS.2-based implementation with pixel back-projection conditioning and exports GLB assets with geometry and PBR textures. The hosted steps below use this site's settings; the local sections follow the upstream repository and require your own GPU environment.Generate your first GLB online
Pixal3D.ai is an independent hosted service running models through Fal, with no affiliation to the research team. Generation requires sign-in. New accounts receive 220 one-time credits valid for 30 days, enough for one default Pixal3D job. Used or expired credits do not reset.
Sign in and upload one image
Use an image you have permission to process, with the whole subject visible and a clear silhouette.
Confirm the model and credit cost
Choose Pixal3D, 1024 resolution, and the Balanced preset (2048 textures) for a 220-credit default job. Check the amount beside Generate after changing settings.
Generate, inspect, and download
Wait for success, download the textured GLB, and check hidden sides, thin structures, and destination import. Manual cleanup may still be needed.
Open the image-to-3D generator · Inspect a GLB for free · Compare Pixal3D and Trellis 2
Official local single-image defaults
What Pixal3D actually is
Pixal3D lifts image features into 3D with explicit pixel back-projection. That direct 2D-to-3D correspondence is the project's defining idea. The official repository separates two implementations: main is the current, improved TRELLIS.2-based version; paper is the Direct3D-S2-based version used for the published paper results.
Do not mix the two branches
paper branch. Installation commands and current output behavior describe main. Treating them as one identical pipeline produces misleading comparisons and broken setup advice.Before choosing local or hosted
| Question | Local official workflow | Hosted Pixal3D.ai workflow |
|---|---|---|
| Setup | Linux, CUDA toolchain, compiled dependencies | Browser upload and managed runtime |
| Hardware | Start from TRELLIS.2 official 24GB NVIDIA baseline | No local GPU required |
| Control | Full repository and local files | Managed settings and account task history |
| Best for | Research, reproducible local pipelines | Evaluating an image before local setup |
Microsoft's official TRELLIS.2 instructions currently say the code is tested on Linux and requires an NVIDIA GPU with at least 24GB memory. Pixal3D adds an on-demand --low_vram mode, but its README does not publish a guaranteed 6GB or 12GB minimum. Measure peak memory instead of treating a community port as an upstream guarantee.
Official installation path
Prepare TRELLIS.2 first
Use its official Linux, CUDA 12.4, Conda, and dependency instructions. Confirm its example runs before adding Pixal3D.
Clone Pixal3D and install its requirements
Keep the environment isolated so NATTEN, CUDA, and renderer versions cannot collide with another ML project.
Compile NATTEN for your GPU architecture
Replace the placeholders with the CUDA architecture and a bounded worker count suitable for the machine.
Run one known sample before personal images
This separates environment failures from image-quality failures.
git clone --recursive https://github.com/microsoft/TRELLIS.2.git
# Follow TRELLIS.2/setup.sh from the official repository, then:
git clone https://github.com/TencentARC/Pixal3D.git
cd Pixal3D
pip install -r requirements.txt
NATTEN_CUDA_ARCH="YOUR_ARCH" NATTEN_N_WORKERS=4 \
pip install natten==0.21.0 --no-build-isolation
pip install https://github.com/LDYang694/Storages/releases/download/20260430/utils3d-0.0.2-py3-none-any.whl
python inference.py --image assets/images/0_img.png --output ./output.glbUse low-VRAM mode deliberately
python inference.py --image input.png --output output.glb --low_vram loads models on demand and defaults to resolution 1024 instead of 1536. It lowers peak memory; it is not a promise that every small GPU will work.Input image checklist
Keep the whole subject visible
Avoid cutting off legs, handles, antennas, or the base of the object.
Use a clean silhouette
Separate the subject from a busy or similarly colored background.
Prefer even lighting
Hard shadows and reflections can be mistaken for geometry or material changes.
Show depth
A three-quarter view usually reveals more shape than a perfectly flat front view.
Expect hidden-side uncertainty
A single image cannot prove the unseen back. Inspect it instead of assuming reconstruction accuracy.
Resolution and failure triage
| Symptom | Likely boundary | First bounded action |
|---|---|---|
| CUDA out of memory | Resolution or simultaneous model residency | Use --low_vram, then try --resolution 1024 |
| NATTEN build failure | Wrong CUDA architecture/toolchain | Verify CUDA_HOME and rebuild for the installed GPU |
| Flash Attention unavailable | Unsupported build or GPU | Use ATTN_BACKEND=sdpa as documented upstream |
| Weak hidden side | Single-view ambiguity | Use a clearer three-quarter source or a verified multi-view workflow |
| Soft material detail | Input/texture resolution | Improve the source before raising output resolution |
Pixal3D GGUF: can you run the weights quantized?
There is no official GGUF build to download. The hosted workspace accepts an image, not GGUF weights, local models, or workflow imports, and the upstream project publishes its checkpoints as .safetensors files in the official model file listing. As of September 28, 2026, neither that listing nor the official repository README documents a supported GGUF workflow. Any GGUF file for Pixal3D is therefore a third-party conversion that needs its own compatibility, license, and output checks, and it is not a hosted option here.
Pixal3D multiview: multi-view inference in the official repository
Multi-view is a real upstream feature, and it is separate from anything this site runs. The repository released multi-view inference code in September 2026, and inference_mv.py conditions the cascade on several views of the same object at once instead of on one image.
- Input: a directory of views plus a
transforms.jsondescribing each view's camera — a 4×4 camera-to-world matrix and the horizontal FOV in radians, in the same Blender/NeRF convention the upstream training renders use. The world is Z-up. - Weights: a separate multi-view set (
ckpts/*_mv, selected bypipeline_mv.json) from the same official model repository. - Flags:
--views_dirselects the views,--num_views Nrestricts the run to the first N of them, and--low_vram,--resolution, andATTN_BACKENDbehave exactly as in the single-image path. - Framing: views are never cropped or rescaled, the first frame is the canonical front view, and every other view is placed relative to it.
- Masks: a view with an alpha channel uses it as the object mask, and a view without one is segmented automatically.
None of that is available in the hosted workspace, which conditions on one source image per job: multi-view stays a local repository workflow that needs your own camera metadata.
Validate the exported GLB
Rotate through 360 degrees
Check the hidden side, bottom, and thin structures.
Inspect topology
Look for holes, disconnected islands, self-intersections, and excessive triangle count.
Inspect PBR channels
Check base color, roughness, metallic, opacity, seams, and color-space settings.
Confirm scale and origin
Normalize units, pivot, orientation, and bounds for the target engine.
Test the destination
Open the final GLB in the actual web viewer, Blender, Unity, or Unreal path.
Where to go next
This page covers the whole project. For the shorter guides: the 3D AI comparison picks a lane, the AI 3D workflow walks the hosted steps, and the 3D pixel guide covers what one source image can and cannot carry into the GLB.
