TRELLIS.2: Native and Compact Structured Latents for 3D Generation
A voxel grid that records where a mesh's surface crosses each cell, so open sheets, sealed interiors, and edges where three or more panels meet all fit.
TRELLIS.2 builds that grid straight from the triangles, stores six material numbers per occupied cell, compresses a 1024³ asset to about 9,600 latent vectors, and generates those vectors from one image with three 1.3-billion-parameter transformers.
Explaining the paperNative and Compact Structured Latents for 3D GenerationMost 3D generators describe a shape as a field that is negative inside and positive outside. A leaf has no inside, the seats of a car sit inside a closed body, and three panels can meet along one edge; the field format loses all three before training starts.
A 3D asset almost always ships as a triangle mesh: a list of vertex positions, a list of triangles that index into it, and texture images glued on through per-vertex UV coordinates. A network cannot read that the way it reads an image. The vertex count changes from asset to asset, the triangles come in no fixed order, and the textures live in their own 2D coordinate system. So every 3D generator converts meshes into some fixed-format array, learns to produce that array, and converts its output back to a mesh.
TRELLIS.2 (Xiang et al., Microsoft, December 2025) replaces that array. Its O-Voxel, short for omni-voxel, is a sparse voxel grid in which each occupied voxel stores a small piece of surface geometry and six material numbers. The paper trains an autoencoder that shrinks a 1024³ O-Voxel to about 9,600 vectors of 32 numbers, and three transformers of about 1.3 billion parameters each that generate those vectors from a single photo. A fully textured 1024³ asset takes about 17 seconds.
The page follows the paper's order: why signed fields lose common shapes; how its predecessor TRELLIS built latents from rendered views; how O-Voxel stores geometry with one vertex per voxel, found by a small least-squares problem; how it stores material; how a sparse autoencoder compresses it 16 times along each axis; how three flow-matching transformers generate the compressed form; and what the experiments show.
Why signed fields lose open and hollow shapes
The usual way to give a shape to a neural network is an iso-surface field. A signed distance function (SDF) returns, for any point in space, its distance to the surface, negative inside the object and positive outside. For a unit sphere, : the centre reads , the point reads , and the surface is the set where . Sample at every corner of a regular grid and the shape becomes a dense 3D array, the 3D analogue of an image, which a convolutional network or a transformer can read. Dora, Direct3D-S2, SparseFlex and TRELLIS all decode to fields of this kind (FlexiCubes, which TRELLIS uses, is also a field method).
The sign needs a consistent answer to "inside or outside?" at every point, and that answer exists only for a closed, watertight surface. Three common kinds of geometry do not have one:
- An open surface, such as a leaf, a sheet of cloth or a single-sided wall, has no interior for the sign to mark.
- A non-manifold surface has an edge shared by three or more faces (a fin welded to a hull, a T-junction of panels) or two parts touching at one point. Around that edge, "which side" has more than two answers.
- A fully enclosed interior, such as the seats inside a closed car body or the gears inside a watch case, is usually deleted on the way into a field. A messy mesh has no reliable sign, so pipelines compute one by flooding the grid from the outside; every cell the flood cannot reach, interior surfaces included, is labelled inside and drops out of the zero level set.
Before training, field-based pipelines therefore run a watertight repair: flood fill or winding-number tests that seal every gap. The repair changes the shape. An open leaf becomes a thin closed slab with two surfaces, a hollow interior becomes solid, and a narrow gap closes.
O-Voxel stores something else: the places where the surface crosses the edges of a voxel grid. Whether a triangle crosses a grid edge is a local test that does not depend on whether the surface ever closes. Drag the gap open in Figure 1 and compare the two readings of the same curve.
At gap zero both panels describe the closed loop. Once the gap opens the field has no inside left, so its zero level set is empty, while the edge crossings describe the open arc with the same accuracy as the closed one.
TRELLIS and its multiview-derived latent
The predecessor, TRELLIS (Xiang et al., 2024), introduced the structured latent (SLAT): a grid in which only the voxels the surface passes through are active, each carrying an 8-number latent vector. The positions of the active voxels give the coarse layout; the vectors give the local detail. Generation runs in two steps, first which voxels are active, then the vectors on them. TRELLIS.2 keeps this structure.
TRELLIS did not compute those vectors from the 3D asset itself. It rendered each training asset from 150 cameras, ran every render through the DINOv2 image encoder, projected each active voxel into each feature map, and averaged the features it found; a sparse transformer autoencoder then compressed those averages into the 8 numbers. The latent can only contain what some camera saw. A sealed interior is never seen, and the colours in the renders already include the lighting they were rendered under, so the latent holds no metallic or roughness values that a renderer could relight.
The word native in the title names the change: TRELLIS.2 encodes the O-Voxel, a 3D object, with a 3D autoencoder, so interior surfaces and material parameters go into the latent directly. The comparison in Table 1 of the paper is direct because both systems use about 9,600 tokens for a shape: TRELLIS at 8 numbers per token reconstructs Toys4K shapes at a normal-map PSNR of 30.29 dB, TRELLIS.2 at 32 numbers per token at 43.11 dB.
How O-Voxel stores geometry
An O-Voxel is a sparse set of feature tuples on an grid:
is the integer grid coordinate of the -th active voxel, describes the surface inside it, its material, and is the number of active voxels. Voxels the surface does not touch are not stored. The geometry part comes from Dual Contouring (2002), shown below next to Marching Cubes, the older algorithm it replaced.
Marching Cubes and Dual Contouring
Marching Cubes reads the sign of the field at the eight corners of each cell. Wherever an edge has one negative and one positive corner, it places a mesh vertex on that edge, at the point where linear interpolation of the two corner values hits zero, and a lookup table joins the vertices of each cell into triangles. Every vertex sits on a grid edge. A sharp corner of the true shape usually falls inside a cell, not on an edge, so Marching Cubes cuts it off, and finer grids only make the cut smaller in absolute size.
Dual Contouring (Ju et al., 2002) places one vertex inside each cell whose edges change sign, and joins the vertices of the four cells around every sign-changing edge into a quad. To decide where in the cell the vertex goes, it uses Hermite data: for every sign-changing edge, the exact crossing point and the surface normal there. Each pair defines a tangent plane, and the vertex goes to the point closest to all of those planes at once. Near a corner the planes of the two faces meet at the corner, so the vertex lands on it. Figure 2 runs both algorithms, in their 2D versions, on the same signed grid.
Marching squares misses the corners by a sizeable fraction of a cell at every resolution, because its vertices cannot leave the grid edges. Dual contouring stays within a few hundredths of a cell.
Both algorithms begin from the sign at every grid corner. TRELLIS.2 keeps Dual Contouring's layout (one vertex per active cell, one face per crossed edge) and drops the sign. A voxel edge counts as crossed when a mesh triangle intersects it, found with an exact triangle-edge test, and the crossing point and normal come from that triangle. An open sheet crosses grid edges the same way a closed surface does. A part sealed inside a shell crosses its own edges. An edge where three panels meet produces three sets of crossings. No step asks which side is inside, and the paper calls the result field-free.
Placing the vertex: equation (2)
Given the crossings in one voxel, TRELLIS.2 places the dual vertex by minimizing a quadratic error function (QEF):
The first sum is the original Dual Contouring energy. Each term is the squared distance from to the plane through with normal , so can slide anywhere along a flat face at zero cost and pays only for leaving it. With two faces at an angle, the only point on both planes is the line (in 3D) or point (in 2D) where they meet, which is the corner.
The other two terms are new. is the distance from to the -th boundary edge of the mesh passing through the voxel, a mesh edge that belongs to only one triangle and so marks the rim of an open surface:
with a point on the edge and its unit direction. Without it, the vertex of a voxel where a leaf ends has only the leaf's own plane to satisfy, slides back toward the crossings, and the reconstructed leaf stops short of its true edge. The last term pulls toward , the mean of the voxel's crossing points. On a flat face every plane term has the same normal, so the first sum is zero along a whole plane and the minimum is not unique; the pull toward picks one point. The released code uses and .
All three terms are quadratic in , so is a quadratic bowl and its minimum is one linear solve, , with a 3×3 matrix in 3D. A 2D example with real numbers: a voxel spanning with a 90° corner at . One face leaves through the left side at with normal , the other through the right side at with . Each plane term adds to and to :
so , 0.045 cells below the corner, pulled slightly toward . With the solve gives exactly . A point-to-point error, , would put the vertex at , half a cell below the corner. Figure 3 solves the same system live and shows as a heat map over the voxel.
Four behaviours are visible in the figure. On the corner, the plane distance puts the vertex on the corner at any angle and the point distance puts it at . At 180° the two normals are equal, has rank 1, and with a whole line of points has the same minimum error; any makes the solve unique. On the open rim, the only crossing is the one where the sheet enters the voxel. With the vertex sits on that crossing, on the voxel's side, 0.71 cells short of the true rim at the default 90°; at it is 0.06 cells from the rim. With two parallel sheets in one voxel, both are crossed on the same two sides, the error is smallest halfway between them, and the reconstruction becomes a single sheet down the middle. The paper lists that last case as a limitation (Appendix F).
What each voxel stores
After the solve, holds seven numbers:
- the dual vertex , in the voxel's local coordinates;
- three edge-intersection flags , one for each of the voxel's three edges that start at its minimum corner (along x, y and z). A cube has 12 edges, but every grid edge is shared by four voxels, so assigning each edge to one owner stores every flag exactly once;
- a splitting weight that decides how quads are cut into triangles. A quad joins four dual vertices and can be cut along either diagonal; in the released code it is cut along the diagonal whose two end voxels have the larger product of . Conversion sets every to 0.5, so the weights only carry information when the autoencoder's decoder predicts them, where the rendering loss of the second training stage can adjust them.
The splitting weights come from FlexiCubes, a field-based method. TRELLIS.2 takes that one piece and none of FlexiCubes' field machinery.
Conversion in both directions is a fixed algorithm with no optimization loop and no rendering. Mesh to O-Voxel takes a few seconds on one CPU; O-Voxel back to mesh takes tens of milliseconds. In pseudocode, following the paper's Algorithm 1 and the released code's defaults:
# mesh -> O-Voxel shape features (Algorithm 1): no field, no rendering
for tri in mesh.triangles:
for edge in grid_edges_hit_by(tri): # exact triangle-edge test
q, n = intersection_and_normal(tri, edge) # Hermite data
for vox in four_voxels_around(edge): # all four become active
vox.A += outer(n, n) # plane term of (2)
vox.b += n * dot(n, q)
vox.q_sum += q; vox.q_count += 1
vox.mark_crossed(edge) # one bit of delta, if owned
for e in mesh.boundary_edges: # open rims only
for vox in active_voxels_hit_by(e):
vox.add_line_term(e.origin, e.dir, w=1.0) # lambda_bound
for vox in active_voxels:
q_bar = vox.q_sum / vox.q_count
vox.add_point_term(q_bar, w=0.1) # lambda_reg
v = solve(vox.A, vox.b) # 3x3 linear solve
store(vox.p, v=v, delta=vox.delta, gamma=0.5) # gamma starts at 0.5and the reverse, Algorithm 2 with the split rule from the code:
# O-Voxel -> mesh (Algorithm 2)
vertex = {p: p + v[p] for p in active} # one vertex per voxel
for p in active:
for axis in (X, Y, Z):
if delta[p][axis]: # p's own edge was crossed
a, b, c, d = four_voxels_around_edge(p, axis) # cyclic order
if not all(k in vertex for k in (a, b, c, d)):
continue
if gamma[a] * gamma[c] > gamma[b] * gamma[d]:
faces += [(a, b, c), (a, c, d)] # cut along a-c
else:
faces += [(a, b, d), (d, b, c)] # cut along b-dThe conversion needs no watertight repair, but it does not keep detail smaller than a voxel. The two-sheet case of Figure 3 is one example; the limitations section comes back to it.
Six material channels per voxel
The same active voxels also store appearance, as six numbers:
These are the parameters of the standard physically based rendering (PBR) metallic-roughness material that game engines and the glTF format use: base colour (three channels), metallic , roughness , and opacity . There is no ambient-occlusion channel; the paper's name "mra" in its loss means metallic, roughness and alpha packed together.
Each parameter controls how light leaves the surface. Roughness sets how spread out reflections are: near 0 the surface is a mirror with a small, sharp highlight; near 1 it is matte with a wide, dim one. Renderers usually feed the reflection model rather than , which spaces the visible steps more evenly along the slider. Metallic is close to binary. At the material is a dielectric (plastic, wood, stone): it has a coloured diffuse body and a weak white specular reflection, about 4% at normal incidence (glTF fixes ). At it is a metal: there is no diffuse body, and the base colour tints the reflection itself, which is why gold reflections look gold. Opacity lets the format hold glass and other translucent surfaces, which shape-only formats had no channel for.
Because these are surface properties and not a picture of the surface under one light, the asset can be lit again in any scene. Figure 4 moves a light around a sphere with one stored material.
A concrete example of what the channels hold (illustrative values): a glass bottle with a steel cap. Voxels on the glass might store , , , ; voxels on the cap , , , . A camera-based latent like TRELLIS's sees the bottle as pixels lit by whatever light the renders used and has no channel to put 0.2 opacity in.
Conversion is sampling and interpolation. Mesh to O-Voxel: project the voxel's centre onto each triangle that intersects the voxel, read the texture at that point through its UV coordinates (at a mipmap level matched to the voxel size), and average the samples with weight , where is the distance from the centre to the projected point. O-Voxel to mesh: at every mesh vertex, or every texel of a newly generated texture map, interpolate the six channels trilinearly from the eight surrounding voxels.
How the SC-VAE compresses 1024³ into 9,600 tokens
A 1024³ grid has 1,073,741,824 cells. A surface fills a thin shell of them, but that shell is still far too many vectors for a transformer to generate, and transformer attention costs grow with the square of the token count. So, as latent diffusion did for images, TRELLIS.2 trains an autoencoder, the Sparse Compression VAE (SC-VAE), and generates in its latent space.
The SC-VAE downsamples 16 times along each axis. A 1024³ O-Voxel becomes a latent on a grid: 262,144 positions, of which about 9,600 are active in Table 1's Toys4K setting (3.7%). Each active position holds 32 numbers, so the whole shape is about 306,000 numbers (Table 1 lists 9.6K tokens and 306K dimensions). At 512³ the latent grid is = 32,768 positions with about 2,200 active. For comparison, SparseFlex at 1024³ downsamples 4× and uses 225,000 tokens; with attention cost growing as the square, times the attention work per layer.
The network is a U-shaped stack of sparse convolutions. Going down, it has 64 channels at full resolution, then 128, 256, 512 and 1024 at 2×, 4×, 8× and 16× downsampling, with 4, 8, 16 and 4 residual blocks at those stages, and a final linear layer that outputs a mean and a log-variance for each of the 32 latent channels (the VAE parameterization). The decoder mirrors it. The whole model has about 800M parameters, 354M in the encoder and 474M in the decoder. It has no attention layers, so it accepts any grid size: it is trained at 256³ and then 512³ and applied to 1024³ with no further training. The next two subsections cover the convolution it uses and its downsampling blocks.
Sparse convolution that keeps the surface thin
An ordinary 3D convolution computes at every cell of the grid, and almost every cell of a surface grid is empty. A sparse convolution computes only at active cells, but the standard version makes an output cell active if any input in its 3×3×3 window is active. Every layer then thickens the surface by one cell in each direction: one isolated active site becomes 27 after one layer and 125 after two. A submanifold sparse convolution (Graham and van der Maaten, 2017) makes an output active only when the centre input is active, so the active set is the same after every layer. Figure 5 stacks layers of both kinds on a thin curve.
Submanifold convolution works only at stride 1; it cannot change resolution. The SC-VAE's resolution changes happen in separate downsampling and upsampling blocks (equations (4) and (5) below). TRELLIS.2 runs all of it on its own sparse-convolution kernels written in Triton (FlexGEMM), which the paper reports as up to 2× faster than widely used sparse-convolution libraries.
Residual autoencoding: a fixed shortcut through each downsampling step
Autoencoders with high spatial compression are hard to train. The DC-AE paper (Chen et al., 2024) argued that this is an optimization problem rather than a capacity problem: good weights exist for a network of that size, but training does not reach them. Its fix, which TRELLIS.2 adapts to sparse voxels, adds a fixed, parameter-free path around every downsampling block that only rearranges numbers between space and channels, and lets the learned block predict a correction on top of it.
When downsampling by 2, each coarse voxel has up to eight child voxels. Their features, each with channels, are stacked into one vector of channels and then averaged in groups down to the coarse width , typically :
Missing children contribute zero vectors. Upsampling reverses the shape of the operation: the coarse vector is split across the eight children, and each child's share is copied within groups back up to channels:
With the encoder's numbers at the 2× stage, and . The eight children stack into 1,024 channels, and averaging groups of four consecutive channels gives 256. On the way up, 256 channels split into eight children of 32 channels each, and copying each channel four times restores 128. The averaging loses information, so the shortcut is not an exact inverse; it gives the learned block a starting point that already has every child's content in the right place. The ablation baseline replaces this shortcut with plain average pooling and nearest-neighbour upsampling, the obvious alternative, which averages the children into one vector before the channels are widened. Figure 6 shows the shortcut and the measured effect of removing it.
The ablation (Table 3) ran on curated Sketchfab assets at 256³. At 16× compression with 32 channels per token (503 tokens), removing the shortcut raises Mesh Distance from 1.032 to 1.747, a 69% increase. At 32× with 128 channels (118 tokens), it goes from 1.405 to 7.394, 5.3 times the baseline. Both settings hold a similar number of latent numbers (503×32 = 16,096 and 118×128 = 15,104), so the 32× result isolates the cost of packing them into fewer, wider tokens, and without the shortcut that cost is large.
The rest of the autoencoder
Early pruning keeps the decoder from creating empty voxels. Each 2× upsampling step could create all eight children of every coarse voxel, and four such steps would turn one latent token into fine voxels, most of them empty. Before each step the decoder instead predicts a mask of which children exist and creates only those.
The residual blocks are slimmer than usual. Each is one submanifold convolution, a LayerNorm, and a point-wise MLP that widens the channels 4× and back, in the style of ConvNeXt, replacing the usual two convolutions. With standard two-convolution blocks the ablation's Mesh Distance is 16% worse and PSNR 0.6 dB lower, at the same runtime.
Training has two stages. The first, at 256³, regresses the O-Voxel features directly:
squared error on dual vertices, binary cross-entropy on the edge flags and on the pruning masks, L1 on the material channels, and the KL term that keeps the latent close to a unit Gaussian. The second stage, at 512³, adds a rendering loss, : the decoded mesh is rendered into a mask, a depth map and a normal map, and compared with renders of the true mesh. Per-voxel errors measure each number but not whether the surface looks right, which the renders do. The appendix gives the weights:
So the mask and depth get L1 only; SSIM and LPIPS (image-similarity measures) are added on the normal map and on the rendered base colour and metallic-roughness-alpha maps. The training cameras use a shallow near plane that cuts through the surface, so the renders also show internal structure and the decoder is penalized for getting it wrong. This is also the stage in which the splitting weights learn useful values.
Geometry and material get separate SC-VAEs with separate latents, so that material can be generated after the shape, or for a shape that already exists. The material SC-VAE follows the shape SC-VAE's subdivision structure while it upsamples, so both decoders produce the same voxels. The shape VAE trains on about 473,000 assets, the material VAE on the roughly 355,000 of them whose materials use the metallic-roughness workflow.
Figure 7 places the reconstruction results of Table 1 on a quality-versus-tokens plot. TRELLIS.2 is not the method with the fewest tokens: Dora uses 2.0K and Direct3D-S2 at 512³ uses 3.0K. Its 512³ model, at 2.2K tokens, still has a higher Toys4K normal PSNR (39.54 dB) than every baseline, including SparseFlex at 225K tokens.
How TRELLIS.2 generates an asset from one image
Generation runs three models in sequence, each a diffusion transformer (DiT) trained with flow matching:
- Sparse structure. Which latent positions are active. This stage follows TRELLIS: the model generates a dense grid of 8-channel latents, and TRELLIS's structure decoder, which the released pipeline reuses unchanged, turns it into a occupancy grid.
- Geometry. A 32-number shape latent on each active position, about 9,600 tokens at 1024³. The shape SC-VAE decodes them into an O-Voxel and then a mesh.
- Material. New in TRELLIS.2. A 32-number material latent on the same positions, conditioned on the image and on the generated shape latents, which are concatenated to the input channel by channel (the input layer reads 32 + 32 numbers per token).
Each stage learns a velocity field that moves samples along straight paths between data and noise. The paper writes the path and the loss as
In this convention is a clean latent at and is Gaussian noise at . The straight path has constant velocity , which points from data toward noise, and the network learns to predict it at every point and time. Sampling starts from noise at and steps backward to , subtracting the predicted velocity. The flow-matching paper of Lipman et al. and rectified flow (Liu et al.) put noise at and data at ; the method is the same with the time axis reversed, so their velocity has the opposite sign. This page uses the TRELLIS.2 convention throughout. A training step in pseudocode:
# one training step of the geometry DiT, in the paper's time convention (9)
x0 = shape_vae.encode(o_voxel) # data latent [~9.6K tokens, 32]
eps = randn_like(x0) # noise [~9.6K tokens, 32]
t = sigmoid(1 + randn()) # logitNorm(1, 1) timestep in (0, 1)
xt = (1 - t) * x0 + t * eps # t = 0 is data, t = 1 is noise
cond = dinov3_L(image) # image tokens, read by cross-attention
if rand() < 0.1:
cond = null_cond # dropped for classifier-free guidance
loss = mse(dit(xt, t, cond), eps - x0) # target points from data to noiseTimesteps are drawn from a logit-normal distribution with mean 1, which puts more of them at the noisy end (). The image condition is dropped 10% of the time so that the model also learns an unconditional velocity, which classifier-free guidance needs at sampling time. The released pipeline samples each stage with 12 Euler steps:
# sampling, as in the released FlowEulerSampler: 12 steps from t = 1 to t = 0
x = randn(num_tokens, 32) # pure noise
ts = warp(linspace(1, 0, 13)) # the release bends the step grid
for t, t_prev in zip(ts[:-1], ts[1:]):
v = guided_velocity(dit, x, t, cond) # guidance strength 7.5 for t >= 0.6
x = x - (t - t_prev) * v # Euler step toward the data end
latent = x # SC-VAE decode -> O-Voxel -> meshIn the released configuration, guidance strength is 7.5 for the structure and geometry stages, applied only while , the noisiest part of the path; the material stage runs at strength 1.0, which is the plain conditional prediction. Figure 8 runs the three stages in order.
Each DiT has about 1.3 billion parameters: width 1536, 30 blocks, 12 attention heads of 128 dimensions, and a feed-forward width of 8192, for about 4 billion across the three. Each block has self-attention over the latent tokens, cross-attention to image features from DINOv3-L (a self-supervised vision transformer, the successor to the DINO line; TRELLIS used DINOv2), and a feed-forward layer. The timestep enters through AdaLN-single, from PixArt-α: one set of shift and scale values is computed from and shared by all blocks, instead of DiT's per-block modulation, which in PixArt-α saved about 26% of the parameters. Positions use rotary embeddings, and queries and keys are RMS-normalized before attention for stability. TRELLIS needed convolutional token packing and skip connections to handle its token counts; with the 16× latent, TRELLIS.2 drops both and uses a plain transformer.
Training used AdamW (learning rate , weight decay 0.01) on 32 H100 GPUs with batch size 256. The structure model trains with 512×512 conditioning images. The geometry and material models train first at 512³ outputs ( latents) and then at 1024³ ( latents), with the conditioning image growing to 1024 pixels.
Following one request end to end: the input is a photo of a glass with a metal rim. Stage 1 denoises a tensor of 4,096 tokens in 12 steps, and the structure decoder turns it into the active cells of a grid, cells on the glass's inner wall included. Stage 2 starts from noise of shape [active cells, 32] and runs 12 guided Euler steps; the shape SC-VAE decoder then upsamples 2× four times, pruning empty children at each step, into a 1024³ O-Voxel, and Algorithm 2 turns that into a mesh. Stage 3 denoises a second [active cells, 32] tensor, reading the stage-2 latents alongside it; the material decoder produces six channels per voxel, and trilinear interpolation writes them onto the mesh. The glass body comes out with low opacity and the rim with metallic near 1, and the asset can be relit in any renderer. The paper reports about 17 seconds for this at 1024³.
Test-time resolution scaling works because the latent is small: the geometry stage can run twice. Max-pool a generated 1024³ O-Voxel down to a sparse structure, use that as the active set, and run the geometry stage again: , so the output is a 1536³ shape, larger than any training resolution. Within the trained range, max-pooling a generated 512³ O-Voxel to a structure and regenerating at 1024³ replaces stage 1's structure with a cleaner one and corrects local errors, at the cost of an extra pass.
Stage 3 also works alone as a texturing model: convert a given mesh to O-Voxel, encode its shape, and generate PBR material for it from a reference image. The paper compares it qualitatively with Hunyuan3D-Paint, which generates multiple views and fuses them, and TEXGen, which works in UV space, and reports fewer ghosting and seam artifacts, plus textures on internal surfaces that cameras never see.
Results: reconstruction and generation
Shape reconstruction in Table 1 encodes and decodes test assets from Toys4K and from 90 recent Sketchfab models. TRELLIS.2 at 1024³ beats Dora, TRELLIS, Direct3D-S2 and SparseFlex on every quality column of both test sets (Mesh Distance and its F-score, Chamfer Distance and its F-score, normal-map PSNR and LPIPS). The 512³ model does not: on Sketchfab, SparseFlex at 1024³ beats it on the Mesh Distance F-score (0.684 vs 0.613) and on normal PSNR (32.12 vs 31.00 dB). Decoding is not always fastest either: TRELLIS decodes in 0.108 s and TRELLIS.2-1024 in 0.301 s, against 3.21 s for SparseFlex-1024 and 13.0 s for Direct3D-S2-1024.
Two of those columns look alike. The paper defines both as a symmetric average of squared nearest distances:
Mesh Distance (MD) samples 1 million points from the whole surface of a mesh, interior parts included, and measures each point's distance to the other mesh surface . Chamfer Distance (CD) renders depth maps from 100 cameras around the object, turns the pixels back into 3D points, samples 1 million of them, and measures point-to-point distances between the two clouds. The cameras see only the outer shell. Both formulas average squared nearest distances in both directions; they differ in their point sets and in whether the nearest distance is to a surface or to a point cloud. MD's points include enclosed parts, CD's do not. Figure 9 builds both point sets for a 2D car body with a seat sealed inside.
Moving the seat by 0.08 (about 4% of the body's width) multiplies Mesh Distance by about 190, and deleting it by about 4,000, while Chamfer Distance stays at 1.00× because no camera ray reaches the seat. In Table 1 the gap between TRELLIS.2 and the best baseline follows the same pattern. On Toys4K, Mesh Distance is 0.0042 for TRELLIS.2-1024 against 0.3132 for SparseFlex-1024, about 75 times lower; Chamfer Distance is 0.5660 against 0.8062, about 1.4 times lower (both ×106). The F-scores count points within a threshold of the other surface, and the comparison is on squared distance, , so (MD) and (CD) correspond to distances of about and .
Material reconstruction has no baseline, because no earlier method encodes PBR materials this way. Rendered PBR attribute maps come back at 38.89 dB PSNR (LPIPS 0.033), and fully shaded renders at 38.69 dB (LPIPS 0.026).
Image-to-3D generation was tested on 100 AI-generated images. TRELLIS.2 scores highest on all four alignment metrics in Table 2: CLIP 0.894, CLIP on normal renders 0.758, ULIP-2 0.477 and Uni3D 0.436. The margins are small; the next best CLIP score is TRELLIS at 0.876. In a user study with about 40 participants, TRELLIS.2 was chosen in 66.5% of overall-quality votes (135 of 202), with Hunyuan3D 2.1 next at 13.3%, and in 69.0% of shape-only votes on normal renders (147 of 213), with Direct3D-S2 next at 12.2%.
Speed: the introduction gives about 3 s for a fully textured 512³ asset, 17 s at 1024³ and 60 s at 1536³, on an H100. The experiments section says all runtime statistics are on an A100, so the GPU behind those three numbers is ambiguous.
Limitations
The paper's Appendix F lists three. First, detail smaller than a voxel is lost. Two parallel surfaces closer than a voxel width are crossed on the same voxel edges, the QEF puts one vertex between them (the two-sheet mode of Figure 3), and the material in that voxel is the average of both surfaces. Second, decoded meshes sometimes have small holes, which the authors attribute to the sparse decoder not guaranteeing a closed surface; standard hole filling repairs most of them. Third, O-Voxel stores geometry and material only, with no parts or semantic structure, which the authors name as future work.
The comparisons also have limits. The ablations run at 256³ on curated Sketchfab assets, not at the 1024³ resolution of the main model. The material numbers have no baseline. The generation metrics measure image-shape agreement with CLIP, ULIP-2 and Uni3D embeddings, and the preference study has about 40 participants and one prompt set.
Questions you might still have
Why not convert to an SDF and repair the mesh first?
Repair changes the shape. Flood fill or winding-number watertighting turns an open sheet into a thin closed slab, fills sealed interiors, and closes narrow gaps, so the field describes a different object. O-Voxel only needs to know where triangles cross grid edges, which works for open, non-manifold and enclosed surfaces without any repair.
How is TRELLIS.2 different from TRELLIS?
TRELLIS built each voxel latent by rendering 150 views, extracting DINOv2 features and averaging them onto the voxel, so it held only what cameras saw and no relightable material. TRELLIS.2 encodes the O-Voxel itself with a sparse convolutional VAE at 16x downsampling (TRELLIS used 4x), stores 32 numbers per token instead of 8, and adds a third generation stage for PBR material.
Is the mesh to O-Voxel conversion lossless?
No. It needs no repair, optimization or rendering, but each voxel keeps one vertex. Detail below the voxel size is lost: two surfaces closer than a voxel collapse into one, and their materials average. At 1024 cubed a voxel is about 0.1% of the bounding-box width.
Why does the dual vertex use distance to planes instead of distance to the crossing points?
Distance to the crossing points puts the vertex at their average, which cuts every corner. Distance to the tangent planes lets the vertex slide along each face for free, so the only zero-cost point is where the faces meet: the corner. For a 90-degree corner with crossings at (0, 0.2) and (1, 0.2), the plane version lands on the corner at (0.5, 0.7) and the point version half a cell below it.
Why not compress 32x and use even fewer tokens?
The ablation tried it at 256 cubed: 118 tokens of 128 channels instead of 503 tokens of 32. With the residual shortcut, Mesh Distance rose from 1.032 to 1.405; without it, to 7.394. The paper settles on 16x with 32 channels.
Is the flow-matching direction the same as in the flow-matching paper?
The method is the same and the time axis is reversed. TRELLIS.2 puts data at t=0 and noise at t=1, so its velocity target eps - x0 points from data to noise and sampling steps from t=1 down to t=0. Lipman et al. and rectified flow put noise at t=0, so their velocity has the opposite sign. The Flow Matching explainer on this site uses the original convention.
Can I use only the texturing stage?
Yes. Convert an existing mesh to O-Voxel, encode its shape with the shape SC-VAE, and run the material stage with a reference image; it generates PBR channels for every voxel, including interior surfaces no camera sees.
Footnotes & further reading
- The paper: Xiang et al., Native and Compact Structured Latents for 3D Generation (arXiv, December 2025). Project page; code at github.com/microsoft/TRELLIS.2 (QEF weights in
o-voxel/o_voxel/convert/flexible_dual_grid.py, sampler intrellis2/pipelines/samplers/flow_euler.py). - The predecessor: Xiang et al., Structured 3D Latents for Scalable and Versatile 3D Generation (TRELLIS, CVPR 2025).
- Ju, Losasso, Schaefer and Warren, Dual Contouring of Hermite Data (SIGGRAPH 2002); Shen et al., Flexible Isosurface Extraction for Gradient-Based Mesh Optimization (FlexiCubes, SIGGRAPH 2023), source of the splitting weights.
- Chen et al., Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models (DC-AE), the space-to-channel residual idea; Graham and van der Maaten, Submanifold Sparse Convolutional Networks (2017), the reference the paper cites for submanifold convolution.
- The generator: Peebles and Xie, Scalable Diffusion Models with Transformers (DiT; explainer: DiT); Lipman et al., Flow Matching for Generative Modeling (explainer: Flow Matching); AdaLN-single from PixArt-α; image features from DINOv3.
- PBR conventions ( for dielectrics, base colour as the metal's specular colour, microfacet ): the glTF 2.0 specification, Appendix B. The paper renders PBR assets with the approximate split-sum renderer from nvdiffrec.
- Runtimes of about 3 s, 17 s and 60 s at 512³, 1024³ and 1536³ are stated for an NVIDIA H100 in the paper's introduction, while Section 4 says all runtime statistics are reported on an A100, and Table 1's decoder times are labelled A100. Read the generation times as approximate.
How could this explainer be improved? Found an error, or something unclear? I read every message.