Hi3DGen: 3D Geometry from Images
Hi3DGen: 3D Geometry from Images
Abstract 1. Introduction
With the growing demand for high-fidelity 3D models from With the rapid advancement of computer vision and graph-
2D images, existing methods still face significant challenges ics technologies, the task of generating 3D models from 2D
in accurately reproducing fine-grained geometric details images has garnered significant attention in both academic
due to limitations in domain gaps and inherent ambigu- and industrial domains. Despite significant advancements
ities in RGB images. To address these issues, we pro- in recent years, existing methods remain inadequate in gen-
pose Hi3DGen, a novel framework for generating high- erating 3D models that sufficiently reflect the geometric de-
fidelity 3D geometry from images via normal bridging. tails present in the input images, especially when dealing
Hi3DGen consists of three key components: (1) an image- with real-world input images, which typically exhibit com-
to-normal estimator that decouples the low-high frequency plex and rich geometric characteristics. Nevertheless, the
image pattern with noise injection and dual-stream training ability to faithfully reproduce these geometric details in 3D
to achieve generalizable, stable, and sharp estimation; (2) generations is of paramount importance, as it directly influ-
a normal-to-geometry learning approach that uses normal- ences the models’ realism, precision, and overall applica-
regularized latent diffusion learning to enhance 3D geome- bility in practical scenarios.
try generation fidelity; and (3) a 3D data synthesis pipeline
Current state-of-the-art techniques for 3D generation
that constructs a high-quality dataset to support training.
from 2D images often rely on deep learning models to learn
Extensive experiments demonstrate the effectiveness and su-
the direct mapping from the 2D RGB image to the 3D ge-
periority of our framework in generating rich geometric de-
ometry. While these methods have shown promising re-
tails, outperforming state-of-the-art methods in terms of fi-
sults [37, 70, 84], their ability in producing fine-grained ge-
delity. Our work provides a new direction for high-fidelity
ometric details is inherently limited by several key factors.
3D geometry generation from images by leveraging normal
First, the scarcity of high-quality 3D training data restricts
maps as an intermediate representation.
the model’s ability to learn detailed geometric features. Sec-
ond, there exists a significant domain gap between the train-
* Equal Contribution. ing images (often rendered from synthetic 3D meshes) and
† Corresponding author: hanxiaoguang@[Link]. test images of various possible styles, leading to suboptimal
1
performance in practical applications. Third, the inherent 2. Related Work
ambiguity in RGB images, caused by lighting, shading, or
complex object textures, further complicates the extraction Datasets for 3D Generation Early 3D datasets typi-
of fine-grained geometric information. cally encompass small-scale objects from a limited category
range [8, 11, 73]. To address this limitation, researchers
To address these limitations, we propose to leverage nor-
endeavor to expand 3D data repositories through scanning
mal maps as an intermediate representation to bridge the
or multi-view photography [16, 31, 51, 62, 71]. This ap-
mapping from 2D RGB images to 3D geometry. Normal
proach leads to the creation of large-scale datasets such
maps, which encode surface orientation information, offer
as MVImgNet [72, 81]. However, the quality of the con-
several advantages for this task. First, by introducing strong
structed data often falls short of the requirements for di-
2D priors to process RGB images into normal maps, we can
rect application in 3D generation tasks. Recently, larger-
effectively alleviate the domain gap between synthetic train-
scale datasets have been constructed by aggregating avail-
ing data and real-world applications, which eases the 2D-
able human-created 3D assets from a wide range of online
to-3D mapping learning. Second, normal maps, as a 2.5D
sources [13, 14]. However, among the 10 million 3D assets
representation, provide clearer geometric cues compared to
in Objaverse-XL [14], 5.5 million are from GitHub [22],
RGB images, thereby having the potential of guiding the
raising license concerns and high quality variety necessi-
geometry learning more effectively, especially in producing
tating costly data cleaning, and another 3.5 million from
fine-grained geometric details.
Thingiverse [57] lack textures required by existing 3D gen-
In this paper, we introduce Hi3DGen, a novel frame- eration pipelines. The remaining objects, mainly from
work for high-fidelity 3D geometry generation from im- Objaverse-1.0 [13], exhibit a severe imbalance, character-
ages via normal bridging. The framework consists of ized by a scarcity of high-quality assets with complex geo-
three key components: (i) an image-to-normal estimator metric structures and rich surface details. This imbalance is
(NiRNE) that achieves generalizable, stable, and sharp nor- a common issue in datasets of human-created 3D meshes,
mal estimation through a noise-injected regressive network resulting in networks generating simplistic 3D models with
with dual-stream training to decouple the representation significant loss of detail. To address this gap, this paper ex-
learning of low- and high-frequency image patterns; (ii) a plores synthesizing 3D data with high semantic variety, ge-
normal-to-geometry learning approach (NoRLD) that em- ometric structure diversity, and surface detail richness, and
ploys normal-regularized latent diffusion learning to pro- utilizes them in the context of 3D generation as a non-trivial
vide explicit 3D geometry supervision during training, sig- complement to human-created 3D assets.
nificantly enhancing generation fidelity; and (iii) a 3D data
Normal Estimation Monocular methods can be primar-
synthesis pipeline that constructs the DetailVerse dataset,
ily divided into diffusion-based and regression-based ap-
containing high-quality synthesized 3D assets, serving as
proaches. Regression-based methods have advanced from
important complementary of humman-created ones, to sup-
early handcrafted features [27, 28] to deep learning tech-
port the training of NiRNE and NoRLD. Our framework
niques [18, 65, 85]. Recent progress includes leveraging
generates rich, fine-grained geometric details, surpassing
large-scale data [17], estimating per-pixel normal proba-
state-of-the-art (SOTA) approaches in terms of generation
bility distributions [2], adopting vision transformers [50],
fidelity, as shown in teaser figure.
and conducting inductive bias modeling [1]. Though con-
Contributions Our key contributions are as follows: ducting deterministic prediction that ensures higher stabil-
ity, regression-based methods struggle with generating fine-
• We propose Hi3DGen, the first framework that leverages grained sharp details. Diffusion-based normal estimation
normal maps as an intermediate representation to bridge has emerged with the adaptation of powerful text-to-image
the gap between 2D images and 3D geometry, address- models [47, 52, 83]. For instance, Geowizard [21] incor-
ing the limitations of existing methods in generating fine- porates a geometry switcher to handle diverse data distribu-
grained details; tions. Considering high-variance results caused by the in-
• We introduce NiRNE, which decouples the low-high fre- herent stochastic nature of diffusion processes [19], strate-
quency learning with noise-injected dual-stream training gies such as affine-invariant ensembling [21, 32] and one-
to achieve robust, stable, and sharp normal estimation step generation [77] have been explored but come with com-
from input images; putational intensity and oversmoothing issues. StableNor-
• We develop a data synthesis pipeline and construct the mal [80] improves estimation stability by reducing diffu-
DetailVerse dataset, which contains high-quality synthe- sion inference variance via a coarse-to-fine strategy, but it
sized 3D assets to support the training of our framework. remains challenged by imperfect stability. Differently, by
We will also release this dataset and hope it can inspire deeply exploring the root causes of the sharpness produced
related research; by diffusion-based methods, we novelly propose a noise-
2
Figure 1. Overview of the proposed normal-bridged 3D geometry generation method. Our Hi3DGen comprises three components: an
image-to-normal estimator, a normal-to-geometry generator, and a synthesized dataset (DetailVerse) construction pipeline.
injected regressive method to enable both sharp and stable troduces a novel method to effectively integrate normal su-
estimations, with a dual-stream training strategy to fully uti- pervision into the diffusion learning of 3D latent codes, ad-
lize training data from different domains. dressing limitations of prior work.
3
methods to encourage learning more high-frequency infor-
mation.
Dual-Stream Architecture Compared to high-frequency
features influencing the prediction sharpness, low-
frequency features, conveying more overall structure
information [9, 24], are important for the generalizability
in low-level vision tasks [39]. To decouple these two kinds
of features, we encode the input image through two inde-
pendent streams: one processes the original image without
noise injection to robustly capture low-frequency details
(clean stream), while the other processes the noise-injected
image to focus on high-frequency details (noisy stream).
The latent representations from both streams are concate-
nated in a ControlNet-style manner [83] and fed into the
decoder for final predictions, in a regression manner. This
design uses noise injection in one stream to encourage
high-frequency representation learning, and also maintains
Figure 2. Left part: Illustration of Noise-injected Regressive Nor-
another clean stream to perceive the original image for
mal Estimation; Right part: Noisy label at high-frequency regions
in real-domain data.
regression, which effectively integrates the strengths of
diffusion-based methods into a regressive method. An
illustration of the method is presented in Fig. 2(c).
ple the low- and high-frequency representation learning for
both generalizability and sharpness, with a domain-specific Domain-Specific Training To encourage the decoupled
training strategy to stimulate the decoupled learning. representation learning in two streams, we design a domain-
specific training strategy to optimize the network by deli-
Noise Injection Considering normal sharpness usually ap-
cately utilizing training data from different domains. Previ-
pears at high-frequency image regions like edges and cavi-
ous methods mix real- and synthetic-domain data in train-
ties, we begin by analyzing from the frequency domain the
ing to enhance the generalizability. However, real-domain
underlying mechanisms that enable sharp normal estimation
data, limited by the collection environment and the preci-
results of diffusion-based methods. Defining the diffusion
sion of scanners, suffer from noisy labels especially at ob-
process with a stochastic differential equation:
ject edges (see a visualized example in Fig. 2 right part),
Z t which hinder accurate learning at high-frequency details. In
xt = x0 + g(s)dwt , (1) contrast, synthetic domain data, constructed via rendering
0
from 3D ground truth, can provide precise high-frequency
where the initial state X0 evolves over time t ∈ [0, T ] labels, while it is limited by the domain gap with real im-
to become xt and wt is a Wiener process (Brownian mo- ages in application. Therefore, we first train the network us-
tion) representing injected random noise. By conducting ing real-domain data to capture low-frequency information
Fourier transformation to this process, we can obtain the for strong generalizability. In the second stage, we fine-
signal-to-noise ratio (SNR) of any frequency component ω tune the noisy stream using synthetic-domain data while
at timestep t: freezing the parameters of the other stream. This allows
the noisy stream to focus on learning high-frequency de-
|x̂0 (ω)|2 tails as a residual component of outputs by the clean stream.
SNR(ω, t) = R t , (2) The domain-specific training not only well utilizes the train-
0
|g(s)|2 ds
ing data from real and synthetic domains according to their
which is only subject to x̂0 (ω) because the power of noise strengths, but also properly encourages the optimization of
is equal over all ω. Since natural images exhibit low-pass dual streams for decoupled representation learning.
characteristics, i.e., |x̂0 (ω)|2 ∝ |ω|−α where α > 0 repre-
sents the attenuation coefficient, the high-frequency compo-
3.2. Normal-Regularized Latent Diffusion
nents in xt has a faster SNR degradation than low-frequency State-of-the-art 2D-to-3D generation methods rely on 3D
ones as the diffusion process progresses. This prompts that latent diffusion, which represents 3D geometries in a com-
the model gets a stronger supervision at high-frequency re- pact latent space so that the 2D-to-3D mapping can be
gions in xt , which encourages the model to focus more on learned more efficiently [37, 70, 74, 84]. However, these
capturing and predicting sharp details. Inspired by this, we methods suffer from easy loss of details or detail-level in-
integrate the noise injection technique into regression-based consistency with the input images (see examples in Fig. ??).
4
Figure 3. An illustration of Normal-Regularized Latent Diffusion.
5
Table 1. Comparison of 3D object dataset statistics. The numbers
X/Y in the third column means the Mean/Medium number.
Dataset Obj # Sharp Edge # Source
GSO [16] 1K 3,071 / 1,529 Scanning
Meta [56] 8K 10,603 / 6,415 Scanning
ABO [11] 8K 2,989 / 1,035 Artists
3DFuture [20] 16K 1,776 / 865 Artists
HSSD [34] 6K 5,752 / 2,111 Artists
ObjV-1.0 [13] 800K 1,520 / 452 Mixed Figure 5. Normal estimation results comparison.
ObjV-XL [14] 10.2M 1,119 / 355 Mixed
DetailVerse 700K 45,773 / 14,521 Synthesis
GPUs (80GB each) for 50k steps with a batch size of 256.
During inference, we set the CFG strength to 3.0 and use 50
SOTA 3D generator, to conduct image-to-3D synthesis. Fi- sampling steps to achieve optimal results.
nally, a rigorous data cleaning process that combines expert
Evaluation Metrics For the evaluation of image-to-
evaluation with automated assessment preserves 700k high-
normal estimation, we basically use normal angle error
quality meshes.
(NE) to measure the overall prediction accuracy, measured
Dataset Statistics We present the model number in the in degrees. We additionally use the metric Sharp Normal
dataset and the mean sharp edge number in each model in Error (SNE) following Dora [10] to give emphasis on sharp
Tab. 1 to show the scale and geometric detail richness of our edges where geometric details are most salient. For eval-
DetailVerse dataset. The sharp edge detection follows the uating normal-to-geometry conversion, we render normal
implementation in Dora-Bench [10]. The synthesized assets maps from 22 viewpoints around each object, which is used
in DetailVerse present rich surface details, as presented by for compute NE and SNE to measure the overall and de-
the examples in the blue block of Fig. 1. tailed geometry accuracy, respectively. More implementa-
tion details are included in the supplementary details.
4. Experiments Competitive methods We compare our NiRNE with
4.1. Experiment Setup SOTA normal estimators across different methodologi-
cal categories. The comparison includes regression-based
Dataset For image-to-normal training, we utilize two methods (Lotus [25] and GenPercept [77]), diffusion-based
complementary datasets. One is a diverse realistic dataset approaches (GeoWizard [21] and StableNormal [80]). Be-
following Depth-pro[4]. Another contains synthetic data sides, Hi3DGen is compared with existing SOTA 3D
consisting of 20M RGB-to-normal pairs created by ren- generation methods including open-sourced CraftsMan-
dering 40 images per asset from 500k DetailVerse assets. 1.5 [37], Hunyuan3D-2.0 [86], Trellis [74], and close-
For normal-to-geometry training, we curate a large-scale sourced Clay[84], Tripo-2.5 [61], and Dora [10]. Note that
dataset comprising 170K cleaned 3D assets from Obja- Dora has not released its testing API, so we compare with
verse [13] and 700K synthesized 3D assets from our De- Dora using the examples on its project page.
tailVerse. We render 40 images per asset following Trel-
lis [74]. For evaluation, the generalization ability of the 4.2. Image-to-Normal Estimation
image-to-normal estimator in real scenes is validated on the
the reconstruction dataset LUCES-MV [41]. All images for Quantitative Results We provide the quantitative compar-
visual comparison and user studies are collected from Hy- ison between our NiRNE on LUCES-MV and other meth-
per3D website [12], Hunyuan3D-2.0 project page [59], and ods in Tab. 2. It validates that NiRNE gets significantly su-
Dora project page [54]. perior normal estimation performance to other regression-
Implementation Details For image-to-normal, we adopt or diffusion-based methods, in both overall normal accu-
GenPercept [77] architecture for Normal Regression net- racy and sharp-region normal accuracy.
work. We initialize the encoder and decoder weights from Qualitative Results A qualitative results is presented in
the Stable Diffusion V2.1 [53], finetuned using the AdamW Fig. 5, which shows that our NiRNE achieves superior es-
optimizer with a fixed learning rate of 3×10−5 . For normal- timation performance in (i) robustness with strong gener-
to-geometry, we build upon the Trellis [74], incorporating alizability on human and object inputs; (ii) stability with
classifier-free guidance (CFG) [26] with a drop rate of 0.1 less wrong details than diffusion-based methods (see error
and AdamW [43] optimizer with a fixed learning rate of maps); and (iii) sharpness especially when compared with
1 × 10−4 . For the normal-to-geometry training stage, we regression-based methods. These results further support our
finetune the Large variant of Trellis using 8 NVIDIA A800 related claims in Sec. 3.1.
6
the 300×6 generations for visual comparison. The evalua-
tion criteria focus on the fidelity of the generated 3D geom-
etry to the input images, which is measured by the consis-
tency in both overall shape and local details. For the parts
of the input images that are not visible, we ask the evalu-
ators to exercise their judgment and imagination to assess
the plausibility of the generated results and their stylistic
consistency with the visible portions. To ensure the com-
Figure 6. Ablations on the importance of normal bridging. prehensiveness and professionalism of the user study, we
invite two groups of evaluators. The first group consist of
50 amateur 3D users, who assess 100×6 randomly sampled
results from the perspective of everyday applications, such
as 3D printing. The second group includes 10 professional
3D artists, who evaluate 20×6 results from the standpoint
of professional use, like 3D modeling and design. The re-
sults are presented in Fig. 8, which shows that our Hi3DGen
achieves the highest generation quality for both amateur
users and professional artists.
Figure 7. High-fidelity 3D results generated by our Hi3DGen. Figure 8. User study results.
Table 2. Performance comparison on image normal estimation.
We use (Diff.) and (Regr.) to indicate diffusion- and regression- 4.4. Ablation Study
based methods, respectively. Bold indicates best results.
Normal Bridge We first validate the effectiveness of using
Method NE ↓ SNE ↓
normal maps to bridge 3D generation. A direct image-to-
(Diff.) GeoWizard [21] 31.381 36.642
geometry generator based on Trellis [74] performs worse
(Diff.) StableNormal [80] 31.265 37.045
than our normal-bridged Hi3DGen, and when using the
(Regr.) Lotus [25] 53.051 52.843
same normal regularization and training data as Hi3DGen,
(Regr.) GenPercept [77] 28.050 35.289
it produces fake details (see the first two columns v.s. the
(Regr.) NiRNE (Ours) 21.837 26.628 last column of Fig. 6). We also validate the influence of
using normal conditions of different accuracy and sharp-
4.3. Normal-to-Geometry Generation ness to the final 3D generation quality. Smoother or wrong
Qualitative Results We give a qualitative comparison normal estimations by other methods lead to a performance
between the generated 3D geometries of the proposed drop, which also proves the importance of using accurate
Hi3DGen and other methods, as shown in Fig. 9. It impres- and sharp estimated normals as the bridge.
sively shows the superiority of our Hi3DGen in generating DetailVerse Data We also validate the value of the pro-
high-fidelity results with rich details that are consistent with posed DetailVerse dataset. By integrating image-normal
the input images, which are easily lost by other methods. training pairs rendered from DetailVerse data, our NiRNE
Besides, our Hi3DGen also produce robust generations with can achieve 0.4 and 1.7 improvements in NE and SNE re-
relatively smooth surface when less details presented in the spectively, as shown in the first two rows of Tab. 3. By us-
input images (e.g. the first and third example in Fig. 9). We ing additional normal-geometry training pairs from Detail-
give more generation results of our Hi3DGen in Fig. 7, with Verse, our NoRLD can achieve higher-fidelity generation
more in supplementary materials. details, as shown in the 3rd and final columns in Fig. 10.
User Study We conducted a user study to evaluate the 3D NiRNE Ablation We conduct ablative experiments to
generation results of our Hi3DGen and 5 other methods in- validate the three components in our NiRNE: the noise
cluding Hunyuan3D-2.0, Dora, Clay, Tripo-2.5, and Trellis. injection technique, the dual-stream architecture, and the
All 3D results for user study are randomly sampled from domain-specific training. Results in Tab. 3 validates the ef-
7
Figure 9. Qualitative 3D generation comparison on samples from Dora’s project page [54].
Method NE ↓ SNE ↓
8
References Vikram Voleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A
universe of 10m+ 3d objects. NeurIPS, 36, 2024. 2, 5, 6
[1] Gwangbin Bae and Andrew J. Davison. Rethinking inductive [15] Zijian Dong, Xu Chen, Jinlong Yang, Michael J Black, Ot-
biases for surface normal estimation. In CVPR, 2024. 2 mar Hilliges, and Andreas Geiger. Ag3d: Learning to gen-
[2] Gwangbin Bae, Ignas Budvytis, and Roberto Cipolla. Es- erate 3d avatars from 2d image collections. In ICCV, pages
timating and exploiting the aleatoric uncertainty in surface 14916–14927, 2023. 3
normal estimation. In ICCV, 2021. 2 [16] Laura Downs, Anthony Francis, Nate Koenig, Brandon Kin-
[3] Maciej Bala, Yin Cui, Yifan Ding, Yunhao Ge, Zekun man, Ryan Hickman, Krista Reymann, Thomas B McHugh,
Hao, Jon Hasselgren, Jacob Huffman, Jingyi Jin, JP Lewis, and Vincent Vanhoucke. Google scanned objects: A high-
Zhaoshuo Li, et al. Edify 3d: Scalable high-quality 3d asset quality dataset of 3d scanned household items. In ICRA,
generation. arXiv preprint arXiv:2411.07135, 2024. 3 pages 2553–2560, 2022. 2, 6
[4] Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain, [17] Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir
Marcel Santos, Yichao Zhou, Stephan R Richter, and Zamir. Omnidata: A scalable pipeline for making multi-
Vladlen Koltun. Depth pro: Sharp monocular metric depth in task mid-level vision datasets from 3d scans. In ICCV, pages
less than a second. arXiv preprint arXiv:2410.02073, 2024. 10786–10796, 2021. 2
6 [18] David Eigen and Rob Fergus. Predicting depth, surface nor-
[5] Baptiste Brument, Robin Bruneau, Yvain Quéau, Jean mals and semantic labels with a common multi-scale convo-
Mélou, François Bernard Lauze, Jean-Denis Durou, and Lil- lutional architecture. In ICCV, pages 2650–2658, 2015. 2
ian Calvet. Rnb-neus: Reflectance and normal-based multi- [19] Martin Nicolas Everaert, Athanasios Fitsios, Marco Bocchio,
view 3d reconstruction. In CVPR, pages 5230–5239, 2024. Sami Arpa, Sabine Süsstrunk, and Radhakrishna Achanta.
3 Exploiting the signal-leak bias in diffusion models. In
[6] Xu Cao and Takafumi Taketomi. Supernormal: Neural sur- WACV, pages 4025–4034, 2024. 2
face reconstruction via multi-view normal integration. In [20] Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang
CVPR, pages 20581–20590, 2024. 3 Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d fur-
[7] Eric R Chan, Connor Z Lin, Matthew A Chan, Koki Nagano, niture shape with texture. IJCV, 129:3313–3337, 2021. 6
Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J [21] Xiao Fu, Wei Yin, Mu Hu, Kaixuan Wang, Yuexin Ma, Ping
Guibas, Jonathan Tremblay, Sameh Khamis, et al. Effi- Tan, Shaojie Shen, Dahua Lin, and Xiaoxiao Long. Geowiz-
cient geometry-aware 3d generative adversarial networks. In ard: Unleashing the diffusion priors for 3d geometry estima-
CVPR, pages 16123–16133, 2022. 3 tion from a single image. In ECCV, pages 241–258. Springer,
[8] Angel X Chang, Thomas Funkhouser, Leonidas Guibas, 2024. 2, 6, 7
Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, [22] Inc. GitHub. Github: Where the world builds software.
Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: [Link] 2025. Accessed: 2025-03-07. 2
An information-rich 3d model repository. arXiv preprint [23] Chun Gu, Zeyu Yang, Zijie Pan, Xiatian Zhu, and Li Zhang.
arXiv:1512.03012, 2015. 2 Tetrahedron splatting for 3d generation. NeurIPS, 37:80165–
[9] Genggeng Chen, Kexin Dai, Kangzhen Yang, Tao Hu, Xi- 80190, 2025. 3
angyu Chen, Yongqing Yang, Wei Dong, Peng Wu, Yanning [24] Guang Han, Kang Wu, Fanyu Zeng, Jixin Liu, and Sam
Zhang, and Qingsen Yan. Bracketing image restoration and Kwong. Dual-stream adaptive convergent low-light im-
enhancement with high-low frequency decomposition. In age enhancement network based on frequency perception.
CVPR, pages 6097–6107, 2024. 4 IEEE Transactions on Computational Imaging, 9:1152–
[10] Rui Chen, Jianfeng Zhang, Yixun Liang, Guan Luo, Weiyu 1164, 2023. 4
Li, Jiarui Liu, Xiu Li, Xiaoxiao Long, Jiashi Feng, and [25] Jing He, Haodong Li, Wei Yin, Yixun Liang, Leheng Li,
Ping Tan. Dora: Sampling and benchmarking for 3d shape Kaiqiang Zhou, Hongbo Liu, Bingbing Liu, and Ying-
variational auto-encoders. arXiv preprint arXiv:2412.17808, Cong Chen. Lotus: Diffusion-based visual foundation
2024. 6, 1 model for high-quality dense prediction. arXiv preprint
[11] Jasmine Collins, Shubham Goel, Kenan Deng, Achlesh- arXiv:2409.18124, 2024. 6, 7
war Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, Tomas [26] Jonathan Ho and Tim Salimans. Classifier-free diffusion
F Yago Vicente, Thomas Dideriksen, Himanshu Arora, et al. guidance. In NeurIPSW, 2021. 6
Abo: Dataset and benchmarks for real-world 3d object un- [27] Derek Hoiem, Alexei A. Efros, and Martial Hebert. Auto-
derstanding. In CVPR, pages 21126–21136, 2022. 2, 6 matic photo pop-up. TOG, page 577–584, 2005. 2
[12] Deemos. Rodin gen-1: A production-level 3d generation [28] Derek Hoiem, Alexei A. Efros, and Martial Hebert. Recov-
model. [Link] 2024. 6 ering surface layout from an image. IJCV, page 151–172,
[13] Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, 2007. 2
Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana [29] Xin Huang, Ruizhi Shao, Qi Zhang, Hongwen Zhang, Ying
Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: Feng, Yebin Liu, and Qing Wang. Humannorm: Learning
A universe of annotated 3d objects. In ICCV, pages 13142– normal diffusion model for high-quality and realistic 3d hu-
13153, 2023. 2, 5, 6, 1 man generation. In CVPR, pages 4568–4577, 2024. 3
[14] Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, [30] Satoshi Ikehata. Scalable, detailed and mask-free universal
Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, photometric stereo. In CVPR, pages 13198–13207, 2023. 3
9
[31] Varun Jampani, Kevis-Kokitsi Maninis, Andreas Engelhardt, [45] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy
Arjun Karpur, Karen Truong, Kyle Sargent, Stefan Popov, Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez,
André Araujo, Ricardo Martin Brualla, Kaushal Patel, et al. Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al.
Navi: Category-agnostic image collections with high-quality Dinov2: Learning robust visual features without supervision.
3d shape and pose annotations. NeurIPS, 36:76061–76084, Transactions on Machine Learning Research Journal, pages
2023. 2 1–31, 2024. 2
[32] Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- [46] Yatian Pang, Tanghui Jia, Yujun Shi, Zhenyu Tang, Junwu
zger, Rodrigo Caye Daudt, and Konrad Schindler. Repurpos- Zhang, Xinhua Cheng, Xing Zhou, Francis EH Tay, and Li
ing diffusion-based image generators for monocular depth Yuan. Envision3d: One image to 3d with anchor views inter-
estimation. In CVPR, pages 9492–9502, 2024. 2 polation. arXiv preprint arXiv:2403.08902, 2024. 3
[33] Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, [47] William Peebles and Saining Xie. Scalable diffusion models
and George Drettakis. 3d gaussian splatting for real-time with transformers. In ICCV, pages 4195–4205, 2023. 2
radiance field rendering. TOG, 42(4):139–1, 2023. 3 [48] Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden-
[34] Mukul Khanna, Yongsen Mao, Hanxiao Jiang, Sanjay hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv
Haresh, Brennan Shacklett, Dhruv Batra, Alexander Clegg, preprint arXiv:2209.14988, 2022. 3
Eric Undersander, Angel X Chang, and Manolis Savva. [49] Lingteng Qiu, Guanying Chen, Xiaodong Gu, Qi Zuo, Mu-
Habitat synthetic scenes dataset (hssd-200): An analysis of tian Xu, Yushuang Wu, Weihao Yuan, Zilong Dong, Liefeng
3d scene scale and realism tradeoffs for objectgoal naviga- Bo, and Xiaoguang Han. Richdreamer: A generalizable
tion. In CVPR, pages 16384–16393, 2024. 6 normal-depth diffusion model for detail richness in text-to-
[35] Black Forest Labs. Flux.1 dev: Open-weight text-to-image 3d. In CVPR, pages 9914–9925, 2024. 3, 1
generation model. [Link] 2025. 5, 2 [50] Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vi-
[36] Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, sion transformers for dense prediction. ICCV, 2021. 2
Jaakko Lehtinen, and Timo Aila. Modular primitives for [51] Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler,
high-performance differentiable rendering. ACM Transac- Luca Sbordone, Patrick Labatut, and David Novotny. Com-
tions on Graphics, 39(6), 2020. 2 mon objects in 3d: Large-scale learning and evaluation of
[37] Weiyu Li, Jiarui Liu, Rui Chen, Yixun Liang, Xuelin Chen, real-life 3d category reconstruction. In ICCV, pages 10901–
Ping Tan, and Xiaoxiao Long. Craftsman: High-fidelity 10911, 2021. 2
mesh generation with 3d native generation and interactive [52] Robin Rombach, Andreas Blattmann, Dominik Lorenz,
geometry refiner. arXiv preprint arXiv:2405.14979, 2024. 1, Patrick Esser, and Björn Ommer. High-resolution image syn-
3, 4, 6 thesis with latent diffusion models. 2021. 2
[38] Yangguang Li, Zi-Xin Zou, Zexiang Liu, Dehu Wang, Yuan [53] Robin Rombach, Andreas Blattmann, Dominik Lorenz,
Liang, Zhipeng Yu, Xingchao Liu, Yuan-Chen Guo, Ding Patrick Esser, and Björn Ommer. High-resolution image syn-
Liang, Wanli Ouyang, et al. Triposg: High-fidelity 3d shape thesis with latent diffusion models. In CVPR, pages 10684–
synthesis using large-scale rectified flow models. arXiv 10695, 2022. 6
preprint arXiv:2502.06608, 2025. 3 [54] Rui Chen. Dora: Sampling and benchmarking for 3d shape
[39] Zhaowen Li, Xu Zhao, Chaoyang Zhao, Ming Tang, and Jin- variational auto-encoders. [Link]
qiao Wang. Transfering low-frequency features for domain 2024. 6, 8
adaptation. In ICME, pages 01–06, 2022. 4 [55] Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul
[40] Minghua Liu, Chong Zeng, Xinyue Wei, Ruoxi Shi, Linghao Joo. Pifuhd: Multi-level pixel-aligned implicit function for
Chen, Chao Xu, Mengqi Zhang, Zhaoning Wang, Xiaoshuai high-resolution 3d human digitization. In CVPR, pages 84–
Zhang, Isabella Liu, Hongzhi Wu, and Hao Su. Meshformer: 93, 2020. 3
High-quality mesh generation with 3d-guided reconstruction [56] Yawar Siddiqui, Tom Monnier, Filippos Kokkinos, Mahen-
model. arXiv preprint arXiv:2408.10198, 2024. 3 dra Kariya, Yanir Kleiman, Emilien Garreau, Oran Gafni,
[41] Fotios Logothetis, Ignas Budvytis, Stephan Liwicki, and Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, et al.
Roberto Cipolla. Luces-mv: A multi-view dataset for near- Meta 3d assetgen: Text-to-mesh generation with high-
field point light source photometric stereo, 2024. 6 quality geometry, texture, and pbr materials. arXiv preprint
[42] Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, arXiv:2407.02445, 2024. 6
Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, [57] Zach Smith. Thingiverse. [Link]
Marc Habermann, Christian Theobalt, et al. Wonder3d: Sin- 2025. Accessed: 2025-03-07. 2
gle image to 3d using cross-domain diffusion. In CVPR, [58] Jingxiang Sun, Cheng Peng, Ruizhi Shao, Yuan-Chen Guo,
pages 9970–9980, 2024. 3 Xiaochen Zhao, Yangguang Li, Yanpei Cao, Bo Zhang, and
[43] Ilya Loshchilov and Frank Hutter. Decoupled weight decay Yebin Liu. Dreamcraft3d++: Efficient hierarchical 3d gener-
regularization. arXiv preprint arXiv:1711.05101, 2017. 6 ation with multi-plane reconstruction model. arXiv preprint
[44] Yuanxun Lu, Jingyang Zhang, Shiwei Li, Tian Fang, David arXiv:2410.12928, 2024. 3
McKinnon, Yanghai Tsin, Long Quan, Xun Cao, and Yao [59] Tencent Hunyuan3D Team. Hunyuan3d 2.0: High-resolution
Yao. Direct2.5: Diverse text-to-3d generation via multi-view 3d asset generation. [Link]
2.5 d diffusion. In CVPR, pages 8744–8753, 2024. 3 2025. 6
10
[60] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, shapenets: A deep representation for volumetric shapes. In
Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, CVPR, pages 1912–1920, 2015. 2
Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. [74] Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng
Llama 2: Open foundation and fine-tuned chat models. arXiv Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong
preprint arXiv:2307.09288, 2023. 5, 2 Yang. Structured 3d latents for scalable and versatile 3d gen-
[61] Tripo AI. Tripo ai - create your first 3d model with text. eration. arXiv preprint arXiv:2412.01506, 2024. 3, 4, 5, 6,
[Link] 2025. 6 7, 2
[62] Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, [75] Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, and Michael J
Thanh Nguyen, and Sai-Kit Yeung. Revisiting point cloud Black. Icon: Implicit clothed humans obtained from nor-
classification: A new benchmark dataset and classification mals. In CVPR, pages 13286–13296, 2022. 3
model on real-world data. In ICCV, pages 1588–1597, 2019. [76] Yuliang Xiu, Jinlong Yang, Xu Cao, Dimitrios Tzionas, and
2 Michael J Black. Econ: Explicit clothed humans optimized
[63] Haochen Wang, Xiaodan Du, Jiahao Li, Raymond A Yeh, via normal integration. In CVPR, pages 512–523, 2023. 3
and Greg Shakhnarovich. Score jacobian chaining: Lifting [77] Guangkai Xu, Yongtao Ge, Mingyu Liu, Chengxiang Fan,
pretrained 2d diffusion models for 3d generation. In CVPR, Kangyang Xie, Zhiyue Zhao, Hao Chen, and Chunhua Shen.
pages 12619–12629, 2023. 3 What matters when repurposing diffusion models for general
[64] Jiepeng Wang, Peng Wang, Xiaoxiao Long, Christian dense perception tasks? arXiv preprint arXiv:2403.06090,
Theobalt, Taku Komura, Lingjie Liu, and Wenping Wang. 2024. 2, 6, 7, 1
Neuris: Neural reconstruction of indoor scenes using normal [78] Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang,
priors. In ECCV, pages 139–155. Springer, 2022. 3 Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3d
[65] Rui Wang, David Geraghty, Kevin Matzen, Richard Szeliski, mesh generation from a single image with sparse-view large
and Jan-Michael Frahm. Vplnet: Deep single view normal reconstruction models. arXiv preprint arXiv:2404.07191,
estimation with vanishing points and lines. In CVPR, 2020. 2024. 3
2 [79] Yuezhi Yang, Qimin Chen, Vladimir G. Kim, Siddhartha
[66] Zehan Wang, Ziang Zhang, Tianyu Pang, Chao Du, Heng- Chaudhuri, Qixing Huang, and Zhiqin Chen. Genvdm: Gen-
shuang Zhao, and Zhou Zhao. Orient anything: Learning erating vector displacement maps from a single image, 2025.
robust object orientation estimation from rendering 3d mod- 3
els. arXiv preprint arXiv:2412.18605, 2024. 5, 2 [80] Chongjie Ye, Lingteng Qiu, Xiaodong Gu, Qi Zuo,
[67] Zijie J Wang, Evan Montoya, David Munechika, Haoyang Yushuang Wu, Zilong Dong, Liefeng Bo, Yuliang Xiu, and
Yang, Benjamin Hoover, and Duen Horng Chau. Diffu- Xiaoguang Han. Stablenormal: Reducing diffusion variance
siondb: A large-scale prompt gallery dataset for text-to- for stable and sharp normal. TOG, 43(6):1–18, 2024. 2, 6, 7
image generative models. In ACL, pages 893–911, 2023. 5, [81] Xianggang Yu, Mutian Xu, Yidan Zhang, Haolin Liu,
2 Chongjie Ye, Yushuang Wu, Zizheng Yan, Chenming Zhu,
[68] Meng Wei, Qianyi Wu, Jianmin Zheng, Hamid Rezatofighi, Zhangyang Xiong, Tianyou Liang, et al. Mvimgnet: A large-
and Jianfei Cai. Normal-gs: 3d gaussian splat- scale dataset of multi-view images. In CVPR, pages 9150–
ting with normal-involved rendering. arXiv preprint 9161, 2023. 2
arXiv:2410.20593, 2024. 3 [82] Zehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sat-
[69] Kailu Wu, Fangfu Liu, Zhihan Cai, Runjie Yan, Hanyang tler, and Andreas Geiger. Monosdf: Exploring monocu-
Wang, Yating Hu, Yueqi Duan, and Kaisheng Ma. Unique3d: lar geometric cues for neural implicit surface reconstruction.
High-quality and efficient 3d mesh generation from a single NeurIPS, 35:25018–25032, 2022. 3
image. arXiv preprint arXiv:2405.20343, 2024. 3 [83] Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding
[70] Shuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng, Jingxi conditional control to text-to-image diffusion models. In
Xu, Philip Torr, Xun Cao, and Yao Yao. Direct3d: Scalable ICCV, pages 3836–3847, 2023. 2, 4, 1
image-to-3d generation via 3d latent diffusion transformer. [84] Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu,
arXiv preprint arXiv:2405.14832, 2024. 1, 3, 4 Anqi Pang, Haoran Jiang, Wei Yang, Lan Xu, and Jingyi Yu.
[71] Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Jiawei Ren, Clay: A controllable large-scale generative model for creat-
Liang Pan, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, ing high-quality 3d assets. TOG, 43(4):1–20, 2024. 1, 3, 4,
et al. Omniobject3d: Large-vocabulary 3d object dataset for 6
realistic perception, reconstruction and generation. In CVPR, [85] Zhenyu Zhang, Zhen Cui, Chunyan Xu, Yan Yan, Nicu Sebe,
pages 803–814, 2023. 2 and Jian Yang. Pattern-affinitive propagation across depth,
[72] Yushuang Wu, Luyue Shi, Haolin Liu, Hongjie Liao, surface normal and semantic segmentation. In CVPR, pages
Lingteng Qiu, Weihao Yuan, Xiaodong Gu, Zilong Dong, 4106–4115, 2019. 2
Shuguang Cui, and Xiaoguang Han. Mvimgnet2. 0: A [86] Zibo Zhao, Zeqiang Lai, Qingxiang Lin, Yunfei Zhao,
larger-scale dataset of multi-view images. TOG, 43(6):1–16, Haolin Liu, Shuhui Yang, Yifei Feng, Mingxin Yang, Sheng
2024. 2 Zhang, Xianghui Yang, et al. Hunyuan3d 2.0: Scaling diffu-
[73] Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Lin- sion models for high resolution textured 3d assets generation.
guang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d arXiv preprint arXiv:2501.12202, 2025. 6
11
[87] Xin-Yang Zheng, Hao Pan, Yu-Xiao Guo, Xin Tong, and
Yang Liu. Mvdˆ 2: Efficient multiview 3d reconstruction
for multiview diffusion. In SIGGRAPH, pages 1–11, 2024.
3
12
Hi3DGen: High-fidelity 3D Geometry Generation
from Images via Normal Bridging
Supplementary Material
6. More Details for the Method
1
Table S4. Image-to-Normal estimation evaluation on Luces-MV (SNE). Comparisons of NiNRE with SOTA photometric stereo techniques.
Bold indicates the second best results and Red indicates best results.
Method Bowl Buddha Bunny Cup Die Hippo House Owl Queen Squirrel Ave.
SDM-UniPS (K=2) 37.65 26.24 29.02 23.70 26.32 31.45 40.68 24.56 27.14 26.10 29.286
SDM-UniPS (K=4) 31.64 20.59 23.23 23.39 25.58 21.91 38.61 22.26 25.97 24.04 25.722
Ours 34.55 21.13 30.45 17.47 27.20 24.64 34.58 25.15 26.82 24.29 26.628
edges where geometric details are most salient. Specifi- Step 2: High-Quality Image Generation With our di-
cally, we compute the Sharp Normal Error (SNE) through verse text prompt collection established, the next step in-
a three-step process: Firstly, we detect salient regions in volved generating corresponding images suitable for 3D as-
the ground truth normal maps through canny. Secondly, set synthesis. The key requirements for these images were:
we dilate these masked regions to ensure complete cover- (i) high visual fidelity with rich details that accurately re-
age of edge features. Finally, we calculate the normal angle flect the textual descriptions; and (ii) specific viewpoints
error within these masked regions. For completeness and and styles that facilitate robust 3D reconstruction.
fair comparison with existing methods, we also report the We integrated the state-of-the-art Flux.1-Dev [35] as our
Normal Error (NE) across the entire normal map, measured image generator. To ensure detailed output, we filtered the
in degrees. For evaluating normal-to-geometry conversion, generated images by ranking their sharpness according to
we render normal maps from 22 fixed, evenly spaced view- the number of sharp pixels, as calculated using Canny edge
points around each object using nvdiffrast [36], which is detection, and retained only the top 50%. For each prompt,
used to compute SNE and NE. we randomly selected a seed to encourage variety, generat-
ing exactly one image per prompt.
7. More Details for the DetailVerse To mitigate geometry distortion in the resulting 3D mod-
To ensure the quality of our synthesized meshes, we im- els, we utilized OrientAnything [66], a robust object orien-
plement a rigorous multi-stage data generation and filtering tation estimation model, to measure the alignment between
pipeline that combines expert evaluation with automated as- the camera view and canonical object orientation. Images
sessment techniques. with angular deviations exceeding 60◦ were rejected to pre-
vent structural distortions and preserve geometric fidelity.
Step 1: Semantic Text Prompt Curation We initiate the
Through this filtering process, we preserved 1 million high-
3D data synthesis process with text prompts rather than
quality images for the subsequent 3D synthesis stage.
image prompts, as textual descriptions enable more pre-
cise control over semantic diversity, thereby ensuring va- Step 3: Robust Image-to-3D Synthesis We employed
riety in the resulting geometries. To collect high-quality Trellis [74], a state-of-the-art two-stage 3D generator, to
text prompts with semantic diversity, we first sourced ap- produce high-fidelity 3D objects from the prepared images.
proximately 14M raw prompts from DiffusionDB [67], cov- Given its superior performance with high-quality inputs, we
ering a wide range of topics relevant to AI generation initially generated a set of preliminary meshes.
applications. We employed a LLaMA-3-8B model [60], To ensure mesh quality, we implemented a rigorous data
fine-tuned with manually annotated examples, to categorize cleaning process combining expert evaluation with auto-
these prompts into four distinct classes: (i) Single Objects; mated assessment. We randomly sampled 10K meshes and
(ii) Multiple Objects; (iii) Scenes; and (iv) Others. Only engaged 10 trained experts to conduct triple-blind quality
prompts from classes (i) and (ii) were retained, yielding ap- assessments. The evaluation criteria primarily focused on
proximately 1M high-fidelity prompt candidates. surface quality, specifically examining whether the rendered
Next, we applied rule-based filtering to preserve geomet- normal maps contained holes or noise artifacts.
ric and semantic attributes while eliminating stylistic modi- Based on these expert annotations, we trained a quality
fiers. Empirically, we observed that input images with near- assessment network using DINOv2 [45] features. Specifi-
isometric viewpoints and CGI-rendered aesthetics signifi- cally, we extracted features from four equiangular rendered
cantly enhance the fidelity of 3D synthesis. Thus, we im- normal maps of each mesh and trained a three-layer MLP
plemented structural prompt standardization to prompting classifier for quality scoring. This trained network was then
the image generation. Specifically, we applying domain- applied to evaluate the entire dataset. Models that received
specific prompt templates to enforce explicit geometric cues positive classifications across all four views were selected
and structural clarity (e.g., “isometric perspective”, “Unreal for training our NoRLD model. Through this compre-
Engine 5 Rendering”, “4K”, “MasterPiece”). This com- hensive quality assurance process, we retained 700K high-
prehensive process yielded approximately 1.5 million well- quality object meshes to form our DetailVerse dataset. A
curated and natural prompts. data gallery is shown in Fig. S13, and better visualizations
2
are presented in the demo video.
9. More Results
More Image-to-Normal Results We compare NiRNE
with SOTA photometric stereo technique (SDM-
UniPS [30]), which works in a different setup that requires
input images under K different lightning conditions. (As
shown in Fig. S12).
More Comparisons We give more qualitative com-
parisons in Fig. S14, which shows our normal-bridged
Hi3DGen can achieve more consistent 3D detailed geome-
tries with input images than existing methods. Better visu-
alizations are presented in the demo video.
3
4
Figure S13. More DetailVerse data exhibition.
Figure S14. More 3D generation results comparison.
Hi3DGen surpasses other methods in generating fine-grained geometric details by using normal maps to bridge 2D to 3D processes, injecting strong geometric priors into the learning process. This approach contrasts with direct image-to-geometry generators, which tend to produce less precise details due to the lack of explicit intermediate representations like normal maps. The Hi3DGen framework's use of noise injection and dual-stream architectures further refines detail accuracy, significantly reducing normal and sharp normal errors compared to other techniques that do not use similar regularization and dataset advancements .
Hi3DGen leverages normal maps as an intermediate representation to input geometric information more precisely into the learning process. Normal maps provide highly annotated orientation data, which simplifies the bridging of 2D images to 3D geometry. By embedding accurate and sharp normals, Hi3DGen enhances visualization and learning fidelity, which was found to significantly boost geometry learning performance compared to methods relying directly on raw image data . The improved normal map estimates act as detailed guides, thereby supporting the generation of intricate and precise 3D details.
Publicly available large-scale 3D datasets often suffer from quality inconsistency, licensing issues, and lack of detailed textures necessary for 3D generation tasks. For example, datasets like Objaverse-XL include a high number of assets with varied quality, leading to costly data cleaning requirements . Hi3DGen addresses these challenges through its DetailVerse data synthesis pipeline, which generates high-quality, consistent 3D assets specifically catered for training. This pipeline ensures robust, domain-specific data that enhances the training of Hi3DGen's components, facilitating more controlled and high-fidelity 3D geometry generation .
Hi3DGen mitigates the domain gap by introducing strong 2D priors through normal maps, which help in bridging the gap between synthetic data and real-world applications by easing the 2D-to-3D mapping learning. The use of noise-injected dual-stream training further enhances the estimation's adaptability and fidelity . Additionally, the DetailVerse dataset augments the framework by providing diverse, high-quality synthetic assets, supporting robust training that aligns more closely with diverse real-world scenarios .
Normal maps offer several advantages in 3D geometry generation. They provide strong 2D priors that help bridge the domain gap between synthetic training data and real-world applications, facilitating the 2D-to-3D mapping learning . Additionally, as a 2.5D representation, normal maps offer clearer geometric cues than RGB images, thus guiding geometry learning more effectively and enabling the production of fine-grained geometric details . In contrast, direct image-to-geometry methods like Trellis produce less accurate results and may generate fake details when not using the same normal regularization and training data as Hi3DGen .
NiRNE achieves robust and sharp normal estimation by employing a noise-injected regressive network with dual-stream training. This approach decouples the representation learning of low- and high-frequency image patterns, enhancing stability and sharpness in the normal estimation . Furthermore, ablative experiments show that each component, including the noise injection technique and dual-stream architecture, is effective in contributing to the model's performance .
Noise injection in the Hi3DGen framework plays a crucial role in stabilizing and sharpening normal estimation. By introducing noise into the regressive network, the framework effectively decouples low-frequency components from high-frequency details, which are crucial for sharp and robust normal map estimation. This technique allows Hi3DGen to maintain accuracy under varying input conditions and contributes significantly to its high-fidelity geometry generation .
Evaluators in the study use judgment and imagination to assess the plausibility and stylistic consistency of generated results with the visible parts of input images. The study involves two evaluator groups: 50 amateur 3D users assessing results for everyday applications and 10 professional 3D artists evaluating for professional use. The outcomes demonstrate Hi3DGen's superior generation quality compared to other methods, as it achieved the highest scores in both amateur and professional evaluations, affirming its effectiveness in generating high-fidelity and consistent 3D geometry .
The DetailVerse dataset enhances the Hi3DGen framework by providing high-quality synthesized 3D assets for training, thus supporting robust normal-to-geometry learning. By integrating image-normal training pairs from DetailVerse, the NiRNE component achieves notable improvements in Normal Error (NE) and Sharp Normal Error (SNE) scores . Moreover, additional normal-geometry training pairs from this dataset allow the NoRLD component to produce high-fidelity generation details, underscoring the dataset's pivotal role in boosting generation accuracy and detail richness .
Ablation studies on Hi3DGen examined the impact of various components such as DetailVerse data (DV), Noise Injection (NI), Dual-Stream architecture (DS), and Domain-Specific Training (DST). These studies revealed that each component significantly improves performance, as reflected in normalized error metrics (NE and SNE). For instance, removing DetailVerse data led to a decrease in generation fidelity, while omitting noise injection or dual-stream architectures resulted in higher error rates, validating the contribution of each feature to Hi3DGen's robust performance .