0% found this document useful (0 votes)
24 views20 pages

Elevate3D: Refining Low-Quality 3D Models

3d texture refinement

Uploaded by

Li Dexter Shuda
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
24 views20 pages

Elevate3D: Refining Low-Quality 3D Models

3d texture refinement

Uploaded by

Li Dexter Shuda
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Elevating 3D Models: High-!

ality Texture and Geometry Refinement


from a Low-!ality Model
NURI RYU, POSTECH, South Korea
JIYUN WON, POSTECH, South Korea
JOOEUN SON, POSTECH, South Korea
MINSU GONG, POSTECH, South Korea
JOO-HAENG LEE, Pebblous, South Korea
SUNGHYUN CHO, POSTECH, South Korea
arXiv:2507.11465v1 [[Link]] 15 Jul 2025

(b) Ours (Elevate3D) (d) Ours (Elevate3D)

Fig. 1. 3D refinement examples from (a) a degraded real-world scan [Downs et al. 2022] and (c) a state-of-the-art image-to-3D generative model [Xiang
et al. 2024]. Our method, Elevate3D, e"ectively refines both texture and geometry while preserving their alignment, as shown in (b) and (d). Inputs for the
experiment: the GSO dataset [Downs et al. 2022], and ©MasaStojanovic/pixabay.

High-quality 3D assets are essential for various applications in computer ACM Reference Format:
graphics and 3D vision but remain scarce due to signi!cant acquisition costs. Nuri Ryu, Jiyun Won, Jooeun Son, Minsu Gong, Joo-Haeng Lee, and Sunghyun
To address this shortage, we introduce Elevate3D, a novel framework that Cho. 2025. Elevating 3D Models: High-Quality Texture and Geometry Re-
transforms readily accessible low-quality 3D assets into higher quality. At !nement from a Low-Quality Model. In Special Interest Group on Computer
the core of Elevate3D is HFS-SDEdit, a specialized texture enhancement Graphics and Interactive Techniques Conference Conference Papers (SIGGRAPH
method that signi!cantly improves texture quality while preserving the Conference Papers ’25), August 10–14, 2025, Vancouver, BC, Canada. ACM,
appearance and geometry while !xing its degradations. Furthermore, Ele- New York, NY, USA, 20 pages. [Link]
vate3D operates in a view-by-view manner, alternating between texture and
how
integrate
to them
together ?

geometry re!nement. Unlike previous methods that have largely overlooked


geometry re!nement, our framework leverages geometric cues from images
re!ned with HFS-SDEdit by employing state-of-the-art monocular geome- 1 Introduction
try predictors. This approach ensures detailed and accurate geometry that
High-quality 3D models are in unprecedented demand, serving as
aligns seamlessly with the enhanced texture. Elevate3D outperforms recent
competitors by achieving state-of-the-art quality in 3D model re!nement,
essential components for various applications and as training data
e"ectively addressing the scarcity of high-quality open-source 3D assets. in computer graphics and 3D vision [He et al. 2025; Shi et al. 2023b;
Tewari et al. 2022]. Despite the growth in large-scale open-source 3D
CCS Concepts: • Computing methodologies → Computer graphics. model datasets [Deitke et al. 2023, 2022; Wu et al. 2023], high-quality
models remain scarce due to high acquisition costs. To address this,
Additional Key Words and Phrases: 3D Asset Re!nement, Di"usion models
we tackle the problem of texture and geometry re!nement of easily
accessible low-quality models, bridging the gap between abundant
Authors’ Contact Information: Nuri Ryu, POSTECH, South Korea, ryunuri@[Link]. low-quality data and the pressing need for high-quality models.
kr; Jiyun Won, POSTECH, South Korea, w1jyun@[Link]; Jooeun Son, POSTECH, Constructing high-quality 3D models from low-quality counter-
South Korea, jeson@[Link]; Minsu Gong, POSTECH, South Korea, gongms@
[Link]; Joo-Haeng Lee, Pebblous, South Korea, joohaeng@[Link]; Sunghyun parts, a process we refer to as 3D model re!nement, is of great
Cho, POSTECH, South Korea, [Link]@[Link]. importance but has received relatively less attention. Traditional
methods, such as mesh subdivision and denoising, primarily focus
on re!ning geometry by suppressing noisy structures or smoothing
surfaces using simple geometric priors. However, these methods
This work is licensed under a Creative Commons Attribution 4.0 International License. fall short of producing high-quality geometric details as they rely
SIGGRAPH Conference Papers ’25, Vancouver, BC, Canada solely on simple priors and do not address texture re!nement.
© 2025 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-1540-2/2025/08 Recently, several 3D model re!nement methods have emerged, ei-
[Link] ther as post-processing steps within 3D model generation pipelines

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
2 • Ryu et al.

or as stand-alone methods for re!ning existing low-quality mod- coarse-to-!ne nature of the di"usion model’s generative process [Ris-
els [Lee et al. 2024; Sun et al. 2024; Tang et al. 2024b; Wu et al. sanen et al. 2023], where low-frequency features are established !rst
2024a; Yang et al. 2024c; Zhang et al. 2023a, 2024]. These approaches and strongly in#uence subsequent high-frequency detail generation.
leverage generative priors, such as single-image and multi-view If these low-frequency features are constrained to match a low-
di"usion models [Ho et al. 2020; Sohl-Dickstein et al. 2015] and quality input, the !nal image inevitably inherits unwanted artifacts.
GANs [Goodfellow et al. 2014], to enhance both texture and geom- Instead, we let the di"usion model freely generate low-frequency
etry or texture alone. Speci!cally, they render multiple views of a features while applying constraints only high-frequency compo-
low-quality model, enhance each view, and update the 3D model. nents to match the input image during the early denoising steps. In
One common strategy for view enhancement is to apply di"usion- this way, the generative path is steered toward the high-quality im-
or GAN-based image enhancement models speci!cally trained to age distribution while minimal high-frequency guidance preserves
re!ne the appearance of 3D models [Wu et al. 2024a; Yang et al. crucial edges and details—su$cient to maintain the input’s unique
2024c; Zhang et al. 2024]. Another popular approach is to employ identity without embedding its artifacts.
SDEdit [Meng et al. 2022], which o"ers high #exibility, enabling After re!ning the texture at a viewpoint, we enhance the ge-
re!nement of low-quality models with various degradations [Sun ometry for that viewpoint using the re!ned texture. We employ a
et al. 2024; Tang et al. 2024b; Yang et al. 2024c; Zhang et al. 2023a]. state-of-the-art monocular normal predictor [Martin Garcia et al.
Despite the advances in 3D model re!nement, limitations in both 2025] to infer detailed surface normals aligned with the updated
texture and geometry still remain. Most recent approaches re!ne texture. The predicted normals may not perfectly match the initial
each view independently, leading to cross-view inconsistencies and 3D geometry. Thus, we employ a regularized normal integration
blurry textures. While multi-view di"usion models may improve scheme to estimate a geometry consistent with the initial 3D ge-
consistency, they are limited in image resolution due to large mem- ometry from the predicted normals. We then stitch the estimated
ory requirements, restricting re!nement quality [Long et al. 2023]. geometry with previously re!ned or unre!ned regions, maintaining
SDEdit-based approaches also su"er from a trade-o" between consistent geometry across viewpoints. This process produces a 3D
!delity to the input and the quality of re!ned 3D models. SDEdit representation faithful to the re!ned textures and supports direct
generates structured noise by interpolating the input image with texture projection onto the enhanced geometry without misalign-
Gaussian noise, and applies the di"usion model’s denoising pro- ment.
cess [Ho et al. 2020; Sohl-Dickstein et al. 2015]. This aligns the image Comprehensive experiments demonstrate that Elevate3D suc-
with the high-quality distribution learned by the di"usion model cessfully produces high-quality 3D models from low-quality ones,
while preserving key features of the input. However, this process including those generated by previous 3D generation models and
imposes a quality-!delity trade-o" based on the noise level. Stronger low-quality scanned models. To summarize, our main contributions
noise enhances conformance to the di"usion model’s distribution are as follows:
but harms !delity. In contrast, lower noise maintains !delity but
• We propose a novel 3D model re!nement framework, Ele-
o"ers only minor enhancements.
vate3D, which alternates between texture and geometry re-
Lastly, previous approaches rely solely on image-based priors
!nement in a view-by-view fashion to produce a high-quality
derived from large-scale image generative models. This exclusive
3D model with well-aligned texture and geometry
reliance on photometric constraints leaves the geometry under-
• We introduce HFS-SDEdit for texture re!nement, leveraging
constrained; even if the re!ned models appear visually appealing
high-frequency guidance to achieve high-quality and high-
from certain viewpoints, their underlying geometric structure may
!delity enhancements while mitigating the limitations of
still be of poor quality [Barron et al. 2022; Yu et al. 2022; Zhang
previous SDEdit-based methods.
et al. 2020]. To address these issues, some approaches incorporate
• We demonstrate that our framework can achieve state-of-the-
separate geometry and texture optimization stages [Sun et al. 2024;
art quality re!nement of 3D models compared with recent
Tang et al. 2024b]. However, decoupling the two processes often
competitors through comprehensive experiments.
results in misaligned texture and geometry [I Ho et al. 2024].
This paper proposes Elevate3D, a novel 3D model re!nement ap-
proach that produces a high-quality 3D model with well-aligned 2 Related Work
texture and geometry. Elevate3D operates iteratively, re!ning tex- 3D Model Generation. Recent advances in 3D model generation
tures and geometries view-by-view. For each view, it !rst enhances leverage di"usion models due to their powerful image priors [Li et al.
the texture and then leverages the re!ned texture to adjust the 2024]. Initially, the SDS loss was introduced to use pre-trained text-
geometry, ensuring alignment. In subsequent views, the texture to-image di"usion models for generating 3D models directly from
re!nement is based on the visible parts of the updated geometry text or image prompts [Melas-Kyriazi et al. 2023; Poole et al. 2023;
and the original geometry. This process ensures precise alignment Tang et al. 2023; Wang et al. 2023; Xu et al. 2023]. Although promis-
between texture and geometry and allows the use of high-quality ing, these methods often produce textures with over-saturated colors
priors trained on large-scale image datasets for both texture and and blurred details. Alongside these optimization approaches, early
geometry, ultimately producing superior 3D models. e"orts also adapted 2D di"usion models to generate multi-view
For texture re!nement, Elevate3D adopts High-frequency-Swapping consistent novel views via inpainting for 3D reconstruction [Kant
SDEdit (HFS-SDEdit), a specialized method that resolves the !- et al. 2023; Ryu et al. 2023]. These were succeeded by a line of works
delity–quality trade-o" in SDEdit. Our key insight leverages the that !ne-tune pre-trained di"usion models to generate multi-view

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • 3

images for 3D reconstruction [Liu et al. 2024; Long et al. 2023; Shi 2024a; Youwang et al. 2024; Zeng et al. 2024]. Texturing approaches
et al. 2023a, 2024]. However, these models require additional com- synthesize textures for 3D models without altering the geometry.
ponents on top of the large-scale di"usion model [Rombach et al. Hence, these methods, like 3D model re!nement approaches, face
2022] to enforce multi-view consistency. This added complexity the limitation of texture-geometry misalignment. Even if an input
increases computational costs and memory requirements, limiting 3D model lacks geometric details, these methods may still produce
both the number of views that can be generated simultaneously and textures with rich details learned from generative priors, leading
their resolution, consequently a"ecting the quality of the generated to inconsistencies between texture and geometry. This highlights
3D models [Long et al. 2023]. In our work, we demonstrate that the need for methods that jointly re!ne both texture and geometry.
our re!nement method can successfully produce high-quality 3D Our proposed Elevate3D addresses the shortcomings of one-sided
models from low-quality ones generated by previous approaches. methods by integrating both aspects within a uni!ed framework.
3D Model Re!nement. Recently, 3D model re!nement methods 3 HFS-SDEdit for Image Refinement
have primarily emerged as components of 3D model generation
pipelines. To improve the low-quality outputs from a coarse 3D gen- HFS-SDEdit, built on top of SDEdit [Meng et al. 2022], is a critical
eration step, recent pipelines often introduce a relatively straight- component of Elevate3D for high-quality, high-!delity texture re-
forward re!nement stage. One widely adopted approach is to use !nement. This section reviews SDEdit then introduces HFS-SDEdit.
SDEdit [Meng et al. 2022] with an image-based loss [Sun et al. 2024; SDEdit. SDEdit is a technique designed to guide the image syn-
Tang et al. 2024b; Yang et al. 2024c; Zhang et al. 2023a]. Specif- thesis process of di"usion models. Speci!cally, for a given reference
ically, a low-quality image is rendered from a coarse 3D model image, such as a stroke image or a low-quality image, the goal of
at an arbitrary viewpoint and then re!ned using SDEdit. The 3D SDEdit is to generate a realistic image that aligns with the learned
model is updated using the re!ned images through an MSE loss. distribution of the di"usion model and adheres to the structure of
However, this approach has signi!cant drawbacks. First, re!ned the reference image. SDEdit leverages the observation that adding
details at each view are not multi-view consistent; details intro- more noise to images in both the reference image domain (I𝐿 ) and
duced in one viewpoint are averaged out across others, resulting the realistic image domain (I) gradually merges them, transform-
in blurry outcomes. Second, the image-based loss under-constrains ing them into the domain of Gaussian noise. Based on this, SDEdit
the geometry, forcing such methods to keep the geometry !xed and proposes a simple, training-free strategy. Let 𝐿𝐿 ↑ I𝐿 represent a
focus solely on texture re!nement. Another line of approaches [Wu reference image in the latent space of a di"usion model. SDEdit then
et al. 2024a; Zhang et al. 2024] generates multi-view images and initalizes a noisy latent 𝐿𝑀𝐿 by adding noise to 𝐿𝐿 as:
re!nes them using o"-the-shelf super-resolution networks such as
𝐿𝑀𝐿 = 𝑀 (𝑁𝑁 )𝐿𝐿 + 𝑂 (𝑁𝑁 )𝑃, (1)
Real-ESRGAN [Wang et al. 2021]. The re!ned images are then used
to synthesize 3D models. However, these methods perform super- where 𝑁𝑁 is a timestep in the denoising schedule of the di"usion
resolution on each image independently, resulting in blurriness due model such that 𝑁𝑁 ↑ {𝑄 ,𝑄 ↓ 1, · · · , 0}. The functions 𝑀 (𝑁) ↑ [0, 1]
to multi-view inconsistency. and 𝑂 (𝑁) ↑ [0, 1] control the noise level at timestep 𝑁. Starting from
Standalone 3D re!nement methods outside of 3D generation this initial noisy latent, SDEdit performs the iterative denoising
pipelines encounter similar challenges. For instance, MagicBoost [Yang process to produce a realistic image 𝐿 0 .
et al. 2024c] employs multi-view di"usion models combined with SDEdit o"ers several distinct bene!ts, including - easy integration
SDS optimization to enhance both geometry and texture. However, into
-
di"usion-based image synthesis pipelines, training-free imple-
artifacts introduced during SDS optimization necessitate additional mentation, and computational e$ciency. Thus, it has been widely
re!nement using SDEdit, leading to the same multi-view consis- adopted for various tasks, e.g., image editing [Wu and De la Torre
tency issues. DiSR-NeRF [Lee et al. 2024] pairs an SDS loss [Poole 2023], video generation [Zhang et al. 2023b], 3D editing [Chen et al.
et al. 2023] with a di"usion-based 2D super-resolution model [Rom- 2024b]. However, it exhibits a !delity-quality trade-o", which will
bach et al. 2022]. Although this method iteratively re!nes a low- be discussed in detail later.
resolution NeRF by enhancing 2D images and synchronizing the 3D HFS-SDEdit. To overcome the !delity-quality trade-o" of SDEdit,
model, it prioritizes consistency in low-resolution features. Conse- our proposed HFS-SDEdit replaces the high-frequency component of
-

quently, high-resolution details generated during re!nement can


See equation (3)
-

the latent representation with that of the reference image to ensure


still be averaged out during 3D synchronization. -
that the synthesized image 𝐿 0 aligns with the structural details of
In contrast, Elevate3D adopts a novel approach to re!nement -
the reference image 𝐿𝐿 . Speci!cally, HFS-SDEdit starts with Eq. (1).
by preserving and distinguishing previously enhanced regions in -
Then, in the subsequent timesteps 𝑁, HFS-SDEdit replaces the high-
each view. This ensures a multi-view consistent result across all frequency component of the latent 𝐿ˆ𝑀 , the output of the standard
viewpoints. Additionally, by alternating between texture and geom- denoising process, by evaluating:
etry re!nement steps, Elevate3D allows the geometry to be re!ned
according to the improved textures. This strategy addresses the limi- 𝐿˜𝑀 = 𝑀 (𝑁)𝐿𝐿 + 𝑂 (𝑁)𝑃, and (2)
tations of earlier methods, which either average out high-frequency 𝐿𝑀↔ = (𝑅 ↓ 𝑆𝑂 ) ↗ 𝐿˜𝑀 + 𝑆𝑂 ↗ 𝐿ˆ𝑀 (3)
details or neglect geometry re!nement altogether. T highfreg of the tent space high-freq of reference image

where 𝐿˜𝑀 is a noised reference image and is a calibrated latent 𝐿𝑀↔


3D Model Texturing. Another relevant task related to ours is 3D whose high-frequency component is replaced with that of 𝐿˜𝑀 . 𝑅 is a
model texturing [Chen et al. 2023; Richardson et al. 2023; Tang et al. Dirac delta function, 𝑆𝑂 is a Guassian kernel with standard deviation

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
4 • Ryu et al.

Naïve SDEdit
High Quality Low Quality
(a) Ours (b) SDEdit(High) (c) Input (d) SDEdit(Low)
𝜖𝜖 ~𝑁𝑁 0,1
𝜖𝜖
Ours

𝑧𝑧𝑡𝑡ℎ
𝑧𝑧𝑡𝑡𝑙𝑙 𝑧𝑧𝑟𝑟

Denoising Time Steps

Fig. 2. Overview of HFS-SDEdit. HFS-SDEdit addresses the quality-fidelity trade-o" in SDEdit. By adding a substantial amount of noise 𝑃 to the low-quality
reference image 𝑄𝑀 in (c) and initiating the denoising process from the noisy latent 𝑄𝑁h , SDEdit removes domain information, enabling the di"usion model to
generate a high-quality image as depicted in (b). However, this approach compromises fidelity to the reference image. Conversely, adding a small amount of
noise and starting the denoising process from the noisy latent 𝑄𝑁l preserves the low-quality domain information, resulting in only minor refinements, as seen
in (d). In contrast, HFS-SDEdit incorporates high-frequency feature injection-based guidance, as detailed in Section 3, allowing for high-fidelity generation
even when starting the denoising process from 𝑄𝑁h . This approach achieves both high quality and high fidelity in the refinement process. Input: ©Wa#s/flickr.

𝑇, and ↗ is the convolution operator. Once 𝐿𝑀↔ is obtained, denoising Texture Refinement

is performed with 𝐿𝑀↔ resulting in 𝐿ˆ𝑀 ↓1 . We perform high-frequency


replacement until 𝑁 reaches a prede!ned timestep 𝑁 stop to ensure
that the resulting image does not reproduce all the details, such as
the degraded details of the low-quality reference image.
HFS-SDEdit is based on the intuition that the domain information Texture Texture

of an image is encoded in the mid-frequency component of the


Geometry Refinement

latent representation. Speci!cally, we may decompose the latent Geometry Geometry

representation into two frequency bands: low-, and high-frequency Coarse Geometry & Refined Geometry
components. Similar to conventional images, the high-frequency Unrefined Texture & Texture

component describes small-scale structures [Park et al. 2023]. On


the other hand, the low-frequency component contains not only Fig. 3. Framework Overview. Given a low-quality 3D model, Elevate3D
large-scale structures but also domain information. alternatingly refines texture and geometry. Input for the experiment:
Fig. 2 illustrates the intuition behind HFS-SDEdit as well as the ©Momentmal/pixabay.
quality-!delity trade-o" of SDEdit with an image enhancement
example. Adding more noise to the reference image 𝐿𝐿 gradually
removes information from high-frequency to low-frequency compo- 4 Elevate3D
nents. Thus, to su$ciently merge the realistic image domain and the Fig. 3 visualizes an overview of Elevate3D. Given a low-quality tex-
reference image domain, or equivalently to remove domain infor- tured mesh, our approach progressively re!nes both the texture and
mation, SDEdit requires the use of large noise to remove a su$cient geometry of the mesh by iterating through a prede!ned camera path
amount of low-frequency components. However, such excessive V = {𝑈 0, . . . , 𝑈𝑅 }. Speci!cally, at the 𝑉-th iteration corresponding
noise also destroys small- and large-scale structures, causing high- to 𝑈𝑆 , we denote the initial partially re!ned model as 𝑊𝑆 , where 𝑊0
quality but low-!delity results. Conversely, SDEdit with small noise is initialized as the input low-quality textured mesh. Our method
removes only the high-frequency component, producing an image then re!nes the unre!ned regions of the texture and geometry of
not only with di"erent small-scale details but also within the same 𝑊𝑆 that are visible at 𝑈𝑆 through texture re!nement and geometry
domain as the reference image, the low-quality image domain. re!nement stages. The texture re!nement stage focuses on enhanc-
To improve both quality and !delity, HFS-SDEdit starts with large ing the unre!ned regions while maintaining consistency with the
noise and injects high-frequency components of the reference image already re!ned areas. Afterward, the geometry re!nement stage ex-
into the synthesis process. This approach e"ectively removes the do- tracts geometric cues from the newly re!ned texture and re!nes the
main information of the reference image, resulting in a high-quality mesh geometry accordingly. This ensures that the geometry accu-
image. The injected high-frequency component not only constrains rately re#ects the details present in the re!ned texture, maintaining
the high-frequency details of a synthesized image, but also guides texture-geometry consistency.
the di"usion model to synthesize a low-frequency component that This view-by-view re!nement strategy enables our method to
--
aligns with the injected details, achieving high-!delity synthesis. utilize prede!ned image and geometry priors, which contribute to
achieving high-quality texture and geometry. Additionally, by lever-
aging the re!ned texture and geometry when processing unre!ned
regions in subsequent viewpoints, our method ensures cross-view
consistency. Also, the geometry re!nement is performed based on

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • 5

the re!ned texture, signi!cantly enhancing texture-geometry con- respectively, where 𝑐 and 𝑈 represent pixel coordinates. Our goal
sistency. In the following, we explain each stage in detail. is to estimate the depth component 𝐿 (𝑐, 𝑈) such that the resulting
surface 𝑎𝑆 satis!es the following two criteria: First, the normals of 𝑎𝑆
4.1 Texture Refinement with HFS-SDEdit must be close to n𝑆 , or equivalently, its tangent vectors 𝑒𝑎𝑆 (𝑐, 𝑈)/𝑒𝑐
The texture re!nement stage improves the texture of the partially and 𝑒𝑎𝑆 (𝑐, 𝑈)/𝑒𝑈 are orthogonal to n𝑆 (𝑐, 𝑈). Second, the estimated
re!ned 3D model 𝑊𝑆 for the current view 𝑈𝑆 by leveraging HFS- depth 𝐿 (𝑐, 𝑈) must be close to the depth 𝑏 (𝑐, 𝑈) rendered from 𝑊𝑆 .
SDEdit with a large-scale pretrained image di"usion model. We Based on these two criteria, we de!ne an energy functional:
∬ $%
begin by rendering 𝑊𝑆 from 𝑈𝑆 to obtain the initial image 𝑋𝑆 (Fig. 4-a). 𝑒𝐿 𝑑𝑇 (𝑐, 𝑈) & 2 % 𝑒𝐿 𝑑 𝑈 (𝑐, 𝑈) & 2 '
This image contains regions re!ned at previous views {𝑈 0, . . . , 𝑈𝑆 ↓1 } 𝑓 (𝐿) = +
𝑒𝑐 𝑑𝑄 (𝑐, 𝑈)
+ +
𝑒𝑈 𝑑𝑄 (𝑐, 𝑈)
𝑏𝑐 𝑏𝑈
and unre!ned regions that still exhibit low-quality textures. ∬ % &2 (5)
To isolate the unre!ned regions in 𝑋𝑆 , we identify pixels that +𝑔 𝐿 (𝑐, 𝑈) ↓ 𝑏 (𝑐, 𝑈) 𝑏𝑐 𝑏𝑈,
are not visible in previous views. Speci!cally, we rasterize normal
vectors for 𝑈𝑆 and compare them with the viewing directions of the where the !rst and second terms on the right-hand-side correspond
previous views {𝑈 0, . . . , 𝑈𝑆 ↓1 } by computing cosine similarities. We to the !rst and second criteria, respectively. 𝑔 is a regularization
detect pixels whose cosine similarity values exceed a threshold 𝑌 parameter that balances the two terms. Minimizing 𝑓 (𝐿) yields a
for all previous views and construct a binary mask 𝑍𝑆 (Fig. 4-b). We re!ned depth map for 𝑎𝑆 . In our implementation, we follow the
set 𝑌 = 0.5, equivalent to a 60↘ angle di"erence. This approach may normal integration method of Cao et al. [2022] to minimize 𝑓 (𝐿).
include already-re!ned pixels in the mask, but it allows regions seen Once we obtain the re!ned geometry 𝑎𝑆 represented by the re-
at oblique angles to be further re!ned. !ned depth map, we update the mesh 𝑊𝑆 . Speci!cally, the re!ned
Once 𝑍𝑆 is obtained, we re!ne the unre!ned regions in 𝑋𝑆 using geometry region 𝑎𝑆 is integrated with the unchanged parts of 𝑊𝑆
HFS-SDEdit. To re!ne only the unre!ned regions speci!ed by 𝑍𝑆 , using Poisson surface reconstruction [Kazhdan et al. 2006]. This
we introduce a slight modi!cation to HFS-SDEdit. Speci!cally, at step generates the updated triangular mesh 𝑊˜ 𝑆 . While Poisson re-
each timestep of the di"usion sampling process, we !rst compute a construction does not strictly preserve the original mesh topology,
noised reference 𝐿˜𝑀 and a latent 𝐿𝑀↔ using Eqs. (2) and (3). To ensure it rarely introduces geometric artifacts in our case. This robustness
that the regions outside 𝑍𝑆 are not overwritten, we blend 𝐿𝑀↔ with stems from our regularized integration in Eq. (5), which constrains
𝐿˜𝑀 using 𝑍,
˜ a downsampled version of 𝑍𝑆 . This blending is done as: the geometric update using the coarse input mesh 𝑊𝑆 via the depth
! " map 𝑏 as a guide. This process ensures that each view’s improved
ẑ𝑀 = m̃ ≃ z𝑀↔ + 1 ↓ m̃ ≃ z̃𝑀 . (4) geometry is seamlessly integrated without disrupting previously
Finally, ẑ𝑀 goes through the denoising step. After iterative denois- re!ned areas of the model. Furthermore, thanks to the geometry re-
ing with HFS-SDEdit, we obtain a re!ned image 𝑋𝑆↔ with improved !nement process leveraging the re!ned texture, Elevate3D ensures
textures faithfully re#ecting 𝑋𝑆 for the unre!ned regions while pre- proper texture-geometry alignment.
serving the textures from 𝑋𝑆 in the already-re!ned regions (Fig. 4-c). Fig. 5 illustrates this geometry re!nement process. Given a par-
tially re!ned mesh 𝑊𝑆 (Fig. 5-a), our geometry re!nement stage
4.2 Geometry Refinement with Refined Texture estimates a re!ned geometry 𝑎𝑆 , visualized via its depth map (Fig. 5-
The geometry re!nement step enhances the geometry of the par- b). This re!ned region is then merged with the other regions of
tially re!ned 3D triangle mesh 𝑊𝑆 by utilizing the re!ned texture 𝑊𝑆 using Poisson reconstruction (Fig. 5-c), resulting in the updated
image 𝑋𝑆↔ from the previous texture re!nement stage. This image not mesh 𝑊˜ 𝑆 (Fig. 5-d). We note that the seams in Fig. 5-b are the result
only exhibits improved textures but also provides valuable cues for of !ltering of unreliable depth values around discontinuities, which
geometric details. These details can be extracted using monocular we provide details in the supplementary material.
geometry estimation models [Ke et al. 2024; Martin Garcia et al. After the geometry re!nement stage, we project the re!ned tex-
2025; Yang et al. 2024b]. ture image 𝑋𝑆↔ onto the updated mesh 𝑊˜ 𝑆 to obtain the textured mesh
In Elevate3D, we start by inferring a normal map n𝑆 from the re- 𝑊𝑆+1 for the next view re!nement iteration. For this, we employ
!ned image 𝑋𝑆↔ using a state-of-the-art normal estimation model [Mar- projection mapping, a form of UV-free texture mapping, similar to
tin Garcia et al. 2025]. We then integrate the estimated normals n𝑆 to recent mesh texturing methods [Chen et al. 2023; Richardson et al.
obtain a re!ned surface 𝑎𝑆 corresponding to the viewpoint 𝑈𝑆 . This 2023; Tang et al. 2024a].
surface is consistent with the re!ned texture image 𝑋 ↔ . However,
because n𝑆 is derived solely from 𝑋𝑆↔ and not conditioned on the 5 Experiments
existing geometry of 𝑊𝑆 , the re!ned geometry 𝑎𝑆 can signi!cantly 5.1 Implementation Details
deviate from that of 𝑊𝑆 . Consequently, directly stitching 𝑎𝑆 onto 𝑊𝑆
In all experiments, we use FLUX1 —an open-source large-scale text-
could introduce severe geometric distortion.
to-image di"usion model trained with the recti!ed #ow-matching
To address this, we introduce a regularized normal integration
formulation [Albergo and Vanden-Eijnden 2022; Lipman et al. 2023;
scheme to estimate 𝑎𝑆 while ensuring consistency with the existing
Liu et al. 2022]. We follow the default parameters and the denois-
geometry of 𝑊𝑆 . Our regularized normal integration scheme is imple-
ing schedule used by FLUX. For all experiments, we use a total of
mented as follows. We assume an orthographic camera model and
𝑄 = 30 denoising steps. We set the initial noise timestep 𝑁𝑁 to 29
rasterize a depth map 𝑏 of 𝑊𝑆 from viewpoint 𝑈𝑆 . We de!ne 𝑎𝑆 and n𝑆
as 𝑎𝑆 (𝑐, 𝑈) = [𝑐, 𝑈, 𝐿 (𝑐, 𝑈)] ⇐ , and n𝑆 (𝑐, 𝑈) = [𝑑𝑇 (𝑐, 𝑈), 𝑑 𝑈 (𝑐, 𝑈), 𝑑𝑄 (𝑐, 𝑈)] ⇐ , 1 [Link]

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
6 • Ryu et al.

Table 1. !antitative Comparison on 3D Refinement. Elevate3D consis-


tently achieves the best scores across various non-reference quality metrics,
highlighting its high-quality 3D refinement. DreamGaussian [Tang et al.
2024b] refines only textures while others refine both texture and geometry.
All refinement time were measured using an NVIDIA RTX A6000 GPU.

Model MUSIQ ⇒ LIQE ⇒ TOPIQ ⇒ Q-Align ⇒ Time ⇑


(a) Initial image 𝐼𝐼𝑖𝑖 (b) Refinement Mask 𝑚𝑚𝑖𝑖 (c) Refined image 𝐼𝐼′𝑖𝑖
DreamGaussian [2024b] 61.667 2.1185 0.4690 2.7416 1 min
DiSR-NeRF [2024] 48.940 1.2869 0.3879 2.6794 6 hrs
Fig. 4. Texture Refinement. Given a partially-refined texture image 𝑉𝑂 in MagicBoost [2024c] 51.646 2.1085 0.3915 2.4992 20 min
(a), the texture refinement stage detects a refinement mask 𝑊𝑂 in (b), and Elevate3D (Ours) 66.527 2.7744 0.5295 3.2151 25 min
produces a refined image in (c) using HFS-SDEdit.

Table 2. !antitative Comparison on 2D Image Refinement. The full-


reference metrics are evaluated against the high-quality source images.
Baseline (LQ) denotes the degraded version of the source images.

Full-Reference No-Reference
Model
PSNR ⇒ SSIM ⇒ LPIPS ⇑ MUSIQ ⇒ QAlign ⇒ LIQE ⇒ TOPIQ ⇒
Baseline (LQ) 20.701 0.521 0.662 21.918 2.018 1.303 0.159
SDEdit (strength = 0.4) 19.214 0.473 0.679 22.863 2.237 1.261 0.174
SDEdit (strength = 0.8) 15.255 0.379 0.746 29.190 2.860 1.321 0.214
NC-SDEdit 17.737 0.442 0.697 25.257 2.476 1.329 0.184
HFS-SDEdit (Ours) 15.588 0.391 0.598 39.519 3.337 2.105 0.283

(a) Partially (b) Refined (c) Region of 𝑀𝑀𝑖𝑖 (d) Updated


Refined Mesh 𝑀𝑀𝑖𝑖 Geometry 𝑆𝑆𝑖𝑖 Invisible from 𝑣𝑣𝑖𝑖 Mesh 𝑀𝑀�𝑖𝑖
Fig. 10 shows a qualitative comparison. As the !gure shows, the
results of previous approaches show blurry textures and less accu-
Fig. 5. Geometry Refinement. Given a partially refined geometry 𝑋𝑂 in
rate geometries that are inconsistent with the textures. In contrast,
(a), we obtain a refined surface 𝑌𝑂 in (b), and stitch it with the other regions
of 𝑋𝑂 shown in (c), resulting in the updated mesh 𝑋˜ 𝑂 in (d). Input for the
our method successfully re!nes input low-quality 3D models, pro-
experiment: the GSO dataset [Downs et al. 2022] ducing detailed textures and geometries, substantially surpassing
previous methods. Our results also exhibit high texture-geometry
consistency, thanks to our geometry re!nement strategy that lever-
ages re!ned textures. We also report a quantitative comparison of
and the frequency swapping threshold 𝑁 stop to 18. We employ a the quality of the re!ned 3D models in Table 1. For quantitative eval-
Gaussian low-pass !lter with 𝑇 = 4. These parameters were selected uation, we render the re!ned models and assess the quality of the
by qualitatively comparing di"erent combinations and choosing rendered images using various image quality metrics: MUSIQ [Ke
the setting that yielded the best qualitative results. A quantitative et al. 2021], LIQE [Zhang et al. 2023c], TOPIQ [Chen et al. 2024a],
comparison for various combinations of 𝑇 and 𝑁 stop is provided in and Q-Align [Wu et al. 2024b]. As reported in the table, our method
the supplementary material. For the depth regularization, we set consistently outperforms other approaches across all quality metrics.
𝑔 = 0.008. We use an orthographic camera, and for selecting the Regarding computation times,the unoptimized prototype implemen-
camera schedule, we adopt a strategy similar to Text2Tex [Chen tation of Elevate3D operates slower than DreamGaussian, which
et al. 2023]: starting with a pre-de!ned set of camera poses V, we exclusively re!nes textures. However, it achieves similar e$ciency
use an automatic view selection scheme to choose the re!nement to MagicBoost and is considerably faster than DiSR-NeRF.
view that covers the largest unre!ned region. We leave the speci!c Elevate3D can also be utilized to produce superior 3D models
details in the supplementary. when combined with state-of-the-art (SoTA) image/text-to-3D syn-
thesis methods. Fig. 11 illustrates the re!nement of 3D models gen-
5.2 Evaluation of Elevate3D erated by TRELLIS [Xiang et al. 2024], a leading 3D model synthesis
We evaluate the 3D model re!nement quality of Elevate3D using a method. A common issue with 3D generation models like TRELLIS
real-world scan dataset, GSO [Downs et al. 2022], where we formed is that they often fail to produce high-quality results for inputs out-
a test set comprising 59 objects. To this end, we degrade the 3D side the domain of their 3D training datasets, which mostly consist
models from the GSO dataset by reducing the number of faces to 20% of synthetic objects. Elevate3D e"ectively enhances these results,
and applying a Gaussian low-pass !lter with 𝑇 = 8 to the textures. as shown in the !gure, producing superior 3D models.
We then compare our method with recent 3D model re!nement
approaches: MagicBoost [Yang et al. 2024c], DiSR-NeRF [Lee et al. 5.3 Evaluation of HFS-SDEdit
2024], and DreamGaussian [Tang et al. 2024b]. DreamGaussian We analyze HFS-SDEdit in the image enhancement task using the
focuses solely on texture re!nement, while the others re!ne both validation set of LSDIR [Li et al. 2023], a large-scale image restoration
texture and geometry. Using each method, we re!ne the degraded dataset. From LSDIR’s high-quality images, we create low-quality
3D models and compare the quality of the re!ned models. images by downsampling and upsampling them by a factor of 8.

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • 7

For text prompts, we use ChatGPT Vision [OpenAI et al. 2024] to


generate descriptions for the images.

Impact of Low-Frequency Component in Di"usion Sampling. We


validate our intuition that the domain information is encoded in (b) Low Quality Input (c) No Swap

the low-frequency component in the latent representation. To this


end, we conduct an experiment as follows. We prepare a low-quality
reference image. Then, we sample three di"erent images from the
same pure Gaussian noise using di"erent strategies. We sample the
!rst sample following the conventional di"usion process. We sam- (a) Mean RAPSD Plot (d) Low-Frequency Swap (e) High-Frequency Swap

ple the second sample similarly, but we replace the low-frequency


component of its latent representation with that of the low-quality Fig. 6. Impact of Low-frequency Component in Di"usion Sampling.
reference. To this end, we modify Eq. (3) as 𝐿𝑀↔ = 𝑆𝑂 ↗𝐿˜𝐿 + (𝑅 ↓𝑆𝑂 ) ↗𝐿𝑀 . The mean radially averaged power spectral density (RAPSD) graphs of dif-
ferent example images shown in (b)-(e) are shown in (a). Image (b): ©patrick
Finally, for the third image, we replace its high-frequency compo-
janicek/flickr.
nent with that of the low-quality reference. We swap the frequency
components only for the !rst four denoising timesteps for both
images when low-frequency features primarily emerge.
Fig. 6 compares the three images with their mean radially av-
eraged power spectral density (RAPSD) graphs. As shown in the
!gure, when the low-frequency component is replaced with that
of the low-quality reference, the di"usion model struggles to syn-
thesize high-frequency details, indicating a shift in the generation (a) Image Before Degradation (b) HFS-SDEdit (Ours) (c) NC-SDEdit
path toward the low-quality domain. Conversely, when only high-
frequency component is swapped, the model still produces detailed
textures regardless of the reference image’s quality. This observa-
tion clearly indicates that the domain information is not in the
high-frequency component but in the low-frequency component.
(d) Image After Degradation (e) SDEdit (Strength 0.4) (f) SDEdit (Strength 0.8)
Image Re!nement Comparisons. We validate the e"ectiveness of
HFS-SDEdit by comparing it with SDEdit and NC-SDEdit [Yang et al. Fig. 7. !alitative Comparison on 2D Image Refinement. The refine-
2024a]. NC-SDEdit is a video enhancement approach that updates ment results in (b), (c), (e), and (f) are obtained from the low-quality image
the low-frequency component of latents to match reference frames in (d), which was degraded from the image in (a). Image (a): ©Mathias
during the di"usion denoising process, enhancing !delity. For com- Appel/flickr.
parison, we apply NC-SDEdit to single images. We evaluate the
!delity of the re!nement results using full-reference metrics against
the high-quality source images: PSNR, SSIM, and LPIPS [Zhang Meanwhile, NC-SDEdit’s low-frequency retention condition inad-
et al. 2018]. For quality assessment, we employ no-reference met- vertently carries over low-quality domain information, leading to
rics: MUSIQ [Ke et al. 2021], LIQE [Zhang et al. 2023c], TOPIQ [Chen subpar results. Consequently, its outputs exhibit both lower !delity
et al. 2024a], and Q-Align [Wu et al. 2024b]. and perceptual quality compared to those of HFS-SDEdit, reinforcing
Table 2 shows the quantitative comparison, where HFS-SDEdit the advantage of our method’s high-frequency-based guidance.
achieves the best performance in no-reference metrics, demonstrat-
ing its capability to produce high-quality outputs. It also obtains the 5.4 Ablation Studies
best LPIPS score among its competitors, indicating that the result- Finally, we conduct ablation studies to justify our design choices.
ing images preserve a high degree of perceptual similarity to the To demonstrate the necessity of texture and geometry re!nement
original images. However, because HFS-SDEdit uses a generative stages, we compare scenarios where only texture or geometry re-
approach to re!ne the original input, it does not achieve the best !nement is performed, as shown in Fig. 8. Without geometry re-
PSNR and SSIM scores. This is common in generative-re!nement !nement, the resulting 3D model retains the crude geometry of the
methods, which prioritize plausible re!nement over exact pixel-level input model (Fig. 8-a). Conversely, without texture re!nement, the
!delity [Blau and Michaeli 2018; Gu et al. 2020, 2022; Yu et al. 2024]. geometry re!nement stage relies on the low-quality input texture,
Despite this, Fig. 7 illustrates that HFS-SDEdit still produces images resulting in minimal geometry improvement (Fig. 8-b). Employing
with convincing !delity and high-frequency details. both texture and geometry re!nement stages yields high-quality
When examining SDEdit across various strengths, we observe textures and geometry that are consistent with each other (Fig. 8-c).
a !delity-quality trade-o". Lower strength values yield better full- Fig. 9 illustrates the e"ect of the regularized normal integration.
reference metrics but lower no-reference scores, indicating that As discussed in Section 4.2, normals predicted solely from a re!ned
!delity is achieved at the expense of quality. Conversely, higher texture may be inconsistent with the input geometry (Fig. 9-a),
strength values improve perceived quality at the expense of !delity. thus directly using them for geometry re!nement causes severe

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
8 • Ryu et al.

Acknowledgments
This work was supported by Pebblous Inc., and the Institute of
Information & Communications Technology Planning & Evalua-
tion (IITP) grants (RS-2019-II91906, Arti!cial Intelligences Graduate
School Program (POSTECH), RS-2024-00457882, AI Research Hub
Project) funded by the Korea government (MSIT).

References
Michael S. Albergo and Eric Vanden-Eijnden. 2022. Building Normalizing Flows with
Stochastic Interpolants. arXiv:2209.15571 [[Link]]
Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman.
(a) Only Texture Refinement (b) Only Geometry Refinement (c) Full Refinement 2022. Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields. In Proceedings
of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5470–
5479.
Fig. 8. E"ect of Texture and Geometry Refinement Stages. Top: tex- Colin Barré-Brisebois and Stephen Hill. 2012. Blending in Detail.
tured meshes, bo#om: geometries. All results are rendered using flat shading. [Link]
Input for the experiment: the GSO dataset [Downs et al. 2022] Yochai Blau and Tomer Michaeli. 2018. The Perception-Distortion Tradeo". In 2018
IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 6228–6237.
[Link]
Xu Cao, Hiroaki Santo, Boxin Shi, Fumio Okura, and Yasuyuki Matsushita. 2022. Bilat-
eral Normal Integration. In ECCV.
Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu, Liang Liao, Wenxiu Sun, Qiong
Yan, and Weisi Lin. 2024a. TOPIQ: A Top-Down Approach From Semantics to Dis-
tortions for Image Quality Assessment. Trans. Img. Proc. 33 (March 2024), 2404–2418.
[Link]
Dave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov, and Matthias
Nießner. 2023. Text2Tex: Text-driven Texture Synthesis via Di"usion Models. In
Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
(a) Input Geometry 𝑀𝑀𝑖𝑖 (b) Refined w/o Regularizer (c) Refined w/ Regularizer 18558–18568.
Hansheng Chen, Ruoxi Shi, Yulin Liu, Bokui Shen, Jiayuan Gu, Gordon Wetzstein, Hao
Su, and Leonidas Guibas. 2024b. Generic 3D Di"usion Adapter Using Controlled
Fig. 9. E"ect of Regularized Normal Integration. (a) Initial geometry. Multi-View Editing. arXiv:2403.12032 [[Link]]
(b) Geometry refinement using normal integration w/o regularization. (c) Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya
Geometry refinement using normal integration w/ regularization (Ours). Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli
VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani,
Input for the experiment: the GSO dataset [Downs et al. 2022]
Ludwig Schmidt, and Ali Farhadi. 2023. Objaverse-XL: A Universe of 10M+ 3D
Objects. In Thirty-seventh Conference on Neural Information Processing Systems
Datasets and Benchmarks Track. [Link]
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli Vander-
Bilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. 2022.
distortions (Fig. 9-b). Our regularized normal integration e"ectively Objaverse: A Universe of Annotated 3D Objects. arXiv:2212.08051 [[Link]]
addresses this issue, resulting in high-quality geometry (Fig. 9-c). Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista
Reymann, Thomas B. McHugh, and Vincent Vanhoucke. 2022. Google Scanned Ob-
jects: A High-Quality Dataset of 3D Scanned Household Items. In 2022 International
6 Conclusion and Future Work Conference on Robotics and Automation (ICRA) (Philadelphia, PA, USA). IEEE Press,
2553–2560. [Link]
In this work, we proposed Elevate3D, a novel 3D model re!nement Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry
Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim
framework that alternates between texture and geometry re!nement Dockhorn, Zion English, and Robin Rombach. 2024. Scaling Recti!ed Flow Trans-
in a view-by-view fashion to produce high-quality 3D models with formers for High-Resolution Image Synthesis. In Forty-!rst International Conference
well-aligned texture and geometry. We introduced HFS-SDEdit for on Machine Learning. [Link]
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley,
texture re!nement, leveraging high-frequency guidance to achieve Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial
high-!delity enhancements while mitigating the limitations of pre- nets. Advances in neural information processing systems 27 (2014).
vious SDEdit-based methods. Through comprehensive experiments, Jinjin Gu, Haoming Cai, Haoyu Chen, Xiaoxing Ye, Jimmy Ren, and Chao Dong. 2020. PI-
PAL: a Large-Scale Image Quality Assessment Dataset for Perceptual Image Restora-
we demonstrated that our framework achieves state-of-the-art qual- tion. arXiv:2007.12142 [[Link]] [Link]
ity re!nement of 3D models compared to recent competitors. Jinjin Gu, Haoming Cai, Chao Dong, Jimmy S. Ren, and Radu Timofte. 2022. NTIRE
2022 Challenge on Perceptual Image Quality Assessment. arXiv:2206.11695 [[Link]]
[Link]
Limitations and Future Work. While our framework produces Yong He, Hongshan Yu, Xiaoyan Liu, Zhengeng Yang, Wei Sun, Saeed Anwar, and Ajmal
high-quality textured meshes, it shares a common limitation with Mian. 2025. Deep learning based 3D segmentation in computer vision: A survey.
similar di"usion-based texturing methods [Chen et al. 2023; Richard- Information Fusion 115 (2025), 102722. [Link]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising di"usion probabilistic
son et al. 2023; Tang et al. 2024a]: re!nement time increases pro- models, Vol. 33. 6840–6851.
portionally with the number of views that need to be generated by Hsuan I Ho, Jie Song, and Otmar Hilliges. 2024. SiTH: Single-view Textured Human
Reconstruction with Image-Conditioned Di"usion. In Proceedings of the IEEE/CVF
the di"usion model. Recent advances in increasing the e$ciency of Conference on Computer Vision and Pattern Recognition (CVPR). 538–549.
di"usion models [Kim et al. 2024; Sauer et al. 2024] o"er potential Yash Kant, Aliaksandr Siarohin, Michael Vasilkovsky, Riza Alp Guler, Jian Ren, Sergey
reductions in our framework’s computational cost. Future work will Tulyakov, and Igor Gilitschenski. 2023. iNVS: Repurposing Di"usion Inpainters for
Novel View Synthesis. In SIGGRAPH Asia 2023 Conference Papers (Sydney, NSW,
explore integrating such models to optimize Elevate3D’s processing Australia) (SA ’23). Association for Computing Machinery, New York, NY, USA,
time while maintaining its high-quality output. Article 16, 12 pages. [Link]

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • 9

Michael Kazhdan, Matthew Bolitho, and Hugues Hoppe. 2006. Poisson surface re- Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe,
construction. In Proceedings of the Fourth Eurographics Symposium on Geometry Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv
Processing (Cagliari, Sardinia, Italy) (SGP ’06). Eurographics Association, Goslar, Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer
DEU, 61–70. McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie
and Konrad Schindler. 2024. Repurposing Di"usion-Based Image Generators for Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David
Monocular Depth Estimation. In Proceedings of the IEEE/CVF Conference on Computer Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard
Vision and Pattern Recognition (CVPR). Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe
Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. 2021. MUSIQ: Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita,
Multi-scale Image Quality Transformer. In 2021 IEEE/CVF International Conference on Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Bel-
Computer Vision (ICCV). 5128–5137. [Link] bute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny,
Geonung Kim, Beomsu Kim, Eunhyeok Park, and Sunghyun Cho. 2024. Di"usion Model Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Eliza-
Compression for Image-to-Image Translation. In Computer Vision – ACCV 2024: beth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond,
17th Asian Conference on Computer Vision, Hanoi, Vietnam, December 8–12, 2024, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder,
Proceedings, Part V (Hanoi, Vietnam). Springer-Verlag, Berlin, Heidelberg, 148–166. Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt,
[Link] David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jes-
Jie Long Lee, Chen Li, and Gim Hee Lee. 2024. DiSR-NeRF: Di"usion-Guided View- sica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens,
Consistent Super-Resolution NeRF. In Proceedings of the IEEE/CVF Conference on Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Na-
Computer Vision and Pattern Recognition (CVPR). 20561–20570. talie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang,
Xiaoyu Li, Qi Zhang, Di Kang, Weihao Cheng, Yiming Gao, Jingbo Zhang, Zhihao Liang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Eliza-
Jing Liao, Yan-Pei Cao, and Ying Shan. 2024. Advances in 3D Generation: A Survey. beth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe,
arXiv:2401.17807 [[Link]] [Link] Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay
Yawei Li, Kai Zhang, Jingyun Liang, Jiezhang Cao, Ce Liu, Rui Gong, Yulun Zhang, Hao Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila
Tang, Yun Liu, Denis Demandolx, Rakesh Ranjan, Radu Timofte, and Luc Van Gool. Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wietho", Dave Willner,
2023. LSDIR: A Large Scale Dataset for Image Restoration. In Proceedings of the Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu,
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. Je" Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech
1775–1787. Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2024. GPT-4 Technical
2023. Flow Matching for Generative Modeling. In The Eleventh International Confer- Report. arXiv:2303.08774 [[Link]]
ence on Learning Representations. [Link] Yong-Hyun Park, Mingi Kwon, Jaewoong Choi, Junghyo Jo, and Youngjung Uh. 2023.
Xingchao Liu, Chengyue Gong, and Qiang Liu. 2022. Flow Straight and Fast: Learning Understanding the Latent Space of Di"usion Models through the Lens of Riemannian
to Generate and Transfer Data with Recti!ed Flow. arXiv:2209.03003 [[Link]] Geometry. In Thirty-seventh Conference on Neural Information Processing Systems.
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and [Link]
Wenping Wang. 2024. SyncDreamer: Generating Multiview-consistent Images Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. 2023. DreamFusion:
from a Single-view Image. In The Twelfth International Conference on Learning Text-to-3D using 2D Di"usion. In ICLR.
Representations. [Link] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen
Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al. 2023. Wonder3D: Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From
Single Image to 3D using Cross-Domain Di"usion. arXiv preprint arXiv:2310.15008 Natural Language Supervision, Vol. 139. 8748–8763.
(2023). Elad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes, and Daniel Cohen-Or. 2023.
Gonzalo Martin Garcia, Karim Abou Zeid, Christian Schmidt, Daan de Geus, Alexander TEXTure: Text-Guided Texturing of 3D Shapes. In ACM SIGGRAPH 2023 Conference
Hermans, and Bastian Leibe. 2025. Fine-Tuning Image-Conditional Di"usion Models Proceedings (, Los Angeles, CA, USA,) (SIGGRAPH ’23). Association for Computing
is Easier than You Think. In Proceedings of the IEEE/CVF Winter Conference on Machinery, New York, NY, USA, Article 54, 11 pages. [Link]
Applications of Computer Vision (WACV). 3588432.3591503
Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. 2023. Re- Severi Rissanen, Markus Heinonen, and Arno Solin. 2023. Generative Modelling with
alFusion: 360deg Reconstruction of Any Object From a Single Image. In CVPR. Inverse Heat Dissipation. In The Eleventh International Conference on Learning
8446–8455. Representations. [Link]
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer.
Stefano Ermon. 2022. SDEdit: Guided Image Synthesis and Editing with Stochastic 2022. High-Resolution Image Synthesis With Latent Di"usion Models. In CVPR.
Di"erential Equations. In International Conference on Learning Representations. https: 10684–10695.
//[Link]/forum?id=aBsCjcPu_tE Nuri Ryu, Minsu Gong, Geonung Kim, Joo-Haeng Lee, and Sunghyun Cho. 2023. 360°
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Reconstruction From a Single Image Using Space Carved Outpainting. In SIGGRAPH
Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Asia 2023 Conference Papers (SA ’23). Association for Computing Machinery, New
Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, York, NY, USA, Article 75, 11 pages. [Link]
Haiming Bao, Mohammad Bavarian, Je" Belgum, Irwan Bello, Jake Berdine, Gabriel Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and
Bernadett-Shapiro, Christopher Berner, Lenny Bogdono", Oleg Boiko, Madelaine Robin Rombach. 2024. Fast High-Resolution Image Synthesis with Latent Adversarial
Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Di"usion Distillation. arXiv:2403.12015 [[Link]] [Link]
Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei,
Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Linghao Chen, Chong Zeng, and Hao Su. 2023a. Zero123++: a Single Image to
Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Consistent Multi-view Di"usion Base Model. arXiv:2310.15110 [[Link]]
Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory De- Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. 2024. MV-
careaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Dream: Multi-view Di"usion for 3D Generation. In The Twelfth International Con-
Steve Dowling, Sheila Dunning, Adrien Eco"et, Atty Eleti, Tyna Eloundou, David ference on Learning Representations. [Link]
Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Ful- Zifan Shi, Sida Peng, Yinghao Xu, Andreas Geiger, Yiyi Liao, and Yujun Shen. 2023b.
ford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Deep Generative Models on 3D Representations: A Survey. arXiv:2210.15663 [[Link]]
Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan [Link]
Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Je" Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015.
Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Deep Unsupervised Learning Using Nonequilibrium Thermodynamics. 2256–2265.
Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Jingxiang Sun, Bo Zhang, Ruizhi Shao, Lizhen Wang, Wen Liu, Zhenda Xie, and Yebin
Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Liu. 2024. DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Di"usion
Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Prior. In The Twelfth International Conference on Learning Representations. https:
%ukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak //[Link]/forum?id=DDX1u29Gqr
Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Jiaxiang Tang, Ruijie Lu, Xiaokang Chen, Xiang Wen, Gang Zeng, and Ziwei Liu. 2024a.
Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, %ukasz Kondraciuk, Andrew InTeX: Interactive Text-to-Texture Synthesis via Uni!ed Depth-aware Inpainting.
Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael arXiv preprint arXiv:2403.11878 (2024).
Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li,

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
10 • Ryu et al.

Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. 2024b. Dream- Generation with Progressive Controllable 2D Repainting. arXiv:2312.13271 [[Link]]
Gaussian: Generative Gaussian Splatting for E$cient 3D Content Creation. In The Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. 2020. NeRF++: Analyzing
Twelfth International Conference on Learning Representations. [Link] and Improving Neural Radiance Fields. arXiv:2010.07492 [[Link]] [Link]
net/forum?id=UyNXMqnN3c abs/2010.07492
Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Haoran Jiang, Wei
Chen. 2023. Make-It-3D: High-Fidelity 3D Creation from A Single Image with Yang, Lan Xu, and Jingyi Yu. 2024. CLAY: A Controllable Large-scale Generative
Di"usion Prior. arXiv:2303.14184 [[Link]] Model for Creating High-quality 3D Assets. ACM Trans. Graph. 43, 4, Article 120
A. Tewari, J. Thies, B. Mildenhall, P. Srinivasan, E. Tretschk, W. Yifan, C. Lassner, V. (July 2024), 20 pages. [Link]
Sitzmann, R. Martin-Brualla, S. Lombardi, T. Simon, C. Theobalt, M. Nießner, J. T. Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. 2018.
Barron, G. Wetzstein, M. Zollhöfer, and V. Golyanik. 2022. Advances in Neural The Unreasonable E"ectiveness of Deep Features as a Perceptual Metric. In CVPR.
Rendering. Computer Graphics Forum 41, 2 (2022), 703–735. [Link] Weixia Zhang, Guangtao Zhai, Ying Wei, Xiaokang Yang, and Kede Ma. 2023c. Blind
cgf.14507 arXiv:[Link] Image Quality Assessment via Vision-Language Correspondence: A Multitask Learn-
Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. 2021. Real- ing Perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and
ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data. Pattern Recognition (CVPR). 14071–14081.
arXiv:2107.10833 [[Link]] [Link]
Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. 2004. Image quality as-
sessment: from error visibility to structural similarity. IEEE Transactions on Image
Processing 13, 4 (2004), 600–612.
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun
Zhu. 2023. Proli!cDreamer: High-Fidelity and Diverse Text-to-3D Generation with
Variational Score Distillation. In Thirty-seventh Conference on Neural Information
Processing Systems. [Link]
Chen Henry Wu and Fernando De la Torre. 2023. A Latent Space of Stochastic Di"usion
Models for Zero-Shot Image Editing and Guidance. In Proceedings of the IEEE/CVF
International Conference on Computer Vision (ICCV). 7378–7387.
Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li,
Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, Qiong Yan, Xiongkuo Min,
Guangtao Zhai, and Weisi Lin. 2024b. Q-ALIGN: teaching LMMs for visual scoring
via discrete text-de!ned levels. In Proceedings of the 41st International Conference on
Machine Learning (Vienna, Austria) (ICML’24). [Link], Article 2216, 15 pages.
Kailu Wu, Fangfu Liu, Zhihan Cai, Runjie Yan, Hanyang Wang, Yating Hu, Yueqi Duan,
and Kaisheng Ma. 2024a. Unique3D: High-Quality and E$cient 3D Mesh Generation
from a Single Image. In The Thirty-eighth Annual Conference on Neural Information
Processing Systems. [Link]
Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Liang Pan Jiawei Ren, Wayne Wu, Lei
Yang, Jiaqi Wang, Chen Qian, Dahua Lin, and Ziwei Liu. 2023. OmniObject3D:
Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and
Generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition
(CVPR).
Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong
Chen, Xin Tong, and Jiaolong Yang. 2024. Structured 3D Latents for Scalable and
Versatile 3D Generation. arXiv preprint arXiv:2412.01506 (2024).
Dejia Xu, Yifan Jiang, Peihao Wang, Zhiwen Fan, Yi Wang, and Zhangyang Wang. 2023.
NeuralLift-360: Lifting an In-the-Wild 2D Photo to a 3D Object With 360deg Views.
In CVPR. 4479–4489.
Fan Yang, Jianfeng Zhang, Yichun Shi, Bowen Chen, Chenxu Zhang, Huichao Zhang,
Xiaofeng Yang, Jiashi Feng, and Guosheng Lin. 2024c. Magic-Boost: Boost 3D
Generation with Mutli-View Conditioned Di"usion. arXiv:2404.06429 [[Link]]
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and
Hengshuang Zhao. 2024b. Depth Anything V2. arXiv:2406.09414 (2024).
Qinyu Yang, Haoxin Chen, Yong Zhang, Menghan Xia, Xiaodong Cun, Zhixun Su,
and Ying Shan. 2024a. Noise Calibration: Plug-and-Play Content-Preserving Video
Enhancement Using Pre-trained Video Di"usion Models. In Computer Vision –
ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024,
Proceedings, Part XXXVI (Milan, Italy). Springer-Verlag, Berlin, Heidelberg, 307–326.
[Link]
Kim Youwang, Tae-Hyun Oh, and Gerard Pons-Moll. 2024. Paint-it: Text-to-Texture
Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based
Rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern
Recognition (CVPR). 4347–4356.
Fanghua Yu, Jinjin Gu, Zheyuan Li, Jinfan Hu, Xiangtao Kong, Xintao Wang, Jingwen
He, Yu Qiao, and Chao Dong. 2024. Scaling Up to Excellence: Practicing Model
Scaling for Photo-Realistic Image Restoration In the Wild. arXiv:2401.13627 [[Link]]
Zehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler, and Andreas Geiger.
2022. MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface
Reconstruction.
Xianfang Zeng, Xin Chen, Zhongqi Qi, Wen Liu, Zibo Zhao, Zhibin Wang, Bin Fu, Yong
Liu, and Gang Yu. 2024. Paint3D: Paint Anything 3D with Lighting-Less Texture
Di"usion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and
Pattern Recognition (CVPR). 4252–4262.
David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao
Gu, Difei Gao, and Mike Zheng Shou. 2023b. Show-1: Marrying Pixel and Latent
Di"usion Models for Text-to-Video Generation. arXiv:2309.15818 [[Link]] https:
//[Link]/abs/2309.15818
Junwu Zhang, Zhenyu Tang, Yatian Pang, Xinhua Cheng, Peng Jin, Yida Wei, Munan
Ning, and Li Yuan. 2023a. Repaint123: Fast and High-quality One Image to 3D

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
(a) Input (GSO)
(b) DreamGaussian
(c) MagicBoost
(d) DiSR-NeRF
(e) Ours (Elevate3D) Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • 11

Fig. 10. !alitative Comparison on 3D Refinement. We compare the 3D refinement results from a low-quality degraded input shown in (a). Dream-
Gaussian [Tang et al. 2024b] refines only the texture, leaving the geometry degraded as seen in (b). MagicBoost lacks a fidelity constraint, resulting in large
deviations from the input as seen in (c). DiSR-Nerf maintains high fidelity but struggles to generate high-frequency details as seen in (d). In contrast, our
method e"ectively refines both texture and geometry while preserving input fidelity while producing high-quality, well-aligned textures and geometries with
high quality as shown in (e). Inputs: the GSO dataset [Downs et al. 2022]

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
12 • Ryu et al.

(a) Input Image (b) Generated Result (TRELLIS) (c) Refined Result (Elevate3D)

Fig. 11. !alitative Results on Refining TRELLIS Outputs. Due to the domain gap between synthetic training data and real-world images, TRELLIS o$en
struggles to generate high-quality results from real-world inputs images such as in (a), as shown in (b). Therefore, we apply Elevate3D to refine TRELLIS’s
outputs. As seen in (c), our method produces realistic textures and accurate geometry, resulting in high-quality refinements. Inputs: ©mec4411/pixabay,
©vimleshtailor/pixabay, ©maja7777/pixabay, ©jacksonmoccelin/pixabay

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • i

A Additional Technical Details A.2 Projection Mapping Implementation Details


A.1 Details on Geometry Refinement and Filtering As mentioned in Section 4.2 of the main paper, after geometry
As described in Section 4.2 of the main paper, the geometry re!ne- re!nement yields the updated mesh 𝑊˜ 𝑆 , we project the correspond-
ment stage aims to !nd a re!ned depth map 𝐿 (𝑐, 𝑈) for the surface ing re!ned texture image 𝑋𝑆↔ along with other relevant textures
patch 𝑎𝑆 by minimizing the following energy functional: onto 𝑊˜ 𝑆 to produce the textured mesh 𝑊𝑆+1 . This projection map-
ping is implemented via a custom OpenGL fragment shader. This
∬ $%
𝑒𝐿 𝑑𝑇 (𝑐, 𝑈) & 2 % 𝑒𝐿 𝑑 𝑈 (𝑐, 𝑈) & 2 ' shader requires several pre-computed data structures passed as
𝑓 (𝐿) = + + + 𝑏𝑐 𝑏𝑈 uniforms to operate. These include arrays storing all source tex-
𝑒𝑐 𝑑𝑄 (𝑐, 𝑈) 𝑒𝑈 𝑑𝑄 (𝑐𝑕 𝑈)
∬ % &2 (6) ture images (Textures[]) corresponding to the camera path views
+𝑔 𝐿 (𝑐, 𝑈) ↓ 𝑏 (𝑐, 𝑈) 𝑏𝑐 𝑏𝑈. V = {𝑈 0, . . . , 𝑈𝑅 }, alongside a status array (IsRefined[]) indicat-
ing which textures have been re!ned by HFS-SDEdit. Additionally,
Here, the !rst term enforces consistency with the estimated normal parameters de!ning the projection for each view are needed: the pro-
map n𝑆 = [𝑑𝑇 , 𝑑 𝑈 , 𝑑𝑄 ] ⇐ and the second term regularizes the solution jection direction (ProjDirections[]) and the orthographic view
towards the depth 𝑏 (𝑐, 𝑈) rendered from the existing mesh 𝑊𝑆 . We and projection matrices (ProjViewMats[], ProjProjMats[]), de-
adapt the normal integration method proposed by Cao et al. [2022] rived from the camera’s pose for that view. Finally, pre-rendered
to perform this minimization. Their core contribution is a bilaterally depth maps (DepthMaps[]) from each viewpoint are crucial for han-
weighted functional designed to handle potential depth discontinu- dling occlusions during the blending process. With these inputs
ities inherent in surfaces estimated from normal maps. Instead of prepared, the fragment shader executes the logic outlined in Algo-
assuming a globally smooth surface, their method operates under rithm 1 for each surface fragment.
the semi-smooth surface assumption, allowing for one-sided discon- The core of the blending logic within Algorithm 1 lies in calcu-
tinuities. They introduce bilateral weights 𝑖𝑍 (𝑐, 𝑈) and 𝑖 𝑎 (𝑐, 𝑈) at lating an appropriate weight (weight) for each texture’s potential
each pixel (𝑐, 𝑈) during optimization. These weights re#ect the local contribution. This weighting is carefully designed to ensure high-
surface continuity; values close to 0.5 indicate local smoothness in quality results:
the respective direction (horizontal for 𝑖𝑍 , vertical for 𝑖 𝑎 ), while • View-dependent Alignment: The smoothstep(0.3, 1.0,
values approaching 0 or 1 suggest a likely discontinuity boundary. ...) function applied to the cosine similarity between the
surface normal and projection direction ensures that views
Filtering of Unreliable Depth Estimate. The depth map 𝐿 (𝑐, 𝑈), ob- nearly perpendicular to the surface contribute strongly, while
tained by minimizing Eq. (6), represents the re!ned geometry 𝑎𝑆 . contributions smoothly fall o" to zero for views at grazing
However, areas near depth discontinuities, indicated by 𝑖𝑍 or 𝑖 𝑎 angles (beyond approximately 72.5↘ ), preventing artifacts
deviating from 0.5, can lead to unreliable depth estimates in 𝐿 (𝑐, 𝑈). from oblique projections.
Directly using these unreliable values during the Poisson surface • Re!nement Priority: By drastically reducing the weight of
reconstruction could introduce geometric artifacts when merging unre!ned textures (⇓10 ↓8 ) compared to re!ned ones (⇓1.0),
𝑎𝑆 with the mesh 𝑊𝑆 . the algorithm ensures that the high-quality details introduced
To mitigate this, we !lter the depth map 𝐿 (𝑐, 𝑈) based on the by HFS-SDEdit are preferentially used in the !nal texture
continuity weights 𝑖𝑍 and 𝑖 𝑎 computed during the optimization wherever a re!ned view provides relevant, visible informa-
process. We identify pixels (𝑐, 𝑈) where the estimated surface is tion.
potentially unreliable or discontinuous by checking if either weight • Occlusion and Transparency: Setting the weight to zero for
signi!cantly deviates from the ideal smooth value of 0.5. We mark occluded fragments (based on depth map comparison using
a pixel as unreliable if: the conceptual checkOcclusion function) or for fragments
projecting onto transparent background regions of the source
𝑖𝑍 (𝑐, 𝑈) < 0.4 or 𝑖𝑍 (𝑐, 𝑈) > 0.6, or textures (via the conceptual sampleTexture function check-
(7) ing alpha) prevents projecting incorrect colors or background
𝑖 𝑎 (𝑐, 𝑈) < 0.4 or 𝑖 𝑎 (𝑐, 𝑈) > 0.6.
details onto the foreground mesh surface.
A binary mask is then created, keeping only those pixels where This combination of projection (represented conceptually by the
both 𝑖𝑍 (𝑐, 𝑈) and 𝑖 𝑎 (𝑐, 𝑈) fall within the con!dence interval [0.4, project function), visibility checks, and carefully designed weight-
0.6]. This mask highlights regions deemed to be reliably estimated ing allows the shader to synthesize a seamless, UV-free texture
and locally smooth. To further ensure robustness and remove poten- on the !nal mesh, integrating the best available information from
tially isolated unreliable pixels or thin artifacts near discontinuity multiple viewpoints.
boundaries, this binary mask is processed with morphological ero-
sion using a 3 ⇓ 3 kernel. The !nal eroded mask de!nes the reliable A.3 Further Technical Details
region of the depth map 𝐿 (𝑐, 𝑈) that is subsequently used in the Pois- Texture Re!nement. To re!ne the texture at a target view 𝑈𝑆 , we
son surface reconstruction step to update the mesh 𝑊𝑆 , ensuring a adopt several techniques from the state-of-the-art mesh texturing
cleaner and more robust integration of the re!ned geometry. The method Paint3D [Zeng et al. 2024] in the texture re!nement process
seams visible in Fig. 5-b of the main paper are a direct result of this with HFS-SDEdit. Speci!cally, we implement a multi-view depth-
!ltering process of removing pixels around detected discontinuities. aware texture sampling strategy by horizontally concatenating the

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
ii • Ryu et al.

ALGORITHM 1: Projection Mapping Fragment Shader Logic


Input: fragPosition, fragNormal // Fragment attributes
Data: Textures[], DepthMaps[], NumTextures, IsRe!ned[] // Texture data
Data: ProjDirections[], ProjViewMats[], ProjProjMats[] // Projection parameters
Data: Epsilon // Depth tolerance
Output: !nalColor // Output fragment color
accumColor ⇔ (0, 0, 0)
totalWeight ⇔ 0.0
normFragNormal ⇔ normalize(fragNormal)
for 𝑏 ⇔ 0 to NumTextures ↓1 do
// Initialize weight for texture j
weight ⇔ 1.0
// Calculate view-dependent alignment weight
alignment ⇔ dot(normFragNormal, ProjDirections[j])
alignmentWeight ⇔ smoothstep(0.3, 1.0, clamp(alignment, 0.0, 1.0))
weight ⇔ weight ⇓ alignmentWeight
// Modulate weight based on refinement status
if IsRe!ned[j] then
// Keep weight if refined
weight ⇔ weight
else
// Down-weight if unrefined
weight ⇔ weight ⇓ 1e-8
end
// Project fragment and get texture coordinates
fragTexUV ⇔ project(fragPosition, ProjViewMats[j], ProjProjMats[j])
// Perform occlusion test using depth maps
isOccluded ⇔ checkOcclusion(fragPosition, fragTexUV, DepthMaps[j], ProjViewMats[j], ProjProjMats[j], Epsilon)
if isOccluded then
weight ⇔ 0.0
end
// Sample texture and check transparency
texColor ⇔ sampleTexture(Textures[j], fragTexUV)
if [Link] < 1.0 then
// This case indicates backround in our implementation
weight ⇔ 0.0
end
// Accumulate weighted color
if weight > 0 then
accumColor ⇔ accumColor + [Link] ⇓ weight
totalWeight ⇔ totalWeight + weight
end
end
if totalWeight > 1e-8 then
![Link] ⇔ accumColor / totalWeight
![Link] ⇔ 1.0
else
// Default color (e.g., black)
!nalColor ⇔ (0, 0, 0, 0)
end
return !nalColor

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • iii

renderings from the initial viewpoint 𝑈 0 and the target viewpoint 𝑈𝑆 , Table S1. Parameter Sweep of 𝑂 in 𝑑𝑃 and replacement threshold 𝑀 stop .
where the camera path is de!ned as V = {𝑈 0, . . . , 𝑈𝑅 }. We render Starting the denoising process from 𝑒 = 30, the swapping is performed
until 𝑀 stop . The table presents the refinement results according to various
depth, RGB, and re!nement mask images, resulting in corresponding
combinations of 𝑂 and 𝑀 stop values. Full-reference metrics include PSNR,
depth, RGB, and mask grid images. In the mask grid, only the right SSIM, and LPIPS, while non-reference metrics include MUSIQ, QAlign, LIQE,
half is used as the re!nement mask. Subsequently, we perform multi- and TOPIQ.
view depth-aware texture re!nement with HFS-SDEdit. For the
initial viewpoint 𝑈 0 , we render the next planned view for re!nement
Full-Reference Metrics Non-Reference Metrics
but utilize only the re!ned image corresponding to 𝑈 0 . For depth 𝑇 𝑁 stop
PSNR ⇒ SSIM ⇒ LPIPS ⇑ MUSIQ ⇒ QAlign ⇒ LIQE ⇒ TOPIQ ⇒
conditioning, we employ a variant of FLUX2 . 22 12.8984 0.3144 0.6427 67.1750 4.5764 4.0468 0.5798
20 13.5825 0.3353 0.6215 58.0347 4.1461 3.2241 0.4679
Re!nement View Selection. Inspired by the automatic camera selec- 2
18 14.2574 0.3562 0.6136 49.3993 3.6700 2.5849 0.3793
tion scheme introduced in the mesh texturing literature, Text2Tex [Chen
16 14.8523 0.3739 0.6084 41.9640 3.2286 2.0822 0.3057
et al. 2023], we choose the camera path V = {𝑈 0, . . . , 𝑈𝑐 , . . . , 𝑈𝑅 }
as follows. We begin by de!ning a sparse camera schedule with 22 14.1419 0.3514 0.6020 54.7414 4.1851 3.0206 0.4206
polar angles [45↘, 90↘, 135↘ ] and corresponding azimuthal angles 4 20 14.8375 0.3702 0.5980 45.7231 3.6959 2.4237 0.3354
[0↘, 45↘, 180↘, 270↘ ]. Since an orthographic camera is used, these 18 15.5876 0.3906 0.5982 39.5193 3.3370 2.1052 0.2828
viewpoints cover most of the visible regions for re!nement. How- 16 16.2478 0.4057 0.5969 34.8861 3.0465 1.8206 0.2390
ever, we employ an automatic camera selection process to address 22 16.2883 0.4014 0.6864 31.1365 3.0541 1.4288 0.2341
areas that remain unre!ned. We sample 100 views from the sphere. 16 20 17.0381 0.4210 0.6840 28.2405 2.7874 1.3988 0.2074
For each viewpoint 𝑈𝑐 , we render the object to obtain a foreground 18 17.7739 0.4390 0.6806 26.0664 2.6046 1.3661 0.1913
mask 𝑍 fg and the re!nement mask 𝑍. We also compute a cosine 16 18.3762 0.4529 0.6752 25.1298 2.4729 1.3231 0.1845
similarity map 𝑍 cos to evaluate viewpoint quality. We then calculate
the ratio 𝑗𝑐 as
𝑍𝑆 ≃ 𝑍 cos Table S2. Comparison with ProlificDreamer. Elevate3D consistently
𝑗𝑐 = achieves the be#er scores across various quality metrics, highlighting its
𝑍 fg high-quality, visually appealing 3D refinement.
for each viewpoint and select the one with the highest ratio for
further re!nement, ensuring we prioritize the most informative Full-Reference No-Reference
Method
angle. Finally, when the ratio across the object falls below 0.02, we PSNR⇒ SSIM⇒ LPIPS⇑ MUSIQ⇒ LIQE⇒ TOPIQ⇒ Q-Align⇒
conclude that su$cient coverage has been achieved and terminate LQ (Baseline) 33.202 0.966 0.057 61.241 2.069 0.457 2.690
the re!nement process. Ours 26.163 0.941 0.070 66.527 2.774 0.529 3.215
Proli!cDreamer 18.037 0.904 0.156 57.307 1.989 0.439 1.893
Geometry Re!nement Detail. Section 4.2 of the main paper de-
scribes inferring the normal map n𝑆 from the re!ned texture 𝑋𝑆↔ . In
practice, our implementation incorporates an additional normal Table S3. !antitative Results on the Ablation. We compare with the
blending step prior to normal integration to further enhance the cases where only texture or geometry refinement is performed.
robustness of the geometry re!nement stage. Speci!cally, before
minimizing the energy functional, we blend two normal maps: the Model
Full-Reference
PSNR⇒ SSIM⇒ LPIPS⇑ MUSIQ⇒
No-Reference
LIQE⇒ TOPIQ⇒ Q-Align⇒
Normal FID⇑

map n𝑆 , inferred from 𝑋𝑆↔ that captures !ne, texture-consistent de- LQ (Baseline) 33.202 0.966 0.057 61.241 2.069 0.457 2.690 60.786
tails, and a base normal map representing the current geometry of FULL
Only Geometry Re!nement
26.163
28.365
0.941
0.954
0.070
0.068
66.527
61.086
2.774
2.032
0.529
0.465
3.215
2.599
52.195
48.043
the mesh 𝑊𝑆 . This blending, performed using the UDN blending Only Texture Re!nement 26.393 0.939 0.067 68.717 3.188 0.572 3.454 55.465

method [Barré-Brisebois and Hill. 2012], helps combine the details


from the re!ned texture’s normal map with the underlying structure Table S4. Experimental Results with a Di"erent Di"usion Backbone.
of the existing mesh geometry. The full-reference metrics are evaluated against the high-quality source
images. LQ (Baseline) means the low-quality reference images degraded
B Additional Experiments from the high-quality source images.

B.1 The E"ect of the Parameters in HFS-SDEdit


Full-Reference No-Reference
Model
As discussed in the main paper, HFS-SDEdit employs a Gaussian PSNR⇒ SSIM⇒ LPIPS⇑ MUSIQ⇒ QAlign⇒ LIQE⇒ TOPIQ⇒
kernel 𝑆𝑂 , and the high-frequency replacement is performed until LQ (Baseline) 20.701 0.521 0.662 21.918 2.018 1.303 0.159
it reaches the time step 𝑁 stop . To analyze the e"ect of di"erent pa- SDEdit (strength=0.4) 18.982 0.457 0.689 23.090 2.204 1.316 0.155
SDEdit (strength=0.8) 14.809 0.372 0.741 31.714 2.854 1.446 0.224
rameter choices for 𝑇 and 𝑁 stop in HFS-SDEdit, we conduct the same NC-SDEdit 17.344 0.424 0.717 24.730 2.347 1.274 0.168
low-quality image re!nement experiment presented in the main Ours 15.853 0.376 0.601 41.754 3.415 1.876 0.295

paper. For assessing re!nement !delity, we utilize full-reference met-


rics, including PSNR, SSIM [Wang et al. 2004], and LPIPS [Zhang
et al. 2018]. To evaluate perceptual quality, we employ non-reference TOPIQ [Chen et al. 2024a], and Q-Align [Wu et al. 2024b]. As shown
metrics such as MUSIQ [Ke et al. 2021], LIQE [Zhang et al. 2023c], in Table S1, for a !xed 𝑇, increasing the number of replacement steps
improves !delity, leading to better full-reference metrics. However,
2 [Link] this also increases the risk of preserving degraded details from the

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
iv • Ryu et al.

Table S5. !antitative Analysis of Fig. 10 in the Main Paper. We com- geometry and texture metrics, while geometry-only and texture-
pare the quantitative results of the objects shown in Fig. 10 in the Main only re!nements perform well in their respective domains but not
Paper.
both. This highlights our method’s holistic improvement across
both geometry and texture aspects.
Full-Reference No-Reference
Method
PSNR⇒ SSIM⇒ LPIPS⇑ CLIP Sim⇒ MUSIQ⇒ LIQE⇒ TOPIQ⇒ Q-Align⇒
LQ (Baseline) 32.086 0.940 0.096 88.640 61.203 2.022 0.429 2.577
Ours
DreamGaussian
25.342
31.559
0.902
0.942
0.098
0.086
90.577
88.387
67.674
61.669
2.756
2.083
0.540
0.457
3.277
2.704
B.4 Experimental Results with a Di"erent Di"usion
Backbone
To examine the robustness of HFS-SDEdit with respect to the di"u-
Table S6. !antitative Comparison on 3D Refinement Using Full-
reference metrics. DreamGaussian [Tang et al. 2024b] refines only textures
sion backbone, we repeated our main experiment using a less pow-
while other methods refine both texture and geometry. erful model, Stable Di"usion 3.5 medium [Esser et al. 2024], under
identical experimental conditions. Results in Table S4 re#ect similar
trends to the main experiments, demonstrating improvements in
Model PSNR⇒ SSIM⇒ LPIPS⇑ CLIP⇒
non-reference metrics and LPIPS [Zhang et al. 2018] scores.
LQ (Baseline) 33.202 0.966 0.057 90.688
Ours 26.163 0.941 0.070 89.886
DreamGaussian [Tang et al. 2024b] 32.720 0.965 0.056 90.890
DiSR-NeRF [Lee et al. 2024] 30.294 0.949 0.088 88.696
B.5 !antitative Analysis of Fig. 10 in the Main Paper
MagicBoost [Yang et al. 2024c] 23.865 0.934 0.106 83.412 Fig. 10 in the main paper qualitatively demonstrates substantial
Proli!cDreamer [Wang et al. 2023] 18.037 0.904 0.156 78.509 visual enhancements achieved by our method over DreamGaus-
sian [Tang et al. 2024b]. However, the corresponding quantitative
analysis using non-reference metrics in Table S5 reveals less dra-
low-quality image, lowering the performance on non-reference met- matic numerical di"erences. This highlights a common limitation
rics. Conversely, increasing 𝑇 allows a broader band of frequencies where such metrics may not fully re#ect perceived visual quality
to be preserved for a !xed number of replacement steps. While this improvements. Despite this, the scores do con!rm a relative im-
improves !delity and enhances the full-reference metrics, it also provement provided by our approach.
risks incorporating low-quality domain information from the low-
quality reference image, thereby negatively impacting non-reference
metrics. The parameter combination of 𝑇 = 4 and 𝑁 stop = 18, used B.6 !antitative Comparison of 3D Refinement with Full
in the main experiment, is observed to provide a balance, achieving Reference Metrics
both high-!delity and high-quality re!nement. We compare our method with recent 3D model re!nement ap-
proaches in terms of various full reference metrics: PSNR, SSIM [Wang
B.2 Comparison with ProlificDreamer et al. 2004], LPIPS [Zhang et al. 2018], and CLIP similarity [Radford
We provide a comparison between our method and Proli!cDreamer [Wang et al. 2021]. As summarized in Table S6, we see that when the initial
et al. 2023]. While Proli!cDreamer was initially developed for text- coarse 3D input is already of moderate !delity, these metrics yield
to-3D synthesis, it can also be applied to re!ne existing 3D mod- similar scores across di"erent re!nement methods. Thus, these met-
els due to its VSD-loss-based framework. We use the geometry rics may not adequately capture actual perceptual enhancements
and texture re!nement stages of Proli!cDreamer for re!nement. for the generative re!nement task that involves generating details
As shown in Table S2, our method signi!cantly outperforms Pro- absent in the reference.
li!cDreamer across all evaluated metrics, including MUSIQ [Ke
et al. 2021], LIQE [Zhang et al. 2023c], TOPIQ [Chen et al. 2024a],
and Q-Align [Wu et al. 2024b]. Proli!cDreamer showed lower !- B.7 Robustness to Normal Prediction Failures
delity re!nements due to the lack of explicit !delity constraints. After the input image (Fig. S1-a) is re!ned with HFS-SDEdit (Fig. S1-
Quantitatively, Proli!cDreamer exhibits substantially lower PSNR, b), we extract extract geometric cues from the re!ned image using an
SSIM [Wang et al. 2004], and higher LPIPS [Zhang et al. 2018] val- o"-the-shelf monocular normal estimator [Martin Garcia et al. 2025].
ues. Qualitatively, we can see in Fig. S3 that direct VSD application However, the normal estimator may occasionally produce failure
often results in artifacts, such as multi-faced Janus e"ects, due to cases. A common failure mode produces overly smooth or detail-
the absence of explicit !delity constraints. less predicted normal maps (Fig. S1-c). Directly Integrating such a
compromised normal map produces #at, featureless surfaces that
B.3 !antitative Ablation Study of Texture and Geometry negate the purpose of re!nement (Fig. S1-e). Our method, however,
Refinement is robust to such failures due to the regularization term in our
We performed a quantitative ablation study to evaluate contributions energy functional Eq. (6). This term explicitly encourages the re!ned
from geometry and texture re!nements as shown in Table S3. For surface 𝑎𝑆 to remain close to the depth map derived from the input
geometry evaluation, we computed the FID between high-quality coarse geometry 𝑊𝑆 (Fig. S1-d). This allows our method to produce
normal maps and those from each re!nement method. The full meaningfully re!ned geometry (Fig. S1-e) without severe distortions,
re!nement achieves balanced and competitive results across both avoiding catastrophic failure during the re!nement iterations.

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • v

B.8 Additional 2D Refinement Results


In Fig. S2, we show more qualitative comparsion results comparing
SDEdit [Meng et al. 2022], NC-SDEdit [Yang et al. 2024a], and HFS-
SDEdit. Our method shows a good balance between !delity and
quality.

B.9 Additional 3D Refinement Results


In this section, we present additional qualitative examples. In Fig. S3,
we show more qualitative comparison results for the re!nement re-
sults of the degraded GSO [Downs et al. 2022] dataset. In Fig. S4, we
(a) Initial Image (b) Refined Image (c) Predicted Normal Map further show more qualitative results for re!ning the 3D generation
results of TRELLIS [Xiang et al. 2024]. We also provide an additional
supplementary video. We strongly suggest the readers also see
the video for a better understanding of our model’s output quality.

(d) Coarse Geometry (e) Geometry Refined w/ (e) Normal Integration


Regularization w/o Regularization
Fig. S1. Robustness to Normal Prediction Failure. Even when the nor-
mal predictor generates poor results, our regularized normal integration
successfully preserves the coarse geometry structure, preventing severe
distortion.

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
vi • Ryu et al.

(a) Before (b) After (c) HFS-SDEdit (d) NC-SDEdit (e) SDEdit (f) SDEdit
Degradation Degradation (Ours) (Strength = 0.4) (Strength = 0.8)

Fig. S2. Additional !alitative Comparison on 2D Image Refinement. The refinement results in (c), (d), (e), and (f) are obtained from the low-quality
image in (b), which was degraded from the image in (a).

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
Elevating 3D Models: High-!ality Texture and Geometry Refinement from a Low-!ality Model • vii

(f) Ours (Elevate3D)

Fig. S3. Additional !alitative Comparison on 3D Refinement We show additional 3D refinement comparisons on the degraded GSO dataset. Among all
methods, our model produces the highest quality textures with a well-aligned geometry.
SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.
viii • Ryu et al.

(a) Input Image (b) Generated Result (TRELLIS) (c) Refined Result (Elevate3D)

Fig. S4. Additional !alitative Comparison on TRELLIS Outputs. We show additional examples of refining 3D generation results of TRELLIS where
it generates the 3d in (b) taking real world input image in (a). As seen in (c), our method produces realistic textures and accurate geometry, resulting in
high-quality refinements.

SIGGRAPH Conference Papers ’25, August 10–14, 2025, Vancouver, BC, Canada.

Common questions

Powered by AI

Elevate3D demonstrates computational efficiency similar to that of MagicBoost and is significantly faster than DiSR-NeRF, despite being slower than DreamGaussian, which only focuses on texture refinement. This comparison indicates that while Elevate3D performs more extensive geometry and texture refinement than DreamGaussian, it achieves reasonable efficiency and speed compared to other comprehensive refinement methods, making it practical for high-quality model generation .

NC-SDEdit's retention of low-frequency components results in carrying over low-quality domain information, leading to lower fidelity and perceptual quality in outputs. In contrast, HFS-SDEdit uses high-frequency guidance to avoid this issue, achieving superior results. The low-frequency retention in NC-SDEdit inadvertently incorporates undesired domain characteristics, while HFS-SDEdit's method results in refined images with higher quality and consistency due to its focus on high-frequency components .

ChatGPT Vision is used in the Elevate3D evaluation process to generate descriptions for images from the LSDIR dataset. These descriptions are part of the process of validating various aspects of the model, including the impact of domain information encoded in the low-frequency component. By generating relevant text prompts, ChatGPT Vision aids in assessing how well the model captures and enhances key visual features, contributing to a comprehensive evaluation of refinement effectiveness .

Elevate3D differs from previous methods by utilizing a geometry refinement strategy that leverages refined textures, resulting in consistent high-quality textures and geometries. Previous approaches often result in blurry textures and geometries less consistent with textures, while Elevate3D achieves high texture-geometry consistency. This is accomplished by alternating between texture and geometry refinement in a view-by-view fashion, surpassing the quality of prior methods. This approach ensures that the refined models exhibit superior detail and accuracy .

Models like TRELLIS often fail to produce high-quality results when dealing with inputs outside the domain of their training datasets, which mainly consist of synthetic objects. Elevate3D addresses these issues by enhancing the results from TRELLIS, even for out-of-domain inputs, through its superior texture and geometry refinement capabilities. This leads to enhanced 3D models with better detail and realism, overcoming the limitation of domain-specific training data .

Regularized normal integration improves Elevate3D’s geometry refinement by addressing inconsistencies between predicted normals from refined textures and the input geometry. By incorporating regularization, the geometry refinement process resolves severe distortions that arise when using solely refined texture normals. This enhancement results in higher-quality, distortion-free geometries, contributing to the overall consistency and accuracy of Elevate3D’s refined 3D models .

The table of quantitative evaluations highlights that Elevate3D consistently outperforms other methods across multiple image quality metrics such as MUSIQ, LIQE, TOPIQ, and Q-Align. This performance superiority is attributed to its advanced texture and geometry refinement strategies, which lead to better-rendered image quality and improved 3D model details. Elevate3D's refined models exhibit higher marks in quality assessments, showcasing its effectiveness compared to traditional state-of-the-art methods .

The need for both texture and geometry refinement in the Elevate3D framework is justified through ablation studies that demonstrate the inadequacies of performing only one type of refinement. Without geometry refinement, models retain the crude geometry of input models, and without texture refinement, the geometry refinement relies on low-quality input textures, resulting in minimal improvement. Combining both refinements ensures high-quality textures and geometries that are consistent and well-aligned, which maximizes the visual and structural quality of the 3D models .

In image enhancement techniques, varying strength values create a trade-off between fidelity and quality. Lower strength values result in better fidelity, maintaining alignment with the reference image metrics. However, this decreases no-reference scores, indicating a drop in perceived quality. Conversely, higher strength values enhance perceived quality, as indicated by improved no-reference scores, but at the expense of fidelity to the original content. This trade-off highlights the inherent challenge of balancing perception with technical accuracy in image refinement processes .

The effectiveness of manipulating the low-frequency component in diffusion sampling for image enhancement is validated through experiments that involve replacing the low-frequency component of a sample's latent representation with that from a low-quality reference. This replacement improved the quality of the sampled images by embedding domain-specific information, as the low-frequency components encode critical domain information. The experiment confirmed that incorporating the domain information via low-frequency retention significantly impacts the perceptual quality of the generated samples .

You might also like