0% found this document useful (0 votes)
12 views13 pages

MicroDiffusion for 3D Reconstruction

MicroDiffusion is a novel method for reconstructing high-quality, depth-resolved 3D volumes from limited 2D microscopy projections using a combination of Implicit Neural Representations (INR) and Denoising Diffusion Probabilistic Models (DDPM). This approach enhances the fidelity of 3D reconstructions by integrating structural coherence from INR with detail enhancement from DDPM, significantly improving image quality over traditional methods. The method has shown remarkable performance in various optical microscopy datasets, demonstrating its potential for advancing volumetric imaging in biological and medical applications.

Uploaded by

김민솔
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views13 pages

MicroDiffusion for 3D Reconstruction

MicroDiffusion is a novel method for reconstructing high-quality, depth-resolved 3D volumes from limited 2D microscopy projections using a combination of Implicit Neural Representations (INR) and Denoising Diffusion Probabilistic Models (DDPM). This approach enhances the fidelity of 3D reconstructions by integrating structural coherence from INR with detail enhancement from DDPM, significantly improving image quality over traditional methods. The method has shown remarkable performance in various optical microscopy datasets, demonstrating its potential for advancing volumetric imaging in biological and medical applications.

Uploaded by

김민솔
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MicroDiffusion: Implicit Representation-Guided Diffusion for 3D

Reconstruction from Limited 2D Microscopy Projections

Mude Hui1 * Zihao Wei2 * Hongru Zhu3 Fei Xia4† Yuyin Zhou1†
1
University of California, Santa Cruz 2 University of Michigan, Ann Arbor
3
Johns Hopkins University 4 Ecole Normale Supérieure de Paris
arXiv:2403.10815v1 [[Link]] 16 Mar 2024

muhui@[Link] zihaowei@[Link] {hongruzhu95, zhouyuyiner}@[Link] fx43@[Link]

Abstract (a) Conventional 3D laser scanning microscopy:


slow but depth-resolved
Acquired 3D stack
Volumetric optical microscopy using non-diffracting Laser beam
beams enables rapid imaging of 3D volumes by projecting x
z
them axially to 2D images but lacks crucial depth informa-
tion. Addressing this, we introduce MicroDiffusion, a pi- (b) Volumetric microscopy with non-diffracting laser beam:
oneering tool facilitating high-quality, depth-resolved 3D fast but depth-unresolved
volume reconstruction from limited 2D projections. While Laser beam Acquired 2D images
existing Implicit Neural Representation (INR) models of- x
z
ten yield incomplete outputs and Denoising Diffusion Prob-
abilistic Models (DDPM) excel at capturing details, our
method integrates INR’s structural coherence with DDPM’s (c) MicroDiffusion-enabled volumetric microscopy with non-diffracting
Laser beam beam: fast and depth-resolved
fine-detail enhancement capabilities. We pretrain an INR x
Reconstructed
model to transform 2D axially-projected images into a pre- z Acquired 2D images 3D volume
liminary 3D volume. This pretrained INR acts as a global MicroDiffusion
prior guiding DDPM’s generative process through a linear
interpolation between INR outputs and noise inputs. This
strategy enriches the diffusion process with structured 3D Proposed method

information, enhancing detail and reducing noise in local-


ized 2D images. By conditioning the diffusion model on Figure 1. Background and concept of MicroDiffusion-enabled
the closest 2D projection, MicroDiffusion substantially en- volumetric microscopy. (a) Conventional 3D laser scanning mi-
hances fidelity in resulting 3D reconstructions, surpassing croscopy, while depth-resolvable due to its point-scanning 3D data
acquisition scheme, suffers from slow imaging speed. (b) Volu-
INR and standard DDPM outputs with unparalleled image
metric microscopy using a non-diffracting laser beam provides fast
quality and structural fidelity. Our code and dataset are
volumetric imaging by axially projecting 3D volumes onto 2D im-
available at [Link] VLAA/ ages but lacks depth information within each acquired 2D image.
MicroDiffusion. (c) Our proposed MicroDiffusion model is employed as a digi-
tal backend for 3D volumetric reconstruction from 2D projections
acquired in (b). MicroDiffusion significantly enhances volumet-
1. Introduction ric imaging performance, providing a synergistic balance between
imaging speeds and depth-resolving capabilities.
Volumetric optical imaging has emerged as a pivotal tool in
biological and medical domains, enabling precise 3D visu-
alization of intricate structures with unprecedented tempo- scanning methods (Fig. 1(a)). This limitation not only re-
ral resolution [13, 46]. Despite its high spatial resolution, stricts clinical diagnosis mostly to 2D imaging, potentially
the predominant approach in optical microscopy, reliant on compromising diagnostic accuracy [6, 12], but also impedes
3D laser scanning, suffers from suboptimal temporal res- the observing of dynamic 3D biological processes [14, 18].
olution due to the slow data acquisition inherent in point- Recent advancements using non-diffracting beams have
* Denotes equal first author expedited laser scanning microscopy by optically project-
† Denotes equal senior author ing 3D volumes as 2D projections for volumetric imag-

1
ing [8, 30, 42]. However, this approach sacrifices depth in- information and imaging speed.
formation within each 2D snapshot [3, 43] (see Fig. 1(b)),
necessitating the development of tools capable of recon- 2. Related Works
structing depth for accurate 3D reconstruction from 2D im-
Laser scanning microscopy with non-diffracting beams
ages. In this paper, we aim to reconstruct 3D volumes from
for volumetric imaging. Laser scanning microscopy has
limited 2D projections obtained through such volumetric
emerged as the gold standard in biomedical imaging. Com-
imaging, striving to expedite optical volumetric imaging
monly used in biomedical applications, imaging modalities
without compromising depth resolvability of 3D volumes.
such as multiphoton microscopy [14], optical coherence mi-
Existing 3D reconstruction methods, such as Implicit croscopy [21], and photoacoustic microscopy [2] all share
Neural Representations (INR) [38], offer a comprehen- at least one laser scanning mode. However, a critical chal-
sive global view from given 2D microscopy projections by lenge in laser scanning microscopy lies in accelerating data
mapping coordinates to a holistic 3D volume using neu- acquisition without sacrificing resolution and depth, espe-
ral networks. However, direct 3D reconstruction using cially with the most widely used point-scanning methods
INR often yields globally coherent yet visually blurry out- that scan a tightly focused 3D laser point to collect volu-
puts, lacking local details. This limitation in spatial res- metric data (see Fig. 1(a)).
olution may stem from the limited number of 2D images To address this, optical strategies using non-diffracting
acquired. Conversely, Denoising Diffusion Probabilistic laser beams, notably Bessel and Airy beams, have been
Models (DDPM) [17], especially with the U-Net architec- proposed. These beams generate an elongated and almost
ture [32], excel in detailed generative modeling, manag- uniform axial point spread function. Scanning with non-
ing spatial hierarchies, and preserving fine-grained details. diffracting beams essentially captures multiple axial layers
Building upon these strengths, we propose MicroDiffusion, in a single lateral scan, as opposed to the point-scanning
a hybrid approach integrating INR’s global structural coher- method (see Fig. 1(b)). For instance, they have been uti-
ence with DDPM’s detail enhancement capabilities. lized for rapid volumetric imaging, enabling the real-time
MicroDiffusion encompasses two key designs as shown capture of dynamic biological processes [2, 3, 30]. Despite
in Fig. 2: 1) INR pretraining, which transforms 2D projec- their speed advantages, these strategies often compromise
tions into a preliminary 3D volumetric output, establishing depth information, yielding 2D projections without detailed
a global structure; and 2) implicit representation-guided information on feature depths. To combine the speed bene-
diffusion, where the pretrained INR acts as a global prior fits of non-diffracting beam scanning methods with the ca-
guiding a diffusion model, enhancing details and reducing pability to discern depth information, we propose a deep
noise in local 2D projections within the 3D volume. Specifi- learning model for this inference (see Fig. 1(c)).
cally, MicroDiffusion employs a linear interpolation of INR
model output with the noise input, rather than starting from Implicit Neural Representations (INR). Implicit neural
a conventional Gaussian noise baseline, to enrich the diffu- representations excel at modeling the forms of 3D objects,
sion process with structured 3D information. Furthermore, generating surfaces for 3D scenes, and capturing detailed
this step enhances image fidelity by conditioning image and 3D structures. Pioneering work such as GQN [9] utilizes
positional embeddings extracted from the closest projec- a generative query network to learn scene representations
tions. Thus, MicroDiffusion generates 3D reconstructions from multiple perspectives. Building on this foundation,
faithfully representing original optical microscopy images. Mildenhall et al. [25] introduce the seminal concept of Neu-
Comprehensive experiments on three optical microscopy ral Radiance Fields (NeRF), which use a multi-layer percep-
datasets showcase MicroDiffusion’s efficacy. Compared to tron to encode 3D scenes for view synthesis. Other works,
the baseline INR, it notably enhances reconstruction qual- such as Poly-INR [37], SIREN [34] and LIIF [4], have em-
ity by up to 15.5% in SSIM, 15.2% in PSNR, and 64.7% in ployed periodic activation functions, significantly enhanc-
DICE on Dendrite dataset, up to 15.0% in SSIM, 3.0% in ing the quality and adaptability of image representations.
PSNR, and 0.3% in DICE on Vasculature dataset, and up to Parallel to these advancements, the application of INR in
1.8% in SSIM, 0.8% in PSNR, and 4.7% in DICE on Neu- medical imaging has shown remarkable potential. For ex-
ron dataset. The resulting 3D stacks demonstrate remark- ample, ARSSR [44] and CoIL [40] have adapted NeRF-like
able resolution, delineating individual dendrites (less than methods for super-resolution in medical images. NeRP [35]
1µm) and preserving coherent 3D structures—an achieve- distinctively combines the inherent image information with
ment unattainable by the naive DDPM approach. These the physics of sparse measurements to enhance medical im-
tangible outcomes establish MicroDiffusion as a pioneer- age reconstruction. Cryodrgn [47] and fpm-inr [48] are no-
ing framework for reconstructing high-quality 3D volumes table for reconstructing 3D volumes from 2D microscopy
from 2D projections in volumetric microscopy using non- images. As a recent advancement, IDM [10] integrates INR
diffracting beams, reconciling the trade-off between depth and diffusion models by employing INR as the decoder of

2
a diffusion model. In contrast, we leverage INR to gener- ric imaging, 2D projections are axially downsampled by a
ate continuous and interpretable 3D representations used as factor of n in Mi , resulting
Pn in each projection Xi being ex-
guidance for a diffusion model. pressed as Xi = n1 k=1 mik . Hence, the problem is to find
a model f : {Xi } → M to reconstruct depth-resolved 3D
Diffusion Models. Diffusion models are currently at the volumes from downsampled 2D projections.
forefront of generative model innovation. The Denois-
ing Diffusion Probabilistic Model (DDPM) [16] can in- 4. Method
crementally convert Gaussian noise into coherent signals.
Subsequent research has expanded on controlling the out- In this section, we begin by revisiting key concepts of Im-
put of these models, primarily categorized into classifier- plicit Neural Representations (INR) and present our INR
guidance [7] and classifier-free guidance [15, 31]. Re- design crafted for optical microscopy reconstruction in
cent studies demonstrate the versatility of diffusion mod- Sec. 4.1. We then delve into MicroDiffusion, our implicit
els in creating content guidance from a variety of sources, representation-guided diffusion model in Sec. 4.2.
including images, text, depth, video, and their combina-
tions [1, 11, 20, 29, 31, 33]. 4.1. Implicit Neural Representation
In 3D reconstruction, considerable efforts are made to Revisit INR for 3D Reconstruction. INR methods uti-
produce 3D models from text prompts or 2D references [5, lize a function, typically a Multilayer Perceptron (MLP) de-
19, 22, 23, 27, 36, 41, 45]. The approach most similar to noted as f_{\text {inr}} , to implicitly represent a 3D field. f_{\text {inr}} operates
ours is Magic123 [28], which utilizes a two-stage, coarse- over continuous 3D space and maps coordinates to a pre-
to-fine framework to generate 3D models with reference im- dicted property, like intensity or occupancy, formulated as
ages. MicroDiffusion differs from Magic123 in several as-
pects. First, Magic123 employs pretrained knowledge from m_{inr} = f_{\text {inr}}(p(z)), (1)
models like Stable Diffusion [27, 31] or Zero1-to-3 [24]
to generate reference views for training Neural Radiance where {z} denotes normalized 3D coordinate within the range
Fields. In contrast, our MicroDiffusion has no such pre- [-1, 1] to ensure uniformity across the input space. p(\cdot )
trained knowledge, and we must directly train the Implicit denotes the positional encoding that transforms 3D coordi-
Neural Representation (INR) from projection. Second, dif- nates into a higher-dimensional space, crucial for capturing
fusion in the Magic123’s fine stage is applied solely to im- high-frequency details during reconstruction. minr repre-
prove the mesh generated from NeRF, without considering sents the property such as the intensity at position z.
the NeRF’s information. In MicroDiffusion, our diffusion Training INR involves a reference dataset, encompass-
model focuses on learning 3D reconstructions, with the INR ing a set of 2D projections from 3D sub-volumes, as de-
acting as a source of prior knowledge for global information scribed in Sec. 5.1, with 3D coordinates and corresponding
and actively contributing throughout the training process. intensities. The training objective is to minimize the recon-
struction error between the predicted intensities m_{inr} and
3. Problem Formulation the actual data sampled at each coordinate.

As depicted in Figure 1(b), a non-diffracting beam creates


a uniform point spread function along the axial direction INR for Volumetric Microscopy Reconstruction. As
with a limited width, offering an n -fold increase in imaging shown in Figure 2 (step 1), we sample 3D coordinates uni-
speed when the optical axial width of the non-diffracting formly from the 3D volume M, followed by positional en-
beam is n times that of a conventional point-like axial pro- coding and the use of an MLP to map these encodings to
file width beam. However, this advantage in volumetric mi- voxel density values. For each reference projection Xi , we
croscopy comes at the cost of depth information, as these compute the coordinates for n neighboring slices (defined
beams optically project 3D volumetric information along as 2D projections of neighboring 3D sub-volumes), and
the axial direction, leading to a lack of depth information concurrently synthesize these slices {mi1 , . . . , min } using
f_{\text {inr}} . Reconstruction loss is measured as the mean squared
in resulting 2D images.
This study aims to develop a model f capable of recon- error (MSE) between the mean of the synthesized slices and
structing a depth-resolved 3D volume from 2D projections the reference projection:
{Xi }, obtained using non-diffracting beams. The objec-
tive is to achieve a reconstructed 3D volume M with image L_{\text {mse}} = \sum _{i=1}^{N} \text {MSE}\left (\frac {1}{n} \sum _{k=1 }^{n} m^i_k, \mathbf {X}_i\right ), (2)
quality comparable to traditional point-scanning methods.
The 3D stacks that can be acquired from point-scanning
methods within the same ith sub-volume are represented where m_k denotes the k-th synthesized slice, and \protect \mathbf {X}_i repre-
as Mi = {mi1 , mi2 , . . . , min }. In non-diffracting volumet- sents the i-th reference projection.

3
Step 2: INR-guided Diffusion Step 1: INR Pretraining Implicit Neural
Representations
Acquired 2D Projections
𝐿𝐿𝑚𝑚𝑚𝑚𝑚𝑚

Reconstructed
3D Coordinates 3D volume
𝑋𝑋𝑇𝑇 𝑋𝑋𝑇𝑇−1 𝑋𝑋�
Denoising
U-Net …

INR
+ +
MicroDiffusion pipeline
Figure 2. Pipeline of MicroDiffusion. Step 1, we pre-train an INR which provides rough reconstructed images. Step 2, the 2D projections
and 3D coordinates are used as the classifier-free guidance of the MicroDiffusion, and the INR output is integrated into the noisy image as
guidance during the diffusion process. Detailed information is available at Sec. 4.
INR Neighbouring-based Inference. During training, During training, the neural network ϵθ (Xt , t) is trained to
we simulate the downsampling process triggered by the op- reconstruct the original data X0 from the noised data Xt .
tical axial projection with a non-diffracting beam, and opti- This is achieved by minimizing ℓ2 loss between the pre-
mizate the target loss by averaging the output over n coordi- dicted noise and the actual noise introduced in the data:
nates. However, this may introduce a distribution shift when
(5)
predicting the density solely from its own coordinate. To
L(\theta ) = \mathbb {E}_{(X,t)} \left [ \| \epsilon - \epsilon _{\theta }(X_t, t) \|^2 \right ] \label {eq:loss1}
mitigate this issue, we incorporate information from neigh- where t is time step in the forward diffusion process. Dur-
boring slices during inference. When considering a particu- ing the generation process, the neural network ϵθ (Xt , t) it-
lar 3D coordinate z, the inference result is obtained through eratively denoises Xt to achieve high-quality output, θ is
a weighted average of n neighboring slices. The weight for the trainable parameters of the model.
the k-th neighboring slice follows a Gaussian distribution: Classifier-free Guidance [7, 15] is a method for steer-
ing the output generation in Diffusion Models. Diverging
g_k = \frac {1}{\sqrt {2\pi }} e^{-\frac { (\frac {k}{n}-0.5)^2}{2}} (3) from the standard approach of diffusion models, this tech-
nique involves training a neural network ϵθ (Xt , t, c) with
While INR reconstruction offers a comprehensive global an additional conditioning c. The goal is to reconstruct X0
view, the reconstructed 3D slices suffer from blurriness, ar- while incorporating a probability puncond that c ← ∅. The
tifacts, and lack of fine details (as demonstrated in Figure 3). loss function L(θ) can be written as:
These issues compromise the spatial resolution and overall
reliability of optical microscopy. To address these limita- L(\theta ) = \mathbb {E}_{(X,t)} \left [ \| \epsilon - \epsilon _{\theta }(X_t, t, c) \|^2 \right ], \label {eq:loss2} (6)
tions and ultimately improve reconstruction quality, we in- where ∅ is the null class. At the generation time, the model
troduce a novel approach where we leverage INR as a global uses a guidance scale ω to balance the influence of the con-
prior to guide a diffusion model, enhancing the details and ditioning information. This is done by interpolating be-
reducing noise in each local 2D slice. tween the model’s predictions with and without the condi-
tioning:
4.2. Implicit Representation-Guided Diffusion
\tilde \epsilon _{t} = \epsilon _\theta (X_t, t, c) + \omega \cdot (\epsilon _\theta (X_t, t, c) - \epsilon _\theta (X_t, t)), (7)
Diffusion Models with Classifier-free Guidance We
employ Diffusion Models [17, 39] to reconstruct 3D vol- where ϵ̃t is the noise distribution that model predicts.
umes. As a likelihood model, Diffusion Model can grad-
ually recover the data from Gaussian noise. The forward Projection and Coordinate Guidance In MicroDiffu-
diffusion process transforms an input X0 to Gaussian noise sion, we use 2D projections and 3D coordinates as condi-
XT ∼ N (0, 1) by T iterations, defined as: tioning information c in Eq. 6. Projections provide content
information of the 3D volume while 3D coordinates provide
q(X_t | X_0) = \mathcal {N}(X_t | \sqrt {\bar {\alpha }_t} X_0, (1 - \bar {\alpha }_t) I), \label {eq:diff1} (4)
3D spatial information.
where XtQrepresents the data with added noise at time step As depicted in Figure 2, we introduce two distinct en-
t
t. ᾱt = s=0 (1 − βs ) and βs represents the noise vari- coders: an image encoder, denoted as Eimg , and a posi-
ance schedule, and N represents the Gaussian distribution. tional encoder, denoted as Epos . Eimg encodes the current

4
projection Xz , while Epos encodes the 3D coordinate z of Algorithm 2 Sampling function of MicroDiffusion
the current projection. The ultimate condition information Require: w: guidance strength; z: 3D Coordinate; γ: in-
c is formulated as follows: terpolation rate; T : Max time step; X: 2D projections.
c = E_{img}(\mathbf {X}_z) \oplus E_{pos}(p(z)), (8) XT ∼ N (0, 1)
where Xz is the reference projection, ⊕ is the concatenate c = Eimg (X) ⊕ Epos (p(z))
function and p(·) is the coordinate embedding function. minr = finr (p(z)) : INR Inference in Sec. 4.1
for t = T to 1 do
Xt′ = γminr + (1 − γ)Xt
INR Prior Integration To leverage global information ϵ̃t = (1 + w)ϵθ (Xt′ , t, c) − wϵθ (Xt′ , t)
and coherent 3D structures in INR to guide the Diffusion ϵt = sample from ϵ̃t
Models, we integrate INR outputs as prior knowledge for X̃t−1 = Xt − ϵt
the diffusion process. Specifically, the INR output minr end for
is linearly interpolated with the noisy image Xt , which is return X0
later to be denoised by Diffusion Models. In MicroDiffu-
sion’s training and testing process, for each noisy image Xt
at time step t, we perform linear interpolation with the INR
Generation process is outlined in Algorithm 2. Here w
output minr pixel by pixel as
is the condition weight controlling whether the model bias
X_t' = \gamma m_{\text {inr}} + (1 - \gamma ) X_t (9) more towards conditional or unconditional generation. Sim-
where Xt′ is the INR-enhanced image, Xt is the noised im- ilar to the training process, we first prepare all model inputs,
age that needs to be denoised by the diffusion model, minr and then have the model predict the noise distribution ϵ̃t at
is the reconstructed output from INR, and γ is the interpo- the current time-step t. We sample a noise ϵt from the noise
lation rate. This approach empowers the diffusion model distribution, subtract it from Xt , and repeat this for T times.
to directly leverage structural information learned by INR, We repeat this process for all the coordinates until the algo-
addressing the learning challenge with a limited number of rithm converges.
input 2D projections and enhancing its capacity to generate
images with correct 3D structures. 5. Experiments
5.1. Datasets
Training and Generation Process MicroDiffusion
We collected experimental data using a conventional mul-
adopts a conditional U-net [32] similar to that in stable-
tiphoton laser scanning microscope, which has been a gold
diffusion [31]. However, in our Denoising U-Net, we
standard imaging tool for modern biomedical study. This
remove the cross-attention mechanisms, and add both
approach is known for creating a three-dimensional, point-
time condition and conditional feature c at each output
like point spread function. By 3D scanning the tightly fo-
of the ResNet block. MicroDiffusion training algorithm
cused Gaussian beam, we generated high-quality 3D vol-
Algorithm 1 Training function of MicroDiffusion ume stacks. These stacks serve as ground truth datasets
for our research problem. Our setup captures 3D volume
Require: X: 2D projections; z: 3D Coordinate; t: Time stacks of various biological structures—such as dendrites,
step; puncond : Probability of being unconditional. neurons, and vasculature—within the shallow layers of the
minr = finr (p(z)) : INR Inference in Sec. 4.1 mouse cortex in a living animal (Fig. 1). These datasets al-
c = Eimg (X) ⊕ Epos (p(z)) low us to test our model with varied 2D projected images
c ← ∅ with probability puncond from diverse 3D biological features of varying densities.
Xt ←sample from q (Xt | X0 ) Subsequently, we generated three synthetic datasets.
Xt′ = γminr +(1 − γ)Xt These datasets simulate the case of fast data acquisition us-
L(θ) = E(X,t) ∥X0 − ϵθ (Xt′ , t, c)∥2

ing a non-diffracting beam, whose point spread function has
Take gradient step on L(θ) a quasi-uniform distribution axially and a predefined width
(Fig. 1b). In later experiments, as we will demonstrate,
is formulated in Algorithm 1, which continues running we varied this width to be different multiples (denoted as
until convergence. ⊕ is the concatenate operation. During step length n) of the Gaussian point spread function’s axial
training, we first encode 3D coordinates and 2D projections width used to scan and generate the ground truth datasets.
into conditional features c. We then generate the INR prior Consequently, the acquired 2D image sequences effectively
and the noised data X_t , and linearly interpolate them with averaged every n frames along the axial direction, with no
an interpolation rate \gamma . After preparing all model inputs, spatial overlapping in between. This approach reduced the
we follow the equation (6) to update the model parameters. volume data acquisition time by a factor of n. We then used

5
both the ground truth and the generated synthetic datasets sure (SSIM), and the Dice coefficient. PSNR and SSIM
to evaluate our model at different step length n values, fo- are calculated slice-wise along the axial direction, with the
cusing on various datasets from the brain. The design of mean value across all slices being reported. For the DICE
these datasets allows us to directly determine the optimal coefficient, we use the OTSU [26] algorithm to determine
step length n value, which will inform both future optimal the threshold for each image to assess the volumetric sim-
experimental data acquisition and hardware optical design. ilarity between the generated 3D structure and the ground
truth. As presented in Table 1, our method demonstrates
5.2. Implementation Details
strong performance, successfully capturing the principal
Depending on the specific imaging modality, the practical structure of the original high-resolution model. Addition-
axial width of a non-diffracting beam, such as a Bessel ally, we observe that the diffusion model’s decoder signifi-
beam, can vary from a few times to tens or hundreds of cantly enhances the performance of the pure INR model.
times that of the point-like Gaussian beam used in con-
ventional 3D laser scanning microscopes [3, 13]. We ini- Dataset Method SSIM ↑ PSNR ↑ DICE ↑

tiate our experiments with a step length n of approximately Interpolation 0.5799 28.78 0.6482
Interpolation - cubic 0.6511 28.85 0.3973
6, which corresponds to roughly an order of magnitude INR 0.5837 25.81 0.4589
Dendrite
in speed-up — an important initial milestone for volumet- Naive Diffusion 0.0297 19.95 0.2869
Interpolation Diffusion 0.6366 27.02 0.5729
ric imaging. The performances of different reconstruction Interpolation - cubic Diffusion 0.4765 21.31 0.3786
models were compared at this setting, and subsequently, an MicroDiffusion 0.6742 29.74 0.7557

ablation study was conducted over the step length value. Interpolation 0.3774 20.42 0.5936
Interpolation - cubic 0.5204 20.52 0.4448
This study aims to further understand the impact of the step INR 0.5032 21.69 0.7136
Vasculature
length of n on reconstruction quality, with the goal of identi- Naive Diffusion 0.0207 14.81 0.3234
Interpolation Diffusion 0.4039 19.09 0.4860
fying the optimal trade-off region between n times speed-up Interpolation - cubic Diffusion 0.2395 16.41 0.2672
and image reconstruction quality. MicroDiffusion 0.5787 22.35 0.7158
For computational efficiency, we downsample all the Interpolation 0.1208 24.12 0.3553
Interpolation - cubic 0.3265 26.50 0.1116
samples to a resolution of 128 \times 128 pixels in the lateral INR 0.4759 26.43 0.6403
Neuron
plane. For the pure Implicit Neural Representation (INR) Naive Diffusion 0.0210 24.08 0.1468
Interpolation Diffusion 0.4426 25.35 0.2425
model and the INR encoder, we map the 3D coordinates to a Interpolation - cubic Diffusion 0.3478 23.79 0.1318
512-dimensional space using a Gaussian-based embedding MicroDiffusion 0.4845 26.66 0.6708
technique. The INR model is optimized using the Adam
optimizer with a learning rate of 10^{-3} over 5000 epochs, Table 1. Main results of the image reconstruction quality across
a process that takes approximately 8 hours on an A-100 different datasets with different biological features: vasculature,
GPU. Additionally, we employ the AdamW optimizer with neurons, and dendrites. For all metrics, higher values indicating
better performance as indicated by the arrows.
a learning rate of 2^{-4} and a weight decay of 10^{-4} . As for
the diffusion model, it is trained over 2000 epochs, taking 5.4.2 Qualitative Results
around 4 hours on a single NVIDIA A-100 GPU. Here, we present the reconstruction results of three meth-
ods: pure INR and MicroDiffusion. Part of the slices from
5.3. Baselines the reconstructed 3D stacks are illustrated in Fig. 3, where
Given the novelty of this task and the absence of exist- we randomly selected three slices from the 3D reconstruc-
ing reference works, we established baseline methods. The tions generated by naive diffusion, the pure INR reconstruc-
initial approach is a straightforward Interpolation method, tion, our MicroDiffusion and compare with ground truth.
in which the generated structure is created through a uni- From these results, it is evident that the reconstructions ob-
formly weighted average of the two adjacent projections, tained via the MicroDiffusion method more closely resem-
weighted according to their distance. And Interpolation - ble the ground truth as the density of the biological features
cubic, which estimates values by using cubic polynomials increases. This result indicates an encouraging possibility
between points. This means that each interpolated curve that volumetric imaging with a non-diffracting beam allows
segment is based on the position and slope (derivative) at its not only well-known volumetric imaging of sparse features
endpoints. The last baseline employs a pure INR method, such as neurons in the cortex [3] but also denser features
which functions as our prior to the diffusion model. such as vasculature and even dense dendrites.
5.4. Reconstruction Results 5.5. Ablation study
5.4.1 Quantitative Results 5.5.1 Ablation on conditional feature
We evaluate our methods using three metrics: Peak Signal- How to encode 3D positional information? We initially
to-Noise Ratio (PSNR), Structural Similarity Index Mea- evaluate two positional encoding methods for MicroDiffu-

6
Naïve Ground (b) 3D Vasculature
(a) INR MicroDiffusion Truth (Ground Truth) method SSIM↑ PSNR↑ DICE↑
Diffusion
w/o feature guidance 0.5122 21.27 0.6512
Dendrite

cross-attention 0.5371 21.25 0.6784


addition (ours) 0.5787 22.35 0.7158
Vasculature

Table 3. Ablation results fusion ablation on vasculature


Reconstructed
(MicroDiffusion)

prior. uniformly-mean means that we use the uniformly av-


Neuron

eraged output of the INR corresponding to the six frames


centered on the current 3D coordinate z. The experimen-
tal results demonstrate that our approach performs the best
when introducing Neighbouring-based Inference, allowing
Figure 3. Qualitative results: (a) Comparative visualization of
the model to obtain a more comprehensive 3D INR prior.
slices from 3D reconstructions with different methods. Observ-
able differences between the INR reconstruction, MicroDiffusion method SSIM↑ PSNR↑ DICE↑
reconstruction, and ground truth are indicated with white arrows. no-neighbouring 0.4315 15.26 0.2025
(b) 3D vasculature. Scale bar: 30 µm. uniformly-mean 0.4996 17.36 0.4086
Neighbouring-based (ours) 0.5787 22.35 0.7158
sion: (1) the sine-cosine based encoding as described by
NeRF [25], and (2) the Gaussian-based encoding used in Table 4. Ablation of neighbouring based inference on vasculature.
NeRP [35]. As shown in Table 2, our experiment results
demonstrate that the Gaussian-based encoding yields supe-
rior results, particularly in rendering clearer textures. We at- How to add INR prior? We investigate the necessity of
tribute this improvement to the intrinsic properties of Gaus- the INR prior in this experiment. We trained a naive diffu-
sian embeddings, which affords a more flexible mapping of sion model that incorporates the INR prior as the projection
positions to a higher-dimensional space. This flexibility en- introduced in 4.2. We use the image encoder Eimg to en-
hances the subsequent learning process within MicroDiffu- code the output of INR minr and concatenate the feature
sion, leading to more detailed and accurate representations. with the other conditions. This allows the model to gener-
ate images that resemble true biological features. However,
Method SSIM↑ PSNR↑ DICE↑ this method performs poorly in acquiring global informa-
Sin-cos 0.5243 21.03 0.6667
tion, as evidenced by the very low DICE result in Table 5.
gassian-based(ours) 0.5787 22.35 0.7158
Method SSIM↑ PSNR↑ DICE↑
Table 2. Ablation of the positional encoding type on vasculature. Naive Diffusion 0.4178 14.97 0.4540
MicroDiffusion 0.5787 22.35 0.7158
How to add conditional feature? We ablate on the way
we incorporate the conditional feature c into the model. w/o Table 5. Ablation of diffusion model INR prior on the vasculature.
feature guidance involves setting all conditions to ∅. cross-
attention involves adding c using cross-attention to replace
5.5.3 Ablation on training method
the self-attention in Denoising U-net, where the image fea-
ture is the query, and the conditional feature c serves as In our pipeline, we adopt a two-stage training process where
the key and value. Experimental results demonstrate that the INR is trained initially and then frozen during the Mi-
adding the conditional feature c is effective and leads to the croDiffusion training. We conducted ablation experiments
best performance. This is likely because we have only one to explore two alternatives: (1) joint-training, where we
conditional feature, and in such a case, the cross-attention jointly train INR and the Denoising U-Net from random
mechanism may not be effective. Therefore, it is better to initialization, and (2) trainable, where we unfreeze the INR
directly add the conditional feature to the output of each during the MicroDiffusion training.
ResNet block in the Denoising U-net. We used the same number of epochs for all methods. For
joint-training, we added the INR loss and applied a decay-
5.5.2 Ablation on INR prior
ing weight to balance the training dynamics. As shown in
How to generate INR prior? Here, we test three INR Table 6, our method achieved the best performance. How-
prior generation methods as shown in Table 4. no- ever, joint-training proved to be too challenging and ad-
neighbouring means that we only utilize the INR output versely affected the MicroDiffusion training process, while
corresponding to the current 3D coordinate z as the INR trainable impaired the ability of INR to provide priors.

7
Ground Truth Step length: 8 Step length: 16 Step length: 32 Step length: 64
Therefore, we chose to train INR first and then train Mi- 0

croDiffusion with it frozen.

method SSIM↑ PSNR↑ DICE↑


260 µm

joint-training 0.5051 21.37 0.6603


trainable 0.5347 21.27 0.6875
freeze (ours) 0.5787 22.35 0.7158 Figure 5. Reconstruction of the depth-resolved sparsely dis-
tributed neuron images and depth-resolved volumetric projections
Table 6. Ablation results of training methold on vasculature with different step lengths.

5.5.4 Ablation on Different Step length


ages. This suggests a potential for significant improvements
We investigate the impact of the step length on MicroDiffu-
in imaging efficiency without substantial loss in image fi-
sion performance. As outlined in our methodology, a larger
delity for sparser featue of interest.
step size results in faster volumetric imaging but makes re-
construction more challenging. Conversely, smaller step
5.6. Discussion and Future Work
size leads to slower imaging but improved reconstruction.
We keep the number of iterations the same and train our In our experiments, we find that when ground truth data
model from scratch. Results are presented in Figure 6. contains Gaussian noise, MicroDiffusion outperforms other
We observed that as the step length increased, all model methods in noise removal. This demonstrates the poten-
metrics gradually decreased. To strike a reasonable bal- tial of MicroDiffusion for denoising 3D volumes acquired
ance between sampling speed and reconstruction quality, by volumetric optical microscopy. In cases where Gaus-
we chose a step length of 6. sian noise is intentionally part of the Ground Truth data and
SSIM, DICE, and PSNR Across Different Step Lengths SSIM should not be removed, it is necessary to investigate how to
DICE
0.8 PSNR
24 use other types of noise for MicroDiffusion training.
0.7
22
0.6 6. Conclusion
SSIM / DICE

0.5 20
PSNR

0.4
In this paper, we introduce MicroDiffusion, an innova-
18 tive 3D reconstruction framework that adeptly addresses
0.3
the challenges of rapid volumetric imaging and the need
0.2 16
for depth-rich visualizations in biomedical research. By
0.1
10 20 30 40 50 60 ingeniously integrating INR with DDPM, MicroDiffusion
Step length capitalizes on limited 2D projections to reconstruct high-
(a) SSIM, PSNR and DICE resolution 3D images, significantly enhancing the capabil-
ities of optical microscopy. Our approach not only accel-
Figure 4. Performance metrics across different step lengths. erates the image acquisition process but also maintains 3D
5.5.5 Reconstruction of sparse neuron dataset at vari- spatial information, allowing for the detailed observation of
ous step lengths complex biological structures with minimal data acquisition
at high speed. The successful application of MicroDiffusion
A natural question that may arise is whether we can further across various datasets, from densely distributed dendrites
increase the step length if our features are sparse in space. to sparsely distributed neurons, underscores its potential as
To address this, Here, we conduct one further experiment a transformative tool in medical diagnostics and fundamen-
aim to assess whether MicroDiffusion can further enhance tal biomedical research. This work paves the way for de-
the speed of volumetric imaging, particularly for samples signing next-generation volumetric optical microscopy, set-
with sparse spatial distribution. In this context, we have ting a new benchmark for the integration of machine learn-
conducted a comparative analysis using the neuron dataset, ing in 3D microscopy volume reconstruction, and opening
which is the most sparse case among all the three datasets. avenues towards high-speed, high-resolution 3D optical mi-
The results of this comparison are illustrated in Figure 5. croscopy.
We evaluated the performance of our reconstruction models
across a range of step lengths, which correspond to vary- Acknowledgement
ing degrees of data acquisition speed. Notably, our findings
indicate that, in the case of sparse neuron dataset, the step This work is partially supported by TPU Research Cloud
length can be extended to approximately 16 without signif- (TRC) program, and Google Cloud Research Credits pro-
icantly compromising the quality of the depth-resolved im- gram.

8
References Deán-Ben, Shy Shoham, and Daniel Razansky. Rapid vol-
umetric optoacoustic imaging of neural dynamics across the
[1] Tim Brooks, Aleksander Holynski, and Alexei A Efros. In- mouse brain. Nature biomedical engineering, 3(5):392–401,
structpix2pix: Learning to follow image editing instructions. 2019. 1, 6
In Proceedings of the IEEE/CVF Conference on Computer [14] Fritjof Helmchen and Winfried Denk. Deep tissue two-
Vision and Pattern Recognition, pages 18392–18402, 2023. photon microscopy. Nature methods, 2(12):932–940, 2005.
3 1, 2
[2] Rui Cao, Jingjing Zhao, Lei Li, Lin Du, Yide Zhang, Yilin [15] Jonathan Ho and Tim Salimans. Classifier-free diffusion
Luo, Laiming Jiang, Samuel Davis, Qifa Zhou, Adam de la guidance, 2022. 3, 4
Zerda, et al. Optical-resolution photoacoustic microscopy [16] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif-
with a needle-shaped beam. Nature Photonics, 17(1):89–95, fusion probabilistic models. In Advances in Neural Infor-
2023. 2 mation Processing Systems, pages 6840–6851. Curran Asso-
[3] Bingying Chen, Xiaoshuai Huang, Dongzhou Gou, Jianzhi ciates, Inc., 2020. 3
Zeng, Guoqing Chen, Meijun Pang, Yanhui Hu, Zhe Zhao, [17] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif-
Yunfeng Zhang, Zhuan Zhou, et al. Rapid volumetric imag- fusion probabilistic models. Advances in neural information
ing with bessel-beam three-photon microscopy. Biomedical processing systems, 33:6840–6851, 2020. 2, 4
optics express, 9(4):1992–2000, 2018. 2, 6 [18] Jan Huisken, Jim Swoger, Filippo Del Bene, Joachim Wit-
[4] Yinbo Chen, Sifei Liu, and Xiaolong Wang. Learning con- tbrodt, and Ernst HK Stelzer. Optical sectioning deep inside
tinuous image representation with local implicit image func- live embryos by selective plane illumination microscopy.
tion, 2021. 2 Science, 305(5686):1007–1009, 2004. 1
[5] Yen-Chi Cheng, Hsin-Ying Lee, Sergey Tulyakov, Alexan- [19] Animesh Karnewar, Andrea Vedaldi, David Novotny, and
der G Schwing, and Liang-Yan Gui. Sdfusion: Multimodal Niloy Mitra. Holodiffusion: Training a 3d diffusion model
3d shape completion, reconstruction, and generation. In Pro- using 2d images, 2023. 3
ceedings of the IEEE/CVF Conference on Computer Vision [20] Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen
and Pattern Recognition, pages 4456–4465, 2023. 3 Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic:
[6] Hye Jin Cho, Hoon Jai Chun, Eun Sun Kim, and Bong Rae Text-based real image editing with diffusion models. In Pro-
Cho. Multiphoton microscopy: an introduction to gastroen- ceedings of the IEEE/CVF Conference on Computer Vision
terologists. World Journal of Gastroenterology: WJG, 17 and Pattern Recognition, pages 6007–6017, 2023. 3
(40):4456, 2011. 1 [21] Kye-Sung Lee and Jannick P Rolland. Bessel beam spectral-
[7] Prafulla Dhariwal and Alex Nichol. Diffusion models beat domain high-resolution optical coherence tomography with
gans on image synthesis, 2021. 3, 4 micro-optic axicon providing extended focusing range. Op-
[8] JJJA Durnin. Exact solutions for nondiffracting beams. i. the tics letters, 33(15):1696–1698, 2008. 2
scalar theory. JOSA A, 4(4):651–654, 1987. 2 [22] Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa,
[9] S. M. Ali Eslami, Danilo Jimenez Rezende, Frederic Besse, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler,
Fabio Viola, Ari S. Morcos, Marta Garnelo, Avraham Ruder- Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution
man, Andrei A. Rusu, Ivo Danihelka, Karol Gregor, David P. text-to-3d content creation. In Proceedings of the IEEE/CVF
Reichert, Lars Buesing, Theophane Weber, Oriol Vinyals, Conference on Computer Vision and Pattern Recognition,
Dan Rosenbaum, Neil Rabinowitz, Helen King, Chloe pages 300–309, 2023. 3
Hillier, Matt Botvinick, Daan Wierstra, Koray Kavukcuoglu, [23] Minghua Liu, Chao Xu, Haian Jin, Linghao Chen,
and Demis Hassabis. Neural scene representation and ren- Mukund Varma T, Zexiang Xu, and Hao Su. One-2-3-45:
dering. Science, 360(6394):1204–1210, 2018. 2 Any single image to 3d mesh in 45 seconds without per-
[10] Sicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu, Yan- shape optimization, 2023. 3
jing Li, Xiaoyan Luo, Jianzhuang Liu, Xiantong Zhen, and [24] Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok-
Baochang Zhang. Implicit diffusion models for continuous makov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3:
super-resolution, 2023. 2 Zero-shot one image to 3d object, 2023. 3
[11] Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat [25] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik,
Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf:
Misra. Imagebind: One embedding space to bind them all, Representing scenes as neural radiance fields for view syn-
2023. 3 thesis. In ECCV, 2020. 2, 7
[12] Jan Goedeke, Peter Schreiber, Larissa Seidmann, Geling [26] Nobuyuki Otsu. A threshold selection method from gray-
Li, Jérôme Birkenstock, Frank Simon, Jochem König, and level histograms. IEEE transactions on systems, man, and
Oliver J Muensterer. Multiphoton microscopy in the diag- cybernetics, 9(1):62–66, 1979. 6
nostic assessment of pediatric solid tissue in comparison to [27] Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Milden-
conventional histopathology: results of the first international hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv,
online interobserver trial. Cancer management and research, 2022. 3
pages 3655–3667, 2019. 1 [28] Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren,
[13] Sven Gottschalk, Oleksiy Degtyaruk, Benedict Mc Larney, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Sko-
Johannes Rebling, Magdalena Anastasia Hutter, Xosé Luı́s rokhodov, Peter Wonka, Sergey Tulyakov, and Bernard

9
Ghanem. Magic123: One image to high-quality 3d object [43] Andres Flores Valle and Johannes D Seelig. Two-photon
generation using both 2d and 3d diffusion priors, 2023. 3 bessel beam tomography for fast volume imaging. Optics
[29] Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, express, 27(9):12147–12162, 2019. 2
Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. [44] Qing Wu, Yuwei Li, Yawen Sun, Yan Zhou, Hongjiang Wei,
Zero-shot text-to-image generation, 2021. 3 Jingyi Yu, and Yuyao Zhang. An arbitrary scale super-
[30] Cristina Rodrı́guez, Yajie Liang, Rongwen Lu, and Na Ji. resolution approach for 3d MR images via implicit neural
Three-photon fluorescence microscopy with an axially elon- representation. IEEE Journal of Biomedical and Health In-
gated bessel focus. Optics letters, 43(8):1914–1917, 2018. formatics, 27(2):1004–1015, 2023. 2
2 [45] Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying
[31] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Shan, Xiaohu Qie, and Shenghua Gao. Dream3d: Zero-shot
Patrick Esser, and Björn Ommer. High-resolution image text-to-3d synthesis using 3d shape prior and text-to-image
synthesis with latent diffusion models. In Proceedings of diffusion models. In Proceedings of the IEEE/CVF Con-
the IEEE/CVF conference on computer vision and pattern ference on Computer Vision and Pattern Recognition, pages
recognition, pages 10684–10695, 2022. 3, 5 20908–20918, 2023. 3
[46] Seok H Yun, Guillermo J Tearney, Benjamin J Vakoc, Milen
[32] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-
Shishkov, Wang Y Oh, Adrien E Desjardins, Melissa J Suter,
net: Convolutional networks for biomedical image segmen-
Raymond C Chan, John A Evans, Ik-Kyung Jang, et al. Com-
tation. In Medical Image Computing and Computer-Assisted
prehensive volumetric optical microscopy in vivo. Nature
Intervention–MICCAI 2015: 18th International Conference,
medicine, 12(12):1429–1433, 2006. 1
Munich, Germany, October 5-9, 2015, Proceedings, Part III
18, pages 234–241. Springer, 2015. 2, 5 [47] Ellen D. Zhong, Tristan Bepler, Joseph H. Davis, and Bon-
nie Berger. Reconstructing continuous distributions of 3d
[33] Chitwan Saharia, Jonathan Ho, William Chan, Tim Sali-
protein structure from cryo-em images, 2020. 2
mans, David J. Fleet, and Mohammad Norouzi. Image super-
resolution via iterative refinement, 2021. 3 [48] Haowen Zhou, Brandon Y. Feng, Haiyun Guo, Siyu Lin,
Mingshu Liang, Christopher A. Metzler, and Changhuei
[34] Changwon Seo, Kyeong-Joong Jeong, Sungsu Lim, and
Yang. Fpm-inr: Fourier ptychographic microscopy image
Won-Yong Shin. Siren: Sign-aware recommendation using
stack reconstruction using implicit neural representations,
graph neural networks, 2022. 2
2023. 2
[35] Liyue Shen, John Pauly, and Lei Xing. Nerp: Implicit neu-
ral representation learning with prior embedding for sparsely
sampled image reconstruction, 2023. 2, 7
[36] J Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner,
Jiajun Wu, and Gordon Wetzstein. 3d neural field generation
using triplane diffusion. In Proceedings of the IEEE/CVF
Conference on Computer Vision and Pattern Recognition,
pages 20875–20886, 2023. 3
[37] Rajhans Singh, Ankita Shukla, and Pavan Turaga. Poly-
nomial implicit neural representations for large diverse
datasets, 2023. 2
[38] Vincent Sitzmann, Julien Martel, Alexander Bergman, David
Lindell, and Gordon Wetzstein. Implicit neural representa-
tions with periodic activation functions. Advances in neural
information processing systems, 33:7462–7473, 2020. 2
[39] Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan,
and Surya Ganguli. Deep unsupervised learning using
nonequilibrium thermodynamics. In International confer-
ence on machine learning, pages 2256–2265. PMLR, 2015.
4
[40] Yu Sun, Jiaming Liu, Mingyang Xie, Brendt Wohlberg, and
Ulugbek S. Kamilov. Coil: Coordinate-based internal learn-
ing for imaging inverse problems, 2021. 2
[41] Stanislaw Szymanowicz, Christian Rupprecht, and Andrea
Vedaldi. Viewset diffusion: (0-)image-conditioned 3d gen-
erative models from 2d data, 2023. 3
[42] Xiao-Jie Tan, Cihang Kong, Yu-Xuan Ren, Cora SW Lai,
Kevin K Tsia, and Kenneth KY Wong. Volumetric two-
photon microscopy with a non-diffracting airy beam. Optics
Letters, 44(2):391–394, 2019. 2

10
Supplementary Material images. This index is particularly adept at capturing per-
ceptual differences, making it a robust tool in image quality
In this supplementary material, we detail the metrics we assessment.
used in Sec. 7, including SSIM in Sec. 7.1, PSNR in
Sec. 7.2, DICE in Sec. 7.3. Furthermore, we conduct an 7.2. PSNR
ablation study to find the best interpolation rate in Sec. 8.
Finally, we visualize the MicroDiffusion reconstruction re- The Peak Signal-to-Noise Ratio (PSNR) is another crucial
sults at various step lengths on the vasculature dataset in metric, predominantly focusing on the ratio between the
Sec. 9. maximum possible power of a signal and the power of cor-
rupting noise. It is articulated as:
7. Details of the metrics
In this section, we delineate the metrics employed for as- \text {PSNR}(x, y) = 10 \log _{10} \left ( \frac {255^2}{\text {MSE}(x, y)} \right ), (14)
sessing the quality of image reconstruction. We consider a
reference image x and a test image y, both being grey-level where MSE(x, y), the Mean Squared Error between the two
(8 bits) images of dimensions M × N , drawn from respec- images, is computed as:
tive sets X and Y . Three evaluation measures are utilized:
the Structural Similarity Index Measure (SSIM), the Peak
Signal-to-Noise Ratio (PSNR), and the Sørensen–Dice co- \text {MSE}(x, y) = \frac {1}{MN} \sum _{i=1}^{M} \sum _{j=1}^{N} (x_{ij} - y_{ij})^2, (15)
efficient (DICE), each detailed below.
7.1. SSIM with xij and yij representing the pixel values at the ij th
The Structural Similarity Index Measure (SSIM) serves as a position. The PSNR values range from 0 to ∞, where a
pivotal metric in quantifying the resemblance between two higher value indicates superior image quality, reflective of
images. It evaluates three fundamental aspects: Luminance, lesser noise interference.
Contrast, and Structure.
7.3. DICE
Luminance is quantified through the mean gray scale
value of the pixels, encapsulated in the equation: The Sørensen–Dice coefficient (DICE), a statistical tool,
quantifies the similarity between two sets. It is particularly
l(x, y) = \frac {2\mu _x\mu _y + C_1}{\mu _x^2 + \mu _y^2 + C_1}, (10) effective in comparing the spatial arrangement of pixel val-
ues. The DICE is defined as:
where µx and µy represent the mean luminance of images
x and y, respectively. The constant C1 prevents a zero de- \text {DICE}(X, Y) = \frac {2 |X \cap Y| + C}{|X| + |Y| + C}, (16)
nominator.
Contrast is gauged using the gray scale standard devia-
where |X ∩ Y | denotes the intersection size of sets X and
tion, as:
Y, and |X| and |Y | are their respective sizes. The coeffi-
c(x, y) = \frac {2\sigma _x\sigma _y + C_2}{\sigma _x^2 + \sigma _y^2 + C_2}, (11) cient ranges from 0 to 1, with 1 indicating perfect agreement
(complete overlap) and 0 denoting no overlap at all. This
where σx and σy denote the standard deviations of the im- metric is particularly beneficial in scenarios where spatial
ages, and similarly, C2 prevents a zero denominator. correlation is a critical aspect of image similarity.
Structure is assessed through correlation coefficients,
formulated as: 8. What is the best linear interpolation rate?
s(x, y) = \frac {\sigma _{xy} + C_3}{\sigma _x\sigma _y + C_3}, (12) As demonstrated in the main paper, to incorporate global
information and coherent 3D structures into the diffusion
with σxy being the covariance between x and y. The con- model, we employ a linear interpolation strategy between
stant C3 ensures non-zero denominators. the Implicit Neural Representations (INR) output and the
The aggregate SSIM value, encapsulated within the noisy image at each time step t. This approach is applied
range [0, 1], is derived as: during both the training and testing phases of MicroDiffu-
sion. Such integration of INR as prior knowledge is pivotal
\text {SSIM}(x, y) = l(x, y) \cdot c(x, y) \cdot s(x, y), (13)
for guiding the diffusion process, particularly when dealing
offering a comprehensive measure of similarity. Notably, with limited 2D projection inputs.
a SSIM score of 0 implies an absence of correlation be- An ablation study focusing on the interpolation rate γ
tween the images, whereas a score of 1 indicates identical was conducted, with the results summarized in Figure 6.

11
Our findings indicate that the effectiveness of γ plateaus be-
yond a threshold of 0.1. Further increments in γ yield min-
imal improvements, as evidenced by a marginal decrease
in both Peak Signal-to-Noise Ratio (PSNR) and Structural
Similarity Index Measure (SSIM) metrics. This trend sug-
gests that, while the incorporation of INR prior is benefi-
cial, an excessive reliance on it, particularly in the absence
of Gaussian noise, can compromise the model’s generaliza-
tion capabilities.
SSIM, PSNR, and DICE Across Different
0.7 22
21
0.6
20
SSIM / DICE

0.5 19

PSNR
18
0.4
17
0.3 16
15SSIM
0.0 0.2 0.4 0.6 0.8 1.0 DICE
PSNR

(a) SSIM, PSNR and DICE

Figure 6. Performance metrics across different linear interpolation


rates.

9. Visualization of MicroDiffusion reconstruc-


tion of vasculature at various Step Lengths
We visualize the reconstructed images in Figure 7, which
clearly demonstrates that the difficulty of reconstruction es-
calates with increasing step length, leading to a noticeable
decline in model performance. This trend is quantitatively
supported by the rapid decrease in metrics such as SSIM,
PSNR, and DICE. Despite this challenge, it is noteworthy
that satisfactory reconstruction quality is still achievable at
step lengths of approximately 6 to 8. This finding is sig-
nificant as it implies the potential to increase the speed of
volumetric imaging by a factor of 6 to 8, enhancing imag-
ing efficiency substantially. Looking ahead, our research
aims to further improve model performance at even higher
step lengths, pushing the boundaries of efficient and high-
quality imaging in MicroDiffusion processes.

12
Depth: 100 µm 150 µm 200 µm 250 µm 300 µm 350 µm 400 µm 0 - 200 µm
0

Ground
truth
200 µm
0

Step
length: 4
200 µm
0

Step
length: 8
200 µm
0

Step
length: 16

200 µm
0

Step
length: 32
200 µm
0

Step
length: 65

200 µm
0

Step
length: 130

200 µm

Figure 7. Reconstruction of the depth-resolved vasculature images and depth-resolved volumetric projections with different step lengths.

13

You might also like