0% found this document useful (0 votes)
35 views9 pages

Generate 360° Panoramas from NFoV Images

A paper about panoramic image generation with AI

Uploaded by

edwardsj2468
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
35 views9 pages

Generate 360° Panoramas from NFoV Images

A paper about panoramic image generation with AI

Uploaded by

edwardsj2468
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

360-Degree Panorama Generation from Few Unregistered NFoV

Images
Jionghao Wang∗ Ziyu Chen∗ Jun Ling
Shanghai Jiao Tong University Shanghai Jiao Tong University Shanghai Jiao Tong University
shanemankiw@[Link] 1252060456@[Link] lingjun@[Link]

Rong Xie Li Song†


Shanghai Jiao Tong University Shanghai Jiao Tong University
xierong@[Link] song_li@[Link]
arXiv:2308.14686v1 [[Link]] 28 Aug 2023

Figure 1: Illustration of our PanoDiff, a novel approach that is capable of synthesizing fine-grained and diverse 360-degree
panorama from a(few) unregistered NFoV (Narrow Field of View) image(s) and text prompts.
ABSTRACT overcome the limitations. Firstly, a two-stage angle prediction mod-
360◦ panoramas are extensively utilized as environmental light ule to handle various numbers of NFoV inputs. Secondly, a novel
sources in computer graphics. However, capturing a 360◦ × 180◦ latent diffusion-based panorama generation model uses incomplete
panorama poses challenges due to the necessity of specialized and panorama and text prompts as control signals and utilizes several
costly equipment, and additional human resources. Prior studies geometric augmentation schemes to ensure geometric properties
develop various learning-based generative methods to synthesize in generated panoramas. Experiments show that PanoDiff achieves
panoramas from a single Narrow Field-of-View (NFoV) image, but state-of-the-art panoramic generation quality and high controlla-
they are limited in alterable input patterns, generation quality, and bility, making it suitable for applications such as content editing.
controllability. To address these issues, we propose a novel pipeline
called PanoDiff, which efficiently generates complete 360◦ panora- CCS CONCEPTS
mas using one or more unregistered NFoV images captured from • Computing methodologies → Computer vision; Computer
arbitrary angles. Our approach has two primary components to graphics.
∗ Indicatesequal contribution.
† Corresponding
KEYWORDS
author. Affiliated with both School of Electronic Information and
Electrical Engineering, Shanghai Jiao Tong University & MoE Key Lab of Artificial 360-degree panorama, generative models, multimodal models, la-
Intelligence, AI Institute, Shanghai Jiao Tong University. tent diffusion, image pose estimation
Jionghao Wang, Ziyu Chen, Jun Ling and Rong Xie are with School of Electronic
Information and Electrical Engineering, Shanghai Jiao Tong University. ACM Reference Format:
Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, and Li Song. 2023. 360-
Permission to make digital or hard copies of all or part of this work for personal or Degree Panorama Generation from Few Unregistered NFoV Images. In
classroom use is granted without fee provided that copies are not made or distributed
for profit or commercial advantage and that copies bear this notice and the full citation Proceedings of the 31st ACM International Conference on Multimedia (MM
on the first page. Copyrights for components of this work owned by others than the ’23), October 29-November 3, 2023, Ottawa, ON, Canada. ACM, New York,
author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or NY, USA, 9 pages. [Link]
republish, to post on servers or to redistribute to lists, requires prior specific permission
and/or a fee. Request permissions from permissions@[Link].
MM ’23, October 29-November 3, 2023, Ottawa, ON, Canada 1 INTRODUCTION
© 2023 Copyright held by the owner/author(s). Publication rights licensed to ACM.
ACM ISBN 979-8-4007-0108-5/23/10. . . $15.00 Panoramic images, which capture an extensive field of view en-
[Link] compassing a full 360◦ horizontal by 180◦ vertical FoV scene, have
MM ’23, October 29-November 3, 2023, Oawa, ON, Canada Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, & Li Song

become increasingly significant across various applications, such as we know, PanoDiff is the first latent-diffusion-based panorama
as environment lighting, VR/AR, and autonomous driving system. outpainting model that can not only handle incomplete panoramas
However, obtaining high-quality panoramic images can be both of various shapes as input but also supports text prompts. Finally,
time-consuming and costly, as they typically necessitate the use of PanoDiff outperforms existing methods in terms of quantitative
specialized panoramic cameras or stitching software to combine and qualitative results on different input types, as demonstrated
images from multiple perspectives. Our method addresses two main through abundant experiments.
limitations regarding previous generating methods, namely input
pattern and generation quality & controllability. 2 RELATED WORKS
For Input pattern. Most previous works [1, 5, 37] only support 2.1 Rotation Estimation
a single FoV region in the center of an incomplete panorama as
input. However, relying solely on a single input region restricts Camera pose estimation has been a long-standing task in computer
flexibility when controlling specific contents of the generated scene. vision [19]. Most prior works focus on identifying various visual
Such flexibility is particularly important for applications that re- cues from pairs of images. For instance, the classic SIFT [22] ex-
quire precise control over the visual elements and ensure accurate tracts scale and direction invariant features from input images and
representations of the intended scene. matches them to find correspondences. Some approaches utilize
In terms of generation quality & controllability. Generating a com- neural networks for more robust feature extraction or graph neu-
plete 360-degree panorama from NFoV images could be viewed as a ral networks for enhanced matching [6, 28, 32]. However, these
large-hole inpainting problem [35]. Previous methods [1, 11, 37] typ- methods do not account for extreme rotations where the input
ically all rely on GAN(Generative Adversarial Networks) [10] based images have little to no overlap, known as wide baseline scenarios.
methods. Besides, previous methods approach this problem as an Some earlier works focus on mining specially handcrafted cues [24]
image-conditioned generation task, leaving their pipeline with little to address this problem, but they are not generalizable for scenes
control and flexibility over the generation results. A more recent without a specific clue. The work that inspired us the most is [2],
work [5] proposes to use additional text guidance for a GAN-based which estimates pair-wise rotation angles in both overlap and wide
image inpainting method [41]. However, diffusion-based image baseline scenarios. However, the estimation network occasionally
generation methods [14] have shown impressive results on various misidentifies an overlapping pair of images as non-overlapping,
generative tasks, and models trained on large datasets [29, 33] have which severely impacts the generation quality later on.
shown better performance and robustness against GAN-based mod-
els [7, 15]. Moreover, GANs have limited mode coverage and are 2.2 Diffusion Models
difficult to scale for modeling complex multimodal distributions. Diffusion models have shown impressive ability in generative tasks [7,
Unlike GANs, likelihood-based models such as diffusion models 14]. Generally, diffusion models represent the image generative pro-
are capable of learning the complex distribution of natural images, cess as a denoising process, where a noise map sampled from white
resulting in the generation of high-quality images, as mentioned Gaussian noise is iteratively denoised using a learned prior distri-
in [7, 29]. bution  ( −1 | ). The objective function for training diffusion
To overcome these limitations, a pipeline capable of accepting models is simplified as
a flexible number of NFoV image(s) as input and generating high-  
simple = , 0 , | −  ( , )| 2 (1)
fidelity panoramas is crucial. Nevertheless, this endeavor presents
two main challenges: 1) estimating relative camera poses and accu- which has been shown to be approximately equivalent to optimizing
rately warping the input NFoV image(s) on the panorama, and 2) the prior probability distribution’s variational lower bound [14].
using a latent-diffusion-model-based method to generate the entire To further improve the efficiency and effectiveness of the dif-
panorama from input partial panorama of various shapes. fusion model, recent work has proposed to conduct the denois-
This paper introduces PanoDiff, a novel pipeline that efficiently ing process in latent space instead of pixel space [29]. In addition,
generates complete 360° panoramas using one or more unregis- they integrate multi-head cross-attention mechanism [3, 36] in de-
tered NFoV images captured from arbitrary angles. The proposed noising U-Net [25, 30] blocks. This approach enables features of
method overcomes the limitations of existing methods, it enables different modalities to guide the generation process[16], e.g. CLIP-
the generation of high-quality panoramas from incomplete 360- ViT [27]. Working on latent space along with some other sampling
degree panoramas that warped from any number of NFoV inputs. schedule [34] could also potentially make the sampling more ef-
This is achieved by addressing two key challenges. First, we pro- ficient. Controlling the pre-trained model parameters by training
pose a robust two-stage angle prediction pipeline that classifies hypernetworks[12, 40] can improve its performance on a more
image pairs based on their overlap before regressing specific angle specific task while preserving its original generative capability.
values. Second, we train a hypernetwork that controls a pre-trained
large latent-diffusion model, and utilizes geometric augmentation 2.3 Diffusion-based Inpainting
schemes during both training and inference sampling phases to The task of inpainting was predominantly been done by GAN-
ensure the geometric properties of the generated panoramas. based models [21, 35, 42]. Recently, diffusion-based methods have
We summarize our contributions as follows. Firstly, the first been proposed and have already achieved promising results [23,
flexible framework for generating panoramic images from single 38]. However, [23, 38] operate the diffusion and sampling process
or multiple NFoV inputs. A two-stage pose estimation module is solely on image space rather than latent space, which limits their
specially designed for relative pose estimation. Secondly, as far generation flexibility. Recently, [31] offered an image inpainting
360-Degree Panorama Generation from Few Unregistered NFoV Images MM ’23, October 29-November 3, 2023, Oawa, ON, Canada

Figure 2: An overview of generating panorama from a few NFoV images. We first calculate their relative rotations based on a
two-stage angle prediction network R (Sec. 3), then project them using backward equirectangular projection Fproj (Sec. 4) to
obtain partial panorama  and visibility mask . Finally, we feed [, ] along with text prompts to our control-based latent
diffusion model FLDM (Sec. 4) and sample iteratively in our rotating schedule (Sec. 4.4.2) to get the final generated panorama.

model finetuned from Stable Diffusion [29] and demonstrates strong image  0 as an anchor and estimate the relative poses [Δ , Δ,
text-controlled inpainting ability. Δ ] of the remaining images with respect to the anchor. Follow-
ing the assumptions made in [2], we consider that cameras are
2.4 Panorama Generation typically upright, allowing us to estimate the absolute pitch an-
With recent advancements in deep generative neural networks, gle (vertical/latitude) instead of the relative angle. We also assume
several works adopt different forms of GAN-based methods to that the camera roll angles are static and remain unchanged at 0,
generate full panoramas [1, 5, 20, 37], such as VQGAN [8] and i.e., ∀ ∈ 0, ...,  − 1,  = 0. Consequently, our network R can be
CoModGAN [41]. However, these methods are limited to single- expressed as:
view input, in contrast to our alterable input patterns. One work
R ( 0,  ) → [Δ→0,  0,  ] (3)
has utilized a few NFoV images as input [11] with accurate camera
positions already given, thus cannot deal with unregistered camera where Δ→0 is the relative longitude angle between the images,
inputs. and  0,  are the absolute lattitude angle for  0 and  , respectively.
Recently, some researchers have utilized diffusion models to
generate panoramas from text prompts alone [4]. However, this
work does not support partial FoV as input, and thus users have 3.2 Two-stage Angle Prediction
limited control over the outcome. In the context of FoV overlapping, we categorize image-pair re-
lationships into two distinct types: overlap and wide baseline. In
3 RELATIVE POSE ESTIMATION overlap scenarios (green pair in Fig. 3), two NFoV images share
a portion of their FoV, resulting in an overlap on the panorama.
3.1 Problem Formulation Conversely, in wide baseline situations (red pair in Fig. 3), the NFoV
Unlike previous panorama generation pipelines, our method takes image pair does not have overlapping FoV, leading to two separate
either single or multiple NFoV images as input. In the case of mul- regions on the panorama. The ultimate goal of our method is to
tiple NFoV (Narrow Field-of-View) images captured at the same create a plausible panorama for further applications. However, a
location but in different poses, our objective is to estimate the rela- single regression network [2] is not robust enough to differenti-
tive poses between them. In the context of omnidirectional images, ate between these two cases. Consequently, images without any
where no camera translation is involved, we formulate the relative overlap might sometimes be predicted to have an overlap, creating
camera poses as rotation angles. Specifically, we view the 360◦ severe artifacts in the input for our controlled LDM, which would
panorama as a spherical map and express the rotation angles as result in unsatisfactory results. Furthermore, the precision of a
the 1) lookat direction, which includes the longitude , latitude , single regression network is insufficient.
and 2) roll angle . For a single image, our pipeline estimates only To address these problems, we propose a two-stage angle pre-
 and places it in lateral center, i.e.,  =  = 0. Formally, for a set diction pipeline that initially performs classification to determine
of  NFoV images  =  0,  1, ...,   −1 , we aim to learn a model R whether the two images have an overlap, and subsequently re-
that predicts their relative angles: gresses the precise angles [Δ→0,  0,  ]. This approach signifi-
R (I) → [ , ,  ] (2) cantly reduces the issue of wide baseline images being mistakenly
predicted to overlap. The overview of our rotation prediction mod-
where  = { }=0,..., −1 ,  = { }=0,..., −1 and  = { }=0,..., −1 , ule can be seen in Fig. 3. We adopt the same backbone structure
respectively. We decompose this task into estimating pair-wise cam- as [2], which consists of a feature encoder and a 4D dense cor-
era relative rotation, which is scalable and can be applied to pair relation module. Within the two-stage pipeline, a total of three
input and sometimes more than two images. We design a pairwise backbones are utilized, each with its own prediction head. To esti-
angle prediction network that estimates the relative angles between mate the relative camera rotation between a pair of images, we first
images. For a set of  NFoV images ( 0,  1, ...,   −1 ), we select one pass the image pair to our classifier to determine if the images are
MM ’23, October 29-November 3, 2023, Oawa, ON, Canada Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, & Li Song

time step , given input latent  and condition , the denoiser


network is formulated as  ( , , ). The denoiser in SD employs a
U-Net bottleneck architecture.
[40] proposed to use combine trainable copies from pre-trained
SD model parameters, and zero-convolution blocks, whose parame-
ters are initialized as all zero. Concretely, suppose there is a neural
network block F (·; Θ) from the pre-trained SD denoiser U-Net,
which takes an input feature map x ∈ Rℎ×× and outputs a fea-
ture y, i.e. y = F (x; Θ). To add more conditions as control signals,
two zero-convolution layers Z1, Z2 and their parameters Θ1, Θ2 ,
Figure 3: Illustration of our two-stage angle prediction net- and a trainable copy of the original neural block and its parameters
work. E , E represent feature encoders, and DΔ , D0 , D as F and Θ are introduced. Now given a new condition signal c,
serve as prediction heads for overlap and wide-baseline angle the new controlled output feature y can be written as:
regression, respectively. The pipeline first estimates whether
y = F (x; Θ) + Z2 (F (x + Z1 (c; Θ1 ); Θ ); Θ2 ) (5)
the input image pair has an overlap, then directs the image
pair to the corresponding regression network. It is worth noting that the original operation F (·; Θ) is locked
and does not produce any gradient for backpropagation. In this
setting, the new set of trainable parameters are Θ , Θ1, Θ2 , where
overlapping or wide apart. Subsequently, we feed the image pair to the initial Θ is copied from Θ, and both Θ1 and Θ2 are initial-
its corresponding regressor based on the classification result. ized as zeros. This allows the controlled input at the first step to
In terms of training, our classifier is trained on our full set of be identical to the original output, i.e., y = y. This property is
training image pairs. A specialized penalty loss function is employed favorable as it maintains the generation ability of the pre-trained
for misclassified cases, which effectively mitigates failure instances. model and makes the updating process of Θ , Θ1, Θ2 more stable.
As for our regressors, we separate our training image pairs into An intuitive illustration of this scheme can be seen in Fig. 4.
two splits based on their overlapping status and train the overlap
branch and wide branch separately on their respective data splits.
That is, for the overlap regressor, we train it using primarily overlap
data pairs, and vice versa for the wide baseline regressor.

4 PANORAMA GENERATION
4.1 Formulation
With Sec. 3, we can formulate a set of image-angle pairs, i.e., {I0, 0,  0 }
and {Ii, Δ→0,  }=1,..., −1 . In this part, our goal is to take the im-
ages and their relative angles to produce a complete 360◦ × 180◦
panorama. We divide this problem into two parts: 1) project the
image set  = {Ii } onto the omnidirectional map based on their Figure 4: An intuitive explanation of integrating a new con-
estimated angles  = { }=0,..., −1 and  = { }=0,..., −1 ; and 2) trol signal into an existing model block.
generate the full panorama using the incomplete input as a control
signal. Specifically, we formulate this problem as:
Fproj ( ,  , ) → , , 4.3 Controlling LDM with Partial FoV
(4) Given a noisy latent space feature  at time step , we want to
FLDM (, , text) → Pano,
learn a denoiser that could predict the noise  as in:
where Fproj and FLDM represent the projection and diffusion gener-
ation operations, respectively. The projection operation Fproj takes  =  ( , , ctext, cP ) (6)
the NFoV image set and their angles as input and produces a partial Here, ctext and cP = [, ] represent the text conditions and input
panorama image  and a visibility mask , indicating which parts panorama conditions, respectively. We introduce two components
of the FoV are missing. Subsequently, the LDM (Latent Diffusion for our partial-FoV-controlled LDM, namely Pano-Denoiser  and
Model) operation FLDM accepts the incomplete , visibility mask , Pano-Encoder E .
and a text prompt as input, iteratively generating the full panorama Our Pano-Denoiser is designed as discussed in Sec. 4.2, where we
Pano. employ control units to introduce control signals to the pre-trained
Stable Diffusion model. In practice, we follow the approach of [40]
4.2 Recap: Controlling Stable Diffusion and add control units to the four encoder blocks and one middle
We propose a generative process based on training a hypernetwork block of the denoising U-Net model used in Stable Diffusion, while
over a pre-trained Stable Diffusion (SD) model[29], a large text-to- utilizing zero-convolutions for the other four decoder blocks.
image generative model utilizing latent diffusion. Latent diffusion In addition to the Pano-Denoiser, we incorporate a shallow Pano-
models iteratively perform denoise sampling in latent space. At Encoder E to construct our controlling condition. In order to feed
360-Degree Panorama Generation from Few Unregistered NFoV Images MM ’23, October 29-November 3, 2023, Oawa, ON, Canada

image-space signals, such as  and , into the latent space as con-


trol signals, we use E to encode them into latent features that
are compatible with the original SD U-Net and consequently our
control unit. The latent features are then passed to our learnable en-
coders and zero convolutions connected to the middle and decoder
parts of the original SD U-Net.

Figure 6: An illustration of our rolling schedule. The target


origin represents the position of the original pixel corre-
sponding to longitude and latitude coordinate ( = 0,  = 0)
on the panorama.

where A () represents an image space projecting operation, simi-


lar to Fproj . This equation implies that when the latent feature 
and all input image-space conditions undergo a transformation by
A (), the predicted noise should also transform accordingly. To
enforce this behavior, we introduce a rotation perturbation for la-
tent features during the training phase. Consequently, the training
Figure 5: An overview of our partial-FoV-controlled latent objective evolves from Eq. 1 into:
diffusion model. The upper part and lower part are the diffu-
L=
sion process and sampling process of the model, respectively.
Note that the text conditions and classifier-free guidance E 0 ,, text , P , [∥A (Δ) −  (A (Δ) , ,  text, A (Δ) P ) ∥ 22 ]
process are omitted from the figure for simplicity.
4.4.2 Rotating Schedule. During the inference process, we employ
a customized schedule as depicted in Fig. 6.
As depicted in Fig. 5, during the diffusion process (upper part As illustrated in Figure 6, the denoising steps are executed while
of the figure), the encoder E of the AutoEncoder transforms the the latent feature  undergoes a scheduled transformation A (Δ )
image into feature space, and then Gaussian noises are added itera- in a step-by-step manner. It is important to note that this schedule
tively to the latent feature, as described in [14]. For the denoising is consistent with the rotation constraint incorporated during the
process, our Pano-Denoiser utilizes the partial panorama  and training phase, as both involve the same type of transformation in
visibility mask  as control signals and denoises the input noisy the latent space. The rotating schedule enhances the robustness
latent  iteratively in our rotating schedule(Sec. 4.4.2) for  steps of our method and facilitates the generation of panoramas with
until  0 is generated. Finally,  0 is decoded by the decoder E of improved geometric integrity.
the AutoEncoder to produce the final panorama.
4.4.3 Circular Padding. As discussed in [11], generative methods
operating in the latent space may lead to geometric discontinuity
4.4 Denoising in 360-Degree
due to border effects of convolution operations. To address this
In this subsection, we focus on denoising 360-degree panoramas issue, during inference time, we implement circular padding to
while preserving their unique geometric characteristics. During mitigate edge effects. Specifically, the right portion of the latent
the training stage, we introduce a rotation equivariance loss to feature is concatenated to the left side, while the left part of the
enforce rotation consistency in the latent space. During the infer- original latent feature is concatenated to the right side. This process
ence stage, we employ a customized rotating schedule to enhance is illustrated in Fig. 7.
the robustness and maintain the geometric integrity of the gener-
ated panoramas. Additionally, we implement a circular padding
technique during inference to mitigate edge effects and prevent
geometric discontinuity.
4.4.1 Rotation Equivariance Loss. Panoramas are captured to depict
a completely spherical environment, which means they exhibit
rotation equivariance. This property implies that if we apply a
transformation based on an  (3) rotation (limited to 2 degrees
of freedom corresponding to longitude and latitude rotations) to a Figure 7: An illustration of our circular padding technique.
panorama image, it should still be able to represent the same scene. Thin slices of the feature, denoted as  and , from both
As our method functions within the latent space, the constraint the left and right ends of the latent feature, are copied and
on our denoiser can be expressed as: concatenated to the right and left sides of the latent feature
 (A () , , ctext, A ()cP ) = A () ( , , ctext, cP ), (7) as ′ ,  ′ , respectively.
MM ’23, October 29-November 3, 2023, Oawa, ON, Canada Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, & Li Song

Table 1: FID↓ results compared with other generation methods for quantitative evaluation.

SUN360 [39] Laval [9]


Methods Single Pair (GT rots) Pair (Pred rots) Single Pair (GT rots) Pair (Pred rots)
SIG-SS [11] 13.06 15.94 16.50 - - -
StyleLight [37] - - - 22.53 34.10 33.88
Omni-Comp [1] 14.69 12.38 12.63 18.91 14.62 14.73
Omni-Adjust [1] 11.27 9.93 10.63 9.60 10.47 10.92
ImmerseGAN [5] 9.26 10.46 10.57 12.69 24.23 24.51
PanoDiff (Ours) 8.61 6.94 7.04 7.72 6.96 7.18

After the sampling process, the decoder generates an image Table 2: Evaluation of relative pose estimation on
of shape ( + 2  ) × ℎ, where   is the extra width from the SUN360 [39] and Laval Indoor [9]. We present the Av-
padded feature. To produce a standard panorama, the extra width erage geodesic error Avg(◦ ↓) and the percentage of pairs with
is removed from the final image. errors under 10 degrees 10◦ (%↑).

5 EXPERIMENTS SUN360 [39] Laval [9]


Pair Type Method Avg(◦ ↓) 10◦ (%↑) Avg(◦ ↓) 10◦ (%↑)
5.1 Implementation Details 1-stage 7.19 90.27 4.21 96.74
Overlap
5.1.1 Data Preparation. We conducted experiments using real- 2-stage 3.58 95.61 3.66 96.88
world 360-degree panoramic image datasets SUN360 [39] and Laval 1-stage 41.52 40.64 32.31 62.34
Wide Baseline
2-stage 27.12 49.77 29.46 62.62
Indoor [9]. SUN360 comprises both indoor and outdoor scenes,
1-stage 24.29 65.54 18.00 79.86
while Laval Indoor comprises solely indoor scenes. For SUN360, All
2-stage 15.31 72.77 16.32 80.07
we randomly selected 2000/500 panoramas for training/testing,
respectively. For Laval Indoor, we followed the approach in [11]
and chose 289 images for testing. Notably, we do not train our is trained on the Laval dataset, the same training set as used in
model using the Laval Indoor dataset as there are already indoor Omni-Dreamer. We inferred these models with their officially re-
scenes within our selected training data from the SUN360 dataset. leased trained models on our test split. Since their training split
We used two input types in our experiments: a single input, is unknown to us, images in our test set could be in their training
which is a single NFoV image with a 90◦ FoV placed at the center set. We also reached out to the authors of ImmerseGAN [5] and
of the panorama, and paired input, which is a pair of NFoVs with acquired their model’s results on our test set.
relative rotation, allowing us to validate our method’s capability
to generate a panorama from multiple input images. We generated 5.2 Quantitative Evaluation
five pairs for each of our training/testing panoramas, resulting in
We examine the performance of our approach from two perspec-
10,000/2,500 pairs of training/testing inputs on SUN360.
tives, namely the accuracy of rotation estimation and the panorama
5.1.2 Training. We train our angle prediction network and latent generation quality.
diffusion model separately. The angle prediction network is trained Rotation Estimation. We evaluate the performance of relative
with the strategies mentioned in 3.2. Our Pano-denoiser and Pano- rotation estimation on SUN360 [39] and Laval Indoor [9] datasets.
Encoder are trained for 7 epochs on pair inputs. For both single For clarity, we separate the input pairs into two categories based on
and pair input, we take in NFoV images of shape 256 × 256 and whether they belong to the ’overlap’ or ’wide baseline’. We compare
produce 1024 × 512 panoramas. During training, we generate input two relative rotation estimation models: our proposed two-stage
text prompts using BLIP [18]. model, and a single-stage model which is designed the same as [2].
Table. 2 reports the evaluation results, which demonstrate that
5.1.3 Metrics. We evaluate our method using two kinds of metrics. the two-stage relative rotation estimation method significantly
Panorama Generation. We use Fréchet Inception Distance (FID) [13, outperforms the single-stage approach in terms of average relative
17] as our quantitative metric since FID can report the visual quality pose estimation error for both ‘overlap’ and ‘wide baseline’. It is
of generated panorama images to some extent. Besides, it has also noted that both methods trained only on the SUN360 dataset.
been adopted by prior studies [1, 11, 37]. Panorama Generation. The primary results are summarized in
Rotation Estimation. Following Cai et al. [2], we evaluate the Table. 1. As demonstrated in the table, our method attains the most
geodesic error of the estimated rotation matrix (denoted by R̂) and favorable FID metrics with all three input types on both datasets.
the ground truth matrix (denoted by R) using arccos(
 (R R̂) −1
). The terms GT rots and Pred rots denote the utilization of ground
2
truth relative angles and predicted angles from our network R ,
5.1.4 Baselines. We compared our method to three previous SOTA respectively. Despite the imperfect nature of the FID metric [17, 26],
methods. Omni-Dreamer [1] is trained on SUN360 with 47938 im- the significant margin of our method’s superiority substantiates its
ages. For the Laval dataset, Omni-Dreamer is finetuned on the model overall effectiveness. It is noteworthy that our model is NOT trained
trained on SUN360 and with 1837 training samples. SIG-SS [11] is on the Laval Indoor dataset, yet it surpasses the performance of
only trained on SUN360 with 50000 training images. Stylelight [37] previous methods that were specifically trained on this dataset.
360-Degree Panorama Generation from Few Unregistered NFoV Images MM ’23, October 29-November 3, 2023, Oawa, ON, Canada

Input SIG-SS Omni-Comp Omni-Adjust ImmerseGAN PanoDiff Ground Truth

Figure 8: Out-painting results on SUN360 Dataset [39]. Omni-Comp and Omni-Adjust denote the CompletionNet and
AdjustmentNet outputs of Omni-Dreamer [1], respectively.
Input StyleLight Omni-Comp Omni-Adjust ImmerseGAN PanoDiff Ground Truth

Figure 9: Out-painting results on Laval Indoor Dataset [9]. We evaluate the models of StyleLight, Omni-Comp, and
Omni-Adjust, by utilizing their officially released models trained on Laval Indoor Dataset. The results obtained from
ImmerseGAN were graciously provided by the authors who ran their model for our evaluation.

5.3 Qualitative Comparison of specific training for handling multiple FoVs. PanoDiff outper-
The visual results are shown in Figure 8. To illustrate the generation forms previous approaches in content consistency, realism, and
quality of our method, we compare our method to previous works texture continuity, maintaining high-quality and realistic content
under three kinds of inputs: single NFoV input(top two rows), the generation.
paired NFoV input with GT rots(middle two rows), and the paired We proceed to look at the generalizing performance of our ap-
NFoV input with Pred rots(bottom two rows). As can be observed proach on Laval Indoor Dataset[9] and present the results in Fig-
from the results, both SIG-SS and Omni-Adjust exhibit inconsisten- ure 9 with both single and paired inputs. Note that our model was
cies in generating boundaries between the input patches and the NOT trained on the Laval Indoor dataset but solely on SUN360.
out-painting content. While the generated images of Omni-Comp Nonetheless, our method still achieves high-quality panorama gen-
show a marginal improvement in boundary issues, they compro- eration with consistent visual properties, e.g., lighting, temperature,
mise the quality and realism of the generated content to achieve and geometrical consistency.
smooth output boundaries. Omni-Adjust produces relatively more 5.4 Ablation Study
realistic content, but it occasionally out-paints the wrong content. We validate the efficacy of key components in our approach, i.e.,
For instance, in the second row, Omni-Adjust wrongly out-paints the denoising strategies, and the circular padding strategy.
a ‘beach’ while the input region indicates a grassland scene. Im- Denoising Strategies in 360-Degree. We conduct quantitative
merseGAN excels at generating satisfactory results with a single comparisons in Table. 3 regarding the usage of our rotation equiv-
FoV input, but struggles with paired FoV input due to the lack ariance loss, and rotating schedule. These denoising strategies assist
MM ’23, October 29-November 3, 2023, Oawa, ON, Canada Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, & Li Song

Inputs ‘Art Gallery’ ‘Seaview Room’ ‘Warehouse’ ‘Garage’

‘Ocean’ ‘Desert’ ‘Snow’ ‘Oasis’

Figure 10: Control panorama generation with text prompts. This demonstrates the capability of our method to generate diverse
high-quality results with various text prompts, while still able to maintain the geometric characteristics of panoramas.

rotate 180° rotate 180° lighting for 3D assets. Figure 12 shows some examples of our gen-
erated panoramas as environmental textures. As can be found, our
method produces diverse panoramas that not only serve as plausible
rendering backgrounds (first two) but also provide environmental
lighting (last two).

w/o Circular Padding Circular Padding

Background
Lighting

Figure 11: Circular padding solves the problem of discontinu-


ity between the left and right sides in panorama generation.
‘Stardust’ ‘Burning’ ‘Tulip Field’ ‘Pond’

our model to understand the geometric characteristics of panoramas


in latent space. As could be seen from the table, both our strategies Figure 12: Examples of panoramic images used as environ-
contribute significantly to the quality of our panorama generation. ment textures.

Table 3: FID↓ results of PanoDiff with different strategies.


Multiple NFoV As Input. Our framework includes a two-stage rel-
‘Equi.’ denotes the usage of our rotation equivariance loss
ative pose estimation module, allowing it to handle multiple NFoV
in Sec. 4.4.1, and ‘Schedule’ denotes our rotating denoising
images (e.g., >2) of the same scene as input. As shown in Figure 1,
schedule in Sec. 4.4.2.
our pipeline can robustly generate high-quality panoramic images
using multiple NFoV images. Please refer to the supplementary for
Strategy Equi. Schedule FID ↓ more generated 360-degree panoramas.
- - 7.88
Choices ✓ - 6.73 6 CONCLUSION
✓ ✓ 6.56
In this paper, we present PanoDiff, a novel framework that gener-
ates 360◦ panoramas from one or more NFoV inputs. The pipeline
Circular Padding For Continuity. When decoding the latent
consists of two main modules, namely rotation estimation and
feature  0 , the inherited border effect caused by the domain gap in
panorama generation. Our two-stage rotation estimation network
latent space and image space frequently occurs. To alleviate this
first classifies the input image pairs into overlap and wide baseline
issue, we implement a circular padding strategy as described in 4.4.3.
scenarios and then performs precise angle prediction. In panorama
In Figure 11, we validate the necessity of circular padding with
generation, we use incomplete partial panoramas along with text
visual result comparisons by rotating 180◦ horizontally from the
prompts as signals to generate diverse panoramas. We hope that
output image. As observed, circular padding seamlessly eliminates
our work inspires further research in panorama generation for ad-
discontinuity between both ends.
vanced applications, including style control, direct HDRI panorama
generation, and related areas.
5.5 Applications
Acknowledgement. This work was supported by the Fundamental
Text Editing. We extend the editing capability of our method. Research Funds for the Central Universities, STCSM under Grant
As depicted in Figure 10, our method not only inherits the power- 22DZ2229005, 111 project BP0719010.
ful text-to-image generation ability from Stable Diffusion but also
preserves the sound geometric properties of panoramas.
Environment Texture. Panoramic images can serve as environ-
ment textures in 3DCG software, which can provide background
360-Degree Panorama Generation from Few Unregistered NFoV Images MM ’23, October 29-November 3, 2023, Oawa, ON, Canada

REFERENCES [22] David G Lowe. 2004. Distinctive image features from scale-invariant keypoints.
[1] Naofumi Akimoto, Yuhi Matsuo, and Yoshimitsu Aoki. 2022. Diverse plausible 360- International journal of computer vision 60 (2004), 91–110.
degree image outpainting for efficient 3dcg background creation. In Proceedings [23] Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte,
of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11441– and Luc Van Gool. 2022. Repaint: Inpainting using denoising diffusion proba-
11450. bilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and
[2] Ruojin Cai, Bharath Hariharan, Noah Snavely, and Hadar Averbuch-Elor. 2021. Pattern Recognition. 11461–11471.
Extreme rotation estimation using dense correlation volumes. In Proceedings of the [24] Wei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun, and Antonio
IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14566–14575. Torralba. 2022. Virtual correspondence: Humans as a cue for extreme-view
[3] Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. 2021. Crossvit: Cross- geometry. In Proceedings of the IEEE/CVF Conference on Computer Vision and
attention multi-scale vision transformer for image classification. In Proceedings Pattern Recognition. 15924–15934.
[25] Ozan Oktay, Jo Schlemper, Loic Le Folgoc, Matthew Lee, Mattias Heinrich, Kazu-
of the IEEE/CVF international conference on computer vision. 357–366.
nari Misawa, Kensaku Mori, Steven McDonagh, Nils Y Hammerla, Bernhard
[4] Zhaoxi Chen, Guangcong Wang, and Ziwei Liu. 2022. Text2light: Zero-shot
Kainz, et al. 2018. Attention u-net: Learning where to look for the pancreas.
text-driven HDR panorama generation. ACM Transactions on Graphics (TOG) 41,
arXiv preprint arXiv:1804.03999 (2018).
6 (2022), 1–16.
[26] Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. 2022. On Aliased Resizing and
[5] Mohammad Reza Karimi Dastjerdi, Yannick Hold-Geoffroy, Jonathan Eisenmann,
Surprising Subtleties in GAN Evaluation. In CVPR.
Siavash Khodadadeh, and Jean-François Lalonde. 2022. Guided Co-Modulated
[27] Alec Radford, Jong Wook Kim, Chris Hallacy, A. Ramesh, Gabriel Goh, Sandhini
GAN for 360° Field of View Extrapolation. In 2022 International Conference on 3D
Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen
Vision (3DV). IEEE, 475–485.
Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From
[6] Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. 2018. Superpoint:
Natural Language Supervision. In ICML.
Self-supervised interest point detection and description. In Proceedings of the
[28] Chris Rockwell, Justin Johnson, and David F. Fouhey. 2022. The 8-Point Algorithm
IEEE conference on computer vision and pattern recognition workshops. 224–236.
as an Inductive Bias for Relative Pose Prediction by ViTs. In 3DV.
[7] Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on
[29] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn
image synthesis. Advances in Neural Information Processing Systems 34 (2021),
Ommer. 2021. High-Resolution Image Synthesis with Latent Diffusion Models.
8780–8794.
arXiv:2112.10752 [[Link]]
[8] Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers
[30] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolu-
for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on
tional networks for biomedical image segmentation. In Medical Image Computing
computer vision and pattern recognition. 12873–12883.
and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference,
[9] Marc-André Gardner, Kalyan Sunkavalli, Ersin Yumer, Xiaohui Shen, Emiliano
Munich, Germany, October 5-9, 2015, Proceedings, Part III 18. Springer, 234–241.
Gambaretto, Christian Gagné, and Jean-François Lalonde. 2017. Learning to
[31] RunwayML. 2021. Stable Diffusion. [Link]
predict indoor illumination from a single image. arXiv preprint arXiv:1704.00090
diffusion.
(2017).
[32] Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi-
[10] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley,
novich. 2020. Superglue: Learning feature matching with graph neural networks.
Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial
In Proceedings of the IEEE/CVF conference on computer vision and pattern recogni-
networks. Commun. ACM 63, 11 (2020), 139–144.
tion. 4938–4947.
[11] Takayuki Hara, Yusuke Mukuta, and Tatsuya Harada. 2022. Spherical Image
[33] Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross
Generation From a Few Normal-Field-of-View Images by Considering Scene
Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell
Symmetry. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022).
Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next
[12] Heathen. [n. d.]. Discussion on Stable Diffusion WebUI. [Link]
generation image-text models. arXiv preprint arXiv:2210.08402 (2022).
automatic1111/stable-diffusion-webui/discussions/2670
[34] Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion
[13] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and
implicit models. arXiv preprint arXiv:2010.02502 (2020).
Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to
[35] Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova,
a local nash equilibrium. Advances in neural information processing systems 30
Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park,
(2017).
and Victor Lempitsky. 2022. Resolution-robust large mask inpainting with fourier
[14] Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic
convolutions. In Proceedings of the IEEE/CVF winter conference on applications of
models. Advances in Neural Information Processing Systems 33 (2020), 6840–6851.
computer vision. 2149–2159.
[15] Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi,
[36] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
and Tim Salimans. 2022. Cascaded Diffusion Models for High Fidelity Image
Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all
Generation. J. Mach. Learn. Res. 23, 47 (2022), 1–33.
you need. Advances in neural information processing systems 30 (2017).
[16] Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv
[37] Guangcong Wang, Yinuo Yang, Chen Change Loy, and Ziwei Liu. 2022. Style-
preprint arXiv:2207.12598 (2022).
light: Hdr panorama generation for lighting estimation and editing. In Computer
[17] Tuomas Kynkäänniemi, Tero Karras, Miika Aittala, Timo Aila, and Jaakko Lehti-
Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022,
nen. 2022. The Role of ImageNet Classes in Fr\’echet Inception Distance. arXiv
Proceedings, Part XV. Springer, 477–492.
preprint arXiv:2203.06026 (2022).
[38] Yinhuai Wang, Jiwen Yu, and Jian Zhang. 2022. Zero-Shot Image Restoration
[18] Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping
Using Denoising Diffusion Null-Space Model. arXiv preprint arXiv:2212.00490
Language-Image Pre-training for Unified Vision-Language Understanding and
(2022).
Generation. In ICML.
[39] Jianxiong Xiao, Krista A Ehinger, Aude Oliva, and Antonio Torralba. 2012. Rec-
[19] Kang Liao, Lang Nie, Shujuan Huang, Chunyu Lin, Jing Zhang, Yao Zhao, Moncef
ognizing scene viewpoint using panoramic place representation. In 2012 IEEE
Gabbouj, and Dacheng Tao. 2023. Deep Learning for Camera Calibration and
Conference on Computer Vision and Pattern Recognition. IEEE, 2695–2702.
Beyond: A Survey. arXiv preprint arXiv:2303.10559 (2023).
[40] Lvmin Zhang and Maneesh Agrawala. 2023. Adding Conditional Control to
[20] Kang Liao, Xiangyu Xu, Chunyu Lin, Wenqi Ren, Yunchao Wei, and Yao Zhao.
Text-to-Image Diffusion Models. arXiv:2302.05543 [[Link]]
2022. Cylin-Painting: Seamless 360 {\deg } Panoramic Image Outpainting and
[41] Shengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong, Xiao Liang, Eric I Chang,
Beyond with Cylinder-Style Convolutions. arXiv preprint arXiv:2204.08563 (2022).
and Yan Xu. 2021. Large Scale Image Completion via Co-Modulated Generative
[21] Guilin Liu, Aysegul Dundar, Kevin J Shih, Ting-Chun Wang, Fitsum A Reda,
Adversarial Networks. In International Conference on Learning Representations
Karan Sapra, Zhiding Yu, Xiaodong Yang, Andrew Tao, and Bryan Catanzaro.
(ICLR).
2022. Partial convolution for padding, inpainting, and image synthesis. IEEE
[42] Shengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong, Xiao Liang, Eric I Chang,
Transactions on Pattern Analysis and Machine Intelligence (2022).
and Yan Xu. 2021. Large scale image completion via co-modulated generative
adversarial networks. arXiv preprint arXiv:2103.10428 (2021).

You might also like