Generate 360° Panoramas from NFoV Images
Generate 360° Panoramas from NFoV Images
Images
Jionghao Wang∗ Ziyu Chen∗ Jun Ling
Shanghai Jiao Tong University Shanghai Jiao Tong University Shanghai Jiao Tong University
shanemankiw@[Link] 1252060456@[Link] lingjun@[Link]
Figure 1: Illustration of our PanoDiff, a novel approach that is capable of synthesizing fine-grained and diverse 360-degree
panorama from a(few) unregistered NFoV (Narrow Field of View) image(s) and text prompts.
ABSTRACT overcome the limitations. Firstly, a two-stage angle prediction mod-
360◦ panoramas are extensively utilized as environmental light ule to handle various numbers of NFoV inputs. Secondly, a novel
sources in computer graphics. However, capturing a 360◦ × 180◦ latent diffusion-based panorama generation model uses incomplete
panorama poses challenges due to the necessity of specialized and panorama and text prompts as control signals and utilizes several
costly equipment, and additional human resources. Prior studies geometric augmentation schemes to ensure geometric properties
develop various learning-based generative methods to synthesize in generated panoramas. Experiments show that PanoDiff achieves
panoramas from a single Narrow Field-of-View (NFoV) image, but state-of-the-art panoramic generation quality and high controlla-
they are limited in alterable input patterns, generation quality, and bility, making it suitable for applications such as content editing.
controllability. To address these issues, we propose a novel pipeline
called PanoDiff, which efficiently generates complete 360◦ panora- CCS CONCEPTS
mas using one or more unregistered NFoV images captured from • Computing methodologies → Computer vision; Computer
arbitrary angles. Our approach has two primary components to graphics.
∗ Indicatesequal contribution.
† Corresponding
KEYWORDS
author. Affiliated with both School of Electronic Information and
Electrical Engineering, Shanghai Jiao Tong University & MoE Key Lab of Artificial 360-degree panorama, generative models, multimodal models, la-
Intelligence, AI Institute, Shanghai Jiao Tong University. tent diffusion, image pose estimation
Jionghao Wang, Ziyu Chen, Jun Ling and Rong Xie are with School of Electronic
Information and Electrical Engineering, Shanghai Jiao Tong University. ACM Reference Format:
Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, and Li Song. 2023. 360-
Permission to make digital or hard copies of all or part of this work for personal or Degree Panorama Generation from Few Unregistered NFoV Images. In
classroom use is granted without fee provided that copies are not made or distributed
for profit or commercial advantage and that copies bear this notice and the full citation Proceedings of the 31st ACM International Conference on Multimedia (MM
on the first page. Copyrights for components of this work owned by others than the ’23), October 29-November 3, 2023, Ottawa, ON, Canada. ACM, New York,
author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or NY, USA, 9 pages. [Link]
republish, to post on servers or to redistribute to lists, requires prior specific permission
and/or a fee. Request permissions from permissions@[Link].
MM ’23, October 29-November 3, 2023, Ottawa, ON, Canada 1 INTRODUCTION
© 2023 Copyright held by the owner/author(s). Publication rights licensed to ACM.
ACM ISBN 979-8-4007-0108-5/23/10. . . $15.00 Panoramic images, which capture an extensive field of view en-
[Link] compassing a full 360◦ horizontal by 180◦ vertical FoV scene, have
MM ’23, October 29-November 3, 2023, Oawa, ON, Canada Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, & Li Song
become increasingly significant across various applications, such as we know, PanoDiff is the first latent-diffusion-based panorama
as environment lighting, VR/AR, and autonomous driving system. outpainting model that can not only handle incomplete panoramas
However, obtaining high-quality panoramic images can be both of various shapes as input but also supports text prompts. Finally,
time-consuming and costly, as they typically necessitate the use of PanoDiff outperforms existing methods in terms of quantitative
specialized panoramic cameras or stitching software to combine and qualitative results on different input types, as demonstrated
images from multiple perspectives. Our method addresses two main through abundant experiments.
limitations regarding previous generating methods, namely input
pattern and generation quality & controllability. 2 RELATED WORKS
For Input pattern. Most previous works [1, 5, 37] only support 2.1 Rotation Estimation
a single FoV region in the center of an incomplete panorama as
input. However, relying solely on a single input region restricts Camera pose estimation has been a long-standing task in computer
flexibility when controlling specific contents of the generated scene. vision [19]. Most prior works focus on identifying various visual
Such flexibility is particularly important for applications that re- cues from pairs of images. For instance, the classic SIFT [22] ex-
quire precise control over the visual elements and ensure accurate tracts scale and direction invariant features from input images and
representations of the intended scene. matches them to find correspondences. Some approaches utilize
In terms of generation quality & controllability. Generating a com- neural networks for more robust feature extraction or graph neu-
plete 360-degree panorama from NFoV images could be viewed as a ral networks for enhanced matching [6, 28, 32]. However, these
large-hole inpainting problem [35]. Previous methods [1, 11, 37] typ- methods do not account for extreme rotations where the input
ically all rely on GAN(Generative Adversarial Networks) [10] based images have little to no overlap, known as wide baseline scenarios.
methods. Besides, previous methods approach this problem as an Some earlier works focus on mining specially handcrafted cues [24]
image-conditioned generation task, leaving their pipeline with little to address this problem, but they are not generalizable for scenes
control and flexibility over the generation results. A more recent without a specific clue. The work that inspired us the most is [2],
work [5] proposes to use additional text guidance for a GAN-based which estimates pair-wise rotation angles in both overlap and wide
image inpainting method [41]. However, diffusion-based image baseline scenarios. However, the estimation network occasionally
generation methods [14] have shown impressive results on various misidentifies an overlapping pair of images as non-overlapping,
generative tasks, and models trained on large datasets [29, 33] have which severely impacts the generation quality later on.
shown better performance and robustness against GAN-based mod-
els [7, 15]. Moreover, GANs have limited mode coverage and are 2.2 Diffusion Models
difficult to scale for modeling complex multimodal distributions. Diffusion models have shown impressive ability in generative tasks [7,
Unlike GANs, likelihood-based models such as diffusion models 14]. Generally, diffusion models represent the image generative pro-
are capable of learning the complex distribution of natural images, cess as a denoising process, where a noise map sampled from white
resulting in the generation of high-quality images, as mentioned Gaussian noise is iteratively denoised using a learned prior distri-
in [7, 29]. bution ( −1 | ). The objective function for training diffusion
To overcome these limitations, a pipeline capable of accepting models is simplified as
a flexible number of NFoV image(s) as input and generating high-
simple = , 0 , | − ( , )| 2 (1)
fidelity panoramas is crucial. Nevertheless, this endeavor presents
two main challenges: 1) estimating relative camera poses and accu- which has been shown to be approximately equivalent to optimizing
rately warping the input NFoV image(s) on the panorama, and 2) the prior probability distribution’s variational lower bound [14].
using a latent-diffusion-model-based method to generate the entire To further improve the efficiency and effectiveness of the dif-
panorama from input partial panorama of various shapes. fusion model, recent work has proposed to conduct the denois-
This paper introduces PanoDiff, a novel pipeline that efficiently ing process in latent space instead of pixel space [29]. In addition,
generates complete 360° panoramas using one or more unregis- they integrate multi-head cross-attention mechanism [3, 36] in de-
tered NFoV images captured from arbitrary angles. The proposed noising U-Net [25, 30] blocks. This approach enables features of
method overcomes the limitations of existing methods, it enables different modalities to guide the generation process[16], e.g. CLIP-
the generation of high-quality panoramas from incomplete 360- ViT [27]. Working on latent space along with some other sampling
degree panoramas that warped from any number of NFoV inputs. schedule [34] could also potentially make the sampling more ef-
This is achieved by addressing two key challenges. First, we pro- ficient. Controlling the pre-trained model parameters by training
pose a robust two-stage angle prediction pipeline that classifies hypernetworks[12, 40] can improve its performance on a more
image pairs based on their overlap before regressing specific angle specific task while preserving its original generative capability.
values. Second, we train a hypernetwork that controls a pre-trained
large latent-diffusion model, and utilizes geometric augmentation 2.3 Diffusion-based Inpainting
schemes during both training and inference sampling phases to The task of inpainting was predominantly been done by GAN-
ensure the geometric properties of the generated panoramas. based models [21, 35, 42]. Recently, diffusion-based methods have
We summarize our contributions as follows. Firstly, the first been proposed and have already achieved promising results [23,
flexible framework for generating panoramic images from single 38]. However, [23, 38] operate the diffusion and sampling process
or multiple NFoV inputs. A two-stage pose estimation module is solely on image space rather than latent space, which limits their
specially designed for relative pose estimation. Secondly, as far generation flexibility. Recently, [31] offered an image inpainting
360-Degree Panorama Generation from Few Unregistered NFoV Images MM ’23, October 29-November 3, 2023, Oawa, ON, Canada
Figure 2: An overview of generating panorama from a few NFoV images. We first calculate their relative rotations based on a
two-stage angle prediction network R (Sec. 3), then project them using backward equirectangular projection Fproj (Sec. 4) to
obtain partial panorama and visibility mask . Finally, we feed [, ] along with text prompts to our control-based latent
diffusion model FLDM (Sec. 4) and sample iteratively in our rotating schedule (Sec. 4.4.2) to get the final generated panorama.
model finetuned from Stable Diffusion [29] and demonstrates strong image 0 as an anchor and estimate the relative poses [Δ , Δ,
text-controlled inpainting ability. Δ ] of the remaining images with respect to the anchor. Follow-
ing the assumptions made in [2], we consider that cameras are
2.4 Panorama Generation typically upright, allowing us to estimate the absolute pitch an-
With recent advancements in deep generative neural networks, gle (vertical/latitude) instead of the relative angle. We also assume
several works adopt different forms of GAN-based methods to that the camera roll angles are static and remain unchanged at 0,
generate full panoramas [1, 5, 20, 37], such as VQGAN [8] and i.e., ∀ ∈ 0, ..., − 1, = 0. Consequently, our network R can be
CoModGAN [41]. However, these methods are limited to single- expressed as:
view input, in contrast to our alterable input patterns. One work
R ( 0, ) → [Δ→0, 0, ] (3)
has utilized a few NFoV images as input [11] with accurate camera
positions already given, thus cannot deal with unregistered camera where Δ→0 is the relative longitude angle between the images,
inputs. and 0, are the absolute lattitude angle for 0 and , respectively.
Recently, some researchers have utilized diffusion models to
generate panoramas from text prompts alone [4]. However, this
work does not support partial FoV as input, and thus users have 3.2 Two-stage Angle Prediction
limited control over the outcome. In the context of FoV overlapping, we categorize image-pair re-
lationships into two distinct types: overlap and wide baseline. In
3 RELATIVE POSE ESTIMATION overlap scenarios (green pair in Fig. 3), two NFoV images share
a portion of their FoV, resulting in an overlap on the panorama.
3.1 Problem Formulation Conversely, in wide baseline situations (red pair in Fig. 3), the NFoV
Unlike previous panorama generation pipelines, our method takes image pair does not have overlapping FoV, leading to two separate
either single or multiple NFoV images as input. In the case of mul- regions on the panorama. The ultimate goal of our method is to
tiple NFoV (Narrow Field-of-View) images captured at the same create a plausible panorama for further applications. However, a
location but in different poses, our objective is to estimate the rela- single regression network [2] is not robust enough to differenti-
tive poses between them. In the context of omnidirectional images, ate between these two cases. Consequently, images without any
where no camera translation is involved, we formulate the relative overlap might sometimes be predicted to have an overlap, creating
camera poses as rotation angles. Specifically, we view the 360◦ severe artifacts in the input for our controlled LDM, which would
panorama as a spherical map and express the rotation angles as result in unsatisfactory results. Furthermore, the precision of a
the 1) lookat direction, which includes the longitude , latitude , single regression network is insufficient.
and 2) roll angle . For a single image, our pipeline estimates only To address these problems, we propose a two-stage angle pre-
and places it in lateral center, i.e., = = 0. Formally, for a set diction pipeline that initially performs classification to determine
of NFoV images = 0, 1, ..., −1 , we aim to learn a model R whether the two images have an overlap, and subsequently re-
that predicts their relative angles: gresses the precise angles [Δ→0, 0, ]. This approach signifi-
R (I) → [ , , ] (2) cantly reduces the issue of wide baseline images being mistakenly
predicted to overlap. The overview of our rotation prediction mod-
where = { }=0,..., −1 , = { }=0,..., −1 and = { }=0,..., −1 , ule can be seen in Fig. 3. We adopt the same backbone structure
respectively. We decompose this task into estimating pair-wise cam- as [2], which consists of a feature encoder and a 4D dense cor-
era relative rotation, which is scalable and can be applied to pair relation module. Within the two-stage pipeline, a total of three
input and sometimes more than two images. We design a pairwise backbones are utilized, each with its own prediction head. To esti-
angle prediction network that estimates the relative angles between mate the relative camera rotation between a pair of images, we first
images. For a set of NFoV images ( 0, 1, ..., −1 ), we select one pass the image pair to our classifier to determine if the images are
MM ’23, October 29-November 3, 2023, Oawa, ON, Canada Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, & Li Song
4 PANORAMA GENERATION
4.1 Formulation
With Sec. 3, we can formulate a set of image-angle pairs, i.e., {I0, 0, 0 }
and {Ii, Δ→0, }=1,..., −1 . In this part, our goal is to take the im-
ages and their relative angles to produce a complete 360◦ × 180◦
panorama. We divide this problem into two parts: 1) project the
image set = {Ii } onto the omnidirectional map based on their Figure 4: An intuitive explanation of integrating a new con-
estimated angles = { }=0,..., −1 and = { }=0,..., −1 ; and 2) trol signal into an existing model block.
generate the full panorama using the incomplete input as a control
signal. Specifically, we formulate this problem as:
Fproj ( , , ) → , , 4.3 Controlling LDM with Partial FoV
(4) Given a noisy latent space feature at time step , we want to
FLDM (, , text) → Pano,
learn a denoiser that could predict the noise as in:
where Fproj and FLDM represent the projection and diffusion gener-
ation operations, respectively. The projection operation Fproj takes = ( , , ctext, cP ) (6)
the NFoV image set and their angles as input and produces a partial Here, ctext and cP = [, ] represent the text conditions and input
panorama image and a visibility mask , indicating which parts panorama conditions, respectively. We introduce two components
of the FoV are missing. Subsequently, the LDM (Latent Diffusion for our partial-FoV-controlled LDM, namely Pano-Denoiser and
Model) operation FLDM accepts the incomplete , visibility mask , Pano-Encoder E .
and a text prompt as input, iteratively generating the full panorama Our Pano-Denoiser is designed as discussed in Sec. 4.2, where we
Pano. employ control units to introduce control signals to the pre-trained
Stable Diffusion model. In practice, we follow the approach of [40]
4.2 Recap: Controlling Stable Diffusion and add control units to the four encoder blocks and one middle
We propose a generative process based on training a hypernetwork block of the denoising U-Net model used in Stable Diffusion, while
over a pre-trained Stable Diffusion (SD) model[29], a large text-to- utilizing zero-convolutions for the other four decoder blocks.
image generative model utilizing latent diffusion. Latent diffusion In addition to the Pano-Denoiser, we incorporate a shallow Pano-
models iteratively perform denoise sampling in latent space. At Encoder E to construct our controlling condition. In order to feed
360-Degree Panorama Generation from Few Unregistered NFoV Images MM ’23, October 29-November 3, 2023, Oawa, ON, Canada
Table 1: FID↓ results compared with other generation methods for quantitative evaluation.
After the sampling process, the decoder generates an image Table 2: Evaluation of relative pose estimation on
of shape ( + 2 ) × ℎ, where is the extra width from the SUN360 [39] and Laval Indoor [9]. We present the Av-
padded feature. To produce a standard panorama, the extra width erage geodesic error Avg(◦ ↓) and the percentage of pairs with
is removed from the final image. errors under 10 degrees 10◦ (%↑).
Figure 8: Out-painting results on SUN360 Dataset [39]. Omni-Comp and Omni-Adjust denote the CompletionNet and
AdjustmentNet outputs of Omni-Dreamer [1], respectively.
Input StyleLight Omni-Comp Omni-Adjust ImmerseGAN PanoDiff Ground Truth
Figure 9: Out-painting results on Laval Indoor Dataset [9]. We evaluate the models of StyleLight, Omni-Comp, and
Omni-Adjust, by utilizing their officially released models trained on Laval Indoor Dataset. The results obtained from
ImmerseGAN were graciously provided by the authors who ran their model for our evaluation.
5.3 Qualitative Comparison of specific training for handling multiple FoVs. PanoDiff outper-
The visual results are shown in Figure 8. To illustrate the generation forms previous approaches in content consistency, realism, and
quality of our method, we compare our method to previous works texture continuity, maintaining high-quality and realistic content
under three kinds of inputs: single NFoV input(top two rows), the generation.
paired NFoV input with GT rots(middle two rows), and the paired We proceed to look at the generalizing performance of our ap-
NFoV input with Pred rots(bottom two rows). As can be observed proach on Laval Indoor Dataset[9] and present the results in Fig-
from the results, both SIG-SS and Omni-Adjust exhibit inconsisten- ure 9 with both single and paired inputs. Note that our model was
cies in generating boundaries between the input patches and the NOT trained on the Laval Indoor dataset but solely on SUN360.
out-painting content. While the generated images of Omni-Comp Nonetheless, our method still achieves high-quality panorama gen-
show a marginal improvement in boundary issues, they compro- eration with consistent visual properties, e.g., lighting, temperature,
mise the quality and realism of the generated content to achieve and geometrical consistency.
smooth output boundaries. Omni-Adjust produces relatively more 5.4 Ablation Study
realistic content, but it occasionally out-paints the wrong content. We validate the efficacy of key components in our approach, i.e.,
For instance, in the second row, Omni-Adjust wrongly out-paints the denoising strategies, and the circular padding strategy.
a ‘beach’ while the input region indicates a grassland scene. Im- Denoising Strategies in 360-Degree. We conduct quantitative
merseGAN excels at generating satisfactory results with a single comparisons in Table. 3 regarding the usage of our rotation equiv-
FoV input, but struggles with paired FoV input due to the lack ariance loss, and rotating schedule. These denoising strategies assist
MM ’23, October 29-November 3, 2023, Oawa, ON, Canada Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, & Li Song
Figure 10: Control panorama generation with text prompts. This demonstrates the capability of our method to generate diverse
high-quality results with various text prompts, while still able to maintain the geometric characteristics of panoramas.
rotate 180° rotate 180° lighting for 3D assets. Figure 12 shows some examples of our gen-
erated panoramas as environmental textures. As can be found, our
method produces diverse panoramas that not only serve as plausible
rendering backgrounds (first two) but also provide environmental
lighting (last two).
Background
Lighting
REFERENCES [22] David G Lowe. 2004. Distinctive image features from scale-invariant keypoints.
[1] Naofumi Akimoto, Yuhi Matsuo, and Yoshimitsu Aoki. 2022. Diverse plausible 360- International journal of computer vision 60 (2004), 91–110.
degree image outpainting for efficient 3dcg background creation. In Proceedings [23] Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte,
of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11441– and Luc Van Gool. 2022. Repaint: Inpainting using denoising diffusion proba-
11450. bilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and
[2] Ruojin Cai, Bharath Hariharan, Noah Snavely, and Hadar Averbuch-Elor. 2021. Pattern Recognition. 11461–11471.
Extreme rotation estimation using dense correlation volumes. In Proceedings of the [24] Wei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun, and Antonio
IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14566–14575. Torralba. 2022. Virtual correspondence: Humans as a cue for extreme-view
[3] Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. 2021. Crossvit: Cross- geometry. In Proceedings of the IEEE/CVF Conference on Computer Vision and
attention multi-scale vision transformer for image classification. In Proceedings Pattern Recognition. 15924–15934.
[25] Ozan Oktay, Jo Schlemper, Loic Le Folgoc, Matthew Lee, Mattias Heinrich, Kazu-
of the IEEE/CVF international conference on computer vision. 357–366.
nari Misawa, Kensaku Mori, Steven McDonagh, Nils Y Hammerla, Bernhard
[4] Zhaoxi Chen, Guangcong Wang, and Ziwei Liu. 2022. Text2light: Zero-shot
Kainz, et al. 2018. Attention u-net: Learning where to look for the pancreas.
text-driven HDR panorama generation. ACM Transactions on Graphics (TOG) 41,
arXiv preprint arXiv:1804.03999 (2018).
6 (2022), 1–16.
[26] Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. 2022. On Aliased Resizing and
[5] Mohammad Reza Karimi Dastjerdi, Yannick Hold-Geoffroy, Jonathan Eisenmann,
Surprising Subtleties in GAN Evaluation. In CVPR.
Siavash Khodadadeh, and Jean-François Lalonde. 2022. Guided Co-Modulated
[27] Alec Radford, Jong Wook Kim, Chris Hallacy, A. Ramesh, Gabriel Goh, Sandhini
GAN for 360° Field of View Extrapolation. In 2022 International Conference on 3D
Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen
Vision (3DV). IEEE, 475–485.
Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From
[6] Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. 2018. Superpoint:
Natural Language Supervision. In ICML.
Self-supervised interest point detection and description. In Proceedings of the
[28] Chris Rockwell, Justin Johnson, and David F. Fouhey. 2022. The 8-Point Algorithm
IEEE conference on computer vision and pattern recognition workshops. 224–236.
as an Inductive Bias for Relative Pose Prediction by ViTs. In 3DV.
[7] Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on
[29] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn
image synthesis. Advances in Neural Information Processing Systems 34 (2021),
Ommer. 2021. High-Resolution Image Synthesis with Latent Diffusion Models.
8780–8794.
arXiv:2112.10752 [[Link]]
[8] Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers
[30] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolu-
for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on
tional networks for biomedical image segmentation. In Medical Image Computing
computer vision and pattern recognition. 12873–12883.
and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference,
[9] Marc-André Gardner, Kalyan Sunkavalli, Ersin Yumer, Xiaohui Shen, Emiliano
Munich, Germany, October 5-9, 2015, Proceedings, Part III 18. Springer, 234–241.
Gambaretto, Christian Gagné, and Jean-François Lalonde. 2017. Learning to
[31] RunwayML. 2021. Stable Diffusion. [Link]
predict indoor illumination from a single image. arXiv preprint arXiv:1704.00090
diffusion.
(2017).
[32] Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi-
[10] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley,
novich. 2020. Superglue: Learning feature matching with graph neural networks.
Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial
In Proceedings of the IEEE/CVF conference on computer vision and pattern recogni-
networks. Commun. ACM 63, 11 (2020), 139–144.
tion. 4938–4947.
[11] Takayuki Hara, Yusuke Mukuta, and Tatsuya Harada. 2022. Spherical Image
[33] Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross
Generation From a Few Normal-Field-of-View Images by Considering Scene
Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell
Symmetry. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022).
Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next
[12] Heathen. [n. d.]. Discussion on Stable Diffusion WebUI. [Link]
generation image-text models. arXiv preprint arXiv:2210.08402 (2022).
automatic1111/stable-diffusion-webui/discussions/2670
[34] Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion
[13] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and
implicit models. arXiv preprint arXiv:2010.02502 (2020).
Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to
[35] Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova,
a local nash equilibrium. Advances in neural information processing systems 30
Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park,
(2017).
and Victor Lempitsky. 2022. Resolution-robust large mask inpainting with fourier
[14] Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic
convolutions. In Proceedings of the IEEE/CVF winter conference on applications of
models. Advances in Neural Information Processing Systems 33 (2020), 6840–6851.
computer vision. 2149–2159.
[15] Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi,
[36] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
and Tim Salimans. 2022. Cascaded Diffusion Models for High Fidelity Image
Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all
Generation. J. Mach. Learn. Res. 23, 47 (2022), 1–33.
you need. Advances in neural information processing systems 30 (2017).
[16] Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv
[37] Guangcong Wang, Yinuo Yang, Chen Change Loy, and Ziwei Liu. 2022. Style-
preprint arXiv:2207.12598 (2022).
light: Hdr panorama generation for lighting estimation and editing. In Computer
[17] Tuomas Kynkäänniemi, Tero Karras, Miika Aittala, Timo Aila, and Jaakko Lehti-
Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022,
nen. 2022. The Role of ImageNet Classes in Fr\’echet Inception Distance. arXiv
Proceedings, Part XV. Springer, 477–492.
preprint arXiv:2203.06026 (2022).
[38] Yinhuai Wang, Jiwen Yu, and Jian Zhang. 2022. Zero-Shot Image Restoration
[18] Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping
Using Denoising Diffusion Null-Space Model. arXiv preprint arXiv:2212.00490
Language-Image Pre-training for Unified Vision-Language Understanding and
(2022).
Generation. In ICML.
[39] Jianxiong Xiao, Krista A Ehinger, Aude Oliva, and Antonio Torralba. 2012. Rec-
[19] Kang Liao, Lang Nie, Shujuan Huang, Chunyu Lin, Jing Zhang, Yao Zhao, Moncef
ognizing scene viewpoint using panoramic place representation. In 2012 IEEE
Gabbouj, and Dacheng Tao. 2023. Deep Learning for Camera Calibration and
Conference on Computer Vision and Pattern Recognition. IEEE, 2695–2702.
Beyond: A Survey. arXiv preprint arXiv:2303.10559 (2023).
[40] Lvmin Zhang and Maneesh Agrawala. 2023. Adding Conditional Control to
[20] Kang Liao, Xiangyu Xu, Chunyu Lin, Wenqi Ren, Yunchao Wei, and Yao Zhao.
Text-to-Image Diffusion Models. arXiv:2302.05543 [[Link]]
2022. Cylin-Painting: Seamless 360 {\deg } Panoramic Image Outpainting and
[41] Shengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong, Xiao Liang, Eric I Chang,
Beyond with Cylinder-Style Convolutions. arXiv preprint arXiv:2204.08563 (2022).
and Yan Xu. 2021. Large Scale Image Completion via Co-Modulated Generative
[21] Guilin Liu, Aysegul Dundar, Kevin J Shih, Ting-Chun Wang, Fitsum A Reda,
Adversarial Networks. In International Conference on Learning Representations
Karan Sapra, Zhiding Yu, Xiaodong Yang, Andrew Tao, and Bryan Catanzaro.
(ICLR).
2022. Partial convolution for padding, inpainting, and image synthesis. IEEE
[42] Shengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong, Xiao Liang, Eric I Chang,
Transactions on Pattern Analysis and Machine Intelligence (2022).
and Yan Xu. 2021. Large scale image completion via co-modulated generative
adversarial networks. arXiv preprint arXiv:2103.10428 (2021).