Geometry-Guided Restoration of Defective 360°
Equirectangular Images Using 3D Gaussian
Splatting
Tausif Ansari, Shreetej Meahram, Hamza Khan
Department of Computer Science
Indian Institute of Information Technology, Vadodara
Email: {202351148, 202352333,
2023523404}@[Link]
Abstract—Image sensors in omnidirectional cameras wear and
tear with time and can result in dead pixels or fixed pattern defective image, producing new replacement content for the
noises that can be seen in an equirectangular view form. damaged areas. In order to perform reliable testing using
Traditional methods like interpolation and planar inpainting ground-truth data, our method is tested on a synthetic indoor
does not work for two reasons; the spherical projection model scene created in Blender 4.6, which is consistent with current
cannot be ig- nored, and there is no scene geometry available. practices for novel view synthesis benchmarking [2], [3].
[15] We propose a geometry-guided restoration pipeline that
reconstructs the scene as a 3D Gaussian Splat [1] from a multi- This paper makes three contributions: an end-to-end
view perspective capture and re-renders a clean equirectangular restora- tion pipeline via 3DGS, achieving 94.70%
view from the exact pose of the damaged camera. Our approach correction on a 20 MP image; a synthetic evaluation dataset
is tested on an artificial indoor scene created in Blender 4.6, with with exact ground-truth; and a cubemap-based equirectangular
defect injection performed programmatically to ensure that
ground-truth data is precisely known. Tested at 6144×3240 rendering strategy using six 90° perspective faces stitched via
resolution with 99,122 injected dead pixels (0.497% of image GLSL reprojection. [19].
area), our method reduces the defect count to 5,251 pixels,
achieving a correction ratio of 94.70%, a global MSE of
0.00284, and a PSNR of 25.47 dB. II. RELATED WORK
Index Terms—360° image restoration, 3D Gaussian Splatting,
equirectangular projection, defective pixel correction, novel view A. Sensor Defect Detection and Correction
synthesis, omnidirectional imaging, synthetic evaluation
Pixel-level defects in CMOS sensors arise from man-
I. INTRODUCTION ufacturing imperfections, radiation damage, and electrical
Omnidirectional cameras are used in virtual reality, pho- stress [12]. Classical algorithms replace a suspect pixel with
togrammetry and 360-degree video recording applications. a neighbourhood-interpolated value when it deviates beyond
They typically stitch two or more wide-angle lenses into a sin- a statistical threshold [10]. These methods work well when
gle equirectangular image covering the full 360°×180° field of defects are sparse and isolated, but performance drops quickly
view. When damage occurs — such as burnout or malfunction once defects cluster. FixPix [11] improves on this with a
of the sensor — this information gets baked into all segmentation network for detection and a vision-transformer
subsequent captures. Defects in equirectangular content are for reconstruction, but like all 2D methods it remains blind
particularly damaging because the projection is non-uniform to scene geometry. Demosaicing [13] — interpolating missing
[15]: a pixel cluster near the equator corresponds to a compact colour values in RAW Bayer mosaics — is structurally related
real-world region, while the same cluster near the poles maps but addresses a different degradation model altogether.
to a heavily distorted and visually prominent area.
Existing repair approaches falls into two categories. Low- B. Image Inpainting
level methods [10], [11] replace each bad pixel from its neigh-
bourhood, losing effectiveness when defects cluster spatially Deep inpainting methods [8], [9] use encoder-decoder or
or span structured surfaces. Inpainting methods [8], [9] create diffusion-based architectures to fill irregular holes with
plausible textures but lack access to the precise geometry of seman- tically plausible content. For equirectangular
the scene and therefore cannot guarantee photometric consis- panoramas, the equirectangular projection introduces severe
tency with the captured image. None of these approaches use polar distortion that standard planar networks cannot handle.
information about the 3D structure of the scene: both methods Lee et al. [7] address this by projecting to a cube map — six
operate purely in the 2D image plane. low-distortion perspective faces — and performing inpainting
Our approach is to re-observe the scene physically: a multi- independently on each face. The approach handles large
view perspective sequence reconstructs it as a 3D Gaussian missing regions but hallucinates content with no reference to
Splat. [1].The 3D model of the scene is then projected from scene geometry, and processing faces independently can leave
the exact sensor pose and in equirectangular format like the seams at their boundaries.
C. Novel View Synthesis The repair mask M is derived analytically from the known
NeRF [2] showed that a multi-layer perceptron trained on injection locations:
multi-view images can synthesise photorealistic novel views M (x, y) = 1 ∃ c : |Id(x, y, c) − Iclean(x, y, c)| > 0 (2)
by modelling volumetric radiance. Extensions such as Mip-
NeRF 360 [3] address unbounded scenes and produce state- This approach eliminates the need for mask estimation and al-
of-the-art quality, but training and rendering remain slow. lows the evaluation to focus entirely on reconstruction quality.
3D Gaussian Splatting (3DGS) [1] represents the scene as C. Stage 2: 3DGS Training via SkySplat
a set of anisotropic Gaussians initialised from a sparse SfM
point cloud. Each Gaussian stores position, scale, rotation, The 240 perspective frames are passed to SkySplat [18],
opacity, and view-dependent colour encoded in spherical har- which runs COLMAP [4] for pose estimation followed by
monics [6]. A tile-based differentiable rasteriser composites 3DGS optimisation [1] with adaptive densification and prun-
the Gaussians at real-time frame rates, while adaptive density ing. Training converges to 321,048 Gaussians exported as
control converges training in minutes rather than hours. For a PLY file containing position, rotation, scale, opacity, and
omnidirectional content, ErpGS [16] extends 3DGS with ge- spherical-harmonic color coefficients [6].
ometric regularisation to suppress large anisotropic Gaussians
near the poles.
D. Structure from Motion
COLMAP [4] remains the standard for incremental SfM: it
detects SIFT keypoints [5], matches them across image pairs,
and recovers camera intrinsics and extrinsic poses via bun-
dle adjustment. The SkySplat add-on [18] wraps COLMAP-
based [4] pose estimation with 3DGS training in a single
Blender workflow, forming the training backbone of our
pipeline.
MEta/Gsplat_raw.png
III. METHODOLOGY
Our pipeline has six sequential stages, combined by a
master Python script.
A. Synthetic Scene and Rendering
All source imagery is generated in Blender 4.6 using the
Cycles path-tracer, which produces physically based lighting
and accurate reflections. The indoor scene was designed to
stress-test reconstruction in several ways. Two of the four
walls are mirrored, generating inter-reflections and depth
ambiguity that challenge geometry estimation. A specular
vase on a central table serves as the primary scene anchor, Fig. 1. Trained 3D Gaussian Splat from the 240-frame circular-arc sequence.
while a sofa, floor lamp, and wall posters provide varied
surface texture. Multiple light sources further exercise the
splat’s view- dependent colour model [6]. D. Cubemap Render from the Matched Pose
A panoramic equirectangular camera renders the clean Since standard 3DGS rasterisers require perspective cam-
ground-truth image Iclean at 6144×3240. For the training eras, we obtain equirectangular output via a cubemap inter-
sequence, a perspective camera (25 mm, constrained to track mediate. Six perspective cameras (90° FoV, 1024×1024) are
the vase) is animated in a full 360° arc at varying elevations, placed at the 360° camera’s position and oriented toward the
yielding 240 frames at 1920×1080 — enough to cover every six canonical cube face directions {±X, ±Y, ±Z}. The six
wall face renders are reprojected to a 6144×3240 equirectangular
panorama Is using the GLSL-based Cubemap to Panorama
B. Defect Injection tool [19], which performs a per-pixel spherical-to-cubemap
Dead pixels are injected into Iclean using a coordinate lookup with bilinear interpolation.
Python/OpenCV [17] script. A corruption rate of 0.5%
E. Global Colour Alignment
selects n = ⌊0.005 × H × W ⌋ pixel locations uniformly at
random and zeros all channels: The arc training renders and the equirectangular cam-
era differ slightly in their photometric integration paths,
Id(xi, yi, c) = 0, c ∈ {R, G, B}, i = 1, . . . , n (1)
producing minor colour discrepancies. We align Is to
MEta/[Link] MEta/SBS_comp.png
Fig. 2. Cubemap face renders and the assembled equirectangular splat image
Is. Fig. 3. Synthetically corrupted image and the corresponding repair mask.
Id per channel using mean-and-variance normalisation via via SkySplat. Dead-pixel injection corrupted 99,122 pixels
[Link] [17]: (0.497%), all zeroed across all channels.
σI ,c B. Quantitative Evaluation
I′ s(·, c) = σ d + I (·,
s c) − µIs,c + µId,c (3)
Is,c Table I summarises the results. The defect count was
ε
where µ, σ are computed over non-defective pixels only and reduced from 99,122 to 5,251, a correction ratio of 94.70%.
ε = 10 . The result is clipped to [0, 255].
−6
Most of the uncorrected pixels fall on mirror and specular
surfaces, where the appearance changes substantially between
F. Defect Mask (Real-Sensor Mode) the arc training views and the 360° viewpoint — a variation
In the synthetic setting, M is exact from Stage 1. For real- the spherical harmonic [6] model cannot fully capture.
sensor deployment, the pipeline supports a sliding-window
sta- tistical detector: a pixel is flagged when any channel TABLE I
deviates by more than τσc from the local mean, with dead QUANTITATIVE RESULTS OF THE RESTORATION PIPELINE
pixels caught(by an additional zero-threshold pass:
Metric Value
1 if ∃ c : |I (x, y, c) − µ (x, y)| > τ σ (x,
y)
d c c
M (x, y) = Image resolution 6144 × 3240
0 otherwise Total pixels 19,906,560
(4) Training frames 240
Splat Gaussian count 321,048
with τ = Initial defect count 99,122
3.5. Remaining defect count 5,251
Correction ratio 94.70%
G. Composite and Seam Blending Defect rate (before) 0.497%
Defect rate (after) 0.026%
The repaired image is formed by feathered compositing of Global MSE (norm.) 0.00284
the aligned splat render into the damaged image:
PSNR 25.47 dB
I r(x, y) = 1 − Mf (x, y) · dI (x, y) + M
f (x, y)·s I′ (x, y)
(5)
where Mf is M convolved with a Gaussian kernel (σ = 5 The MSE and PSNR are computed on pixel values nor-
px), implemented via NumPy broadcasting inside the
malised to [0, 1] against the known Iclean. The 25.47 dB
Blender Python API (bpy).
PSNR reflects concentrated error in the specular and mirror
IV. RESULTS AND DISCUSSION regions — areas where the spherical harmonic [6] model
A. Experimental Setup falls short when the test viewpoint departs from the training
distribution. The 99.5% of untouched pixels contribute zero
The ground-truth image Iclean is 6144×3240 (19,906,560 error; the reported PSNR measures reconstruction difficulty in
total pixels), rendered in Blender 4.6 Cycles. The 240-frame the recovered region, not overall image quality. For reference,
arc sequence was used to train a 321,048-Gaussian splat Wang et al. [14] note that PSNR can be a misleading global
indicator when errors are spatially concentrated, which is with physically complex materials, the 321,048-Gaussian splat
precisely the case here. trained by SkySplat [18] on 240 perspective frames achieves
a 94.70% defect correction rate, reducing defect count from
99,122 to 5,251 pixels across a 20-megapixel image. The
25.47 dB PSNR reflects reconstruction error concentrated in
specular and mirror regions — the expected limit of low-order
spherical harmonics [6] at viewpoints not well covered by
training.
The next step is replacing the cubemap intermediate with a
native equirectangular renderer [16] , which would eliminate
face-boundary seams. Beyond that, diffusion-based inpaint-
ing [9] could handle sub-regions occluded in all training
views, and testing on physically damaged sensor hardware
remains the final validation.
MEta/[Link] ACKNOWLEDGMENT
We are grateful to our mentors, Dr. Pratik Shah and Dr.
Pramit Majumdar, for their guidance throughout this project.
Their feedback helped us refine our approach and avoid
several pitfalls. We sincerely appreciate the time and insight
they invested in improving both our understanding and the
final outcome of this work.
REFERENCES
[1] B. Kerbl, G. Kopanas, T. Leimku¨hler, and G. Drettakis, “3D Gaussian
splatting for real-time radiance field rendering,” ACM Trans. Graphics,
vol. 42, no. 4, pp. 1–14, Jul. 2023.
[2] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R.
Fig. 4. Visual comparison before and after restoration over the defective Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural
region. radiance fields for view synthesis,” in Proc. ECCV, 2020, pp. 405–421.
[3] J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman,
“Mip-NeRF 360: Unbounded anti-aliased neural radiance fields,” in
C. Discussion Proc. IEEE/CVF CVPR, 2022, pp. 5470–5479.
[4] J. L. Scho¨nberger and J.-M. Frahm, “Structure-from-motion revisited,”
The key advantage over inpainting is that replacement in Proc. IEEE/CVF CVPR, 2016, pp. 4104–4113.
pixels come from a real 3D reconstruction: depth, perspective, [5] D. G. Lowe, “Distinctive image features from scale-invariant keypoints,”
Int. J. Comput. Vision, vol. 60, no. 2, pp. 91–110, 2004.
and lighting context are all preserved, not guessed from [6] R. Green, “Spherical harmonic lighting: The gritty details,” in Proc.
neighbours. The synthetic evaluation makes this verifiable: Game Developers Conf. (GDC), 2003.
every repaired pixel has a known ground-truth value, so the [7] J. Lee, W. Oh, and Y. Kim, “A 360-degree panoramic image inpainting
network using a cube map,” Computers, Materials & Continua, vol. 66,
reported metrics reflect actual reconstruction error rather than no. 1, pp. 1019–1033, 2020.
a perceptual proxy. [8] F. Li, X. Qin, C. Zhao, and H. Wang, “A review of image inpainting
The principal failure mode is view-dependent surface ap- methods based on deep learning,” Applied Sciences, vol. 13, no. 20,
p. 11189, 2023.
[Link] and mirror surfaces look quite different [9] R. Suvorov, E. Logacheva, A. Mashikhin, A. Remizova, A. Ashukha, A.
from the arc training angles versus the 360° viewpoint. The Silvestrov, N. Kong, H. Goka, K. Park, and V. Lempitsky, “Resolution-
low-order spherical harmonic model [6] handles smooth view- robust large mask inpainting with Fourier convolutions,” in Proc.
IEEE/CVF WACV, 2022, pp. 2184–2193.
dependent colour, but not sharp highlights — which accounts [10] C.-Y. Cho, T.-M. Chen, W.-S. Wang, and C.-N. Liu, “Real-time photo
for the residual 5.3% error. A secondary source of error sensor dead pixel detection for embedded devices,” in Proc. DICTA,
is the cubemap reprojection, which introduces mild seam 2011, pp. 164–169.
[11] A. Grover, M. Karthikeyan, and P. Ienne, “FixPix: Fixing bad pixels
artefacts at cube-face boundaries. These are attenuated by using deep learning,” arXiv preprint arXiv:2310.11637, 2023.
feathered compositing but motivate future adoption of a native [12] J. Nakamura, Ed., Image Sensors and Signal Processing for Digital Still
equirectangular renderer such as ErpGS [16]. Deployment on Cameras. Boca Raton, FL: CRC Press, 2005.
[13] M. Gharbi, G. Chaurasia, S. Paris, and F. Durand, “Deep joint demo-
a real damaged sensor requires only swapping Stage 0 for a saicking and denoising,” ACM Trans. Graphics, vol. 35, no. 6, pp.
physical capture and Stage 5’s exact mask for the statistical 191:1–
detector; all other stages are camera-agnostic. 191:12, 2016.
[14] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image
V. CONCLUSION quality assessment: From error visibility to structural similarity,” IEEE
Trans. Image Process., vol. 13, no. 4, pp. 600–612, Apr. 2004.
We have shown that sensor-defective 360° equirectangular [15] J. P. Snyder, Map Projections — A Working Manual, U.S. Geological
Survey Professional Paper 1395. Washington, DC: U.S. Government
images can be restored by re-rendering damaged regions from Printing Office, 1987.
a 3D Gaussian Splat [1] trained on a circular-arc capture [16] S. Ito, N. Takama, K. Ito, H.-T. Chen, and T. Aoki, “ErpGS:
sequence. Evaluated on a synthetic Blender 4.6 indoor scene Equirectan- gular image rendering enhanced with 3D Gaussian
regularization,” arXiv preprint arXiv:2505.19883, 2025.
[17] G. Bradski, “The OpenCV Library,” Dr. Dobb’s Journal of Software
Tools, 2000. [Online]. Available: [Link]
[18] R. Collins, “SkySplat: A Blender add-on for 3D Gaussian Splat-
ting,” GitHub repository, 2024. [Online]. Available: [Link]
rcollinsfx/SkySplat
[19] D. Makovetsky (danilw), “GLSL Cubemap to Panorama
(Equirectangular) Converter,” Web tool, GitHub Pages, 2024.
[Online]. Available: [Link] to
panorama js/cubemap to [Link]