Structured3D: Photo-realistic 3D Dataset
Structured3D: Photo-realistic 3D Dataset
4
The Pennsylvania State University
[Link]
1 Introduction
Inferring 3D information from 2D sensory data such as images and videos has
long been a central research topic in computer vision. Conventional approach
to building 3D models typically relies on detecting, matching, and triangulating
local image features (e.g., patches, superpixels, edges, and SIFT features). Al-
though significant progress has been made over the past decades, these methods
still suffer from some fundamental problems. In particular, local feature detection
is sensitive to a large number of factors such as scene appearance (e.g., texture-
less areas and repetitive patterns), lighting conditions, and occlusions. Further,
the noisy, point cloud-based 3D model often fails to meet the increasing demand
for high-level 3D understanding in real-world applications.
*: Equal contribution.
†: The work was partially done when Jia Zheng interned at KooLab, [Link].
2 J. Zheng et al.
3D design (a) Cuboid Manhattan world Semantics (b) Configuration Semantic labels (c)
Fig. 1: The Structured3D dataset. From a large collection of house designs (a)
created by professional designers, we automatically extract a variety of ground
truth 3D structure annotations (b) and generate photo-realistic 2D images (c).
floorplans, and room layouts), and also exploit their relations in future data-
driven approaches (e.g., the wireframe formed by intersecting planar surfaces in
the scene).
To create a large-scale dataset with the aim of facilitating research on data-
driven methods for structured 3D scene understanding, we leverage the avail-
ability of professional interior designs and millions of production-level 3D ob-
ject models – all coming with fine geometric details and high-resolution textures
(Fig. 1(a)). We first use computer programs to automatically extract information
about 3D structure from the original house design files. As shown in Fig. 1(b),
our dataset contains rich annotations of 3D room structure including a variety
of geometric primitives and relationships. To further generate photo-realistic 2D
images (Fig. 1(c)), we utilize industry-leading rendering engines to model the
lighting conditions. Currently, our dataset consists of more than 196k images of
21,835 rooms in 3,500 scenes (i.e., houses).
To showcase the usefulness and uniqueness of the proposed Structured3D
dataset, we train deep networks for room layout estimation on a subset of the
dataset. We show that the models trained on both synthetic and real data outper-
form the models trained on real data only. Further, following the spirit of [27,8],
we show how multi-modal annotations in our dataset can benefit domain adap-
tation tasks.
In summary, the main contributions of this paper are:
– We create the Structured3D dataset, which contains rich ground truth 3D
structure annotations of 21,835 rooms in 3,500 scenes, and more than 196k
photo-realistic 2D renderings of the rooms.
– We introduce a unified “primitive + relationship” representation. This rep-
resentation enables us to efficiently capture a wide variety of semi-global or
global 3D structures and their mutual relationships.
– We verify the usefulness of our dataset by using it to train deep networks for
room layout estimation and demonstrating improved performance on public
benchmarks.
4 J. Zheng et al.
(a) Plane [16] (b) Wireframe [12] (c) Cuboid [10] (d) Room layout [39]
(e) Floorplan [17] (f) Abstracted 3D shape (wireframe [32] and cuboid [28])
2 Related Work
like wireframes and room layouts, how to reliably detect them from raw sensor
data remains an active research topic in computer vision.
In recent years, synthetic datasets have played an important role in the suc-
cessful training of deep neural networks. Notable examples for indoor scene un-
derstanding include SUNCG [25], SceneNet RGB-D [20], and InteriorNet [15].
These datasets exceed real datasets in terms of scene diversity and frame num-
bers. But just like their real counterparts, these datasets lack ground truth
structure annotations. Another issue with some synthetic datasets is the de-
gree of realism in both the 3D models and the 2D renderings. [38] shows that
physically-based rendering could boost the performance of various indoor scene
understanding tasks. To ensure the quality of our dataset, we make use of 3D
room models created by professional designers and the state-of-the-art industrial
rendering engines. Table 2 summarizes the differences of 3D scene datasets.
Room layout estimation. Room layout estimation aims to reconstruct the
enclosing structure of the indoor scene, consisting of walls, floor, and ceiling.
Existing public datasets (e.g., PanoContext [37] and LayoutNet [41]) assume
a simple box-shaped layout. PanoContext [37] collects about 500 panoramas
from the SUN360 dataset [33], LayoutNet [41] extends the layout annotations
to include panoramas from 2D-3D-S [3]. Recently, MatterportLayout [42] col-
lects 2,295 RGB-D panoramas from Matterport3D [6] and extends annotations
to Manhattan layout. We note that all room layout in these real datasets is
manually labeled by the human. Since the room structure may be occluded by
furniture and other objects, the “ground truth” inferred by humans may not be
consistent with the actual layout. In our dataset, all ground truth 3D annotations
are automatically extracted from the original house design files.
The main goal of our dataset is to provide rich annotations of ground truth 3D
structure. A naive way to do so is generating and storing different types of 3D an-
notations in the same format as existing works, like wireframes as in [12], planes
6 J. Zheng et al.
(a) Primitives: junctions and lines (b) Primitives: planes (c) Relationships: R1 and R2
Fig. 3: The ground truth 3D structure annotations in our dataset are represented
by primitives and relationships. (a): Junctions and lines. (b): Planes. We high-
light the planes in a single room. (c): Plane-line and line-junction relationships.
We highlight a junction, the three lines intersecting at the junction, and the
planes intersecting at each of the lines. (d): Cuboids. We highlight one cuboid
instance. (e): Manhattan world. We use different colors to denote planes aligned
with different directions. (f ): Semantic objects. We highlight a “room”, a “bal-
cony”, and the “door” connecting them.
as in [16], floorplans as in [17], and so on. But this leads to a lot of redundancy.
For example, planes in man-made environments are often bounded by a number
of line segments, which are part of the wireframe. Even worse, by representing
wireframes and planes separately, the relationships between them are lost. In
this paper, we present a unified representation in order to minimize redundancy
while preserving mutual relationships. We show how the most common types
of structure studied in the literature (e.g., planes, cuboids, wireframes, room
layouts, and floorplans) can be derived from our representation.
Our representation of the structure is largely inspired by the early work of
Witkin and Tenenbaum [31], which characterizes structure as “a shape, pattern,
or configuration that replicates or continues with little or no change over an
interval of space and time”. Accordingly, to describe any structure, we need to
specify: (i) what pattern is continuing or replicating (e.g., a patch, an edge, or
a texture descriptor), and (ii) the domain of its replication or continuation. In
this paper, we call the former primitives and the latter relationships.
and at the same time belong to one of the three classes in the Manhattan world
model.
Discussion. The primitives and relationships we discussed above are just a
few most common examples. They are by no means exhaustive. For example,
our representation can be easily extended to include other primitives such as
parametric surfaces. And besides cuboids, there are many other types of regular
or symmetric shapes in man-made environments, where type corresponds to a
different symmetry group.
Our representation of 3D structures is also related to the graph representa-
tions in semantic scene understanding [13,2,30]. As these graphs focus on seman-
tics, geometry is represented in simplified manners by (i) 6D object poses and
(ii) coarse, discrete spatial relations such as ”supported by”, ”front”, ”back”,
and ”adjacent”. In contrast, our representation focuses on modeling the scene
geometry using fine-grained primitives (i.e., junctions, lines, and planes) and
relationships (in terms of topology and regularities). Thus, it is highly com-
plementary to the scene graphs in prior work. Intuitively, it can be used for
geometric analysis and synthesis tasks, in a similar way as scene graphs are used
for semantic scene understanding.
(a) (b)
Fig. 4: Comparison of 3D house designs. (a): The 3D models in our database are
created by professional designers using high-quality furniture models from world-
leading manufacturers. Most designs are being used in real-world production.
(b): The 3D models in SUNCG dataset [25] are created using Planner 5D [1],
an online tool for amateur interior design.
Since the measurements are highly accurate and noise-free, other types of
relationship such a Manhattan world (R3 ) and cuboids (R4 ) can also be eas-
ily obtained by clustering the primitives, followed by a geometric verification
process. Finally, to include semantic information (R5 ) into our representation,
we map the relevant labels provided by the professional designers to the geo-
metric primitives in our representation. Fig. 3 shows examples of the extracted
geometric primitives and relationships.
Fig. 6: Photo-realistic rendering vs. real-world decoration. The first and third
columns are rendered images.
5 Experiments
5.1 Experiment Setup
To demonstrate the benefits of our dataset, we use it to train deep neural net-
works for room layout estimation, an important task in structured 3D modeling.
12 J. Zheng et al.
Real dataset. We use the same dataset as LayoutNet [41]. The dataset consists
of images from PanoContext [37] and 2D-3D-S [3], including 818 training images,
79 validation images, and 166 test images. Note that both datasets only provide
cuboid layout annotations.
Our Structured3D dataset. In this experiment, we use a subset of panoramas
with the original lighting and full configuration. Each panorama corresponds to
a different room in our dataset. We show statistics of different room layouts in
our dataset in Table 3. Since the current real dataset only contains cuboid layout
annotations (i.e., 4 corners), we choose 12k panoramic images with the cuboid
layout in our dataset. We split the images into 10k for training, 1k for validation,
and 1k for testing.
Evaluation metrics. Following [41,26], we adopt three standard metrics: (i)
3D IoU: intersection over union between predicted 3D layout and the ground
truth, (ii) Corner Error (CE): normalized `2 distance between predicted corner
and ground truth, and (iii) Pixel Error (PE): pixel-wise error between predicted
plane classes and ground truth.
Baselines. We choose two recent CNN-based approaches, LayoutNet [41,42]1
and HorizonNet [26]2 , based on their performance and source code availability.
LayoutNet uses a CNN to predict a corner probability map and a boundary map
from the panorama and vanishing lines, then optimizes the layout parameters
based on network predictions. HorizonNet represents room layout as three 1D
vectors, i.e., boundary positions of floor-wall, and ceiling-wall, and the existence
of wall-wall boundary. It trains CNNs to directly predict the three 1D vectors.
In this paper, we follow the default training setting of the respective methods.
For specific training procedures, please refer to the supplementary materials.
Table 4: Quantitative evaluation under different training schemes. The best and
the second best results are boldfaced and underlined, respectively.
PanoContext 2D-3D-S
Methods Config.
3D IoU (%) ↑ CE (%) ↓ PE (%) ↓ 3D IoU (%) ↑ CE (%) ↓ PE (%) ↓
s 75.64 1.31 4.10 57.18 2.28 7.55
r 84.15 0.64 1.80 83.39 0.74 2.39
LayoutNet [41,42]
s+r 84.96 0.61 1.75 83.66 0.71 2.31
s→r 84.77 0.63 1.89 84.04 0.66 2.08
s 75.89 1.13 3.15 67.66 1.18 3.94
r 83.42 0.73 2.09 84.33 0.64 2.04
HorizonNet [26]
s+r 84.45 0.70 1.89 84.36 0.59 1.90
s→r 85.27 0.66 1.86 86.01 0.61 1.84
datasets with our synthetic data boosts the performance of both networks. We
refer readers to supplementary materials for more qualitative results.
Performance vs. synthetic data size. We further study the relationship
between the number of synthetic images used in pre-training and the accuracy
on the real dataset. We sample 1k, 5k and 10k synthetic images for pre-training,
then fine-tune the model on the real dataset. The results are shown in Table 5.
As expected, using more synthetic data generally improves the performance.
Domain adaptation. Domain adaptation techniques (e.g., [27]) have been
shown to be effective in bridging the performance gap when directly applying
models learned on synthetic data to real environments. In this experiment, we
do not assume access to ground truth layout labels in the real dataset. We adopt
LayoutNet as the task network and use PanoContext and 2D-3D-S separately.
We apply a discriminator network to align the output features of the LayoutNet
for two domains. Inspired by [8], we further leverage multi-modal annotations
in our dataset by adding another decoder branch to the LayoutNet for depth
prediction. We concatenate the boundary, corner, and depth predictions as the
input of the discriminator network. The results are shown in the Table 6. By
incorporating additional information, i.e., depth map, we further boost the per-
formance on both datasets. This illustrates the advantage of including multiple
types of ground truth in our dataset.
Limitation of real datasets. Due to human errors, the annotation in real
datasets is not always consistent with the actual room layout. In the left image
14 J. Zheng et al.
Table 6: Domain adaptation results. NA: non-adaptive baseline. +DA: align lay-
out estimation output. +Depth: align both layout estimation and depth outputs.
Real: train in the target domain.
PanoContext 2D-3D-S
Methods
3D IoU (%) ↑ CE (%) ↓ PE (%) ↓ 3D IoU (%) ↑ CE (%) ↓ PE (%) ↓
NA 75.64 1.31 4.10 57.18 2.28 7.55
+DA 76.91 1.19 3.64 70.08 1.36 4.66
+Depth 78.34 1.03 2.99 72.99 1.24 3.60
Real 81.76 0.95 2.58 81.82 0.96 3.13
of Fig. 7, the room is a non-cuboid layout, but the ground truth layout is labeled
as cuboid shape. In the right image, the front wall is not labeled as ground truth.
These examples illustrate the limitation of using real datasets as benchmarks.
We avoid such errors in our dataset by automatically generating ground truth
from the original design files.
6 Conclusion
In this paper, we present Structured3D, a large synthetic dataset with rich
ground truth 3D structure annotations of 21,835 rooms and more than 196k
photo-realistic 2D renderings. Among many potential use cases of our dataset,
we further demonstrate its benefit in augmenting real data and facilitating do-
main adaptation for the room layout estimation task.
We view this work as an important and exciting step towards building intel-
ligent machines which can achieve human-level holistic 3D scene understanding.
In the future, we will continue to add more 3D structure annotations of the
scenes and objects to the dataset, and explore novel ways to use the dataset to
advance techniques for structured 3D modeling and understanding.
Acknowledgement. We would like to thank [Link] for providing the
database of house designs and the rendering engine. We especially thank Qing
Ye and Qi Wu from [Link] for the help on the data rendering. This
work was partially supported by the National Key R&D Program of China
(#2018AAA0100704) and the National Science Foundation of China (#61932020).
Zihan Zhou was supported by NSF award #1815491.
Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling 15
References
20. McCormac, J., Handa, A., Leutenegger, S., Davison, A.J.: Scenenet RGB-D: can
5m synthetic images beat generic imagenet pre-training on indoor segmentation?
In: ICCV. pp. 2697–2706 (2017) 5
21. Purcell, T.J., Buck, I., Mark, W.R., Hanrahan, P.: Ray tracing on programmable
graphics hardware. ACM Trans. Graph. 21(3), 703–712 (2002) 10
22. Ros, G., Stent, S., Alcantarilla, P.F., Watanabe, T.: Training constrained deconvo-
lutional networks for road scene semantic segmentation. CoRR abs/1604.01545
(2016) 12
23. Silberman, N., Hoiem, D., Kohli, P., Fergus, R.: Indoor segmentation and support
inference from RGBD images. In: ECCV. pp. 746–760 (2012) 4, 5
24. Song, S., Lichtenberg, S.P., Xiao, J.: SUN RGB-D: A RGB-D scene understanding
benchmark suite. In: CVPR. pp. 567–576 (2015) 4, 5
25. Song, S., Yu, F., Zeng, A., Chang, A.X., Savva, M., Funkhouser, T.A.: Semantic
scene completion from a single depth image. In: CVPR. pp. 1746–1754 (2017) 5,
9, 10
26. Sun, C., Hsiao, C.W., Sun, M., Chen, H.T.: Horizonnet: Learning room layout with
1d representation and pano stretch data augmentation. In: CVPR. pp. 1047–1056
(2019) 2, 12, 13
27. Tsai, Y.H., Hung, W.C., Schulter, S., Sohn, K., Yang, M.H., Chandraker, M.:
Learning to adapt structured output space for semantic segmentation. In: CVPR.
pp. 7472–7481 (2018) 3, 13
28. Tulsiani, S., Su, H., Guibas, L.J., Efros, A.A., Malik, J.: Learning shape abstrac-
tions by assembling volumetric primitives. In: CVPR. pp. 2635–2643 (2017) 2,
4
29. Wald, I., Woop, S., Benthin, C., Johnson, G.S., Ernst, M.: Embree: a kernel frame-
work for efficient CPU ray tracing. ACM Trans. Graph. 33(4), 143:1–143:8 (2014)
10
30. Wang, K., Lin, Y.A., Weissmann, B., Savva, M., Chang, A.X., Ritchie, D.: Planit:
Planning and instantiating indoor scenes with relation graph and spatial prior
networks. ACM Trans. Graph. 38(4) (2019) 8
31. Witkin, A.P., Tenenbaum, J.M.: On the role of structure in vision. In: Beck, J.,
Hope, B., Rosenfeld, A. (eds.) Human and Machine Vision, pp. 481–543. Academic
Press (1983) 6
32. Wu, J., Xue, T., Lim, J.J., Tian, Y., Tenenbaum, J.B., Torralba, A., Freeman,
W.T.: 3d interpreter networks for viewer-centered wireframe modeling. IJCV
126(9), 1009–1026 (2018) 2, 4
33. Xiao, J., Ehinger, K.A., Oliva, A., Torralba, A.: Recognizing scene viewpoint using
panoramic place representation. In: CVPR. pp. 2695–2702 (2012) 5
34. Xiao, J., Russell, B., Torralba, A.: Localizing 3d cuboids in single-view images. In:
NeurIPS. pp. 746–754 (2012) 3
35. Yang, F., Zhou, Z.: Recovering 3d planes from a single image via convolutional
neural networks. In: ECCV. pp. 87–103 (2018) 2
36. Yu, Z., Zheng, J., Lian, D., Zhou, Z., Gao, S.: Single-image piece-wise planar 3d
reconstruction via associative embedding. In: CVPR. pp. 1029–1037 (2019) 2
37. Zhang, Y., Song, S., Tan, P., Xiao, J.: Panocontext: A whole-room 3d context
model for panoramic scene understanding. In: ECCV. pp. 668–686 (2014) 3, 5, 12
38. Zhang, Y., Song, S., Yumer, E., Savva, M., Lee, J.Y., Jin, H., Funkhouser, T.:
Physically-based rendering for indoor scene understanding using convolutional neu-
ral networks. In: CVPR. pp. 5287–5295 (2017) 5
39. Zhang, Y., Yu, F., Song, S., Xu, P., Seff, A., Xiao, J.: Large-scale scene under-
standing challenge: Room layout estimation (2016) 3, 4
Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling 17
40. Zhou, Y., Qi, H., Zhai, S., Sun, Q., Chen, Z., Wei, L.Y., Ma, Y.: Learning to
reconstruct 3d manhattan wireframes from a single image. In: ICCV. pp. 7698–
7707 (2019) 2, 3, 4
41. Zou, C., Colburn, A., Shan, Q., Hoiem, D.: Layoutnet: Reconstructing the 3d room
layout from a single RGB image. In: CVPR. pp. 2051–2059 (2018) 2, 3, 5, 12, 13
42. Zou, C., Su, J., Peng, C., Colburn, A., Shan, Q., Wonka, P., Chu, H., Hoiem,
D.: 3d manhattan room layout reconstruction from a single 360 image. CoRR
abs/1910.04099 (2019) 3, 5, 12, 13