Segment Any Mesh
George Tang William Zhao Logan Ford David Benhaim Paul Zhang
MIT MIT Backflip AI Backflip AI MIT
arXiv:2408.13679v2 [[Link]] 9 Mar 2025
Figure 1. Our method, Segment Any Mesh, lifts per view 2D segmentations into a mesh part segmentation. We show examples on our
diverse object dataset curated from a 3D generative model. Segment Any Mesh generalizes well to novel shapes and object classes.
Abstract pare our method with a robust, well-evaluated shape anal-
ysis method, Shape Diameter Function, and show that our
method is comparable to or exceeds its performance. Since
We propose Segment Any Mesh, a novel zero-shot mesh
current benchmarks contain limited object diversity, we also
part segmentation method that overcomes the limitations
curate and release a dataset of generated meshes and use it
of shape analysis-based, learning-based, and contemporary
to demonstrate our method’s improved generalization over
approaches. Our approach operates in two phases: multi-
Shape Diameter Function via human evaluation.
modal rendering and 2D-to-3D lifting. In the first phase,
multiview renders of the mesh are individually processed
through Segment Anything to generate 2D masks. These 1. Introduction
masks are then lifted into a mesh part segmentation by as-
sociating masks that refer to the same mesh part across Mesh part segmentation has numerous applications ranging
the multiview renders. We find that applying Segment Any- from texturing and quad-meshing in graphics to object un-
thing to multimodal feature renders of normals and shape derstanding for robotics. Popular mesh part segmentation
diameter scalars achieves better results than using only un- methods are either learning-based or employ more tradi-
textured renders of meshes. By building our method on tional shape analysis methods. For example, a robust shape
top of Segment Anything, we seamlessly inherit any fu- analysis for mesh segmentation is Shape Diameter Func-
ture improvements made to 2D segmentation. We com- tion (ShapeDiam) [23, 24], which computes a scalar value
1
per face that represents its local thickness. These values are • We propose Segment Any Mesh, a novel, zero-shot
clustered to get segmented regions. method that lifts masks produced by applying SAM on
However, learning-based methods are limited by the lack multiview renders into a 3D mesh part segmentation
of diverse segmentation data [2, 31] while traditional meth- • We show performance increases by working with masks
ods do not work well beyond a limited number of mesh fused from multimodal renders, specifically rendered sur-
classes [6]. Recently, large 2D foundation models, such face normals and ShapeDiam scalars.
as Segment Anything (SAM) [11] and SAM2 [22], have • We benchmark our method against the robust, well-
achieved state-of-the-art results for image segmentation. evaluated Shape Diameter Function and demonstrate our
This has revived interest in approaching 3D segmentation method’s effectiveness on existing datasets.
from aggregating masks produced using these 2D models • Since these existing benchmarks do not exhibit much ob-
on multiview renders of the mesh. Previous and contempo- ject diversity or complexity, we also curate a dataset from
rary 2D to 3D lifting methods for mesh segmentation, how- a custom 3D generative model and use it to demonstrate
ever, are restricted to semantic component segmentation as our method’s generalization.
opposed to part segmentation since they require an input
text description for the part to be segmented [2, 31], mak- 2. Related Work
ing them unable to partition duplicate objects (e.g. arms, Zero-Shot 2D Segmentation Advances in 2D image seg-
hands). mentation have been driven by the development of large-
We propose Segment Any Mesh, a zero-shot approach scale foundation models, which leverage extensive datasets
for mesh part segmentation requiring only an input mesh. and substantial compute [7, 15, 17, 21, 29]. The abil-
Our method consists of two phases, multimodal rendering ity to generalize across datasets without finetuning enables
and 2D to 3D lifting. During the first phase, renders from these models as powerful tools for application in down-
different angles individually are fed into SAM to produce stream tasks. Among these models, Segment Anything
multiview masks1 . In addition, we also render the IDs of Model (SAM) [11] and its successor SAM2 [22] are cur-
the triangle faces of the mesh from each respective view. rently state-of-the-art. SAM2 builds on SAM by improv-
During the lifting phase, we use the multiview masks and ing the quality of segmentation and extends segmentation
face IDs to construct a match graph, which is used to as- to dense video, though we do not leverage this capability in
sociate 2D region labels with their corresponding 3D part. our method.
Specifically, we run community detection on this graph to
obtain a rough mesh part segmentation, which we further
Zero-Shot 3D Segmentation In the realm of 3D segmen-
postprocess to obtain the final segmentation.
tation, adjacent domains such as point cloud [5, 16, 18], and
We operate in the untextured setting and find that un-
neural radiance fields (NeRFs) [4, 12, 20, 27, 30], have seen
textured renderings do not result in SAM producing suffi-
significant progress. However, mesh part segmentation is
ciently detailed 2D masks that distinguish mesh parts. As
still at large. Learning-based methods for mesh segmenta-
an alternative, we find that SAM can operate on inputs from
tion [9, 13, 14, 19, 25] are hindered by the limited availabil-
other modalities. Specifically, we show that by feeding both
ity of diverse segmentation datasets. Datasets like CoSeg
normal renders and ShapeDiam scalars renders of a geome-
[28] and Princeton Mesh Segmentation [6] contain only a
try into SAM we can obtain detailed 2D masks with higher
few object classes, restricting the generalizability of models
quality mesh part segmentations compared to relying on a
trained on them.
single modality.
Shape analysis methods have long been used for 3D
We benchmark our method by comparing it against Sha- mesh segmentation. One robust method uses the Shape Di-
peDiam, which performs well on existing mesh part seg- ameter Function through the following steps: first, compute
mentation benchmarks. We show our method is compara- the ShapeDiam scalar for each face of the mesh (the local
ble to or exceeds the performance of ShapeDiam on these thickness of the shape); second, apply a one-dimensional
benchmarks. Due to the limited object diversity of objects Gaussian Mixture Model (GMM) to cluster these values
in these benchmarks, we curate and release a dataset of di- into k groups; third, perform an alpha expansion graph cut
verse object models created through a generative model. We [3] to segment the mesh based on these clusters; and finally,
show through human preference evaluation that the quality split disconnected regions with the same hierarchical label
of segmentations produced by Segment Any Mesh greatly into distinct part labels. k reflects the object part hierar-
exceeds those produced by ShapeDiam. chy rather than the number of object parts [23, 24]. While
To summarize, our contributions are effective in certain scenarios, these traditional methods are
1 We utilize SAM2 for image segmentation, which produces better re- often limited in their ability to handle complex and diverse
sults than SAM. We do not employ video segmentation since the pose object classes, and they require careful parameter tuning to
changes between views are larger than what SAM2 expects. achieve optimal results.
2
Figure 2. Our method applies Segment Anything to renderings of different modalities across views to get masks for each modality.
These masks are then fused per view to get the per view 2D segmentation mask. FaceIDs, which are view consistent, are used to match
segmentations across views.
Contemporary zero-shot 3D segmentation methods [1, 2, SAM’s prediction IOU threshold to 0.5 for all our exper-
8, 31] are based on lifting the outputs of 2D foundation iments, lower than the default 0.8. We find that although
models and perform much better than previous multiview this allows noisy masks to permeate each view, they are fil-
3D segmentation [10, 26] works due to a strong frontend tered out by the 2D to 3D lifting phase, while details are
2D segmenter. However, they rely on input vocabulary and preserved. We also render face IDs, which are used for con-
thus are limited to segmenting one part of an object at a structing the match graph in the next phase. Specifically,
time, effectively performing semantic rather than part seg- let mi,binary denote the binary masks accrued for view i, and
mentation. Furthermore, often parts of a mesh (e.g. CAD the corresponding camera pose pi . Let Fj be the rendering
parts or components of architecture) do not have a suitable function for modality j for J modalities, and → be concate-
corresponding textual description. These approaches, while nation. Then mi,binary an be expressed as
useful, do not fully address the need for automatic, compre-
hensive part segmentation in 3D meshes.
SAM(F1 (M, pi )) → ... → SAM(FJ (M, pi ))
3. Mesh Part Segmentation via 3D Lifting
3.1. Multimodal Rendering
Given more than one modality, we perform mask fusion.
In order to utilize SAM’s powerful visual prior, we trans- We sort all SAM binary masks from different modalities
late geometric information into visual information. Given by area for each view. The masks are fused into a single
a mesh M , we render n views over varying modalities in instance segmentation mask, mi,instance by overlaying from
a regular icosahedral layout around the given mesh. We the largest area binary masks on the bottom to smallest area
then apply SAM to each render and, for each view, fuse binary masks on top (the mask area essentially defines a
masks from different modalities into a single 2D segmenta- rasterization order for the binary masks). After mask over-
tion mask. laying, we remove islands and holes with an area less than
We explore rendering three modalities and their combi- ASAM = 1024 pixels, which we use for all our experiments.
nations as inputs to SAM: untextured renderings, rendered Masks that fall in the background are removed using the
surface normals, and rendered ShapeDiam scalars. We set face ID renders.
3
Figure 3. Overview of the Segment Any Mesh pipeline. In the Multimodal Renderer phase, surface normals and ShapeDiam scalars are
rendered and fed into the SAM, and face IDs are also rendered. In the Lifting 2D to 3D phase, the match graph, which associates 2D labels
-
that refer to the same mesh part, is constructed using the 2D masks and face IDs. Finally, community detection and postprocessing are
applied to get the mesh part segmentation.
3.2. Lifting Segment Anything ωCD = 1, to get node communities. The resulting nodes’
We define our unweighted match graph G = (V, E), where labels in each respective community refer to the same mesh
each node corresponds to a 2D region label in a rendered part segmentation label. We then project the mesh part seg-
view mask while an edge indicates if two nodes potentially mentation labels onto the faces, with the canonical part seg-
refer to the same mesh part. For all pairs of 2D region labels mentation label per face set as the most referred label. Ties
are broken arbitrarily, though we observe they are negligible
(r1 , r2 ) ↑ E across all mi,instance , we calculate the overlap
ratio of their 2D masks, R, to determine whether an edge in quantity.
should be added between their respective nodes. Specifi- 3.3. Mesh Segmentation Refinement
cally, let OF(r1 ) be the number of faces the projections of
r1 and r2 onto the mesh share and F(r) be the number of We first remove holes (unlabeled space) with an area less
faces r’s projection occupies. If than Amesh = pAmesh · Nfaces , where the face count fraction
! " pAmesh is 0.025 in all experiments. We then handle islands by
OF(r1 , r2 ) OF(r1 , r2 ) expanding the frontier by one face each iteration for Ismooth
min , > ωR
F(r1 ) F(r2 ) iterations. We set Ismooth to a large number (64 in all exper-
and the condition OF(r1 , r2 ) > ωC , where ωC is an over- iments) so that the remaining mesh has no holes. We then
lap threshold to filter noise, we connect nodes r1 and r2 . split disconnected regions and, as in ShapeDiam follow up
We set ωC = 32 in all our experiments. We observe that by applying an alpha expansion graph cut [3] for 1 itera-
the optimal ωR for different meshes can vary but a some- tion2 , which we use in all our experiments. We weight the
what decent threshold can be approximated dynamically as alpha expansion graph cut cost term by a factor of ε, which
follows: we discretize the overlap ratio R ↑ [0, 1] over all we vary in our experiments depending on the dataset (see
region pairs into a histogram H with resolution rH = 100 section Implementation).
bins. We set ωR to the bin where the prefix sum of the num-
ber of edge candidates is at least fraction pωR of the total 4. Experiments
number of edge candidates. We vary pωR in our experiments Datasets We first conduct a human evaluation study for
depending on the dataset (see Section Implementation). Segment Any Mesh vs ShapeDiam as well as ablations on
# b
% the modalities used. We curate a diverse set of meshes from
$
ωR = min b : H(i) > pωR Npairs a custom 3D generative model consisting of 75 diverse wa-
i tertight meshes— daily objects, architecture, and artistic
We apply Leiden community detection on the match objects. We release the dataset to the public domain.
graph using igraph’s implementation with resolution pa- 2 For all our experiments, we use a smoothing value of ω = 4 as done
rameter 0. Afterwards, we filter out communities with size in [24]
4
Figure 4. We show comparisons between Segment Any Mesh and ShapeDiam for meshes from our curated dataset.
5
For the human evaluation protocol, we concatenate ren- Model Rank Mean (↓) Rank Std (↓)
dered videos of the segmented meshes color mapped with
the same set of visually distinct colors, from largest region ShapeDiam 1.82 0.27
area to smallest region area, and ask the human evaluator Ours 1.18 0.27
to rank the videos. The order of videos in the concatenation
Table 1. Human evaluation rankings for Segment Any Mesh vs
was randomized across meshes. In addition, the ablations of
ShapeDiam.
the modalities fed to SAM included untextured renderings,
rendered surface normals, rendered ShapeDiam scalars, and
combining rendered surface normals and rendered Shape- Implementation We run all experiments on a single
Diam scalars. We do not benchmark on all combinatorially Nvidia A10 GPU. For our human evaluation study, we re-
possible cases since 1) we do not notice much improvement cruited a community of m = 17 participants (n = 5 trials
using untextured renderings as a modality compared to ren- per mesh). The workforce was asked to rank the segmen-
dered surface normals and ShapeDiam scalars, and 2) the tations from best to worst. We show the interface and eval-
cognitive load of ranking more than 4 concatenated videos uation criteria in the supplementary material. For Prince-
will lead to inaccurate evaluation. ton Mesh Segmentation Evaluation, we use the provided
We separately compare our method against ShapeDiam segmented meshes for ShapeDiam while we implement our
on traditional mesh segmentation benchmark datasets, in own ShapeDiam for use on the other two datasets.
which ShapeDiam achieves decent results and represents We now discuss how we choose our parameters. Larger
the upper ceiling of previous nonlearning approaches. ε means more smoothing and fewer components, which is
We benchmark on the CoSeg dataset and the Prince- desirable for coarse segmentations. Similarly, lower ωR cor-
ton Mesh Segmentation Dataset against human-annotated relates with fewer edges being added to the match graph
ground truth. Unlike our curated dataset, the CoSeg dataset and fewer components. For the human evaluation study and
is composed of 8 classes of objects while the Princeton Princeton Mesh Segmentation Benchmark, we set Segment
Mesh Segmentation has 19, and they are more useful for Any Mesh’s ε = 6 and ωR = 0.125 for human evalua-
measuring segmentation consistency for meshes belonging tion and ωr = 0.05 for CoSeg. For ShapeDiam, we ob-
same class as opposed to generalization. serve the best performance with k = 5 GMM clusters and
ε = 15 for human evaluation and Princeton Mesh Seg-
Metrics For the Human Evaluation Study, each mesh had mentation Benchmark, while for CoSeg, k = 3 (see CoSeg
n = 5 trials, and we computed the mean/std rank (starting dataset visualizations on its webpage) and ε = 15.
at 1) across all trials and meshes for each method.
For CoSeg and Princeton Mesh Segmentation, we use Human Evaluation Study Table 1 shows the mean and
the 7 metrics introduced in Princeton Mesh Segmentation. standard deviation of human rankings for Segment Any
We briefly introduce them below. Mesh vs ShapeDiam. We can see that Segment Any Mesh
1. Cut Discrepancy (↓) sums the distances from points consistently ranks above ShapeDiam. Qualitative results for
along the cuts in the generated segmentation to the near- Segment Any Mesh on the generated dataset are shown in
est cuts in the ground truth segmentation, and vice versa. Figure 1 while comparisons between Segment Any Mesh
This provides an intuitive measure of the separation be- and ShapeDiam are shown in Figure 4. Table 2 shows the
tween the cuts. mean and standard deviation for rankings of ablations on
2. Hamming Distance (↓) measures the overall region- input modalities to SAM, we see that combining surface
based difference between two segmentations. normal and ShapeDiam scalar value renderings are evalu-
3. Hamming Distance (Rf) (↓) assuming second segmenta- ated as the best. Figure 6 shows a comparison between the
tion is ground truth, measures false positive labels mesh part segmentations produced by using different input
4. Hamming Distance (Rm) (↓) assuming second segmen- modalities.
tation is ground truth, measures true negative labels
5. Rand Index measures the likelihood that a pair of faces Traditional Benchmarks We now show Segment Any
are either grouped within the same segment or separated Mesh performs as well as ShapeDiam on traditional bench-
into different segments across two segmentations. marks, which ShapeDiam is overfitted to. Note that metrics
6. Local Consistency Error (↓) accounts for nested, hierar- should only be taken as an estimate of segmentation qual-
chical similarities in segmentations ity since mesh part segmentation can be quite subjective.
7. Global Consistency Error (↓) accounts for nested, hier- Table 3 shows we match the performance of ShapeDiam,
archical similarities in segmentations, forcing all local and Figure 5 shows example outputs of our method on the
refinements to be in the same direction CoSeg dataset. Table 4 shows we also exceed the perfor-
For exact formulations, please reference [6]. mance of ShapeDiam, and Figure 7 shows example outputs
6
Figure 5. Qualitative examples of Segment Any Mesh vs ShapeDiam outputs on the CoSeg dataset. We provide additional comparisons
with ShapeDiam in the supplemental.
Rendering Modality Rank Mean (↓) Rank Std (↓) Metric ShapeDiam Ours
Ours (only norm) 2.57 0.58
Ours (only shape diameter function) 2.52 0.75 Cut Discrepancy 0.39 0.37
Ours (untextured) 2.59 0.70 Hamming Distance 0.16 0.18
Ours 2.33 0.59 Hamming Distance (Rf) 0.10 0.08
Hamming Distance (Rm) 0.23 0.28
Table 2. Human evaluation rankings for different modality com- Rand Index 0.21 0.22
binations used as input to Segment Anything. We only perform 4 Local Consistency Error 0.06 0.05
ablations as any number beyond that results in a ranking task that
Global Consistency Error 0.09 0.08
is too difficult.
Table 3. Quantitative results for ShapeDiam vs Segment Any
Mesh on the CoSeg dataset for automatic determination of the
number of segmentation regions.
Figure 6. Comparison of mesh part segmentations produced by
mark provides mesh segmentations for ShapeDiam that use
varying input modalities to Segment Anything. From left to right,
rendered surface normals only, rendered ShapeDiam scalars only, the mode number of human segmentation regions as the tar-
untextured rendering, and rendered surface normals combined get number of GMM clusters, k. We benchmark our method
with rendered ShapeDiam scalars (best). against these mesh segmentations to reflect the scenario
when the target number of segmentation regions is roughly
known. Specifically, we search ε—the weight of the cost
of our method on the Princeton Mesh Segmentation bench- for alpha graph expansion. We select the smallest ε that
mark. results in the number of segments produced by our method
Additionally, the Princeton Mesh Segmentation Bench- within a predefined error margin of the mode value. We set
7
Figure 7. Qualitative examples of Segment Any Mesh vs ShapeDiam outputs on the Princeton Mesh Segmentation Benchmark dataset. We
provide additional comparisons with ShapeDiam in the supplemental.
Metric ShapeDiam Ours Metric ShapeDiam Ours
Cut Discrepancy 0.34 0.31 Cut Discrepancy 0.28 0.30
Hamming Distance 0.21 0.17 Hamming Distance 0.16 0.17
Hamming Distance (Rf) 0.17 0.17 Hamming Distance (Rf) 0.14 0.14
Hamming Distance (Rm) 0.26 0.17 Hamming Distance (Rm) 0.18 0.20
Rand Index 0.22 0.21 Rand Index 0.17 0.18
Local Consistency Error 0.10 0.07 Local Consistency Error 0.08 0.08
Global Consistency Error 0.16 0.12 Global Consistency Error 0.13 0.13
Table 4. Quantitative results for Princeton Mesh Segmentation Table 5. Quantitative results for Princeton Mesh Segmentation
Benchmark for ShapeDiam vs Segment Any Mesh for automatic Benchmark for ShapeDiam vs Segment Any Mesh given a ref-
determination of the number of segmentation regions. erence number of segmentation regions determined by labelers.
ωR = 0.35 while ε is variable in the range [1, 15]. Table 5 tors. Lifting tackles the problem from a 2D view perspec-
shows our method matches the performance of ShapeDiam. tive, aligning with humans when analyzing object parts and
affordances.
5. Conclusion In the future, an interface can be developed refining the
segmentations, which can further improve quality for tar-
Segment Any Mesh performs well on existing benchmarks geted tasks such as object understanding, texturing, and re-
and generalizes for diverse classes of meshes. This is due to topology. We can also leverage this human-in-the-loop pro-
our method being based on lifting rather than learning from cess to curate and distill a labeled segmentation dataset into
limited diversity segmentation data or local shape descrip- a MeshCNN model [9].
8
References [17] Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao
Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang,
[1] Ahmed Abdelreheem, Abdelrahman Eldesokey, Maks Ovs- Hang Su, Jun Zhu, and Lei Zhang. Grounding dino: Marry-
janikov, and Peter Wonka. Zero-shot 3d shape correspon- ing dino with grounded pre-training for open-set object de-
dence, 2023. 3 tection, 2024. 2
[2] Ahmed Abdelreheem, Ivan Skorokhodov, Maks Ovsjanikov, [18] Björn Michele, Alexandre Boulch, Gilles Puy, Maxime
and Peter Wonka. Satr: Zero-shot semantic segmentation of Bucher, and Renaud Marlet. Generative zero-shot learning
3d shapes, 2023. 2, 3 for semantic segmentation of 3d point clouds, 2023. 2
[3] Y. Boykov, O. Veksler, and R. Zabih. Fast approximate en- [19] Francesco Milano, Antonio Loquercio, Antoni Rosinol, Da-
ergy minimization via graph cuts. IEEE Transactions on Pat- vide Scaramuzza, and Luca Carlone. Primal-dual mesh con-
tern Analysis and Machine Intelligence, 23(11):1222–1239, volutional neural networks, 2020. 2
2001. 2, 4 [20] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik,
[4] Jiazhong Cen, Jiemin Fang, Zanwei Zhou, Chen Yang, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf:
Lingxi Xie, Xiaopeng Zhang, Wei Shen, and Qi Tian. Seg- Representing scenes as neural radiance fields for view syn-
ment anything in 3d with radiance fields, 2024. 2 thesis, 2020. 2
[5] Runnan Chen, Xinge Zhu, Nenglun Chen, Wei Li, Yuexin [21] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy
Ma, Ruigang Yang, and Wenping Wang. Zero-shot point Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez,
cloud segmentation by transferring geometric primitives, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mah-
2023. 2 moud Assran, Nicolas Ballas, Wojciech Galuba, Russell
[6] Xiaobai Chen, Aleksey Golovinskiy, and Thomas Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael
Funkhouser. A benchmark for 3D mesh segmentation. Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Je-
ACM Transactions on Graphics (Proc. SIGGRAPH), 28(3), gou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr
2009. 2, 6 Bojanowski. Dinov2: Learning robust visual features with-
[7] Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexan- out supervision, 2024. 2
der Kirillov, and Rohit Girdhar. Masked-attention mask [22] Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang
transformer for universal image segmentation, 2022. 2 Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman
[8] Dale Decatur, Itai Lang, and Rana Hanocka. 3d highlighter: Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt-
Localizing regions on 3d shapes via text descriptions, 2022. ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-
3 Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feicht-
enhofer. Sam 2: Segment anything in images and videos,
[9] Rana Hanocka, Amir Hertz, Noa Fish, Raja Giryes, Shachar
2024. 2
Fleishman, and Daniel Cohen-Or. Meshcnn: a network with
an edge. ACM Transactions on Graphics, 38(4):1–12, 2019. [23] Bruno Roy. Neural shape diameter function for efficient
2, 8 mesh segmentation. In ACM SIGGRAPH 2023 Posters.
ACM, 2023. 1, 2
[10] Evangelos Kalogerakis, Melinos Averkiou, Subhransu Maji,
[24] Lior Shapira, Ariel Shamir, and Daniel Cohen-Or. Consistent
and Siddhartha Chaudhuri. 3d shape segmentation with pro-
mesh partitioning and skeletonisation using the shape diam-
jective convolutional networks, 2017. 3
eter function. The Visual Computer, 24:249–259, 2008. 1, 2,
[11] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, 4
Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White-
[25] Dmitriy Smirnov and Justin Solomon. Hodgenet: Learning
head, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and
spectral geometry on triangle meshes, 2021. 2
Ross Girshick. Segment anything, 2023. 2
[26] Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik
[12] Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitz- Learned-Miller. Multi-view convolutional neural networks
mann. Decomposing nerf for editing via feature field dis- for 3d shape recognition, 2015. 3
tillation, 2022. 2 [27] George Tang, Krishna Murthy Jatavallabhula, and Antonio
[13] Juil Koo, Ian Huang, Panos Achlioptas, Leonidas Guibas, Torralba. Efficient 3d instance mapping and localization with
and Minhyuk Sung. Partglot: Learning shape part segmenta- neural fields, 2024. 2
tion from language reference games, 2022. 2 [28] Yunhai Wang, Shmulik Asafi, Oliver van Kaick, Hao Zhang,
[14] Xiao-Juan Li, Jie Yang, and Fang-Lue Zhang. Laplacian Daniel Cohen-Or, and Baoquan Chen. Active co-analysis of
mesh transformer: Dual attention and topology aware net- a set of shapes. ACM Trans. Graph., 31(6), 2012. 2
work for 3d mesh classification and segmentation. In Com- [29] Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li,
puter Vision – ECCV 2022, pages 541–560, Cham, 2022. Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo,
Springer Nature Switzerland. 2 Yaqian Li, Shilong Liu, Yandong Guo, and Lei Zhang. Rec-
[15] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. ognize anything: A strong image tagging model, 2023. 2
Visual instruction tuning, 2023. 2 [30] Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, and An-
[16] Minghua Liu, Yinhao Zhu, Hong Cai, Shizhong Han, Zhan drew J. Davison. In-place scene labelling and understanding
Ling, Fatih Porikli, and Hao Su. Partslip: Low-shot part seg- with implicit scene representation, 2021. 2
mentation for 3d point clouds via pretrained image-language [31] Ziming Zhong, Yanxu Xu, Jing Li, Jiale Xu, Zhengxin Li,
models, 2023. 2 Chaohui Yu, and Shenghua Gao. Meshsegmenter: Zero-shot
9
mesh semantic segmentation via texture synthesis, 2024. 2,
3
10
Supplementary Material
A. Additional Details for Human Evaluation Study
We use AWS Sagemaker as an evaluation interface. For the Segment Any Mesh versus shape diameter function comparison,
we render a video of the meshes side by side, randomized, and ask the participants to rank which segmentations they prefer
in order of (a), (b), (c), (d) where the letter corresponds to the respective segmentation index in the video. Below are the
interfaces participants interacted with
Figure S1. Interace for Segment Any Mesh vs Shape Diameter Function experiment.
A1
Figure S2. Interface for Segment Any Mesh vs Shape Diameter Function experiment. AWS Sagemaker does not support text input, so we
resort to 24 labels corresponding to the ranking permutations.
We also include examples of frames from videos shown to the participants
Figure S3. Frames from a video shown to participants for Segment Any Mesh vs Shape Diameter Function preference experiment.
B. Additional Visualizations
We also provide Segment Any Mesh vs Shape Diameter Function comparisons for CoSeg and Princeton Mesh Segmentation
Benchmark.
A2
Figure S4. Frames from a video shown to participants for SAM input modality ablation experiment.
A3
Figure S5. Segment Any Mesh vs Shape Diameter Function on the CoSeg dataset.
A4
Figure S6. Additional visualizations for Segment Any Mesh vs Shape Diameter Function on the CoSeg dataset.
A5
Figure S7. Segment Any Mesh vs Shape Diameter Function on the Princeton
A6 Mesh Segmentation Benchmark for automatic determination
of the number of segmentation regions i.e. not given the reference number of segmentation regions k.
Figure S8. Additional visualizations for Segment Any Mesh vs Shape Diameter Function on the Princeton Mesh Segmentation Benchmark
for automatic determination of the number of segmentation regions.
A7