0% found this document useful (0 votes)
4 views3 pages

Roto 2.0:: The Robot Tactile Olympiad

The document introduces roto 2.0, a GPU-parallelized benchmark for tactile-based reinforcement learning (RL) that standardizes testing across four robotic morphologies. It emphasizes 'blind' manipulation using only proprioception and tactile sensing, achieving significant performance improvements in tasks like Baoding ball rotation. By open-sourcing the benchmark, the authors aim to lower entry barriers for researchers and encourage focus on fundamental algorithmic challenges in tactile-based RL.

Uploaded by

cy133177
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views3 pages

Roto 2.0:: The Robot Tactile Olympiad

The document introduces roto 2.0, a GPU-parallelized benchmark for tactile-based reinforcement learning (RL) that standardizes testing across four robotic morphologies. It emphasizes 'blind' manipulation using only proprioception and tactile sensing, achieving significant performance improvements in tasks like Baoding ball rotation. By open-sourcing the benchmark, the authors aim to lower entry barriers for researchers and encourage focus on fundamental algorithmic challenges in tactile-based RL.

Uploaded by

cy133177
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

roto 2.

0: The Robot Tactile Olympiad


Elle Miller1 , Jayaram Reddy2 , Ayush Deshmukh1 , Trevor McInroe1 , David Abel1 , Oisin Mac Aodha1 , Sethu Vijayakumar1
arXiv:2605.21429v1 [[Link]] 20 May 2026

Fig. 1. The roto 2.0 Benchmark Suite: A standardised RL framework across four distinct dexterous morphologies (from L to R): ORCA Hand, Shadow
Lite, Allegro Hand, and the Shadow Dexterous Hand. The suite facilitates “blind" tactile manipulation tasks, such as Baoding ball rotation and ball bouncing.

Abstract—Tactile-based reinforcement learning (RL) is cur- arm and single-point contact tasks, while VTDexManip [12]
rently hindered by fragmented research and a focus on over- focuses on pre-training policies with pre-collected datasets. To
saturated orientation tasks. We introduce v2 of the Robot Tactile fill this gap, we introduce v2 of the Robot Tactile Olympiad
Olympiad (roto 2.0), a GPU-parallelised benchmark designed
to standardise tactile-based RL across four distinct robotic (roto 2.0), an RL benchmark built upon GPU-parallelised
morphologies (16-DOF to 24-DOF). Unlike prior benchmarks, Isaac Lab [13]. Originally introduced in [14] for the Shadow
roto focuses on end-to-end “blind” manipulation, utilising only Hand, we expand the suite here to include four distinct
proprioception and tactile sensing without state information or dexterous morphologies: the anthropomorphic Shadow Dex-
distillation. We demonstrate a significant performance leap, with terous Hand (24-DOF), Shadow Dexterous Hand Lite (16-
our blind agents achieving 13 Baoding ball rotations in 10
seconds, an order of magnitude faster than current state-of- DOF), Allegro Hand (16-DOF) and ORCA Hand (17-DOF).
the-art speeds. By open-sourcing our environments and robustly Crucially, we demonstrate that with our RL training pipeline,
tuned baselines, we reduce the barrier to entry and enable “blind” policies (utilising only proprioception and tactile data)
researchers to prioritise fundamental algorithmic challenges over can master sophisticated manipulation without the need for the
tedious RL tuning. Website: [Link] teacher-student distillation or explicit pose estimators common
I. I NTRODUCTION in prior work [3], [15], [16]. While the current state-of-
Real-world robotic manipulation requires the ability to the-art for Baoding ball rotation achieves a maximum of 3
interact robustly in unstructured environments where visual rotations in 10 seconds [17], our blind agents exploit the raw
line-of-sight is frequently obstructed. To achieve this, robots potential of tactile-proprioceptive loops to achieve significantly
must learn to “feel.” While Reinforcement Learning (RL) has higher throughput of up to 15 rotations. As demonstrated
revolutionised locomotion across complex terrains [1], tactile- in [14], integrating self-supervised forward dynamics to aid
based manipulation lags significantly behind. Progress is cur- representation learning can further push this boundary to 25
rently hindered by a fragmented landscape: most labs work rotations per 10 seconds, nearly closing the gap between blind
in isolation using unique combinations of sensors and robots, policies and state-based agents. From these results, we argue
making cross-validation difficult. Furthermore, the community that establishing robust haptic foundations is a prerequisite to
has largely over-saturated the task of in-hand orientation, e.g. integrating visual modalities, rather than a secondary addition
[2], [3], [4], [5], [6], [7], [8], [9]. While impressive, a single to vision-centric systems. By open-sourcing roto, we aim to
task fails to capture the broader spectrum of challenges that reduce the barrier-to-entry for researchers interested in tactile-
tactile observations pose, leaving the true utility of tactile based RL and help focus community efforts on high-impact
feedback an open debate. The difficulty of tactile-based RL research directions instead of RL tuning. Contributions:
stems from a trifecta of complexity: manipulation is inherently • The roto benchmark: A set of tactile-based environ-
hard [10], on-policy RL is notoriously difficult to tune, and ments with integrated hyperparameter optimisation and
thus effectively combining the two with sparse and discon- robustly tuned baselines.
tinuous tactile observations is a significant undertaking. This • Morphological diversity: A cross-platform evaluation
difficulty is exacerbated by a lack of standardised tactile- of four robot hands to study how hardware complexity
based RL benchmark environments across complex tasks and influences tactile policy convergence.
morphologies. Tactile-Gym 2.0 [11] is limited to a 4-DOF • Performance breakthroughs: We provide “blind” poli-
1 Universityof Edinburgh, UK. Email: [Link]@[Link] cies that achieve state-of-the-art speeds in complex ma-
2 National University Singapore, Singapore nipulation, providing a new ceiling for tactile intelligence.
II. M ETHODOLOGY Baoding
800
Shadow(states)
RL. We use a customised implementation of Proximal Pol- 700 ORCA(states)
icy Optimisation (PPO) [18] from SKRL [19] to incorporate Allegro(states)
600 ShadowLite(states)

Mean evaluation return


observation stacking, self-supervision, separated environments Shadow(blind)
for continuous evaluation, and various training tricks [20]. 500 ORCA(blind)
Allegro(blind)
We use 8, 092 parallelised environments for training and 100 400 ShadowLite(blind)
for evaluation. For each combination of robot (n = 3), task 300
(n = 2), and observation setting (blind vs. state-based), we
200
perform a hyperparameter sweep across seven PPO parameters
using 40 trials (8 warm-up runs) to ensure robust baselines. 100
• Bounce: The agent must bounce a ball as many times 0
0 25 50 75 100 125 150 175
as possible in 10 seconds (600 timesteps). A bounce is Timesteps (M)
defined as a contact event after a period of at least 5 Bounce
timesteps (∼ 83ms) without contact.
• Baoding: The agent must rotate two balls (55g) around 800

Mean evaluation return


each other in-hand as many times as possible within 10
seconds (600 timesteps). We use 1.5 inch diameter for 600
Shadow Hand and ORCA, 2 inches for Allegro and 1.2
inches for Shadow Lite. 400
MDP. The blind agents receive a history length k = 4
of proprioceptive and binary tactile observations, and are 200
joint-position controlled, see Table I for details. We define
the each task-relevant robot link as a tactile sensor. The 0
0 25 50 75 100 125 150 175
state-based agents additionally receive the object position(s) Timesteps (M)
and linear velocity. For Bounce, the agent is rewarded with
rbounce = 10 for every successful bounce. For Baoding, Fig. 2. Mean evaluation returns across 5 seeds for state-based and blind
agents in the Baoding and Bounce tasks.
we specify two static target positions and define the reward
as rdist1 + rdist2 + rrotation . The dense distance rewards starfish-like pose. Our blind agents demonstrate high sample
rdist1 , rdist2 encourage the balls to the targets. When the efficiency, approaching 80 bounces by 200M steps. Despite the
centers of both balls are within 1.0 cm of the targets, the targets vast differences in hardware, we find that performance trends
switch and the agent receives a bonus reward rrotation = 10. remain consistent across the full-hand morphologies, excepting
The episode terminates if any object is out of reach or the the Shadow Lite. The Baoding task reveals a more significant
maximum episode length T = 600 is reached. The physics performance gap between state-based and blind agents. State-
simulation runs at 240 Hz, the control policy at 60 Hz. based agents achieve a throughput of up to 35 rotations in
10 seconds. Blind agents exhibit much lower performance
TABLE I with high variance; while a top-performing Shadow Hand
O BSERVATION AND ACTION SPACES
seed achieved 13 rotations, others failed to converge. This
Type Description Shadow Shadow Lite Allegro ORCA stochasticity highlights a core challenge in tactile-based RL:
Tactile obs. binary contacts 17 14 20 17
Proprio. obs. joint positions 20 16 16 17
efficient feature extraction [14].
joint velocities 20 16 16 17
joint command error 20 13 16 17 IV. D ISCUSSION & C ONCLUSION
last action 20 13 16 17
Total obs. single timestep 97 72 84 85 We introduce roto 2.0, a multi-morphology benchmark
k = 4 timesteps 388 288 336 340
Actions joint positions 20 13 10 17 for blind dexterous manipulation. While our high-speed sim-
ulated policies represent a “performance ceiling” that exceeds
current real-world hardware limits, they provide a benchmark
III. E XPERIMENTAL RESULTS for developing the next generation of robust RL pipelines.
We evaluate the learning efficiency and asymptotic per- Our results indicate that while blind policies can approach
formance of four distinct morphologies for state-based and privileged performance in simple tasks like bouncing, com-
blind agents. The mean evaluation returns are summarised plex or multi-object manipulation (Baoding) remains an open
in Figure 2; we refer the reader to the project page for the challenge. We identify high-priority research directions for the
policy videos. In the Bounce environment, state-based agents community: beyond sparse binary contacts to richer forms
approach the theoretical maximum return (1, 000 reward, cor- of tactile information, ML methodologies with inductive bi-
responding to 100 successful bounces). While the hands con- ases for tactile data, and expanding tasks beyond hands e.g.
verge to similar success rates, they exhibit hardware-specific whole-body humanoid manipulation [21]. We are actively
strategies, e.g. the ORCA hand adopts an “outstretched” investigating the sim-to-real transferability of these policies
and welcome community contributions to the roto suite to [17] Y. Yuan, H. Che, Y. Qin, B. Huang, Z.-H. Yin, K.-W. Lee, Y. Wu, S.-
accelerate the arrival of the “locomotion moment" for robotic C. Lim, and X. Wang, “Robot synesthesia: In-hand manipulation with
visuotactile sensing,” in ICRA, 2024.
touch. [18] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Prox-
imal policy optimization algorithms,” arXiv preprint arXiv:1707.06347,
2017.
R EFERENCES [19] A. Serrano-Muñoz, D. Chrysostomou, S. Bøgh, and N. Arana-
Arexolaleiba, “skrl: Modular and flexible library for reinforcement
[1] N. Rudin, J. He, J. Aurand, and M. Hutter, “Parkour in the wild: learning,” Journal of Machine Learning Research, vol. 24, no. 254, pp.
Learning a general and extensible agile locomotion policy using 1–9, 2023. [Online]. Available: [Link]
multi-expert distillation and rl fine-tuning,” 2025. [Online]. Available: [20] E. Miller, “The art of robot reinforcement learning,”
[Link] [Link], 2026. [Online]. Available: [Link]
[2] O. M. Andrychowicz, B. Baker, M. Chociej, R. Józefowicz, B. McGrew, [Link]/p/the-art-of-robot-reinforcement-learning
J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, [21] C. Sferrazza, D.-M. Huang, X. Lin, Y. Lee, and P. Abbeel, “Humanoid-
S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba, bench: Simulated humanoid benchmark for whole-body locomotion and
“Learning dexterous in-hand manipulation,” The International Journal manipulation,” 2024.
of Robotics Research, vol. 39, no. 1, Jan. 2020. [Online]. Available:
[Link]
[3] L. Sievers, J. Pitz, and B. Bäuml, “Learning Purely Tactile In-Hand
Manipulation with a Torque-Controlled Hand,” in 2022 International
Conference on Robotics and Automation (ICRA), May 2022, pp. 2745–
2751. [Online]. Available: [Link]
[4] T. Chen, J. Xu, and P. Agrawal, “A system for general in-hand object
re-orientation,” Conference on Robot Learning, 2021.
[5] H. Qi, A. Kumar, R. Calandra, Y. Ma, and J. Malik, “In-Hand Object
Rotation via Rapid Motor Adaptation,” in Conference on Robot Learning
(CoRL), 2022.
[6] H. Qi, B. Yi, S. Suresh, M. Lambeta, Y. Ma, R. Calandra, and
J. Malik, “General In-hand Object Rotation with Vision and Touch,”
in Proceedings of The 7th Conference on Robot Learning. PMLR,
Dec. 2023, pp. 2549–2564, iSSN: 2640-3498. [Online]. Available:
[Link]
[7] L. Röstel, J. Pitz, L. Sievers, and B. Bäuml, “Estimator-
Coupled Reinforcement Learning for Robust Purely Tactile In-
Hand Manipulation,” in 2023 IEEE-RAS 22nd International
Conference on Humanoid Robots (Humanoids). Austin, TX,
USA: IEEE, Dec. 2023, pp. 1–8. [Online]. Available:
[Link]
[8] M. Yang, C. Lu, A. Church, Y. Lin, C. Ford, H. Li, E. Psomopoulou,
D. A. W. Barton, and N. F. Lepora, “Anyrotate: Gravity-invariant in-
hand object rotation with sim-to-real touch,” in Conference on Robot
Learning (CoRL), 2024.
[9] S. Suresh, H. Qi, T. Wu, T. Fan, L. Pineda, M. Lambeta, J. Malik,
M. Kalakrishnan, R. Calandra, M. Kaess, J. Ortiz, and M. Mukadam,
“Neural feels with neural fields: Visuo-tactile perception for in-hand
manipulation,” Science Robotics, p. adl0628, 2024.
[10] M. T. Mason, “Toward Robotic Manipulation,” Annual Review of Con-
trol, Robotics, and Autonomous Systems, no. 1, 2018.
[11] Y. Lin, J. Lloyd, A. Church, and N. F. Lepora, “Tactile gym 2.0: Sim-to-
real deep reinforcement learning for comparing low-cost high-resolution
robot touch,” IEEE Robotics and Automation Letters, vol. 7, no. 4, pp.
10 754–10 761, 2022.
[12] Q. Liu, Y. Cui, Z. Sun, G. Li, J. Chen, and Q. Ye, “Vtdexmanip:
A dataset and benchmark for visual-tactile pretraining and dexterous
manipulation with reinforcement learning,” in The Thirteenth
International Conference on Learning Representations, 2025. [Online].
Available: [Link]
[13] M. Mittal, C. Yu, Q. Yu, J. Liu, N. Rudin, D. Hoeller, J. L. Yuan,
R. Singh, Y. Guo, H. Mazhar, A. Mandlekar, B. Babich, G. State,
M. Hutter, and A. Garg, “Orbit: A unified simulation framework for
interactive robot learning environments,” IEEE Robotics and Automation
Letters, vol. 8, no. 6, pp. 3740–3747, 2023.
[14] E. Miller, T. McInroe, D. Abel, O. Mac Aodha, and S. Vijayakumar,
“Enhancing tactile-based reinforcement learning for robotic control,” in
NeurIPS, 2025.
[15] Z.-H. Yin, B. Huang, Y. Qin, Q. Chen, and X. Wang, “Rotating
without Seeing: Towards In-hand Dexterity through Touch,” Mar. 2023,
arXiv:2303.10880 [cs]. [Online]. Available: [Link]
10880
[16] L. Yang, B. Huang, Q. Li, Y.-Y. Tsai, W. W. Lee, C. Song, and J. Pan,
“TacGNN: Learning Tactile-Based In-Hand Manipulation With a Blind
Robot Using Hierarchical Graph Neural Network,” IEEE Robotics and
Automation Letters, vol. 8, no. 6, pp. 3605–3612, Jun. 2023. [Online].
Available: [Link]

You might also like