TacDyn-WAM: Learning Implicit Tactile Dynamics
in a Heterogeneous Visuo-Tactile World Action Model

Enyi Wang1, Mingxin Wang1,4, Quan Shi1, Hetian Guo1, Hongyu Wang1, Xi Wang1, Bin Qian1,4, Yupeng Zheng2, Wenxuan Song3, Houde Liu4, Yong Xu1, Cheng Chi5, Wenchao Ding6,7, Yilun Chen7, Yan Wang1,†
1Institute for AI Industry Research (AIR), Tsinghua University 2Institute of Automation, Chinese Academy of Sciences 3The Hong Kong University of Science and Technology (Guangzhou) 4Tsinghua University 5School of Information, Renmin University of China 6Fudan University 7TARS Robotics

† Corresponding author

Abstract

World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising. Such prediction can become unreliable under deployment drift: small changes in contact position or force may substantially alter tactile pixels even when the underlying contact evolution remains predictable. We introduce TacDyn-WAM, a heterogeneous visuo-tactile world action model that predicts implicit tactile dynamics rather than reconstructing future tactile observations. It learns TacRep, a dynamics-aware tactile target space trained through masked spatio-temporal prediction on tactile clips and regularized by relational structure distillation. A visual expert and an Implicit Tactile Dynamics Expert predict future visual and tactile representations in separate target spaces while interacting through joint attention; the tactile expert predicts future representations and their changes at multiple horizons in a single forward pass, and a read-only tactile memory supplies the current tactile state. On UniVTAC, TacDyn-WAM achieves an average success rate of 81.5% using only the provided demonstrations, reaching state-of-the-art-level performance and remaining competitive with models pretrained on large-scale visuo-tactile trajectories. Ablations confirm the benefits of both tactile pathways and TacRep over pixel-reconstruction and static alternatives. On five real-robot tasks, TacDyn-WAM reaches 71.0% average success, and modest-scale pretraining raises it to 85.0%, further validating our method.

Model Architecture

Predicting dynamic trends in contact evolution rather than future tactile pixels.

TacDyn-WAM architecture: visual and tactile experts predict separate futures through joint attention, with tactile memory supporting the action expert.
TacDyn-WAM overview. The Visual Latent Generation Expert predicts future visual latents, while the Implicit Tactile Dynamics Expert (ITDE) predicts future TacRep representations and their changes at multiple horizons in a single forward pass. The two experts use separate parameters and target spaces, coupled through joint attention. A read-only Tactile Understanding Memory supplies the current tactile state as keys and values, enabling the Action Expert to combine present contacts with predicted visual and tactile futures.

Learning TacRep

TacRep learning through masked tactile dynamics prediction and relational structure distillation.
Tactile Dynamics Prediction (TDP) learns contact evolution by predicting masked spatio-temporal features of four-frame tactile clips against an exponential moving average (EMA) target encoder. Relational Structure Distillation (RSD) regularizes local patch relations using a frozen DINOv2 teacher, helping improve robustness to out-of-distribution (OOD) tactile images.

Implicit Tactile Dynamics Expert

The Implicit Tactile Dynamics Expert predicts future representations and their changes at multiple horizons.
Future tactile representations and their changes are predicted at multiple horizons in a single forward pass, without iterative denoising.

Staged Training

Attention masks and trainable modules in Stage 2: Tactile World Grounding, Stage 3: Tactile–Action Alignment, and Stage 4: Joint Training.
Progressive integration of tactile modules. After Stage 1 learns TacRep, Stage 2 grounds the tactile world model, Stage 3 aligns tactile information with actions, and Stage 4 jointly trains the model. The diagram shows attention connections in Stages 2–4: bold outlines mark newly unlocked connections, flames indicate trainable modules, and snowflakes indicate frozen modules.

Simulation Experiments

Eight UniVTAC tasks, 100 rollouts per task. TacDyn-WAM is trained on the provided demonstrations only.

SOTA-level performance

81.5%average success rate

+25.7 pointsover InternVLA-A1

7 of 8 tasksat least 72% success

Success rates (%) on UniVTAC
MethodInsert
Hole
Insert
Tube
Lift
Can
Pull-out
Key
Put
Bottle
Lift
Bottle
Grasp
Classify
Insert
HDMI
Avg.
Vision-only policies
π₀.₅25746353410049841.4
StarVLA-α526965518832682456.1
InternVLA-A1705658673758881255.8
Xiaomi-Robotics-0969813801221456954.3
GigaWorld-Policy129032213820016.5
LingBot-VA429605800173831.4
Fast-WAM66980733521721948.0
Visuo-tactile policies
ACT+UniVTAC245629463171992848.0
VITaL25348473272100640.5
RDP237512184184——42.2
TacForcing697963484390——65.3
Tactile-WAM2085102045555030.0
Policies with large-scale visuo-tactile pretraining · shown for context
FTP-164796548479799462.9
N₀-VTLA9599889960991002583.1
N₀-TWAM (20% data)789680254651955265.4
N₀-TWAM999893798758946884.5
Our method · provided demonstrations only
TacDyn-WAM (ours)919787977297991281.5

Gray rows use large-scale visuo-tactile trajectory pretraining and are shown for context, as in the paper. “—” denotes an unreported result; averages for incomplete rows use their reported tasks.

Real-World Experiments

We evaluate five contact-rich tasks—Stack Cups, Remove Plug, Insert Plug, Unscrew Cup Lid, and Wipe Whiteboard—on a Franka Research 3 arm with Xense tactile sensors on both fingers. Each task uses 60 demonstrations and 20 evaluation trials per policy.

Real-world success rates (%)
MethodStack
Cups
Remove
Plug
Insert
Plug
Unscrew
Cup Lid
Wipe
Whiteboard
Avg.
LingBot-VA401045155032.0
InternVLA-A1552045107040.0
FTP-1654565208055.0
TacDyn-WAM806080409571.0
TacDyn-WAM (pretrained)9585905510085.0

Modest-scale tactile pretraining. Using a 6,000-trajectory OmniViTac subset, with our 300 task demonstrations additionally used in Stages 1–2, pretraining improves all five tasks and raises average success from 71.0% to 85.0%. All baselines are fine-tuned per task on the same demonstrations.

Real-Robot Demo