Taming VLAs

under Robot Execution Errors: Self-Compensation and Stress Testing
Scroll
arXiv 2026
Hyunwoo KangPOSTECH
Suha KwakPOSTECH

* Equal contribution

Robots don’t move exactly as commanded.
Let the VLA learn from that gap, while it runs.

Self-compensation at deployment.

Online LoRA updates from the command–execution residual. No rewards, no labels.

+30pp on real robots.

On a new and a 1-year-old Piper arm, for both π0 and π0.5.

Generalizes to unseen objects.

16% → 64% success on bowls and plates absent from fine-tuning.

RoboStress benchmark.

Four joint-level error models, seven deployment scenarios in simulation.

01

Abstract

Vision-language-action (VLA) policies often fail when a robot's executed motion deviates from their commanded action. Such execution errors arise from the robot's mechanics and operating conditions, such as wear and payload changes. We propose self-compensating VLA, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands. Without task rewards or labels, it updates the policy online using the residual between the action commanded by a VLA and the motion executed by the robot. To stress-test VLA robustness across execution conditions that are impractical to cover with physical robots alone, we introduce RoboStress, a controlled simulation benchmark. It combines established joint-level models of friction, backlash, compliance, and gravity-compensation error into seven deployment scenarios whose execution errors depend on the robot's state and motion history. On RoboStress, self-compensating VLA achieves higher average task success than both the base policies and methods that build in robustness during training. On two physical robot arms with different usage histories, it raises the average task success rate by more than 30 percentage points on each arm, and the gains extend to objects not seen in the task demonstrations.

02

Same robot, same VLA

A year-old Piper arm (Robot B). Left: the base policy. Right: the same policy with self-compensation.

“Pick up the gray bowl on the plastic cabinet and place it on the plate”π0

Base
Ours

“Open the top drawer of the cabinet”π0.5

Base
Ours
03

Learning from the gap

The executed motion misses the commanded target; after an online update, the compensated command lands on target.

Wear, payload, and heat make the executed motion drift from the command. That drift is visible through proprioception alone, so it can serve as a free training signal during deployment.

The pretrained VLA stays frozen. Only small LoRA adapters on the action expert are updated, and the robot’s low-level controller is left untouched.

Observation camera images proprioception instruction VLA π₀ / π₀.₅ VLM + action expert FROZEN LoRA TRAINABLE Robot executes unchanged low-level controller wear · payload · heat → error command apolicy executed Δx 01 Residual ηobs = Δx − apolicy 02 Compensated target atarget = apolicy − ηobs 03 Online update flow-matching loss + anchor to init UPDATES LoRA ONLY
01 MEASURE

Residual from proprioception

After each executed step, compare the achieved motion with the command. No reward, label, or error model is needed.

02 COMPENSATE

A corrected target

Subtracting the residual from the command gives the command that would have produced the intended motion.

03 ADAPT

Online LoRA update

Rank-4 LoRA is trained toward that target during deployment, so later chunks pre-compensate.

04

Two real arms, four tasks

A new AgileX Piper arm (Robot A) and one used for a year (Robot B). Both show real command–execution mismatch (mean normalized residual 32.9% and 35.0%).

Robot Backbone Task

Base
Ours

Average success over 4 tasks

25 episodes per task · top-right: gain of Ours over Base

Per-task results
TaskRobot ARobot B
π0.5π0π0.5π0
BaseRVLAOursBaseRVLAOursBaseRVLAOursBaseRVLAOurs
1485284445276445680365660
2366076402880364872325284
3484876446872565284524484
4404868324060364460324456
Avg.435276404772435074384971

Tasks 1–3: pick up the gray bowl (next to the plate / on the plastic cabinet / on the gift box) and place it on the plate. Task 4: open the cabinet’s top drawer. RVLA = RobustVLA.

05

Unseen objects

Bowls and plates never seen in the fine-tuning demonstrations. Robot A, π0.5, 5 variants × 5 episodes.

Seen gray bowl: start and end frames Unseen wooden bowl: start and end frames Unseen plum-colored bowl: start and end frames Unseen light-blue bowl: start and end frames
06

RoboStress

A fleet of worn, hot, or heavily loaded robots is impractical to assemble. RoboStress simulates them with four joint-level error models, each injected where it physically arises.

Friction

Stribeck model: sticky at low speed.

torque

Gravity comp.

Pose-dependent load error.

torque

Backlash

Dead zone on direction reversal.

position

Compliance

Joints deflect under load.

position
EE command→ OSC→ torque + friction, gravity→ position + backlash, compliance→ physics

Seven deployment scenarios

Same LIBERO episode, clean (top) vs. under the scenario (bottom).

Clean
RoboStress
Scenario definitions
ScenarioPhysical causeFric.Grav.Back.Comp.Severity over episodeJoints
Heavy Payloadgrasped load✓✓while graspingall
Thermal Drift-Stribecktemperature✓linearly increasingall
Thermal Drift-Backlashtemperature✓linearly increasingall
Aged Transmissiongear wear✓✓fixedall
Aged Joint-Uniformgear + bearing wear✓✓✓fixedall
Aged Joint-Shouldergear + bearing wear✓✓✓fixedJ1–J3
Aged Joint-Elbowgear + bearing wear✓✓✓fixedJ4

Closer to real errors than Gaussian noise

End-effector deviation replaying a pick-and-place with a 1 kg bowl. Shared y-axis; shaded = before grasp; thin = repeats, thick = mean.

Real Piper arm holding a weighted bowl
Real
Simulated arm holding a bowl
Sim
Real robot
reference
RoboStress
DTW to real 0.027
Gaussian action noise
DTW to real 0.388
07

Numbers on RoboStress

Success rate over the four LIBERO suites (40 tasks, 20 episodes each). Training-time robustness (DR, RobustVLA) gives inconsistent gains; self-compensation is best in all seven scenarios on both backbones.

08

Citation

@article{lee2026taming,
  author  = {Lee, Sohyun and Baek, Yoonjae and Won, Jaesang and Kim, Jinnyeong and
             Kang, Hyunwoo and Baek, Seung-Hwan and Laptev, Ivan and Kwak, Suha},
  title   = {Taming {VLAs} under Robot Execution Errors: Self-Compensation and Stress Testing},
  journal = {arXiv preprint arXiv:2609.37334},
  year    = {2026},
}