Self-compensation at deployment.
Online LoRA updates from the command–execution residual. No rewards, no labels.
Robots don’t move exactly as commanded.
Let the VLA learn from that gap, while it runs.
Online LoRA updates from the command–execution residual. No rewards, no labels.
On a new and a 1-year-old Piper arm, for both π0 and π0.5.
16% → 64% success on bowls and plates absent from fine-tuning.
Four joint-level error models, seven deployment scenarios in simulation.
Vision-language-action (VLA) policies often fail when a robot's executed motion deviates from their commanded action. Such execution errors arise from the robot's mechanics and operating conditions, such as wear and payload changes. We propose self-compensating VLA, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands. Without task rewards or labels, it updates the policy online using the residual between the action commanded by a VLA and the motion executed by the robot. To stress-test VLA robustness across execution conditions that are impractical to cover with physical robots alone, we introduce RoboStress, a controlled simulation benchmark. It combines established joint-level models of friction, backlash, compliance, and gravity-compensation error into seven deployment scenarios whose execution errors depend on the robot's state and motion history. On RoboStress, self-compensating VLA achieves higher average task success than both the base policies and methods that build in robustness during training. On two physical robot arms with different usage histories, it raises the average task success rate by more than 30 percentage points on each arm, and the gains extend to objects not seen in the task demonstrations.
A year-old Piper arm (Robot B). Left: the base policy. Right: the same policy with self-compensation.
“Pick up the gray bowl on the plastic cabinet and place it on the plate”
“Open the top drawer of the cabinet”

Wear, payload, and heat make the executed motion drift from the command. That drift is visible through proprioception alone, so it can serve as a free training signal during deployment.
The pretrained VLA stays frozen. Only small LoRA adapters on the action expert are updated, and the robot’s low-level controller is left untouched.
After each executed step, compare the achieved motion with the command. No reward, label, or error model is needed.
Subtracting the residual from the command gives the command that would have produced the intended motion.
Rank-4 LoRA is trained toward that target during deployment, so later chunks pre-compensate.
A new AgileX Piper arm (Robot A) and one used for a year (Robot B). Both show real command–execution mismatch (mean normalized residual 32.9% and 35.0%).
25 episodes per task · top-right: gain of Ours over Base
| Task | Robot A | Robot B | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| π0.5 | π0 | π0.5 | π0 | |||||||||
| Base | RVLA | Ours | Base | RVLA | Ours | Base | RVLA | Ours | Base | RVLA | Ours | |
| 1 | 48 | 52 | 84 | 44 | 52 | 76 | 44 | 56 | 80 | 36 | 56 | 60 |
| 2 | 36 | 60 | 76 | 40 | 28 | 80 | 36 | 48 | 72 | 32 | 52 | 84 |
| 3 | 48 | 48 | 76 | 44 | 68 | 72 | 56 | 52 | 84 | 52 | 44 | 84 |
| 4 | 40 | 48 | 68 | 32 | 40 | 60 | 36 | 44 | 60 | 32 | 44 | 56 |
| Avg. | 43 | 52 | 76 | 40 | 47 | 72 | 43 | 50 | 74 | 38 | 49 | 71 |
Tasks 1–3: pick up the gray bowl (next to the plate / on the plastic cabinet / on the gift box) and place it on the plate. Task 4: open the cabinet’s top drawer. RVLA = RobustVLA.
Bowls and plates never seen in the fine-tuning demonstrations. Robot A, π0.5, 5 variants × 5 episodes.
A fleet of worn, hot, or heavily loaded robots is impractical to assemble. RoboStress simulates them with four joint-level error models, each injected where it physically arises.
Stribeck model: sticky at low speed.
torquePose-dependent load error.
torqueDead zone on direction reversal.
positionJoints deflect under load.
positionSame LIBERO episode, clean (top) vs. under the scenario (bottom).
| Scenario | Physical cause | Fric. | Grav. | Back. | Comp. | Severity over episode | Joints |
|---|---|---|---|---|---|---|---|
| Heavy Payload | grasped load | ✓ | ✓ | while grasping | all | ||
| Thermal Drift-Stribeck | temperature | ✓ | linearly increasing | all | |||
| Thermal Drift-Backlash | temperature | ✓ | linearly increasing | all | |||
| Aged Transmission | gear wear | ✓ | ✓ | fixed | all | ||
| Aged Joint-Uniform | gear + bearing wear | ✓ | ✓ | ✓ | fixed | all | |
| Aged Joint-Shoulder | gear + bearing wear | ✓ | ✓ | ✓ | fixed | J1–J3 | |
| Aged Joint-Elbow | gear + bearing wear | ✓ | ✓ | ✓ | fixed | J4 |
End-effector deviation replaying a pick-and-place with a 1 kg bowl. Shared y-axis; shaded = before grasp; thin = repeats, thick = mean.


Success rate over the four LIBERO suites (40 tasks, 20 episodes each). Training-time robustness (DR, RobustVLA) gives inconsistent gains; self-compensation is best in all seven scenarios on both backbones.
@article{lee2026taming,
author = {Lee, Sohyun and Baek, Yoonjae and Won, Jaesang and Kim, Jinnyeong and
Kang, Hyunwoo and Baek, Seung-Hwan and Laptev, Ivan and Kwak, Suha},
title = {Taming {VLAs} under Robot Execution Errors: Self-Compensation and Stress Testing},
journal = {arXiv preprint arXiv:2609.37334},
year = {2026},
}