CARE certifies VLA inference speedups with failure-risk guarantees

Vision-language-action (VLA) models have advanced quickly, but running them at every control step is still expensive. Earlier work speeds up VLA inference with techniques such as action chunking and visual-token pruning, and usually judges the result on latency and average task success. The authors of this paper argue that averages hide a real risk: acceleration may discard information and break tasks the original policy would have solved.
The difficulty is that such failures are hard to measure. Action deviations compound over a closed-loop trajectory, so a task failure only shows up across a full episode. The paper therefore defines an acceleration-induced failure through paired rollouts from identical initial conditions: the case where the reference policy succeeds but the accelerated policy fails.
On that definition the authors build CARE, an approach for certified accelerator selection. CARE runs paired rollouts on a calibration set and uses them to give finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It then deploys the fastest candidate that passes certification, and falls back to the reference policy if none qualify. Because it relies only on terminal outcomes and measured compute, CARE applies unchanged across different acceleration mechanisms. Sequential testing and failure-triggered reference rollouts keep the certification process affordable.
The main evaluation uses four LIBERO suites with OpenVLA-OFT. There CARE certifies speedups of 9.0 to 10.8 times while guaranteeing, at 95% confidence, that at least 85.8% of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to 75% of trials, whereas CARE stays within budget. Its sequential form uses 78.9% fewer rollouts than exhaustive evaluation.
The authors also report that CARE generalizes beyond the main setup: to flow-step reduction for pi_0.5, and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.
Key facts
- CARE selects VLA inference accelerators using paired rollouts on a calibration set, with finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget.
- On four LIBERO suites with OpenVLA-OFT, CARE certifies 9.0 to 10.8 times speedups while guaranteeing, at 95% confidence, that at least 85.8% of reference-solved episodes are preserved.
- Under tight budgets, selectors without guarantees exceed the budget in up to 75% of trials; CARE stays within budget.
- CARE's sequential form uses 78.9% fewer rollouts than exhaustive evaluation.
- If no candidate qualifies, CARE falls back to the reference policy; the paper also reports generalization to pi_0.5 flow-step reduction and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.
Why it matters
Speeding up VLA models is usually judged by latency and average task success. The paper's point is that an average can look fine while the accelerated policy quietly fails on tasks the original policy would have solved. CARE turns that into something measurable: a failure is counted only when the reference succeeds and the accelerated policy fails from the same start. It then attaches a statistical guarantee to the choice of accelerator instead of an average score.
Who it affects
Mainly researchers and engineers who accelerate or deploy VLA policies and need to choose between speed-up techniques such as action chunking, visual-token pruning or flow-step reduction. Because CARE only uses terminal outcomes and measured compute, the authors say it applies unchanged across different acceleration mechanisms, so it is not tied to one technique.
How to use it
According to the abstract, the workflow is: collect paired rollouts of the reference policy and each candidate accelerator on a calibration set, set a failure-risk budget, and let CARE deploy the fastest candidate that is certified within that budget. If none qualify, the reference policy stays in place. No code release, dataset release or date is mentioned.
How solid is it
The results come from the paper's own abstract, which names no authors or institutions. The headline numbers are specific: 9.0 to 10.8 times speedups, a guarantee of at least 85.8% of reference-solved episodes preserved at 95% confidence, and 78.9% fewer rollouts with the sequential form. The evaluations named are four LIBERO suites with OpenVLA-OFT, plus pi_0.5 flow-step reduction and Qwen3.5-9B and Llama-3.1-8B agents in Crafter. The abstract does not state the budget value that yields the 85.8% figure.
Risks and caveats
The speedups are given only as multiples; no absolute latency, hardware or wall-clock figures are given. No real-robot results are mentioned, since the evaluations named are LIBERO suites and Crafter. The 75% figure is an upper bound on how often selectors without guarantees exceed the budget under tight budgets, not a typical rate. The guarantee is statistical and tied to the calibration set and the chosen budget, and certification still needs rollouts, even though the sequential form reduces their number.
“acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics”
— CARE paper abstract