HomeBlogBlog Detail

Scaling the Science of Robot Decisions with Ray on Anyscale

By Ian Jordan, PhD   |   October 9, 2026

Closed-loop evaluation of a robot foundation model on Ray is a solved pattern. It incorporates one policy service, simulators as Ray tasks, and a hundred rollouts from one script. We wrote it up in Scale Robot Policy Evaluation with Ray. This post is about what you can learn once evaluation is that cheap. 

We take a frozen real-data vision-language-action (VLA) policy, LeRobot's pi0.5 laundry-folding checkpoint, and run it inside a physics twin of the robot it was trained on. Then we use the twin as an instrument. We perform fast, controlled experiments that change one thing between runs and measure how the frozen policy's decisions change. That answers two questions at once. What does the fine-tuned policy actually attend to? And what does a simulator, or a data-collection or synthetic-data pipeline, have to get right for the policy to keep working?

The mechanics are simple here for demonstration, but effective. We take an observation the policy saw in the twin, replace one of its three images with a real frame from the same phase of the task, keep everything else identical, and count what the policy plans. That located the policy's lift-or-release decision in a few hundred pixels at the jaw tips of its wrist cameras. It told us what the twin's grasp had to look like. And it produced our first zero-shot fold-overs of a cloth shirt by a frozen VLA in simulation, with no real pixels and no fine-tuning. The whole study is one Ray program, and our notebook reproduces its core in six minutes on four L4s.

LinkThe problem

You have a VLA that works on the real rig. You move a camera, or put it in a simulator, and it stops working, and nobody can tell you which pixels mattered. Fine-tuning on more data is the usual answer, and it is a blind one. Ultimately, you don’t know what to collect. Benchmarks like LIBERO-Plus show the failure (viewpoint or initial-state changes drop success from 95% to under 30%) but not the cause (Fu et al., 2025). 

The field's answer over the last eighteen months has been to open the box in one of two ways. Activation-level work hands you a direction in a 2,000-dimensional space. Behavior-level work in simulators hands you a success rate. Neither hands a hardware engineer a decision about where to put a camera, and neither answers the question a simulator engineer has first, which is “what must change in the scene for a frozen policy to keep doing its job?” That needs a controlled experiment. An experiment that uses the same state, the same images, and the ability to change one thing to then analyze what the policy does. A real robot cannot give you that, because no two resets match. What does allow for this is a physics twin. Where in the network the decision lives is the other question, and the activation-level tools are the next step once the twin exists (see references in Further Reading).

Video and splat twins cannot provide that counterfactual for cloth, because there is no physics of the pinch. Furthermore, there is nothing that tells you what the fabric in between the jaws does when the hand lifts. And laundry is where rigid-object evaluation stops working. The shirt has 25,000 degrees of freedom and every grasp changes its shape. So we built a physics twin of this policy’s own rig, from the dataset’s frames.

LinkThe setup: a frozen pi0.5 and a twin of its own rig

The policy is lerobot/folding_latest, pi0.5 fine-tuned by LeRobot on real bimanual OpenArm t-shirt folding. It is a 3B-parameter PaliGemma VLM with a 300M flow-matching action expert. It sees three cameras (a head camera at 640 x 480 and two wrist cameras at 1280 x 720), a 16-dimensional joint state, and the instruction "Fold the T-shirt properly". We never modify or fine-tune it. Every call is the public checkpoint in bf16.

The twin is Isaac Lab 3.0 with the Newton physics backend. The two OpenArm follower arms are MJWarp articulations from the vendor URDF. LeRobot's printed jaws and covers come from their published parts. The dataset's three cameras were placed by fitting their reprojection against hand-read pixel positions in their frames. The table is theirs, and the shirt is a Newton VBD cloth with 25,610 nodes. The policy runs closed loop at 30 Hz on the twin's rendered images and joint state. No real pixel is ever placed into an observation a rollout consumes.

There is one proxy, and it matters. The grasp. A friction pinch on simulated fabric of this thickness does not hold, and every published simulated folding result avoids that the same way we do. When the gripper closes, the cloth nodes between the pads are attached to the gripper frame and carried. Ours adds a pinch fold. The fabric lying across the open jaws is folded into a bunch between the pads, because that is what the wrist camera has to see (we will get to why). A pinned grasp always holds, so the twin validates reaching, folding strategy and visuomotor control, not grasp robustness. That sentence appears in the notebook where the proxy is introduced.

LinkThe architecture: one policy service, three twins, one notebook

Everything below is one notebook on one cluster. A policy service on one GPU, three twins across three GPUs, and the probes as concurrent clients of the same service. The architecture is the one from Scale Robot Policy Evaluation with Ray, which explains why the simulator and the policy cannot share a process or a GPU in more detail. We'll review briefly below:

A policy service on one GPU, three twins across three GPUs, and the probes as concurrent clients of the same service.
A policy service on one GPU, three twins across three GPUs, and the probes as concurrent clients of the same service.

Policy as a service. Here, the replica answers each request with the whole 30-step action chunk and keeps no per-client state:

@serve.deployment(ray_actor_options={"num_gpus": 1}, max_ongoing_requests=64)
@serve.ingress(app)
class PI05PolicyServer:
    def __init__(self):
        self.policy = Pi05Twin(ckpt)                     # lerobot's PI05Policy, bf16, frozen

    @serve.batch(max_batch_size=8, batch_wait_timeout_s=0.05)
    async def _batched(self, obs_list):                  # concurrent requests -> one forward pass
        ...
        return [chunk_i for chunk_i in chunks]           # each (30, 16): absolute joint targets, degrees

    @app.post("/predict")
    async def predict(self, request):                    # npz in: base, left_wrist, right_wrist, state16, task
        return await self._batched(decode(await request.body()))

Stateless chunks are what let many clients share one replica. Each client applies the deployed cadence itself (execute steps 3 through 18 of the chunk, then re-query, exactly what the in-process loop did). Concurrent calls, whether from three simulators or from a probe issuing eighty at once, are batched through a single forward pass. Without a shared replica, three simulators would each load 3B parameters on their own GPU, and one 24 GB L4 cannot hold Isaac Sim, the cloth solver and the policy together.

Twins as tasks. Each twin is a Ray task holding one GPU. The twin itself runs as a subprocess, because Isaac Sim's Kit event loop must not live inside a Ray worker:

@ray.remote(num_gpus=1, num_cpus=4, max_retries=0)
def run_twin(name, policy_url, start_from, seconds):
    cmd = [sys.executable, "-u", "twin/rollout.py", "--policy-url", policy_url,
           "--start-from", start_from, "--seconds", str(seconds), "--out", f"{RUNS}/{name}"]
    return subprocess.run(cmd, ...)                      # the repo is already here: runtime_env working_dir

refs = [run_twin.remote(hold, POLICY_URL, f"data/start_states/{hold}.npz", 12) for hold in HOLDS]

ray.init(runtime_env={"working_dir": "."}) ships the whole repo to every worker, the vendored twin, its assets and the saved holds included. max_retries=0 is deliberate. A silently retried twin would re-run from scratch while overwriting its own frames.

Saved holds. A twin boot costs a minute or three (Kit, the USD scene, the cloth build, RTX warm-up). A settled shirt costs another two hundred simulation steps, and a control step of this cloth costs about a second of physics on an A100. The notebook's twins therefore start from saved holds, with the cloth's node positions, the arm state, the pinned nodes and the gripper hysteresis restored exactly. 

The holds were selected because both hands were pinned, meaning the proxy had attached cloth nodes to each gripper, and the wrist views looked correct by eye. A later applied, stricter test asks whether the jaw was actually over the fabric when it closed. Only the left jaw passes in the shipped states. The right hand’s pin attached nodes from several centimeters away, which the proxy allows and a real gripper cannot. That is the pin proxy’s known failure, and it is why these segments exercise the restore path and the shared replica but are not evidence of a two-hand grasp.

python twin/rollout.py --policy-url http://head:8000 --seconds 60 --seed 9 --save-state-at 600,750,...,1800
python tools/pick_holds.py runs/holds_fwsafe_s9 --out holds_sheet.png        # the numbers, and the frames to look at

The same script also replays a recorded action sequence (--replay-from actions.npy). That is the study's physics-variant loop, a twin change re-tested against a recorded trajectory in eight minutes instead of a thirty-minute policy rollout. Replay is not how the holds were made. Against a 25,000-node cloth an open-loop replay drifts from the closed-loop run within seconds, and the first attempt produced pinned nodes trailing the jaw by 17 to 46 cm, which the jaw-on-cloth rule rejected.

The probe. At a saved moment we have the exact inputs the policy saw in the twin. We replace one image with a real frame from the same phase of the real task (episode 0 at 17.5 to 20 s, both hands holding), keep the state and the other two images as they were, and query the service many times. To be precise about what "real" means here, the real frame is not the same pose or shirt shape as the twin moment. What is identical is everything else within a moment, so the swap is the only difference between conditions, and the counts absorb the sampler's randomness. That mismatch is a limitation, and it is why the next step narrows the swap to a few hundred pixels at the jaw tips:

conditions = {"ALL_SIM":     dict(base=sim.base, left_wrist=sim.lw, right_wrist=sim.rw),
              "base=real":   dict(base=real.base, left_wrist=sim.lw, right_wrist=sim.rw),
              "wrists=real": dict(base=sim.base, left_wrist=real.lw, right_wrist=real.rw),
              "lw=real":     dict(base=sim.base, left_wrist=real.lw, right_wrist=sim.rw),
              "rw=real":     dict(base=sim.base, left_wrist=sim.lw, right_wrist=real.rw)}
rows = probe_grid(client, moments, conditions, repeats=4)     # 4 moments x 5 conditions x 4 = 80 calls, concurrent

Each returned chunk is scored by two counts. A release is a gripper commanded below -20 degrees anywhere in the chunk. A planned lift is a hand's forward-kinematics height at least 20 cm above the table at the chunk's end with no release. The second is a statement about the policy's plan, that it commands the hand up with both grippers held shut, not about the cloth. A chunk cannot tell you whether the shirt would rise. Only the closed-loop twins below measure that. Counts, never a single call. A pi0.5 chunk is a flow-matching sample, and the notebook's first experiment asks the same moment four times and plots four different hand trajectories, sometimes tens of centimeters apart. The study's rule was at least 12 samples per condition before any sentence about the policy was written.

LinkResults: the policy reads the jaw tips

Here is the grid at four holds from the fold configuration (both hands pinned, with the right-hand caveat above), 16 samples per condition, from the notebook's run for this post:

condition           n  plans lift  release  rel L  rel R
ALL_SIM            16           0       14     11     13
base=real          16           0       16      8     16
wrists=real        16          13        0      0      0
lw=real            16           0       16      0     16
rw=real            16           0       16     16      7
What the policy was shown: the twin's wrist views (as the policy saw them) and the real frames swapped in for them.
What the policy was shown: the twin's wrist views (as the policy saw them) and the real frames swapped in for them.

Read it row by row. With the twin's own images the policy lets go (14 of 16 releases, 0 planned lifts). Swap in the real head-camera frame and nothing changes. Swap in both real wrist frames and you get 13 of 16 planned lifts and no releases. Swap in only the real left wrist frame and the left hand never releases (0 of 16) while the right still does (16 of 16). Swap in only the right wrist and the pattern mirrors (left 16 of 16, right 7 of 16). The left hand's own view decides the left release. The study's numbers over more moments agree. With the twin's wrists, planned lifts 0 of 16 and releases 8 to 16 of 16. With real wrists, planned lifts 11 to 16 of 16 and releases 0 to 3. The real frame differs from the twin’s in more than the bunch, so this grid only localises the decision to the wrist cameras. The next experiment isolates what in the frame.

What in the wrist frame? The real frame at a hold shows a pinched bunch of fabric between the jaw tips. The twin's frame showed a thin sliver. Take the real frame and remove only the bunch (gap) and the planned lift dies (0 of 16 in the notebook's run, 0 of 8 in the study). Keep only the real jaw-tip patch, a few hundred pixels, inside the twin's frame (realjaw) and it plans lifts like the full real frame (11 of 16 against 12 of 16, and 6 of 8 against 5 of 8 in the study). Thirty renderer variants (lights, materials, camera, blur, contrast) moved the encoder distance by under 0.03 (0.167 to 0.198, against 0.073 for the real frames), and the 18 image-level variants we scored by lift count (color, sharpness, lens distortion, framing, jaw silhouettes, painted gaps and wads, tuft shapes) planned 0 to 3 lifts of 12 each, against 0 to 1 for the unedited twin frame. The real bunch grayed, darkened, blurred or flattened still planned 9 to 12 of the 12. The cue is geometry not texture. The decision lives in the jaw-tip patch, per hand, and the cue is a bunch of fabric between the pads.

The policy's own encoder confirms it, deterministically. Embed the jaw-tip crop with the SigLIP vision tower inside pi0.5, mean-pool, and take the cosine distance to the real hold frames. The study's real hold frames sit at 0.073, other real moments at 0.09 to 0.15, our renders at 0.15 to 0.22. The notebook's left-wrist crops read twin 0.166, real 0.086, jaw-tip patch 0.089, bunch removed 0.166. Over 65 wrist-image variants with at least 8 samples each, that distance ranks the variants' lift rates with a Spearman correlation of -0.82. A deterministic number that predicts a stochastic count is what turns a scan into a search. The notebook computes it through the same replica's /embed route.

Two things to take from the grid. One geometric cue decided more than thirty render variants, so test what the policy reads before you pay for fidelity. Furthermore, the policy’s own encoder gives a deterministic distance that ranks variants, so the search can run without a rollout. 

LinkFrom finding to fold: what changed in the twin

The finding is a hardware fact and a twin fact at once. The pinch fold at the jaws exists because of it. Fabric lying across the open jaws is folded into a bunch between the pads when the gripper closes, and the wrist camera then sees what the real one sees. Two more changes came from looking rather than probing. The cloth must not move on its own. A simulated shirt can shake with no arm near it. Its 25,000 nodes sit on stiff springs and are stepped by an iterative solver with a capped number of iterations, which leaves a residual error each substep. With stiff enough springs that residual rings at the mesh's own resonances instead of decaying, and numerical damping and the rest-sleep hooks are what remove it. A stiffness change that made the sleeves and collar wave at 12 to 15 mm/s was rejected on exactly that ground. The shipped configuration's median node speed is 0.0 mm/s, and every run records it, so the claim is measured rather than eyeballed. And the table geometry and layout come from their footage, not from a bench.

With that configuration, one policy, one twin, 42 seeds of 120 or 150 seconds, 41 valid. Both hands pinned, flattening and squaring in all 41 runs where no arm went through the table (the gate every run is checked against; the one failure is among the 2 off-table seeds). A fold-over moment (footprint at or below 0.70 of the flat shirt) in 21. A fold still present at the end in 10. The shirt dragged off the near edge of the table in 2. The notebook recomputes the table from the shipped per-run summaries with the rule printed next to it. The metric-only rule gives 22, 11 and 2, and the study's by-eye count judged one end state a crumple. The best run folds the shirt in half between 100 and 119 seconds, and the failure modes get equal screen time.

The best of 42 seeds (fwsafe s9) as five head-camera frames from 0 to 120 s. The best of 42 seeds (fwsafe_s9) as five head-camera frames from 0 to 120 s. The right hand folds the shirt in half between 100 and 119 s.
The best of 42 seeds (fwsafe s9) as five head-camera frames from 0 to 120 s.

A word on what a "lift" is. In the pile experiments the hand rose 30 to 50 cm while the pinned cloth stayed on the table. Hand height is not a carry. Every figure in the notebook overlays the hand height against the cloth's own lift fraction and footprint, and the sentence "the policy lifted" is only ever written about the cloth.

LinkThe payoff: a multi-day study as one Ray program

This is the spine of the piece, so here is the accounting. The multi-day study behind these numbers was one program on one 8-GPU node. Every rollout, every substitution scan and every metric was a Ray task launched from one shell with working_dir packaging. That gave us 42 seeds and dozens of twin variants overnight, and the replayed action sequence cut the physics-variant loop from thirty minutes to eight. The notebook is the same three primitives at tutorial size. One Ray Serve replica on one L4 holds the policy. Three Isaac Lab twins on the other three L4s run as tasks querying it over HTTP. The substitution probes are concurrent clients of the same replica, and every GPU is released when its phase ends. The evaluation post ran a hundred rollouts of one policy. Here the same kind of service answers a few hundred counterfactual queries and boots twins on demand, in the same six minutes on the development box.

Measured in our example notebook code (8 x A100 development box, budgeted to work on 4 x L4):

cell                                             wall s
3 serve.run (returns at once)                       4.4     (deployment thread; the replica loads for 156 s in the background)
4 fan out the twins (returns at once)               0.0     (three Kit boots start on three GPUs)
5 wait for the replica + the stochastic chunk     162.0     (the wait: the twins boot and restore their holds meanwhile)
6 substitution probe (80 calls)                    13.8     (21 batches of ~4 on the replica, 667 ms per batch)
7 bunch vs sliver + encoder distance               11.3
8 collect the live rollouts (3 twins, 5 s each)   157.8     (335 s after fan-out; 0.9 steps/s per twin, ~600 ms per shared inference)
9 twin changes + 42-seed table                      0.4
10 shutdown                                         2.4
TOTAL (sum of cells)                              353.0     = 5.9 min on 8 x A100-40G (three GPUs for twins, one for the policy)

Two infrastructure details carried the budget. The replica batches concurrent requests, so eighty probe calls cost seconds rather than the minutes eighty sequential samples of a 3B-parameter model would. And the twins do not step until the service answers /stats, so the weight load, the three Kit boots and the probes all overlap.

LinkWhat this means for your own policy

The jaw-tip finding is specific to this checkpoint and this rig. The way we got it is not, and that is the takeaway. Before you collect more data, build a simulator, or generate synthetic frames for a fine-tune, you can ask the policy what it reads, in a few hundred cheap queries against a served replica, and let the counts decide. Four general lessons came out of doing that here and Ray makes it easy.

Test what the policy attends to before you spend on fidelity. Thirty renderer variants changed nothing and one geometric cue changed everything. A channel-by-channel substitution tells you which inputs carry the broad decision, and the policy's own encoder gives you a deterministic distance to search against, before you pay for a photoreal scene or another week of teleop.

Let the finding set the simulator's priorities, and the data pipeline's. Whatever the policy reads is what a twin must get right and what sim-augmented or synthetic data must contain. For this policy it was what the wrist camera sees at the jaw tips. For yours it may be a different camera, a different patch, or a state channel. The same style of probing finds it.

Know what your twin can certify. Ours certifies reach, grasp and flatten, and cannot yet certify a pick from a pile because the cloth solver cannot carry pile fabric. Every twin has such a line, and the probe plus the closed-loop runs tell you where it is, so you know which failures to trust in sim and which to take to the real rig.

Measure the baseline before you fine-tune. Seed statistics over one configuration, with the rule printed next to them, are the number a fine-tune has to beat. Without them, "it works better now" is an anecdote.

None of this needs the specific answer we found. It needs a frozen policy behind a service, a simulator you can boot as a task, and a loop cheap enough to run overnight.

LinkRun it on Anyscale

Along with this blog post is a runnable notebook, the vendored twin with its assets, the saved holds, the observation dumps, the probe harness as a reusable tool for any LeRobot policy, the 42-seed table, and the videos. SETUP.md has the Anyscale compute config that runs it in about six minutes on four A100s. The same notebook runs on eight A100s by changing the config, not the code. The checkpoint is public. You need your own Hugging Face token for the gated PaliGemma tokenizer.

The jaw-tip finding and the fold videos are what you get when this workload becomes cheap to iterate. A frozen 3B-parameter policy queried tens of thousands of times, a cloth simulator booted on demand, and a scientific loop fast enough to run overnight, as one program on one cluster.

LinkFurther reading

Explore Anyscale today

Build, run, and scale any AI workload on Ray with a multi-cloud platform built for production AI.