Skip to main content
Run Policy and DAgger collection can drive the arms with your own model, whatever framework it uses (JAX, a different torch version, a VLA served elsewhere, or a hand-written controller). Wrap it in a small server with the almond_axol.policy SDK and select --policy_type custom. The robot streams it joint state and camera frames, and executes the action chunks it returns. The robot needs only the state/action layout, camera configuration and timing contract; it never loads the model. The endpoint (your server) owns preprocessing, inference, instruction selection, postprocessing and any model-specific history or ensembling. The robot keeps executing its accepted plan while inference runs, then adopts the new plan’s remaining rows at their original execution times. Network waits do not own the motor loop. Timing and continuity explain the boundary contract and its limits.

Write the policy

Install almond-axol wherever the model runs. The base install is enough; LeRobot and torch aren’t needed. Subclass Policy and start it with serve():
infer returns T future targets, one row per control tick, in spec.action_names order. See Action chunks for the accepted shapes and State and action names for the layout. If you trained on a dataset from Data Collection, the names match its observation.state and action features. The robot executes each chunk exactly as published and switches to a new one at the row that’s due when it arrives. To keep that hand-over smooth, use one of these:
  • Continue the motion in progress. Condition on obs.plan, or at least start near obs.plan[0].
  • Ensemble on the server. Use serve(..., ensemble=0.01) for an ACT-style chunked policy trained with temporal ensembling.
serve() handles the rest of the interface: it echoes the session contract, caches published plans for continuation, clears them on reset, and relays exceptions to the operator. To implement the endpoint in another language, follow the interface reference directly. The model, checkpoint and instruction live in your policy. On the robot, --task only labels recorded episodes, and --policy_path is not used for custom policies.

Check it without the robot

With the server running, play the robot’s side of a session against it:
policy.check uses the robot’s own scheduler on a virtual clock and simulates 3 ticks of reply latency. It reports late replies that would trigger recovery, the endpoint’s round-trip time, and the largest jumps where one plan hands over to the next. Match --cameras, --cartesian and the chunk settings to your robot configuration. check_policy() does the same from Python, for unit tests.

Run on the robot host

Replace the endpoint address and camera serial with your configuration:
The command connects to the endpoint; it does not spawn a model server. Omitting --server_host connects to an endpoint already listening on 127.0.0.1. The robot and endpoint must agree on the selected camera, state and action schemas before an episode can start. In the control panel, assign cameras, select Run Policy, choose policy type custom, and enter the task label. Configure the endpoint under Settings → Inference and leave the policy path blank. The ordinary episode controls and recording options remain available. See run-policy for the command’s controls and settings. Custom policies use plan_config.request_interval (default 5 rows) to schedule inference and plan_config.max_adoption_offset_steps (default 6) as the reply deadline. The LeRobot aggregate_fn, chunk_size_threshold, temporal-ensemble and chunk-alignment settings do not alter published custom-policy plans. Ensembling belongs at the endpoint; downstream execution filtering remains on the robot.

Collect corrections with DAgger

Hold either grip to operate that arm; release both to request a fresh policy plan. The robot holds the last operator command while waiting, then blends back in joint space after IK. Background predictions can continue during operator control, but they never command motors. Source changes retain the same dataset episode and fence capture so intervention labels remain aligned. DAgger has its own endpoint settings; its shared settings inherit from collect-data, not the run-policy inference category. See collect-dagger for configuration, recording, startup homing and terminal soft shutdown.

Failure behavior

A malformed reply, schema mismatch or transport failure stops the custom policy path. Expired plans follow plan_config.late_policy: abort or hold while obtaining fresh plans after a reset. The robot never retimes an expired row to make it executable and does not automatically reconnect to resume motion. See Lifecycle and failure for the exact behavior and the distinction between a confirmed send and physical tracking.