Skip to main content
The SDK for running a model that isn’t a LeRobot checkpoint. You write a Policy and start it with serve(). Then run Run Policy or DAgger collection with policy type custom. The robot sends each observation (joint state and RGB camera frames) to your infer() and executes the action chunks it returns. serve() implements the endpoint side of the custom policy interface for you. It agrees the session contract with the robot and caches every plan it publishes. It turns the robot’s continuation reference into the rows still executing (Observation.plan), clears that history whenever the robot resets, and can optionally ensemble overlapping chunks. Your code only sees observations and returns chunks. It needs only Axol’s base install (numpy, websockets, OpenCV and Pillow). LeRobot and torch are not required.

Quick start

A plain function works too, for stateless models:

Policy

Base class for a custom policy. Override infer; the other methods are optional hooks. Set these class attributes to have the session refused before anything moves: The model, checkpoint and task instruction belong to your policy: pass them to its constructor. The robot sends neither its task label nor its policy path.

Observation

What infer receives: one observation per request, captured with the same exposure-time alignment as data collection.

PolicySpec

The robot’s session contract, passed to Policy.setup (an alias of PlanSpec).

State and action names

The names follow the robot’s configuration, so you can train on a dataset recorded with collect-data and use its observation.state / action names as-is. Arm joints are radians. The gripper is normalized: 0.0 is closed and 1.0 is fully open. The gripperless SKU drops the gripper.pos entries. The observation and action layouts can differ, e.g. joint observations with Cartesian actions.

Action chunks

infer returns a chunk of future targets:
  • a (T, D) array-like (numpy, torch or nested lists) whose D columns follow PolicySpec.action_names;
  • a single (D,) action (a one-row chunk); or
  • a list of {action_name: value} dicts.
Row k is the target for the k-th tick after the request was taken. The robot keeps executing its current plan while you infer. When your reply arrives, it switches to it at the row that is due by then: if 3 ticks passed, it starts at your row 3. Rows are never retimed. Rows beyond actions_per_chunk are dropped. Every command then goes through the robot’s IK, velocity/acceleration shaping and contact safeguards. The robot executes your chunks as published, without blending them. Two ways to keep the hand-over smooth:
  • Continue the motion in progress. Condition your next chunk on obs.plan, the rows it will replace (real-time chunking). At minimum, start near obs.plan[0].
  • Let the server ensemble. Pass ensemble= to serve() for chunked policies trained with ACT-style temporal ensembling (below).
Replies must arrive within max_adoption_offset_steps rows (default 6, i.e. 200 ms at 30 Hz) and within 10 s overall. The robot then recovers: by default it holds, calls reset, and requests a fresh plan. setup has up to 5 minutes. Check an endpoint with check_policy before running it on the robot.

serve

Serves policy (a Policy or a callable obs -> chunk) until Ctrl+C. It accepts one robot at a time. An exception in your code is logged and sent back to the robot, which stops the rollout and shows the message. ensemble=k temporally ensembles overlapping chunks before publishing them. Each row is a weighted average of every prediction covering that tick, weighted exp(-k·i) with i = 0 the oldest. 0.01 is ACT’s default. Grippers and rotation-vector dims follow the newest prediction instead of being averaged. Leave it at None for policies that already produce smooth plans.
The interface is unauthenticated plaintext, like the LeRobot inference server. Bind 127.0.0.1 when the model runs on the robot’s own machine. Otherwise, keep the server on an isolated, trusted network and firewall the port so only the robot’s IP can connect: anyone who can reach it can impersonate the policy and send arbitrary actions.

PolicyServer

The non-blocking form of serve: run it on a background thread, e.g. in a notebook or test.

check_policy

Exercises an endpoint without a robot, e.g. in a unit test or CI. It plays the robot’s side of a session on a virtual clock, using the same scheduler as run-policy. That scheduler decides when to request and which rows of a delayed reply to adopt, and recovers from late replies. The simulated arm tracks its commands perfectly. axol policy.check is the same from the command line.
delay_steps is the simulated reply latency in ticks, independent of this machine’s speed. default_spec() builds the contract run-policy sends with default settings. Its arguments are cartesian, cameras, image_shape, fps, actions_per_chunk, request_interval and max_adoption_offset_steps. CheckReport has these members:

Robot side and wire format

To implement the endpoint in another language, follow the custom policy interface directly. These types are its Python form, and PlanPolicyClient is the robot’s client:
The SDK speaks interface version 2, the only version the robot accepts.