Policy and start it with serve(). Then run Run Policy or DAgger collection with policy type custom. The robot sends each observation (joint state and RGB camera frames) to your infer() and executes the action chunks it returns.
serve() implements the endpoint side of the custom policy interface for you. It agrees the session contract with the robot and caches every plan it publishes. It turns the robot’s continuation reference into the rows still executing (Observation.plan), clears that history whenever the robot resets, and can optionally ensemble overlapping chunks. Your code only sees observations and returns chunks.
It needs only Axol’s base install (numpy, websockets, OpenCV and Pillow). LeRobot and torch are not required.
Quick start
Policy
Base class for a custom policy. Override infer; the other methods are optional hooks.
Set these class attributes to have the session refused before anything moves:
The model, checkpoint and task instruction belong to your policy: pass them to its constructor. The robot sends neither its task label nor its policy path.
Observation
What infer receives: one observation per request, captured with the same exposure-time alignment as data collection.
PolicySpec
The robot’s session contract, passed to Policy.setup (an alias of PlanSpec).
State and action names
The names follow the robot’s configuration, so you can train on a dataset recorded withcollect-data and use its observation.state / action names as-is.
Arm joints are radians. The gripper is normalized:
0.0 is closed and 1.0 is fully open. The gripperless SKU drops the gripper.pos entries. The observation and action layouts can differ, e.g. joint observations with Cartesian actions.
Action chunks
infer returns a chunk of future targets:
- a
(T, D)array-like (numpy, torch or nested lists) whoseDcolumns followPolicySpec.action_names; - a single
(D,)action (a one-row chunk); or - a list of
{action_name: value}dicts.
k is the target for the k-th tick after the request was taken. The robot keeps executing its current plan while you infer. When your reply arrives, it switches to it at the row that is due by then: if 3 ticks passed, it starts at your row 3. Rows are never retimed. Rows beyond actions_per_chunk are dropped. Every command then goes through the robot’s IK, velocity/acceleration shaping and contact safeguards.
The robot executes your chunks as published, without blending them. Two ways to keep the hand-over smooth:
- Continue the motion in progress. Condition your next chunk on
obs.plan, the rows it will replace (real-time chunking). At minimum, start nearobs.plan[0]. - Let the server ensemble. Pass
ensemble=toserve()for chunked policies trained with ACT-style temporal ensembling (below).
max_adoption_offset_steps rows (default 6, i.e. 200 ms at 30 Hz) and within 10 s overall. The robot then recovers: by default it holds, calls reset, and requests a fresh plan. setup has up to 5 minutes. Check an endpoint with check_policy before running it on the robot.
serve
policy (a Policy or a callable obs -> chunk) until Ctrl+C. It accepts one robot at a time. An exception in your code is logged and sent back to the robot, which stops the rollout and shows the message.
ensemble=k temporally ensembles overlapping chunks before publishing them. Each row is a weighted average of every prediction covering that tick, weighted exp(-k·i) with i = 0 the oldest. 0.01 is ACT’s default. Grippers and rotation-vector dims follow the newest prediction instead of being averaged. Leave it at None for policies that already produce smooth plans.
PolicyServer
The non-blocking form of serve: run it on a background thread, e.g. in a notebook or test.
check_policy
Exercises an endpoint without a robot, e.g. in a unit test or CI. It plays the robot’s side of a session on a virtual clock, using the same scheduler as run-policy. That scheduler decides when to request and which rows of a delayed reply to adopt, and recovers from late replies. The simulated arm tracks its commands perfectly. axol policy.check is the same from the command line.
delay_steps is the simulated reply latency in ticks, independent of this machine’s speed. default_spec() builds the contract run-policy sends with default settings. Its arguments are cartesian, cameras, image_shape, fps, actions_per_chunk, request_interval and max_adoption_offset_steps.
CheckReport has these members:
Robot side and wire format
To implement the endpoint in another language, follow the custom policy interface directly. These types are its Python form, andPlanPolicyClient is the robot’s client:
The SDK speaks interface version 2, the only version the robot accepts.
