Skip to main content
run-policy --policy_type custom and collect-dagger --policy_type custom use this interface for compressed observations and accepted-plan continuation. The almond_axol.policy SDK’s serve() implements the endpoint side of this interface; read on to implement it yourself (for example, in another language). The handshake carries version: 2; the robot rejects any other version and never downgrades. The desktop owns the model, instruction text, normalization, conditioning, and its prediction cache. The robot owns capture, the dispatch clock, and replacement of future targets. Every reply contains absolute robot-space targets in the agreed action layout, units and robot coordinate frames. Names and shapes are validated on the wire; matching numerical units and frame conventions is also an endpoint responsibility. The observation and action layouts may be different (for example, joint observations with Cartesian actions).

Setup

PlanPolicyClient.connect(spec) sends hello. A successful server echoes the entire validated specification in ready; the client refuses any mismatch.
hello and ready both contain protocol: "axol-policy", version: 2, ordered state/action names, cameras, and the four scheduling fields above. Each camera has name, prepared shape: [height, width, 3], and codec: "png". Camera preparation and action-name conventions must agree with the remote endpoint; it must reject incompatible specifications before sending ready. The generic robot client sends dispatch_feedback: true in hello; ready must echo it. This enables the dispatch reference described below without naming a policy architecture. PlanSpec.dispatch_feedback can represent low-level clients without this capability; those clients omit the capability and dispatch-reference fields from the wire. An explicit dispatch_feedback: false is not a wire representation. An endpoint without dispatch-feedback support rejects it during setup, so the robot client and remote endpoint must agree on this capability. The handshake does not silently downgrade dispatch feedback. max_adoption_offset_steps is the greatest permitted age of a reply in control steps, relative to its locally saved request origin. It never delays an early response. The wire schema permits None to remove that additional deadline; rows still expire at their original scheduled times. The generic robot runtime requires an integer deadline inside the horizon. The initial implementation requires PNG for every image, preserving RGB pixels exactly after configured preparation.

Recurring request and reply

Each WebSocket message is binary: a big-endian uint32 JSON-header byte count, that many UTF-8 JSON bytes, then the binary payload. WebSocket compression is disabled because image data is already PNG. An infer header is:
The image payload concatenates complete PNG byte streams in negotiated camera order. Timestamps are integer nanoseconds in one robot monotonic clock; they are not comparable to the desktop clock. The decoded image representation is an RGB uint8 array. Camera count/order/shape, PNG format, lengths, finite state values, and all message fields are validated. continuation identifies the actual accepted prediction and its first remaining row. It means the queued targets are an unchanged suffix of that prediction, not merely that it was the last one generated. It is null on bootstrap or after a reset. The endpoint must reject a missing cache entry or exhausted suffix instead of silently using a different plan. Robot-side temporal blending or editing of queued targets is incompatible with this compact reference. Low-level control/filtering remains downstream of those targets. When dispatch_feedback is enabled, every infer header must also contain last_dispatched. Its value is either null or an object with prediction_id and a zero-based row. It identifies the last published target row whose local dispatch path completed successfully in the current reset generation. This is local send completion, not a hardware acknowledgment. It does not report a measured pose, successful physical tracking, or the applied setpoint after trajectory filtering or inverse kinematics. The referenced value is the exact prefilter row previously sent by the desktop. With dispatch feedback disabled, the field must be omitted. The client freezes both references with the request’s saved row origin, using the latest confirmed dispatch without waiting for a send in progress. A send still in progress is not reported as completed, even if its row has already been reserved. A failed dispatch is never reported as completed; completion from an invalidated generation cannot repopulate the reference. last_dispatched: null means there is no confirmed dispatch reference belonging to the current scheduler/reset generation, including recovery. A dispatch from an invalidated generation may still finish physically and remain omitted. Null does not imply that no command was sent since the wire reset or that the robot has stopped physically. The two references may name different plans. For example, after dispatching X[12], the client may accept a delayed plan Y whose next remaining row is Y[3]. Before dispatching that row, an observation reports continuation={prediction_id: "Y", from_row: 3} and last_dispatched={prediction_id: "X", row: 12}. Y[2] was skipped, so subtracting one from continuation.from_row would not identify the previous command. Desktop adapters must retain both cited plans when they need this history. Fixed-budget adapters receive no latency field. An adaptive adapter may opt into the optional top-level integer delay_steps; it is conditioning input, never a command to delay adoption. The generic runtime caps its estimate at actions_per_chunk; that value means a full horizon or longer. Instruction text stays on the desktop. The actions reply header contains:
Its payload is row-major little-endian float32 values. An optional max_adoption_offset_steps may tighten the negotiated deadline for this reply. There are no absolute robot ticks on the wire. The robot associates row zero with the request origin saved locally, drops elapsed or irrevocable rows, and adopts the remaining suffix immediately. Fixed-prefix conditioning and real-time chunking require the same robot-side replacement operation.

Lifecycle and failure

After connecting, call client.reset(episode) before the first inference. reset carries an integer episode; reset_ok echoes it. Before acknowledging, the endpoint must clear all cached predictions and recurrent state. Request IDs are nonempty strings, unique within an episode. Only one request may be outstanding on a client. On a hold, task change, or new episode, invalidate pending predictions in the local scheduler immediately. Then serialize reset after any outstanding request. Do not let a pre-reset reply repopulate the robot queue. Changing desktop instruction text must follow this episode boundary too. The generic runtime’s late_policy="blocking_refresh" invalidates the accepted plan and dispatch reference on a late reply or exhausted horizon. It serializes a desktop reset after pending inference drains, then holds while requesting a fresh unconditioned plan. It executes up to request_interval rows, invalidates again, and repeats. This clears desktop history at each refresh; blocking mode persists until the scheduler itself is reset. A late response does not automatically resume asynchronous continuation. late_policy="abort" instead fails the session. An error reply has type: "error" and message. The endpoint must reject malformed requests, unknown continuation or dispatch references, and invalid model output. The client closes its connection on an error reply, malformed response or timeout. After a transport failure, start a new connection and explicitly reset it; a late reply cannot be mistaken for a subsequent request. The client does not automatically reconnect and resume motion.

Robot control loop

Select the interface with axol run-policy --policy_type custom --server_host HOST --server_port 8765 --task TASK. The normal robot and camera configuration still applies. task labels the recording; instruction selection stays at the remote endpoint and is not sent in each observation. Observation acquisition and WebSocket request/reply run outside the periodic control thread. There is one outstanding request, one accepted plan and no queue of speculative future plans. Compression and network waits do not hold the dispatch lock. A short lock protects the cursor, accepted plan and reset generation; hardware sends run outside it. The controller reserves a row under that lock, so a replacement cannot overwrite a row already being sent. A request saves the next dispatch tick as its local origin. If its origin is o and the next unreserved tick when the reply arrives is k, the client adopts the suffix starting at max(0, k - o). Row zero is never retimed to the reply’s arrival. During bootstrap there is no accepted plan: the robot holds and the cursor stays fixed, so its first reply starts at row zero. After acquisition, the client freezes its origin and references; image preparation, encoding and transport consume the request’s adoption budget. Reusing a camera frame does not refresh its capture timestamp. Freshness is checked before inference. A missed control deadline invalidates the plan instead of replaying a burst of missed targets. For example, at 30 Hz with a 30-row horizon and a request every 10 rows: Here each plan’s identifier is the request_id whose reply published it. The numerical example assumes sends have completed before the next snapshot; a send still in progress remains unconfirmed in last_dispatched. The scheduler copies accepted arrays and marks them immutable. Custom policies bypass the LeRobot chunk alignment and temporal ensemble code. Their targets still pass through the existing execution path: joint limits, Cartesian IK when configured, configured trajectory filtering, the realtime controller and configured contact safeguards. For independent layouts, robot_config.action_space selects joint or cartesian without forcing the observation layout to change. Omitting it preserves the earlier coupled behavior of observe_cartesian.

Why the custom policy interface supports different policy families

The remaining queued targets are completely described by (prediction_id, from_row) because the robot executes one unchanged suffix. A remote endpoint that retains the exact published plans can reconstruct that suffix without receiving a duplicate action matrix on every request. It must retain both referenced plans when continuation and confirmed dispatch name different IDs. Relative actions must be converted to the negotiated absolute target layout before publication; model preprocessing, normalization and recurrent state remain outside the robot runtime. For real-time chunking, an endpoint can condition its next prediction on the accepted suffix and use a fixed latency budget or the optional delay estimate. The client continues dispatching the old plan while inference runs, then skips the expired prefix of the new plan. The robot does not need to know whether the endpoint used hard prefix conditioning, inpainting or another sampling method. For temporal ensembling, the endpoint can retain overlapping raw predictions, compute the desired weighted targets and publish the final resulting plan. The reference must name those final published targets, not an unprocessed model chunk. The robot neither stores ensemble weights nor re-ensembles the reply. request_interval=1 permits inference on every control tick, but matching a per-tick ensemble requires sufficient end-to-end throughput: one outstanding request does not make an arbitrarily slow policy run at the control rate. These are sufficient execution and feedback semantics, not a guarantee that arbitrary replies are smooth. The endpoint must produce compatible boundary targets and obey the declared layout and timing. Filters and local limits constrain motor commands downstream; last_dispatched is not their output or measured tracking. Timestamped state is the physical feedback. The protocol does not promise position/velocity/acceleration continuity across arbitrary model outputs, nor continuity of policy history after recovery. In blocking-refresh mode the reset deliberately drops that history.

DAgger ownership and recording

collect-dagger --policy_type custom --hold_to_intervene true uses the same wire messages and scheduler. A single local control loop owns both policy and operator sends. Either held grip revokes pending policy work immediately and gives the operator that arm; the other arm holds. Loss of the operator link while grips remain held keeps that ownership and holds the last command. With shadow_inference=true, observations and inference continue during operator control, but replies are discarded. Releasing both grips starts a new local generation, serializes a remote reset after any outstanding reply, and holds the last operator command until a fresh plan arrives. A joint-space handover after IK starts from that held command and constrains velocity and acceleration. Repeated intervention revokes that transition too. No late pre-intervention prediction can retake control. Recording continues within the same dataset episode, with per-row intervention labels and capture fences at ownership changes. record_joint_actions=true records resolved joint commands independently of the policy’s wire action layout. home_on_start=true homes before the first episode; start_from_current_pose=true preserves the pose set during scene reset for subsequent starts. Terminal q from idle or an active episode parks through rest and zero before disabling motors. Contact, fault and interrupted-cleanup paths retain the existing support-preserving behavior.