run-policy --policy_type custom and collect-dagger --policy_type custom
use this interface for compressed observations and accepted-plan continuation.
The almond_axol.policy SDK’s serve() implements the
endpoint side of this interface; read on to implement it yourself (for
example, in another language). The handshake carries version: 2; the robot
rejects any other version and never downgrades.
The desktop owns the model, instruction text, normalization, conditioning,
and its prediction cache. The robot owns capture, the dispatch clock, and
replacement of future targets. Every reply contains absolute robot-space
targets in the agreed action layout, units and robot coordinate frames. Names and shapes
are validated on the wire; matching numerical units and frame conventions is
also an endpoint responsibility. The observation and action layouts may
be different (for example, joint observations with Cartesian actions).
Setup
PlanPolicyClient.connect(spec) sends hello. A successful server echoes the
entire validated specification in ready; the client refuses any mismatch.
hello and ready both contain protocol: "axol-policy", version: 2,
ordered state/action names, cameras, and the four scheduling fields above.
Each camera has name, prepared shape: [height, width, 3], and codec: "png".
Camera preparation and action-name conventions must agree with the remote
endpoint; it must reject incompatible specifications before sending ready.
The generic robot client sends dispatch_feedback: true in
hello; ready must echo it. This enables the dispatch reference described
below without naming a policy architecture. PlanSpec.dispatch_feedback
can represent low-level clients without this capability; those clients omit
the capability and dispatch-reference fields from the wire.
An explicit dispatch_feedback: false is not a wire representation. An
endpoint without dispatch-feedback support rejects it during setup, so the
robot client and remote endpoint must agree on this capability. The handshake does
not silently downgrade dispatch feedback.
max_adoption_offset_steps is the greatest permitted age of a reply in
control steps, relative to its locally saved request origin. It never delays
an early response. The wire schema permits None to remove that additional
deadline; rows still expire at their original scheduled times. The generic
robot runtime requires an integer deadline inside the horizon. The initial implementation requires PNG
for every image, preserving RGB pixels exactly after configured preparation.
Recurring request and reply
Each WebSocket message is binary: a big-endian uint32 JSON-header byte count, that many UTF-8 JSON bytes, then the binary payload. WebSocket compression is disabled because image data is already PNG. Aninfer header is:
continuation identifies the actual accepted prediction and its first
remaining row. It means the queued targets are an unchanged suffix of that
prediction, not merely that it was the last one generated. It is null on
bootstrap or after a reset. The endpoint must reject a missing cache entry or
exhausted suffix instead of silently using a different plan. Robot-side temporal
blending or editing of queued targets is incompatible with this compact
reference. Low-level control/filtering remains downstream of those targets.
When dispatch_feedback is enabled, every infer header must also contain
last_dispatched. Its value is either null or an object with
prediction_id and a zero-based row. It identifies the last published
target row whose local dispatch path completed successfully in the current
reset generation. This is local send completion, not a hardware acknowledgment.
It does not report a measured pose, successful physical tracking,
or the applied setpoint after trajectory filtering or inverse kinematics.
The referenced value is the exact prefilter row previously sent by the
desktop. With dispatch feedback disabled, the field must be omitted.
The client freezes both references with the request’s saved row origin,
using the latest confirmed dispatch without waiting for a send in progress.
A send still in progress is not reported as completed, even if its row has
already been reserved. A failed dispatch is never reported as completed;
completion from an invalidated generation cannot repopulate the reference.
last_dispatched: null means there is no confirmed dispatch reference
belonging to the current scheduler/reset generation, including recovery.
A dispatch from an invalidated generation may still finish physically and
remain omitted. Null does not imply that no command was sent since the wire
reset or that the robot has stopped physically.
The two references may name different plans. For example, after dispatching
X[12], the client may accept a delayed plan Y whose next remaining row is
Y[3]. Before dispatching that row, an observation reports
continuation={prediction_id: "Y", from_row: 3} and
last_dispatched={prediction_id: "X", row: 12}. Y[2] was skipped, so
subtracting one from continuation.from_row would not identify the previous
command. Desktop adapters must retain both cited plans when they need this
history.
Fixed-budget adapters receive no latency field. An adaptive adapter may opt
into the optional top-level integer delay_steps; it is conditioning input,
never a command to delay adoption. The generic runtime caps its estimate at
actions_per_chunk; that value means a full horizon or longer. Instruction
text stays on the desktop.
The actions reply header contains:
max_adoption_offset_steps may tighten the negotiated deadline for this
reply. There are no absolute robot ticks on the wire. The robot associates
row zero with the request origin saved locally, drops elapsed or irrevocable
rows, and adopts the remaining suffix immediately. Fixed-prefix conditioning and
real-time chunking require the same robot-side replacement operation.
Lifecycle and failure
After connecting, callclient.reset(episode) before the first inference.
reset carries an integer episode; reset_ok echoes it. Before acknowledging,
the endpoint must clear all cached predictions and recurrent state. Request IDs are nonempty strings,
unique within an episode. Only one request may be outstanding on a client.
On a hold, task change, or new episode, invalidate pending predictions in the
local scheduler immediately. Then serialize reset after any outstanding
request. Do not let a pre-reset reply repopulate the robot queue. Changing
desktop instruction text must follow this episode boundary too.
The generic runtime’s late_policy="blocking_refresh" invalidates the
accepted plan and dispatch reference on a late reply or exhausted horizon.
It serializes a desktop reset after pending inference drains, then holds
while requesting a fresh unconditioned plan. It executes up to
request_interval rows, invalidates again, and repeats. This clears desktop
history at each refresh; blocking mode persists until the scheduler itself
is reset. A late response does not automatically resume asynchronous
continuation. late_policy="abort" instead fails the session.
An error reply has type: "error" and message. The endpoint must reject
malformed requests, unknown continuation or dispatch references, and invalid
model output. The client closes its connection on an error reply, malformed
response or timeout.
After a transport failure, start a new connection and explicitly reset it;
a late reply cannot be mistaken for a subsequent request. The client does
not automatically reconnect and resume motion.
Robot control loop
Select the interface withaxol run-policy --policy_type custom --server_host HOST --server_port 8765 --task TASK.
The normal robot and camera configuration still applies. task labels the
recording; instruction
selection stays at the remote endpoint and is not sent in each observation.
Observation acquisition and WebSocket request/reply run outside the periodic
control thread. There is one outstanding request, one accepted plan and no
queue of speculative future plans. Compression and network waits do not hold
the dispatch lock. A short lock protects the cursor, accepted plan and reset
generation; hardware sends run outside it. The controller reserves a row under
that lock, so a replacement cannot overwrite a row already being sent.
A request saves the next dispatch tick as its local origin. If its origin is
o and the next unreserved tick when the reply arrives is k, the client
adopts the suffix starting at max(0, k - o). Row zero is never retimed to the
reply’s arrival. During bootstrap there is no accepted plan: the robot holds
and the cursor stays fixed, so its first reply starts at row zero. After acquisition,
the client freezes its origin and references; image preparation, encoding and
transport consume the request’s adoption budget. Reusing a camera frame does
not refresh its capture timestamp. Freshness is checked before inference.
A missed control deadline invalidates the plan instead of replaying a burst
of missed targets.
For example, at 30 Hz with a 30-row horizon and a request every 10 rows:
Here each plan’s identifier is the
request_id whose reply published it.
The numerical example assumes sends have completed before the next snapshot;
a send still in progress remains unconfirmed in last_dispatched.
The scheduler copies accepted arrays and marks them immutable. Custom policies
bypass the LeRobot chunk alignment and temporal ensemble code. Their targets
still pass through the existing execution path: joint limits, Cartesian IK
when configured, configured trajectory filtering, the realtime controller and
configured contact safeguards. For independent layouts, robot_config.action_space selects joint
or cartesian without forcing the observation layout to change. Omitting it
preserves the earlier coupled behavior of observe_cartesian.
Why the custom policy interface supports different policy families
The remaining queued targets are completely described by(prediction_id, from_row) because the robot executes one unchanged suffix. A remote endpoint
that retains the exact published plans can reconstruct that suffix without
receiving a duplicate action matrix on every request. It must retain both
referenced plans when continuation and confirmed dispatch name different IDs.
Relative actions must be converted to the negotiated absolute target layout
before publication; model preprocessing, normalization and recurrent state
remain outside the robot runtime.
For real-time chunking, an endpoint can condition its next prediction on the
accepted suffix and use a fixed latency budget or the optional delay estimate.
The client continues dispatching the old plan while inference runs, then
skips the expired prefix of the new plan. The robot does not need to know
whether the endpoint used hard prefix conditioning, inpainting or another
sampling method.
For temporal ensembling, the endpoint can retain overlapping raw predictions,
compute the desired weighted targets and publish the final resulting plan.
The reference must name those final published targets, not an unprocessed
model chunk. The robot neither stores ensemble weights nor re-ensembles the
reply. request_interval=1 permits inference on every control tick, but
matching a per-tick ensemble requires sufficient end-to-end throughput: one
outstanding request does not make an arbitrarily slow policy run at the
control rate.
These are sufficient execution and feedback semantics, not a guarantee
that arbitrary replies are smooth. The endpoint must produce compatible
boundary targets and obey the declared layout and timing. Filters and local
limits constrain motor commands downstream; last_dispatched is not their
output or measured tracking. Timestamped state is the physical feedback. The
protocol does not promise position/velocity/acceleration continuity across
arbitrary model outputs, nor continuity of policy history after recovery.
In blocking-refresh mode the reset deliberately drops that history.
DAgger ownership and recording
collect-dagger --policy_type custom --hold_to_intervene true uses the same wire messages and scheduler. A single
local control loop owns both policy and operator sends. Either held grip
revokes pending policy work immediately and gives the operator that arm;
the other arm holds. Loss of the operator link while grips remain held keeps
that ownership and holds the last command.
With shadow_inference=true, observations and inference continue during
operator control, but replies are discarded. Releasing both grips starts a
new local generation, serializes a remote reset after any outstanding reply,
and holds the last operator command until a fresh plan arrives. A joint-space
handover after IK starts from that held command and constrains velocity and
acceleration. Repeated intervention revokes that transition too. No late
pre-intervention prediction can retake control.
Recording continues within the same dataset episode, with per-row intervention
labels and capture fences at ownership changes. record_joint_actions=true
records resolved joint commands independently of the policy’s wire action
layout. home_on_start=true homes before the first episode;
start_from_current_pose=true preserves the pose set during scene reset for
subsequent starts. Terminal q from idle or an active episode parks through
rest and zero before disabling motors. Contact, fault and interrupted-cleanup
paths retain the existing support-preserving behavior.