Submit a training + exam job
Validates the spec, spawns a detached job process, and returns
immediately. Job directories are append-only, re-POSTing an
existing job_id returns 409; pick a new id.
Lineage (train.parent) is an explicit input, never ambient state:
warm-start CMA-ES from a prior champion either inline
({"genome6": [6 floats]}) or by reference
({"job_id": "<prior job>"}, refused with 422 if that job has no
champion). Semantics: x0 = parent genome6, sigma0 = 2.0; absent
parent: x0 = zeros, sigma0 = 6.0 (cold start).
Exam-only certification: train.gens = 0 skips training entirely
and certifies a champion you supply. Three sources (one required for
gens=0): train.parent (a genome6), policy.format=genome9 (a raw
9-float controller), or policy.format=python_module (YOUR OWN policy
as a sandboxed Python module, see the policy field). The exam stage
is byte-identical either way; dream_F is null for supplied policies.
Bring your own architecture: a python_module policy implements
reset(seed) + act(obs) where obs is the same observation dict as
the local gym (proximity, collision, …) and the action is the
wheel-fraction vocabulary. A policy trained in the local gym drops into
the exam unchanged. The module runs sandboxed in the exam container
(no network, resource- and time-limited); ship a Python module doing
numpy inference over your exported weights, there is no ONNX detour.
Use train.T = 60 for dream_F scores comparable to campaign history;
T = 20 scores are not comparable.
Authorizations
Single per-partner bearer key, provisioned by Alakazam.
Body
^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$Optional. Certify a champion you supply instead of training one (requires train.gens=0). The policy replaces ONLY the champion arm's controller; world, spawns, episodes, tick, remap, scoring and the anti-exploit arms stay frozen.

