> ## Documentation Index
> Fetch the complete documentation index at: https://docs.alakazam.gg/llms.txt
> Use this file to discover all available pages before exploring further.

# Submit a training + exam job

> Validates the spec, spawns a detached job process, and returns
immediately. Job directories are **append-only**, re-POSTing an
existing `job_id` returns `409`; pick a new id.

**Lineage** (`train.parent`) is an explicit input, never ambient state:
warm-start CMA-ES from a prior champion either inline
(`{"genome6": [6 floats]}`) or by reference
(`{"job_id": "<prior job>"}`, refused with `422` if that job has no
champion). Semantics: `x0 = parent genome6, sigma0 = 2.0`; absent
parent: `x0 = zeros, sigma0 = 6.0` (cold start).

**Exam-only certification**: `train.gens = 0` skips training entirely
and certifies a champion you supply. Three sources (one required for
gens=0): `train.parent` (a genome6), `policy.format=genome9` (a raw
9-float controller), or `policy.format=python_module` (YOUR OWN policy
as a sandboxed Python module, see the `policy` field). The exam stage
is byte-identical either way; `dream_F` is `null` for supplied policies.

**Bring your own architecture**: a `python_module` policy implements
`reset(seed)` + `act(obs)` where `obs` is the same observation dict as
the local gym (`proximity`, `collision`, ...) and the action is the
wheel-fraction vocabulary. A policy trained in the local gym drops into
the exam unchanged. The module runs sandboxed in the exam container
(no network, resource- and time-limited); ship a Python module doing
numpy inference over your exported weights, there is no ONNX detour.

Use `train.T = 60` for `dream_F` scores comparable to campaign history;
`T = 20` scores are not comparable.




## OpenAPI

````yaml /train-v1.yaml post /jobs
openapi: 3.1.0
info:
  title: Alakazam Train API
  version: 1.0.0
  summary: >-
    Train policies inside a world model, certify them in a frozen physics
    oracle.
  description: >
    The **Train API** runs the dream-training pipeline as reproducible jobs:


    1. TRAIN: CMA-ES evolves a 6-parameter controller *inside the doom world
       model* (the same public checkpoint as
       [`alakazamworld/doom-dungeon-hg`](https://huggingface.co/alakazamworld/doom-dungeon-hg)
       on Hugging Face, byte-identical, sha256-verified).
    2. EXAM: the champion is replayed in a frozen Webots oracle (two physical
       worlds, fixed contract, anti-exploit control arms). The oracle is the
       sole scoreboard; dream fitness is never a capability claim.

    One `POST /jobs` does both and leaves a machine-readable `job.json` you can

    poll at any time, jobs are asynchronous, kill-tolerant, and append-only.


    ### Honest verdict semantics

    The exam verdict has four bars per world (contacts, clearance, coverage,

    path). Warm-started jobs reproduce the historical best-known profile,
    contact-safe

    (0 contacts in 40/40 episodes), but no policy in program history has

    passed all four bars. A `pass: false` verdict with clean

    contacts is the expected state of the art, not an error.


    ### Reproducibility

    Training is bit-deterministic for a given spec seed **on a fixed platform**;

    champions differ across platforms (macOS-arm64 vs Linux-amd64 float paths).

    Exam results are deterministic to aggregates (third-decimal physics drift).


    ### Availability & timings

    The serving VM runs on-demand (it is started for sessions, not 24/7).

    Measured wall-times, serving VM (e2-standard-2; CPU generation varies

    per boot): smoke job (`pop 6, gens 2, T 20`) 12-17 min; real-scale

    (`pop 10, gens 8, T 60`) 3-4 h; exam alone 4-6 min. Poll `GET
    /jobs/{job_id}`, never block on the POST.
servers:
  - url: https://api.alakazam.gg/train
    description: >
      The Train jobs + exam API. Bearer-key auth; the serving VM behind it is
      start-on-demand, so coordinate a session window with us before calling.
security:
  - bearerAuth: []
tags:
  - name: Jobs
    description: Submit and poll dream-training jobs.
  - name: Simulation gym
    description: >
      **Beta, self-hosted runner.** Drive a world headlessly as a training
      environment (gym) and stream the SNN-controller observation contract: two
      virtual proximity sensors, a terminal collision spike, object labels, and
      an optional square RGB camera. The same schema is emitted offline by the
      dataset exporter (JSONL episodes), so controllers trained on recorded data
      plug straight into the live loop. Each session holds a real world-model
      GPU stream, sessions are metered by wall-time; create only as many
      parallel environments as your session budget allows, and DELETE them when
      done.
paths:
  /jobs:
    post:
      tags:
        - Jobs
      summary: Submit a training + exam job
      description: |
        Validates the spec, spawns a detached job process, and returns
        immediately. Job directories are **append-only**, re-POSTing an
        existing `job_id` returns `409`; pick a new id.

        **Lineage** (`train.parent`) is an explicit input, never ambient state:
        warm-start CMA-ES from a prior champion either inline
        (`{"genome6": [6 floats]}`) or by reference
        (`{"job_id": "<prior job>"}`, refused with `422` if that job has no
        champion). Semantics: `x0 = parent genome6, sigma0 = 2.0`; absent
        parent: `x0 = zeros, sigma0 = 6.0` (cold start).

        **Exam-only certification**: `train.gens = 0` skips training entirely
        and certifies a champion you supply. Three sources (one required for
        gens=0): `train.parent` (a genome6), `policy.format=genome9` (a raw
        9-float controller), or `policy.format=python_module` (YOUR OWN policy
        as a sandboxed Python module, see the `policy` field). The exam stage
        is byte-identical either way; `dream_F` is `null` for supplied policies.

        **Bring your own architecture**: a `python_module` policy implements
        `reset(seed)` + `act(obs)` where `obs` is the same observation dict as
        the local gym (`proximity`, `collision`, ...) and the action is the
        wheel-fraction vocabulary. A policy trained in the local gym drops into
        the exam unchanged. The module runs sandboxed in the exam container
        (no network, resource- and time-limited); ship a Python module doing
        numpy inference over your exported weights, there is no ONNX detour.

        Use `train.T = 60` for `dream_F` scores comparable to campaign history;
        `T = 20` scores are not comparable.
      operationId: createJob
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/JobSpec'
            examples:
              smoke:
                summary: Smoke job (cold start)
                value:
                  job_id: my-smoke-001
                  train:
                    pop: 6
                    gens: 2
                    T: 20
                    seed: 4242
                  exam:
                    episodes: 20
              lineage:
                summary: Real-scale job warm-started from a prior champion
                value:
                  job_id: my-real-002
                  train:
                    pop: 10
                    gens: 8
                    T: 60
                    seed: 61
                    parent:
                      job_id: my-real-001
                  exam:
                    episodes: 20
              own_policy:
                summary: Certify YOUR OWN policy (sandboxed python module)
                value:
                  job_id: my-cert-001
                  train:
                    pop: 0
                    gens: 0
                    T: 0
                    seed: 0
                  exam:
                    episodes: 20
                  policy:
                    format: python_module
                    module_b64: <base64 of a policy .py or a .zip bundle>
                    entry: policy
                    class: Policy
      responses:
        '200':
          description: Job spawned (asynchronous, poll the `poll` path).
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/JobSpawned'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '409':
          description: Job id already exists (job dirs are append-only).
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '422':
          description: >
            Spec validation failed (shape, `job_id` pattern, dangling parent
            job, or `gens=0` without `train.parent`).
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
components:
  schemas:
    JobSpec:
      type: object
      required:
        - job_id
        - train
        - exam
      properties:
        job_id:
          type: string
          pattern: ^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$
        train:
          type: object
          required:
            - pop
            - gens
            - T
            - seed
          properties:
            pop:
              type: integer
              description: CMA-ES population size (ignored when gens=0).
            gens:
              type: integer
              description: >
                Number of generations. 0 = exam-only certification of the parent
                genome (train.parent then REQUIRED).
            T:
              type: integer
              description: Dream episode length (use 60 for history-comparable dream_F).
            seed:
              type: integer
              description: Training seed (bit-deterministic per platform).
            parent:
              type: object
              description: Optional lineage, exactly one of genome6 / job_id.
              properties:
                genome6:
                  type: array
                  items:
                    type: number
                  minItems: 6
                  maxItems: 6
                job_id:
                  type: string
                  pattern: ^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$
                note:
                  type: string
        exam:
          type: object
          required:
            - episodes
          properties:
            episodes:
              type: integer
              description: >-
                Episodes per world. Use the full 20, Webots startup dominates,
                reduced-episode runs are barely cheaper.
        policy:
          type: object
          description: >
            Optional. Certify a champion you supply instead of training one
            (requires train.gens=0). The policy replaces ONLY the champion arm's
            controller; world, spawns, episodes, tick, remap, scoring and the
            anti-exploit arms stay frozen.
          required:
            - format
          properties:
            format:
              type: string
              enum:
                - genome9
                - python_module
            genome9:
              type: array
              items:
                type: number
              minItems: 9
              maxItems: 9
              description: Required when format=genome9, a raw 9-float controller.
            module_b64:
              type: string
              description: >
                Required when format=python_module, base64 of a policy `.py`
                file or a `.zip` bundle, size-capped. The module implements
                reset(seed)+act(obs); obs matches the local-gym observation
                dict, action is the wheel-fraction vocabulary. Runs sandboxed
                (no network egress, resource/time-limited) in the exam
                container.
            entry:
              type: string
              description: Module name to import (default `policy`).
            class:
              type: string
              description: Policy class name (default `Policy`); a class with reset/act.
    JobSpawned:
      type: object
      required:
        - job_id
        - status
        - poll
        - log
      properties:
        job_id:
          type: string
        status:
          type: string
          const: spawned
        poll:
          type: string
          description: 'Path to poll: /jobs/{job_id}'
        log:
          type: string
          description: 'Path to tail: /jobs/{job_id}/log'
    Error:
      type: object
      properties:
        detail:
          type: string
          description: Human-readable reason.
      required:
        - detail
  responses:
    Unauthorized:
      description: Bad or missing bearer token.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: Single per-partner bearer key, provisioned by Alakazam.

````