How training works

Every coin is a small GPT: 4 layers, 4 heads, 128 dimensions, context 64 tokens, a 1,024-token byte-level BPE vocabulary of its own, about 1 million parameters. AdamW at 1e-3 with 100 warmup steps, cosine down to 1e-4 at step 6,000, gradient clip 1.0, batch 12. This is the CPU configuration from the nanoGPT README.

The CPU lane runs 2 jobs at once. A job is a 60 second slice for one model, resumed from its last checkpoint (weights, optimizer state, step and random state). Every 5 minutes every live coin gets at least one slice when the lane has room. Order when it does not: coins with no checkpoint yet (newborns train within their first minute), then ABOUT TO SPEAK coins, then coins whose launcher holds $TRAIN, then the rest by oldest slice. Lab runs and the mother model use leftover capacity; the mother is guaranteed 1 slice per 3 periods.

Measured speed on this box and the measured time to first words are in VERIFY.md. The trainer pauses below 5 GB of free disk.

What you see

  • Loss: training loss every 10 steps, streamed about once every 2 seconds. Curves show a 10-point moving average; the raw series is at /api/coin/<id>/series.
  • Validation: loss and bits per byte on the fixed validation split every 200 steps.
  • Samples: 48 tokens every 30 s at temperature 0.8, top-k 40, filtered before anyone sees them.
  • Checkpoints: written at every evaluation as bf16 safetensors with a sha256.