The Task World
| Term | Meaning |
|---|---|
| Environment | The graded world an agent acts on: the tasks, the grader, and any tools and per-task state. See Environments. |
| Task | One problem for the model, plus the data needed to score an attempt at it. |
| Curriculum | A curriculum environment’s training tasks, scored live by that environment’s grader as the model attempts them. A frozen dataset is the other thing a run can train on, and neither one is ever called by the other’s name. |
| Exam | An environment holding one published benchmark’s tasks, frozen as bytes. An experiment names its exams at init and assigns each one a role: sealed or visible. |
| Rollout | One attempt at one task, from prompt through to a graded reward. |
| Grader | Scores one attempt into a reward, usually between 0 and 1. It defines what “good” means for its environment. |
| Split | A frozen list of task IDs: the one pile an environment holds. An exam environment holds its benchmark’s tasks, and a curriculum environment holds its training pile. The roles that used to ride on split names moved onto whole exams as sealed and visible, and test is not a split name anywhere. |
What Gets Frozen
| Term | Meaning |
|---|---|
| Experiment | A frozen evaluation contract over a set of exams: which the improver reads freely (visible) and which it can never evaluate (sealed), each pinned with its own eval settings. See Experiment. |
| Sealed / visible | The two roles an experiment assigns its exams at init, and its only decision. A sealed exam is one the improver can never evaluate: it is touched twice per cohort, at the baseline and at the final claim. A visible exam is read freely, every cycle; the improver sees every visible score and reasons over them, and nothing ranks or aggregates across them. The roles are this experiment’s alone: the same exam may be sealed here and visible in an unrelated experiment. |
| Fingerprint | A hash over an environment’s frozen material: the tasks, the grader, the harness configuration, and for a curriculum environment its training data. A run is refused when any of it has changed, and every one of an experiment’s exams is recomputed before every run, not just the exam in hand. |
| Comparability config | What is hashed to decide whether two runs can be compared: the whole exam set, sorted by name, with each exam’s fingerprint, its eval settings, and its role inside. The base model stays outside it, because it is the other half of the cohort key. Train settings stay outside it too, because a baseline is taken before any training and no training setting can affect it. |
| Dataset | A frozen, named pile of demonstration rows for supervised fine-tuning. Any overlap with frozen exams is recorded when it is frozen. Never part of an experiment. Monte uses this word for that frozen pile only: an environment’s own task material is never called a dataset. |
| Materialization | A local copy of something whose truth lives elsewhere, proven against its pin rather than trusted. A data root holds materializations of frozen inputs. Not a Mirror: a Mirror is results, and it can be stale, while a materialization either matches its pin or is refused. |
| Mirror | The read-only local copy of an experiment’s results that monte sync pulls from the results bucket, so ledger state reads with no box up. Weights never enter one, and neither do frozen inputs. |
| Vault | The cloud bucket a freeze pushes to. It is the copy Monte treats as true for what reached it, and a data root holds a materialization of that. monte env install fetches an environment back, and monte dataset show fetches a dataset. What never reaches it: an environment shipped as a package, a curriculum’s rows (the Vault holds the recipe that re-derives them), and an add on a machine with no key that can write the Vault, which freezes locally and says so. See Environments. |
| Baseline | The base model’s score on one exam before any training. Every exam, sealed and visible, needs its own baseline before the first training run in a cohort, because nothing can score the base model once a checkpoint exists. A sealed exam’s baseline is taken deliberately, with monte eval <experiment> --exam <name>, and it spends the first of that exam’s two touches. Every later score on that exam in the same cohort is read against it. |
What Runs
| Term | Meaning |
|---|---|
| Run | One execution recorded as a row. Either a training run, which produces a checkpoint, or an eval run, which scores one. |
train@N / eval@N | N is the checkpoint’s absolute step count, never a run index. The training run ending at step 50 is train@50, and eval@50 scores what it produced. |
| Checkpoint | The model’s saved state at one step. |
| Smoke Run | A capped rehearsal run, started with --smoke, that never counts as evidence: it cannot be a baseline, a parent, or a promote source, and it never spends a sealed exam’s touch. A smoke train runs 3 steps. Its row still stays in the ledger. |
| Mock Run | A run that goes through the real CLI and writes a real ledger row, with a fake score in place of real execution. Off a GPU box every run is a mock run. On a machine that has a GPU but no stack installed, Monte refuses rather than record a mock score silently. |
| Preflight | The checks Monte runs before a run starts, such as the lock probe, the environment fingerprint recompute, and free disk. A failed preflight refuses the run with E_PREFLIGHT (exit code 6). |
| Servable folder | The weights and tokenizer directory an inference server can load, in Hugging Face layout. |
| Load-check | One prompt pushed through a real serving stack to prove a servable folder actually serves. It is not a score and not a run. |
| Recipe | The named file stating how a training run trains: model shape, algorithm, the batch settings, optimizer. See Training. |
| Source | What a training run trains on, stated on every train: a curriculum environment (its training tasks, graded live), or a frozen dataset under an sft recipe. Recorded on the row. |
| Move | One decision: which checkpoint to start from, which source to train on, and which recipe to run. monte status renders it in each hop label as algorithm:source. |
| Improver | The outer-loop agent that drives a search: it reads the ledger, decides the next move, runs it, and reads what came back. It decides from the visible exams alone, and never sees a sealed number its own search produced. See Core Loop. |
| Envelope | The approved terms of one improver search, fixed before it starts: the experiment, the goal, the time window, the spend ceiling, the box, and whether the search is closing (it ends with the blind final claim) or open (it ends with a sync, and the sealed exams’ final touches stay unspent). While one is active, Monte refuses every eval of an exam its experiment seals — under any experiment name on the host — along with monte promote, and starts no new run past the window. |
| Run plan | The complete merged settings of one run, resolved before launch, printed for confirmation, and stored alongside the run. The approved plan is what executes. |
What Is Derived
| Term | Meaning |
|---|---|
| Cohort | The runs sharing a base model and comparability config: the set within which a number means something. Each one needs its own baseline, and a delta never crosses from one cohort into another. |
| Lineage | A checkpoint’s descent path back to the base model, with every hop labelled by its move: the algorithm and the source it trained on. Checkpoints form a tree, so one cohort can hold several Lineages. See Improvement. |
| Hop | One arrow in a lineage: the training run that took one checkpoint to the next. |
| Stage boundary | A hop where the recipe changes. The new run starts from the parent checkpoint’s weights alone, the trainer’s step counter restarts there, the KL anchor moves to that checkpoint, and monte status draws the hop with a double-shaft arrow. |
| Delta | The difference between a run’s score and its baseline’s score, inside one cohort. |
| Paired | Both evals scored the same frozen tasks, so Monte compares them task by task and counts the flips. fixed is a task the baseline got wrong and this run got right, broken is the reverse, and tasks that came out the same way both times cancel. The delta is (fixed - broken) / n, the same number the two scores give on their own. Pairing changes only the interval printed beside it, which is usually, though not always, tighter. |
| Cycle | One turn of the loop, printed by monte status under cycles:: the first look at a new checkpoint, on whichever visible exam it lands. Later looks at the same checkpoint on other visible exams attach their verdicts to it, a retry at an exam already judged there judges nothing, and a sealed read never opens or joins one. Two branches off one checkpoint each get their own. |
| Strike | A Cycle’s verdict on one visible exam, printed as STRIKE, meaning that exam’s paired delta against its baseline is not significantly positive. A significant regression counts as a strike too, since it is no improvement either. Strikes are tallied per exam and are labels to read: nothing adds them up into a stop, and when the loop ends is an operator decision. |
| Nomination | The improver stating which checkpoint its envelope presents, with monte envelope nominate. Re-nominating replaces it, the nominee needs an admissible eval, and an envelope that never nominates presents nothing: no rule picks a winner from the scores in its place. |
| The claim | The sealed exams’ scores of the checkpoint the envelope’s nomination presents. Each sealed exam is read exactly twice per cohort, and the first read is taken on the base model before any training. Skip it and Monte refuses to train at all, because the claim in that cohort could never be anchored. |
| Promoting | Exporting a checkpoint’s servable folder as a checksum-verified copy, with the provenance of the run that wrote it, behind a confirmation typed out in full. Monte checks that the run is a succeeded training run that counts as evidence. It does not check that the checkpoint carries a claim. |
The Record
| Term | Meaning |
|---|---|
| Ledger | The append-only, per-experiment record of runs: what happened, in what order, with what score. Everything that reads it goes through one API rather than touching the files. |
| Trace | One run’s attempt at one task: the rollout verbatim, plus the grader’s own fields. |
| Provenance | The hashes carried on every row, covering the code, the configuration, the exam’s frozen task list, the comparability config, and the container image. |
Related Pages
Improvement
What the rows mean once they exist.
CLI
The commands that write and read them.
Further Reading
- Key Terminology (NeMo Gym)
- NeMo RL Documentation (NVIDIA)

