Skip to main content
Monte is in closed beta. Interfaces and these docs change quickly, and commands can break between versions.
Off a GPU box, every run is a mock run: the real CLI, a real ledger row, and a fake score instead of real execution. That makes a laptop a safe place to learn the loop. This walkthrough needs uv and read access to Monte-Inc/internal-monte-platform. uv fetches its own Python, and the repo is never cloned.
1

Install the CLI

This puts monte on the PATH, so every command below is bare monte.The --with line adds the demo environment. Monte ships no environments of its own: each one is a separate package, and installing it is what registers it. Keep the quotes around that URL.If the shell cannot find monte, run uv tool update-shell and open a new shell.
2

List the Installed Environments

demo appears as a curriculum, and demo-a through demo-d as exams, all mock. A curriculum holds training tasks, graded live by its own grader as the model attempts them: that is the material training runs on. An exam holds a fixed pile of tasks to score against. Mock means they can be scored without a GPU box, which is why every run here is a mock run. The demo package exists to exercise the loop itself.
3

Create an Experiment

This freezes the exam set: each exam’s fingerprint, its eval settings, and its role. --visible names the exams the loop scores freely, and --sealed names the ones it can never evaluate, held back for the final claim. Which exam is sealed is the experiment’s decision, not the environments’. Every score in this experiment comes from those exams’ graders. The name hello is arbitrary.Nothing about training is frozen here. The method and the data are stated on each run, in the next steps.
4

Run the Baseline Eval

The CLI prints the run plan and asks launch?. The mock run takes about 30 seconds. With exactly one visible exam, --exam defaults to it, so this scores demo-a and records its baseline, the score every later run is compared against.
5

Take the Sealed Baseline

Run this now, before any training. A sealed exam is never a default, so the baseline eval above did not touch demo-b: it must be named deliberately.A sealed exam is read exactly twice per cohort, and a cohort is the set of runs that share one base model and one comparability config. Every run in this walkthrough is in the same cohort. The first read is this one, on the base model. The second is the final claim after training. Having both reads is what allows the improvement to be stated on tasks the loop never saw.Every exam needs its baseline before training, not just the sealed one. Skip this step and monte train refuses with E_NO_BASELINE, naming demo-b and this exact command.
6

Rehearse with a Smoke Chunk

A chunk is the fixed unit monte train trains: 50 steps by default, and 3 steps for a smoke. A step count is never stated.Three things travel with every training run, and Monte infers none of them. --recipe is the method: grpo_math_1B ships with the CLI, and monte config lists the rest. --source is what the run trains on, here the demo curriculum. The third is the starting point, and --smoke is the one case that states none: a smoke rehearses the from-scratch path.A smoke run is a capped sanity pass. It is never evidence: it cannot be a baseline, a parent, or a promote source.
7

Train a Real Chunk

Now the starting point is stated. --from-base opens a fresh branch from the base model at step 0. The other way to start is --parent <run>/<step>, which continues an existing checkpoint. monte status prints the exact spelling.This chunk is 50 steps and it does count as evidence.
8

Score the New Checkpoint

With no target flag, an eval scores the cohort’s newest checkpoint on the one visible exam. That is the chunk just trained, so this row lands as eval@50.
9

Read the Ledger

The output holds the run rows, one block per cohort with a line per exam, the lineage: lines, the cycles: line, and the checkpoints: list. Smoke rows carry a [not evidence] marker.The cohort block reads like this:
Take it line by line.cohort Qwen/Qwen2.5-1.5B-Instruct@mock-m1 names the base model and the revision it is pinned to, and 2 exam(s) anchored says both exams have their baselines. One line follows per exam. demo-a: 0.6667 -> 1.0 is the visible exam’s score at the baseline and at the newest eval, and eval@50 says which eval produced that newest score.delta +0.3333 is the difference between those two scores. paired, 2 fixed / 0 broken says how Monte checked it: both evals scored the same frozen tasks, so it compared them task by task. Two tasks the baseline got wrong came out right, none went the other way, and tasks that came out the same way both times cancel. ± 0.3772 is the interval around the delta. Pairing changes that interval, not the delta itself.demo-b (sealed) is the cohort line that tracks the seal: the baseline is recorded and one of the two touches is spent, and this line prints no score until the claim is recorded. The run rows above it still print every row’s score, the sealed baseline’s included.lineage: prints the lineage, the path this checkpoint took back to the base model. Each arrow is one training run, labelled with the move it made: the algorithm, then the source it trained on. Here one run took the base model to step 50, training with grpo on the demo source.cycles: prints each Cycle with one verdict per visible exam, and a Cycle is one turn of the loop: the first eval that counts as evidence at a new checkpoint. STRIKE means that exam’s paired delta is not significantly positive: 0.3333 minus 0.3772 is below zero, so the interval reaches under zero and the gain could be noise.(1 strike — demo-a: 1) tallies the strikes so far per exam, not a countdown. A strike is a label to read. Nothing adds strikes up into a decision, nothing stops the loop at some number of them, and when to stop is an operator decision.presents: says which checkpoint this search puts forward for the final claim. Nothing stands here because nothing was nominated: a checkpoint is presented by a nomination, and no rule picks a winner from the scores in its place.

Stay Current

The beta changes quickly. Pull the newest CLI with:
To remove it, run uv tool uninstall monte.
A real score needs a GPU box, and running on one needs a checkout of the platform repo rather than this tool install. Those instructions are being rewritten and are not on this site yet.