Skip to main content
This is the first agentic loop Monte runs, and it is deliberately the simplest one that works. One search, one box, one move at a time, in sequence. Everything harder — several arms compared at once, more than one box, a search that spans experiments — is a later step, and the shape here is what those steps build on. The other pages call the thing that works the loop a driving agent. The improver is Monte’s own: an outer-loop agent that reads an experiment’s ledger, decides the next Move, runs it, and reads what came back. It runs unattended for a fixed stretch of time on its own control plane, renting the GPU boxes its runs execute on and driving them from there. That makes the interesting question not how it decides, but what it is allowed to do while nobody is watching. The answer is the envelope, and the platform enforces it rather than the prompt.

One Turn of the Loop

The improver works one experiment at a time, and a turn of its loop is four steps.
  1. Read the ledger. monte status prints the per-cohort deltas against the baseline, the Cycles with their strike verdicts, and the checkpoint tree. Checkpoints print as step N with a written by <run> column beside them; a reference to one is <run>/<N>, assembled from those two columns.
  2. Decide one move: which checkpoint to start from, which source to train on, and which recipe to run.
  3. Train one chunk with those three coordinates stated in full, since Monte infers none of them.
  4. Score the new checkpoint on the visible exams, read the paired deltas, and start again.
The move is where the search happens, and it takes four shapes. Continuing an arm keeps the recipe and starts from that arm’s newest checkpoint. Branching keeps the recipe and starts from an earlier checkpoint. That is the move when a recipe looked healthy up to a point and then stopped. A fresh arm starts from the base model under any recipe. A stage boundary continues a checkpoint under a different recipe, so the weights carry over while the trainer’s step counter restarts. That is how a supervised stage becomes the foundation for a reinforcement learning one. Writing recipes is part of the work rather than a setup step that happens first. The improver reads what the last move scored. It can respond by authoring a new recipe, not only by picking a different one off the shelf. A new recipe is a new named file rather than an edit to an existing one. A recipe is pinned by the hash of its content, so editing one in place changes its identity. The next run on that arm then becomes a stage boundary instead of the continuation it was meant to be. Every one of those decisions is read off the visible exams. The improver sees every visible score and reasons over all of them — more visible exams is more information, and nothing merges them into one number. It never produces a sealed number: an eval of a sealed exam is refused while its envelope is active, and the per-cohort line in monte status prints a sealed exam without its score until the claim is recorded. The run rows and monte show still print every recorded row’s score, the sealed baseline’s included, so the mask is on the summary, not the record. What the improver does state is a nomination, the checkpoint its envelope presents. The final claim on the sealed exams is made after it has stopped, by machinery it does not drive, on that nominated checkpoint and nothing else. That is what makes the claim blind: the improver picks without ever running a sealed eval of its own.

The Envelope

An envelope is the approved terms of one search, fixed before it starts and never renegotiated while it runs. It is the unit of approval and the unit of audit, so every action the improver takes is judged against the envelope it ran under. The claim field is one bit, decided at approval. It is the difference between a search that ends in a number and one that does not. A closing envelope ends with the blind final claim, which spends each sealed exam’s last touch. An open envelope ends with a sync alone, and those touches stay unspent for a later envelope to finish the work. Neither one is decided mid-flight. The box is not part of the envelope’s lifetime. One GPU box is acquired for the search program and held. An envelope is an approved window of time on that box, not a box that lives and dies with the search. That separation exists because GPU capacity is genuinely scarce. Making every approval acquire its own box turns each one into a gamble, and pays the setup cost again each time.

What the Platform Refuses

Three refusals are live for as long as an envelope is active, and they hold regardless of which command asked or which harness ran it.
  • An eval of a sealed exam is refused, in every spelling, --smoke included, and under every experiment name on the host — a fresh experiment that shows the same exam is no way around the seal. A smoke run skips the two-touch budget, so it is the one spelling that would otherwise work.
  • monte promote is refused. Promotion is a human act on audited evidence, and it waits for the search to close.
  • No new run starts past the window’s close. The window is the approved spend, so a run started after it is spend nobody approved.
Each of those is keyed to one envelope’s id. The promote and window refusals bind only the experiment the envelope names; the sealed refusal is deliberately wider, covering the exams that experiment seals under any name on this host. There is still no global mode: a seal is what this search cannot see, never a claim on the exam everywhere, so the same exam may be sealed here and visible in an unrelated experiment, and a second search on a second box is never refused by this one. Active means the record says active, and it does not mean the window is still open. An envelope whose time has run out stays active until the finalizer marks it closed. A finalizer that never ran therefore leaves the sealed exams refused rather than quietly available. Closing the record by hand is the way out of that, and it is a deliberate act with a typed confirmation.

Three Layers, Not One

The refusals above are the middle layer of three, and there are three because the outer two are known to be incomplete. The harness guard reads each shell command before it runs and blocks the ones the search may not make. It sees through an ssh invocation to the command quoted inside it. It is a belt, and its limits are written down. A command that hides an exam behind an environment variable goes past it. So does one that writes a script and then runs it. Monte itself is the wall, and it is where the three refusals above live. They fire inside the platform on the terms of the envelope record, so what the harness misses the platform still refuses. The launcher is the clock. It kills the improver’s process at the moment the window closes, and a run still in flight at that instant dies as an aborted run. Finishing the current run first would make the ceiling soft, so nothing is given that grace. Underneath all three, the account the improver runs as on the box is fenced. It has a full login shell and the whole monte CLI, and it has neither sudo nor any credential. That combination is deliberate: an account with sudo could read the key that spends money. It could also stop the watchdog that ends the search, or rewrite the terms of its own envelope. This one can read its envelope and can never write it.

How a Search Ends

After the window closes, a finalizer runs a fixed sequence on the box.
  1. Stop the improver. Nothing the search could still influence happens while a move is in flight.
  2. Mark the envelope closed. This is what releases the test and window refusals, so the claim in the next step is not refused by its own envelope. It is also the one step that stops the sequence if it fails, because everything after it would run into the guards it was supposed to release.
  3. Make the blind claim, on closing envelopes only, as an eval of the nominated checkpoint on each sealed exam. A search that nominated nothing presents nothing: the claim and the push are skipped, the sealed touches stay unspent, and the envelope still ends.
  4. Push the claimed checkpoint’s weights to Hugging Face, where weights live.
  5. Sync the results tree to the durable tier, which is what survives the box.
  6. Sync the improver’s transcript, as its own step rather than a tail on the one above: the two trees have independent value and independent failure modes, so a results push that fails must not be the reason the reasoning is lost.
  7. Release the box to idle, rather than terminating it, since the box outlives the envelope.
A step that refuses or fails is recorded and the sequence continues past it. A failed claim should not also cost the search its sync, and an unreleased box would stay out of normal idle reaping. The transcript is synced because it and the ledger answer different questions. The ledger records what was tried and what each attempt scored. The transcript records why the improver moved as it did, which is what a later search reads before deciding where to pick up.

Reading and Ending an Envelope

list prints every envelope record this host can see, each with its state, its claim bit, its experiment, and its window. show prints one in full: the approved terms exactly as they were approved, plus the state Monte keeps alongside them. nominate records which checkpoint this envelope presents. It is the improver’s verb: the last nomination stands, re-nominating replaces it, and the nominee needs an admissible eval, because a presented model is measured, never merely liked. The finalizer claims and pushes that checkpoint after the search stops. finalize runs the sequence above. It belongs to the box watchdog rather than to a person, and it is idempotent, because that watchdog will call it more than once. An already-closed envelope is reported and not re-finalized, so a sealed exam’s last touch cannot be spent twice. close is the escape hatch, and it takes --confirm <envelope-id> typed out. It marks an envelope consumed without finalizing it, which releases the refusals and skips the blind claim entirely. Reach for it when a finalizer ran on another host, or never ran at all, and an envelope is holding the sealed exams hostage.

What Is Not Here Yet

One envelope runs at a time, on one box, and its search is serial: one run, then the next. Several boxes comparing recipes concurrently under a single envelope is a planned expansion rather than an option available today. The design leaves room for it, but the platform has not shipped it.

Improvement

The loop the improver drives.

Experiment

What sealed and visible exams are for.

Glossary

The vocabulary this page uses.