One Turn of the Loop
The improver works one experiment at a time, and a turn of its loop is four steps.- Read the ledger.
monte statusprints the per-cohort deltas against the baseline, the Cycles with their strike verdicts, and the checkpoint tree. Checkpoints print asstep Nwith awritten by <run>column beside them; a reference to one is<run>/<N>, assembled from those two columns. - Decide one move: which checkpoint to start from, which source to train on, and which recipe to run.
- Train one chunk with those three coordinates stated in full, since Monte infers none of them.
- Score the new checkpoint on the visible exams, read the paired deltas, and start again.
monte status prints a sealed exam without its score until the claim is recorded. The run rows and monte show still print every recorded row’s score, the sealed baseline’s included, so the mask is on the summary, not the record. What the improver does state is a nomination, the checkpoint its envelope presents. The final claim on the sealed exams is made after it has stopped, by machinery it does not drive, on that nominated checkpoint and nothing else. That is what makes the claim blind: the improver picks without ever running a sealed eval of its own.
The Envelope
An envelope is the approved terms of one search, fixed before it starts and never renegotiated while it runs. It is the unit of approval and the unit of audit, so every action the improver takes is judged against the envelope it ran under.
The claim field is one bit, decided at approval. It is the difference between a search that ends in a number and one that does not. A closing envelope ends with the blind final claim, which spends each sealed exam’s last touch. An open envelope ends with a sync alone, and those touches stay unspent for a later envelope to finish the work. Neither one is decided mid-flight.
The box is not part of the envelope’s lifetime. One GPU box is acquired for the search program and held. An envelope is an approved window of time on that box, not a box that lives and dies with the search. That separation exists because GPU capacity is genuinely scarce. Making every approval acquire its own box turns each one into a gamble, and pays the setup cost again each time.
What the Platform Refuses
Three refusals are live for as long as an envelope is active, and they hold regardless of which command asked or which harness ran it.- An eval of a sealed exam is refused, in every spelling,
--smokeincluded, and under every experiment name on the host — a fresh experiment that shows the same exam is no way around the seal. A smoke run skips the two-touch budget, so it is the one spelling that would otherwise work. monte promoteis refused. Promotion is a human act on audited evidence, and it waits for the search to close.- No new run starts past the window’s close. The window is the approved spend, so a run started after it is spend nobody approved.
Three Layers, Not One
The refusals above are the middle layer of three, and there are three because the outer two are known to be incomplete. The harness guard reads each shell command before it runs and blocks the ones the search may not make. It sees through anssh invocation to the command quoted inside it. It is a belt, and its limits are written down. A command that hides an exam behind an environment variable goes past it. So does one that writes a script and then runs it.
Monte itself is the wall, and it is where the three refusals above live. They fire inside the platform on the terms of the envelope record, so what the harness misses the platform still refuses.
The launcher is the clock. It kills the improver’s process at the moment the window closes, and a run still in flight at that instant dies as an aborted run. Finishing the current run first would make the ceiling soft, so nothing is given that grace.
Underneath all three, the account the improver runs as on the box is fenced. It has a full login shell and the whole monte CLI, and it has neither sudo nor any credential. That combination is deliberate: an account with sudo could read the key that spends money. It could also stop the watchdog that ends the search, or rewrite the terms of its own envelope. This one can read its envelope and can never write it.
How a Search Ends
After the window closes, a finalizer runs a fixed sequence on the box.- Stop the improver. Nothing the search could still influence happens while a move is in flight.
- Mark the envelope closed. This is what releases the
testand window refusals, so the claim in the next step is not refused by its own envelope. It is also the one step that stops the sequence if it fails, because everything after it would run into the guards it was supposed to release. - Make the blind claim, on closing envelopes only, as an eval of the nominated checkpoint on each sealed exam. A search that nominated nothing presents nothing: the claim and the push are skipped, the sealed touches stay unspent, and the envelope still ends.
- Push the claimed checkpoint’s weights to Hugging Face, where weights live.
- Sync the results tree to the durable tier, which is what survives the box.
- Sync the improver’s transcript, as its own step rather than a tail on the one above: the two trees have independent value and independent failure modes, so a results push that fails must not be the reason the reasoning is lost.
- Release the box to idle, rather than terminating it, since the box outlives the envelope.
Reading and Ending an Envelope
list prints every envelope record this host can see, each with its state, its claim bit, its experiment, and its window. show prints one in full: the approved terms exactly as they were approved, plus the state Monte keeps alongside them.
nominate records which checkpoint this envelope presents. It is the improver’s verb: the last nomination stands, re-nominating replaces it, and the nominee needs an admissible eval, because a presented model is measured, never merely liked. The finalizer claims and pushes that checkpoint after the search stops.
finalize runs the sequence above. It belongs to the box watchdog rather than to a person, and it is idempotent, because that watchdog will call it more than once. An already-closed envelope is reported and not re-finalized, so a sealed exam’s last touch cannot be spent twice.
close is the escape hatch, and it takes --confirm <envelope-id> typed out. It marks an envelope consumed without finalizing it, which releases the refusals and skips the blind claim entirely. Reach for it when a finalizer ran on another host, or never ran at all, and an envelope is holding the sealed exams hostage.
What Is Not Here Yet
One envelope runs at a time, on one box, and its search is serial: one run, then the next. Several boxes comparing recipes concurrently under a single envelope is a planned expansion rather than an option available today. The design leaves room for it, but the platform has not shipped it.Related Pages
Improvement
The loop the improver drives.
Experiment
What sealed and visible exams are for.
Glossary
The vocabulary this page uses.

