Skip to main content
An experiment is a fixed standard for scoring a model. Two runs under the same experiment answered the same tasks under the same rules, so they can be compared directly. An experiment is built on a set of exam Environments, and it fixes which exams run, each exam’s sampling rules, and what counts as correct.

Why an Experiment Exists

Evaluation has many variables to set. Each task can get one attempt or several, the decoding temperature can change, a different subset of the tasks can run, or answers can be marked more or less strictly. Each of those returns a different number for the same model, and none of them is wrong. They answer different questions about how well the model does the task. An experiment holds those variables constant, so within one experiment the model is the only variable that changed between two runs. Any number of experiments can be defined. When the variables go unrecorded, old results become unreliable. Nothing explains why two numbers differ, so the experiments that produced them cannot be compared.

What an Experiment Fixes

An experiment fixes three things at creation.
  1. The exams. Which exam environments run, each pinned by its fingerprint: an exact task list, not a count and not a name.
  2. The eval settings. How the model is sampled: the decoding, the temperature, and how many attempts each task gets. The settings belong to each exam, so two exams in one experiment can differ.
  3. The grading. Which grader scores the answers, and the prompt the model saw.
An experiment fixes the evaluation only, not the training. The Recipe stays free, and so does the source, what a run trains on, because those are what the driving agent changes between runs. An experiment names no model either: each run binds its own, so many models can compete under one experiment. Monte checks the frozen material before every run. If any of it changed, that is a different experiment, and Monte refuses to score the new run against the old ones.

Sealed and Visible

An experiment assigns each of its exams one of two roles at init, and that assignment is the only decision it makes. Which pile a model trains on is not part of it: the training data is stated per run.
  • Visible. The driving agent evaluates these freely, every cycle, and reasons over all of them. More visible exams is more information: nothing ranks them, and nothing merges their scores into one number.
  • Sealed. The driving agent can never evaluate these. Each sealed exam is touched twice per cohort: the baseline, and the final claim.
The driving agent picks each next move from what the visible exams show, so the loop fits itself to those tasks over time. A visible score therefore overstates the model, which is why it is not the number to report. A sealed exam stays out of the loop entirely, and it gets two reads: the first is taken deliberately, as an eval of the base model before any training begins, and the second is the final claim, after which Monte refuses a third. Every exam, sealed and visible, needs that first read before training starts. Skip one and Monte refuses to train, because nothing can score the base model once a checkpoint exists. Improvement covers how those reads happen.

Environments

The graded world an experiment freezes.

Improvement

How baselines and deltas judge improvement.

Glossary

The record every run writes to.