Why an Experiment Exists
Evaluation has many variables to set. Each task can get one attempt or several, the decoding temperature can change, a different subset of the tasks can run, or answers can be marked more or less strictly. Each of those returns a different number for the same model, and none of them is wrong. They answer different questions about how well the model does the task. An experiment holds those variables constant, so within one experiment the model is the only variable that changed between two runs. Any number of experiments can be defined. When the variables go unrecorded, old results become unreliable. Nothing explains why two numbers differ, so the experiments that produced them cannot be compared.What an Experiment Fixes
An experiment fixes three things at creation.- The exams. Which exam environments run, each pinned by its fingerprint: an exact task list, not a count and not a name.
- The eval settings. How the model is sampled: the decoding, the temperature, and how many attempts each task gets. The settings belong to each exam, so two exams in one experiment can differ.
- The grading. Which grader scores the answers, and the prompt the model saw.
Sealed and Visible
An experiment assigns each of its exams one of two roles at init, and that assignment is the only decision it makes. Which pile a model trains on is not part of it: the training data is stated per run.- Visible. The driving agent evaluates these freely, every cycle, and reasons over all of them. More visible exams is more information: nothing ranks them, and nothing merges their scores into one number.
- Sealed. The driving agent can never evaluate these. Each sealed exam is touched twice per cohort: the baseline, and the final claim.
Related Pages
Environments
The graded world an experiment freezes.
Improvement
How baselines and deltas judge improvement.
Glossary
The record every run writes to.

