Skip to main content
An evaluation runs a model on a set of tasks, scores what comes back, and reports how the model performed. The result shows what is working, what is not, and where the next effort should go. Fine-tuning a model on new data, or adding a tool to its agent harness, is meant to make it do the work better. The change on its own does not confirm that. An evaluation settles it objectively: it puts a number on the model’s performance, so the change can be called an improvement or not, and by how much. That only holds if the two numbers were produced the same way. If the tasks, the sampling rules, or the grader changed between them, the difference could have come from any of those rather than from the model, and the comparison no longer says anything about the change itself. A standardized evaluation removes that doubt, which is why every evaluation in Monte runs against a fixed Experiment.

Model Capability or Agent Capability

A score covers the model and the agent harness together. The harness picks the tools, holds the prompt, and decides when a task is done. Each of those moves the number. To read one of them, hold the other still.
  • Hold the harness, change the model. The score reads model capability.
  • Hold the model, change the harness. The score reads harness quality. A new tool can raise accuracy and triple the token cost at the same time.
An experiment freezes the harness in each exam’s fingerprint. Two runs that score the same exam in the same cohort used the same harness, so the difference between their scores belongs to the model. Two runs on different exams share no such guarantee: each exam carries its own tasks and its own eval settings, and no comparison crosses between them. Changing the harness changes the fingerprint, which makes a different experiment. Monte does not score across experiments, so it does not compare harnesses today.

What Makes an Evaluation Good

Coverage. The tasks have to look like the work that matters. A narrow set of tasks gives a high score that says little about real use. The grader. The grader defines what good means for the Environment. A noisy or biased grader scores the wrong thing, and training then optimizes for it. Programmatic checks such as an exact match, a test suite, or code that runs are steadier than a model that judges a model. Representative tasks. Public benchmarks allow comparison against published results. Bespoke tasks show whether the model works for the intended use. Most useful environments have both. Repeats. One attempt per task is noisy. More attempts separate a real gain from variance. The experiment fixes the number of attempts in each exam’s eval settings, so it cannot drift between runs. Reproducibility. The same run has to be scored under the same conditions every time. The fingerprint hashes the environment’s frozen material: the tasks, the grader, and the harness configuration. Monte refuses a run when any of it changed.

Environments

The graded world, and how one rollout runs.

Experiment

What gets frozen before a model is scored.

Improvement

What the score is used for once it exists.

Further Reading