Model Capability or Agent Capability
A score covers the model and the agent harness together. The harness picks the tools, holds the prompt, and decides when a task is done. Each of those moves the number. To read one of them, hold the other still.- Hold the harness, change the model. The score reads model capability.
- Hold the model, change the harness. The score reads harness quality. A new tool can raise accuracy and triple the token cost at the same time.
What Makes an Evaluation Good
Coverage. The tasks have to look like the work that matters. A narrow set of tasks gives a high score that says little about real use. The grader. The grader defines what good means for the Environment. A noisy or biased grader scores the wrong thing, and training then optimizes for it. Programmatic checks such as an exact match, a test suite, or code that runs are steadier than a model that judges a model. Representative tasks. Public benchmarks allow comparison against published results. Bespoke tasks show whether the model works for the intended use. Most useful environments have both. Repeats. One attempt per task is noisy. More attempts separate a real gain from variance. The experiment fixes the number of attempts in each exam’s eval settings, so it cannot drift between runs. Reproducibility. The same run has to be scored under the same conditions every time. The fingerprint hashes the environment’s frozen material: the tasks, the grader, and the harness configuration. Monte refuses a run when any of it changed.Related Pages
Environments
The graded world, and how one rollout runs.
Experiment
What gets frozen before a model is scored.
Improvement
What the score is used for once it exists.
Further Reading
- Evaluation (NeMo Gym)
- LLM evaluation: Why testing AI models matters (IBM)

