- Exam roles. The improvement loop optimizes against the exams it can see, so a visible exam’s score cannot be the number to report, and a sealed exam is held back for that. The roles live on the Experiment, never here: an environment makes no exam decision.
- A fingerprint. The tasks, the grader, and the harness configuration each change a score without looking like a change to the experiment. The fingerprint hashes the environment’s frozen material, for a curriculum environment its training data included, and refuses a run when any of it changes.
Parts of an Environment
An environment has three parts, and the agent harness sits beside them. Task material. The tasks, plus the data needed to score each attempt. One task is one problem for the model: a coding issue to fix, a tool-calling scenario to work through, a support conversation to resolve. The prompt config that renders a task into a model call belongs here too. An environment carries exactly one pile: an exam holds a published benchmark, and a curriculum holds a training pile, never both. Grader. It scores one attempt into a reward, usually between 0 and 1, by checking the output against the task’s own data: an expected answer, a test suite, or a rubric. It defines what “good” means for this environment, and it serves two jobs at once, because its scores become both the eval metric and the training reward. Tools and state. The per-task world that the agent changes as it acts. Each attempt starts from a clean state, so a reward belongs to that attempt alone. State ranges from none at all to a file system, a database, or a repository that the agent edits. Environments are single-turn today: one question in, one answer out, no tools and no state. Agent harness. The harness works with an environment, but the environment does not own it. A model on its own does stateless inference, and the harness is what makes it an agent: it loops model calls, routes tool use, manages context, and decides when a task is done. A simple harness only loops until the task completes, while others add planning, memory, and skills. Each run states which model it uses. Every environment must have a grader, and every environment names its domain: Gym’s owndomain: field, transcribed verbatim as free text. Ungraded rows enter Monte only one way: frozen as a dataset and stated as an sft run’s source, never as part of an environment or an experiment. One environment can serve several experiments: the same tasks, frozen under different settings, sealed in one and visible in another.
Any environment from the NeMo Gym catalog can be measured, or a new one written as an installable package.
Where a Frozen Input Lives
A frozen environment lives in the Vault, a cloud bucket that outlives any box, and so does a frozen dataset. That copy is the one Monte treats as true, and what follows holds for both kinds. The copy on a data root is a materialization of it.monte env add freezes the files, hashes them, and pushes them to the Vault. monte env install fetches them back and proves every hash before it writes them, which is what makes a local copy safe to discard and fetch again. A curriculum environment’s rows are the one exception in what travels: the Vault holds the recipe that re-derives them — the resolved server config, the pinned upstream revisions, and the grader’s code tree — and install re-derives the rows and holds them to the sha the freeze recorded. A dataset works the same way, with monte dataset show as the command that fetches it. Losing a box, or the disk its data root sits on, therefore costs no frozen input that reached the Vault. Freezing is the only step that writes there, and after it the flow runs one way, from the Vault to a data root to the Gym checkout, so no copy can change the one above it.
Whether an input reached the Vault depends on credentials, and the push and the fetch do not resolve the same ones. Writing goes through the read-write key, reading through the ordinary one, and the monte sync, push section names both. A freeze that finds no key it can write with still freezes on the data root alone, and says so in a note. An environment shipped as an installable package is never pushed at all, since the push belongs to monte env add rather than to its install(). Those copies are held locally, and reading monte env list --vault against a bare monte env list is what identifies them.
The reason any of this matters is the fingerprint. Freezing the same tasks a second time runs Gym’s prepare again, and prepare reads live sources, so the files come back different. Different files are a different fingerprint, which is a different experiment, and every baseline taken under the old one stops meaning anything. The bytes cannot be made a second time, so they have to outlive the machine that made them.
That is the rule the Vault exists to enforce: a data root is where inputs are materialized, not where they live. Frozen inputs belong to the Vault, checkpoints to Hugging Face, and a run’s Traces, results, and events to the results bucket. A run launched with --push, which is what a remote run runs with, sends its own results as it goes, and anything else waits for an explicit monte push.
How a Graded Rollout Runs
A rollout runs across three servers: one for the model, one for the agent harness, and one for the environment’s grader.- The harness reads one task from the exam’s frozen task list.
- The harness sends the task to the model server, with the prompt already applied, and collects the reply.
- The harness sends the reply to the grader.
- The grader returns a reward and the answer that it extracted.
- Monte records one Trace for each task, then writes the score to the ledger.
Related Pages
Experiment
Which exams are sealed, and what the fingerprint pins.
monte env
Freeze an environment, list what is registered, install its assets.
Further Reading
- Environments (NeMo Gym)

