Skip to main content
Training uses an algorithm to update a model’s weights from a signal. The algorithm determines how examples, rewards, or another model’s feedback change the model during a run. The resulting behavior is saved in a checkpoint and can be evaluated independently of the training process. Monte covers post-training: starting from a capable base model and specializing it for work represented by an environment or dataset. Each training run states a recipe, a source, and a starting point. It produces a checkpoint for later evaluation, and the ledger records the choices that produced it.

Training Methods

Supervised Fine-Tuning: Supervised fine-tuning adapts a model from completed examples, each pairing an input with the response the model should produce. Its objective increases the likelihood of that response for the input, so the model learns the patterns represented in the demonstrations. It works well when the desired behavior can be shown directly, such as a task procedure, output format, or tool-calling convention. In Monte, those demonstrations are frozen as a dataset, and the run records that dataset as its source. Reinforcement Learning: Reinforcement learning optimizes a model’s behavior from feedback on the outputs it generates. Rather than train only against a fixed target response, it samples attempts, evaluates them with a reward function, and updates the model to make higher-reward behavior more likely. It fits work where success can be defined by an executable or programmatic check instead of a single reference answer. In Monte, a curriculum environment provides the tasks and grader; the grader evaluates each attempt, and the resulting reward is used by the training algorithm. Distillation: Distillation trains a student model to reproduce behavior from a teacher model. The teacher can provide target outputs, rankings, or feedback, giving the student a learning signal without requiring every example to be written by hand. It is useful for turning the behavior of a stronger model into training material while keeping the work grounded in the tasks that matter. In Monte, a curriculum environment supplies the tasks and grader around the distillation run, so the work being trained remains the work being evaluated.

Recipes

A recipe is the versioned file that states how a run trains: model shape, algorithm, batch geometry, and optimizer. Monte pins it by content hash and records that hash on the run. The source is stated separately with --source, so the method and the material it trains on remain distinct. monte status labels each lineage hop as algorithm:source. For example, grpo:demo is a reinforcement-learning run on the demo curriculum. Improvement explains how Monte reads those runs back as a lineage.

Environments

Where the reward comes from during training.

Experiment

What stays fixed while recipes change.

Improvement

What a trained checkpoint is used for next.