OpenEnv documentation

Training with OpenEnv

You are viewing main version, which requires installation from source. If you'd like regular pip install, checkout the latest stable version (v0.8.0).
Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Training with OpenEnv

OpenEnv is a contract between environments and the code that uses them, not a trainer. An environment exposes reset(), step() and state() (and its tools over MCP), runs as a service anywhere, and computes its own reward, so any training framework that can call it can train on it. This page maps the ways to train on OpenEnv environments and the frameworks that support it.

Ways to Train

You want toWhat the environment givesWorked examples
Train a model with RL, with the trainer running the episodeObservations, tools and a reward. The trainer generates every turn, so it has the tokens and logprobs. This is the white-box pathWordle, a reasoning model, web tasks in BrowserGym with TRL, 2048 with Unsloth
Train the model behind a real agent (OpenCode, Claude Code, Codex, …)Harbor runs the agent on a task, captures every model call, and returns a TrainingTrace with token ids, logprobs and the verifier’s reward. This is the black-box pathHarbor, trained with TRL’s AsyncGRPOTrainer
Warm-start a model with supervised dataopenenv collect runs a teacher in the environment and saves reward-labeled rollouts as a datasetCollecting rollouts for SFT
Evaluate, not trainScores from the environment’s rubricInspect AI, an agent inside the environment

None of these paths ties the environment to a framework. openenv.core does not depend on any trainer, and a TrainingTrace or a collected dataset is plain data that any trainer can read. The worked examples use the frameworks named above because that is where the examples exist today. Harnesses in OpenEnv compares the white-box and black-box paths in more detail.

Integrations

These frameworks and platforms train on OpenEnv environments. If your project supports OpenEnv, open a PR to add it here and to the README.

FrameworkExample
TRLOpenEnv integration guide: environment_factory with GRPOTrainer, several environments at once, and harness training through Harbor
torchforgeGRPO on BlackJack
Unsloth2048 with gpt-oss
SkyRLTraining on OpenEnv environments
ARTOpenEnv integration
OumiOpenEnv GRPO notebook
Lightning AIOpenEnv templates
MilesGRPO on Terminal-Bench 2

Your Own Training Loop

Every integration above comes down to the same calls. To plug OpenEnv into a framework that has no integration yet, drive the client from its rollout code:

from openenv import AutoAction, AutoEnv

env = AutoEnv.from_env("my-env")
Action = AutoAction.from_env("my-env")

with env.sync() as client:
    for episode in range(num_episodes):
        result = client.reset()
        while not result.done:
            action = policy(result.observation)  # your model picks the next action
            result = client.step(action)
            # result.reward is the environment's reward for this step

The Task API lets a trainer list an environment’s tasks and pick which one each episode runs, and Rewards covers how environments compute the reward.

Update on GitHub