An end-to-end framework for creating, deploying and using isolated execution environments for agentic RL, with a simple Gymnasium-style API.
- One API for every environment.
reset(),step()andstate(), sync or async, over a WebSocket. - Isolated, deployable environments. Each environment is a Docker image that runs locally, on a cloud sandbox, or as a Hugging Face Space.
- 40+ ready-to-use environments, from games and coding sandboxes to browsers, finance and simulators, in the environment catalog.
- Train with your framework. TRL, Unsloth, SkyRL, ART, Oumi, torchforge, Miles and more (see Integrations).
- Train real coding agents. The Harbor integration runs Claude Code, Codex, OpenCode, mini-swe-agent and other harnesses on Harbor tasks, and captures the exact tokens for RL.
- Serve tools to agents. MCP environments expose their tools to agents over
/mcpin production mode, while training keeps usingstep()and rewards. - Rewards and evals built in. Compose rewards with rubrics and evaluate with Inspect AI.
Install the OpenEnv package:
pip install openenvInstall an environment client (e.g., Echo):
pip install git+https://huggingface.co/spaces/openenv/echo_envThen use the environment:
import asyncio
from echo_env import CallToolAction, EchoEnv
async def main():
# Connect to a running Space (async context manager)
async with EchoEnv(base_url="https://openenv-echo-env.hf.space") as client:
# Reset the environment
result = await client.reset()
print(result.observation.metadata["message"]) # "Echo environment ready!"
# Send messages
result = await client.step(
CallToolAction(
tool_name="echo_message",
arguments={"message": "Hello, World!"},
)
)
print(result.observation.result["data"]) # "Hello, World!"
print(result.reward)
asyncio.run(main())Synchronous usage is also supported via the .sync() wrapper:
from echo_env import CallToolAction, EchoEnv
# Use .sync() for synchronous context manager
with EchoEnv(base_url="https://openenv-echo-env.hf.space").sync() as client:
result = client.reset()
result = client.step(
CallToolAction(
tool_name="echo_message",
arguments={"message": "Hello, World!"},
)
)
print(result.observation.result["data"])For a detailed quick start, check out the docs page.
Any training framework that can call an environment can train on it. Training with OpenEnv maps the ways to train and the frameworks that support each one.
- Environments as tools (white-box). The trainer runs the multi-turn tool loop and the environment supplies the tools and the reward. In TRL,
GRPOTrainertakes anenvironment_factory: start with the Wordle GRPO tutorial or TRL's OpenEnv guide. - Real agent harnesses (loop-owning). The agent runs its own loop, and OpenEnv's Harbor integration captures every model call as a framework-neutral
TrainingTrace. TRL'sAsyncGRPOTrainertrains on those captures today: seeexamples/async_grpo_harborand The ultimate guide to multi-harness RL.
openenv init my_env # scaffold an environment
openenv validate my_env --level static --skip-build # quick check against the OpenEnv contract
openenv push my_env # deploy it to Hugging Face SpacesSee Your First Environment and Packaging & Deploying. openenv import wraps an existing environment from ORS/OpenReward or Verifiers.
A few to start with:
| Environment | What it is |
|---|---|
| Echo | Minimal MCP environment, for learning the API and testing a deployment |
| Coding | Sandboxed Python execution with stdout, stderr and exit codes |
| TextArena (Wordle and more) | Text games for multi-turn RL |
| OpenSpiel | Board and card games from DeepMind's OpenSpiel |
| BrowserGym | Web navigation tasks (MiniWoB++, WebArena, ...) |
| Harbor | Harbor task datasets through coding-agent harnesses, with token capture for training |
Browse all of them in the environment catalog, or on the OpenEnv Hub organization.
OpenEnv works with a growing ecosystem of RL frameworks and platforms. If your project supports OpenEnv, open a PR to add it here.
| Framework | Example |
|---|---|
| TRL | OpenEnv guide (GRPO with environment_factory, and harness training) |
| Unsloth | 2048 with gpt-oss |
| SkyRL | SkyRL example |
| ART | ART integration |
| Oumi | GRPO notebook |
| torchforge | GRPO BlackJack |
| Miles | Terminal-Bench-2 GRPO |
| Lightning AI | Templates |
- Documentation: concepts, guides and tutorials
- Core Concepts and the CLI reference
- Tutorials, and the Zero to Hero tutorial from our GPU Mode lecture
- RFCs, the proposals behind major changes
- Contributing: development setup, tests and the PR process
Note
OpenEnv is in early development, so APIs may still change. Bug fixes are welcome. For larger changes, open or claim an issue first so the change can be discussed.
OpenEnv is governed by a technical committee that coordinates project direction, major technical decisions, RFCs, and release planning through the public issue tracker, pull requests, and RFC process. Current committee members: Meta-PyTorch, Reflection, Unsloth, Modal, Prime Intellect, Nvidia, Mercor, Fleet AI, Microsoft, Hugging Face, RadixArk, and Nebius.
The project is also supported by a broader community of organizations. If you would like to add your project or organization here, please open a pull request for maintainer review.
Supporters include: Meta-PyTorch, Hugging Face, Scaler AI Labs, Patronus AI, Surge AI, LastMile AI, Unsloth, Reflection, vLLM, SkyRL (UC-Berkeley), Lightning AI, Axolotl AI, Stanford Scaling Intelligence Lab, Mithril, OpenMined, Fleet AI, Halluminate, Turing, Scale AI, Scorecard, Snorkel AI, SGLang, Miles, Nebius
And we'd also like to acknowledge the team at Farama Foundation as the OpenEnv API was heavily inspired by the work you all have done on Gymnasium. Cheers!
BSD 3-Clause License (see LICENSE file)