Skip to content
HChuggingface.co·

Welcome RL Environments to the hub

AI summary

Hugging Face Hub now supports Reinforcement Learning (RL) Environments, offering a dedicated space for these environments to enhance agentic AI systems. This integration allows users to measure and improve agent performance. Developers are encouraged to publish and tag their RL environments, including necessary files, a run command, and the reward rule, to facilitate broader usage and collaboration. The platform supports various tasks such as coding, tool use, games, and robotics, and welcomes contributions for new frameworks.

Why this one

This marks the first time Hugging Face Hub has dedicated a space for RL environments, unlike its previous focus solely on models and datasets.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 28, 2026, 00:00 UTC

IngestedOffset at this time: UTC+0Oct 5, 2026, 15:00 UTC

Published
Sep 28, 2026, 00:00
Ingested
Oct 5, 2026, 15:00
Source type
Official
Tier
First-party
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, Hugging Face Hub now has a special place for RL Environments.

An environment gives an agent a task, responds to its actions with observations, and scores the outcome. The resulting rewards can measure an agent's performance during evaluation or provide a learning signal during training. For an introduction to this interaction loop, see our blogpost on environments . Within the environment, the agent will perform a set of tasks that are represented as datasets. Therefore, environments can be split into broadly two parts: tasksets and runtimes. In this release, we are focusing on the tasksets.

An RL environment on the Hub is a dataset repo that shows up in the new RL Environments filter . The Use this dataset button gives you the command to run it in that framework. There is no new repo type, no registry, and no sign-up. There are already environments in Harbor, Verifiers, and NVIDIA NeMo Gym.

Stop building environment registries

Every RL paper or framework uses its own way to find environments. Custom hubs, runtime registries, independent task datasets, or a GitHub list of tasks with a custom loader. This means that many of the published environments are siloed: if you publish an environment for one framework, users of the other three can’t load it. If you want to train on an environment from another framework or a new paper, you’ll need to port it by hand.

We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule, which are data with a runtime on top. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it."

From a dataset repository to an agent run The Hub stores task data but the frameworks store runtime and verfier code. A framework loads the data and supplies runtime or verifier implementations when they are not included in the repository. At runtime an agent exchanges actions and observations with an environment. A verifier scores the outcome and produces rewards for evaluation or training.

Dataset repository on the Hub Tasks and data · Runtime and verifier files, when included

Framework loads the files

Execution on your machine or a supported cloud backend

Agent

Environment State, tools, and task execution

Actions Observations

Outcome

Verifier

Reward Evaluate or train Score runs or update the model

Task data lives on the Hub. Runtime configuration and verifier code can live in the repo or the framework.

The frameworks keep doing what they are good at. The Hub does what it is good at, which is hosting, discovery, and versioning. Nobody has to own the catalogue. In fact, catalogues can run on other platforms too, powered by the hub.

The dataset repository hosts your environment files. The framework runs them locally or on a supported cloud backend. Hugging Face Jobs can run cloud workloads, and Hugging Face Sandboxes , built on Jobs, provide interactive command execution. The tags describe compatibility and generate loading commands; adding a tag does not start a job or sandbox.

What shipped

The RL Environments filter. Go to huggingface.co/datasets?other=rl-environment . Every dataset with the rl-environment tag appears there, whatever framework it works with.

Framework tags. Four environment frameworks are registered as dataset libraries:

Tag Framework

harbor Harbor