Welcome RL Environments to the hub
Hugging Face Hub now supports Reinforcement Learning (RL) Environments, offering a dedicated space for these environments to enhance agentic AI systems. This integration allows users to measure and improve agent performance. Developers are encouraged to publish and tag their RL environments, including necessary files, a run command, and the reward rule, to facilitate broader usage and collaboration. The platform supports various tasks such as coding, tool use, games, and robotics, and welcomes contributions for new frameworks.
This marks the first time Hugging Face Hub has dedicated a space for RL environments, unlike its previous focus solely on models and datasets.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月28日 00:00 UTC
收录当时偏移:UTC+02026年10月5日 15:00 UTC
- 发布
- 2026年9月28日 00:00
- 收录
- 2026年10月5日 15:00
- 来源类型
- 官方发布
- 档位
- 当事方
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, Hugging Face Hub now has a special place for RL Environments.
An environment gives an agent a task, responds to its actions with observations, and scores the outcome. The resulting rewards can measure an agent's performance during evaluation or provide a learning signal during training. For an introduction to this interaction loop, see our blogpost on environments . Within the environment, the agent will perform a set of tasks that are represented as datasets. Therefore, environments can be split into broadly two parts: tasksets and runtimes. In this release, we are focusing on the tasksets.
An RL environment on the Hub is a dataset repo that shows up in the new RL Environments filter . The Use this dataset button gives you the command to run it in that framework. There is no new repo type, no registry, and no sign-up. There are already environments in Harbor, Verifiers, and NVIDIA NeMo Gym.
Stop building environment registries
Every RL paper or framework uses its own way to find environments. Custom hubs, runtime registries, independent task datasets, or a GitHub list of tasks with a custom loader. This means that many of the published environments are siloed: if you publish an environment for one framework, users of the other three can’t load it. If you want to train on an environment from another framework or a new paper, you’ll need to port it by hand.
We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule, which are data with a runtime on top. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it."
From a dataset repository to an agent run The Hub stores task data but the frameworks store runtime and verfier code. A framework loads the data and supplies runtime or verifier implementations when they are not included in the repository. At runtime an agent exchanges actions and observations with an environment. A verifier scores the outcome and produces rewards for evaluation or training.
Dataset repository on the Hub Tasks and data · Runtime and verifier files, when included
Framework loads the files
Execution on your machine or a supported cloud backend
Agent
Environment State, tools, and task execution
Actions Observations
Outcome
Verifier
Reward Evaluate or train Score runs or update the model
Task data lives on the Hub. Runtime configuration and verifier code can live in the repo or the framework.
The frameworks keep doing what they are good at. The Hub does what it is good at, which is hosting, discovery, and versioning. Nobody has to own the catalogue. In fact, catalogues can run on other platforms too, powered by the hub.
The dataset repository hosts your environment files. The framework runs them locally or on a supported cloud backend. Hugging Face Jobs can run cloud workloads, and Hugging Face Sandboxes , built on Jobs, provide interactive command execution. The tags describe compatibility and generate loading commands; adding a tag does not start a job or sandbox.
What shipped
The RL Environments filter. Go to huggingface.co/datasets?other=rl-environment . Every dataset with the rl-environment tag appears there, whatever framework it works with.
Framework tags. Four environment frameworks are registered as dataset libraries:
Tag Framework
harbor Harbor