Baseten on Hugging Face Inference Providers 🔥
热度趋势
百分比基于当前可用热度信号,而非评论数或独立用户人数。
官方发布带来DeepSeek 模型更新信号,适合跟踪能力变化、生态影响和后续落地。
Baseten 现已成为 Hugging Face Hub 上受支持的推理提供商。用户可以利用 Baseten 进行聊天补全,例如使用 "deepseek-ai/DeepSeek-V4-Flash-0731:baseten" 模型生成带有记忆化的 Python 函数。Hugging Face 鼓励用户就此新集成提供反馈。
We're thrilled to share that Baseten is now a supported Inference Provider on the Hugging Face Hub!
Baseten joins our growing ecosystem, enhancing the breadth and capabilities of serverless inference directly on the Hub's model pages. Inference Providers are also seamlessly integrated into our client SDKs (for both JS and Python), making it super easy to use a wide variety of models with your preferred providers.
Baseten is an AI infrastructure platform that covers serverless AI, training and more. With a catalog of many frontier models, Baseten makes it easy for developers to integrate a wide range of AI capabilities into their applications with minimal setup.
Baseten supports a broad spectrum of model types - from LLMs to text-to-speech and more. As part of this initial integration, Baseten is launching support for conversational and text-generation tasks on Hugging Face, enabling access to popular open-weight LLMs such as Kimi K3, latest DeepSeek V4 Flash, GLM-5.2, and many more. Support for additional tasks will roll out soon!
See the full list of models supported by Baseten here.
Follow Baseten on Hugging Face: https://huggingface.co/baseten.
How it works
In the website UI
- In your user account settings, you are able to:
- Set your own API keys for the providers you've signed up with. If no custom key is set, your requests will be routed through HF.
- Order providers by preference. This applies to the widget and code snippets in the model pages.
- As mentioned, there are two modes when calling Inference Providers:
- Custom key (calls go directly to the inference provider, using your own API key of the corresponding inference provider)
- Routed by HF (in that case, you don't need a token from the provider, and the charges are applied directly to your HF account rather than the provider's account)
- Model pages showcase third-party inference providers (the ones that are compatible with the current model, sorted by user preference)
From the client SDKs
Baseten is available through the Hugging Face SDKs - huggingface_hub (>= 1.26.1) for Python and @huggingface/inference for JavaScript.
The following examples show how to use the latest DeepSeek V4 Flash through Baseten. Use a Hugging Face token to authenticate - the request will be routed to Baseten automatically.
From your favorite Agent Harness
Hugging Face Inference Providers are integrated in most Agent Harnesses - including Pi, OpenCode, Hermes Agents, OpenClaw, and more. This means you can plug baseten-hosted models straight into your favorite tools without any extra glue code. Browse the full list of integrations here.
from Python
import os from openai import OpenAI
client = OpenAI( base_url= "https://router.huggingface.co/v1", api_key=os.environ[ "HF_TOKEN"], )
completion = client.chat.completions.create( model= "deepseek-ai/DeepSeek-V4-Flash-0731:baseten", messages=[ { "role": "user", "content": "Write a Python function that returns the nth Fibonacci number using memoization." } ], )