返回
Hhackernews·altertable
爆款 · 6.4×36
·10小时前·其他 · 官方 API

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

查看原文
OpenAIQwen模型发布模型访问限时活动开源代码

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

OpenAI 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。

爆款判定
判定依据
热度约为该来源近期上榜条目中位水平的 6.4 倍
指标对比
348 vs 中位 54(20 条基线样本)
检出时间
09/03 21:00
AI 摘要

Qwen 3.8 27B 模型现已在 Cerebras 公共端点上提供,其速度约为每秒 1500 个令牌。该模型拥有 270 亿个参数,并支持免费用户 64k 和付费用户 128k 的上下文。作为对比,OpenAI GPT OSS gpt-oss-120b 模型拥有 1200 亿个参数,速度约为每秒 3000 个令牌,并支持 65k/131k 的上下文。

Browse all models available on Cerebras public endpoints.

Models on Cerebras public endpoints are available on the free trial and pay-as-you-go tiers, subject to rate limits and pricing . For additional model families, reserved capacity, higher throughput, and production SLAs, see Dedicated Endpoints .

Available Models

Model Name Model ID Parameters Context (free / paid) Speed (tokens/s) OpenAI GPT OSS gpt-oss-120b 120 billion 65k / 131k ~3000 Qwen 3.8 27B qwen-3.8-27b 27 billion 64k / 128k ~1500

Model Compression

This section provides transparency about the compression state of each model available on our platform. We host a variety of open-source models from the community. We do not currently host pruned models on our public endpoints. All models served through our public endpoints are the original, unpruned versions. While we conduct research on pruning techniques like REAP (Router-weighted Expert Activation Pruning), these pruned models are shared with the research community on Hugging Face but are not available through our shared API. You can read more about REAP in our research blog . All of our public models are unpruned. Cerebras uses selective weight-only quantization only during storage to preserve maximal quality. This means that the weights are stored in partial 16-bit / 8-bit / 4-bit, in-line with industry standards. For quality, sensitive layers are stored at full precision with dequantization on the fly, so operations are done in high precision. The activations, attention, and kv cache remain in full precision and unquantized.

Frequently Asked Questions

Will you change a model's architecture without notice?

No. We are committed to serving the original models for all existing endpoints, without modification. We do not alter model architectures via pruning on our hosted portfolio. If we explore additional compression techniques (like pruning) in the future, these would be offered as separate endpoints with pruning-specific names, ensuring complete transparency and allowing you to choose which version best fits your needs.

Where can I find your REAP pruned models?

Our REAP pruned models are available on Hugging Face for research and experimentation purposes: Cerebras REAP Collection . These models demonstrate our pruning research but are not served through our production API.

What are compression, quantization, and pruning?

Compression is an umbrella term for techniques that reduce model size or computational requirements. Common compression techniques include:

- Quantization: Reducing the precision of numbers used to represent model weights (e.g., converting from FP16 to FP8). This reduces memory usage without changing the model’s architecture.

- Pruning: Permanently removing parts of a model, like layers or experts, to reduce model size. This changes the model’s architecture and creates a different model.