返回
Hhackernews·bratao
爆款 · 3.6×52
·3小时前·官方发布 · 官方 API

Gemini 3.8 Flash and 3.8 Flash Cyber

查看原文
官方公告Gemini模型发布模型访问

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

官方发布涉及Gemini 模型访问、订阅权益规则,适合跟踪产品开放节奏和用户影响。

爆款判定
判定依据
热度约为该来源近期上榜条目中位水平的 3.6 倍
指标对比
338 vs 中位 93.5(20 条基线样本)
检出时间
09/02 16:01
AI 摘要

Gemini 3.8 Flash 是一款专为关键企业自主性设计的新模型,在量化和专业领域表现出色。它在 Vals Finance Agent V2 和 Harvey's Legal Agent Benchmark 等基准测试中超越了 3.7 Flash 和其他前沿模型。3.8 Flash 在 HLE-Verified 上取得了 54.9% 的成绩,展现了其在 STEM、人文和专业领域处理多步骤推理的能力。此外,Gemini 3.8 Flash Cyber 通过 Fairwind Program 提供,为政府机构和关键基础设施运营商提供优先访问权限。

Sep 02, 2026

|

Our newest Gemini models deliver next-generation intelligence for agentic workflows and cybersecurity.

Raluca Ada Popa

Gemini Security Lead, Google DeepMind

Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7. Gemini 3.8 introduces 2 variants:

- Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains. It is available at the same introductory price

1

as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.

- Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program .

While tailored for different deployment environments, both of today's releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models. The significant coding and reasoning gains across this shared core were driven by a number of innovations, including rigorous training in the highly demanding domain of cybersecurity.

Gemini 3.8 Flash: built for long-horizon coding and autonomous agents

Gemini 3.8 Flash delivers substantial gains from 3.7 Flash, often approaching the performance of higher-cost frontier models.

On DeepSWE v1.1 (Long-Horizon Software Engineering) 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost.

Additionally, 3.8 Flash exhibits the dependability required for critical enterprise autonomy, across specialized knowledge domains. In quantitative and professional fields that require advanced analysis and reporting, 3.8 Flash outperforms 3.7 Flash and other frontier models in benchmarks like Vals Finance Agent V2 and Harvey's Legal Agent Benchmark . 3.8 Flash also achieves a 54.9% on HLE-Verified, demonstrating its ability to handle multi-step reasoning across STEM, humanities, and professional fields.

These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.

For applications where compute efficiency is the primary constraint, developers can utilize lower effort levels to minimize token overhead or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads.

Gemini 3.8 Flash Cyber: expert cyber performance

Gemini 3.8 Flash Cyber, available to a set of trusted defenders via the Fairwind Program , provides a decisive advantage in today’s complex cybersecurity landscape, with the Flash speed and cost that enables quick iteration.

Autonomous vulnerability discovery

On the standard industry benchmark for finding vulnerabilities, CyberGym, Gemini 3.8 Flash Cyber demonstrates frontier-level performance in autonomous vulnerability discovery. It surpasses both 3.5 Flash Cyber as well as significantly larger frontier models.

To better capture real-world defensive needs which are not limited to just C/C++ codebases like in CyberGym, we also evaluated Gemini 3.8 Flash Cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex codebases spanning 20 programming languages. Here, the model showcases an impressive leap over our previous models and reaches a success rate exceeding 70%.

Automated patching

With Gemini 3.8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers. This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation.

Gemini 3.8 Flash and 3.8 Flash Cyber · BuzzRadr