跳到正文
AI 脉动

VOL.2026.08.12 · 30 篇报道 · AI 日报

AI 日报 — 2026-08-12

星期三 · 30 篇报道 · 约 9 分钟读完

今日主线

今日AI领域呈现出模型能力提升与复杂AI智能体实际部署的双重焦点。DeepSeek V4 Pro和Qwen3.8-2.4T等新模型正在突破性能极限,而对内省意识和安全漏洞的研究则凸显了大型语言模型日益增长的复杂性。与此同时,从医疗咨询系统到3D世界生成器等先进智能体的出现,展示了AI在现实世界中不断扩展的实用性,尽管政策和行业变动预示着监管和竞争环境的持续演变。

01模型与开源11 篇

  1. DeepSeek V4 Pro 0813 (on OpenRouter)
    日榜第 1 名0 个来源热度 61
  2. Qwen3.8-2.4T
    日榜第 2 名0 个来源热度 55
  3. Emergent Introspective Awareness in Large Language Models
    日榜第 5 名0 个来源热度 43
  4. What sort of maths are LLMs good at?
    日榜第 13 名0 个来源热度 32
  5. llama.cpp
    日榜第 15 名0 个来源热度 31
  6. Everything announced at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and tons of Gemini features

    Google unveiled the Pixel 11 series, Pixel Watch 5, and a new Pixel Tag at its Made by Google 2026 event. The event also highlighted new Gemini-powered features across its devices. The Pixel Watch 5 starts at $399 for the 41mm model and $429 for the 45mm model, with a Stephen Curry edition available for $579.

    日榜第 23 名0 个来源热度 27

02Agent 与工具5 篇

  1. WorldClaw Agentic 3D open-world generation at scale
    日榜第 8 名0 个来源热度 40
  2. AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.

    Google Research and Google DeepMind are advancing AMIE, their research medical AI system, towards real-time clinical video consultations. Built on Gemini and Project Astra with a multi-agent architecture, AMIE can now interpret visual and auditory cues, guide virtual physical exams, and reason diagnostically in real time. This system demonstrates expert-level AI capabilities in this setting, offering a glimpse into the future of health AI, though further research is needed before real-world clinical deployment.

    日榜第 14 名1 个来源热度 31
  3. LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge

    LFM2.5-VL-3B is a vision-language model designed for on-device and real-time applications, offering enhanced vision capabilities for edge computing. This model, which understands documents and screens, grounds objects, and can call tools, provides direct answers rather than reasoning to maintain speed. Benchmarks show LFM2.5-VL-3B (3.1B) achieving an average score of 69.4, outperforming LFM2-VL-3B (3.1B) at 57.2 and gemma-4-E2B-it (5.1B) at 52.0 across various tasks including MMStar, MME, RealWorldQA, and OCRBench v2 (En).

    日榜第 17 名0 个来源热度 30

03应用落地3 篇

  1. Premium seats are coming to ChatGPT Business

    ChatGPT Business is introducing Premium seats, offering a limited-time promotion for the first 10,000 eligible customers. These customers can receive $100 in workspace credits (2,500 credits) for each Premium seat added, up to a maximum of 5 seats. This promotion concludes on August 20, and interested parties can find more details regarding eligibility and how the promotion works in the help center article.

    日榜第 11 名0 个来源热度 36
  2. Daybreak models are now available on AWS

    OpenAI has announced that its Daybreak models are now available on AWS, expanding on earlier availability of OpenAI frontier models and Codex. This integration allows enterprises to access Daybreak capabilities through Amazon Bedrock. Users can utilize Daybreak Red and Daybreak Blue models via the Amazon Bedrock console or the Responses API using the bedrock-mantle endpoint, following enrollment in Daybreak Access.

    日榜第 19 名1 个来源热度 28
  3. Testing ads in ChatGPT

    OpenAI is testing ads in ChatGPT for logged-in adult users on the Free and Go subscription tiers in the U.S., with plans to expand to more markets. Ads will not appear on Plus, Pro, Business, Enterprise, and Education tiers. The company states that ads will not influence ChatGPT's answers, conversations will remain private from advertisers, and users will retain control over their experience. This initiative aims to support broader access to powerful ChatGPT features while maintaining user trust.

    日榜第 24 名0 个来源热度 27

04政策&风险1 篇

  1. Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes

    Anthropic has implemented watermarking on Claude's outputs, embedding invisible code to identify AI-generated text. This decision aligns with the EU AI Act's Transparency Code, which mandates labeling AI-generated or edited content for computer systems. While European regulators may approve, some Claude users are expressing dissatisfaction with this new policy, particularly concerning its implications for their use of the AI in professional or academic settings.

    日榜第 22 名0 个来源热度 27

05行业动态10 篇

  1. Pixel Watch 5
    日榜第 3 名0 个来源热度 47
  2. Google launches Pixel 11 Pro Fold
    日榜第 4 名0 个来源热度 44
  3. Pixel 11 Pro Fold
    日榜第 7 名0 个来源热度 42
  4. Grok Bot
    日榜第 20 名0 个来源热度 28