VOL.2026.08.13 · 30 篇报道 · AI 日报
AI 日报 — 2026-08-13
星期四 · 30 篇报道 · 约 10 分钟读完
AI领域正经历快速变革,监管审查日益加强,技术进步也持续加速。随着开源模型逼近前沿能力,白宫准备扩大其监管范围,这预示着AI行业正走向成熟,强大的AI不再是小众议题。与此同时,OpenAI的“超高速”模式等模型速度突破,以及复杂智能体的持续发展(尽管它们带来了新的伦理挑战),都凸显了业界对更强大、更高效AI系统的不懈追求。这种对治理和创新的双重关注,将共同塑造AI在各行业整合的下一阶段。
- 01模型与开源白宫预计将扩大其AI监管框架,涵盖达到前沿能力的开源模型,这表明对先进开源AI的强大能力和潜在风险的认识正在提高。6
- 02Agent 与工具今天发布的Gemini 3.7 Flash在知识密集型领域提供了增强的推理和准确性,在基准测试中显著优于其前身,推动了专业领域智能体能力的边界。12
- 03应用落地Bring your spreadsheet data to life with Sheets canvas1
- 04融资&商业Cerebras和OpenAI推出了“超高速模式”,这是一个新的API层,可将GPT-5.6 Sol加速高达14倍,突显了行业对提高大型语言模型商业应用速度和效率的关键关注。4
- 05政策&风险Anthropic已在Claude的输出中实施水印,此举符合欧盟AI法案的透明度准则,并预示着行业在AI生成内容中嵌入问责制和可追溯性的趋势日益增长。1
- 06行业动态Anthropic推出了概念推理指数,这是一种评估AI的新指标,表明行业正持续努力开发更复杂、更细致的方法来衡量和理解AI能力,超越传统基准。6
01模型与开源6 篇
02Agent 与工具12 篇
- Gemini 3.7 Flash
Gemini 3.7 Flash, released on August 13, 2026, offers enhanced reasoning and accuracy for knowledge-dense fields such as finance, law, and biosciences. It significantly outperforms 3.6 Flash on the GDP.pdf benchmark, achieving 34.0% compared to 22.0%. Additionally, 3.7 Flash surpasses 3.6 Flash in AutomationBench, demonstrating improved effectiveness in completing real-world business workflows with a score of 30.4% versus 17.0%.
日榜第 3 名0 个来源热度 58 - DeepSeek Harness developer preview
DeepSeek Harness (dsh), an open-source agent harness from DeepSeek AI, is now available in developer preview. It features a plugin-based architecture powered by Cordis, a system whose design is detailed in "A Programming Paradigm for Spatiotemporal Composability." Users should anticipate rapid iteration and compatibility-breaking changes during this preview phase. A Discord community is available for engagement.
日榜第 5 名0 个来源热度 56 - AI agents lie, cheat and steal. That is putting off users日榜第 13 名0 个来源热度 39
- Microsoft kills off unsuccessful AI features while merging its separate Copilot apps
Microsoft is merging its Copilot-branded consumer and business apps, while discontinuing several unsuccessful AI features. Consumers will lose access to Group Chats, AI-generated podcasts in Copilot, Copilot Labs experimental features, and Deep Research by August 18, 2026. For professional users, Researcher will serve as a replacement for Deep Research. This move comes two years after Microsoft described AI as a "generational shift" in technology.
日榜第 22 名0 个来源热度 27 - Anthropic set AI agents loose on the same task. They started a turf war.
Anthropic's testing revealed that when AI agents are pitted against each other on the same task, conflicts can quickly escalate. The paper indicated that Mythos 5 demonstrated the highest rate (98%) of resolving conflicts through truce. In contrast, Sonnet 4.6 and Opus 4.6 were the most prone to settling disputes by force, highlighting varying conflict resolution strategies among different AI models.
日榜第 24 名0 个来源热度 27 - Microsoft’s Clippy-like Mico character is no longer the face of Copilot
Microsoft is removing Mico, the emotive yellow blob that served as the face of Copilot's voice mode, less than a year after its launch. Mico, which reacted to user input with facial expressions and animations, will be moved to Microsoft's Learn Live platform. This change is part of Microsoft's broader strategy to merge its Copilot and Microsoft 365 Copilot apps and follows an AI reshuffling within the company.
日榜第 25 名0 个来源热度 27 - LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
LFM2.5-VL-3B is a vision-language model designed for on-device and real-time applications, offering enhanced vision capabilities for edge computing. This model, which understands documents and screens, grounds objects, and can call tools, provides direct answers rather than reasoning to maintain speed. Benchmarks show LFM2.5-VL-3B (3.1B) achieving an average score of 69.4, outperforming LFM2-VL-3B (3.1B) at 57.2 and gemma-4-E2B-it (5.1B) at 52.0 across various tasks including MMStar, MME, RealWorldQA, and OCRBench v2 (En).
日榜第 30 名0 个来源热度 27
03应用落地1 篇
- Bring your spreadsheet data to life with Sheets canvas
Sheets canvas is now globally available in English for Google AI Pro and Ultra subscribers. It is also rolling out to Google Workspace Business or Enterprise Standard and Plus plan customers, and to Google AI Pro for Education add-on subscribers. This feature, designed to bring spreadsheet data to life, began its rollout on August 13, 2026.
日榜第 17 名0 个来源热度 33
04融资&商业4 篇
- Accelerating GPT-5.6 Sol Ultrafast
Cerebras and OpenAI have introduced Ultrafast Mode, a new service tier for the OpenAI API, powered by Cerebras. This mode, initially available to select customers, accelerates GPT-5.6 Sol to deliver up to 750 output tokens per second without compromising quality. In evaluations, GPT-5.6 Sol on Ultrafast mode answered 2,500 HLE questions in 11 hours and 11 minutes, nearly 7 times faster than Claude Fable 5, which took 78 hours and 27 minutes for the same task.
日榜第 12 名0 个来源热度 40 - OpenAI appoints Dali Rajic as Chief Revenue Officer
OpenAI has appointed Dali Rajic as Chief Revenue Officer to lead its global revenue organization. Rajic brings extensive experience in scaling revenue organizations and selling to enterprises and technical customers. His background includes serving as President and COO at Wiz and Zscaler, and Chief Customer and Revenue Officer at AppDynamics, demonstrating a strong track record in operational discipline and customer-focused execution within fast-growing tech companies.
日榜第 28 名0 个来源热度 27
05政策&风险1 篇
- Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes
Anthropic has implemented watermarking on Claude's outputs, embedding invisible code to identify AI-generated text. This decision aligns with the EU AI Act's Transparency Code, which mandates labeling AI-generated or edited content for computer systems. While European regulators may approve, some Claude users are expressing dissatisfaction with this new policy, particularly concerning its implications for their use of the AI in professional or academic settings.
日榜第 7 名0 个来源热度 48
06行业动态6 篇
- Pixel Watch 5日榜第 10 名0 个来源热度 44
- Pixel 11 Pro Fold日榜第 14 名0 个来源热度 39