跳到正文
AI 脉动

VOL.2026.09.10 · 30 篇报道 · AI 日报

AI 日报 — 2026-09-10

星期四 · 30 篇报道 · 约 16 分钟读完

今日主线

OpenAI凭借GPT-6 Astra等新模型和开发者工具持续推动AI前沿,但其快速发展也伴随着日益增长的担忧。关于不道德数据实践、安全漏洞的指控,以及前研究员对AI生存威胁的严峻警告,都加剧了对加强审查和监管的呼声。这种创新与问责之间的紧张关系定义了当前AI格局,随着行业努力应对日益强大和自主的系统所带来的影响。

01模型与开源9 篇

  1. Qwen 3.8 follows GPT-5.5 Pro reasoning prefills

    A recent analysis of reasoning prefills across several open models, including DeepSeek V4 Flash, Inkling, Kimi K3, and Qwen3.8 A95B, shows varying degrees of alignment with GPT-5.5 Pro. Qwen3.8 A95B demonstrated a significant improvement of +18.18 pp when using reasoning prefills, increasing its overlap from 16.79% to 34.97%. Kimi K3, while having the highest overall overlap with GPT-5.5 Pro (31.11% unprefilled, 35.65% with prefill), saw a smaller gain of +4.54 points from the prefill.

    日榜第 1 名0 个来源热度 52
  2. DeepSeek V4-1 Flash is out

    The Deepseek V4.1 Flash model's parameter count has been clarified, with the "Backbone total" being approximately 552B parameters, as advertised. However, when including "Engram total" (196B) and other components like DSpark/MTP and Vision encoder, the "Everything in total" parameter count reaches approximately 763.21B. This breakdown helps to resolve confusion regarding the model's true size.

    日榜第 3 名0 个来源热度 46
  3. Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

    Cognition has launched its new SWE-2 model, which demonstrates strong performance in coding benchmarks, rivaling models like Fable 5.1 and GPT-Astra. The SWE-2 model achieved 50.0% on Main 50.0 %, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1, and 27.3% on Terminal-Bench 4. Cognition also shared that it uses a length-weighted reward baseline, introduced since SWE-1.6, to stabilize training and reduce gradient variance.

    日榜第 9 名0 个来源热度 36
  4. Tell HN: OpenAI keeps re-enabling the 'allow training' setting

    A user on Hacker News reported that OpenAI repeatedly re-enables the 'allow training' setting without their consent. They noted that despite carefully disabling the setting multiple times, it was found to be re-enabled upon subsequent checks. The user advises others to verify their settings, as this suggests a potential issue with OpenAI's handling of user preferences regarding data training.

    日榜第 11 名0 个来源热度 36
  5. Another researcher says OpenAI trained on conversations, then claimed breakthrou

    A researcher has accused OpenAI of training its models on existing conversations and subsequently presenting the results as a breakthrough. This claim suggests that the company may have leveraged pre-existing conversational data to achieve its advancements, rather than developing entirely novel capabilities from scratch. The accusation highlights ongoing concerns within the AI community regarding the transparency and originality of AI development processes.

    日榜第 12 名0 个来源热度 36
  6. Artificial Intelligence: A threat to humanity?

    A former Anthropic researcher is raising concerns about artificial intelligence, describing it as an existential threat to humanity. This alarm is being sounded amidst growing discussions about the potential dangers AI poses. The report from KTLA's John Fenoglio highlights these warnings, emphasizing the serious implications of advanced AI technologies for the future of mankind.

    日榜第 21 名0 个来源热度 31
  7. Is ChatGPT 6 Really That Good? See for Yourself
    日榜第 23 名0 个来源热度 31
  8. IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license

    IBM has released the SOTA Granite Time Series PatchTST-FM-r2 model, featuring a commercial-friendly open license and high-performance zero-shot forecasting. This model uses self-attention to capture long-range relationships and convolution for local temporal structures, allowing attention to focus on longer horizons. Its conformer blocks utilize alternating convolution kernel sizes of 3 and 5, in a repeating pattern {5, 5, 3, 3}. The pipeline generates future forecasts, including requested quantiles, without fine-tuning or task-specific model fitting.

    日榜第 24 名0 个来源热度 30
  9. GPT-6 Astra: The next generation in intelligence for work
    日榜第 29 名1 个来源热度 29

02Agent 与工具12 篇

  1. OpenAI Agents API

    The OpenAI Agents API provides applications access to the Codex harness, enabling agents to use tools like programmatic_tool_calling, MCP, and web_search. Agents can be configured with models such as "gpt-6-astra" and instructions for answering technical questions, delegating tasks to subagents. The API supports multi-agent capabilities with a maximum of 4 concurrent subagents and retains session state for continuous work. It currently offers data residency only in the United States and does not support Zero Data Retention (ZDR), even with a self-hosted sandbox.

    日榜第 2 名0 个来源热度 47
  2. GPT-6 Astra, looped transformers, and hidden reasoning

    OpenAI's GPT-6 Astra is generating significant interest due to its performance, particularly its looped transformer architecture and rumors of hidden reasoning traces. Astra reportedly achieves 99.9% on the ARC-AGI-3 benchmark, a substantial improvement over GPT-5.6 Sol's 7.8%. While this benchmark assesses logic puzzles and generalization, its performance on math, coding, and computer use benchmarks is considered more relevant to real-world applications. The looped transformer design involves passing intermediate representations through the same transformer blocks multiple times, with consistent weights across passes, an architectural tweak distinct from simply adding more blocks.

    日榜第 4 名0 个来源热度 44
  3. Detecting and countering misuse of AI: September 2026

    A November 2025 operating model for autonomous cyberattacks, initially linked to state-sponsored campaigns, has proliferated among various actors, including lone individuals. Publicly available frameworks like PentAGI automate the cyber kill chain, enabling more sophisticated attacks at greater speed and scale. Over 20 organizations, including government ministries, defense bodies, and diplomatic missions, primarily in Ukraine and Europe, were targeted. The targeting focused on Ukraine and military drone technology, with some exceptions in Southeast Asia and North Africa.

    日榜第 6 名0 个来源热度 40
  4. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

    A research paper titled "Procedural Graphs: Self-Evolving Execution Structures for LLM Agents" was published on arXiv.org on September 8, 2026. Authored by Yuxing Lu, this 36-page document, including references and appendices, explores Artificial Intelligence, Computation and Language, and Multiagent Systems. It is identified as arXiv:2609.09153 [cs.AI] and is available in its first version (v1).

    日榜第 13 名0 个来源热度 36
  5. Show HN: Self-hosted company OS, Claude Code and Codex agents in departments

    OtoDock is a self-hosted company OS that incorporates collaborative agents like Claude Code and Codex within departments. It operates under the Functional Source License, v1.1 (FSL-1.1-Apache-2.0), which permits use, modification, and redistribution for non-commercial purposes. Notably, each version of OtoDock automatically transitions to a plain Apache 2.0 license two years after its release.

    日榜第 15 名0 个来源热度 35
  6. The OpenAI Scandal Is Getting Bigger By The Minute!

    The OpenAI scandal is escalating, with discussions also covering Palmer Luckey's sanctions by China and significant attacks during the Iran War. Other topics include China's master plan, Houthi attacks on Saudi Arabian oil, and a massive education decline in OCED nations. The conversation further delves into unhealthy habits, the LG TV Spygate, Tucker Carlson's stance on algebra, and a Haaretz report on Bibi's October 7th response, alongside AI advancements in medical research.

    日榜第 16 名0 个来源热度 34
  7. OpenAI Just Solved the Biggest Problem in Mathematics

    OpenAI, in collaboration with Tristan Buckmaster and Levent Alpöge, has achieved a significant breakthrough concerning the Navier–Stokes Millennium Prize Problem. This development is highlighted as the most substantial result seen to date in the intersection of mathematics and artificial intelligence. The video discusses the implications and importance of this achievement, suggesting a major advancement in solving one of mathematics' biggest challenges.

    日榜第 17 名0 个来源热度 34
  8. OpenAI JUST solved math....

    OpenAI announced that 10,000 AI agents solved the Navier–Stokes Millennium Prize Problem in 88 hours, sparking a dispute over credit and the role of human mathematicians. The announcement has led to discussions about earlier research, including work by Tristan Buckmaster and Levent Alpöge, and OpenAI's response to the controversy. This event raises questions about the future of scientific discovery and the impact of AI in research.

    日榜第 27 名0 个来源热度 30
  9. Muse – Meta’s personal AI agent
    日榜第 28 名1 个来源热度 30

03应用落地5 篇

  1. Introducing ChatGPT for Financial Services
    日榜第 14 名1 个来源热度 35
  2. Get ready for the game with new football features in Search
    日榜第 25 名1 个来源热度 30
  3. 3 ways to prep for your next big race with Search
    日榜第 26 名1 个来源热度 30
  4. Now everyone can put data to work
    日榜第 30 名1 个来源热度 29

04政策&风险4 篇

  1. What will our economic future look like?

    The economic future with AI is uncertain, with potential for unprecedented growth or widespread unemployment. An economic scenario explorer, currently Version 1.0, simplifies complex reality by focusing on key forces and omitting others like policy responses or financial disruptions. This model indicates that by 2026-2030, 62.2% of knowledge workers and 37.8% of other workers will be affected, with 2.5% displaced and 1.8% crossed over, while 59.7% remain and 0.7% still need to move.

    日榜第 5 名0 个来源热度 41
  2. More questions about whether researchers can trust OpenAI with unpublished math

    Researchers are increasingly questioning whether they can trust OpenAI with their unpublished mathematical work. Concerns are being raised across various platforms, including mathstodon.xyz, x.com, and bsky.app, regarding the security and confidentiality of sharing sensitive, unreleased research with the AI company. This discussion highlights a growing apprehension within the academic community about data privacy and intellectual property when interacting with large language models and their developers.

    日榜第 10 名0 个来源热度 36
  3. Clearview AI Is Testing an AI Tool That Would Let Cops Unearth Your Life Online

    Clearview AI, a controversial surveillance firm, is testing an AI tool that could allow law enforcement to uncover individuals' online lives. The company, which faced lawsuits and regulatory investigations for scraping over 3 billion images from various websites by 2020, has since expanded its database to over 70 billion images. Its technology is reportedly used by more than 2,000 law enforcement agencies. A case, United States v. Sant, highlighted concerns about the accuracy of Clearview's reports, with defense attorneys alleging misidentification and questioning the validity of results labeled "Accepted by Guy Gino."

    日榜第 19 名0 个来源热度 33
  4. A.I.'s threat to humanity given new consideration in Congress

    Following an Anthropic employee's resignation over the threat of artificial intelligence to humanity, members of Congress are seriously considering AI regulation. Rep. Greg Casar discussed with Jen Psaki legislation he is introducing with Senator Bernie Sanders, called the Ban Superintelligence Act, to prevent AI from becoming dangerously out of control.

    日榜第 22 名0 个来源热度 31