Skip to content

AI Pulse

LIVE18/20
reddit.com

BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B

AI summaryAREX-2 is a 27B-parameter agent model from the Beijing Academy of Artificial Intelligence (BAAI), based on a Qwen3.8-compatible multimodal architecture. It features long-horizon self-improvement, feedback-driven reflection, and cross-domain performance, learning to refine solutions over multiple test-time rounds. Trained on machine-learning and algorithmic-programming tasks with verifiable feedback, AREX-2's self-improvement behavior transfers to deep research, sustaining productive iteration as the task budget grows.

reddit.com

DeepSeek now trained on Ascend 950

AI summaryDeepSeek is now training its models on Ascend 950. This development follows a statement made 26 months ago by Liang Wenfeng, who emphasized the necessity of stepping onto the frontier. The move indicates DeepSeek's commitment to utilizing advanced hardware for its model training, aligning with the earlier sentiment about pioneering new territories in technology.

reddit.com

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes [R]

AI summaryA new method, CO2Jump, for concurrent image understanding and generation, utilizes Self-Correcting Coupled Markov Jump Processes. Evaluated on image editing, maze solving, and nonograms, it introduces datasets like JEdit-1M, JMaze-200K, and JNono-200K. CO2Jump demonstrated monotonic improvement in both editing quality and grounding across 8–512 sampling steps, outperforming other samplers in tasks requiring joint accuracy of textual answers and generated images.

reddit.com

LessThink-Qwen3-4B: the same model, with far less thinking [P]

AI summaryA developer has post-trained the Qwen3-4B model, creating "LessThink-Qwen3-4B" which significantly reduces token usage for reasoning by 44% while maintaining its original knowledge and answer style. This entire process was achieved using a single GPU. The developer invites interested individuals to explore this new model further on their website.

reddit.com

1248 GB/s on 5060ti (+40%) with +5500 memory overclocks

AI summaryA Reddit user achieved 1248 GB/s on a 5070ti (corrected from 5060ti) with +5500 memory overclocks, representing a 40% increase. This was made possible by a new unlock for higher memory overclocks called mlock, which reportedly allows GDDR7 to have significant headroom. Such advancements could greatly benefit local inference, particularly for token generation.

reddit.com

Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine

AI summaryThe 'Strata' inference engine significantly outperforms llama.cpp, achieving 51 tokens/second (t/s) with the ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/IQ3_XXS model at 43k context depth on a laptop with 12GB VRAM and 64GB RAM. This speed is more than double the 23 t/s reached by stock llama.cpp using the same quantization. The developer fixed initial bugs, making it possible to run frontier models locally on modest hardware, demonstrating impressive advancements in local inference capabilities.

reddit.com

Mona Lisa SVG Challenge Opus 5.5 vs Sol 6.1

AI summaryA developer community is hosting a "Mona Lisa SVG Challenge" with strict rules. Participants are forbidden from using reference images, tracing, vectorizing, or auto-converting any image. All shapes must be self-authored, either directly or through custom code that doesn't take an image as input. The final SVG file must be under 100 KB, with a canvas size of 600 x 900 (viewBox="0 0 600 900"). The challenge compares "Opus 5.5" and "Sol 6.1" versions.

reddit.com

Opus 5.5 feels way dumber after today’s outage, anyone else noticing this?

AI summaryUsers are reporting a significant decrease in the performance of Opus 5.5 following a recent outage. One user, who previously found Opus 5.5 impressive for legal work, writing Minecraft mod scripts, and troubleshooting NAS setups, noted that after Claude went down and the service returned, Opus 5.5 felt "noticeably, almost ridiculously stupid," like a "quantized version or a completely different model." The user emphasized that the drop in quality has been "really really really really noticeable."

reddit.com

GPT-6.1 Sol beat Pokemon Red's first gym in 244 turns, a new record on PokeBench, for $2.60

AI summaryGPT-6.1 Sol set a new record on PokeBench by defeating Brock in Pokémon Red's first gym in 244 turns. This achievement, costing $2.60, surpassed the previous record of 246 turns held by GPT-6 Astra, which cost $13.49. The run was conducted with a fresh save and a 1000-turn budget, with milestones read directly from game memory, ensuring an objective evaluation without human or model judgment.

reddit.com

Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models

AI summaryQwen-family LLMs are increasingly serving as the foundational architecture for modern audio models, as evidenced by a recent analysis of over 100 audio models. Specifically, 32 audio model families utilize a Qwen-family architecture, with 20 of these explicitly employing the Qwen3 LLM. This trend highlights Qwen's growing prominence as the most common language backbone in this domain, with further analysis detailing which building blocks power various audio model types.

reddit.com

exl3 now in ninfer-ext

AI summaryThe ninfer-ext fork of ninfer now supports exl3, with two models released. This development was announced on reddit.com by a developer from the dev_community. Users can find the models at huggingface.co/jabbatheduck/ninfer-ext-models, indicating an expansion of capabilities for the ninfer-ext project.

YouTube·

A OpenAI lançou +20 NOVIDADES e mudou o ChatGPT pra sempre (DEVDAY)

AI summaryOpenAI's DevDay 2026 introduced over 20 new features, significantly changing ChatGPT. Key announcements included GPT-6.1 Sol, a new $500 Pro plan, and GPT-6 Astra Ultrafast. Other updates featured dots for always-on agents, Codex in the cloud with a new CLI, Code Review, Codex Security Cloud, and the Decisions API. ChatGPT also gained plugins, Space, Pages, collaborative slides, integration with Slack and Microsoft Teams, a meeting plugin, shareable profiles, and an OpenAI Marketplace.

reddit.com

My tldr for OpenAI dev day

AI summaryA developer noted that Anthropic consistently releases new features monthly, contrasting this with OpenAI's dev day. They highlighted Claude Code's significant improvements over codex in recent months. The developer also mentioned discontinuing the use of Codex in favor of pi due to its superior speed and token efficiency.

reddit.com

Best model for blender?

AI summaryA user is seeking recommendations for a local model to generate game assets or 3D printer models, specifically for use with Blender. The primary requirements are that the model should fit within 64GB of RAM and perform well, with speed not being a critical factor. The user is exploring options within a developer community to find suitable suggestions for their creative projects.

reddit.com

I don't understand the constant bashing of GPT 6/6.1 Sol's performance relative to 5.6 and Astra. Why are we ignoring the elephant in the room, the absurd price + efficiency gains? It costs half as much as Sonnet 5.5, 1/3rd of 5.6 Sol, 1/5th of Opus 5.5, and 1.7th of Astra.

AI summaryA discussion on reddit.com questions the criticism of GPT 6/6.1 Sol's performance compared to 5.6 and Astra, highlighting significant price and efficiency gains. The user points out that GPT 6/6.1 Sol costs half as much as Sonnet 5.5, one-third of 5.6 Sol, one-fifth of Opus 5.5, and 1.7 times less than Astra, suggesting these cost benefits are being overlooked.

reddit.com

Is Opus 5.5 entering a “nerfed” phase? LiveNerf baseline update

AI summaryAn open-source project has been released to independently measure changes in Opus 5.5's performance following its release. This tool tracks daily performance using established benchmarks like SWE-bench and SciCode, employing a standardized scientific methodology to ensure measurable and reproducible changes over time. The project aims to determine if Opus 5.5 is entering a "nerfed" phase, allowing the community to monitor its performance evolution.

reddit.com

That didn't backfire at all

AI summaryA developer community discussion highlights a company's product naming strategy for GPT-6 models. Initially, the company could have used simple names like GPT-6 Luna, GPT-6 Terra, GPT-6 Sol, and GPT-6 Astra. Instead, they opted for more complex names such as GPT-6 Terra Sol and GPT-6 Sol Astra Minor, which seemingly backfired, leading to a simplified release of GPT-6.1 Sol. The community expresses gratitude for continued competition in the market.

reddit.com

Opus 5.5 still leads on AA intelligence, but GPT-6.1 Sol shifts the cost frontier

AI summaryA comparison of reasoning effort settings reveals that Opus 5.5 maintains its lead in Artificial Analysis's Intelligence Index. However, GPT-6.1 Sol significantly alters the cost frontier. The analysis considers costs per benchmark task, encompassing both caching and reasoning, rather than subscription prices. The data for this comparison is sourced from artificialanalysis.ai.

reddit.com

What are chinese labs doing differently?

AI summaryChinese AI models are reportedly improving rapidly despite lower spending compared to American labs. One suggested explanation is their acquisition of handcrafted expert data from US companies such as SurgeAI and Mercor. Some believe that access to such data should be restricted, similar to chips and GPUs, to help the US maintain its lead in the AI race.

reddit.com

My expectation were low but holy crap they really did get caught off guard by Anthropic.

AI summaryA Reddit user expressed significant disappointment with OpenAI, stating that the company was caught off guard by Anthropic. The user likened OpenAI's performance to a "generational Atlanta Hawks Falcons level fumble," particularly in the coding domain. This perceived misstep has led the user to believe that OpenAI has conceded ground to Anthropic until its next release cycle, even before the Fable 5.5 launch, resulting in an "easy cancellation" for them.

reddit.com

Recommended replacements for glm 4.7 flash

AI summaryA user on reddit.com is seeking recommendations for replacements for GLM 4.7 flash, which they find performs better than Qwen 3.6 35 a3b, especially for tool calling and world knowledge, despite its age. They are using a strix halo and are looking for a newer model that is comparable in size and performance to GLM 4.7 flash, suggesting their Qwen setup might be incorrect.

reddit.com

Inference Engines will become a series of one-offs

AI summaryInference engines are predicted to become a series of one-off implementations, such as ninfer, dwarfstar, Splash, llamAmpere, and gufo. This is because tasks like optimizing "tok/s go up" for a specific hardware/model combination are fully specified, making them ideal for 100% autonomous AI implementations with trivial correctness tests, eliminating human bottlenecks. A separate growth dimension involves projects like Freetoken and BeeLlama, which frontrun general inference engine features, though these are less model- or hardware-specific.

OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less

AI summaryOpenAI has launched GPT-6.1 Sol, a new model that reportedly offers intelligence levels nearly matching GPT-6 Astra, but at one-fifth the cost for input and output tokens. Unveiled at OpenAI’s DevDay, GPT-6.1 Sol demonstrates improved factual accuracy, especially with difficult prompts, reducing factual errors from 11.4% to 7.7% at low reasoning effort. Across all reasoning settings, its error rate remains within 1.9% of GPT-6 Astra, making it a cost-effective alternative for agentic coding, computer use, and professional work.

huggingface.co

NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction

AI summaryNVIDIA Kumo Tabular sets a new accuracy-efficiency frontier for tabular prediction. The model learns to predict labels of remaining rows using cross-entropy loss for classification and quantile loss for regression. Training occurs in three stages, starting with tables of 1,024 rows and up to 100 columns, then varying context from 400 to 10,240 rows, and finally extending to 60,000 rows. Kumo Tabular-Small/Medium/Large saw approximately 35/71/137 million artificial tables.

reddit.com

Will Nvidia Vera Rubin actually make LLM pre-training faster? And are 10T+ parameter models next?

AI summaryA discussion on Reddit questions whether Nvidia's Vera Rubin platform will significantly accelerate LLM pre-training, noting that many advertised gains appear linked to low-precision formats and inference rather than pre-training. The conversation also explores the potential for 10T+ parameter models, considering if data, power, and cost limitations will shift focus towards Mixture-of-Experts (MoE) and improved data quality over simply increasing model size.

reddit.com

We finally start praising Opus 5.5 and Claude goes down

AI summaryA developer community on reddit.com noted that after praising Opus 5.5, Claude experienced downtime. The sentiment expressed was that Anthropic, the developer of Claude, seemed to have intervened once users were content with Opus 5.5, leading to Claude's unavailability. This suggests a perceived correlation between the positive reception of one AI model and issues with another from the same developer.

reddit.com

Goat Riding a Tractor Compare Models - turns out that level of effort Max seems to be a deciding factor (mostly)

AI summaryA user compared different models for generating an image of a "Goat Riding a Tractor," noting that the effort level of the "Max" version was a significant factor. They spent an additional $2.78 in usage credits to try the Fable Max model, hoping for a better result than the Medium version. The user manually assembled the final image because they ran out of tokens and were unwilling to spend more credits on a Claude-stitched version, eager to share the results.

reddit.com

Sonnet 5.5 did this. Opus 5.5 quality with half price.

AI summaryA 30-second, 1080p, 60 fps video featuring Snoo, kinetic type, a subreddit marquee, and various Reddit-themed animations was created using Sonnet 5.5. The project involved 14 sub-agents and consumed 739k tokens, 101.6M cache reads, and 3.07M cache writes, costing $35.40 with Sonnet 5.5. This was approximately 43% cheaper than if Opus 5.5 had been used for the same tokens, primarily because cache reads, which constitute most of the token usage, cost the same for both models.

reddit.com

OpenAI researcher: "[Navier-Stokes] surprised the fuck out of us." ... "Last 3 months = hell" ... "Suddenly we weren’t dealing with just a small jump in capabilities; we were talking about a different sport altogether."

AI summaryAn OpenAI researcher stated that the Navier-Stokes project "surprised the fuck out of us," describing the last three months as "hell." They noted that the team was suddenly dealing with more than just a small jump in capabilities, suggesting a fundamental shift in their work, akin to a "different sport altogether." This indicates a significant and unexpected advancement in AI capabilities related to the Navier-Stokes equations.

reddit.com

Opus 5.5 vs Sonnet 5.5: 3D steampunk whale modeling

AI summaryA user compared Claude Opus 5.5 and Sonnet 5.5 for 3D steampunk whale modeling, noting that Sonnet 5.5 was released during the project. Both models were used with structured prompts, and the whales were created in Blender before being transferred to three.js for browser display. Opus 5.5 processed 2.50M output tokens and 342M total tokens, costing approximately $156, while Sonnet 5.5 handled 1.68M output tokens and 297M total tokens, costing around $109.

reddit.com

Reflection 70B was released two years ago (September 2024)

AI summaryReflection 70B, an open-source LLM, was released in September 2024, two years prior to the current discussion. It was announced as a model that supposedly "destroyed" GPT-4o, making it a significant invention in the LLM space. This release is highlighted as a particularly cool development, even when compared to other notable LLMs like jev, OpenClaw, or TurboQuant.

reddit.com

Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide

AI summaryA user successfully ran Qwen 3.8 27B Q4 XS with 100K context on a 16GB AMD RX 7800 XT GPU, achieving approximately 30 t/s decode speed. This setup, which many believed infeasible on 16GB VRAM, utilized specific llama-server parameters. Key configurations included `--n-gpu-layers 999`, `--ctx-size 100096`, `--cache-type-k q8_0`, and `--cache-type-v q5_1` to optimize performance and memory usage for the Qwen3.8-27B-UD-IQ4_XS.gguf model.

hackernews·Breakout · 2.3×

Language models for text classification: From bag-of-words to Jev

AI summaryThe Jev AI model has recently gained significant attention within technical communities. This model is part of a broader evolution in language models for text classification, moving beyond earlier approaches like bag-of-words. Key advancements in recurrent neural networks (RNNs) include Long short-term memory (LSTM) networks, introduced in 1997, and gated recurrent units (GRUs), introduced in 2014, which utilize learned gates for information management. More recently, xLSTM: Extended long short-term memory was introduced in 2024.

openai.com4 sources · openai.com / simonwillison.netBreakout · 6.3×

Introducing GPT-6.1 Sol

AI summaryOpenAI has launched GPT-6.1 Sol, offering "Near-Astra intelligence for a fifth of the price." This new model is available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex, and developers can access it via the OpenAI API as gpt-6.1-sol. API pricing is set at $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. An Ultrafast version, offering up to 8x faster token generation, is also planned for Codex.

wired.com

The Next Evolution of AI Is Learning From Your Dodgy Gaming Skills

AI summaryA British startup is using video game data to train new AI models, aiming to overcome the lack of real-world physics data for "world models." Unlike large language models (LLMs) trained on vast text corpora, world models require combined visual and action data. The startup has licensed nearly 1 million hours of data from game studios and plans to compensate individual players in the future, leveraging even unskilled player actions to teach AI about navigating 3D environments and manipulating objects with appropriate force and torque.

huggingface.co

Accelerating vision-language models with LFM2.5-VL-DSpark

AI summaryHugging Face has released an experimental DSpark draft model for their LFM2.5-VL-3B vision-language model. This new model, LFM2.5-VL-DSpark, incorporates a speculative decoding path to accelerate performance. It achieves a significant speedup with only a minimal increase in memory footprint, while maintaining the original output quality. This advancement aims to accelerate vision-language models on edge devices and beyond.

openai.com

Introducing MentalHealthBench

AI summaryOpenAI has introduced MentalHealthBench, developed with over 80 mental health experts from 22 countries, to improve AI responses in sensitive conversations. This initiative builds on previous work like HealthBench and aims to ensure AI models prioritize user safety and well-being, especially as over a billion people use ChatGPT weekly. Enhancements include strengthened responses, expanded access to crisis resources, and the addition of Trusted Contact and ChatGPT for Teens with extra protections.

openai.com

Better prompt caching for GPT-6

AI summaryGPT-6 allows persistent agents to handle complex tasks, utilizing API requests that build upon previous turns. OpenAI caches shared context to reduce response times and offers developers up to 90% discounts on cached input tokens. A "cache_miss" can occur if "tools_changed," resulting in missed tokens. Developers can improve their setup by following the prompt caching guide or using Codex.

openai.comBreakout · 6.0×

Introducing GPT-6 Sol and Luna

AI summaryOpenAI introduced GPT-6 Sol and Luna on September 29, 2026, as more cost-efficient alternatives to their predecessors. While GPT-6 Astra remains the top model for computer use, GPT-6 Sol achieves a similar score to Claude Opus 5 on OSWorld 2.0 offline at 80% lower cost. GPT-6 Luna (max) also surpasses GPT-5.6 Sol (medium) at one tenth of its cost, offering significant performance-to-cost improvements.

huggingface.co

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

AI summaryThe UK AI Security Institute (AISI) is utilizing EvalEval's infrastructure to openly share evaluation results, enhancing the reproducibility and verifiability of evaluation science. These results encompass six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4, alongside data from Cyber CTFs and The Last Ones cyber evaluations. This release supports AISI's paper, "How Inference Compute Shapes Frontier LLM Evaluation," which investigates the impact of inference-time compute and evaluation protocols on benchmark performance.

huggingface.co

Transformers now runs llama.cpp quants

AI summaryHugging Face Transformers now supports running GGUF models efficiently, allowing users to load checkpoints sized for their laptop's memory using familiar Transformers APIs. This integration enables generating text on personal machines by picking a GGUF from the Hub and loading it with from_pretrained. The underlying kernels can also be integrated into other Transformers models and loading workflows, potentially extending to computer vision, audio, and multimodal models, reusing compatible attention, normalization, and matrix multiplication kernels.

openai.com

Advisory Group on Mathematics and Artificial Intelligence

AI summaryOpenAI has been training a new internal model since August 28, which has successfully resolved over 100 long-standing open problems in mathematics, including the Navier–Stokes Millennium Prize problem. The rapid progress of this model has surprised OpenAI's mathematicians, prompting internal discussions on how to best inform and prepare the broader community for these advancements. Melanie Matchett Wood from Harvard is involved in these discussions.

huggingface.co

tokenizers v1: encode, decode and scaling, measured

AI summaryTokenizers v1 significantly improves performance, encoding text 3 to 30 times faster than v0.23 on an Apple M4 Max with a single thread, depending on the model family (e.g., t5-base to gpt2). It also demonstrates strong scalability, achieving 76% of linear scaling across eight workers. Crucially, v1 maintains exact token ID consistency with the previously released library, ensuring no changes in output despite the performance enhancements.