LLM regression in reading comprehension?
一位用户观察到大型语言模型(LLM)在阅读理解方面可能存在退步,他们通过免费使用Qwen 3.8 max和Gemini 3.1 PREVIEW Temp 1.0等模型进行对比。尽管Qwen表现良好但速度较慢,而Gemini则显得有些过时。该用户对开源模型K3很感兴趣,并提到Google AI Studio为Gemini提供了慷慨的免费使用额度。他们正在寻求其他“智能”模型的建议。
- 发布
- 2026年9月6日 10:10
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
时间以 UTC 显示
更多信息
I only use free tiers of these large models to offset compute while my own system runs and for "different" points of view, since what pops ups suggestions seems to vary a lot sometimes, even when building based on the latest research.
But now I've really struck out with GLM 5.3. So far it feels like an regression over 5.2. It has a hard time reading and following instructions, and is somewhat overly certain in it's statements. I worked on a project recently with it but it became unbearable. From a clean slate the first message can be okay and have great research and ideas but it just veers off course almost immediately.
I use Qwen 3.8 max and Gemini 3.1 PREVIEW Temp 1.0 as competing alternatives or as an ensemble to judge overall quality. Gemini is getting a little out of date (flash 3.8 seemed promising) but Qwen has been great so far, but a little slow and maybe overbearing.
Anyone else having problems? Or suggestions for these top "intelligent" models? I haven't been able to access K3 even though its open source, was impressed with the older models so would be neat to try for free. Also Google AI studio is what i use for free for the gemini stuff, probably pretty well known, but the free tier is pretty generous