返回
RCreddit.com
16
·13小时前·开发者社区 · RSS

A hunch: Qwen3.8-27B's general knowledge got pruned (good, if true)

查看原文
模型发布

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

这条记录涉及编程工具或代码能力更新,适合开发者评估工作流变化和可复用价值。

AI 摘要

一位Reddit用户推测Qwen3.8-27B的通用知识可能已被修剪。该用户通过测试图像提示发现,之前的Qwen3.6-27B和Qwen3.6-35B A3B版本在识别其家乡历史地点时表现不一致,这表明模型在该特定领域的知识基础本就薄弱。用户认为,如果知识修剪属实,这可能是一个积极的进展,尽管他们承认样本量小且测试方法不够严谨。

I'm always testing an image prompt with a picture of a historic place in my hometown – a small but well known 250,000 people town in Germany. I'll just ask the model, in which City this photo has been taken.

With the 3.6 generation of both the 27B and the 35B A3B variants, the models sometimes got the right answer and sometimes they didn't. So the signal for this particular knowledge was already weak.

The 35B variant got it right more often but at least, the models reasoning showed my City most of the times, even if it hallucinated the wrong final answer.

Both models could be easily nudged to the right answer with a few hints and then produced some little extra insight about the history or scene and its surroundings, that was mostly true.

Qwen3.8-27B on the other hand barely knows the city at all and has absolutely no clue about related popular, historic facts regarding the scenery or the surrounding buildings.

Nudging isn't very fruitful as well and if told the real name of the city, reasoning shows, that the model only agrees, because the user says so.

I have the feeling, that Qwen labs maybe pruned useless general knowledge for more coding knowledge and agentic skill.

All models ud q4_k_xl variants, image-min-tokens 2048, with and without reasoning.

Anyone else with this feeling?

Disclaimer: My hunch could be very well absolute bullshit. Sample size way to low and methodically sloppy af.

A hunch: Qwen3.8-27B's general knowledge got pruned (good, if true) · BuzzRadr