Why not just have one less feature before softmax? [D]
A discussion explores reducing the number of inputs to a softmax function from N to N-1. This is based on the observation that softmax outputs have N-1 degrees of freedom because their sum must equal one. The proposal suggests calculating the last logit as the negative sum of the others, assuming logits before softmax sum to zero. This could theoretically remove "unnecessary" parameters and potentially speed up model convergence, though the practical benefit might be negligible.
This discussion uniquely explores a theoretical optimization of softmax inputs from N to N-1, unlike typical discussions focusing on its output properties.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月29日 05:26 UTC
收录当时偏移:UTC+02026年9月29日 08:00 UTC
- 发布
- 2026年9月29日 05:26
- 收录
- 2026年9月29日 08:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them may there be) and maybe speed up model convergence a little (my intuition might be wrong about that). Is there any good reason not to do it besides any benefit being negligable in almost all situations?