Why not just have one less feature before softmax? [D]
A discussion explores reducing the number of inputs to a softmax function from N to N-1. This is based on the observation that softmax outputs have N-1 degrees of freedom because their sum must equal one. The proposal suggests calculating the last logit as the negative sum of the others, assuming logits before softmax sum to zero. This could theoretically remove "unnecessary" parameters and potentially speed up model convergence, though the practical benefit might be negligible.
This discussion uniquely explores a theoretical optimization of softmax inputs from N to N-1, unlike typical discussions focusing on its output properties.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 29, 2026, 05:26 UTC
IngestedOffset at this time: UTC+0Sep 29, 2026, 08:00 UTC
- Published
- Sep 29, 2026, 05:26
- Ingested
- Sep 29, 2026, 08:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Sync delayed
Tier is a per-source editorial setting, not a per-item score.
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them may there be) and maybe speed up model convergence a little (my intuition might be wrong about that). Is there any good reason not to do it besides any benefit being negligable in almost all situations?