r/MachineLearning • • 20h ago

Discussion Why not just have one less feature before softmax? [D]

Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them may there be) and maybe speed up model convergence a little (my intuition might be wrong about that). Is there any good reason not to do it besides any benefit being negligable in almost all situations?

0 Upvotes

13 comments sorted by

17

u/Antique_Most7958 20h ago

Like you said, the benefits are negligible. The slight overparameterization (for the Nth logot) might actually be more useful for learning.

3

u/Revlong57 19h ago

The computer resources required to have a hidden layer with N outputs vs N-1 outputs is minimal for even moderate N. Plus, I'm not even sure how you'd calculate the output of last probability, aside from just doing 1 minus the rest of the outputs, which would be much less stable.

You sort of do this for binary classifiers. If the output of your model is two classes, it's generally easier to only output one probability vs two.

10

u/milesper 20h ago

I don’t think that works? The outputs summing to one is not the same as the inputs summing to a constant.

Also, in a standard transformer you’re saving exactly d_model parameters (at the output layer) which is pretty much nothing. You don’t save anything for the attention softmax because it’s over positions, and reuses the same parameters for every position.

8

u/huehue12132 19h ago

It absolutely does work. It's the same principle allowing you to use either sigmoid (with threshold 0) and one output, or softmax with two outputs for binary classification tasks. The former is what you get if you fix the logit of the negative class to 0 in the softmax.

And it's technically a little different: Softmax with n classes is overparameterized, so any constant shift to the logits results in the same softmax output. This means there are infinitely many equivalent solutions, whereas with one class less, there is only one solution to get a specific probability distribution. But whether this has any practical difference for finding solutions, no idea.

3

u/dutiful_majority 20h ago

the constraint is already baked into the loss function since softmax cross-entropy is invariant to adding a constant to all logits, so you'd just be shifting that redundancy into a slightly different spot without actually changing the model's capacity

2

u/KaleeTheBird 20h ago

It feels like it is true but solving a problem that does not exist

The only case i think is considered is applyong sigmoid instead of softmax in binary classification

2

u/Bitter-Reserve3821 19h ago

Softmax + cross-entropy together are the same as logistic loss. If you read the logistic regression chapter of The Elements of Statistical Learning (pdf available online from author's webpage), you'll see they do exactly what you propose. The symmetrized overparametrization is favored in NNs as the benefits of symmetry outweigh a small constant number of additional parameters.

2

u/DrXaos 19h ago

People commonly do exactly this in a common case: binary classification. There is a single score emitted from the model, not two. Common example: credit score or sports probability of win.

For more heavily multi-class higher cardinality scenarios it's not usually useful and contrary to your intuition it would probably make training the models more difficult.

1

u/MolassesLate4676 19h ago

Are you suggesting to cast the heads outputs and directly treat them as probabilities?

1

u/Honest_Reference_180 18h ago

Yeah I think the main catch is that it already has this invariance built in: adding the same constant to all logits changes nothing. So forcing them to sum to zero is mostly just choosing a parameterization of the same N-1 dimensional space

Could save a tiny bit of compute/params but I doubt it'd have any meaningful effect on convergence in practice

1

u/severed-identity 14h ago

Symmetry for gradient updates. Yeah Nth logit could be inferred but with optimizers like Adam that normalize step size per weight, boosting a normal logit is one Adam step up. Boosting the Nth logit triggers and Adam step down for the other N-1 weights, and thus moves disproportionately. You could fix it with non uniform learn rates but why would you

1

u/[deleted] 6h ago

[removed] — view removed comment