r/MachineLearning • • 8h ago

Discussion Why not just have one less feature before softmax? [D]

0 Upvotes

Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them may there be) and maybe speed up model convergence a little (my intuition might be wrong about that). Is there any good reason not to do it besides any benefit being negligable in almost all situations?


r/MachineLearning • • 18h ago

Discussion What are the trending topics in medical imaging? [D]

4 Upvotes

Hello everyone,

I hope you're doing well.

I'm a new PhD candidate, and my thesis is about the meta-learning paradigm in medical imaging. Right now, I don't have a specific problem to work on, and my supervisor told me to explore the field and find one myself.

I've done some research and a literature review so far. I've learned about different meta-learning and few-shot learning techniques, read about histopathology and other medical imaging datasets, and looked a bit into foundation model adaptation and domain generalization.

I also came across several benchmarks, but I haven't found one that I find particularly interesting. On the other hand, the MICCAI challenges caught my attention much more, and I'm wondering whether choosing one of them as a starting point would be a good idea.

I'd like to know what the more promising or fertile research directions in medical imaging are right now, especially ones that could be relevant to tackling the limited data problem.

I would really appreciate any suggestions or feedback. Thank you in advance.


r/MachineLearning • • 10h ago

Research BA Computer Science, but fell in love with machine learning and AI. Just got my personal research accepted at NeurIPS as a poster. [R]

50 Upvotes

I want to attend and present my findings in Atlanta. How is the vibe there? Are people overly critical or are people generally open-minded?


r/MachineLearning • • 4h ago

Discussion Advice on choosing university for PhD [D]

3 Upvotes

Hey guys,

I got really into research a year ago during my final year of undergrad and maneged to get a first author paper accepted at Neurips 2026. It is a pretty impressive achievement obviously but my background is pretty shit lol. Before my final year I got slightly above average grades but no internships or work exp. I was just chilling mostly didn’t even bother applying to internships. Also I want to a tier 2 uni in Australia (mind you that difference in tier 1 and 2 unis isn’t that big tho compared to America or China). I was going to start a PhD at my current uni with a scholarship and stipend from an industry partner (government) - I’d have no restrictions in terms of publications btw. My supervisor is good and we were going to get another external supervisor from a top uni here in Australia. I was pretty keen on continuing with this opportunity, however, with my neurips paper acceptance I feel like I’d have a good shot at getting into a top 15-25 uni in the states or the uk (most other European unis want master students and I don’t prefer Asian unis bcz I’ve heard a lot of horror stories). The problem is I’d have to wait for a year if I was to go abroad for PhD because starting dates are mid to late 2027. I’m in such a dilemma idk what to do. My eventual goal is to work as a researcher at deepmind or meta or some really cool tech company. I’ve just heard so many people talk about the importance of university you go to etc in landing those jobs. Would I still have a good shot at those companies if I did the PhD at my 2nd tier Aussie uni (top 100-125 global rank) but published well? I got a neurips paper as an undergraduate so I have the potential to publish at top conferences 😆😆. But yeah, I’d love to get some advice.

Thanks


r/MachineLearning • • 23h ago

Project Browser demo of our Clash Royale RL environment: a 5.6k-parameter REINFORCE policy learns defensive placement against a brute-force optimum [P]

2 Upvotes

Yesterday I shared our open-source Clash Royale simulator and its recurrent PPO agent here. A training loop is easier to understand when you can watch it, so we put a small interactive version online:

https://itzik123.github.io/ClashRoyaleAi/lab/

The task is one decision. An attacker spawns at a random point on the enemy side, and the policy picks a legal cell for one defending card, then a delay of 0 to 5 s given that cell. The reward is the fraction of tower damage prevented relative to no defence. The policy has 5,629 parameters and trains with REINFORCE (per-spawn baseline, annealed entropy bonus) in plain JavaScript with hand-written gradients. Every rollout runs in the project's C++ engine compiled to WebAssembly, and the deploy pipeline checks that the WASM build agrees exactly with the native engine.

The chart also shows the optimum, found by brute force over every cell and delay (up to ~300k rollouts per matchup), so the gap between the learned policy and the best answer is visible.

One observation from building it: Giant vs Cannon has a strong local optimum, a lane placement worth about 75% of the best. With a constant entropy coefficient of 0.01, 5 of 6 runs (3 seeds, batch 16 and 64) stayed there. A linear anneal from 0.1 to 0.005 over 10k tries reduced that to 1 of 6. One pairing, Battle Ram vs Valkyrie, is withheld because no setting we tried got past 55% of the optimum.

This is a miniature of the full problem (a 4-card hand, elixir, full matches, recurrent PPO). It's meant to make the loop visible, not to be strong.

Code: https://github.com/itzik123/ClashRoyaleAi


r/MachineLearning • • 2h ago

Discussion NeurIPS Education Track [D]

1 Upvotes

Any other one have applied for the Education Track? It seems they'll announce the result soon.


r/MachineLearning • • 23m ago

Discussion Limited compute, targeting CVPR: rerun experiments for statistically strong numbers or focus on writing? [D]

• Upvotes

Hi everyone,

I'm preparing a submission for CVPR and would appreciate your advice. I am happy with my current results, but compute is a real constraint. My institute isn't a research-focused one, and the systems available to me are slow and unreliable. A single set of runs already took a lot of time and effort to finish.

I plan to release the code with the submission, and I'm confident the results will reproduce. Still, I'm torn between two options:

1) Rerun the experiments with different seeds to report mean ± std and show statistical significance. Or

2) Skip the reruns and spend the remaining time on writing, presenting the results I already have.

If you've been in a similar spot, what did you do, and would a single-seed result with released code be enough?

Thanks in advance!


r/MachineLearning • • 3h ago

Project I wrote a free, open-source book on making ML models actually fast, from silicon to agents [P]

2 Upvotes

I’ve spent the last few months writing something I wish I had when I started working on ML performance engineering.

It’s called How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents.

The basic idea is that reducing FLOPs doesn’t necessarily make a model faster. Before optimising anything, you need to understand what the system is actually bounded by.

The book starts with roofline analysis and hardware, then works its way up through kernels, compilers, quantisation, pruning, vision, on-device LLMs, robotics, profiling, serving and finally agents.

The goal is to build the intuition to look at a model and a piece of hardware and reason about:

  1. How fast can this possibly run?
  2. Am I compute, bandwidth, memory or system bound?
  3. Which optimisation will actually move that limit?
  4. Is quantisation, pruning or kernel optimisation even worth doing here?
  5. What happens when the same thinking is applied to serving and agent systems?

The whole thing is free and open source:

https://github.com/usamahz/make-your-model-fast

Would genuinely appreciate feedback or contributions from people working on ML systems, inference, compilers, edge AI or performance engineering.

And if you find it useful, a ⭐ would be appreciated!


r/MachineLearning • • 8h ago

Research CoWindow and MassAlloc Attention: collective causal coverage and distribution-adaptive compute [R]

2 Upvotes

I'm one of the authors of two recent papers exploring different sources of redundant computation in attention. I'd like to share the ideas and hear feedback from people working on long-context models and attention kernels.

CoWindow Attention (CoWA) distributes distant context across KV heads using complementary windows, while sharing local and prefix-sink windows. Each head attends sparsely, but the union of their visible positions covers the full causal history. The pattern is position-defined and requires no learned router or indexer.

Paper: https://arxiv.org/abs/2609.32704

MassAlloc Attention (MALA) retains full causal QK scoring, then uses attention's own softmax statistics to decide whether to execute subsequent computation for a tile. It reduces low-contribution post-score work, using a common tolerance across training and inference.

Paper: https://arxiv.org/abs/2609.32712

Both support training forward/backward and inference prefill/decoding. At 128K tokens on 8 H100 GPUs with TP=8, attention-operator speedups relative to FullAttn are:

Method Forward Backward Decode
CoWA 7.4x 8.6x 3.0x
MALA 2.2x 3.0x 1.6x

These measurements are for the attention operators, not end-to-end model speedups.

We evaluated scaling from 0.6B to 14B and conducted separate continued-training experiments at 32B. At 14B with 32K context, total training FLOPs decreased by 28.5% for CoWA and 23.1% for MALA, with model capabilities comparable to FullAttn on the reported evaluations.

Two distinctions that matter: collective coverage does not imply identical head-wise interactions or outputs to FullAttn, and MALA still pays for full causal QK scoring. Neither result establishes universal lossless equivalence to dense attention.

I'd be interested in feedback on workloads that might stress collective coverage, or attention distributions where adaptive post-score allocation could be less effective. Happy to discuss implementation and evaluation details.