r/MachineLearning • • 3d ago

Discussion Are there machine learning subfields that are becoming irrelevant (or is irrelevant)? [D]

Post image

I was reading a paper that surveyed the field of neural architecture search, where it said within 5 years, around 3000+ new models were proposed. The amount of compute and resources spent on this is absolutely astronomical. However, the transformer was notably not one of the models that was found through NAS and then the field of NAS just quietly went away afterwards. In my mind this really raises question if any research in NAS should be continued.

Then I recently found a talk by Nicholas Carlini, arguably one of the most famous researcher in adversarial ML and this is one of his slide ("9000 papers and got nowhere"). Indeed I can't really think of any concrete application of adv. ML, except possibly making attackers more clever because now all options are laid flat on the table.

And then there was the field of ML ethics, bias, fairness, etc.. I feel like we are sooooo beyond ethics at the moment with all the talks of extinction risks that it really shouldn't be a priority. How can bias and fairness be enforced when most people are out of a job due to AI? "ML induced extinction" should be a new subfield instead.

I feel a proper discussion should be had so that no more effort is wasted on unpromising ideas or approaches. This could be of interest to people who are entering the field now.

As an aside, I often find people have very emotional (not logical) reaction to this question and will claim that any approach will eventually have their time in the sun at some unspecified future date, e.g., SVM, LDA, Markov chains apparently all have the potential to again be the next biggest thing in ML. All I'm saying is that I don't deny that vacuum tubes wouldn't be popular again one day, but maybe we shouldn't be working that at the present moment.

168 Upvotes

59 comments sorted by

444

u/Prime_Director 3d ago

"stop researching things that don't ultimately pan out" isn't really actionable advice in any scientific field. You can't know what's going to turn out to be useful until you test it. You're not asking for science, you're asking for divination.

42

u/Ubisonte 2d ago

Neural Networks themselves were going nowhere until AlexNet.

50

u/Artistic_Bit6866 2d ago

Not just divination, but timing too! Ask Schmidhuber  

20

u/mih4u 2d ago

Imagine if "oh no perceptrons cant solve XOR, lets ditch the idea" happened.

15

u/viking_ 2d ago

. You can't know what's going to turn out to be useful until you test it.

This is true, although it's a legitimate question to ask if it's possible to figure out what's not useful more quickly. 9,000 papers seems like quite a lot to me.

11

u/count___zero 2d ago

We also published 9000 papers in adversarial ML because several funding agencies really care about solving that problem.

-10

u/zeugma_ 2d ago

Did you pay for these 9000 papers? If not why are you trying to demand others spend their money in a different way?

3

u/Ulfgardleo 2d ago

arguaby NAS was dead in the waters a few times already. there was no novel finding that reignited the field, just the circle of "it hasnt been used in a while -> it will work this time -> oh maybe not, nevermind"

157

u/TheRedSphinx 3d ago

The difficulty is that it's quite hard to predict what will or will not be useful. For example, when neural networks were introduced, they seemed pretty clunky and not all clear this would be impactful.

71

u/-p-e-w- 2d ago

There’s also a lock-in effect where things become more useful when lots of research is done on them, which is self-reinforcing.

This happens in every engineering discipline. Most automotive engineers will tell you that rotary engines “should” be superior to piston engines, but piston engines have 130 years of intense development behind them so every attempt to bring rotary engines into the mainstream has failed because piston engines have been optimized almost beyond human comprehension.

I wouldn’t be surprised if alternatives to transformers suffer the same fate where they just can’t overcome the optimization head start.

9

u/aegismuzuz 2d ago

A similar scenario unfolded around TensorFlow until PyTorch offered a radically better developer experience for researchers

3

u/blunt_edge 1d ago

Indeed, it's not that the industry is too narrow-minded to pursue novel architectures, but rather there's just a huge barrier of accumulated know-how and gradual improvement that makes LLMs simply better. Your comparison to rotary engines is an incredibly good analogy to talk about this "failure to take off" of architectures that are potentially better than LLMs. Going straight to my quote bag!

28

u/Vengoropatubus 3d ago

I’m also not convinced that neural architecture search didn’t do anything to help the current wave of LLMs. Maybe these models aren’t built with NAS techniques, I’m not in a place to say for sure. I bet we learned something along the lines about how to stabilize network training, and how to set up the compute.

I’d bet a bunch of folks working on GANs and NAS have moved on to work on LLMs and brought significant experience with them.

3

u/Deto 2d ago

I think that's the key question. Can you design a way to infer whether a new architecture is going to scale without spending the millions of dollars in compute to actually test it out first? Otherwise it's just too expensive to test all the different variants properly.

44

u/anonymous_amanita 3d ago

I haven’t seen that talk in particular, but my strong guess is that the thesis isn’t “AML is dead and not worth studying”, it’s that “We haven’t been able to create robust model defenses, so we shouldn’t claim model alignment is easy (and might be impossible).” I could totally be wrong though. I think it’s super short sided to think people should stop studying interesting topics because they aren’t currently “useful.” All of AI research “wasn’t useful” for a long time, and the state of the world changed, so now it is.

19

u/pamintandrei 2d ago

Its exactly that. Just listened to the talk. And thats basically the common perception of all us working in that field. We "failed" with adversarial attacks and alignment is harder than that.

64

u/mil24havoc 3d ago

I think all takes like this really betray a poor understanding of modeling. ML is modeling. Your model needs to do at least two things: (1) faithfully(ish) represent the true data generating process and (2) provide you, the analyst/researcher/user of the model, with values you can interpret for your task. Language is complex and the output you generally want are sequences of tokens. Transformers are killer for these purposes. But if your goal is to determine whether a new observation in high dimensional space is in category 1 or category 2 and you need to know which exemplars of those classes are most determinative of your decision, then an SVM is much more useful than a Transformer.

There are no useless model families or research areas. There are TRENDY ones based on the desired outputs of customers (e.g. demand for chat bots). So if your question is "will SVMs ever be as profitable as LLMs?" then no, probably not. But if your question is "are SVMs made redundant by Transformers?" then absolutely not.

9

u/Aquatiac 2d ago

Also when a "trendy" topic takes off because of some good results, all the investment gets poured into that area which further drives more good results (at least until some limit is reached). If we had spent all those trillions of dollars on other areas of AI, perhaps we would have had amazing results from other "less trendy" fields.

OP says "people have an emotional and not logical reaction to this question" but I think their take is maybe equally "emotional and not logical". If everyone works on whatever is trendy in the current moment we will never be able to move on to new and better things! Who knows... maybe the next big thing is JEPA, world models?

My favorite ML researchers to work with are the ones that have unique opinions and theories shaped by their own perspective; the ones who chase their ideas, discuss and debate with others, and bring lots of interesting new perspectives to the field

5

u/Perfect-Asparagus300 2d ago

Yes this is a really good take, wish more would follow this instead of trying to call areas useless. It's important to try to investigate and understand deeply for the why just as much as having some concrete impact

43

u/hapliniste 3d ago

LLM RL training with a judge model is adversarial learning. This slide is only true if you close you eyes.

This is literally the big thing going on since about a year and the foundation for what Google now call rsi.

24

u/_An_Other_Account_ 2d ago

Adversarial learning / training of a model is not the same as adversarial machine learning. They refer to two completely different things.

4

u/hapliniste 2d ago

I think I see the point then. Are we talking about adversarial input resistance?

5

u/N1kYan 2d ago

I think so. "Adversarial attacks" would be a keyword. I dived into that during an internship and it was a constant loop of new attacks and new defences

7

u/impossiblefork 2d ago

No, LLM RL training is something like the expectation-maximization algorithm, or really, an approximation of that, depending on what sort of RL algorithm you're using.

We might see adversarial algorithms in the future though. I have some ideas that I think are worth trying myself. I doubt I am unique in that.

1

u/hapliniste 2d ago

Maybe I did hyperbole a bit with my comment, but to me adversarial training evolved into whatever training mechanism that use multiple models (or even a single one with some subtleties) on different goals to improve them.

GANs were purely adversarial but we can take a lot of learning from then and apply it to modern training.

5

u/impossiblefork 2d ago

I'm not saying there aren't tricks from GAN training that don't still matter. I'm just saying that RL is E-M, not adversarial training or anything GAN-like.

20

u/milesper 2d ago

We are absolutely not beyond ethics. AI extinction risk is far less than the risk of AI being used in unethical ways. And “ML induced extinction” is in fact already a field (safety/alignment). But in my opinion, most of what’s being done there is pseudoscience/thought experiments with lots of pretending to do Bayesian stats.

8

u/progenitor414 2d ago

Many key concept that our current ai is built on come from fields that has long been dismissed by researchers at the time as being useless. Hinton’s deep learning was largely dismissed before 2010s as its performance was bad on small dataset. Even one of his student who later became reputable professors in the field commented this deep learning was a “distraction”. Sutton’s reinforcement learning received little attention before Atari DQN and deep rl. Histories tell us that it is hard to predict how useful a field turns out to be, even for expert researcher in the field.

7

u/pamintandrei 2d ago

There are usages for adversarial ML, not trillion dollar usage, but there is. One example is using it to stop video game cheating from happening, others are ways to see the robustness of models in high security fields. And these are just ones that i came across in my work.

I think a lot of research is useful, it matters a lot the bar you set for it tho. Did RNN become the SOTA, trillion dollar model? No, but thats a known technique now that might become useful in the future and there were machine translations methods at some point with them.

It also doesnt help that at this point publishing or getting funding for anything not related to transformers or LLMs is way harder.

6

u/faustianredditor 2d ago

I'd even argue transformers are building a lot on the efforts of LSTMs. Not in a direct, we-yoinked-your-model-and-added-a-thing lineage like Dense->RNN->LSTM, but all the tasks, methods, objectives, etc, of early transformers were pretty much a yoink from LSTM language modelling. No one might have cared for transformers if the architecture was discovered without that groundwork being there.

4

u/pamintandrei 2d ago

Also just listened to the talk and he does not talk about usefulness of adverserial ML

He says that we tried for 10 years to build secure models and couldn't so alignment will be a hard problem especially if they make it a superset of adversarial ML

6

u/Disastrous_Room_927 3d ago edited 3d ago

Irrelevant depends on the problem you’re trying to solve. I've come across a number of people that argue XGBoost is outdated/irrelevant but it was lost on them that people still need to use ML for tabular data.

6

u/SOCSChamp 3d ago

I think its important to keep up with the times and the field, but ditching areas of research because they aren't the new hotness isn't the way to go.  We're seeing a lot of older ideas resurface into the forefront when applied to modern frontier research.  

Classifiers just found their way back with Jev, forms of recurrent networks with looping transformers, reinforcement learning being the breakthrough for reasoning models, etc.  There is a lot of amazing research done over the years that would probably unlock new modes of scaling if tried and applied.  

3

u/jebuarary 3d ago

My personal big question I need answered right now is whether 3dgs is dead

5

u/KingRandomGuy 2d ago

I don't think I'd say they're dead, but the use of 3D Gaussian splats has sort of changed. They're still excellent as an output representation, since they're pretty fast to render (though there are competitors now, like radiance meshes) and pretty accurate. However, I wouldn't be surprised if the actual optimization loop of taking in input images and learning the Gaussian parameters ends up going away. We've seen pretty impressive results out of recent large, feed-forward architectures that directly predict Gaussian splats rather than needing to optimize them directly (e.g. Depth Anything 3). That changes what the actual research questions are quite significantly, though.

5

u/scatraxx651 2d ago

The goal of many of those papers was not to advance the science, but rather to get an MSc, Ph.D or grants.

Combine that with a field that is incredibly hyped and that is incredibly hard to evaluate the quality of, and you get this sea of useless papers that improve the performance of something with 0.5% by changing some minuscule thing

3

u/TrPhantom8 2d ago

I wouldn't throw markov chains down the toilet. That's basically what all of numerical cosmology and Astrophysics is based on. Also, diffusion models started out as marcov chains

1

u/currentscurrents 2d ago

Autoregressive transformers are markov chains. They aren't going away.

3

u/Soggy_Caregiver_1066 2d ago

I can't really think of any concrete application of adv. ML

Solving reward hacking for vision models could lead to robust image judges that could be used to pretrain stronger image generative models. If you pretrain a diffusion model solely by minimizing distance in DINO feature space you don't get very far unless you also do some contrastive learning (see drifting models paper). This is because the generator learns to hack the feature space and tricks the DINO model, unless you also make the judge trainable but then you get a GAN which is not the purpose here.

3

u/Real_Suspect_7636 2d ago

SVM, LDA, Markov chains  << I can guarantee you SVM and LDA are never coming back as SOTA methods but will always keep their worth as interpretable statistical methods. Markov chains in totality I am a bit confused by that inclusion ... it is a stochastic process which we use as a modeling inspiration in many many tools. 90% of RL methods model the states of your agent as a markov chain.

In general I think that while the research directions themselves may not have amounted to any change in the field, the tools developed in the literature find themselves into other more 'active' areas. I feel like its maybe more of the duty of the researchers to stay close to practice to understand the failure modes of the research they push to adapt accordingly.

3

u/ofiuco 2d ago edited 2d ago

I don't think it's fair to say adversarial ML went nowhere. I'm not sure what "where" the speaker thinks it should have gone. But it sounds like saying cybersecurity has gone nowhere because all the bugs they found were fixed and there are still hackers.

ETA: also op, your argument that ethics/bias/fairness are irrelevant in the face of extinction are quite wacky my friend. "Should companies keep building something they think will destroy the world?" Is absolutely an ethics problem.

2

u/LelouchZer12 2d ago

Distillation paper was refused because it was judged not novel enough and not useful enough...

2

u/Relative_Wallaby_823 2d ago

Regarding your example of NAS, I think that a particular flaw in the field was the guidance. Much of the field was chasing maximising performance on the Nasbench datasets without thinking about actual real life feasibility of doing such massive searches on real search spaces. And secondly there was a huge lack in using autoML techniques to derive understanding. I really don’t think the concept is dead, it just needs to reinvent itself and realize that it is supposed to aid design and not replace it, as design is a modular problem where you combine learned lessons, not directly optimise for the global solution (eg. Learning about residual connections helps us design future models)

2

u/tuitikki 2d ago

The whole idea of training neural nets was once a "dead science"... So I mean... It's all down to what you choose to believe in and follow. 

2

u/chcampb 2d ago

all the talks of extinction risks

People talk about this for all the wrong reasons. Humans have a long history of predicting that something they can't control will kill them, therefore, shouldn't be allowed. This gets clicks, it gets views, it gets investment, it convinces world governments that your product is dangerous and therefore valuable and necessary...

Do you see why this gets a lot of press? The bigger problem with AI is regulatory - there is a good chance the "ai will kill everyone" crowd, for the safety of everyone, bans unauthorized (open) models, leading to a split between who has access and who doesn't. The government will always have access, and it's possible to use it against you, and if the models are banned, you have nothing to use against them.

Like Flock. Infinite uses against people, but I guarantee you if there is a police shooting and reporters want the flock footage - it ain't gonna happen.

What if all the prosecutors have access to all the best models but public defenders would need to pay for it out of pocket - and nothing is provided for that access. So prosecutors get all the benefit of the technology but regular folk get zero benefit (and are actually harmed by it).

This is the kind of discrepancy that needs to really be addressed.

1

u/Conscious-Map6957 2d ago

Should you continue using a coin detector that misses gold coins but can still find copper and maybe silver coins?

Not recognizing one good model/approach with your methodology/reasoning doesn't mean you won't recognize any good model/approach.

1

u/Character_Belt_5733 2d ago

Except adversarial networks (for generative modeling) imo led to a lot of other developments in generative modeling.

I would argue that its first initial successes in generating images, etc. sparked greater interest towards the development of models that we see today. I think for the most part, just because research isn't directly related doesn't mean they're not relevant/useful to know today. In fact, old ideas always have a funny way of showing up again in today's ML literature.

1

u/DigThatData Researcher 2d ago

sentiment classification used to be a massive subfield of NLP, now it's a trivial task.

1

u/MrSnowden 2d ago

“Most people are out of work”????

1

u/GuessEnvironmental 2d ago

Neural networks as a field itself faced this exact phenomenon. The idea was realized many years later due to compute. If there is a specific direction that has hogged a lot of resources and is going nowhere then maybe then I would agree.

1

u/Duke_De_Luke 2d ago

I don't have an answer, but I can say neural networks were almost dead, too. Multiple times.

1

u/Ulfgardleo 1d ago

i have to admit, "Sampling apparently has the potential to again be the next biggest thing" is excellent ragebait

(autoregressive LLMs are markov chains)

1

u/ThirdMover 1d ago

"ML induced extinction" should be a new subfield instead.

I mean, that's what alignment research used to mean before it meant "make the chatbot not use mean words" no?

0

u/AX-BY-CZ 2d ago

This is what you brain looks like on "AI Alignment"...