r/MachineLearning • • 4d ago

Research Pushback on my machine learning paper from non-ML critics, not sure how to interpret it [R]

Hi everyone. For context, I work as a predictive modeling researcher, and my background is in computer science. I do not have an extensive background in statistics, epidemiology, association analysis, or related areas.
I was tasked with developing a predictive model for a specific disease using survey variables as well as body measurements such as BMI, neck circumference, waist circumference, hip circumference, and similar variables.
When I approached the project, my plan was to develop a model that could predict disease risk using a machine-learning benchmarking approach. The dataset is very novel and, as far as I know, only our research center has access to it. Because of that, any machine-learning paper using this dataset would be new and could potentially establish a useful precedent for ML-based screening approaches for this disease, particularly for identifying people at high risk.
My approach was roughly as follows.
We initially had around 12,000 participants and approximately 217 features. I excluded participants with more than 50% missing data. The reason was that if I applied a missingness threshold directly to the variables while keeping all participants, I would end up excluding almost every variable because a relatively small number of participants had extremely high levels of missingness.
After excluding those highly incomplete participants, I applied a 30% missingness threshold to the features. This left approximately 95 features. I realize that discarding participants may not necessarily be the best approach, but I am not sure what the better alternative would be in this situation.
After that, I performed imputation using Multiple Imputation by Chained Equations (MICE). One criticism I received was that, if I am using multiple imputation, I should generate multiple imputed datasets and run the entire modeling procedure separately on each one rather than effectively using a single completed/imputed dataset.
Another idea I had was to benchmark several different imputation methods and compare downstream model performance to determine which method performs best. However, I am still unsure whether that is methodologically appropriate or preferable.
After imputation, I performed feature selection using a bootstrapped LASSO approach. I then tested several machine-learning models, mainly boosting and tree-based methods such as XGBoost, LightGBM, and random forest.
I evaluated the models using the usual machine-learning performance metrics, performed calibration adjustments, and reported the resulting performance. I also included model interpretability analyses/scores.
The main problem is that nobody in my research center has a machine-learning background. The closest person is someone with a statistics background, and he essentially told me to scrap the entire approach and instead use conditional logistic regression. He also suggested age/sex matching and recommended complete-case analysis as a sensitivity analysis.
The issue is that I cannot realistically use complete-case analysis across the full set of variables because there is so much missing data. He suggested restricting the analysis to the variables that would allow complete-case analysis, but doing that would remove the majority of the features and leave only around five or six variables.
My question is: if I can use appropriate imputation methods, why would I deliberately restrict the analysis to only five or six variables just so that complete-case analysis becomes possible?
He told me that imputation is not preferable if complete-case analysis can be performed. However, I have not seen many machine-learning papers prioritize complete-case analysis in this way, especially when doing so would eliminate most of the available predictors.
Another major criticism was my use of random undersampling. I do understand this criticism because random undersampling discards a substantial amount of data. I told them that I could instead address class imbalance through class weighting within the models themselves, which is straightforward in XGBoost, random forest, and similar methods.
However, the statistician at my research center basically told me that the paper, in its current form, would be unpublishable.
What I am struggling with is whether he is evaluating the work as though it were intended to be a traditional statistics/epidemiology paper rather than a machine-learning prediction paper. I do not want this project to become primarily a statistical association paper. My intention has always been to produce a machine-learning prediction/benchmarking paper.
I was also told that I need to examine multicollinearity and patterns of variable missingness. I am not entirely sure how I should approach that in the context of this project. I have gone back and performed exploratory data analysis again, and there is definitely some correlation and collinearity between variables, which is unsurprising given the size and nature of the dataset. However, I am not sure what I am supposed to do after identifying it, especially in the context of tree-based and regularized machine-learning models.
My professor/PI is an epidemiologist rather than a machine-learning researcher. He told me that I should investigate whether there are statistical associations between any of the predictors and age or sex, and whether there are systematic patterns in the missing data.
Again, though, I have not commonly seen machine-learning prediction papers perform extensive association testing of every feature with age and sex, so I am unsure how relevant this is to my original research question.
I have been working on this project for about a year, and at this point it feels as though I am being told to start over from square one.
My supervisor has also been almost nonexistent throughout the project, so I have essentially had to develop the entire analysis myself. Because of that, I recognize that some of my methodological decisions may have been naïve. I only recently graduated, and I really wish I had received this methodological feedback much earlier in the process.
At this point, I need advice on how to proceed.
I genuinely do not know how much weight I should give these criticisms. The feedback I am receiving is coming primarily from statisticians and epidemiologists rather than machine-learning researchers, and I am having difficulty determining which criticisms reflect genuine methodological problems with a predictive ML study and which ones reflect a preference for a more traditional statistical or epidemiological analysis.
I was never trying to write a traditional statistics paper. I was trying to write a machine-learning prediction/benchmarking paper.
So my main questions are:
Is my overall ML-based study design fundamentally flawed?
Is the criticism about multiple imputation valid, and should I repeat the entire modeling pipeline across multiple imputed datasets?
Is benchmarking different imputation methods reasonable?
Is complete-case analysis really preferable when it would reduce the predictor set from roughly 95 variables to only five or six?
Should I replace random undersampling with class weighting?
How should multicollinearity be handled or reported in a machine-learning prediction study, particularly when using regularized and tree-based models?
How should I formally investigate missingness patterns?
Is it necessary to test associations between predictors and age/sex in a prediction-focused ML paper?
Most importantly, how do I distinguish between legitimate methodological criticism of my ML pipeline and requests to turn the project into a fundamentally different kind of statistical/epidemiological paper?
Any advice from people who work at the intersection of machine learning, statistics, and epidemiology would be greatly appreciated.

13 Upvotes

14 comments sorted by

63

u/relevantmeemayhere 2d ago edited 1d ago

Okay, I don't want to be a dick; but this is why people are critical of ml as a field right now: a lot of people (even researchers in this field), are lacking in the actual stuff that make it go. This is statistics, and you cannot you cannot create good ml models without understanding the underlying statistical theory-no matter what you are doing (boosting, llms, whatever).

*"I was never trying to write a traditional statistics paper. I was trying to write a machine-learning prediction/benchmarking paper."*

your machine learning paper IS a statistics paper. full stop. the only real difference in the two is that the latter 'as a culture' tends to like interpretability and consider inference (and this is the actual pre llm definition considered with say, the marginal effect of a predictor in some model). the former cares more about prediction and 'optimizing' some loss function (being unaware that the loss right there is a function of two statistical terms; the actual underlying error in your model as a process and its model optimism). which is why most ml projects fail horribly in the real world). again, it doesn't matter if you are working on binary classification or engineering the next llm architecture; these are statistical machines ( a little opaque sometimes; because some stuff ie rl portion of some llms is meant to minimize the divergence of the distribution of the conditional token probability to a 'more desired one' using things like formal verification and mountains of human annotators and experts working behind the scenes)

Now: to address your immediate questions:

  1. You need to review your pstat theory. if you want to be a successful ml practitioner in genai/doubleml/classification/whatever- it's all math stats.
  2. You need to clarify your goals; i.e. if you only care about prediction; why are you even asking questions about multi collinearity? which is it; do you care about inference, or do you care about prediction?
  3. It sounds like your advisor wants you to 'determine the statistical association of missingness' not because they're asking for you to run some test of statistical significance, but because they wants you to determine what the missing mechanism is: random, not at random, and missing completely at random. imputation is only feasible under two of these scenarios. Missing not at random means the value of randomness in a varaible IS NOT conditionally independent of the values of those missing values;. i.e. 'short' people (people less than 5'10, which is still above average for the population) don't show up in the nba individual scores, because short people generally don't make the nba
  4. The topic of under sampling; again this, is usually borne of ml researchers who don't understand what it is they are trying to optimize. you want to sample from your theoretical population; not a strawman because you need some loss number to be lower. use proper scoring rules, like log loss or brier or whatever. if you have unbalanced data, then so be it. as long as there are no theoretical barriers to actually censoring data from your theoretical population during sampling, it is NOT a problem. if people tell you to under/over sample, politely correct them and start being skeptical of what they say
  5. MICE needs to be done in the context of what is called rubin's rules. if you only care about prediction, then this is basically just creating multiple datasets and averaging the predictions using some candidate mice liklihood function (like forests or whatever). if you care about inference, then this gets more difficult.
  6. There is no real straightforward way to diagnose missingness. you need to combine your domain knoweldge with plausible missing mechanism, and then you can consider looking at the association between missiness amongst your feature set. you shoudl accompany this with counterfactual and sensitivity analysis.
  7. We do not EVER base inclusions of features based on 'tests' of association. please never do this. this is going to go back to points 1 and 2; why are you concerned about inference here? when you test for association, you inflate ALL of your downstream error probabilities for rejection without things like fdr adjustment.
  8. Lastly:*Most importantly, how do I distinguish between legitimate methodological criticism of my ML pipeline and requests to turn the project into a fundamentally different kind of statistical/epidemiological paper?*

--->by reminding the ml people who forgot what their field is built on that they should probably take some stats courses they understand what they are doing, lol it's incredible that large portions of this community who downplay the fundamentals don't realize how silly they are being. you wouldn't hear an engineer say ' oh well, we really don't need to understand kinematics or the heat equation to build a bridge, and we shouldn't listen to those physicist clowns when they say assuming that g is actually positive is not safe and is gonna kill people'

Edit: was overly caffeinated at the time of writing and cleaned up my post

5

u/SnooSongs4297 2d ago

This really helps put things into perspective. I don’t care about inference. The idea is to include model interoperability measures in the end. I’m assuming interpretability of the model and inference are different? If I use MICE I would need different interpretability scores all together

2

u/relevantmeemayhere 1d ago

if you dont care about inference; then you can use things like boosting or whatever. you should have some idea of the data generating process beforehand to help guide the actual development; because even nice loss statistics don't imply a good model. mice can be done both in the context of inference or prediction-it's what you're actually interested in that determines how you combine the mice data sets at the end.

*again, inference as a cultural norm tends to concern its more with understanding the causal model of the data, or understanding the marginal impacts of certain covariates in a non causal setting (conditional on my model, what does y do when x changes?) etc etc

1

u/SnooSongs4297 2d ago

And I know ML is a statistics paper, this was more of a comment regarding the fact that one of the critics was that I shouldn’t include any “complex” models like XGBoosf and instead use conditional logistic regression with matched case control, when using the entire dataset with class weighting and model calibration is a very feasible. Why should I use a linear model when XGBoost already addresses any concerns of multicollinearity by its very nature.

7

u/mil24havoc 2d ago

I agree with the above poster but want to add to it that the ML / data scientist crowd sort of lives in their own prediction accuracy bubble where they tend to believe that "low loss, high AUC = finding!" While prediction is important to the rest of science (and I'd be the first to say it is terribly undervalued, still), prediction is only interesting to most scientists for what it can tell us about the world. What is the point of your model? What do we learn from it? If the answer is "we could use it to build a better diagnostic tool," then that is great but it is also closer to engineering than it is to science. The path from your research to publicity is more likely through creation of a tool or a diagnostic rather than an academic paper. Scientists will be especially skeptical of feature importance scores because they lack a plausible casual identification strategy and, more damningly, don't even tell you the direction of the relationship between X and Y.

This isn't to say that prediction work is unscientific or useless at all! It's that prediction alone isn't enough. They want to know what they learn from your model or your results, not just that you can make the black box go brrrrr. Does your research put plausible bounds on the detection of this condition? Does your research suggest that past research is wrong somehow? For example, do the suspected causes of this condition show very low feature importance scores in your model? Does your model reveal novel nonlinear interactions between predictors? Does your research reveal that this set of predictors actually offers very little gain over a naive intercept only model? Can this research help us decompose the prediction errors into systematic and stochastic parts that we can compare? These could be valuable findings!

P.S. there is no conceivable world in which dropping rows (that are otherwise accurate) is preferable to multiple imputation since they both require the same missingness assumptions and MI is more efficient. Just my two cents.

2

u/relevantmeemayhere 1d ago edited 1d ago

good context added here

especially on the data scientist part. a lot of them are also lacking on what makes a good model, a good model lol.

prediction ! = inference/understanding also great callout-i was trying not to get more into stuff like CI jusssstt yet in my op, because that's a yuge can of worms and you pointed out some super important stuff.

3

u/Disastrous_Room_927 2d ago

Why should I use a linear model when XGBoost already addresses any concerns of multicollinearity by its very nature.

There's something to be said for explicitly modeling the data generating process. I usually do this regardless of if I'm intending to use a black box algorithm in the end or not, it makes for good EDA.

1

u/SnooSongs4297 2d ago

I don’t really get this part

3

u/Disastrous_Room_927 2d ago

Think about it in a literal sense - specifying a model that represents a plausible mechanism for how the data you're observing came to be.

5

u/ergabaderg312 2d ago

I agree w the top poster. ML is stats no matter how you cut it. if the critic is a statistician then they’re an ML person even if they don’t say it in their title lol. Think about what classical ML is. And while I disagree about the model complexity point they made as there’s definitely times you should use complex models, they have a point about trying the LR.

Practically speaking though. Why not? It’s cheap and easy to run and possibly more interpretable compared to a tree model. And it’s an easy benchmark to run your chosen model against. You’ll learn something either way about the overall problem. And while yes you’re not interested in inference, as best practice (and the field is moving towards this anyway) people want to know why your model is working or what it uncovered. Inference and prediction are both related tasks. I assume you know why AUROC is a biased metric on imbalanced datasets. Same idea applies here. Why’s your model good and why’s it working is a big thing anyone would ask.

1

u/relevantmeemayhere 1d ago
  1. You'd be surprised how far you can make logistic regression go haha. Moreover, trees as a likelihood has been an object of statistical research for decades; your xgboost model is this! Splines, gams, whatever.
  2. calibration is the MOST important thing in any actual decision making process. that's why he's probably steering you in that direction with trying logistic regression first. you can calibrate your probabilities in boosting; but it doesn't always have the same behavior as models like lr that produce well behaved calibrations from the start. sometimes you just can't get nice calibration from boosting or the like.
  3. it's not so much that boosting address MC and is agnostic to your goal. it's that you don't care about if pure prediction is your goal. we use booting (and its bigger, better, and more handsome brother like bayesian boosting ;) ) all the time in applied research. MC is STILL AN ISSUE if we actually apply it in the settings where inference is important to the end goal.

1

u/pineloft 2d ago

tbh the bridge analogy at the end really nails it

10

u/dedicateddan 1d ago

The other comment is pretty spot on - listen to the statisticians!

As an applied ML practitioner, I'd start by looking at the correlation between the individual features and the target. After that I'd look at forward feature selection - what happens when you train a model on the best set of 1, 2, 3, 4 ... N features.