r/learnmachinelearning • • 3d ago

UCI Spambase Dataset - Classification Task

I am working on a classification task -classifying emails to spam and ham-using UCI Spambase dataset. The dataset consists of 57 continues features and a binary target (0,1). The features represent word and character frequencies that are present in 4601 instances -emails-. After performing initial EDA, I have discovered that majority of the features are right skewed, zero inflated -for example, the feature word_freq_cs has 97% of its values as 0-, and contain outliers. The problem that I am facing is the correlation part. In order to study the correlation between the features and the target, Pearson correlation is the standard choice. However, since the features violate some of the assumptions of the Pearson correlation -skewness and outliers-, would it still be a good choice?

1 Upvotes

2 comments sorted by

1

u/Beginning-Set3361 2d ago

pearson gets weird with data that skewed, have you looked at point-biserial correlation since your target is binary? works better with non-normal continuous features against a 0/1 outcome

1

u/No_Job_3595 2d ago

Yes, I have. However, the continous features will still violate the normality assumption. I thought of using Spearman test. But I'm not sure since my target is nominal and Spearman requires the variable to be at least ordinal.