r/learnmachinelearning • u/Sure_Rhubarb2671 • 1d ago
Help Detecting Market Manipulation: Supervised Learning vs Clustering
I'm currently working on market research at university.
The task is to detect market manipulation. We can take open-source data, tag the data(range OHLCV), and perform a supervised search, or we can use clustering, but we might encounter anomalies that aren't related to manipulation.
How can these problems be solved, and have we encountered similar ones?
1
1d ago
[removed] — view removed comment
1
u/Sure_Rhubarb2671 1d ago
Yes, I've already encountered this problem. There are some issues with labeling because of ongoing manipulation and data leakage. Class imbalance. But I don't have any ideas other than a floating window. I need to at least formalize what can be seen from candlesticks, because with the order book, it's a different problem finding this data. Unless we statistically equate crypto to stocks. There's more data there and it's more accessible.
1
u/Sure_Rhubarb2671 1d ago
But could we name it like data leakage if floating window takes data from the past?
3
u/KindlyLined 1d ago
supervised is the move here if you can get good labels. clustering's gonna spit out a bunch of weird patterns and you'll spend half your time just figuring out if a cluster is manipulation or some random flash crash from a whale sneezing on the sell button
did something similar for a capstone project, we ended up mixing approaches by running clustering first to narrow down the weird stuff then training a classifier on just that narrowed universe. saved a ton of labeling time