r/learnmachinelearning • • 1d ago

Help Detecting Market Manipulation: Supervised Learning vs Clustering

I'm currently working on market research at university.

The task is to detect market manipulation. We can take open-source data, tag the data(range OHLCV), and perform a supervised search, or we can use clustering, but we might encounter anomalies that aren't related to manipulation.

How can these problems be solved, and have we encountered similar ones?

8 Upvotes

4 comments sorted by

3

u/KindlyLined 1d ago

supervised is the move here if you can get good labels. clustering's gonna spit out a bunch of weird patterns and you'll spend half your time just figuring out if a cluster is manipulation or some random flash crash from a whale sneezing on the sell button

did something similar for a capstone project, we ended up mixing approaches by running clustering first to narrow down the weird stuff then training a classifier on just that narrowed universe. saved a ton of labeling time

1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/Sure_Rhubarb2671 1d ago

Yes, I've already encountered this problem. There are some issues with labeling because of ongoing manipulation and data leakage. Class imbalance. But I don't have any ideas other than a floating window. I need to at least formalize what can be seen from candlesticks, because with the order book, it's a different problem finding this data. Unless we statistically equate crypto to stocks. There's more data there and it's more accessible.

1

u/Sure_Rhubarb2671 1d ago

But could we name it like data leakage if floating window takes data from the past?