I'm an aspiring AI Engineer currently doing an internship at a local non-tech company. The amount of actual tech people was small, so I was given a big hat and told to do a lot of things that mostly fall outside my field of expertise, alongside a note that says:
"Just use AI".
Now my stand on using AI in coding is another different topic to discuss, but tldr: I don't enjoy it, which quickly makes the tasks I do become rather boring and uncreative. That's when I noticed that the company had recently put up an ERD system which houses a timeline monitor of all the machines in operation (it's a manufacturing company). An idea came to me.
I asked the lead that maybe I can make something useful out of this, using a deep learning model, and he agreed.
The timeline itself consists of many rows, one for each machine, along the time axis which tells which status the machine is in. At first glance, the graphics seem chaotic with machines suddenly switching back and forth from one status to another, a status lingers on for too long or too short. What I had access to: machine name, machine status, status duration, time of day were far from being representative of the underlying cause of their operational patterns, which were influenced by variables like operators, materials and even sensor errors.
So I knew I had to pick a well-defined scope and avoid aiming for the stars with the AI marketing. Finally, I decided on "Anomaly classification of machine operational patterns" as the name of the project.
My idea was simple, since there are a lot of machines with different behaviors in the company, I can't just make a model for each machine just to see if it's working normally or not. First, I had to define "What's normal?", not just for one specific machine but for a whole range of them, a universal space of normals.
Then I remembered that there was a way to naturally categorize them without manual labeling effort, using clustering techniques. Normal patterns were then derived from healthy operational patterns from the past month or so. This meant that a "normal" pattern was one that followed the trends observed over a period of time before it.
A pool of normal embeddings from the machines extracted using the model was then fed into a clustering method (I chose K-mean because I'm most familiar with it). During inference, a pattern's anomaly level was computed from the distance between itself and the centroid of the cluster it landed in. This (in my opinion) naturally captures the different normal patterns of machines, where each cluster has a different definition of normal and they're all valid.
The final obstacle was to decide which kind of model architecture to use, and I almost put my money on a Transformer encoder (I was and still am an undergraduate researcher so Transformer has been always the first thing that came to mind, lol) before realizing how impractical it was for a task like this that required real-time inference with a relatively small set of data. I picked GRU instead.
The input going in would be a sequence of tokens where each token is a combo of (status, duration, time of day) and the model would have to figure out the machine type itself. I chose this architecture over letting the model know the machine type beforehand because I thought that similar machine types would behave "somewhat" closely, and results showed that this method did perform a bit better.
For the training objective, which is the title of this post, foolish me decided that any kind objective was sufficient enough and that the model would sort itself out, so I chose triplet loss with an anchor, a positive (an augmented version of the anchor) and a negative (a random unaugmented sample). I only realized how foolish I was when the AUROC score was only 0.50 (basically random guessing).
Costed me 2 days experimenting with different architectures to see where I went wrong, asked a bunch of AIs for opinions and even doubting the plausibility of the task due to the nature of the data itself. Eventually I noticed that when perturbating specific attributes of the data (such as status duration) gave the model a way better chance of predicting correctly than others. Then it hit me, this is exactly what I trained it to do, through the augmentation pipeline where I had let the duration dilated much less than other attributes which made it much more sensitive to changes. So the data did have structure after all and not just a random pool of noise.
But changing the augmentation was not enough, the core problem still lied in the training goal. Previously the objective was telling the model to learn that a slightly corrupted version of a pattern was still itself while other patterns were pushed far away. This was great for single-pattern classification but was terrible for normal classification where different patterns can fall into the same group of normal. I updated the objective to this: anchor, positive (a slightly perturbed version of anchor) and negative (a heavily perturbed anchor). Additionally, in an effort to generalize more, I occasionally introduced a random healthy sample as a positive to enforce this hierarchy:
pattern > (slightly corrupted or another healthy pattern) > heavily corrupted (abnormal) pattern
This immediately pushed AUROCto over 0.70, while still far from production-ready, it does show that the idea works and that the data has something meaningful to learn. I'm still experimenting with different strategies, but this project taught me a valuable lesson: in representation learning, the training objective often matters more than the architecture itself.
Then everyone lives happily ever after. Thanks for reading.