Getting Predictive Maintenance To Actually Work On The Factory Floor
Most people think Artificial Intelligence In Operation Management is about slapping a model onto a dashboard and calling it smart. It's not. It's about cleaning data until your hands bleed and then dealing with the moment the model predicts a bearing failure at 3 AM on a Saturday when nobody's around to verify anything. I spent six months building a predictive maintenance pipeline for a mid-size automotive parts plant. The initial deployment looked solid on paper. Accuracy numbers were fine. Precision was acceptable. Then we hit the real world and everything fell apart.Practical First Steps With Artificial Intelligence In Operation Management
Start with sensor coverage, not algorithms. I have seen more projects fail because someone tried to run a deep learning model on temperature data from a single sensor mounted three feet away from the actual heat source. The signal never reached the machine. You need vibration, thermal, acoustic, and power draw readings at the component level. For a CNC spindle, that means accelerometers directly on the bearing housing, not bolted to the machine frame. The data collection infrastructure usually eats 60 to 70 percent of your budget. I learned that the hard way. We used MQTT to pull sensor data into a time-series database on Prometheus, then fed it into a feature store built on Feast. The pipeline itself ran on Apache Kafka with about twelve minutes of end-to-end latency from sensor to prediction. That was acceptable for preventive maintenance. It would have been useless for real-time anomaly detection. The models that actually ship out to production tend to be surprisingly simple. A gradient-boosted tree like XGBoost or LightGBM on engineered features will beat a convolutional neural network on raw vibration waveforms in almost every operation management scenario I have touched. The CNN needs more data than you will realistically collect. The tree needs clean features and it runs on a CPU, which matters when you are deploying to edge devices. Feature engineering is where the work actually happens. Raw vibration data is noise. You need RMS values, kurtosis, crest factor, spectral centroid, and FFT bins extracted at regular intervals. My team spent three weeks just deciding whether to compute features in sliding windows of 0.5 seconds, 1 second, or 2 seconds. The 1-second window with a 50 percent overlap ended up giving the best tradeoff between responsiveness and false positives.One thing nobody tells you: train your model on failure data that actually happened, not on synthetic degradation. We tried generating fake wear patterns for a heat exchanger using a physics-based model. The predictions were garbage. Real degradation in that unit was driven by calcium carbonate buildup from the cooling water, which the physics model did not account for because the water chemistry changed seasonally. We ended up pulling five years of maintenance logs and matching them against sensor readings from the same dates. The model trained on real failure events caught issues the synthetic approach missed entirely.
The biggest practical pitfall is concept drift. Your model will degrade. A compressor that runs at 80 percent load in summer and 40 percent in winter will look completely different to your model across seasons. We solved this with a retraining cadence of every ninety days and a rolling baseline that updates weekly. The model does not need perfect accuracy forever. It needs to stay within a known error band. When the error band widens, you retrain. Another issue that comes up constantly is false positive burnout. If your system flags a problem and it turns out to be nothing three times in a week, the maintenance crew stops reading the alerts. Ours started at about eight false positives per week. We narrowed the detection threshold and added a confirmation rule that required two consecutive abnormal readings before triggering an alert. That dropped false positives to about one per week without missing any actual failures over a four-month period.