Most predictive models answer the same question.

Will this fail?

It's a useful question. It's often the wrong one.

A maintenance engineer doesn't schedule work because something might fail one day. They need to know when the risk becomes high enough to act.

Years ago, I worked on a predictive maintenance problem where conventional machine learning ran into practical limits almost immediately.

Failures were rare. That was good for the factory. It wasn't good for the model.

Most equipment was still running. Many components were replaced during routine maintenance before they actually failed. Others were replaced even earlier because nobody wanted to be the person who let one run too long.

They weren't failures. They weren't healthy forever either.

Eight components shown as timeline bars. Two ended in failure. Three were replaced before failing. Three were still running when the observation window closed. A binary classifier uses two of the eight; survival analysis uses all eight.
Eight components. Two failures. Six that still know something.

Discarding those observations meant losing most of the operational history. Labelling them as healthy wasn't true.

The problem wasn't the algorithm. It was the question.

The maintenance team wasn't really asking,

"Will this fail?"

They were asking,

"Given everything we've seen so far, when does replacing this component actually make sense?"

That's when we turned to Survival Analysis.

Instead of forcing every observation into the same label, each one contributed what it actually knew: equipment still running, equipment replaced early, and equipment that had genuinely failed.

The model didn't just estimate remaining useful life. It informed maintenance decisions.

On the test machines, planned replacements became less frequent without a corresponding increase in failures. Components that would previously have been replaced conservatively stayed in service longer because there was now evidence that they could.

Maintenance didn't become riskier. It became more informed.

I moved on before the project reached production, so I never measured the long-term operational impact. But it convinced me that we had stopped asking the wrong question.

We spend a lot of time chasing newer algorithms. Sometimes the better move isn't finding a newer one. It's remembering an older one that was built for the question we're actually trying to answer.

Everyone predicts failure. Few predict time.