Sign In
Back to blog
Pat O'Donnell

Catching Defects Before the Batch Becomes Scrap

Abstract visualization of a defect signal being detected on a production sensor waveform

What happens in the 90 seconds between when a defect pattern starts forming and when your line monitoring system generates an alert? On most continuous lines, a lot of material runs through.

The 90-second window isn't arbitrary. It reflects what we measured when auditing alert lag on several wire drawing and extrusion lines during early Shelfmark pilots. The sensor generates a signal. The historian logs it. The control system runs its threshold check on whatever polling interval it's configured for. The alert goes out. Your operator reads it and decides what to do. The line hasn't stopped. The material hasn't stopped. And by the time anyone acts, the defect-producing condition has already traveled a considerable distance down the process.

Before getting into how to close that window, it is worth being precise about what we mean by "alert lag," because the term gets used to describe several different problems that require different solutions. Alert lag in this context means the elapsed time from when the first sensor signal that would, in retrospect, indicate a defect-generating condition occurs, to when the operator who can do something about it has information that could guide an action. That is the full chain: detection, threshold evaluation, alert routing, and human reception. Each step adds delay, and in a continuous manufacturing context, each second of delay is meters of material.

The threshold problem

Most line monitoring systems are built around hard thresholds on individual channels. Temperature exceeds X. Diameter falls below Y. Tension goes above Z. These thresholds are set based on what you know a defect looks like at its worst, calibrated against historical events, and reviewed when something goes wrong.

The problem is that defects on continuous lines don't usually begin with a single channel going outside its individual threshold. They begin with a correlated pattern across multiple channels, each of which may remain inside its individual bounds. A temperature drift that is within tolerance, combined with a pressure variation that is within tolerance, combined with a slight dimensional change that is within tolerance, can produce a surface defect that no individual channel threshold would have flagged until it was too late.

This is why we built Shelfmark around multi-channel pattern detection rather than single-channel thresholding. The signal for most continuous-line defects exists in the correlation between channels, not in any single channel value. Seeing that correlation requires holding all three streams in memory together and scoring the combined pattern against a baseline, not checking each one independently against a number.

What the alert lag actually costs

On a typical wire drawing line running at 200 meters per minute, 90 seconds of alert lag means 300 meters of potentially compromised material. On a plastic extrusion line at 60 meters per minute, that's 90 meters. For most customers, material cost per meter in these ranges runs from a few cents to several dollars depending on the substrate and the spec. The cost is not just material. It's also the time to identify, quarantine, and disposition that material, plus the impact on the production run that was supposed to use it.

The goal of fast detection isn't perfection. There will always be some lag between when a condition begins and when it's actioned. The goal is to push that lag as close to the physical detection limit as possible, which means eliminating the polling interval, the threshold-check delay, and the multi-step alert chain that adds time between the sensor signal and the operator's awareness.

We are not suggesting that detection speed is the only variable that matters in a quality monitoring system. Accuracy matters more than speed if fast detection comes with a high false-alarm rate, because operators stop responding to a noisy alert system and the detection capability becomes worthless. The goal is fast detection with low false-alarm rates, which requires better signal models, not just faster polling.

A scenario worth thinking through

Consider a copper alloy wire drawing line running 24 strand with a target diameter of 0.45 mm. The line is running at 400 meters per minute aggregate throughput. A lubrication feed anomaly on die 7 begins at 14:32:15. The anomaly does not cause an immediate break: it causes a gradual tension increase at that die position that propagates as a surface condition variation in the drawn wire. The single-channel tension threshold on die 7 is set at the level that historically precedes breaks, which is higher than the level that precedes surface condition issues.

At 14:33:40, the tension channel on die 7 crosses the single-channel alert threshold. The alert is generated and routed to the operator's screen. Total elapsed time since onset: 85 seconds. The operator checks the alert, walks to the die position, and intervenes. By this point, the affected material extends approximately 560 meters from the point of onset.

The multi-channel pattern model scores this condition at high defect probability at 14:32:28, based on the combination of the die 7 tension trend and a correlated temperature signature at the same die position. The alert reaches the operator at 14:32:40. Total elapsed time: 25 seconds. The affected material at that point: approximately 167 meters. The difference in affected material length: roughly 393 meters. At the copper alloy spec the line is running, that is a meaningful cost difference per event.

The calibration challenge

Faster detection without better calibration just produces more false alarms. This is the failure mode we see in plants that have deployed vibration or optical monitoring systems without good baseline calibration: the system fires constantly, operators start ignoring it, and the alerts stop adding value. The system becomes a liability rather than an asset.

Calibration at Shelfmark means establishing a process baseline from two to three weeks of production data, building a multi-channel pattern model from that data, and setting alert thresholds relative to departure from baseline rather than departure from a fixed number. This lets the system account for the natural variation in your process, which is different from every other customer's process, and focus on signals that represent genuine anomalies rather than normal operating range.

When the process changes, intentionally or not, the baseline needs to update. Grade changes, recipe adjustments, line speed modifications, seasonal shifts in input material properties: all of these can change what normal looks like. Monthly review calls are built into every Shelfmark tier specifically to review the alert history and update the model where the process has drifted in ways that need to be incorporated into the new baseline.

The 90-second goal, and what comes after it

On the pilot lines where Shelfmark has been running, we've measured alert-to-operator times in the 12 to 18 second range for the most common defect patterns. That's still not zero, but it's a very different number than 90 seconds, and it makes a very different impact on how much material runs through between the condition starting and the operator being informed.

The investment in faster detection pays off most clearly in the first defect event that happens on a line where Shelfmark is running. Not because it's the biggest event, but because it's the one where you see the difference between what the line monitoring system caught and what Shelfmark caught, and you can measure the gap in meters.

Achieving detection times in the 12-to-18 second range requires closing each step in the alert chain: continuous scoring rather than polled threshold checks, direct push notification rather than historian-dependent relay, and a trained model that can distinguish a genuine anomaly from a normal process fluctuation with high confidence. The 90-second lag is not a fundamental property of the technology. It's the accumulated delay of a system that wasn't designed with detection speed as a primary requirement. The question for any continuous line operation is whether the cost of that lag justifies the investment to close it.

Continue reading

More from the Shelfmark blog

See all articles