When a Weld-Inspection
Model Is Most Reliable

on the Defects That Matter Least

Cognitive AI is The Next Scientific Frontier in Machine Intelligence

From Explainability
to Cognition

The first generation of modern AI, statistical AI, focused on optimizing performance through scale: more parameters, more data, deeper networks. The second generation, explainable AI (XAI), sought to interpret model outputs, using saliency maps, feature attributions, and slice discovery to reveal how models behave. While valuable, these approaches remain diagnostic. They help humans analyze errors after the fact, but do not change how models make decisions.

Cognitive AI represents a third generation. It embeds reasoning within the system itself, enabling models to:

Map

the geometry of success and failure in training data.

DETECT

when an input falls into regions of ambiguity or uncertainty.

TRIGGER

adaptive interventions when predictions are unreliable.

Rather than functioning as a black box with a static confidence threshold, Cognitive AI actively monitors its own decision-making and adjusts dynamically. It operationalizes explainability into an ongoing cognitive process.

From Explainability
to Cognition

The first generation of modern AI, statistical AI, focused on optimizing performance through scale: more parameters, more data, deeper networks. The second generation, explainable AI (XAI), sought to interpret model outputs, using saliency maps, feature attributions, and slice discovery to reveal how models behave. While valuable, these approaches remain diagnostic. They help humans analyze errors after the fact, but do not change how models make decisions.

Cognitive AI represents a third generation. It embeds reasoning within the system itself, enabling models to:

Map

the geometry of success and failure in training data.

DETECT

when an input falls into regions of ambiguity or uncertainty.

TRIGGER

adaptive interventions when predictions are unreliable.

Rather than functioning as a black box with a static confidence threshold, Cognitive AI actively monitors its own decision-making and adjusts dynamically. It operationalizes explainability into an ongoing cognitive process.

Case study: automated radiographic weld inspection

Automated radiographic weld inspection replaces part of manual X-ray review with a deep learning model. On the line, welds are X-rayed, and a convolutional object-detection model, a YOLO-family or DETR-family architecture is typical, scans each radiograph for the defect classes defined by the governing weld-quality standard: porosity, slag inclusions, lack of fusion, incomplete penetration, undercut, and cracks.

Models of this kind reach overall accuracy in the mid-eighties on held-out test data, within the range a across the published literature for the task, and on the strength of that figure a model is trusted to triage incoming welds, passing the ones it classifies as sound, and flagging the rest for a human radiographer.

The figure is real. The conclusion drawn from it does not follow, and the gap between the two is where this kind of system fails.

Why the aggregate accuracy is dangerous here

Weld defects are not equally common, and they are not equally consequential. Porosity, gas pockets trapped in the weld metal, is relatively frequent and appears in radiographs as rounded, comparatively high-contrast features. A model trained on a representative dataset sees many examples and learns to detect it well; reported precision for porosity routinely exceeds eighty-five percent. Lack of fusion, similarly, is common enough and visually distinct enough to be learned reliably.

Cracks are the opposite case on every axis that matters. They are among the rarest defects in any real production dataset because good welding procedures are specifically designed to prevent them. They are the hardest to detect: a crack often presents as a thin, low-contrast line, sometimes narrower than two pixels in the radiograph, with irregular shape and fine branching that blends into background texture. Above all, they are the most consequential defect type: a crack is a propagating discontinuity that grows under load and is the classic initiation site for catastrophic fracture. The governing standard, ISO 5817, reflects this with a zero-tolerance rule: cracks are unacceptable at every quality level, regardless of size.

Put those three facts together and the problem becomes clear.

The model is most reliable on the defects that are common, easy to see, and comparatively tolerable. It is least reliable on the one defect class that is rare, visually subtle, and never permitted.

The mid-eighties accuracy figure is an average dominated by the porosity and fusion cases the model handles well. It quietly absorbs a much lower detection rate on cracks, exactly the failure that, in a pressure vessel or a structural member, the entire inspection exists to prevent.

From Explainability
to Cognition

The first generation of modern AI, statistical AI, focused on optimizing performance through scale: more parameters, more data, deeper networks. The second generation, explainable AI (XAI), sought to interpret model outputs, using saliency maps, feature attributions, and slice discovery to reveal how models behave. While valuable, these approaches remain diagnostic. They help humans analyze errors after the fact, but do not change how models make decisions.

Cognitive AI represents a third generation. It embeds reasoning within the system itself, enabling models to:

Map

the geometry of success and failure in training data.

DETECT

when an input falls into regions of ambiguity or uncertainty.

TRIGGER

adaptive interventions when predictions are unreliable.

Rather than functioning as a black box with a static confidence threshold, Cognitive AI actively monitors its own decision-making and adjusts dynamically. It operationalizes explainability into an ongoing cognitive process.

From Explainability
to Cognition

The first generation of modern AI, statistical AI, focused on optimizing performance through scale: more parameters, more data, deeper networks. The second generation, explainable AI (XAI), sought to interpret model outputs, using saliency maps, feature attributions, and slice discovery to reveal how models behave. While valuable, these approaches remain diagnostic. They help humans analyze errors after the fact, but do not change how models make decisions.

Cognitive AI represents a third generation. It embeds reasoning within the system itself, enabling models to:

Map

the geometry of success and failure in training data.

DETECT

when an input falls into regions of ambiguity or uncertainty.

TRIGGER

adaptive interventions when predictions are unreliable.

Rather than functioning as a black box with a static confidence threshold, Cognitive AI actively monitors its own decision-making and adjusts dynamically. It operationalizes explainability into an ongoing cognitive process.

Cognitive AI is The Next Scientific Frontier in Machine Intelligence

From Explainability
to Cognition

The first generation of modern AI, statistical AI, focused on optimizing performance through scale: more parameters, more data, deeper networks. The second generation, explainable AI (XAI), sought to interpret model outputs, using saliency maps, feature attributions, and slice discovery to reveal how models behave. While valuable, these approaches remain diagnostic. They help humans analyze errors after the fact, but do not change how models make decisions.

Cognitive AI represents a third generation. It embeds reasoning within the system itself, enabling models to:

Map

the geometry of success and failure in training data.

DETECT

when an input falls into regions of ambiguity or uncertainty.

TRIGGER

adaptive interventions when predictions are unreliable.

Rather than functioning as a black box with a static confidence threshold, Cognitive AI actively monitors its own decision-making and adjusts dynamically. It operationalizes explainability into an ongoing cognitive process.

From Explainability
to Cognition

The first generation of modern AI, statistical AI, focused on optimizing performance through scale: more parameters, more data, deeper networks. The second generation, explainable AI (XAI), sought to interpret model outputs, using saliency maps, feature attributions, and slice discovery to reveal how models behave. While valuable, these approaches remain diagnostic. They help humans analyze errors after the fact, but do not change how models make decisions.

Cognitive AI represents a third generation. It embeds reasoning within the system itself, enabling models to:

Map

the geometry of success and failure in training data.

DETECT

when an input falls into regions of ambiguity or uncertainty.

TRIGGER

adaptive interventions when predictions are unreliable.

Rather than functioning as a black box with a static confidence threshold, Cognitive AI actively monitors its own decision-making and adjusts dynamically. It operationalizes explainability into an ongoing cognitive process.

Why the model gives no warning

The natural assumption is that the model will at least be uncertain when it encounters a crack it cannot handle, and that low confidence will flag the case for review. It will not, and the reason is structural.

A crack that lies in a sparse, underpopulated region of the model's learned representation space does not produce a distinctive "I am unsure" signal. The model resolves the radiograph to the nearest familiar representation, and for a faint, thin crack against textured weld metal, the nearest familiar representation is frequently sound weld. The model returns a confident "no defect" because confidence reflects how strongly the model prefers that classification, not whether the input resembled anything it was adequately trained on. A confident pass on a genuinely sound weld and a confident pass on a missed hairline crack are, at the output, identical. Nothing in the score distinguishes them.

This is why retraining does not resolve the problem cleanly. Adding more crack examples helps where those examples land, but cracks are heterogeneous: orientation, length, branching, base material, and exposure all vary, and the next missed crack is the one whose appearance the expanded dataset still did not cover. The model's confidence on that new case will again be high. The failure does not announce itself; it has to be looked for, in a place the output does not expose.

Making the competence boundary visible, and enforcing it on the line

The problem here is that the team cannot see where the model's competence ends, and so cannot tell a trustworthy pass from a dangerous one. This is the gap Squint Vision Studio is built to close.

Rather than reading the model's reliability off its accuracy figure, Squint's Discovery tools map the model's decision-making space and distinguish the regions where its reasoning is well supported, the trusted regions, from the regions where inputs are sparse or overlapping, and error becomes likely. For weld inspection, this surfaces directly what the aggregate score hides: the crack class, and particularly the thin and atypical cracks, occupy ambiguous regions where the model's judgment cannot be relied on, even when its confidence is high. The team can see the boundary of competence before the model is trusted on the line, rather than discovering it through a field failure.

That visibility becomes an operational control. From the analysis, the team designs a watchdog, a module that recognizes when a given radiograph is being classified in one of those ambiguous regions, and exports it into the deployed system. When the model classifies a weld as sound but does so in a region where its reasoning is unreliable, the watchdog flags that specific decision for the human radiographer, regardless of the confidence the model attached to it. The common, well-supported porosity and fusion cases continue to flow through automatically. The rare, subtle, zero-tolerance crack cases, the ones the model is worst at and the standard least forgives are routed to the person best equipped to catch them.

It is worth being exact about what this addresses and what it does not. Mapping the model's competence boundary catches the cracks that fall in sparse or ambiguous regions of its decision space, the atypical, thin, or awkwardly-presented defects the model was never adequately trained to resolve, which is where the dangerous misses concentrate. It is not a guarantee of catching a crack whose radiographic appearance falls squarely inside the region the model treats as sound weld: if a defect genuinely looks, to the model, like well-supported normal material rather than like unfamiliar territory, a boundary map keyed to where the model's judgment is unsupported will not necessarily flag it. The boundary is a strong safeguard against the model being trusted where its reasoning is unreliable; a defect that presents as ordinary sound weld is a distinct and harder problem that a competence map does not claim to fully solve.

Cognitive AI is The Next Scientific Frontier in Machine Intelligence

From Explainability
to Cognition

The first generation of modern AI, statistical AI, focused on optimizing performance through scale: more parameters, more data, deeper networks. The second generation, explainable AI (XAI), sought to interpret model outputs, using saliency maps, feature attributions, and slice discovery to reveal how models behave. While valuable, these approaches remain diagnostic. They help humans analyze errors after the fact, but do not change how models make decisions.

Cognitive AI represents a third generation. It embeds reasoning within the system itself, enabling models to:

Map

the geometry of success and failure in training data.

DETECT

when an input falls into regions of ambiguity or uncertainty.

TRIGGER

adaptive interventions when predictions are unreliable.

Rather than functioning as a black box with a static confidence threshold, Cognitive AI actively monitors its own decision-making and adjusts dynamically. It operationalizes explainability into an ongoing cognitive process.

From Explainability
to Cognition

The first generation of modern AI, statistical AI, focused on optimizing performance through scale: more parameters, more data, deeper networks. The second generation, explainable AI (XAI), sought to interpret model outputs, using saliency maps, feature attributions, and slice discovery to reveal how models behave. While valuable, these approaches remain diagnostic. They help humans analyze errors after the fact, but do not change how models make decisions.

Cognitive AI represents a third generation. It embeds reasoning within the system itself, enabling models to:

Map

the geometry of success and failure in training data.

DETECT

when an input falls into regions of ambiguity or uncertainty.

TRIGGER

adaptive interventions when predictions are unreliable.

Rather than functioning as a black box with a static confidence threshold, Cognitive AI actively monitors its own decision-making and adjusts dynamically. It operationalizes explainability into an ongoing cognitive process.

What the shift changes

Without this, an inspection program is trusting a mid-eighties accuracy figure to stand in for a guarantee it cannot make: that the model is reliable on the defect it most needs to catch. The figure says the model is usually right. It does not say the model is right on cracks, and in fact the model is reliably worst exactly there. An inspection system built on the aggregate number is strongest where the stakes are lowest and weakest where they are highest, and nothing in its output reveals the inversion.

Seeing where the model's competence ends turns that inversion from an invisible liability into a managed one. The model keeps the throughput advantage on the cases it genuinely handles. The cases it cannot be trusted on are identified as such and sent to human judgment. The inspection is no longer governed by an average that flatters the model. It is governed by a map of where the model can be believed, which, for a weld that has to hold under load, is the only thing worth governing by.

The scenario described here is representative, grounded in published research on automated radiographic weld inspection; the defect taxonomy, model architectures, detection difficulty, the ISO 5817 zero-tolerance rule for cracks, and the reported porosity precision reflect the state of the field. Squint Vision Studio's capabilities, decision-space mapping and exportable watchdog modules are described as designed.