The Weavers Test

A diagnostic for systems that use AI: what to strengthen, and what to watch

Short paper · working note

Purpose

The Weavers Test is a way of assessing any system that uses AI. It does not return a single verdict of safe or unsafe. It returns a diagnostic: a reading of where the system is strong and where it is weak, and — the part that matters most — a sorting of its weaknesses into two kinds. Some weaknesses can be engineered away by applying simple rules. Others cannot be removed because they are intrinsic to what the system is, and the only response to them is sustained attention. The test’s real output is this distinction: here is what to strengthen, and here is what to never stop watching.

This is more useful than a gate. A gate tells you stop or go. A diagnostic tells you what to do, and concentrates attention on the one or two places where vigilance is the only defence rather than diffusing it across everything at once.

The five criteria

The test applies five criteria to a system. The distinction running through them — between errors in outcomes and errors in the system — is deliberate, because the two are detected and corrected by different means and most assessments blur them.

#CriterionWhat it asks of an AI-using system
1Risk and rewardCan the level of risk and reward the system carries be gauged?
2Errors in outcomesCan errors and their impacts be identified in the system’s outputs?
3Errors in the systemCan errors be identified within the system itself, not only its outputs?
4Resilience to outcome errorsIs the system resilient when an output turns out to be wrong?
5Correcting system errorsCan errors within the system be corrected once found?

Each criterion can be graded by severity (for example on a one-to-five scale). A grade is only meaningful if its levels are defined by an observable property rather than an impression — for instance, criterion 2 at its most severe level might be defined as: outcome errors cannot be detected without independently re-deriving the result. Sharp definitions are what allow two assessors to score the same system and converge, rather than each expressing a prior intuition.

Relationship to existing frameworks

The five criteria are real and each has an established literature, but they live in five different fields, and no single existing framework assembles them into one per-system diagnostic. That assembly is the contribution.

  • Risk and reward — graded by risk-management frameworks (such as the NIST AI Risk Management Framework) and regulatory risk tiers, which categorise a system’s risk but do not give a per-system, per-axis reading.
  • Errors in outcomes vs. errors in the system — addressed separately by observability (detecting wrong outputs) and interpretability (locating why they occurred); the test treats them as one paired question.
  • Resilience to outcome errors — the territory of safety engineering: fault tolerance, graceful degradation, defence-in-depth, and formal safety cases.
  • Correcting system errors — the concern of model governance and corrigibility: can the system be patched, rolled back, or corrected without resistance.

The closest single relatives are safety cases (a structured argument that a system is acceptably safe) and assurance levels such as the automotive ASIL and aviation DAL, which scale required scrutiny to the severity of failure. The Weavers Test borrows their core idea — scrutiny proportional to consequence — and adds what they lack: an explicit reading of whether a system’s errors are detectable and correctable by procedure, or only by deep human expertise. A system whose errors can be found only by one knowledgeable practitioner is more dangerous than one whose errors trip an automated alarm, even when both produce identical outputs. No standard framework grades that property, and it is often the decisive one.

Worked example: applying the test to the Weavers itself

The test is illustrated here on the Weavers, using the recent Southport regeneration analysis as the case. Turning the instrument on itself is also the honest first move: a test worth trusting should be willing to indict the thing that made it.

Initial reading (unconstrained use)

  • Risk and reward — high (level 5). The initial assumption was that replacing the town’s tourism offer with a similar one would deliver the regeneration required. A different perspective has the potential to identify an alternative or parallel approach — so both the reward of a better path and the risk of missing it are large.
  • Errors in outcomes — hard to detect. The limitation in the initial perspective could only be found by a knowledgeable practitioner with in-depth knowledge of the Weavers, of AI assistants, and of the town. The error is real but low in visibility.
  • Errors in the system — hard to locate. Because the Weavers and ARIA hold many registers and details, identifying a particular inaccuracy is difficult without in-depth knowledge.
  • Resilience to outcome errors — poor. If the Weavers had reinforced an existing strategy, a change in town strategy would later be required to recover — a costly correction. This is the most dangerous failure mode: a confident, articulate endorsement of a flawed frame.
  • Correcting system errors — moderately simple. A mis-stated register entry can be fixed directly.

Revised reading (with the disciplines of use applied)

The initial reading scored an unconstrained Weavers. The disciplines that actually govern its use change several scores materially.

  1. Used to raise questions, never to validate. The Weavers is used to raise questions for investigation, not as a validation mechanism. A question carries no endorsement, so the worst failure mode — confident confirmation of a flawed frame — is largely designed out. Resilience rises substantially, conditional on the discipline being held: the moment an output is read as confirmation rather than as a question to investigate, the failure mode returns.
  2. Registers kept simple and reviewed. The registers are continually reviewed so they do not become too large or complex to hide their impacts — the instrument refusing to grow its own vine. System-error detectability rises, because complexity in which errors can hide is capped.
  3. The output carries its own provenance. The Weavers can identify which registers and information it used, and the language itself points to the areas drawn on. Errors become localisable — a wrong output can be traced to its source and inspected. This improves detection of register-level errors. It does not reveal whether a correct register was applied soundly, which is a different class of error.
The residual the test isolates The disciplines push every avoidable risk down toward acceptable. One risk remains and cannot be removed: a skilled practitioner may produce a sophisticated inversion that traces cleanly to a simple, correct register, reads as genuine, and is in fact decorative — an articulate question that the frame did not actually need. This error lives in the practitioner’s judgement, not in any register, so no rule reaches it. It is the recognition function itself, which is tacit and human. The only mitigation is the practitioner’s continual honesty.

That residual is not a flaw to be hidden; it is the honest floor of the instrument, and locating it precisely is the test working as intended. A well-disciplined Weavers is low-risk on every axis except the one that is its own deepest nature — and on that axis the defence is vigilance, not engineering.

Where the test is most useful

The same exercise is sharper still on two classes of system of current concern: the use of AI and natural language in business intelligence, and AI-driven systems such as digital twins.

Both share a dangerous property the test is built to catch: they produce fluent, authoritative outputs whose errors are low in detectability. A natural-language business-intelligence system that misreads a request returns a confident, plausible number — the outcome error is nearly invisible, and the underlying misinterpretation can be spotted only by someone who knows both the data and the question. A digital twin renders a confident, complete-looking model whose blind spots are hidden because the rendering looks complete. Applied to these systems, the test returns high risk specifically on the detection criteria — which is the useful, discriminating result: the danger is not that they fail loudly, but that they fail silently and with authority, the worst failure profile there is.

In each case the test does the same work: it clears away the risks that rules genuinely handle, so attention is not diffused, and it concentrates vigilance on the irreducible residual where vigilance is the only defence. For the Weavers that residual is decorative judgement; for natural-language BI and for digital twins it is the fluent, authoritative, wrong output that no amount of tracing catches. Naming that one place precisely — and refusing to let it hide behind the risks that have been handled — is the test’s purpose.

A note on the test itself

The test embodies the principle it applies: it raises the questions that direct investigation rather than delivering a verdict, and it points attention at the barely-visible consequence rather than letting it be lost behind the impressive, handled risks. That makes it subject to its own findings. A test that scores systems will feel rigorous, and feeling rigorous is what genuine and decorative assessments share. Its proof is practical: whether two practitioners scoring the same system independently converge, and whether a score ever changes a decision. If the test makes someone strengthen what they would have shipped, or watch what they would have trusted, it is real. If it produces an elegant risk profile that everyone admires and no one acts on, it is the frame admiring itself — which is the one thing it exists to catch.

Working note. The five criteria, the system-versus-outcome distinction, the detectable-by-procedure-versus-by-expertise dimension, and the strengthen-versus-watch sorting are the load-bearing elements. The worked example is illustrative; the intended targets are AI in business intelligence and AI-driven digital twins.