ASCERTAINING AND PROMOTING THE BENEFITS OF AI
Anthropic’s Current Approaches — and What the Weavers Offers
A Briefing Paper
David Sutton CITP MBCS | July 2026
| Provenance — and why it is stated first This paper was produced with Claude Fable 5, Anthropic’s current frontier model, working with the full Weavers instrument attached as context: the Weavers Main document v51 including Appendix 1 (the thirty insights from the UK Industry 4 Transformation Strategy) and Appendix 2 (the Industry 4 model: eight global problems, the 17 UN Sustainable Development Goals, the seven Industry 4 domains); the ARIA Integrated Register v1.8; and the Cultural Values and Ethics Register v0.2. The human practitioner held the recognition function throughout: the AI generated, the practitioner tested, corrected, and selected. The provenance is stated first because the Ethics register’s second founding inversion applies directly to this document: can AI support the development of its own ethics — and by extension, the assessment of its own maker’s benefit claims? An Anthropic model analysing Anthropic’s benefit programmes is the goldfish bowl in live operation. The paper does not claim to have escaped that condition. It claims something more useful: that naming the condition, attaching an externally built analytical instrument, and holding human judgement over every conclusion is currently the most honest available response to it. Readers should weigh what follows accordingly. |
| Executive Summary Anthropic has built the most developed benefit-ascertainment apparatus of any frontier AI laboratory: systematic measurement of how its AI is actually used across the economy, a published policy framework for sharing the gains, a funded deployment programme aimed at those the market will not reach, and independent research enablement. This is genuine, well-resourced, and structurally serious work. Read through the Weavers, however, the apparatus measures adoption and distributes access — and benefit is neither of those things. The frame cannot yet see whether adoption is above or below the line, whether access builds sovereign capability or begins a dependency cascade, whether the golden thread from ‘raise the floor’ to specific programme decisions actually holds, and whether success is being measured against technology uptake or against the global problems and the SDGs that define what success means. These are not criticisms of intent. They are the questions the current frame cannot ask from inside itself. The Weavers offers the missing layer: a tested practitioner instrument for exactly this gap — the five golden thread tests as a programme audit, the dependency cascade and sovereign capability as design vocabulary, Appendix 2’s grounding of benefit in problems and SDGs rather than adoption rates, ARIA’s learning method for fluency programmes, and the Ethics register’s Data Protection parallel for making benefit claims independently examinable. |
1. Why Benefit Ascertainment Is Now a First-Order Question
For most of the current AI transition, the serious institutional work has concentrated on risk: safety research, evaluations, governance frameworks, red-teaming. Benefit has been asserted rather than ascertained — claimed in mission statements, assumed in adoption metrics, promised in visions of transformed medicine and education. That asymmetry is now closing. The laboratories, led visibly by Anthropic, are building apparatus to measure, distribute, and argue for the benefits of their systems with the same seriousness previously reserved for risk.
This matters for two reasons. First, whoever defines how benefit is measured defines what the technology is for — the measure becomes the mission. Second, the window in which those definitions are being set is the transition moment itself. The Ethics register carries the Data Protection parallel: every architecture decision made at a transition moment is a clock-setting decision, and resetting the clock after it is embedded costs orders of magnitude more than setting it correctly now. What is true of data architecture is equally true of benefit-measurement architecture. The definitions of benefit being institutionalised in 2026 will govern what the systems are optimised toward for years.
2. Anthropic’s Current Approaches — A High-Level View
Anthropic’s benefit work, as publicly documented at July 2026, runs on four strands, underpinned by one structural position.
Measurement. The Anthropic Economic Index uses a privacy-preserving analysis system to track how Claude is used across the economy — by occupation, task, geography, and interaction type — with the stated purpose of understanding AI’s economic impacts early enough for workers, employers, and policymakers to prepare. Its most consequential analytical distinction is between augmentation (collaborative use that extends the person’s capability) and automation (delegated use that replaces the person’s activity). Successive reports through 2025–26 introduced ‘economic primitives’ — task complexity, user and AI skill, autonomy, success rates — and in mid-2026 a linked survey connecting several thousand users’ stated experience of AI and work to their actual usage patterns.
Policy design. In June 2026 Anthropic published an Economic Policy Framework: a three-tier plan calibrated to the unemployment rate, preceded by three institutional prerequisites — measurement of AI adoption and its effects, a dedicated government analysis capability, and modernised delivery infrastructure. Its stated position is that the central challenge is not stimulating growth but ensuring the gains are widely shared, with a commitment to ‘pay our fair share’ if AI generates transformative returns.
Deployment for benefit. A dedicated Beneficial Deployments team runs the delivery side: a partnership with the Gates Foundation committing $200 million across global health, life sciences, education, and economic mobility, framed as extending AI’s benefits where markets alone will not; Claude Corps, a $150 million fellowship placing around 1,000 trained fellows into several hundred nonprofits for a year; discounted and credit-based access programmes for nonprofits, education institutions, and scientific laboratories; and AI fluency curricula co-developed with sector bodies. The team’s organising question, in its own language, is how to use AI to raise the floor, not just the ceiling.
External research enablement. The Economic Futures programme, extended to the UK and Europe in late 2025, funds independent research on AI’s economic impacts, alongside open data releases from the Index intended to let researchers and the public investigate the effects directly.
The structural position. Anthropic is constituted as a public benefit corporation with a stated dual mission — developing AI safely, and deploying it so the benefits reach all of society — and its safety research programme (alignment, interpretability, dual-use safeguards) is positioned as the enabling condition of benefit rather than a competing priority.
3. What Holds — The Weavers Reading in Anthropic’s Favour
An honest Weavers analysis begins by recognising what the current approaches get structurally right, because several of them independently arrive at positions the framework argues from first principles.
The measurement-first discipline matches Insight 30’s demand for a wider map: Anthropic is attempting to see the extended system — usage, geography, task structure, perception — before prescribing responses, and its policy framework makes measurement an explicit prerequisite of intervention. The augmentation/automation distinction is a genuine, operationalised step toward the question the Weavers considers decisive — whether AI extends capability or replaces it. The ‘raise the floor’ framing is a close cousin of the blue flower. And the willingness to publish findings that complicate the company’s own interests — uneven adoption, concentration of benefit among the already-advantaged, workers’ fears — is the flame taking on honest learning rather than concealing it. These are not decorative virtues. They are the preconditions of everything the Weavers would add.
4. What the Frame Cannot Yet See
Each of the following gaps is structural rather than accidental: it is a question the current apparatus cannot ask from inside its own frame, which is precisely the class of question the Weavers exists to surface.
4.1 Adoption is being measured; benefit is being inferred
The Index measures usage — tasks, occupations, interaction types. The benefit programmes measure reach — credits issued, fellows placed, organisations onboarded. Neither measures outcomes against an external definition of what success means. Appendix 2 supplies that definition and insists on it: the eight global problems are the demand side, the 17 SDGs are the internationally agreed success criteria, and the test of any initiative is which SDG it moves, by how much, and for whom. Key Message 2’s discipline applies directly: if the answer is ‘none’ or ‘we haven’t measured’, the initiative has not yet been justified. AI is the multiplying factor, and multiplying zero produces zero — an access programme that multiplies a weak underlying capability produces scaled weakness, not benefit.
4.2 Access is not sovereign capability — the dependency cascade applies to beneficiaries
The dependency cascade (Insight 10) describes how organisations outsource first the doing, then the understanding, then the direction, until they can no longer see what they have lost — because seeing would require the capability that has gone. Benefit programmes built on donated access are, by default, Phase 1 of this cascade for their recipients. A nonprofit whose core service now runs on subsidised frontier-model credits has gained capability and acquired a dependency in the same motion; the question is whether anything in the programme design builds its ability to understand, audit, challenge, and if necessary replace what it now depends on (Insight 24). Claude Corps embeds a trained fellow for a year — the Weavers question is what remains when the fellow leaves. Insight 32 names the design element the programmes currently lack: a deliberate transfer mechanism, through which the capability to work above the line is built in the recipient rather than performed on their behalf.
4.3 Augmentation/automation is a proto-measure of a deeper distinction
The Index’s augmentation category counts collaborative interaction. The Weavers’ above/below-the-line distinction asks something the interaction pattern alone cannot reveal: whether the person is working capably within the prevailing frame, or is able to question the frame itself — to see what is missing, name the elephant, identify the vine. A user can iterate collaboratively all day below the line. The published finding that output quality tracks the depth of what the user brings (Insight 20’s territory) points at this: the benefit of frontier AI concentrates where above-the-line capability already exists. If that is true, then a benefit strategy that distributes access without developing the orientation will widen the very gap it is designed to close — and the Index, measuring interaction types, will not see it happening.
4.4 The goldfish bowl — benefit claims assessed from inside the system
The Ethics register’s first founding inversion asks whether an organisation can develop ethics that examine the system it exists within, when examining that system honestly would require questioning its own position. Anthropic ascertaining the benefits of Anthropic is this inversion applied to benefit: the measure-maker, the deployer, and the beneficiary of the conclusions are the same institution. The register’s answer is not cynicism but architecture — the Data Protection parallel. What made data ethics operational was not aspirational principles but an independent function with protected standing and its own reporting line, and enforceable individual rights rather than organisational obligations. The equivalents for AI benefit — an independent function empowered to test benefit claims, and rights held by the people the benefits are claimed for — do not yet exist anywhere in the landscape. The clock-setting moment for building them is now.
4.5 The golden thread from mission to programme has not been traced
‘Raise the floor’ is a stated purpose. The five golden thread tests ask whether specific decisions connect to it. Does each deployment build or erode the recipient’s sovereign capability? Does it enable knowledge to flow across the sector, or create hub-and-spoke dependence on the provider? Is it designed for how the least-resourced organisation actually works, tested against frontline reality? What failure patterns is it at risk of repeating, and who can halt it if warning signs appear? Is it accountable to evidence and outcomes the beneficiaries themselves can see and challenge? A programme portfolio that cannot yet answer these thread by thread has a purpose and a set of activities, but not yet a traceable connection between them — and the Weavers’ finding across domains is consistent: where the thread cannot be traced, change is drift, not evolution, however well-resourced.
5. What the Weavers Offers
The Weavers is not a competing benefits programme. It is the practitioner instrument for the layer the current apparatus does not cover — built over three years through exactly the working method (deep domain expertise combined with frontier AI, Insight 23) that Anthropic’s own research identifies as the highest-value use of its systems. Five contributions follow directly from the analysis above.
A benefit audit instrument. The five golden thread tests, applied decision by decision across a benefit portfolio, convert ‘raise the floor’ from an aspiration into a testable design discipline. Each test produces a yes or a no; a programme that fails a thread is redesigned, not exempted.
An outcome grounding. Appendix 2’s three-tier model — problems, SDGs and the AI stack, Industry 4 domains — read from the top down, supplies the external success criteria that adoption metrics cannot: benefit defined against the problems the transformation exists to solve, with the blue flower as the binding test — is the least-included, most-complex-needs participant flourishing as a result? Not served, not reached: flourishing.
A design vocabulary for deployment. Appendix 1’s thirty insights give benefit programmes the concepts their design currently lacks: the dependency cascade and sovereign capability as the axis on which access either empowers or captures; structural rather than aspirational cooperation, so that knowledge flows between recipient organisations instead of radiating from the provider; design for reality rather than the idealised participant; and the transfer mechanism (Insight 32) as the difference between performing capability for a community and building it in one.
A learning method for fluency at scale. ARIA carries the theory of learning that AI fluency programmes need and conventional curricula cannot supply: begin at the impressive output, use AI to close the capability gap, work backwards into understanding — with sovereign capability as the added discipline that keeps the method from becoming Phase 1 dependency. For programmes such as Claude Corps and the fluency curricula, this is the difference between teaching tool use and developing the orientation that makes the tools worth having.
An honest ethics architecture for benefit claims. The Ethics register’s founding inversions and Data Protection parallel offer the governance design for the goldfish bowl problem: an independent examining function built before the architecture is embedded, and benefit expressed as enforceable expectations held by beneficiaries rather than commitments held by the provider. The register does not resolve its own hardest questions — it holds them open deliberately — which is precisely what makes it usable by an institution that cannot credibly claim to have resolved them either.
6. Propositions
For Anthropic, or any frontier laboratory serious about ascertaining rather than asserting benefit, the analysis yields five propositions. First: adopt an external outcome frame — the SDGs and the problem layer — as the published success criteria for beneficial deployments, and report against it. Second: add a sovereign-capability test to every access programme: what does the recipient understand, control, and retain if the provider withdraws? Third: extend the Index’s augmentation measure toward the above/below-the-line distinction — whether users are being enabled to question frames, not only to iterate within them. Fourth: establish the independent examining function for benefit claims now, at the clock-setting moment, on the Data Protection model rather than the advisory-board model. Fifth: treat practitioner-built instruments such as the Weavers as evidence and as method — the corpus itself demonstrates the benefit class the laboratory most wants to promote: one experienced practitioner, working above the line with a frontier model, producing in weeks what interdisciplinary teams would take months to reach.
| The Kokomo — held open If the institutions best placed to measure AI’s benefits are the institutions that produce the AI, what would it take for the definition of benefit to be held by the people the benefits are claimed for — and would the laboratories recognise that transfer of the definition as the largest benefit they could deliver? |
David Sutton CITP MBCS
davesutton19@gmail.com
Produced with Claude Fable 5 (Anthropic), with the Weavers Main v51 (including Appendices 1 and 2), the ARIA Integrated Register v1.8, and the Cultural Values and Ethics Register v0.2 attached as working context. Human recognition and judgement throughout by the author.
Anthropic material drawn from public sources at July 2026: the Anthropic Economic Index and reports (2025–26), the Economic Policy Framework (June 2026), the Gates Foundation partnership announcement (May 2026), the Claude Corps announcement (June 2026), and the Economic Futures programme (UK and Europe launch, November 2025). Figures are public commitments as announced.