The agentic value method · Version 1.1 · 1 August 2026
The general method, and how to take it to your own sector
Sector in. On stated assumptions, a labour cost, a headcount and a cycle time out, with a
stamp on every input and a worked example carried end to end. This page is the method itself.
Banking and insurance is one application
of it, and there is nothing in that study a reader cannot rebuild here for a different sector.
It produces no savings promise. Every figure is a band, and a finance director can halve
any of it in good faith. The point is to make that halving an argument about one named assumption rather
than about the whole number.
The chain, in order
- 1QualifyFour yes/no gates decide whether a function is in scope
- 2VolumeA four-tier anchor ladder, where tier four means drop the function
- 3ThroughputPer head, sourced in a fixed order, with a named default
- 4CostFully loaded, in three pay tiers, scaled to the firm's own disclosure
- 5ShareCounted by where the rule lives, not picked
- 6Cycle timeBefore is sourced or the time delta is not reported
- 7LinesCash out, loss avoided, revenue enabled, never summed
- 8PhasingA curve off a programme that finished and published its years
The tailoring formula
Five lines of arithmetic, and where every term comes from
Nothing below is specific to banking or to insurance. Substitute your own sector's numbers in
order, stop at the first term that does not resolve, and the chain tells you whether the sector
carries a priceable line at all.
- addressable population
- headcount × share of the workforce in qualifying functions
- wage bill in that work
- addressable population × fully loaded cost per head, by pay tier
- comes off, a year
- wage bill × share taken
(low / mid / high) × (1 − running cost)
- lands by year n
- comes off × the phasing curve
22 / 43 / 70 / 91 / 100
- the clock, separately
- value in flight × days removed ÷ 365 × cost of funding
The gates decide what enters the first line, and only the gates. A written rule
held by both parties, a deterministic output under it, a volume anchor at tier one to three, and a
disputant working to a clock. A function that fails any one of the four is dropped rather than
estimated. A derived figure then carries the weakest stamp among its inputs, which is why every line
above resolves to COMPOSED_HYPOTHESIS the moment the share is applied. That is
the answer, not a hedge.
Step one. Qualify the function
Decompose by who disputes the number, not by org chart. A function qualifies when all four gates
return yes. Four reads off documents, no judgement call.
Work that fails the deterministic gate is routed rather than excluded. It carries a receipt instead of
a re-derivation, and it never enters the cost figures.
Step two. The volume anchor
Take the highest tier available. Never blend two tiers into one number.
Tier four is a real outcome, not a failure. A method that lets an unanchored line
through permits anything. Six of the ten sectors below land at tier four on volume, and the honest
consequence is that the cost line is dropped in all six rather than filled with a plausible number.
Step three. Throughput per head
Sourcing order, stopping at the first that resolves.
Productive hours per full-time equivalent per year: 1,800 by default, permitted range 1,600 to
2,000, stamped COMPOSED_HYPOTHESIS. It is a convention. Moving it inside the
range moves every headcount by about a tenth, which is worth saying before a reader finds it. The
sensitivity is set by rule rather than chosen: the low case runs at twice the default throughput, the
high case at half.
Step four. Fully loaded cost per head
One blended figure is the weakest form of this step. A processing clerk does not cost what a
model validator costs, and the two sit in different proportions in every workstream. Split the population
into clerical, professional and managerial in the published ratio 1 : 1.73 : 3.58, then scale all three
by a single factor so the whole population reproduces rank one above. The level stays the firm's own
disclosure. Only the shape is imported, and the tier mix per workstream is the assumption a reader
should argue with3.
A per-firm cost never reaches a client surface. It is converted to a verified sector average or a share
of a published base, and no firm is named beside a figure.
Step five. The share the arithmetic settles
Not a percentage anybody picks. It is a count of qualifying task classes, tiered by where the rule
lives, so the band states rule provenance rather than confidence.
Judgement work never enters any case. Materiality, scope, hardship and a call between two defensible
readings stay with the person who answers for them.
Catch rate and false-positive rate are never stated. Neither has been
measured. Both carry TARGET until the benchmark runs, and implying either on a
client surface is the one error in this method that cannot be walked back.
Step six. Cycle time
Before is sourced or dropped. Three acceptable sources, in order: a published turnaround
statistic, a clock written into a rule, a disclosed service level. An assumed before-time makes the
after-time meaningless, so a function with no sourced before-time reports a cost delta and no time
delta.
After is bounded below by the measured verify time and is never quoted as a total cycle. A clean
verify runs at 0.335s median across 40 pilot bundles, 0.214s to 0.413s, wall clock including
interpreter startup INTERNAL_BENCHMARK. Bundles are small and no scaling curve
has been fitted.
The only defensible sentence: the check stops being the long pole. What remains
on the clock is the exception queue and the judgement. Anyone claiming a whole cycle collapses to
seconds is claiming the judgement disappeared.
Step seven. The three line classes
Each class carries its own label wherever it appears. A headline that sums them without saying so is
the failure this step exists to prevent.
Step eight. Phasing
A run rate is not a first-year figure, and the gap between the two is where most of these estimates
break. Take the curve from a programme that finished and published its year-by-year realisation. Never
assert one.
Anything faster than the faster of the two needs its reason written on the page. Neither ramp was an AI
programme, and no AI programme has published one, so the curve is an analogue and is labelled as one.
No programme publishes a first-quarter realisation either, so quarter one is taken as an eighth of the
year-one share, back-loaded inside the year. It is the weakest cell on the
curve4.
Price the build as a quote, stage by stage. Never borrow a merger's cost-to-achieve ratio.
An integration pays redundancy and merges two system estates; a deployment does neither, and where
the checker is open source there is no licence to price at all. What is left is the producer-side
emitter layer, and it is quoted in stages: a pilot on one workstream in one market, the plumbing
built once, the first workstream into production, then a lower unit price for every workstream
after it. Cost tracks the number of system estates the inputs live in, not the size of the saving.
Severance for the roles released is a separate class under step seven and never enters the cash
line COMPOSED_HYPOTHESIS.
Ten sectors, one gate at a time
What the four gates return, sector by sector
Every one of the ten passes the written-rule gate, the deterministic gate and the disputant gate. The
volume anchor is where they separate, and it separates them hard. Four carry an anchor at tier one or
two and can be priced. Five have nothing published and the cost line is dropped. One has not been
searched and says so rather than being promoted.
- AI assurance gate 3 fails · tier 4, dropped
- Carbon and ESG all four pass · tier 2
- Engineering gate 3 fails · tier 4, dropped
- Finance all four pass · tier 2 · 29%
- Healthcare all four pass · tier 2
- Infrastructure gate 3 fails · tier 4, dropped
- Insurance all four pass · tier 1 · 32%
- Pharmaceuticals gate 3 fails · tier 4, dropped
- Public sector gate 3 not searched
- Telecommunications gate 3 fails · tier 4, dropped
Gate 1 a written rule both parties can hold. Gate 2 a deterministic output
under it. Gate 3 a volume anchor at tier one, two or three. Gate 4 a disputant working to
a clock. Every sector clears the first, second and fourth. The third is where they separate. Open one
of the ten below for the qualifying functions, the anchor, the disputant and the honest limit.
-
01
AI assurance
no published share
Evaluation scoring, usage metering and retrieval-corpus builds re-derive from the item-level results under a declared rule. The copilot answer itself never does, and carries a receipt instead.
The four gates what each one returns
- Written rule, yes, and the weakest yes of the ten. The scoring rule and the grounding policy are written and versioned, though usually by the same firm that runs the model.
- Deterministic, yes on the scoring leg. A reported score is a pure function of the item-level results. The generation step fails by construction and is routed to a receipt.
- Volume anchor, no. No annual count of agent outputs is published, per firm or per sector. Tier four, so the cost line is dropped rather than estimated.
- Disputant, yes. The assurance function and, from 2 August 2026, a person told they are talking to a machine.
What the arithmetic gets and what it does not
- No cash line. The chain stops at gate three, so the sector is carried as a receipt obligation and not as a saving.
- What is published is a residual error rate, roughly 0.7 to 1.5% unsupported claims on the easiest public benchmark. That sizes the problem and is never a share taken.
- Over 40% of agentic projects are forecast cancelled by the end of 2027 on cost and inadequate controls. A loss-avoided class, kept out of any cash figure.
The anchor
No volume anchor published
tier 4 · UNVERIFIED
The clock
AI Act Art 50 · in force 2 Aug 2026
Art 12 record-keeping 2 Dec 2027 · WEB_SOURCE
The receipt settles that the answer cites real, unaltered sources under a declared policy. The model is never re-run. Nobody can promise identical output twice, so nothing here does. What changes is detectability, not the error rate.
-
02
Carbon and ESG
no published share
Scope one to three aggregation, disclosed datapoint derivations and intensity metrics re-derive from activity data and the declared factor set. A proxy-estimated supplier gap carries a receipt.
The four gates what each one returns
- Written rule, yes. The reporting standard and the factor set are published by somebody other than the reporter, and the assurance provider holds the same copy.
- Deterministic, yes. Tonnes are activity times factor under a declared aggregation.
- Volume anchor, yes, tier two. Ongoing assurance runs about €320,000 a year with about €287,000 to stand up, and limited assurance sits at 20 to 30% of the financial-audit fee.
- Disputant, yes. The limited-assurance provider, working to the reporting deadline.
What the arithmetic gets and what it does not
- A cash line is available, because the anchor is already denominated in money. Throughput is skipped for the same reason the worked example skips it.
- The share stays a classification. No operator has published a result for how much of an assurance workpaper an agent takes over.
- What changes is scope, not speed. A re-derivation tests the whole computation where a tie-out samples it.
The anchor
~€320k a year, ~€287k to stand up
tier 2 · WEB_SOURCE
The clock
First FY2027 limited-assurance engagement
assurance stays limited · WEB_SOURCE
The share
No operator figure published
COMPOSED_HYPOTHESIS
It settles that the tonnes are the committed data times the committed factors under the committed rule. It will not launder a guess into an assured number, and an estimate stays labelled an estimate with its method on the record.
-
03
Engineering
no published share
The safety-factor arithmetic over the solved deck re-derives, together with load-case combinations, code margins and mass rollups. The solver run itself is matched inside a declared tolerance and never claimed byte-exact.
The four gates what each one returns
- Written rule, yes. The material allowables, the design code and the acceptance margin are written by somebody other than the analyst.
- Deterministic, yes on the sign-off, no on the field. Parallel floating-point reductions are not associative, so an identical stress field is off the table and is never claimed.
- Volume anchor, no. No annual count of design releases is published. Tier four, so the cost line is dropped.
- Disputant, yes. The design review board, and in a regulated submission the reviewer working against the credibility record.
What the arithmetic gets and what it does not
- No cash line. Cost and cycle deltas specific to simulation sign-off are unresolved, and they are carried that way.
- What the gate does buy is a named failure mode. A wrong material card silently voids the safety factor, and no revision log or signature catches it.
- Re-verifying that a number came from a specific deck is a manual forensic exercise today, which is why it essentially never happens before a field failure.
The anchor
No volume anchor published
tier 4 · UNVERIFIED
The clock
The design-release gate
computational-modelling guidance final 16 Nov 2023 · WEB_SOURCE
Byte-exact re-derivation covers the arithmetic over the exact files that were solved. A solver re-run is judged inside the tolerance you declare. Whether the model is valid physics stays with the analysts, made auditable rather than automatic.
-
04
Finance
29%
Interest accrual, trade settlement, net asset value and the deterministic capital and reserving runs re-derive. A financial-crime alert or a credit decision carries a receipt instead.
The four gates what each one returns
- Written rule, yes. The control test script, the day-count convention and the reporting template are written elsewhere and both sides hold the same copy.
- Deterministic, yes. A recomputed control result either matches the asserted one or does not.
- Volume anchor, yes, tier two. One control-testing programme runs 15,580 staff hours and about $2.3m a year at a large filer, in-scope systems up from 17 to 40 while automated controls fell from 21% to 17%.
- Disputant, yes. The external auditor against the filing date, and the validation function against the supervisory standard.
What the arithmetic gets and what it does not
- The full chain resolves. This is the sector the worked example below is drawn from, and the only one where every term has a source outside the building.
- 29% of the wage bill in scope on the middle case, in the application study. Eight of its twelve kinds of work sit on a published operator result, four rest on analogy and are marked there.
- A cross-check exists, which is the part worth keeping. A supervisor's own examination cost-recovery rate lands within 8% of the rate the survey implies.
The anchor
15,580 staff hours · $2.3m a year
tier 2 · WEB_SOURCE
The clock
SR 26-2 superseded SR 11-7 · 17 Apr 2026
plus the annual filing cycle · CONFIRMED_EXTERNAL
The share
29% of the wage bill in scope
middle case · COMPOSED_HYPOTHESIS
Deterministic compute only, and integrity rather than correctness throughout. Whether the accrual convention or the monitoring scenario is the right one stays the firm's call, made auditable.
-
05
Healthcare
no published share
Coverage determinations, claims adjudication, member cost-share and risk-adjustment submissions re-derive from the committed policy and the member's own record. A triage model carries a receipt.
The four gates what each one returns
- Written rule, yes. The coverage criteria and the fee schedule are licensed or filed documents, and the provider can hold the same copy.
- Deterministic, yes. Cost-share is a pure function of the allowed amount and the plan terms.
- Volume anchor, yes, tier two. About 53 million determinations in one national programme in 2024, manual handling at $10.97 a transaction against $5.79 electronic, and about 13 hours a week per physician.
- Disputant, yes. The member and the provider on appeal, against a seven-day standard and a 72-hour urgent window.
What the arithmetic gets and what it does not
- The strongest anchor set of the ten. Three independent burden and cost figures, denominated in events, dollars and hours.
- No operator has published a share taken, so the classification stands alone and is the first thing a pilot should measure.
- 80.7% of appealed denials are overturned. Much of that is new documentation on appeal rather than an arithmetic failure, so it sizes the exposure and is never read as a share.
The anchor
~53m determinations · $10.97 against $5.79
tier 2 · WEB_SOURCE
The clock
CMS-0057-F operative · 1 Jan 2026
a specific reason for every denial · WEB_SOURCE
The share
No operator figure published
COMPOSED_HYPOTHESIS
It settles that the determination followed the declared policy over the declared record. It does not judge whether the coverage policy is good medicine. Clinicians own that, and the record makes their decision contestable on the day it issues.
-
06
Infrastructure
no published share
Safety-interlock predicates, protection-relay settings and settlement arithmetic re-derive from the reviewed source, offline and inside the air-gap. A predictive-maintenance advisory carries a receipt.
The four gates what each one returns
- Written rule, yes. The interlock predicates and the setpoints are declared and reviewed before anything reaches a controller.
- Deterministic, yes. The compiled artefact either is the build of the reviewed source or is not, and a declared predicate either holds over every trace or returns a counter-example.
- Volume anchor, no. No annual count of control-logic sign-offs is published. Tier four, so the cost line is dropped.
- Disputant, yes. The functional safety manager, at the factory and site acceptance tests.
What the arithmetic gets and what it does not
- No cash line. The downtime figures that circulate are a loss-avoided class, the strongest of them is about a decade old, and none is carried here.
- The bottleneck moved. It did not disappear. A copilot drafts a function block in seconds where an engineer took half an hour, so assurance is now what sets the pace.
- Interlock correctness today rests on finite-sample testing, not on a proof over every trace. That gap is what a declared predicate closes.
The anchor
No volume anchor published
tier 4 · UNVERIFIED
The clock
Acceptance testing, then 2 Dec 2027
high-risk obligations on a fixed date · WEB_SOURCE
It settles that the logic is the committed build satisfying the predicates that were declared. It does not say the hazard analysis is complete or the plant safe. Supplementary evidence, never a replacement for a certified assessment.
-
07
Insurance
32%
Solvency capital under the standard formula, deterministic reserving and roll-forward, premium rating from filed tables and treaty settlement all re-derive. A fraud referral carries a receipt.
The four gates what each one returns
- Written rule, yes. The filed rate table and the policy wording are written before the work, and the supervisor holds the same copy.
- Deterministic, yes on actuarial compute. An internal-model simulation re-derives only with a pinned seed, and that condition is stated wherever the figure appears.
- Volume anchor, yes, tier one. Carriers disclose headcount, staff cost and policy counts in their own returns, which is what the application study's insurer figure is built from.
- Disputant, yes. The supervisor, inside a statutory six-month decision window on a complete internal-model application.
What the arithmetic gets and what it does not
- The full chain resolves, and off tier-one disclosure instead of a sector aggregate. The cleanest fit of the ten on both rates that matter.
- 32% of the wage bill in scope on the middle case, against 29% at the bank, and 61% of the workforce in scope against 41%.
- Compare the percentage, never the total. The absolute figure scales with payroll, which says more about the firm's size than about the work.
The anchor
The firm's own disclosed headcount and staff cost
tier 1 · CONFIRMED_EXTERNAL
The clock
Six-month supervisory window
model calibrated to 99.5% over one year · WEB_SOURCE
The share
32% of the wage bill in scope
middle case · COMPOSED_HYPOTHESIS
It settles that the figure is the committed calculation over the committed inputs. Whether the reserve is adequate or the calibration sound is a different question and stays exactly where it was.
-
08
Pharmaceuticals
no published share
Efficacy and safety tables specified in the analysis plan re-derive over hash-locked data in a pinned environment, seeds included, alongside batch-record and stability arithmetic. A drafted narrative carries a receipt.
The four gates what each one returns
- Written rule, yes. The statistical analysis plan is written and locked before the data are.
- Deterministic, yes, conditionally. The environment has to be pinned. An unpinned version or option set returns an inconclusive verdict rather than a pass, which is the whole reason the third verdict exists.
- Volume anchor, no. No annual count of locks, amendments or pivotal tables is published. Tier four, and a market-size projection is not a substitute for one.
- Disputant, yes. The inspector and the reviewing statistician, against the submission date.
What the arithmetic gets and what it does not
- No cash line. The double-programming step is close to universal and is nowhere counted, so the population cannot be sized from published material.
- The old way is already a re-derivation done by hand. Two programmers re-code every pivotal result and reconcile, repeated at every lock and every amendment.
- Data integrity is a leading theme in a rising warning-letter count, which is a trigger and not a volume, so it is carried as one.
The anchor
No volume anchor published
tier 4 · UNVERIFIED
The clock
Database lock ahead of a submission
and every protocol amendment after it · WEB_SOURCE
It settles that the table is the deterministic result of the committed derivation over hash-locked data in a pinned environment. It does not validate the trial design or the choice of method, and it is not yet validated for production use in a regulated environment.
-
09
Public sector
no published share
Eligibility determinations, benefit amounts, overpayment recalculations and indexation runs re-derive from the published rule and the citizen's declared facts. A fraud-risk score carries a receipt.
The four gates what each one returns
- Written rule, yes, and the cleanest of the ten. The rule is published legislation, so the claimant can hold the identical copy.
- Deterministic, yes. A benefit amount is a pure function of the declared facts under that rule.
- Volume anchor, not searched. Determination counts may sit in a national statistical return and none has been read for this file, so the cell is carried open rather than dropped. That is a different outcome from tier four and is never promoted into one.
- Disputant, yes. The claimant, inside the appeal window.
What the arithmetic gets and what it does not
- No cash line yet. It is the one sector of the ten where the anchor is probably obtainable and simply has not been obtained.
- Improper payments run about $162bn in one national programme year. A loss-avoided class under step seven, so it never enters a cash figure.
- What changes is timing, not outcome. The exact rule becomes inspectable and challengeable at issuance instead of reconstructed years later.
The anchor
Not searched, carried open
UNVERIFIED
The clock
Annex III high-risk · 2 Dec 2027
plus the appeal window · WEB_SOURCE
It settles that the determination followed the published rule as declared. Whether the rule itself is just belongs to the legislature and the courts. A rule that is internally consistent and still unlawful re-derives, which is correct behaviour and is also the limit.
-
10
Telecommunications
no published share
Usage rating, mediation, interconnect and roaming settlement, invoicing and taxation all re-derive from the declared records and a rate card pinned by digest. A fraud flag carries a receipt.
The four gates what each one returns
- Written rule, yes, and the roaming leg is the strongest in the set. The caps are external and published to the cent: €1.30 a gigabyte in 2025, €1.10 in 2026 and €1.00 from 1 January 2027, voice at €0.019 a minute and text at €0.003, read against the primary text.
- Deterministic, yes. Rating and cap application are arithmetic with nothing stochastic in them.
- Volume anchor, no. Record counts and wholesale volumes are not published per operator. Tier four, so the cost line is dropped.
- Disputant, yes. The roaming partner, inside the record rejection-and-return window. Miss the window and the charge stands.
What the arithmetic gets and what it does not
- No cash line, but the cleanest two-party structure of the ten. The host both rates the traffic and bills for it, and no feature release fixes that.
- Fraud runs about $28.3bn a year globally, of which about $5.04bn is international revenue-share fraud. Loss avoided, never added to a cash line.
- The transferred-record formats sit behind a member gateway and could not be read, so no field spelling from them is quoted as confirmed.
The anchor
No volume anchor published
tier 4 · UNVERIFIED
The clock
The monthly transfer-and-return window
caps to 30 June 2032 · CONFIRMED_EXTERNAL
A wholesale customer with no engineering capacity will not run a checker, however free it is. Where both sides are content with the clearing house, the arithmetic is not in dispute and there is nothing to sell. Ask whether they have ever built a shadow rating engine.
Every figure above sizes a problem.
All of them come from public regulation, published surveys and operator disclosures rather than from a
result of ours. Four of the ten resolve into a cash line. Five are dropped at the volume gate and one is
carried open, and that ratio is the most useful thing on this page: most sectors do not have the anchor,
and the method says so instead of filling the cell.
The chain, run once in full
The worked example, carried end to end
Finance. Internal control testing and the evidence run behind it, chosen because both anchors are
published at sector level, so a reader outside the building can check every step.
Every worked example should carry a cross-check. A chain nobody can check
independently is a chain nobody believes. Here two figures from unrelated publishers land within 8% of
each other, which is a sanity check rather than an identity.
For an eleventh sector
The template, and a blank cell is a stop rather than a gap
Copy the block, complete it in order, and stop at the first line that does not resolve. Nothing below
is optional. A cell filled with a plausible number is worse than an empty one, because a reader who
finds it unaided stops believing the rest of the page.
Fill-in template
SECTOR:
FUNCTION:
G1 written rule, held by both parties ....... [ ] which document:
G2 deterministic under that rule ............ [ ]
G3 volume anchor, tier 1/2/3 ................ [ ] value + stamp:
G4 disputant + clock ........................ [ ] which clock:
VOLUME ...................................... value / stamp / tier
THROUGHPUT .................................. value / stamp / source rank 1-3
HEADCOUNT ................................... volume / throughput / 1,800
COST PER HEAD ............................... value / stamp / source rank 1-3
FUNCTION COST ............................... headcount x cost per head
SHARE low / mid / high ...................... % / % / % + the task-class table
CROSS-CHECK ................................. an independent rate or ratio, and the gap
CYCLE BEFORE ................................ value / stamp / which of the three sources
CYCLE AFTER ................................. the check, never the whole cycle
LINES ....................................... cash out / loss avoided / revenue enabled
RUNNING COST ................................ 20 / 10 / 5% of the human cost replaced
PHASING ..................................... 22 / 43 / 70 / 91 / 100, labelled an analogue
BUILD ....................................... pilot / plumbing / first workstream / each after
STAMPS: CONFIRMED_EXTERNAL WEB_SOURCE INTERNAL_SOURCE INTERNAL_BENCHMARK
TARGET COMPOSED_HYPOTHESIS UNVERIFIED
A derived figure carries the weakest stamp among its inputs.
Catch rate and false-positive rate are permanently TARGET. Never state either.
Two lines that stop the work rather than slow it. A function whose rule
lives in somebody's head fails gate one and drops out. A function with nothing published at tier one to
three fails gate three and drops out. In most sectors those two together take out a meaningful slice of
the work, and the slice they take out is the honest answer to what an agent can be held to.
What version 1.1 closed, and what is still open
What this method cannot do
The line every figure sits on
The agent's own output is never re-run. Language models are not bit-reproducible, even at
temperature zero. What the agent carries is a show-its-work receipt: the sources used, the criteria
included and excluded, the model and its version. What re-derives exactly is the deterministic
arithmetic downstream of it, recomputed offline on the doubter's own machine by the verifier's own
code.
Integrity, never correctness. A record settles that a figure is the deterministic result of the
committed computation over the committed inputs. Whether the rule behind that figure was the right rule
stays where it already is. Nothing here maps to a compliance claim, and no page of this method makes
one.