NEXIVERIFY
The agentic value method  ·  Version 1.1  ·  1 August 2026

The general method, and how to take it to your own sector

Sector in. On stated assumptions, a labour cost, a headcount and a cycle time out, with a stamp on every input and a worked example carried end to end. This page is the method itself. Banking and insurance is one application of it, and there is nothing in that study a reader cannot rebuild here for a different sector.

It produces no savings promise. Every figure is a band, and a finance director can halve any of it in good faith. The point is to make that halving an argument about one named assumption rather than about the whole number.

The chain, in order
  1. 1QualifyFour yes/no gates decide whether a function is in scope
  2. 2VolumeA four-tier anchor ladder, where tier four means drop the function
  3. 3ThroughputPer head, sourced in a fixed order, with a named default
  4. 4CostFully loaded, in three pay tiers, scaled to the firm's own disclosure
  5. 5ShareCounted by where the rule lives, not picked
  6. 6Cycle timeBefore is sourced or the time delta is not reported
  7. 7LinesCash out, loss avoided, revenue enabled, never summed
  8. 8PhasingA curve off a programme that finished and published its years
The tailoring formula

Five lines of arithmetic, and where every term comes from

Nothing below is specific to banking or to insurance. Substitute your own sector's numbers in order, stop at the first term that does not resolve, and the chain tells you whether the sector carries a priceable line at all.

addressable population
headcount × share of the workforce in qualifying functions
wage bill in that work
addressable population × fully loaded cost per head, by pay tier
comes off, a year
wage bill × share taken (low / mid / high) × (1 running cost)
lands by year n
comes off × the phasing curve 22 / 43 / 70 / 91 / 100
the clock, separately
value in flight × days removed ÷ 365 × cost of funding
TermWhere the number comes from TierStamp
Headcount The firm's own disclosed staff number. Almost every large regulated firm publishes one, even where it publishes no functional split. 1CONFIRMED_EXTERNAL
Share of the workforce in qualifying functions The four gates, applied function by function. Where the firm publishes no split, a supervisory survey's functional share of payroll stands in for it. 2WEB_SOURCE
Fully loaded cost per head Disclosed staff cost divided by disclosed headcount, then split across three pay tiers in a published occupational ratio and rescaled so the weighted average reproduces the disclosure. The level is the firm's own. Only the shape is imported. 1COMPOSED_HYPOTHESIS
Share taken A count of qualifying task classes, tiered by where the rule lives. Low, middle and high are three provenance tiers, not three confidence guesses. COMPOSED_HYPOTHESIS
Running cost 20, 10 and 5% of the human cost each workstream replaces, across the three cases. The agent's own compute plus the supervision that stays. Deducted from every figure before it is shown. COMPOSED_HYPOTHESIS
The phasing curve The mean of two published year-by-year ramps, from cost programmes that finished and disclosed each year against a stated total. Neither was an AI programme, so the curve is an analogue and is labelled as one wherever it appears. 1CONFIRMED_EXTERNAL
Value in flight The reader's own balance sheet. It is the one input nobody outside the firm can supply, which is why no figure on any of these pages multiplies it out. 1supplied by the reader
Days removed A published turnaround statistic, a clock written into a rule, or a disclosed service level. An assumed before-time makes the after-time meaningless, so with none of the three the cost delta is reported and the time delta is not. 1–3WEB_SOURCE
Cost of funding Stated rather than sourced, at 4.0% in the first application. Every clock figure moves one-for-one with it, so a reader who disagrees rescales the line in one step. COMPOSED_HYPOTHESIS
The gates decide what enters the first line, and only the gates. A written rule held by both parties, a deterministic output under it, a volume anchor at tier one to three, and a disputant working to a clock. A function that fails any one of the four is dropped rather than estimated. A derived figure then carries the weakest stamp among its inputs, which is why every line above resolves to COMPOSED_HYPOTHESIS the moment the share is applied. That is the answer, not a hedge.

Step one. Qualify the function

Decompose by who disputes the number, not by org chart. A function qualifies when all four gates return yes. Four reads off documents, no judgement call.

GateTestFails when
Written ruleThe rule was written by somebody other than the party doing the work, and both sides can hold the same copyThe rule lives in an expert's head
DeterministicThe figure is a pure function of committed inputs under that rule. Same inputs, same answer, any machine, any dayDiscretion, or a private source that cannot be re-obtained
Volume anchorAn annual count of events exists at tier one, two or three belowNothing published and no peer ratio
ClockSomebody downstream challenges the figure against a deadline: an exam cycle, an appeal window, a filing dateNobody ever asks

Work that fails the deterministic gate is routed rather than excluded. It carries a receipt instead of a re-derivation, and it never enters the cost figures.

Step two. The volume anchor

Take the highest tier available. Never blend two tiers into one number.

TierWhere it comes fromStamp
DisclosedThe firm's own annual report, regulatory return or statistical return states the countCONFIRMED_EXTERNAL
Regulator aggregateA supervisor publishes the sector total. Divide by the firm's disclosed share of premium, assets or revenueWEB_SOURCE
Peer ratioA published peer discloses events per unit of revenue or per employee. Apply that ratio to the target's disclosed base, and state ratio and base separatelyCOMPOSED_HYPOTHESIS
NothingDrop the function
Tier four is a real outcome, not a failure. A method that lets an unanchored line through permits anything. Six of the ten sectors below land at tier four on volume, and the honest consequence is that the cost line is dropped in all six rather than filled with a plausible number.

Step three. Throughput per head

Sourcing order, stopping at the first that resolves.

RankSource
FirstA regulator or professional body's own burden study. The strongest source, because it was built to be argued with
SecondAn industry body benchmark with a stated sample
ThirdA survey figure already denominated in hours, which makes throughput unnecessary
FourthNothing. Return to tier four above and drop the function

Productive hours per full-time equivalent per year: 1,800 by default, permitted range 1,600 to 2,000, stamped COMPOSED_HYPOTHESIS. It is a convention. Moving it inside the range moves every headcount by about a tenth, which is worth saying before a reader finds it. The sensitivity is set by rule rather than chosen: the low case runs at twice the default throughput, the high case at half.

Step four. Fully loaded cost per head

RankSource
FirstThe firm's own disclosed staff cost divided by disclosed headcount. Almost every large regulated firm publishes both, even where it publishes no functional split
SecondA supervisory survey's functional share of payroll, multiplied by disclosed total staff cost. Compliance runs 6–10% of payroll at the largest banking institutions and 11–15.5% at the smallest quartile WEB_SOURCE
ThirdA national sector wage index times a loading factor. Default loading 1.30 on base salary, permitted range 1.25 to 1.45 COMPOSED_HYPOTHESIS

One blended figure is the weakest form of this step. A processing clerk does not cost what a model validator costs, and the two sit in different proportions in every workstream. Split the population into clerical, professional and managerial in the published ratio 1 : 1.73 : 3.58, then scale all three by a single factor so the whole population reproduces rank one above. The level stays the firm's own disclosure. Only the shape is imported, and the tier mix per workstream is the assumption a reader should argue with3.

A per-firm cost never reaches a client surface. It is converted to a verified sector average or a share of a published base, and no firm is named beside a figure.

Step five. The share the arithmetic settles

Not a percentage anybody picks. It is a count of qualifying task classes, tiered by where the rule lives, so the band states rule provenance rather than confidence.

CaseWhat is countedWhy the tier holds
LowRe-derivable arithmetic and evidence assembly under a rule pinned by digest, held in the same version by both partiesThe rule cannot be swapped, so nothing is contested but the arithmetic
MidLow, plus work under an internal written rule the firm controlsIt re-derives, but the firm still writes the rule it is graded against
HighMid, plus two-party reconciliations where each side holds its own copy of the ruleHighest value, longest build, most exposed to a mismatch between the copies

Judgement work never enters any case. Materiality, scope, hardship and a call between two defensible readings stay with the person who answers for them.

Catch rate and false-positive rate are never stated. Neither has been measured. Both carry TARGET until the benchmark runs, and implying either on a client surface is the one error in this method that cannot be walked back.

Step six. Cycle time

Before is sourced or dropped. Three acceptable sources, in order: a published turnaround statistic, a clock written into a rule, a disclosed service level. An assumed before-time makes the after-time meaningless, so a function with no sourced before-time reports a cost delta and no time delta.

After is bounded below by the measured verify time and is never quoted as a total cycle. A clean verify runs at 0.335s median across 40 pilot bundles, 0.214s to 0.413s, wall clock including interpreter startup INTERNAL_BENCHMARK. Bundles are small and no scaling curve has been fitted.

The only defensible sentence: the check stops being the long pole. What remains on the clock is the exception queue and the judgement. Anyone claiming a whole cycle collapses to seconds is claiming the judgement disappeared.

Step seven. The three line classes

Each class carries its own label wherever it appears. A headline that sums them without saying so is the failure this step exists to prevent.

ClassWhat it isRule
Cash outCost that leaves the payroll or the external fee lineThe only class allowed in a cash figure
Loss avoidedLeakage, penalty, rework, restatementNeeds a sourced anchor of its own or it does not appear. Never in a cash figure
Revenue enabledNew business the record makes possibleNever cash, never in a headline, always carried with the volume it depends on

Step eight. Phasing

A run rate is not a first-year figure, and the gap between the two is where most of these estimates break. Take the curve from a programme that finished and published its year-by-year realisation. Never assert one.

Published rampYr 1Yr 2Yr 3Yr 4Yr 5
One large European bank, cost programme, share of its eventual total CONFIRMED_EXTERNAL14%31%60%83%100%
One large Swiss bank, integration cost saves, share of its stated ambition CONFIRMED_EXTERNAL30%56%79%100%100%
Modelled, the mean of the two COMPOSED_HYPOTHESIS22%43%70%91%100%

Anything faster than the faster of the two needs its reason written on the page. Neither ramp was an AI programme, and no AI programme has published one, so the curve is an analogue and is labelled as one. No programme publishes a first-quarter realisation either, so quarter one is taken as an eighth of the year-one share, back-loaded inside the year. It is the weakest cell on the curve4.

Price the build as a quote, stage by stage. Never borrow a merger's cost-to-achieve ratio. An integration pays redundancy and merges two system estates; a deployment does neither, and where the checker is open source there is no licence to price at all. What is left is the producer-side emitter layer, and it is quoted in stages: a pilot on one workstream in one market, the plumbing built once, the first workstream into production, then a lower unit price for every workstream after it. Cost tracks the number of system estates the inputs live in, not the size of the saving. Severance for the roles released is a separate class under step seven and never enters the cash line COMPOSED_HYPOTHESIS.
Ten sectors, one gate at a time

What the four gates return, sector by sector

Every one of the ten passes the written-rule gate, the deterministic gate and the disputant gate. The volume anchor is where they separate, and it separates them hard. Four carry an anchor at tier one or two and can be priced. Five have nothing published and the cost line is dropped. One has not been searched and says so rather than being promoted.

Gate 1 2 3 4 Volume anchor Share taken AI assurance tier 4, dropped no published share Carbon and ESG tier 2 no published share Engineering tier 4, dropped no published share Finance tier 2 29%, middle case Healthcare tier 2 no published share Infrastructure tier 4, dropped no published share Insurance tier 1 32%, middle case Pharmaceuticals tier 4, dropped no published share Public sector not searched no published share Telecommunications tier 4, dropped no published share

Gate 1 a written rule both parties can hold. Gate 2 a deterministic output under it. Gate 3 a volume anchor at tier one, two or three. Gate 4 a disputant working to a clock. Every sector clears the first, second and fourth. The third is where they separate. Open one of the ten below for the qualifying functions, the anchor, the disputant and the honest limit.

Every figure above sizes a problem. All of them come from public regulation, published surveys and operator disclosures rather than from a result of ours. Four of the ten resolve into a cash line. Five are dropped at the volume gate and one is carried open, and that ratio is the most useful thing on this page: most sectors do not have the anchor, and the method says so instead of filling the cell.

The chain, run once in full

The worked example, carried end to end

Finance. Internal control testing and the evidence run behind it, chosen because both anchors are published at sector level, so a reader outside the building can check every step.

StepWhat resolvedStamp
QualifyThe test script is written by somebody other than the tester, a recomputed control result either matches or does not, hours are published, and the external auditor disputes it against the filing datePasses all four gates
Volume15,580 staff hours per programme per year, tier twoWEB_SOURCE
ThroughputSkipped. The anchor is already in hours. Implied headcount at the 1,800-hour default: 8.7 full-time equivalentsCOMPOSED_HYPOTHESIS
Cost$2.3m average programme cost, up 44% year on year. Implied blended rate $147.62 an hourWEB_SOURCE
Cross-checkA supervisor publishes its own examination cost-recovery rate at $137 an hour. The implied rate sits 7.8% above it, and the survey figure carries non-labour cost the recovery rate does not, so the implied rate is an upper boundCONFIRMED_EXTERNAL
ShareLow 35%, middle 55%, high 70% of the hours. Scoping, materiality, deficiency severity and the management conclusion are excluded from every caseCOMPOSED_HYPOTHESIS
Cycle timeBefore: an annual cadence, and a 438-day average restatement period when the annual check missed something. Two clocks, kept apart. After: the check itself runs in under half a second on a pilot bundleWEB_SOURCE INTERNAL_BENCHMARK
LinesCash out: $1.27m and 4.8 full-time equivalents a year on the middle case. Loss avoided, stated separately and never added in: about $16m average negative net-income impact, for more than half of restating companies in one yearCOMPOSED_HYPOTHESIS
Every worked example should carry a cross-check. A chain nobody can check independently is a chain nobody believes. Here two figures from unrelated publishers land within 8% of each other, which is a sanity check rather than an identity.
For an eleventh sector

The template, and a blank cell is a stop rather than a gap

Copy the block, complete it in order, and stop at the first line that does not resolve. Nothing below is optional. A cell filled with a plausible number is worse than an empty one, because a reader who finds it unaided stops believing the rest of the page.

Fill-in template
SECTOR:
FUNCTION:
  G1 written rule, held by both parties ....... [ ]  which document:
  G2 deterministic under that rule ............ [ ]
  G3 volume anchor, tier 1/2/3 ................ [ ]  value + stamp:
  G4 disputant + clock ........................ [ ]  which clock:

  VOLUME ...................................... value / stamp / tier
  THROUGHPUT .................................. value / stamp / source rank 1-3
  HEADCOUNT ................................... volume / throughput / 1,800
  COST PER HEAD ............................... value / stamp / source rank 1-3
  FUNCTION COST ............................... headcount x cost per head
  SHARE low / mid / high ...................... % / % / % + the task-class table
  CROSS-CHECK ................................. an independent rate or ratio, and the gap
  CYCLE BEFORE ................................ value / stamp / which of the three sources
  CYCLE AFTER ................................. the check, never the whole cycle
  LINES ....................................... cash out / loss avoided / revenue enabled

  RUNNING COST ................................ 20 / 10 / 5% of the human cost replaced
  PHASING ..................................... 22 / 43 / 70 / 91 / 100, labelled an analogue
  BUILD ....................................... pilot / plumbing / first workstream / each after

STAMPS: CONFIRMED_EXTERNAL  WEB_SOURCE  INTERNAL_SOURCE  INTERNAL_BENCHMARK
        TARGET  COMPOSED_HYPOTHESIS  UNVERIFIED
A derived figure carries the weakest stamp among its inputs.
Catch rate and false-positive rate are permanently TARGET. Never state either.
Two lines that stop the work rather than slow it. A function whose rule lives in somebody's head fails gate one and drops out. A function with nothing published at tier one to three fails gate three and drops out. In most sectors those two together take out a meaningful slice of the work, and the slice they take out is the honest answer to what an agent can be held to.

What version 1.1 closed, and what is still open

HoleHow it closes
A leakage line with no anchorStep seven. A loss-avoided line needs a tier one to three anchor of its own or it is absent, and it never enters a cash total
Headcount assumed, because firms publish a total staff number and no functional splitStep four, second rank. The functional share of payroll comes from a supervisory survey, so the split is sourced at sector level even where the firm publishes nothing. An assumption becomes a citation with a stated range
A phasing curve nobody sourced, and one blended cost per headStep eight and the pay-tier split in step four. Two banks have published a full year-by-year ramp against a stated total, and the pay ratio comes off a published occupational wage table instead of a guess
Still openEvidence-assembly hours have no published time-and-motion source in any of the ten sectors, so every share in step five rests on a classification. A first engagement recording hours by task class is the single highest-value measurement available to this method

What this method cannot do

It measures integrity, never correctnessA faithfully re-derived wrong-but-committed rule re-derives, and the figure counts as agreement. That is correct behaviour and it is also the limit
It cannot price a rule that lives in somebody's headThe first gate fails and the function drops out. In most sectors that is a meaningful slice of the work
It produces no savings promiseEvery cost figure is a band on stated assumptions, and a finance director can halve any of them in good faith
It says nothing about a rule that is itself wrongPerfect application of a bad rule is perfect application of a bad rule

The line every figure sits on

The agent's own output is never re-run. Language models are not bit-reproducible, even at temperature zero. What the agent carries is a show-its-work receipt: the sources used, the criteria included and excluded, the model and its version. What re-derives exactly is the deterministic arithmetic downstream of it, recomputed offline on the doubter's own machine by the verifier's own code.

Integrity, never correctness. A record settles that a figure is the deterministic result of the committed computation over the committed inputs. Whether the rule behind that figure was the right rule stays where it already is. Nothing here maps to a compliance claim, and no page of this method makes one.