Nobody decides to stop checking invoices. What happens is that the close takes a week, three documents per bill is forty minutes an hour of it, and at some point the honest description of the process becomes: check a portion, and assume the rest are like it.
That is sampling, and as a response to a time budget it is entirely rational. As a control it has a specific and predictable failure.
Sampling finds what is common, and misses what is rare
A sample tells you about the distribution. It is good at "our freight charges have crept up 3% across the board" — a pattern that shows up in almost any subset you pull.
The things a three-way match exists to catch are not distributed like that. They are rare and individually expensive:
- one duplicate bill, paid twice, at full face value
- one pallet-handling fee added at the bottom of one invoice, then every invoice
- one line nobody ordered
- one short delivery where the order and the bill agree with each other and disagree with the dock
A 10% sample finds a 10%-frequency problem reliably and a one-off almost never. And the one-off is the one that costs the money.
The second failure is quieter
A sample that comes back clean tells you the sampled bills were clean. It gets reported as "the close is clean", because that is the sentence the process was there to produce.
Nobody is lying. The claim just quietly widens somewhere between the check and the summary, and there is no artifact anywhere that records how wide the check actually was.
The interesting thing about full coverage is what it removes
The argument for checking all of them is not "be thorough". It is that once every bill is compared, the exceptions are the only output — so the queue is a complete statement rather than a sample of one. Nothing was skipped, which means nothing has to be assumed about what was skipped.
That only works if the checking is free. Comparing three documents line by line is arithmetic: quantities against quantities, unit prices against the agreed price, totals against price times quantity received. It is exactly the kind of work that should have finished before anyone woke up.
MatchRail's demo book is a month of purchasing — fifty-one bills, of which forty-two agree and nine do not. The nine are what you see. The forty-two are the point: they were all checked, and you did not have to look at any of them.
What full coverage does not mean
It does not mean the machine decides. Everything outside tolerance is queued for a person, with all three documents in front of them and the disagreement written out in words — "billed 24, received 21", "ordered at $12.50, billed at $14.10". You are reading the disagreement, not taking a tool's word for it.
And nothing is written back to your ledger on a tap. An approved correction is queued with an expiry, waits behind a kill window, writes its own audit row, and is re-checked against the current documents before it posts. If the figure moved after you approved it, it does not post at all.
Full coverage on the checking. None on the deciding. That is the trade the week you do not want is supposed to buy you.