Program Integrity Math
Let's say that you are a government responding to a disaster – a flood, for example – and you want to direct resources to people who have been badly affected.
You think that there are 1,000 households that fall into this category and you'd like to give them $10,000 each. Helpfully, your Treasury department agrees and they give you a budget of $10,000 x 1,000 = $10 million.
But there is a problem. Your 1,000-households-affected number was just an educated guess: you don’t know who they are exactly. If you offer the $10,000 to all comers, many people who weren’t affected will be happy to take your money. This is an outcome the Treasury and your boss are very clear that they do not want.
The default strategy to deal with this is something like the following: Ask everyone affected by the disaster to apply for the money and to provide evidence that shows they are indeed eligible. Since you want to make sure fraudsters can't do the same thing, this will be onerous: forms, documents, photographs, signatures, and so on. You need to assign staff to review those submissions one by one and decide if they are legitimate.
There are other options too:
- You might ask people to apply for the money, but rather than insisting on evidence, you just make them promise that they are eligible. Now you task your team to audit some small proportion – 10%, say. If you catch anyone lying on their application then you throw the book at them.
- You don't ask people to apply at all. Instead, you use all your data – maps, satellite imagery, tax returns, just walking around – to make your best guess at who is affected, and then just send them the money.
- You might give up on the idea of identifying the very most deserving group altogether. You might instead set a low threshold for eligibility, and then select your recipients at random from those who meet it. If budgets allow, you could even decide to make the scheme universal and give the money to everyone.
This dilemma – we might call it the program integrity problem – comes up all the time, in and outside of government. As well as disaster relief, versions of it apply when allocating research grants, COVID support, asylum decisions, insurance premiums, and unemployment benefits.
What are we trying to solve?
How should we choose between our different integrity regimes?
We are essentially trying to balance and minimise two different sorts of cost:
- On the one hand you don't want to make mistakes. You want to avoid both giving money to people who aren't eligible (false positives), and refusing it to people who are (false negatives).
- On the other hand, you want to minimise the administrative burden, both on citizens applying for the money and on officials overseeing it. If this is too high relative to the amount of money being handed out, the whole exercise becomes a waste of resources.
As we increase the administrative burden – more evidence, more checks, more data – we should in theory (although not always in practice, as discussed below) be able to cut the error rate. Perhaps the relationship looks something like this:
In reality we obviously don’t know the exact shape of these curves. Perfect optimisation is not the goal. But we can aspire to choosing a design that gets the trade-off roughly correct for the context it’s deployed in. Some back-of-the-envelope maths can help us get there.
Program Integrity Math: a worked example
To start with, we need an estimate of how many false positives and false negatives we can expect from our regime. Ideally here you would have historical data – perhaps a survey of the population to work out who is applying, and then a deep dive into a sample of applications to see how good your scrutiny really is at telling eligible from ineligible.
But even a rough guess can help. Let’s say the population of our flooded town is 10,000 households, and of these 10% have been badly damaged and are eligible for our scheme. Let’s further say that all of these households will apply. Of the remaining 9,000 who aren’t eligible, we guess that 5% will apply.
Here's our population:
= 100 households = eligible applicants = ineligible applicants
Let’s say our application approval process is 90% accurate (more precisely, both its sensitivity and specificity – its likelihood of accepting eligible and rejecting ineligible applicants, respectively – are 90%).
A little spreadsheet work tells us that we will get 1,450 applications in total, of whom 1,000 are eligible and 450 not. 900 of the 1,000 eligible will be accepted, and 100 rejected (our false negatives). 405 of the ineligible group will be rejected, and 45 accepted (our false positives).
So our applicant pool looks like this:
= true positive (people who are elibile and received the money)
= true negative (people who are inelibile and were refused)
= false positives
= false negatives
Now we need to put a number on the social cost of a false positive and a false negative. In our flood example, perhaps we sit down with the minister and agree that the cost of a false positive (which undermines trust in the system) is $5,000 and of a false negative (which means that someone who deserves the money isn't getting it) is $3,000.
Multiplying out the numbers and cost of both types of error gives us a cost of $300k for the false negatives, and $225k for the false positives, for a total cost of errors of $525k.
For our final step, we need to estimate the administrative costs. Let’s say it costs $100 to process each application, and $200-worth of applicants’ time to submit one. If we only care about the social cost of submission for eligible applicants, our admin cost would $100 x 1,450 total applications = $145k for processing, plus $200 x 1,000 eligible applications = $200k for the applications, so $345k total admin cost:
| Cost | Number | Total Cost | |
| False Positives | $5,000 | 45 | $225k |
| False Negatives | $3,000 | 100 | $300k |
| Total Cost of Errors | $525k | ||
| Submitting (Eligible) Applications | $200 | 1,000 | $200k |
| Checking All Applications | $100 | 1,450 | $145k |
| Total Cost of Admin | $345k |
This puts the cost of errors at c.5% and of admin at c.3% of the $10m we are giving out. These are reasonable numbers, and the costs of errors and the costs of admin are within the same ballpark. It looks like our program integrity regime is calibrated about right.
Avoiding the worst failure modes
By plugging in different numbers, we can also illustrate situations in which our setup would no longer make sense. For example:
- What if the proportion of eligible households was much lower, say 1% of a 100,000-household population? Keeping the other numbers constant, now only 17% of our applicants would be eligible. Our false positive cost becomes $2.5m, and the total error plus admin cost would be 37% of the value of the money being given out. This is now very high! In this context we would need to toughen up our acceptance threshold. We might want to use cruder filtering mechanisms (e.g. only households in certain badly affected postcodes can apply at all) or harsher penalties for fraudulent applications to rebalance the applicant pool towards eligible households.
- What if rather than $10,000, we only had $1,000 to give out to each household? Even if our estimated cost of false positives and false negatives came down by the same ratio, in this scenario the admin costs of the scheme would predominate, at c.35% of the value of the money given out. Here it would be worth thinking about a less onerous way of doing the scrutiny: cutting steps from the application, or switching to an automatic allocation or lottery model.
Writing even rough numbers down helps avoid a program integrity regime that is totally mismatched to the circumstances.
Collecting easy wins
Once we’ve got the regime roughly right, we can of course still do better.
Our assumption above was that, even where it was overkill, all our admin and scrutiny was doing something useful. In reality this is often not the case: governments regularly impose pointless make-work on citizens which does little to deter or detect fraud, as research on ‘sludge’ has shown.
In these cases, doing proper statistical validation of each step and getting rid of pointless stuff can cut the administrative burden with little downside (in our flood example, maybe submitting the photo does all the fraud prevention heavy lifting, and most of the other parts of the form can be safely junked). Sensible checklists and decision rules can improve the speed and quality of officials’ checks.
There are generally similar cheap wins available on the application side as well. Simple nudges can cut the proportion of ineligible applications: clear and persuasive explanations of who is eligible and why; or reinforcing the social norm that most people don’t cheat the system. Just don’t bother getting people to sign at the top of the form.