Author(s): Hossain Pazooki
Originally published on Towards AI.
Causal estimators for loan decisions cannot be graded on real lending data, because the data never contains the answer. Planting the truth in a synthetic world — and then trying hard to break your own harness — can be.
A biased loan approval model doesn’t just cut into lender profits. It skews the incentives that govern access to capital: who gets to start a business, who gets to refinance out of trouble, whose default history becomes the training data for the next model. And the uncomfortable part is that the bias is largely unmeasurable from inside the system that produces it. You cannot ask historical data what would have happened to the applicants you declined, and you cannot ask it whether your model’s counterfactual reasoning — if this borrower had linked a bank feed, what would their default probability be? — is right or wrong. The label you would need to check the claim was never generated.
I think this is where AI tooling can genuinely level the playing field. Not by promising an unbiased model — nobody can deliver that — but by making the failure modes of the modeling pipeline measurable and falsifiable: demonstrating precisely where a lender’s causal claims stop being trustworthy, and ruling out the scenarios the data cannot rule out on its own.










