The Problem We Couldn't Ignore
A buyer files a dispute: "Item not as described. The bag in the photo was leather. What arrived is plastic." The seller responds: "Item matches the listing exactly. Buyer is lying."
Two conflicting narratives. One order record. One blurry product photo. Somewhere in there is the truth — but no human agent will read it carefully at 11 PM on a Friday.
Online marketplaces collectively handle hundreds of millions of disputes every year. The resolution process at most platforms is, charitably, a keyword filter with a confidence score hardcoded to sound decisive. The systems that call themselves "AI-powered" are almost always rule engines that optimize for throughput over correctness. On genuinely ambiguous cases, they guess.
We kept asking: what would it look like if the AI actually reasoned over a dispute the way a careful adjudicator would? And — critically — what would it look like if it knew when to stop and say "I'm not sure"?







