Fawad is CEO and co-founder of Penguin Ai, a healthcare AI startup that sells to payers, providers and revenue cycle management companies. gettySomewhere in most health systems, there's an AI tool an internal team built and almost nobody uses. The demo was impressive. The board saw it, and the innovation team presented it at a conference. Eighteen months later it still runs. But nobody owns it, and no one can say exactly when people stopped trusting it.The post-mortem almost always includes the same phrase. We should have bought instead of built. It’s a reasonable conclusion. It's also the wrong one, which is why the organizations that swear off building tend to repeat the failure with a new tool a year later.The Build Never The Hard PartBuilding has never been cheaper. A capable team can stand up a working prototype in a few weeks. The demo will look polished, because demos always do. The hard part comes after, when the tool has to survive real workflows and earn the trust of clinicians. Most tools don't survive that test. When researchers inventoried every AI project at one large academic health system, they counted 87. Only four had made it into sustained clinical service. The other 83 sat in development or early implementation, some of them for years. Those numbers should unsettle anyone whose AI strategy assumes the pilot phase is the risky part.I watched this play out from the buyer's chair for years. As chief data officer at Kaiser Permanente, UnitedHealthcare and Optum, I spent hundreds of millions of dollars on technology and learned the hard way that most of it was never built for healthcare. The failures I saw almost never happened during development. They happened later, in daily operations, and they took down builds and buys at the same rate.What Kills These ProjectsGeneral-purpose AI models know healthcare's vocabulary. They don't know its logic. A model that can summarize a discharge note will still fall apart the moment you ask it to convert clinical text into structured, billing-ready data. In one recent benchmark, leading models attempting exactly that task topped out at a score of 0.363 out of 1. The same results appear in head-to-head testing, where models pretrained on clinical data outperformed general-purpose ones across most healthcare workflows. You can't prompt a model into understanding a prior authorization it has never seen.Internal teams usually figure this out about a month in. They respond the only way they can. They start teaching healthcare to the model themselves. Payer rules get hand-coded, while clinical edge cases get patched one at a time, usually after someone downstream catches the mistake. The tool that took three weeks to demo needs six months of tutoring before anyone else will rely on it. The second killer arrives after the launch announcement, and it's a bit quieter than the first. Models degrade. Payer policies change and clinical practice moves, and a system trained on last year's patterns slowly stops matching the reality of care. Researchers studying how health systems maintain their AI describe a "responsibility vacuum," where monitoring is ad hoc and maintenance is nobody's job. Even organizations that do monitor often can't see the problem. In one study of nearly 240,000 chest X-rays, standard performance metrics failed to flag drift that was clinically obvious.So, the tool still runs, but the analyst who understood its logic left, and the staff who once used it, have gone back to the manual process they trust. Failure Becoming More ExpensiveStates have begun regulating how AI touches coverage and care decisions, and scrutiny is coming for any system that influences a clinical or financial outcome. A recent review of hospital AI governance found that while most U.S. hospitals run predictive models, only about half assess those systems for bias and roughly two-thirds check them for accuracy. There’s also a human cost. In that same 87-project health system, clinicians described innovation fatigue as the accumulated weariness of contributing hours to pilots that never return value to their work. Every abandoned build makes the next one harder to launch because clinicians have learned to expect abandonment. A Better Question Than Build Or BuyWhen an internal build fails, many organizations will switch sides in the build-versus-buy debate. I'd push leaders toward two different questions. Does the foundation underneath the system understand healthcare, with payer logic and clinical rules built in? And who will own it in year two? Both questions apply to anything you build or buy. Without the right foundation, everything you build costs 10 times more and scales much slower. The person closest to the work, like the nurse who has worked through a 200-page case file for a knee replacement and knows exactly whether the documentation should be approved, is best positioned to catch what the system missed. Her judgment is the scarcest asset in healthcare technology, and for years the industry has built its tools for everyone but her.The inventory I mentioned earlier counted 87 projects and four survivors. Multiply that across a few thousand health systems and we see the scale of these failures. Healthcare has never lacked people who understand the work. The waste comes from asking them, year after year, to build on foundations that don't.Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?
Why Build Vs. Buy Is The Wrong Question For Healthcare AI
Every abandoned build makes the next one harder to launch because clinicians have learned to expect abandonment.






