Building Trust in Healthcare AI

You Can’t Vibe Code Trust

Leading healthcare tech companies of the future will act like AI partners, not AI vendors.

The customer shouldn’t have to see the gears.

The new trust gap

A few days ago, a customer of ours (the CEO of a skilled nursing company) texted me yet another new “AI MDS analyzer” to check out and give him an opinion on.

I never even had to tap the link to the website to already know it was vibe coded AI slop. 

The preview image that unfurled inside the text thread was a default preview URL for a platform called Lovable, not a preview image for the actual website itself, but for the vibe coding platform it was created on. The image just screamed “AMATEUR” on something being marketed as enterprise AI.

I just shook my head and kinda laughed. But it crystallized something I’ve been chewing on for a year: the thing that used to make software companies great has moved. And most of the market, both the builders and the buyers alike, hasn’t noticed yet.


“The thing that used to make software companies great has moved. And most of the market…hasn’t noticed yet.”

Programming by vibes

Some quick context: In February 2025, Andrej Karpathy coined the term "vibe coding," which is where you describe what you want in plain English, let AI write the software, and (in his words) “forget the code even exists.” By November, Collins Dictionary had named it Word of the Year. That’s how fast this went from a tweet to a movement.

Lovable is the movement’s poster child. Just type a prompt, get a working full-stack app with a database, auth, and a deploy button. The Swedish company went from launch to $100M in ARR in eight months, the fastest ramp in software history, crossed $400M earlier this year with 146 employees, and is reportedly in talks to raise at a $12 billion valuation. Cursor, Claude Code, Codex, Windsurf, Replit, Bolt, v0… all the same wave. The capital markets noticed: AI captured roughly half of all global venture funding in 2025 — about $211 billion, up 85% in a single year.

First, let me be clear about something, because it matters for everything that follows: I love these tools. Our engineers at Carebility ship with Claude Code, Codex, and Cursor every day. I vibe code prototypes on weekends for fun. Collapsing the distance between an idea and a working application is the most exciting thing to happen to software in my career. But rapidly generated software is not the problem.

The problem is what happens when prototype-quality systems get marketed into regulated healthcare; into provider facilities full of protected health information, survey windows, OIG audits, and the reimbursement that keeps the lights on.

The data here is not subtle. Veracode has been testing AI-generated code for two years across more than 150 models, and the pattern is remarkable: models now produce syntactically correct code over 95% of the time, and introduce a known security flaw about 45% of the time, a rate that hasn’t improved across any testing cycle. Last year, researchers scanned 1,645 apps built on Lovable and found 170 with critical flaws exposing real user data (and the platforms themselves are candid that security is ultimately the builder’s responsibility). Which is fair! They’re prototyping tools. The security, the governance, the operational discipline…that was always going to be somebody’s job. The vibes don’t do it for you.

So here’s the actual danger, and it isn’t the tools: a polished prototype and a serious AI system look identical for the first fifteen minutes of a demo. Buyers can no longer tell the difference by looking. That’s genuinely new. And in healthcare, the cost of guessing wrong isn’t a bad quarter. It’s a breach notification with residents’ names in it.

When software was the hard part

I started with SimpleLTC as CTO in March of 2010, and it’s worth remembering how different the terrain was.

In those days, SaaS was still an argument we had to win. "Wait, our residents’ data will live... on the internet?" Browser extensions felt genuinely innovative in 2015, and most people had never even heard of them. RPA was a dark art practiced by specialists, and it broke every time a vendor moved a button. Integrations were priced like construction projects (months of work and five-figure invoices). Getting an MDS assessment from a facility’s computer into CMS’s systems, reliably, at scale, was a moat all by itself.

The difficulty was the barrier to entry. If you could build enterprise healthcare software that actually worked, you’d already outrun most of the field. We rode that for years, and by the time my tenure ended after Netsmart acquired SimpleLTC, we had processed more than 60 million MDS assessments and had 8,500+ facilities using our products, and every single one carried someone’s care plan and someone’s reimbursement. The software was hard, and that hardness protected us.

That world is gone. Today a motivated product manager can generate in an afternoon what took my 2012 team a quarter. Software creation is being commoditized in real time, and I think that’s wonderful.

But it means the moat isn’t where we left it.

You can’t vibe code the next SimpleLTC

I hear a version of this from founders and investors almost weekly: with today’s tools, someone could rebuild what you built at SimpleLTC in ninety days.

They’re half right. You could rebuild the beautifully simple UI screens in ninety days (the forms, the dashboards, the workflow). But the screens were never the company. The company was fifteen years of earned trust: every edge case in the RAI manual, every state-by-state quirk, every submission-window fire drill, every last-minute PBJ submission at 11:59 Eastern time on Valentine’s Day. And here’s the uncomfortable part for those of us who lived that era: the new version of "hard" in the AI era is even harder than ours was.

Because the scarce discipline now is AI systems engineering, and it looks much less like old-school web development.

It’s about evaluation frameworks: running thousands of tests to check the AI’s judgment, not just its code. It’s managing model drift so the system doesn’t get dumber between March and July. It’s layered architecture to stop prompt injections, and zero-retention BAAs to keep PHI where it belongs. It’s hard engineering, not just a clever prompt.

It’s model orchestration: choosing the right tool for the job and having a fallback for when an API goes down. It’s managing tokens as COGS and designing for humans-in-the-loop so nurses can make fast, informed calls. And it’s a bulletproof audit trail, because "the AI did it" won’t fly with a surveyor or the OIG auditor.

You can vibe a demo. You cannot vibe an audit trail.

Nowadays almost anyone can generate software, but almost no one can operate secure, governed AI systems. That gap is the whole game now.


You can vibe a demo. You cannot vibe an audit trail.

Your customers buy outcomes; stop selling them tokens

There’s a second thing serious AI companies will have to do, and this one is a choice, not a capability: start owning customer risk instead of passing it through.

Look at how AI is being priced right now. Input tokens. Output tokens. Cached tokens. Context-window tiers. Model classes. Credit packs that expire. Overage math. It reads like a 2004 cell phone bill, and it quietly pushes every ounce of uncertainty onto the customer (model price changes, provider swaps, even the vendor’s own inefficiency).

Now picture the actual buyer in post-acute care. She’s an administrator running a 120-bed building, or a regional MDS director covering fourteen facilities with three coordinator seats she can’t fill. She thinks in PPD, PDPM per diems, census, and survey risk and lookback periods. She has never once thought in tokens, and she never should. Asking her to forecast token consumption is asking her to underwrite your bad engineering.

This industry has never bought software by the sip. It buys per bed, per facility, per assessment — units that map to how the business actually earns money. And what she’s really buying from an AI company isn’t software at all. It’s work. A completed, accurate, defensible MDS assessment — the artifact that sets the per diem, survives the audit, and lets a nurse go back to nursing. That’s the unit that matters, so that’s the unit to price. What we’re repricing is a labor budget, not an IT budget — and labor has never been billed by the token.

So there are two pricing models on the table, and they reveal everything about the company behind them.

The metered model says: pay for our AI’s consumption, whatever it turns out to be. If the model gets expensive, that’s your problem. If our prompts are bloated, your problem. If we routed a simple task to a frontier model out of laziness…still your problem.

The outcome model says: pay for completed work. We’ll carry the variance and token cost.

Here’s what I’ve come to believe running an AI company in this space: absorbing token risk isn’t generosity. It’s the forcing function that produces great engineering. The moment I eat the tokens, I am ruthlessly incentivized to route small tasks to small models, cache aggressively, kill wasteful agent loops, and measure cost-per-completed-assessment like the COGS line it is. If the customer eats the tokens, my inefficiency becomes her invoice, and I never have to get better.

Your pricing page is your engineering philosophy on public display.

That’s why at Carebility we price completed work, full stop. Not because it’s easy, but because it’s hard. Because it makes us carry exactly the risk we’re best positioned to engineer away. That’s what a partner does. A vendor sends you a meter reading as an invoice.

How to tell true AI systems engineering from AI slop

This isn’t a takedown of anyone, because the honest truth is that demos can’t tell you who’s serious anymore, but some simple questions can. If you run facilities and you’re evaluating AI, these are the ones I’d ask:
  • Security architecture. Where exactly does PHI travel? Is there zero data retention at the model layer, and does a BAA cover the inference path? Is the AI inference and patient data residence always in the US? Ask for the data-flow diagram. Serious teams will have one.
  • Compliance reality. Don’t accept the badge on the website; ask for the SOC 2 report and read the period and scope. Type II attests to months of real operation, not intentions. Then ask whether an independent firm has pen-tested the product. Ask to see the latest pen test.
  • Evaluation framework. How do you test the AI’s judgment, continuously, before and after every model change? What happened the last time your model provider shipped an update?
  • Model flexibility. Are you welded to a single model, or can you orchestrate across several (with a fallback) when pricing, performance, or availability shifts?
  • Auditability. Show me the log. Can you replay exactly what the AI did on a given assessment, and why, in a form I could hand a surveyor?
  • Human oversight. Where does the licensed clinician sit in the loop? Can they see, edit, and override, and does the record show who owned the final call?
  • Governance. What is the agent not allowed to do? Who decides that, and where’s the kill switch?
  • Healthcare depth. Who on your team has actually held the job this product touches? Can they argue coding logic from the RAI manual without a search bar?
  • Operational reliability. What happens at 7 a.m. when the model API is down? What’s your uptime history, your incident process, your status page?
  • Pricing philosophy. Do I pay for consumption or for completion? When your model costs change, whose problem is that?
  • Outcome ownership. What do you actually stand behind — and what happens when the system gets one wrong?
Most of these questions don’t require a deep technical background. All of them are hard to fake. Teams doing real AI systems engineering will light up when you ask, because somebody finally asked. The others will steer you back to the slick demo.

Software is now abundant; trust is even more scarce

Ten years ago, great software changed industries. I’m lucky I got to live it. The difficulty of building was the moat, and the companies that crossed it earned a decade of loyalty on the other side.

That era is over, and I’m glad it is. Software is becoming abundant. Anyone can generate it, and the tools get better every quarter.

But trust didn’t get easier to generate. Trust got scarcer, precisely because everything around it got faster, shinier, and harder to evaluate from the outside.

So here’s my bet on the next decade. The companies that win in healthcare AI won’t think of themselves as AI vendors, shipping software and a meter. They’ll act as AI partners. They’ll know the customer’s workflow down to Section GG ADL coding. They’ll absorb complexity instead of exposing it (the models, the tokens, the drift, the orchestration), all disappearing behind one simple promise. They’ll own risk instead of transferring it. And they’ll do the least glamorous thing in technology: quietly, securely, reliably produce outcomes, month after month, until nobody remembers what the work felt like before.

In the last era, we sold software and the customer did the work. In this one, the winners will do the work and sell the outcome.

Not tokens. Not credits. Not prompts. Not licenses.

Finished work you can trust.


“In the last era, we sold software and the customer did the work. In this one, the winners will do the work and sell the outcome.”