What an AI Audit Actually Finds

The audit's real output: a ranked map of what to build, what to buy, and what to skip.
When someone asks me to "add AI," they almost never have a model problem. They have a which problem is even worth it problem. An AI audit is how you find out — before you spend six months and a budget line on the wrong one.
I've spent twelve years building ML and AI in production: healthcare integration platforms, clinical tooling, edge systems, generative products. Along the way I've watched a lot of AI initiatives stall. The ones that failed almost never failed because the model wasn't good enough. They failed earlier than that, in ways an audit catches cheaply. Here's what a real audit actually surfaces.
1. Half your ideas shouldn't be built
The single most valuable output of an audit is usually a "no." Not every problem is an AI problem, and plenty of the ones that are have a $20/month tool that already solves them. A good audit ranks your list of ideas by leverage — cost of the problem versus cost of the solution — and is honest about which ones are a database query, a form, or an existing SaaS in disguise.
If an audit only ever says "yes, build everything," it isn't an audit. It's a sales pitch.
2. The problem is written wrong
Most AI briefs are written as solutions — "we want a chatbot," "we want to use LLMs." A workable brief is written as an outcome: hours reclaimed, revenue moved, risk reduced, a specific decision made faster. The audit's job is to rewrite each idea as an outcome a non-engineer would sign off on, with a way to measure it. If you can't measure success, you can't ship responsibly, and you certainly can't tell whether the thing worked.
3. The data isn't where you think it is
This is where audits find the most surprises. The model needs data it can reach — and again and again, that data is trapped in a PDF, a legacy system, one person's spreadsheet, or a vendor you don't control. Or it exists, but nobody knows how often it's wrong. The audit answers three questions bluntly: does the data exist, can the model actually get to it, and is it good enough to trust. In regulated spaces — healthcare especially — it also draws the line the model must never cross around PII/PHI.
Related Articles

Your Support Team Is Claude Code on a Timer
An L2 support engineer that's just Claude Code on a /loop, with all its state living in Slack reactions. No app, no database, no deploy — and it only pings you when it's real.

Inside a Production Voice Agent: How the Stack Actually Ships
Production voice-AI has converged on a pattern: graph-based conversations, separated decision and response prompts, synthetic-call regression testing, and per-component latency budgets. Why the stack looks the way it does — and what most teams are still missing.

Prompt Engineering Didn't Die. It Got Unrolled.
Everyone keeps announcing the death of prompt engineering. They are describing the symptom, not the shift. The loops you used to run by hand — refine, retry, verify, learn — moved out of your head and into infrastructure. Four of them, simultaneously.
Building something like this?
I help teams ship AI in production — audits, consulting, custom agents, and eval systems. Start with an AI Audit (from $5k) for an honest read on what to build.



