Skip to main contentSkip to navigation
    Provider reviewing coding accuracy from AI documentation
    Implementation & Buyer Guidance

    Why 99% Accuracy in Medical Coding Isn't the Right Metric

    AI coding vendors compete on accuracy rates, and the metric is real. What it consistently fails to measure is whether the inference that produced the code was correct before the code was assigned, and that distinction is where denial risk lives.

    Percy Bhathena, VP Product & AI
    5/22/2026
    8 min read

    Percy Bhathena is VP of Product & AI at DeliverHealth, overseeing InstaNote and InstaCode, the company's ambient documentation and AI medical coding platforms.

    Every AI coding vendor in healthcare seems to be competing on accuracy. The number varies by vendor, but the pitch remains consistent: our engine codes at x% accuracy, which is ultimately better than human coders, meaning you can reduce your coding staff and trust the system to handle the rest. 

    There’s a problem with that pitch. The accuracy number, as typically presented, doesn’t measure what buyers think it measures. The difference between what it measures and what matters in medical coding can determine whether a claim gets paid or denied. 

    Understanding that gap isn’t a reason to avoid AI coding tools. It’s a reason to evaluate them more carefully. 

    What Accuracy Actually Measures, and What It Misses 

    When an AI coding vendor reports accuracy, they’re typically measuring whether the engine produced the correct code given a specific encounter. That sounds precise. However, it leaves significant room for interpretation. 

    Accuracy can be measured at the encounter level: did the AI correctly identify what happened? It can be measured at the code category level: did the AI identify the right code family? It can be measured at the E&M level: did the AI capture the right complexity and time-based components? Each of these measurements can produce a high accuracy score while missing the specific code that determines whether a claim is approved, rejected, or flagged for audit. 

    The gap lives in inference, and inference is where medical coding gets complicated. 

    The Soccer Field Problem 

    Let’s review a scenario that illustrates the issue clearly. 

    AdobeStock_1995780338.jpeg

    A child presents a knee injury. The encounter note states the patient was hurt while playing soccer. The AI correctly identifies the injury type, the body part affected, the injury severity, and the encounter duration. It produces a code for a sports-related knee injury. So far, so good, right? 

    Then it infers. Because the encounter mentions a soccer game, it is inferred that the injury occurred on a soccer field. 

    There is a specific ICD-10 code for a soccer field injury. There is a different code for a soccer injury that did not occur on a soccer field. If the patient was actually playing soccer in the backyard, on the street, not on a designated soccer field, the claim carries the wrong code. That code, not the injury type, not the encounter length, not the severity, is what the payer adjudicates.  

    You can measure accuracy and say: it got the encounter right, then length, the severity, you can say all of those things are right, and say you got it 100%. But, if you’re being truthful, that last little bit is hard for the LLMs.

    This isn’t an edge case. It’s a structural feature of how large language models work. It applies across encounter types, specialties, and code categories wherever inference is required to bridge the gap between what was said and what must be specified. 

    Probabilistic Models in a Deterministic Rule System 

    Medical coding has approximately 70,000 rules. When you first encounter that number, it sounds like exactly the problem AI should be able to solve. Rules mean structure. Structure means logic. Logic is what computer systems are designed for. 

    The reality is more complicated than that. Code rules are deterministic. If condition X, then apply code Y. The clinical language that feeds those rules is not. Encounter notes are written by humans in natural language, with all the ambiguity and contextual inference that natural language entails. 

    Large language models are trained to predict the most probable output. When a note says “soccer game,” the most probable inference is “soccer field.” That inference is correct for the vast majority of the time. But medical coding doesn’t grade on probability. It grades on specificity. A plausible inference that produces the wrong code produces the wrong claim. 

    “Probabilistically, if you’re playing a soccer game, you’re playing it on a field. That’s how the models think. And that’s what makes this level of complexity so easy to miss when you’re looking at top-line accuracy numbers.” 

    This is why accuracy rates need to be interrogated rather than accepted. An engine can be right about almost everything and still produce code that results in a claim being denied. 

    What “Getting It Right” Actually Looks Like 

    The vendors who are honest about this problem have converged on the same architecture: autonomous coding with a human review layer. 

    Not because AI can’t code. It can accurately and quickly handle a high volume of straightforward encounters without human involvement. The efficiency gains are real. But the encounters where inference is required, where clinical language is ambiguous, where a single unspecified detail determines the correct code, those still benefit from a trained coder who understands context in ways a probabilistic mode doesn’t. 

    The value of this architecture isn’t a concession to technology’s limitations. It’s a design choice that yields better outcomes throughout the full claims cycle.  

    I think when we provide autonomous coding with a service component on it, that’s where you can provide the most value. You’re still providing efficiency, but you’re also being very truthful about what the engines can do today, and where the human matters. 

    An autonomous engine processes the bulk of claims at speed. A human coder reviews the cases where the engine’s confidence is low or where specificity matters most. The result is faster turnaround, lower denial rates, and documentation that holds up under audit scrutiny, which matters for any organization operating under payer scrutiny.  

    • Why this matters for the revenue cycle
      Denial reduction is the most universal ROI metric in AI medical coding, across segments and specialties. A coding engine that produces high accuracy rates but doesn’t account for inference errors will generate denials at a higher rate than those top-line numbers suggest. The downstream metric is what the payer does with the code, not what the engine produced. 

    Evaluating a Vendor: The Questions That Matter 

    When an AI medical coding vendor presents their accuracy numbers, these are the questions worth asking before accepting them at face value: 

    1. How is accuracy measured? At the encounter level, the code category level, or the specific code level? The answers are different, and the differences are significant. 

    1. What happens when the engine is uncertain? Does it flag cases for human review, or does it produce a code at whatever confidence level it has? 

    1. How does the system handle inference cases where the clinical note implies but doesn’t explicitly state a relevant detail? Is there a process for reviewing those cases before the claim goes out? 

    1. What is the denial rate, not just the accuracy rate? An engine can score high on accuracy and still produce claims that get denied. The payer’s decision is the downstream metric that matters. 

    1. Who owns the error when a code is wrong? The answer to these questions tells you a great deal about how confident the vendor actually is in their autonomous claims. 

    These aren’t adversarial questions. They’re the due diligence that protects your organization’s revenue and compliance position. Vendors who can answer them clearly and specifically are the ones worth moving forward with. 

    The Metric That Matters 

    The medical coding market is moving toward autonomy, and that direction is broadly correct. Autonomous AI coding produces real efficiency gains at scale; faster processing, reduced coder workload, shorter turnaround on claims. These are meaningful outcomes for health systems operating under margin pressure. 

    But autonomy without a human review layer isn’t a feature. It’s a gap in the accountability chain. The inference errors that probabilistic models produce aren’t hypothetical. They're structural characteristics of how the technology works. In a domain where a single code determines whether a claim is paid or denied, that gap has a dollar value. 

    The metric to hold vendors to isn’t accurate in isolation. It's accuracy plus denial rate plus audit exposure plus accountability, the full picture of what the engine produces under real clinical conditions. Vendors who can speak to all those measures honestly are the ones building for outcomes, not just for demos. 

    DeliverHealth’s InstaCode combines an autonomous AI coding engine with flexible human-in-the-loop deployment options, designed to deliver efficiency at scale while maintaining the code-level accuracy that determines claim outcomes. Learn more about InstaCode and contact DeliverHealth for a quick 20-minute conversation about your coding options. 

    About the author:

    Percy Bhathena leads product and AI strategy at DeliverHealth, where he oversees the development of InstaNote and InstaCode — DeliverHealth's ambient documentation and AI medical coding platforms. With a background spanning enterprise technology and healthcare AI, Percy focuses on building tools that augment clinical workflows rather than replace the human judgment at their center. He speaks and writes on the intersection of AI, clinical practice, and responsible product development.

    Tags

    AI Medical Coding
    Medical Coding
    InstaCode
    Accuracy

    Stay Updated

    Subscribe to our newsletter for the latest healthcare AI insights and company updates.