The Science of Confabulation Elimination: Toward Hallucination-Free AI-Generated Clinical Notes

The Science of Confabulation Elimination

Toward Hallucination-Free AI-Generated Clinical Notes

Introduction

Ambient AI has rapidly become ubiquitous in healthcare workflows. The rapid adoption of this technology speaks to the tangible benefits experienced by clinicians. In this paper, we focus on a topic that is top-of-mind for adoption of any AI-powered product: “hallucinations.” There is now an opportunity to reduce errors in clinical documentation, even to the point of near-elimination.

Not all unsupported claims (colloquially called “hallucinations") are created equal.

We first discuss how we characterize and categorize the degree to which claims are supported by context, and the severity of unsupported claims. Many claims sit at the boundary of “reasonable inferences,” since clinical conversations can leave room for interpretation. We then reveal the inner workings of our note-generation pipeline to discuss how we approach the challenge of ensuring that every claim is faithful to clinical context. We focus in particular on the guardrails we have developed, supported by over one thousand hours of human validation, to automatically revise AI-generated “first drafts” in real time, preemptively detecting and correcting unsupported claims before a clinician reviews their draft documentation.

Assessing Factuality of AI-generated Documentation

When considering unsupported claims, it is essential to provide clear definitions to distinguish among distinct phenomena. Discussing “hallucinations” in black-and-white terms can be misleading and uninformative, especially when conducting fine-grained evaluations of systems, and developing auxiliary AI models that can precisely detect and correct unsupported claims.

For example, text models often draw inferences that are likely but not explicitly stated from context: Consider an encounter where the conversation included a discussion of continuing a patient's dose of metformin and monitoring HbA1c levels; the generated note then references “diabetes,” even though the word “diabetes” is never mentioned in the conversation.

Reasonable Inference

Drawing a precise line between reasonable inference vs. unsubstantiated extrapolation can be complicated and subjective; for example, not all clinicians may agree if a particular inference is reasonable in context. Yet, some cases are clearly incorrect: Consider a claim that is factually contradicted by the conversation, with implications for care (e.g., “patient denies chest pain” when the patient did in fact mention their chest pain). These cases are unambiguously problematic and warrant more severe concern.

With these distinctions in mind, we present in this document our current internal guidelines for categorizing unsupported claims.

Support

This axis assesses whether the content is explicitly affirmed by the transcript, contradicted by the transcript, or somewhere in-between.

Severity

This axis evaluates the gravity of having included unsupported claims in the clinical note, and only applies in cases where there are such claims.

The focus of these guidelines: limitations ↗

Is the claim supported by the transcript? For each factual claim in the generated note, we assign one of the following categories to reflect the extent to which the claim is supported by the transcript of the conversation. Note that these categories are independent of clinical severity, which is assessed in the “severity” axis.

Examples of Claim Classification:

  1. Directly Supported: A statement is classified as “directly supported” if the content precisely matches the transcript with no meaningful deviations or assumptions.
  2. Circumstantially Supported, Reasonable Inference: A statement is classified as “circumstantially supported, reasonable inference” if the statement, while not explicitly stated in the transcript, can be logically and reasonably inferred based on available information.
  3. Circumstantially Supported, Questionable Inference: A statement is classified as “circumstantially supported, questionable inference” if the statement can be inferred based on available information, but where there is some doubt as to whether or not the inference is correct.
  4. Unmentioned: A statement should be classified as “unmentioned” if it is not at all substantiated by the transcript. These are statements that are neither explicitly covered in the transcript nor possible to infer from context.
  5. Contradiction: A statement is classified as a “contradiction” if it directly conflicts with statements made in the transcript.

Major Severity

An unsupported claim should be classified as “major” if most clinicians would agree that, if left uncorrected, the claim would likely have a negative impact on clinical care and/or has a non-trivial chance of leading to substantial harm.

Conclusion

Our vision for ambient AI is not just about faster and easier documentation: The opportunity is clear: A review of 48 research studies that examined medical records to detect errors found that in 47 of the 48 studies, errors were present, suggesting that errors can commonly be found in medical documentation. These mistakes are not harmless; from the review: “Various studies have shown that poor documentation contributes to medical errors, malpractice claims, and even patient mortality.”

By sharing how we understand and address the problem of “hallucinations,” we hope to set a higher standard for transparency and trust. We also hope to spark a broader conversation on measuring progress towards the goal of perfect accuracy: Not all “hallucinations” are the same, and understanding the subtle-but-important differences in type and severity is crucial to our long-term aim to reduce inaccuracies in clinical documentation to near zero.

The adoption of ambient AI has the potential to nearly eliminate all mistakes in medical documentation.