Acting Assistant Inspector General Julie Kroviak, MD,

WASHINGTON, DC — While VA’s investment in artificial intelligence is growing year over year, a preliminary report from the department’s Office of the Inspector General raises concerns about the agency’s ability to ensure patient safety. Specifically, the OIG is concerned with the department’s use of generative AI in clinical care and documentation and the lack of a formal mechanism to identify, track and resolve risks to patients.

This past Oct. 14, OIG began a national review of VHA’s use of generative AI. What investigators discovered was concerning enough for Acting Assistant Inspector General Julie Kroviak, MD, to release a preliminary advisory to the undersecretary for health in advance of the complete report.

VHA currently authorizes two general purpose generative AI chat tools for use with patient health information: VA GPT and Microsoft 365 Copilot Chat. Clinicians provide patient information along with a prompt and output from the AI chat tool can be used for medical decision-making and eventually input to the patient record.

“However, generative AI can produce inaccurate outputs, including omissions, which may affect diagnosis and treatment decisions,” Kroviak warned.

VA itself has also recognized the dangers posed by AI in clinical decision-making. Guidance released by VA’s National AI Institute (NAII) and the department’s chief technology officer states that “this technology introduces new risks and unknown consequences that can have a significantly negative impact on the privacy and safety of veterans.” Major risks recognized by VA include misinformation, bias and discrimination and threats to data privacy and security.

According to VHA policy, the Office of Quality Management and the National Center for Patient Safety are meant to establish and provide operational oversight of VHA patient safety programs. However, OIG investigators found that VHA was not coordinating with that office in the fielding of AI chat tools for clinical use. Instead, VHA’s AI efforts were driven by an informal collaboration between the acting director of NAII and the Chief AI Office at VA’s Office of Information Technology.

“The OIG is concerned about VHA’s ability to promote and safeguard patient safety without a standardized process for managing AI-related risks,” Kroviak said. “Moreover, not having a process precludes a feedback loop and a means to detect patterns that could improve the safety and quality of AI chat tools used in clinical settings.”

VA Inspector General Cheryl Mason has since spoken publicly on the issue, noting that the risks of AI use in a clinical setting have been well documented.

“The concern is that AI hallucinates. There’s research done on this,” Mason said during a recent appearance on the Federal Drive podcast. AI hallucinations refer to when an AI model generates factually inaccurate or entirely fabricated information.

“There’s been many, many legal cases on this,” Mason said. “When you use an AI chat in clinical documentation, if the process is hallucinating, that can have an impact on patient diagnosis and management.”

In response to the preliminary report, VA spokespersons have stressed that clinical decisions still rest in the hands of providers, not AI, and that best practices require physicians to proof the work being generated by chat tools.

Despite policy that requires human proofing of AI-generated data, studies have found that errors still slip into the patient record. A report released by Stanford University early this year examined the state of clinical AI and found that clinicians exposed to flawed AI recommendations experience a significant downgrade in diagnostic accuracy.

“Automation bias poses significant patient safety risks that require robust validation frameworks and interface safeguards that actively prevents routine deference to [large language model AI] recommendations,” the report stated.

Meanwhile, VA’s inventory of AI use cases continues to expand, growing from 128 in 2023 to 227 in 2024 and 367 in 2025. Of those 367, 253 sit in VHA. Of those, 112 are fully deployed, 11 are in the pilot stages, 75 are in predeployment, and 55 have been retired. The vast majority of the 123 deployed use cases touch on patient health, ranging from extracting patient health information from unstructured clinical notes to using a patient’s health record to predict their risk for suicide.

As for VA GPT, according to VA the system has 95,000 users onboarded, with 70% reporting improved job satisfaction and an average savings of 2 to 3 hours per week.