AI Discharge Summaries: What the 2026 Study Found

The 2026 Stanford study supports a narrow answer: an AI workflow can draft hospital-course text that hospitalists often use, but review still finds omissions and inaccuracies, while measured time savings may be small. It does not prove that every discharge-summary product is safe, fast, or ready for unattended use.

Read the study this way

  • The pilot involved 11 attending hospitalists at one inpatient medicine unit.
  • Physicians used AI text in 219 of 384 discharge cases, a 57% use rate.
  • Detailed feedback on 100 summaries found more omissions than hallucinations.
  • Audit-log time changed far less than clinicians felt it changed.
  • Hospital Copilot says its module drafts from the stay’s record; it does not claim the Stanford implementation.

The study was narrower than the headline

The JAMA Network Open paper evaluated MedAgentBrief, a custom agentic workflow at one Stanford Health Care inpatient medicine unit. Eleven attending hospitalists participated in a single-arm prospective quality-improvement pilot from August 1 through October 11, 2025 (Grolleau et al., 2026). There was a before-period for some measures, but no randomized control group.

MedAgentBrief did not write a discharge record from ambient room audio. It pulled free-text history-and-physical notes and daily progress notes, then generated a one-line patient summary, a narrative hospital-course overview, and a problem-based summary. The study says the workflow did not access structured laboratory results, medication lists, or flowsheet data. Physicians received source-linked output and could use, edit, or discard it (Grolleau et al., 2026).

That distinction changes the buying question. A product can be called an AI discharge-summary tool while taking very different inputs: a discharge conversation, selected notes, the whole longitudinal chart, structured medication data, or text pasted by a clinician. The study’s outcome belongs to the tested pipeline, not to the label.

The numbers, with their denominators

During 384 discharges for 331 patients, the workflow generated 1,274 daily hospital-course summaries. Physicians incorporated AI text into final discharge documentation in 219 cases, or 57% of the discharges (Grolleau et al., 2026). That is adoption evidence. It is not a 57% success rate.

Detailed feedback covered 100 summaries: 88 that were used and 12 that were not. Within those reviews, physicians reported omissions in 25, inaccuracies in 20, and hallucinations in two. The categories could overlap. Eighty-eight were rated as having no harm potential, and one was initially rated likely to cause moderate harm, although later adjudication found the disputed directive appropriate (Grolleau et al., 2026).

The denominator deserves a red circle. Feedback covered only part of the used set and a much smaller portion of the unused set. The authors acknowledged that selection could bias the observed safety profile. A buyer shouldn’t turn the reviewed-sample percentages into a universal error rate.

Common claimWhat the paper supportsWhat it does not support
”Hospitalists used it”AI text appeared in 57% of final discharge documents during the pilotEvery hospitalist will use it at the same rate
”The summaries were safe”Most of the 100 reviewed outputs were rated as having no harm potentialOutputs were error-free or safe without review
”It reduced burnout”Mean work-exhaustion score fell in 10 paired responsesThe tool caused the change in a controlled trial
”It saved time”Five of seven physicians with matched logs had median reductions up to 2.9 minutesA reliable large time saving across users
”It generated the discharge summary”It drafted the hospital-course portion from free-text notesIt reconciled structured medications, labs, or every discharge field

This evidence-to-claim table is the part worth taking into procurement. It keeps a good study from becoming a bad sales sentence.

The omission problem is the useful finding

Omissions appeared in 25 of the 100 reviewed summaries. The paper grouped them around stable chronic conditions, unresolved diagnostic uncertainty, and details that were present but not given enough weight (Grolleau et al., 2026). That failure mode makes sense for a hospital course. A polished summary can omit the thing the outpatient clinician needed most.

So review by fluency is weak. Read against sources. For a synthetic trial case, put the draft beside the admission note, active problem list, daily notes, consultant recommendations, medication record, pending studies, and follow-up plan. Mark every statement that lacks a source and every source fact that never reaches the draft.

Do the medication pass separately. The Stanford workflow did not use structured medication lists, and this post does not claim that Patient Square reconciles discharge medications. A hospital’s medication-reconciliation and order workflows remain whatever the approved local system requires.

Time felt different from time measured

The mismatch between perception and audit logs may be the paper’s best operational lesson. Among seven physicians with matched baseline data, five saw median documentation-time reductions of up to 2.9 minutes, while two saw increases of up to 1.5 minutes. The before-versus-pilot difference did not meet the paper’s statistical threshold. Yet clinicians reported perceived savings in 67 of the 100 feedback responses, including 32 that estimated more than 15 minutes saved (Grolleau et al., 2026).

Perceived effort matters. So does clock time. Measure both without pretending they are the same variable. A hospital can ask users how draining the task felt, then pull active documentation time from the record system and measure time to signed completion. The intervention may feel easier without creating capacity elsewhere in the shift.

Burnout scores fell from a mean 1.75 to 1.20 among ten paired respondents, while the study did not detect a change in cognitive-burden scores (Grolleau et al., 2026). Because this was a small single-site pilot without a concurrent control group, the clean wording is “associated with,” not “caused.”

Patient Square claims a draft from the stay record

Patient Square is an AI clinical platform. Practice Copilot brings the whole practice under one AI copilot: an ambient AI Medical Scribe that hands back a structured SOAP note, ICD-10 suggestions, and a prescription draft minutes after the visit, plus a bundled AI EHR, scheduling, and messaging as you move up the plan. Hospitals get Hospital Copilot.

The Hospital Copilot page says its AI Discharge Summary is drafted from the stay’s own record during the stay, then reviewed and signed. That establishes a broad source and review boundary. It does not identify which note types or structured fields enter the draft, which record systems connect, whether statements link to source notes, or how authentication works. None of the Stanford performance numbers can be transferred to Patient Square without a direct evaluation.

This is a better sales boundary, frankly. A hospital should bring its own source map and a set of fictional or properly authorized cases to the demo. Ask the implementation team to show what enters the module, what stays outside it, where the draft lands, and what happens when the output omits an unresolved diagnosis.

CMS says entries prepared by scribes must be authenticated by the treating practitioner (CMS MLN). An AI summary has the same practical finish line: the accountable clinician reviews and authenticates the record under hospital policy.

A twelve-case evidence packet

Don’t run a demo on one tidy three-day admission. Build twelve fictional or approved cases that include a long stay, diagnostic uncertainty, a late consultant recommendation, a pending test, a changed medication, an important stable condition, a readmission, and an ordinary short stay. Some cases should contain more than one trap.

For each output, record source-supported statements, clinically material omissions, incorrect statements, unsupported additions, medication or follow-up discrepancies, edit time, and final disposition: use, rewrite, or discard. Have hospitalists review independently before discussing the outputs as a group. The local disagreement is information too.

If your team wants to test that boundary,

Book a demo for US clinics

Prefer a separate page? Open booking in a new tab.

. A polished paragraph isn’t the acceptance criterion. Traceability is.

FAQ

Common questions

What did the 2026 Stanford AI discharge-summary study test?

It tested MedAgentBrief, a custom workflow that generated draft hospital-course summaries from free-text notes at one Stanford inpatient medicine unit. Eleven attending hospitalists participated. It was a prospective, single-arm quality-improvement pilot, not a randomized comparison and not an evaluation of every commercial discharge-summary product.

Did AI save hospitalists time on discharge summaries?

Not reliably in the study's audit logs. Five of seven physicians with matched data had median reductions of up to 2.9 minutes, while two had increases of up to 1.5 minutes. The before-versus-pilot difference did not meet the paper's statistical threshold. Perceived savings were much larger.

Were the AI-generated hospital-course summaries error-free?

No. Among the 100 summaries with detailed feedback, physicians reported omissions in 25, inaccuracies in 20, and hallucinations in two. Those groups can overlap. The sample of reviewed summaries was not a random sample of all generated outputs, so the percentages need that denominator and limitation.

Can an AI-generated discharge summary be signed without review?

No. The Stanford workflow gave physicians source-linked drafts they could use, edit, or discard, and its safety outcome explicitly asked about the unedited text. A hospital should require review against source notes, medications, results, pending work, follow-up, and its own authentication policy before the record is signed.

Does Patient Square claim to reproduce the Stanford workflow?

No. Hospital Copilot says its module drafts from the stay's own record during the stay for review and signature. It does not claim Stanford's three-stage method, specific source-note retrieval, FHIR connection, inline citations, Epic workflow, or study outcomes. A procurement demo should verify each required input and control.

Sources

  1. Grolleau F, Liang AS, Keyes T, et al. Physician-Reported Safety Outcomes of AI-Generated Hospital Course Summaries. JAMA Network Open. 2026;9(5):e2616556.
  2. CMS MLN: Complying with Medicare Signature Requirements, including entries prepared by scribes (April 2024).
  3. Patient Square: US Hospital Copilot product scope (reviewed September 2026).
  4. Patient Square: US security, BAA, encryption, audio handling, and SOC 2 status (reviewed September 2026).