Skip to content

Visual language agent analyzes abdominal ultrasound waves and predicts surgery, with three-hospital tests putting diagnostic accuracy ahead of multi-task model benchmarks

Share
Visual language agent analyzes abdominal ultrasound waves and predicts surgery, with three-hospital tests putting diagnostic accuracy ahead of multi-task model benchmarks

Listen to this article

Read by Anchor

The journal “Nature Communications” published a research study led by a multi-center medical team from the schools of medicine at Johns Hopkins, Stanford, Colorado, Central South, Taipei, and Brown, in which they developed a multi-task visual-language agent to automate interpretation of ultrasound images of the right upper quadrant of the abdomen and to provide clinical decision support. The model aims to address the chronic challenge in ultrasound imaging, a core component in emergencies involving liver and biliary diseases, but it remains subject to variation in examiner skill and time pressure in critical-care departments.

The team trained the model on a large database from a single academic medical center that included 9,189 cases comprising 594,099 ultrasound images. To test the model’s ability to generalize clinically without relying solely on the original environment’s data, the system was subjected to an independent external evaluation on two datasets from separate institutions: the first from the University of Colorado, containing 1,704 cases with 90,616 images, and the second from Stanford University, containing 108 cases with 3,240 images.

The agent links multi-frame visual understanding with report-based linguistic reasoning to perform three concurrent diagnostic tasks.These tasks include classifying 16 different clinical appearances, generating diagnostic radiology reports, and providing surgical decision support by predicting the need for cholecystectomy based on combined imaging findings and clinical data.

The evaluation results showed the model achieved a macro-average area under the receiver operating characteristic curve of 0.820 on the internal dataset, 0.794 on the University of Colorado data, and 0.775 on the Stanford data, while the internal micro-average was 0.865 on the clinical outcome classification task. In double-blind assessments of the quality of generated reports after review and revision, radiologists had difficulty distinguishing reports drafted with the model’s assistance from those written entirely by human consultants, and the system succeeded in predicting cholecystectomy decisions with accuracy that supports surgical case triage.

For the health sector in the Gulf, Egypt, and the Levant, this development touches the core of operational efficiency in emergency departments and primary-care centers. Hospitals in the region suffer from night-time triage pressure and a shortage of radiology consultants available to read immediate ultrasound examinations for acute abdominal pain. Introducing agent models that combine frame reading, draft report writing, and surgical pathway prediction shifts the physician’s role from full manual documentation to rapid clinical verification. The study data show that the drop in generalization accuracy across independent institutions, from 0.820 to 0.775, requires regional hospitals to establish local data calibration frameworks and not rely on off-the-shelf model outputs without direct human supervision.

Don't miss the next story

Subscribe for updates