LLMs Have Shown Mixed Results in Medical Meta-Analyses

Researchers evaluated four major language models on their ability to autonomously perform ophthalmic statistical reviews.

Updated on Sept. 25, 2026 in Artificial Intelligence

Bold flat-color editorial illustration showing a stylized ophthalmic lens over an abstract data grid, representing automated medical research analysis.
A new research study suggests that major large language models demonstrate inconsistent performance when tasked with autonomously executing complex ophthalmic statistical meta-analyses. AI Illustration. Upload story photo >

Live Poll

Do you trust artificial intelligence to perform complex scientific data analysis without human oversight?

A research study published on September 25, 2026, tested the ability of four large language models—GPT 5.6, GPT 5.3, Claude Sonnet 5 High Effort, and Gemini 3.1 Pro—to perform automated meta-analyses. The findings indicate varied success in extracting complex statistical data from medical literature.

Why it matters

The evaluation assessed whether AI can reliably automate the labor-intensive process of synthesizing evidence from clinical literature. Success in this field could accelerate systematic research, though the current discrepancies in statistical output highlight a need for validation.

GPT 5.6 demonstrated the highest performance, reproducing 85.0% of 2x2 tables and 80.0% of I2 values (a measure of study heterogeneity). Claude Sonnet 5 High Effort also achieved 85.0% cell accuracy, while Gemini 3.1 Pro matched 40.0% of pooled odds ratios.

The players

GPT 5.6

A high-capability large language model developed by OpenAI that focuses on advanced reasoning and complex data extraction.

Claude Sonnet 5 High Effort

An optimized language model from Anthropic designed for high-precision tasks and document analysis.

Gemini 3.1 Pro

A multimodal large language model from Google optimized for integration with complex informational datasets.

GPT 5.3

An iteration of the OpenAI GPT series used here as a comparative benchmark for data extraction performance.

The details

The evaluation required models to process primary study PDFs and follow prompts to perform specific data extraction and statistical calculations. Researchers compared the outputs against established statistical reference standards to measure accuracy in reconstructing 2x2 tables—grids used to display frequency counts for binary variables—and pooled odds ratios. Models were tested across five distinct meta-analyses sourced from ophthalmic literature to determine their viability as research tools.

Timeline

  1. The article detailing the model evaluation was published on September 25, 2026.

The Tech Race

This study situates modern LLMs against the rigorous Cochrane Handbook for Systematic Reviews of Interventions, the gold standard for evidence synthesis. It marks an early effort to determine if AI can meet the professional benchmarks required for peer-reviewed medical consensus.

This research suggests that researchers currently cannot rely on AI for autonomous statistical synthesis without manual verification. Users should treat AI-generated statistical output as a draft that requires oversight against primary source documents.

The takeaway

The study indicates that while LLMs demonstrate high accuracy in basic cell extraction, they still struggle with the complex, multi-step calculations required for authoritative meta-analysis. Future developments should track subsequent evaluations in other medical disciplines to determine if these error rates persist across different types of clinical data.

Further reading

For more on the current state of automated research tools, visit Artificial Intelligence.

Live Poll

Do you trust artificial intelligence to perform complex scientific data analysis without human oversight?

LLMs Have Shown Mixed Results in Medical Meta-Analyses | Highwise Tech