RADAR, a generalist AI trained on over 400,000 abdominal CTs, achieves expert-level diagnosis across 18 structures and 146 findings, with high accuracy and robustness.
- 1 on Bluesky
- 1 opens in JClub
Moving in Nature, Bioinformatics, Cancer Discovery, Cell, JAMA, Nature Medicine, Science.
Rebuilt
RADAR, a generalist AI trained on over 400,000 abdominal CTs, achieves expert-level diagnosis across 18 structures and 146 findings, with high accuracy and robustness.
AI model Oncoformer accurately predicts cancer development, diagnoses tumors, and stages cancer using health records and imaging data across diverse populations.
MIRA, an autonomous AI agent in a sandboxed EHR, surpasses physicians in diagnostic accuracy and makes safe, guideline-concordant clinical decisions across multiple cases.
Large language models like GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 outperform specialized clinical AI tools across multiple medical benchmarks, including real-world queries.
Agentomics autonomously develops state-of-the-art biomedical machine learning models, outperforming existing agentic systems and matching human expertise on 11 of 20 datasets.
A new LLM-based system, AMIE, matches primary care physicians in disease management reasoning and surpasses them in treatment precision and guideline adherence in multi-visit scenarios.
Reliability failures in AI for healthcare include erroneous outputs, population disparities, and performance decline, threatening safe deployment across 20+ failure modes.
Newest first · last 60 days
Synthetic data generation could revolutionize gastrointestinal medicine by enabling more accurate AI-driven diagnosis and treatment with privacy-preserving data.
The Multimodal Anonymizer achieves near-complete deidentification of multimodal hospital data with over 98% sensitivity and preserves critical clinical information at rates above 99%.
A ReAct agentic AI system achieves 93.4% accuracy in natural language querying and analysis of TCGA clinical data, surpassing rule-based and LLM baselines.
Physicians using conversational AI in Latin America achieved significantly higher validity scores (median 2.83 vs. 2.46) than usual search methods in clinical case responses.
42.7% of surveyed obstetricians and gynecologists in France reported using AI in routine practice, mainly for administrative and informational tasks.
General-purpose multimodal large language models in histopathology exhibit errors and hallucinations in over 80% of outputs, impacting diagnostic safety and reliability.
Algorithm D achieved the highest AUROC of 0.849 and sensitivity across all AI-positive rates in a large US screening cohort, outperforming other commercial and open-source models.
RoB2-Evaluator, a rule-constrained GPT-5.2 system, achieved 90% agreement with human reviewers and 94.3% stability in Cochrane Risk of Bias 2 assessments across 10 RCTs.
Implementing EHR-based risk models with imaging and blood tests detects pancreatic cancer early in 8.2% of high-risk adults aged 50-84.
Large language models can serve as human proxies in various scientific and applied contexts, each role requiring distinct validity criteria.
AI is revolutionizing biomedical research by enabling integration of diverse biological data and addressing previously intractable questions.
RADAR, a generalist AI trained on over 400,000 abdominal CTs, achieves expert-level diagnosis across 18 structures and 146 findings, with high accuracy and robustness.
Reliability failures in AI for healthcare include erroneous outputs, population disparities, and performance decline, threatening safe deployment across 20+ failure modes.
Agentic AI systems like Co-Scientist, Robin, and Biomni can autonomously generate hypotheses, design experiments, analyze data, and refine reasoning, advancing scientific discovery.
Single-barrier reforms in algorithmic recourse systems achieve less than 0.02% improvement, with cross-layer interactions accounting for 87.6% of potential gains in digital health contexts.
Multi-billion parameter AI models still struggle with surgical tool detection in neurosurgery, showing limited improvements despite extensive scaling efforts.
Journals this month
Moving areas, week to 3 Oct 2026