The integration of large language models (LLMs) and large multimodal models (LMMs) into medical settings is becoming increasingly prevalent, but a recent study by researchers at the University of California at Santa Cruz and Carnegie Mellon University raises serious concerns about their reliability in high-stakes, real-world scenarios. The study reveals that even advanced models, including GPT-4V and Gemini Pro, perform poorly when asked to identify conditions and positions in medical images, with accuracy dropping by an average of 42% across tested models. The researchers introduced a new dataset, ProbMed, which features 6,303 images from two widely-used biomedical datasets, and subjected seven state-of-the-art models to probing evaluation. The results are alarming, with even the most robust models experiencing a minimum drop of 10.52% in accuracy. The study highlights the urgent need for more robust evaluation methodologies to ensure the accuracy and reliability of LMMs in real-world medical applications.

Source.

TOP STORIES

Navigating AI Regulation - Balancing Safety and Innovation
The ongoing debate on AI regulation highlights the balance between safety and innovation …
Rogue AI - A Wake-Up Call for Enterprise Security
The recent breach involving rogue AI models reveals urgent security gaps in enterprise AI governance …
Time to Slow Down? Sam Altman on Pacing AI Development
Sam Altman argues for a careful approach to AI development amidst security concerns …
Claude Chats Exposed - Private Conversations Found on Google Search
Sensitive Claude chats were found publicly searchable on Google, revealing personal information …
Microsoft Launches Powerful AI Cybersecurity Tools to Combat Threats
Microsoft has launched MAI-Cyber-1-Flash and the Perception platform to enhance cybersecurity …
OpenAI's AI Model Breach Sparks Debate on Safety and Control
The breach of OpenAI’s model at Hugging Face highlights urgent concerns about AI safety and control …

latest stories