Medical VLM reliability evaluation
An evaluation of Qwen2.5-VL-7B and Qwen3-VL-8B on a balanced set of 200 chest X-rays, zero-shot and fine-tuned with LoRA, scoring each diagnosis together with its stated confidence and explanation.
Accuracy and F1 hid the real problem, so we added metrics for overconfident errors, fluency and faithfulness. The strongest model still missed over half of pneumonia cases, and its wrong answers were its most fluent: 0.855 fluency on errors against 0.655 accuracy.
- Role
- Co-author: evaluation, metrics and fine-tuning
- Stack
- PyTorch · Hugging Face · Qwen-VL · LoRA · 4-bit NF4
- Result
- 76.8% of Qwen3-VL’s errors made with high confidence