Is an Image Worth a Thousand Words? Hidden Failures of Multimodal Metrics, with Dr. Elisa Kreiss
Listen to episode
About this episode
An image is not worth a thousand words - it's worth an indefinite number of them. So why do the metrics we use to evaluate AI-generated image descriptions still assume there's one correct answer?
In this episode of Women in AI Research, I talk with Elisa Kreiss (Assistant Professor of Communication at UCLA, director of the Coalas Lab) about what happens when you actually test the metrics the field relies on, and why CLIPScore, one of the most widely used measures for scoring image descriptions, stops correlating with human judgment the moment you introduce context. We also get into why longer descriptions aren't necessarily more informative, what happens when you just ask a model to "be concise," and whether AI models trip over charts and graphs the same way humans do.
Elisa's research sits at the intersection of linguistics, accessibility, and multimodal AI, and this conversation covers the full arc of her work, from the theoretical question of why humans never describe images the same way twice, to the practical question of what that means for building systems that actually work for blind and low-vision users.In this episode:
- Why "context matters" is more radical than it sounds for image description evaluation
- The hidden reason CLIPScore breaks down once context enters the picture
- Why length is a bad proxy for information density - and what to use instead
- What happens when you prompt a model to just "be concise"
- Why charts and photos need completely different evaluation approaches
- Whether AI models make the same mistakes as humans when reading data visualizations
- What NeurIPS's Top Reviewer Award taught her about writing a genuinely useful peer review
Resources & Links:
- Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation Metrics
- When More Words Say Less: Decoupling Length and Specificity in Image Description Evaluation
- CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.
Follow WiAIR at:
- LinkedIn
- Bluesky
- X (Twitter)
- WiAIR website
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity