Women in AI Research (WiAIR)
Women in AI Research (WiAIR)

Is an Image Worth a Thousand Words? Hidden Failures of Multimodal Metrics, with Dr. Elisa Kreiss

19 August 2026 1:17:30 WiAIR

Listen to episode

About this episode

An image is not worth a thousand words - it's worth an indefinite number of them. So why do the metrics we use to evaluate AI-generated image descriptions still assume there's one correct answer?


In this episode of Women in AI Research, I talk with Elisa Kreiss (Assistant Professor of Communication at UCLA, director of the Coalas Lab) about what happens when you actually test the metrics the field relies on, and why CLIPScore, one of the most widely used measures for scoring image descriptions, stops correlating with human judgment the moment you introduce context. We also get into why longer descriptions aren't necessarily more informative, what happens when you just ask a model to "be concise," and whether AI models trip over charts and graphs the same way humans do.


Elisa's research sits at the intersection of linguistics, accessibility, and multimodal AI, and this conversation covers the full arc of her work, from the theoretical question of why humans never describe images the same way twice, to the practical question of what that means for building systems that actually work for blind and low-vision users.In this episode:

  • Why "context matters" is more radical than it sounds for image description evaluation
  • The hidden reason CLIPScore breaks down once context enters the picture
  • Why length is a bad proxy for information density - and what to use instead
  • What happens when you prompt a model to just "be concise"
  • Why charts and photos need completely different evaluation approaches
  • Whether AI models make the same mistakes as humans when reading data visualizations
  • What NeurIPS's Top Reviewer Award taught her about writing a genuinely useful peer review

Resources & Links:

  • Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation Metrics
  • When More Words Say Less: Decoupling Length and Specificity in Image Description Evaluation
  • CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models

🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.


Follow WiAIR at:

  • ⁠⁠LinkedIn⁠⁠
  • ⁠⁠Bluesky⁠⁠
  • ⁠⁠X (Twitter)⁠⁠
  • ⁠⁠⁠WiAIR website⁠

Want to find AI jobs?

Join thousands of AI professionals finding their next opportunity

We respect your inbox. Unsubscribe at any time.

© 2026 Women in AI Research (WiAIR). All rights reserved.

Common Questions

Frequently asked questions

Quick answers about how DevFound's AI matching, resumes, and referrals work.

DevFound's AI Copilot ingests your profile, goals, and live job data to deliver curated matches in seconds. Every match includes a resume variant, suggested referrals, and interview prep so you can act immediately. The more feedback you provide, the sharper the Copilot becomes.

AI-led job searches shrink the hours spent sifting through boards and formatting resumes. DevFound pairs automation with your personal outreach, so you reserve energy for interviews and negotiation. Traditional networking still matters, but AI gives you a lift before you even send a message.

Modern AI roles expect comfort with production-grade code, data fluency, and practical ML tooling. The strongest candidates pair deep technical chops with storytelling—translating model impact to product, GTM, and exec partners. Continuous learning keeps you ahead as stacks evolve.

DevFound rewards active seekers. Keep your profile fresh, respond to match quality prompts, and enable alerts so you never miss a role. The AI prioritizes companies and teams that align with your feedback, accelerating both introductions and interview invites.

High-density tech hubs continue to host the deepest AI talent pools, yet distributed teams are catching up fast. Use DevFound filters to hone in on onsite, hybrid, or fully remote roles and watch openings expand across time zones.

DevFound aggregates thousands of remote AI openings and flags the nuances—core hours, async culture, and visa needs—up front. The Copilot also recommends how to position your distributed work experience so hiring managers know you can thrive on a remote team.