We still don't understand LLMs today.
Listen to episode
About this episode
How do LLMs actually work—and why do they remain black boxes?
In this episode of Merge, Shriyash “Yash” Upadhyay, co-founder of Martian, joins CodeRabbit to explore LLM interpretability, AI research, code review benchmarks, and the search for the “steam engine of AI.”
We discuss:
- Why we still don’t fully understand how LLMs work
- How Martian is researching machine intelligence
- Why code review is a crucial test of AI code generation
- How precision and recall shape AI code-review performance
- Why static AI benchmarks eventually become unreliable
- How real-world developer behavior can improve evaluations
- What more reliable and interpretable AI could unlock
- The tools and programming languages Yash uses in his own work
- How aspiring researchers can get started in machine learning
Today’s language models can generate code, solve complex problems, and power increasingly autonomous systems. But without understanding why they succeed, when they will fail, and how their internal mechanisms produce their outputs, building AI systems we can truly trust remains difficult.
Could interpretability provide the scientific foundation for the next generation of AI?
Learn more about Martian: https://withmartian.com/
Learn more about CodeRabbit: https://coderabbit.ai/
Subscribe for more conversations about AI, software engineering, code review, and the future of developer tools.
#LLM #AIInterpretability #ArtificialIntelligence
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity