AGENTS.md vs Skills: MLOps, Evals & Agent Governance with Maria Vechtomova
Listen to episode
About this episode
Agent code can look productive right up until a dependency changes, an eval misses the real failure mode, or an over-permissioned tool turns a routine task into a security incident. So what does it actually take to operate AI systems responsibly?
In this episode, Mehdi, Dumky de Wilde, and Maria Vechtomova connect MLOps and LLMOps to agent evals, MCP governance, regenerated software, security, and the engineering practices that still matter when outputs are non-deterministic.
All links and note : https://motherduck.com/podcast
Chapters:
00:00 Meet Maria Vechtomova
01:01 From MLOps to forward-deployed engineering
01:59 Principles first, Databricks second
06:18 What changes from MLOps to LLMOps
09:10 Deterministic tools for non-deterministic systems
09:50 Who maintains regenerated software?
12:38 Hiring for critical thinking with AI
17:38 MCP skills, files, extensions, and stateless servers
20:43 The missing governance layer for agent tools
24:15 Turning deployment pain into reusable practices
28:44 Testing and evaluating LLM systems
31:59 Why AGENTS.md beat skills in Vercel’s evals
40:13 When an agent accidentally hacks Hugging Face
45:00 Skills and the software supply chain
47:53 Consulting that leaves teams stronger
50:42 Wrap-up
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity