Skip to main content

AGENTS.md vs Skills: MLOps, Evals & Agent Governance ft. Maria Vechtomova

August 24, 202651:05

Hosted by Mehdi Ouazza, Dumky de Wilde · With Maria Vechtomova (Co-founder, Cauchy)

Agent code can look productive right up until a dependency changes, an eval misses the real failure mode, or an over-permissioned tool turns a routine task into a security incident. Maria Vechtomova joins Mehdi Ouazza and Dumky de Wilde to connect MLOps and LLMOps with agent evals, MCP governance, regenerated software, and AI security.

Maria Vechtomova on LinkedInCauchy

Chapters
  • 0:00Meet Maria Vechtomova
  • 1:01From MLOps to forward-deployed engineering
  • 1:59Principles first, Databricks second
  • 6:18What changes from MLOps to LLMOps
  • 9:10Deterministic tools for non-deterministic systems
  • 9:50Who maintains regenerated software?
  • 12:38Hiring for critical thinking with AI
  • 17:38MCP skills, files, extensions, and stateless servers
  • 20:43The missing governance layer for agent tools
  • 24:15Turning deployment pain into reusable practices
  • 28:44Testing and evaluating LLM systems
  • 31:59Why AGENTS.md beat skills in Vercel's evals
  • 40:13When an agent accidentally hacks Hugging Face
  • 45:00Skills and the software supply chain
  • 47:53Consulting that leaves teams stronger
  • 50:42Wrap-up

$sudounlock--all-shows

Links and show notes on every episode.