ML Engineer: "You've been lied to about open source"
Listen to episode
About this episode
Download our free AI Implementation guide non-technical leaders: https://www.ninetwothree.co/guides/ai-playbook
___
Self-hosting an LLM sounds like a way to stop paying for tokens
forever. The token bill turns out to be a small part of the cost.
Nay Cook-Nelson talks with Vitalijus, a machine learning engineer
with a PhD in natural language processing and 7+ years in data
science. He has deployed self-hosted models for tens of thousands
of businesses, starting years before GPT existed, back when
training a model from scratch on your own data was the only
option in machine learning.
⚡ IN THIS EPISODE
▸ What "self-hosted" means in 2026, and how it differs
from a serverless API call
▸ The 15 things you inherit the moment you leave a
third-party provider
▸ Why an H100-class GPU at $25K to $40K only gets you
to your first few concurrent users
▸ How a model priced 50% lower per token can cost more
in production
▸ Hybrid routing, and how to evaluate before you swap
models
▸ The team you need: ML engineer plus DevOps or MLOps,
on duty for the life of the project
▸ Where self-hosting is the only option: sovereignty,
edge, IoT, and zero-connectivity environments
▸ The token pricing trap, and what model deprecation
does to a product built on prompts
🕒 CHAPTERS
00:00 What you get from this episode
00:53 Seven years in data science, a PhD in NLP, and
self-hosted deployments before GPT
01:58 Self-hosting vs. serverless APIs
03:37 Why every ML model used to be trained from
scratch, and what BERT changed in 2018
06:13 Model distillation, and how DeepSeek got
competitive at a fraction of training cost
07:50 What you inherit: serving, scaling, load
balancing, autoscaling, failover, monitoring,
batching, caching, quantization, GPU
utilization, rollbacks, upgrades, guardrails
10:35 A 120B open-weight model needs an H100-class
GPU at $25K to $40K. Then a fifth user shows up
12:45 Verbose models eat your margin. Prompt caching
cuts cost up to 90%, and you build it yourself
13:48 Hybrid routing by use case instead of one
provider for everything
14:53 Three ways to evaluate an LLM system: humans in
the loop, performance metrics, models as judges
17:10 Zero-retention agreements, and when proprietary
or medical data rules out an API
19:07 Who you need to hire, and why it is a department
rather than two people
20:38 A healthcare client self-hosted Langfuse,
provisioned 7 to 8 resources for one dashboard,
then moved to a paid zero-retention tier
23:26 When self-hosting is the only option: 100%
uptime, aircraft, IoT, EV charging in garages
with no signal
24:50 US-hosted open-weight inference, and how to get
cheaper tokens without buying GPUs
26:17 Infrastructure carries about 95% of this decision
27:14 The token pricing trap. Under 5% separates the
top model from the tenth, per Stanford's AI Index
30:27 A version bump does not mean a better model for
your workload
📌 THE TAKEAWAY
Open-weight models are good enough for most workloads, so
picking one is the straightforward part. Infrastructure,
evaluation, and staffing decide whether the move pays off.
Vitalijus's rule: know your load, know your user count, build
a baseline with real evaluation, then experiment. Swapping
models without that baseline will cost you.
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity