Return On Intelligence: Implement AI Without Wasting Time & Money
Return On Intelligence: Implement AI Without Wasting Time & Money

ML Engineer: "You've been lied to about open source"

27 August 2026 27:04 NineTwoThree AI Studio

Listen to episode

About this episode

Download our free AI Implementation guide non-technical leaders: https://www.ninetwothree.co/guides/ai-playbook
___
Self-hosting an LLM sounds like a way to stop paying for tokens
forever. The token bill turns out to be a small part of the cost.

Nay Cook-Nelson talks with Vitalijus, a machine learning engineer
with a PhD in natural language processing and 7+ years in data
science. He has deployed self-hosted models for tens of thousands
of businesses, starting years before GPT existed, back when
training a model from scratch on your own data was the only
option in machine learning.

⚡ IN THIS EPISODE

▸ What "self-hosted" means in 2026, and how it differs
from a serverless API call
▸ The 15 things you inherit the moment you leave a
third-party provider
▸ Why an H100-class GPU at $25K to $40K only gets you
to your first few concurrent users
▸ How a model priced 50% lower per token can cost more
in production
▸ Hybrid routing, and how to evaluate before you swap
models
▸ The team you need: ML engineer plus DevOps or MLOps,
on duty for the life of the project
▸ Where self-hosting is the only option: sovereignty,
edge, IoT, and zero-connectivity environments
▸ The token pricing trap, and what model deprecation
does to a product built on prompts

🕒 CHAPTERS

00:00 What you get from this episode
00:53 Seven years in data science, a PhD in NLP, and
self-hosted deployments before GPT
01:58 Self-hosting vs. serverless APIs
03:37 Why every ML model used to be trained from
scratch, and what BERT changed in 2018
06:13 Model distillation, and how DeepSeek got
competitive at a fraction of training cost
07:50 What you inherit: serving, scaling, load
balancing, autoscaling, failover, monitoring,
batching, caching, quantization, GPU
utilization, rollbacks, upgrades, guardrails
10:35 A 120B open-weight model needs an H100-class
GPU at $25K to $40K. Then a fifth user shows up
12:45 Verbose models eat your margin. Prompt caching
cuts cost up to 90%, and you build it yourself
13:48 Hybrid routing by use case instead of one
provider for everything
14:53 Three ways to evaluate an LLM system: humans in
the loop, performance metrics, models as judges
17:10 Zero-retention agreements, and when proprietary
or medical data rules out an API
19:07 Who you need to hire, and why it is a department
rather than two people
20:38 A healthcare client self-hosted Langfuse,
provisioned 7 to 8 resources for one dashboard,
then moved to a paid zero-retention tier
23:26 When self-hosting is the only option: 100%
uptime, aircraft, IoT, EV charging in garages
with no signal
24:50 US-hosted open-weight inference, and how to get
cheaper tokens without buying GPUs
26:17 Infrastructure carries about 95% of this decision
27:14 The token pricing trap. Under 5% separates the
top model from the tenth, per Stanford's AI Index
30:27 A version bump does not mean a better model for
your workload

📌 THE TAKEAWAY

Open-weight models are good enough for most workloads, so
picking one is the straightforward part. Infrastructure,
evaluation, and staffing decide whether the move pays off.
Vitalijus's rule: know your load, know your user count, build
a baseline with real evaluation, then experiment. Swapping
models without that baseline will cost you.

Want to find AI jobs?

Join thousands of AI professionals finding their next opportunity

We respect your inbox. Unsubscribe at any time.

© 2026 Return On Intelligence: Implement AI Without Wasting Time & Money. All rights reserved.

Common Questions

Frequently asked questions

Quick answers about how DevFound's AI matching, resumes, and referrals work.

DevFound's AI Copilot ingests your profile, goals, and live job data to deliver curated matches in seconds. Every match includes a resume variant, suggested referrals, and interview prep so you can act immediately. The more feedback you provide, the sharper the Copilot becomes.

AI-led job searches shrink the hours spent sifting through boards and formatting resumes. DevFound pairs automation with your personal outreach, so you reserve energy for interviews and negotiation. Traditional networking still matters, but AI gives you a lift before you even send a message.

Modern AI roles expect comfort with production-grade code, data fluency, and practical ML tooling. The strongest candidates pair deep technical chops with storytelling—translating model impact to product, GTM, and exec partners. Continuous learning keeps you ahead as stacks evolve.

DevFound rewards active seekers. Keep your profile fresh, respond to match quality prompts, and enable alerts so you never miss a role. The AI prioritizes companies and teams that align with your feedback, accelerating both introductions and interview invites.

High-density tech hubs continue to host the deepest AI talent pools, yet distributed teams are catching up fast. Use DevFound filters to hone in on onsite, hybrid, or fully remote roles and watch openings expand across time zones.

DevFound aggregates thousands of remote AI openings and flags the nuances—core hours, async culture, and visa needs—up front. The Copilot also recommends how to position your distributed work experience so hiring managers know you can thrive on a remote team.