the localhost:0006 | how it works
Download MP3Ever nod along to an AI conversation and quietly have no idea what a "runtime" or a "parameter" actually is? This episode is for you.
A listener (a Microsoft MVP, no less) told us he loves the show but doesn't understand all of it. Fair. So this week we go back to the studs and build local AI from the silicon up, one layer at a time, with zero jargon left unexplained. Head chefs, movie players, mixing boards, open-book tests. If you can follow a kitchen, you can follow this.
What you'll learn:
- Why memory, not the chip, decides which AI models your laptop can run
- CPU vs GPU vs NPU, explained with a restaurant kitchen
- What a model actually is (spoiler: it's just a file)
- What "8 billion parameters" really means
- Mixture of Experts, and why it fits like a big model but runs like a small one
- Context windows, tokens, and why the model "forgets" you
- MCP, agents, and why the app matters more than you think
- Robert's closet full of old Surface laptops now running as one AI cluster with NVIDIA PAIR
Chapters
00:00 Cold open: the comment that started it
00:50 Intros (from a Sprinter van in Iowa)
02:26 The stack: five layers, all on your machine
03:20 Layer 1: Silicon. Head chef, prep cooks, and the food processor
08:50 Layer 2: Runtime. The VLC for AI
11:04 NVIDIA PAIR: one address, every machine in the house
15:46 Layer 3: The model. It's just a file
19:53 What's a parameter? Knobs on a mixing board
21:46 Mixture of Experts: pull the L encyclopedia
28:58 Layer 4: Context. The open-book test
33:06 Layer 5: The app. Agents, 11-hour coding sessions, and the harness
38:15 MCP: USB-C for AI
42:17 Wrap: the whole stack in five lines
About the localhost
the localhost is a weekly podcast about local AI: running models on your own hardware, in your own house, with your own data. No cloud required. We're five friends from the Windows and devices world who spend our free time (and a lot of our electricity) figuring out what actually works, and then explaining it in plain English. 127.0.0.1. There's no place like home.
The hosts
Frank Buchholz: Former CTO of Surface Marketing at Microsoft, now running Cadence 3, a consultancy focused on local and on-device AI. Recording this week from his Sprinter van.
Chauncey Larsen: Microsoft. Diehard tech fan, Lego nerd, and the guy whose electricity bill proves it.
Robert Henry: Microsoft, Windows Incubation. Owner of a closet full of Surface devices that are now an AI cluster. Will be wearing a different hockey jersey every week until further notice.
Jacob Rhoades: Microsoft. Thinks about devices as endpoints for a living and runs more local models than anyone should.
Neil Misak: Out this week, somewhere over the Atlantic on his way back from London. Back next episode.
Chauncey, Robert, Jacob, and Neil work at Microsoft. Opinions here are their own and not their employer's.
Shout out to Lars Berlau at HP Poly for the comment that made this episode happen. Lars, let us know if it landed.
If you've been nodding along without following, you're exactly who we want to hear from. Drop a comment and tell us what to explain next.
Listen everywhere: https://thelocalhost.show
#LocalAI #AI #NPU #NVIDIA #DGXSpark #Surface #CopilotPlusPC #Ollama #LMStudio #MixtureOfExperts #MCP #Podcast
A listener (a Microsoft MVP, no less) told us he loves the show but doesn't understand all of it. Fair. So this week we go back to the studs and build local AI from the silicon up, one layer at a time, with zero jargon left unexplained. Head chefs, movie players, mixing boards, open-book tests. If you can follow a kitchen, you can follow this.
What you'll learn:
- Why memory, not the chip, decides which AI models your laptop can run
- CPU vs GPU vs NPU, explained with a restaurant kitchen
- What a model actually is (spoiler: it's just a file)
- What "8 billion parameters" really means
- Mixture of Experts, and why it fits like a big model but runs like a small one
- Context windows, tokens, and why the model "forgets" you
- MCP, agents, and why the app matters more than you think
- Robert's closet full of old Surface laptops now running as one AI cluster with NVIDIA PAIR
Chapters
00:00 Cold open: the comment that started it
00:50 Intros (from a Sprinter van in Iowa)
02:26 The stack: five layers, all on your machine
03:20 Layer 1: Silicon. Head chef, prep cooks, and the food processor
08:50 Layer 2: Runtime. The VLC for AI
11:04 NVIDIA PAIR: one address, every machine in the house
15:46 Layer 3: The model. It's just a file
19:53 What's a parameter? Knobs on a mixing board
21:46 Mixture of Experts: pull the L encyclopedia
28:58 Layer 4: Context. The open-book test
33:06 Layer 5: The app. Agents, 11-hour coding sessions, and the harness
38:15 MCP: USB-C for AI
42:17 Wrap: the whole stack in five lines
About the localhost
the localhost is a weekly podcast about local AI: running models on your own hardware, in your own house, with your own data. No cloud required. We're five friends from the Windows and devices world who spend our free time (and a lot of our electricity) figuring out what actually works, and then explaining it in plain English. 127.0.0.1. There's no place like home.
The hosts
Frank Buchholz: Former CTO of Surface Marketing at Microsoft, now running Cadence 3, a consultancy focused on local and on-device AI. Recording this week from his Sprinter van.
Chauncey Larsen: Microsoft. Diehard tech fan, Lego nerd, and the guy whose electricity bill proves it.
Robert Henry: Microsoft, Windows Incubation. Owner of a closet full of Surface devices that are now an AI cluster. Will be wearing a different hockey jersey every week until further notice.
Jacob Rhoades: Microsoft. Thinks about devices as endpoints for a living and runs more local models than anyone should.
Neil Misak: Out this week, somewhere over the Atlantic on his way back from London. Back next episode.
Chauncey, Robert, Jacob, and Neil work at Microsoft. Opinions here are their own and not their employer's.
Shout out to Lars Berlau at HP Poly for the comment that made this episode happen. Lars, let us know if it landed.
If you've been nodding along without following, you're exactly who we want to hear from. Drop a comment and tell us what to explain next.
Listen everywhere: https://thelocalhost.show
#LocalAI #AI #NPU #NVIDIA #DGXSpark #Surface #CopilotPlusPC #Ollama #LMStudio #MixtureOfExperts #MCP #Podcast
