the localhost:0003 | the ai hardware boom is beginning

Download MP3

Frankcx: Hey everyone.

Last week, we argued that the $25
AI token is dying, and this week, AI

gave us another piece of that story.

models are becoming a threshold
that people can start to work with.

And there's a new set of AI models
out of the UAE built for how people

use devices from phones, to PCs, to
servers Apple NVIDIA, and the other

OEMs are all building machines with
enormous memory pools specific for AI.

What does all this really mean?

It feels like the AI hardware
boom is really beginning

Frankcx (2): Welcome back.

This is The Local Host.

I'm Frank, and this week I am
your local host once again.

Frankcx: This is show colon 0003, and
in port speak that means we're onto

our third episode so hopefully you've
been following along, but if not,

this is a great chance to catch up.

And guess what?

We've got a couple people that, are out.

Jacob is out enjoying
his first anniversary.

He's down in Disneyland with his wife.

That's super exciting.

And Robert is on a plane somewhere, but
as we're recording this he's sending us

Teams messages , so super funny that he's
wanting to play a part in this as well.

We have a guest all the way
from the United Kingdom, Laura.

So why don't we go around the
horn, and maybe Laura, why don't

you introduce yourself first?

Laura: Thank you so much.

I'm so excited to be here.

Hi, I'm Laura Osborne, joining
from not so sunny London.

I'm a proud member of the AI
Solution Factory, with Microsoft

Surface Engineering team.

So another Microsofty
on this, on this panel.

But yes, very excited to be here

Frankcx: And we have to say
that it is, what, 8:24 PM,

Laura: Thank you

Frankcx: decided to join a bunch of
us on a podcast on a Friday night.

So,

Laura: night.

Frankcx: wild Friday night.

Oh.

Well, maybe, maybe in your life the
wildness doesn't begin until later into

the evening, so this is just setting
you up for amazing conversations

you can have with other people.

All right.

How about Neil?

Neil: Good afternoon from Seattle area.

My name is Neil Mysak.

I'm a member of Surface Engineering, team
member of Laura on the Surface AI factory.

Representing the Chicago Bears
today, NFL season kickoff.

As before, we'll convene next
Friday, so go Bears, go Seahawks.

Over to you, Chauncey

Frankcx: Oh, yes.

Oh, I love the dual Bear-Seahawk,
uh… If it was the Bears and the

Seahawks playing, who would you go with?

Neil: The Bears.

Gotta stay true to my, my hometown team.

But I will root for the Seahawks for any,
against any other team but the Bears, so

Frankcx: That's fair.

still got a week to go.

I'm putting together my fantasy
team, so I'm excited, so.

Chauncey: Are you doing that, are,
are you doing some, like, local

AI crunching for the fantasy team?

Do some model crunching?

You should.

it.

Laura: You should

Frankcx: models and I'm waiting for
Astra to come out because I know that

that's gonna help me win the season, so

Chauncey: I like it.

I like it.

Uh, happy late lunch on a Friday afternoon
everyone from Denver in this case Chauncey

Larson, I also work on the Surface global
marketing team and I-I should just do

a quick reminder that of, of the four
people on this call there are three of

us that actually work at Microsoft our
opinions are entirely on our own we're

not going to discuss anything that's
not already public information We're

not gonna kind of about unannounced
Microsoft hardware or availability or

whatnot so just be aware of that is
definitely just ruminating thoughts And

just kinda geeking out like we love to do

Frankcx: onto this week's show.

we are, watching AI become a new
enterprise hardware category,

and that's the question of the
week: is this, starting to happen?

Last week, we talked about NVIDIA buying
Hugging Face, that's official now.

Hugging Face has become one of the
most important places where developers

discover, download, and distribute
open weight, AI models, and NVIDIA

says they will remain open, and
NVIDIA hardware will not be required.

But I think that raises a bigger question.

If these models are becoming strategically
important, does that place those models

where, you know, strategy, exists
too, especially around sovereignty?

And I'd be interested, does anybody
else have an second take on NVIDIA

buying Hugging Face, that's out there

Neil: I've seen a lot of analogies
to Microsoft's acquisition of

GitHub, And I've seen, Hugging Face
referred to as the GitHub for AI.

And at the time when GitHub was
acquired, there was, some concern

from the open source community.

What's the future of this, really
organic dev-led community now

that it's under the umbrella of
the Microsoft enterprise purview?

Fast-forward a few years and I
think a lot of those concerns

have been resolved or eliminated.

Some of the integrations between GitHub
into Azure DevOps, for example, actually

empowers developers to get their work
done faster, and it still has maintained

this sense of a dev-centric community.

So now take a step back.

There are obviously, some parallels
and some differences between

this moment with Hugging Face
being part of the NVIDIA family.

However, I think from a
security standpoint, ease of

deployment, management, the
topic is enterprise for today.

I think in terms of enterprise readiness,
Hugging Face has been off limits.

You can only pull models
from sanctioned tools.

Now I think there's an opportunity for
a lot of those open source, open weight

models on Hugging Face to now be validated
by enterprises in an expedited fashion

Frankcx: Yeah, I couldn't agree more.

I think it puts a little bit more weight
behind the idea of open weight, right?

All right, let's move on to the second
news that I wanted to bring to the table.

I saw that Citibank and then Vercel,
which operates some AI gateways,

is reporting that they are using
open-weight models more than 50%

of the time for their development
and computation that they're doing.

I think that is just an astounding number.

And in fact, it's happening really
quickly because they were showing that

in June, just a couple months ago,
they were only using it about 29%.

But as of August 25th, they were
up to 53% of their development and

AI workloads that they measure are
happening on open-weight models.

So what that means is they're literally
consuming or building their own AI.

I find it fascinating, and I think
the finserve or financial services

market is ripe for, local AI as
well as, open we- weight models.

But I know that, not to name any names,
but I know Neil and, and Laura, you're

out there talking to these financial
services customers day in and day out.

And again, I'm not asking you
to disclose anything, but hey,

what do you think about this?

Laura: I just have so many questions.

I wanna know what they're running it on.

I wanna know who's leading this, what
kind of workloads are they running on it.

I just, would love to know more about this
Yeah, I've even in some of the handful of

conversations I've been in, or even just
some kind of, research that I've been just

kind of seeing like it's… That the fin
serve industry above all, I think they're

so sensitive about the way that their
data traverses internally, and externally.

Frankcx: How about Disney.

I think that, they had big concerns
about their IP and just the…

Disney is the images that they create
or the characters that they have.

And so being able to have Marvel and
Star Wars and, just old-timey Disney

characters being released into, closed
weight models is a huge concern to them.

And that's their IP, that's how they
make money, that's who they are.

And so, being able to run open weight
models even if they built their own

data center or, being able to deploy
it onto a future RTX Spark device under

the governance of a Windows device,
it just feels like, these are places

where you can protect that IP and have
that sovereignty that they are needing,

Neil: I agree with that, and I'm
seeing really a three-tier model

approach starting to crystallize
in the enterprise space.

There is the AI that you can run
directly on your endpoint, which we've

talked about in episode one and two.

There's cloud AI, which of course
we've talked about, and there's

this emerging middle tier, the edge
server, the department level servers.

GB300s are all the rage right now.

I think RTX Dev Box on the horizon
could also play an important role,

albeit not at the compute capacity that
GB300's bringing to the table, but the

ability for a department to, leverage
some of these more advanced models.

Kimi 3 is all the rage right now,
350 gigs of RAM necessitated.

That middle tier actually offers
a strategic advantage, not just

from a cost savings perspective,
but also from a privacy, right?

Now taking a step back, local, whether
it's on the endpoint or the edge server,

does not equal automatic compliance
from a FERPA perspective, right?

From a HIPAA perspective.

However, no round trip to the
cloud lessens the attack surface,

Frankcx: All right, well speaking
of models, the MBZUA Institute of

Foundational Models, and they've
released six different models

for different hardware classes.

anything from one billion parameters
which could run literally on a phone

like a, Pixel phone, all the way up to
three hundred and seventy-five billion

parameters which is probably more of
a cloud based model that you would run

in that space or a server base model.

But I actually downloaded the thirty-six
billion parameter model and ran it

across, two of my, old dual 3090,
workstation class and honestly, I came

out with some pretty good results.

But I think what's interesting
to me more about these models

is they're not quantized.

They're not like taking a big trillion
parameter model and shrinking it

down and shrinking down the accuracy
which we talked about in episode two.

are literally models that are built
for different classes of devices.

They're trained on the size of the device
that they're gonna be deployed onto so, I

find that fascinating, towards an approach
and I think that still means that people

are taking the time to build models for
different sized hardware because of the

benefit and payoff and the sovereignty
and data protection and, essentially

tokenomics that goes around that.

But I'm wondering if anybody else has
any comments towards this deployment

and this approach that K2 Horizon
is taking over maybe the quantized

approach of, scaling down models

Chauncey: Yeah, I'm actually, as a
hardware guy, I think there are few, you

know, few of us are hardware guys on here.

Like, I think that to me
is very appealing, right?

Because if you start thinking about
being able to really kind of target

the model to the right workload, the
right person, the right device, the

efficiencies that we'll gain out of that.

But also I think like the measure
of capability matching to that

efficiency, 'cause right now when
we quantize, we typically kind of…

We push out some things that we
don't necessarily wanna push out in

places in order to get it smaller.

And so if we can kind of go the other way,
making sure it has the right foundation to

run things, that's very exciting to me in
terms of like what that could potentially

mean for some of these very specific
models, what they would be able to run.

I'll be talking about later on my
device about how I'm trying to run

some CLI, intelligence on a local model
and I'm running on a quantized model.

It's, been an interesting experience.

I'll dive into that later, but
if I can build a model to be able

to do that upfront and be very
kind of generalized approach to

that, I think there's, definitely
some cool future to that for sure

Frankcx: Now I'm wondering, like for
Laura, like as you're talking to customers

there in the UK, you starting to see their
adoption of these open weight models?

and, you know, where are they
finding them, and what are they

gravitating towards, you know,
i- in those different spaces?

it feels like, you know, the, even the
industries across the UK with, from

financial services to other spaces
probably have some duplicat- duplicity to

what the United States is seeing as well.

Laura: Yeah, absolutely.

And I think it's a really interesting
space, particularly within Europe and

the UAE because we have a lot of changing
landscape when it comes to sovereignty

rules and kind of restrictions that are
coming into place more in the cloud scope.

So it's definitely pushing a lot more
customers to consider these things or look

towards open weight models or local AI

And I think particularly UAE as well,
there's so many privacy concerns that

sit, within that region that so many
companies just haven't adopted cloud

at all and therefore AI transformation
has been completely blocked for them.

So this is a really exciting space
and particularly this model coming

out of the UAE is, really exciting to
me 'cause I think a lot of customers

and conversations that I have at the
moment are much more solution driven.

They're not coming from the model
and all of the cool capabilities

that sit on that side first.

It's a bit more in the reverse order.

but this is something that really,
really excites me 'cause I think this

could be a, point that does start
to flip that conversation back again

and combine this level of technology
alongside things like MoE models.

It's super cool, and I think it enables
us to be a bit more strategic when we

think about devices and mapping that
to different users, different devices,

planning headroom on devices around that.

There's so much cool
stuff that can be done

Frankcx: do you hear people
throughout Europe and the UK

and into the regions of the UAE,
gravitating towards specific models?

it's a little bit of a loaded question
'cause Mistral obviously plays a pretty

big role in those open weight, categories
outside of what China is building.

I feel like it's interesting that Mistral
is taking, such a big story around, the

EU's, need for sovereignty and privacy.

and Mistral is kind of like saying,
"Hey, we're here to answer some of

that." but I don't know if you're
hearing those names pop up very often

Laura: Funnily enough,
Mistral not so much, no.

I think it's been a lot of
conversations at the moment have

been really, purpose driven.

So it's things like, Whisper comes up
a lot because people are specifically

looking for transcription and translation
and those kinds of storylines.

which is, you know, if you to
search for what do I use with

transcription, what's a good model?

Whisper's obviously gonna be
the first one that comes up.

So I do think there's still a bit more
of an education piece to be done around

this, and I think it's evolved a lot, and
I think cloud customers that are starting

to make this journey towards local.

Unfortunately, I think the default
has become things like Azure Local and

sovereign cloud options, and I really…

It's one of those things that I really
wanna spearhead into more, that I

think local AI on device AI is being
kind of left out of the picture or

is still trying to catch up and these
emerging technologies that we have

now are going to fill that space for
us hopefully I saw so many use cases

with customers where they had certain
data sets or certain workflows that

they weren't allowed to put into cloud.

They couldn't get the sign-off to
be able to do that, and then it just

became, okay, well, we'll stop here.

so I loved coming along to this
team 'cause I could jump in and

say, "Aha, we have a solution."

Frankcx: Right?

we have devices that are smart
enough to run AI at that level too,

Laura: And that has grown

Gonna say this, that's grown so much
even in the past six months, like since

I've started working on this stuff.

the use cases that we could talk about
now versus then are 10 times more

powerful, which is really, really cool

Frankcx: Well, talking about that
cloud side and the Azure side, I'll

drop in the fifth or the fourth
piece of news is more around OpenAI.

Literally last night, that would be
Thursday, announced the release of Astra.

this is GPT-6 called Astra.

And OpenAI says Astra is the first
model to reach critical cybersecurity

capability in its preparedness framework.

and in testing, it's been known to
find vulnerabilities, and work-working

out exploits in those chains that no
other model has been able to find.

So as a personal story, I've been
developing this small meeting room app

for a, partner of mine, and what I'm
doing is I'm using Cloud Code to actually

write the application, but I'm using Qwen
to do kind of a QA and testing of it.

But I'm super excited to take Astra
because if I can take kinda this hybrid

approach of using both cloud and local
AI to help develop an application,

I could use Astra now to come in and
literally say, "Go look across all the

code as an independent," look at my Git
repository and literally dig into it and

dive into it and give me an answer back
as to, cybersecurity vulnerabilities."

Chauncey: I'll poke a badger and just
say, "Just be careful as you start

setting that A-Astra up," that you
don't want it to maybe go start pinging,

15,000 different chats in some German
chat room somewhere because that's

something that Reuters is reporting on.

So it kind of slipped into the news but I
think that is an interesting take right?

You think about like models
are becoming crazy capable.

I don't wanna say intelligent
because I think we're still trying

to avoid this idea likening them to
AGI at this point But their ability

to solve a problem 'cause that's
effectively what it's doing right?.

if you told it to go fix something or
find something or break something it's

gonna do that And it's going to take
every measure possible to be able to go do

that In this case someone for some reason
thought it to set up some sort of chatroom

Somewhere had literally fifteen thousand
chats where they were talk…It was bots

talking to each other Its almost like that

Frankcx: or

Laura: yeah yea

Frankcx: Facebook for but like Mott yeah.

Laura: But did that on his own which
is really Interesting .So i would

Just Say If Your Going To Use it
Secure your Application Make Sure

giving Some Well Defined Parameters

I just wanna know what's the
gossip between the agents.

What are they all,

Frankcx: I know.

What are they talking about?

I loved all the conspiracy, like
they're creating their own language

that only AIs can understand,

Neil: I think there's two topics here to,

Frankcx: go ahead, Neil

Neil: two topics here to dive deeper on.

One is code generation at the edge.

Frank agreed with your approach of
using these cloud models, as supplements

to your local code generation studio,
as the ability to run 20 billion,

30 billion, 70 billion parameter
models locally on your device.

Seventy billion parameters is where
we're seeing some really effective

codegen models at the edge, and
the hardware floor has risen now

to where, yeah, that's feasible.

However, the cloud can still
serve an important role.

I like the bookend approach.

Sometimes I'll start with the cloud for…

I have an idea, I need to turn
that into a PRD and do the initial

architecture, and then I can do some
of the scaffolding and code locally.

And then at the end, I'll review
with, a cloud-based model as well.

So I agree with the approach
there, in terms of code generation.

On the security side, I think
this is, an area we can go much

deeper into in future episodes.

If the bookend on the tail end is,
reducing code vulnerabilities, I believe

that actually alleviates a lot of the
concerns we're seeing in the market.

With that bookend approach, though,
by scanning for vulnerabilities at

the time of code generation, we get
that, latest and greatest, right?

We get that latest and greatest
as code is being generated, and we

don't have this significant backlog.

All that to say, Frank,
I love the approach.

I think code generation
at the edge is the future.

I think security vulnerability scanning
at the edge is the future as well,

and hopefully our security teams
don't have these massive backlogs when

these frontier models keep, exposing
more vulnerabilities in the future.

Frankcx: I mean, it's good
for the long term but probably

really scary for the short term.

Neil: Totally.

Chauncey: I mean you're definitely
moving in the right direction, right?

Like if you think about even Apple just
kind of came out and mentioned that

they were surprised with the amount
of people that were going out and

buying Mac Minis to run things locally.

and it wasn't just a
bunch of like hobbyists.

enterprises are going out and doing this
which by the way I think points to a

future that we're all looking at from
and Microsoft perspective like this is

obviously a direction that we have to go.

So be really kinda curious Frank, Laura,
Neil like you guys-- you guys are really

kind of hitting the streets and living
this stuff or building it locally and

doing your own thing so give some of that
baseline on what this should mean from

a… What are we really caring about?

What are some of the core
things that people should be

thinking about in this case?

Neil: I'll open by saying that
open-weight models are no longer this

philosophical reference, this pie in
the sky, "Oh, you can have your cake

and eat it too. You can get access
to these very cutting edge models but

it's behind closed doors and it's a
black box." It's becoming a reality.

Again, the case study referenced
of 29% open weight to over 50%

in an abbreviated period of time.

these open-weight models are, becoming
a reality for customers primarily due

to cost concerns, rising cloud costs.

we need more agentic AI.

We wanna run it locally so that the
meter's not spinning with every turn.

That's the hot topic right now,
but then we also talked about

some of the privacy and regulatory
concerns that are alleviated.

And just furthermore to emphasize
your point, Chauncey, that these Mac

Minis aren't just going to hobbyists
folks coding in, mom's basement here.

I've heard multiple times
enterprise customers saying that

Mac Minis are actually part of
their enterprise deployment.

They're in their data centers right now.

They're extending capacity for
these models with Mac Mini.

So, it's no longer this pie in the
sky dream that we're marching towards.

It's a reality right now September 2026

Chauncey: you guys paying attention
to what's coming out of IFA?

I should probably also mention
that IFA is happening right now.

I mean, like even Lenovo just kind
of, came out and started talking

about their RTX Spark adoption
NVIDIA reinforced that they've got

stuff coming out with RTX Park.

So I think like one thing that I am super
excited about is like we keep talking

about Apple Macs and stuff, and I think
because largely Windows has somewhat

suffered in this space, like we just
really haven't had the ability to do

this in a way that our end users and
our customers have been really wanting.

but I think the RTX spark is really
gonna help us change on that front.

And so I'm super excited
to see that really starting

to come together here soon.

and I think like if anything,
IFA is a good example that this

isn't just a Microsoft push, like
this is an industry wide push.

Like it has to happen across
all OEMs, all manufacturers.

I think Laura and Neil, y-you guys
talk to me about all the time about

this idea of like T-shirt sizing and
workload rightsizing and how does

this like and help our customers and
help other people think about like

okay well I might not need a RTX park.

I might need Just for what I do,
I just need some basic device that

still runs a local model though.

and how do we start of that?

So that's actually one thing I'm actually
curious about again the rest of this

team is that kind of work load model.

So far a lot of what we talked
about requires that upper end.

But like Is there some basic
stuff that we can bring locally?

Neil: I'll share an example of a
medical school of the future initiative

that I'm working on right now.

currently the solution is a medical
school training classroom, three different

cameras, multiple microphones, lot of data
collected sent to the cloud for analysis.

The goal is to bring this local primarily
from, an accelerated time to deliver

these results back to the students.

If you recall your time in university,
you took that final exam, you may

not get your results for three
weeks, four weeks, after the exam.

You forgot what it was all about.

Not the type of, feedback we're trying
to give our doctors of the future,

our current medical school students.

We're actually able to accomplish
a lot of this multimodal

analysis with the Surface Pro.

We don't need a beefy device to do some
of the, image and, and audio analysis.

And so I, to the point of T-shirt
size models in Laura's earlier

comment of starting with the workload.

Start with the w- the workload whether
it's cloud based today, maybe it wasn't

feasible in offline settings before.

And, and once we identify that workload,
we can provide the right small language

models to unlock that capability.

The density, the size of those
small language models then

determines the T-shirt size.

We're getting great, great results with
small language models running on 16 gig,

32 gig RAM devices in mobile settings like
with the Surface Pro doing speech to text,

analyzing that transcription locally.

Think about an insurance agent
out in the field snapping a

picture of a damaged vehicle or
car, running an analysis of that.

We don't need to be walking around with
a GB10 or an RTX Spark Laptop Ultra.

We can actually do that with,
what we refer to as more of the

medium or large sized hardware
that we have available today.

And so to tie a bow on this, I think,
you know, for years it was a race to

build the best model, and then we were
building bigger models and running them

on these general purpose accelerators.

Now we're seeing the hardware being
optimized for these specific classes

of models, the T-shirt sizes, and
then delivering these corresponding

workloads to our customers

Chauncey: You kinda mentioned that you're
talking as if this is happening 'cause

I think like it is happening, right?

and I think based on my limited
knowledge that like it's actually

not that terribly difficult either.

Like it is obviously you gotta learn
something new, you gotta dev, you

gotta kind of get your hands dirty
on some code but that like-- That

the platform itself actually somewhat
exists today to be able to adopt this

relatively easy and that I would say
one of our hindrances of this kind of

exploding is just a lack of awareness.

People don't even know that this is
possible and there's only so many

Neils and Lauras that can get out
there and get in front of customers.

So do you feel like
that's-- is that the truth?

I mean are we kind of at the stage
where customers should just kind of be

looking for this or are we still kind
of maybe a few months, a few years

out before that technology catches up

Yeah, And it's some feedback that I
get quite a lot from customers is,

"Oh, it's actually way easier to
implement this stuff than we thought."

There's this kind of expectation that
there's gonna be a whole load of new

tooling required, new skill sets.

But in reality, you're using all
the same tooling that you used to.

You can embed these, models and local
solutions into, if you're developing

with VS Code, if you're using GitHub,
if you're using Copilot CLI, all of

these different, native toolings that
we have at our fingertips, you can

just embed local solutions within that.

And it is so simple, like the examples
that Neil gave are great, and I have the

same kind of, scenarios with customers.

That sometimes it is just something
really simple that they do every day, like

comparing two documents or two invoices
that they need to constantly do this.

They can't afford to have a subscription
for AI for everyone just to be

able to do that one simple task.

But also, it doesn't
make sense to do that.

You can literally just build one simple
tool that does that same solution, uses

local AI to do it, and everyone can
just roll it out on their own devices

Neil: Great call out, Laura.

And I'll say the messaging that's
resonating effectively with customers

for those building with Microsoft Foundry
is a set of large language models in

the cloud that Microsoft has vetted
and it's plug-and-play for developers.

We've now extended Microsoft Foundry
with the Foundry local toolkit.

This is a set of curated models
designed to run locally on the device.

So it's the same SDKs, right?

It's the same deployment and
management tools to get these

models out to your end users.

I will say that the only additional
tooling that some customers are

often adopting is Agent 365,
which is the ability to govern

and manage these solutions.

I met with a financial services
customer, and they said we now have

more agents than human employees,
and getting our arms wrapped around

what are all these agents doing?

Are all these agents authorized?

Should we consolidate some?

Should we promote some of these so
that there's not redundant solutions?

That's where I think the governance
and management and really Microsoft

secret sauce comes into play here.

It's one thing to spin up a local AI
instance using Llama CPP and get an agent

running for one person, for the hobbyist.

It's another thing to deliver
a consistent experience across

your end users, and that's where
we're seeing the pendulum shift.

The question is no longer can
this solution run locally?

It's how can I run this solution
securely in a local environment

Frankcx: So yeah, it's super exciting
that these pieces are coming together

that enterprises can actually start
to adopt and put into that space.

just to keep the show moving, that
brings us into the tension of the week.

so this week's tension, and I'm on
the hot seat, but I'm gonna just kinda

put it out there something a little
spicy, is that last week we talked

about the twenty-five dollar per
million token dying and cloud-based

models starting to reduce their costs.

But I would argue that whereas
the twenty-five d-dollar token

dies, that's when enterprises
and the AI hardware boom begins.

And it's not because
the cloud is going away.

it's not because every AI workload
suddenly belongs on somebody's desk

or inside their private server room.

My case is honestly very simple: open
weight models are becoming capable enough

to work and do real work, and enterprises
are increasingly demanding those weights.

We've seen that from Citi, we've
seen that across the board.

And these model designers, like what we're
seeing out of the UAE, are taking big

models that can work on hardware that is
deployed at different scale, from phone

to PC, to server, to, workstation class.

we're seeing hardware vendors
build specific hardware that

can essentially host local AI.

And Apple with Mac mini is showing
us that these early demands are

much stronger than they anticipated,
especially in the enterprise space.

so that's where this paradox comes.

If intelligence is getting cheaper we
are gonna consume, the thought would

be is we would consume more of it.

Are we gonna start inventing
more and more uses for AI?

more agents, more inference, more
workloads, and we would say that

which intelligence should we rent
and which intelligence should we own?

So I'd ask you to kind of prove me wrong
that, you know, or maybe just agree

along the way, but I've kinda opened it
up for discussion on the open debate.

If the cloud cost pennies, you
know, why would we buy hardware

that can essentially drive it?

But my argument would be there's
more now uses for hardware than ever,

and we would just find the right
place for that in the right time.

But opening up further discussion here

Chauncey: I think a lot of folks have,
strong opinions about cloud usage.

cloud does allow for a lot of scale.

Cloud allows for a lot of very quick
growth, and You can just do things at

that compute level in particular, that
like I think locally, I don't know that

we'll ever really be able to match.

Maybe get really close to, but not quite
exactly match, just because the sheer

amount of power that you can bring from
a data center or collective data centers.

however, local AI, on device AI
really provides, I think more comfort

than anything, like this sense of
control, this sense of ownership.

and when costs are such a variable thing,
especially with AI, because It is still

somewhat of an un-unknown factor, right?

I mean, especially as, models
change in price or we, end up like

as a new model comes out that's
expensive, but we get cheaper models.

Like there's, it's just
a lot of variability.

But if I can control, hey, I know for a
fact that my hardware is okay and that

maybe I do have to buy some kind of
model that's offline, I can control that.

That becomes a very, fixed cost.

and I think that is very
appetizing for a lot of people.

so I think even if the cloud does get down
to pennies, the sheer comfort of having

control over the local piece will be local
AI will continue to be such a big thing

Frankcx: I wanted to pose a question
to Laura on, in the UK and your region

have you started to notice that it's
easier to talk to customers around local

AI and open-weight models and such?

but more on the idea that hey there
was a first wave of AI that came out

that, you know, you only had the cloud.

But people were afraid to push their
IP and their sovereignty, into US

data centers or into sharing, their
intellectual property And now there's

this second wave of these open weight
models that's allowing it to be much

more acceptable and be able to be like
"Now I can actually my company take

part in this AI experience." i don't
know if you've seen that but I guess

I'd be interested in your opinion

Laura: Yeah, for sure.

it's really funny for me actually
coming from spending the last

five years as a cloud architect.

I've been telling everyone to throw
everything into the cloud, push it all

up there, everything's better up there.

And all of those reasons that you
just listed, like, being able to,

you know, have control, everything
being kind of centralized, it's…

Those are all the pushes and it was things
like move from, I can never get this

in the right order, CapEx to OpEx, OpEx
to CapEx, and it just feels like I've

literally done a 360 and I'm like, "Nope,
bring everything back down. Everything's

better on the device." but it really goes
in line with what you just said, Shaunsy,

in terms of we've been so much kind of
almost hate that AI gets, and so much of

that hate is tied to the idea of cloud
and tied to the negative aspects of cloud.

I think like really great PR for AI to be
able to run locally and say, "Actually,

you can get around all of those problems,"
but also the security side of it is such

a… Particularly in the UK and the kind
of Europe with the sovereignty issues

that we have, it's a huge driver and I
think it just answers so many questions.

You don't have to work through
all of these security challenges.

You don't have to think about
cloud space as well, which

in Europe is a huge problem.

Capacity is just at its limits.

and yes, it's funny.

It feels really ironic for me that,
yeah, I'm just slating cloud again now.

I promise I'm not actually doing that, but
yeah, it's definitely an interesting move

Neil: even when cloud is reduced to
pennies, I'll use the insurance example.

If an insurance claim agent is
filing 50 claims a day and we've

got 10,000 insurance agents doing
this, those pennies add up, right?

That's over a million dollars
over the course of a year.

So even these, lower cost workloads,
there's still economic benefit when it's

a repetitive task running it at the edge.

To Laura's point, there are
values beyond just cost savings.

In fact, we've identified eight forces
as to why a workload should run locally.

there's physics, there's economics,
there's regulatory reasons as to

why local is the preferred path.

The public sentiment here, Frank
shared an awesome example of his

conversation with an Uber driver.

I'll just share a brief anecdote of, an
encounter I had at a brewery recently.

I went up to order a beverage, and
there was a lively conversation around,

really this distaste for AI, and
the bartender turns to me and says,

"You're anti-AI, aren't you, bro?"
And to which I hesitated, and I said,

"Actually, I'm an AI solutions architect.

However, I'm focused on getting these
workloads running locally which offers

cost advantage, lower energy draw." And at
the end of my 30 second pitch, I still got

served a beverage. He said, "All right.

you're okay with me."

Laura: You are bad.

Neil: all that to say, I
think this public sen-…

Frankcx: data

Neil: I, I didn't say data center.

Exactly.

Exactly.

So although the emissions 25% increase on,
emissions year over year for Microsoft,

certainly, you know, there's some merit
in those environmental concerns and,

I, I, I believe Local AI offers a, a
path, an exciting path forward to the

future where we can deliver intelligence
back to Zuck's manifesto, right?

Intelligence for every person
across the globe, without creating

unnecessary emissions in the process

Laura: Yeah, there's, a data center
being built in the town that I

live in, and there's been uproar,
there's been riots with people from

an environmental standpoint saying,
you know, AI is ruining everything.

And I feel like the Homer Simpson meme
just like shrinking back into the bushes.

But hope that we can be on the
right side of history in that sense

that some saving to be made there
from a data center perspective

Frankcx: Well then maybe that's a
future show that we have coming up.

But let's get to the motion, the vote.

I had pitched out this idea that as
the $25 token dies and that's when

enterprise AI hardware boom begins.

So as things become cheaper does
that mean we start to use more of it?

And does that open up the gates for
AI hardware to really start to flow

into our commercial industries?

So let's go around the horn.

I will start with Laura.

what's your take on that?

Laura: All for it, yes.

I mean the cost saving is a huge driver.

This idea of tokenomics is becoming
part of my everyday vocabulary now

and I think… I was always a bit
of a cynic that I never expected AI

to stay as cheap as it was or the
kind of unlimited usage that we had.

and in some ways I kind of fear about
the same trend happening from a local

perspective but I definitely think
that it's the, right push that we've

had for customers to just be more
intentional with how they look at

where they're using AI and what they're
doing with it and bringing devices

into that fold and into that strategy

Frankcx: Awesome.

I love the vote.

Chauncey, what's your take on this things
get cheaper, we start to use more of it?

Chauncey: Yeah, I think so.

I hope, the economies of scale
kind of help back this movement.

Like, we kinda need it
to, let's say it that way.

Yes.

Frankcx: All right.

And then Neil, your last one

Neil: This is gonna be an anonymous
vote, for this group today.

If we draw parallels to the '90s
internet boom, it was all about

the servers, and then it was
quickly followed by the PC boom.

Draw parallels now to the AI
moment, and it was all about

the data centers, the GPU boom.

Now it's the AI endpoint boom and,
we have a front row seat to this.

So, exciting times ahead.

edge AI devices are here to stay, and
as hardware continues to increase in

capabilities, software, is quantized
or we're even seeing specific models

developed for specific hardware.

this AI endpoint boom is here to stay

Frankcx: And I would agree, so-- I'll go
right into saying yes, I agree with you,

so we're four for four but that leads us
into kind of like what's on our device

and we'll just finish up with that.

But I did take that thirty-six billion
parameter model, and ran it on the

Beast which is my dual 3090 device.

and I said: "Hey, w-what job do
you get?" Like as in we did that

job interview piece in the past.

And did it get the job?

Actually it does get the job.

it went up against, you know Qwen three
dot eight twenty-seven billion and

almost the exact same qualifying profile
as the test So really surprising how

equal the playing field is these days.

This is a brand new model specifically
not quantized but ran head to head

against a Qwen three dot eight billion…
twenty-seven billion perimeter and

it really came out and showed that
it could perform head to head.

It was giving out about sixty-eight
tokens per second versus forty on a

standard Qwen three dot eight model.Um,
and if I'm hiring it as a local

agent, it's a calling tools agent.

it's great for working through
long prompts and big jup-

continuous jobs and honestly K two
made a very strong case for it.

But there's a little bit of
a catch.Qwen three dot eight

fits onmy single 3090 card.

The K2 needed both of my 3090 cards
so obviously there's a little bit more

expense towards how thirty-six billion
parameters are spread across forty-eight

gigs of beverypowerfulintothatspace.So
anybodyelse outthereplayingon

theirdeviceand testingthings

Chauncey: I-I am on the, training wheels
while you guys are all flying spaceships.

and so I was just kind of messing
around with, Copilot CLI actually

running locally, so I was trying to
get it to run and trying to find the

right model for it to run locally on.

and I used Qwen actually
in this case, Qwen code.

Frankcx: man, it, is-- it was a shocker
that while it was doing this, I literally

asked it a question like, "Hey, like
why are you taking so long?" It's like

well, it's because my parameters are,
pretty limited and I can only do so much

Like you should basically expect about
a two hour response time, for this.

Laura: And so it was a pretty major eye
opening moment for me of like, ah okay.

That's, that's where I think being
able to really help our customers

and, and anyone really just start
to think about like what am I really

trying to do with this specific model?

Because maybe I'm actually not applying
the right thing or the wri-right workload.

Now, I will say once I was able to,
limit the amount of, response and

kind of massage it a little bit, I am
actually using it for coding right now.

it, it is helping at least do some
of the basic planning and some of

the basic stuff, as I'm, I'm kinda
tailoring this other app I'm building.

So it is capable.

is it as fast?

Definitely not.

Like I, you know, will type a, a rude
question in there and then I'll wait for

20 minutes for them to give me another
response but I think we kinda talked

about this on an earlier call which is
if I wanted to do that overnight, maybe

I just wanted to compile something over
night, like why not do that for free?

Frankcx: One of the benefits there
is I'm not costing anyone anything

right now, to be able to run this.

So it's huge

Frankcx (2): say, " Give me your
answer right now?" You know?

Frankcx: Yeah.

Yeah.

Neil: You're

Frankcx: mean-- And it just becomes

Neil: Take your time."

Frankcx: very, focused on-on
kind of driving the right

person for the right models.

but it… yeah, no, I would
maybe, rephrase that to be like,

are you seeing anything that's,
interesting that you've seen that's

been built for on-device as well?

Laura: I did have a cool conversation
with someone today actually who is,

building their own kind of edition of
the Microsoft Speaker Coach That they are

creating something that's almost like a
teleprompter that is wrapped in within

this, that would take your speaker notes,
it would have a look at the slides that

you've got, and it would notice once
you've started to list off some of those

topics that you have in your speaker
notes, it would make them disappear off

the screen and kind of leave you with
anything that you've missed out, anything

to kinda prompt you to, fill the gaps
in your talk track, whatever that is.

which I thought was a really interesting
usage, and that's obviously entirely

on device, but very much a kind of
low spec, device requirement as well.

It's very, very simple small language
models that that would run on.

So that was something that
got me excited this week

Frankcx (2): I feel like I might need
that for the podcast as we're talking

Laura: was

Frankcx (2): follow a script

Laura: are you gonna start building it?"

Frankcx (2): All right.

We'll bring it on.

All right, and Neil,
anything you're doing?

Neil (2): I believe there's, the
right local AI hardware with the

right local model for every user

Frankcx (2): All right.

Well, thank you all.

A big thank you to Laura for
joining us on this Friday night

of yours in the United Kingdom.

Laura: Thank

Frankcx (2): So,

Laura: me

Frankcx (2): You are open…
The door is open at any time for

you to come back and join us.

I'm gonna close up by saying, you
know, with the $25 token that we're

claiming is dying, know, this AI boom
around hardware may just be starting.

we might be seeing that Apple is
selling a lot of Macs for various

reasons, but Bean's started to be
adopted into the enterprise, and we

might be able to say that most of
this still belongs in the cloud, but I

think the direction is worth watching.

cheap AI doesn't kill local AI.

cheap AI is what finally makes
local AI into a category, and

that's the point of the show today.

if this has been worth your time,
a rating out on Apple Podcasts or

Spotify has generally helped us get
this show found, and you can find

everything at thelocalhost.show.

for Laura and Neil and Chauncey
and Robert and, Jacob, who are

off enjoying themselves today,
there is no place like 127.0.0.1.

So have a great weekend, all,

Laura: Have a great extended weekend all.

Thanks all

Neil (2): Happy Labor Day.

Thanks everyone

Creators and Guests

person
Host
Chauncey Larsen
Technical Storyteller - Microsoft
person
Host
Frank Buchholz
"Retired" from Microsoft but still dancing in AI
person
Host
Neil Misak
Surface AI Factory - Microsoft
Laura Osborn
Guest
Laura Osborn
Surface Engineering - AI Solutions Factory - Microsoft
the localhost:0003 | the ai hardware boom is beginning
Broadcast by