the localhost:0003 | the ai hardware boom is beginning
Download MP3Frankcx: Hey everyone.
Last week, we argued that the $25
AI token is dying, and this week, AI
gave us another piece of that story.
models are becoming a threshold
that people can start to work with.
And there's a new set of AI models
out of the UAE built for how people
use devices from phones, to PCs, to
servers Apple NVIDIA, and the other
OEMs are all building machines with
enormous memory pools specific for AI.
What does all this really mean?
It feels like the AI hardware
boom is really beginning
Frankcx (2): Welcome back.
This is The Local Host.
I'm Frank, and this week I am
your local host once again.
Frankcx: This is show colon 0003, and
in port speak that means we're onto
our third episode so hopefully you've
been following along, but if not,
this is a great chance to catch up.
And guess what?
We've got a couple people that, are out.
Jacob is out enjoying
his first anniversary.
He's down in Disneyland with his wife.
That's super exciting.
And Robert is on a plane somewhere, but
as we're recording this he's sending us
Teams messages , so super funny that he's
wanting to play a part in this as well.
We have a guest all the way
from the United Kingdom, Laura.
So why don't we go around the
horn, and maybe Laura, why don't
you introduce yourself first?
Laura: Thank you so much.
I'm so excited to be here.
Hi, I'm Laura Osborne, joining
from not so sunny London.
I'm a proud member of the AI
Solution Factory, with Microsoft
Surface Engineering team.
So another Microsofty
on this, on this panel.
But yes, very excited to be here
Frankcx: And we have to say
that it is, what, 8:24 PM,
Laura: Thank you
Frankcx: decided to join a bunch of
us on a podcast on a Friday night.
So,
Laura: night.
Frankcx: wild Friday night.
Oh.
Well, maybe, maybe in your life the
wildness doesn't begin until later into
the evening, so this is just setting
you up for amazing conversations
you can have with other people.
All right.
How about Neil?
Neil: Good afternoon from Seattle area.
My name is Neil Mysak.
I'm a member of Surface Engineering, team
member of Laura on the Surface AI factory.
Representing the Chicago Bears
today, NFL season kickoff.
As before, we'll convene next
Friday, so go Bears, go Seahawks.
Over to you, Chauncey
Frankcx: Oh, yes.
Oh, I love the dual Bear-Seahawk,
uh… If it was the Bears and the
Seahawks playing, who would you go with?
Neil: The Bears.
Gotta stay true to my, my hometown team.
But I will root for the Seahawks for any,
against any other team but the Bears, so
Frankcx: That's fair.
still got a week to go.
I'm putting together my fantasy
team, so I'm excited, so.
Chauncey: Are you doing that, are,
are you doing some, like, local
AI crunching for the fantasy team?
Do some model crunching?
You should.
it.
Laura: You should
Frankcx: models and I'm waiting for
Astra to come out because I know that
that's gonna help me win the season, so
Chauncey: I like it.
I like it.
Uh, happy late lunch on a Friday afternoon
everyone from Denver in this case Chauncey
Larson, I also work on the Surface global
marketing team and I-I should just do
a quick reminder that of, of the four
people on this call there are three of
us that actually work at Microsoft our
opinions are entirely on our own we're
not going to discuss anything that's
not already public information We're
not gonna kind of about unannounced
Microsoft hardware or availability or
whatnot so just be aware of that is
definitely just ruminating thoughts And
just kinda geeking out like we love to do
Frankcx: onto this week's show.
we are, watching AI become a new
enterprise hardware category,
and that's the question of the
week: is this, starting to happen?
Last week, we talked about NVIDIA buying
Hugging Face, that's official now.
Hugging Face has become one of the
most important places where developers
discover, download, and distribute
open weight, AI models, and NVIDIA
says they will remain open, and
NVIDIA hardware will not be required.
But I think that raises a bigger question.
If these models are becoming strategically
important, does that place those models
where, you know, strategy, exists
too, especially around sovereignty?
And I'd be interested, does anybody
else have an second take on NVIDIA
buying Hugging Face, that's out there
Neil: I've seen a lot of analogies
to Microsoft's acquisition of
GitHub, And I've seen, Hugging Face
referred to as the GitHub for AI.
And at the time when GitHub was
acquired, there was, some concern
from the open source community.
What's the future of this, really
organic dev-led community now
that it's under the umbrella of
the Microsoft enterprise purview?
Fast-forward a few years and I
think a lot of those concerns
have been resolved or eliminated.
Some of the integrations between GitHub
into Azure DevOps, for example, actually
empowers developers to get their work
done faster, and it still has maintained
this sense of a dev-centric community.
So now take a step back.
There are obviously, some parallels
and some differences between
this moment with Hugging Face
being part of the NVIDIA family.
However, I think from a
security standpoint, ease of
deployment, management, the
topic is enterprise for today.
I think in terms of enterprise readiness,
Hugging Face has been off limits.
You can only pull models
from sanctioned tools.
Now I think there's an opportunity for
a lot of those open source, open weight
models on Hugging Face to now be validated
by enterprises in an expedited fashion
Frankcx: Yeah, I couldn't agree more.
I think it puts a little bit more weight
behind the idea of open weight, right?
All right, let's move on to the second
news that I wanted to bring to the table.
I saw that Citibank and then Vercel,
which operates some AI gateways,
is reporting that they are using
open-weight models more than 50%
of the time for their development
and computation that they're doing.
I think that is just an astounding number.
And in fact, it's happening really
quickly because they were showing that
in June, just a couple months ago,
they were only using it about 29%.
But as of August 25th, they were
up to 53% of their development and
AI workloads that they measure are
happening on open-weight models.
So what that means is they're literally
consuming or building their own AI.
I find it fascinating, and I think
the finserve or financial services
market is ripe for, local AI as
well as, open we- weight models.
But I know that, not to name any names,
but I know Neil and, and Laura, you're
out there talking to these financial
services customers day in and day out.
And again, I'm not asking you
to disclose anything, but hey,
what do you think about this?
Laura: I just have so many questions.
I wanna know what they're running it on.
I wanna know who's leading this, what
kind of workloads are they running on it.
I just, would love to know more about this
Yeah, I've even in some of the handful of
conversations I've been in, or even just
some kind of, research that I've been just
kind of seeing like it's… That the fin
serve industry above all, I think they're
so sensitive about the way that their
data traverses internally, and externally.
Frankcx: How about Disney.
I think that, they had big concerns
about their IP and just the…
Disney is the images that they create
or the characters that they have.
And so being able to have Marvel and
Star Wars and, just old-timey Disney
characters being released into, closed
weight models is a huge concern to them.
And that's their IP, that's how they
make money, that's who they are.
And so, being able to run open weight
models even if they built their own
data center or, being able to deploy
it onto a future RTX Spark device under
the governance of a Windows device,
it just feels like, these are places
where you can protect that IP and have
that sovereignty that they are needing,
Neil: I agree with that, and I'm
seeing really a three-tier model
approach starting to crystallize
in the enterprise space.
There is the AI that you can run
directly on your endpoint, which we've
talked about in episode one and two.
There's cloud AI, which of course
we've talked about, and there's
this emerging middle tier, the edge
server, the department level servers.
GB300s are all the rage right now.
I think RTX Dev Box on the horizon
could also play an important role,
albeit not at the compute capacity that
GB300's bringing to the table, but the
ability for a department to, leverage
some of these more advanced models.
Kimi 3 is all the rage right now,
350 gigs of RAM necessitated.
That middle tier actually offers
a strategic advantage, not just
from a cost savings perspective,
but also from a privacy, right?
Now taking a step back, local, whether
it's on the endpoint or the edge server,
does not equal automatic compliance
from a FERPA perspective, right?
From a HIPAA perspective.
However, no round trip to the
cloud lessens the attack surface,
Frankcx: All right, well speaking
of models, the MBZUA Institute of
Foundational Models, and they've
released six different models
for different hardware classes.
anything from one billion parameters
which could run literally on a phone
like a, Pixel phone, all the way up to
three hundred and seventy-five billion
parameters which is probably more of
a cloud based model that you would run
in that space or a server base model.
But I actually downloaded the thirty-six
billion parameter model and ran it
across, two of my, old dual 3090,
workstation class and honestly, I came
out with some pretty good results.
But I think what's interesting
to me more about these models
is they're not quantized.
They're not like taking a big trillion
parameter model and shrinking it
down and shrinking down the accuracy
which we talked about in episode two.
are literally models that are built
for different classes of devices.
They're trained on the size of the device
that they're gonna be deployed onto so, I
find that fascinating, towards an approach
and I think that still means that people
are taking the time to build models for
different sized hardware because of the
benefit and payoff and the sovereignty
and data protection and, essentially
tokenomics that goes around that.
But I'm wondering if anybody else has
any comments towards this deployment
and this approach that K2 Horizon
is taking over maybe the quantized
approach of, scaling down models
Chauncey: Yeah, I'm actually, as a
hardware guy, I think there are few, you
know, few of us are hardware guys on here.
Like, I think that to me
is very appealing, right?
Because if you start thinking about
being able to really kind of target
the model to the right workload, the
right person, the right device, the
efficiencies that we'll gain out of that.
But also I think like the measure
of capability matching to that
efficiency, 'cause right now when
we quantize, we typically kind of…
We push out some things that we
don't necessarily wanna push out in
places in order to get it smaller.
And so if we can kind of go the other way,
making sure it has the right foundation to
run things, that's very exciting to me in
terms of like what that could potentially
mean for some of these very specific
models, what they would be able to run.
I'll be talking about later on my
device about how I'm trying to run
some CLI, intelligence on a local model
and I'm running on a quantized model.
It's, been an interesting experience.
I'll dive into that later, but
if I can build a model to be able
to do that upfront and be very
kind of generalized approach to
that, I think there's, definitely
some cool future to that for sure
Frankcx: Now I'm wondering, like for
Laura, like as you're talking to customers
there in the UK, you starting to see their
adoption of these open weight models?
and, you know, where are they
finding them, and what are they
gravitating towards, you know,
i- in those different spaces?
it feels like, you know, the, even the
industries across the UK with, from
financial services to other spaces
probably have some duplicat- duplicity to
what the United States is seeing as well.
Laura: Yeah, absolutely.
And I think it's a really interesting
space, particularly within Europe and
the UAE because we have a lot of changing
landscape when it comes to sovereignty
rules and kind of restrictions that are
coming into place more in the cloud scope.
So it's definitely pushing a lot more
customers to consider these things or look
towards open weight models or local AI
And I think particularly UAE as well,
there's so many privacy concerns that
sit, within that region that so many
companies just haven't adopted cloud
at all and therefore AI transformation
has been completely blocked for them.
So this is a really exciting space
and particularly this model coming
out of the UAE is, really exciting to
me 'cause I think a lot of customers
and conversations that I have at the
moment are much more solution driven.
They're not coming from the model
and all of the cool capabilities
that sit on that side first.
It's a bit more in the reverse order.
but this is something that really,
really excites me 'cause I think this
could be a, point that does start
to flip that conversation back again
and combine this level of technology
alongside things like MoE models.
It's super cool, and I think it enables
us to be a bit more strategic when we
think about devices and mapping that
to different users, different devices,
planning headroom on devices around that.
There's so much cool
stuff that can be done
Frankcx: do you hear people
throughout Europe and the UK
and into the regions of the UAE,
gravitating towards specific models?
it's a little bit of a loaded question
'cause Mistral obviously plays a pretty
big role in those open weight, categories
outside of what China is building.
I feel like it's interesting that Mistral
is taking, such a big story around, the
EU's, need for sovereignty and privacy.
and Mistral is kind of like saying,
"Hey, we're here to answer some of
that." but I don't know if you're
hearing those names pop up very often
Laura: Funnily enough,
Mistral not so much, no.
I think it's been a lot of
conversations at the moment have
been really, purpose driven.
So it's things like, Whisper comes up
a lot because people are specifically
looking for transcription and translation
and those kinds of storylines.
which is, you know, if you to
search for what do I use with
transcription, what's a good model?
Whisper's obviously gonna be
the first one that comes up.
So I do think there's still a bit more
of an education piece to be done around
this, and I think it's evolved a lot, and
I think cloud customers that are starting
to make this journey towards local.
Unfortunately, I think the default
has become things like Azure Local and
sovereign cloud options, and I really…
It's one of those things that I really
wanna spearhead into more, that I
think local AI on device AI is being
kind of left out of the picture or
is still trying to catch up and these
emerging technologies that we have
now are going to fill that space for
us hopefully I saw so many use cases
with customers where they had certain
data sets or certain workflows that
they weren't allowed to put into cloud.
They couldn't get the sign-off to
be able to do that, and then it just
became, okay, well, we'll stop here.
so I loved coming along to this
team 'cause I could jump in and
say, "Aha, we have a solution."
Frankcx: Right?
we have devices that are smart
enough to run AI at that level too,
Laura: And that has grown
Gonna say this, that's grown so much
even in the past six months, like since
I've started working on this stuff.
the use cases that we could talk about
now versus then are 10 times more
powerful, which is really, really cool
Frankcx: Well, talking about that
cloud side and the Azure side, I'll
drop in the fifth or the fourth
piece of news is more around OpenAI.
Literally last night, that would be
Thursday, announced the release of Astra.
this is GPT-6 called Astra.
And OpenAI says Astra is the first
model to reach critical cybersecurity
capability in its preparedness framework.
and in testing, it's been known to
find vulnerabilities, and work-working
out exploits in those chains that no
other model has been able to find.
So as a personal story, I've been
developing this small meeting room app
for a, partner of mine, and what I'm
doing is I'm using Cloud Code to actually
write the application, but I'm using Qwen
to do kind of a QA and testing of it.
But I'm super excited to take Astra
because if I can take kinda this hybrid
approach of using both cloud and local
AI to help develop an application,
I could use Astra now to come in and
literally say, "Go look across all the
code as an independent," look at my Git
repository and literally dig into it and
dive into it and give me an answer back
as to, cybersecurity vulnerabilities."
Chauncey: I'll poke a badger and just
say, "Just be careful as you start
setting that A-Astra up," that you
don't want it to maybe go start pinging,
15,000 different chats in some German
chat room somewhere because that's
something that Reuters is reporting on.
So it kind of slipped into the news but I
think that is an interesting take right?
You think about like models
are becoming crazy capable.
I don't wanna say intelligent
because I think we're still trying
to avoid this idea likening them to
AGI at this point But their ability
to solve a problem 'cause that's
effectively what it's doing right?.
if you told it to go fix something or
find something or break something it's
gonna do that And it's going to take
every measure possible to be able to go do
that In this case someone for some reason
thought it to set up some sort of chatroom
Somewhere had literally fifteen thousand
chats where they were talk…It was bots
talking to each other Its almost like that
Frankcx: or
Laura: yeah yea
Frankcx: Facebook for but like Mott yeah.
Laura: But did that on his own which
is really Interesting .So i would
Just Say If Your Going To Use it
Secure your Application Make Sure
giving Some Well Defined Parameters
I just wanna know what's the
gossip between the agents.
What are they all,
Frankcx: I know.
What are they talking about?
I loved all the conspiracy, like
they're creating their own language
that only AIs can understand,
Neil: I think there's two topics here to,
Frankcx: go ahead, Neil
Neil: two topics here to dive deeper on.
One is code generation at the edge.
Frank agreed with your approach of
using these cloud models, as supplements
to your local code generation studio,
as the ability to run 20 billion,
30 billion, 70 billion parameter
models locally on your device.
Seventy billion parameters is where
we're seeing some really effective
codegen models at the edge, and
the hardware floor has risen now
to where, yeah, that's feasible.
However, the cloud can still
serve an important role.
I like the bookend approach.
Sometimes I'll start with the cloud for…
I have an idea, I need to turn
that into a PRD and do the initial
architecture, and then I can do some
of the scaffolding and code locally.
And then at the end, I'll review
with, a cloud-based model as well.
So I agree with the approach
there, in terms of code generation.
On the security side, I think
this is, an area we can go much
deeper into in future episodes.
If the bookend on the tail end is,
reducing code vulnerabilities, I believe
that actually alleviates a lot of the
concerns we're seeing in the market.
With that bookend approach, though,
by scanning for vulnerabilities at
the time of code generation, we get
that, latest and greatest, right?
We get that latest and greatest
as code is being generated, and we
don't have this significant backlog.
All that to say, Frank,
I love the approach.
I think code generation
at the edge is the future.
I think security vulnerability scanning
at the edge is the future as well,
and hopefully our security teams
don't have these massive backlogs when
these frontier models keep, exposing
more vulnerabilities in the future.
Frankcx: I mean, it's good
for the long term but probably
really scary for the short term.
Neil: Totally.
Chauncey: I mean you're definitely
moving in the right direction, right?
Like if you think about even Apple just
kind of came out and mentioned that
they were surprised with the amount
of people that were going out and
buying Mac Minis to run things locally.
and it wasn't just a
bunch of like hobbyists.
enterprises are going out and doing this
which by the way I think points to a
future that we're all looking at from
and Microsoft perspective like this is
obviously a direction that we have to go.
So be really kinda curious Frank, Laura,
Neil like you guys-- you guys are really
kind of hitting the streets and living
this stuff or building it locally and
doing your own thing so give some of that
baseline on what this should mean from
a… What are we really caring about?
What are some of the core
things that people should be
thinking about in this case?
Neil: I'll open by saying that
open-weight models are no longer this
philosophical reference, this pie in
the sky, "Oh, you can have your cake
and eat it too. You can get access
to these very cutting edge models but
it's behind closed doors and it's a
black box." It's becoming a reality.
Again, the case study referenced
of 29% open weight to over 50%
in an abbreviated period of time.
these open-weight models are, becoming
a reality for customers primarily due
to cost concerns, rising cloud costs.
we need more agentic AI.
We wanna run it locally so that the
meter's not spinning with every turn.
That's the hot topic right now,
but then we also talked about
some of the privacy and regulatory
concerns that are alleviated.
And just furthermore to emphasize
your point, Chauncey, that these Mac
Minis aren't just going to hobbyists
folks coding in, mom's basement here.
I've heard multiple times
enterprise customers saying that
Mac Minis are actually part of
their enterprise deployment.
They're in their data centers right now.
They're extending capacity for
these models with Mac Mini.
So, it's no longer this pie in the
sky dream that we're marching towards.
It's a reality right now September 2026
Chauncey: you guys paying attention
to what's coming out of IFA?
I should probably also mention
that IFA is happening right now.
I mean, like even Lenovo just kind
of, came out and started talking
about their RTX Spark adoption
NVIDIA reinforced that they've got
stuff coming out with RTX Park.
So I think like one thing that I am super
excited about is like we keep talking
about Apple Macs and stuff, and I think
because largely Windows has somewhat
suffered in this space, like we just
really haven't had the ability to do
this in a way that our end users and
our customers have been really wanting.
but I think the RTX spark is really
gonna help us change on that front.
And so I'm super excited
to see that really starting
to come together here soon.
and I think like if anything,
IFA is a good example that this
isn't just a Microsoft push, like
this is an industry wide push.
Like it has to happen across
all OEMs, all manufacturers.
I think Laura and Neil, y-you guys
talk to me about all the time about
this idea of like T-shirt sizing and
workload rightsizing and how does
this like and help our customers and
help other people think about like
okay well I might not need a RTX park.
I might need Just for what I do,
I just need some basic device that
still runs a local model though.
and how do we start of that?
So that's actually one thing I'm actually
curious about again the rest of this
team is that kind of work load model.
So far a lot of what we talked
about requires that upper end.
But like Is there some basic
stuff that we can bring locally?
Neil: I'll share an example of a
medical school of the future initiative
that I'm working on right now.
currently the solution is a medical
school training classroom, three different
cameras, multiple microphones, lot of data
collected sent to the cloud for analysis.
The goal is to bring this local primarily
from, an accelerated time to deliver
these results back to the students.
If you recall your time in university,
you took that final exam, you may
not get your results for three
weeks, four weeks, after the exam.
You forgot what it was all about.
Not the type of, feedback we're trying
to give our doctors of the future,
our current medical school students.
We're actually able to accomplish
a lot of this multimodal
analysis with the Surface Pro.
We don't need a beefy device to do some
of the, image and, and audio analysis.
And so I, to the point of T-shirt
size models in Laura's earlier
comment of starting with the workload.
Start with the w- the workload whether
it's cloud based today, maybe it wasn't
feasible in offline settings before.
And, and once we identify that workload,
we can provide the right small language
models to unlock that capability.
The density, the size of those
small language models then
determines the T-shirt size.
We're getting great, great results with
small language models running on 16 gig,
32 gig RAM devices in mobile settings like
with the Surface Pro doing speech to text,
analyzing that transcription locally.
Think about an insurance agent
out in the field snapping a
picture of a damaged vehicle or
car, running an analysis of that.
We don't need to be walking around with
a GB10 or an RTX Spark Laptop Ultra.
We can actually do that with,
what we refer to as more of the
medium or large sized hardware
that we have available today.
And so to tie a bow on this, I think,
you know, for years it was a race to
build the best model, and then we were
building bigger models and running them
on these general purpose accelerators.
Now we're seeing the hardware being
optimized for these specific classes
of models, the T-shirt sizes, and
then delivering these corresponding
workloads to our customers
Chauncey: You kinda mentioned that you're
talking as if this is happening 'cause
I think like it is happening, right?
and I think based on my limited
knowledge that like it's actually
not that terribly difficult either.
Like it is obviously you gotta learn
something new, you gotta dev, you
gotta kind of get your hands dirty
on some code but that like-- That
the platform itself actually somewhat
exists today to be able to adopt this
relatively easy and that I would say
one of our hindrances of this kind of
exploding is just a lack of awareness.
People don't even know that this is
possible and there's only so many
Neils and Lauras that can get out
there and get in front of customers.
So do you feel like
that's-- is that the truth?
I mean are we kind of at the stage
where customers should just kind of be
looking for this or are we still kind
of maybe a few months, a few years
out before that technology catches up
Yeah, And it's some feedback that I
get quite a lot from customers is,
"Oh, it's actually way easier to
implement this stuff than we thought."
There's this kind of expectation that
there's gonna be a whole load of new
tooling required, new skill sets.
But in reality, you're using all
the same tooling that you used to.
You can embed these, models and local
solutions into, if you're developing
with VS Code, if you're using GitHub,
if you're using Copilot CLI, all of
these different, native toolings that
we have at our fingertips, you can
just embed local solutions within that.
And it is so simple, like the examples
that Neil gave are great, and I have the
same kind of, scenarios with customers.
That sometimes it is just something
really simple that they do every day, like
comparing two documents or two invoices
that they need to constantly do this.
They can't afford to have a subscription
for AI for everyone just to be
able to do that one simple task.
But also, it doesn't
make sense to do that.
You can literally just build one simple
tool that does that same solution, uses
local AI to do it, and everyone can
just roll it out on their own devices
Neil: Great call out, Laura.
And I'll say the messaging that's
resonating effectively with customers
for those building with Microsoft Foundry
is a set of large language models in
the cloud that Microsoft has vetted
and it's plug-and-play for developers.
We've now extended Microsoft Foundry
with the Foundry local toolkit.
This is a set of curated models
designed to run locally on the device.
So it's the same SDKs, right?
It's the same deployment and
management tools to get these
models out to your end users.
I will say that the only additional
tooling that some customers are
often adopting is Agent 365,
which is the ability to govern
and manage these solutions.
I met with a financial services
customer, and they said we now have
more agents than human employees,
and getting our arms wrapped around
what are all these agents doing?
Are all these agents authorized?
Should we consolidate some?
Should we promote some of these so
that there's not redundant solutions?
That's where I think the governance
and management and really Microsoft
secret sauce comes into play here.
It's one thing to spin up a local AI
instance using Llama CPP and get an agent
running for one person, for the hobbyist.
It's another thing to deliver
a consistent experience across
your end users, and that's where
we're seeing the pendulum shift.
The question is no longer can
this solution run locally?
It's how can I run this solution
securely in a local environment
Frankcx: So yeah, it's super exciting
that these pieces are coming together
that enterprises can actually start
to adopt and put into that space.
just to keep the show moving, that
brings us into the tension of the week.
so this week's tension, and I'm on
the hot seat, but I'm gonna just kinda
put it out there something a little
spicy, is that last week we talked
about the twenty-five dollar per
million token dying and cloud-based
models starting to reduce their costs.
But I would argue that whereas
the twenty-five d-dollar token
dies, that's when enterprises
and the AI hardware boom begins.
And it's not because
the cloud is going away.
it's not because every AI workload
suddenly belongs on somebody's desk
or inside their private server room.
My case is honestly very simple: open
weight models are becoming capable enough
to work and do real work, and enterprises
are increasingly demanding those weights.
We've seen that from Citi, we've
seen that across the board.
And these model designers, like what we're
seeing out of the UAE, are taking big
models that can work on hardware that is
deployed at different scale, from phone
to PC, to server, to, workstation class.
we're seeing hardware vendors
build specific hardware that
can essentially host local AI.
And Apple with Mac mini is showing
us that these early demands are
much stronger than they anticipated,
especially in the enterprise space.
so that's where this paradox comes.
If intelligence is getting cheaper we
are gonna consume, the thought would
be is we would consume more of it.
Are we gonna start inventing
more and more uses for AI?
more agents, more inference, more
workloads, and we would say that
which intelligence should we rent
and which intelligence should we own?
So I'd ask you to kind of prove me wrong
that, you know, or maybe just agree
along the way, but I've kinda opened it
up for discussion on the open debate.
If the cloud cost pennies, you
know, why would we buy hardware
that can essentially drive it?
But my argument would be there's
more now uses for hardware than ever,
and we would just find the right
place for that in the right time.
But opening up further discussion here
Chauncey: I think a lot of folks have,
strong opinions about cloud usage.
cloud does allow for a lot of scale.
Cloud allows for a lot of very quick
growth, and You can just do things at
that compute level in particular, that
like I think locally, I don't know that
we'll ever really be able to match.
Maybe get really close to, but not quite
exactly match, just because the sheer
amount of power that you can bring from
a data center or collective data centers.
however, local AI, on device AI
really provides, I think more comfort
than anything, like this sense of
control, this sense of ownership.
and when costs are such a variable thing,
especially with AI, because It is still
somewhat of an un-unknown factor, right?
I mean, especially as, models
change in price or we, end up like
as a new model comes out that's
expensive, but we get cheaper models.
Like there's, it's just
a lot of variability.
But if I can control, hey, I know for a
fact that my hardware is okay and that
maybe I do have to buy some kind of
model that's offline, I can control that.
That becomes a very, fixed cost.
and I think that is very
appetizing for a lot of people.
so I think even if the cloud does get down
to pennies, the sheer comfort of having
control over the local piece will be local
AI will continue to be such a big thing
Frankcx: I wanted to pose a question
to Laura on, in the UK and your region
have you started to notice that it's
easier to talk to customers around local
AI and open-weight models and such?
but more on the idea that hey there
was a first wave of AI that came out
that, you know, you only had the cloud.
But people were afraid to push their
IP and their sovereignty, into US
data centers or into sharing, their
intellectual property And now there's
this second wave of these open weight
models that's allowing it to be much
more acceptable and be able to be like
"Now I can actually my company take
part in this AI experience." i don't
know if you've seen that but I guess
I'd be interested in your opinion
Laura: Yeah, for sure.
it's really funny for me actually
coming from spending the last
five years as a cloud architect.
I've been telling everyone to throw
everything into the cloud, push it all
up there, everything's better up there.
And all of those reasons that you
just listed, like, being able to,
you know, have control, everything
being kind of centralized, it's…
Those are all the pushes and it was things
like move from, I can never get this
in the right order, CapEx to OpEx, OpEx
to CapEx, and it just feels like I've
literally done a 360 and I'm like, "Nope,
bring everything back down. Everything's
better on the device." but it really goes
in line with what you just said, Shaunsy,
in terms of we've been so much kind of
almost hate that AI gets, and so much of
that hate is tied to the idea of cloud
and tied to the negative aspects of cloud.
I think like really great PR for AI to be
able to run locally and say, "Actually,
you can get around all of those problems,"
but also the security side of it is such
a… Particularly in the UK and the kind
of Europe with the sovereignty issues
that we have, it's a huge driver and I
think it just answers so many questions.
You don't have to work through
all of these security challenges.
You don't have to think about
cloud space as well, which
in Europe is a huge problem.
Capacity is just at its limits.
and yes, it's funny.
It feels really ironic for me that,
yeah, I'm just slating cloud again now.
I promise I'm not actually doing that, but
yeah, it's definitely an interesting move
Neil: even when cloud is reduced to
pennies, I'll use the insurance example.
If an insurance claim agent is
filing 50 claims a day and we've
got 10,000 insurance agents doing
this, those pennies add up, right?
That's over a million dollars
over the course of a year.
So even these, lower cost workloads,
there's still economic benefit when it's
a repetitive task running it at the edge.
To Laura's point, there are
values beyond just cost savings.
In fact, we've identified eight forces
as to why a workload should run locally.
there's physics, there's economics,
there's regulatory reasons as to
why local is the preferred path.
The public sentiment here, Frank
shared an awesome example of his
conversation with an Uber driver.
I'll just share a brief anecdote of, an
encounter I had at a brewery recently.
I went up to order a beverage, and
there was a lively conversation around,
really this distaste for AI, and
the bartender turns to me and says,
"You're anti-AI, aren't you, bro?"
And to which I hesitated, and I said,
"Actually, I'm an AI solutions architect.
However, I'm focused on getting these
workloads running locally which offers
cost advantage, lower energy draw." And at
the end of my 30 second pitch, I still got
served a beverage. He said, "All right.
you're okay with me."
Laura: You are bad.
Neil: all that to say, I
think this public sen-…
Frankcx: data
Neil: I, I didn't say data center.
Exactly.
Exactly.
So although the emissions 25% increase on,
emissions year over year for Microsoft,
certainly, you know, there's some merit
in those environmental concerns and,
I, I, I believe Local AI offers a, a
path, an exciting path forward to the
future where we can deliver intelligence
back to Zuck's manifesto, right?
Intelligence for every person
across the globe, without creating
unnecessary emissions in the process
Laura: Yeah, there's, a data center
being built in the town that I
live in, and there's been uproar,
there's been riots with people from
an environmental standpoint saying,
you know, AI is ruining everything.
And I feel like the Homer Simpson meme
just like shrinking back into the bushes.
But hope that we can be on the
right side of history in that sense
that some saving to be made there
from a data center perspective
Frankcx: Well then maybe that's a
future show that we have coming up.
But let's get to the motion, the vote.
I had pitched out this idea that as
the $25 token dies and that's when
enterprise AI hardware boom begins.
So as things become cheaper does
that mean we start to use more of it?
And does that open up the gates for
AI hardware to really start to flow
into our commercial industries?
So let's go around the horn.
I will start with Laura.
what's your take on that?
Laura: All for it, yes.
I mean the cost saving is a huge driver.
This idea of tokenomics is becoming
part of my everyday vocabulary now
and I think… I was always a bit
of a cynic that I never expected AI
to stay as cheap as it was or the
kind of unlimited usage that we had.
and in some ways I kind of fear about
the same trend happening from a local
perspective but I definitely think
that it's the, right push that we've
had for customers to just be more
intentional with how they look at
where they're using AI and what they're
doing with it and bringing devices
into that fold and into that strategy
Frankcx: Awesome.
I love the vote.
Chauncey, what's your take on this things
get cheaper, we start to use more of it?
Chauncey: Yeah, I think so.
I hope, the economies of scale
kind of help back this movement.
Like, we kinda need it
to, let's say it that way.
Yes.
Frankcx: All right.
And then Neil, your last one
Neil: This is gonna be an anonymous
vote, for this group today.
If we draw parallels to the '90s
internet boom, it was all about
the servers, and then it was
quickly followed by the PC boom.
Draw parallels now to the AI
moment, and it was all about
the data centers, the GPU boom.
Now it's the AI endpoint boom and,
we have a front row seat to this.
So, exciting times ahead.
edge AI devices are here to stay, and
as hardware continues to increase in
capabilities, software, is quantized
or we're even seeing specific models
developed for specific hardware.
this AI endpoint boom is here to stay
Frankcx: And I would agree, so-- I'll go
right into saying yes, I agree with you,
so we're four for four but that leads us
into kind of like what's on our device
and we'll just finish up with that.
But I did take that thirty-six billion
parameter model, and ran it on the
Beast which is my dual 3090 device.
and I said: "Hey, w-what job do
you get?" Like as in we did that
job interview piece in the past.
And did it get the job?
Actually it does get the job.
it went up against, you know Qwen three
dot eight twenty-seven billion and
almost the exact same qualifying profile
as the test So really surprising how
equal the playing field is these days.
This is a brand new model specifically
not quantized but ran head to head
against a Qwen three dot eight billion…
twenty-seven billion perimeter and
it really came out and showed that
it could perform head to head.
It was giving out about sixty-eight
tokens per second versus forty on a
standard Qwen three dot eight model.Um,
and if I'm hiring it as a local
agent, it's a calling tools agent.
it's great for working through
long prompts and big jup-
continuous jobs and honestly K two
made a very strong case for it.
But there's a little bit of
a catch.Qwen three dot eight
fits onmy single 3090 card.
The K2 needed both of my 3090 cards
so obviously there's a little bit more
expense towards how thirty-six billion
parameters are spread across forty-eight
gigs of beverypowerfulintothatspace.So
anybodyelse outthereplayingon
theirdeviceand testingthings
Chauncey: I-I am on the, training wheels
while you guys are all flying spaceships.
and so I was just kind of messing
around with, Copilot CLI actually
running locally, so I was trying to
get it to run and trying to find the
right model for it to run locally on.
and I used Qwen actually
in this case, Qwen code.
Frankcx: man, it, is-- it was a shocker
that while it was doing this, I literally
asked it a question like, "Hey, like
why are you taking so long?" It's like
well, it's because my parameters are,
pretty limited and I can only do so much
Like you should basically expect about
a two hour response time, for this.
Laura: And so it was a pretty major eye
opening moment for me of like, ah okay.
That's, that's where I think being
able to really help our customers
and, and anyone really just start
to think about like what am I really
trying to do with this specific model?
Because maybe I'm actually not applying
the right thing or the wri-right workload.
Now, I will say once I was able to,
limit the amount of, response and
kind of massage it a little bit, I am
actually using it for coding right now.
it, it is helping at least do some
of the basic planning and some of
the basic stuff, as I'm, I'm kinda
tailoring this other app I'm building.
So it is capable.
is it as fast?
Definitely not.
Like I, you know, will type a, a rude
question in there and then I'll wait for
20 minutes for them to give me another
response but I think we kinda talked
about this on an earlier call which is
if I wanted to do that overnight, maybe
I just wanted to compile something over
night, like why not do that for free?
Frankcx: One of the benefits there
is I'm not costing anyone anything
right now, to be able to run this.
So it's huge
Frankcx (2): say, " Give me your
answer right now?" You know?
Frankcx: Yeah.
Yeah.
Neil: You're
Frankcx: mean-- And it just becomes
Neil: Take your time."
Frankcx: very, focused on-on
kind of driving the right
person for the right models.
but it… yeah, no, I would
maybe, rephrase that to be like,
are you seeing anything that's,
interesting that you've seen that's
been built for on-device as well?
Laura: I did have a cool conversation
with someone today actually who is,
building their own kind of edition of
the Microsoft Speaker Coach That they are
creating something that's almost like a
teleprompter that is wrapped in within
this, that would take your speaker notes,
it would have a look at the slides that
you've got, and it would notice once
you've started to list off some of those
topics that you have in your speaker
notes, it would make them disappear off
the screen and kind of leave you with
anything that you've missed out, anything
to kinda prompt you to, fill the gaps
in your talk track, whatever that is.
which I thought was a really interesting
usage, and that's obviously entirely
on device, but very much a kind of
low spec, device requirement as well.
It's very, very simple small language
models that that would run on.
So that was something that
got me excited this week
Frankcx (2): I feel like I might need
that for the podcast as we're talking
Laura: was
Frankcx (2): follow a script
Laura: are you gonna start building it?"
Frankcx (2): All right.
We'll bring it on.
All right, and Neil,
anything you're doing?
Neil (2): I believe there's, the
right local AI hardware with the
right local model for every user
Frankcx (2): All right.
Well, thank you all.
A big thank you to Laura for
joining us on this Friday night
of yours in the United Kingdom.
Laura: Thank
Frankcx (2): So,
Laura: me
Frankcx (2): You are open…
The door is open at any time for
you to come back and join us.
I'm gonna close up by saying, you
know, with the $25 token that we're
claiming is dying, know, this AI boom
around hardware may just be starting.
we might be seeing that Apple is
selling a lot of Macs for various
reasons, but Bean's started to be
adopted into the enterprise, and we
might be able to say that most of
this still belongs in the cloud, but I
think the direction is worth watching.
cheap AI doesn't kill local AI.
cheap AI is what finally makes
local AI into a category, and
that's the point of the show today.
if this has been worth your time,
a rating out on Apple Podcasts or
Spotify has generally helped us get
this show found, and you can find
everything at thelocalhost.show.
for Laura and Neil and Chauncey
and Robert and, Jacob, who are
off enjoying themselves today,
there is no place like 127.0.0.1.
So have a great weekend, all,
Laura: Have a great extended weekend all.
Thanks all
Neil (2): Happy Labor Day.
Thanks everyone
