the localhost:0002 | the $25 token is dead

Download MP3

Frankcx: Hey, everyone.

Last week, I said that AI conversations
were moving from cost to control.

Well, cost just crashed too.

The best AI is getting dramatically
cheaper, and the AI you can

run on your own device is also
getting dramatically better.

this week's question is simple: If being
smart gets cheap, are the big AI companies

still gonna charge a premium for?

From cost to control to what's left.

find out

Welcome back to The Local Host.

I'm Frank, and this week,
I am your local host again.

This is show colon 0002 port speak.

hey, let's do a little bit of around the
horn with, Jacob, Chauncey, and Robert.

And we know Neal is, on a
important call, so he will join us.

You know, customers come first, but
he'll, join us as soon as he's available.

So hey, Jacob, how are

Jacob: Hi, I'm doing great, Frank.

Yeah, I'm Jacob.

I work for Microsoft, so
that's an important disclosure.

I love telling stories about
on-device productivity, including AI.

Excited to be here.

Frankcx: Awesome.

Chauncey: Awesome.

Yep.

Chauncey Larson, tech nerd,
also worked for Microsoft

Frankcx: And our hockey fan from Florida.

Robert: It almost seems like an oxymoron.

I am a little bit of both.

And I'm Robert, and, I too am, employed
by Microsoft under full disclosure,

I'm gonna do my best to match Frank's,
electric energy on today's show

Chauncey: You are bringing the best
hat of the crew though, that's for sure

Frankcx: yeah, that, the
Copilot, "I am your father."

Chauncey: Love it

Frankcx: Hey, let's get into it.

It's been a wild week.

I swear every day I look at the news and I
see something else that's coming in around

AI and it's just an ever-evolving, piece.

But we have this section of the
show we call the rundown, which

is kind of news of the week.

Hopefully you're listening to this and
you understood that, as we're recording

this, it's been reported that NVIDIA,
is thinking of buying Hugging Face.

What's Hugging Face?

Well, Hugging Face is essentially where
all the AI models are stored, from

cloud-based models that can be deployed
up to a trillion parameters, all the way

down to what can run on a Copilot+ PC.

If you want to find it, there's a little
slider bar that says, "I have this much

memory," or, "This is the kind of device
I have," and it'll tell you essentially

the model that you can download.

It's literally where… It's
the App Store of, models.

And it's super interesting that NVIDIA,
I always say is essentially the company

that's selling the shovels for you to
go out and mine AI, is also the company

that's now becoming the claim where you
come in and tell, the AI where it is that

you've, found AI where you can go get it.

hey, Neil.

We were just talking about,
NVIDIA buying Hugging Face.

Let's, see if anybody has any
opinions towards this announcement

Jacob: The discourse online has
been pretty rampant about this.

Kind of felt like it was a matter of time.

It seems like, people are
generally positive to Nvidia.

I think it's a better option than
some of the other model providers, and

model creators because Nvidia has been
moving in an open source direction

already with some of their in-house
models, Nemotron and things like that.

So I think it's par for the course.

Probably a good thing, but that's my take.

Robert: Does it dilute their focus though?

I mean, I don't think
Hugging Face makes money,

Chauncey: Yet.

It doesn't make money yet

Jacob: I heard they
actually do make money.

They do cover their costs, but they
have a ton of investment from a

whole bunch of different areas, so
that's not their focus, I don't think

Chauncey: do they just make money?

'Cause that might be also
the other thing, right?

Frankcx: I think it's
a double down, right?

I mean, I think it's a double down into
the idea of what an open weight model

is and why it's important to NVIDIA.

I mean, NVIDIA is known as the company
that essentially builds cloud-based

AI, but they're also, with their
announcements and deployments of DGX

Spark and RTX Spark coming, you know,
the idea would be is that now you have

models that can run wherever you want
intelligence to be, whether that's on

your device or in your data center.

And Hugging Face is the tool that
deploys those different solutions.

So to me, it aligns really
well with what NVIDIA is doing.

But, you know, I think, time
will tell if they change

Jacob: Well I

mean, there's just a risk they
lose their users, like the

community forks and like…

Chauncey: Yeah

Jacob: I think they've just
got the users and the models,

and so that's like the market.

Neil, you were gonna say something?

Neil: I saw an analogy
that NVIDIA's chips, right?

These are the ovens, and Hugging Face,
they have the library of recipes,

and if you control the oven and
the recipes, you can start, really

delivering outsized value to the market.

I think it's also important to note that
this isn't, NVIDIA's first attempt, to,

Chauncey: Hmm

Neil: they actually invested, in
Hugging Face back in 2023, alongside

Salesforce, Google, Amazon, and IBM.

Last year, NVIDIA wanted to follow
on with another 500 million.

Hugging Face said, "No, we don't wanna
give outsized control to a single

investor." Potential hot take here would
be, did the, breach of Hugging Face,

lead potentially to this acquisition?

Frankcx: Oh

Neil: need, we do need support from

Hmm

to,

Chauncey: Interesting

Neil: some of the security and,
and maximize, the, the value

of this entity we've created"?

So couple talking points there.

This isn't, out of the blue.

NVIDIA actually has, quite a extensive
history, including being on Hugging

Face's cap table, three years ago

Robert: Does that become a
hotspot for the regulators if they

own the recipes and the ovens?

Chauncey: I mean there's still--
Yeah, but there's still a decent

amount of competition here.

Maybe not great competition, right?

But there are still other
platforms out there.

So what grounds do the hawks
have to stand on in this case?

Jacob: there's nothing preventing
those users from going to another

Chauncey: Yeah

Jacob: And that's why I think,

Chauncey: Or even starting one

Jacob: especially for things
like, uncensored models that

strip away some of the guardrails.

Like for that portion of the local AI
community, I think they will go elsewhere,

away from the corporate entities

Frankcx: Well, leaning into what Neil
was saying around they have a history

of doing this, I used to work for
this company called Silicon Graphics

that was making big time, you know,
graphics engines, and then NVIDIA

comes along and literally scales it
so that they give it to everybody.

So it's like they have a history
of finding out ways to scale their

business to bring everybody into
the market, and not just somebody

that can own a data center.

So I think that that's a big play for
an NVIDIA here is like their customer

isn't just the big giant data centers,
it's literally you and me, and, you know,

anybody that can afford a 3090 card that's
used or buy a new RTX Spark device, you

know, that's who they're playing for.

So I like it personally.

I think it helps our story, for local AI,

Robert: hopefully it'll

Chauncey: fingers

Robert: the prices of RAM down

Frankcx: Right.

Robert: Maybe they can

Chauncey: Mm-hmm.

Frankcx: You're right.

Yeah, I mean, that is an interesting…
Maybe that's a show coming up of like, is

local AI gonna take some of the heat and
pressure off a data center deployment?

You know, like, how do we balance that?

But there will always be, in my mind,
cloud AI, but, what level of it balances

in the future is up to be seen, So moving
on to the next topic of news that I'm

bringing to the table today is around
Qwen, which is our favorite Chinese model.

At least it's one of mine.

It's pretty capable.

It's got amazing intelligence.

It scales or quantizes,
if I'm saying that right.

But the interesting thing that
Qwen, it had released a model

3.8 a little over a week ago.

We benchmarked it.

But it also released last week or
this week this interesting model

called Qwen3.8-Flash.NEXT, which is
actually a preview towards Qwen4.

But it has some really interesting memory
capabilities, and I know Jacob, you've

played around a little bit with this, so
maybe you can kind of tell us what you

Jacob: I'm very, very passionate and
interested in just like experimentation,

'cause at this point I'm just learning,
trying to be a sponge as much as possible.

So the moment that it dropped, I had my
agents working on the best way to figure

out how to install it, 'cause I knew
that it was 170, 180 billion parameters,

and I was like, that seems like too much
for what I can fit on a consumer card.

it seems like it's just too big.

I've got that external 3090 that I
was thinking about running it on.

And, I realized that there's a lot that
I'm still learning, about this lookup

table, these Ngram tables that are
essentially like available for being

able to… Well, they just don't need
the same availability as the rest of the

model does, is the way I think of it.

They can use slower memory, they
can use system RAM, and I heard some

folks in the community talking about
being able to use SSD storage for it.

so that was what was cool for me, and
I had some success with running the

bulk of those model weights on SSD in
the Ngram tables so that it could look

up and then a small amount on VRAM

Robert: Now, is this an MOE model, Frank?

Is it a mixture of experts model

Jacob: it is an MoE model technically.

I think there are 6 billion
parameters that are active at a time.

I'll have to double-check.

And then there's 51 billion parameters
that are in that n-gram table, and then

there's another 120 are available that
should be on some fast RAM as well,

but it doesn't need to be VRAM because
they're not active all the time being MoE

Frankcx: Yeah, it feels like even though
you look at the size of the model and

you're like, "That can't run on my
device," the way it loads the portions

of the model that you need are kind
of in the chunks that you're asking

for, not like everything all at once.

So yeah, it's a,

Robert: Which is the whole
premise of a MoE model, right?

It's a combination of a bunch
of smaller models together,

Chauncey: Mmhmm

Robert: and it could pull
from any one of them.

there's a component that lives in
there that says, "Oh, let's use this

model for this task and this model
for that one," which is awesome.

It's super cool

Jacob: Yeah, I think
there's something like

Frankcx: that,

Jacob: 512 experts, and each one of those
experts is a certain size in this model.

Frankcx: Right

Jacob: I'm still learning,
but it's very interesting.

Frankcx: I tried to run it on the CPU
on my device, which is a Xeon-based

device, and it was getting like four
tokens per second, which is like

barely… Like, if you think of a
token as like three quarters of a word,

that's like somebody is slurring their
words, like they're having trouble

ran it on the 3090, it got up to
50 tokens per second, but because

it was just chunking along, it
had trouble finishing its tasks.

So, there's definitely some
tuning that has to be done.

I think in the end, what I found
is it's not built for the 3090.

It has this weird lookup table that's
26 gigs and, a 3090 only has 24 gigs

of VRAM, so it just doesn't fit right.

It'll be interesting when, RTX Spark or
DGX Spark or, other devices that ha- Like,

Neil, I'd be interested if you tested
it on a Mac that has unified memory, how

that would run, because it seems like it's
more geared towards that type of hardware

than it is something like a 3090 card.

Neil: Got my homework for next week, so

Frankcx: All right

Neil: received some feedback folks
liked the, our ability to translate

technical topics into business terms.

I think pausing for a moment on Mixture
of Experts, all the rage right now.

Maybe this analogy fits, maybe it doesn't.

I like to view Mixture of Experts as,
analogous to the Encyclopedia Britannica

bookshelf at my grandparents' house.

You've got

Frankcx: Oh, right

Neil: of all

Chauncey: All right

Neil: topics.

If you know

Frankcx: B through C-A

Neil: You don't need to open all
of the books at the same time.

So, just, again, wanna pause for a
moment, define Mixture of Experts.

Frankcx: love it

Neil: this is going to allow us to
leverage larger parameter models

without, running into some of the
memory bandwidth constraints that we've

seen with, previous, models before
Mixture of Experts became, top of mind

Robert: And I'll share
some stuff later on,

Chauncey: Oh, nice

Robert: this week with MoE models

Frankcx: Oh, nice.

Neil, given that you're talking about
unified memory and kind of memory

management, I think you've been
doing some stuff with Apple as well.

So what can you tell us about
some of the Mac stuff or the news

that you're excited about for Mac?

Neil: So big announcements
this week from the Apple world.

They announced M5, M6 processors, Mac
Minis, which became all the rage with the

OpenCL moment earlier this calendar year.

We're continuing to see Apple
double down on this local AI moment.

The twenty twenty-six was
supposed to be the year of agents.

I think one can make an argument that it's
becoming the year of local and hybrid AI.

Um, Apple continues to, um, innovate
at the silicon layer, uh, dropping

down to a two you look at some of the
models like, Kimi V3, right, and three

hundred and fifty gigs, you see almost
a direct parallel between the, the

type of silicon and, and these types of
MoE models that we're seeing released.

Ran some tests on an M3 Max,
versus the thirty ninety.

So Frank, thanks for sharing
those benchmark tests.

I, I think it boils down to
what are you trying to get

out of these local AI models?

Ran those five scenarios, and the
consistent takeaway was anything

that's bandwidth bound, memory bound,
Mac is outpacing the competition.

When it comes to raw compute, NVIDIA's,
you know, taking the cake, right?

Apple is two to six X slower.

So again, that's a, a couple generations
back, but you're seeing Apple continue

to double down on these, bandwidth,
memory bound constraints, and you're

seeing NVIDIA to, you know, continuing
to outpace the competition when

it comes to, to raw compute power.

So it'll be interesting to see the
shift in paradigm and, again, lots of,

excitement for Mac enthusiasts, this week.

Curious any other takes from the crew
here on, Apple's recent announcements?

Chauncey: I mean, I would say as a,
fellow hardware nerd and just tracking

these other guys, like, this is only just
beneficial to the entire ecosystem, right?

If Apple's going this way,
Windows has to follow.

Like, as an ecosystem, of course, if
you think about Surface and Dell and

Lenovo and everything else that's out
there, like, it's so critical that

we all continue to match this space.

So I'm always excited when new things
come out because that means that there's

gonna be new things on the horizon for
everyone else and we get some, I mean,

it's just gonna continue to get better.

Just imagine the amount of power we
have now and what we're able to achieve,

which all you guys are talking about
so far of, like, testing locally.

What we'll be able to do in
just a year or two years it's

just exponential from here out.

Robert: I

wouldn't exactly say
Windows is following, right?

But just to be clear, I mean, there's
some pretty amazing things that are

happening in the Windows ecosystem Look,
Neil and I have had these conversations.

It was just cloud, then it was cloud and
device, and now it's cloud, edge, and

device in sort of a three-tiered model.

once you get to the phone, you

know, bets are off

Frankcx: I think the thing that I see
is, look at what happened this past week.

Two things: NVIDIA looking at Hugging
Face, that's a double down on local

AI and open-weight models, and
then Apple, another multi-trillion

dollar company, doubles down on
local AI themselves by releasing

these Macs that are in that space.

So I would argue, and I'm not with
Microsoft anymore, but I would argue

that the governance that Microsoft is
gonna bring to these local AI spaces, I

think that's why a lot of people should
be excited about what RTX Spark could

bring, 'cause you're bringing this
capability of what CUDA and speeds are,

but you're bringing the governance of
what you need, because many enterprise

companies are scared to death of running,

Robert: And they should be, yeah.

Well, not Linux, but

Of running agents for sure

Chauncey: Yeah

Jacob: I was talking to a customer

Frankcx: being able to have that
governance that Microsoft can bring over

top of what local AI can, bring from an
intelligence standpoint, there's some

exciting things just around the corner

Jacob: I was talking to a company in Japan
this week, and they were pretty frank

with me, like: "Hey, right now we can't
run agents at all in our organization."

And they are very AI forward.

So AI forward, they actually have
a partnership now to run local

models on all of their laptops.

So they're running local models
everywhere, but those local models

are not oper-- They're-- it's all
just connecting to the cloud or to

specific outputs and not agented tasks.

So they're like: "We want to run like
Scout or things like OpenClaw or, you

know, a whole host of other agents."
And, I think there's a continued story

that, Windows will continue to add value
with Microsoft Execution Containers,

Agent 365, all of that goodness.

Frankcx: Well, you guys will
get a kick out of this story.

You may not know this, but I've been
playing around with doing a little bit

of Uber driving myself, just on like a,

Robert: In the sprinter?

Frankcx: I want, I wanna get down
into Cap Hill and see what the

cool kids are doing, so I'm driving
them around, dad in his Volvo.

I picked up this guy, and we had this
conversation, and I'm like: "Hey, what

do you do?" And he's like: "Oh I run
an aerospace company." I'm like: "Oh,

that's pretty cool." I was like: "Are
you guys using much AI?" He's like:

"I use it to write emails, but I can't
use it for engineering because we're

doing contracts with Boeing and the
defense side that I can't share the IP

of what we're doing with the cloud."

And I was like: "Well, do you know much
about local AI?" And he's like: "Well,

tell me." And I'm like, here's this
Uber driver telling him about local AI.

He used to be in this band called
Fifth Angel, and he was the lead

singer for it, and it's this Seattle
hair metal band that came out

Robert: may- maybe his next
band will be called Air Gap

Chauncey: Ooh, nice.

That's good.

Robert: Hair, you got a

Chauncey: That's good.

Robert: gap

Frankcx: gap.

Chauncey: Robert coming up
with the band names, man

Frankcx: Yeah.

Does anyone else have any
news they wanna share?

All right, if not, it feels
like we should do a baseline.

Baseline to me is like, Chauncey, I know
you're interested in making sure we're all

talking speak that everybody understands.

Maybe you can give us
a little bit of a clue

Chauncey: This is also referred to
as the section of the podcast where

it's like, I have no idea what you
guys are actually talking about.

Can you please explain it?

So can you do a couple things?

One, I know some of it.

Like what's the difference
when we're-- Well, first of

all, like what is quantization?

So you kind of mentioned
quantizing earlier.

Like, I think it's probably really
worth diving a little bit into the

details of what quantization means.

And then I would actually love
to know your thoughts on the

difference in cost per token and
how we're thinking about that too.

Frankcx: Yeah, I mean, I hear a lot of
people talking about, the $25 token or the

$6 token or, like, you know, what is that?

Nobody's paying $6 for a token.

Well, it comes down to
it's $6 per million tokens.

And what the funny part is that this past
week, you know, Grok, which I think is

Neil's favorite AI, like it's, I called
it last week your drunk uncle, but your

drunk uncle sobered up and became a pretty
capable, model at $6 per million tokens,

which honestly is like a fifth of the
cost of what, OpenAI and Claude charge

for their, frontier models at $25 a token.

But I think the point is, is that
because of the ability to start

running these models locally, it's
starting to drive down, you know…

And when Grok comes out with a $6 per
million, you know, token cost, that

really makes everybody else look at
it like, "Hey, what are you doing

there? And how come you're not charging
as much as we are?" Well, you know,

they're not the premier, you know,
AI, but they are showing themselves

to have just as much capability.

And Neil, maybe you can talk a little bit
about Grok, but it'd be like, you know,

it feels like there are models out there,
because they're all sharing algorithms,

they're all becoming kind of equally smart

Neil: Yeah, it's becoming more of
a commodity value-based discussion.

The days of unlimited
tokens are over, right?

As we mentioned, currently work at
Microsoft and we have set token limits.

And so continuing to route every prompt
through frontier models, Opus 5, GPT

Sol 5.6 no longer becomes practical.

So, shifting some of the
development needs to Groq 4.6

stretches my token budget further.

And if it can do, as good or close to
as good of a job, if I can get more

done with my existing token budget,
that's why, Groq, the drunken uncle,

that sobered up, is now becoming, one
of my favorite models to work with.

not all tokens are equal, dependent
upon which models you choose.

And I think Groq offers, great
value at this moment in time

Chauncey: I'm really interested
to see where we go with, model

choice and making it easier.

I mean, like the average person-- Your
developers, obviously, they're gonna

know, like, "Oh, I should probably
choose something cheaper," and whatnot.

But like the average person's just gonna
go, "Oh, this says it works better.

I'm gonna click on work better."
Right? And, but ultimately, like I

would say companies would value either
Microsoft or whoever making this kind

of orchestration layer that says, "Hey,
for sure, I'm building a PowerPoint,

I'm gonna have this level of model.

But if I'm just answering a simple
question, I should be able to like

dynamically switch to another model."

I think that's where I'm really
excited to see some of this happen.

And I know we have some auto
settings in, Cowork and in

Scout, but not quite there yet.

It's still a pre-selected handful
of models versus if it can really

start dynamically going across
whatever your organization says.

Like, I want only these
six models to ever run.

I think that's when we're gonna start
seeing some pretty cool things happen.

Robert: it goes down a whole
other layer that general

users don't understand, right?

Chauncey: Well, yeah.

Yeah

Robert: know, sonnet or whatever,
then there's like medium, max,

Chauncey: Yep.

Robert: extra large, extra max,

Frankcx: Right

Robert: big max,

Chauncey: Plus

Frankcx: wondering, to bring it
back to Chauncey's original question

around quantization, I'm wondering
who can take a shot at that.

And when you quantize
something, does it make it less

Jacob: Well I'll do my first pass, and
then I'll get corrected by Robert or

Neil, is probably gonna be my choice here.

The way I think of quants is I think
of it like any other compression.

Like you're making the model and
the weights take up less space,

a result, you are potentially
cutting some of the detail.

Just like when you have a JPEG
image that's compressed, you may

be cutting out some of the detail.

It's doing some calculations of,
"Hey, we're grouping these pixels

together. They look the same."
It's saving you space on disk.

It's saving you space in RAM.

And so if a model like was a seven billion
parameter model, like its full weights

would probably be like double that.

They would probably be like 14 gigs.

if you, compress it down to Q8,
then it might be seven gigs.

So it might cut the space in half, that's
kinda how I think about it, is like you

may be trading off some precision, you may
not be able to trust the outputs in the

same way, in a world where you have models
where you're running them to do a task,

and then you have another model checking
its work, then maybe it makes sense.

Chauncey: So

Robert: Yep,

Chauncey: it

Robert: The trade-off is accuracy.

How do you

Chauncey: Hmm

Robert: more, you know, tighter
without losing the accuracy?

And that's the trick that goes with it.

And look, we're even seeing
some models get down to one-bit

quantization, which is insane.

But that's where the technology is going.

But that, that's always, as I understand
it, that's the challenge is how do you

make it and use less, but still maintain
the accuracy and the credibility of the,

Chauncey: S-

Robert: the answers you get?

Chauncey: how do you choose then?

Like, if I have this option of a frontier
version that's gonna be obviously the most

powerful, the most intelligent, and like,
that's kind of like what I would think I

would want all the time versus something
that is more efficient, that is gonna be

smaller, that I can maybe run locally.

Like, if there is kind of that
more gut reaction of just like,

well, then how do I choose?

Like, what would be your
kind of answer then?

Jacob: I would say the highest quant
that you can the space that you have.

So like the benchmarks for like 3.8, 27
billion parameters that we talked about

last week, and, I think Robert, you might
have some more to share about this week.

Like those benchmarks were generally ran
at full precision, like no quant at all.

So you want it to be as precise
as possible, but there's a

trade-off in speed and efficiency
and what can physically fit.

That's my take.

Frankcx: Yep.

Well, getting to that, I think it's time
to go into the tension of the week, and

I know that Robert, you said you'd be
willing to be on the hot seat, so I'm

Robert: Why not, eh?

Give it a go.

Frankcx: it

Robert: I'm always anxious
to get into the fight.

So listen, we've been talking
about AI running it locally.

Now we can run it ourselves.

The models keep getting better.

it, just obviously saves money than
using the ones… But the big cloud

models keep getting cheaper as well.

so the big question is: where
are the AI companies eventually

gonna get their money?

What are they gonna charge for?

Wendell at Level 1 Tech
says the answer is speed.

Not smarter answers, just
the same answer, but faster.

the example I used, earlier off air
was, you know, I'm gonna order something

from Amazon, it'll come in two days.

If I want it in a day, I'm gonna
pay, two ninety-nine to get it.

So we're already starting
to see it, in the AI world.

If you wanna wait, you pay less.

If you want it right now, you pay more.

Same AI, just a different clock.

And so, you know, the twenty-five
dollar, token is dead, right?

I think within 18 months, the only
thing that frontier labs are-- really

gonna be able to sell is the speed,
and not so much of the intelligence.

I experienced some of this this
week, in loading, a thirty billion

parameter model, onto one of my
devices, and I was blown away, at the

speed and the accuracy and, just the
capability of it and how far it's come.

So, in plain English, soon being
smarter will not be enough to

charge more for, you're gonna have
to be faster, and that's my take.

Frankcx: Yeah.

I mean, you know, for me, I think
of it as like Neil's… He's Mr.

Analogy, so I'll bring in the analogy
of like, to me, like a cloud AI

is like hiring a consulting firm.

They can bring in a lot of people.

They can do a lot of work.

They can scale very well.

They're expensive.

They don't know you very well, so
they have to take the time to learn

you, you know, what you're doing.

if you think about local AI,
could literally go out and hire

an MBA from business school and
bring them in and train them.

They're smart.

They're just as smart as who you could
bring in from a consulting group,

but they just can't do as much work.

They don't have the same
volume that they can do.

You could hire more and more of
them, but then you have to take

on the responsibility of managing
them and taking care of them.

for me, it's this, type of analogy
that the AI is getting smart enough

that you can run it locally, and
then you're able to control privacy.

That person isn't gonna take that
information and share it with

some other consulting agency.

It's sovereign, meaning that,
it's controlled as to where

it goes and where it lives.

It's local, so if something happened,
like, there was a storm and, things got

knocked out, that consulting company
that, flies in every week would not be

able to get to you, but that business
person can come to work that day.

So there's a lot of
advantages of having local.

But that doesn't mean that the
cloud's ever gonna go away.

It's gonna be there for when you
need it and the tough problems.

But I would argue that AI
being balanced is the way we're

looking at AI in the future

Chauncey: To your point about, the
analogy too, going back to speed,

a bigger team of individuals may or
may not come to that answer faster.

We can maybe throw more people at
that project, so you get that ability

to get either a more cohesive answer
or a faster answer, because of

that ability to throw more at it.

Is that kinda the mindset here?

Maybe?

Frankcx: Yeah, I would agree

Jacob: I don't know if I believe the
assertion fully, Robert, because

Robert: not say it with enough conviction?

Chauncey: Ding, ding, let's get it on

Jacob: I feel like everyone's just trying
to acquire users, and so it's a race

to the bottom temporarily while they're
trying to prove that their models or

their values, is where the user should be.

it's to the point a couple weeks ago,
there's this Ox Alpha, like kind of

like a stealth model on OpenRouter,
so you can choose if you're, running

a cloud model via OpenRouter, you can
choose this Ox Alpha and it gives you a

anonymous model that you don't know what
it is, but it's free to run and use.

So you're, using a model, you don't know
what it is, you don't have any trust.

It's in the cloud, it's not
local, but, it's free to use.

And so they're trying to acquire

Robert: Model roulette
is what that's called.

Jacob: What's that?

Robert: That's called model roulette.

Jacob: Exactly.

Robert: Good Lord

Jacob: but everyone wants users, so, like,
they're gonna continue lowering the cost,

but I don't think that that's sustainable.

I'm not into the economics of it with
these companies, but, like, it just

doesn't feel like that's sustainable.

At some point, they're gonna, like,
entrench that they've got their

users and the prices go back up

Neil: I'll offer a
perspective in terms of speed.

Robert, the argument that speed is
going to be what these frontier labs

offer, how does that shift when we
get into asynchronous workflows?

If I can craft my prompt to build the
code I need overnight, why would I

pay a premium for a frontier model?

Why would I pay the outsource
consultant, leveraging the example?

Yes, maybe my MBAs take, much longer
to execute that task, but if I can set

them up proactively and achieve the
desired result, certainly under some

time constraint pressures, yeah, maybe I
need to outsource and execute a sprint.

Does speed become less of a factor
when we start talking about,

orchestration and autopilot workflows
that can accomplish work when

Robert: But is,

Neil: lid is shut?

Robert: is that a builder
perspective or is that a general

everyday user perspective?

Neil: I would say right now
it's a builder perspective.

I'm seeing AI from a knowledge
work perspective, a more one in

one out, a more sophisticated Ask
Jeeves, Google search, if you will.

Once we start talking asynchronous
workflows, which developers are, are

Robert: Yeah.

Yep

Jacob: Yeah

Neil: Right to that,

Jacob: You

Neil: think it's the developer
workflows where speed doesn't matter.

I'm used to crafting a prompt, letting my
device cook for 10 minutes, 15 minutes.

My benchmark test took an
hour and a half this morning,

Hmm Hmm

acceptable.

So,

Jacob: Yeah.

It's like,

Neil: may not matter for developer

Jacob: Frank, if you're tossing, one
of those tasks like, you know, build an

app, build whatever it is that you're
building these days, like your four tokens

per second example of that, 180 billion
parameter model that you're running.

Like, could theoretically just say,
"Hey, I'm gonna be gone for the

weekend," like build me something

Frankcx: Right

Jacob: come back and you've got something
that's smarter and better and free so

Frankcx: Especially these days when
you're saying, like, you're kicking off

seven different Claude Costa sessions
and you're like, "Well, just go to

Robert: Fleet.

You've got a fleet, yeah

Frankcx: You're like,
"Can you keep working on

Robert: Does,

Frankcx: please?"

Robert: Does this equation change based on
what we saw from NVIDIA and Hugging Face?

Now, maybe not in that exact
scenario, but thinking about

consolidation in the industry does
that change your on any of this?

Chauncey: Wait, so just kind
of rehashing that thought.

Does NVIDIA potentially acquiring Hugging
Face affect this speed versus like

They're only gonna

be building for speed

Robert: Yeah, but not that
specifically, but just in

general industry consolidation.

I mean, look at all the model
providers we have, the hardware

manufacturers, the data centers.

Like for example, Microsoft makes models,
build data centers, they provide services.

I mean, are we gonna see more of that?

And is that gonna affect your answer?

Frankcx: there's some tools out
there that are like these routers

now that can essentially look at…
They're like intelligence routers.

They look at the question that's being
asked, and then they determine what

model or what tool, this can run on.

And I feel like within an industry, if
you had your own ability to route to

local because you knew that job that
you were requesting based on the prompt

could be handled by a local AI or a
cheaper model, then you would, right?

So I feel like router, types of,
intelligence is gonna be involving

local AI in the future as well

Robert: And does

Chauncey: Hmm

Robert: equation change, I'm
gonna throw another curve ball

here, with a three-tier model?

If you have your device running a
local model, you have a server a

server or blade in your house that
can run close to a frontier model,

that change things even further?

Chauncey: from a revenue perspective, like

Robert: a, from a, what
are they gonna charge for?

Chauncey: Yeah.

Yeah, well that's, that's, yeah.

Yeah

Robert: I could probably get some
pretty serious speed off of it

Jacob: You're,

Chauncey: I mean it's, wonder
even, like, how we ended up making

a decision, if you think about
even analogous to Office, right?

At one point, we had Office running
locally, and, everyone questioned why the

heck we moved, we being Microsoft, moved
to going to a cloud subscription-based

model for, and now that that subscription
model is king for everyone, right?

That is not just, a Microsoft thing.

Subscription is king across the board.

And so how do we start looking
at models in that same fashion?

Are we reverting?

I don't think so.

I think there's still gonna be a play of
individuals that are gonna want access to

the newest, the latest, the greatest, the
fastest of what the cloud has to offer

and to that provides that scale, right?

Frankcx: Yeah.

Chauncey: where can the cloud grow?

It can grow

Frankcx: Yep.

No, I think that's another show,
because I feel like there's this

whole idea of, like, butts in seats
equals subscriptions times the number

of people you have, and that doesn't
equal an AI model when it comes to

Chauncey: Right

Frankcx: But, hey, let's, let's
do a little vote around the horn.

I'll jump in first.

So I'm gonna support, Robert's,
theory that the $25 token is, a

thing of the past, especially since
Grok and others are already leading.

I don't know what that means towards
the economy and all the investment

banking that takes place around AI.

Hopefully they can pay their bills,
but, but that's the way I feel about it.

Jacob?

Jacob: I think that it's gonna get
expensive again before it gets cheaper.

Local is here, we all think it, but I
think the market's not gonna be fully

there and the costs are gonna go up
again, especially as licenses, like open

source licenses are in the mix of like
what local models are able to do legally.

I think that some of the Chinese AI firms
are going to… Well, we're seeing them

add more, and not as open of licensing.

So you're gonna see more commercial
customers want to-- they'll be paying

more for it because it's in way
higher demand before it goes down.

So I don't agree with
you yet, Robert, yet.

Frankcx: All

Robert: On the side

Frankcx: Chauncey?

Chauncey: Yeah, I'm on
the same boat as Jacob.

Demand's gonna be high enough to continue
to maintain a higher price for now.

Robert: Must be in marketing

Frankcx: Yeah, they want the

Chauncey: Or, yeah.

Frankcx: And then, Neil

Neil: Speed matters when
we're talking security.

If we have bad actors using these
frontier models, we need access to

these frontier models to combat.

I'm going to agree with the stance
that frontier models deliver speed.

The use cases for speed vary.

I think there will continue to
be a premium for frontier models

specifically to combat rising threat
actors leveraging the same models, us

Jacob: And people will
pay for that, all right?

Frankcx: All right, Robert,
you get the last word.

Robert: I'm

Frankcx: know where your

Robert: I'm holding firm.

Now, what we might see is different
tiers of speed you pay for.

I think regardless of where it
goes, and I do think speed is gonna

be, the first, shoe to drop here.

Over the next six to 12 months, I think
we're gonna see some significant changes

in how this looks, is good for us
'cause it gives us stuff to talk about

Frankcx: Yep.

right, let's, wrap up with on my device.

Anybody have anything cool
that they're working on?

Robert: Here's the short story.

Frankcx: problem

Robert: week I talked about my, struggle
to try and find out how we can get,

on-device coding, a code agent running.

And I of tripped into it or
backed into it by mistake.

The short version of the story is
I loaded up a twenty-seven billion

parameter model, connected it in,
thank you Jacob, into Copilot CLI.

I typed in, "Let's make something,"
and then I got distracted.

I turned on autopilot, which essentially
gave it full, permission to do

whatever it wanted, and I walked away.

And I came back, and it said, " Well,
you didn't respond in ten minutes, so

I built you something." And it built an
online coding agent, and it said, " You

should use a-- this MoE thirty billion
parameter three point eight Qwen

model." and, and the rest is history.

There's a lot more to it, but it

Chauncey: Oh, that's awesome.

Frankcx: All right, gentlemen, I
hope you all have a terrific weekend.

So is no place like 127.0.0.1.

Robert: hang with you fellas

Frankcx: now

Chauncey: See you.

Jacob: Thanks, team

Chauncey: Thanks guys

Creators and Guests

person
Host
Chauncey Larsen
Technical Storyteller - Microsoft
person
Host
Frank Buchholz
"Retired" from Microsoft but still dancing in AI
person
Host
Jacob Rhoades
Technical Storyteller and Educator - Microsoft
person
Host
Neil Misak
Surface AI Factory - Microsoft
person
Host
Robert Henry
Surface AI Factory - Microsoft
the localhost:0002 | the $25 token is dead
Broadcast by