the localhost:0002 | the $25 token is dead
Download MP3Frankcx: Hey, everyone.
Last week, I said that AI conversations
were moving from cost to control.
Well, cost just crashed too.
The best AI is getting dramatically
cheaper, and the AI you can
run on your own device is also
getting dramatically better.
this week's question is simple: If being
smart gets cheap, are the big AI companies
still gonna charge a premium for?
From cost to control to what's left.
find out
Welcome back to The Local Host.
I'm Frank, and this week,
I am your local host again.
This is show colon 0002 port speak.
hey, let's do a little bit of around the
horn with, Jacob, Chauncey, and Robert.
And we know Neal is, on a
important call, so he will join us.
You know, customers come first, but
he'll, join us as soon as he's available.
So hey, Jacob, how are
Jacob: Hi, I'm doing great, Frank.
Yeah, I'm Jacob.
I work for Microsoft, so
that's an important disclosure.
I love telling stories about
on-device productivity, including AI.
Excited to be here.
Frankcx: Awesome.
Chauncey: Awesome.
Yep.
Chauncey Larson, tech nerd,
also worked for Microsoft
Frankcx: And our hockey fan from Florida.
Robert: It almost seems like an oxymoron.
I am a little bit of both.
And I'm Robert, and, I too am, employed
by Microsoft under full disclosure,
I'm gonna do my best to match Frank's,
electric energy on today's show
Chauncey: You are bringing the best
hat of the crew though, that's for sure
Frankcx: yeah, that, the
Copilot, "I am your father."
Chauncey: Love it
Frankcx: Hey, let's get into it.
It's been a wild week.
I swear every day I look at the news and I
see something else that's coming in around
AI and it's just an ever-evolving, piece.
But we have this section of the
show we call the rundown, which
is kind of news of the week.
Hopefully you're listening to this and
you understood that, as we're recording
this, it's been reported that NVIDIA,
is thinking of buying Hugging Face.
What's Hugging Face?
Well, Hugging Face is essentially where
all the AI models are stored, from
cloud-based models that can be deployed
up to a trillion parameters, all the way
down to what can run on a Copilot+ PC.
If you want to find it, there's a little
slider bar that says, "I have this much
memory," or, "This is the kind of device
I have," and it'll tell you essentially
the model that you can download.
It's literally where… It's
the App Store of, models.
And it's super interesting that NVIDIA,
I always say is essentially the company
that's selling the shovels for you to
go out and mine AI, is also the company
that's now becoming the claim where you
come in and tell, the AI where it is that
you've, found AI where you can go get it.
hey, Neil.
We were just talking about,
NVIDIA buying Hugging Face.
Let's, see if anybody has any
opinions towards this announcement
Jacob: The discourse online has
been pretty rampant about this.
Kind of felt like it was a matter of time.
It seems like, people are
generally positive to Nvidia.
I think it's a better option than
some of the other model providers, and
model creators because Nvidia has been
moving in an open source direction
already with some of their in-house
models, Nemotron and things like that.
So I think it's par for the course.
Probably a good thing, but that's my take.
Robert: Does it dilute their focus though?
I mean, I don't think
Hugging Face makes money,
Chauncey: Yet.
It doesn't make money yet
Jacob: I heard they
actually do make money.
They do cover their costs, but they
have a ton of investment from a
whole bunch of different areas, so
that's not their focus, I don't think
Chauncey: do they just make money?
'Cause that might be also
the other thing, right?
Frankcx: I think it's
a double down, right?
I mean, I think it's a double down into
the idea of what an open weight model
is and why it's important to NVIDIA.
I mean, NVIDIA is known as the company
that essentially builds cloud-based
AI, but they're also, with their
announcements and deployments of DGX
Spark and RTX Spark coming, you know,
the idea would be is that now you have
models that can run wherever you want
intelligence to be, whether that's on
your device or in your data center.
And Hugging Face is the tool that
deploys those different solutions.
So to me, it aligns really
well with what NVIDIA is doing.
But, you know, I think, time
will tell if they change
Jacob: Well I
mean, there's just a risk they
lose their users, like the
community forks and like…
Chauncey: Yeah
Jacob: I think they've just
got the users and the models,
and so that's like the market.
Neil, you were gonna say something?
Neil: I saw an analogy
that NVIDIA's chips, right?
These are the ovens, and Hugging Face,
they have the library of recipes,
and if you control the oven and
the recipes, you can start, really
delivering outsized value to the market.
I think it's also important to note that
this isn't, NVIDIA's first attempt, to,
Chauncey: Hmm
Neil: they actually invested, in
Hugging Face back in 2023, alongside
Salesforce, Google, Amazon, and IBM.
Last year, NVIDIA wanted to follow
on with another 500 million.
Hugging Face said, "No, we don't wanna
give outsized control to a single
investor." Potential hot take here would
be, did the, breach of Hugging Face,
lead potentially to this acquisition?
Frankcx: Oh
Neil: need, we do need support from
Hmm
to,
Chauncey: Interesting
Neil: some of the security and,
and maximize, the, the value
of this entity we've created"?
So couple talking points there.
This isn't, out of the blue.
NVIDIA actually has, quite a extensive
history, including being on Hugging
Face's cap table, three years ago
Robert: Does that become a
hotspot for the regulators if they
own the recipes and the ovens?
Chauncey: I mean there's still--
Yeah, but there's still a decent
amount of competition here.
Maybe not great competition, right?
But there are still other
platforms out there.
So what grounds do the hawks
have to stand on in this case?
Jacob: there's nothing preventing
those users from going to another
Chauncey: Yeah
Jacob: And that's why I think,
Chauncey: Or even starting one
Jacob: especially for things
like, uncensored models that
strip away some of the guardrails.
Like for that portion of the local AI
community, I think they will go elsewhere,
away from the corporate entities
Frankcx: Well, leaning into what Neil
was saying around they have a history
of doing this, I used to work for
this company called Silicon Graphics
that was making big time, you know,
graphics engines, and then NVIDIA
comes along and literally scales it
so that they give it to everybody.
So it's like they have a history
of finding out ways to scale their
business to bring everybody into
the market, and not just somebody
that can own a data center.
So I think that that's a big play for
an NVIDIA here is like their customer
isn't just the big giant data centers,
it's literally you and me, and, you know,
anybody that can afford a 3090 card that's
used or buy a new RTX Spark device, you
know, that's who they're playing for.
So I like it personally.
I think it helps our story, for local AI,
Robert: hopefully it'll
Chauncey: fingers
Robert: the prices of RAM down
Frankcx: Right.
Robert: Maybe they can
Chauncey: Mm-hmm.
Frankcx: You're right.
Yeah, I mean, that is an interesting…
Maybe that's a show coming up of like, is
local AI gonna take some of the heat and
pressure off a data center deployment?
You know, like, how do we balance that?
But there will always be, in my mind,
cloud AI, but, what level of it balances
in the future is up to be seen, So moving
on to the next topic of news that I'm
bringing to the table today is around
Qwen, which is our favorite Chinese model.
At least it's one of mine.
It's pretty capable.
It's got amazing intelligence.
It scales or quantizes,
if I'm saying that right.
But the interesting thing that
Qwen, it had released a model
3.8 a little over a week ago.
We benchmarked it.
But it also released last week or
this week this interesting model
called Qwen3.8-Flash.NEXT, which is
actually a preview towards Qwen4.
But it has some really interesting memory
capabilities, and I know Jacob, you've
played around a little bit with this, so
maybe you can kind of tell us what you
Jacob: I'm very, very passionate and
interested in just like experimentation,
'cause at this point I'm just learning,
trying to be a sponge as much as possible.
So the moment that it dropped, I had my
agents working on the best way to figure
out how to install it, 'cause I knew
that it was 170, 180 billion parameters,
and I was like, that seems like too much
for what I can fit on a consumer card.
it seems like it's just too big.
I've got that external 3090 that I
was thinking about running it on.
And, I realized that there's a lot that
I'm still learning, about this lookup
table, these Ngram tables that are
essentially like available for being
able to… Well, they just don't need
the same availability as the rest of the
model does, is the way I think of it.
They can use slower memory, they
can use system RAM, and I heard some
folks in the community talking about
being able to use SSD storage for it.
so that was what was cool for me, and
I had some success with running the
bulk of those model weights on SSD in
the Ngram tables so that it could look
up and then a small amount on VRAM
Robert: Now, is this an MOE model, Frank?
Is it a mixture of experts model
Jacob: it is an MoE model technically.
I think there are 6 billion
parameters that are active at a time.
I'll have to double-check.
And then there's 51 billion parameters
that are in that n-gram table, and then
there's another 120 are available that
should be on some fast RAM as well,
but it doesn't need to be VRAM because
they're not active all the time being MoE
Frankcx: Yeah, it feels like even though
you look at the size of the model and
you're like, "That can't run on my
device," the way it loads the portions
of the model that you need are kind
of in the chunks that you're asking
for, not like everything all at once.
So yeah, it's a,
Robert: Which is the whole
premise of a MoE model, right?
It's a combination of a bunch
of smaller models together,
Chauncey: Mmhmm
Robert: and it could pull
from any one of them.
there's a component that lives in
there that says, "Oh, let's use this
model for this task and this model
for that one," which is awesome.
It's super cool
Jacob: Yeah, I think
there's something like
Frankcx: that,
Jacob: 512 experts, and each one of those
experts is a certain size in this model.
Frankcx: Right
Jacob: I'm still learning,
but it's very interesting.
Frankcx: I tried to run it on the CPU
on my device, which is a Xeon-based
device, and it was getting like four
tokens per second, which is like
barely… Like, if you think of a
token as like three quarters of a word,
that's like somebody is slurring their
words, like they're having trouble
ran it on the 3090, it got up to
50 tokens per second, but because
it was just chunking along, it
had trouble finishing its tasks.
So, there's definitely some
tuning that has to be done.
I think in the end, what I found
is it's not built for the 3090.
It has this weird lookup table that's
26 gigs and, a 3090 only has 24 gigs
of VRAM, so it just doesn't fit right.
It'll be interesting when, RTX Spark or
DGX Spark or, other devices that ha- Like,
Neil, I'd be interested if you tested
it on a Mac that has unified memory, how
that would run, because it seems like it's
more geared towards that type of hardware
than it is something like a 3090 card.
Neil: Got my homework for next week, so
Frankcx: All right
Neil: received some feedback folks
liked the, our ability to translate
technical topics into business terms.
I think pausing for a moment on Mixture
of Experts, all the rage right now.
Maybe this analogy fits, maybe it doesn't.
I like to view Mixture of Experts as,
analogous to the Encyclopedia Britannica
bookshelf at my grandparents' house.
You've got
Frankcx: Oh, right
Neil: of all
Chauncey: All right
Neil: topics.
If you know
Frankcx: B through C-A
Neil: You don't need to open all
of the books at the same time.
So, just, again, wanna pause for a
moment, define Mixture of Experts.
Frankcx: love it
Neil: this is going to allow us to
leverage larger parameter models
without, running into some of the
memory bandwidth constraints that we've
seen with, previous, models before
Mixture of Experts became, top of mind
Robert: And I'll share
some stuff later on,
Chauncey: Oh, nice
Robert: this week with MoE models
Frankcx: Oh, nice.
Neil, given that you're talking about
unified memory and kind of memory
management, I think you've been
doing some stuff with Apple as well.
So what can you tell us about
some of the Mac stuff or the news
that you're excited about for Mac?
Neil: So big announcements
this week from the Apple world.
They announced M5, M6 processors, Mac
Minis, which became all the rage with the
OpenCL moment earlier this calendar year.
We're continuing to see Apple
double down on this local AI moment.
The twenty twenty-six was
supposed to be the year of agents.
I think one can make an argument that it's
becoming the year of local and hybrid AI.
Um, Apple continues to, um, innovate
at the silicon layer, uh, dropping
down to a two you look at some of the
models like, Kimi V3, right, and three
hundred and fifty gigs, you see almost
a direct parallel between the, the
type of silicon and, and these types of
MoE models that we're seeing released.
Ran some tests on an M3 Max,
versus the thirty ninety.
So Frank, thanks for sharing
those benchmark tests.
I, I think it boils down to
what are you trying to get
out of these local AI models?
Ran those five scenarios, and the
consistent takeaway was anything
that's bandwidth bound, memory bound,
Mac is outpacing the competition.
When it comes to raw compute, NVIDIA's,
you know, taking the cake, right?
Apple is two to six X slower.
So again, that's a, a couple generations
back, but you're seeing Apple continue
to double down on these, bandwidth,
memory bound constraints, and you're
seeing NVIDIA to, you know, continuing
to outpace the competition when
it comes to, to raw compute power.
So it'll be interesting to see the
shift in paradigm and, again, lots of,
excitement for Mac enthusiasts, this week.
Curious any other takes from the crew
here on, Apple's recent announcements?
Chauncey: I mean, I would say as a,
fellow hardware nerd and just tracking
these other guys, like, this is only just
beneficial to the entire ecosystem, right?
If Apple's going this way,
Windows has to follow.
Like, as an ecosystem, of course, if
you think about Surface and Dell and
Lenovo and everything else that's out
there, like, it's so critical that
we all continue to match this space.
So I'm always excited when new things
come out because that means that there's
gonna be new things on the horizon for
everyone else and we get some, I mean,
it's just gonna continue to get better.
Just imagine the amount of power we
have now and what we're able to achieve,
which all you guys are talking about
so far of, like, testing locally.
What we'll be able to do in
just a year or two years it's
just exponential from here out.
Robert: I
wouldn't exactly say
Windows is following, right?
But just to be clear, I mean, there's
some pretty amazing things that are
happening in the Windows ecosystem Look,
Neil and I have had these conversations.
It was just cloud, then it was cloud and
device, and now it's cloud, edge, and
device in sort of a three-tiered model.
once you get to the phone, you
know, bets are off
Frankcx: I think the thing that I see
is, look at what happened this past week.
Two things: NVIDIA looking at Hugging
Face, that's a double down on local
AI and open-weight models, and
then Apple, another multi-trillion
dollar company, doubles down on
local AI themselves by releasing
these Macs that are in that space.
So I would argue, and I'm not with
Microsoft anymore, but I would argue
that the governance that Microsoft is
gonna bring to these local AI spaces, I
think that's why a lot of people should
be excited about what RTX Spark could
bring, 'cause you're bringing this
capability of what CUDA and speeds are,
but you're bringing the governance of
what you need, because many enterprise
companies are scared to death of running,
Robert: And they should be, yeah.
Well, not Linux, but
Of running agents for sure
Chauncey: Yeah
Jacob: I was talking to a customer
Frankcx: being able to have that
governance that Microsoft can bring over
top of what local AI can, bring from an
intelligence standpoint, there's some
exciting things just around the corner
Jacob: I was talking to a company in Japan
this week, and they were pretty frank
with me, like: "Hey, right now we can't
run agents at all in our organization."
And they are very AI forward.
So AI forward, they actually have
a partnership now to run local
models on all of their laptops.
So they're running local models
everywhere, but those local models
are not oper-- They're-- it's all
just connecting to the cloud or to
specific outputs and not agented tasks.
So they're like: "We want to run like
Scout or things like OpenClaw or, you
know, a whole host of other agents."
And, I think there's a continued story
that, Windows will continue to add value
with Microsoft Execution Containers,
Agent 365, all of that goodness.
Frankcx: Well, you guys will
get a kick out of this story.
You may not know this, but I've been
playing around with doing a little bit
of Uber driving myself, just on like a,
Robert: In the sprinter?
Frankcx: I want, I wanna get down
into Cap Hill and see what the
cool kids are doing, so I'm driving
them around, dad in his Volvo.
I picked up this guy, and we had this
conversation, and I'm like: "Hey, what
do you do?" And he's like: "Oh I run
an aerospace company." I'm like: "Oh,
that's pretty cool." I was like: "Are
you guys using much AI?" He's like:
"I use it to write emails, but I can't
use it for engineering because we're
doing contracts with Boeing and the
defense side that I can't share the IP
of what we're doing with the cloud."
And I was like: "Well, do you know much
about local AI?" And he's like: "Well,
tell me." And I'm like, here's this
Uber driver telling him about local AI.
He used to be in this band called
Fifth Angel, and he was the lead
singer for it, and it's this Seattle
hair metal band that came out
Robert: may- maybe his next
band will be called Air Gap
Chauncey: Ooh, nice.
That's good.
Robert: Hair, you got a
Chauncey: That's good.
Robert: gap
Frankcx: gap.
Chauncey: Robert coming up
with the band names, man
Frankcx: Yeah.
Does anyone else have any
news they wanna share?
All right, if not, it feels
like we should do a baseline.
Baseline to me is like, Chauncey, I know
you're interested in making sure we're all
talking speak that everybody understands.
Maybe you can give us
a little bit of a clue
Chauncey: This is also referred to
as the section of the podcast where
it's like, I have no idea what you
guys are actually talking about.
Can you please explain it?
So can you do a couple things?
One, I know some of it.
Like what's the difference
when we're-- Well, first of
all, like what is quantization?
So you kind of mentioned
quantizing earlier.
Like, I think it's probably really
worth diving a little bit into the
details of what quantization means.
And then I would actually love
to know your thoughts on the
difference in cost per token and
how we're thinking about that too.
Frankcx: Yeah, I mean, I hear a lot of
people talking about, the $25 token or the
$6 token or, like, you know, what is that?
Nobody's paying $6 for a token.
Well, it comes down to
it's $6 per million tokens.
And what the funny part is that this past
week, you know, Grok, which I think is
Neil's favorite AI, like it's, I called
it last week your drunk uncle, but your
drunk uncle sobered up and became a pretty
capable, model at $6 per million tokens,
which honestly is like a fifth of the
cost of what, OpenAI and Claude charge
for their, frontier models at $25 a token.
But I think the point is, is that
because of the ability to start
running these models locally, it's
starting to drive down, you know…
And when Grok comes out with a $6 per
million, you know, token cost, that
really makes everybody else look at
it like, "Hey, what are you doing
there? And how come you're not charging
as much as we are?" Well, you know,
they're not the premier, you know,
AI, but they are showing themselves
to have just as much capability.
And Neil, maybe you can talk a little bit
about Grok, but it'd be like, you know,
it feels like there are models out there,
because they're all sharing algorithms,
they're all becoming kind of equally smart
Neil: Yeah, it's becoming more of
a commodity value-based discussion.
The days of unlimited
tokens are over, right?
As we mentioned, currently work at
Microsoft and we have set token limits.
And so continuing to route every prompt
through frontier models, Opus 5, GPT
Sol 5.6 no longer becomes practical.
So, shifting some of the
development needs to Groq 4.6
stretches my token budget further.
And if it can do, as good or close to
as good of a job, if I can get more
done with my existing token budget,
that's why, Groq, the drunken uncle,
that sobered up, is now becoming, one
of my favorite models to work with.
not all tokens are equal, dependent
upon which models you choose.
And I think Groq offers, great
value at this moment in time
Chauncey: I'm really interested
to see where we go with, model
choice and making it easier.
I mean, like the average person-- Your
developers, obviously, they're gonna
know, like, "Oh, I should probably
choose something cheaper," and whatnot.
But like the average person's just gonna
go, "Oh, this says it works better.
I'm gonna click on work better."
Right? And, but ultimately, like I
would say companies would value either
Microsoft or whoever making this kind
of orchestration layer that says, "Hey,
for sure, I'm building a PowerPoint,
I'm gonna have this level of model.
But if I'm just answering a simple
question, I should be able to like
dynamically switch to another model."
I think that's where I'm really
excited to see some of this happen.
And I know we have some auto
settings in, Cowork and in
Scout, but not quite there yet.
It's still a pre-selected handful
of models versus if it can really
start dynamically going across
whatever your organization says.
Like, I want only these
six models to ever run.
I think that's when we're gonna start
seeing some pretty cool things happen.
Robert: it goes down a whole
other layer that general
users don't understand, right?
Chauncey: Well, yeah.
Yeah
Robert: know, sonnet or whatever,
then there's like medium, max,
Chauncey: Yep.
Robert: extra large, extra max,
Frankcx: Right
Robert: big max,
Chauncey: Plus
Frankcx: wondering, to bring it
back to Chauncey's original question
around quantization, I'm wondering
who can take a shot at that.
And when you quantize
something, does it make it less
Jacob: Well I'll do my first pass, and
then I'll get corrected by Robert or
Neil, is probably gonna be my choice here.
The way I think of quants is I think
of it like any other compression.
Like you're making the model and
the weights take up less space,
a result, you are potentially
cutting some of the detail.
Just like when you have a JPEG
image that's compressed, you may
be cutting out some of the detail.
It's doing some calculations of,
"Hey, we're grouping these pixels
together. They look the same."
It's saving you space on disk.
It's saving you space in RAM.
And so if a model like was a seven billion
parameter model, like its full weights
would probably be like double that.
They would probably be like 14 gigs.
if you, compress it down to Q8,
then it might be seven gigs.
So it might cut the space in half, that's
kinda how I think about it, is like you
may be trading off some precision, you may
not be able to trust the outputs in the
same way, in a world where you have models
where you're running them to do a task,
and then you have another model checking
its work, then maybe it makes sense.
Chauncey: So
Robert: Yep,
Chauncey: it
Robert: The trade-off is accuracy.
How do you
Chauncey: Hmm
Robert: more, you know, tighter
without losing the accuracy?
And that's the trick that goes with it.
And look, we're even seeing
some models get down to one-bit
quantization, which is insane.
But that's where the technology is going.
But that, that's always, as I understand
it, that's the challenge is how do you
make it and use less, but still maintain
the accuracy and the credibility of the,
Chauncey: S-
Robert: the answers you get?
Chauncey: how do you choose then?
Like, if I have this option of a frontier
version that's gonna be obviously the most
powerful, the most intelligent, and like,
that's kind of like what I would think I
would want all the time versus something
that is more efficient, that is gonna be
smaller, that I can maybe run locally.
Like, if there is kind of that
more gut reaction of just like,
well, then how do I choose?
Like, what would be your
kind of answer then?
Jacob: I would say the highest quant
that you can the space that you have.
So like the benchmarks for like 3.8, 27
billion parameters that we talked about
last week, and, I think Robert, you might
have some more to share about this week.
Like those benchmarks were generally ran
at full precision, like no quant at all.
So you want it to be as precise
as possible, but there's a
trade-off in speed and efficiency
and what can physically fit.
That's my take.
Frankcx: Yep.
Well, getting to that, I think it's time
to go into the tension of the week, and
I know that Robert, you said you'd be
willing to be on the hot seat, so I'm
Robert: Why not, eh?
Give it a go.
Frankcx: it
Robert: I'm always anxious
to get into the fight.
So listen, we've been talking
about AI running it locally.
Now we can run it ourselves.
The models keep getting better.
it, just obviously saves money than
using the ones… But the big cloud
models keep getting cheaper as well.
so the big question is: where
are the AI companies eventually
gonna get their money?
What are they gonna charge for?
Wendell at Level 1 Tech
says the answer is speed.
Not smarter answers, just
the same answer, but faster.
the example I used, earlier off air
was, you know, I'm gonna order something
from Amazon, it'll come in two days.
If I want it in a day, I'm gonna
pay, two ninety-nine to get it.
So we're already starting
to see it, in the AI world.
If you wanna wait, you pay less.
If you want it right now, you pay more.
Same AI, just a different clock.
And so, you know, the twenty-five
dollar, token is dead, right?
I think within 18 months, the only
thing that frontier labs are-- really
gonna be able to sell is the speed,
and not so much of the intelligence.
I experienced some of this this
week, in loading, a thirty billion
parameter model, onto one of my
devices, and I was blown away, at the
speed and the accuracy and, just the
capability of it and how far it's come.
So, in plain English, soon being
smarter will not be enough to
charge more for, you're gonna have
to be faster, and that's my take.
Frankcx: Yeah.
I mean, you know, for me, I think
of it as like Neil's… He's Mr.
Analogy, so I'll bring in the analogy
of like, to me, like a cloud AI
is like hiring a consulting firm.
They can bring in a lot of people.
They can do a lot of work.
They can scale very well.
They're expensive.
They don't know you very well, so
they have to take the time to learn
you, you know, what you're doing.
if you think about local AI,
could literally go out and hire
an MBA from business school and
bring them in and train them.
They're smart.
They're just as smart as who you could
bring in from a consulting group,
but they just can't do as much work.
They don't have the same
volume that they can do.
You could hire more and more of
them, but then you have to take
on the responsibility of managing
them and taking care of them.
for me, it's this, type of analogy
that the AI is getting smart enough
that you can run it locally, and
then you're able to control privacy.
That person isn't gonna take that
information and share it with
some other consulting agency.
It's sovereign, meaning that,
it's controlled as to where
it goes and where it lives.
It's local, so if something happened,
like, there was a storm and, things got
knocked out, that consulting company
that, flies in every week would not be
able to get to you, but that business
person can come to work that day.
So there's a lot of
advantages of having local.
But that doesn't mean that the
cloud's ever gonna go away.
It's gonna be there for when you
need it and the tough problems.
But I would argue that AI
being balanced is the way we're
looking at AI in the future
Chauncey: To your point about, the
analogy too, going back to speed,
a bigger team of individuals may or
may not come to that answer faster.
We can maybe throw more people at
that project, so you get that ability
to get either a more cohesive answer
or a faster answer, because of
that ability to throw more at it.
Is that kinda the mindset here?
Maybe?
Frankcx: Yeah, I would agree
Jacob: I don't know if I believe the
assertion fully, Robert, because
Robert: not say it with enough conviction?
Chauncey: Ding, ding, let's get it on
Jacob: I feel like everyone's just trying
to acquire users, and so it's a race
to the bottom temporarily while they're
trying to prove that their models or
their values, is where the user should be.
it's to the point a couple weeks ago,
there's this Ox Alpha, like kind of
like a stealth model on OpenRouter,
so you can choose if you're, running
a cloud model via OpenRouter, you can
choose this Ox Alpha and it gives you a
anonymous model that you don't know what
it is, but it's free to run and use.
So you're, using a model, you don't know
what it is, you don't have any trust.
It's in the cloud, it's not
local, but, it's free to use.
And so they're trying to acquire
Robert: Model roulette
is what that's called.
Jacob: What's that?
Robert: That's called model roulette.
Jacob: Exactly.
Robert: Good Lord
Jacob: but everyone wants users, so, like,
they're gonna continue lowering the cost,
but I don't think that that's sustainable.
I'm not into the economics of it with
these companies, but, like, it just
doesn't feel like that's sustainable.
At some point, they're gonna, like,
entrench that they've got their
users and the prices go back up
Neil: I'll offer a
perspective in terms of speed.
Robert, the argument that speed is
going to be what these frontier labs
offer, how does that shift when we
get into asynchronous workflows?
If I can craft my prompt to build the
code I need overnight, why would I
pay a premium for a frontier model?
Why would I pay the outsource
consultant, leveraging the example?
Yes, maybe my MBAs take, much longer
to execute that task, but if I can set
them up proactively and achieve the
desired result, certainly under some
time constraint pressures, yeah, maybe I
need to outsource and execute a sprint.
Does speed become less of a factor
when we start talking about,
orchestration and autopilot workflows
that can accomplish work when
Robert: But is,
Neil: lid is shut?
Robert: is that a builder
perspective or is that a general
everyday user perspective?
Neil: I would say right now
it's a builder perspective.
I'm seeing AI from a knowledge
work perspective, a more one in
one out, a more sophisticated Ask
Jeeves, Google search, if you will.
Once we start talking asynchronous
workflows, which developers are, are
Robert: Yeah.
Yep
Jacob: Yeah
Neil: Right to that,
Jacob: You
Neil: think it's the developer
workflows where speed doesn't matter.
I'm used to crafting a prompt, letting my
device cook for 10 minutes, 15 minutes.
My benchmark test took an
hour and a half this morning,
Hmm Hmm
acceptable.
So,
Jacob: Yeah.
It's like,
Neil: may not matter for developer
Jacob: Frank, if you're tossing, one
of those tasks like, you know, build an
app, build whatever it is that you're
building these days, like your four tokens
per second example of that, 180 billion
parameter model that you're running.
Like, could theoretically just say,
"Hey, I'm gonna be gone for the
weekend," like build me something
Frankcx: Right
Jacob: come back and you've got something
that's smarter and better and free so
Frankcx: Especially these days when
you're saying, like, you're kicking off
seven different Claude Costa sessions
and you're like, "Well, just go to
Robert: Fleet.
You've got a fleet, yeah
Frankcx: You're like,
"Can you keep working on
Robert: Does,
Frankcx: please?"
Robert: Does this equation change based on
what we saw from NVIDIA and Hugging Face?
Now, maybe not in that exact
scenario, but thinking about
consolidation in the industry does
that change your on any of this?
Chauncey: Wait, so just kind
of rehashing that thought.
Does NVIDIA potentially acquiring Hugging
Face affect this speed versus like
They're only gonna
be building for speed
Robert: Yeah, but not that
specifically, but just in
general industry consolidation.
I mean, look at all the model
providers we have, the hardware
manufacturers, the data centers.
Like for example, Microsoft makes models,
build data centers, they provide services.
I mean, are we gonna see more of that?
And is that gonna affect your answer?
Frankcx: there's some tools out
there that are like these routers
now that can essentially look at…
They're like intelligence routers.
They look at the question that's being
asked, and then they determine what
model or what tool, this can run on.
And I feel like within an industry, if
you had your own ability to route to
local because you knew that job that
you were requesting based on the prompt
could be handled by a local AI or a
cheaper model, then you would, right?
So I feel like router, types of,
intelligence is gonna be involving
local AI in the future as well
Robert: And does
Chauncey: Hmm
Robert: equation change, I'm
gonna throw another curve ball
here, with a three-tier model?
If you have your device running a
local model, you have a server a
server or blade in your house that
can run close to a frontier model,
that change things even further?
Chauncey: from a revenue perspective, like
Robert: a, from a, what
are they gonna charge for?
Chauncey: Yeah.
Yeah, well that's, that's, yeah.
Yeah
Robert: I could probably get some
pretty serious speed off of it
Jacob: You're,
Chauncey: I mean it's, wonder
even, like, how we ended up making
a decision, if you think about
even analogous to Office, right?
At one point, we had Office running
locally, and, everyone questioned why the
heck we moved, we being Microsoft, moved
to going to a cloud subscription-based
model for, and now that that subscription
model is king for everyone, right?
That is not just, a Microsoft thing.
Subscription is king across the board.
And so how do we start looking
at models in that same fashion?
Are we reverting?
I don't think so.
I think there's still gonna be a play of
individuals that are gonna want access to
the newest, the latest, the greatest, the
fastest of what the cloud has to offer
and to that provides that scale, right?
Frankcx: Yeah.
Chauncey: where can the cloud grow?
It can grow
Frankcx: Yep.
No, I think that's another show,
because I feel like there's this
whole idea of, like, butts in seats
equals subscriptions times the number
of people you have, and that doesn't
equal an AI model when it comes to
Chauncey: Right
Frankcx: But, hey, let's, let's
do a little vote around the horn.
I'll jump in first.
So I'm gonna support, Robert's,
theory that the $25 token is, a
thing of the past, especially since
Grok and others are already leading.
I don't know what that means towards
the economy and all the investment
banking that takes place around AI.
Hopefully they can pay their bills,
but, but that's the way I feel about it.
Jacob?
Jacob: I think that it's gonna get
expensive again before it gets cheaper.
Local is here, we all think it, but I
think the market's not gonna be fully
there and the costs are gonna go up
again, especially as licenses, like open
source licenses are in the mix of like
what local models are able to do legally.
I think that some of the Chinese AI firms
are going to… Well, we're seeing them
add more, and not as open of licensing.
So you're gonna see more commercial
customers want to-- they'll be paying
more for it because it's in way
higher demand before it goes down.
So I don't agree with
you yet, Robert, yet.
Frankcx: All
Robert: On the side
Frankcx: Chauncey?
Chauncey: Yeah, I'm on
the same boat as Jacob.
Demand's gonna be high enough to continue
to maintain a higher price for now.
Robert: Must be in marketing
Frankcx: Yeah, they want the
Chauncey: Or, yeah.
Frankcx: And then, Neil
Neil: Speed matters when
we're talking security.
If we have bad actors using these
frontier models, we need access to
these frontier models to combat.
I'm going to agree with the stance
that frontier models deliver speed.
The use cases for speed vary.
I think there will continue to
be a premium for frontier models
specifically to combat rising threat
actors leveraging the same models, us
Jacob: And people will
pay for that, all right?
Frankcx: All right, Robert,
you get the last word.
Robert: I'm
Frankcx: know where your
Robert: I'm holding firm.
Now, what we might see is different
tiers of speed you pay for.
I think regardless of where it
goes, and I do think speed is gonna
be, the first, shoe to drop here.
Over the next six to 12 months, I think
we're gonna see some significant changes
in how this looks, is good for us
'cause it gives us stuff to talk about
Frankcx: Yep.
right, let's, wrap up with on my device.
Anybody have anything cool
that they're working on?
Robert: Here's the short story.
Frankcx: problem
Robert: week I talked about my, struggle
to try and find out how we can get,
on-device coding, a code agent running.
And I of tripped into it or
backed into it by mistake.
The short version of the story is
I loaded up a twenty-seven billion
parameter model, connected it in,
thank you Jacob, into Copilot CLI.
I typed in, "Let's make something,"
and then I got distracted.
I turned on autopilot, which essentially
gave it full, permission to do
whatever it wanted, and I walked away.
And I came back, and it said, " Well,
you didn't respond in ten minutes, so
I built you something." And it built an
online coding agent, and it said, " You
should use a-- this MoE thirty billion
parameter three point eight Qwen
model." and, and the rest is history.
There's a lot more to it, but it
Chauncey: Oh, that's awesome.
Frankcx: All right, gentlemen, I
hope you all have a terrific weekend.
So is no place like 127.0.0.1.
Robert: hang with you fellas
Frankcx: now
Chauncey: See you.
Jacob: Thanks, team
Chauncey: Thanks guys
