AD
Episode
517
Interview
Web News

AI’s Biggest Problem Isn’t Intelligence Anymore

Recorded:
September 23, 2026
Released:
September 26, 2026
Episode Number:
517

AI models are already remarkably capable, but their intelligence doesn’t matter much if they’re too expensive, limited, or confusing to use. In this edition of the Web News, Matt and Mike discuss why cheaper and faster models, higher usage limits, and a better user experience could be more important than winning the next benchmark - and why the future of AI may be about making today’s technology genuinely usable.

Listen

Also available on...
...and many more, check your podcast app!

Who’s in This Episode?

Show Notes

Is more intelligence all we want from our AI? For the past few years it seems that chasing more intelligence has been the primary goal of AI companies and for good reason - they've been cutting down on hallucinations and increasing their usability quite dramatically model over model. Now that we have some seriously useful AI at the consumer level, is it time to finally start to focus on UX? 

Questions/Topics to Discuss & Resolve

  • Does AI need more intelligence? Or is it time for better UX? 
  • Is it time to make AI more consumer friendly? 


Transcript

This transcript is machine generated, there may be errors.

[00:00:00]

Mikhail: All right. Welcome to the Web News. So today's gonna be a little bit different. Uh, there were some pretty big AI releases, so I will start off with that. But really what I wanna talk about in this episode is About how AI isn't just about intelligence anymore, in my opinion, right? Like, I, I think there's more to it now, now that the new models have gotten solid.

They're not... Like, you know, some of them are smarter, some of them are dumber. It's fine. But they've gotten to a point where it's almost like diminishing returns at the top end, unless you're doing something very crazy. Uh, I'm starting to focus more on other things outside of intelligence in AI models now, and that's what I wanna talk about in this episode.

But before we get down there, uh, I wa- I wanna kind of preface this with the n- the, the latest releases and what's prompted me to talk about this. L- This week there was a day, I think it was yesterday, actually. Yeah, yesterday, which was Tuesday, S- September 22nd. There was three new model releases at the same time.

Uh, Opus 5.5, which is, like, the latest [00:01:00] generation Claude model, and then there was two GPT OpenAI releases. Uh, they released Sol 6 and Luna 6. And the big theme with all of these releases isn't, like, they're pushing the intelligence boundary. What they're doing with these releases is they're making these more usable.

They're, they're making them slightly cheaper, they're making them slightly faster and, and in some cases actually significantly cheaper. So for example, uh, the GPT re- release, the GPT-6 Sol release and the GPT-6 Luna release, they've actually cut both the price, the token price in half. So it's, that's a pretty significant drop.

Um, but they haven't touted that they've increased intelligence by much. In some cases, actually, in some benchmarks it's showing that GPT-6 Sol is actually slightly below 5.5. In someone, in some of them are b- uh, it's higher but, you know, w- we're seeing it kind of as almost the same level of intelligence.

But decreased the pricing in half, and also it, it... they've also decreased the [00:02:00] price per task by a significant portion, too. And that really is the cusp of it as well is, like, some of these models might be cheap. Like a DeepSeek model might be cheap, but its cost per task sometimes is actually very high because to get to the same output as an, a more intelligent model, it needs to do way more work.

It needs to think for way longer. It needs to check 15 different things. It goes down different rabbit holes, and then goes back and goes down and like... So it uses more tokens per task than a high-end model which can get to the solution a lot faster. So that's what they're kind of touting in this one.

Opus 5.5 is a little bit different, where its intelligence is actually quite advanced, more advanced than Opus 5, but it's still in the same realm of, like, a Fable or an Astra, so it's not the best model, right? Like, it's not like all of a sudden the best no matter what model. It's just, like, better intelligence and they've reduced cost by a little bit.

They've given us more usage. So again, it, it makes the subscriptions, the Claude subscriptions, [00:03:00] the GPT subscriptions, the Coda subscriptions more usable. And I've really noticed- That being more important for me. And for, like, Matt, I'm just curious about you. You, you've, you've just started to dive into the agentic stuff over the past few weeks, I would say, like, you know, doing more and more agent work.

I wonder if you've run into any limits yet. Have you, have you hit any, you know, top of the line limits? Do you see yourself worrying about that yet?

Matt: Uh, so full disclosure, I have-- I'm on like the ChatGPT like $20

Mikhail: Угу

Matt: plan. Um, and I don't agentic code or agentic... I don't use an AI agent all day, but I do use it pretty frequently, and I've only hit the limit once, and the limit was hit when Astra came out and I was using Astra 'cause it defaulted to Astra for me for just all the normal tasks I would, I would normally give it. And then I kinda had a brief chat with you, and then I think you even mentioned it in one of our podcast recordings that, you know, maybe you wanna go down [00:04:00] to a different level, a different, uh, model. And so I just went back to what I was on, which was, I forget now, 5.6 or whatever it was. And then, uh, 'cause like the tasks I was doing was a part of a, a greater project, and so the, the tasks themselves hadn't gained complexity to the point where like Astra was desperately needed or something like that. And so I've, I switched back quite a few of the agentic tasks I give it to 5.6, and I haven't hit the limit again. And again, I'm only on the $20 or $25 a month plan. So I'm, I'm not, I'm not giving it tasks that take more than... Like a lot of my tasks are sub five minutes, at most 15 minutes. I'm not giving it stuff that runs all day, runs all night, checks a bunch of stuff, goes into a server, and like flips around and does this. I'm not doing that to, to be clear. I still do a fair bit of stuff manually or, uh, the way I kinda work is I'm like, "I wanna do this, but I also want this done, so I'll get [00:05:00] the AI to do this while I do this," and then I check in with it and then, "Okay, now let's do another two tasks," and that's kind of how I, I work with it.

Now I'm sure that'll change over time 'cause my workflow changes all the time, but that's been w- that's what's been working for me. And I mean, we, like even before this episode, we went through kind of like what we had done in the last two weeks, so 10 business days, and we've actually done quite a bit. And, you know, part of that's 'cause of the agents, and part of that's 'cause we've been just kinda hitting the ground, but hitting the ground running. But, uh, that's where I'm at with that. I've only ever hit the limit

Mikhail: No. But I, I-- okay, so this is a good example though, because the Astra limit is exactly where most of us have sat for a little while now. So Astra is the top-of-the-line model, and on, on Claude's side, it's Fable. So those models are super intelligent, but to be honest, unless you have infinite money, they're not that useful.

And that's really what this, like, kind of like the cusp of [00:06:00] everything that I'm talking about, because I've, I've done the same thing where I started in a similar way to you, where I was just kind of like sending small tasks to a- to agents. They would run for anywhere between one to five minutes, and that was how I was using them.

And with that kind of workflow, I was never really running, running into any sort of usage problems. Now, I was on the higher plan, and I was using it quite heavily in that sense, but I was never running into usage problems. As I've trusted agents more and more, I've given them longer and longer-running tasks, and I think that's a good natural progression, 'cause then you under- start to understand what they can and can't do.

Once I started trusting them, the tasks became stuff that like, "Hey, I need you to implement this. I need you to plan this implementation, double-check it with this top-end model like Astra, make sure that the implementation makes sense, implement it with this model, like the, the Sol 5.6 model, right? Then maybe kick off a couple of the lower-cost workers to do another, like, smaller implementation.

Review it with the [00:07:00] top-end model. Once you're happy with the review and the cycle has completed a couple times and you're happy with everything, deploy it here. Here's, you know, a way to deploy it, uh, safely. Then monitor the deployment here." Like, I'll give it this, like, really long-running thing because I've s- I've, I've done all of those steps individually on my own with those, like, kind of five to 10-minute tasks that now I can combine them all together, and that's where my usage just died.

Like, I just burned through. And even though that, that workflow worked, uh, really well for me, I was really struggling, uh, due to the fact that it just burned through my usage. So it wasn't-- It was great, it was intelligent, but it wasn't really usable. So I had to kind of almost revert back to my older, the older way of, like, doing smaller things just to save some usage and wait for my reset to hit or whatever, and I hated that.

Like that, that's a bad user experience. And with these new model releases that are coming out and the cheaper open source models that are coming out, I think that that's where the next evolution is going to be. We've [00:08:00] talked-- Or n- not we, but the, the AI companies have now come out and said, like, "Hey, we need to slow down frontier development," right?

Because the intelligence is getting so high, like we're, we're having these weird cases of, you know, them hacking other, uh, you know, external things. It's becoming dangerous. Whether that's marketing speak or not, I don't know. I, I kind of-- I, I'm leaning on the side of them being actually cautious rather than it just being marketing at this point.

Uh, but regardless- It seems that, that most of these companies are taking it seriously. So I think what they're focusing on now is let's make what we have now really usable, and I'm here for it. Because now with these new models, I, I tested them yesterday, I haven't run into that issue. I can get them to orchestrate, I can get them to do a full end-to-end workflow, I can get them to review, and I- it barely eats up like, you know, half of what it used to.

It's eating up even less than half of what it used to, so I can actually do multiple of those workflows now back to back without having to wait for a reset. It's now become usable, and I can [00:09:00] do more complicated things than I could before, even if the intelligence of these models hasn't gotten better.

Because the workflows that I've unlocked now, having done it so many times, are becoming more demanding

Matt: The word I kinda wanna used is, kinda wanna use, and I think it's not really right, is things are starting to become consumerized in an, in an extent, but kinda not. Because a person that hears about GPT-6 Astra and then gets access to it is gonna just turn it on and be like, "It's the best one. Of course it is," and then just kinda go with it. Meanwhile, I'll compare it to like an iPhone, where it's there's like iPhone Pro, iPhone Pro Max. I mean, the person that is the, like on a budget, they're gonna go for the Pro at most, and I think there's other lower models of iPhone and things as well. Uh, so like, they're not g- ever gonna get the Pro Max experience, and that's what I kinda I guess I mean by it a little bit, [00:10:00] where the press and the news cycle and the hype is around something like a GP six-- GPT-6 Astra, but the general population knows that they're not really gonna use that. And it kinda stinks a little bit because with something like an iPhone, it's like if I just don't buy the Pro Max, I just don't get the Pro Max. With something like ChatGPT, it's like if even if I purchase a lower plan, maybe I'll get a little bit of later access, but I'll get pretty quick access to Astra, but then I have these, these l- usage limitations on

Mikhail: Yeah

Matt: now. Now the people that are on a free plan, they're used to limitations, they're used to, you know, whatever, v-various models and various, uh, AI apps have different limitations, prompting limits, conversation limits, et cetera, et cetera, et cetera. Amount of time that you've used it, amount of compute time. You can't use the-- our agents, you can only use our chat. The list goes on and on and on for different li- ways that things are limited. So the people on the free tiers are probably already accustomed to [00:11:00] this, where they know they'll get Astra way later, uh, and so they don't care, and they're not, you know, overly using AI. But it's just weird. It, it'd be almost like they sold us an iPhone...

It's almost like they s- like if you're on a paid plan, you get access to Astra. It's like, "Here's the iPhone Pro Max, but one of those cameras only works for six sh- six shots a day, and then we're taking that camera away."

Mikhail: I, I mean, it's a classic upsell technique, right? Like, they're trying to make you buy the, the more expensive plan, I'm guessing. But I do kind of agree with the fact that, like, they should almost ... And people might get mad at this. They should almost limit the use of the higher-end models in the lower-tier plans 'cause, like, the Astra 1 doesn't make any sense.

Like, you could ... You ... If you give it a complicated enough prompt, it will burn through all your usage in one prompt. That's how, like, costly those models are.

Matt: Yeah,

Mikhail: And for the majority of people, it won't give you the output, uh, uh, like the different type of output than the, [00:12:00] like, the, the, the slightly lower-tier model, like the, the Sol model, right?

Now, if you're doing something super complicated, maybe it should prompt you that you're doing something super complicated, and you should upgrade your plan or something like that. I don't know. But I do agree that there should be some ... Th- they need to think through it a little bit in a usability sense 'cause they haven't addressed this yet.

That, that's not part of this new, uh, this new release. Like, it's still ... Astra's still ludicrously expensive, and it will remain that way probably for a little while un- unless they release, like, a new, like, more optimized version of it, which I hope they do because Astra is a really good model. Um,

Matt: Yeah, the

Mikhail: mm-hmm,

Matt: were, were very good,

Mikhail: yeah, it's a

it's very good

Matt: feedback. Like, it, even on fi- 5.6, it wasn't like I was getting a terrible result from my prompts, but I would maybe get a, a prompt a result that I needed, needed a bit of massaging, like another couple, like, clarification prompts, 30 seconds of work each type of thing, just to clean some things up, whatever. [00:13:00] but the Astro one maybe got it on the first try, or only needed one additional prompt. But you're right, like, it's, if it's, if it's burning through the tokens, if it's burning through the usage, w- like, what's the, what's the point? And, and I noticed that it, like, the instant, and there's, like, the think longer, and there's, like, all these, like, different, like, user controls that you can do. But, like, it's, it's kind of interesting because it's, it's handing, wanna say technician controls, but it's handing sort of, like, controls to people on a technology that's supposed to not really need the controls, if that makes sense. Like, this is an intelligent thing. Why am I gating it? Like, I'm going in and being like, "No, no, no, just use the instant model for this.

No, no, no, use the longer model for this." some stuff I get where you wanna, like, "Hey, I wanna engage the deep research mode on this," right? Uh, and, and there's also other things that it, it does do automatically, where I've been in ChatGPT in the chat mode, and then it said, "Hey, we think that this might be something that would [00:14:00] be useful for work." When it goes to work, uh, it, you know, opens up work and it has more capabilities. In that particular case, it changed the model. Now, whether my model in work was set differently than the model in chat, I don't, I don't remember that. But thing is, is that it, it does have a little bit of, like, intelligence

Mikhail: Угу

Matt: "Hey, you should switch to this.

Hey, let's, you know, hey, let's try it here." And then if I'm doing something that's starting to bump up on coding, it's like, "Hey, maybe you should use Codex." Like, it knows a little bit. But a lot of controls for something that's supposed to be very consumerized. Like, the thing that you're describing, the thing that I'm describing isn't quite up on the, the high-end spectrum of, of, of token maxing,

Mikhail: Угу

Matt: it, it's touching...

It's a little bit of token maxing.

Mikhail: Угу

Matt: a little bit of, like, "Hey, let's..." Like, like you're saying, like, "Let's, let's distribute our work," right? It's sort of like, "Give the intern this, give the s- the senior guy this, give the tech lead this."

Mikhail: Mm-hmm. Yeah, for sure. Like it, it, [00:15:00] it's a whole different paradigm, and I think they're struggling with the UX of it, honestly. Uh, trying to make it more consumerized and going down rabbit holes like the co-work rabbit hole and the work versus ChatGPT versus Codex rabbit hole. Like, all of that is still, I would say, not optimized and still fairly confusing.

Matt: Well, I think the ChatGPT Classic versus

Mikhail: yeah

Matt: Codex, that has a cl- that has a Codex mode. But by the way, both ChatGPT Classic and ChatGPT yeah, the ChatGPT app, which contains the ChatGPT mode, that contains chat, but also contains work like... You know what I mean? Like, we're starting to get...

It's not rocket science. 10, 15 minutes in

Mikhail: Mm-hmm.

Matt: it out.

Mikhail: Yeah

Matt: for something that's just supposed to be like, "I'm an artificial intelligence. Hello, what do you need?" Not, "What do you need? and what model do you need it for? And how long do, how long will you allow me to work? And do you have any resets left in case we..."

You know? Like, like, there shouldn't be follow-up [00:16:00] questions. Like, they should be working toward not having follow-up questions that aren't about the prompt itself,

Mikhail: And, and I think they--

Matt: it

Mikhail: I do think they are, and I, I, I just think it'll take a little bit more time. Like, for example, there is gonna be, unfortunately, another OpenAI app released along the, alongside Codex and ChatGPT. Like, there's gonna be a separate one, um, that's gonna be more similar to like a GroqBot or a Claude, uh, OpenClaude or a Hermes or a Muse Spark.

Like the-- It's gonna be more of like a self-contained agent, uh, that has its own compute. But it's gonna be, have a lot of overlap with ChatGPT and Codex, obviously. So like it, it-- I feel like it's gonna get more confusing before it gets better, is what I'm trying to say. So strap in. Uh, the UX of this whole experience is still definitely not solved.

I think we're still in the them experimenting on us phase. Uh, but I think they are working towards it, and at least from a usability [00:17:00] perspective, from an engineering perspective, like for engineering point of view, they're starting to work on like, "Hey, let's make these things faster. Let's make them a little bit cheaper.

Let's make it so that people c-can actually use our real- the really complicated workflows that are really cool, uh, without burning their usage in 30 minutes, uh, but still, and still get the output that they would receive." So that's what I'm kind of excited about with these new releases and where, where this is going.

I, I think I mentioned this, you know, maybe a few months ago now that like I hope that they start focusing on speed and, and cheapness versus intelligence because the models are already pretty good, and it looks like that's where at least the in- this early stage focus is on. So pretty cool.

Matt: And I think this points, this points out something too is like, uh, like first of all, the AI, uh, industry is in its infancy still.

Mikhail: Угу

Matt: early. They're still figuring things out. Um, and every industry has this. Let's all have separate apps. Let's, let's all put the apps together. Let's all have separate again.

Let's, let's put them all [00:18:00] together again, you know, back and forth, back and forth. We saw this with companion apps for gaming, uh, for PlayStation. PlayStation had like three freaking apps now. Then it went back down to one, and like back and forth, you know, whatever. Uh, we, we've seen this with the internet, where like now it's just very obvious.

Hey, make an account, you know, get into the app, or get into... Or, or you wanna host a podcast, look up a podcast service, make an account, get going kind of thing. Um, I'm sure like the UX for that was, you know, slowly but surely getting figured out back in the day as well. So like, I mean, we... not necessar- I'm not necessarily complaining like, "Hey, what the hell?"

Like, "Let's get this done." I, I acknowledge that this is, it, it is early, and like these are more comments than they are, I don't have my pitchfork out or anything like that. It's sort of more, "Hey, this is where we need to start working towards," because it's confusing for consumers. Um, and, and this is also kind of where something like Gemini on the phone is probably... It makes a lot of sense. Like my, even my parents will just hold down their power button, have the Gemini just, and they just talk, and they don't [00:19:00] go like, "Hey, do you need an instant answer? Do you need like a, a moderate plan? Do you want Gemini Four? Do you want Gemini Six? Do you want Ge-" I'm just throwing versions out

Mikhail: Yep

Matt: you know what I mean?

Like, it's not asking all these questions. Like the, like my parents are like, "What?" Like, "I just wanted to know how long to cook this meat pie for." Like

Mikhail: Yep.

Matt: So kind of where something like, know, like a Gemini does have a bit of an edge there because it's like hold the power button, talk to it.

Okay. You know, and, and, and away it goes. So I mean, I, I mean... A- and also I would like to say one final thing is, as we keep pushing models forward, like you're saying, you're handing it to, to these models. You're handing it to models that are like a year-ish old, not three, four, five, six years old.

Mikhail: No, ja

Matt: matter...

Obviously, not six years old, but the point of the matter is, is that as we keep models forward, the old ones become mundane. I don't mean bad. What I mean is we know how to... We know we've min-maxed their compute.

Mikhail: Угу

Matt: out how to make them super efficient on the hardware we have, whether it's a hardware efficiency, a software, or a software efficiency, or both. And we see that in the [00:20:00] gaming industry when a console comes out and then a slim one comes out later that has the same power. It's like this is smaller. It's the same. It pro- it provides the same graphic, it, it, you know, et cetera, et cetera. Oftentimes it's cheaper, but gaming industry is kind of in a thing right now.

But regardless, the point is, is that it's good because if Astro's amazing, I mean, if Astro's amazing at making posters, it will kind of, quote-unquote, "always" be amazing at making posters. as we go to G- you know, GPT-7 and GPT-8, and like the years go on, that awesome poster maker is always gonna be there or like, you know, re- reasonably will always be there and it's easy to do that thing.

Like if, if you, if you've mastered poster making or you've mastered a certain task, you've mastered the task. And so we don't necessarily need a new model. And so then if we figure out how to make the tokens for Astro really cheap or make it use like no tokens, there you go. So then even people on the free tier will have an amazing [00:21:00] experience with their answers, their posters, and whatever else they're making.

I just, I just pulled posters out of thin

Mikhail: Yeah, yeah, of course.

Matt: it's good at p- making

Mikhail: Yeah. No, I, I mean, like th- this is what... Yeah, that's exactly kind of like the cusp of the, of the point of this episode is that we've, we've reached an intelligence point where these models are starting to become useful in many different ways. If we can just optimize them, they start to become even more useful, and hopefully that happens with Astra, and it probably will, uh, due to like how technology works.

So we're, we're in an interesting spot because intelligence, yeah, it's, it's... There's still some non-diminishing returns on the intelligence side, so we still need some more intelligence, but we're starting to see the light at the end of the tunnel of like, okay, these things are smart enough now. We really do-- Like if we, if we stopped developing the intelligence today, we'd still have a really cool kickass system that would potentially be safer because, you know, it's not gonna go into ASI or whatever.

And, uh, and we'd be [00:22:00] able to maybe min-max the performance of these systems. If, if Astra is 10 times faster and 10 times cheaper, damn could I do a lot more with it today than I could, you know,

Matt: I

Mikhail: tomorrow than I could with it today. So that's, I think, a good point. Um, yeah, I'm curious what everyone else thinks.

Like, what, what are you most excited about with the future of models? Is it the intelligence or is it them becoming a little bit more usable?

Matt: Yeah, let us know in the, in the comments. Let us know on, on the Spotify comments, on the YouTube comments. Let it- Hit us up on the old social medias with, uh, with your answers there. We'd be interested to hear your opinion on, I mean, do you use Astra? Do you not? What do you think about using old models for certain tasks like Mike does, kinda delegating out certain tasks to certain models to kind of token max.

We'll call it token

Mikhail: Yep

Matt: in a way. Uh, and, uh, yeah, that's it. That's the web news for this week. Hope you enjoyed it. you next time

Mikhail: Goodbye