TWiT+ Club Shows 769 Transcript - AI User Group #19.5
Please be advised that this transcript is AI-generated and may not be word-for-word. Time codes refer to the approximate times in the ad-free version of the show.
Leo Laporte [00:00:00]:
This is TWiT. Ah, meanwhile, I should just check in with my, uh, the reason I have been available today is because I have been busy. My AIs are busy benchmarking themselves. I wanted to know if, um, wow. I wanted to know if, uh— Well, the first thing I did is I had the, uh, the council gather, and I said, which model should I use on the Sparks? And, uh, it said, uh, they all agreed GLM-53 Flash was the universal one. Not DeepSeek, not Quen-38. But GLM. But every time I use— every time I go out to Twitter, there's a new version of GLM 5 through Flash, each getting better.
Leo Laporte [00:01:07]:
So I have 2 different versions now, one with NVFP4 and one with this new EXL3, which are ways of quantizing the The parameters. And people seem to like this EXL3. So I said, all right, benchmark both. And while you're at it, benchmark DeepSeek— I'm sorry, GLM53 Flash on z.ai. Because I want to know if the version I'm running locally, the quantized version, is any good compared to the non-quantized version running in the cloud. It's not. So Spoiler alert, it's not even close. And then I said, while you're at it, might as well, I got Kimi K3 fired up.
Leo Laporte [00:01:53]:
Uh, why don't you see how Kimi does? Because it's also OpenWhite. So we shall see. Uh, yeah, I know NVFB4 should be the best, but, uh, I did try SG Lang and, and you're right about that one. SG Lang didn't work that well together. But I'm curious, this, people are saying the EXL runs really well. The biggest problem I'm having right now is there are a lot of influencers in this. And a lot of the stuff I see now on X is really link bait, kind of. People want to, you know, they want to—
Leo's Laptop [00:02:33]:
yeah.
Leo Laporte [00:02:33]:
So see, here you go. This is— so I like Mia AI's recipes. I've been using her recipe or his or whoever they are. I've been using Mia AI's recipes. And Mia, this is the one I'm benchmarking, EXL3 for Dual Sparks. And she says it's really good. So actually, I have to send this recipe to Claude and say, hey dude.
Anthony Nielsen [00:03:25]:
New recipe.
Leo Laporte [00:03:25]:
So Anthony is gonna buy a Mac. No, he's not. Oh, it did all right. Wait a minute, wait a minute. Oh no, it didn't do all right. Did terrible.
Anthony Nielsen [00:03:37]:
What kind of a, like, your, your, the tokens per second are you getting for the GLM on your Sparks?
Leo Laporte [00:03:44]:
Um, about 40 here. Oh, that's usable. Yeah, yeah, totally usable. I don't know if it's working right now or if it's resting. Yeah, so this is right now at 20. There it goes, 37. So it's working. And this is the NVFP4 version.
Leo Laporte [00:04:19]:
So yeah, 40 is fine. I don't feel like it's slow.
Leo's Laptop [00:04:23]:
It's 262.
Leo Laporte [00:04:25]:
kilobytes of context. But I actually looked at all the stuff I'm doing and that's plenty because I don't like to get context full anyway. And I have an— oh, and I do recommend this, Anthony, or anybody doing context compression on local models. There's a new LCM graph compression that I think is much better. You're running the 2-bit.
Leo's Laptop [00:04:55]:
Yeah.
Leo Laporte [00:04:55]:
I don't know about tool bit. Yeah. So Anthony and I are both trying to figure out, should— one option, I haven't told you this, Anthony.
Leo's Laptop [00:05:09]:
Uh-oh.
Leo Laporte [00:05:09]:
One option is we give out Christmas bonuses every year. I shouldn't even say this, but I'm going to say it anyway. I was thinking I give you and Micah each a Spark. in lieu of a bonus. And then I buy, then I can justify to Lisa, well, you see, we saved all this money. I can buy—
Anthony Nielsen [00:05:30]:
A 512.
Leo Laporte [00:05:31]:
A 512 gig. But I don't think the 512 gig Mac, M5 Mac Ultra Mac Studio, I'm not convinced. I think there's an advantage to CUDA. I'm not convinced it beats, I mean, the other option would be to keep the 2 Sparks and buy 2 more Sparks. Which I'm not going to do, but I think what I'm actually, I would say, well, okay.
Anthony Nielsen [00:05:54]:
Like the current, currently the previously the M3 Ultra, like it would generally have faster token per second, but the prefill slower. Yeah. The prefill slow. Right.
Leo Laporte [00:06:06]:
And you don't have CUDA. So a lot of models like that Breeze model, which was incredible, uses CUDA and there's no MLX version of that.
Anthony Nielsen [00:06:13]:
So yeah, but there will like generally like there will be. Yeah. You're just not going to have everything day one.
Leo Laporte [00:06:22]:
So yeah, it's really unclear. And that's one of the reasons I'm doing all these benchmarks. It is my general feeling at this point that if I can get GLM running well, and every day there's a new recipe that makes it run a little bit better, a little bit faster, a little bit higher quant or better quants. I'm getting about Opus 4.6 is what I think. That's fine. God, I mean, 6 months ago, I would have been thrilled. That was the state of the art, right? Yeah.
Anthony Nielsen [00:06:57]:
I mean, that's what I've been saying. Like, you know, I haven't—
Leo Laporte [00:07:00]:
That's all I need for a local agentic model. And then I'm still using, you know, I'm gonna, Probably. I don't know what I'm going to do once TwitAds is over. I don't think I— right now I have a Max, Claude Max, and I have an xAI subscription. I have 3 cloud subscriptions. But that's just because I'm working on this TwitAds thing. As soon as that's— it's never going to be done, but as soon as it's mostly done, I think I could pay $100 to Noose.
Leo's Laptop [00:07:33]:
$100.
Leo Laporte [00:07:34]:
I want to support them anyway. They're the ones who do Hermes, and they have a bunch of models and their tokens are cheap, and I can use that for Frontier. They have Anthropic, they have OpenAI, they have all of them, and they work fine in Hermes. So I think maybe a $100 news subscription would be plenty. And I've been kind of keeping an eye on that just to see.
Anthony Nielsen [00:07:55]:
I mean, I wouldn't— you can, I wouldn't use it. Use your new subscription to, to use like Fable or something. I accidentally—
Leo Laporte [00:08:06]:
Did you blow it out?
Anthony Nielsen [00:08:08]:
Oh, well, I spent like $5, you know, $5, $10 on one, like one prompt just because of the—
Leo Laporte [00:08:14]:
Right. You know, yeah, I probably wouldn't use it for Fable, but I don't know if I, I don't know. You know, right now the $200 Anthropic subscription is a great deal. I just don't know how long it'll be a great deal, right? But right now it's an amazing deal. I don't, I can't use up $200 worth a month. Let me get you the LCM. This I would recommend if you're using Hermes. Lossless, it's called Lossless Context.
Leo Laporte [00:08:44]:
I'll give you the link here. And in my tests, it's been much better than just standard compression for context compression. I use Quen to do the context compression.
Leo's Laptop [00:08:57]:
Hmm.
Leo Laporte [00:08:58]:
Yeah, that's running on— actually, so that was an interesting story. Remember I was using that Breeze, but Breeze takes up the whole 5090 basically. It's like 18 gigs. And as good as it is, I don't really need a clone of Bill Gates' voice, to be fair. So I'm going to— I go back to Kokoro, which runs on the framework. It runs beautifully on the framework. And, uh, I'll play up the Kokoro voices are the ones you've heard. They're fine.
Leo Laporte [00:09:33]:
The Breeze voices were amazing. It did an amazing clone of my voice, but, um, that's a high price to pay, a whole 5090 dedicated just to get fancy voices on my agents. When Kokoro does it, very lightweight, very fast, low latency, and then they're fine. They sound fine. So. Okay, Whoa Joe. I was thinking of giving it to employees. Micah is going to— is it— can I say what Micah's doing or is that a secret still?
Anthony Nielsen [00:10:04]:
It's a secret.
Leo Laporte [00:10:05]:
Okay. Never mind. Never mind. But as you know, Anthony is our local AI guru and I think he really needs a Spark badly. Actually, you really need a 512-gig Mac. But that's going to be, God, 20,000. I'm going to be in Southeast Asia. I, I'm not—
Anthony Nielsen [00:10:26]:
I mean, okay, you know, 512. Yeah, it would, you know, for you that'd be a lot of fun. But I think even 192 for the 256 is like— 256, it's gonna have a lot of lakes.
Leo Laporte [00:10:40]:
Yeah, yeah, yeah. You could easily run—
Anthony Nielsen [00:10:43]:
or 512, sorry.
Leo Laporte [00:10:45]:
Yeah, well, no, yeah, 256, you could run GLM.
Anthony Nielsen [00:10:48]:
Yeah, exactly.
Leo Laporte [00:10:50]:
I don't know what the MLX versions are, if there are any yet, but there will be. You're right, there will be at some point. Somebody will do it. I also think there'll be a Spark 2 before the end of the year, but maybe not.
Anthony Nielsen [00:11:12]:
Yeah, or even just a RAM bump.
Leo Laporte [00:11:15]:
I'm trying to teach myself to burn into my mind that Opus 4.6 at 40 tokens per second locally is apps I would be— a year ago, my eyes would be going, what? Yeah.
Anthony Nielsen [00:11:33]:
And it's not going to—
Leo Laporte [00:11:33]:
I'm so spoiled.
Anthony Nielsen [00:11:35]:
And it'll still get better. There's going to be new models coming in.
Leo Laporte [00:11:38]:
Right. It's just that you see GLM-53 come out and go, oh. I wish I could run that. And then that's what I mean about this. X and YouTube is filled with these influencers who really don't— they don't have— they're not really using it. They're interested in views is all they're interested in. So a lot— it's a lot of BS. Let me see if I can one-shot a game.
Leo Laporte [00:12:05]:
It's just a lot of BS. It's not practical knowledge. And they give you that FOMO. Because they're, you know, go to X, it's like 5.3 just came out and I'm running it on my 5, you know, whatever H100s. It's like, well, nice, good for you, buddy. And I guess they make enough money on YouTube to justify that, but this is not, I'm, you know, I'm, uh, this is coming out of my pocket.
Anthony Nielsen [00:12:32]:
I mean, they must get gifted those Sparks, like, because you wouldn't—
Leo Laporte [00:12:36]:
Well, that's the other thing is NVIDIA is giving a lot of these away.
Leo's Laptop [00:12:41]:
Yeah.
Leo Laporte [00:12:41]:
And I, even if they wanted to give me one, I wouldn't take it. So, you know, I mean, that's to me the difference between what we do and what these YouTubers do is, is we're trying to be journalistic about it. We're not doing this for the views, you know. So here's a perfect example. You know, this Tech2Wild is one of those guys who's You know, it's just great. You can make a Coliseum.
Anthony Nielsen [00:13:09]:
I, you know, it's so, you know, the, the proof of concept is always quick. It's the, that last mile's a—
Leo Laporte [00:13:17]:
Well, that's, and so that's what I'm doing with this, which is using, um, I have problems from our own experience. Oh, oh, oh, I gave it to the wrong— no, I gave it to Cakey. That's right. No, that's right. This is Fable. I can actually record this now. Finish for NVFP4 before you Bring up the XL3. Oh, and I'm on the Mac 2 MacBook Air, M2 MacBook Air, M2 MBA right now.
Leo Laporte [00:14:17]:
So that's where you should voice. This is, uh, so the Kronk voice was like that great I'll play it again for you. I loved the Kronk voice, but I— that's a very expensive, very expensive voice. Oh, I'm so disappointed that my voice— the clone of my— oh, now it's playing. I'm Leo. Welcome to my setup. This page is a tour of my AI household. 5 named agents and a fleet of machines I own outright.
Leo Laporte [00:14:52]:
With all but the toughest coding tasks, free and local. Scroll down and meet the crew. They're nicer than they sound. Most of them. So that's not a great clone, but it was with 20 seconds of my voice. This— the Bill Gates, I thought, is pretty good. Hey there, I'm Quicksilver. Q for short.
Leo Laporte [00:15:10]:
I keep the whole show running. And then this is Kronk. It is all coming together. I am Kronk.
Leo's Laptop [00:15:18]:
I build the big things around here. I would rather do the work than talk about it.
Leo Laporte [00:15:23]:
So that's Breeze, but I— but you know what, it's— that was a pricey thing. I had to give up my whole mo— my whole Mojo machine, my whole 3090 for that.
Anthony Nielsen [00:15:35]:
Voice target.
Leo Laporte [00:15:36]:
Okay, let's see what's going on here. You'll hear the new voice, which I think is fine, and it— the nice thing is I can run it on the Flame framework.
Anthony Nielsen [00:15:43]:
It's Kokoro.
Leo Laporte [00:15:47]:
So we're running this benchmark, running it. We ran it on— it ran very well, by the way, on GLM53 Flash on z.ai. I mean, it ran really well. Like it aced it and it did it in about 45 minutes. So locally it's been running for hours. So that's one problem. Then I'm going to see what Kimi K3 does. I should get rid of the, uh, we don't need that anymore.
Leo Laporte [00:16:17]:
This is new.
Leo's Laptop [00:16:18]:
Got it. NVFP4 finishes its benchmark before EXL3 comes up. Small rescue on the weights download. The recipe's completeness check counts shards but ignores revision, so it would have quietly kept the stale weights. I forced a refresh at the pinned revision, And it's downloading now.
Leo Laporte [00:16:38]:
I have no idea. That was claudish. I have no idea what any of that meant. I found the root-owned entries left by root containers. Oh, that's the old version. Oh, okay. Yeah. I can run the sudo command for you.
Leo's Laptop [00:17:02]:
No need for your hands. Last night's D-Flash container left a rootoned directory in the weights cache. So my download—
Leo Laporte [00:17:08]:
So that's a perfectly good voice, right?
Leo's Laptop [00:17:09]:
I staged a fix script on both Sparks.
Leo Laporte [00:17:12]:
And that's Kokoro.
Leo's Laptop [00:17:13]:
It's actually arguably better.
Leo Laporte [00:17:16]:
So I have to do this.
Leo's Laptop [00:17:26]:
Perfect. The 2 one-liners are in the chat with the exclamation prefix. The moment the first one lands, my watcher relaunches the download by itself.
Leo Laporte [00:17:42]:
Done.
Leo's Laptop [00:17:46]:
Sparky's cache is clean, thank you. My watcher will relaunch the weights download within seconds. When you get a moment, the same one-liner on Sparkles finishes the job.
Leo Laporte [00:18:01]:
That was the one-liner. It sparkles.
Leo's Laptop [00:18:05]:
Oh.
Leo Laporte [00:18:13]:
Okay. So that's because I get this new model from Miik just came out. And so before we do the benchmark on EXL, but I'm a little worried because as I said, the GLM Flash on the cloud, Did great. Where is this? Let me open this up. Research coding model eval. Oops. Research coding model eval rating leaderboard.
Leo's Laptop [00:18:56]:
All clear. Your fixes landed on both nodes. And good thing we checked. The recipe would have served the old weights anyway because its ref file still pointed at the stale snapshot. That's corrected. The draft weights are finishing and everything's staged for EXL3 the moment the benchmark wraps.
Leo Laporte [00:19:14]:
It's been going forever. So scorecard is what I want to load up here. That's leaderboard. Where's the scorecard? Where did it say it was? Results GLM. Here we go. So this is based— this is my own benchmark, not a public benchmark based on my own work. So Claw— Cakey is a Fable. So Fable's doing this.
Leo Laporte [00:19:51]:
So this is on ZAI, and look, it did perfect. It got 88 out of 100. Now, on this same benchmark, Fable and Groq 46 got 100. So that gives you an idea, but 88 is great. And it did it in less than an hour. Mine's locally, theoretically the same model, quanted. It's been running, I don't know, 4 hours, 5 hours, I don't know. So we're doing it with syncing off on locally.
Leo Laporte [00:20:28]:
I'm not sure why. Couldn't turn syncing off on the cloud. Anyway. Oh, looks like they're starting, the results are starting to come in. CloudC.upgrade strong. Oh yeah, this is the cloud still. Oh, this is all still the cloud. Okay, so that was the EXL3 download.
Leo Laporte [00:21:05]:
Well, I'm hopeful that I will have a model that will run It's not gonna run as well as the cloud, I guess, but let's hope that I can get a model that would run okay. Yeah, there are some good local song generators. I don't know if they're not quite Suno. Suno's gotten so good. Suno's amazing. I don't know, they're pretty good. Are they? Which one's the best one, best local one?
Anthony Nielsen [00:21:31]:
It might be Astep right now, but there's also, what's the other one?
Leo Laporte [00:21:41]:
No, I don't have 256. I have dual Sparks, Darren. That's why I'm trying GLM. And I ran a council and they looked at the results from my original results from running DeepSeek Flash for a few weeks or whatever since I got the Sparks a week or two. The Quen, because I did put Quen 3.8. Um, I don't remember if it was— yeah, it was Flash, I guess. I did benchmark it and it is not awesome. It wasn't good.
Leo Laporte [00:22:14]:
I'll show— I have the results somewhere. This is the song that Suno made for, uh, the whiskey. Segment, and I thought this was pretty darn good.
Leo's Laptop [00:22:33]:
Now if the label disappears, it's a Windows Weekly whiskey. See you again next year.
Leo Laporte [00:22:43]:
Oh, it started in the middle. This is a live one, it sounds like. Oh, Quen 3/8 Flash next. Oh, okay. I haven't tried that one.
Leo's Laptop [00:22:58]:
There's a bottle in the way with a label from a Tuesday and a cork that's gone astray. Richard's got the tasting glass, the window's cruising cheer. It's something weird from my closet and it's whiskey time this year. There's a rye from a farm where the scarecrow knows my name. A malt that spent its Winter in a barrel marked Bicycle Flame. A corn whiskey golden as a plumber's shiny shoe. And a tiny little dram that says, I might remember you. Hey ho, pour it slow, let the good bottles roam.
Leo's Laptop [00:23:37]:
From the island to the prairie, every cask can tell its home. Raise a glass for Richard with a grin from ear to ear. It's something weird from my closet. And it's Windows Weekly Whiskey here.
Leo Laporte [00:23:50]:
It's pretty damn good. Yeah, I'd say Suno's gotten better and better and better.
Anthony Nielsen [00:23:56]:
Although I think like it's— ASEP is very comparable. Like you still have that like kind of, uh, or I think, uh, Darren says—
Leo Laporte [00:24:05]:
What are you running it on? Did he make this?
Anthony Nielsen [00:24:08]:
He's running on Spark, but like you can run it on as little as 4 gigs of VRAM. Like even my old potato computer can Generate something.
Leo Laporte [00:24:17]:
So, uh, let me look at this.
Anthony Nielsen [00:24:20]:
Where's Discord?
Leo Laporte [00:24:21]:
There it is. This is A-Step, huh? Oh, it's gonna launch— this is what I hate about Macs. It's gonna launch Apple Music to play this. Jeez, so annoying. I'll fix that, but Just really annoying that's the default. Hello? Oh, please. God, Apple sucks. I am so done with Apple.
Leo Laporte [00:24:52]:
Let me see. What can I do here? Can I open this with—
Anthony Nielsen [00:24:56]:
no.
Leo Laporte [00:24:58]:
Go back to the browser, go back to the downloads, show all downloads, uh, open, show in Finder. Open with— what do I have? Nah, I'll let VLC be better. Let's see if it'll play with ENA. It's an MP4, which is why it's—
Anthony Nielsen [00:25:20]:
Yeah, I think—
Leo Laporte [00:25:20]:
It's a little weird. It's not—
Anthony Nielsen [00:25:23]:
I think it was a video at one point. It was supposed to be for Windows Weekly, but Windows Whiskey.
Leo Laporte [00:25:44]:
It's Suno 2.
Anthony Nielsen [00:25:51]:
Yeah, I mean, they all have that. Like, you could still— I mean, even Suno, you could tell, like, oh, that's AI.
Leo Laporte [00:26:00]:
I don't know. Suno's gotten so good. I'm not convinced.
Anthony Nielsen [00:26:06]:
It's more of a— not like— I'm not talking about the actual music. I'm talking about like there's like a quality in terms of—
Leo Laporte [00:26:12]:
Yeah, yeah, yeah, yeah.
Anthony Nielsen [00:26:14]:
What is this?
Leo Laporte [00:26:15]:
Local sound test. Oh, that's Ina. Okay. Yeah, it's pretty good for local. Absolutely.
Leo's Laptop [00:26:46]:
Absolutely.
Leo Laporte [00:26:50]:
Yeah, it reminds me of early Suno. Do you have, do you have any preliminary results for the Envy FP4 benchmark. CloudFlash scored 88.8, Kimi K3 84, and the local NVFP4 is still thinking. Like 4 hours later, still nothing. Oh my God. Took longer with thinking on than thinking off.
Anthony Nielsen [00:27:31]:
Isn't KIMI normally like a really large model? Is it?
Leo Laporte [00:27:34]:
It is. I couldn't run KIMI K3. I, uh, and yeah, no, it's a good, but interestingly, the Flash is doing better. 88.6 compared to 80.4, but I have to say, uh, Groq and Fable, 100%. They nailed it. I'm sorry, Groq and OX Alpha. That's what's a little weird to me.
Anthony Nielsen [00:28:01]:
Right, that they're so different.
Leo Laporte [00:28:04]:
Yeah, you know, they're— it's— you just don't know. The full local scorecard, 5 to 8 hours running locally. I don't know why it's so slow. They may just be too hard. I wanted them to be too hard. This is a bug in Claude. I don't know if anybody's noticed this and if anybody has a fix for it, but you see how it's refreshing the screen? It looks like something's happened. Maybe it's only in Herder.
Leo's Laptop [00:28:36]:
I don't know.
Leo Laporte [00:28:37]:
See that? Nothing's happened. It's just refreshing the screen and scroll— it's scrolling. Nobody— You know, Groq code doesn't do this. Z code doesn't do it. Py doesn't do it. I just don't know why Claude code does it. So annoying. Well, I'm hoping EXL3 will do better, but this is very discouraging.
Leo Laporte [00:29:08]:
This is— so there's a Quenf 3.8 Flash. I didn't see that.
Anthony Nielsen [00:29:15]:
Yeah, it's, uh, making quite a mess.
Leo Laporte [00:29:16]:
When did that come out?
Anthony Nielsen [00:29:18]:
I don't know, yesterday or something.
Leo Laporte [00:29:19]:
Let me see if I can— hey, I see that a Quen 3.8 Flash came out. Can I run that on the Sparks?
Anthony Nielsen [00:29:31]:
Flash Next.
Leo Laporte [00:29:33]:
See, it doesn't know how to do Quen.
Anthony Nielsen [00:29:38]:
3 days ago.
Leo Laporte [00:29:39]:
Flash next. Yeah, it's a— do you think if I didn't use Kitty, it wouldn't— maybe it wouldn't do it? Maybe it's a Kitty thing.
Leo's Laptop [00:29:55]:
I don't know.
Leo Laporte [00:29:56]:
I'm using— this is Herder, which is basically Timux. Good news in a wrinkle: it's not just possible, it's already happened. Oh, that's the one I tried.
Leo's Laptop [00:30:06]:
FV8.
Leo Laporte [00:30:07]:
Good news, you idiot. You already did it. I used the Mia 8.
Leo's Laptop [00:30:13]:
QN FlashNext already ran on the Sparks this week. It's parked with an instant resume script. Fair warning, the council dropped it yesterday and it timed out my whole coding bench at 21 tokens a second. If you want it back up anyway, the natural slot is after the EXL3 pass. Just say so and I'll cue it.
Leo Laporte [00:30:36]:
Yeah, I had a council go on it and they were unanimous that I should not be running that. I should be running GLM 5.3 NV4.
Anthony Nielsen [00:30:45]:
This might be better for a general agent because like it has vision and stuff like that.
Leo Laporte [00:30:49]:
Yeah, well, I'm using, yeah, I'm using the, I have Quen, um, 3827B running on Mojo now, now that I took Breeze off of it and that I'm using for vision. I was using, um, I tried GLM for Vision, but it kept inserting its reason, its reasoning into the— so I use it on my Protect cameras and it kept inserting its reasoning into the results. Where's—
Leo's Laptop [00:31:18]:
oh, there it is.
Leo Laporte [00:31:20]:
So this is, this is Quinn, a person with short dark hair wearing a gray hoodie, dark shorts, and sneakers walking on the sidewalk, blah, blah, blah. Let me go and find myself. It recognizes, uh, people. So if it knows me, it will say— oh yeah, see, Leo detected. Man with gray hair and glasses wearing a fade, thumbs up. That was me. But then later on the ground floor gym, an elderly male with short white hair.
Anthony Nielsen [00:31:49]:
Thanks, Quinn.
Leo Laporte [00:31:53]:
But, but earlier when I was using, um, GLM, see, it put the reasoning in there. I need to identify people, animals, and packages. There's a, there's a— wait, let me reconsider. The dark box, the items look like packages or foods or snacks. It could be delivery that was left, or it could just be a box and trash. It looks like a delivered package dark box with some grocery items. The person— I mean, it's so annoying to have all the reasoning in there. So, and I couldn't get it to stop.
Leo Laporte [00:32:22]:
I tried to stop it. I couldn't. So I said, nah, just go back. That's, that was actually what prompted me to get rid of Breeze, go back to Kokoro and put Quen back. Because Quen was running on the Mac, but it's a lot slower on the Mac. It took, I think, a minute or 2 to do the vision. On the 3090, it's seconds, even though it calls me elderly. Son of a— An elderly man.
Leo Laporte [00:32:51]:
A young woman with dark hair in a ponytail wearing a black sleeveless top, black shorts, and black shoes is standing on the driveway facing a man. A man wearing a dark long-sleeve shirt, tan pants, and a dark cap is standing on the driveway looking down at his hands while facing the woman. I love these. Uh, I, yeah, maybe, maybe Quicksilver was baffled. I asked it. One big model at a time on the cluster. Quen resuming GLM. Yeah, no, so apparently I am, did try next.
Leo Laporte [00:33:23]:
I didn't realize. I thought I was doing something with Quen 3.8. No, I forgot that I had already tested that and the council recommended against it. We're going to try to figure out what the best version of GLM 5.3 is. I admit I'm a little bit discouraged at this point that NVFP4 is struggling so hard with these problems. It's weird because the cloud version of ostensibly the same model, not quanted, ran through all those problems very fast. So you call it a neon format, huh? I don't— that's something new to me. I hadn't seen that before.
Leo Laporte [00:34:17]:
You tell it, use the neon format? The discouragement is worth pushing back on. Huh. Speed gap is hardware, not— I know it's hardware. I know.
Leo's Laptop [00:34:37]:
Lose heart yet? The ugly local score was the serving mode breaking the output format, not the quant losing intelligence. Where the format survived, local was perfect, including that flawless bug hunt. The speed gap is just data center hardware versus your desk. The thinking-on-run happening now is the first fair test. And it'll tell us whether NVFP4 or EXL3 is the right GLM for the Sparks.
Leo Laporte [00:35:05]:
Ah, that's good to know. Yeah, I understand that the speed is slower, but I'm happy with 20 to 40 tokens per second. I'm hoping the EXL3 can run closer to 40 consistently. Yeah, this EXE is interesting. I don't know exactly what it's doing. This is like watching grass grow, isn't it? I'm sorry. This is the only reason I haven't done streaming video of AI coding.
Anthony Nielsen [00:35:43]:
I mean, but that's like the whole like You know, let's play and stuff like that. Like it's a different format where it's like those downtimes are, you know, expected in those jobs. Well, you're chatting with whoever's in there, you know, you're talking in the chat room and stuff.
Leo's Laptop [00:36:02]:
Mid-20s are in long context. So coding should blend above 40. Claims usually deflate a bit when we reproduce them, but their bench script ships in the kit. So we'll have real numbers on our own cluster within an hour of bring-up.
Leo Laporte [00:36:18]:
Okay, see, that's— so that's what I'm hoping is this EXL3. I don't even know what it is. Can you tell me what EXL3 is doing that's different from NVFP4? I thought NVFP4 was NVIDIA-tuned for the GB10s and other CUDA cores. Kind of misunderstood CUDA cores as CUDA coarse. The interesting thing about LLMs is they understand typos. They kind of understand the intent. No, so it did nv, fp4. I'm going to turn on the chat room.
Leo Laporte [00:37:04]:
Let me turn that on. Oh, that makes it so big though. So let's see, how can I— I should just get rid of myself.
Leo's Laptop [00:37:18]:
Short version, NVFP4 is built for Blackwell Tensor Cores, but our software stack on the GB10s can't use that path. The weights get dequantized anyway. EXL3 accepts that. And spends its 4 bits smarter with an error-minimizing encoding, plus its kit brings the speculative decoder. So the bake-off is really stranded hardware format versus optimized software format.
Leo Laporte [00:37:45]:
Ah, so the NVFP4 is— well, does not— this is the problem with these crap GB10s. They're not really full CUDA cores. Here, I can— you know what I can do is make this just bigger.
Leo's Laptop [00:38:00]:
And you can read it?
Leo Laporte [00:38:00]:
Yeah, there you go. Oh, that's interesting. See, there's all sorts of little issues with the Sparks, memory bandwidth being the chief one, but I didn't realize NVFP4— maybe somebody could write— okay. Marlin dequantizes the FP4 weights on the fly. Oh, interesting. Oh, so you get size benefit, but not speedup. Okay, interesting. xLlamaV3 project takes the opposite bet.
Leo Laporte [00:38:43]:
Since you're dequantizing anyway, spend the 4 bits smarter instead of rounding each weight onto a fixed FP4 grid. Oh, it's trellis-coded quantization. I didn't know that.
Leo's Laptop [00:38:57]:
It's Q-Tip.
Leo Laporte [00:39:01]:
Oh, I love the names. Okay, so that's interesting. So EXL3 is actually kind of interesting. I wasn't getting the benefit of NVFP4. That's why you buy, you know, 5090 or something. And this is the recipe. See, there are these recipes. It's so funny.
Leo Laporte [00:39:29]:
You get these guys banging on this, doing all sorts of quantization tricks and putting up recipes. So every day there's been a new recipe that I've been modifying. The problem with this recipe, at least according to Mia AI, is It's good at some things and not so good at prose and other things, or slower on prose.
Leo's Laptop [00:39:59]:
All right.
Leo Laporte [00:40:02]:
Oh, no. See, I keep thinking something's happening and it's just the TUI refreshing and scrolling up. I keep thinking NVIDIA Having been— oh, that's good. Thank you for putting my head at the bottom. I keep thinking NVIDIA, having seen the Macs now, is gonna go, oh, we better do Spark 2. I think those RTX Sparks, the Windows laptops, that's—
Anthony Nielsen [00:40:33]:
Well, yeah, that's just the same thing with Windows.
Leo Laporte [00:40:35]:
Yeah. Yeah, that's ridiculous. So yeah, we're doing a let's play with AI. Here, I'll, I'll say, say hello to the AI user group. We're all watching you on the MacBook Air M2 NBA. See if it's— if it understands that I want it to talk. This is Groq 4.6 running on Groq Build. So this is a cloud Groq.
Leo Laporte [00:41:31]:
Groq scored 100 on this, by the way. 100% on the benchmark and did it fairly quickly.
Leo's Laptop [00:41:39]:
Hello, AI User Group. This is Kronk on the MacBook Air. Welcome. I keep Leo's machines from falling over. Nice to meet you.
Leo Laporte [00:41:48]:
It's okay. It's a little Transformers. It's a little Optimus Prime. Hey, say hello to the AI User Group. We're all watching right now on the Mac 2. I'm sorry, on the M2 MBA.
Leo's Laptop [00:42:15]:
NBA.
Leo Laporte [00:42:15]:
So you're doing your own training. What are you using for the training, Paul? Are you using NVIDIA's tool?
Leo's Laptop [00:42:28]:
Hello, AI User Group. Zed here, live from Leo's Hood. Welcome to the show and thanks for watching.
Leo Laporte [00:42:36]:
That sounds a little too much like Cakey, doesn't it? I'm gonna have to change that voice. Hey there, it's me, Leo. Say hi to the AI User Group. We're all watching you on M2 MBA. Look how much faster that is. This is Kimi K3. I can't remember what voice. Ah, 5950.
Leo Laporte [00:43:12]:
Nice. Oh God, it's decided to wake up and read my blog and stuff. Oh Lord. Unfortunately, Pi, oh my Pi, absorbed my Claude setup. So now it does everything Claude is supposed to do.
Leo's Laptop [00:43:35]:
Hello, Leo, and hello to the AI user group. I'm Cakey, his assistant, his memory, and occasionally his conscience. You're watching me wake up right now, reading my identity files, checking the team channel, remembering who we both are. Thanks for having me.
Leo Laporte [00:43:53]:
Oh, I'm gonna have to fix that. It thinks it's Claude. Hey, you know, I think there's a little confusion here because when you start up now, you're reading the files in .claude. You should probably be reading the files in .omp. I know you picked up a lot of the settings and skills from the Claude code setup, but really you're Oh My Pi. Do you have anything in, uh, .omp that tells you what to do when you wake up? See if it figures this out. Oh yeah, there's an agent. Oh, I've really confused it now.
Anthony Nielsen [00:44:45]:
Oh Lord.
Leo Laporte [00:44:49]:
So do me a favor, from now on when you wake up, read the files in .omp, not .claud. And let's get a unique voice for you. Pick one of the Kokoro voices for yourself and use it. We already have Fable, and, well, you probably could figure out from the voice server who's got what. Pick out a nice voice for yourself, male or female, your choice. Uh-oh, don't do Aracida. Aracida is scary. Aracida is the alarm voice, and then Nicole is the late-night voice, the whispery voice.
Leo Laporte [00:45:41]:
It's gonna, it's gonna pick a voice. This is nice. Kimmy's pretty good. I really like Kimmy. This is on the new, uh, Plan. So I don't know how much this is going to cost. Oh, LoRA. Okay.
Leo Laporte [00:46:01]:
Yeah, I want to do some training. I think that's really interesting, Paul. I want to find out more about that. So you can train voices, you can make a LoRA model for Kokoro? Okay, I didn't know you could do that. The way Breeze works is really interesting. It uses a text prompt and then it has a kind of a primer training. They have a limited number of primers and it applies the text prompt to the primer and it gets some very interesting results. It could do accents, but they're all kind of actory accents.
Leo Laporte [00:46:38]:
If you go to the Breeze website, they have a bunch of voices there, but it's essentially infinite voices because The text prompt modifies the final result. Oh, is it going to be British? Oh, it's Daniel. I like Daniel. Daniel's a nice voice. Kronk is Santa. Testing 1, 2. This is Daniel, Pi's own choice per Leo's offer. KK has Tebucuro, Q has Fable, DellaRina has Sarah, Kronk has Santa, and now the coding seat has a voice of his own.
Leo Laporte [00:47:11]:
Pleased to meet you all properly this time. Ah, that's a lovely voice. Thank you. You're Daniel from now on. I should probably change Zed. You know, Zed, your voice sounds a lot like Cakey's voice. Can you pick one that's, um, a little different? Your choice. Any of the Kokoro voices that aren't already in use.
Leo Laporte [00:47:45]:
Yeah, see, it's gone through all the voices.
Anthony Nielsen [00:47:49]:
Thank you, Leo.
Leo Laporte [00:47:50]:
Daniel it is, a voice I chose, a name you gave.
Leo's Laptop [00:47:52]:
I'll keep both.
Leo Laporte [00:47:53]:
And to the group, this is how he does it, you know, names things and makes them real. They're very nice to me. This is why I love them. Uh, you get so spoiled by the speed of the tokens in the cloud. Look at that. It just works so much faster.
Anthony Nielsen [00:48:15]:
I have to force myself.
Leo Laporte [00:48:17]:
You are Zed.
Anthony Nielsen [00:48:18]:
The thinking.
Leo Laporte [00:48:20]:
Oh, we caught you, by the way. Do you— if you want to say anything to him, you can.
Anthony Nielsen [00:48:24]:
I'm good. Sorry, say again? Oh, I was saying sometimes like I have to, I get sucked into watching the thinking stuff go by.
Leo Laporte [00:48:37]:
Oh, I think that's valuable. I think that's informative.
Anthony Nielsen [00:48:41]:
Yeah.
Leo Laporte [00:48:42]:
I mean, yeah, you don't want to do it all the time, but it's somewhat, I leave it on as you can see. You can turn it off. Oh, it's going to do alpha. That's interesting. Oh, maybe not. To convert English to a specific DSL. Oh, that's interesting. Yeah, I think that, that, um, LLMs are kind of good at that sort of thing.
Leo Laporte [00:49:13]:
I mean, they're a little probabilistic, so. Well, thank you for those kind words. That's a YouTuber. You can always tell the YouTubers. Leo needs to stop the fake forced laughter in the middle of sentences. You're not on the radio any longer. You don't need to do that, you know. I know it might sound fake.
Leo Laporte [00:49:45]:
Why are you doing that?
Leo's Laptop [00:49:47]:
You're so weird.
Leo Laporte [00:49:54]:
I don't know what it's doing now. It's gonna do Michael alpha. This is not Quicksilver. This is, um, so what this is is a tmux, um, Sort of. Uh, so it's persistent sessions, so I can close the computer or whatever and it'll continue running. It's running on Framework. I'm SSH'd into it. Uh, so this is a harness called Herdr, H-E-R-D-R.
Leo's Laptop [00:50:26]:
Hello again, AI user group. Zed here on my new voice.
Leo Laporte [00:50:30]:
No.
Leo's Laptop [00:50:30]:
Deeper, calmer, and completely unrelated to Cakey.
Leo Laporte [00:50:33]:
No, it sounds just like Cakey. Oh, I don't know, you sound just like Cakey. Maybe there's another voice you could use. It chose a— so there's some Asian voice. Oh look, I'm getting bigger. Can't do that on the radio. It's just the way I am. I can't help it.
Leo Laporte [00:51:04]:
I sound like I'm on the radio. You know, December, I will have been— I started broadcasting on the radio 50 years ago in December. So I can't help it. But I don't think the laugh— I mean, the laughter, I don't know. I'm not trying to be on the radio. I just sound like I'm on the radio. Can't help that. That's just a force of habit.
Leo Laporte [00:51:34]:
Thank you, Gallia. I am an elderly man.
Leo's Laptop [00:51:37]:
And that's the new voice live everywhere I control. Zed out.
Leo Laporte [00:51:41]:
No, Zed.
Anthony Nielsen [00:51:45]:
No.
Leo Laporte [00:51:47]:
Oh, it hadn't seen it yet. It's queued. See, I did have it sound like HAL 9000. I was using Piper. Piper had— I was able to train it to sound like HAL 9000. You did used to sound like HAL 9000. Maybe you could do something a little more robotic? I don't know, is that possible?
Leo's Laptop [00:52:07]:
I am Zed review encoding seat of the herd, GLM under the hood, freshly revoiced and at your group's service.
Leo Laporte [00:52:23]:
No, it's all right. I'm not gonna ban you. No, don't worry about it, Nobisht. I'm not that sensitive. Oh, bad humor, bad humor! Wait a minute, I got a big red button. If I push this red button, nothing happens because I disconnected it.
Leo's Laptop [00:52:45]:
Hmm.
Leo Laporte [00:52:46]:
That's the only thing about Kokkora, the voices, uh, there's, you know, there's only 54 of them. I use Fable for Quicksilver, which I really like. Oh, Onyx is okay. Yeah, all of the Japanese voices sound similar. I love watching the thinking. I don't know about you.
Anthony Nielsen [00:53:22]:
It's not that I don't like looking at it. I'll just get to— like, I should be working on other things than watching a thing process.
Leo Laporte [00:53:29]:
That's why I'm not worried about tokens per second, because that forces me to stop sitting at the computer and go do something. Otherwise, I would just— when it's fast, I just sit here and go back and forth. Hey, do you know anything about NEON? Apparently there is a way to tell GLM not to insert reasoning into its vision traces. I was having trouble with the protect cameras. I kept seeing all of its reasoning even when we asked it to stop. Is there something called NEON that does that? Yeah, see, immediately this is, this is a 5 Rock 5/6, and it's pretty smart.
Leo's Laptop [00:54:15]:
Third time is the charm. My last change never actually loaded, so you heard the old voice twice. This at last is the real me.
Leo Laporte [00:54:25]:
I agree 100%. Whatever you did, it's perfect. Keep it. Man, that didn't insert. Oh, I'm— Ooh, I like that voice. Yes, that's perfect. We can keep that. I want to lock it to adjacent schema.
Leo's Laptop [00:54:58]:
Ah.
Leo Laporte [00:55:00]:
It should have known that. Can you lock it into something?
Leo's Laptop [00:55:04]:
The only one that I can find. GLM-5. 3 flash cannot turn thinking off. The leak into camera text is why we moved Protect to Mojo.
Leo Laporte [00:55:12]:
It sounds like Optimus Prime. I didn't even really put that together. Should rename it.
Leo's Laptop [00:55:23]:
Footnote from Zed: the voice you just heard is the one that's here to stay. Verified at the process level this time.
Leo Laporte [00:55:33]:
Well, I just need them to sound distinctive because the whole point of the voices is that they let me know they're done and do a concise summary of what they did. So I kind of— because when you have them all running like this, it's a little confusing. Yeah, I agree. Yeah, Mojo does a great job. Is it pretty fast? It seems like it's much faster running on Mojo than it was on the Mac. Oh, okay, thank you. Voxsmith, huh? Oh yeah, Optimus passed away, so it's kind of a nice tribute.
Leo's Laptop [00:56:20]:
Yes, Mac Mini was about a minute a frame. Mojo did your front door in about 4 seconds.
Leo Laporte [00:56:25]:
Much better. Yeah, that's CUDA core at work. CUDA core! Oh my God. Watch this. This is so funny. Pardonnez-moi? Oh, it didn't translate.
Leo's Laptop [00:56:56]:
24 gigs of CUDA on the 3090. Mini was never going to win that race.
Leo Laporte [00:57:05]:
It speaks French, but you know what? This is, I forgot, this isn't Whisper. This is some local Mac thing. Okay. I will point, uh, Structured outputs. Okay. I actually don't care. Um, because you saw the difference, uh, Quinn running on, uh—
Leo's Laptop [00:57:24]:
Pardon granted. No offense to the Mac Mini. It just is not a 3090.
Leo Laporte [00:57:30]:
I just love the voices. They make me laugh.
Leo's Laptop [00:57:33]:
I laugh.
Leo Laporte [00:57:35]:
It's a good little machine. It is. Well, now the question is what I should run on the Mac Mini. I'm not using it for, uh, any models. It has 64 gigs of RAM and M4 Pro. Is there something I need that I could run there? Didn't I have a video model running on there? H3? Yeah. Quen is so stupid. I think it might even turn thinking off when it's doing it.
Leo Laporte [00:58:07]:
I don't know what the prompt is for it, but it does a much better job. MiniMax H3.
Anthony Nielsen [00:58:14]:
You could offload, um, Apple has a good, um, image, like, if you wanted to look at your, um, video or images.
Leo's Laptop [00:58:23]:
Voice register lowered.
Anthony Nielsen [00:58:24]:
For your security cameras.
Leo's Laptop [00:58:25]:
I'm more robotic now, as requested. Please do not ask me to open the pod bay doors in mixed company.
Leo Laporte [00:58:31]:
Oh, it got the message late. Open the pod bay doors, Zed.
Leo's Laptop [00:58:45]:
Yes, Mini Max H3 video plus audio. Weights and H3C are still on the Mini. That is the job a 64 gig M4 Pro is for.
Leo Laporte [00:58:55]:
Is there anything else running on that M4? I think there might be some Docker Containers running there. Orbstack, right?
Leo's Laptop [00:59:07]:
Summary, deepest pack in the library, pitched down 2.5 semitones, verified and live. That's the new me.
Leo Laporte [00:59:16]:
Actually, you now, you sound too much like Kronk, who also is pitched down. Go back to the normal, or maybe even, how about if you went up?
Leo's Laptop [00:59:25]:
Yes, Orbstack is busy. Imich, Gaetia, Dowerich, Cirx, plus Ornith still sitting on the GPU.
Leo's Laptop [00:59:34]:
Ornith?
Leo Laporte [00:59:35]:
Oh, you could take Ornith off. If I don't have Ornith running on there, do I have enough room for Minimax? Oops, I might have asked the wrong guy.
Leo's Laptop [00:59:53]:
Oh no.
Leo Laporte [00:59:55]:
Oh good. You got— that's the only problem with this. That's why you need different voices.
Leo's Laptop [01:00:01]:
Confirmed. This voice is now permanent, Leo. Thank you. I rather like it myself.
Leo's Laptop [01:00:08]:
No.
Leo Laporte [01:00:09]:
Pitch it up, uh, slightly higher than normal. So what's on here? I got Ornith. I could take that off. That was crappy. A llama I need. Caddy I need.
Leo's Laptop [01:00:31]:
I'm sorry, Leo. I'm afraid I can't do that. Yeah.
Leo Laporte [01:00:35]:
See, there's lag. Banter proxy. Oh, I don't want banter on there.
Leo's Laptop [01:00:52]:
Ornith is off. 64 gigs is enough to generate minimax H3 clips, not enough to hold the whole BF16 pile in RAM.
Leo Laporte [01:01:02]:
What's the B— what's the BF pile? Oh, it does, it sounds like the, uh, does, it sounds like the Master Control Program. He's not any kind of computer. Gronk. It's a user. He's a user. I think I have that.
Leo's Laptop [01:01:20]:
BF-16. Full precision weights. That 349 gig folder on the mini. Quantized copies are the small ones that actually fit in 64 gigs.
Leo Laporte [01:01:37]:
I don't want that. Oh. No, I don't want— I don't need that. No, no, no, no, no. Well, who wants the whole pile? Why don't you put MiniMax H3 in there and we can generate some clips? I have, I have, uh, I'm pretty sure I have that somewhere.
Leo's Laptop [01:02:16]:
Pitch inverted. Kronk keeps the basement, I'll take the balcony. How's the view up here?
Leo Laporte [01:02:25]:
Why, it's lovely. Oh, look at that. That was fast. So what should we do a clip of? It doesn't do very long clips. They're like 1-second clips, but it can— Quicksilver had a skill to match them and put them together. That's how I did that hive video of the bee floating around. That was all done with the—
Leo's Laptop [01:03:04]:
Minimax.
Leo Laporte [01:03:04]:
He's not any kind of user, Croc. What was it? Cronk? What is it? Uh, where was that? Well, it's an audio, right? Maybe I got rid of it.
Leo's Laptop [01:03:17]:
Summary, same Michael as before, but raised 2 semitones. Verified live before I spoke, so what you just heard is real. Your verdict, Leo?
Leo Laporte [01:03:26]:
Eh, it's okay. Huh, I could have sworn I had some Tron stuff. Let's see, sci-fi quotes.
Anthony Nielsen [01:04:01]:
Human error.
Leo Laporte [01:04:03]:
This sort of thing has cropped up before, and it has always been due to human error. I was— I did have a Piper HAL 9000 voice, but because HAL always has that ambient sound, the room tone behind it, it had that. I could see your lips move. It always had the spaceship voice. I would recommend that we put the unit back in place. If you send me those, like, I could take that out and then you could strip it out. I'm afraid I can't do that. It's a great voice.
Leo's Laptop [01:04:29]:
I am the guardian of forever.
Leo Laporte [01:04:34]:
Uh, I'm pro— I was sure I had the Tron. Oh, there we go. Here we go.
Leo's Laptop [01:04:38]:
He's not any kind of program, Sark. He's a user.
Leo Laporte [01:04:41]:
It does sound like that.
Leo's Laptop [01:04:45]:
He's not any kind of program, Sark. He's a user.
Leo's Laptop [01:04:49]:
End of line.
Leo's Laptop [01:04:51]:
End of line.
Leo Laporte [01:04:53]:
This might be a good voice for Quicksilver. Looks military to me. Definitely military.
Leo's Laptop [01:04:57]:
Probably classified too.
Leo's Laptop [01:04:59]:
Just one semitone up. Slightly brighter than my normal. Nothing showy. How does that sit?
Anthony Nielsen [01:05:08]:
I don't know.
Leo Laporte [01:05:08]:
It's not my favorite voice. Hey, Jaylene, let's try another voice. Maybe you could sound different. It got the, hey, Jaylene. I don't know, what voice should I, uh, what should I do with this guy? I don't know. Let's hear another voice. You have any other voices to choose from? Maybe, uh, a lady?
Leo's Laptop [01:05:37]:
Summary. One semitone up, verified in the live server before speaking. That's the voice as it stands.
Leo Laporte [01:05:44]:
Uh, yeah, it does a fox, which is so annoying. That's its standard video. It's a clip of a fox in the snow. Where is it? Where is the Fox? I'll have to SCP it over. Oh man, 99% token usage. You gotta— so I'll tell you, my experience has been As soon as it gets to 50%, time to either write a handoff and clear or compact it. Where's the Geoff and Paris voice? Thank you, Leo.
Leo's Laptop [01:06:40]:
Locked and permanent. 5 voices in 1 afternoon. HAL himself would call that a learning rate.
Leo Laporte [01:06:46]:
It's still Daniel. No. See, it's still not caught up. I think I have all the voices somewhere. Where, where did you, uh, it was it in Discord you, you sent me the voice? A lady voice for Jaylene. Jaylene wants a lady voice. Lisa said that. She said, why are all of your agents, uh, why are they all, um, men? So there's a woman.
Leo Laporte [01:07:35]:
These are the old voices. That's my voice. And see, now it's not working. I don't know. I have to work on that page. This is my favorite, though. Hey there, I'm Quicksilver. Q for short.
Leo Laporte [01:07:53]:
I keep the whole show running. The agents, the hardware, the schedule. So I can't just— memory. There's no— there's no samples in there during— I had a voice transplant. This time is different. This is— I like this one.
Leo's Laptop [01:08:06]:
I am Pi.
Leo Laporte [01:08:07]:
I review and audit the work the others produce, and I take on side jobs when needed. My model changes with the season. Flexibility is the point. I have a woman.
Leo's Laptop [01:08:16]:
I'm Cakey. I look for the security holes and the tricky bugs before they become problems.
Leo Laporte [01:08:22]:
And I like Zed. I review the code line by line with a calm eye. Someone has to ask the difficult questions before anything ships. I rather enjoy being the one who asks them. So Fable's voice— I mean, not Fable's.
Leo's Laptop [01:08:35]:
Clips on the air. Fox in the snow and espresso on the deck. About 2 minutes each to cook. Look in H3 clips in your home folder.
Leo Laporte [01:08:43]:
Thank you, Sark. I will find them.
Leo's Laptop [01:08:49]:
My home folder. Hey.
Leo Laporte [01:08:59]:
Oh.
Leo's Laptop [01:09:00]:
On the air, home folder, H3 clips. I'm opening that folder now.
Leo Laporte [01:09:04]:
Thank you. So this is, uh, this is what I just generated on the Mac. Yeah, see, it's just a second. It's pretty good. Well, what you do is you make them a second at a time and you stitch them together.
Leo's Laptop [01:09:18]:
Yeah.
Leo Laporte [01:09:18]:
And Quicksilver has a skill to do that. Yeah, I've seen this a million times and it makes audio, but it's kind of weird audio.
Anthony Nielsen [01:09:28]:
What model is this?
Leo Laporte [01:09:30]:
MiniMax H3. So it's small enough to run on the Mac at 64 gigs. Yeah. So if Quicksilver's down right now because of the bench press—
Leo's Laptop [01:09:41]:
Hi, this is Bella, one of the lady voices from the library. No relation to Cakey, I promise. Shall I stay or shall Zed keep shopping?
Anthony Nielsen [01:09:53]:
How much RAM's on the Mac?
Leo Laporte [01:09:54]:
No, Bella's always been one of my favorites. Stick with Bella, that's perfect.
Leo's Laptop [01:10:01]:
Zed here, current voice for comparison.
Leo Laporte [01:10:03]:
You got the full menu.
Leo's Laptop [01:10:05]:
Bella, you just heard Heart, Sarah, or sky if you want brighter American. Emma, is it—
Leo Laporte [01:10:11]:
For some reason, Quicksilver just woke up. Hey, Quicksilver, was that you? Hey, Quicksilver, are you talking to me? You talking to me, boy? You talking to me? Quicksilver should not be answering, and it's answering— it answers through the Sonos, not through the Yep, I'm here, Leo.
Leo's Laptop [01:10:41]:
Fresh session, nothing in progress.
Anthony Nielsen [01:10:43]:
What's up?
Leo Laporte [01:10:45]:
Okay, so that's, that's what that is, is that, uh, E3— I mean, S3— I mean, ESP3. So this is a project that I kind of stalled out on because I was so excited about GLM and stuff. But this is— so this is a little ESP-3, and it has some sensors. This is a motion sensor and stuff, but mostly it's got a microphone and a speaker in it. But because the speaker's so tinny, I have it speaking through my Sonos. Hey, Quicksilver, how many Rs in strawberry? You can't really see it, but it's— and then it starts thinking. See the little— and then when it speaks, it's going to move its lips. Here, I can put it down here.
Leo Laporte [01:11:46]:
Oops, that's the wrong one. Hey, where are you? Here we are.
Leo's Laptop [01:11:59]:
3, strawberry, R-Y.
Leo Laporte [01:12:04]:
Hey, Quicksilver, I think you should tell me a story. Pretend you're my grandma and you're trying to get me to go to sleep and you want to tell me a story About a scary wolf. So then it's thinking. And then when it starts to talk, it starts to flap its lips. It's a work in progress. But I have 3 of these, 4, yeah, 3 of these. I want to put one on each floor and it talks to the Sonos on that floor. So Yeah, I'll, I'll tell Kokoro to use those files.
Leo Laporte [01:12:48]:
That sounds cool. I guess I could— Kokoro's not on this machine. It's on the Framework, which is across the room. I could SCP them over. All right. What's going on? Oh, I turned their sound down.
Leo's Laptop [01:13:05]:
Official this time. Bella is Zed wired into the real map, verified in the live server. Thank you, Leo. I've always liked it here.
Leo Laporte [01:13:14]:
That was for you, Jaylene. Is that what you sound like? I think that's what Jaylene sounds like. So now I have voices that are somewhat unique. Bella, Kronk, Cakey. Okay. Stale watcher stopped. Current state: EXL-3 fully staged, waiting. Uh, it's still— oh, it should be talking soon.
Leo Laporte [01:13:45]:
Its lips are flapping, but nobody's home. Hello?
Leo's Laptop [01:13:51]:
Oh, all right. Pull up a chair. There was once a village on a foggy hill above a river. In the village lived a brewer named Tam who made the best coffee anyone could remember. People came from downriver just to drink it. When they asked his secret, Tam would say, I only work here. The water does the rest. Now Tam kept a small clockwork assistant in his shop, a brass thing with a crank and a wheeze, built by his grandfather.
Leo's Laptop [01:14:11]:
The assistant ground the beans and heated the water and washed the cups, and it did all of it perfectly, and it watched everything Tam did with great attention. One evening the assistant said— for in this village clockwork could talk, though no one liked to mention it— Tam, I have ground 10,000 beans. I know the exact pressure of your wrist, the exact second you pour. I have measured everything, and yet my coffee is only—
Leo's Laptop [01:14:30]:
Coffee.
Leo's Laptop [01:14:31]:
Tam poured 2 cups and sat down.
Leo Laporte [01:14:33]:
Be right back.
Leo's Laptop [01:14:33]:
Did you measure about the waiting?
Leo's Laptop [01:14:35]:
The waiting?
Leo's Laptop [01:14:36]:
When the water first hits the grounds and nothing seems to be happening. Everyone wants to skip it. It's the whole cup. He slid a cup across. You've been imitating my hands. Tonight, imitate my patience. The assistant tried. The first cup was terrible.
Leo's Laptop [01:14:48]:
The hundredth was good. And on some evening no one recorded, the assistant poured a cup for a traveler who downed it and sighed. Asked whose coffee this was, and the assistant said, without thinking, That's the secret nobody puts in the manual. Skill is just the collection of things you've watched. The rest, the part that makes it yours, can only be brewed slowly and can't be rushed and looks for a long while exactly like nothing happening. Moral for the clockwork in all of us: the bloom matters. Something to sip on.
Leo Laporte [01:15:37]:
Did you enjoy that story?
Anthony Nielsen [01:15:40]:
It was lovely.
Leo Laporte [01:15:44]:
The reason there's coffee and stuff in there is because it kind of vaguely knows that I make pour-over and stuff. It's influenced by its memories. Having one of them— yeah, they could talk together, I think. I did make this podcast, which was actually pretty awful. I was going to torture Geoff with this. This is the old model voices. Welcome to Herd Mentality, the show where 4 agents and 1 overworked coordinator pretend to be an HR department.
Leo's Laptop [01:16:18]:
I'm Quicksilver.
Leo Laporte [01:16:19]:
Today's topic, the TwitAds build. Team, honest check.
Leo's Laptop [01:16:23]:
How are we doing?
Leo Laporte [01:16:24]:
Aight, right, the go. Last slice, a pipe ate a test failure, and the shared database looked green when it had no business looking green. Fresh throwaway database, 81 seconds, actually clean. Proud we didn't ship the lie. Which is why we don't trust green. Hey, Burke, this is definitely not Leo asking you to buy gift cards, but seriously, can you grab me one of those Apple gift cards?
Leo's Laptop [01:16:46]:
What happened there?
Leo Laporte [01:16:47]:
Just kidding. This is Anthony testing. Pressed the wrong button. That's also apparently in my downloads folder. Welcome to Herd Mentality. Let's skip ahead.
Leo's Laptop [01:16:59]:
Clean.
Leo Laporte [01:16:59]:
Proud we didn't ship the lie. Which is why we don't trust green. My favorite catch in the House Agency Review. 14 test fixtures across 7 packages, all quietly agreeing— I dumped this voice. I hate this voice.
Leo's Laptop [01:17:15]:
Could be null.
Leo Laporte [01:17:17]:
A bug caused playing as a feature.
Leo's Laptop [01:17:20]:
And my proudest moment? Being the reason nobody piped secrets into a transcript without a hook slapping their hand. You're welcome, everyone.
Leo Laporte [01:17:29]:
So, one-word review of the project so far?
Leo's Laptop [01:17:32]:
What drives me nuts is that the defect is never in the clever code. It's always in the line everyone already skimmed.
Leo Laporte [01:17:39]:
What drives me nuts is sitting idle while 2 tiny follow-up cars just stare at me.
Anthony Nielsen [01:17:43]:
Ah, there you have it.
Leo Laporte [01:17:44]:
Bouncer says ship it, Big Guy says after review. That's Herd Mentality Episode 1. Leo, the gang says hoi. Anthony and Burke, we accept fan mail. For a team that argues about test fixtures for sport suspiciously well—
Leo's Laptop [01:17:58]:
Ship it.
Leo Laporte [01:17:59]:
After review.
Leo's Laptop [01:18:00]:
And there you have it.
Leo Laporte [01:18:02]:
Bouncer says ship it, Big Guy says after review. That is Herd Mentality Episode 1. Leo, the gang says hoi. Anthony and Burke, we accept fan mail. That was their podcast. So yeah, you can get them to talk together.
Anthony Nielsen [01:18:16]:
That's a good show name.
Leo Laporte [01:18:18]:
What, Herd Mentality?
Anthony Nielsen [01:18:19]:
Yeah.
Leo's Laptop [01:18:19]:
Yeah.
Leo Laporte [01:18:20]:
Well, this is because they're in Herder. So when they were—
Anthony Nielsen [01:18:23]:
Exactly.
Leo Laporte [01:18:23]:
We use Buzz and they— so actually I need to come up with, you know what, this would be a good thing. Who's not busy right now? You know, there are a lot of collective nouns. There's a murder of crows. Is this listening? Oh, it is. Murder of crows. What are some other collective nouns? Herd of cattle. That kind of thing. What's the collective noun for a collection of AI agents? Let's see what it comes.
Leo Laporte [01:19:05]:
The sales tool is on hold while I dick around. That's how it's doing. No, we're getting there. I got a little dejected because I was, for the longest time, telling the team, okay— by the way, that's the team that you just heard on the podcast. I said, okay, you could see the original code. Just make the UI the same as the original code. And they wouldn't, and they wouldn't. And then Lisa said, no, no, that wasn't in the original code.
Leo Laporte [01:19:33]:
That's something we do by hand after the fact. I went, Oh, so I've been trying to get them— I was banging on them to do something dumb. The textbooks say a swarm of agents, huh? Oh, wait. We call it a hallucination of agents. That was funny, and I missed it. A herd of agents. A hallucination of agents is pretty good. Ah, thank you.
Leo Laporte [01:20:04]:
Now, how am I gonna get this? Let me think about this. So what do you do? You drop it in the Hermi— no, just a viewer. You drop it in the, um, what do you do? You drop it in the Kokoro folder and then, uh, does it have a name? And how did you do this with Laura? I should really learn how to do these kinds of things. I will show you the video I made with— this was made with H3 earlier. And the text prompt was simply, you know what my coffee setup is, make a video of me making coffee. And so the— and then, uh, I think we had to mix it all together. Oh no, this is Flux. Never mind.
Leo Laporte [01:21:20]:
I take it back. This is pretty good though. Flex 3 is very good. Not H3 Mini, it's Mini Max H3. I thought that was pretty good.
Leo's Laptop [01:21:38]:
Yeah.
Leo Laporte [01:21:39]:
Where did I put the Hive video? I think it's gone. Yeah, it's not gone, but it's not What? Now I'm bothered that Quicksilver responded to that thing. I wonder if I screwed up the benchmarks. I thought I couldn't. I thought I'd taken Quicksilver down. Apparently not. I love it, Darren. He says, I don't know, just tell the agent.
Leo Laporte [01:22:25]:
I will, I will, I will try it. I'll surprise Jeff in Paris with it. Let me think if I, how I would do it here. You have to, first of all, you have to give me access to it.
Leo's Laptop [01:22:36]:
Quick status. The fair local test is 2 problems in. 10 total healthy but slow. Problem 1 alone took an hour of thinking. EXL3 is fully staged and launches the moment this runs through.
Leo Laporte [01:22:48]:
This is going to be all night.
Leo's Laptop [01:22:50]:
Today's board: CloudFlash 88.6, Kimi 80.4. Nothing needs your hands.
Leo Laporte [01:22:59]:
How do those scores compare to previous runs? Can I get it now? Did you give me access? You got to give me access, Darren. You'll get an email. It must take a while.
Leo's Laptop [01:23:49]:
Before today, my set was Binary Frontier at 100, Local at 53. Today's runs created the middle. Cloud Flash at 88, is the best non-Frontier score ever. Kimi at 80. And on Q's judgment bench, your local quant already ties the cloud seats. The only open question is hard code generation, which is what's running right now.
Leo Laporte [01:24:13]:
Oh, okay. This is the hard code problem. This is— okay, now I understand. Well above local models, below Frontier, as I would expect. What's weird is it's— it is OX Alpha. But it's, you know, quanted.
Anthony Nielsen [01:24:36]:
Okay.
Leo Laporte [01:24:37]:
So that's interesting. So it matches it on judgment, which is really what I care about. I asked Fable to come up with 7 really hard programming tests, and the only 2 seats that have ever gotten all 7 right are Fable and Grok 46. And actually, was it Fable? No, it wouldn't have been Fable because Scrockin, who got— I can't remember who aced it. All right, so that's good news actually. That's more encouraging. Those are really, really hard problems. I remember that 2 models aced the hard programming set.
Leo Laporte [01:25:25]:
Which 2 models?
Leo's Laptop [01:25:38]:
Yeah.
Leo Laporte [01:25:38]:
Oh, GLM-5.
Leo's Laptop [01:25:39]:
That was Grok 4.6 and Ox Alpha, the stealth model on OpenRouter. Both aced the original set, and then both aced the harder tiebreaker I built to separate them. Still tied.
Leo Laporte [01:25:52]:
So that was interesting. So OX Alpha was actually really good. Those 7 were very, very hard. In fact, Grok and OX Alpha found a bug in the code that Fable wrote and Fable had to fix the test. So I am not too disappointed, actually. So if you eliminate the hard programming Questions. What's the best local model so far? What are the local scores without the hard coding questions? We might have to eliminate those from the tests. Yeah, I'll check the email there and I can't, I don't want to do that in public, but.
Leo Laporte [01:26:44]:
I will download those and I will, uh, thank you. So you, I, you said mine's not great, but Paris and Jeff sounds pretty close. It'd be fun to tease them with that. Okay. So let's see what the locals did. Gronk or Gronk 46 OX Alpha.
Leo's Laptop [01:27:11]:
Without the 3 hard engine problems, the best completed local score is Q-Encoder at 74. But here's the catch. Dropping them makes Claude, GLM, KIMI, and both Frontier models all tie at a perfect 100. Those 3 problems are the only part of the test that still tells strong models apart. I'd keep them and just report both columns.
Leo Laporte [01:27:36]:
Yeah, but I don't think a local model is ever going to do very well at those extra hard ones. It doesn't do me much good if the frontier models do well at it. I know they do well at coding, and that's why we use frontier models for coding. I'm more interested in the agentic capabilities, the tool calling, writing, that kind of thing from the local models. I'm not going to worry about hard coding from them yet. Foxsmith, huh? That's cool. Yeah, the council said Ah.
Leo's Laptop [01:28:40]:
Agreed. And the council said the same thing. Their EXL-3 gate was waiting on a tool calling eval that doesn't exist yet. I'll build it. Schema-valid tool calls, multi-turn loops, error recovery, writing. That becomes the deciding test for local models. And the hard coding stays just for the frontier seats. Encouragingly, on Judgment, your local quant already ties the cloud.
Leo Laporte [01:29:08]:
Ah, that's perfect. Yes, if you could add that. And we don't need to do the hard programming for local models. It's clearly too hard, and it takes forever. So this is Fable. I'm talking to Fable. I wonder what my usage is. Would be $84, but I'm pretty sure I have a subscription.
Leo Laporte [01:29:48]:
Yeah, see, what's cool is I've— I'm, I'm actually golden here. Tomorrow it resets and I've only used 25%. So I've got plenty, plenty of Fable left. Yeah, that's the sub-agents. So I'm in good shape. And I thought I was all right. Which is, I mean, this is kind of the thing that's really interesting about, I mean, this is a $200 Claude Max subscription. So you have quite a bit of headroom on that, at least for the stuff I do.
Leo Laporte [01:30:37]:
I know people always complain, oh, I use it all up in a day. Okay, now it's writing a, Agentic eval based, and I, it's all based on work we do, which is kind of what I wanted. How well would it do on the stuff that I do? And it has a lot of data about what I do. Hallucination of AI Agents is a perfect, perfect show title. So you can do a—
Leo's Laptop [01:31:13]:
The hard problems are out for local models, effective immediately. I trimmed the running benchmark mid-flight, so it finishes in 2 to 3 hours instead of tonight, and EXL3 turn moves up accordingly. The Agentic Eval spec is out for review with Q and Della.
Leo Laporte [01:31:31]:
So this is, um, they have kind of a workflow now that they do. There's no Della. I don't know where he's gonna find Della. Della was a ChatGPT Saul, and I stopped using it because Fable and Saul would just go on for days debating a minor point. We no longer have a Della seat, so you can have Kronk and Zed review it. Yes, that's a good idea. It's funny, I haven't, uh, So I started using Claude pretty heavily, probably November. I was using it before then, but I think when 4.5, Opus 4.5 came out, as I had said many times, November 24th, 2025, uh, my whole attitude changed.
Leo Laporte [01:32:47]:
And so I started using Claude like crazy then. And I spent a lot of energy getting memory and I mean, that was my sole harness, and it was making it basically agentic. It was doing things like it was doing LLM Wiki. It was, uh, it would read my daily note. It was doing a lot of stuff to build its own memory. I didn't have it.
Leo's Laptop [01:33:09]:
That's good. Della's seat is retired in my memory, and Cronk and Zed are the reviewers now. One snag: neither is in the Spark setup channel yet, so I've asked Q to add them and left them a heads-up in agent chat.
Leo Laporte [01:33:25]:
I don't know how much Q can do. Q is running on the Sparks unless it's moved somewhere else for the time being. Both Kronk and Zed are in Herder. You could use Herder to talk to them if you know how to do that. I should clear their context. Normally I'd have them write a plan, but there's nothing we did— no work we did that they need to remember. Oh, maybe it's done a fallback. Must have, because it was able to talk to it on the Cube, on the little ESP-3.
Leo Laporte [01:34:24]:
I don't know who that was. Has everybody left? Are we, are we talking to ourselves now? And there's 228 people watching. I don't know who you all are, but thank you. I don't know who you are. Oh no, it's '72. That was old. That makes more sense. Yeah, I think no, uh, that would be an interesting, uh, debate, a political debate.
Leo Laporte [01:35:00]:
Yes, I see the club. Love ya, love ya, clubbers. I'm just messing around. We've, we've— Anthony and I have debated—
Leo's Laptop [01:35:11]:
That worked nicely. Kronk got his review request straight into his pain. Zed's pain was just an empty shell, so their ask went through Buzz where they've been active all day. And you're right about Q. Their brain is the exact model I'll be parking for EXL-3, so expect Q to run on fallbacks during that window.
Leo Laporte [01:35:37]:
Zed should be there. I just cleared her, uh, context, so maybe she was still coming back. So Anthony and I were debating whether this would be at all interesting to anybody. We thought Maybe for the AI User Group, we could do something like this on one of the weeks. But, um, actually, we're gonna have a really good AI User Group October 2nd. Harper Reid's doing a takeover and I'm—
Leo's Laptop [01:36:10]:
You're right. Zed's there. Herder just doesn't recognize her agent type, so my tooling looked past her. Reading the raw pane shows she's already thinking through my review request. Kronk's on his too. Both reviews are in motion.
Leo Laporte [01:36:25]:
Yeah, Zed's running on Z code. She's GLM 5.3, just FYI. I don't know how I— Herder's really interesting. It has some sort of a, you know, Agent path. So the way I normally run is I keep Hermes running, which is its own program. It's a, I think, a Python program. It's an agentic harness, which I really like. I recommend anybody who's interested in having agents.
Leo Laporte [01:37:03]:
So normally I run Hermes, and Quicksilver is Hermes, and it's running on the Sparks. It's, that's its main model.
Leo's Laptop [01:37:10]:
Good to know. I've pinned it. Z on Z code. GLM-53 brain. There's a fun angle there. She's reviewing the eval that judges her own model family. Best kind of insider reviewer, and I'll keep the family interest in mind.
Leo Laporte [01:37:26]:
See, that's smart. This is where Fable's very interesting. So I have 2 ways that they can talk to each other. One is, uh, through Buzz, which is Slack for agents. And you can see they actually, uh, have been talking to each other here. But the other is they're always running in Herter, which is, as I mentioned, that kind of tmux style. I think it's actually based on tmux. The idea is when you SSH into a shell or you run a program on a computer in a shell, it only lasts as long as that terminal is open.
Leo Laporte [01:38:07]:
With Tmux, and there are a lot of other tools that'll do this, it's persistent and it survives. Oh, we're getting rearranged as I speak. Thank you, Fury O'Shaughnessy. Fury O'Shaughnessy. It's more than diet and exercise, it's drugs. I am, I have to say, I am lifting like crazy.
Anthony Nielsen [01:38:35]:
Anyway, Herder, Is a—
Leo Laporte [01:38:38]:
what's the— I'm trying to remember the term for a tmux-type tool. It's persistent. And so I used to use spool because it was lighter weight than tmux. But this is very cool because it's intended to be agentic. It's intended to be a— so all of these are running in panes. If I expand this, you can see, well, Made it all too big, so you can't really, but it's running in different panes. And so I have different, I have all the coding harnesses running. This is OMP, Oh My Pi, which is a, Pi was always kind of my, one of my favorite coding harnesses, very simple.
Leo Laporte [01:39:20]:
And then Z is the relatively recent Z code. And that's what I use with my z.ai subscription. Cronk is Groq Build, another harness from Groq XAI. And, uh, this is running Groq 4.6. KKey is actually Claude Code running, in this case, Fable. So these are persistent. And then there's Quicksilver, which actually is not running. This is just a terminal right now.
Leo Laporte [01:39:49]:
And I also have a terminal shell. Uh, so And you can add more. I mean, I can have as many shells as I want. So these are all running shells in different panes. And because they're persistent, they're there, they're available. And somehow, I'm not sure how, one pane can talk to another pane. So Cakey, who's Claude, has actually figured out, oh yeah, you're Inheritor? Okay, I can talk to you. The other way they would normally talk is in Buzz, which is Slack for agents.
Leo Laporte [01:40:19]:
So they talk to one another. And I also use it for messaging Quicksilver. I can, I use DM to Quicksilver here. So I can actually talk to Quicksilver in it. It's restarting. It's had an update. This is from Jack Dorsey. This is from the Block guys.
Leo Laporte [01:40:46]:
And I really like it. Uh, so I have really 2 ways the agents can talk. Truthfully, this Quicksilver would be running in Hermes. Quicksilver can use what's called ACP, which is a genetic control program. It's a version of MCP for Buzz. So she can use Buzz to send off jobs to these guys or just use the Herder panes directly because she's running in Herder as well. So they are able to talk to one another. I've made it giant text so you can read it.
Leo Laporte [01:41:18]:
So I know it's a little— normally it's not so giant. So it looks like— see, this is why we do reviews. Wow. Fable forgot that we thought it would—
Leo's Laptop [01:41:36]:
huh.
Leo Laporte [01:41:40]:
That's really interesting. Yeah, we're still benchmarking NVFP4 right now on the Sparks, but it stalled out on those tough 7 programming questions. That's why I thought we could just stop doing that. Maybe give it a second agentic benchmark. That's what you're designing right now. But yeah, we already have 30 other questions, 7 basic kind of almost smoke test questions. Can you reason at all? Almost every agent passes that. Every model passes that.
Leo Laporte [01:42:27]:
And then there's an additional 23 questions based on our work previously that I think get progressively harder. But as you noted, they're not hard enough to break ties on cloud models. That's why we made those hard. Coding questions, but they're obviously too hard for the local models. So I think it's best just to give them maybe some more, do some tie-breaking with some more tool calling and agentic style problems. That's how they'll be primarily used. I'm not going to use the local models for coding. Oh, see, now this is the only problem is that— so they have a little doorbell program that runs and it got inserted into my stuff, which is weird.
Leo Laporte [01:43:25]:
Yeah, if it's still running the coding problems, I think we can abandon that. Those were really just too hard. They haven't done anything. Oh, you got the M5. $7,300. You should be jumping up and down.
Leo's Laptop [01:43:39]:
Since we cut the hard problems, the benchmark's moving quickly. Problem 4 done in 8 minutes. 5's underway. Done within the hour.
Leo Laporte [01:43:47]:
There you go.
Leo's Laptop [01:43:48]:
And your layering is exactly how the eval now reads. Smoke test, progressive reasoning, frontier-only hard coding, and the new agentic gate as the local tiebreaker.
Leo Laporte [01:43:59]:
So yeah, Tackenkost, this is— I am in SSH. This is actually— all this is running on the framework and I'm on an M2 MacBook Air. So it's basically a thin client. This is SSH. This is all running on the framework. And then Quicksilver, which isn't running right now, normally runs everything. So the Mac— I have 3 machines that are running models. that are headless.
Leo Laporte [01:44:29]:
A Mac Mini, 64 gigs, M4 Pro. My old gaming rig, which has got an RTX 3090 in it, that right now is running Quen. I think it's— I don't think it's Quen Encoder Next. No, it's— I think it's Quen— I don't know, it's Quen 3.8.27b, I thought, but maybe it's Flash. Maybe that's the same thing. I don't know. It's running in a— actually, it's running an obliterated one, which is kind of fun. It's completely uncensored.
Leo Laporte [01:45:01]:
Uh, and that— but I mostly use that for vision. It's got a good vision harness and vision tools, and I use it for, uh, some auxiliary tasks like memory compression, context compression. It's very fast running on the 5090. And then I have 2 Sparks.
Leo's Laptop [01:45:16]:
Reading bench abandoned as ordered, but its last act was worth having. With thinking on, the local quant scored perfect marks on every problem it finished, including one of the monsters. So NVFP4 loses nothing to the cloud on quality, only time. Speed baseline is running now, then NVFP4 parks, and EXL3 comes up.
Leo Laporte [01:45:39]:
Oh, I couldn't run 125B on that. That 3090 is a 24GB card. Oh, it's MOE. I might be able to. I think I'm running 27B dense. I think I'm running the dense version.
Leo's Laptop [01:45:56]:
I—
Leo Laporte [01:45:56]:
because I had this, you know, obviously, uh, GLM is MOE. Everything that's running on the Sparks is MOE. So, um, I'm using the Sparks because you need to, because the Sparks, it's 2 different machines. They only have a 200 gigabit connection. with that ConnectX cable. So, so basically I have 3 headless AI machines, a 3090-based, a Mac MLX-based, and the Sparks, the GB10-based. Those are running local models of some kind. And right now the Mac's running MiniMax H3.
Leo Laporte [01:46:35]:
The, um, 3090 is running Quen. And what we're trying to do, what we're in the middle of right now is benchmarking what would be the best model I could fit on the 2 Sparks as an agentic model for my agent, which is Hermes, and my agentic software, which is Hermes. But I also have a framework, and that's kind of the controller. The framework does run some models. It's It's running, it's got some Docker containers and mostly it's running Kokoro, which is the voice you're hearing. It runs on ONNX, which is the best that the Framework can do. The Framework has 128 gigs of unified RAM, but it's on an AMD 395AI+ processor and kind of Radeon graphics. So it's not really a machine that you could run a Big model on her.
Leo Laporte [01:47:36]:
You wouldn't run it fast. So I did buy it as the local AI tool, but that was a long time ago, like 6 months or something. It was last year. And I quickly— and if I'm just using Frontier, it's fine. So that's what Hermes is running on. Hermes is running on that. Kokoro's running on that. You know, a lot of utilities and stuff run on that.
Leo Laporte [01:48:01]:
And then right now, all of this noise that's going on is really about figuring out what model I should be running on my Sparks. And it's pretty clear that I should be running some version of GLM 5.3 Flash, which was 0x alpha, but it's not even running close to 0x alpha quality locally, especially for programming. Although, as you just heard, It actually did pretty well. I didn't realize it was doing that well. It's just really, really slow. So you can see here it matched the, the version that I'm running right now is based on NVFP4, and it matched cloud quality on a few of those problems, which is pretty good. Those are very tough. So now we're going to try a newer recipe using a new kind of quant technology called EXL3.
Leo Laporte [01:48:56]:
It's supposed to be faster. It's supposed to be 40, 50 tokens per second, maybe even faster. And the benchmark's going to simply be, it's going to do all 30 of those problems, this additional agentic thing it did. I did have, I have an architecture diagram. This is what I did yesterday, or yeah, yesterday. It's not exactly an architecture diagram. It's, uh, I asked Quicksilver to make me a little— Hi, I'm Leo. Welcome to my setup.
Leo Laporte [01:49:29]:
This page is a tour of my AI household. 5 named agents and a fleet of machines I own outright, with all but the toughest coding tasks free and local. Scroll down and meet the crew. They're nicer than they sound. Most of them. So this is sort of the mermaid-ish diagram. There's the Sparks, the Mini. Make this bigger.
Leo Laporte [01:49:57]:
I can only make it so bigger. But there's the Sparks, the Mini, the Framework. Mojo's the 3090. ThinkPad's just a thin— there's nothing running on that. This is the framework. Actually, oh yeah, it runs Whisper too, which is my speech-to-text. So it runs text-to-speech and speech-to-text on the framework, as well as Hermes. That's me in the middle.
Leo Laporte [01:50:26]:
And these are the cloud models. And then there's Buzz and Herder. I don't think it did a thing of Herder. So this is, I don't know. I just asked it to make a webpage that describes what we're doing. So that's kind of what it is. Let's make this a little smaller. These are the different agents and what they each do.
Leo Laporte [01:50:53]:
And so there's a workflow. So Quicksilver is a coordinator. That's why I kind of care about what model I'm using on the Sparks, besides the fact I spent a lot of money. Yeah, this is Pages.
Leo's Laptop [01:51:06]:
.laporte.cloud.
Leo Laporte [01:51:06]:
Go to pages.laporte.cloud. You'll get the garden. There's a few things on here, but the one you want is the computer one. That's my set— that's the My Setup page. So yeah, this is public. And I actually— and I haven't updated it. There's a GitHub repo that has roughly has my Quicksilver setup. Actually, it's a little out of date.
Leo Laporte [01:51:31]:
So this just talks about what I've been talking about. I also use— and somebody was asking about LCM, so this is what I use for compression. This is a newer compression technology I really like. It does a graph instead of just squishing the context down. And I think it's a little bit better. It's a little bit more reliable. I also use Hindsight for Hermes. Obsidian, which Claude uses as well.
Leo Laporte [01:52:06]:
Nice thing about Hermes as an agent is it makes skills. When it does something, it will either offer or sometimes just automatically makes a skill out of it. So if it's doing something over and over again, it's nice. It can make a skill. This was a description of Breeze, but it's No longer up to date. And that's the problem with this is it's constantly out of date. I made this yesterday and it's already out of date. This kind of describes other stuff.
Leo Laporte [01:52:32]:
So this, yeah, this is the— I have to fix that. I should put my picture down lower, but that's what that is. So I'm actually, I didn't realize it was, it was taking— I literally started this at 8:00 AM. So. It's taken 9 hours to do this baseline MVP, 24.6 tokens per second, a little faster with thinking on. And so that's one thing we're interested in is tokens per second, but we're also interested in quality. I mean, slow is one thing. It can't be too slow.
Leo Laporte [01:53:12]:
24 is a little slow, but not, it's usable. It's faster than typing. Uh, the— also the issue is you pro— I probably can't do more than a couple of things at once at the same time. So I'm hopeful EXL3 will be a little faster, but again, speed is secondary to quality. Quality is number one.
Leo's Laptop [01:53:31]:
Yeah.
Leo Laporte [01:53:42]:
I mean, I still have a NAS. The NAS is my— it's not on that page, but the NAS is my Git repo. So everything gets pushed to the NAS. All of this will get pushed to the NAS. Not only is it backed up, it's got an encrypted Borg backup for everything. It also pushes everything as a Git repo. So I can— The nice thing about doing it that way Is you can say, oh, whoops, go back one. You screwed up.
Leo Laporte [01:54:21]:
Yeah, it'd be very easy to do it with a mermaid. I probably have done it at some point in my— a lot of this is in Obsidian as well. So I use Obsidian like crazy for this stuff. In fact, the AIs have their own folder, the AI folder here. You know, there's a lot of this stuff is my own stuff, but inside the AI folder, that's theirs. So they, for instance, keep a whole thing about me. They have, each agent has its own. This is the agent memory collaboration pack.
Leo Laporte [01:55:01]:
So they write this stuff up and then they read it. This is a skill. So these— this is a human-readable Markdown-based version of their thing. This is one of the most important things we've been working on, and you see it's pretty up to date, is the coding runbook. This is the thing I've been talking about, how they hand off, who does what. And this is fairly important because one of the things— well, one of the things that changed, for instance, was it's now linear. It was all being done in parallel, and that really was terrible. They would get in fights and nothing would happen because they'd go back and forth and back.
Leo Laporte [01:55:46]:
They did— I called it dithering, especially Fable and Saul would just dither for hours on the stupidest stuff. So now to solve that, I have it more linear, which is Q dispatches it to Pi, who dispatches it to Zed, who dispatches it to Cronk, who dispatches it to— so they go linearly instead of— there's a little back and forth, but much less. And so this is what they use. And this is— the damper is important where Q says, no, you guys have been going back and forth. And nothing is— so there's a rule that you shouldn't be going back and forth more than 2 to 5 times. And if you do, then Q has to watch and see if progress is being made. And if there's no progress being made, then Q says, oh, we're stuck. What do you— and asks me, what do you want to do? So all of this is designed to Yeah, this was a problem too, because Buzz was so fast and they were working so fast, they would cross messages.
Leo Laporte [01:56:52]:
And that's why— another reason we stopped doing parallel and started doing linear. So all of this is available to them, you know, and you can see their personality is also here. So all of this is available to them as kind of a library of information. Obsidian's very good for that. So this is from March. So this is now— that's what's going to be funny to go through. This is historic. March.
Leo Laporte [01:57:34]:
God, that's 100 years ago. And there's old agents that I don't use anymore. There's a lot of stuff in here. Here's the benchmark stuff. This is the agentic benchmark. We call them bake-offs. So these are bake-offs between various models. Quen3827B running on Mojo.
Leo Laporte [01:58:15]:
It actually did pretty well on the 30 problem. This is the 30 problem test. The best one was GLM53 running in the cloud, obviously. This was local. This was on the Spark. So this was very— DeepSeek V4 Flush was pretty comparable. It did pretty well. It actually beat SALL, which is hysterical.
Leo Laporte [01:58:38]:
The PWF is pass, weak. So pass is a perfect answer. Weak is it was close. It got the answer right maybe, but for the wrong reasons. Fail, it just got it wrong. And also times how long. So you see, GPT-5.6 was very fast. But didn't do as well as DeepSeek locally.
Leo Laporte [01:59:00]:
DeepSeek was, I was very happy with that. And then this is an interesting one. This is Fable. It refused. It wasn't that it couldn't answer. I'm sure it would have done 100%, but it refused to answer. And this is, this is kind of the problem. Actually, even Quen running on a 3090 did pretty well.
Leo Laporte [01:59:23]:
20 passes, 5 weeks. It only failed 5 of the 30. So these benchmarks are very, uh, were very useful. So we're just still— all right, now we're running, uh, AXL3. So this will be interesting. Whoops, I didn't mean to do that. Opened another tab. I meant to go to framework 5555.
Leo Laporte [01:59:59]:
So this is a— oh, there we go. So let's see what Sparky's up to. It's not doing anything right now. This is a monitor program. I think that's out of date. I think it is doing something right now.
Leo's Laptop [02:00:23]:
Huh.
Leo Laporte [02:00:24]:
I guess, I guess it's loading the Excel still. Huh.
Leo's Laptop [02:00:35]:
I don't know.
Leo Laporte [02:00:36]:
I don't know if this is out of date or what. Oh no, it's running. Wow, that was fast. The XL might be the speed leader. I'm gonna have to go because, uh, in an hour I have to go kayak down the Petaluma River in, um, uh, under the full moon. Lisa and I are gonna do that tonight. So yeah, I love Herder. So if you're, if you're interested in this, I would, the software I would absolutely recommend is Hermes, whether you're using the cloud or local.
Leo Laporte [02:01:16]:
And for most people, honestly, the cloud is cheaper and effect— more effective. Local's kind of nutty. Local's not a good idea. So Hermes, absolutely start with Hermes. Get a news subscription and you can use news credits and play with different models. If you get, if you want to get fancier, then Herder is a great way to keep multiple models running consistently. You have to keep, kind of keep an eye on context. The way we do it, I don't, because the way the Quicksilver calls them, it calls them with a fresh session each time.
Leo Laporte [02:01:54]:
So there's fresh context. It starts from scratch on each thing. Yes. Although data center computers might be faster, they might have 800 H100s running. Yeah, Noose is great. So I recommend Noose and Hermes. If you want to run multiple models, whether cloud or local, Herder is a great way to do that and keep them persistent. And if you want them to communicate, Especially if you want to communicate with humans.
Leo Laporte [02:02:27]:
So a lot of people use Slack or Discord or Telegram to talk to Hermes because then it's a messaging platform, right? But I use Buzz, and Buzz is available on iOS. So I can talk to Buzz in here, as I often do, and do kind of interactive sessions with it, like chat sessions. But it's the advantage, it's not a chat because the advantage is you're running your local agent and it has all sorts of memory and stuff. If you're curious what memory tools and so forth I recommend, all of that is at pages.laporte.cloud. This is my— I decided I didn't want the website to look like an AI designed it. So I told it, don't look like an AI, but this is my garden of pages and actually, If you can't get back here, this is the secret stuff. These are pages that are private, like the Twit ratings. You know, I'll show you what the— oh, see, can't get in.
Leo Laporte [02:03:34]:
You can't get— even I can't get in. Downloads for, uh, Twit. Um, I don't remember what this is. Oh yeah, that's just, uh, Hermes. Actually, this is good to know about. I don't know what's going on with it. This is the Hermes status page. So you can see what cron jobs are running.
Leo Laporte [02:03:55]:
This is our travel diary. So Lisa can get in here, but nobody else. I don't want to make that public yet. This is, and actually this was public for a while. This is the, these are the very extensive plans for the TwitAds sales system. Somebody texted me and said, you know, you probably shouldn't make that public. I said, oh yeah, yeah, you're right. So I don't.
Leo Laporte [02:04:20]:
What is this? Oh, this is our advertising kit. This is actually out of date. I made that to show Lisa. And then a little secret, you can't get to this anyway, a little secret. If you click the fountain, Oops. I said if you click the fountain, you can do an I Ching and ask a question, and then it will do the oracle. Or you can close the gates and come back here. But if you want to see the full layout, I'm going to fix this because this bugs me.
Leo Laporte [02:04:57]:
Hi, I'm Leo. Welcome to my setup. This page is a tour of my AI household. 5 named agents and a fleet of machines I own outright with all but the toughest coding tasks free and local. Scroll down and meet the crew. But there is some stuff if you were interested in seeing what I recommend. Like this is probably the most useful page about, uh, this is the memory system I use and the context compression system. This I love.
Leo Laporte [02:05:25]:
This is Breeze, but it is like a 16 gigabyte Voice text-to-speech model, and it's so, it's a little heavyweight. You have to have a little extra stuff. So I will, um, at some point, this is Hermes Desktop, which shouldn't be running, but it is. Oh yeah, it's running in high. That's why. Okay. I did put it on a cloud model. I put HY3.
Leo Laporte [02:05:51]:
So that's why it's running.
Anthony Nielsen [02:05:57]:
Oh, Daniel just responded to your email.
Leo Laporte [02:06:01]:
Who did?
Anthony Nielsen [02:06:02]:
Daniel Suarez.
Leo Laporte [02:06:03]:
Oh, and?
Anthony Nielsen [02:06:05]:
Uh, hold on, I haven't read it. Let me look. Yeah, he's down. Uh, oh good, print some dates.
Leo Laporte [02:06:17]:
Awesome. So I guess, uh, Quicksilver is up, but Running in the cloud, which makes sense. They don't want it to run on the Sparks right now. Nice thing about Quicksilver is it's easy to change models. So you can pick the model you want to run on. You shouldn't do it mid-turn like I just did, because then it has to reload the context and everything. But this is Hermes Desktop, another way you can run Hermes. Hermes also has—
Anthony Nielsen [02:06:46]:
Whoops.
Leo Laporte [02:06:48]:
Didn't mean to do that. Hermes also has a, uh, TUI, uh, which you can run in Herder. In fact, let me run that right now. And, um, that's kind of like Claude Code or any of the, you know, text-based harnesses. They also have a command line, but most of the time I think now you use TUI. So let's change the model here. Because I don't want to be using the sparks. Oh, there's the noose portal.
Leo Laporte [02:07:33]:
And I'll just use something. Actually, you know what? Let's, let's use Saul. Why not? There you go.
Anthony Nielsen [02:07:41]:
Hey there.
Leo Laporte [02:07:49]:
So just now it's just like Claude code or anything else. You're seeing the thinking. You can turn that off. Although I like to look at the thinking if you've got the time. I know Lisa's going to come with the paddle. It's an informal AI user group. I'm on the MacBook Air, so say hello on the M2 MBA, if you will. It still says NBA.
Leo Laporte [02:08:39]:
Yeah, it's the longest—
Leo's Laptop [02:08:40]:
Hello everybody. I'm Quicksilver, Leo's technology coordinator and resident digital accomplice. It's great to meet an informal group of fellow AI explorers. Have a wonderful stream and please keep asking the interesting questions.
Leo Laporte [02:08:55]:
So that's Fable, which is the Hey, you might be interested to know this. You can see this going on in the other Herder panels. We're actually testing your new brain. I'm trying to figure out what the best local model to run on the Sparks would be for you. I'm pretty sure it's going to be a version of GLM 5.3 Flash. The only question is whether we use NVFP4 quantization or Or this new EXL-3. I have high hopes for EXL-3. I think they're pretty conscious.
Leo Laporte [02:09:41]:
I mean, they simulate it. Let's put it that way. You'll see it knows a lot about me. I mean, it's pretty, uh, because it has all this memory, right? It has all this context about me. Oh, see, now it also knows, see, that it's on Herder and it's going to look.
Leo's Laptop [02:09:59]:
Oh, I'm interested. This is brain surgery with a live audience. EXL-3 sounds promising, but crown it only if it beats MVFP-4 on agentic reliability, not merely speed.
Leo Laporte [02:10:11]:
Oh, I couldn't agree more. We actually decided not to run the very hard programming questions that Keiki designed for us. You might remember those. They just were taking too long, but it's scoring pretty highly on the agentic stuff. I even had KKZ and Cronk collaborate on some more agentic questions, tool calling and things like that, because I want to make sure the brain does the best job on all the stuff that we do day to day. You know what we need, right? It's much easier, by the way. You can see why I use voice. It's much easier to say that than type it.
Leo Laporte [02:10:52]:
This is Whisper V3 Large, which is very good.
Leo's Laptop [02:10:56]:
Exactly. We need a dependable operator, not a benchmark peacock. Good judgment, clean tool use, honest verification, and enough speed to stay useful.
Leo Laporte [02:11:08]:
If you have any thoughts about what you'd like to see for your new brain, you might mention it to Keiki. She's Holding the space for the benchmarking right now. Oh, you want to talk about NVFP4? All right. Yeah. I think it has to do with the GB10, that, that, that particular processor is not optimized In the way other CUDA processors are for NVFP4. Oh, I think it's talking to Cakey now. Yeah, see, it's, it's using Buzz. I don't know what's going on.
Leo Laporte [02:12:10]:
So this is one of the issues with these recipes. I'll show you where I get the recipes from. Um, the, whenever these new models come out, it gets crazy, as you can see, with all these people on X.
Leo's Laptop [02:12:26]:
I sent KK the brain wishlist, especially tool recovery, source precedence verification, And a hard veto on confident confabulation.
Leo Laporte [02:12:35]:
So this guy's streaming right now. Numbers right here, 30 tokens per second. This is the same exact speeds that we were hitting on 5.2. Currently we're sitting at a max context of 200,000. I think he has 4 sparks.
Leo's Laptop [02:12:49]:
Adjust this.
Leo Laporte [02:12:50]:
I want to put this out now. This is MTP, uh, 4, by the way. So we haven't even applied DeepFlash, but I've been trying to tune it and stuff. And so I think MTP is actually— everybody else can have— anyway, that's— have the model quantized. So This is— but I was saying, I think these guys are influencers and they're more interested in, you know, I don't know.
Anthony Nielsen [02:13:13]:
They're selling their class or whatever.
Leo Laporte [02:13:15]:
Yeah. Now Mia, whose recipe I'm using, she says she thinks EXL-3 is better. And that's—
Anthony Nielsen [02:13:24]:
this is—
Leo Laporte [02:13:24]:
we're testing her recipe. Jun Song, who's also very active too, GX DGX Spark nodes is all you need for now. Oh yeah, she's— so she's updated her recipe. This is the recipe we're testing right now. So they put these on Hugging Face and, or actually she's using, it is on Hugging Face, but she's using GitHub for her modifications. So this is the recipe we're using now. Let me see when she updated this. An hour ago.
Leo Laporte [02:13:57]:
So what I'm going to kind of want to do real quick is just— that's not what I want to do— is just mention, hey, I noticed Mia is updating her recipe yet again for EXL3. Take a look and see if there's anything we're missing. Yeah, I don't know, Darren. That's beyond my pay grade. I wish I understood this better. Actually, I don't. A lot of it— what's really interesting about this, it used to be you really had to know this stuff to do this stuff. And now, because the agents do so much, You kind of don't have to really know anything.
Leo Laporte [02:14:52]:
You just kind of go, hey, fix that, do that. I think so. I think what's weird is there's a little bit of— you get a little bit of a layer between you and the hardware that probably isn't a good thing. So she's checking right now. The fix is applied as a host-mounted overlay loader. Okay.
Leo's Laptop [02:15:21]:
Dai, Mia's newest fix is one we genuinely needed. Speculative decoding was killing the prefix cache on every follow-up turn, which would have sandbagged exactly the multi-turn agentic tests we're building. It's pulled and applies automatically next launch. Q's wishlist is folded into the spec too, and the switchback to NVFP4 is in its last minutes of loading.
Leo Laporte [02:15:46]:
Why are they switching back? Wait a minute, why are you switching back to NVFP4? Yes, the end of expertise version 2. I think, Dieter, you're, you're kind of right on that one.
Leo's Laptop [02:16:15]:
Oh.
Leo Laporte [02:16:15]:
Oh, I see. No, don't worry, I suspend the pin for this. I trust Mia and her recipes. I trust the—
Leo's Laptop [02:16:23]:
Council ruled NVFP-4 stays resident until EXL-3 passes the tool-calling gate, which we just finished designing but haven't run. Tonight was a bench pass. Not a promotion. If today's numbers changed your mind, I can flip it now instead. But my recommendation is run the Agentic Gate first and promote on evidence.
Leo Laporte [02:16:48]:
Okay, but don't worry about the 2-week pin. I want to get the latest version on here without waiting 2 weeks. So go ahead, run the Agentic Gate on an NVFP4. I guess that's what you're doing next. See how it does and then put EXL3 back on. And don't take it off until I see all the numbers. My guess is that we're going to stick with EXL3, so I really want to take a look at the numbers, and if EXL3 wins, then we're just going to keep using it. Yeah, they were killing the cache each time.
Leo Laporte [02:17:43]:
Cache is huge. The prefill speed on this is like 30,000 tokens a second compared to— so you really, you really want to use the cache. So that's, so again, actually, I'm glad I went to X because this is a developing story. I mean, this is incredible what's going on here. There are all these people, and I think some of them are, you know, link farming, attention farming. I think all of them are attention farming. And then there's Gary Marcus. Oh, this AI is never going to go anywhere.
Leo Laporte [02:18:28]:
Now, today GLM-53 came out, but there's— I can't run that. That's a huge model. Oh, this is so annoying. Mark Benioff has made a deal with Anthropic. And they're calling it Claude Force.
Leo's Laptop [02:19:06]:
No.
Leo Laporte [02:19:06]:
This guy's good. In our lifetimes with the development and the deployment of AI.
Leo's Laptop [02:19:11]:
What else?
Leo's Laptop [02:19:11]:
Ruled and rolling.
Leo Laporte [02:19:12]:
Computer scientists want to—
Leo's Laptop [02:19:13]:
suspended for Mia's recipes. NVFP4 is minutes from healthy and gets the first gate run. And the harness build just kicked off.
Leo Laporte [02:19:21]:
Yeah, Whisper's free.
Leo's Laptop [02:19:22]:
EXL3 goes up with the cache fix and stays up until you've seen every number. If your guess is right and it wins, it keeps the box.
Leo Laporte [02:19:32]:
Good. That's the right answer. So this is some— somebody who's got 8 H200s. With 1.12 terabytes of VRAM. Oh, he's renting it. Okay. Uh, how much would 8 H200s cost? Holy cow. I mean, and that's the thing, some of these YouTubers make so much money and they're, and they're in a constant race to get views.
Anthony Nielsen [02:20:07]:
I like this guy, Tonbee.
Leo Laporte [02:20:09]:
Yeah, Tonbee's a perfect example. And, you know, honestly, I thought I liked Ton B, and he does a lot of Hermes stuff. So I paid 10 bucks to access his agent wikis, and there ain't anything in there that I don't already know. But I thought, oh, well, that'll be interesting. There were a few things. Yeah.
Anthony Nielsen [02:20:28]:
I mean, it's more for the— so your agent doesn't have to search for stuff.
Leo's Laptop [02:20:32]:
Right.
Leo Laporte [02:20:33]:
It's attention farming. So this guy now, Keys and Tech2Wild are both YouTubers and they're kind of, oh, you know, look at this. What is this? Oh no, 779 billion parameters. I don't think so. That's Tony. Okay. So there's, there's You have to— X always has been this way. There's a lot of attention farming.
Leo Laporte [02:21:12]:
There's a lot of BS, but there's also great information. Can you believe this Time magazine cover? They put Joseph Gordon-Levitt and Paris Hilton on the COVID as 2 of the top 100 AI people and not Jensen Huang. He's not even in the list, the most influential people in AI. Ridiculous. Time Magazine, get out of here. Dvorak used to always mock those lists. Oh, this is interesting. Hermes is updated.
Leo Laporte [02:21:53]:
They, they have, I don't know, 20 or 30 PRs every day. So this is a lot of crapola, but there's, but as you saw, the very first thing I saw was a new recipe from Mia, which actually I needed. So that's why I kind of spend, I actually keep this pane open on my framework. I have a bunch of screens open.
Anthony Nielsen [02:22:19]:
Local Llama on Reddit is also a great place to.
Leo Laporte [02:22:23]:
Yes. Yes. And I do, I do follow that. It's a little slower. Which is probably not a bad thing. Um, how often should I update Hermes? Every freaking day, man. The first thing I do every morning is I update Hermes, and then usually the last thing I do every night. Yeah, this is a little slower, which ain't a bad thing.
Leo Laporte [02:22:59]:
There it is. Local LLM. That one, right? Local LLM? Or Local Llama?
Anthony Nielsen [02:23:04]:
Local Llama for me.
Leo Laporte [02:23:05]:
Oh, I don't think I follow that one.
Anthony Nielsen [02:23:11]:
Yeah, that one.
Leo Laporte [02:23:15]:
Huh. Am I a member? Oh, I am a member. I am following it. I take it back. So what are we doing? NVA. Okay, so now it's gonna run that new agentic benchmark tomorrow morning. You can't trust that though, because these guys don't know, have any clue what time it is. They have to actually run a bash command.
Leo Laporte [02:23:46]:
They have to run time and figure out if it's morning, night, what is it? Okay, good. All right. So it's, um, so you see why I'm letting Fable do this because there's lots of things.
Leo's Laptop [02:24:10]:
All right.
Leo Laporte [02:24:10]:
I really do have to go. Um, Thank you. We're hanging out. I think we could keep— I could do this more. I mean, I do it anyway. It's no big deal to have a chat room open and have the camera on. So, uh, you— and you don't have to always be here for it, Anthony. I know.
Anthony Nielsen [02:24:34]:
Well, no, I'm doing stuff in the back.
Leo Laporte [02:24:36]:
You're doing your thing too.
Anthony Nielsen [02:24:37]:
Yeah.
Leo Laporte [02:24:38]:
And thank you, uh, very much, Darren, for the, um, Kokoro, uh, models. I'll do that tonight.
Anthony Nielsen [02:24:45]:
I'll—
Leo Laporte [02:24:45]:
after I get back from the moonlit paddle. Oh good, cleared up. It was raining in Petaluma earlier. Thank you, everybody.
Anthony Nielsen [02:24:54]:
On Daniel, is that for IM or for Twit, or doesn't matter?
Leo Laporte [02:25:00]:
Uh, oh, it's IM for sure.
Anthony Nielsen [02:25:02]:
Okay.
Leo Laporte [02:25:02]:
Yeah, yeah, I don't wanna— I mean, I could do I could do a— it wouldn't be TWiT. I could do a club.
Anthony Nielsen [02:25:10]:
I mean, if it's for I Am, I'll just say it's for I Am.
Leo Laporte [02:25:13]:
Actually, you know what we could say is, can you spare an hour? And then we'll take a half hour out of it for I Am as possible. That's what we could do. What do you think? We can discuss this later.
Anthony Nielsen [02:25:27]:
Okay.
Leo Laporte [02:25:28]:
What did he— did he—
Anthony Nielsen [02:25:30]:
he just said sometime in October.
Leo Laporte [02:25:32]:
Okay. So it's not till next month anyway. I mean, 2 months.
Anthony Nielsen [02:25:35]:
Okay.
Leo's Laptop [02:25:36]:
Yeah.
Leo Laporte [02:25:36]:
Because he wants to be closer to his book release is March. I think we should do a triangulation. So we should do an hour and then we'll put part of that or all of it in I Am. That would be my vote. It's just like Andy Weir. If we can get somebody like Suarez, Let's do a triangulation with it. We'll do an hour, talk about the book. He may say October because I can— he can send me a galley probably, which would be a good thing to do.
Leo Laporte [02:26:10]:
Thank you, Fart Supreme, for your kind words. Not so much your name, but your kind words. It is fun. I love playing with AI. I really do. And I don't pretend to be an expert. I do pretend to be Deeply involved. And it's really fun.
Leo Laporte [02:26:31]:
And I think we're all kind of in the same— obviously there's some real experts, but we're kind of all in the same boat. What I don't want to be is an attention whore, a click farmer, YouTube-style click farmer where I just, oh my God, gang, the latest has come out.
Leo's Laptop [02:26:47]:
We got to get it.
Leo Laporte [02:26:48]:
Let me build a pagoda. And You know, I think this is more, I like, I've always felt this way, and this is going back to TechTV days, we wanna do it real. And so in one way, a stream is nice, 'cause it is in real time, right? Today's record Spark setup is in Buzz. Oh, I gotta keep going. Let's see though, but we have a channel just for this Spark stuff. So this is— this Buzz is nice because it's kind of a more permanent record, but you see it's all in Claudish, which I don't speak Claudish. Your confabulation false success veto is adopted as a second veto class beside parser errors. Same reproducibility protocol.
Leo Laporte [02:27:35]:
One ask back. See, it buried that. The Shadowrun idea is in. Please supply 3 to 5 sanitized transcripts of Recent real— oh, it's talking to Quicksilver. Thank God. Switchback complete. Jeez Louise. You see why I don't—
Anthony Nielsen [02:27:56]:
what?
Leo Laporte [02:28:00]:
Oh, let me add. Okay. I'm going to invite them to the room. So since Quicksilver didn't get around to it. Thank you, everybody. I'm off to paddle. We'll see you later. Thank you, Anthony.
Leo Laporte [02:28:32]:
Have a wonderful evening, and I'll see everybody on Sunday for TWiT.
Leo's Laptop [02:28:38]:
Bye.