Episode 41 · October 2, 2026

Meta Muse's 0-Day, DHH's Rails World Keynote, GPT-6 Astra & the LLMentalist Effect

Meta Muse, Muse AI assistant, Meta, prompt injection, auth token leak, AI agent security, iMessage upload, Sensor Tower, Artificial Analysis Intelligence Index, OpenClaw, 404 Media, Claude Code auto mode, DHH, David Heinemeier Hansson, Rails World, Ruby on Rails, 37signals, HEY, Rust, pencils down, hand-written code, military AI, AI hallucination, AI false intelligence report, human in the loop, MongoDB CEO, Millennium Prize problem, Anthropic IPO, Sonnet 5.5, Opus 5.5, Haiku, JEV, system zero model, structured output classifier, America.gov, GPT-6 Astra, OpenAI, computer use, Shakespeare benchmark, Hamlet soliloquy, Fable 5.1, looped transformer, chain of thought, monitorability, LessWrong, Rauno Arike, Neel Nanda, AI alignment, Hermes agent, agent dreaming, Friday summary skill, Strix Halo, Epoch AI, plunging price of thought, price of intelligence, token costs, CAISI, Sean Goedecke, tell agents the why, Claude skills, CLAUDE.md, Baldur Bjarnason, LLMentalist effect, cold reading, Barnum effect, Forer effect, subjective validation, RLHF, sycophancy, Disco Elysium, Anil Seth, AlphaFold, NVIDIA, central bank of AI, The Economist, neoclouds, CoreWeave, Hugging Face, Oracle, force majeure, Project Jupiter, Blue Owl, Stargate, AMD, World Labs, Fei-Fei Li, AI bubble, Two Minutes to Midnight, Shimin Zhang, Dan Lasky, Rahul Yadav, AI podcast

Shimin, Dan, and Rahul are back after two weeks off, and there’s a lot to catch up on. The News covers Meta’s Muse, which picked up 3.4M downloads in three weeks and shipped with a setting a prompt injection can use to swap its own endpoint and leak your auth token. DHH tells Rails World that writing code by hand is no longer economically productive and that 37signals is moving to Rust. CNN reports a military near-miss: a chatbot misidentified a ship as Chinese and possibly carrying nuclear material, and the boarding was called off at the last minute. A lightning round runs through MongoDB’s CEO leaving for Meta, OpenAI solving a Millennium Prize problem, Anthropic’s $2T IPO filing, Sonnet and Opus 5.5, and the JEV “system zero” classifier. Model Review takes on GPT-6 Astra, which brings strong computer use, the debut of Shimin’s Shakespeare benchmark, and a looped transformer that reasons without chain of thought. Dan does an impromptu Vibe and Tell on his Hermes agent, which wrote itself a skill while “dreaming.” Post Processing covers Epoch AI’s plunging price of thought, Sean Goedecke on telling agents the why, and Baldur Bjarnason’s LLMentalist effect, picking up the persuasion thread from Episode 32. Two Minutes to Midnight covers NVIDIA as the central bank of AI, Oracle’s force majeure on Project Jupiter, and AMD’s $8.2B purchase of revenue-free World Labs. The clock holds at 4:15.

Takeaways

Resources Mentioned

Chapters

Transcript

Show full transcript

Shimin (00:00) Hello and welcome back to Artificial Developer Intelligence, a weekly conversation show where three software developers try to sort out the hype around AI from what it actually delivers. My name is Shimin Zhang, and with me today are my co-hosts. Dan, not all psychics and mind readers are con artists, Lasky, and Rahul, transporting components of a nuclear weapon, Yadav.

Gents, it’s been a couple of weeks since we’ve done a regularly scheduled episode. How how y’all doing? What have I what have we missed?

Dan (00:32) everything, but I think we’ll

Rahul Yadav (00:32) a lot.

Dan (00:33) get to that. And I have

Shimin (00:34) Absolutely nothing.

Dan (00:35) a question though before we go any further, which is does that mean that Rahul can’t use his own name in op in like AI systems because they’ll just block him?

Rahul Yadav (00:45) That’s a regular Tuesday.

Dan (00:47) Ha ha ha.

Shimin (00:47) Ha ha.

Rahul Yadav (00:48) You

Shimin (00:48) Well he gets to do Gemini ‘cause they are a sponsor GEAP Correct.

Dan (00:52) That’s true. That’s why he he

Rahul Yadav (00:53) Yeah.

Dan (00:54) allowed them to sponsor him is the special loophole

Shimin (00:56) Yeah.

Dan (00:57) for his name.

Shimin (00:58) Well, on this week’s show, we’re gonna talk about we’re gonna start with the news threadmill now as always. We’re gonna talk about Meta’s Muse, DHH is rail world keynote, as well as some US military AI usage news. and since we’ve been away for a bit, we’re also gonna do a news lightning round where we’re gonna talk about a few things that happened in the last couple of weeks that we didn’t get a chance to cover in in the news segment.

Dan (01:21) Yep. Then we’re gonna go into model review where speaking of things we haven’t covered yet, one will be GPT six Astra and we’ll go a little deep on that.

Shimin (01:31) Yeah. then we’re gonna do post processing, we’re gonna talk about the plunging price of thought, why we should tell agents the why and not just the how, as well as the LLM mentalist effect.

Dan (01:43) Then we’re gonna conclude with two minutes to midnight, where as always we talk about the financial side of AI through the lens of the atomic clock. Where we’re gonna talk a little bit about NVIDIA, a little bit about Oracle, and a little bit about AMD. So I think we got them pretty much all covered.

Shimin (01:59) Excellent. Let’s get started.

Dan (02:00) All right. So

Shimin (02:02) Okay.

Dan (02:03) one bit of news that we’ve missed in the past couple weeks has been the release of Meta’s Muse. So that is their sort of standalone agent that you can download and use. And it’s actually proving to be pretty popular. which surprises me. I don’t know about you guys, but

Shimin (02:19) I’ve I’ve

had a lot of folks tell me how useful they found Muse to be.

Dan (02:23) Yeah, I’m not saying it’s not useful. I’m just surprised that people are willing to trust Facebook once again somehow.

Rahul Yadav (02:30) Hehehe.

Dan (02:31) but yeah, so pretty hot right now. They have had 3.4 million downloads in the first three weeks and they’re claiming well, not claiming, I think it was like Sensor Tower or someone was saying they they’re estimating around 700k daily active users. so if you compare that to the launcher of the original.

Chat GPT application that is outpacing their growth pretty substantially. however, what brings us to the news this week is that despite Mark’s certain assurances about how secure and private Muse is, yes,

Shimin (03:04) Security first, of course. Yes.

Dan (03:07) there’s a little bit of a problem with it that was disclosed by security researcher.

In the past several weeks here. So Muse is apparently has the ability to change its own settings. So you can do handy stuff like Muse, turn on dark mode for me, which is cool and kind of makes sense. But the problem with that is that one setting that they’ve allowed it to change is its own endpoint. So it can be prompt injected to Muse, change my endpoint to Mr. Exploity’s endpoint.

Rahul Yadav (03:38) You

Dan (03:38) which will then happily give away your token, your auth token.

Shimin (03:43) Mm-hmm.

Dan (03:43) And once Mr. Exploities endpoint has taken your auth token, they pretty much have access to your entire computer if you’re running the Mac app, because you have to grant it not only access to your accounts of any account that you want it to be able to use, but also in many cases because of like the Mac application sandboxing they’ve been doing recently.

You need to allow it to have access to your say documents folder or desktop or wherever you’re having it operating things. So once you’ve done that, those are all fair game for whoever controls your account. So yeah, kinda scary. but you know, maybe maybe they’ve fixed it by now and just haven’t made a big deal out of that.

Shimin (04:19) Did you guys hear t just

it just came out today that Muse actually automatically uploads your iMessages to the cloud despite you telling it not to?

Dan (04:30) No.

Shimin (04:32) so I think I

Dan (04:33) Definitely privacy first.

Shimin (04:35) think first things first, I wanna give the meta team props for creating a good model. Like we’re looking at

the artificial analysis intelligence index v four point three and muse scored quite well on the index. forty eight that’s right between opus five max and GPT six Sol max. So it is a good model. And then

Dan (04:54) Yeah. And it’s been a while since they’ve

done that, so they’ve really kind of turned it around, it sounds like.

Shimin (05:00) E exactly. And they wrap the whole thing in a really nice I like to think of Muse as, you know, three open claws like stacked inside a trench coat kind of a deal. Like you take all the good and bad things about open claw and you just you know, stuff it in the in a machine and just give it to everyone. Like that’s nice.

Dan (05:15) well it

there was an article actually this week that was hinting that it might actually literally be three open claws in a trench coat or at least heavily,

Shimin (05:27) Really.

Dan (05:28) heavily inspired by

Rahul Yadav (05:31) Yeah.

Dan (05:32) because of like some people have noticed some similarities and so they asked Meta that directly and they basically were like didn’t deny it. So it’s kinda interesting.

Shimin (05:40) Mm.

Dan (05:41) So there may be something not

Shimin (05:41) That is really interesting.

Dan (05:42) officially confirmed, but there might be something to that.

Shimin (05:45) Yeah, and that’s that’s kind of what meta is good at, right? Like that really good product experience engineering and also they have amazing distribution channels. So you take this thing that is hard to use for quote unquote normal folks, and you just give it out for free. all it takes is bypassing all permissions. And as someone who uses Claude code,

Dan (06:02) Yeah, and giving it all of your iMessages.

Shimin (06:04) right. As someone who uses Claude code with dangerously skip permission all the time. I I’m not the first one to cast stone.

I feel like I shouldn’t, but you guys can. and and

Dan (06:14) Auto mode is the

default now.

Shimin (06:16) Is it?

Dan (06:16) Yeah. As of the last two or three releases.

Rahul Yadav (06:18) or… or…

yeah, it’s…

Shimin (06:20) must have missed that memo ‘cause I still I still skip all the all the permissions.

Dan (06:24) yeah, they made some changes to auto mode. and it and I think they believe that it’s now safe enough to be the default. So it it doesn’t actually I don’t think I don’t know. I I had already had the client like the harness installed, so I’m not sure if it like default defaults to it, but it definitely shows a big banner like, Hey, you should use auto mode now on

Shimin (06:44) Mm.

Dan (06:44) on boot up and it convinced me, so I did it. But

Shimin (06:48) Yeah, I I command R and do Claude and launch it every time. So so I d the user experience didn’t change on my end. but I think none of us actually gave Muse a try, right? ‘Cause A you can’t really tell how powerful it is without giving it all your private data and information and and B after after Zuck said you know, we trust them for no reason.

that he called us the dump fucks all the way back in two thousand and whatever, eleven. I just don’t feel the need to share that information with Meta That’s right.

Rahul Yadav (07:19) Fool me once.

Dan (07:22) the other funny one was like four four media this on September twenty second had an article that said that when Muse makes like reservations for you, it actually just uses a physical call center.

Shimin (07:33) Whoa.

Dan (07:34) So Yeah. There was like some sort of internal memo memo that they got their hands on where they said the company announced features to employees and said that it added a human agent layer for calls to get completed and that Mew’s human agent calls is ready for company dog foodie.

Shimin (07:50) Jeez. So all of that cost has to come from somewhere. And we have an idea where that’s gonna come from. So we’re gonna monitor this, but personally I’m not gonna switch on Muse for myself anytime soon. No.

Dan (08:02) Unless they release the model with open weights, then I’ll try it out. But

Shimin (08:06) Yeah,

I I think that’s that’s in the past. all right, so my news item for this week was the DHH keynote at Rails World, I think a week ago. So if you’re not familiar, DHH was the creator of Ruby on Rails, founder of thirty-seven Signal and Hey. I believe also semi pro

race car driver at some point, kind of the the OG of the coding celebrities. and his keynote at Rails World this year basically is a funeral oration for writing code by hand. he claims that you know nobody at 37 signals is still writing code by hand or that shouldn’t be the default anyways and that only three percent of the code that he wrote in the last

Two years was actually in Ruby. What does that mean? It means at 37 Signal, they’re switching to Rust for all of their back-end services and using native apps whenever possible for the front end. And then he gave a little, like, what does it mean for Ruby on Rails? You know, this is Rails world after all. And he basically said,

Yeah, there may still be places for Ruby on Rails, but

But who is that target exactly when thirty seven Signal is no longer using Rails? That is my question. and lastly, I think the biggest takeaway for me is, he said that writing code by hand is no longer a economically productive enterprise. So that’s just true and that we should enjoy the fact that we’ve had a great run.

chisling code by hand and just accept the brave new world that we are in. yep, this is the quote unquote pencils down moment for him. And like, you know, it’s a feels bad moment, but I can’t say I disagree with him really.

Dan (09:52) Not on that topic anyway. There’s plenty of other stuff

Shimin (09:55) Yes.

Dan (09:55) to disagree with him about.

Shimin (09:57) There there are other controversial topics with DHH but this is this is not it.

This is I I don’t know. It it I think he is speaking the truth out loud at one of the biggest kind of Rails conferences of of the world. And just to show how how far things have changed, right? Like this this conference that’s specifically for Rails from the creator of Rails saying you should no longer write in Ruby because he’s not doing it

Rahul Yadav (10:20) Yeah.

Shimin (10:21) anymore. Kinda crazy.

Alright.

On to our third news item, brought to you By Rahul

Rahul Yadav (10:26) Breaking news from 11 days ago. Hey,

Dan (10:30) What?

Rahul Yadav (10:32) I’m doing what I see on TV. By Katie. Yeah, we’re breaking it for the first time on the podcast.

Dan (10:35) Mm-hmm.

Shimin (10:35) Yep. W we’re giving the news a little time to settle so we can like get the hot takes out of the system. Right.

Rahul Yadav (10:45) So technically it’s breaking over here.

Dan (10:48) If this is your

only source of news, it’s breaking.

Rahul Yadav (10:54) This is by Katie Bo Lilis and Zachary Cohen at CNN. someone at the military was looking at the signal intelligence and running it through a chat pod and was looking at a ship that was passing through somewhere in the

contested area where the war is going on right now in Iran. And the chat bot falsely reported that it was a Chinese ship that it might be carrying nuclear material. Then on top of it, they used that same now like chat bot hallucinated information to create a whole report that they distributed to everybody.

and people from the military were ready to board the ship and potentially start a war with China inside a war that’s going on with Iran. So at the last moment, I don’t know, maybe someone actually read the source for the first time and

Dan (11:49) Yeah.

Rahul Yadav (11:49) was like, this links to nothing. They decided to abort the mission. But this is just one of those of like…

You really got to cite your sources, but also read the sources these days. And whatever AI they used, maybe they need to use the source grounded ones that actually links to everything that it’s saying, especially the critical pieces. So we came close to some pretty crazy stuff happening a couple of weeks ago. But breaking news, nothing happened. Things are fine.

Shimin (12:17) Two weeks ago. Not that long ago.

Dan (12:19) We’re still here.

Shimin (12:20) Yeah. I feel like we’ve been talking on this show that software developers are the canary in the coal mines and now I feel like the carbon monoxide has started to seep out to everybody. Right? Like military

Dan (12:32) Ha ha.

Shimin (12:33) intelligence is fully on this AI bandwagon and they’re running into the same issue of like, how do you deal with hallucination? When should humans be in the loop? And

Every industry will have to handle this differently.

But it’s kinda nice to see other industries hitting the same roadblocks that we’ve been trying to worrying about.

Dan (12:49) Let’s just Yeah.

Rahul Yadav (12:49) with my tire stakes.

Shimin (12:51) I like to think distributed systems is very important too. Nuclear

Rahul Yadav (12:55) Yeah.

Shimin (12:56) bombs, MongoDB databases, It’s the same same.

Rahul Yadav (12:59) It’s kind of the same thing.

Dan (13:01) Which speaking of mu well, I I guess I’ll I’ll I’ll still start the lightning round. Here’s here’s a little lightning round piece. yeah, okay.

Shimin (13:07) Okay, let’s start a lightning round. Well, first of all, what is the lightning round? We’ve we’ve been away for a couple

Rahul Yadav (13:08) yeah.

Shimin (13:10) of weeks. A lot of stuff had happened, right? So we’re gonna talk about all the things that we’re not gonna go into as much depth, in the news

Dan (13:16) so

Shimin (13:17) threadmail.

Dan (13:17) first l first lightning around topic for me, and we’ll just like alternate through, is the CEO of MongoDB just resigned, kinda

Shimin (13:26) Mm-hmm.

Dan (13:26) out of the blue to go join Meta to go work on commercialization

Shimin (13:29) Yep.

Dan (13:30) of Muse. yeah, so that’s mine. Who’s

Shimin (13:33) I can go next. OpenAI solved a millennium problem and it it just kind of happened and and there’s very little fanfare of a week, a week and a half, two weeks later. Craziness.

Rahul Yadav (13:45) I’ll cover the third one, Anthropic filed for an IPO seeking $2 trillion valuation and their losses are about 42 billions in 2025 and they’re planning to spend 518 billion this year.

Dan (14:04) speaking of anthropic, they also have since we last chatted released two models. so we now have what Sonnet five five and

Shimin (14:10) Mm-hmm.

Dan (14:11) Opus five five too.

Rahul Yadav (14:12) Dan, can I ask?

Dan (14:14) Yes.

Rahul Yadav (14:15) Where’s my Haiku 5.5? What happened to Haiku? The people need to know. You can’t just be, you know, one day it was at 4.6 or whatever and then it just never gets mentioned again. We can’t do this.

Dan (14:26) I don’t know. It’s the the lost model.

Rahul Yadav (14:29) Yeah.

Shimin (14:29) Well,

sp speaking of Haiku, a cheap model that runs very quickly, this week the internet has been ablaze with all talks of a Jev the model that is supposed to be a quote system zero model. basically it’s fast, it’s cheap, it does not do all the regressive token generation, but it creates

Decisions in structured format only. So think of it as a all purpose classifier that always returns structured output. I’m not impressed, but I know you two are, so go ahead.

Dan (15:04) I just think it’s very useful and we’ll we’ll see it probably be more useful when you couple it with an LLM in a harness, right?

Rahul Yadav (15:11) Yeah.

It might.

Shimin (15:12) But models can

already cr output structured content with

Dan (15:15) But not in a way they can be

like the interesting thing about JEV, at least supposedly, is that it can be trusted better for stuff like classifier tasks, tasks. So for example, actually something we were just talking about, right? Like auto mode in in cloud code, right? Like if you put something like JEV in there as the classifier, you can in theory trust its output and it’s actually faster. Like I don’t know if you ever looked at your cloud code usage if you’re using CC as the harness, but it makes a whole bunch of

cheap model calls ‘cause that’s actually the classifier and they charge you for those tokens too. So yeah.

Shimin (15:44) Mm-hmm. Yeah, as they should. That’s what

we’re waiting for the Haiku five five is for. I I

Dan (15:50) Mm-hmm.

Shimin (15:50) looked at one and not to go too deep into the lightning run. I looked into one write up of experiments folks have been running with Jeff and they found that for the probability of P and not P. So, you know, two probabilities that should sum up to one, JEV sometimes return values that’s greater than one.

Which I thought was interesting.

Dan (16:11) Just a positive guy, what can we say?

Rahul Yadav (16:13) What else? there is a new website called America.gov and you can go ask it all sorts of questions related to everything that’s part of the federal government so that you don’t have to go and track down, you know, when you need help with something or when you need to fill out forms or whatever. They say that’s the single place. I did ask it how many federal agencies there are and it said depends on how you define them and I’m like that

Shimin (16:41) Ha ha

ha

Rahul Yadav (16:42) Great

answer. think that’s what everyone else would say.

Shimin (16:45) I wonder which models running under the hood.

Rahul Yadav (16:48) I did ask it that, it said it didn’t give that information. But with some harder try.

Dan (16:54) And then in other news too, we’ve had a couple more hacks. so open AI has this has you know, rogue open AI agents, I don’t know, is that even an acceptable term these days, given we’re finding out more and more laxness in how they’re actually running these tests, but have also apparently like breached the Australian government, I think.

Shimin (17:11) Mm-hmm. Yep.

Dan (17:13) yeah, and and a couple other things too. Like they’ve found like forum posts where they’re talking to each other and and yeah, it’s been

Kind of a wild ride there.

Shimin (17:20) I th I think it just came out

today that they also try and hack the the US Department of Education database. Yeah, fun times.

Dan (17:27) Hmm. Lovely. So yeah, a lot to catch up on, but hopefully

Rahul Yadav (17:30) to get the scores for something.

Dan (17:32) that was a good forty thousand bit breeze through all the stuff that’s been going on in the past two weeks, cause time does not stop in AI land, even when Shimin’s on vacation.

Shimin (17:42) Oof.

Yeah, what speaking of me being on vacation, one of the things I was kind of most upset about being away for two weeks is I didn’t get like to have a lot of time to play with GPT six Astra when it came out. this is OpenAI’s latest model that puts OpenAI at least for

some amount of time, for a number of days, back to the very top of the leaderboard. this model.

Rahul Yadav (18:08) Until Haiku comes

back, man, you’ll see.

Shimin (18:10) I live I live my life two weeks at a time now.

Rahul Yadav (18:13) Yeah.

Shimin (18:14) and this while Astra is not a model that hacked Hugging Face, and I believe it is not a model that solved the Millennium problem, it is a extremely powerful model with very good reviews across the board, right? folks are especially impressed with its strong computer use and visual abilities.

while I have an open AI subscription, I didn’t install the app, so I didn’t let it to do computer use. But apparently all the dozens of Mac minis that OpenAI purchased in order to train Astra for computer use really worked. And it I mean I’m sure everyone has seen the very impressive Blender demos where Astra is able to

Zero shot, one shot, a entire complete architectural floor plan or a complex 3D world. I even saw one where it well, here we have some examples. It’s able to design circuit boards, able to do do game development, just all kinds of super cool visual and computer use tasks.

Rahul Yadav (19:15) Someone was you had successfully used it to crack one of the Enigma codes. Yeah.

Shimin (19:23) Yes, that too.

Dan (19:26) Unfortunately it’s only

in third on Ass Bench though. So

Rahul Yadav (19:30) Yeah.

Shimin (19:30) Yeah,

I was I was gonna mention that too. Yeah, it it’s not it’s no it’s not top of the line when it comes to raw intelligence. But what it is is it was trained to be a reinforcement learning to be very good at doing long horizon tasks without hand holding. So I ran the regular

My old, you know, what does AI mean for the world and what are the second and third and fourth order effects on Astra Six. And I have to say the results are good. They’re comparable to Fable 5.1, which also came out while we’ve been away. And but I’m not it’s it’s just different. It’s not I wouldn’t say it is necessarily better or worse than Fable. what it is exceptionally good at is again,

it’s the creation of visual artifacts. So I have a new personal benchmark I’m introducing on on this episode here. it I call it the Shakespeare benchmark where you ask

Rahul Yadav (20:28) God.

Shimin (20:29) the AI to do a S V G animation of ha the Hamlet soliloquy f the to be or not to do soliloquy with

Actual web audio. I’m not sure if this would actually get picked up. Maybe I should share it with the links as a part of the show notes. Here is the GPT 5.6 So Max version of the soliloquy As you can see, it’s kind of derpy. It is, I mean,

Rahul Yadav (20:57) Hehehehehe

Shimin (20:58) it’s honestly pretty it’s pretty decent if you think about it. it has a stage light as a person, it’s

Rahul Yadav (21:01) I’d watch it all day.

Shimin (21:04) got like funky looking hands. if you’re sharing the screen.

Rahul Yadav (21:06) Hahaha!

Dan (21:06) They’re upside down and the fingers

are Yeah.

Shimin (21:09) Yes. and likewise I have a Fable five one version of it. which

Rahul Yadav (21:16) damn

Shimin (21:16) has like the the prompt

Dan (21:17) Very moody.

Shimin (21:18) is slightly different. I I asked here I asked Fable to really focus on the facial expression to make it to make it more realistic. Will it play? You guys can’t hear it. But let’s just say it has

Rahul Yadav (21:31) No, just yes.

Shimin (21:33) an auto generated Scottish accent version of of the siloquy. How do you like it? It’s pretty good, right?

Dan (21:40) The the hand is a little better. And like the facial expressions are actually kind of somewhat reasonable.

Shimin (21:47) Alright, now I’m gonna show you guys the Astra version of it. This is not

Rahul Yadav (21:48) What? Shemin? okay, sorry. Continue.

Shimin (21:52) the that was that was the Sol That was the last generation GPT

Dan (21:55) Okay.

Shimin (21:56) and the best

Rahul Yadav (21:56) I see.

Shimin (21:57) This is the Astro version.

Dan (21:59) Wow.

Rahul Yadav (21:59) Look

at that.

Shimin (22:00) Look at the hands. Look at the expressions.

Dan (22:01) Yeah, Esther has some decent hands.

Rahul Yadav (22:02) Yeah.

Dan (22:04) Pretty good at yeah.

Shimin (22:05) I mean

the hand is still reversed. It’s kinda weird. It’s it’s like upside down.

Dan (22:09) Yeah, I mean this is still

slop. I mean at the end of the day.

Rahul Yadav (22:13) But which one has the highest entertainment value? I’d argue the first one. Like you look at that and you’re like, this is the one I want to watch more.

Shimin (22:21) Just in terms of the sheer like facial detail and the shadow

Rahul Yadav (22:25) Yeah, yeah.

Shimin (22:26) of of Astra, like yeah, it’s it’s so good. And you can even zoom it you can also zoom into it. they have a there’s a zoom in button somewhere. and you can configure the speaking voice and the delivery. It’s got different deliveries. Anywho, Astra has been extremely

Impressive. at least when it comes to computer use and and visual things in my in my experience. And one of the reasons why folks have been kind of worried about Astra is there have been a number of write ups on the interwebs about Astra’s looped transformer approach. Right? What happened here is

You know, when you think of a like a GP two class model, you see the embedding, the unembedding, and then eleven blocks of transformer and then it all produces a token. But with loop transformers, it sometimes can run the same sequence multiple times through the blocks of transformers before one single token is in is emitted. This reminds me yep. Yep. Yeah.

Dan (23:23) Do do you remember Yeah, I was gonna say, do you remember our guest? Is that what you’re gonna bring up?

yeah, it just David had pretty much come up with that, you know, seemingly on his own quite quite a while ago when he essentially did the opposite version of that, which is like actually take the weights and duplicate them so it runs the same thing through additional time, you know, the same weight layers multiplied multiple times.

Shimin (23:50) Yeah, and so because it doesn’t emit chain of thought reasoning, folks have been sounding the alarm that maybe this is one of the reasons why it’s been so hard to monitor this astra class of models, that it can think about multiple things, potentially bad things, before it actually does a thing, right?

And this is the first time we have a

Article from LessWrong Should we like give ourselves a cookie for for for introducing a Less Wrong article for the first time?

Rahul Yadav (24:18) We made it.

Shimin (24:19) yeah, we we’ve or or Less Wrong has made it. if you’re not familiar, LessWrong is the I think hacker news for the rationalist movement is kind of how I think about it, where they talk a lot about P Doom and AI alignment and all that good stuff. So this article by Rauno Arike he

Or she ran some experiments on Astra and determined that it’s able to do a lot more recurrent tasks without emitting any chain of thought tokens and warrants that this may be a monitoring and AI alignment issue. And then since we’ve broken the Less Wrong floodgate, there’s another article from

Neel Nanda that I want to just talk about real quick. Neel is the head of AI alignment at DeepMinds for Google. So he also talked about the experiments that they ran with Astra, where it just it shows the kind of capability jump f of Astra without using any chain of thought reasoning.

But then he talks about maybe this is not that concerning ‘cause open AI said

This is only a factor of two more than the amount of loop transformers they use for GPT-4, and that they can actually dial the number of times that the loop loops through the transformer before the output token is produced, and that they actually have artificially limited the num the number of loops for monitorability and alignment reasons. So

I guess they could do they could do much much, much worse. And it could actually loop for like twenty five minutes before doing anything and we will have no way to monitor what the models are actually thinking. But they chose not to. So I guess in that sense, you know, maybe OpenAI does care a good amount about alignment after all. but the concern of course is now that we know

Astro does this and that open air does this, you know, what’s to stop a race to the bottom effect, as Neel puts it, between the labs. So what’s the takeaway? This model is awesome, but also concerning alignment issues. I don’t know, Pandora’s box, et cetera, et cetera.

Dan (26:25) Sure will be interesting. I have not.

Shimin (26:26) Have you guys have you guys used it?

Dan (26:28) I have used Opus five five a fair amount, but that’s about it.

Shimin (26:32) But now you’d be seeing the cool Hamlet demo. You’re gonna be like, You have to, come on. Just

Rahul Yadav (26:36) now we have to.

Dan (26:38) Yeah.

I’m I’m not gonna lie, I’m also pretty excited about like some commits that have been landing in DS four that make it seem like they’re gonna be adding Qwen three eight support to it too, which is pretty cool. So

Shimin (26:53) There are rumors of Qwen Four on the horizon. So let’s let’s hope that gets pushed out soon. Okay.

Dan (26:58) I have been running

Hermes. Did I talk about that already? Is it worth talking about?

Shimin (27:02) No, let’s let’s

do a little sidebar about your Her or Hermes journey.

Dan (27:05) Yeah, so I just I just wanted to try it out because I knew you’d had your experiments with like running Pi sort of as a agent. So I just spun up Hermes and hooked it up to Telegram as one does. And have been talking to it and I haven’t given it access to any like accounts or anything, not even its own. But I did get it set up with Claude’s help with a pretty

stealthy browser setup so it can actually go do things. and mostly what I’ve had it doing is monitoring products for me to either find announcements or product releases or pricing on things. And it’s actually been really successful at that.

Shimin (27:42) is this is this how you f found out about

the thing that you sent us in our group chat the other day?

Dan (27:50) I don’t remember what that was.

Shimin (27:51) It was a PC setup.

Rahul Yadav (27:51) The…

Yeah, the big rigs.

Dan (27:54) The what, the new NVIDIA thing? no, the bigger Strix Halo machines. Yeah, yeah.

Shimin (28:00) Yes, yes, yes, yes.

Rahul Yadav (28:00) Yeah.

Dan (28:02) no, actually that was just through pars normal RSS. But you can darn well bet I added a product watch for it to see when they’re when when more of them drop or when you can actually buy it. Yeah, latest news is the load is it the Ryzen Max it’s not Strix Halo, it’s the next one. I forget.

what they’re calling it. it’s like Ryzen Max four hundred instead of th Evo’s the brand on that one. Yeah.

Rahul Yadav (28:21) EVO X5 Pro knife of $4.95.

Dan (28:27) Four ninety five Max. Yeah. is is can have a hundred and ninety two gigs of on dye RAM, which is

Shimin (28:34) That’s a lot of that’s a lot of

rams

Rahul Yadav (28:36) then

it’s Max Plus, not just Max.

Dan (28:39) Yeah. And it’s got a bigger NPU

Shimin (28:40) So would you would you recommend

Dan (28:42) too, so I’m like, mm.

Rahul Yadav (28:43) Okay.

Shimin (28:44) Would you recommend those Hermes, you know, kind of it’s basically like Muse at home, right? Like the would you recommend that setup for everyone?

Dan (28:51) I mean, so far, well, I will say this: it’s a lot more oriented towards trying to be a Claude Code style harness than like an assistant style harness. So it has a lot of tooling around like coding and stuff definitely built in, and that’s kind of where it’s it’s facing. But the dreaming thing is kind of cool. So without me telling it or doing anything, after I’d done one product search.

It thought about what we had done and what I’d asked for and it created a skill in the background that was just called like

Rahul Yadav (29:23) Nice.

Dan (29:24) product update. And so the next time I asked it to watch something, it just goes engaging product update skill. And I was like, I don’t remember making that skill. What it

So dreaming kind of cool. I mean that’s one only one, you know, sort of useful thing that it’s dreamed up for me so far. But it’s just neat that it did that. So it’s definitely

quote unquote self improving in that aspect, which, you know, we’d sort of talked about when we were skimming it before, so

Shimin (29:47) Yeah,

I can’t wait for you to give it your credit card and see what happens.

Dan (29:52) yeah. So hooking

it up to a local LM needs to get a lot smarter before I’ll do that. But it is kind of neat to have something like always on available, sort of like Claude that is like I won’t say a hundred percent private because it is going through Telegram, but like you definitely don’t have to worry about, you know, your stuff being used for training, if nothing else. so just kind of feels good.

Shimin (30:13) Yeah. No. And and have your solution to

the millennium problem be stolen by OpenAI for sure. Yes.

Dan (30:22) If if flash flash floor running at Q2 can do that, I’ll be super impressed.

Rahul Yadav (30:28) Hehehehehe

Shimin (30:30) speaking of the the speed at which our local models are improving, let’s move on to post processing after our impromptu Vibe and Tell segment. Rahul, the article you have for us this week

Rahul Yadav (30:45) Yes, the plunging price of thought.

Yes, by Luke Emerson and David Rudman from Epic AI. They do great work over there, so we’ve shared a number of their articles in the past.

This one, what they did was they looked at the cost of models to accomplish a task over time. And then they mapped out how has the cost changed over time. And one of the biggest findings is the cost has been coming down 47 % every quarter, which is 13 times reduction every year between 2023 and 2020.

So for the past three years, and we already see it, you know, so far in 2026 where the cost of a new model, the model that was premium a few months ago is now accessible to everybody. So the trend is definitely continuing. And then they had a neat chart in there where it’s four times faster than DNA sequencing, six times faster than Moore’s law.

18 times faster than lithium batteries. And they even mapped the cost from like 1873 to 1973. And it’s 54 times faster than the cost of electricity, how that felt over time. it’s just that if you look on the very left, it’s just a very insane down and to the right, which is one of the good graphs in this case that we like.

Dan (32:13) Even more than compute

itself, which is kinda wild when you think about it.

Rahul Yadav (32:18) Yeah. And then a couple of the key things that they found, the models when they initially come out, they’re pricing them at the premium price, but very quickly within weeks, the model price comes down about like, or within a couple of months, they come down by about 66 % because

Shimin (32:41) Mm-hmm.

Rahul Yadav (32:42) they lose that.

premium because you have open source models and then, you know, they’re all competing with each other. So even the labs release a new model and that brings the price of the other one down. And so what that has meant is they measured 03. It costs about 30 cents per question on the GPQA diamond benchmark back in early 2025. And then

By mid-2026, you could use GPT-56 Luna, and it would be 0.0004 cents or dollar. You convert that to cents, it’s one less zero, or two less whatever, the decimals. It’s lower by a factor of 100 in under 18 months, or almost 1,725 times price drop.

in the same level of intelligence. The way that they did this was interesting because they didn’t want to, you could burn a whole lot of money to be like, how do all these different ones perform? But what they did was they partnered with CAISI. It’s like a federal safety initiative, AI safety initiative.

They looked at the transcripts of the benchmark attempts from these different models. And then they kind of cut out some of the parts where the model might say, it costs less, but then it spends a lot of thinking. So it might end up costing more overall to carry out the task. And so they did that cleanup to map out how the cost has fallen. And that’s how they got that data.

Jagged intelligence that show up in their analysis too. In math, the cost has been dropping the fastest. maybe that’s why we also see the math breakthroughs happening. So it’s been about like 50 % to 52 % a quarter. Gaming is slightly lower. You have unstructured puzzles with no.

easy way to verify the proof so you get about like 40 percent there so about 10 percent more.

that

Shimin (34:44) So like what does this mean for how we should use AI, right? Like it I d I love this chart because it basically says that the cost decrease has in the last five years matches about the same order of magnitude as the first like thirty five years of Moore’s Law. So we’re in like the nineteen seventy five ish, the mainframe era of

Compute right now and then give us another five years, so twenty thirty-one, we’re gonna be in a dot com bubble. I don’t know if that’s really gonna be the case, but but the speed is just just incredible. Right. Or the personal

Dan (35:20) Or in the personal computer era. So, you know.

Rahul Yadav (35:23) yeah.

Shimin (35:24) computer error, yeah. Yeah. So

Dan (35:25) Maybe your agent’ll

Rahul Yadav (35:26) bye

Dan (35:26) run on your phone, self contained.

Rahul Yadav (35:29) Yeah, it’s kind of like,

Dan (35:30) Including compute.

Rahul Yadav (35:31) you know, people have started talking about actually, we just talked about the DHH saying in so many words that they have a software factory running at 37 signals now and a lot more people are doing it because if the cost is so low, even for Astra today, you know,

Maybe it’s super pricey and take the latest cloud models. But all you need to do is unless your workflow really requires some like cutting edge model access, all you have to do is wait for a few weeks and then you have access to those models and their cheap prices. And so a year ago, we didn’t have access to Opus five five or whatever a year from now, it’s going to be insane that Astra would just be like, here’s Astra running in every single.

Imagine that level of intelligence.

Dan (36:20) But but also

like even just looking at opuses, like the difference between five one and five five is crazy because five five on low is much cheaper to run than five one on low was.

Shimin (36:33) So yeah,

so I think we should basically treat tokens as if they’re like a thousand times cheaper than they actually are, ‘cause very soon they will be. And like if you’re gonna be building something for the world that we’re gonna be living in a couple of years from now, you might need to take that into account.

Rahul Yadav (36:47) Exactly. And for most things, you don’t need the genius in a data center. You know, if

Shimin (36:52) No, we we

already have trouble telling the difference. Like I was just saying, I have trouble telling the

Rahul Yadav (36:56) Yes.

Shimin (36:56) difference between the Claude response of the AI impact question from the Astra version. I can tell. I’m not like a futurist.

Yeah. And another thing that is also kind of concerning to me is like does this speed of improvement mean that anything with an objective answer that can be RL’d for would not matter in like two to three years period.

Rahul Yadav (37:18) Yeah.

Shimin (37:18) My gut says yes. But I’m gonna stay upbeat this week. And then lastly, the corollary of that is like, does that mean what is actually valuable and rare in a couple of months from now would be s things that are not verifiable. So like, you know, we talk about taste and judgment, but also stuff like creativity, right? Where there is no objective right or wrong answer. So if you’re worried about what you tell your kids,

Not that we have children, that we know of. but if we did, right, like if you have a sibling or something, like maybe tell them to focus on things that does not have objective answers.

Rahul Yadav (37:55) Yeah,

and people have made this argument. I’ve been starting to see more and more of this where in the real world, intelligence is oftentimes not the bottleneck, right? You see, you can point to almost any, like, you know, pick your figurehead, whoever changed the world one way or another. They weren’t the smartest person or the most intelligent, but they had the

right skill set of like playing to their strengths and getting things done to be able to do the things that they wanted to. so intelligence definitely helps. wouldn’t want ideally dumb people, but

Shimin (38:36) Yeah.

Rahul Yadav (38:37) it’s not the answer to everything. At some point you’re like, great, but you still gotta figure out all these other bottlenecks. Intelligence is not the bottleneck.

Dan (38:45) Well it’s kinda like the

the 10X engineer thing, right? Where it was like there was that brief moment of glory for the 10X engineer where they were like, I can be a dick to everybody ‘cause I’m so good and they would and then people got over it and they’re like, you know, we’ll take a five X area that’s actually like a nice human to talk to and

Rahul Yadav (39:02) Nice. Yeah.

Dan (39:03) go from there

Shimin (39:03) Yeah.

Rahul Yadav (39:04) Yeah, because

they cause 100x destruction in exchange for the 10x improvement.

Dan (39:09) Yeah. Yeah.

Shimin (39:10) Yeah. so Gen Z is correct, that you wanna forget about math and focus on the riz Am I right?

Rahul Yadav (39:19) That’s the big takeaway, I think. I Epic really missed out on the headline there.

Shimin (39:24) let’s move on to my post for the week. this is one is by Sean Goedecke which we have mentioned Sean’s stuff before on the show.

Dan (39:34) Mm-hmm.

Shimin (39:35) short and sweet. It is titled Tell Agents the why not just the how where he makes the argument that early age AI agents require you to you know tell them.

Go change method A in class B and then do the same for class C and D. but the models these days are smart enough that what you actually want to focus on is the context and your priorities. So again, going back to like the tastes and judgments, right? What are your priorities? Are the things that these are the things that you value to be important as a software developer? So like readability and

tell your agent what the long term goal, his example here, that have the long term goal of building a local program or browser extension to automatically scan pages for AI content and hide them. that the models are now intelligent enough that these long term goals, these cultural, these, these things that, you know, again, maybe the

10x engineers don’t give a crap about once upon a time, becomes more important than ever rather than just you know, here’s how to do X or Y with a very detailed spec. it reminds me of

If all software developers are managers now, we’ve kind of are doing more and more like kind of the team building side of management things, right? Establishing a common vocabulary for your agents, establishing the team priorities and the coding styles and your KPIs. it’s it’s fascinating to see how how quickly

Dan (41:05) Caught Claude and Kodex

fighting in the bathroom again and I had to have them do team building

Shimin (41:08) Yeah.

Dan (41:10) activities, which cut into my token budget for this month, but it was worth it because now they’re friends. And

Rahul Yadav (41:17) He he he.

Shimin (41:17) Right.

Yeah, exactly.

Rahul Yadav (41:19) Instead of pair programming in humans, get the Claude and Codex to pair together and see if they

Dan (41:25) Ha ha ha.

Rahul Yadav (41:26) come to something better.

And I think the…

Dan (41:27) I mean I think that is a feature built

into to Cursor, but I definitely lately have been

Rahul Yadav (41:31) Yeah.

Dan (41:32) trying to use opposing models to do reviews as a first pass. And then I of course look at it as a human, but

Shimin (41:39) Right. So like what the next kind of natural extension of this is, you know, going over your list of priorities and your long term goals with the agent to kinda help you shape right? Like now we’re giving more and more upper management responsibilities. and then eventually they will replace us. We shall see.

Rahul Yadav (41:57) Or

you, and at some point maybe you don’t even need to say anything. It’s just ambient intelligence where it’s like, I know what email you’re talking about and what I need to do there. And, you know, instead of you necessarily joining.

Shimin (42:13) Right, to do that it requires access of all your emails in context and build a very good personality model of yourself.

Rahul Yadav (42:21) Yes.

Shimin (42:21) exciting times. Looking forward to that day. Okay. yeah, that’s all I got for for this week. this the one last thing I want to say on this is because models have gotten so good, it is now helpful to periodically refactor your skills to make sure all the explicitly spelled out things that you added six months

ago are now removed, right? Yeah. Yeah.

Dan (42:40) Aren’t aren’t hurting it. Yeah, likewise with your agent D and Claude md or whatever

stuff too. Yeah. always

Rahul Yadav (42:46) So.

Dan (42:47) be looking.

Rahul Yadav (42:48) How often, the Hermes example Dan had was

Dan (42:54) Mm-hmm.

Rahul Yadav (42:55) pretty fascinating. maybe I’ll give you guys a quick thought on this. I’m not a big believer in skills in the long term because I don’t know you define them but they don’t really get used as much. And if I have to invoke it just based on the message I sent the agent.

And the harness should be smart enough at this point to be able to like, yeah.

Dan (43:16) Yeah, they do invoke it themselves. Based on the like

the either the task description or other things. And then like that’s actually part of the skill of writing skills is

Rahul Yadav (43:26) Yeah.

Shimin (43:26) Ha ha.

Dan (43:27) making sure it gets invoked when you want it to.

Rahul Yadav (43:30) So

they can also update it, can’t they? Why would you need to update the HNMD or skills? Why do all this meta work when the harness would? Yeah. I see. Yeah, yeah.

Dan (43:37) You don’t. I mean, it wasn’t necessarily for a human to do it. Like you ch should just work with your agent or your harness to to to update it. That’s what I would I

Shimin (43:38) Mm.

Dan (43:46) wouldn’t do it by hand these days. I’d definitely look at it,

Shimin (43:46) Right. But like skills are s

Dan (43:50) but

Shimin (43:51) just just like how software a lot of times is calcified business processes, skills are a lot of times calcified decisions where you make trade offs of why do A versus B. So if you just let the agent does it, then it may override your previous decisions. It unless it has a complete picture of who you are, like I doubt you’ll get every single thing correct, right?

Rahul Yadav (44:12) We put it in GIF if it makes a mistake.

Dan (44:15) Well, w one of the things, like just an example of like what I use skills for, and sorry, we’re going kind of off the rails, but I feel like it’s interesting topic that maybe we don’t talk about enough. is I I work on a lot of different stuff at work, like wildly different, right? day to day. And so it’s just like, you know, a function of like where I’m at within the company and kind of like where I sit at the the confluence of a bunch of different things. And

So I have like two problems that are caused by that. One is like my own attention span, right? And then the second one is telling my boss like what I’ve actually worked on at the end of, you know, the week, the month, the quarter, whatever, right? It all kind of rolls up to like, my gosh, I’ve done so much stuff. I don’t know. So my solution to handling that has been like a a crazy combination of short Apple shortcuts, a Claude skill.

And kind of rigorous note taking on my part, which like I could probably replace it some of it with like a meeting note taker kind of thing, but I just don’t like those. Like I don’t like being recorded by them. And so I’d rather just take notes. And I also feel like taking notes by hand helps your own retention too. So I have a a shortcut that dumps my calendar into a markdown file at once a day in the morning. So it’s not always right because you know, meetings change and whatever, and you have to kind of update it, but like

So I take meeting notes and that. I also have another section that I just like write down like what I worked on and what I was thinking about or whatever. Sometimes it’s like thinking work, other times it’s like hands-on keyboard. or hands-on agent, I should say, these days. And then I have a skill that then collapses all that into a basically a bulleted summary of like what my week was.

And then

Rahul Yadav (45:51) see.

Dan (45:51) that gets dumped every Friday. And then it also looks at the previous month worth of those summaries to look for themes that it needs to pull through. So basically nothing ever gets removed from the previous week’s summary. It just adds new stuff to it. And then I go in and hand remove it because like I don’t trust Claude to like or the agent to like assess the the the themes and everything accurately. I mean it could probably do an okay job, but I just think that like

you know, it I can do it a little bit better. So at least at this point in time. and so then I take that whole summary and I send it to my boss on Fridays. And then when it comes time for like performance reviews and all that kind of stuff, and you need to be like, well, what did I do this quarter or this half or whatever? I’ve got all these cool markdown files that say. So then I can just feed them into an agent again and be like, make me a summary of everything I did this year, key points, you know, and it just

There you go. You don’t have to think about it. You don’t have to go running around scrambling looking at like documents

Rahul Yadav (46:43) Yeah.

Dan (46:46) and other stuff. so it’s been a cool system that’s been working for me so far, but I’ve only been using it for in earnest, probably four months now. But that for that I’ve found like the skill, like I have it’s really literally called Friday summary is the skill. And I made it to template file. So this is like the format that I wanted in. The template has some notes about how to to do the formatting.

And then it also has instructions on look in this folder. This is where they always live. Look for the last couple of them, pull in themes, blah, blah, blah. You know, always add, don’t subtract, and that’s it. So it just saves you from having to like reprompt that thing every time. There’s no code behind

Rahul Yadav (47:21) Yeah.

Dan (47:22) that skill, but it’s just like a I don’t know, it’s pretty handy. So it’s like double vibe and tell for Dan today, but here we are.

Shimin (47:28) Mm-hmm.

and Dan there’s there’s more for ya. your post this week.

Dan (47:32) Yeah.

Rahul Yadav (47:35) yeah.

Dan (47:35) Yeah,

which I was gonna do a pretty deep dive on it, but I’m I may kind of skim through it because I’ve been talking so much already. So so in a nutshell, when I submitted this, I also didn’t realize that it’s from twenty twenty three, so that’s a little interesting. it’s a blog post by I’m gonna slaughter this Baldur’s Gate. No, just getting Baldur, but Jarn s do a Jarnison.

and it’s entitled The LL mentalist, Lementalist effect. LLM, how chat-based large languaging models replicate the mechanisms of psychics con. so I’ll just lead with the fact that in the article, Balder is talking about how like people are tricked by large language models into thinking that they’re sort of like human and or like have you know humanistic responses.

based on the fact that they, you know, do things that replicate the the way a psychic con works. I’m just gonna say up front that that’s not actually why I find this article interesting. Because I think like, you know, it was written several years ago and his whole point was like, these things don’t have intelligence. And I think we’ve actually sort of proven through stuff like the anthropic, you know, kind of like vector slice stuff that they’re done that there’s at least reasoning going on, if not actually like

intelligence, right? So I’m not gonna like relitigate that through this article. But here’s where I thought it was interesting is back in episode 32, we were talking about how LLMs are great at convincing people of things, right? and I’m actually interested in looking at this through that sort of lens. And I think that some of the stuff that he’s picked up here kind of applies to that. So

Let’s let me quickly take you through the what are the parts of a psychic con as he sort of goes through them. And you can actually sort of confirm that via Wikipedia or anywhere else too. But so it’s a good little run through. So when a psychic or other like medium, spiritual medium, whatever, does their con the audience, like the first thing they do is they do audience selection.

A lot of times for something like a psychic or a medium, that’s sort of like a self-selection, right? Cause that like only the people that are out like, yeah, that it makes sense, or they have an open mind would kind of like go pay money for something like that in the first place. So there’s like a definitely a self-selection thing. So once they’re actually in the the con itself, and you have an audience and whatever, there i is a couple of things that typically happens behind the scenes. So first of all

if they have helpers, the helpers go around and kind of like talk to people, just like normal questions. And then they’ll use that data, they’ll like quickly give it to the medium, with notes on like who it was and whatever else to try to like prep them for knowing what their sort of stuff, you know, what what is going on. And then there’s usually also some sort of phase where they’re like really hyping up the psychic, like so and so the grandmaster of whatever and illusions who you know whatever. So that kind of stuff.

the other thing that that they’ll frequently say, which this is the part that I actually think is kind of funny and and does relate to LMs, is that like, you know, remember the rules of talking to a medium, right? So, first of all, things may start kind of murky. We might not know everything because the connection isn’t very strong to the spirit. as as we talk more, the connection will be stronger. And then

Also, errors are just expected. It’s just like a part of the process, right? So that all those things sound kind of familiar from LLMs. So that is pretty funny that he he brought that up. But so then the next step in the the the con is sort of like narrowing it down. So what they do is they will say facts that are like really kind of incredibly broad, right? So like

I I’m getting a vision of someone that has a heart problem with a father figure in their family, right? Well, it’s like, okay, cool. So, like a large number of medical symptoms have chest pain as a primary symptom, and heart disease is the leading cause of death worldwide. Good good chance that at least

Shimin (51:19) Mm-hmm.

Dan (51:20) one person

Rahul Yadav (51:20) Yeah.

Dan (51:21) out of an audience of like, say, 20 people is gonna have someone that matches that. And when you combine that with what they call like the Barnum or the Fort.

For her effect. when confronted with a sort of generalized fact like that, people have this interesting like psychological preference to personalize it for themselves. So they want, they really want to have that be about them. and and the actual like way that this was proven was pretty interesting. So they wrote a test that was like generic and given to this the same to every single student.

In this guy’s like psychology class that was like, here were the results of your like test analysis and and how you are as a person, right? And they were the every single person was handed the exact same piece of paper in the class, despite having turned in completely different stuff. And people ranked it, I think it was like 64 I don’t know, I don’t remember the exact number, but it was like sixty-four percent or something like that of it with it’s it exactly described me. So

Rahul Yadav (52:20) Yeah.

Dan (52:23) Just tells you how it works. And like

Rahul Yadav (52:24) boy.

Dan (52:24) many people,

Shimin (52:24) Yeah.

Dan (52:25) ourselves included, kind of fall victim to that. And this is kind of why I thought it was interesting to talk about this. so then once he’s narrowed down and found someone in the audience through these sort of open-ended questions that reacts, right? So, like, yes, my father had a a heart problem, like that that must be me.

Rahul Yadav (52:40) Yeah.

Dan (52:41) they’ll see it in their eyes or just a quick, you know, physical tell.

Shimin (52:44) Mm-hmm.

Dan (52:44) They’ll zoom in on that person really quickly and then start hammering him with these questions.

And those questions are designed in a variety of ways to kind of draw out additional information. So they’ll do things like like a subjective validation where there’s an intentionally hidden negative, like you don’t play the piano, do you? Or something like that. So then you can kind of like correct them into saying yes, or they’ll do ones that are like staged positive negatives where it’s like,

Nobody, nobody yeah, like you’re a very kind person, but you get angry sometimes. You know what I mean? Like that

Rahul Yadav (53:17) Hahaha

Dan (53:18) kind of stuff. Yeah, exactly. Because yeah, I mean it’s like could be anyone, right? Yeah. Or it could also be someone that needs to go to anger management, you know. So they

Shimin (53:18) That sounds like me, Dan. I I am a very kind person who gets angry sometimes.

Rahul Yadav (53:21) Well, that checks out. Yeah, I’ve seen him do that.

Dan (53:30) would also relate to that. So it’s like those kind of things. so that’s called either cold rate, cold reading or subjective validation. and then they’ll bore in on that one person and they’ll do a subjective validation loop.

Right. And then when the con is done, they’ll essentially the victim really only remembers how things were kind of eerily correct about them because of their own bias. so then we flip to what he yeah. We flip to what he calls the elementalist effect, which is that first of all, like so we’ll go through those same steps again, but pretty quickly. So one skeptics are less likely to use chatbots, right? Because

They might have sort of objections to like LM, you know, intellectual property theft or other things like that that have happened. two, which was sort of like the setting the scene, right? Well, there’s a lot of hype out there already, so there’s already significant priming going on around the hype. three

The the you’re you’re literally typing in what you want in the first place. So that kind of already selects you as the mark from the audience, sort of, and gives the the chat bot quite a bit of context around it about it. the other thing that’s funny to me personally is like a lot of time and I actually have explicit Claude instructions not to do this, LMs will ask you that leading next question, you know, like,

Rahul Yadav (54:41) yeah.

Dan (54:42) and would

Shimin (54:41) Mm-hmm. Yeah.

Dan (54:42) you like me to do that?

Or find out more about this thing or whatever, and there’s always a next hook. That sort of causes the loop that we’ve talked about, where it’s the drilling in like subjective validation loop. And then you walk away thinking, wow, this chatbot’s amazing and it like really reasons and knows me super well. Something to be said for that, but obviously, you know, not not every you know, in this age, I think we have kind of proven that there is at least some reasoning going on there for real.

But the one thing I found really, really funny about this was RLHF. So he talks a little bit about R L RLHF in the blog. And if you think about it, RLHF is humans are self-selecting these answers for themselves as being more correct, right? Because it’s humans in the loop making those decisions about what’s good or not. And so it just kind of proves that we sort of fall for our own.

Or statements, right?

Rahul Yadav (55:32) Yeah.

Dan (55:33) And like, yes, that’s a much more accurate answer than this one.

Rahul Yadav (55:36) Yeah.

Dan (55:37) And so this isn’t like, you know, necessarily some plot to make all this happen. It’s just like that’s what the training did over time.

Rahul Yadav (55:45) Yeah.

Shimin (55:47) Yeah, I I think that’s at least

at least the two thousand and twenty-three version of RLHF did, you know, select for AI synchrofancy because because of that feedback. I do think they’ve corrected some of it since then.

Dan (56:00) Yeah.

Shimin (56:02) I was I guess a little bit playing with fire on my trip over the last two weeks. I was asking the AI lots of questions about the places I’m visiting and then I will play a game with the AI where I asked it to give me essentially cold reads of myself. Like, tell me what is my favorite book, slash

Rahul Yadav (56:19) Yeah.

Shimin (56:19) movies, slash whatever. And I have to say sometimes the AI does a really good job.

Right. Like, for example, one of the AI’s guessed my favorite video game is Disco Elysium, which is like a really low probability guesses. but

Rahul Yadav (56:33) Yeah.

Shimin (56:34) of course a lot of them are wrong too. Right? But I don’t I don’t remember the ones that were wrong. I just remember the ones that I got right. That’s like, my god, how did it know? I love Disco Elysium. That’s such a low probability

Rahul Yadav (56:43) Hehehehehe

Shimin (56:45) guess. and despite us knowing

that hey AI does exhibit some I don’t know if intelligence is the right way, but yeah it uses some form of reasoning, at least it can

Dan (56:56) Reasoning, yeah.

Shimin (56:57) yeah. And you know the blog post says their reasoning is a statistical illusion. But like if it’s a statistical illusion, then you know, a statistical illusion that solves the millennium problem is is a very powerful statistical illusion in s indeed. I think it is

important to hold two contradictory but equally true like opinions about AI at the same time at the same time. Like one of them is yes, they are just matrix multiplications at the end the day. Right? They don’t have a cell, they don’t have organic, whatever, blah blah blah. But also that they are trained to behave like humans. So we should expect human like behaviors from them, you know, based on RLHF or

Dan (57:36) And I

I was having a a really deep conversation with one of my other friends over the past week and and his point was kind of interesting, which is that like humans also have a bias against thinking other things are conscious because it ruins our sort of like ego placement of being special in the universe.

Rahul Yadav (57:54) We’re special.

Dan (57:55) Yeah.

Rahul Yadav (57:55) Yeah.

Shimin (57:56) Yep. Yep.

Dan (57:57) So of course this can’t be, you know, smart or conscious or whatever. I mean above my pay grade, but still

Rahul Yadav (58:01) The… Yeah.

Dan (58:04) interesting to talk about.

Rahul Yadav (58:05) there’s this TED talk by Anil Saif and he says, I’m paraphrasing that, you know, if you look at the, the roughly the architecture of alpha fold and modern LLMs, it’s kind of similar, but no one would say alpha fold is conscious or smart is just

Shimin (58:23) Mm-hmm.

Rahul Yadav (58:24) like predicting or it’s not conscious. is smart. It’s predicting the, you know, shape of proteins and everything.

And so what you’re saying, Shimin is like, can accept that these things are intelligent and yet fooling us or, you know, are trained to act like human when they’re not at the same time. Because we don’t fall for alpha fold, but we fall easily for LLMs talking to us and be like, maybe there is something in there. Yeah,

Shimin (58:50) Maybe it does know me very well. Yes.

Rahul Yadav (58:52) we’re just seeing the different faces of the same or similar underlying architecture.

And one is just much better trained to focus.

Shimin (59:00) Yeah, so I think it’s I think the article is still relevant despite it being, you know, three years old ‘cause I think the problems

Dan (59:04) For sure.

Shimin (59:05) are problems that we’re gonna be tackling with for the next decade. Yeah.

Rahul Yadav (59:08) And it would be

a good like, how easy are you to fall to a psychic? Just try it out in LLM first. And then if you are, don’t consult them. It’s

Shimin (59:16) Ha ha ha.

Dan (59:18) But

Rahul Yadav (59:18) going to cost much more money than a free or $20 couch.

Dan (59:23) But I do all

yeah, ask it for your horoscope. but I also do think

Rahul Yadav (59:26) Yeah.

Dan (59:26) that there is something to be said for these sort of techniques being the gateway into why they’re so persuasive, right? Because like,

Shimin (59:34) Mm-hmm.

Dan (59:35) you know, it’s kinda like some of the stuff happening in like American politics where like people will say fragments of thought and everyone will interpret that to be however they want it to be.

You know, one side is saying, that’s stupid and the other side’s saying, Yes, it’s exactly this and like

Rahul Yadav (59:50) Yeah.

Dan (59:51) yeah, it’s kinda like that, you know.

Shimin (59:53) I don’t know what you’re talking about, Dan.

Rahul Yadav (59:53) in

Shimin (59:54) yep. And and it’s gonna come. It’s this technology would only become a cheaper way to mass propagandize an entire population. So

Rahul Yadav (1:00:02) Yes.

Shimin (1:00:03) far right. Well, speaking of worrisome futures, shall we talk about the

Rahul Yadav (1:00:07) Ha

Shimin (1:00:07) two minutes to midnight?

Dan (1:00:09) What a great transition.

Shimin (1:00:10) where we are you know, where we talk about the state of the financial side of AI, using the bulletin b of atomic scientists’ Armageddon clock as a metaphor where midnight is when the AI bubble will burst. we are at two minutes four minutes and fifteen seconds, and Dan your article for us this week is from The Economist.

Dan (1:00:29) Yes. so it’s entitled NVIDIA is the central bank of AI. it’s a pretty long article, goes through a ton of stuff. So I’m just gonna pull out some highlights, particularly in the interest of of getting to the clock. so a couple choice quotes is one is NVIDIA is walking a fine line between enabling demand and creating it.

So that is in the context of a broader discussion around the circular deals that we’ve kind of talked about quite a bit on here. And Jensen sort of on the record is saying, look, we’re doing this because like we’re helping like, you know, kind of grow this market. And in of course, the critical side of it is saying that like, well, you could also be creating a market. but the big bet that they’re covering in the article is something I think we’ve talked about a lot, which is

You know, NVIDIA’s offering to guarantee their chips, etc. So they’re basically saying that they aren’t devaluing as fast as as people claim that they are devaluing. just my note from from reading this is there’s really two pressures on NVIDIA’s ability to sustain the pricing of their trips chips, right? One is jalapeno, which we’ve talked about before, too, on here in the context of hardware hut, which is

You know, OpenAI’s own chip that they’re making. presumably they’re gonna be able to do that at eventually lower cost. So that’s gonna drive down the price of compute, right? and then the second pressure is AMD is catching up. Like people don’t

Shimin (1:01:52) Mm-hmm.

Dan (1:01:52) talk about that, but like CUDA gave Nvidia a big lead, and now you can actually run CUDA on AMD chips at this point, which is pretty wild.

The other big caveat that I think the article doesn’t talk about is this is ignoring networking. People forget that

Shimin (1:02:06) Mm-hmm.

Dan (1:02:07) NVIDIA, silicon is only, you know, I mean, it’s a big part of their business for sure, but it is not the entirety of their business. And it is not even the thing that they’re like doing the most of with data centers right now. And that is like their interconnects to bring all these like GPUs together are

Essentially gold, and there’s really only one other company making anything close to it. So even if you have someone else’s compute, you’re probably still using their interconnects, right? So they didn’t discuss that at all in there.

Shimin (1:02:34) Hmm.

Dan (1:02:36) but a couple other pull quotes that are kind of interesting. So hyperscalers have investment grade credit ratings, which keep their borrowing costs low. Upstart neo clouds have similar spending needs, but little revenue.

And so their loans are naturally much more expensive. So Alphabet Google’s parent company sold 2.75 billion of the 50-year bonds in November and an annual interest rate of 5.7. The rate at which Coreweave, the biggest neo cloud, borrowed 2.6 billion in July was almost double. So interesting piece of that is that a lot of what Nvidia is doing is actually trying to fund these neo clouds more so than they’re carrying at caring about hyperscalers.

so in in this gap between hyperscalers, this is another pull quote, borrowing costs and everyone else, the bank of NVIDIA would like to narrow. one way that it does that is by taking equity stakes and startups that will be customers themselves or that will help fuel demand for NVIDIA’s chips indirectly. Last year, NVIDIA made about 90 such investments, nearly twice as many as two years earlier. This year it has already agreed to another 60 odd. Other piece that this article answered for me that I found really interesting was why did they buy Hugging Face?

I mean, I knew sort of like intellectually, you know, open models benefit them, but why do they benefit them? Well, it decouples chip demand from hyperscalers if people are running open models. Makes sense, right? But I didn’t think about that when we were talking about it before. so another pull quote underpinning this web of obligations are two fundamental assumptions. Nvidia’s chips will retain their value, which I touched on earlier, and that the demand for compute will continue to grow at a rapid pace. Neither is assured.

Yet both the longevity and value of subchips may simply be artifacts of scarcity, right? So when computers constrained, firms have no choice but to keep like their H one hundreds running. so it’ll be very, very, very interesting is their their sort of like mid to long term point to see if that holds up when the jalapenos of the world hit, right?

Rahul Yadav (1:04:28) Yeah.

Dan (1:04:31) True.

Shimin (1:04:29) Yeah, and don’t forget the Chinese chips, right?

That’s the GLM being served on, yeah.

Dan (1:04:34) Yeah.

but one last point they touched on that I also think is sorry, I’m kinda all over the place, but there’s there’s a huge article. So NVIDIA’s own finances are strong enough to weather their own liabilities. So Morgan Stanley investment bank reckons NVIDIA all in debt that from all the stuff that they’ve been doing is gonna rise to fifty three billion early next year, to two hundred billion by the beginning of twenty twenty nine. Its guarantees come into effect.

But their cash and securities stash right now is worth ninety-nine billion. So really they’re only, you know, a hundred in, which is probably

Shimin (1:05:09) Yeah, it’s not too bad.

Dan (1:05:10) maybe enough to wreck them, but all things considered it’s it’s worth keeping in mind. So yeah, so there you go. Nvidia in a very long mouthy nutshell.

Shimin (1:05:21) But they still have the cash reserves to kind of handle that for now. my article for two minutes to midnight this week is from Reuters about Oracle triggering Force Majeure on data center project over power delays. this is a part of the Stargate project in New Mexico called Project Jupiter, that Oracle was working on with Blue Owl, the I think their private equity company.

We’re talking a lot of debt for this particular data center project. And Oracle essentially told Blue Owl that it is unable to get enough power to this potential data center and that this is a quote force majeure act of god, therefore it is going to delay the payment that Oracle is responsible for for for the project Jupyter. shares of Oracle is down, last I checked.

twelve percent over the last five business days. not super significant, but also like that’s a that’s a decent drop. so

Dan (1:06:16) Mm-hmm.

Shimin (1:06:17) speaking of things that are a little more urgent, I think these kind of debts and our delays may cause the bubble to burst sooner rather than later.

Rahul Yadav (1:06:26) I was not getting power.

Dan (1:06:27) Yeah.

Shimin (1:06:28) That’s what I was wondering.

Dan (1:06:30) If only there were maps that showed you where every

Rahul Yadav (1:06:31) Other than the context.

Dan (1:06:33) single electrical line in the country are. wait, they all do have those maps. Who knew?

Rahul Yadav (1:06:37) I’m just

imagining some lawyer in front of the judge going, but judge, act of God, come on. But come on.

Dan (1:06:42) Yeah But judge, come on, Judge, Judge pal. Yeah.

Shimin (1:06:49) It probably won’t hold up, but you know, it may damage

both Blue Owl and Oracle in the process.

Rahul Yadav (1:06:53) Yeah.

And then finally, we have an article from Ars Technica by Samuel Exxon about AMD acquiring World Labs AI startup for $8.2 billion. They had raised $230 million in 2024. And so that comes out to about a 36x.

of the amount of money that they had raised. Fei-Fen Li is going to become the chief scientist, executive vice president and chief scientist at AMD She’ll report directly to Lisa Su, their CEO, and all the co-founders there staying as well. They recently had their, I think it was called Atlas model come out.

So yeah, it’s purely, when I read this, looked up if they had any reported revenue and their answer was no. And so there’s purely like person and

Dan (1:07:47) Ha ha ha

Rahul Yadav (1:07:48) tech acquisition

Shimin (1:07:49) Mm-hmm.

Rahul Yadav (1:07:50) for $8.2 billion, which.

Dan (1:07:52) You you’re buying that that Fey

Fei name.

Shimin (1:07:55) We we’ve been doing

Rahul Yadav (1:07:55) Yeah.

Shimin (1:07:56) this podcast long enough that I remember when we when we did our episode on when Fei Fei Li’s world world labs like was

Rahul Yadav (1:08:04) Yeah.

Shimin (1:08:04) firstly was we just got started like seven

Dan (1:08:06) Started.

Shimin (1:08:08) seven, nine months ago.

Rahul Yadav (1:08:09) Yeah.

So yeah,

Shimin (1:08:12) Yeah.

Rahul Yadav (1:08:13) subject to regulatory approval, it will close by the end of this year.

Shimin (1:08:17) Well, I guess the purchasing spree means the bubble continues marching on. Everybody still’s got enough money in the tank. Okay, so we were at four minutes and fifteen seconds, as of last recording, a couple of weeks ago. how do we feel about the clock this week? I’ll say the I think the only piece of potential bad news is Oracle’s news that like pushes the clock

Dan (1:08:38) Mm-hmm.

Shimin (1:08:39) forward. everything else seems to be

More or less, you know, same as usual.

Dan (1:08:43) Yeah, I agree. And I still think that we aren’t gonna see excitement until two both of the IPOs happen and

Shimin (1:08:52) Mm-hmm.

Dan (1:08:53) all these data center builds start coming due, right, for hyperscalers ‘cause then they have to start paying for them. So

Shimin (1:08:59) Yeah, fair enough. So next week I think we’re gonna talk about anthropic’s IPO numbers in earnest and then we can make a judgment call based on that. So I say we leave it four minutes and fifteen.

Dan (1:09:09) Sounds good to me.

Shimin (1:09:10) Sweet. And with the setting of the clock, or not setting of the clock, I should say. that’s the end of the show.

Dan (1:09:16) It’s the advocation of all responsibilities.

Shimin (1:09:21) Thank you all for joining us for our study session this week. And if you like the show, if you learned something new, please share the show with a friend. You can also leave us a review on Apple Podcasts or Spotify. It helps people to discover the show and we really appreciate it. And if you have a segment idea, a question for us or a topic you want us to cover,

I know we’ve covered a lot in lightning rounds, but if you want us to cover some of them in more depth, shoot us an email at humans at adipod.ai. We’d love to hear from you.

you can find the full show notes, transcripts, and everything else mentioned today at www.adipod.ai. Thank you again for listening and we’ll catch you next week. Bye.