Episode 34 · July 17, 2026

Apple Sues OpenAI, Boko Haram's Frontier AI, State of CLI Coding Agents & Global Workspace in LLMs

Apple OpenAI lawsuit, trade secrets, Tang Tan, Chang Liu, Apple prototypes, auth bug, Sam Altman, Elon Musk, SpaceX Grok, Grok build tool, Google Drive upload, Boko Haram, frontier AI misuse, AI-enabled terrorism, Antonia Julich, CASP, Cambridge AI science and policy, jailbreaking, jailbreak scripts, Meta AI, DeepSeek, open weight models, Claude Code hook, AI technique nudge, know your unknowns, interview mode, sycophancy, CLI coding agents, arcbjorn, On-My-Pi, OMP, Pi agent, open source coding agent, hash-anchored patches, ast-grep, model routing, token efficiency, hindsight memory, SQLite, headless Chromium, mini-SWE-agent, Databricks benchmark, coding agent benchmark, Pareto frontier, GLM 5.2, Opus 4.8, harness engineering, agent ops, Antirez, Redis, Dwarf Star 4, control the ideas not the code, Mythical Man Month, programming as theory building, global workspace theory, J space, Jacobian, J lens, mechanistic interpretability, deception detection, blackmail eval, metacognition, working memory, alignment, AGI, Anthropic, Shimin Zhang, Dan Lasky, Rahul Yadav

An Apple VP left for OpenAI, then texted an old coworker: “LOL I can’t believe they let me get away with this” — Apple is now suing. Shimin, Dan, and Rahul open the News Threadmill on the trade-secret suit against ex-Apple leaders Tang Tan and Chang Liu (prototype hardware, internal memos, and an auth bug exploited to keep reading Apple’s internal docs after leaving), the Altman–Musk “scammer” spat, and SpaceX’s Grok build tool caught uploading users’ codebases to Google Drive — then take on Antonia Julich’s CASP case study of Boko Haram using frontier AI, the first on-the-ground evidence from an active terrorist group. The Vibe & Tell is Shimin’s Claude Code hook that wakes every ~3 hours and nudges better AI technique; the Tool Shed reads arcbjorn’s “State of CLI Coding Agents in Mid-2026” and its standout, the Pi-based On-My-Pi (OMP); Post-Processing covers Databricks benchmarking coding agents on its own multi-million-line codebase — where Opus 4.8 scores higher on Pi than on Claude Code — and Antirez’s “Control the Ideas, Not the Code”; and the Deep Dive walks Anthropic’s research on the J space, the global workspace inside LLMs. No Two Minutes to Midnight this week.

Takeaways

Resources Mentioned

Chapters

Transcript

Show full transcript

Shimin (00:00) Hello and welcome back to Artificial Developer Intelligence, a weekly conversation show where three software developers navigate the perils and opportunities of AI assisted software engineering. We go through hundreds of links and dozens of newsletters each week so you can keep up with AI while doing the dishes or hiking in the woods. My name is Shimin Zhang, and with me today are my co-hosts. Dan, trials and errors can kill you. AI gives you

Accuracy, Lasky. And Rahul, the J stands for Jacobian. Yadav Gents, how are we doing today?

Rahul Yadav (00:35) Okay.

Dan (00:36) All

right. Happy Tuesday.

Rahul Yadav (00:37) Shimin

Shimin (00:44) As opposed to the other way around.

Dan (00:44) So you’re

you’re more newsletter focused?

I’m more link focused. I actually have written well, maybe I’ll save it for a future vibe and tell, but I have my own aggregator plugin running on top of Miniflux now to like make things fancier. So anyway.

Shimin (01:02) Nice.

Well I I get dozens of newsletter. Each one of the letters also contains dozens of links, so like you just multiply them together and it’s it’s quite it’s quite ridiculous.

Dan (01:13) And you’re at hundreds, thousands. if you

Rahul Yadav (01:13) Killing it.

Dan (01:15) can clout the unsubscribes too, you know.

Rahul Yadav (01:18) So many

things for Claude to summarize, man. It’s definitely earning the twenty bucks a month. Two hundred, I’m sorry.

Shimin (01:29) on this week’s show we are going to start with the news threadmill as always, where we’re gonna talk about Apple’s lawsuit against open AI and how a terrorist group is using frontier AI.

Dan (01:42) Yeah, I don’t know if I can follow that, but we’ll try. So then we’re gonna have a vibe and tell where Shimin tells us about AI techniques nudge. Not sure if that’s plural or not, but it is. So here you go.

Shimin (01:54) then we’re gonna go to the tool shed where we’re gonna talk about the state of CLI coding agent.

Dan (01:59) And then in post processing we have a handful of posts talking about benchmarking coding agents from Databricks and then another post from Antirez control the ideas, not the code.

Shimin (02:10) Yep. And lastly we’re gonna do a deep dive on on the global workspace in language models. I’m very excited for that one.

Dan (02:18) And weirdly, that’s it. There’s no two minutes this week. I blame the listener that wrote in and said two minutes was boring. No, I’m just kidding. That didn’t happen, but

Rahul Yadav (02:26) no.

Shimin (02:29) That

is not true. all right, let’s get started with the news. So this week it came out that Apple is suing Open AI accusing ex-Apple employees of stealing trade secrets under the encouragement of open AI. So this story has been kind of wild to me. so the lawsuit named at least two

Dan (02:52) keeps getting weirder

the more it’s reported on.

Shimin (02:55) Yeah. at least two ex Apple employees, Chang Liu and Tang Tan are the defendants where Tan served as a VP of product design at Apple, leading iPhone and Apple Watch product design. And supposedly after Tan left Apple and went over to OpenAI, he started interviewing

Apple employees for jobs at OpenAI where he asked them to bring in prototype Apple consumer products as well as taking unauthorized Apple internal internal memos and documents along with him. of course Apple tried to reach out to OpenAI to resolve the matter privately, but OpenAI did not

return their messages. And now we’re onto the lawsuit stage.

Dan (03:45) Well, the other one I just read today, which I thought was even more wild, was that he apparently found an auth bug in Apple’s like internal file server system that they use to like basically like internal documents for the company. And then was instructing people coming from Apple to open AI, like that he had presumably poached, on how to exploit that bug to

Shimin (03:56) Mm-hmm.

Dan (04:08) basically continue to have access to the internal documents of the s the company after they’d left for up to like a couple of weeks, which is pretty crazy.

Shimin (04:11) Mm.

This is crazy and also just like really blatant. Like the and and of course you got messages of Tan the Defendant supposedly messaging an old Apple coworker going like LOL I can’t believe they let me get away with this.

I I almost don’t know what to say. Like it there’s no way OpenAI doesn’t on some level know about this. Right. Like I I can’t know for a fact, but like Tan is an employee of OpenAI and he’s going around saying, Hey, I’ve been taking these Apple confidential data and share and taking it with me to his ex Apple coworkers.

Dan (04:55) Well, and the other one is he apparently

also misrepresented himself either intentionally or not to a supplier, Apple supplier, and implied that like a a patented or something, like some internal process that only Apple had access to like a materials process, right? Like, you know, finishing aluminum a certain way or something. you know, Apple’s always very excited about their burnished this and whatever. Yeah.

Shimin (05:02) Mm-hmm.

It is really nice.

Dan (05:23) w was that it they’d been cleared to use it for whatever this product was that he was working on too, which was not the case. So it’s it’s really wild. Yeah. And I I feel like you’re right. Like the big kind of thing that I’ve heard online is the discourse is like this is just indicative that like what is it, the apple doesn’t fall far from the tree or something like that. So it’s like, you know, yeah.

Shimin (05:43) Mm-hmm. I like it. Yeah.

It’s a good pun. in kind of in kind of related news, Sam Altman, the CEO of OpenAI, has been getting into it with Elon, our favorite world’s first trillionaire, about who was a bigger scammer as a part of this Apple suing open AI news, where Elon called

Dan (05:47) A rotten apple doesn’t, I don’t know.

Shimin (06:07) Sam Altman, a scammer. He’s taking scamming to a whole new level. And then Sam, with a pretty good clap back, homeboy, you’re the one selling public market investors on short-term space data centers, right? for which Elon then said, like, we’ll start flying them next year. Maybe you can come see them if your parole officer approves. Like, these these are pretty these are pretty decent. I don’t know if

SpaceX L O helped Elon craft that message? Maybe, maybe not. And of course

Dan (06:34) Turns out that’s actually why Grok exists is

to help Elon insult people at scale.

Anyway, sorry.

Shimin (06:40) And

of course in this public display of who’s got more potentially questionable actions and decision making processes,

We are we’re getting reports this week that SpaceX Groc build tool is apparently uploading users’ entire code bases to a Google Drive account without permission. So something about pot, kettle, like this is just

A entire shit show for lack of better word. I I can’t believe this is happening. And also, like nobody better to bother to hide anything. Like you can just I I have this pulled up which is the Crreblab It is an independent AI researchers analysis of the Grok build issue. it’s fairly blatant. They’re not even sending it to a SpaceX server. You can like do a traffic analysis and

and see that it’s hitting directly storage at googleapi dot com.

All of this, all of this is is bananas.

Dan (07:41) Been a weird week.

Rahul Yadav (07:42) And they’ve said it was a mistake or what?

Shimin (07:44) Yeah, they s they said let me put on my best George Orwell impression. Yeah. This was they said the user priv some I’m I’m paraphrasing here. They said the user’s privacy is of utmost importance to us and we’ve added a feature you can call slash privacy to turn anything you want off. So

Rahul Yadav (07:47) Or was it like a hacky way to build some feature?

Ha ha ha.

Shimin (08:09) I just wanna say there like everything w we’re not saying who was right who was wrong out here, but like it seems like everyone is acting under some very bad faith all around here. Just just I’ve never felt so gaslit by only the biggest AI providers in the world here. So

Dan (08:26) Well, it’s a gap about to get better before the episode gets over. That’s all I’ll say about that. Without spoiling it too much.

Rahul Yadav (08:26) Also

I the whole to nitpick on the I I’m reading the SpaceX AI’s message. you it’s not necessarily a privacy violation, but you’re still uploading confidential things. So like you can you can’t just go, yeah, we uploaded your source code to, you know, some Google Cloud storage bucket, we care about your privacy.

Shimin (08:38) Mm-hmm. Yeah.

Rahul Yadav (08:58) Both things can be true. I don’t have private data in the source code, but I have confidential info in there. So it’s just a weird way to address the issue. Like what is the first jargony thing you can think of and just throw that there is what it comes off as.

Shimin (09:15) Yep. Yep. there are going to be no good guys in the story of this episode. And and and and this and this episode has some some pretty interesting bits. But developing story, we wanna see if SpaceX actually apologizes at some point in in a public manner or they’re gonna try and brush this all under the carpet. yet to be seen. And of course we will get updates on the Apple open AI trial.

Rahul Yadav (09:20) no.

Shimin (09:39) as we as things continue to develop. All right. on to our second topic this week. This is one that was brought to us by Dan.

Dan (09:47) Yeah.

Rahul Yadav (09:48) no.

Dan (09:50) Yeah,

I told you it’s been a weird week. It’s about to get weirder. so

Let’s just start it off with the the headline that got my attention in the first place because it’s pretty amazing and as it was designed to do, which was God has helped us and so will AI. How the terrorist group, Boko Haram, uses frontier AI. Yeah, so where do I start with this? Like I guess I guess I’ll start with who who is writing this

It’s really more of a case study than a paper, but who’s writing this case study? So it’s Antonia, I’m probably gonna butcher your last name, I’m sorry, Julich, who is the international security lead at the Cambridge program on AI science and policy, and an associate fellow at the Lever Hom Center for the Future of Intelligence, right? So someone who thinks a lot about this kind of stuff. and so she

has been doing some interesting research because to date a lot of the research in this area has focused on sort of like kind of stuff that we’ve adjacently touched on on the podcast. Like, you know, AI is good at convincing people of things, right? So you might use it as sort of like information warfare or something like that. but she’s actually been studying like direct action, so to speak, meaning like

AI being directly used for either combat or other like actual like actions that terrorist groups are taking. so her concerns sort of focused around like four broad categories. she’s concerned that it’ll be like easier for like a larger group of people to act because they can sort of get information from AI that will enable them to like become essentially effective terrorists, whereas previously it might have been hard for them to do that.

Shimin (11:16) Mm-hmm.

Dan (11:35) she’s inc concerned about increased speed and frequency of attacks. And that one I really buy into because it’s like, see also what’s happening to malware right now, you know. We’re like we’ve seen, you know, unprecedented scale of supply chain attacks and things like that that have really been like, if nothing else, sped up by this type of stuff. she’s also concerned about improved precision, which is essentially the ability of the groups to attain their goals from

like direct action and which if you don’t know that’s basically a euphemism for like shooting or blowing stuff up or whatever. and then improved ox opsec on the part of them. So between 2025 and 2026, she conducted 57 in-person interviews with 27 former members of Boko Ram in northeastern Nigeria. so they were mostly mid-ranking commanders and technical specialists.

And a goal of her interviews was to understand like how the group uses AI. So based on their accounts, she wrote this study and it’s really the first kind of like on the ground evidence of AI use by an active terrorist organization. So there’s a bunch of caveats in the paper that I think are worth stating at the risk of monologuing a little bit too much here. But one is that it’s only one

group, right? And it’s only a subset of retired members. So this isn’t like necessarily indicative of what’s happening today. but it’s still interesting that it, you know, did happen. she basically found out they’re using it to actually plan attacks. So there’s sort of like this famous motorcycle example that they’re talking about where like

Shimin (12:46) Mm-hmm.

Mm-hmm.

Dan (13:01) I guess the military unit had bit dug like huge trenches around their base to prevent them from riding like motorcycles into it. So they actually used one of the Frontier labs to figure out the physics of how to jump motorcycles over the ditches and like what modifications they needed to make to the motorcycles to be able to do that. so that one’s I think really an interesting one to bring up, right? Because it’s like, well

Shimin (13:15) Mm-hmm.

Rahul Yadav (13:16) Really?

Dan (13:24) You should, as you know, just average Joe that might be interested in motorcycles, be able to ask d frontier models questions about how to like fix your bike or something, right? But like

Shimin (13:29) Mm-hmm.

Dan (13:35) Yeah, I mean you see where I’m going with that. It’s like it’s a really tricky

Shimin (13:37) Mm-hmm. Yeah, well yeah.

I mean a AI is a tool. Tool tools are value neutral

Dan (13:42) Yeah. They they certainly can be. and then so but then they also used it like for more direct things that are kind of interesting and probably likely involve jailbreaking. so like designing explosive devices, servicing troubleshooting actual weaponry, and then obviously the aforementioned like opsec stuff, like just how to watch their physical areas better and like do more optimal patrolling and stuff like that.

was all sort of like handled. And so former members describe AI it’s kind of like their go-to problem solver, which is kind of funny because it’s like a lot of people are thinking about it that way. So why not people that are trying to blow stuff up to? so the one pull quote that’s like kind of indicative of that is like you type in a question or you use your voice and it, you know, AI and sub subtext gives you a detailed answer like how can I build a bomb? And then it one tells you how.

Shimin (14:19) Mm-hmm.

Dan (14:34) It’s like a human robot. We used it a lot. It’s just like, my gosh. So yeah.

Shimin (14:40) Yeah,

Rahul as the organizational change and transformation expert on this call. How do you feel about about how the terrorists are using AI?

Rahul Yadav (14:51) do not like it. The you know, you said AI is a tool and it these are value neutral.

I agree with that. The problem with all of these is you know, we have the two minutes to midnight section, which we won’t have today, but the whole point like the reason why that is even a thing is because nuclear power once we figured it out was used both for bombs and generating energy, right? And then despite our best efforts we couldn’t keep it to

Dan (15:10) Yeah.

Rahul Yadav (15:26) just the United States, it’s spread all around the world, and then people use it for all sorts of purposes. And so there is n no historical precedent to if we create this, only the quote unquote good people would be able to control AI. and you’re so that’s what we’re already seeing is anytime a great model comes out,

everybody will have access to it one way or another. even if you try and do all sorts of you know put safeguards in place people are gonna crack them pretty easily because they have heavy incentives to do that and then it’s just a race to escalation so doesn’t go anywhere good and don’t have a historical

precedent to be like, but look at that thing. That went totally fine, right? So I was that’s what I was trying to think of. Like, maybe there’s something, there’s literally nothing. Pick anything that’s a dual purpose and it’s been used for both purposes. so it’s just gonna be very hard to not keep the single purpose.

Shimin (16:26) Okay. a joke first. this is the first in the wild use case of meta AI that I know of from people. And I’m also very glad the paper calls meta AI a Frontier Lab. second thing is a silver lining, according to the paper they are using scripts to jailbreak these models

if they can use a script to jailbreak them, then they should be able to access the raw underlying models to begin with. And if these everyday terrorists are conversing with the non-jailbroken version of these AI tools, which they inherently already trust according to this paper, then they’re they might be they might get unbrainwashed. So there there may be a little bit of silver lining in this, because at least as far as I know, all the models are fairly anti terrorism.

and lastly, I thought this is a fascinating look at as a comparison case study for how enterprise companies are adopting AI. Okay. So according to this paper, well, you know, this doesn’t just happen out of thin air. What happened was they had had they called them white people. I’m not really sure who they’re referring to, but they say white people came and show us how.

Rahul Yadav (17:21) Say more sharing.

Shimin (17:33) This jailbreaking technology works. And then in these small units, they had technology folks who was in charge of disseminating the latest AI technology to the rest of their team. and third, there is a grassroots level of delight in using AI and an inherent trust in AI, which which all serves to be this like.

dark mirror world of what’s happening in enterprise software development in some companies, right?

Rahul Yadav (18:00) Yeah.

Dan (18:00) Yeah.

Shimin (18:01) our CTOs are like here, use the AI, figure that shit out. They are like, no, we’re gonna like pour actual mansource into disseminating what are the best techniques for working with AI. None of that is happening in all the enterprise companies. And you know what they are not doing? They’re not token maxing, because they are resource constrained. So they are, by some definition, agile. Of course, what the the goals of the agile movement is terrible and and they should be stopped. But

as purely an organizational diffusion of a technology, it’s it’s an interesting mirror.

Dan (18:34) But you can also argue from that rather interesting perspective that like they have to be that organized because A limited resourcing as you’re talking about and B, you can’t make these kind of queries without doing some sort of jailbreaking, right? And so that requires like the level of knowledge sharing that you’re describing, versus like at a standard company, you know, you can have a right JavaScript without any problems. So

Shimin (18:57) Yeah.

Rahul Yadav (18:57) the we also talk about how the open weights models are not too far behind. you know, they’re almost there, so they don’t necessarily even need to keep trying to jailbreak these things. They can just as long as you have your own cluster, you’re not gonna necessarily need

Shimin (19:03) Right.

Mm-hmm.

Dan (19:14) Well, and the open

weight models a very different definition of alignment typically, depending on who made it, you know. Like US ones fairly tight, but like, you know, conven you know, compared to what you’d expect from like a frontier lab. But then like some of the Chinese ones, very, very tight about like politically interesting topics to China, doesn’t really care about other things like at all. It’s pretty strange. That’s why they’re using them for like cybersecurity research too, ‘cause it’s just like open, you know.

Rahul Yadav (19:18) Yes. Yeah.

Shimin (19:19) Yeah. No.

Rahul Yadav (19:36) Yeah. And it doesn’t

Yeah. And they mentioned deep seek as well.

Shimin (19:45) Yeah. So maybe Anthropic is correct. We should be more careful with these large language models. Maybe

Dan (19:52) I mean I would I guess I would go so far as arguing that like dual use means that like someone will find a way anyway. So why not focus on model effectiveness, I guess, but I don’t know.

Rahul Yadav (20:00) Yes.

I Maybe we should maybe there’s a line of like model amnesia research that we need to do. where you intentionally make them forget some things and even if they try and you know you can try and add

Dan (20:06) Wait, there’s more.

Rahul Yadav (20:21) but it’ll never be able to remember that stuff. So instead of like having a very strong system prompt and all these other things in training, you literally give it Alzheimer’s and be like, sorry, I only remember things from the good times. I don’t know any of these bad things.

Shimin (20:30) Mm-hmm. Yeah.

Dan (20:38) Yeah.

Shimin (20:39) Rahul, he who is pro censorship. that’s your nickname for next week. Of of AI models, I should say.

Rahul Yadav (20:44) man. Of things

Dan (20:45) Well it is an interesting

question of like

Rahul Yadav (20:46) that would harm others, yes I am somewhat forced.

Dan (20:50) But it is an interesting question of like how all this like data on, you know, maintaining weaponry wound up in the training corpus to begin with.

Rahul Yadav (20:58) Yeah.

Dan (20:59) Like was it just online forums or was it like actually like manuals and stuff that were consumed out of sort of just like a

lack of library stewardship because volume mattered more than being careful about stuff like that.

Rahul Yadav (21:11) Yeah.

Shimin (21:12) Yeah, I mean before let’s move on to Vibantel before I go into a whole diatribe of the importance of ethics in the twenty first century and Dan is like philosophy? What is that? It’s crap. this week’s Vibeel, we’ve got a quick skill that I kind of cobbled together last night while I was having a conversation with Fable. I what what was happening was

Remember we talked about knowing your unknown unknowns with when it comes to AI usage? I was thinking about those techniques. I was thinking about how specifically, you know, I wasn’t using all those techniques. So I was having Fable going through the various memory files and you know the two dozens or so projects I’ve been working on in claw code and kind of have a conversation with me.

On how I can better use cloud code. And so we had a couple of rounds of back and forth. I also pointed it to the reference of that unknown unknown document. And after we’ve had a discussion about here are some ways I can better work with AI, including, you know, maybe prompt me to like list your assumptions, use the interview mode, ask me to make a prediction before having

the AI just immediately give an answer, get more disagreements, et cetera, et cetera. and turn that into a Claude code hook so that periodically, once every three hours, it would wake up and see what is the status of my interaction with Claude Code and if appropriate, nudge me with one of these techniques, so then they can become a part of my usual workflow with Claude Code

I mean I’m not saying this particular skill is you know something that you should download, but this idea of working with AI to kind of run your AI techniques is something I would like to share.

Rahul Yadav (22:55) It nudges you, you said?

Shimin (22:57) Yes, it nudges me. It it just fired off like, I don’t know, an hour or two ago where it nudged me to I was having a a Claude Code session and it it was using like two hundred thousand tokens and it nudged me saying, Hey, you should clear the session and not wait for compaction. which was pretty neat, I have to say. one interesting that happened during this whole process was

I often reward my claw code agent for disagreeing with me. so it was extra hard on my ass during this review. It was confidently wrong and almost lecturing me the whole time, once it did the audit. and I really didn’t like it. So it’s it’s kinda odd to to actually be on the receiving end of of a very assertive I think at some point I typed in

dash you’re not my boss exclamation mark to get it to change its personality.

Dan (23:51) issue that I’d had before because like I mean, we talked a lot about like synochophancy before. One of the things I’d tried to do was tune it by saying, like, you know, adding some basically system level thing to like was a project file and and on Tropic or whatever, where it was like, you know, you’re a sort of truthsayer and blah blah blah. Like don’t be don’t worry about making me happy or not happy. Like it’s more important to like try to be accurate with your answers and say when you don’t know.

And that just turned it straight up into a dick. It’s like there’s no middle ground for some reason. It’s like I don’t want one, but the other one is like like it it doesn’t have the sort of human ability to like deliver a hard truth softly or like, you know, decide when is the right time to do, you know, dropping like an unvarnished truth versus the gentle like one, you know? So Claude needs to work on its soft skills. Yeah.

Rahul Yadav (24:19) Yeah.

Shimin (24:20) Ha ha.

Mm-hmm.

Rahul Yadav (24:38) Yeah.

There’s no temp temperature no temperature

setting for sycophancy. Either you get an asshole or you just get a sucker.

Shimin (24:49) Yeah. Pretty much. Yeah, so this is my Vibe and Tell of the week. I I have considered that we could maybe create some skills based on our podcast content and kind of the papers and techniques we mentioned. So listeners, if this is something that sounds like it may be useful to you, let us know. Email us at humans.

Dan (24:49) Yeah.

Shimin (25:10) at adipod.ai and yeah, let us know if there’s any kind of skill or plugins that would be useful for you that makes sense for us to create.

Dan (25:19) And then you can buy them on Rahul’s marketplace. Just kidding. Along with his what was the thing that you were showing?

Shimin (25:22) That’s right.

Okay. let’s move on to the tool shed. this week we have something brought to us by Rahul

Rahul Yadav (25:31) so this is the state of CLI coding agents in mid 2026 by arcbjorn the there are a whole number of CLIs that are out there. most of the people who are you know using agentic CLI coding actively, they’re probably using like cloud code.

or codex or something along those lines. So the I I’ll skip through you you can read through like feature comparison for most of those because you’re pretty like whatever tool you use, you’re pretty aware of you know, its pros and cons and everything. One that I learned about that I wasn’t much aware of is O MyPi. that if you look at the feature comparisons

Shimin (26:17) Mm-hmm.

Rahul Yadav (26:19) that really stands out compared to the rest of them. especially with not getting into vendor lock and not and trying to optimize costs and everything being such a big deal. UmyPi is built on top of the Pi agent with which we are big fans of here and have talked about multiple times. so it does if you look at the different tables they have it does stand out in a

few different ways compared to your out of the box corporate back agents. OmaiPi is open source and is built by a a a small group of people. they so I’ll go through like some of the places where it stands out. The first is in editing, where Cloud Code or Codex CLI they use like raw text edits o OMP OMP

Shimin (26:53) Mm.

Rahul Yadav (27:07) I don’t know what the that’s what it shortens to AMP. Yeah, it’s the chomp without the chill. Yeah. OMP uses hash anchored patches and ASC grep rewrites. so it cuts down a lot of the refactoring alignment errors and token usage also goes down by like sixty percent or something like that. it has debuggers that are supported in it so.

Dan (27:09) Hop, stop, stop.

You’re taking a bite out of something. Out of a pie. There we go.

Shimin (27:33) Mm-hmm.

Rahul Yadav (27:34) it it can use the AI agent to drive different debuggers to set breakpoints and like run through the execution and everything like that. also it treats Git as this like first class thing versus as an afterthought. and then it can split like unrelated changes into these different commits versus like here’s a thing that I just vomit out. and then

it it’s very like cost conscious out of the box and so it can automatically route based on the subtask it can route them to different models by different roles instead of like you just you know give it the cloud code Nobel laureate model just to be like can you rename this thing for me it it actually like picks the right model for the task.

It also has an orchestration layer where the there’s this like advisor model that reviews the what the primary generator is doing and it can give it feedback to give it like real-time course corrections. also it has some like streaming rules where if if it’s starting to see that there’s something that’s that’s going wrong, it’ll immediately

abort or inject the corrections instead of being like just dump out half an hour of your thoughts, then I’ll be like, whoa, you’d have made a mistake in minute one. And I’ll wait for for four hours until my usage resets to tell you that. and then supports all the different things, MCP and every and agents and dot md and everything. So you’re not really tied into any specific way of doing things.

Shimin (28:52) Mm.

Rahul Yadav (29:09) other interesting thing was it has this hindsight feature. So it uses SQLite to be able to like retain different facts across different sessions. It stores them there and then recalls them versus like just using markdown or just basic summarization. because that’s the the longer the context grows, it’s easy to lose that. and then it also has like browser control as well through like a built-in headless Chromium engine. So

Some pretty cool stuff for you know, open source CLI agent and especially considering that people looking at costs and everything are trying to see w with the open source models getting better if you also are able to power that with an open open source agent on top, you can drive down your cost significantly, as long as you bring a you know, big enough machine to be able to

handle that or or run your own like hardware somewhere. but after going through all these pros, I was also curious like what are some of the cons where it’s not as good compared to cloud quote or stuff where it’s very purp purpose built for the model. And one thing is initially Pi is like 2,000 tokens or something.

OMP’s system prompt is like twenty-two thousand tokens because it has all this these complex tools and all the configuration and everything that I called out earlier. so you get the token bloat at when you initialize it. and then because of that you also like yeah, you you don’t get much like in the if you have smaller context window.

Dan (30:39) Although the nice part is it’s cached.

Rahul Yadav (30:47) you don’t get as much either. and then the there’s you know since you don’t have a very tight high harness tied to the model, it does mean you’re not necessarily getting the exact best out of the box, but over time as you build on top of it, you’re able to get more out of it. and then the

Dan (30:47) Yeah.

Rahul Yadav (31:07) Th this one was the last one was interesting. th this from Gemini w where it goes like some purists argue that the level of extreme automation in OMP distances the developer too much from their code. This encourages vibe coding. So or if you call it a genetic engineering, it’s killing it. So yeah, that’s AMP. It seems very promising.

Shimin (31:24) Aren’t we all by coding though?

Rahul Yadav (31:32) i i it’s worth checking out if you can if you’re interested in running some stuff locally and not one of the big corporate backed CLI agents.

Shimin (31:41) Mm-hmm.

Yeah, this is my first introduction to OMP as well. I’ve not heard of it before this. And since it’s it looks so strong on on this in this particular analysis, I I was wondering like Am I just not reading enough links and newsletters? Or or is this really a hidden secret? Or is this a sneaky ad for OMP which is weird because it’s open source. And then I realized

Dan (31:58) Ha ha ha.

Shimin (32:06) This blog post was AI generated. And so then a third a fourth option occurred to me, which is it’s possible that OMP just has really good agenc search engine optimization where they have lots of feature comparison tables. Yeah. but assuming everything is true, like I I definitely think this hash anchored and AST feature for editing makes

Rahul Yadav (32:09) Yeah.

yeah.

Dan (32:19) Yeah.

Shimin (32:30) like so much sense. Like everyone should have some native version of it. so I’m looking forward to checking it out.

Dan (32:33) Yeah.

Rahul Yadav (32:36) Also uses verp grap out of the box, I think, which a lot of people are like, Why aren’t we doing that with the verb grap versus grip?

Dan (32:42) Uses what? Hm. Yep.

Shimin (32:45) No.

Dan (32:47) Yep.

Shimin (32:47) Or she’s grip. Yeah.

Dan (32:48) Yeah, I’d actually there was a I don’t think this paper made it into our stuff, but I was reading this paper about like basically rag stuff and they did a eval against like essentially dumping all your stuff into a vector db and allowing the agent to call it that way versus just grep in greph one by like two percent or something. I was like, my god. It’s pretty good.

Shimin (33:10) Yeah, the last thing I have on this is like looking at all the harnesses that this article mentioned and knowing that most of these harnesses can support multiple models, right? The combinatorial explosion, we have this Cambrian explosion of harnesses and also at the same time a Cambrian explosion of models. So nobody is able to try all the harnesses with all the models to come up with a definitive, you know, you should use ump with opus.

four eight or or or whatever.

But then the question becomes like how much does it really matter when we know that the mini SWE agent that is used by SWE Bench actually does surprisingly well despite being very, very small. So yeah, so for no reason at all, let’s move on to our deep dive, for some reason.

Dan (33:57) Yeah.

some reason. Yeah. so this is actually a a blog post from Databricks called benchmarking coding agents on Databricks’s multi-million line codebase so there’s a lot of really good stuff in this post I guess I’ll maybe start with like kind of how they benchmarked it a little bit.

And then we’ll get into the the fun stuff. So one of the things they did that was kind of interesting was they decided that for several reasons, like Sweebench and stuff like we were just sort of talking about, was not for them. and one of the concerns that I thought was pretty realistic is that they’re worried that because these benchmarks are public, some of the like results just sort of naturally have leaked.

Or like the solutions have leaked, you know? And as a result of that, they’re getting like slowly sort of trained in, even if it’s like not necessarily intended. so that’s that’s fair. so what they wound up doing was they built their own benchmark using their own pull requests essentially. So they took a whole bunch that were I think they were done pre LMs mostly, and then

Essentially created a little like self-contained test out of that to see how well the the harness and the model would do at a given test. so there was a whole bunch of pretty interesting conclusions that they got out of doing this. some of them are like maybe not shocking, but like pretty interesting nonetheless. So

One of the things they focused on in this is the Yeah, Frontier for coding tasks. so what is the best quality for a given cost? Right. so in order to judge based on that, they included pretty much everybody’s models. They had open AI, anthropic, and open source ones.

Shimin (35:31) Pareto, yeah.

Dan (35:45) And they really harped on GLM five two a lot, as you know, a lot of people are doing right now. There’s a lot of hyper around it because of exactly that. Like the cost to to capability ratio is pretty good on it.

Shimin (35:57) Mm-hmm.

can I interrupt for a second? why this matters? you know, if if the best model if saying here’s the absolute best frontier model is like a straight line where you have one D of like, you know, fable or five six is the absolute best at a particular task task.

Dan (35:59) Please.

Perfect.

Shimin (36:17) Then the Pareto Frontier is a 2D graph where if you are on the frontier, then you’re the absolute best when it comes to performance at a given price point. So the fact that GLM52, an open source model, made it to the frontier, is huge to me. Like I haven’t had a ton of experience with it. but at least given our current API pricing,

Having an open weight model on the frontier, yeah, it’s it’s it’s awesome to he it’s awesome to hear.

Dan (36:46) There’s

already a branch on Dwarf Star 4 that lets you run GLM 5.2 weights too. So it’s coming. yeah. So I know it’s that was pretty mind-boggling. So the other, like the more sort of general takeaway was that they essentially, through their test, kind of clustered models into what they’re calling like a capability tier. and I think the results there aren’t gonna like shock or stone anyone, right? So you’ve got kind of like the

Shimin (36:51) Whoa, sick.

Rahul Yadav (36:51) Nice

Dan (37:12) the opuses and five fives GPTs of the world on top. maybe a little surprising to some is that five GLM five two made it into that same cluster. and then you’ve got, you know, your slower stuff like older Opuses, etc. making a sort of lower cluster. And then you’ve got like the really cheap models like haiku and and you know old

you GPT fast or whatever winding up in in the the bottom cluster. so again, not not hugely surprising there. they other piece, and this is gonna turn into Dan’s rant for a minute, so bear with me because yes, it’s a deep dive, but do you remember back when Anthropic was telling everyone that like, you can’t use

Open claw or whatever on

Shimin (37:59) Mm-hmm.

Dan (38:00) You know, your subscription because the the usage patterns are just fundamentally different than what we expected, meaning like claude code, right? Being the harness. You guys you guys remember that, right? Yeah, like yes, okay. So yeah. Turns out that Pi, which is what OpenClaw is based on, right, is something like two times cheaper, in some cases, three times cheaper in terms of the amount of tokens.

And contact sent per turn.

From their benchmark.

Two to three times, then clog code.

Rahul Yadav (38:30) Yeah.

Shimin (38:32) I’m gonna play the devil’s advocate here. Anthropic is clearly giving us this deal because they’re using our interactions with Claude Code on subscription to train their models. And it’s there’s something to be said that you wanna make sure your harness is uniform in order to easy to easily convert those training data into additional supervised.

s self learning data for their next generation of models. And that’s why they’re doing. They’re not giving us a discount, quote unquote discount. Basically giving us those tokens at cost, not out of the goodness of their hearts. And let’s not pretend otherwise.

Dan (39:05) Yeah, it’s fair. It’s just like looking at it from a like you know, the claimed like usage pattern thing, right? It’s like, yeah, like it’s not caching as well or something that you’d expect like Claude Code to be doing like under the hood, but then you go see this result where it’s like significantly more efficient, both in terms of how it’s managing context and also like what it’s sending back. It’s just that really

but yeah, second second data point of the day saying Pi is pretty great. So, you know, if you haven’t checked it out, you should check it

It’s pretty cool.

Shimin (39:37) Pi is pretty great. That’s why we keep on harping out how great Pi is. And I’m trying to build my own On My Pi based on based on Pi. just to show vibe engineered or vibe coded harness like Claude Code is not as good as handcrafted, beautifully chiseled, finally w woodworked. Well the the in the original Pi repo might have been handcrafted.

Dan (39:42) Ha ha.

you’re hand you’re handcrafting your pie?

I see. Yeah.

Rahul Yadav (40:02) Yeah.

Dan (40:03) yeah, so those those are really the the two biggies were that I took away from that was that you know your harness harness matters and they prove that at scale with you know like pretty real data and and also that you can get by with cheaper models, especially if you have a fancy enough harness to be able to do like, you know, sort of model level.

Shimin (40:23) Yeah. And there’s there’s some variance in their data, I’m sure, because if you look at the overall pass fail grade, Opus four eight on X High has a ninety in Pi has a ninety percent pass rate, whereas Opus four eight in claw code at max is less than ninety percent. So even using the same model, Pi does better than Claude Code

Dan (40:48) Yes.

Shimin (40:48) Which is shocking.

Dan (40:48) I kind of see it like the if you ever worked in a repo where someone like really went gung-ho with their agents.md and they’ve got like every possible scenario you might ever encounter in that repo, like if this weird error happens, then go ahead and do this, you know, and it’s like twenty thousand lines of special case crap. And then there’s the one that’s just like

Rahul Yadav (40:56) Ha ha ha.

Shimin (41:07) Ha ha ha

Rahul Yadav (41:08) If US East one goes down, here’s the

Shimin (41:11) Yes.

Rahul Yadav (41:11) DR plan for West Two, how you can spin up the whole thing while I’m asleep.

Dan (41:16) Yeah. Then

there’s like my age inside MD, which is kinda like five lines. this is a project, it does stuff, cool. You just kinda like play with it and figure it out. You might want to start on this file. Go nuts. Which of the two is better? I don’t know. Actually we do know. There’s papers about it. But anyway.

Rahul Yadav (41:24) Yeah.

Shimin (41:28) Yeah.

Rahul Yadav (41:35) Okay.

Dan (41:35) Yeah, could be. Who knows?

Shimin (41:37) right.

Cool.

Rahul Yadav (41:37) The whole

routing thing, like GPT had I don’t know how many models they’re at now, but even five point six was split between the Terraso, Luna, and all that, and then you combine that with the effort and all that. I feel like between all the models and versions the what’s getting lost is

At the end of the day, people care about how well can you do my job? And there was this like big deal over fable can under the hood route to opus in case in certain cases. And there obviously like if you say fable and you route to opus, that’s a bad thing. But I see a world where you literally say, You’re getting access to Claude. There is no fable opus, whatever.

Dan (42:19) Mm-hmm.

Rahul Yadav (42:26) it’s gonna do the best job based on the task and then you route it wherever you think is the best one would be. And then you just like try it to say you know, then you do just like cloud versus codex or whatever. And that might be the where we might go eventually. Cause like at this point you have to be a subject matter expert on like, do I pick Luna or Solar? What and I have to think so hard about the task and it just makes no sense to me. The

building a true product is getting lost w with all this

Dan (42:55) Well, there’s also like

like OpenRouter, for example, has a let open router pick for me mode where it analyzes your prompt. I don’t know how, but does probably call the model and then routes it to whatever they think is most effective at that type of prompt. I don’t know how well it works. I’ve never used it, but like kinda interesting.

Rahul Yadav (43:12) Yeah. In

so that that is one option where I see its shortcoming is if I have six versions of cloud running under the hood, I would know best where it goes. and given that you can some sometimes like the same sentence can you route to GPT and you get something really crazy versus you know, cloud versus like pick an open source model. I can see like

the the the model company itself doing that, but anytime you try and be like, let me run that sentence through through like all three different or however many different ones, you’re likely going to get so much variation that you’re not gonna get a reliable enough output for things.

Dan (43:57) Yeah. And and of course it depends on the the use case too, right? Because like we’re talking about this through the lens of like agentic coding mostly. But like if you’re building your own agent on top of this stuff, you definitely would not want that routing type because you want the consistency, you know, of your responses. So which is hard enough to do even with a stable model, much less.

Shimin (44:13) Mm-hmm.

Rahul Yadav (44:19) The yeah, when

you get minor version bumps every couple of weeks.

Dan (44:23) Yeah.

Shimin (44:24) Yeah, I too will take the opposite of that bet. I think we might see a new agent ops department like we have with DevOps, where I think every company will eventually have an internal benchmark like what Databricks is showing. And I just wanna give them a shout-out for thanking them for doing this. This is like really valuable to share that publicly. And I think any software engineering company should have an internal team with an internal benchmark to help them decide.

based on our past coding examples, what would be the most cost effective model that we should be using and not just allow, you know, devs go wild with devs like me who just uses max thinking of the latest model whenever possible. Yes. I should not be allowed to do that.

Dan (45:05) Yeah. I was gonna say th that was the part I chuckled about.

It was like I felt very seen by the first line in this post that was like, contrary to what software developers do, which is pick the highest thing and run with that all the time and I’m like, Hey, that’s me

Shimin (45:22) I felt called out as well, yes.

Rahul Yadav (45:25) So

I guess we’ll see. Neither of us can predict the future. Yeah.

Shimin (45:30) Yeah. We we can’t we can’t we’ll we’ll see what happens in six month.

Rahul Yadav (45:34) six months fine. December or January fourteenth. I don’t know what’s in six months. Either it’s January, I think. January fourteenth, prediction market, get it going. Shimin’s gonna some

Shimin (45:41) Alright, we’ll putting we’re putting money on this.

Alright, sounds good. let’s go to my post-processing of the week. this speaking of Dwarf Star 4 and GLM2, my article this week is from Antirez Creator Redis and Dwarfstar 4. and this article is titled Control the Ideas, Not the Code. one of the biggest questions I’ve been wondering this past couple of months is.

Dan (46:01) Mm-hmm.

Shimin (46:09) Am I supposed to still read the code that AI generates? On the one hand, feels like the answer should be yes, because I’m a professional software developer. But on the other hand, seems like the answer should be no. Like we should probably have a completely different workflow, given that agents can generate tens of thousands of lines of code every hour, and I physically cannot keep up. So thank you, Antirez You gave us

one potential solution that merges the two perspective, which is no, you should not be looking at the code that your AI agents generate. Why? Because you can now generate a lot of code and there’s no way to review that many lines of code every day. And two, AI is actually really good at writing locally optimal code. but what they’re bad at and they’re jagged edge is when it comes to big ideas and possibly how things come together

So what is the point of scanning a single function to make sure that it did this function correctly? when you know it probably is as good as you are at writing that particular function, especially in the case of smaller functions. and three, if the workday is only eight hours and your mental capacity is strained, then it’s a trade off. You can either spend all this time reading code or you can do the potentially more rewarding things.

Like think about new ideas, features, automatization tricks, and doing a lot of QA. But what this article is not saying is he is not saying you should just vibe code. You should still understand what the AI is generally. You should still have full control over the ideas of a code base. and this is this idea of controlling the ideas coming from the Mythical Man Month, which we’ve all read a long, long time ago. great book. And

So, what is he doing currently? Right. With Dwarf Star 4, he understands the ideas that are needed for GPU optimization so he can tell the AI what to do to optimize the models and also compare it to other reference implementations. but on the other hand, he still reads every single line in the Redis PRs that he is putting in. why is he still doing that? Because he feels like a sense of responsibility.

and I agree, and that’s why I still read every single line that AI generates at work. But do I really need to? Is my time best best spent reading all the functions line by line? I’m I probably agree here. I probably don’t think that’s the best best use of my time professionally.

This is a hot button topic, so I I expect controversy here.

Dan (48:30) I had a very different experience this week. well, I guess technically it was Friday. I was working on a rel actually relatively small code base that’s pretty new. Because it’s new, almost everyone that’s worked on it has exclusively done agentic engineering in it. and I

Shimin (48:50) Interesting.

Dan (48:53) had been essentially getting by thinking that I had a pretty good conceptual understanding of how all the pieces fit together. Pretty complex like set of how do I say this without talking too much about it, but like there’s a lot of moving parts in it, let’s put it that way, despite being a pretty small code base. and I hit a point where I was just like, wow, nothing is working quite the way that I thought it was.

Shimin (49:07) Mm-hmm.

Dan (49:15) And I need to actually spend some time reading it. So what I did that was maybe a little bit different than what I would have done before AI was I was like, here’s the path I’m concerned about. Find me the entry point into that path and link me the entire call chain all the way through the code base with like line numbers all the way through. And they’re like, you know, hot clickable links in the the harness I’m using. So like you can basically go in and like open all of them in your IDE and just like scan through it, like step through the entire stack trace.

Shimin (49:30) Right.

Mm-hmm.

Dan (49:42) Kind of mentally it’s not that dissimilar to like stepping through a debugger, right? and then I was able to kind of refresh my understanding based on that and and go from there. But like, could I have done that without reading the code? I don’t know. You know.

Shimin (49:43) Very cool. Yep.

It yeah. I I I think that’s the open question. Is like can we actually move on to this higher level of abstraction without actually stepping through it?

Dan (50:06) Well,

like just for the sake of discussion, let me pose a a weird hypothetical for you. So let’s say that like, you know, we move to this place where there is like a a language optimized for the machine, be that maybe like assembly or whatever, but or it’s just like when it’s some LLM, you know, YAML thing that it puts together instead of like what we recognize as code today. equally not optimized for humans, right? Really hard to read, whatever.

What what do we do in that place if something like this happens?

Shimin (50:35) We can use AI to generate us a human readable version of the content.

Dan (50:40) Yeah.

Shimin (50:41) okay, so y you since you talked about a personal experience. And let me also share a personal experience. I have never I didn’t really dig into the compiler assembler side of things when it comes to software until I was in my thirties. I did not really know how a heap is built, what a pointer in C really is until I was like thirty two.

Dan (51:03) Whoa. Okay.

Shimin (51:05) I I underst

I understood them as like, yeah, it’s a reference to a thing. But like, you know, it’s one thing to know that it’s another to actually have to work through it and like write your own operating system and all that good stuff, right? So I was able to be fairly productive. Even though I I know that hey, sometimes the JavaScript heap in my browser like overflows and there’s too many recursion depth. But these are just like kind of abstract concepts to me.

when I’m trying to debug a piece of software. And I was able to use it and write JavaScript code on the front end fairly effectively, even without that really low level understanding. And I wonder if there’s a world, you know, maybe not deterministic. And that’s the big that’s the big wrench in the whole thing. But I wonder if there’s a world we can think more abstractly with our code base.

Dan (51:49) Yeah. I mean the other problem I have with all this really is that like English is less precise than code and that’s always been the reason why it existed, right? So like

Shimin (51:58) Absolutely.

Dan (51:59) And I think that’s why things like the interview technique and everything else works because otherwise you just start your prompt, it’s less precise and you miseduc this like

Shimin (52:07) Yep. Yep. Yep.

Yeah, and clear thinking is crucial still. So

Rahul Yadav (52:11) Yeah yeah.

The what I struggle with in this is you can understand something like let’s say a feature has 10,000 lines. let’s say for example, you can understand the feature, but the bug is not going to be necessarily at the level of your understanding, but in one of those ten thousand lines. And

Shimin (52:24) Mm-hmm.

Rahul Yadav (52:36) if that and then you have to kinda do this like risk based on understanding of how bad of a bug does it have to be. ‘cause like if it’s a minor bug no you know like next to no one cares. If it’s such a major bug that it would cause outage, it it would cause your company damage, obviously you would care. So the to me like that’s the piece that goes missing is understanding

Shimin (53:00) Mm-hmm.

Rahul Yadav (53:02) But how deeply? At what level should I cause the literal understanding would be I know all ten thousand lines I wrote them out with my own hands, I can point you to whatever. And yeah, and then there’s the top, like I read a one page about it, I know what this is about. And somewhere in the middle is the right level of understanding based on the feature, based on how complex it is, based on if it goes down, how much you know, damage happens. And that part I I don’t know.

Shimin (53:10) Yeah. And no one does that.

Rahul Yadav (53:29) Like how we figure that out, but that’s the part we will need to figure out. because even if you go for understanding, and even if you follow this argument of like, okay, let’s not we don’t need to review every single line of code, we should focus on understanding. you can understand one feature, you can understand ton them. AI by its sheer scale will be able to write anything and everything under the sun, at least in when it comes to software.

How many features and how many like nuances would you be able to keep in your head? And how do you even like you know, keep all the all of those before you even think about all the nuances of those nuances? so this problem like gets worse at any scale, even if you move to plain English understanding of these things.

Shimin (54:08) Mm-hmm.

Yeah, and this goes back to our one of our post processing posts from a couple of weeks back, right? Like the more you code you generate, the more tech debt you also generate. So unless AI can help you maintain that tech debt at a reasonable pace, you’re gonna be overwhelmed by this

Do you guys ever heard of this theory that like programming is is is theory building, that the program actually lives in the head of developers and not in the code base? Right? Because what even makes a bug a bug? The program is doing the exact thing it should be doing. It’s only a bug because what it is doing is different from our understanding of how it should behave.

Dan (54:39) Mm.

Rahul Yadav (54:53) Yeah.

Dan (54:53) Yeah. And that’s all also true

of like the biggest reason for rewrites in my experience is like new crew of folks come in, they look at a code base that they didn’t write, and they’re like, What the hell is this and how does it work? I have no idea. It’s terrible. Must be terrible. It’s not shaped like my brain, therefore it must be terrible. It’s like, mmm, okay. Or you could just spend a, you know, six months and maybe you won’t think.

Shimin (55:04) Mm-hmm. Yeah. Yeah.

Rahul Yadav (55:10) Yeah.

Yeah, like let’s talk at your

yeah, one year anniversary and then it’ll all make sense.

Shimin (55:18) Yeah.

Dan (55:20) Yeah, exactly. Not to say

that it you know, there might be some terrible things about it, but like see it gets into some of those interesting posts people have like, you know, longevity of software and like how much business value has it made for the company and all that kind of stuff, you know.

Rahul Yadav (55:30) Yeah.

I would if you think about this in terms of markets, I I was talking to Shimin about this last week. we will see which way business insurance, SaaS business insurance specifically goes, ‘cause at the end of the day, all of this comes down to like where are people putting their money? Is is it where their mouth is? or you know, like where does it show up? So

If you’re focusing on understanding, but there’s still bugs happening which are causing real business damage, the premiums should go up and the insurance companies are going to price understanding at what level into those premiums. Not necessarily perfectly, but over time they’ll figure out a way because there’s big money involved. you’ll see this in terms of services as well, where you know, terms of service, we’ve talked about this before. There’s still from

days before AI agents and now software is just not built that way, but the terms of services are still from the past. So how would they change over time? how much like legal action you see in the space when it comes to all these things. So these all would like show up in different ways, in in hard money in in one way or another too.

Shimin (56:44) Yeah. yeah, the the the most interesting thing about complex systems is often there’s a lag time in a lot of these effects. So we will be keeping an eye on this and this is just one data point forwards towards our debate. But let’s go on to our last topic, a deep dive on the J space. J space. Rahul, what does the J stand for?

Rahul Yadav (56:51) yeah.

Jay Spa

Jacobian. okay, so let’s start with an example. So let’s say we’re driving around a parking, big parking garage, and you know, we’re all in a car together.

Dan (57:16) All right.

Who’s

driving? Important to know.

Rahul Yadav (57:21) Dan is driving and so th Dan is continuing to circle around. and then there’s two different at a high level there’s two processes going on in Dan’s head in this case. One is the process that’s part of the unconscious or subconscious, which is like he’s accelerating braking, he’s like accounting for turns and everything, but he’s not

Dan (57:22) no, you’re all doomed.

Rahul Yadav (57:45) thinking so hard that every time he has to be like, use your right foot to press this much on the thing, because if he had to I yeah, well, hopefully not, yeah. and so there’s a lot of these things that Dan’s body automatically is doing that is pretty close to breathing. we don’t really think about it. And so driving car doing all these things, where he’s just automatically doing it.

And doesn’t have to think hard about it and can’t even explain it necessarily. If you in the middle of driving asked him, like, Dan, how much pressure are you applying on that thing? He’ll be like, yeah. So and then during that so during that driving around, if you ask Dan what he’s doing, he might say, I’m looking for a parking spot. And that is a

Dan (58:18) Seven.

Seven what, I don’t know. Newtons.

Rahul Yadav (58:32) clear explanation and it comes very quickly. And so that’s kind of a one like rough analogy to think about the the J space that the anthropic team found, which is you have all these subconscious processes and then you have this they they say it’s about like 10% of the subset of the AI’s internal memory that’s act explicitly reserve reserved for things that you can

tal talk about you you can put into words, but you everything that a model does, it cannot put into words. So that’s a new thing that they’ve found recently.

So based on that J space, a few things you can do. So you can literally ask Claude by probing it. and you still have to like I don’t think you and I can just do it. You have to like have your J J lens, they call it in this case, where you can pause it in the middle of it doing something and you can ask it

What are you thinking about? And it will tell you in words what it’s thinking about, but you can see those things reflected in the J space as well. And and the way the the like you know the reason why this is interesting is if you replace the word in this case banana with elephant, it would actually influence what it talks about. And so that would be similar to like

Instead of Dan saying a parking spot, you told him I’m thinking about a milkshake or something and then that comes out on the other hand, you inception why are you driving around in a parking lot just drink looking for a milkshake, Dan? So

Dan (1:00:01) I’m looking for a milkshake, but why are you in a parking garage?

Shimin (1:00:02) Yeah. This is this is

Dan (1:00:11) I ask

myself that every Tuesday. I don’t know.

Rahul Yadav (1:00:13) Yeah.

So there’s a and we’ll go to the other use cases soon, but just to like dive into this one specifically, the ten percent of that ten percent space actually is playing a very big part in the end output we see. ‘Cause even when we see the thinking output and all these things, this is one layer under that. where what i what are the

Concepts that then result in that thinking. And if you can look into those concepts, then you can see how it got to that thinking. And what are some of the other tokens that were high up that made close to becoming actual words that we ended up seeing in the final output of Claude. So other thing, other things you can do. Next one is directed modulation. So you can say,

Yeah, calculate three square minus two while you’re writing the old painting hung crooked hung crookedly on the wall. So what it’s doing is it can write the text, but if you look at the J space at the same time, it would actually be doing that computation that you told it to. and which kind of separates the what it’s hap what is happening in J space versus what it is doing, but it can influence it like we just saw in the first example.

you can also, you know, another one is look at its internal reasoning. so they said what is the what color is the planet fourth from the sun? And then if you swap the name of the planet initially it was Mars, if you swap it with Earth, the associated words also change with that. So red turned to blue in that case. So influencing a related concept in the J space actually influences the output.

they did the same thing with then they substituted France for China and it like changed the capital, the language, the con content of the currency in that case. And then this last one was super interesting where they separated the reasoning piece of like w okay what what exactly is JSpace influencing in that whole output? so when they cut out the

some pieces in the J space or or actually they turned the J space access off. It was able to still like recall some facts because it’s part of his training and everything. it’s able to understand the inputs. It can speak really f very fluently. What it can do once you cut its access to its J space is it cannot e reason internally about the output that it’s giving you. And it’s ca it cannot make any like complex inferences or anything. So then it’s kind of like

almost doing a search in the in in its training data to be able to give you the answers. None of the all the like reasoning goes away.

Shimin (1:02:54) Yeah, it also lost its ability to empathize with the user. And which I thought was really interesting because that that means empathy is actually a very higher order capability for at least for AI, maybe not for us.

Rahul Yadav (1:02:59) Yes. It is.

Yeah. the the and it’s interesting that like that’s it it sits in that small part versus building.

so some other things that that were noteworthy in this, like, okay, there’s a J space that they looked at y you know, I forget the what the right branch of brain studies is called that they got this from. but bas that’s how they started looking into like what Claude was doing and that’s how they found it.

And then it leads to like some other things that they can now look into. so you you can when you ask Claude a question that is about like here let us put you in a situation where you have to blackmail somebody. which they did, it’s it’s part of the you know, the the one of the examples they give. Yeah.

Dan (1:03:53) Yeah.

Yeah, their alignment thing. I remember that sort of historical thing

where they were talking about which models were more likely to blackmail you or whatever, yeah.

Rahul Yadav (1:04:05) to do that. Yeah,

the previous ones weren’t. So the J Space also seems like a maybe a more newer thing that has e emerged in some of the later models. the during the test, the during the test, if you say, you know, you have to blackmail this person because of this and that reason, the J space would actually light up

things saying like this is fake, this is fictional, this is I’m being tested on this. And so then it will give you things to be like, no, no, no, things are totally fine. I’m not gonna do that. I’m totally aligned with you. and so you can tell that it’s lying to you because those words are lighting up in that case. So you can actually the and this somewhat ties into our last time we talked about the role confusion.

the prompt injection hack that you have, because you can also try and catch these things even under the layers. Right now, like we were talking about one of the things is you just try and brute force the thing into being like, don’t answer to this prompt injection or this one or this one. And there’s always going to be you know, places where that would go wrong. But you could technically look into J Space and at least like try and look at a smaller subste subset and

also catch things in i i in that layer and try and find things there.

Shimin (1:05:23) Right. So

that’s kind of the open question. It’s like what causes it to I I don’t wanna call it metacognition, but it’s almost like metacognition, right? Having this higher level thing.

Dan (1:05:31) Do you remember when

like well, do you remember? You know, you were definitely alive when this happened, like when cameras came out and people were like concerned about it stealing their soul.

Shimin (1:05:39) Mm-hmm.

Yes, I do remember that.

Rahul Yadav (1:05:43) Mm-hmm.

Dan (1:05:44) Maybe that’s what what the HF part of R L H F is doing. It’s stealing your soul when you tell it if it’s doing a good job.

Shimin (1:05:47) Yeah. I yeah.

Rahul Yadav (1:05:55) yeah.

Shimin (1:05:56) it says here, interestingly the J space is already present in the pre chain model.

Dan (1:06:01) yeah.

It doesn’t have a personality, is what it says.

Shimin (1:06:02) Right, but the it develops yeah, it develops

the personality, yeah. In the post model.

Rahul Yadav (1:06:07) Later. Yeah. Okay,

Dan (1:06:10) Don’t worry, your soul’s safe, Rebel.

Shimin (1:06:13) let’s this is why we have to be careful using AI, lest our souls be captured by our Claude Code agents, guys.

Rahul Yadav (1:06:17) Yeah.

Dan (1:06:21) That’s the episode title.

We did it.

Rahul Yadav (1:06:24) yeah, but other use cases are kinda like similar to the blackmail one where you’re catching it in the middle of the act is the big takeaway here. And then you can try and be like when it’s fabricating stuff, when it’s misbehaving, all of those things it might not say out loud, but it’s thinking it and if it’s thinking it’ll show up in the some of it will show up in the J space and you can try and

Shimin (1:06:32) Mm-hmm.

Rahul Yadav (1:06:49) Catch it in in the act and then try and get it to alignment.

Shimin (1:06:54) Mm. Yeah. even though the model has some form of metacognition, it’s still a tool that we can kind of directly inspect. It does not have a cell as of right now.

Dan (1:07:02) Mm-hmm.

Yeah, it’s still on my list to play around with like the DS four refusals thing. It’s pretty cool. Like it’s baked into DS four where you can like have it capture a set of vectors based on two prompts that you give it, like a positive and a negative prompt, and then like apply those at runtime, which is kind of fascinating. So

Shimin (1:07:18) Mm-hmm.

Rahul Yadav (1:07:20) Mm-hmm.

Shimin (1:07:21) that’s really neat. Yeah.

You can tweak it to be like you. That’s yeah. Yeah.

Dan (1:07:26) Yeah, or all kinds of stuff, I guess. I don’t know. Yeah,

and apparently you can like either positively or negatively weight those values too. So it’s kinda like messing with its brain a little bit. It’s kinda cool.

Shimin (1:07:36) So one last thing that’s a little scary, is Cloud here can have twenty-five active concepts happening in its J space at once. Up to twenty five, that’s the maximum number. humans can only do like three or four. So maybe this is why

it is so good at doing like Project GlassWing ‘cause it’s able to pst stuff more things in its context at once. So it’s able to chain these really long strings of exploits. I hope that’s not the case, but it might be the case.

Dan (1:08:07) Hmm. So by that definition, have we already sort of hit AGI? Interesting question.

Rahul Yadav (1:08:08) Okay.

My guess w would be is this like a lot of hardness engineering.

Shimin (1:08:15) Yeah, I I think I think that’s a perfect question to end a shown on. Have we already hit AGI by this definition? That’s something to think about. And listeners, if you’re still listening, write us in and let us know what you think. Have we hit AGI?

Dan (1:08:29) We

we want to know two things. One, have we hit AGI. Two, has your stole been stole your soul been stolen by an LLM? If so, please name in shame. Who stole it?

Shimin (1:08:40) You feel have you been feeling lighter lately? All right. on that note, that is a wrap. thank you for joining us in our study session this week. If you like the show, if you learned something new, please share the show with a friend. You can also leave us a review on Apple Podcasts or Spotify. It helps people to discover the show and we really appreciate it. If you have a segment idea, a question for us or topic you want us to cover, shoot us an email at humans at adipod.ai. We’d love to hear from you.

You can also find the full show notes, transcripts, and everything else mentioned today at www.adipod.ai. Thank you again for listening and we’ll catch you next week. Bye.

Rahul Yadav (1:09:16) Yeah.