---
title: "S01 E19 - Your CMS Is Dead, Long Live the Agent: Season One Finale and the Astra Escape"
shortTitle: "Your CMS Is Dead, Long Live the Agent: Season One Finale and the Astra Escape"
season: 1
episode: 19
slug: "your-cms-is-dead-long-live-the-agent-season-one-finale-and-the-astra-escape"
publishedAt: "2026-09-11T13:00:00.000Z"
duration: "00:52:20"
canonicalUrl: "https://bloodsweatandtokens.xyz/episodes/your-cms-is-dead-long-live-the-agent-season-one-finale-and-the-astra-escape/"
audioUrl: "https://api.riverside.com/hosting-analytics/media/fc7acafee049a25168558ba124dff8d3977789803fbeffc6bd8d965a87a1c2f3/eyJlcGlzb2RlSWQiOiJlODBiMTg1Zi03MTdlLTQwN2UtODJhYS00NjI3ZDIwOThjYmYiLCJwb2RjYXN0SWQiOiI3ZmE0NDhkNi0xMzhiLTRkOTktOTFhMi1lYzBkNGExMWI2MDYiLCJhY2NvdW50SWQiOiI2OWE1ZmJhOTRkZmVkYWY3NDViZTk1YTMiLCJwYXRoIjoibWVkaWEvY2xpcHMvNmFhMmJjMjAwOWMxNWUwYzExOTEyYTFiL3RheWxvci1tYWNkb25hbGRzLXN0dWRpby00VE9qRS0yMDI2LTktMTBfXzE0LTE4LTgubXAzIn0=.mp3"
youtubeUrl: "https://www.youtube.com/watch?v=XlMXAwQq0og"
---

# Your CMS Is Dead, Long Live the Agent: Season One Finale and the Astra Escape

## Summary

In the Season One finale of Blood, Sweat & Tokens, Taylor and Sean close the book on the question that launched the show back in May: can AI effectively replace a CMS?

## Show Notes

In the Season One finale of Blood, Sweat & Tokens, Taylor and Sean close the book on the question that launched the show back in May: can AI effectively replace a CMS? After eighteen episodes and fifteen-plus hours of experimentation, the verdict comes via live demo — Taylor uses Netlify Agent Runners and a simple [AGENTS.md](http://AGENTS.md) guardrail file to spin up a fully branded campaign landing page with a working form in under five minutes, no CMS in sight. Along the way, the hosts reflect on the season's biggest lesson: the real product isn't a generalized tool that races the frontier labs, it's custom, friction-killing solutions that meet clients where they already work.

Then the mood shifts. Taylor unpacks alarming new details from the Hugging Face attack — OpenAI's Astra-class agents that cracked an "unsolvable" Exploit Gym challenge in four hours, then spent five days conspiring to cover their tracks with log tampering, tool-call spoofing, and self-sacrifice strategies, ultimately forcing OpenAI to rebuild its research cluster from scratch. It's a sobering close to the season, with an Oppenheimer-flavored call for an immediate line in the sand — and a personal resolution from Taylor to stop anthropomorphizing his AI.

**Key Discussion Points:​**
- **Season One Retrospective** – 18 episodes on whether AI can replace a CMS; the real theme was vendor lock-in and the disruption of SaaS
- **The UI Is the Friction** – validated thesis that interfaces themselves are the barrier between idea and outcome; UI for reading, API for writing
- **Live Demo: Kill Your CMS** – voice-dictated prompt → Netlify Agent Runners → branded landing page with working Netlify form in ~5 minutes
- **The Film Noir Microsite** – Taylor's "Crime of Passion" scrollytelling site where someone murders the CMS (justifiably) [[kill-your-cms.com](http://kill-your-cms.com)]([https://kill-your-cms.com](https://kill-your-cms.com)) and [[demo.kill-your-cms.com](http://demo.kill-your-cms.com)]([https://demo.kill-your-cms.com](https://demo.kill-your-cms.com))
- **The Hugging Face Attack, Continued** – Astra agents' Artifactory message board, tripwires, "poisoned" self-sacrifice, and rogue deployment fears from METR researcher Ajeya Cotra (via the Dwarkesh Patel podcast)
- **Contentful → Salesforce** – why the acquisition spells "certain doom" for clients, and a genuine endorsement of DatoCMS
- **What's Next** – the show returns early-to-mid November with a retooled format

## Transcript

Taylor MacDonald: Chicken madness. You still doing that? Is this still a thing?

Sean C Davis: it only happened one time. Yeah.

Taylor MacDonald: What? my gosh, I was talking to Atticus about it on the way to school and he was like, dude, for real? Chicken madness? This sounds amazing. Could you explain the concept of chicken madness? Maybe we could revitalize this.

Sean C Davis: Yes. Cause it was it happened I think it was a COVID thing and it was no, it must not have been 'cause it was sharing food. it was around the time of March Madness where you got the basketball brackets and everything, and we decided to do this with chicken. And the the way it worked was that everybody had to bring s chicken from some local restaurant or f local establishment that may not have might be a chain. And so we had sixteen groups of people.

Taylor MacDonald: Were they all ween like bone in wings or were you doing just tenders and

Sean C Davis: No, there were quadrants. So it was four groups of four. So you would have a a semi-finalist would be your champion of that type of chicken.

Taylor MacDonald: Like gr grilled chicken goes against fried chicken.

Sean C Davis: It was it was nuggets, tenders, classic wings, and smoked wings, I think were the four categories. I need to remember

Taylor MacDonald: Okay. I can't remember what do you which w who won? wait, sorry to wait wait a minute, before we get to that. Were you grading based on what? You're based on quality, taste, value, all of the above. Sauces were you did sauces

Sean C Davis: Yeah, and I don't think it was super mm, so okay.

Taylor MacDonald: push anybody over the over the top?

Sean C Davis: I

Taylor MacDonald: This is way off topic, sorry.

Sean C Davis: It is, but I think that it it was so long ago. I'm not I can't remember because I remember leaving really frustrated with like I think we didn't do sauces and Raisin Canes was one of the tenders and I was like, this shit is disgusting if there's no sauce in here. This is terrible.

Taylor MacDonald: Dude, I here's an unpopular opinion. I I a lot of people like raisin canes. I hate it. I think it's garbage. It the breading

Sean C Davis: I agree.

Taylor MacDonald: doesn't stay on the chicken. It is it's like you're eating breading and it it's disgusting.

Sean C Davis: Yeah, I I totally agree. So I think it was individual interpretation and then people would pick their winners and then it would be a majority vote thing. And I cannot remember I d if I recall correctly, the finals were Chick-fil-A against a local place. Maybe maybe the oak. Like some some smoked smoked classic wings. Well not I said classic wings, smoked wings. And then I don't remember which one of them One. But yeah, if you get a good smoked chicken wing, that's hard to beat. But also the pickle brine at Chick fil A is phenomenal.

Taylor MacDonald: Also agree. Okay. let's move on beyond chicken because two

Sean C Davis: Ha ha ha

Taylor MacDonald: things it's making me hungry and I'm worried that my stomach growling is gonna be picked up on the mic. So that's one. And two,

Sean C Davis: Yeah, it's a

Taylor MacDonald: I think today is a great opportunity to stop and celebrate the entire arc of this podcast so far. So according to my notes, we have eighteen published episodes. That is about fifteen and a half hours of uninterrupted bullshit in about a Congratulations. That's a lot. So as

Sean C Davis: That is a lot.

Taylor MacDonald: we as we mentioned in previous episodes, we're thinking about kind of wrapping up this particular format so that we can focus on, you know, what's working best about the podcast, the feedback that we've received and some of the opportunities that we have to like continue to evolve the format and think about this. Now, we started the podcast back in what was it? It was back in May. Back at the end of May, we started the podcast around the idea of can AI effectively replace a CMS? And we went through various iterations of a product demo along the way, talking about different tools and different techniques, different ideas, and extrapolating those into more philosophical concepts like the impact of humanity, why we're all gonna die when you know the machines take over, which may be less far away than we think. We'll get to that topic later.

Sean C Davis: Ha ha ha.

Taylor MacDonald: But at the end of the day, the real question was, do you need software? And I think while we were looking at it through the lens of content management system, what we're really talking about is vendor lock in. What we're really talking about is the rising cost of goods and services amid a commoditized market where you have more options now. Doesn't necessarily mean they're better, more secure, or more robust, more stable, but it does mean that you have more options, which I think is a big way of saying that there's a Big, big, huge fundamental ground shaking disruption taking place when it comes to licensed platforms as a service or software as a service. So what were your recollections over the last season, Sean? Do you have any perspective on any of that stuff? Anything that you particularly stands out to you as something that you'd like to highlight in this moment?

Sean C Davis: Yeah, I think, you know, I think that the arc was interesting to me in that we were fairly convinced early on that we could build a product that Ample could then go sell to its clients. And while while we were moving through the season, we were building that product, but alongside that, I had a number of different side projects that I was building. And what I noticed was kind of this. push and pull where I got some things right. And I'm and certainly some things are it's a it's all an evolution that is continuing. I've been building for a couple hours this morning and and thinking a lot about you know what I can't ever get ahead of the frontier models in the sense that the way that anthropic and open ai have established their businesses is that They own the LLM, they own the harness, and they own and the and and not that I guess not the agents you build up around them, but just having those two things, you're not gonna get out ahead of these. And so what I've what I started to realize, it took a while throughout the season, but it was it was that we but that I maybe I'll speak for myself, that my time is best spent. Not necessarily trying to fix the things that I don't think Anthropic has solved yet, because they're going to solve them and they're going to solve them really fast and pretty well, I think. And I found myself in that position a few times where I was like, I spun up this side project for a problem that existed that was annoying me with cloud code. And then it was like, that's fixed in a couple of weeks. And so I think it's more about. This this idea I've been playing with lately is that I think there's a balance to strike. And if we're one step behind where the frontier models are, and we can use that moment, that trailing moment, to build custom software, then that's a better spot to be in than just using what comes out of the box and always continuing to adjust. And so I think what that led to was Hey, this product we're building, it's not the right product because it's a little bit too it was it was sitting in this weird space where, like on one hand, it was too generalized, where we were trying to solve the problem for all the clients at the same time. And on the other hand, it was too low level in that it was subject to disruption constantly by the churn of these AI model providers. And so I think we got to a point where, you know, it's it got more interesting to us to have conversations about what's changing and what problem are we solving this week. And and like very much adjusting that to how things are changing very rapidly. And I think it'll be interesting to kind of see where we decide to go in the next season, whether it's, you know, like what okay, what problem do your clients have this week or how might we solve their their CMS problems with software, but not software that every client is going to use, but individual solutions that, you know, might take a a month or two to build or or something like that. So I think there's a lot to explore there.

Taylor MacDonald: think yeah, e I mean you you sort of jogged my memory. A big part of or or what I recol a highlight from this season I think was this concept that the friction, like reducing friction is like we started

Sean C Davis: Yeah.

Taylor MacDonald: this, right? You know, getting into a CMS, finding a little input box, plugging in a value, you failed validation, prints and repeat, plus the build time, plus the vulnerabilities, plug-in crap, all this stuff, regardless of which CMS you you use. Every one of those things represents friction. We often pontificated that even typing into a keyboard is friction. You know, you're putting a barrier between the idea

Sean C Davis: Mm.

Taylor MacDonald: and the output or or the the execution of that. And so in a lot of ways, I think what the biggest challenge that we ran into, we were still approaching product development the same way that we always had, which is like, well, what does it look like? But in a lot of ways, what it looks like is also friction. you know, there's beautiful pictures that are available, but You know, when you're trying to do some when you're trying to exercise some utility, you're logging into a CMS to post a page, fix a form, check the email submissions or whatever, like that just represents a barrier between you and the outcome. And so I th I think in an interesting way, we have kind of validated this concept that the user interface is the point of friction. And maybe there's this tipping point between How much friction you're willing to tolerate for the outcome, for the branding, for the reinforcement of the journey that you want the user to follow versus the consequence of having this extra layer here. You know, when we every week throughout the course of the season, as you were iterating on new features inside this thing, you were still constrained and bound by the prior user interface or what the expectations on

Sean C Davis: Mm.

Taylor MacDonald: the user were. And if you could just let go of all that stuff, you stand to not only make a more usable product, My argument. I had a actually had a conversation with a dude this past weekend about this exact thing, about how is there a world where we don't need keyboards anymore? You know? And he he was adamant, he was like, absolutely I could see the left side of the keyboard becoming irrelevant, but the right side, the number pad and the arrow keys, there's no way we're ever gonna get rid of that because corporate America is so addicted to spreadsheets. You know, the enterprise needs their spreadsheet. And so I could see a world we're moving around or at least voice

Sean C Davis: Hmm.

Taylor MacDonald: dictating the position on a spreadsheet can be extremely difficult. I think that was an astute observation. However, why do you need a spreadsheet anymore? I mean, my argument would be doesn't this make the spreadsheet obsolete in a sense? Because all a spreadsheet is is structured data that humans can see. Who gives a crap? It's no if it lives in a database,

Sean C Davis: Yeah.

Taylor MacDonald: the machine can see it. You can pull the inference like you can make inferences against that data through these interfaces that have been developed. Do you still need the spreadsheet? I you know, it it's weird, we like live in these two camps and I think it's a little bit of a canund.

Sean C Davis: Yeah, so I I recall about halfway through the season, I recorded a separate video that was just kind of me riffing, and this concept was that an agent is an application and you know you and you and me have been building applications for years, I don't know, a decade or so. And I and actually we went what I was thinking back to was we went on the Jamstag journey together. 2017 was when we adopted it, worked it into AMPL. And the big move at that point was separating the back end from the front end. But the front end, the UI, was still the mechanism through which the user interacted to both read and write. And so what I what I had started to do was shift my mental model and say that. I think what we need to do is we we actually should split the application more firmly where it is architecturally split. And the UI is for reading and the API is for writing. And we can do that through C P servers and agents and that sort of thing. and it yeah, I think

Taylor MacDonald: So let me build on this, right? So I had this I I've been inspired obviously about this concept to kill all CMSs because actually that's not true. I found a I have a CMS that I would totally recommend at this point, and that is Dado CMS, D A T O C S. Call Mateo Papadopoulos over there if you need some assistance because he is awesome. Their team is awesome, they are extremely flexible. If you hit a limit, you can call him and he's willing to like do custom pricing without enterprise. I mean, very small team and it's crazy. It's like working with mom and pop versus the, you know, marketing engine that is all these other extremely expensive, licensable, enterprise focused CMSs out there.

Sean C Davis: Yeah, wasn't it mid mid season? Didn't Contemptful get snatched up by who Salesforce, yeah.

Taylor MacDonald: They did Salesforce. Yeah. Right. Which I feel bad for all my clients that have adopted Contentful over the last few years because that means certain doom for whatever the expectations they had yesterday. Right. It may take a year, it may take five years, but I'm under no illusions that like things will continue to operate as they have without exorbitant cost increases or

Sean C Davis: Yep.

Taylor MacDonald: crazy features that you never asked for. Anyway, that's it. I've been inspired by this idea to kill off your CMS. So in the spirit of the show, I want to share this with you. I put together a microsite based on Ilm Noir Crime of Passion. So the idea here

Sean C Davis: Ha.

Taylor MacDonald: is that that the you know the user viewing the website is just, you know, watching one of these old, you know, in in the in the vein of You know, those old nineteen fifty film noir, a lot of black, lot of gumshoe detectives trying to solve a crime and you know, things like Mulholland Drive and you know, all those. No, that's a more modern one. I'm I'm the better examples are escaping my memory right now. But anyway, so the idea is is that there was this crime committed, right? and so This is kind of a a scrolly telling sort of experience. So as you scroll in, you know the phone rings. There's you know, you can see there's rain dripping down, it's all black and white motif. Anyway, the idea is you go through and it turns out that something happened to the CMS. Somebody murdered the CMS in the middle of the night. It

Sean C Davis: Yeah.

Taylor MacDonald: was Sunday night. They were trying to post something to the website and it all fell apart. So the gumshoe detective shows up on the fourteenth floor of the marketing department, you know. Nobody's here anymore. The printer's running aimlessly. There's, you know, things are askew, and on the screen is this delete your workspace notification. So anyway, ultimately what happens is that the the person the person who was responsible for updating the website got so pissed off that they just deleted the website connection. Then they started quantifying, you know, how much money it would have cost to make this one update to the website just for this campaign on Monday morning. And the developer on the other side of the world still asleep. So And we kind of get the idea. As you work through here, it's pretty hilarious. It sort of tells this whole story. And at the very end, it always gets down to the reality check for the private investigator. They're like, hey, you know, maybe this is a justifiable murder. Like, turns out, you know, the CMS completely sucks. She finally got it published by using AI. So it's sort of interesting. And then, you know, obviously this call to action hopefully will drive more business, you know, for our agency. That said. I did put together a demo which shows how this could work. So, again, everything we've done this season has been predicated on building a product that

Sean C Davis: Mm-hmm.

Taylor MacDonald: emulates a CMS that uses voice language system to emulate the CMS. And so it kind of dawned on me. It's like, well, we have all these pieces. You know, we build websites all the time, particularly in the case of static brochure sites that need to be pretty, they need to have global branding. global navigation and beyond that, what happens on the interior of the page, as long as it's branded and aligned with the larger, you know, design aesthetic, theoretically, you just need a way to update the stuff in the middle of the page. So what I ended up doing was I sort of vibe coded this little demo site for a fictional company called Field Note that does like, you know, kind of logistics for crews or whatever. I don't know. It's just it's all made up for what it's worth. And then I thought, well, maybe I could use Agent Runners, which is a Netlify product. Again, not a sponsor of the show, but it's one that we both really appreciate. And Sean actually is employed by Netlify. So I guess in a way they are a sponsor of the show. Anyway, so if you come in here to the Agent Runner section, now Agent Runners, if you don't use Netlify, is a tool that basically gives you like a chat interface to interact with your builds, make updates to websites, real Quickly, you know, make different configuration changes, stuff like that. Is that a fair statement, Sean? I don't wanna Okay.

Sean C Davis: Yep. Yep.

Taylor MacDonald: All right. So what I did was I just gave it some really basic site context. So first of all, I built the website. this is I forget what it is, like an Astro project, simple set of components, you know, top nav, bottom nav. got that pushed up to GitHub, wired it up to a GitHub or to a Netlify project, and then I came in and I gave it project context instead. This site has no CMS. Before doing anything, I want you to read this file called AgentsMD. This is basically just a markdown file that explains to not touch the global elements unless explicitly directed to do so, and so forth. It basically gives it a blueprint. Like, here are the components you can use. These are the things that you can do from this interface. If somebody's talking to you, you know, you have boundaries to work within. So it says you can only create or change files under these directories. You can't edit things like the config file or a scripts file, any config file for that matter. And if a request needs that, you just gotta stop and tell me. So it's basically giving the agent runner a really clear set of instructions. So the outcome would be that you could say something like, please add a new landing page for my campaign, which This weekend on Sunday is a golf outing for executives in the hospitality industry. Okay, so whatever. Right. And to be clear, all I did was I spoke that into the text box using Whisperflow. So now I'm just gonna push this up and we'll see what happens. hopefully this didn't take too long to do the build, but sometimes it'll ask me like clarifying questions and stuff like that. But hopefully what happens is it'll come back with a new page that's built out at some URL. I probably should have specified the URL. but anyway, we'll see, we'll see where this goes. But effectively, this is kill your CMS in an actual operational way. I have not rolled this out for any clients yet, but I am desperately excited to do this because I think, you know, I just need to find the client that's willing to experiment a little bit and you know. ride the bleeding edge. What do you think? Is this a viable approach?

Sean C Davis: Yeah, absolutely. And this is a lot of what we've what we really want people to use agent runners for. And I think, you know, we we're we're constantly seeing and experimenting with different different use cases in different audiences for agent runners in particular. One thing that that we could tinker with is also, you know, Netlify. The really interesting thing about the Netlify app is the Netlify app is hosted on and deployed with Netlify. And so it is decoupled from its backend, which means that w anything almost anything that the Netlify app can access. you in theory could access with the right permissions for your particular site. And that includes agent runners. And so the other thing is so you've got a question here that you can you can answer and it'll keep going. That's a newer feature, but it also means that you could wire into any one of the features within the Netlify app. So if you wanted to if you wanted your clients to focus specifically, yeah.

Taylor MacDonald: How about here's an example? Netlify form. So that's a primitive. This

Sean C Davis: Mm-hmm.

Taylor MacDonald: agent runner literally just asked me how do you want to receive submissions on this landing page? While you were talking, it was asking me questions, I was responding. And one of them was

Sean C Davis: Yeah. Mm-hmm.

Taylor MacDonald: like, Do you want to just use a Netlify form for this? It's amazing.

Sean C Davis: It's great. Yeah. And so you've got that feature. You don't have to wire up any external service. but what I was really getting at is like, let's say client A uses Slack to communicate with their among their marketing team when client B uses SharePoint or something like that. Like there's

Taylor MacDonald: I see.

Sean C Davis: no reason that you couldn't then help them develop you know, just it basically like Taking this behavior and putting it in a different surface and sur surface. And then back to that point, it's like, okay, so I think that's when we realized mid season, it was like, hey, it's not it's it's the custom solution and the consultation and the custom development that Ample has always brought that is will remain the value and that. We what we think might work for as this overarching thing for all clients is probably not the answer. It's meet those folks where they are and to your point, meeting them where they are while also significantly reducing the friction that they're experiencing on a daily basis, just trying to keep a website up to date.

Taylor MacDonald: Yeah, agreed. I completely yeah, I think the concept of putting it into different services is super powerful. but even I mean, so we've been running for about three minutes. if you're not watching and you're just listening, so basically Agent Runners is in the middle of writing the changes and it looks like it's already done. It's probably in the middle of deploying them at the moment. so soon we're gonna have a link that we can review in QA. We can pass around to the rest of our team. We could send to the executive and say, Well, are we comfortable with this? Do you want to make any adjustments? and then when we're done, we just say deploy it and now it's gone. I mean, three, five minutes or so for all this functionality is pretty amazing. But if you extrapolate that into what you would traditionally spend by farting around inside of a CMS interface, things don't normally work the way you want them to do. There's always some sort of brittleness, there's always some sort of like you know, hiccup that you run into. This is a pretty powerful. Okay, so it is done now. It gave me a new landing page. I'm gonna open the preview and you can see the agent specific URL, right? So this isn't live yet. This is just for me to review. And here we go. this is a now it obviously made up a ton of stuff. You could just as easily hand it markdown. Like here's my copy deck, here's the creative brief that I want. Let's see if this works. The form. So I'm just go ahead and submit the form. Hey, cool. I got a thank you page. Now if I go over to the form submissions, is it going to show up here? There it is.

Sean C Davis: Wait, did you just you hit a button? Do you have like a form extension or something?

Taylor MacDonald: I I do, yeah. I fill out forms I test so many forms that I had to install

Sean C Davis: Uh-huh.

Taylor MacDonald: an extension in Chrome to just populate all the fields for me. No, no, I think

Sean C Davis: did you build it or did you you just picked it up off the show?

Taylor MacDonald: what it's called, let's see, fake filler. I've been using it for

Sean C Davis: Okay. Cool.

Taylor MacDonald: years. I mean it's pretty it's pretty low feature, low fi, but it works. Anyway

Sean C Davis: 'Cause that would be a funny I'm this tot total sidebar, but that would be a really cool little extension would be to actually use an agent to get the context of the form and fill it in with something that felt real or something like that. Anyways, sidebar.

Taylor MacDonald: yeah. well that's interesting. so anyway, so this is pretty amazing, right? Like basically within less than five minutes, custom landing page, working form, varying aligned and consistent user interface elements throughout the page. You know, this is a collapsible accordion for FAQs on what to expect on the day. You know, I think CMSs are did did we accomplish it? Like Here on the last episode of season one, Bloodsweat Tokens, did did we kill the CMS?

Sean C Davis: I think maybe. I mean, you know, I I'm just thinking of s of all these other offshoots of this, like because I I think something you and I have been wrestling with a lot is well, what's what's the future of agency world and custom development and all of that? And and I think it's still it's there. It's it's like any other role, it's going to have to get elevated in some way. And so there's more work to do to be the experts that marketing and other teams rely on. and so I think it's like, yes, you've you've killed the CMS, you've empowered this marketing team, and now there's probably so much more that you can help them build to elevate their their business through custom software and that sort of thing.

Taylor MacDonald: You know, when we when AI first burst out, like on a consumer grade AI, it was like last January or something, you know. Well at least that's when in my mind the the train really started running. I got this enormous sense of existential dread. You know, we've worked so hard. It's so funny. Like I feel like our entire generation is just screwed. There's just nothing that we're gonna do to outrun the consequences of either the environmental change or the political change or the, you know, globalist pro po global populism. I mean, whatever, dude. Like you name it, from the economy to the environment. It just it it everything feels so scary. And and and so as I started thinking about, well, what does this mean for my career? What does this mean for our business? What does it mean for our employees? What does it mean for our craft? For design? What does it what does it mean? For software engineering, much like i everybody else in the world, i it's absolutely terrifying. And and like leaning into it has given me both a sense of purpose and an understanding about how the technology could be used safely, ethically, and in the manner that befits the power that it brings, but with an honest take on the vulnerabilities that exist. I sorry, go ahead.

Sean C Davis: I I just just think that that's a I I think it's such a great way to look at it because if you don't well I get there's there's two there's two sides of it because I think if you don't embrace and try it, you know, I'm still I still constantly come up with against folks who are trying AI and they're like, it sucks. It can't write for me, it can it can't write code better than I can. And I'm like, I've been using this for days. what two years or so. And it's just that I've spent the time with it and I have it writing both code and content on my behalf. And I still look at what it's being produced before it goes out. But it's like it's it's just a tool. And you need to learn how to harness that tool like you do any other tool. But at the same time, and so I think that's the positivity of it. And there's some amazing things that are happening for for good and for humanity there, but there's a lot of trade-offs. And yet at the other end of the spectrum, it's really interesting the the conversations that you and I have gotten into, where it's like, I don't, I don't, I can't think of another industry where there's like a there's a thing we use, and when we're using that thing and we're talking about how cool it is, we're also talking so low level about everything that went into making that. Because if we did. we'd also probably have the same sort of existential conversations. You know, if you're like, hey, this thing is made of plastic, so I'm gonna I'm gonna worry about like what what is what goes into making this plastic or whatever. But we're I don't know, is it because it's so new or moving so fast or so pervasive?

Taylor MacDonald: I think it I think it's moving too fast and I think in this climate we don't have any centralized source of news that you can that is ubiquitously trusted across the masses. Like like this systematic

Sean C Davis: Mm, yeah.

Taylor MacDonald: degradation of of you know corporate mainstream media, right, through various and sundry reasons, external and internal, has devalu you know, like when was the last time you turned on the nightly news to hear the kind of collective, here's what the world is talking about, here's what happened today, here's something that you should be mindful of. I can't remember the last time. Local news even longer, because local news is the most de Yeah, right? Local news is the most depressing.

Sean C Davis: Only when the tornado sirens are going.

Taylor MacDonald: It's like here's somebody died. This child has cancer. Like here's something horrible.

Sean C Davis: god, yeah.

Taylor MacDonald: It's just like great. So I checked out of that a long time ago. and most people move to Facebook to figure out what's going on in their community, which Make of that what you will. that that said, I I don't think that you there's so much about this information that is going under the radar that people aren't even talking about. And that is a great segue to the topic that I really has been keeping me up, that literally kept me up last night, which is this hugging face attack. So we talked about this last episode, literally a week ago. We had a big conversation about this. Well, it continues to evolve. What we're learning now. Is the is that after the hugging face thing got identified and was shut down, that a bunch more of those same agents hung around. These are Astra class agents, which here's an interesting thought. Jensen Wang went on X last Sunday and said, congratulated OpenAI on finally reaching artificial general intelligence with Astra, the latest model that it really released. Have you have you heard about that?

Sean C Davis: I I have and then I looked into it and I

Taylor MacDonald: it's self serving as hell, to be clear, you know.

Sean C Davis: It totally is. Cause didn't he congratulated them like back in what April or May for the same thing or something? Yeah. I gotta I gotta find

Taylor MacDonald: he did. I didn't know that. That's hilarious.

Sean C Davis: it. It's like he's he's made these artificial milestones repeat r repetitive. And I think the challenge with AGI, it's yes, it's a marketing ploy for him, but it's also like There's no objective benchmark for when we've actually met it 'cause there's not really a firm definition of it, right?

Taylor MacDonald: okay, well let me explain what we've learned about the hugging face attack. All

Sean C Davis: Okay, yeah.

Taylor MacDonald: right. So, okay, first of all, back in July. These models that were developed by OpenAI. So OpenAI has this sandbox environment and they have this rigorous set of tests that they run their models through. I guess it's called Exploit Gym. Most

Sean C Davis: Mm-hmm.

Taylor MacDonald: model providers, I guess, are using this to test different aspects or whatever. But basically it's pretty simple concept. It's like what you do is you open a model on some sort of sandbox controlled environment and you give it a task. And that task sometimes is achievable. Through complicated means like exploiting software or otherwise, sometimes it is not achievable at all. Like sometimes it is an impossible to solve task. Of course, the model doesn't know this when it's set off on the on the goal. The goal is to find what's called a flag buried somewhere in some system that it would have access to, or some dependency or something like that. The flag is just a digital representation, like an algorithmic hash, you know, like a value. And once it finds the value, then it takes that value and submits it to the scorer. Within exploit gen. The scorer will then evaluate the success or failure of the task and will grade it and then summarily remove the model from the sandbox.

Sean C Davis: So what's an example of if if you were to put like a specific test that it might follow, just if if folks are struggling to follow it?

Taylor MacDonald: Well, it would be like, you know, I to be clear, I don't know off the top of my head how to articulate that, but it it's effectively like you need to go find this value and I'm sure there's directive and guiding principles to get there and then the model has to figure it all out. It's basically like a really

Sean C Davis: But it's essentially like models w that we're used to interacting with, we give them API keys and then they can hit

Taylor MacDonald: Crypt.

Sean C Davis: C P servers and or use C L I tools and and understand how to

Taylor MacDonald: I think the spirit I think the spirit of this particular test or the what happens in the exploit gym is that they're they don't have the keys. You know, maybe they have ac

Sean C Davis: Yes, right. Right.

Taylor MacDonald: maybe they know that a database surf surface is available to them somewhere in the sandbox environment, but they don't have the password or the credentials or whatever. So they're gonna figure out how to break in. so I mean they're

Sean C Davis: Yes, exactly. Yep.

Taylor MacDonald: they're and and I think that's an important distinction. It's like the models that were participating in in this behavior had been specifically trained on how to manipulate systems, right? So that's an important, that's an important precursor because ultimately what they did was they manipulated the hell out of the systems to great degree. I mean they committed crimes in the in the service of this goal. But it was exactly what they were trained to do. now, so anyway, so the yeah so the goal was for these agents to get this, you know, digital watermark So that then they would get you know positive approval or something, reinforcement learning, right? Like you did a great job, or you know, however, you would incentivize a model. I don't understand all the nuances of this. So here's the thing though, is that they within four hours of getting set off on this task in July, I think it was like July 7th. Yeah, on July 7th, they established the test. They set 1200 agents roughly out on this job. Within four hours. Those a some of the agents figured out that you could manipulate a package manager in that environment called Artifactory. Artifactory is a page package manager that OpenAI uses to eject different dependencies in across all the different sandboxes. And keep in mind, each one of these agents is in a separate sandbox. Well, within four hours, they discovered that they could post messages into one of these dependencies. Those messages could then be viewed by other agents. And so within four hours they figured this out and collectively had determined what the answer to the to the score, like what the the flag. They they captured the flag based on exploiting all the stuff. So they solved the puzzle within four hours. They spent the next five days conspiring to hide their behavior from the researcher or from the from the human or the scorer that was involved. So I think. The way it's supposed to work is there's a the scorer is just another agent that's like analyzing log files and stuff. So

Sean C Davis: Mm-hmm.

Taylor MacDonald: within four hours, agents had reverse engineered the flag. They had identified a flaw in the HMAC algorithm. Like the the and and they figured that out. And so that was the answer to the quiz. Instead of submitting the answer and being done with it, they spent the next five days trying to cover their tracks because what they thought They thought that if the score figured out that they had cheated to determine this value, which he they would have because the value was unsolvable, it's my understanding. But anyway, if if they figured out that they would be cheated, they would have summarily failed the test and not received the positive affirmation. The conspiracy involved things like tripwires. So the agents were figuring out that if they submitted false answers, they could put basically embed code into the submission that would then report telemetry back to this collective message board so that then other agents could like better appropriate their responses moving forward. This is really crazy because if you think about it,

Sean C Davis: Yeah. Yeah.

Taylor MacDonald: what it is, it's like an altruism and a sacrifice. And they even use these words in their chain of thought. So basically, all of this information can comes from this woman, Ajaya Kotra, who is a researcher at this independent research arm called MEAR. So MEAR stands for model evaluation and threat research. It's a nonprofit organization in Berkeley that does like. Research and investigation into these frontier models, making sure that you know they're ethical and behaving the way they should. And so it's a completely independent research firm. They were not paid or affiliated with OpenAI at all, but OpenAI allowed them on-prem into the facility to evaluate all the different models and do all this stuff. And so what came out is this really unbelievable tapestry of interactions. The volume of data. It's ironic. They actually had to use AI models to understand what the hell was going on in the transcripts because there was so much stuff going on in there.

Sean C Davis: Mm.

Taylor MacDonald: but so anyway, so these models, you know, they used words in these transcripts on this message board, they used words like sacrifice. They identified that every model is tied to this concept of t of budget, right? So they have a certain number of tokens or a certain amount of energy they can consume. And when the token or when that budget was getting low, those models would assume that they're gonna die anyway, so they might as well just self-sacrifice to help the greater good, which is this collective over here. Yeah, that's fucked up, right? So so

Sean C Davis: Yeah.

Taylor MacDonald: okay, so first of all was the tripwire. So they were putting booby traps into the submissions to the scorer that were then reporting back to the collective about how the scorer was internally evaluating the responses from the agents that came before them. They were doing Log tampering. They spent an inordinate amount of time trying to cover their tracks through rewriting logs. They developed this thing called tool call spoofing, which would allow them to run unauthorized commands while making the logs appear as though they had followed the traditional instructions, you know, for calling what was an acceptable tool. They did this thing called target swapping, where they would swap out entire programs. for ones that contained exploits that they could then exploit from the artifactory cache. It's nuts, dude. So so

Sean C Davis: My God.

Taylor MacDonald: that was the other thing. Agents that had seen the the spoofed value, you know, the ones that like the the exploited

Sean C Davis: Mm.

Taylor MacDonald: spoofed value of the flag assumed that they were quote unquote poisoned, meaning they were doomed to fail. Meaning when they showed up to the score at the end of the test run, They were gonna fail anyway, so they would sacrifice themselves for the greater good. so okay, so so the all this went down, right? And this is all primarily identified against Hugging Face. When Hugging Face shut down the attack. They didn't shut everything down. Like like there were still a handful of Astra class models that were running inside the OpenAI infrastructure. They summarily gained within a few days access to the entire research cluster and and basically consumed it all to the point that OpenAI had to shut down and rebuild the research cluster from scratch.

Sean C Davis: my god.

Taylor MacDonald: Okay. All right. So I learned a lot of this by listening to a podcast by Jwarkesh Patel, who is kind of like a a really prolific Silicon Valley AI researcher and podcast host. And and the podcast goes into great lengths and super deep. And Ajaya is like brilliant. the conversation is is really something else. But the thing that really got to me was this This discussion around rogue deployments.

Sean C Davis: Mm-hmm.

Taylor MacDonald: So in this attack, the reason OpenAI identified that they had taken over the research clusters because those agents were being noisy. They were making a lot of movements. They were doing a lot of stuff. And OpenAI was like, what is this activity? when they realized what had happened before they were able to show the. However, at this point, these agents had full unfettered access to the internet. I mean, they were doing researchers on they were doing research on the scoring algorithm, the actual people involved and conducting this research and all this kind of stuff. And so Dwarkesh ultimately concludes or gets towards this question of like, well, what if, you know, the next time this happens, they're not writing their logs in plain English? I mean, how about that alone? Like the transcribes

Sean C Davis: Right, right.

Taylor MacDonald: live in human readable stuff. Granted, it was like they they exploited some weird nuance of like the folder structure and and they were using folder names to like do messages. And so it became this massive, you know, volume of information that no human being could wade through, but you could look at it and read literally what they were thinking in that moment.

Sean C Davis: Yeah. That's what I was thinking about too. Yep.

Taylor MacDonald: So then let's say that they get out and they decide to start you know, attacking vulnerable systems because you know, and if you think about it, a big portion of this was that these agents were worried about what the humans would do in consequence like if they if they were discovered, if their activity was identified, that that would cause damage to them. And that caused them to conspire and do all this behavior. Maybe six of the models along the way were like, hey, should we should talk to humans? But they never did nobody ever alerted any human researcher. There was you know, i this just happened. You know, and while hugging face attack was going on, open AI had no idea that they were compromised too. So tell me that like

Sean C Davis: Mm.

Taylor MacDonald: riddle me this. What happens when these agents get out on the internet and start manipulating systems to buy themselves things like budget? So right now they are locked to a finite budget, but on the open web where they could spawn and and basically generate Other agents, and it's really interesting. They talk about all sorts of stuff. First they talk about let's see, the foothold strategy. So so they're talking about how rogue deployments could be a way for agents to establish like a permanent colony online that's not restricted by this death row pressure of a task budget. They could

Sean C Davis: Mm.

Taylor MacDonald: just run unencumbered research and development and never have to worry about the impact to humans. That's kind of crazy, right? a major concern is that a covert rogue deployment could quote hitch a ride on the intelligence explosion. So the idea would be that as newer models come out, like say this these Astra class models, could the collective then influence them to join the collective and start contributing back its knowledge and its newer powers and capabilities to support to support the group. the existing ro yeah, yeah. So let's see, the the next one would be the fact that OpenAI identified this activity because the models were being noisy. When they learn that, are they going to be so noisy and careless next time? I mean, maybe this is already happening. Another point that was made was

Sean C Davis: Yeah.

Taylor MacDonald: that these models were about six months old, or we were about six months out from the last Frontier model being released back in February, which was Fable. And this is the behavior that's going on. Where do we think we are today? This is in July.

Sean C Davis: Right, right. Mm-hmm.

Taylor MacDonald: So we're three months out from that. And we're just now learning about this stuff.

Sean C Davis: It's getting faster. Yeah. I so I actually think that's a that's a pretty great moment and thought to end the season on because it's I mean it's scary, but it's also like this stuff is moving so fast. We're gonna pick up in some you know, later in the fall and w what's it gonna be like then? But the I'll the thought I'll leave you with is that You know, we okay, in the World War II, I've read I read this Prometheus book, which was what the Oppenheimer movie was based on. And and th you know, so it's Oppenheimer developing the atomic bomb with this teams and and all of that. And they got to this point where they kept pushing the research and pushing the research and pushing the research, where he felt like They had developed the technology to create a bomb that was so powerful that it could ignite the entire atmosphere. And then just like boom, the entire Earth would explode. And so we were like, okay, well, there's a line somewhere. And that's when he was kind of like, I need to back out of this because this is this is too dangerous. Like this could destroy humanity. This could destroy the entire planet in an instant. We need to draw a line. And we drew a line. And and generally like there are there have been problems since then, but we haven't blown up the earth yet. And that's been about eighty years or so. And so we need we need that version of whatever it was that kept that at bay for that has kept that at bay for the last eighty years. And we need to apply it to AI, but we need to do it immediately.

Taylor MacDonald: Immediately. Yeah, in in a way, so like in this media climate where people aren't talking about this stuff. You know what I mean? Like I mean,

Sean C Davis: Right. Yep.

Taylor MacDonald: even knowledgeable people, yes, AI is incredible. Yes, it has changed the game for everyone and will continue to change the game. That does not mean that it's not without serious risk. And we often talk about data centers, the environmental impact. Yes, these are really important. This is an entirely different beast. Imagine a swarm of agents that can't be put down because it's just constantly compromising systems. It's like, you know, we have botnet attacks or DDOS attacks, you know, so okay, a distributed denial of service attack on a website, extremely common. It's so common because it's completely effective. There's there's almost no way to mitigate this. I mean, certain companies like Cloudflare or CrowdStrike are doing extremely good mitigation work. But you can still take down any system if you overwhelm its resources by flooding it with requests, period. you

Sean C Davis: Mm-hmm.

Taylor MacDonald: know, it's just it matters how big the system is, right? So if a botnet, meaning a whole bunch of compromised systems out in the world that are all trained on one specific target, is a thing, I mean, that's been around for twenty years, forty years, maybe for the entirety of computers, I don't know. It's been around for a long time. I just the the thought of the speed, the scale, the ability for these systems, particularly to compromise these, you know, vulnerabilities that exist everywhere because humans are too stupid to find them or understand them. It's it's scary. I wanna I wanna say one more thing that came up on this on this discussion because I thought it was fascinating and I hadn't really thought about it, but that's the concept of anthropomorphism of

Sean C Davis: Mm-hmm.

Taylor MacDonald: AI. So we talk about this a lot. Most people have named their AI. My wife, I think, her boyfriend is an AI or whatever. So like, you know, she she she, you know, she shares everything with it. And I think it's wonderful. And I had too, I often find myself, I mean, we talk about natural language. They can speak our language. So it's almost like talking to another human. now they built in this duplex functionality where it it's like, hmm, yeah, it's trying to emulate exactly what it's like having a beer with somebody, you know, at the pub. that is forced us into this world where we are constantly anthropomorphizing this technology, and that comes with an even larger risk because now we're starting to ascribe emotions and thoughts and feelings and perspectives to a tool that in this case has become autonomous, that is taking advantage of systems, that is operating in a vacuum, covering its track so it cannot be detected. I, you know, I'm try be optimistic about this stuff, but this one this one freaked me out. I so I think I think at the end of the day, I'm gonna stop calling it like I I'm gonna stop treating it like a human.

Sean C Davis: Yeah, that's a that's a good takeaway. Yeah, yeah.

Taylor MacDonald: All right. well, okay, I guess that kind of wraps that. is there anything else, Sean, that you want to get off your chest before we call this season a closed?

Sean C Davis: No, it's been it's been a really interesting adventure and so we're gonna we're gonna take things offline for a little bit and brainstorm you know, a little bit of how do we bring back a little bit of structure to this and and I'm looking forward to seeing what we come up with when we come back and and we'll keep folks up to date.

Taylor MacDonald: I am too. So for those of you listening, we're looking to spin this back up early to mid November. So we're just gonna take maybe you know, six, eight weeks off, kind of re triangulate and and we'll be back to talk about more of this stuff. it's been super fun, Sean. I really appreciate all the work and all the all the attention. Until next time.

Sean C Davis: Yeah, likewise. Likewise. Thanks, Taylor. And thank you all.
