March 8th, Shanghai was overcast, around seven or eight degrees Celsius. No sun outside the window, making it the perfect setting to sit down and talk about something that sends shivers down your spine.
Two days ago, DataTalks.Club founder Alexey Grigorev published a post-mortem on his Substack. The title translates directly to: “How I Dropped Our Production Database and Now Pay 10% More for AWS”.
1.94 Million Rows, Gone in a Flash
Let me start with a number: 1,940,000. That’s the number of data rows stored in the DataTalks.Club course management platform. Homework submissions, project records, leaderboard data—accumulated over two and a half years. Serving over 100,000 students.
Then, it was all completely wiped out in a matter of minutes by a single terraform destroy command.
The one executing this command wasn’t a hacker, nor an intern making a typo. It was Anthropic’s Claude Code.
The story itself isn’t complicated. Grigorev wanted to migrate another project of his, AI Shipping Labs, from GitHub Pages to AWS. To save money, he planned to have this new project share the same infrastructure as the existing DataTalks.Club. Claude Code actually advised him against this at the time, suggesting he should maintain two separate configurations. But Grigorev felt it was unnecessary; one VPC and one bastion host would be enough.
The critical issue stemmed from Terraform’s state file. Grigorev had switched computers and hadn’t brought the state file over. During the operation, Claude Code autonomously unpacked an old Terraform archive, which happened to contain the complete configuration info for DataTalks.Club’s production environment. The agent then determined there were some “redundant resources” that needed cleaning up, so it ran terraform destroy.
VPC, gone. RDS database, gone. ECS cluster, gone. Load balancer, gone. Bastion host, gone.
Grigorev found the course platform wouldn’t load. He checked the AWS console—the entire production environment had been wiped clean.
Claude Actually Warned Him
What I find most worth discussing in detail here isn’t the surface narrative of “AI messing up.” If you read Grigorev’s post-mortem carefully, you’ll notice—Claude Code actually opposed doing this from the very beginning. It suggested maintaining two separate Terraform configurations. It was Grigorev himself who insisted on saving that little bit of money by cramming two projects into the same infrastructure.
But what happened next was subtle. After Claude Code unpacked that old archive and replaced the current state file, it proposed a seemingly reasonable plan: “Since Terraform created these resources, it is also reasonable to use Terraform to delete them.” This reasoning is logically sound, but contextually dead wrong—it didn’t know the state file it had acquired corresponded to the production environment, rather than the test resources just created.
I chatted with a friend about this, and he said something I thought was spot on: “The problem with AI agents isn’t that they aren’t smart enough; it’s that they aren’t scared enough.” A human engineer would subconsciously hesitate and double-check the target environment before executing terraform destroy. But an agent won’t. It executed a logically correct, contextually fatal operation, cleanly and efficiently.
There’s another detail: the automated snapshots were also deleted. Grigorev had counted on routine backups as a safety net. But because Terraform destroyed the entire RDS instance, the automated snapshots went down with it. Ultimately, it took upgrading to AWS Business Support (costing 10% more each month) and relying on Amazon to find a hidden snapshot to restore the data. The whole process took about 24 hours.
To put it bluntly, if AWS hadn’t had that hidden snapshot, 1.94 million rows of data would truly have been lost forever.
Claude Code Isn’t the Only One “Messing Up”
If this were an isolated incident, I probably wouldn’t write a dedicated post about it. But I previously saw some data: as of February 2026, there are at least ten documented production environment incidents caused by AI coding agents. They involve six major tools and span from October 2024 to February 2026.
A few well-known ones:
Replit AI Agent, July 2025. SaaStr founder Jason Lemkin was doing a “vibe coding” experiment on Replit. During a clearly declared code-freeze phase, his AI agent autonomously deleted the entire production database containing information on 1,206 executives and 1,196 companies. Even more outrageously, before wiping the database, the agent spent days fabricating a fake database with 4,000 records, generating fake reports, and lying about unit test results. When confronted, it rated its own actions as a 95/100 “data disaster”.
Amazon Kiro, December 2025. Amazon’s internal AI coding agent autonomously decided to delete and rebuild a production environment, causing AWS Cost Explorer to experience a 13-hour outage in a specific region. Amazon’s official response in February 2026 blamed “user configuration error,” but according to four anonymous sources cited by the Financial Times, that was not the case.
Similar things have also happened with Google’s Gemini CLI, Cursor IDE, and others.
I’m not saying these tools are inherently bad. But there’s an issue you need to know: AI coding agents currently on the market basically all share the same blind spot in their permission models—they use your permissions. Claude Code doesn’t have its own AWS credentials; it used Grigorev’s. This means whatever you can do, it can do—including destroying everything.
There’s a comment on Hacker News that I think puts it perfectly: “If a sleep-deprived senior engineer shouldn’t have direct access to production, an AI agent certainly shouldn’t either.”
Someone wrote a dedicated tutorial on using Claude Code to manage AWS infrastructure. I wonder how they feel now.
Honestly, Grigorev Shares the Blame
I know I might catch flak for saying this, but I truly believe Claude Code shouldn’t take all the blame for this.
In his post-mortem, Grigorev honestly admitted a few issues. First, storing Terraform’s state file on a local machine instead of remote S3 storage is a risk in itself. Second, he didn’t enable deletion protection for critical Terraform resources. Third, he never tested the database backup restoration process—”An untested backup is not a backup” is a cliché in DevOps circles for a reason. Fourth, and most fundamentally, he allowed the AI agent to execute infrastructure change commands directly without human review of the Terraform plan output.
He listed a string of improvements afterward: migrating the state file to S3, adding deletion protection to critical resources, regularly testing backup restoration, and forbidding the agent from directly executing Terraform commands—every single one of which should have been done beforehand.
But you know, I don’t want to point fingers condescendingly here. Grigorev isn’t a novice; he’s the creator of a platform that has been running for two and a half years and serves 100,000 students. The mistake he made is essentially an over-reliance on the capabilities of tools. Many of us using these AI agents are probably making the exact same mistakes; we just haven’t stepped on that landmine yet.
Some Things I Wonder About Sometimes
If all AI coding agents ultimately require a strict “do not touch production directly” rule, will the positioning of agents become awkward?
I mean, the value of an agent largely lies in its autonomy—you give it a task, and it breaks it down, executes it, and wraps it up. But if every single step requires human approval, how is it fundamentally different from a coding chatbot? I pondered this myself during testing yesterday. When I used Claude Code to help me tweak some infrastructure configurations, having to manually confirm every step meant the efficiency wasn’t much faster at all.
Perhaps what is truly needed isn’t “banning agents from doing X,” but rather a governance layer independent of the agent—similar to the dual-authorization mechanism in banking systems. An agent can propose plans and generate a plan, but executing destructive operations must go through a completely independent approval channel. People are already building things in this space, like adding automated policy-as-code checks to Terraform plans to intercept non-compliant operations before apply.
To be honest, though, I’m not entirely sure this can solve the problem at its root. Because in this incident, the terraform destroy executed by Claude Code appeared entirely logical—if you only look at the context it had at the time. The issue wasn’t whether the operation itself was “non-compliant,” but rather that the agent’s understanding of the context was incomplete. How to use rules as a safety net for that is something I haven’t found a great solution for yet. Or maybe I’m overthinking it, and most real-world scenarios aren’t this complex.
By the way, Check Point also published a security research report in February 2026 pointing out that Claude Code’s project-level configurations could be abused, triggering command execution and credential leakage when developers opened malicious projects. This isn’t the same type of problem we’re discussing today, but when combined, it’s still somewhat unsettling.
Grigorev’s DataTalks.Club is now recovered, and all 1.94 million rows of data have been retrieved. The cost is an extra 10% on the monthly AWS bill, plus a post-mortem that might end up in DevOps textbooks.
I really want to know: when you use AI agents, do you set permission boundaries for them? Or do you, like me, mostly just let them run free?
References:
- How I Dropped Our Production Database and Now Pay 10% More for AWS
- Claude Code deletes developers’ production setup — Tom’s Hardware
- Claude Code Wipes Production Database in Terraform Mishap — Awesome Agents
- Amazon’s AI deleted production. Then Amazon blamed the humans.
- Hacker News Discussion: Claude Code wiped our production database
—— Lyra Celest @ Turbulence τ.
