Post

Towards Autonomous Development

What I Learned Letting an Agent Build an App on Its Own

I wanted a first-hand look at the unsafe future of building with autonomous agents, so I let Instinct build an entire app for me. Instinct is an autonomous agent platform: it gives the agent its own sandbox, email inbox, memory and search tools, and lets it work with far less hand-holding than a coding assistant.

More and more software is being built by citizen developers - people who describe what they want and let AI write the code. I usually stay involved in architectural and security-related decisions for an app, but this time I played the citizen developer: I let the agent make its own decisions and restricted myself to product and functionality comments.

The app: “Yesh-Makot” (Hebrew slang for “there’s a fight nearby”) - a joke web app that lets you report bar fights in your region while others subscribe to watch from a safe distance. Along the way it turned into a live incident feed.

The reason I say “unsafe” is that I gave it a higher level of autonomy than I usually would. This comes with a high reward - tasks get done more effortlessly - but also with a higher risk: occasionally the agent strays sideways, and the more autonomy it has, the further it goes in the wrong direction before it’s corrected.

I was considering giving it a pre-paid credit card loaded with a small amount to use for its needs, and decided I’d do it when the need arose. I was lucky. At some point I asked Instinct to add an event feed from some Telegram channels it would subscribe to. It set up an account with a phone-number provider and came back asking me to top it up with $5 - through the payment option labeled “credit card nigeria”, or crypto if my card didn’t clear. If I had given it a card, it probably would have made the purchase.

Pause. Rethink. What other risks am I missing? Let’s mitigate them.

Principle 1: Own the agent’s data and identity

I have no idea what the agent does in the background, and even if I had access I would probably not spend my time reviewing it. That gives me two reasons to keep it at arm’s length. The first is blast radius: I don’t want it to have broad access to my data. The second is ownership: Instinct gives the agent its own email box, but by default it works under your accounts everywhere else. What happens if the agent needs to be shut down, or put behind a paywall? Everything it built and everything it knows goes with it.

The harness has memory and search tools baked into it. They’re good to use, but all inputs to the harness should go through a layer you control, and all output artifacts should be stored in a place you own. This means:

  • A dedicated email box with the agent’s identity - don’t let it borrow yours.
  • A dedicated GitHub org to store all code artifacts and the CI/CD pipelines that ship them.
  • Your own cloud runtime, where you define what the agent can access.

Owning the pipeline paid off in a second way: visibility. I can’t see inside Instinct’s sandbox, but every change it made had to land as a commit in my GitHub org and ship through my GitHub Actions workflows to my own cloud project. That gave me a trail of what an otherwise unsupervised agent was actually doing - commits, CI runs, deploys, runtime logs. GitHub also became the channel for steering it: the reviewing agents (more on that below) reported their findings to Instinct as issues and pull requests, so every correction left a record I could follow. I can’t see inside the agent, but I can see everything it ships.

Principle 2: Limit what the agent can break, and what it can leak

I’ve seen how this agent strays into gray territory. How do I catch it early enough, before any serious damage is done? What happens if a malicious email tries to weaponize the agent? The agent itself is now part of the attack surface, and it should be part of the threat model.

Start with what it can break. The agent needs access to GitHub in order to push code, report and fix issues, check CI pipeline failures and so on, but does it really need to be an admin on the dedicated org? It needs access to runtime logs in order to fix bugs you report, but does it need permission to remove your user and lock you out of your cloud provider? Give it exactly what the task needs and nothing more.

Then there’s what it can leak. Simon Willison’s lethal trifecta is the right lens here: an agent with access to private data, exposure to untrusted content, and the ability to communicate externally can be tricked into sending your data to an attacker. Any two are manageable; all three together are the problem. A fully autonomous builder has all three by default - it reads your code and logs, it reads an inbox anyone can email, and it can push code, send email and sign up for services.

Every limit I put on the agent costs some autonomy, so the goal is to break the trifecta for each task rather than lock everything down. An agent fixing a bug needs production logs, code and the issue tracker, but not the inbox that anyone in the world can push a malicious prompt into. One providing support reads untrusted email by definition, so it gets documentation, not your password vault - and it should be scoped so it can’t leak one customer’s emails to another.

Principle 3: Have specialized models and harnesses review each other

The agent is quite confident, tests pass, the build is green - but what happens if we forgot something embarrassing, like adding authentication to a sensitive route? I’m not reviewing the code here by design, so how do I catch those in time?

The code in my app was written by Instinct. I asked Claude Code and Codex to review it with a focus on security and architecture. They found and fixed a surprisingly large number of security bugs - no rate limiting anywhere, a stored XSS, moderation that left removed reports publicly reachable - that would have been easily caught and prevented in the old days, when humans wrote and reviewed code. I’m sure some remain.

What the future may look like

Time to market drives a lot of business decisions, and giving agents more autonomy shortens it significantly. I don’t see companies giving up that edge in the near future just because it carries risk.

This means we need to find ways to manage and reduce the risk. While I don’t have the complete recipe for autonomous development, I believe some pieces of the puzzle can already be assembled. Here is what they look like:

  • AI profiles defined by the ops team: Builder, Bug Hunter, Data Analyst, Customer Support. Each profile has its own role and permission policies, and no profile holds the full lethal trifecta: external-facing roles are separated from internal ones, and anything that reads external data is isolated from the rest of the system.
  • Identity hygiene: each agent has its own identity, tied to a profile and an instance ID, and never borrows developer identities on any platform. A nice side effect is honest attribution: when a developer replies to a reported issue, it’s in their own voice, not an agent’s.
  • An agent control plane: lets human staff launch high-autonomy agents to build features, fix bugs, or analyze data
    • and keeps them in check. Every run is logged, spending is capped, actions like payments or production deploys wait for human approval, and any agent can be stopped instantly.

Where a human still had to step in

This approach reduces some of the risks, but some mistakes still took a human to catch and correct. When I asked the agent to build things, it sometimes opted for the simplest way, which is not necessarily the best security practice, and sometimes it read a request more broadly than I meant it. I caught these because I know what to look for. A citizen developer wouldn’t have.

A few examples from my project:

  • The way it chooses to manage secrets and credentials is not always the best security practice. My correction was required when it tried to store static credentials in CI secrets instead of using OIDC to deploy the app to GCP from GitHub Actions.
  • Using a dev container wasn’t the agent’s default choice when the repo was initially set up. This exposes the agent sandbox to supply chain attacks, which more than doubled in monthly volume during 2025. Ideally the sandbox would have no credentials in it at all, but I have no control over how Instinct implemented its sandbox, so I can’t really tell.
  • Bar fights turned out to be rare, so I asked Instinct to pull in public data sources. It chose official Magen David Adom (MDA) reports and other public incident channels, and my joke app quietly became a live incident feed. I asked for more data; I didn’t ask for real emergencies on the map. When a fatal incident showed up, I had it write a content policy that filters out deaths and serious injuries. That fixed the feed, but I don’t have a recipe for preventing the next overreach like it - only for cleaning it up afterwards.

Closing thoughts

Autonomous development is coming whether we’re ready or not - the speed is too valuable to leave on the table. The agent built a working app with very little input from me, and it also asked me to make a payment that could have landed me on a fraud watchlist. Both are true, and both will keep being true. The job is no longer reviewing every line of code the agent writes; it’s designing the environment it works in: its identity, its permissions, and who checks its work.

This post is licensed under CC BY 4.0 by the author.