Get Into Agentic Development From the Shallow End

If you’re anything like me, enabling that “Approve for me” setting in your agentic coding tool of choice feels like a big step. Waiting for approval prompts isn’t the best use of my time, but I’m still skeptical enough that I don’t trust the agent to decide entirely for itself when it needs my input.

ChatGPT action approval setting selector

Dipping My Toes In

Reading Rob Bell’s post about the risks of permission fatigue and his approach to avoid it in Claude Code helped me realize I didn’t have to dive in all at once. His deny/approve/assess model offered a more gradual path: block operations I never want an agent performing, allow routine operations I understand, and continue to review the rest.

I implemented a similar model for Codex using its command rules, sandbox configuration, and a guard hook. Allowing routine local Git operations and project-scoped file editing noticeably reduced the prompts in my normal workflow. Blocking recursive deletion and access to .env files, while requiring review before pushes gave me more confidence that the agent wouldn’t do serious damage while operating more independently.

Once I had accounted for those obvious extremes, I was left with another question: how should I maintain and expand the system as my workflow and my confidence in using agents for development changed?

At the time, I was working heavily on a mobile application and doing a lot of automated and manual testing with device simulators. Maestro tests and simulator interactions became the next obvious batch of operations to allow. But I suspected future patterns wouldn’t always be so easy to spot.

Turning Prompts Into Data

What do engineers do when faced with a problem? Collect some data and create a system to solve it!

With help from Codex, I added an observability layer to the permission system. It records each approval request, groups related operations, and tracks whether the requested action was subsequently executed. Instead of guessing when prompts felt excessive, I could start making those decisions from evidence collected through my work.

The first version used Codex hooks to record each permission request and whether the requested operation was subsequently executed. Because the hooks could not directly observe whether I clicked “deny,” I conservatively grouped every other resolved request as “not executed or denied.” I stored the results in SQLite and had a command that generated a report ranked by frequency.

After my first few rounds of running this report I noticed a limitation of simply sorting through the recorded prompts. Effectively identical operations appeared separately because their worktree names, pull requests, or branches differed. Adding a rule for every exact command would replace permission fatigue with allowlist-maintenance fatigue.

Grouping Prompts by Intent

For the next iteration on the system, I had the report group the requests by intent rather than the full command text, so numerous filename-specific rows became one useful category.

One run of the report categorized 248 prompts and surfaced several useful patterns:

Total prompts: 248 Approved and executed: 236 Not executed or denied: 12 count approved not-run rate candidate grouped-prompt-family 98 98 0 100% yes apply_patch /private/tmp/<job>/… 36 29 7 81% no node with runtime assignments 25 23 2 92% no git mutating operation 19 19 0 100% no GitHub API calls 4 3 1 75% yes GitHub CLI read-only operations

Turning Patterns Into Proposals

I then added a step to the report to highlight the top candidates for broader permission, favoring frequent, consistently approved, lower-risk operations. As you can see from the result above, even though I had approved every recorded GitHub API request, a broad gh api rule could also permit mutations, so the report held it back.

Top 2 human-review candidates These are suggestions only; this report did not change your configuration. 1. temporary validation-file writes — 98/98 approved (100%); current coverage: not determined Evidence group: apply_patch /private/tmp//… (source or test) Caution: This grants workspace-style writes anywhere below /private/tmp. # Add this entry inside the existing [sandbox_workspace_write] section. writable_roots = ["/private/tmp"] 2. read-only GitHub CLI inspection — 3/4 approved (75%); current sample command decision: allow Evidence warning: 1 matching prompt(s) were denied or did not execute; investigate them before adoption. Evidence group: GitHub CLI read-only operations Evidence note: compound prompt; the proposal covers only the safe leading command Caution: This deliberately excludes gh api because later flags can change its HTTP method. [generated rule omitted]

The suggestions exposed different risks. Allowing writes throughout /private/tmp would reduce worktree prompts but permit changes in a directory shared by many applications, so I could instead place my worktrees beneath a narrower dedicated directory. A representative gh pr view command, meanwhile, was already allowed. The report couldn’t explain every historical prompt, but it kept me from blindly adding a duplicate rule.

Safely Wading In

Running this report does not immediately adopt anything. Instead, the system gives me a feedback loop that allows me to iterate on the level of autonomy and feels more grounded than going on vibes.

  1. Work normally and continue reviewing unfamiliar operations.
  2. Collect the resulting approval history.
  3. Group requests by the capability they represent.
  4. Identify frequently approved, low-risk candidates, and inspect the prompts behind them.
  5. Review and test the proposed rule.
  6. Add it to my configuration only when I’m comfortable with it.

The result isn’t a completely autonomous agent, and that isn’t really my goal.

Novel, destructive, ambiguous operations, or anything involving credentials still need my attention. The routine work I have repeatedly reviewed can gradually move out of the way.

Thinking about autonomy as a feedback loop instead of a switch has made agentic development feel much more approachable to me. I don’t have to dive in the deep end by deciding upfront how much I trust the agent in every possible situation. I can wade in from the shallow end, observe how the agent behaves, and grant it more freedom as I build confidence.

Conversation

Join the conversation

Your email address will not be published. Required fields are marked *