Governance is the deployment decisions, not a policy document

When a business talks about "AI governance," it's easy to picture a compliance binder. The part that actually determines whether a Claude deployment is safe to run is much more concrete: what permissions the agent has, which of its actions require a human to sign off first, and how you know - on an ongoing basis, not just at launch - whether it's still meeting the bar you set for it. Those three things are design decisions made while building the deployment, not paperwork added after.

Permission scoping: narrower than it feels like it needs to be

The instinct when connecting an agent to internal systems is to grant it whatever access makes the integration easiest. The better default is the opposite: scope every connection - every MCP server, every custom tool, every data source - to the minimum the agent's actual job requires, and scope it per user where the underlying system supports it, so an agent acting on behalf of a specific employee can only see and touch what that employee could already access themselves. An internal knowledge assistant connected to a wiki, a CRM, and Slack doesn't need write access to any of them if its job is answering questions.

A useful test: if this agent's credentials leaked, or it made a serious reasoning error, what's the actual blast radius given its current permissions? If the honest answer is "it could do real damage," the scoping is too broad before you've even reached the question of whether the model itself was reliable.

Approval gates: for the actions you can't easily undo

An approval gate is a checkpoint where the agent must stop and get explicit human sign-off before proceeding, reserved for actions that are costly or difficult to reverse - sending an external email, modifying a financial record, deleting data, publishing something publicly. Gate those specifically. Gating everything, including low-stakes reversible steps, defeats the purpose of using an agent in the first place and trains people to rubber-stamp approvals without really checking them, which is worse than no gate at all.

Gate this

Sending a message to a customer, modifying a database record, executing a payment, deleting a file, publishing content - anything hard to undo once it happens.

Don't gate this

Reading a document, drafting a response for review, searching internal data, running a calculation - reversible, low-consequence steps that don't need a human in the loop every time.

Outcomes: define what "working" means before it runs unsupervised

The third piece is measurement. Before an agent runs on new, real cases without a human checking every output, it needs a defined outcome or rubric - a concrete description of what a successful run looks like - checked against historical cases where you already know the right answer. This isn't a one-time sign-off. Agents deployed into a live process should be monitored on an ongoing basis, the same way any other production system is, because the inputs it sees in production will eventually diverge from what it was validated against.

Building a governance layer alongside an agent deployment, not after it's already live? Talk to us about scoping it properly from the start.

How the three pieces work together

ControlQuestion it answersSet at
Permission scopingWhat can this agent actually touch?Integration/deployment time
Approval gatesWhich of its actions need a human first?Workflow design time
Outcomes/monitoringIs it still meeting the bar, in production, over time?Ongoing, not one-time

Key takeaways

  • Governance is a set of deployment decisions - permission scope, approval gates, outcome measurement - not a document written after the agent is already live.
  • Scope every tool and data connection to the minimum the agent's job requires, per user where the underlying system supports it - test it against "what's the blast radius if this goes wrong."
  • Reserve approval gates for costly or hard-to-reverse actions specifically; gating everything trains people to rubber-stamp, which defeats the point.
  • Define a concrete outcome or rubric, validate against historical cases before unsupervised use, and keep monitoring afterward - production inputs drift from what was originally validated.