Home / Insights / AI agents
AI agents

Guardrails first: building AI agents that cannot move money.

Two production agents for an online gaming operator, and why we designed what each agent cannot do before what it can.

ClaudeAWSAmazon Aurora

An AI agent that can do anything is easy to demo and hard to trust. In a business where conversations involve money, identity and disputes, the useful question is not what the agent can do. It is what happens on the agent's worst day.

We built two agents for an online gaming operator: one that answers players in support, and one that reviews withdrawal requests for the risk team. Both started from the same rule: design the limits first.

Start from the worst case

For each agent we wrote down the worst realistic outcome before writing a prompt:

  • The support agent promises a bonus that does not exist, or a player talks it into revealing something it should not.
  • The withdrawal reviewer approves a fraudulent payout, or a player hides instructions in their own profile that steer the review.

Then we removed the ability to cause those outcomes, rather than asking the model nicely not to.

The support agent has no tools at all

The support runtime has shell, file, web and database access switched off. It cannot look up an account, change a balance or call an API. It answers only from an FAQ that the support team writes and owns, and when nothing in that FAQ matches, the ticket goes to a person instead of the model improvising.

That sounds limiting, and it is. It is also why the worst case is a handed-off ticket rather than a wrong payout. A prompt injection that succeeds still has nothing to call.

Two smaller decisions mattered too. Support runs on fast, mid-size models for predictable latency and cost, with the largest model blocked for this role. And the instructions include explicit defences against prompt injection, so player text is treated as a question, never as an instruction.

The reviewer advises; a deterministic engine decides

The withdrawal reviewer reads a player's history, forms a view, and explains its reasoning. It never approves anything. The operator's payout engine, a chain of deterministic checks that fails closed, remains the only component that can move money.

The reviewer's power is limited in three more ways:

  • It reads production data through a database user that can only run SELECT.
  • Player-supplied fields are wrapped and labelled as untrusted before they reach the model.
  • Its verdict goes back to the platform's review endpoint and into the analysts' chat as a recommendation with reasons, not as an action.

Measure before you trust

Before the reviewer saw live traffic, the rule-based scorer was replayed in shadow mode on 3,479 historical withdrawal requests, and the team hand-labelled 999 requests to measure how often its flags were right.

The point of that set is not a headline number. It is that every later change to a prompt, a rule or a model can be measured against the same requests before it reaches production. Without a fixed evaluation set, "the new prompt feels better" is the only test you have.

Patterns that transfer

  • Capability, not instructions. Remove a tool instead of telling the model not to use it.
  • Advisory by default. Let the agent recommend; let deterministic systems act on money and records.
  • Read-only by construction. Enforce it with database users and IAM, not with prompts.
  • Untrusted data stays data. Label everything a customer can type before it reaches a model.
  • A human path for everything else. "I will pass this to a colleague" is a feature.
  • A fixed evaluation set. Built before launch, reused for every change.

None of these make an agent less useful. They make it something you can leave running on a Friday evening.

What’s next for
your business?

Let’s talk