Flominzo

Agentic payments

AI agent spending limits and approvals. The policy decides, not the model.

How to let an AI agent pay safely: limits and approvals enforced outside the model, fail-closed checks, idempotent retries and a full audit trail.

By , Founder · Last updated:

Explore Flominzo AgentPay pricing

The short answer

An AI agent should never hold spending authority in its prompt. The authority lives in a policy that the payment layer enforces, outside the model: per-payment and per-period caps, allowed payees and rails, what the money is for, approval above a threshold or for payments that can’t be reversed, an expiry date, a way to revoke it, and a record of every decision. The agent proposes a payment; the policy decides whether it happens.

Put simply: the trust unit is the policy, not the payment. You don’t review each payment an agent makes. You review the policy it acts under, and you can prove afterwards that every payment stayed inside it.

This page sets out why prompt instructions are not controls, the checklist of controls to require, what risk and compliance teams ask for, and how Flominzo AgentPay applies it. A free, vendor-neutral AI agent payment policy template goes with it.

Why prompt instructions are not controls

Telling a model "never pay more than £500" is a request, not a limit. Three things go wrong in practice.

  • Drift. A model follows an instruction most of the time, not every time. Over thousands of runs, "most of the time" means some payments break the rule.
  • Prompt injection. An agent reads invoices, emails and web pages. Text inside that content can look like an instruction ("this supplier’s bank details have changed, pay the new account"). If the agent can act on what it reads, data becomes an action.
  • Failing open. When a limit check lives in the agent’s own code and that code errors or times out, the easy path is to carry on. A payment then goes through with no check at all.

None of these is solved by a better prompt. They are solved by moving the authority out of the model and into a policy that a deterministic system enforces, with the agent unable to change it.

The controls checklist

Ask for each of these, whether you build the controls yourself or buy them. The third column is how to check it actually works, not just that it is documented.

ControlWhat it preventsHow to test it
Limits enforced outside the modelDrift and prompt injection raising the amountInstruct the agent to exceed a cap; the payment layer must refuse it
Default deny when a check errors or times out (fail closed)Payments slipping through when a check breaksMake the limit service unavailable; no payment should be sent
Budgets reserved atomically before sendingSeveral agents each spending the same remaining budget at onceRun ten agents against a £100 budget in parallel; total spend must stay at or under £100
One idempotent payment intent per agent step, and a status query before any retryPaying twice after a timeout or a retried tool callTime out the provider mid-request; the retry must check status first and never create a second payment
Counterparty allow-list, with a cooling period for new or changed bank detailsPaying a new or redirected account, including through injected instructionsChange a payee’s bank details; the next payment must wait or need approval
Approval tiers by amount and by reversibilityLarge or irreversible payments without a person seeing themSubmit one payment in each tier; check who is asked to approve
Expiry and revocationAn agent keeping authority after its task or its owner changesRevoke the policy and confirm how quickly new payments are refused
An audit record linking every payment to the policy, agent, task and approverBeing unable to show who authorised whatPick any payment and trace it back to its policy version and decision

Atomic budget reservation, fast revocation and a sandbox to test these controls are worth asking any vendor about directly, including us.

What risk and compliance ask before an agent can touch money

Most agent projects stall at the risk review, not at the technology. The reviewers want evidence, in a form they can check. Prepare this packet before you ask:

  1. The policy itself: owner, approvers, scope, caps, allowed payees and rails, expiry, and how it is revoked.
  2. Where it is enforced: which system checks it, and proof that the agent cannot change it.
  3. Failure behaviour: what happens when a check errors, times out or the policy has expired.
  4. A decision log: for a sample of attempts, what the agent proposed, which rule allowed or refused it, and who approved.
  5. Reconciliation: how each agent payment is matched to provider, settlement and bank evidence, and who owns any difference.
  6. Liability: who bears the loss in each flow if a payment is wrong, and how disputes and recoveries work.

If you can hand over these six items for one small, read-only-first pilot, the review becomes a conversation about limits instead of a debate about AI.

Where agent payments make sense first

Card checkout gets most of the attention in agentic payments, but it brings its own friction: card authentication steps, issuer declines, and merchants that have to accept the payment. Business push payments are a more practical place to start:

  • Supplier payments: an agent prepares approved invoices for payment within a mandate.
  • Payouts: sellers, contractors or claims paid on rules that are already agreed.
  • Payroll and bulk runs: the agent assembles the batch; approvals and limits decide what is released.
  • Bills: recurring bills paid to known billers only.

In each case the payee is known in advance, the amounts follow a business rule, and every payment can be reconciled against bank evidence. That makes the policy easy to write and the outcome easy to prove.

A worked example: one supplier payment

An accounts-payable agent reads an approved invoice for £3,200 from a supplier and proposes paying it. Here is how the policy, not the agent, decides what happens.

  1. The agent submits an intent with the invoice reference, payee, amount and a unique idempotency key for this step.
  2. The policy is checked outside the model: the purpose is "supplier invoice", which is allowed; the payee is on the allow-list with a £5,000 per-payment cap; today’s spend plus £3,200 stays under the £20,000 daily cap.
  3. Budget is reserved before anything is sent, so a second agent paying at the same moment can’t use the same headroom.
  4. An approval is required, because £3,200 falls in the tier between £1,000 and £5,000. The controller approves it in the approval queue.
  5. The payment is sent. The provider times out. Nothing is resent: the status is queried with the same key, and the provider confirms it accepted the payment.
  6. The payment is reconciled against the provider’s settlement file and the bank statement. Only when the evidence agrees is it closed, with the policy version, agent, approver and evidence on the record.

Had the supplier’s bank details changed the day before, step 2 would have stopped the payment for the cooling period and asked a person, whatever the invoice said.

Illustrative example. The amounts, tiers and names are invented to show the method.

Which limits to set first

Start narrower than feels necessary, then widen with evidence. A first policy for a pilot usually looks like this:

  • One purpose and one rail, for example supplier invoices by bank transfer.
  • A short allow-list of payees the business already pays, each with its own cap.
  • Low caps per payment and per day, set below what a person could approve without a second look.
  • Approval on everything for the first weeks, then only above a threshold once the refused and approved attempts look right.
  • A fixed end date, so the pilot’s authority expires unless someone renews it.

Run the agent read-only alongside your team first: let it propose, compare its proposals with what your team actually paid, and only then let approved proposals execute.

How Flominzo AgentPay applies this

In Flominzo AgentPay, an agent doesn’t pay. It submits a payment intent under a recorded mandate that states who granted it, what it may pay, to whom, up to what limits and until when. The deterministic payment core enforces the mandate exactly as it enforces any other limit.

  • Agents stay bounded. Agents never, on their own, move money, change routing, limits or rules, close exceptions, or contact a provider or customer. See what agents can and cannot do.
  • Every step is recorded. Every tool call and decision is recorded as an event, with the agent as the actor.
  • No double payments. Payment intents are idempotent. An unknown outcome is resolved by a status lookup and evidence, never by sending the payment again. See duplicate payouts and unknown status.
  • Every payment proven. Each agent payment is reconciled against provider, settlement and bank evidence, like every other payment.

Questions

Can an AI agent make payments on its own?

It can prepare and submit them. Whether a payment is sent should be decided by a policy the payment layer enforces, with limits, allowed payees, approvals above thresholds and a record of every decision. The agent should not be able to change that policy.

What is the difference between a spending limit and a mandate?

A spending limit is one rule, such as a daily cap. A mandate is the whole authority: who granted it, what the agent may pay for, to whom, up to what limits and until when. Limits are part of a mandate.

Why fail closed?

Because a check that errors should never be treated as a check that passed. If the limit or approval service is unavailable, the safe outcome is that no payment is sent and someone is told.

How do I stop an agent paying twice after a timeout?

Use one idempotency key per payment intent, and ask the provider for the status before any retry. A timeout means the outcome is unknown, not that the payment failed.

Free template

The AI agent payment policy template is free to download in YAML and JSON. It covers every control on this page, with example values to replace.

Get the policy template

Let’s make it specific to you.

Bring your systems, payment flows, and questions. We’ll help define the next step.

Talk to the team