Why per-token AI pricing breaks at transaction volume
Per-token pricing is a good deal for experiments and a bad deal for alert queues. The reason is structural: risk work scales with every transaction you process.
Pay-per-use AI is attractive at the start. There is nothing to run, the first invoices are small and a pilot can begin in an afternoon. The trouble comes later, and in financial services it comes faster than in most industries.
Risk work scales with volume
Fraud alerts, sanctions hits and transaction-monitoring cases grow with payment volume. So do the tokens needed to triage them: the alert, the customer’s history, the relevant playbook and the draft disposition. Double the transactions and you roughly double the AI bill, month after month.
That makes per-token spend hard to forecast and awkward to defend. The more useful the AI becomes, the more it costs, and heavy users end up rationed.
Where the break-even point sits
Every AI workload has a break-even point. Below it, paying per token is cheaper. Above it, owning capacity is cheaper, because a GPU costs the same whether it answers a thousand requests or a million. For high-volume, repetitive work like alert triage, most institutions cross that line early.
At very low volumes, pay-per-use can be cheaper. The point is not to avoid it, but to know which side of the line each workload is on.
How to keep owned capacity lean
- Small models first. Most alerts can be handled by small, fast models. Large models are reserved for complex cases.
- Capacity that follows your day. Computing scales with business hours, batch windows and month-end peaks.
- Shared hardware. Several models and tasks share the same processors.
- Managed AI only where it pays. Public research tasks can still use managed models, per token, under a monthly cap.
Run the numbers on your own queues
The break-even calculator shows where your volumes sit. Cost is only one reason for private AI in banking, though: client confidentiality and banking secrecy often justify it before the numbers do.