Architecture · · 5 min read

Hybrid AI in financial services: what goes private, what goes managed

Banks don’t have to choose between sovereignty and the most capable managed models. A hybrid architecture sends each task to the model that fits it, and the rule for deciding is simpler than most teams expect.

Most conversations about AI in banking start with a false choice. One camp wants everything on private infrastructure, because client data and banking secrecy leave no room for doubt. The other wants the most capable managed models, because the quality gap is real. Both are right about their half of the work.

The rule: follow the data, not the department

The question that decides where a task runs is not “which team is asking?” but “what will the model see?”. If a request touches client, account or transaction data, or your proprietary risk models, it runs on private models in your own cloud. If it only touches public material, it can use a managed model such as Claude, GPT or Grok.

That gives a clear split for most financial workloads:

  • Private: fraud alert triage, AML case summaries and SAR/STR drafts, KYC and beneficial-ownership review, credit memos, and explanations of your risk scores.
  • Managed: tracking new regulation and supervisory guidance, summarising published market research, and drafting from public templates.

When in doubt, it stays private

Real requests are messy. An analyst might paste a customer name into a question about new guidance. That is why routing is done by a policy check on every request, before any model sees it, rather than by trusting users to pick the right tool. Anything involving client data, or anything the check is unsure about, goes to the private side.

The safe default is private. Managed AI is something you switch on for specific task types, not something you have to remember to avoid.

Why not just pick one side?

Private-only means your compliance team reads new regulation without the strongest models. Managed-only means client data leaves your environment and every screened transaction becomes a per-token cost. Hybrid gives each task the right model: confidentiality where it matters, frontier capability where it helps, and the lower cost for each workload.

What to ask any vendor

Ask where the policy check runs, what happens when it is unsure, whether you can see every routing decision in an audit log, and whether you can switch managed AI off entirely. If the answers are vague, the architecture probably is too.