AWS Builder Center
Why banks should train their own foundation models (Part 1 of 3)

Why banks should train their own foundation models (Part 1 of 3)

Learn why and how banks should train foundation models on proprietary event data — with architecture patterns, business impact, and regulatory guidance for UK banking.

Data and Ai lead at AWS
Learning level: Advanced (400)
Author: Faris Haddad
Are you managing hundreds of machine learning models across your bank? A foundation model trained on your own banking data could replace them all — delivering better performance, unified governance, and faster time-to-value. Yet most banks continue building task-specific models from scratch, repeating feature engineering for every new use case and wondering why each model takes six months to reach production.
In this post — the first in a three-part series — I explain why banks should train their own foundation models on proprietary banking event data. I cover the architecture at a conceptual level, the business impact, and the regulatory considerations. In Part 2 , I show you how to build the data foundations. In Part 3 ,  I walk through the model architecture, training, and production deployment with code examples.
I have also added a supplemental part 4  to show you how to use Amazon Nova Forge   to build a banking-specific language model that works in tandem with the custom foundation model from Parts 1–3. By the end, you'll have a hybrid architecture where structured embeddings from your custom FM feed into a domain-adapted Nova model — giving you both quantitative precision and language fluency in a single system.

The problem: model fragmentation in banking

Every bank generates billions of events: card transactions, digital engagement, communications, product interactions, and risk signals. Today, these events feed into siloed, task-specific models. Each model learns its own representation of customer behavior from scratch, with no shared understanding across tasks.
I've seen this pattern repeatedly in large financial institutions. A fraud team builds a transaction model. A credit team builds another. A marketing team builds a third. All three need to understand the same customer behavior — but they learn independently, from different feature sets, using different architectures. The result? Three models that each capture a fraction of what a single model could learn from the full picture.
This fragmentation creates three problems:
  • Redundant effort — Feature engineering is repeated for every new use case. A fraud model and a credit model both need to understand transaction patterns, but they learn independently.
  • Inconsistent governance — Each model requires its own validation framework, fairness assessment, and regulatory documentation. For a Tier 1 bank, this means hundreds of separate model risk reviews.
  • Diminishing returns — Task-specific models trained on limited labeled data plateau quickly. They can't use the billions of unlabeled events sitting in your data lake.
Foundation models solve this by learning general-purpose representations from massive unlabeled data, then adapting to specific downstream tasks with minimal additional training. One model. One governance framework. Eight or more downstream tasks.

Why this works: the transfer learning opportunity

The insight behind foundation models is simple: pre-train once on your full event corpus, then fine-tune cheaply for each downstream task. This is the same principle that made BERT transformative for natural language processing (Devlin et al., 2019) and that drives modern computer vision systems.
For banking, the opportunity is compelling. Revolut's PRAGMA model (2026) demonstrates what's possible. Pre-trained on 24 billion banking events from 26 million user records, PRAGMA outperformed task-specific baselines across six downstream tasks:
  • Credit scoring: +130.2% Precision-Recall Area Under Curve (PR-AUC), +12.4% Receiver Operating Characteristic Area Under Curve (ROC-AUC)
  • Fraud detection: +16.7% precision, +64.7% recall
  • Communication engagement: +79.4% PR-AUC, +20.4% ROC-AUC
  • Product recommendation: +40.5% mean Average Precision (mAP)
  • Recurrent transactions: +5.8% F1
  • Lifetime value: +1.8% PR-AUC, +2.6% ROC-AUC
Critically, Low-Rank Adaptation (LoRA) fine-tuning (Hu et al., 2022) matched or exceeded full training from scratch while updating only 2–4% of parameters. This means you can adapt a single backbone to eight or more tasks without the cost of training eight separate models.

Why now?

Until PRAGMA's publication in April 2026, there was no public evidence that foundation models worked for banking data at scale. The research existed for text (BERT, 2019) and images (ViT, 2020), but banking event sequences were unproven territory. That's no longer the case.
The playbook now exists. The first banks to execute it will have a 2–3 year head start on their competitors — not just in model performance, but in the organizational capability to iterate quickly on new use cases. When your competitor launches a new product recommendation engine in three weeks using LoRA fine-tuning, and yours takes six months of feature engineering, the gap compounds fast.

The industry is already moving

PRAGMA isn't an isolated research result — it's the most visible signal of a broader industry shift that was already underway. By mid-2026, at least six major financial firms have published or publicly demonstrated transaction foundation models in production:
  • Stripe built a payments foundation model pre-trained on its global transaction network, using it to improve fraud detection and merchant risk scoring across millions of merchants who lack sufficient individual transaction history for traditional per-merchant models.
  • Nubank developed NuFormer  (2025), a transformer pre-trained on the Brazilian neobank's 100M+ customer transaction sequences. NuFormer powers credit limit decisions, fraud scoring, and personalised product recommendations from a single backbone — the model Nubank credits with compressing a 6-month model development cycle down to weeks for new use cases.
  • Visa published TransactionGPT (2025), pre-trained on anonymized card network data across billions of cardholders. Visa reports significant improvements in authorization accuracy and real-time fraud decisioning at network scale.
  • HSBC developed FraudTransformer (2025) specifically for AML and transaction fraud, pre-trained on internal customer event sequences. HSBC noted that the shared representation significantly reduced the false positive rate for cross-border transaction alerts — a persistent pain point for correspondent banking teams.
  • Trustly published an Open Banking Foundation Model (2025) pre-trained on open banking payment sequences across European markets, demonstrating transfer learning across different national banking systems within a single backbone.
  • Featurespace released NPPR (2024), a transaction sequence model deployed with several UK and European banks, showing consistent double-digit fraud recall improvements over gradient-boosted tree baselines without any task-specific feature engineering.
NVIDIA's independent benchmark validation (June 2026) reported a near-50% lift in Average Precision on the IBM TabFormer fraud dataset using a transaction foundation model approach over a strong XGBoost baseline — confirming that the gains reported by individual firms are reproducible on standardised public benchmarks, not just internal data.
This is no longer a research phenomenon. It's a production pattern. For a traditional bank still running gradient-boosted trees with hand-crafted feature pipelines, the competitive gap is already opening.

Why not just use a general-purpose large language model?

This is the first question most technology leaders ask. If GPT-4 or Claude can process text, can't they process banking events too?
The answer is no — and understanding why is key to the architecture decisions in Parts 2 and 3.
General-purpose large language models (LLMs) are trained on text: words, sentences, documents. Banking event data is structurally different in three ways that make text-trained models fundamentally unsuited:
  1. Tabular structure, not narrative. A transaction record is a set of key-value pairs — merchant, amount, channel, MCC code, timestamp. Serialising this as text ("On 14 March at 09:32, a payment of £47.50 was made to Tesco...") works for a human reader but destroys the structural information the model needs. The relationships between fields within an event — amount × merchant × channel × time-of-day — are what carry predictive signal. A text LLM treats these as sequential tokens; a properly designed banking model treats them as structured co-occurrences.
  2. Continuous time, not discrete tokens. Banking events happen irregularly in time. The gap between a customer's last transaction and an unusual one is itself a signal — a £500 cash withdrawal at 2am after 30 days of inactivity means something very different from the same transaction the day after payday. Text LLMs use positional embeddings designed for fixed token sequences. Banking event models require temporal encodings (log-seconds, calendar features) that capture continuous time explicitly.
  3. No transferable vocabulary. A text LLM's vocabulary — 50,000–100,000 subword tokens — has zero overlap with banking semantics. Merchant Category Codes, transaction channels, risk flags, and product types don't exist in any text corpus the model was trained on. Fine-tuning a general LLM on banking data means teaching it an entirely new language from scratch, while dragging along billions of parameters that have nothing to contribute. Training your own model on your own data, with a domain-specific tokenizer, is both cheaper and more accurate.
Amazon Bedrock is an excellent choice for text-heavy banking use cases: document summarisation, regulatory filing analysis, customer communication generation. But for the core quantitative banking tasks — fraud, credit scoring, AML, lifetime value — the data is structured, temporal, and proprietary. That's exactly the domain where training your own foundation model delivers the step-change improvement the PRAGMA results demonstrate.

The architecture: how to build a bank foundation model

So how do you actually build one of these? A bank foundation model requires four key components. I describe each at a conceptual level here, with full implementation details in Part 2  and Part 3 .

Multi-source event tokenization

Banking data is heterogeneous. A single customer generates card transactions (with amounts, merchants, and categories), digital engagement events (app navigation, feature usage), communications (emails, push notifications), product events (account openings, plan changes), and risk events (alerts, flags).
The key architectural decision is how to represent this heterogeneity without losing information. PRAGMA introduces a key-value-time tokenization scheme that standardizes all event types into a uniform format. Each data point is represented by three components:
  • Semantic type (key): What field this is (for example, "merchant_category" or "channel")
  • Value: The field's content, encoded differently by type — percentile buckets for numerical values, single tokens for categorical values, and Byte Pair Encoding (BPE) subwords for text
  • Temporal coordinate: When the event occurred, encoded as log-seconds since the previous event plus calendar features (hour of day, day of week)
This approach preserves the structural nature of tabular data without inflating sequence lengths the way naive text serialization would. Think of it as giving the model a structured language for banking events — one that's rich enough to capture nuance but compact enough to process efficiently.

Privacy-first data preparation

Banks operate under strict data protection requirements. Before any data enters the training pipeline, you need a pseudonymization layer that removes direct identifiers while preserving the statistical patterns the model needs to learn.
The essential components work together as a pipeline. First, Hash-based Message Authentication Code (HMAC)-SHA256 hashing replaces customer identifiers with non-reversible tokens — the model can still track a customer's event sequence without ever seeing their real identity. Second, names, addresses, and contact details are suppressed entirely. Third, monetary amounts are converted to percentile buckets — the model learns that a transaction is "high value relative to this customer's history" without knowing the exact pound figure. Finally, a differential privacy budget (ε ≤ 4.0) ensures individual records can't be reverse-engineered from model weights.
For storage, Apache Iceberg tables with row-level deletion capabilities support General Data Protection Regulation (GDPR) right-to-be-forgotten requests without retraining the entire model. On AWS, you can implement this using Amazon S3  for storage, AWS Glue  for extract, transform, load (ETL) processing, and Apache Iceberg for GDPR-compliant table management.

Three-encoder architecture

The model architecture uses three Transformer encoder blocks that process different aspects of a customer's data:
  1. Profile State Encoder — Processes static customer attributes (account tenure, product holdings, risk band) and life-long events (first top-up, first international transaction). Uses Rotary Position Embeddings (RoPE) (Su et al., 2021) to encode temporal distance.
  2. Event Encoder — Processes individual banking events independently, learning within-event relationships between fields. Calendar features capture daily and weekly cycles.
  3. History Encoder — Fuses the profile state representation with the sequence of event representations, learning cross-event patterns and long-range dependencies.
Why three encoders instead of one? Because banking data has fundamentally different temporal scales. Your customer segment changes over months. Individual transactions happen in seconds. The three-encoder design lets each component specialize in its temporal domain before the History Encoder learns how they interact.
The model is pre-trained using masked language modeling (MLM): randomly mask tokens, then train the model to reconstruct them. This forces the model to learn the relationships between fields, events, and temporal patterns without requiring any labeled data.
PRAGMA scales this architecture from 10 million to 1 billion parameters, with the scaling behavior being task-dependent. Credit scoring benefits substantially from larger models, while simpler tasks like recurrent transaction detection see diminishing returns beyond 100 million parameters. Start with 100 million and scale based on measured gains — don't assume bigger is always better.

Downstream task adaptation

Once pre-trained, the backbone adapts to specific tasks in two ways:
  • Embedding probes — Freeze the backbone, extract embeddings, and train a simple linear classifier on top. This takes minutes and is useful for rapid experimentation.
  • LoRA fine-tuning — Update 2–4% of parameters using Low-Rank Adaptation. This takes hours rather than weeks and consistently outperforms training from scratch.
For production deployment on AWS, serve the model using Amazon SageMaker  real-time endpoints for latency-sensitive tasks like fraud detection and Amazon SageMaker Feature Store  for low-latency feature retrieval. The Feature Store InMemory tier  — available since 2024 — provides sub-millisecond read latency for fraud scoring use cases where every millisecond matters. Use batch transform for portfolio-level scoring tasks like credit risk assessment.

What this enables: business impact

A foundation model trained on your bank's data unlocks value across multiple dimensions. Let me translate the technical capabilities into business outcomes.

Downstream use cases

Based on PRAGMA's published results and industry benchmarks, a large retail bank can expect the following impact areas:
  • Credit loss reduction — Better risk discrimination means fewer defaults slip through. For a bank with a £10 billion+ lending book, even modest improvements in PR-AUC translate to £12–18 million per annum in reduced credit losses.
  • Fraud prevention — Higher recall with maintained precision means catching more fraud without overwhelming investigators with false positives. Expected value: £8–14 million per annum.
  • Model maintenance savings — Consolidating 200+ models into a single backbone with task-specific heads reduces engineering effort. Expected savings: £4–7 million per annum.
  • Faster time-to-value — New use cases that previously took 6–9 months to develop can be deployed in weeks using LoRA fine-tuning on the existing backbone.

Addressing the AML limitation

Here's something I respect about the PRAGMA paper: the authors are refreshingly honest about where their model fails. It underperforms on anti-money laundering (AML) tasks because AML detection requires cross-record network analysis that a single-record encoder can't capture. Fraud is about what one person does. Money laundering is about what a network of people do together.
Banks can address this by incorporating graph embeddings from Amazon Neptune Analytics  into the architecture. Neptune Analytics provides built-in openCypher graph algorithms — including Louvain community detection, PageRank, and Strongly Connected Components — that run nightly on your transaction network without any infrastructure management. Pre-computing these features and injecting them as additional input tokens extends the foundation model to capture the relational signals that AML requires. This isn't optional for most banks — it's essential.

Operational efficiency

Beyond direct financial impact, a foundation model fundamentally simplifies your operating model. For a large retail bank with hundreds of task-specific models: instead of hundreds of separate governance processes, you have one backbone validation plus lightweight downstream head assessments. Instead of hundreds of monitoring dashboards, you have one unified system tracking data drift, performance, and fairness. Instead of multiple teams doing feature engineering in parallel, you have one team maintaining the backbone and multiple teams fine-tuning task heads in days rather than months.
The total projected value at steady state: £30–51 million per annum, with an 18-month payback period.

Regulatory and compliance considerations

But can you actually deploy this in a regulated environment? Yes — and in some ways, a foundation model makes compliance easier than the status quo.

Model risk management (PRA SS1/23)

A foundation model should be classified as a Tier 1 Material Model under the Prudential Regulation Authority (PRA) Supervisory Statement SS1/23, given its influence across multiple downstream decisions. This sounds daunting, but consider the alternative: validating 200+ individual models, each with its own risk profile. A single backbone with a four-layer validation framework is actually simpler:
  1. Data layer — Validate input data quality, representativeness, and privacy controls
  2. Backbone layer — Validate the pre-trained model's stability, convergence, and representation quality
  3. Downstream layer — Validate each task-specific head independently against its own performance and fairness criteria
  4. Monitoring layer — Continuous validation of data drift, performance drift, and fairness metrics in production
My strong recommendation: engage your PRA supervisor during development, not after deployment. Proactive engagement builds trust and avoids costly redesigns at the validation stage.

Fairness and Consumer Duty (FCA)

The Financial Conduct Authority (FCA) Consumer Duty requires that outcomes are fair across customer segments. For a foundation model, this means systematic fairness testing across protected characteristics defined in the Equality Act 2010 (age, gender, ethnicity, disability, and others).
Implement demographic parity, equalized odds, and calibration metrics for each downstream task. Here's the advantage of a shared backbone: fairness interventions at the representation level benefit all downstream tasks simultaneously. Fix a bias in the backbone, and you fix it everywhere — rather than hunting for the same bias across 200 separate models.

Data privacy (UK GDPR)

Three technical controls address GDPR requirements:
  • Differential privacy (ε ≤ 4.0) ensures individual records can't be reconstructed from model parameters. Note that ε = 4.0 is a relatively permissive budget — it represents the PRAGMA paper's setting. Production deployments targeting stronger privacy guarantees often use ε ≤ 1.0; tighter budgets typically require more data or larger models to offset the accuracy cost
  • Row-level deletion via Apache Iceberg supports right-to-be-forgotten without full retraining
  • Pseudonymization pipeline removes direct identifiers before data enters the training environment

Getting started

Ready to assess whether this is right for your bank? In Part 2 of this series ,  I show you how to build the data foundations — the pseudonymization pipeline, BPE tokenizer, and GDPR-compliant storage layer. In Part 3 , I walk through the model architecture, training, and production deployment with complete code examples.
If you want to assess your readiness now, ask yourself these questions:
  • Do you have 18+ months of event data across multiple source systems?
  • Can you access transaction, digital engagement, and product event data in a unified data lake?
  • Do you have a model governance framework that can accommodate a shared backbone architecture?
  • Is your PRA supervisor open to proactive engagement on novel model architectures?
If you answered yes to at least three of these, you're ready to start. If you answered yes to all four, you're already behind — because your competitors are asking the same questions.
Try Amazon SageMaker  in the AWS Management Console to explore the ML infrastructure you'll need for this journey.

Conclusion

Remember those 200+ models I mentioned at the start? For a large retail bank, each one represents a team that spent months building something that a foundation model could deliver in weeks — with better performance. The evidence from PRAGMA demonstrates that banking event sequences admit transferable representations, just as text and images do. A single pre-trained backbone can outperform hundreds of task-specific models while simplifying governance, reducing costs, and accelerating innovation.
The investment is significant but bounded: based on the author's experience with comparable large-scale ML programmes, approximately £28–34 million over 30 months for a large retail bank is a reasonable indicative range, with projected annual value of £30–51 million at steady state. Your costs will vary based on data volume, existing infrastructure, and team composition.
The question isn't whether banks should train foundation models. It's whether you can afford to keep paying the fragmentation tax while your competitors eliminate it.
Are you exploring foundation models for your bank? I'd love to hear what's driving your decision — or what's holding you back. Share your thoughts in the comments below.

References

  1. Ostroukhov, M., Mikhailov, R., Iashin, V., et al. (2026). "PRAGMA: Revolut Foundation Model." arXiv:2604.08649.
  2. Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2019). "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding." NAACL-HLT.
  3. Hu, E.J., Shen, Y., Wallis, P., et al. (2022). "LoRA: Low-Rank Adaptation of Large Language Models." ICLR.
  4. Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). "Attention Is All You Need." NeurIPS.
  5. Su, J., Lu, Y., Pan, S., et al. (2021). "RoFormer: Enhanced Transformer with Rotary Position Embedding." arXiv:2104.09864.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article