Graph foundation models in production
Part 3 of a three-part series on Avra's Graph Foundation Models: the measured performance gains from a live client deployment — where relational embeddings were the single largest driver of improvement — and what we're building next.
- Jonas RodriguesResearch Scientist
Part 3 of a three-part series on Avra’s Graph Foundation Models. Start with Part 1: From enterprise data to graph embeddings.
The first two posts built the machine. Part 1 turned a client’s relational database into graph embeddings; Part 2 opened up the Graph Foundation Model that produces them. This post asks the question that ultimately settles whether any of it matters: does it work on real client data?
One of the most common sayings in machine learning deployment is that “the real test of a model is how it performs in production”. At Avra, we treat large discrepancies between validation performance and production performance as an evaluation-design problem until proven otherwise. Enterprise data is usually plentiful — especially in the relational setting Avra operates in — but it must be partitioned, sampled, and validated in a way that reflects the decision the model will actually support.
This is not just a philosophical point but a practical one: in our client deployments, models have achieved production performance close to, and in some cases above, validation performance. That is evidence that temporal validation, leakage controls, and relational sampling discipline matter as much as architecture.
Measured performance gains
The clearest evidence comes from deployed client work.
One of our early partners is a large Latin American financial institution. With millions of transactions processed daily, they have a wealth of relational data that can be leveraged to build powerful predictive models for applications such as credit risk assessment, fraud detection, and customer segmentation.
This client partnered with Avra early and experienced both sides of what Avra offers: the immediate lift from using pre-trained GFM embeddings as drop-in features, and the larger paradigm shift of training a custom relational foundation model tailored to their data.
The results were clear. Using GFM embeddings as additional features for an existing XGBoost model already produced a significant boost in both performance and stability. Going further — training an end-to-end relational foundation model specifically tailored to their entity and relationship tables, with joint training of a downstream prediction head — delivered an additional lift over the baseline that used no GFM embeddings at all.
The customized relational embedding also proved useful outside the specific task it was trained for, showing strong performance when used as features for other downstream models in different business units within the organization. That is evidence of the generality and robustness of the learned relational representations. This highlights one of the key benefits of a graph foundation model like Avra’s: the ability to learn general-purpose relational representations that transfer across tasks and business units.
What drove the gain
The most important observation: relational features, not tabular features alone, were the largest single driver of improvement.
GFM pre-trained embeddings contributed a +3.3 percentage-point ROC-AUC gain from LKG-derived relational context, with zero client-specific training. Custom relational embeddings added a further +4.3 percentage points on top. The two signal types are complementary: pre-trained embeddings contribute economy-wide relational context learned from the broader Brazilian network of relationships; custom embeddings capture client-specific structural patterns, layering micro-level graph signals on top of the macro-level pre-training foundation.
What this establishes:
- The GFM transfers. A +3.3 percentage-point gain from pre-trained embeddings means the LKG has encoded signals that are genuinely portable to a new distribution. This is the pre-training dividend.
- Relational embeddings bring value even to highly engineered models. The client model before adding Avra’s intelligence was already strong, with a rich set of features engineered by an experienced data science team. That relational embeddings provided a significant boost on top of that is strong evidence of the unique value graph-based representations bring to enterprise predictive modeling.
- The two signals are complementary, not redundant. Pre-trained embeddings and custom relational embeddings capture different aspects of the same underlying graph. In this deployment, the gains were additive. The ceiling on performance is set by how fully a client can leverage the relational signal available in their data.
The same pattern holds across applications. On a separate fraud track using their data, relational embeddings drove gains across every metric without joint model changes:
| Credit risk — incremental gain over baseline | Improvement |
|---|---|
| GFM pre-trained embeddings (zero custom train.) | +3.3 pp ROC-AUC |
| Custom relational embeddings (on top) | +4.3 pp ROC-AUC |
| Total relational lift | +7.6 pp ROC-AUC |
| Fraud detection — gain from relational embeddings | Improvement |
|---|---|
| ROC-AUC | +5.3 pp |
| PR-AUC | +8.5 pp |
| KS statistic | +13.6 pp |
What we’re working on
A few directions that shape our next research interests:
GFM v3: individuals (CPFs) as first-class entities. Avra started with companies as the primary entity type — deliberately, because new companies are created every day with little to no historical data, making the cold-start problem acute and the case for relational intelligence clearest. GFM v3 extends that foundation by modeling natural persons as first-class graph citizens. The web of relationships between companies and individuals is dense and informative: business owners, directors, guarantors, and retail customers all exist in the same graph. Modeling both simultaneously unlocks new applications in personal credit, retail fraud, and AML, while also enriching company-level tasks with individual-level signals that a company-only model cannot access.
Scale via graph training infrastructure. Avra’s internal graph training library is the infrastructure backbone for GFM v3. Its high-performance sampler handles temporal filtering and bidirectional neighborhood traversal with the performance margin to train on economy-scale graphs. This lets us continue pushing the boundaries of what is achievable with relational representations, at whatever data volume clients bring.
Even more flexible. The goal is to make it as simple as possible for enterprises to plug in their tables and let Avra do the rest: training customized foundation models and downstream applications that are production-ready, adapted to their specific data, constraints, and schema. Schema evolution here is not a corner case. Much like the LKG, where new nodes and edges are added every day, enterprise data is in constant flux — new data sources, changing schemas, evolving relationships. Building a system that handles schema evolution and adapts to new data modalities without extensive manual intervention is a key focus for us: architectures and training routines flexible and robust enough to accommodate changes in the data without sacrificing performance or requiring significant rework.
Join us
Are you eager to put your hands on economy-scale graphs beyond typical public benchmarks? Do you feel challenged by the idea of training on a graph where the combinatorial explosion of data is a real issue to tackle? Do you want to be part of a team where each new metric, milestone, and success becomes a reason to pursue the next hard problem? Check out our careers page — we would love to chat.
Avra is an AI Frontier Lab building a platform for relational intelligence focused on enterprise decision-making, leveraging economy-wide relational data to produce performance gains that other foundation-model approaches often miss.
- Platform · Sep 4, 2026
LLMs vs. RFM: why ChatGPT can't predict your customer's next move
LLMs read, summarize, and reason over language, but they can't tell you which customer will churn or default. Why prediction over structured business data needs a different kind of foundation model.
- Research · Aug 11, 2026
Inside the Graph Foundation Model
Part 2 of a three-part series on Avra's Graph Foundation Models: why heterogeneous graph neural networks, what the proprietary Large Knowledge Graph is, and the engineering decisions behind the model — Matryoshka embeddings, sampling at scale, and hard constraints.
- Research · Aug 3, 2026
From enterprise data to graph embeddings
Part 1 of a three-part series on Avra's Graph Foundation Models: how a client's relational database becomes a temporal graph, and from there the embeddings that downstream models consume.