All posts
Research·

Inside the Graph Foundation Model

Part 2 of a three-part series on Avra's Graph Foundation Models: why heterogeneous graph neural networks, what the proprietary Large Knowledge Graph is, and the engineering decisions behind the model — Matryoshka embeddings, sampling at scale, and hard constraints.

  • Jonas RodriguesResearch Scientist
16 min read

Part 2 of a three-part series on Avra’s Graph Foundation Models. Start with Part 1: From enterprise data to graph embeddings.

Part 1 followed a client’s relational database through the pipeline that turns it into graph embeddings, and stopped at the model that produces them. This post is about that model: the Graph Foundation Model itself. Why build it on graph neural networks — and why heterogeneous ones? What is it pre-trained on? And what engineering makes it work at the scale of an entire economy?

Why heterogeneous graphs

The Avra GFM is built on top of Graph Neural Networks (GNNs), a family of architectures designed to learn from graph-structured data. Although extensive research in GNNs has produced a variety of architectures (GCN, GraphSAGE, GAT, R-GCN, and others), they are not all equally suited for the kind of graph Avra operates on. Many foundational GNN architectures assume a homogeneous graph, where nodes and edges share a common type. That can be effective for certain applications (e.g. social networks, citation networks) but it is not a good fit for the complexity of enterprise data.

The graphs Avra operates on are not homogeneous. They contain multiple entity types (companies, people, cities, lawsuits) and multiple relationship types (ownership, transactional, geographic, economic, and many more domain-specific relationships). A homogeneous GNN would treat all nodes and edges as the same, losing critical distinctions between different types of entities and relationships. Hence, Avra uses a heterogeneous GNN architecture that captures this diversity explicitly, learning different parameters for different types of nodes and edges in order to capture the rich semantics present in the data.

In the Avra LKG, not all edges carry the same information. An ownership relationship between two companies implies something different from a transactional relationship, which implies something different from a judicial-regulatory link. A homogeneous GNN that processes all edges with the same learned parameters would wash out these distinctions.

Avra’s GFM solves this by making type-awareness a first-class property of the architecture. Every entity type and relationship type in the graph gets its own dedicated learned capacity, so an ownership edge and a transactional edge are processed differently by construction — the model captures the unique semantics of each relationship present in the graph rather than averaging them away. Such an architecture is more expressive and a better inductive bias for the kind of data Avra operates on, enabling it to differentiate between complex patterns in the data — leading to richer embeddings and better performance on enterprise decision-making tasks.

homogeneous graphsame node type, same edge typeall nodes and edges treated uniformlyheterogeneous graphdifferent node and edge typesCompanyPersonLawsuitCitySuppliertyped nodes, relations, and parameters

Figure 1 — Left: a homogeneous graph, where all nodes and edges are treated the same. Right: a heterogeneous graph, where distinct node and edge types are modeled explicitly.

The Large Knowledge Graph: Avra’s proprietary foundation

The GFM does not arrive at its understanding of the Brazilian economy from a generic graph benchmark or a customer’s first-party data. It is pre-trained on the Avra Large Knowledge Graph (LKG): a proprietary, temporal graph assembled and curated over years, covering the interconnected web of economic actors and relationships that constitute Brazil’s economy.

The LKG spans many entity types representing major economic roles in the Brazilian economy: people, companies, transactions, geographic units, enforcement records, and more. The number of relationships is larger still, connecting these entities into a dense, interconnected graph.

That density is the key to what the GFM can do. For entities where local information is already plentiful, graph context amplifies the most relevant signals, sharpening what tabular features already capture. For entities where local information is scarce (new companies, thin-file individuals), relational learning is often the only alternative: their profile and position in the economy is only legible through the inductive bias captured via their connections.

Building this graph required not only data collection and curation work, but years of ontology design: deciding what entity types exist, which relationships are semantically meaningful, how to handle temporal versioning, and how to absorb new data sources without breaking existing representations. The moat is not just model code. Reproducing a comparable foundation requires the same institutional investment in data sourcing, curation, and ontology design accumulated over years.

Schema, architecture, and evolution

This raises a subtlety that is easy to overlook: the GFM architecture and the LKG schema are deeply coupled. A graph foundation model organizes its learned capacity around the entity and relationship types present in the LKG. The model’s capacity is literally shaped by the LKG schema, with new entity types or relationship classes requiring new learned parameters.

This coupling is a strength: a richer LKG schema translates directly into a more expressive model, capable of capturing distinctions that a homogeneous or schema-agnostic architecture would wash out. Other foundation-model families can also be represented as learning over structured objects — sequences, grids, patches, or tokens. The GFM architecture starts from the more general case: irregular, typed, semantic relationships.

However, this coupling also means that schema evolution is a first-class engineering concern. The LKG schema of enterprise data is not static. New entity types and relationship classes are added as new data sources are integrated, as the economic landscape evolves, and as client needs drive the addition of new types of information. This dynamic nature requires the model to accommodate new types without degrading its representations of existing ones. In practice, the GFM architecture must allow new parameters corresponding to new entity and relationship types to be added with a degree of flexibility and modularity — without requiring a complete retraining of the model or a degradation of performance on existing types.

Making this adaptation seamless — absorbing new schema elements without sacrificing existing representations, handling schema drift in upstream data sources, and adapting to new client datasets with minimal manual intervention — is one of the core engineering challenges in building and maintaining a GFM. It requires careful design of the model architecture, training procedures, and data pipelines so the model can evolve alongside the LKG schema while maintaining high-quality representations across all entity and relationship types. This has become a specialty of Avra’s research team, built up over years of experience with the LKG and GFM.

Key engineering decision 1: Matryoshka embeddings

The GFM produces high-dimensional embeddings (e.g. 1024d, 2048d, 4096d, depending on customer needs). Not every downstream application benefits from such a high-dimensional representation. Some tasks, especially those with limited labeled data, may only support a smaller embedding size (e.g. 128d or 256d) before overfitting sets in.

The immediate question: how do you produce versatile embeddings capable of supporting multiple downstream tasks with different data regimes, without training separate models for each embedding size?

A common workaround is to train a single high-dimensional embedding and then apply dimensionality reduction techniques (e.g. PCA, UMAP, t-SNE) to produce smaller embeddings for specific tasks. That can be useful for analysis and visualization, but it is a poor default for production features. Projection methods can discard task-relevant signal, introduce another fitted artifact to govern, and make it harder to preserve a stable representation across many downstream models.

After careful consideration of enterprise client needs and the nature of GFM training, Avra concluded that the most effective path to versatile, reproducible embeddings is to bake Matryoshka Representation Learning (MRL) into the GFM training itself, as a first-class property of every embedding the model produces.

The Matryoshka property means the embedding is structured such that any prefix of the vector (e.g. the first 128, 256, or 512 dimensions) is itself a well-formed, informative representation of the entity, albeit at a coarser level of granularity than the full embedding. Each additional block of dimensions adds progressively finer-grained signal, allowing downstream models to choose the appropriate level of detail for their specific task and data regime.

# Matryoshka loss: sum over nested prefix lengths
dims = [64, 128, 256, 512, 1024]
total_loss = sum(
    task_loss(embedding[:, :d], labels)
    for d in dims
)

The result: any prefix of the embedding vector is itself a well-formed, informative representation. Slicing the first 128 dimensions is semantically valid; slicing an arbitrary interior block such as dimensions [100:228] is not.

This approach creates a clean division of responsibility. The complexity of enforcing the Matryoshka property is entirely Avra’s concern: it lives in the training objective, not in the client’s workflow. The customer-facing interface is simple: choose a prefix length, get a valid representation. No dimensionality reduction, no tuning of projection methods, no risk of destroying the semantic structure of the embedding.

64d prefix · AUC 0.72128d prefix · AUC 0.85 ★256d prefix · AUC 0.800–6364–127128–255256–511512–10231024d →practical dimension guide16–64d · rule-based systems, extreme latency128d · real-time inference256–512d · tree models (start here)1024d · deep learning, vector retrieval

Figure 2 — A 1024-dimensional Matryoshka embedding. The first 64 dimensions carry the strongest structural signal; each additional block adds progressively finer-grained context. AUC values are an illustrative validation profile; the optimal prefix depends on available labeled data. Prefix truncation (taking the first d dimensions) is always safe; slicing arbitrary interior blocks is not.

An illustrative validation profile shows why this matters:

Prefix lengthROC-AUC
64d0.72
128d0.85 ★
256d0.80

The right prefix is not a hyperparameter to guess. It is an empirical question: sweep across prefix lengths (64, 128, 256, 512, 1024), measure validation performance, and choose the smallest prefix that gives the desired lift. Because GFM weights are frozen and only the downstream head changes, this sweep is cheap. If performance starts to degrade, that is a useful pruning signal, but nearby prefixes should still be checked.

Key engineering decision 2: Coping with massive scale

Even though Avra’s Knowledge Graph is already a large dataset, enterprise data may include billions of nodes and edges, especially for clients with large transaction volumes, extensive supply chains, or complex ownership structures. Despite remarkable advances in scaling deep learning models, training at the scale of relational data presents unique challenges rooted in its combinatorial nature.

To paint a clear picture of the challenges and the solutions, imagine a simplified scenario: a graph of 10 people, with edges representing friendships. We select a node — say, Alice — and check her friendships (the 1-hop neighborhood). Alice is friends with Bob and Carol. Now, if we check the friends of Alice’s friends (the 2-hop neighborhood), Bob is friends with Dave and Eve, while Carol is friends with Frank and Grace. As we keep expanding the neighborhood, the number of nodes we need to visit grows exponentially — it practically doubles with each additional hop. This combinatorial explosion is characteristic of relational data: both its greatest strength and its biggest challenge.

DaveEveFrankGraceBobCarolAliceseed · 1 node1-hop · 2 nodes2-hop · 4 nodes

Figure 3 — From a single seed node, the reachable neighborhood grows exponentially with each additional hop. Alice (1 node) → Bob and Carol (2 nodes) → Dave, Eve, Frank, and Grace (4 nodes). In a real enterprise graph with thousands of neighbors per node, this explosion reaches billions of reachable nodes within 3–4 hops.

If one were to follow every possible relationship starting from a given node, the number of nodes encountered would be too large to store or process in memory. If each node has on average 10 neighbors, then after 3 hops we would have to consider 10³ = 1,000 nodes; after 5 hops, 10⁵ = 100,000 nodes; and after 10 hops, 10¹⁰ = 10 billion nodes. This exponential growth quickly becomes unmanageable as the neighborhood expands.

In deep learning, this explosion is particularly problematic because training repeats the process over many seed nodes and mini-batches. Even if a single node’s neighborhood could fit in memory, the combined memory requirements for processing many sampled neighborhoods would quickly exceed available resources, making training impractically slow. To train a GNN on such a large graph, one needs to balance the trade-off between the richness of the relational signal captured by larger neighborhoods and the computational feasibility of processing them.

This is a fundamental challenge of training GNNs on large graphs, and it requires careful engineering that other model families usually do not face in this form. This single problem hints at how powerful GNNs can be: if capturing the relational signal in large neighborhoods were as simple as processing other data modalities, the results relational models have achieved across domains would not have been so surprising.

This scenario illustrates why the model needs to access the relevant parts of the graph efficiently. Loading the entire neighborhood of a node into memory would quickly exhaust it, especially for graphs as large as the LKG. That’s why GNN training relies on neighborhood sampling: instead of looking at the entire neighborhood, we sample a subset of neighbors for each node during training. This allows training on large graphs without loading the entire graph into memory at once.

In more detail: sampling starts from the node we wish to generate an embedding for (usually called the seed node). We sample a fixed number of neighbors from the seed’s immediate neighborhood (1-hop neighbors). Then, for each sampled neighbor, we repeat the process to sample their neighbors (2-hop neighbors), and so on, up to a certain number of hops or until a specified budget of selected nodes is reached. This produces a local subgraph of controlled size around the target node — manageable enough to fit in memory, yet rich enough to capture the relational context needed for meaningful representation learning.

full graphsample 1-hopsample 2-hopseedsampled 1-hopsampled 2-hopnot sampled

Figure 4 — Neighborhood sampling builds a manageable subgraph around a seed node. Left: the full graph, all neighbors reachable. Center: sampling selects 2 of 3 one-hop neighbors; the unsampled neighbor stays white. Right: for each sampled one-hop neighbor, one two-hop neighbor is selected. The result is a local subgraph of fixed, controlled size.

Key engineering decision 3: Supporting constraints

Supervised learning on enterprise data almost always involves a temporally structured dataset: features built from the past, labels derived from the future. The central concern is data leakage — the model must never see information that would not have been available at prediction time. When predicting whether a company will default within 90 days, the model must not have access to any events that occur after the prediction date, including whether the company actually defaulted.

Consider a dataset of companies with features updated at different frequencies: monthly financial statements, daily transaction data, and quarterly ownership records. We want to predict whether a company is in good standing over the next 90 days. The model must only see financial statements, transactions, and ownership data up to the anchor date: the point in time at which the prediction is made. Any data after that date, whether used as a feature or as a label, must be strictly partitioned. This temporal discipline directly determines the integrity of the sampling process.

time →anchor timemonthly financials (−12m)daily transactions (−90d)ownership (−1q)← historical featuresavailable at anchor timelabel window (+90d)future labels →must not enter the input graphpost-anchor edge: excluded

Figure 5 — Training example construction at an anchor timestamp. Features are computed from historical data (left of anchor); labels are derived from the future window (right). The model must never see data from the right side when constructing its input features, including through graph edges created after the anchor date.

The temporal nature of enterprise data imposes one of the most fundamental constraints when working with graph-structured data for predictive modeling. Instead of freely sampling any part of the graph, the sampling process must respect the temporal constraints of the data, so that the information present in the sampled subgraph is consistent with what would be available at prediction time. When we sample the neighborhood of a node for training, we must only include nodes and edges that would be available at the anchor date, and exclude anything that would only become available afterwards. This is a critical aspect of training GNNs on temporal graph data, and it requires careful engineering to make the sampling process both efficient and compliant with the temporal constraints.

Temporal constraints are not the only constraints that matter for enterprise applications, though. In many cases, there are domain-specific invariants the model must respect. In credit risk modeling, there is a fundamental financial invariant: the probability of defaulting within 60 days must be at least as large as the probability of defaulting within 30 days. Cumulative default probability is non-decreasing over time by definition — if a company defaults within 30 days, it has also defaulted within 60 days.

Left unconstrained, a neural model will occasionally violate this invariant. Violations are both incoherent and operationally expensive: lenders reject models that produce non-monotonic forecasts because they cannot be integrated into standard risk frameworks or regulatory filings.

A soft loss penalty, such as a hinge loss that penalizes out-of-order PDs during training, can reduce the frequency of violations but cannot eliminate them. A model trained with a soft constraint will still occasionally produce non-monotonic sequences at inference time, which is unacceptable in a regulatory context.

For cumulative risk outputs, hard constraints are the safer production pattern. Avra enforces monotonicity through structural output constraints and calibrated post-processing checks so that delivered PD term structures are non-decreasing by construction. No matter what the underlying model produces, the exposed output must satisfy the invariant.

This reflects a broader design principle: Avra is a platform to accelerate business decision-making, not another layer of complexity for a client’s model validation team to audit. By encoding domain constraints as hard structural properties — not soft training incentives — the models Avra delivers are easier to validate, monitor, and integrate into standard risk workflows. Clients still run their normal model-risk and compliance review, but the invariant is handled by design instead of left as a recurring exception to detect after the fact.

monotone sequence ✓0.250.50PD30d60d90d180d365dviolated sequence ✗0.250.50PD30d60d90d180d365dPD(90d) < PD(60d)

Figure 6 — Left: a valid monotone PD sequence — cumulative default probability rises with each horizon. Right: a violated sequence where PD(90d) < PD(60d) (highlighted).


The architecture only earns its keep on real data. Part 3 is the proof point: what happened when this model was deployed at a large Latin American financial institution — where relational signal, not tabular features, turned out to be the single largest driver of improvement — the fraud-detection results, and the research directions we’re pursuing next.

Continue to Part 3: Graph foundation models in production.