All posts
Research··Updated Aug 3, 2026

From enterprise data to graph embeddings

Part 1 of a three-part series on Avra's Graph Foundation Models: how a client's relational database becomes a temporal graph, and from there the embeddings that downstream models consume.

  • Jonas RodriguesResearch Scientist
8 min read

This is Part 1 of a three-part series on how Avra builds Graph Foundation Models.

LLMspath · single token typet1t2t3t4attention over full sequenceGPT · LLaMA · BERThomogeneous nodes,no graph structureCNNs / ViTsgrid · single patch typelocal receptive fieldResNet · ViT · ConvNeXthomogeneous nodes,grid-structured topologyGraph Foundation Modelsgeneral · multiple entity typesCPFCitySuitCNAECNPJnode typesCompany / CNPJPerson / CPFCity / StateLawsuit / CNAEtopology: sequence → lattice → general heterogeneous graph

Figure 1 — Foundation models occupy distinct positions in the space defined by graph topology and node-feature heterogeneity. Node shades denote entity type: a single tint for single-type models, the full blue scale for the GFM’s heterogeneous graph. Avra’s GFM operates on the general case: arbitrary topology, multiple entity types, typed edges.

Every enterprise decision in Brazil happens against a backdrop the model almost certainly doesn’t see.

A company applies for trade credit. Its balance sheet looks fine: moderate revenue, no visible delinquencies. What the balance sheet doesn’t show: its controlling shareholder sits on the board of three other firms, two of which defaulted six months ago, and one of which is under active judicial scrutiny. Capturing that context requires looking not at a row in a table but at an interconnected network spanning every economic participant: company, person, or city.

This is the problem Avra is built to solve. The Graph Foundation Model (GFM) is pre-trained on a temporal Large Knowledge Graph (LKG) spanning Brazilian companies, individuals, places, legal events, and the relationships between them. Ownership, supply, judicial, regulatory, and geographic connections are modeled as typed edges. The resulting model learns a structural and semantic language of the Brazilian economy that standard tabular models usually miss.

This series explains how that model is built. In this first post, we follow the data: how a client’s relational database becomes a graph, and from there the embeddings that downstream models consume. Part 2 goes inside the Graph Foundation Model itself — the architecture and the engineering decisions behind it. Part 3 turns to the numbers: the measured gains from real client deployments.

Step 1: Relational databases are already graphs

Most enterprise data lives in relational databases. Customers, transactions, accounts, contracts — each in its own table, connected to others via primary and foreign key relationships. This structure is already a graph, just encoded in a form that makes its topology invisible to a model.

The transformation to an explicit graph is mostly mechanical:

  • Entity tables become node types. A customers table becomes a Customer node; a seller table becomes a Seller node.
  • Foreign-key relationships become edge types. If a transactions table carries both a customer_id and a seller_id, each row implies an edge between a Customer and a Seller.
  • Timestamp columns are preserved as node and edge metadata. The resulting graph is temporal by construction.

When a client’s schema is not a clean entity-relationship diagram (e.g. overlapping identifiers, implicit relationships, loosely coupled systems), Avra specialists work with domain experts to define an explicit ontology: what entities exist, which relationships are semantically meaningful, and what each edge type represents. This work pays down. A well-constructed ontology is the foundation for a high-quality graph embedding, and it produces a durable governance artifact that teams can reason about long after the initial model ships.

The client graph built this way is used for relational fine-tuning (described in Step 4). The GFM itself is pre-trained on Avra’s proprietary LKG, not the client graph.

GFMpre-trained on the Avra LKGClient RDBentity–relationshipGraph mappingontologyRelational fine-tuningper workspaceMatryoshka emb.64 / 128 / … / 1024Downstreamapplication

Figure 2 — End-to-end data flow. The client’s relational database is translated to a graph via ontology-driven schema mapping. The GFM, pre-trained on the Avra LKG, provides frozen embeddings that combine with the client’s workspace graph for relational fine-tuning and downstream model training.

Step 2: Why relational data is harder than it looks

Flattening a relational database into feature vectors is deceptively lossy. The core problem is combinatorial.

In a graph, the number of candidate paths grows quickly with neighborhood size. A bank with 500k customers and 10M transactions does not need to be anywhere near fully connected for the 3-hop search space to become impractical: high-degree entities, shared counterparties, and repeated transactions can turn a single seed into millions or billions of candidate paths before temporal and semantic filters are applied. No feature engineering process can enumerate that space exhaustively. Traditional ML on relational data handles this by aggregating: count transactions in the last 30 days, sum balances, compute variance. These aggregations are human-defined and inevitably incomplete, as they encode what the feature engineer knew to ask for, not the full relational signal.

A graph model does something different. It learns to propagate information along typed edges and aggregate at each node, discovering which structural patterns are actually predictive. The embedding it produces is a compressed summary of this learned propagation.

In the economic context, the stakes for getting this right are concrete. A company connected to a defaulted peer through a shared guarantor carries structural risk that does not appear in any of its own database rows. A merchant whose primary supplier just lost a judicial dispute may face supply chain disruption within weeks. These signals exist in the graph. They do not exist in a flattened table.

Step 3: What an embedding actually is

An embedding is a fixed-length vector: a point in a high-dimensional space that encodes an entity’s position in the learned representation of the graph.

It is not a lookup table. Individual coordinates do not map to interpretable features like “default probability”, “supplier risk score”, or “propensity to buy a certain product”. The entire vector is a compressed summary of the entity’s relational context. This is analogous to how a compressed archive is an efficient encoding of a document rather than a human-readable text. The signal is recovered when you feed the embedding into a machine learning model, which learns to interpret it into valuable knowledge that improves predictions for a given application.

For example, two companies with identical balance sheets but different ownership networks will have significantly different embeddings based on their ownership structure. This is what the model buys you: the ability to distinguish entities that look identical to a tabular model but behave differently once you consider the complex relationships present in enterprise data.

Step 4: Relational training per workspace

Avra’s GFM is a general-purpose foundation for other machine learning models. It has learned the structural language of the Brazilian economy, but it has not seen any client-specific data. Adapting it to a client’s prediction task is the relational training step.

Avra logically isolates this adaptation per workspace, the tenant-isolated unit within a customer organization. Each workspace gets its own tailored relational foundation model — a customized GFM trained on that workspace’s first-party data, with strict temporal validation to prevent label leakage. After training, the workspace-specific GFM produces customized graph embeddings for that workspace’s entities, which are then fed into downstream models: predictive models trained for business decision-making in customer-facing applications.

There are two main approaches for leveraging Avra’s intelligence in a workspace: late fusion and full relational training. The choice depends on the client’s data landscape and compute budget.

Late fusion (recommended starting point). Uses the pre-trained GFM embeddings to augment the client’s existing features, and trains a downstream model directly. In this scenario, GFM weights are frozen; only the customer-facing model learns. This is fast, robust to data scarcity, and immediately unlocks relational signal.

Full relational training. When a client has a large first-party database with substantial labeled data and rich relational structure, Avra can train a foundation model specialized to the customer’s reality. A personalized foundation model is trained end-to-end on a personalized graph, created by merging the client’s data with Avra’s intelligence from both the LKG and the pre-trained GFM. This produces a workspace-specific embedding that better reflects the client’s entity relationships. Finally, a downstream model can be trained on top of these embeddings. This approach is more computationally intensive and requires more data, but it can yield superior performance by leveraging the full power hidden in the client’s relational data.

Late fusion is the right starting point for clients seeking to quickly unlock relational signal with minimal engineering effort, or for those with limited labeled data. Full relational training is ideal for clients who want to maximize performance and are willing to explore the full potential of what their data and Avra have to offer.


We’ve followed a client’s data from relational tables to trained embeddings — but we’ve treated the model that produces them as a black box. Part 2 will open it up: why Avra builds on heterogeneous graph neural networks, what the proprietary Large Knowledge Graph is (and why its schema is fused to the model), and the three engineering decisions — Matryoshka embeddings, neighborhood sampling at billion-node scale, and hard temporal and monotonicity constraints — that make the model hold up in production.

Part 2: Inside the Graph Foundation Model — coming soon.