A document model stores records as the main objects. A graph model also stores the connections between them, as explicit links, called edges. Both can hold the same information, so my question was what changes in practice when you store connections and can query them directly. I tested it in my final-year project: I built a graph next to an existing document store of real lending data, and asked the same six questions of both.
The short version. For questions you already know how to ask, a graph doesn't reveal information that wasn't already in your data. What changes is where the work of connecting records happens: once, in a tested pipeline, instead of again in every query. I think that's worth paying for when the same connecting logic keeps being rebuilt across reports and tools, and not when a single document already holds the answer.
In my comparison, five of the six document queries rebuilt the same ID by gluing two fields together; the graph stored that link once.
What matters is whether connections are stored and checked, or rebuilt by every query. A graph is one way to store them, and a relational database with declared foreign keys gets much of the same benefit. The design worked for the questions I tested. I can explain why it helps with consistent numbers and audits, but whether that's worth the cost is a decision each company has to make based on its own needs.
Company data is also increasingly queried by AI tools, not only by people. Stored, named links could matter for them too, and that's where this article ends, and my next project starts.
1. The premise nobody states
Graph databases are often sold with the promise that they will reveal connections hidden in your data. I believed that for a while, but then I realised the connections weren't hidden anywhere. They were already in the data; the graph just stores them differently. There is also a simpler promise: when connections have their own names, questions about them become easier to ask. So that's the promise I wanted to test. There are three main reasons people want a graph:
- Retrieval. Questions that follow connections, like what is linked to what, through what, and how many steps apart, are easier to write when the links are stored.
- Readability. Named links show how the data fits together, in a way both people and programs can read. That matters more now that AI tools are asked to query company data, because without named links they have to guess how records connect. Section 7 comes back to this.
- Algorithms. These are methods that look at the whole network at once, like finding groups, the most connected players, or patterns that repeat over time. Section 6 looks at one of these, temporal motifs.
Many of these algorithms work on the network as a matrix, a table marking which pairs of records are connected. Some graph databases, like FalkorDB, store graphs that way underneath.
That also shows something important. The same connections can live in documents, relational tables, a matrix or a graph database, and moving them doesn't add or lose any information. This stops being true once time matters, because a single matrix can't say that A paid B before B paid C. For that you also need the dates, and section 6 comes back to this.
That's why I think the useful question is what you buy by moving the work of connecting records, who pays for it, and when. I set out to show that the graph version was better than the document version it came from, and ended up with a more useful answer than the one I wanted.
2. Setting it up so the answer meant something
Comparing two data models is harder than it looks. If you put the document version in one database and the graph version in another, you also change the query language and how fluently you write each one. Once you time anything, you also change how each database plans its queries. Any difference you find could come from any of those.
I was lucky here. I didn't choose ArangoDB for the experiment, and I didn't build the document store either: both were already there, behind the commercial system my data came from. My part was the graph side. I designed the model, built the pipeline that loads it and the checks that verify it, and wrote the queries for both versions. Because ArangoDB stores documents and graphs side by side and queries both with AQL, I could build the graph alongside the existing documents from the same records and compare the two in one language. That removes the problems above. I compared the shape of the queries, not their speed, so the control that matters is the language: if both sides use the same one, differences in structure can't come from syntax. Section 5 explains why performance is out of scope.
This has limits. AQL may not be the most natural language for either model. And my six comparisons all had a fixed number of steps, because those were the questions I was asking of this data. None needed shortest paths, or whether one thing connects to another no matter how many steps away. That's my main point again: the question decides what the model has to do well. It also means my result doesn't cover those questions, and they're exactly where a graph may have a clearer advantage. Questions like knock-on effects, where a problem at one party spreads to others connected to it, could turn out to be that kind.
The two schemas were also built for different jobs. The document store takes in records from several platforms that all name things differently and keeps working when one of them changes, and it does that well. My graph was built to answer questions about how a process connects. That means part of what I measured is the difference in purpose, and this isn't a general claim about documents versus graphs.
3. How I decided what counted as "better"
I built the graph, and I wanted it to win, so I couldn't just pick the version that looked better to me. Before comparing anything, I set three tests and used them on every pair of queries:
- Simplicity: things you can count. How many query blocks there are, how many collections the query uses, and how many joins or steps it needs. A step means following one link from one record to the next.
- Clarity: can you tell what the query is for just by reading it, without a comment explaining it?
- Easy to change: if the question changes a little, can you make a small change, or do you have to rewrite the whole query?
These are my own tests, and they're still judgements, so I linked each one to something I could point to in the query. They also line up with an established standard: clarity and ease of change correspond to analysability and modifiability, two parts of maintainability in the ISO/IEC 25010 software quality standard.
I also kept track of three tricks the document queries needed, because the connections weren't stored:
- gluing two text fields together to rebuild an ID
- searching inside a list stored inside a record
- connecting records by comparing names
Two of these have names in the relational world. Storing a list inside one field is what Karwin (2010) calls Jaywalking, and leaving relationships undeclared is his Keyless Entry.
I didn't count lines of code, because the number of lines depends on how the data is stored, not on how hard the question is. A document that already holds everything about a record takes one line to fetch. The graph stores the same details as separate records, so the query needs a line for each one. A short query can simply mean the data was kept in one place, so line counts would make the graph look worse even where its query is the safer one.
Counting steps is fairer, but it still treats every step as equal, and section 4 explains why that matters.
4. What I found: the work is conserved
I compared six pairs of queries. Each pair asked the same question twice, once of the documents and once of the graph. The questions were of three kinds:
- Looking up one thing: for example, everything about one shipment.
- Following the whole path: for example, from the purchasing policy, through the orders placed under its rules, to the shipment that delivered each one.
- Totals and breakdowns: for example, how much was spent with each supplier, split by product, month and department.
The data came from a commercial system, so I can't show the real details. The example below uses a made-up domain instead: a buyer, its suppliers, purchase orders and shipments. The queries have the same shape as the real ones, but the names and details are invented.
In the example, the buyer has one purchasing policy for its approved suppliers, with one rule per supplier that sets that supplier's prices and terms. Every order is placed with one supplier and has to follow that supplier's rule. Once the order is placed, it is shipped.
In the graph, each of these is a record, and each link between them is stored once with a name:
- the policy has rules (
POLICY_HAS_RULE) - each rule governs the orders placed under it (
RULE_GOVERNS_ORDER) - each order is placed with a supplier (
ORDER_PLACED_WITH_SUPPLIER) - each order is fulfilled by a shipment (
ORDER_FULFILLED_BY_SHIPMENT)
The documents don't store these links, so every query has to recompute them.
Here is one of the path questions, written both ways: for a given purchasing policy, which fulfilled orders were placed under it, which rule covered each order's supplier, and which shipment delivered each order?
Document version
LET policy = DOCUMENT("policies", "P-4471")
FOR order IN purchase_orders
FILTER order.policy == policy._key
FILTER order.status == "fulfilled"
LET rule = FIRST(
FOR r IN policy.rules
FILTER LOWER(r.supplier_name) == LOWER(order.supplier_name)
RETURN r
)
FILTER rule != null
LET shipment = DOCUMENT(
"shipments",
CONCAT(order.carrier_code, "_",
TO_STRING(order.carrier_consignment_id))
)
FILTER shipment != null
RETURN {
order: order._key,
supplier: order.supplier_name,
rule: rule.rule_ref,
shipment: shipment._key,
delivered_at: shipment.delivered_at
}
In plain words:
- Begin with the policy record.
- Identify all fulfilled orders referencing this policy.
- For each order, locate the corresponding rule within the policy's rules list by matching supplier names.
- Construct the shipment's unique identifier by concatenating two specific fields from the order.
- Use this constructed ID to retrieve the associated shipment record.
Three fragile things are happening here:
- A manufactured key. The shipment's ID is built by gluing two fields together, so the query depends on a naming convention that exists only in people's heads. It also assumes each order has exactly one shipment, because one ID can only find one document.
- A name match. The order is matched to its rule by comparing supplier names in lowercase, so a trailing space or a renamed company breaks it without any error.
- A nested-list filter. The rule lookup searches a list stored inside the policy, so it depends on that list keeping the same shape across records from different sources.
Graph version
LET policy = DOCUMENT("policy", "P-4471")
FOR rule IN 1..1 OUTBOUND policy POLICY_HAS_RULE
FOR order IN 1..1 OUTBOUND rule RULE_GOVERNS_ORDER
FILTER order.status == "fulfilled"
FOR supplier IN 1..1 OUTBOUND order ORDER_PLACED_WITH_SUPPLIER
FOR shipment IN 1..1 OUTBOUND order ORDER_FULFILLED_BY_SHIPMENT
RETURN {
order: order._key,
supplier: supplier.name,
rule: rule.rule_ref,
shipment: shipment._key,
delivered_at: shipment.delivered_at
}
In plain words: the query starts from the same policy and follows its named links. It goes to the policy's rules, then from each rule to the orders it governs (keeping only fulfilled ones), then from each order to its supplier and its shipment.
There's no manufactured key, no name matching and no nested list, and the language is the same in both. The work was done once, when the data was loaded, and the links were built. That's what I mean by the work being conserved: the same logic still exists, but it runs once in the pipeline, and every query reuses the result.
Most of that work went into one link, RULE_GOVERNS_ORDER. Building it
means doing the same supplier-name match the document query does. It also means
deciding, once, what happens when two rules name the same supplier. The document
query answers that with FIRST(), which silently takes whichever match
comes first, depending on the order the documents happen to be in. The pipeline has
to answer it explicitly, write the answer down and test it.
What the six pairs showed
Manufactured keys, nested-list filters and name matches all appeared only on the document side. Five of the six document queries rebuilt an ID by gluing fields together, and none of the graph queries did. A document store built for analysis could also store those IDs when the data is loaded, but then it pays the same cost in its own pipeline, which is my point.
The other two tricks aren't visible in the data itself. Nothing says that
rules[].supplier_name is supposed to match
purchase_orders.supplier_name. That knowledge lives in the head of the
analyst who wrote the query, and it leaves with them. An AI tool asked to query the
same documents would have to guess it too.
The two models also start in different places. Five of the six document queries had to begin from an internal mapping record. This separate document links records to each other, because that's where the storage makes the connection findable. Graph queries start from the thing the question is about, such as the policy.
On the step counts from section 3, the graph actually did worse: in most pairs it took more steps. What differed was the kind of step. Every graph step does the same simple thing, following one named link to the next record, so you can read what the query is for from the link names. The document versions mix lookups, status filters, nested fields, manufactured keys and name matching.
Ease of change showed most clearly in a query that broke spending down four ways. Adding a fifth way to the graph version would take one more link to follow and one more entry in a list. In the document version, I'd have to find the right field in records from different sources and check that they were named the same way.
What the pipeline costs
The moved work isn't free. I built the pipeline in four stages: extract, normalise, link, load. I gave every record a fixed ID made from source fields that don't change, so the same input always produces the same record. Rerunning the pipeline updates records instead of duplicating them, and it rejects records without an ID. After loading, I ran separate checks to confirm that every link points to existing records and that the main chain the questions depend on has no breaks. Both checks came back clean.
The prototype has one weakness: when something looks wrong, it doesn't stop. It keeps going. If two sources disagree, for example, one says an order was placed on 3 March, and the other says 4 March, it picks one and writes the conflict to a log. If a link points to a record that doesn't exist, it skips that link instead of rejecting it. The post-load checks would only notice the gap if it happened to be on a path they test. In production, I'd make both of these stop the load, because quietly working around bad data is exactly the problem this approach is meant to catch. So the pipeline has a real cost. My argument is that it's cheaper to pay that cost once, in the pipeline, than again in every query.
Storing links has one more benefit: you can ask questions about the link itself, like which rule applied, whether it held on a given date, or how many steps separate two records. In the document version there's nothing to ask, because the link only exists while the query that builds it is running.
The pair that went against me
One comparison contradicted my thesis, and it's the one I'd want to see if I were reading this. The question was "give me everything about one shipment". The shipment document already holds its carrier, origin, destination, customs classification and insurance, with its milestones in a nested list, so the document version is a single fetch. The graph version needs a separate step for every detail that became its own record, plus one for each of three event streams: planned milestones, status changes, and tracking scans. For that question, it's worse in every way that matters.
I think that's a fact about my modelling choice more than about graphs. I built the schema around a process, how an order moves from policy to rule to supplier to shipment, so questions about the process became short, named paths. A question about one thing's details isn't a process question, and the same split that made the process easy to follow scatters one shipment's details across many steps. The document version made the opposite choice: it kept the shipment together and left the process implicit.
The split isn't pure cost, though. Pulling the three event streams out of the shipment is what makes them queryable as separate series. So comparing what was planned with what actually happened becomes a simple walk along links, instead of digging through a nested list.
So the real boundary, I think, is between a model built around processes and a model built around things, and the question you ask decides which one looks clumsy. That's uncomfortable, because it means a comparison like mine partly compares two modelling decisions, and not only two technologies. It's also the most useful thing I learned: the model follows the question, and if you model around one kind of question, expect to pay on another.
5. Why moving it is worth paying for
If the logic is conserved, why move it at all? This is the part that answers the business question.
Logic that gets rebuilt in every query gets rebuilt slightly differently each time. In my six document queries, the same ID-building logic appears five times. My copies match because I wrote them all, so the drift below is how it usually happens, not something I measured.
In a typical company, the same logic also ends up in a dashboard, a monthly report, a reconciliation script written in a hurry, and an Excel file that was meant to be temporary. Those are four versions, written by different people at different times, and nothing in the system knows they're supposed to agree. When a source changes a field, or someone handles a missing value differently, they drift apart, and months later two numbers disagree in a meeting. Keeping it in one pipeline means you write and test the logic once, so it can only be wrong in one place. The benefit is control and audit.
I didn't compare speed. I had limited time, so I focused on these internal business benefits. Speed matters when choosing a database, but it's a separate question and would make a good project on its own. LDBC's FinBench, a benchmark built for financial graph workloads such as anti-fraud, would be the place to start. A quick test wouldn't have answered it anyway. Both versions sit in the same database, but my dataset fit entirely in memory, where almost any query is fast. So the differences would mostly be noise, and they wouldn't show how either version performs on a large production system. I set this scope before running anything, and nothing here says the graph is faster.
History matters too, especially in regulated lending. Regulators and auditors often ask what was in force when a decision was made, or what state a loan was in on a given date. When history is kept only as dates on records, answering that means rebuilding it inside each report, and different reports can end up doing it slightly differently. When each change is stored as its own record, linked to what it changed, the question becomes a simple filter. Most of that benefit comes from storing changes as records, which a history table can also do. The graph adds a named link from each change to what it changed, which makes these questions easier to ask. That's a much smaller claim than meeting any regulation, and it's the only one my evidence supports.
A temporal graph goes one step further, because time belongs to the whole graph. You can ask for "the network as it was on a date" directly, without writing that query each time. I didn't use one in this comparison, but Raphtory (Steer et al., 2024), which I used in my thesis (section 6), supports this kind of view. Whether one helps is something I'd like to test in my next project (section 8).
6. Beyond retrieval: temporal motifs
Section 1 gave three reasons people want a graph. Sections 4 and 5 covered the first, retrieval, with questions I already knew how to ask, and section 7 comes back to readability. Here I wanted an example of algorithms, methods that look for patterns across the whole network, to show what having the graph makes possible.
I ran a first version of this on the lending data in my thesis. Here I share what I took from modelling data as a temporal graph and using Raphtory: what time on a link adds, how to tell which questions motifs can answer, and how to set the analysis up properly.
What time on a link adds
In a temporal graph, every link carries the time it happened. Not just "A paid B", but "A paid B at 10:42 on 3 March". That makes several kinds of work possible; three matter here. You can look back and see the network as it was on a date, as in section 5. You can find sequences: small patterns of events in a set order, within a set time. And you can train models on the whole history to predict what comes next, such as which link will form, a field called temporal graph machine learning. I focus on sequences found with temporal motifs because they need no training, and each count is made of real events you can point to. Raphtory's motif functions return the counts, though, not the events behind them, at least in the version I checked, so listing those takes extra work. That matters when you have to explain a result to an auditor or a regulator.
What temporal motifs are
A temporal motif is a small pattern of events that happen in a set order within a set time. For example: money comes into an account, and within an hour it goes out again. Paranjape, Benson, and Leskovec (2017) defined the version I used, and Arnold et al. (2024) applied it to cryptocurrency payments with Raphtory, a tool built for networks that change over time. Each link is one real event between two parties: who paid whom, and when.
How to read the counts
Arnold et al. give two rules for reading motif counts:
- Don't trust one total. One very busy account or one unusual month can make up most of the number. So look at each account separately, and at how the counts change from month to month.
- Compare with chance. A business account that pays and gets paid all day will show patterns by accident. So you also count the same payments, between the same accounts, with their times shuffled, many times over. This is the baseline. If the real count sits inside the range the shuffles produce, it's probably luck. If it's well above, something real is going on.
The shuffle keeps how many payments happen in each period and only changes who paid when. So it can tell you whether particular accounts line up in time, but not whether there is an overall rhythm, because every shuffled version keeps the same one.
Is it a motif question?
The method only helps if the question fits it. A simple test: can you tell the situation as a short story, in two or three steps, with a stopwatch? "Money comes in, and within an hour it goes out again." If you can, it's probably a motif question. Four things have to be true:
- Real parties act on each other: who paid whom.
- The order matters: this happened, then that.
- The gap matters: within minutes or days, not just at some point.
- You want to know whether it happens more often than chance, which is what the baseline tells you.
Some questions sound similar but aren't motif questions:
- Is there a rush at month-end? That's a rhythm. The baseline keeps it in every shuffled version, so count per day instead.
- Which originator (the company that made the loans) has the most arrears? That's a total, not a sequence.
- Are arrears rising? That's a trend in one number over time.
Where it's used
This is where the method is aimed commercially. Fraud and anti-money-laundering teams look for exactly these sequences. Most patterns fall into the three families Raphtory counts: stars around one account, back-and-forth between two, and triangles among three.
- Pass-through. An account receives money and sends almost all of it on within hours, a classic sign of an account used to move money rather than hold it. Busy business accounts also receive and pay all day, so the baseline asks whether the gap between money in and money out is shorter than chance would give.
- Fan-in, fan-out. Many small deposits from different people reach one account within a day, then leave as one large transfer. Or one stolen card is used at many shops within minutes.
- Back and forth. Two accounts trade the same asset back and forth to fake trading volume. Real trading partners also trade often, so the question is whether the reversals come faster than chance.
- Round trip. Money leaves A and comes back to A through B and C, making turnover look bigger than it is. VAT carousel fraud works this way.
Banks already catch many of these with simple rules, like "flag if more than X leaves within Y hours". What motifs add is counting every pattern at once, for every account, across several time windows, and checking each count against chance.
Where motifs fall short
- A count isn't a reason. A motif shows that a pattern happens more often than chance, not why. A payment processor or a payroll account passes money through all day, for perfectly good reasons. Motifs show where to look; you still have to explain what you find.
- The time window is a choice. A pattern that stands out within an hour can disappear within a day, and with enough windows to choose from, you can find the result you wanted. So I'd pick the windows from the business question before looking at the counts, and report all of them.
- Chance depends on the baseline. The shuffle keeps some things fixed and mixes others, and a different shuffle can give a different answer. Gauvin et al. (2022) describe many of these reference models; choosing one is part of the analysis, not a detail.
- With many accounts, some stand out by luck. Check thousands of accounts against the shuffled range and a few will land outside it by chance alone. The more you test, the stronger the evidence each result needs.
- They only see small patterns. Raphtory's motif functions count three events among up to three parties. A scheme that moves money through ten accounts over a week needs another method, such as time-respecting paths (section 7).
Applying it in Raphtory
The good news is that Raphtory does the heavy lifting. It counts the patterns for each account or for the whole network, and for the whole network it can try several time windows in one run. It even comes with a built-in baseline: one function returns your event table with only the times shuffled.
What it can't do is decide what your data means. That part is still yours:
- Define what counts as one event, and do it once, in the pipeline. It's the same point as section 5.
- Run the shuffle many times, so you can see the range chance produces.
- Check events that share a timestamp. Raphtory keeps them in the order they were loaded, unless you set the order yourself, so it's worth testing how much that order matters. Arnold et al. did this for payments in the same blockchain block and found it moved the counts by at most about 3%.
7. What I'd tell someone deciding
I wouldn't migrate by default. Start by building a separate analytical layer next to the document system of record (the main database the business treats as the official record), fed by a pipeline. Let the document store keep the job it's good at: taking in data from sources that all name things differently, and keeping each source's own view of it. A full migration could make sense later, if most of the questions a business asks turn out to be process questions, or if keeping the layer in step with its source costs more than it saves. I didn't test a migration, so those are conditions to check, not findings.
I'd reach for the graph in two situations:
- when the same connecting logic keeps being rebuilt across many reports and tools
- when "what was true on this date?" is a regular requirement rather than an occasional favour
Model it around the questions you actually intend to ask, because the modelling decision does more work than the technology choice. Don't reach for a graph because the diagram looks better, and don't expect a query to reveal anything. Everything a graph query shows you was already there, unless you built the links yourself from your own matching logic, and then it's also showing you that logic.
The same goes for choosing a method. Start by asking what you actually want to find out, because each kind of question needs something different built:
- "How did this case unfold?" To follow one case from start to end, break it down, or check what was true on a date, you need the whole process modelled, with every kind of link. That's the graph in sections 4 and 5.
- "Does this pattern show up more than it should?" To find a few events between two or three parties, close together in time, you don't need the whole process. You need a clean list of the events: who did what to whom, and when. That's what motifs are for (section 6).
- "Where did the money go?" To follow money from one account over many steps, you need the same list of events but a different method: time-respecting paths, which follow a link only if it happened after the previous one. Raphtory supports these too.
Why not just rebuild the network each time?
If a graph is just a matrix you can rebuild anywhere, why store it at all? You could build the network from the documents with a script and run the analysis in a graph library. At prototype scale, that works, and for algorithms it can be exactly the right choice. It's how I ran my own motif analysis.
What matters is whether building the network is a lasting, tested piece of work or something rebuilt for each analysis. A one-off export brings back the problem from section 4, with the manufactured keys and name matches now living in Python scripts. My own motif script is an example.
A stored graph buys three things:
- the ability to ask new questions along links you've already built, without writing new matching logic
- checks on the data as it's written
- one place where the meaning of each connection is defined
I think those are worth something, but on their own they're not worth a migration.
What I didn't evaluate is how often the question is asked and what a rebuild costs. Building a network usually means reading all the data. If the question comes up once a quarter for a report, that's fine. If it's asked hundreds of times a day, or sits behind something a person is waiting for, rebuilding every time is a waste.
On the other side, a stored graph has to stay in step with its source, and one rebuilt nightly from data that changes hourly is wrong in a different way. So it depends on how often the question is asked, how expensive a rebuild is, and how fresh the answer has to be. I measured none of those, so they belong in the decision next to the structural argument.
Named links and AI tools
This is the readability reason from section 1. Named links are documentation a program can read. Unwritten conventions, like the glued key or the name match standing in for a join, have to be rediscovered by every analyst who touches the data. More and more, they also have to be rediscovered by every AI tool asked to query it.
An AI tool can infer those conventions by spotting patterns, and it often will, confidently, but it can't reliably tell you which parts of its answer were inferred. A graph query can only follow connections that were built and checked when the data was loaded. That's a narrower ability, and an easier one to audit.
This is strongest against a document store where connections are conventions. It's much weaker against a relational database with declared foreign keys, or a semantic layer that defines business terms for machines, because both can state the same structural facts without a graph. My model went beyond structure: it captured when each change happened, and facts the pipeline derived from the data using rules, so you could trace which rule applied and when. A relational database or a semantic layer can store all of that too. What the graph adds is that these facts are connected to the records they describe by named links.
That leaves a practical question. Most companies won't move their whole database into a graph just so AI tools can read it more easily, and I wouldn't recommend it. But the separate layer I recommend above, fed by a pipeline and kept next to the system of record, might be enough. AI tools could read its named, checked links instead of guessing from the raw documents. That only works if the layer itself can be trusted: it has to record what happened and when, record how each fact was derived, and keep checking itself against the source, so any answer can be verified rather than just believed. I haven't tested whether that actually makes their answers more reliable.
8. What I'm building next
That question is where my next project starts. My thesis model already captured time and rule-based facts, so this builds on it rather than starting over. I'm building Liquet in public, with synthetic data for a made-up wealth manager: its own cash records and its custodian's monthly statement. The question is the one every finance team answers at month-end: does our cash book agree with the bank, and if not, why?
The records stay in a document database, like the system I studied, and a graph layer gives each connection a clear name and meaning. Rules do the matching; an AI agent investigates what's left, explains each difference with evidence, and a person approves it. The principle is simple: rules for what you can state, an agent for what you can't, and a graph as the shared, checkable context for both. The agent runs on an LLM that calls a small set of tested tools rather than writing its own queries, so you can swap the model without changing the design.
Then I'll test whether the agent is more reliable and more traceable working through the graph layer than from plain tables. Time is part of it from the start, because every reconciliation is about what was true, and what was known, on a given day. My second article, Model the Job, Not the Outcome, shows how I designed that graph.
9. The conclusion I can actually defend
I started this project wanting to show that the graph was better. What I can actually defend is a smaller claim: the graph moved the work of connecting records into one tested pipeline, without adding any information that wasn't already there. Whether that pays off depends on how often a business rebuilds that logic, and how often it needs to ask what was true on a past date.
The prototype ran on a narrow slice of one portfolio, so I hold these claims at three levels of confidence. The design works for the questions I tested. The business case looks promising but is untested. And I know what data would settle it: many investors and originators, not one. A lot of technical writing quietly treats the first as if it proved the second.
Temporal motifs are where I see the most commercial promise, and the most room to get it wrong. The patterns fraud and anti-money-laundering teams look for, like money passing straight through an account or coming back to where it started, are motif-shaped, and Arnold et al. suggest that motif counts for each account may prove useful as features for this work. Much of the published research uses cryptocurrency or synthetic data, and that's to be expected: bank transaction data is confidential, so it's rarely available for public research. Real-world testing is more likely to happen inside companies that build for these cases. Pometry, the company behind Raphtory, won the BIS Innovation Hub's 2025 Analytics Challenge by using Raphtory to detect financial-crime patterns such as layering, where large incoming payments are quickly spread across many accounts. Results from client work usually stay private, which I think is why the public evidence looks thinner than the commercial interest.
Tools have made the method much easier to reach. With Raphtory, counting the patterns for every account or the whole network, trying several time windows, and building the baseline takes a few function calls rather than a research project. The tool can't pick the question or read the answer. If you're thinking of using motifs, start by writing your question as a short story with a stopwatch. If it doesn't fit, pick another method. If it does, choose your time windows and your baseline before you look at the counts, and treat every count as a place to look, not a verdict.
If there's one thing I'd want a reader to take away, it's that the question comes first. It decided what my model had to do well, it decides which method fits, and a comparison like mine is only as fair as the questions behind it.
So I finished with a better question than "should I use a graph?". Once connections are explicit, checked and readable by machines, can an AI system use them more reliably than it can rebuild them for itself? And can a graph layer next to the existing database give it that, without moving everything? That's what I'm building next, with reconciliation as the test case.
A note on how I wrote this: the research, the graph model, the pipeline, the queries and the analysis are my own work from my thesis. I used AI tools (Claude) to help edit and restructure the text, and to check my summaries of the cited papers and the Raphtory documentation.
References
- Arnold, N. A., Zhong, P., Ba, C. T., Steer, B., Mondragon, R., Cuadrado, F., Lambiotte, R. and Clegg, R. G. (2024). Insights and caveats from mining local and global temporal motifs in cryptocurrency transaction networks. Scientific Reports, 14, 26569.
- Gauvin, L., Génois, M., Karsai, M., Kivelä, M., Takaguchi, T., Valdano, E. and Vestergaard, C. L. (2022). Randomized reference models for temporal networks. SIAM Review, 64(4), 763–830.
- ISO/IEC 25010, Systems and software quality models.
- Karwin, B. (2010). SQL Antipatterns: Avoiding the Pitfalls of Database Programming. Pragmatic Bookshelf.
- Paranjape, A., Benson, A. R. and Leskovec, J. (2017). Motifs in temporal networks. In Proceedings of the Tenth ACM International Conference on Web Search and Data Mining (WSDM), 601–610.
- Steer, B. et al. (2024). Raphtory: The temporal graph engine for Rust and Python. Journal of Open Source Software, 9(95), 5940.