Graph databases differ widely in speed and costs for queries and updates

Graph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines

DatabasesArtificial Intelligence

Summary

Not all graph databases work equally well for searching and updating connected data, even if they seem similar on the surface. The authors tested eight different database systems using a large biomedical network and a variety of common query types. They found that some systems were faster for small, local searches, while others were better for big scans or complex joins. Most importantly, the time it takes to load the data into the database varied hugely and often dominated the overall cost for tasks with fewer queries. This means choosing the right system depends a lot on how often you query versus how often you update your data.

What this means in practice

  • For enterprise data engineers: Optimize database choice by balancing data ingest speed against query demands for large connected datasets in biomedical or similar domains.
  • For database platform architects: Design graph database systems that better balance ingest throughput and query performance based on task-specific workload characteristics.

Authors

Donald Nguyen, Gurbinder Gill, Hadi Ahmadi, Christopher J. Rossbach

Abstract

Graph databases are frequently positioned as categorically necessary for connected-data workloads, yet the systems dimension along which they actually differ - query planning, indexing, and data-readiness cost - is rarely isolated from vendor framing. We construct a synthetic, biomedical-shaped property graph (1.02 million nodes, 5.34 million total node and edge rows) and a twenty-query workload spanning neighborhood lookups, bounded paths, set intersections, anti-joins, grouped aggregation, top-k ranking, temporal filters, full scans, and relational joins. We benchmark Corvic AI - a purpose-built columnar query engine underlying Corvic's ontology management layer ("memories")- against seven purpose-built or graph-extension database systems (LoraDB, Ladybug, DuckPGQ, Memgraph, Neo4j, HugeGraph, and FalkorDB) at three graph scales spanning three orders of magnitude. We report query latency geomeans, bulk-ingest throughput, point-update latency, and answer correctness for each system, and we derive a simple total-cost-of-ownership model that expresses the ingest/query trade-off as a function of query volume. Our central finding is that no system in this sample is categorically fastest: a native graph engine (Ladybug) outperforms Corvic AI on narrow, bounded-neighborhood shapes, while Corvic AI is faster on shapes that scan or join a large fraction of the graph, and a system implementing graph query syntax via SQL/PGQ (DuckPGQ) is measurably slower purely due to query-plan choice. The dominant cost differential in our data is not query latency but the cost of making data queryable at all: bulk-ingest throughput varies by three orders of magnitude across engines (5.0k-4.3M rows/s), a gap that a simple crossover-point calculation shows dominates total cost for any workload with fewer than roughly 105 queries per data refresh.