Why most people pick the wrong graph database and then blame the model

I spent three weeks debugging a Cypher query that kept timing out on a production Neo4j instance before I realized the problem wasn't the query at all. The real issue was that I had indexed the label instead of the relationship type, and the query planner was doing a full table scan across 40 million nodes every time. This happens constantly when you treat a graph database like a relational one with extra steps. An open source graph database is a storage engine where the primary data structure is the graph itself. Nodes represent entities, edges represent relationships, and both can carry properties. That sounds simple enough on paper. The hard part is figuring out when this actually matters versus when you just need a well-normalized Postgres schema and are done with it.

Choosing an open source graph database for your stack

The three most relevant projects right now are Neo4j Community Edition, JanusGraph, and ArangoDB. Neo4j Community is what most people reach for first because the documentation is good and the Cypher query language is expressive. JanusGraph is what you pick when you need to span multiple backend storage engines and handle massive fan-out patterns at scale. ArangoDB is the multi-model option if you want graph traversal without giving up document storage. Here is the part nobody talks about upfront. Neo4j Community Edition has a hard limit: it runs single-threaded for write operations and there is no native sharding. If your write throughput exceeds roughly two thousand transactions per second on a decent machine, you are already past what that edition handles gracefully. I learned this the hard way when a real-time recommendation pipeline started queuing transactions and the latency spiked from 12 milliseconds to 4 seconds overnight. The workaround was moving the write path to JanusGraph on top of Cassandra while keeping Neo4j for the read-heavy traversal queries that power the analytics dashboard.

How graph traversal actually works under the hood

Unlike a SQL database that rewrites your query into a join plan and hopes the optimizer picks the right indexes, a graph database follows pointers. When you ask for nodes connected to a starting node through a relationship of type MANAGES, the engine walks the edge directly. There is no join. There is no Cartesian product. The cost is proportional to the number of relationships you traverse, not the total size of the database. This is why graph databases excel at recommendation engines, fraud detection networks, and access control graphs. They also fall apart immediately when you try to use them for bulk analytics across entire tables. I once saw a team try to run a monthly revenue aggregation by traversing every transaction relationship in the graph. It took 47 minutes. The same query on a columnar store with the proper indexes ran in 8 seconds. Sometimes the right tool is just not a graph database.

Get the Full Details

Open Source Graph Cayley – An Open Source Graph Database In Go
Open Source Graph Cayley – An Open Source Graph Database In Go

Practical setup and the mistakes that slow you down

Starting with Neo4j Community Edition is straightforward. You download the tarball or use the Docker image, set the initial password, and you are running. The default configuration works for development. It does not work for anything that resembles production. The first thing I always change is the heap size. The default 512 megabytes is fine for a few thousand nodes. It is useless for a real dataset. I typically allocate between 4 and 8 gigabytes depending on the working set size, and I set the page cache to about 70 percent of available memory if this is a dedicated server. Indexing is where most people burn themselves. Neo4j uses index-only scans when possible, which means a query like MATCH (n:User {userId: "abc123"}) RETURN n is fast if you have a uniqueness constraint on User.userId. Without that constraint, it falls back to a full node scan. I always create constraints before loading data. Loading millions of nodes without constraints and then adding them afterward forces a full re-scan of the entire dataset, which can take hours depending on your hardware. Another thing that trips people up is the difference between labels and indexes. Labels are just tags. They do not enforce uniqueness or speed up lookups unless you add an index or constraint on top of them. I see this mistake repeatedly in codebases where someone creates a large label and expects it to behave like a primary key. It does not. You need an explicit CREATE CONSTRAINT or CREATE INDEX statement for that.

Query patterns that matter

The most common productive use case I run into is shortest-path and k-hop traversal for access control. If you model your permissions as a graph where users connect to roles and roles connect to resources, a single Cypher query can determine whether someone has access through any combination of direct assignment and group membership. In a relational model, this requires recursive CTEs or application-level logic that traverses layers one at a time. The graph version is one query. For fraud detection, the pattern is slightly different. You are looking for tightly connected clusters that share attributes like device ID, IP address, or billing information. I use a combination of GDS (Graph Data Science library) and iterative traversal for this. The community detection algorithms in GDS are fast because they operate entirely in memory. The catch is that they require the subgraph to fit in the allocated heap, so you typically run them on sampled or partitioned data rather than the full database.

When to walk away from a graph database

There are honest scenarios where a graph database is the wrong call. If your data is predominantly flat with occasional joins, a relational database will be simpler and faster. If you need complex analytical aggregations across billions of rows, a columnar warehouse is the better fit. If your relationships are static and never change, there is no point in paying the write amplification cost that graph databases incur. My rule of thumb is this: if your primary query pattern involves following relationships more than three hops deep, or if the relationship structure itself is the thing you are analyzing, use a graph database. If your queries are mostly point lookups by ID or range scans on a timestamp, stick with what you have. The graph model adds overhead that only pays off when relationships are the core of your workload.

GitHub - memgraph/documentation: The official documentation for Memgraph open-source graph database.
GitHub - memgraph/documentation: The official documentation for Memgraph open-source graph database.

Common pitfalls with an open source graph database

The biggest pitfall is assuming that import speed scales linearly with hardware. Loading data into Neo4j using the built-in CSV importer is fast, but only up to a point. Once you start exceeding the page cache capacity, the engine starts flushing to disk on every transaction, and the import time can increase by an order of magnitude. I solved this by batching imports with AUTO INDEX turned off during load and rebuilding indexes in bulk after the import completes. This cut my 3-hour import down to roughly 25 minutes for a dataset of about 12 million nodes and 45 million relationships. Another pitfall is ignoring query plan visualization. The EXPLAIN and PROFILE clauses in Cypher are not optional. I have wasted days chasing slow queries that looked fine on paper because the query planner was choosing a plan that scanned the entire relationship index instead of using a node lookup. Running PROFILE on any non-trivial query will show you exactly where the bottleneck is. Most of the time it is an unexpected full scan, and the fix is a missing index or a constraint that should have been there from the start. There is also the matter of backup and recovery. Neo4j Community Edition does not support online backups. If you need to back up a large database without taking it offline, you are looking at JanusGraph with a distributed backend or moving to the Enterprise edition. This is a hard requirement for any production system and it is easy to overlook until you actually need it and your database is down for four hours while you copy files.

The ecosystem around these tools has matured significantly over the past few years. The Graph Data Science library, the Cypher query language, and the broader tooling around data modeling have all reached a point where a well-designed graph database is a reliable production choice. It is still not a universal solution. It is also not something you should adopt just because it sounds interesting. Pick it when your queries are relationship-driven and your data model benefits from that shape. Otherwise, you are just adding complexity for no return.