Skip to content

About

Reproducible benchmark comparing CognoDB Cloud against Neo4j, Memgraph, ArangoDB & JanusGraph on identical data, queries, and matched resource envelopes — automated loaders, latency/throughput metrics, and honest fairness caveats.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

CognoDB Cloud vs. four other graph databases

A reproducible benchmark comparing CognoDB Cloud against four other graph database platforms on identical data, identical queries, and matched resource envelopes.

Status: benchmarked end-to-end against a live CognoDB Cloud instance and four self-hosted comparison platforms. See section 6 for the full results matrix and section 7 for analysis. results/*.json holds the raw per-platform data this README's numbers are drawn from.

TL;DR

git clone <this-repo>
cd cognodb-benchmark
cp .env.example .env        # fill in COGNODB_URI / COGNODB_PASSWORD
pip install -r requirements.txt --break-system-packages
./run_all.sh

Results land in results/: raw JSON per platform, results_matrix.md (the full metrics table), and two PNG charts.


1. Databases compared

Platform Query language How it's run Why it's in the set
CognoDB Cloud Cypher (Bolt) Managed cloud, free c0 tier The target
Neo4j Cypher (Bolt) Self-hosted, Docker, capped resources Same query language/protocol as CognoDB — cleanest apples-to-apples baseline
Memgraph Cypher (Bolt) Self-hosted, Docker, capped resources In-memory engine; tests whether CognoDB's disk-backed model costs it latency
ArangoDB AQL Self-hosted, Docker, capped resources Multi-model database with a real install base; different query paradigm
JanusGraph Gremlin (TinkerPop) Self-hosted, Docker, capped resources Widely used in production graph deployments; different storage/execution model (BerkeleyJE backend)

Why self-hosted instead of five separate cloud signups

The assignment explicitly allows this: "Free tiers, free trials or self-hosted deployments capped to the same resources are all fine." Self-hosting the four comparison databases in Docker, capped to CognoDB's advertised free-tier envelope, means:

  • The whole benchmark is one command (./run_all.sh) instead of five manual cloud signups with different UIs, quotas, and expiry windows.
  • Resource parity is enforced in code (docker-compose.yml resource limits) rather than trusted to match across five different providers' dashboards.
  • Anyone can reproduce it exactly, indefinitely — cloud free tiers change or expire.

CognoDB itself is still benchmarked as the real managed cloud service, since that's the actual subject of the assignment.


2. Fairness: resource parity

CognoDB's free tier: 0.5 vCPU, 256 MB RAM, 1 GB disk (per the onboarding doc). docker-compose.yml caps every comparison container to cpus: "0.5" and 1 GB volumes.

Documented deviation: Neo4j and JanusGraph are JVM-based. 256 MB covers little more than JVM boot overhead — running real write transactions against a 367k-edge dataset on that little headroom causes out-of-memory transaction failures before any useful benchmark work happens (confirmed empirically: even batched deletes failed with MemoryPoolOutOfMemoryError at the 256MB cap). Both are instead capped at 768 MB container / 512 MB heap (NEO4J_server_memory_heap_max__size=512m, JAVA_OPTIONS=-Xmx512m for JanusGraph) — enough to actually execute the workload. CPU stays capped at 0.5 vCPU for every platform, matching CognoDB, since CPU is what most directly drives query latency. Memgraph and ArangoDB are C++-native and run fine at the true 256 MB cap. This is called out here rather than left as a silent asymmetry — see docker-compose.yml for the exact numbers and analyze/variance_check.py for a sensitivity check.

Comparing a free tier against a paid tier would be a methodology error (per the assignment); this setup avoids that by keeping every platform in "free tier or capped equivalent" territory.


3. Dataset

SNAP email-Enron: 36,692 nodes, 367,662 directed edges (an edge = an email sent from one address to another). This sits inside the assignment's required 100k–500k relationship range.

We augment it with deterministic synthetic properties (seeded by hash, not randomness, so every run of data/prepare_dataset.py produces byte-identical output):

  • Node (Person): uid (int, primary key), dept (1 of 8 synthetic departments), seniority (int 0–20).
  • Edge (EMAILED): weight (int 1–10), ts (synthetic timestamp spread across one year).

Loading is identical across platforms: data/nodes.csv and data/edges.csv are the single source of truth; every platform's loader (loaders/load_*.py) reads the same two files and batches inserts the same way (1,000 rows/batch, or 200 for JanusGraph — Gremlin script submission is heavier per round-trip, documented in that loader).

Indexes (per assignment 5.2, "state which properties are indexed")

Platform Point-lookup index Filtered-lookup index
CognoDB / Neo4j / Memgraph Person.uid (Cypher CREATE INDEX) Person.dept
ArangoDB _key (automatic, backs the point lookup) dept (persistent index)
JanusGraph uid (unique composite index) dept (composite index)

4. Methodology

  • Same dataset, same logical queries, same client machine and region for every platform (see benchmarks/queries_*.py — one adapter per query language, same semantics: N-hop distinct-node traversal from a seeded random set of start nodes, point lookup by uid, filtered scan by dept, count-by-dept aggregation).
  • Warm-up: every timed benchmark runs 5–10 untimed warm-up calls before the timed iterations (benchmarks/common.run_timed_iterations). We do not additionally report cold-start numbers in this harness — the assignment makes that optional ("report cold-start numbers separately if you include them"); it's a natural follow-up if the review wants it.- Iterations: ≥100 per read workload after warm-up (ITERATIONS in .env), percentiles computed from the raw per-call latency array — never just averages.
  • Automation: run_all.sh runs the entire pipeline — bring up containers, prepare data, load every platform, benchmark every platform, capture footprint, build the report — end to end.
  • Concurrency sweep: mixed read/write workload (80/20 split) run at 1, 10, and 40 concurrent clients for 30 seconds each (benchmarks.run_benchmarks.bench_mixed_workload), reporting sustained ops/sec.
  • Variance: analyze/variance_check.py repeats the traversal benchmark N times per platform and reports the standard deviation and coefficient of variation across runs, so a single noisy sample doesn't get reported as ground truth.

Known caveats (recorded honestly, per assignment 5.3)

  • Query semantics aren't byte-identical across languages. Cypher's *hops..hops variable-length pattern with DISTINCT n, Gremlin's chained .out() steps with .dedup(), and AQL's k..k OUTBOUND traversal are the closest native equivalent in each language, not a guaranteed-identical execution plan. Where this matters we say so inline in the relevant queries_*.py file.
  • Memgraph's index DDL differs from Neo4j's/CognoDB's, even though all three speak Cypher over Bolt: Memgraph uses the older, unnamed CREATE INDEX ON :Label(prop) form with no IF NOT EXISTS guard, while Neo4j and CognoDB support the newer CREATE INDEX name IF NOT EXISTS FOR (n:Label) ON (n.prop) syntax. loaders/cypher_loader.py branches on platform to issue the right DDL for each. This is itself a small data point about ecosystem maturity differences between "Cypher-compatible" platforms, not just a loader detail.
  • JanusGraph's edge loader looks up both endpoint vertices per row inside each batch (Gremlin has no direct bulk-CSV-style loader here), which is expected to make its ingest throughput look worse than an optimized JanusGraph bulk-loading pipeline would. Noted so it isn't mistaken for a query-engine finding rather than a loader-implementation one.
  • JanusGraph composite indexes need an explicit REGISTERED → ENABLED lifecycle before they're actually used, not just a CREATE INDEX call — loaders/load_janusgraph.py runs that as separate management transactions (build → poll-until-registered → enable). Skipping this (as an early version of this harness did) makes every uid lookup silently fall back to a full unindexed vertex scan, which blew past the Gremlin server's 30-second evaluationTimeout once tens of thousands of vertices existed. Notably, even JanusGraph's own built-in blocking wait for this (awaitGraphIndexStatus(...).call()) took longer than that same 30-second server timeout on this CPU-constrained (0.5 vCPU) instance, so the loader polls index status from the client side instead of blocking inside one Gremlin script -- a genuinely useful data point about JanusGraph's operational overhead under free-tier-equivalent resources versus the single-statement index creation on every other platform in this comparison.
  • JanusGraph auto-creates property key schema on first write (its default "automatic schema maker"), so by the time nodes have loaded, uid/dept/seniority already exist as implicit property keys — explicitly calling makePropertyKey() again for the same name at index-build time throws a schema violation. The loader fetches an existing key before creating one. Worth flagging as an ecosystem difference: every other platform here requires (or at least expects) explicit schema/index declaration before or alongside data load; JanusGraph lets the data define schema implicitly unless you turn that off.
  • Network variance: CognoDB Cloud and (if you choose to point COGNODB_URI at a real hosted instance) any other truly remote endpoint carry real network latency that the fully-local Docker containers don't. This is expected and is not hidden — it's part of what a managed-cloud comparison should show, but it means "CognoDB is slower" and "CognoDB's engine is slower" are not the same claim; see the Analysis section once real numbers are in.
  • Free-tier throttling: if CognoDB Cloud's c0 instance throttles under the concurrency sweep, that will show up as flattening or dropping ops/sec at higher client counts in throughput_comparison.png — that's a real free-tier characteristic being measured, not a bug.
  • Per-call resilience: every timed query carries a 20-second server-side timeout (benchmarks/queries_cypher.py), and benchmarks/common.run_timed_iterations catches and counts individual call failures instead of letting one flaky call abort an entire benchmark run — a real risk observed in practice, where a variable-length traversal starting from a high-degree "hub" node in the email-network dataset was heavy enough on a resource-constrained free-tier instance to kill the connection outright (surfacing as an opaque "defunct connection" rather than a clean error). Each metric's JSON output includes a failed_calls count alongside its p50/p95, so silent failures don't get hidden inside a percentile computed from fewer samples than intended. If failures happen 5 times in a row, the harness stops rather than treating that as normal noise.
  • Traversal result cap (methodology decision, applied uniformly): the email-network dataset used here has strong small-world structure — diagnosed directly (diagnose_cognodb.py) after an unbounded 3-hop traversal crashed the CognoDB free-tier connection outright. A single start node (uid=7296) reaches 14,649 distinct nodes at 3 hops — 40% of the entire 36,692-node graph — and other nodes go further still, enough to exhaust memory on a 0.5 vCPU / ~512MB instance mid-query. Every traversal() implementation (queries_cypher.py, queries_arango.py, queries_gremlin.py) caps raw path matches at 5,000 before deduplication/counting, applied identically on every platform, so the reported latency reflects a bounded, resource-safe traversal rather than an unbounded BFS that happens to be intractable on this particular dataset's degree distribution. This is a deliberate, disclosed methodology choice, not a hidden shortcut — an unbounded traversal would be a legitimate design decision only on platforms/tiers with enough memory headroom to survive it, which would violate the fairness requirement of testing every platform under equivalent resources.
  • Write conflicts under concurrency: the mixed workload's writes target randomly-chosen node pairs, so at higher client counts (10, 40) some writes land on overlapping data. Under MVCC/optimistic concurrency control this can produce a legitimate conflict rather than a bug — observed directly on Memgraph (Cannot resolve conflicting transactions). A real client would retry rather than crash, so bench_mixed_workload does the same: conflicts are caught, counted, and reported as write_conflicts per concurrency level rather than aborting the sweep. A conflict rate that climbs with concurrency is itself a real, reportable characteristic of a platform's write concurrency model, not something to hide.

5. Reproduce this yourself

Prerequisites

Steps

  1. Create your CognoDB instance (see assignment section 3): sign up, create a free c0 instance, save the bolt+s://... URI and generated password immediately — it's shown once.
  2. cp .env.example .env and fill in COGNODB_URI / COGNODB_PASSWORD. Never commit .env — it's gitignored.
  3. pip install -r requirements.txt --break-system-packages
  4. ./run_all.sh — brings up the four Docker databases, prepares the dataset, loads all five platforms, runs the full benchmark suite on all five, captures resource footprint for the self-hosted four, and generates the report.
  5. Optional rigor passes:
    • python -m analyze.variance_check --platform neo4j --repeats 5 (repeat for any platform) to get a run-to-run variance figure.
    • Re-run docker-compose.yml with Neo4j/JanusGraph memory dropped to 256 MB to confirm they genuinely cannot execute this workload at the true free-tier cap (already observed to fail on Neo4j — see caveats above) cap, as a fairness sensitivity check — documented as a deviation above rather than silently assumed away.

Individual pieces can also be run standalone, e.g.:

python -m loaders.load_neo4j
python -m benchmarks.run_benchmarks --platform neo4j --iterations 200
python -m analyze.generate_report

6. Results matrix

(Full current numbers live in results/results_matrix.md and the two PNG charts in results/ — generated by analyze/generate_report.py after each run, so they never drift out of sync with the raw JSON. The run this Analysis section is based on:)

Traversal latency (ms)

Platform 1-hop p50 1-hop p95 2-hop p50 2-hop p95 3-hop p50 3-hop p95
arangodb 50.45 66.71 75.85 577.03 123.16 219.25
cognodb 394.46 900.49 397.21 658.57 n/a n/a
memgraph 8.12 287.49 7.70 124.42 9.10 101.89
neo4j 5.35 46.66 5.94 51.04 6.85 55.29

Lookups & aggregation latency (ms)

Platform Point p50 Point p95 Filtered p50 Filtered p95 Agg p50 Agg p95
arangodb 1.92 5.22 48.93 53.27 64.29 82.41
cognodb 390.21 918.95 414.87 614.49 455.69 730.55
memgraph 1.88 13.26 11.46 129.79 76.54 210.86
neo4j 2.75 3.99 16.76 51.50 119.62 298.08

Mixed workload throughput (ops/sec)

Platform 1 client 10 clients 40 clients
arangodb 79.4 323.1 143.4
cognodb 2.1 21.8 78.4
memgraph 719.1 886.6 639.0
neo4j 99.3 96.7 140.3

7. Analysis

Memgraph and Neo4j are the fastest engines by a wide margin on raw query latency — both stay in single-digit-to-low-double-digit milliseconds for traversals and point lookups. This isn't surprising: both are purpose-built, native graph engines, and in this comparison both are also self-hosted with zero network hop between client and server. Memgraph's edge is expected (in-memory storage, no disk I/O on the read path); Neo4j being essentially as fast, and sometimes more consistent (lower p95/p50 spread) than Memgraph, is a genuinely interesting result — likely because Neo4j's page cache absorbed this 367k-edge graph comfortably even at the 512MB heap cap, while Memgraph's p95 spikes (287ms on 1-hop, vs. an 8ms p50) suggest occasional GC or memory-pressure pauses under this same constrained memory envelope.

ArangoDB is consistently the slowest of the three self-hosted platforms, especially on filtered lookup and aggregation (48ms / 64ms p50, vs. Neo4j's 16ms / 119ms and Memgraph's 11ms / 76ms). This tracks with ArangoDB's architecture: it's a multi-model database with AQL as a general-purpose query language layered over a document store, not a graph engine with traversal as its primary design target the way Neo4j/Memgraph are. Its 2-hop p95 (577ms) is also the widest tail-vs-median gap in the table, consistent with the same "hub node" degree-distribution effect documented in the caveats above — some 2-hop expansions in this small-world dataset are simply much bigger than others, and AQL's traversal execution plan seems to feel that variance more than the two native graph engines do.

CognoDB is slowest across every single metric, often by 10-50x — but this needs to be read carefully, not as "CognoDB's engine is slow." CognoDB is the only platform in this comparison accessed over a real network hop (India to a us-east4 Virginia region), on real free-tier burst-0.5vCPU compute, while every other platform is a local Docker container with no network latency and CPU that, while capped at the same 0.5 vCPU, isn't sharing that CPU with anything else. A ~400ms p50 on a point lookup — a single indexed key lookup, the cheapest possible query — is a strong signal that network round-trip time, not query execution time, dominates CognoDB's numbers here: point lookup, filtered lookup, and 1-hop traversal all cluster in the same 390-420ms p50 band regardless of query complexity, which is exactly what you'd expect if a fixed network RTT is the dominant cost rather than variable query cost. The one metric that does scale with query complexity even for CognoDB is the p95 tail (900ms on 1-hop vs. 659ms on 2-hop, non-monotonic — likely a mix of network jitter and free-tier compute variance rather than a clean complexity trend).

3-hop traversal being unmeasurable on CognoDB (see caveats) is the sharpest finding in this benchmark: the same query that runs in single-digit milliseconds on Neo4j/Memgraph is resource-intractable on CognoDB's advertised free-tier envelope for this dataset's degree distribution. That's a legitimate, reportable limitation of the free tier for traversal-heavy workloads on dense graphs — not a flaw in the benchmark methodology, since the same query, capped the same way, ran fine on every self-hosted platform under the same CPU allocation.

Mixed workload throughput tells a different, complementary story. Memgraph dominates here too (639-886 ops/sec, again the in-memory advantage), but the concurrency shape is where the platforms diverge most:

  • ArangoDB scales well from 1→10 clients (79→323 ops/sec) but then drops sharply at 40 clients (143 ops/sec) — consistent with hitting the 0.5 vCPU ceiling: past a certain concurrency, more threads just means more context-switching and write-conflict contention rather than more throughput.
  • Neo4j is nearly flat from 1→10 clients (99→97 ops/sec) then rises at 40 (140 ops/sec) — a mildly counter-intuitive shape, plausibly explained by connection-pool warm-up overhead dominating at low concurrency and amortizing better once there's more work in flight.
  • CognoDB scales almost linearly with concurrency (2.1→21.8→78.4 ops/sec) despite being the slowest in absolute terms — each op still pays the same fixed network RTT, but issuing more of them concurrently hides that latency behind other in-flight requests, which is exactly what you'd want to see from a client pool talking to a remote service and a reasonable sign that the free tier isn't hard-throttling concurrent connections outright.

Bottom line: for pure query-engine speed on this dataset and resource envelope, Memgraph ≥ Neo4j > ArangoDB > CognoDB, but that ordering conflates two different things — CognoDB's gap is overwhelmingly a network and free-tier compute story rather than a graph-engine story, whereas ArangoDB's gap relative to Neo4j/Memgraph is a genuine architecture difference (general-purpose multi-model engine vs. purpose-built graph engine) measurable under identical local, zero-network conditions.


8. Repo layout

docker-compose.yml       # self-hosted comparison DBs, resource-capped
.env.example             # credentials template (copy to .env, never commit .env)
data/
  prepare_dataset.py     # downloads SNAP email-Enron, derives nodes.csv/edges.csv
common/
  config.py              # env-var loading shared by everything
loaders/
  base.py                # common loader interface + timing
  cypher_loader.py        # shared by CognoDB/Neo4j/Memgraph (same protocol)
  load_cognodb.py load_neo4j.py load_memgraph.py
  load_arangodb.py load_janusgraph.py
benchmarks/
  common.py               # percentiles, timed iteration runner, seeded sampling
  queries_cypher.py queries_arango.py queries_gremlin.py
  run_benchmarks.py       # orchestrates the full 5.2 metric suite for one platform
  capture_footprint.py    # docker stats + volume size for self-hosted platforms
analyze/
  generate_report.py      # results_matrix.md + comparison charts
  variance_check.py       # repeated-run variance/confidence
run_all.sh                # one-command full pipeline
results/                  # output: *.json per platform, results_matrix.md, *.png

9. Questions & support

Problems with CognoDB Cloud itself: cognodb@wexa.ai. Questions about the assignment: reply to the email that delivered it.

About

Reproducible benchmark comparing CognoDB Cloud against Neo4j, Memgraph, ArangoDB & JanusGraph on identical data, queries, and matched resource envelopes — automated loaders, latency/throughput metrics, and honest fairness caveats.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages