A high-performance, concurrent in-memory key-value store written in Java.
vkdb is an in-memory key-value database server inspired by Redis. It handles thousands of concurrent connections using Java virtual threads, persists data via an append-only log with background compaction, and supports real-time key-change notifications via a pub/sub mechanism.
Benchmarked on a single node (MacBook), 1000 ops per client, with durable fsync-before-ack writes and group commit enabled:
| Metric | 10 Clients | 500 Clients |
|---|---|---|
| SET-only throughput | 18,077 ops/sec | 21,694 ops/sec |
| GET-only throughput | 60,655 ops/sec | 56,697 ops/sec |
| Mixed (70R/30W) throughput | 36,593 ops/sec | 50,445 ops/sec |
| SET p50 / p99 latency | 0.52 ms / 0.88 ms | 23.3 ms / 33.8 ms |
| GET p50 / p99 latency | 0.16 ms / 0.23 ms | 8.72 ms / 15.37 ms |
| Errors | 0 | 0 |
500 concurrent clients × 1000 ops each = 500,000 operations completed with zero errors in every workload.
Writes stay durable — every SET/SETX/DEL is fsynced before it is acknowledged —
but group commit batches all writes waiting at a given moment into a single fsync, so
throughput scales with concurrency instead of being capped at one fsync per write. Under
500 concurrent writers, SET throughput is ~5× higher than the naive one-fsync-per-write
approach (4.3k → 21.7k ops/sec) and mixed throughput ~3.7× higher (13.6k → 50.4k), while
the durability guarantee is unchanged. Reads never fsync and stay in the ~57k ops/sec range.
- Concurrent Connections: Thousands of simultaneous clients via Java virtual threads
- Durable Writes (sync-before-ack): Each write is appended to the log and fsynced before the server replies, so an acknowledged write survives an immediate crash
- Group Commit: Concurrent writes waiting at the same moment are batched into a single fsync, so write throughput scales with load without weakening durability
- Binary Append-Only Log: Length-prefixed binary records with a versioned file header; values may contain any bytes (spaces,
=, newlines) without corrupting the log - Background Compaction: Periodic log compaction collapses overwrites and drops deleted/expired entries
- Key Expiration (TTL): Time-to-live on keys, enforced both lazily on read and by a background sweep
- Pub/Sub Notifications: Subscribe to key changes with
NOTIFY; receiveCHANGEDevents in real-time (each event is an immutable snapshot, so concurrent writes can't clobber payloads) - Atomic Transactions:
BEGIN/COMMITvalidates the whole batch first and aborts entirely on any invalid command — no partial application - Master–Replica Replication: A replica connects with
REPLICAOF, receives a point-in-time snapshot, then streams live writes - Authentication: PBKDF2-hashed passwords (fail-closed — a hashing failure never falls back to plaintext) with per-user login
- Thread-Safe I/O: Per-socket write locks prevent message corruption under concurrent notifications
- Resource Bounds & Backpressure: Configurable caps on concurrent connections, request size, and the write queue prevent unbounded memory/thread growth under load
- Memory Limits & Eviction: Cap the number of stored keys with a
noevictionorallkeys-lrupolicy - Observability: Built-in
/healthand/metricsHTTP endpoints (ops counts, key count, connections, queue depth, evictions) with no extra dependencies - Configurable: All limits and intervals load from a
vkdb.conffile — no recompiling to tune a deployment - Login Rate Limiting: Repeated failed logins lock a username for a cooldown, blunting brute-force attacks
- Java 21 or higher
- Gradle (optional — the included
./gradlewwrapper handles it)
The project builds with Gradle (a wrapper is included, so no local Gradle install is needed).
git clone https://github.com/vmskonakanchi/vkdb.git
cd vkdb
# Build the server fat jar
./gradlew :server:shadowJar
# Start server (default port 6969)
java -jar server/build/libs/vkdb-server-1.0.jar
# Or specify a port
java -jar server/build/libs/vkdb-server-1.0.jar -p 7070./gradlew :client:shadowJar
java -jar client/build/libs/vkdb-client-1.0.jarPoint a second instance at a master to replicate its writes:
java -jar server/build/libs/vkdb-server-1.0.jar -p 6970 \
-rh localhost -rp 6969 -ru <user> -rpw <pass>vkdb> REGISTER myuser mypass
REGISTERED
vkdb> LOGIN myuser mypass
LOGIN SUCCESSFUL!
vkdb> SET user1 JohnDoe
SAVED
vkdb> GET user1
JohnDoe
vkdb> SETX session abc123 60000 # expires in 60s
SAVED
vkdb> DEL user1
DELETED
vkdb> BEGIN
START
vkdb> SET count 1
SAVED TO BATCH
vkdb> SET count 2
SAVED TO BATCH
vkdb> COMMIT
COMMITTED
vkdb> DISCONNECT
BYEClient 1:
vkdb> NOTIFY user1
OKClient 2:
vkdb> SET user1 JaneSmith
SAVEDClient 1 receives:
CHANGED user1 JaneSmithThe server reads vkdb.conf (a key=value properties file) from the working
directory, or a path given with --config. Anything omitted falls back to a
built-in default; --port on the command line overrides the file. See
vkdb.conf.example for the full annotated template.
| Key | Default | Description |
|---|---|---|
port |
6969 | TCP listen port |
maxConnections |
10000 | Concurrent clients before new connections are rejected |
maxRequestBytes |
1048576 | Largest accepted single command |
writeQueueCapacity |
100000 | Bounded group-commit queue; writers block (backpressure) when full |
maxKeys |
0 | Max stored keys; 0 = unlimited |
evictionPolicy |
noeviction | noeviction (reject writes) or allkeys-lru (evict when full) |
metricsPort |
9100 | HTTP port for /health and /metrics; 0 disables it |
loginMaxAttempts |
5 | Failed logins per username before lockout; 0 disables |
loginLockoutMs |
60000 | Lockout duration / failure window |
compactionIntervalMs |
200000 | Background compaction interval |
cacheCheckIntervalMs |
10000 | TTL sweep interval |
curl http://localhost:9100/health # -> OK
curl http://localhost:9100/metrics # -> plaintext vkdb_* counters/gauges/metrics exposes uptime, role, key count, active/accepted/rejected connections,
write-queue depth, per-command op counts, evictions, and rejected-write counters
in a Prometheus-style text format.
| Command | Format | Description |
|---|---|---|
| REGISTER | REGISTER <USER> <PASS> |
Register a new user (no login required) |
| LOGIN | LOGIN <USER> <PASS> |
Authenticate |
| WHOAMI | WHOAMI |
Show current user |
| SET | SET <KEY> <VALUE> |
Store a value |
| SETX | SETX <KEY> <VALUE> <TTL> |
Store with expiry (milliseconds) |
| GET | GET <KEY> |
Retrieve a value |
| DEL | DEL <KEY> |
Delete a key |
| BEGIN | BEGIN |
Start a transaction |
| COMMIT | COMMIT |
Execute batched commands |
| NOTIFY | NOTIFY <KEY> |
Subscribe to changes on a key |
| KEYS | KEYS |
List all (non-expired) keys |
| ALL | ALL |
List all entries as key value ttl |
| REPLINFO | REPLINFO |
Show replication role and connected replicas |
| DISCONNECT | DISCONNECT |
Close connection |
Note on values: a value may contain spaces (e.g. SET greeting hello big world). For
SET the value is the rest of the line; for SETX the TTL is the final token, so the
value is everything between the key and the TTL.
flowchart TB
client([Client]) -->|"command (writeUTF)"| accept["Accept Loop"]
accept -->|"one virtual thread per client"| handler["ClientHandler"]
subgraph server["Server"]
handler -->|"GET / KEYS / ALL"| db[("ConcurrentHashMap<br/>(database)")]
handler -->|"SET / SETX / DEL<br/>append + fsync BEFORE ack"| persist["Persistence"]
persist --> db
handler -->|"on change<br/>immutable NotifyEvent"| nq(["Notification Queue"])
nq --> notify["Notify Thread"]
db -.->|"every 10s"| ttl["TTL Expiry Thread"]
persist -.->|"periodic compaction"| compact["Compaction Thread<br/>forward scan + atomic swap"]
end
notify -->|"CHANGED key value"| subscribers([Subscribed Clients])
handler -->|"forward write commands"| replicas([Connected Replicas])
aol[("append-log.vdb<br/>magic + versioned binary records")]
persist --> aol
aol -.->|"replay on startup"| db
Key design decisions:
- Virtual threads (Project Loom): One thread per client, scales to thousands without thread-pool tuning.
- ConcurrentHashMap: Lock-free reads, O(1) average operations.
- Sync-before-ack durability: Each write is appended to the AOL and fsynced before the client is acknowledged.
SAVEDtherefore means durably saved. This replaced the earlier background write queue (which could lose acknowledged-but-unwritten data on a crash). - Group commit: To avoid one fsync per write serializing all writers, writers enqueue their record (a node in a linked
BlockingQueue) and block until it is durable; a single writer thread drains everything currently queued, writes the whole batch, and issues one fsync for the batch, then releases all the waiting writers together. Each writer is still only toldSAVEDafter the fsync that covers its record — and gets an error, not a false ack, if the batch write fails. This lifts write throughput ~5× under load while keeping the same durability guarantee. - Binary AOL: Records are length-prefixed (
op, key, value, ttl) behind a magic+version header. Length-prefixing removes the delimiter fragility of the old text format, so values may contain arbitrary bytes. - Compaction forward-scan: The log is replayed forward into a map (last-write-wins, DEL removes, expired dropped), then atomically swapped in — simpler and correct for variable-length binary records.
- Immutable notification events: A
NotifyEventsnapshots key, value, and the subscriber set at publish time, so two rapid writes to the same key can't overwrite each other's payload on the queue. - Atomic transactions:
COMMITvalidates every buffered command before applying any, aborting the whole batch on the first invalid one. - Per-socket write lock: Prevents notification + response interleaving on the same connection.
- Replication: A replica sends
REPLICAOF, the master streams a point-in-time snapshot then forwards live writes; the master tracks connected replicas (surfaced viaREPLINFOand the web UI). Note: this is best-effort async fan-out — a replica disconnected mid-write is not guaranteed to catch up beyond its next full snapshot, and user accounts are not replicated.
java -cp "server/build/classes/java/test:server/build/libs/vkdb-server-1.0.jar" \
com.vkdb.benchmark.VkDbBenchmark --clients 500 --ops 1000The test suite is self-contained — the integration tests build the server jar and spawn their own server instances on ephemeral ports, so no manual setup is needed:
./gradlew :server:testThis runs unit tests (AolCodec, SaveItem, PasswordHasher) and end-to-end
integration tests (auth, CRUD, space-in-value, TTL/lazy expiry, atomic transactions,
notifications, REPLINFO, and group-commit durability across a restart).
vkdb is a from-scratch learning project, not a drop-in replacement for a mature data store. It aims to be correct and honest about its guarantees. Here is where it stands for a single-node deployment.
Handled:
- Durable writes (fsync-before-ack) with group commit
- Crash recovery via binary append-only log replay
- Atomic transactions, TTL expiry, pub/sub notifications
- Resource bounds (connections, request size, bounded write queue with backpressure)
- Memory limits with
noeviction/allkeys-lrueviction - Health/metrics endpoints and configurable limits
- PBKDF2 auth (fail-closed) with login rate limiting
Intentionally out of scope (not implemented):
- TLS / encryption in transit — traffic is plaintext, so run only on a trusted network or behind a TLS-terminating proxy. Do not expose it directly to the internet.
- Authorization model — any authenticated user can read/write/delete any key. There are no per-key ACLs or roles.
- Robust replication / HA — replication is best-effort async fan-out: a replica that misses writes while disconnected is not guaranteed to catch up beyond its next full snapshot, and user accounts are not replicated. It is not a failover-grade high-availability solution.
- Non-blocking compaction — compaction holds a lock for the duration of a full log scan, which can stall writes on a very large log.
- Horizontal write scaling — a single fsync-serialized log caps write throughput; there is no sharding.
For a trusted internal deployment with bounded load, the handled list above covers the essentials. For anything customer-facing or on an untrusted network, the out-of-scope items (especially TLS and an authorization model) would need to be addressed first.
MIT — see LICENSE.
- Inspired by Redis
- Built with Java 21 virtual threads (Project Loom)