Skip to content

Repository files navigation

vkdb

License Java

A high-performance, concurrent in-memory key-value store written in Java.

Overview

vkdb is an in-memory key-value database server inspired by Redis. It handles thousands of concurrent connections using Java virtual threads, persists data via an append-only log with background compaction, and supports real-time key-change notifications via a pub/sub mechanism.

Performance

Benchmarked on a single node (MacBook), 1000 ops per client, with durable fsync-before-ack writes and group commit enabled:

Metric 10 Clients 500 Clients
SET-only throughput 18,077 ops/sec 21,694 ops/sec
GET-only throughput 60,655 ops/sec 56,697 ops/sec
Mixed (70R/30W) throughput 36,593 ops/sec 50,445 ops/sec
SET p50 / p99 latency 0.52 ms / 0.88 ms 23.3 ms / 33.8 ms
GET p50 / p99 latency 0.16 ms / 0.23 ms 8.72 ms / 15.37 ms
Errors 0 0

500 concurrent clients × 1000 ops each = 500,000 operations completed with zero errors in every workload.

Writes stay durable — every SET/SETX/DEL is fsynced before it is acknowledged — but group commit batches all writes waiting at a given moment into a single fsync, so throughput scales with concurrency instead of being capped at one fsync per write. Under 500 concurrent writers, SET throughput is ~5× higher than the naive one-fsync-per-write approach (4.3k → 21.7k ops/sec) and mixed throughput ~3.7× higher (13.6k → 50.4k), while the durability guarantee is unchanged. Reads never fsync and stay in the ~57k ops/sec range.

Key Features

  • Concurrent Connections: Thousands of simultaneous clients via Java virtual threads
  • Durable Writes (sync-before-ack): Each write is appended to the log and fsynced before the server replies, so an acknowledged write survives an immediate crash
  • Group Commit: Concurrent writes waiting at the same moment are batched into a single fsync, so write throughput scales with load without weakening durability
  • Binary Append-Only Log: Length-prefixed binary records with a versioned file header; values may contain any bytes (spaces, =, newlines) without corrupting the log
  • Background Compaction: Periodic log compaction collapses overwrites and drops deleted/expired entries
  • Key Expiration (TTL): Time-to-live on keys, enforced both lazily on read and by a background sweep
  • Pub/Sub Notifications: Subscribe to key changes with NOTIFY; receive CHANGED events in real-time (each event is an immutable snapshot, so concurrent writes can't clobber payloads)
  • Atomic Transactions: BEGIN/COMMIT validates the whole batch first and aborts entirely on any invalid command — no partial application
  • Master–Replica Replication: A replica connects with REPLICAOF, receives a point-in-time snapshot, then streams live writes
  • Authentication: PBKDF2-hashed passwords (fail-closed — a hashing failure never falls back to plaintext) with per-user login
  • Thread-Safe I/O: Per-socket write locks prevent message corruption under concurrent notifications
  • Resource Bounds & Backpressure: Configurable caps on concurrent connections, request size, and the write queue prevent unbounded memory/thread growth under load
  • Memory Limits & Eviction: Cap the number of stored keys with a noeviction or allkeys-lru policy
  • Observability: Built-in /health and /metrics HTTP endpoints (ops counts, key count, connections, queue depth, evictions) with no extra dependencies
  • Configurable: All limits and intervals load from a vkdb.conf file — no recompiling to tune a deployment
  • Login Rate Limiting: Repeated failed logins lock a username for a cooldown, blunting brute-force attacks

Quick Start

Prerequisites

  • Java 21 or higher
  • Gradle (optional — the included ./gradlew wrapper handles it)

Build & Run

The project builds with Gradle (a wrapper is included, so no local Gradle install is needed).

git clone https://github.com/vmskonakanchi/vkdb.git
cd vkdb

# Build the server fat jar
./gradlew :server:shadowJar

# Start server (default port 6969)
java -jar server/build/libs/vkdb-server-1.0.jar

# Or specify a port
java -jar server/build/libs/vkdb-server-1.0.jar -p 7070

Client

./gradlew :client:shadowJar
java -jar client/build/libs/vkdb-client-1.0.jar

Replica

Point a second instance at a master to replicate its writes:

java -jar server/build/libs/vkdb-server-1.0.jar -p 6970 \
  -rh localhost -rp 6969 -ru <user> -rpw <pass>

Usage

vkdb> REGISTER myuser mypass
REGISTERED
vkdb> LOGIN myuser mypass
LOGIN SUCCESSFUL!
vkdb> SET user1 JohnDoe
SAVED
vkdb> GET user1
JohnDoe
vkdb> SETX session abc123 60000    # expires in 60s
SAVED
vkdb> DEL user1
DELETED
vkdb> BEGIN
START
vkdb> SET count 1
SAVED TO BATCH
vkdb> SET count 2
SAVED TO BATCH
vkdb> COMMIT
COMMITTED
vkdb> DISCONNECT
BYE

Notifications

Client 1:

vkdb> NOTIFY user1
OK

Client 2:

vkdb> SET user1 JaneSmith
SAVED

Client 1 receives:

CHANGED user1 JaneSmith

Configuration

The server reads vkdb.conf (a key=value properties file) from the working directory, or a path given with --config. Anything omitted falls back to a built-in default; --port on the command line overrides the file. See vkdb.conf.example for the full annotated template.

Key Default Description
port 6969 TCP listen port
maxConnections 10000 Concurrent clients before new connections are rejected
maxRequestBytes 1048576 Largest accepted single command
writeQueueCapacity 100000 Bounded group-commit queue; writers block (backpressure) when full
maxKeys 0 Max stored keys; 0 = unlimited
evictionPolicy noeviction noeviction (reject writes) or allkeys-lru (evict when full)
metricsPort 9100 HTTP port for /health and /metrics; 0 disables it
loginMaxAttempts 5 Failed logins per username before lockout; 0 disables
loginLockoutMs 60000 Lockout duration / failure window
compactionIntervalMs 200000 Background compaction interval
cacheCheckIntervalMs 10000 TTL sweep interval

Observability

curl http://localhost:9100/health    # -> OK
curl http://localhost:9100/metrics   # -> plaintext vkdb_* counters/gauges

/metrics exposes uptime, role, key count, active/accepted/rejected connections, write-queue depth, per-command op counts, evictions, and rejected-write counters in a Prometheus-style text format.

Command Reference

Command Format Description
REGISTER REGISTER <USER> <PASS> Register a new user (no login required)
LOGIN LOGIN <USER> <PASS> Authenticate
WHOAMI WHOAMI Show current user
SET SET <KEY> <VALUE> Store a value
SETX SETX <KEY> <VALUE> <TTL> Store with expiry (milliseconds)
GET GET <KEY> Retrieve a value
DEL DEL <KEY> Delete a key
BEGIN BEGIN Start a transaction
COMMIT COMMIT Execute batched commands
NOTIFY NOTIFY <KEY> Subscribe to changes on a key
KEYS KEYS List all (non-expired) keys
ALL ALL List all entries as key value ttl
REPLINFO REPLINFO Show replication role and connected replicas
DISCONNECT DISCONNECT Close connection

Note on values: a value may contain spaces (e.g. SET greeting hello big world). For SET the value is the rest of the line; for SETX the TTL is the final token, so the value is everything between the key and the TTL.

Architecture

flowchart TB
    client([Client]) -->|"command (writeUTF)"| accept["Accept Loop"]
    accept -->|"one virtual thread per client"| handler["ClientHandler"]

    subgraph server["Server"]
        handler -->|"GET / KEYS / ALL"| db[("ConcurrentHashMap<br/>(database)")]
        handler -->|"SET / SETX / DEL<br/>append + fsync BEFORE ack"| persist["Persistence"]
        persist --> db
        handler -->|"on change<br/>immutable NotifyEvent"| nq(["Notification Queue"])
        nq --> notify["Notify Thread"]

        db -.->|"every 10s"| ttl["TTL Expiry Thread"]
        persist -.->|"periodic compaction"| compact["Compaction Thread<br/>forward scan + atomic swap"]
    end

    notify -->|"CHANGED key value"| subscribers([Subscribed Clients])
    handler -->|"forward write commands"| replicas([Connected Replicas])

    aol[("append-log.vdb<br/>magic + versioned binary records")]
    persist --> aol
    aol -.->|"replay on startup"| db
Loading

Key design decisions:

  • Virtual threads (Project Loom): One thread per client, scales to thousands without thread-pool tuning.
  • ConcurrentHashMap: Lock-free reads, O(1) average operations.
  • Sync-before-ack durability: Each write is appended to the AOL and fsynced before the client is acknowledged. SAVED therefore means durably saved. This replaced the earlier background write queue (which could lose acknowledged-but-unwritten data on a crash).
  • Group commit: To avoid one fsync per write serializing all writers, writers enqueue their record (a node in a linked BlockingQueue) and block until it is durable; a single writer thread drains everything currently queued, writes the whole batch, and issues one fsync for the batch, then releases all the waiting writers together. Each writer is still only told SAVED after the fsync that covers its record — and gets an error, not a false ack, if the batch write fails. This lifts write throughput ~5× under load while keeping the same durability guarantee.
  • Binary AOL: Records are length-prefixed (op, key, value, ttl) behind a magic+version header. Length-prefixing removes the delimiter fragility of the old text format, so values may contain arbitrary bytes.
  • Compaction forward-scan: The log is replayed forward into a map (last-write-wins, DEL removes, expired dropped), then atomically swapped in — simpler and correct for variable-length binary records.
  • Immutable notification events: A NotifyEvent snapshots key, value, and the subscriber set at publish time, so two rapid writes to the same key can't overwrite each other's payload on the queue.
  • Atomic transactions: COMMIT validates every buffered command before applying any, aborting the whole batch on the first invalid one.
  • Per-socket write lock: Prevents notification + response interleaving on the same connection.
  • Replication: A replica sends REPLICAOF, the master streams a point-in-time snapshot then forwards live writes; the master tracks connected replicas (surfaced via REPLINFO and the web UI). Note: this is best-effort async fan-out — a replica disconnected mid-write is not guaranteed to catch up beyond its next full snapshot, and user accounts are not replicated.

Running Benchmarks

java -cp "server/build/classes/java/test:server/build/libs/vkdb-server-1.0.jar" \
  com.vkdb.benchmark.VkDbBenchmark --clients 500 --ops 1000

Running Tests

The test suite is self-contained — the integration tests build the server jar and spawn their own server instances on ephemeral ports, so no manual setup is needed:

./gradlew :server:test

This runs unit tests (AolCodec, SaveItem, PasswordHasher) and end-to-end integration tests (auth, CRUD, space-in-value, TTL/lazy expiry, atomic transactions, notifications, REPLINFO, and group-commit durability across a restart).

Production Readiness & Scope

vkdb is a from-scratch learning project, not a drop-in replacement for a mature data store. It aims to be correct and honest about its guarantees. Here is where it stands for a single-node deployment.

Handled:

  • Durable writes (fsync-before-ack) with group commit
  • Crash recovery via binary append-only log replay
  • Atomic transactions, TTL expiry, pub/sub notifications
  • Resource bounds (connections, request size, bounded write queue with backpressure)
  • Memory limits with noeviction / allkeys-lru eviction
  • Health/metrics endpoints and configurable limits
  • PBKDF2 auth (fail-closed) with login rate limiting

Intentionally out of scope (not implemented):

  • TLS / encryption in transit — traffic is plaintext, so run only on a trusted network or behind a TLS-terminating proxy. Do not expose it directly to the internet.
  • Authorization model — any authenticated user can read/write/delete any key. There are no per-key ACLs or roles.
  • Robust replication / HA — replication is best-effort async fan-out: a replica that misses writes while disconnected is not guaranteed to catch up beyond its next full snapshot, and user accounts are not replicated. It is not a failover-grade high-availability solution.
  • Non-blocking compaction — compaction holds a lock for the duration of a full log scan, which can stall writes on a very large log.
  • Horizontal write scaling — a single fsync-serialized log caps write throughput; there is no sharding.

For a trusted internal deployment with bounded load, the handled list above covers the essentials. For anything customer-facing or on an untrusted network, the out-of-scope items (especially TLS and an authorization model) would need to be addressed first.

License

MIT — see LICENSE.

Acknowledgements

  • Inspired by Redis
  • Built with Java 21 virtual threads (Project Loom)

About

A lightweight, Java-based, open-source key-value store with notification support.

Topics

Resources

Contributing

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages