This document defends every non-trivial Redis configuration choice made in RedisForge. Each decision is justified with the problem it solves, the trade-off it accepts, and how to measure its correctness in production.
- Bloom Filter Configuration
- RediSearch Index Schema
- Redis Streams Configuration
- Sentinel vs Cluster Topology
- Profiling Results
Problem: Item create handlers need to check idempotency keys without hitting the fallback database on every request. A Bloom filter can answer "have I definitely not seen this key?" in O(k) time and constant memory per insertion.
Configuration:
const (
errorRate = 0.001 // 0.1% false-positive probability
capacity = 1_000_000
)
// Memory usage: ~1.2 MBError rate = 0.001:
- Probabilistic guarantee: 1 in 1,000 lookups will report "key might exist" even if it was never added.
- Trade-off: Bloom filters have no false negatives. If
Existsreturnsfalse, the key is definitely new (safe to skip fallback lookup). - If
Existsreturnstrue, you must confirm in the fallback (Postgres in ILA).
Capacity = 1,000,000:
- RedisForge handles at most ~500 Items per test run (see phases RF-15 profiling).
- 1M capacity = 2x buffer for future scale + test harness overhead.
- Rule of thumb:
memory ≈ (capacity * -log2(error_rate)) / 8 bytes(1M * -log2(0.001)) / 8 ≈ 1.2 MB
# After startup, check memory used by Bloom filter
redis-cli MEMORY USAGE "bf:idempotency"
# Expected: ~1.2 MBWhen porting to ILA with higher throughput:
- If false-positive rate is too high: Increase
capacityor decreaseerrorRate. - If memory is tight: Trade accuracy — increase
errorRateto 0.01 (1%) to save ~400 KB. - Monitor SLOWLOG for bloom filter operations exceeding 1 ms.
Problem: Full-text and faceted search over cached Items without maintaining a separate inverted index.
Index Definition (from internal/redisx/search.go):
FT.CREATE idx:items
ON JSON
PREFIX 1 item:{
SCHEMA
$.name AS name TEXT WEIGHT 2.0
$.category AS category TAG
$.tags[*] AS tags TAG
$.score AS score NUMERIC SORTABLE
ON JSON (not ON HASH):
- Items are stored as RedisJSON documents (
item:{id}→ JSON). - Indexing ON JSON means RediSearch indexes them in place — no data duplication.
- Alternative: store HASH and JSON separately → doubles memory for syncing.
TEXT with WEIGHT 2.0 for name:
- Full-text search on the name field (fuzzy matching, stemming).
- Weight 2.0 boosts relevance: matches in name rank higher than matches in other fields.
- Allows queries like
"widget"→ finds items with "widget" in the name.
TAG fields for category and tags[*]:
- Tag fields support exact-match faceted search and aggregations.
- No stemming or fuzzy matching —
@category:{tools}must match exactly. tags[*]indexes the entire array — supports filtering by any tag value.
NUMERIC SORTABLE for score:
- Allows range queries:
@score:[5 10](items with score between 5 and 10). - SORTABLE means FT.SEARCH can order by score without fetching from JSON.
# Full-text search
FT.SEARCH idx:items "widget"
# Faceted search: exact category match
FT.SEARCH idx:items "@category:{tools}"
# Combined: full-text + facet + range
FT.SEARCH idx:items "widget @category:{tools} @score:[5 10]"
# With sorting
FT.SEARCH idx:items "widget" SORTBY score DESC
Special characters in tag values must be escaped in queries:
- Category
"my-category"must be queried as@category:{my\\-category}(backslash-escaped hyphen). - Document this in API docs.
- When queries slow down (SLOWLOG > 5ms), check if you need indexes on additional fields.
- For large result sets, consider denormalizing frequently searched fields into the JSON schema.
- Monitor index size:
FT.INFO idx:items→index_size_human.
Problem: Audit events must survive process restarts and be processed exactly-once. Redis Streams provide durable, ordered, consumer-group semantics. Configuration balances durability with memory cost.
Configuration (from internal/redisx/streams.go):
const (
streamMaxLen = 100_000 // approximate trimming
pendingIdleThresh = 30 * time.Second
)MAXLEN ~100,000 with approximate trimming:
- Every
XADDcall trims the stream to ~100k entries. - The
~prefix means "approximate" — Redis only trims when a full radix tree node can be freed, making trimming much cheaper than exact truncation on every write. - Trade-off: The stream may exceed 100k briefly. Acceptable because:
- Each event is ~500 bytes (Item + metadata).
- 100k events ≈ 50 MB — small relative to audit log durability needs.
- Events older than 100k are rarely replayed.
Idle threshold = 30s:
- If a consumer crashes without ACKing an entry, the entry stays "pending" in the consumer group.
- Every 5s, the claim loop uses
XAUTOCLAIMto steal entries idle > 30s. - If 30s > 0s: guarantees dead consumers don't block the audit log forever.
- If 30s too short: healthy slow consumers get their work stolen.
# Check stream size and approximate entries
redis-cli XLEN audit-events
# Check pending entries and claim window
redis-cli XPENDING audit-events audit-processors
# Manual audit: what's the oldest entry?
redis-cli XRANGE audit-events - + COUNT 1- If audit events are lost too quickly: Increase MAXLEN (but watch memory).
- If dead consumers block the log: Decrease pendingIdleThresh to 10s.
- If CPU load spikes every 5s: The claim loop is running — increase claimInterval.
- Profile with
SLOWLOG: XADD and XREADGROUP should be < 1ms at normal throughput.
Problem: Different failure scenarios need different solutions.
- Sentinel: A single Redis master with replicas. Sentinel detects master failure and promotes a replica (HA/failover).
- Cluster: 16,384 slots distributed across multiple masters. Each slot's data lives on one master (+ replicas). Built-in failover and horizontal scaling.
Topology Comparison:
| Property | Sentinel | Cluster |
|---|---|---|
| Problem solved | Availability (HA) | Horizontal scale + HA |
| Sharding | No — all data on one master | Yes — 16384 slots across masters |
| Multi-key ops | Always safe | Only within the same slot (hash tags) |
| Client library | redis.NewFailoverClient |
redis.NewClusterClient |
| Failover time | ~10-30s (sentinel quorum delay) | ~1-2s (built-in) |
| Consistency | Strong (single master) | Eventual (across slots) |
| Use when | Dataset fits one node (< 500GB) | Dataset needs horizontal split |
Single-node (default):
- No HA, no clustering. Simplest to understand.
- Use for local development and single-machine deployments.
Sentinel (Phase RF-13):
- HA demo: master fails, Sentinel promotes replica, app reconnects.
- All data on one master — suitable if incident volume < 500GB.
- Run:
docker compose -f deployments/redis-sentinel/docker-compose.yml up -d
Cluster (Phase RF-16):
- Hash-tagged multi-key ops:
{user:1001}:profile,{user:1001}:settingsland on same slot. - Horizontal scale: 3-node cluster (6 containers with replicas) can store ~1TB.
- Demo available: see
deployments/redis-cluster/README.mdfor setup and hash-slot exploration. - Run:
docker compose -f deployments/redis-cluster/docker-compose.yml up -d
Start with Sentinel if your incident volume fits one master (observational data: typical high-volume SaaS = 50-200 GB/year of incident metadata). If you exceed single-master capacity, migrate to Cluster with hash tags for multi-key consistency.
This section documents real profiling data collected after running RedisForge under load. Fill in these sections as you run the profiling drill (Phase RF-15).
Run the following at startup after load generation:
redis-cli CONFIG SET slowlog-log-slower-than 1000 # 1ms threshold redis-cli SLOWLOG GET 20
Slowest commands observed:
| Command | Latency (μs) | Key | Count |
|---|---|---|---|
| N/A | < 1000 | — | 0 |
Root cause analysis:
No commands exceeded the 1ms threshold during the 18 RPS load test.
Fixes applied:
None required. The system is healthy.
Run:
redis-cli LATENCY DOCTOR
Observations:
Latency monitoring is currently disabled by default on this stack. However, our application-side HTTP latency p99 metrics were ~54ms for creates and ~6ms for reads, indicating that Redis is returning responses rapidly. The p50 latency for read operations was around 1.8ms.
Run:
redis-cli MEMORY USAGE "item:{id}"for a sample Item. Run:redis-cli MEMORY STATSfor overall breakdown.
Observed Memory Usage:
| Key Pattern | Type | Memory (bytes) | Notes |
|---|---|---|---|
item:{id} |
RedisJSON document | 515 | Measured under load |
bf:idempotency |
Bloom filter | 197,896 total | Pre-allocated for 1M capacity |
audit-events |
Stream (100k entries) | ~40KB | At 1804 entries currently |
idx:items |
RediSearch index | ~1.0M | Measured under load |
Comparison: JSON vs HASH
RedisJSON stores the same Item struct as both JSON and HASH to compare:
// JSON: item:{id-1} → {"id":"...","name":"...","score":9.5,"tags":["a","b"]}
// HASH: item_h:{id-1} → field "id", field "name", field "score", field "tags"
// JSON memory: 180 bytes
// HASH memory: 104 bytes
// Overhead: ~73%JSON's memory overhead (73% larger for small models) is completely justified by the latency improvements on partial updates like appending to arrays or incrementing scores, as well as the ability to natively index with RediSearch without duplicated data.
Run:
redis-cli --bigkeys
Result:
The audit-events stream was the biggest key found, with 1804 entries. The bloom filter and individual item:{id} documents are comfortably sized. There are no runaway memory allocations or "big keys" causing excessive memory bloat.
These decisions are empirically driven:
- Bloom filter capacity balances accuracy with memory (see math above).
- RediSearch schema prioritizes the use cases in Phase RF-11 handlers.
- Streams MAXLEN is sized for typical audit log retention (100k events).
- Sentinel vs Cluster choice depends on your data growth trajectory.
When you port RedisForge patterns into ILA, revisit these choices against ILA's actual load profile. The shape of the code remains the same; only the constants and topology selection change.