Skip to content

feat: add Redis-backed editor data storage for high availability - #46

Open
nvanlaerebeke wants to merge 42 commits into
Euro-Office:mainfrom
nvanlaerebeke:feature/redis-editor-data-storage
Open

nvanlaerebeke wants to merge 42 commits into
Euro-Office:mainfrom
nvanlaerebeke:feature/redis-editor-data-storage

Conversation

@nvanlaerebeke

Copy link
Copy Markdown

Summary

This pull request adds Redis-backed storage for the Euro Office Co-Authoring service to support high-availability deployments.

The shared Redis state allows multiple server instances to participate in the same editing session. It covers the coordination state required by Co-Authoring, including:

  • Active users and connection presence.
  • Document and editing locks.
  • Co-authoring state and messaging, including chat-related state.
  • Force-save coordination.
  • Connection and usage statistics.
  • Graceful shutdown coordination.

The backend supports both standalone Redis and Redis Cluster. Redis is used for shared state and does not replace the existing RabbitMQ pub/sub or task queue mechanisms.

Redis storage is not enabled by default. It can be selected with:

{
  "services": {
    "CoAuthoring": {
      "server": {
        "editorDataStorage": "editorDataRedis"
      }
    }
  }
}

Testing

The implementation has been tested with Redis and is working as expected.

The test coverage includes Co-Authoring presence, locking, messaging, force-save behavior, statistics, expiration, timeout recovery, health checks, and loading the backend through the packaged configuration.

Local validation completed successfully:

  • Redis test suite: 5 tests passed.
  • ESLint passed.
  • Prettier passed.
  • git diff --check passed.

Implementation note

The initial implementation was largely AI-generated. It has been reviewed, cleaned up for this pull request, and re-tested; everything currently appears to be working as expected.

This is submitted as a candidate implementation for review. If the maintainers do not want to merge it in its current form, it may still serve as a useful working base for implementing Redis-backed high availability in the future.

@moodyjmz

Copy link
Copy Markdown
Member

Hi, this is an interesting PR and converges with #44. I feel 44 has more depth, but this is wider in feature set. They also use differing libraries, node-redis and ioredis. Being as your submission is AI, I would take the opportunity to run AI to compare both PRs. My questions are around depth, features and libs. I will update here once I have done that.

@moodyjmz moodyjmz added the enhancement New feature or request label Sep 21, 2026
@nvanlaerebeke

Copy link
Copy Markdown
Author

FYI, the reason I didn't use ioredis was that the node-redis is I think the "more" official one.

Updating that package to 5.x also gives support for redis sentinel etc and I thought it wasn't productive to have multiple packages that basically do the same.

I'll see if I can bring this PR closer to #44

@moodyjmz

Copy link
Copy Markdown
Member

Hi @nvanlaerebeke I would agree having multiple packages doing the same is not productive. I think I would also agree with you about a shift to node-redis and would be more of the mind to change 44 to use that than vice versa.

@moodyjmz

Copy link
Copy Markdown
Member

@nvanlaerebeke Would it be useful for you if I run a comparison and share it? (44 and 46 - they point at different things, but might be useful and I can burn some tokens so you don't have to)

@moodyjmz

Copy link
Copy Markdown
Member

@nvanlaerebeke @MonaAghili — comparison of #46 against #44, read at cf0d7200. There are three independent efforts on this now, and the technical content below is the part that matters regardless of which base is chosen.

Convergent

Four non-obvious editorDataMemory semantics, each a trap a straight port falls into, all handled correctly in #46:

  • _checkAndLock is re-entrant with TTL refresh. A plain SET NX denies the current holder, and the client re-locks on every save cycle, breaking roughly every second save. The LOCK script handles it.
  • lockNotification relies on NaN !== NaN for non-re-entrancy. Routed through a shared owner-comparing script the token serialises to "NaN", which equals itself, and the throttle stops throttling. SET NX PX is correct.
  • addForceSaveTimerNX is first-write-wins. A plain ZADD slides the autosave deadline forward under continuous editing so it never fires. ZADD NX is correct.
  • Unlock returns the numeric c_oAscUnlockRes enum and callers branch on the number. tonumber is correct.

A clean-room derivation from editorDataMemory produced the same four independently; #44 implements two and delegates two. Also convergent: Redis with Lua for compare-and-set, one hash tag per document, encoded key components, shutdown key used verbatim without re-prefixing or tenant-scoping.

Divergent

Surface. #46 implements the full surface including EditorStat. #44 Redis-backs ten EditorData methods and delegates seventeen. Neither has per-method failure-path coverage or a two-process suite; #46's tests are integration-shaped over the full surface, #44's unit-shaped over ten methods.

#44's partial state is not merely incomplete, it is harmful: isSaveLock (DocsCoServer.js:3700, :3705-3714) compares a client's document-wide change index against the force-save record's index and denies the save lock until the connection's unsynced timestamp is expire.saveLock old, warns at :3716, resets at :3725. With force-save state per-replica and locks shared, every replica compares against a record another replica armed. Sharing part of the surface is worse than sharing none.

Test infrastructure. redisEditorDataTests.yml stands up a real cluster in CI — a redis:7-alpine service for standalone, plus six nodes in run steps with --cluster-enabled yes and redis-cli --cluster create --cluster-replicas 1. #44 has no cluster tier; none of the twelve workflows on its branch mentions redis or a cluster. verify.js is client-agnostic at the EditorData API and portable to either implementation.

#44's redis suite does run in CI inside the existing unit job against an ephemeral server, and #44 has additionally been exercised against a real multi-node cluster under sustained multi-replica load outside CI. A cluster CI job establishes that the code runs correctly on a cluster; a field run establishes that it survives one. Both are needed.

Failure policy. #46 throws from every method with no call site changed (16 files, none of DocsCoServer.js, canvasservice.js, gc.js). One path interacts badly: a throw inside commandSfcCallback after the file is assembled leaves the document row at SaveVersion/UpdateVersion, createSaveTimer will not re-arm because its mask requires Ok (DocsCoServer.js:1558), and receiveTask acks regardless (canvasservice.js:2088-2092). What re-drives a row left in those states is not established. #44 returns decided values — false from lock(), LOCKED from unlock() — matching callers written for values rather than exceptions. Neither branch justifies the per-method return at the call site, which is the open gap.

Findings

1. Both cross-document indexes are on one slot. {editor:index}:documents and {editor:index}:forcesavetimer (editorData.js:40-41) both hash to slot 1073. Every document in a deployment writes index entries to one master; losing that master stalls expiry and force-save deployment-wide. Throughput is not the issue — one ZADD per connection per documentsCron tick and a single-key pop. The issue is that the key cannot be scaled out and its failure domain is the whole deployment. #44 shards the equivalent sixteen ways over a hash of (tenant, docId), keyed on both so a single-tenant deployment does not collapse into one shard. Transferable regardless of client library.

2. The command timeout outlives the transport. base.js:61-62 sets 30 s command, 15 s connect. The socket dies at around 25 s (comment at the saveChanges lock site in DocsCoServer.js), so a stalled command outlasts the connection it belongs to.

3. Sentinel deployments run standalone silently. The orchestrated entrypoint emits sentinel configuration only under iooptions (build/scripts/orchestrated/docker-entrypoint.sh:132-138, DocumentServer repo) — sentinels, group name, sentinelPassword. #46 removes that block and never reads it, so a sentinel-configured operator gets a standalone connection and no error.

The reverse applies to #44: the standalone entrypoint writes credentials only under redis.options (entrypoint.sh:211-213, as username/password/database) and #44 never reads options, so it attempts to connect without credentials — on a protected server that fails AUTH and fails closed. Both are config plumbing. normalizeNodeOptions is the better of the two answers and transfers as-is.

4. disableOfflineQueue is never set, so it sits at node-redis's default of on: during a reconnect after a socket drop, commands queue in memory rather than failing fast, and a caller that should have been told "no" waits. #44 sets the ioredis equivalent off deliberately. On a cluster it must go in defaults rather than rootNodes, since per-node connections do not inherit rootNodes configuration.

Destroying the client on timeout is correct and #44 has no equivalent. It is the only mechanism that turns a black-holed connection into a reconnect: when a peer dies without TCP teardown — SIGKILL, an idle drop in a NAT or load balancer, a partition — there is no FIN and no RST, the socket reads as open, the reconnect strategy never runs, and detection falls to TCP keepalive minutes later. Without _abortClient() every later command is written to the same dead socket and fails at the deadline for that window. With it, the next command rebuilds inside one connect budget, and on a cluster re-discovers slots by the same path. Its cost is collateral rejection of concurrent in-flight commands when the connection was healthy and one command was slow; the lever for that is the budget (finding 2) and the hot index (finding 1), not the mechanism.

ioredis's commandTimeout only rejects the promise — Command.js:191-197, setTimeout → this.reject(new Error("Command timed out")) — and touches nothing on the connection. #44 sets no reconnectOnError. So #44 currently has the half-open exposure with no recovery, and existing field measurements could not have surfaced it because no node died during them. On migration #44 needs this mechanism.

5. POP_EXPIRED (scripts.js:95-101) removes entries before the reply reaches the caller, so a lost reply is a claimed-then-discarded batch. #44's sweep has the same defect in a different form; it should be fixed once.

Client library

node-redis, per Redis's own recommendation for new projects.

Version: 6.x, with one decision to make explicitly. Current is 6.2.1; the 5 line is at 5.12.1; 4.7.1 is tagged maintenance-v4. 4.7.0 has no sentinel — it exports createClient and createCluster only — and sentinel is a supported topology, so createSentinel in 5 is the floor. 6.0 makes RESP3 the default protocol, which needs an explicit RESP: 2 or a measured decision rather than silent inheritance; an equivalent ioredis 6 bump was deferred for the same reason. 6.2 also fans out cluster multi-key commands across slots transparently, so a cross-slot key layout stops announcing itself — slot assertions in tests should check the slot directly rather than infer safety from a command succeeding.

withTimeout is not redundant on 6.x. commandOptions.timeout applies only to commands not yet written to the socket. Once a command is on the wire it has no per-command deadline. socketTimeout is not a substitute: it is an idle timeout, fires on a healthy idle connection, and the default reconnect strategy then closes the client permanently. Measured on 5.12.1 — timeout: 300, in-flight GET under CLIENT PAUSE 1500, resolved at 1599 ms. The queued window is covered natively; the in-flight window only by a caller-side race. The 30 s budget is the part worth revisiting (finding 2); while _abortClient() is on that path, the budget is also the blast radius.

Cluster options belong in defaults, not rootNodes. Per-node connections do not inherit rootNodes config; credentials, TLS and timeouts must be in defaults to reach every connection. ioredis hardcodes the equivalent, node-redis leaves it configurable.

Command resend. node-redis does not replay a command already written — on socket close, sent commands reject rather than re-running — so the "was the lock actually taken?" hazard is closed on standalone and cluster without an option, where ioredis needed one.

The exception is sentinel. RedisSentinel.execute re-runs a master command that rejected while the client is not ready, up to maxCommandRediscovers (default 16), and a command written to a socket that then closes rejects exactly that way. The logic is unchanged between 5.12.1 and 6.2.1. 6.2.1 does fix the adjacent case — a lost master no longer holds the caller indefinitely, rejecting after the same bound — which is a further argument for 6. maxCommandRediscovers: 0 disables the re-send, at the cost of failing in-flight master commands during a failover rather than retrying; on a lock path that is the wanted behaviour.

This lands harder on #46's surface than #44's. A re-run is harmless where the operation is idempotent, and most of #44's ten are — a re-granted re-entrant lock, a repeated presence write. On the full surface it is not: incrEditorConnectionsCountByShard double-counts, feeding licence counting, and addMessage RPUSHes a duplicate. This is from reading the client source, not running it; neither branch has a sentinel test rig, and #46's cluster job is the pattern for building one.

Evidence. Existing field measurements were taken against ioredis on a real cluster and do not describe a node-redis build. Equivalents are now identified for the offline queue, the resend behaviour and the timeout, which scopes re-measurement to MOVED following and the in-flight window rather than the whole failure contract.

Cluster readiness has no client-side signal: isOpen is true from the start of connect(), before slot discovery, so a health check needs its own flag. (Keep the existing 'error' listeners; without one the emitter throws and the uncaughtException handler exits the process.)

Config surface. #46 removes the iooptions* block, #44 uses it, and an entrypoint change in the DocumentServer repo (PR #363) emits one of the two shapes. Whichever lands, that PR has to match it.

Open, not delivered

  • A call-site audit exists and is not yet public: twelve places where a wrong answer from this store causes an irreversible action — deleted change history, a stub force-save record overwriting a real one, a client told a lock is released while it stands. Derived from editorDataMemory and its consumers, so independent of both implementations, and it is what a per-method failure policy should be justified against.
  • The definition of done is spread across several documents. One requirement neither branch meets: a suite exercising presence, locks, messages, force-save and the mutexes across at least two processes sharing one Redis, with the in-memory backend as a differential oracle. verify.js is close in shape but runs in-process.
  • Fail-open versus fail-closed per method is undecided and blocks finishing the failure policy on either branch.

Neither implementation has been read line by line beyond the areas above, and #46's cluster job has not been run here, so the findings are for verification rather than acceptance.

@MonaAghili

Copy link
Copy Markdown

Independent deep-dive on #46 against editorDataMemory.js as the behavioral baseline, cross-checked with @moodyjmz's comparison above (read at cf0d7200) and a second, independently-derived review pass. Reviewed at acc43970 (current head). Read-only analysis, no code changes.

Where this overlaps with the comment above, I've marked it; the rest is new.

Summary

The hard behavioral work is largely right: lock re-entrancy, the numeric unlock enum, SET NX PX for lockNotification, ZADD NX first-write-wins for the force-save timer, and the presence version-token protocol are all implemented correctly and hold up under the standard race traces (two replicas racing a lock, renew-then-release, crash-then-delayed-unlock). None of that is in doubt.

But five separate things block merge, three of which aren't in the discussion yet:

Blockers

1. Issue #267 itself is not demonstrably fixed. The new module is split into six files under DocService/sources/editorDataRedis/ (base.js, core.js, editorData.js, editorStat.js, index.js, scripts.js), but DocService/package.json's pkg.scripts (and the matching lists in FileConverter/package.json and AdminPanel/server/package.json) still name only editorDataRedis.js. None of the five new files are listed, and no CI job builds or runs a pkg binary with editorDataStorage: "editorDataRedis". Issue #267 is a MODULE_NOT_FOUND in a pkg binary — this PR fixes the source-tree gap but leaves the packaging gap completely untested. tests/redis/editorDataRedis.pkg-smoke.js requires from the source tree and never touches a built binary despite its name. #44, for comparison, lists all 8 new files in pkg.scripts explicitly.

2. deleteKey is implemented against a key this module doesn't own.

// editorStat.js:228-230
EditorStat.prototype.deleteKey = async function (key) {
  await this._command(['DEL', key]);
};

editorDataMemory's deleteKey is a documented no-op. This implementation does a real, unconditional DEL on whatever key the caller passes, unprefixed, in whatever logical DB REDIS_SERVER_DB_KEYS_NUM names — a database this codebase has no ownership of and no way to know the layout of. Call sites pass bare document ids (DocsCoServer.js:1531, :2191) or a task-result mask key (canvasservice.js:495). If that shared logical DB is ever used by another component whose keys happen to collide with a document id, this silently deletes someone else's data with no audit trail. It exists only because contract.tests.js:37 asserts method-name parity with editorDataMemory, which forces Redis to grow some implementation to pass the test — even though the safe answer here is to not implement it at all, matching memory's no-op.

3. Force-save Lua scripts cjson.decode/cjson.encode the whole record on every state transition.

scripts.js:157  local ok, value = pcall(cjson.decode, raw)
scripts.js:164  local updated = cjson.encode(value)

This round-trips the entire changeInfo/convertInfo payload through Lua's JSON codec on every force-save transition. Three concrete corruptions follow: an empty Lua table is ambiguous, so any empty array nested inside changeInfo/convertInfo gets re-encoded as an empty object instead of an array; numbers get truncated at cjson's 14 significant digits; and START_FORCE_SAVE_SCRIPT additionally does value.convertInfo = nil unconditionally on every start, dropping the field entirely rather than preserving it the way editorDataMemory.js:208-217 does. That last one is currently masked because today's call sites happen to already have convertInfo == null at that point, but it's a landmine for the next caller that doesn't. Keeping this payload opaque (store it as an untouched string field, only comparing/writing the small scalar fields in Lua) avoids all three failure modes at once.

4. Sentinel support is silently inert. The orchestrated DocumentServer entrypoint writes sentinel settings under iooptions (docker-entrypoint.sh:132-138). This PR removes iooptions from config and reads a new, differently-shaped optionsSentinel block instead (base.js:71). A sentinel-configured deployment gets a plain standalone connection with no warning at any log level — it looks healthy until the master fails over, at which point every write starts rejecting with READONLY.

5. cleanDocumentOnExit can leave a document's full state behind for up to 7 days, and removeLocks can silently fail to release a lock.

  • cleanDocumentOnExit (scripts.js:282-296) is a no-op — deletes nothing, not even locks/messages/saved-status — if any live presence entry remains (HLEN(presence:hash) > 0), unlike editorDataMemory.js:234-243's unconditional delete. A ghost presence entry from a dead replica is enough to trigger this, and the document's index entry may already be consumed by that point, making it invisible to gc going forward.
  • removeLocks (scripts.js:129-136) deletes by comparing the stored value byte-for-byte against the caller's JSON-stringified copy, not by lock id like editorDataMemory.js does. If a lock is mutated between read and remove (e.g. _recalcLockArray rewriting lock.block.* on a spreadsheet row/column insert, DocsCoServer.js:3595-3603), the comparison fails silently and the lock is never released — the client is told it's free while Redis still holds it.

Also worth flagging (not blocking, but real)

  • POP_EXPIRED_SCRIPT (scripts.js:95-101) has no batch cap — after an outage/backlog it can pop tens of thousands of expired members in one EVAL, blocking the single-threaded server for the duration. A LIMIT-bounded pop (gc's next tick collects the remainder) closes this.
  • Shard staleness silently reuses the presence-expiry setting instead of its own config knob — an operator tuning presence TTL for its stated purpose unknowingly also changes how fast a crashed replica's connection counts decay from /info. They happen to share the same default today, but that's coincidence, not design.
  • addPresence/updatePresence write the global index (ZADD) as a second, non-atomic round trip after the main script — a crash in that window leaves live presence untracked by the index.
  • Client/topology: this PR swaps in node-redis 6 with Cluster/Sentinel support — a real scope and dependency change (pulls in @redis/search, @redis/bloom, @redis/json, @redis/time-series) that the maintainer thread is already discussing. Worth an explicit decision recorded somewhere before merge, since it's the fork everything else (finding 4 above included) hangs off of.
  • 3DPARTY.md wasn't updated for the redis 4.7.0 → 6.2.1 swap and doesn't list the new @redis/* sub-dependencies it pulls in.
  • Installed node_modules/redis in this checkout is still 4.7.0 against a 6.2.1 manifest/shrinkwrap — redis.createSentinel doesn't exist at 4.7.0, so the sentinel path can't load locally; worth confirming what version any "tests passed" run actually executed against.

Test gaps

  • contract.tests.js only checks method name/arity parity across backends, never behavior — this is literally what forces deleteKey to exist. A real shared behavioral contract would have caught the deleteKey, removeLocks, cleanDocumentOnExit, and convertInfo issues above.
  • No multi-process test exists anywhere in either PR — verify.js's concurrency races run several clients in one Node process, which doesn't reproduce two independent DocService replicas (independent crash behavior, independent SHARD_IDs, independent gc timers).
  • No real-server sentinel test at all (only mocked option-normalizer unit tests) — this is what let the sentinel config mismatch through.
  • CI runs redis:7-alpine only; worth adding a Valkey leg given that's the actual deployment target, and cjson's empty-table behavior is build-dependent.

Bottom line

Competent and the most complete of the candidates — full EditorData/EditorStat surface, no call-site changes, real Redis Cluster CI job, and it gets the non-obvious editorDataMemory semantics right where a naive port wouldn't. But it doesn't yet close the issue it's filed against, implements a method against a key it doesn't own, uses a serialization pattern that can silently corrupt force-save state, and has two silent behavioral regressions (cleanDocumentOnExit, removeLocks) in exactly the paths that matter for correctness. Needs another pass before merge — the packaging/deleteKey/cjson issues are all straightforward fixes; the client-library/topology question needs an explicit ruling first, since it changes the shape of a good chunk of the rest.

@nvanlaerebeke

Copy link
Copy Markdown
Author

I've pulled in the #44 yesterday and applied the changes from this pull request on it.

Still have to review it and I'll take the posted comments into account as well, that way it starts from the #44 pull request as a base, that will be for this evening though.

@moodyjmz

Copy link
Copy Markdown
Member

@nvanlaerebeke @MonaAghili I will run a check against Mona's deep dive - the referenced issue is not actually in this repo. It is here: Euro-Office/DocumentServer#267

@nvanlaerebeke thanks for reacting so quickly to replies here, really appreciate it

@moodyjmz

Copy link
Copy Markdown
Member

@nvanlaerebeke @MonaAghili — both reviews verified against acc43970. Claims in each do not survive at that head, including most of the client-library section of my earlier comment, which was read at cf0d7200.

If you are working on this tonight

You said you had pulled #44 and were applying this on top. Four things that are worth having before you start, rather than after.

Sign off while you rebase. DCO is failing on all four commits and it is the only completed check on this PR. git rebase --signoff folds it into the pass you are already doing; found later it costs a second history rewrite.

Client library is settled: node-redis. #44 moves to it. You should not end up carrying both clients through the rebase.

Do not spend time on these — each is something one of the two reviews above could reasonably send you doing, and none of them is needed:

  • listing the six files under editorDataRedis/ in pkg.scripts (pkg traces the shim; verified by execution)
  • making deleteKey a no-op (it is the only reason its proxy exists; it is a decision about the env var, not a defect)
  • reverting cleanDocumentOnExit to an unconditional delete (the HLEN guard is correct multi-replica behaviour)
  • removing withTimeout because 6.x has a native timeout (it covers queued commands only)

CI is now released. Checks on this PR had never run — they were held behind the fork-approval gate. Both are approved and executing against acc43970 as of this comment, so results will appear above shortly. Anything red there is real, not a gating artefact.

If you only get to a few things: removeLocks (defect 1) and the sentinel config shape (defect 2) are the two that are wrong rather than incomplete.

TL;DR

  1. acc43970 predates my earlier comment by ~14 hours. The PR was already on node-redis 6.2.1 with createSentinel; that comment's version section is void. Corrections below.
  2. No CI had ever executed for this PR — every check on all four commits sat action_required behind the fork gate. Now approved and running against acc43970. The one check that had completed is DCO, failing: no Signed-off-by on any commit.
  3. Suite now run here: 41/41 standalone, 41/41 on a real six-node cluster built from this PR's workflow recipe. Node 20.20.2, node-redis 6.2.1, Redis 8.10.2.
  4. The six files under editorDataRedis/ do not need listing in pkg.scripts.
  5. Client library: node-redis. Settled; the rebase proceeds on that basis.
  6. Six defects below. Fail-open versus fail-closed is open, with options set out rather than ruled on.

Corrections — my earlier comment

Claim Status at acc43970
"4.7.0 has no sentinel; the floor is node-redis 5" Void. Already 6.2.1 with createSentinel. Only the RESP pin remains.
F2 "the socket dies at around 25 s" Wrong. The 25 s comments (DocsCoServer.js:3333, :3537) are the browser sockjs connection, not the Redis transport. Constants (30 s command, 15 s connect) correct; 30 s now also overrides node-redis 6.2.1's 5000 ms default on all topologies.
F3 sentinel Substance stands, description obsolete. No sentinel code existed at cf0d7200; at acc43970 it exists with a config shape that does not match the entrypoint. The review above is current.
"Cluster readiness has no client-side signal" Addressed at acc43970 — isReady awaited on all topologies.
"#44 lists 8 files, #46 lists one" 7 files, and the comparison is void — pkg traces #46's shim.
"#46's cluster job has not been run here" Now run.

Unchanged: slot 1073 (re-confirmed by CLUSTER KEYSLOT on the live cluster), disableOfflineQueue unset, POP_EXPIRED, sentinel re-send bounded at maxCommandRediscovers 16, withTimeout not redundant on 6.x (commands-queue.js:409-436 removes the timeout listener at socket-write time).

Corrections — the review above

B1, pkg.scripts. Listing the six files is unnecessary. Verified by execution with @yao-pkg/pkg 6.22.0 (the packer e2e.yml:92 uses): a shim in pkg.scripts has its literal require('./editorDataRedis/index.js') traced transitively and every file included; a control with a non-literal require fails MODULE_NOT_FOUND.

#267's stack fails in Common/sources/notificationService.js:36 — a dynamic require crossing from Common into DocService/sources/ — and Common/package.json has no pkg block. That is a separate bug, independent of both PRs and of which backend ships, and warrants its own issue against DocumentServer.

Neither PR is scoped to #267. A complete, tested Redis backend answers it; the packaging gap is fixed on its own terms. What belongs in this PR is the demonstration: boot a packaged DocService with editorDataStorage: "editorDataRedis" in the existing E2E pkg step.

(A bare #267 in this repository resolves to Euro-Office/server#267, which does not exist; it needs the DocumentServer/ prefix.)

W4, dependency growth — inverted. origin/main already resolves @redis/search, @redis/bloom, @redis/json, @redis/time-series and @redis/graph via its existing redis@4.7.0. 6.2.1 drops graph; this PR drops ioredis. The tree shrinks.

B3, cjson — three mechanisms reproduce, none has a live trigger. Empty-array-to-object and 16-digit truncation reproduced directly, with identical output on Redis 8.10, redis:7-alpine and Valkey 8.1 — not build-dependent, which also removes the stated rationale for a Valkey CI leg. changeInfo contains no arrays, InputCommand has no array fields, time is 13 digits. The convertInfo = nil divergence (scripts.js:163 vs editorDataMemory.js:208-217) is real and worth fixing as a latent trap.

B2, deleteKey — facts confirmed, severity overstated. editorStat.js:228-230 performs a real DEL where memory's is a no-op; five call sites, not three. But editorStatProxy is constructed only when REDIS_SERVER_DB_KEYS_NUM is set (DocsCoServer.js:156-159), deleteKey is the only method invoked on it, every call is gated by preStopFlag or follows a cache purge, and no shipped entrypoint sets that variable. The proxy exists to perform that delete. This is a decision on whether to keep the hook — a no-op makes the env var dead config — not a correctness defect.

B5a, cleanDocumentOnExit — divergence confirmed, trigger overstated. The script sweeps expired presence (scripts.js:283-289) before the HLEN guard, and the guard is correct multi-replica behaviour: an unconditional delete would drop another live replica's locks. The orphan state requires an index score older than a live presence score, via the non-atomic index write or clock skew. The 7-day bound is correct.

W6, installed 4.7.0 — local to that checkout; the shrinkwrap pins 6.2.1 and the workflow runs npm ci. The underlying point, that no run had been verified, was correct and is now answered.

Confirmed as stated: B5b removeLocks, B4 sentinel config shape, the POP_EXPIRED batch cap, the shard-staleness TTL reuse, the non-atomic index write.

Contradictions between the two reviews

Topic Resolution
pkg.scripts Do not list the six files; pkg traces the shim.
Sentinel The review above is current; my "no sentinel support" described the older head.
withTimeout Keep it. 6.x's native commandOptions.timeout covers queued commands only.
cleanDocumentOnExit Do not revert to an unconditional delete; fix the orphan path.

Remaining work

Defects:

  1. removeLocks — compare by lock id (HDEL), not by serialised value. _recalcLockArray mutates locks between read and remove and removeUserLocks publishes releaseLock regardless, so a client is told a lock is released while Redis holds it.
  2. Sentinel config shape — either this PR reads iooptions or the DocumentServer entrypoint emits optionsSentinel. One side moves, and a sentinel-configured deployment that lands standalone must fail loudly.
  3. DCO sign-offs on all commits.
  4. POP_EXPIRED — bound the batch and stop losing entries on a lost reply; one fix.
  5. Force-save Lua — keep the payload opaque, compare and write scalar fields only, stop clearing convertInfo on start.
  6. 3DPARTY.md:49 still records 4.7.0.

Not demonstrated rather than broken: #267 end-to-end (boot a packaged DocService with editorDataStorage: "editorDataRedis"); a real-server sentinel test; a multi-process differential suite. None of the three is covered by the jobs now running.

Client library

node-redis. Redis recommends it for new projects, two packages doing the same job is not worth carrying, and this PR is already there. #44 moves. The rebase should not end up carrying both clients.

Open under it: node-redis 6 defaults to RESP3 (DEFAULT_RESP = 3). This PR runs RESP3 unpinned and passes 41/41, but RESP3 against a cluster is unmeasured, so an explicit RESP: 2 is the conservative setting until it is.

Open: fail-open versus fail-closed

EditorData has 29 methods and a Redis failure means something different in each. editorDataMemory never fails, so the contract is silent on it. Three options:

A — throw from every method (this PR today). Uniform, no call-site changes, and safe where a plausible-looking value is dangerous: getdelSaved returning null on a network error causes canvasservice.js:1257-1258 to treat it as success and delete the document's change history. Against: wrong for methods whose callers were written for values, and it produces the commandSfcCallback path where a row is left at SaveVersion/UpdateVersion with createSaveTimer unable to re-arm.

B — return decided values (#44, for its ten). Matches existing call-site expectations; denial is safe for the lock methods and a delayed save costs little. Against: for readers there is often no safe value — getLocks returning {} is indistinguishable from "no locks", and sendAuthInfo (DocsCoServer.js:3402-3405) then tells every client the document is unlocked.

C — per-method. Most work, and the answer space is three-valued: deny (lock methods); throw or return a distinct sentinel (getdelSaved, getLocks); return a specific enum member (_checkAndUnlock must return Locked, since callers branch on Locked !== res and undefined takes the wrong branch); return an explicit unknown (presence, which #44 marks so callers can distinguish it from empty).

#44's existing policy covers the ten methods it implements, so that pattern is safe for those during the rebase. The methods beyond them are where this bites.

Views welcome, including on whether per-method is worth it for a first enablement or whether uniform throw is an acceptable start with per-method to follow.

@moodyjmz

Copy link
Copy Markdown
Member

@j-base64 just to add you here

@moodyjmz

Copy link
Copy Markdown
Member

CI results on acc43970 — the first run this PR has had.

Job Result
E2E (overlay on nightly), including the pkg binary build green
Redis editor data tests — standalone green
Redis Cluster editor data tests red, in setup

The cluster failure is the harness, not the code. The nodes start correctly via docker run redis:7-alpine, but the readiness loop and redis-cli --cluster create are invoked on the runner host, and ubuntu-latest does not ship redis-cli:

line 13: redis-cli: command not found

Add sudo apt-get install -y redis-tools before that step, or route the calls through docker exec redis-7000 redis-cli …. The suite itself passes on a six-node cluster once redis-cli is present — that is where the 41/41 figure above came from.

This job had never executed before today, so it is not a regression; the fork gate meant nothing had surfaced the missing tool.

On the pkg question: the Build server binaries step is green against the restructured six-file module, which supports the point that the files under editorDataRedis/ do not need listing. It builds with the default config, so it still does not boot with editorDataStorage: "editorDataRedis" — that demonstration remains outstanding.

Amendment to my first comment: I described this cluster job as better than anything on our side. The design still is, and the finding does not change that — it had just never run.

DCO remains the only other red, and git rebase --signoff covers it in the pass you are already making.

@moodyjmz

Copy link
Copy Markdown
Member

Merged main in to keep this current — one file, AdminPanel/server/sources/routes/fonts/router.js, nothing near the redis work. If it is in the way of your rebase, force-push straight over it; the merge commit is expendable.

Re-approved the checks against the new head.

@nvanlaerebeke
nvanlaerebeke force-pushed the feature/redis-editor-data-storage branch from 1f9f803 to 913378a Compare September 23, 2026 17:43
nvanlaerebeke and others added 4 commits September 23, 2026 19:50
Assisted-by: OpenAI Codex
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
…rage implementation

Assisted-by: OpenAI Codex
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Assisted-by: OpenAI Codex
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Align the header with the rest of the codebase, which is AGPL version 3
only and attributes copyright to Euro-Office contributors.

Assisted-by: ClaudeCode:claude-opus-5
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Julius Knorr <jus@bitgrid.net>
@nvanlaerebeke
nvanlaerebeke force-pushed the feature/redis-editor-data-storage branch from 913378a to ee22902 Compare September 23, 2026 17:51
Assisted-by: OpenAI Codex
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
@nvanlaerebeke

nvanlaerebeke commented Sep 23, 2026 •

Copy link
Copy Markdown
Author

I suggest using the following configuration structure to support:

  • Single Redis
  • Redis Cluster
  • Redis Sentinel

The structure follows the current redis package API and the normalization logic in server/DocService/sources/editorDataRedis/base.js.

{
  "services": {
    "CoAuthoring": {
      "redis": {
        "name": "redis",
        "prefix": "ds:",

        "host": "redis.example.com",
        "port": 6379,

        "options": {
          "username": "redis-user",
          "password": "redis-password",
          "database": 0,
          "socket": {
            "tls": true,
            "rejectUnauthorized": true
          },
          "commandOptions": {
            "timeout": 30000
          }
        },

        "optionsCluster": {
          "rootNodes": [
            {
              "url": "redis://redis-cluster-0.example.com:6379"
            },
            {
              "url": "redis://redis-cluster-1.example.com:6379"
            },
            {
              "url": "redis://redis-cluster-2.example.com:6379"
            }
          ],
          "defaults": {
            "username": "redis-user",
            "password": "redis-password",
            "socket": {
              "tls": true,
              "rejectUnauthorized": true
            }
          },
          "commandOptions": {
            "timeout": 30000
          }
        },

        "optionsSentinel": {
          "name": "mymaster",
          "sentinelRootNodes": [
            {
              "host": "sentinel-a.example.com",
              "port": 26379
            },
            {
              "host": "sentinel-b.example.com",
              "port": 26379
            }
          ],
          "database": 0,
          "nodeClientOptions": {
            "username": "redis-user",
            "password": "redis-password",
            "socket": {
              "tls": true,
              "rejectUnauthorized": true
            }
          },
          "sentinelClientOptions": {
            "username": "sentinel-user",
            "password": "sentinel-password",
            "socket": {
              "tls": true,
              "rejectUnauthorized": true
            }
          },
          "commandOptions": {
            "timeout": 30000
          }
        }
      }
    }
  }
}

The example shows all supported topology blocks for reference. In an actual deployment, only one topology should be configured; the other blocks should remain empty.

Deployment Configuration
Single Redis host, port, and options
Redis Cluster optionsCluster.rootNodes and optionsCluster.defaults
Redis Sentinel optionsSentinel.name, sentinelRootNodes, nodeClientOptions, and sentinelClientOptions

Topology selection

The current implementation selects the topology as follows:

  1. If optionsSentinel is non-empty, Sentinel mode is selected.
  2. Otherwise, if optionsCluster.rootNodes is non-empty, Cluster mode is selected.
  3. Otherwise, standalone Redis is selected.
  4. If both Cluster and Sentinel are configured, the implementation throws an error.

Configuration details

  • options contains standalone Redis connection settings.
  • optionsCluster.rootNodes contains the Redis Cluster discovery nodes.
  • optionsCluster.defaults contains settings inherited by discovered cluster nodes, such as credentials and TLS.
  • Redis Cluster does not support a non-zero logical database.
  • optionsSentinel.name is the Sentinel master/group name.
  • optionsSentinel.sentinelRootNodes contains the Sentinel hosts and ports.
  • nodeClientOptions contains credentials and TLS settings for Redis master/replica connections.
  • sentinelClientOptions contains credentials and TLS settings for Sentinel connections.
  • optionsSentinel.database selects the logical database for Redis master/replica connections.
  • commandOptions belongs at the topology level and controls command behavior such as timeouts.
  • prefix is used by the editor-data adapter when constructing Redis keys.
  • name identifies the Redis connector and must currently be set to "redis".

Legacy configuration

iooptions is a legacy configuration namespace from the previous ioredis-based implementation.

The older/intermediate fields options.sentinels and sentinelPassword are not used by the current implementation.
The current implementation expects the native optionsSentinel structure described above.

The default configuration is defined in server/Common/config/default.json.

The normalization logic is implemented in server/DocService/sources/editorDataRedis/base.js.

@moodyjmz

Copy link
Copy Markdown
Member

Correction to my comment above, on the mechanism rather than the conclusion.

I wrote that with force-save state per replica, "every replica compares against a record another replica armed". That is backwards. Each replica compares the client's document-wide change index against its own force-save record, and only saves handled by that replica ever advance it — so a participant's replica arms a record early, every later save goes through a different replica, and the local record never moves. getForceSave returns the local value; resetForceSaveAfterChanges on that same replica is what armed it.

The conclusion is unchanged: sharing the locks while leaving force-save per-replica is worse than sharing neither. The deployment that reported this had the mechanism right before we did.

Nothing in this PR is affected — the surface is already complete, which is the right answer under either reading.

@j-base64

Copy link
Copy Markdown

Really nice work @nvanlaerebeke 🙌, and thanks for pushing this forward. Having the full surface plus standalone/cluster/sentinel in one place looks like an interesting base to build on.

With that in mind, a few things that might be worth a look:

  • Failover double-grant. A lock granted on a primary can be granted again on a promoted replica, because the promotion can lag the write that took the lock. We reproduced this against a real primary/replica pair. WAIT doesn't close it: in a run where the key was in fact lost on promotion it still returned 1 (reported success), and on the shipped replica-less default it returns 0 in under a second, so treating 0 as "unsafe, deny" would deny every save on a stock single-node install. There's no safe threshold to act on. This applies to any Redis-backed backend here, feat: add Redis-backed editor data storage for high availability #46 included, and it's a different mechanism from the sentinel maxCommandRediscovers resend already discussed above. Might be worth recording explicitly as a known, out-of-scope-for-now, pre-production gate: a first enablement can ship without it, it just shouldn't be described as failover-safe.

  • A fail-open vs fail-closed data point, in this PR's favor. In commandSfcCallback, getdelSaved is read at canvasservice.js:1258 as null == savedVal || '1' === savedVal. That's a loose ==, so both null and undefined satisfy it, and a truthy requestRes routes into cleanDocumentOnExit(..., deleteChanges=true), which deletes the document's changes (sqlBase.deleteChanges there also isn't awaited). A backend that fails open by returning absence on a Redis error would hit this and drop recovery changes, whereas this PR's throw propagates and never reaches it. So the uniform-throw policy is actually the safe choice on this specific path; worth keeping in mind when weighing per-method return values for the reader methods.

  • The two-process differential gap, made concrete. Both @moodyjmz and @MonaAghili flagged the missing multi-process test; confirming it from the diff, it's absent at the code level. There's no child_process / fork / worker_threads / spawn anywhere in the PR, so the cross-replica tests (e.g. the presence "visible from another replica" case) run two store instances inside one Node process. That catches basic shared visibility, but not independent crash, gc/cron timers, or shard identity, which is where the remaining cross-replica edges live: for example the orphan window where a crash between the main write and the non-atomic index update can leave state that cleanDocumentOnExit's presence guard then holds onto. contract.tests.js also checks method name/arity parity only (.length equality), not behavior, so a cross-replica behavioral regression wouldn't be caught there either. The lock-id fix that just landed is the kind of edge careful reading catches one at a time; a two-process differential suite (two independent processes on one shared Redis, with the in-memory backend as the oracle) would cover that whole class in one place, and being client- and implementation-agnostic it would outlive whichever base ends up chosen, so it might be worth tracking as its own follow-up. Concrete shape: kill one replica mid-operation and assert the other's view of presence and locks.

Keen to see where this goes 😊

@nvanlaerebeke

nvanlaerebeke commented Sep 24, 2026 •

Copy link
Copy Markdown
Author

Thank you everyone for the positive feedback and for taking the time to review the implementation in such detail.
The reviews helped me understand a lot of concerns that I would have missed on my own as my own knowledge doesn't go that in depth.

I’m still learning this part of the codebase, so I’ll try to take this as far as I can with my current knowledge. I also appreciate any further guidance to get this merged.

I'm pushing euro office at my work place as a replacement for office for the web.

Regarding the Redis configuration I previously suggested, the DocumentServer entrypoints expect the iooptions/ "old" configuration structure.

What would be the preferred way to handle this?

  1. Should I open a corresponding pull request in the main DocumentServer repository to update the entrypoint shell scripts?
  2. Should I leave the entrypoint unchanged for now and continue with the current PR until the implementation is further along?
  3. Alternatively, should the server implementation temporarily support the existing entrypoint format for compatibility?

I'm not a fan of "3", redis is currently not part of the codebase at all so it's better to start without any "legacy" as there should be no deployments with that configuration as those won't start due to the missing storage provider.

Use the redis-7000 container for Redis Cluster readiness checks and
cluster creation, avoiding a host-side redis-cli dependency.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Bound POP_EXPIRED batches and lease claimed entries until the GC worker
acknowledges successful processing. Unacknowledged claims are reclaimed after
the lease expires, preventing lost Redis responses or worker crashes from
dropping document-presence and force-save work.

Keep all claim keys in the existing Redis Cluster hash slot and add regression
coverage for batching, timeouts, lost responses, partial processing, and stale
acknowledgements.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Keep force-save payloads unchanged while Redis Lua updates only the
small state fields. This prevents empty arrays becoming objects, large
numbers being rounded, and convertInfo being cleared unexpectedly.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Replace the yes pipeline with redis-cli supported --cluster-yes option.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Build and run the packaged DocService binary with the Redis editor-data
backend, verifying module resolution, health, and cleanup in CI.

Also extract shared Redis test helpers for context creation, delays, and
public method discovery.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Apply Prettier formatting fixes required by the server format check.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Share compatible Redis clients through an explicit connection manager,
preserve isolation for separate databases and topologies, and ensure
connections drain and close exactly once during shutdown.

As a consequence of the ownership refactor, clean up the Redis storage
implementation by splitting configuration, transport, lifecycle, codecs,
keys, settings, and EditorCommon responsibilities into focused modules.
Remove the obsolete base.js façade and update callers and tests.

Add lifecycle, module-boundary, and shutdown coverage.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
@nvanlaerebeke

Copy link
Copy Markdown
Author

Took me a bit to get it into a state I liked, I've opened up a PR (Euro-Office/DocumentServer#386) on the DocumentServer project with the changes for the sentinel configuration:

Euro-Office/DocumentServer#386

It also includes easy to run commands to validate the setups for both Redis and Valkey in:

  • Standalone
  • Cluster
  • Sentinel

So at the time of writing I've been able to test each topology from a locally build image both with redis and valkey.

I'll go over the remaining items next

Validate the connected client reference before dispatching Redis commands
so concurrent aborts return RedisUnavailableError instead of TypeError.
Add deterministic regression coverage for the connection-abort race.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Add authenticated and unauthenticated Redis configuration tests, failover
integration coverage for Cluster and Sentinel, and failure-policy tests for
connection loss during document cleanup.

Make dedicated topology CI jobs fail when their required failover configuration
is missing, and ensure unauthenticated tests remove inherited credentials.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Replace unsafe Redis GETDEL handling with durable claim/read/ack semantics.
Preserve claims through successful cleanup, recover abandoned claims during
terminal document cleanup, and fail closed on unknown Redis outcomes.

Add focused tests for retries, concurrency, cleanup races, lost responses,
and canvasservice failure handling.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Default optionsSentinel to an empty object when absent for compatibility
with older DocumentServer configurations.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
@moodyjmz

Copy link
Copy Markdown
Member

Thanks for the connection-layer work and the failure-handling tests. One regression in the connection layer needs fixing:

A single failed connect disables Redis for the life of the process. At 85c8fb3c, the catch in _connect() sets this.closed = true (redisConnection.js:231). _withOperation() (:348-350) rejects any call while closed is true, before it tries to connect, and closed is only reset in _createClient(), which is only reached through _withOperation(). RedisConnectionManager.acquire() keeps returning the same object. From then on every store sharing that connection (editorData, editorStat, info, notificationService) throws RedisUnavailableError until the process restarts.

Fix: keep closed for deliberate shutdown (close(), SIGTERM). On a failed connect, drop the client and let the next operation connect again.

Details
  • Came in with 6b817086 ("refactor(redis): centralize editorDataRedis connection ownership"). base.js before it had no closed flag.
  • Repro (standalone): docker stop Redis, one call (ECONNREFUSED after ~240 ms), docker start. Every later call fails with Redis connection is closed, with no connect attempt. A fresh RedisConnection in the same process gets PONG.
  • It also latches at startup: if editorData.connect() fails once, the error is logged and nothing retries (DocsCoServer.js:4316-4323).
  • Current tests don't catch it. The connect-failure cases in connection.tests.js stop after the failure. The reconnect cases in editorDataRedis.tests.js / expiration.tests.js pause for less than the 15 s connect timeout, so their reconnect succeeds. failure.integration.tests.js stops Redis and never restarts it.
  • A test that fails at 85c8fb3c and passes with L231 removed (needs a Redis on the test port):
test('a failed connect() does not prevent the next command() from connecting', async () => {
  const connection = new RedisConnection();
  connection.connector = 'redis';
  connection.cluster = false;
  connection.client = fakeClient({
    isOpen: false,
    isReady: false,
    connect: async () => {
      throw new Error('Connection timeout');
    }
  });

  await assert.rejects(connection.connect(), /Connection timeout/);
  try {
    assert.equal(await connection.command(['PING']), 'PONG');
  } finally {
    await connection.close();
  }
});

fakeClient is an EventEmitter with isOpen/isReady set and the given properties merged in.

@moodyjmz

Copy link
Copy Markdown
Member

Proposal: land #46 as a stack of smaller PRs rather than one. It is currently +7,533 lines across 56 files, and at that size a problem in one area holds up everything else. Feedback welcome, from the author and from anyone else working on this code, before anything gets cut.

A stack keeps one branch to test: the top one carries everything below it, so anyone building #46 today would build that branch instead.

Rough shape, bottom first:

  1. Connection core, failure-handling tests, CI
  2. Locks
  3. Document state
  4. Force-save
  5. Presence and the shared sweep
  6. Shutdown handling, enablement, docs

Open questions:

  • Is this the right cut? Different boundaries, or fewer PRs, are fine if they keep each PR reviewable on its own.
  • Where the branches live. Each PR's base must be a branch in Euro-Office/server, so the lower branches have to exist here, not only on a fork.
Details
  • Before cutting: agree the per-method failure policy (which methods throw and which return a default when Redis is unavailable; getdelSaved must throw), and freeze the EditorCommon call surface (_eval/_command/_commands) and its error contract. Every PR above the first calls that seam. Changing it later touches all of them.
  • One module per method group. Right now all groups sit in editorData.js and scripts.js, so every restack would conflict across sibling PRs.
  • Outside the stack: the editorDataMemory.js, canvasservice.js and savedState.js changes in their own PR, landed first.
  • Mechanics: restack with git rebase --update-refs from the top branch and push --force-with-lease. No squash merges; delete each branch on merge so GitHub retargets the next PR.
  • Review order: the first PR gets reviewed to merge. The rest stay as drafts until the one below them merges.

@MonaAghili

MonaAghili commented Sep 29, 2026 •

Copy link
Copy Markdown

I re-checked 85c8fb3 against the discussion above and only list what hasn't been raised yet. Items marked reproduced were run against node-redis 6.2.1 (the pinned version) with this PR's own modules, using a local TCP proxy for outages and latency and a small fake Sentinel; the rest were traced in the code. Entrypoint and dev-stack points are on DocumentServer #386.

Not raised yet

1. Regression on the default in-memory backend: every final save is treated as failed (reproduced)
setSaved() is only written for a non-'1' status, so on a normal save nothing is stored and the in-memory getdelSaved() returns undefined (editorDataMemory.js#L194-L199). consumeSavedState() only treats null as "nothing saved" (savedState.js#L12-L21); the old code used null == savedVal. I ran the real commandSfcCallback with editorDataMemory, mocking only the database, storage and HTTP. After the integrator answered {error: 0}, the file was copied to forgotten-files storage, the task status was restored to SaveVersion, and updateVersion went out with success: false. This hits every deployment on the default configuration.
Suggested fix: savedValue == null (or make the memory backend return null), plus a test that uses the real memory backend instead of mocking getdelSaved.

2. SIGTERM, SIGINT and uncaughtException now hang while any editor is connected (reproduced)
shutdownRedis() awaits server.close() (server.js#L499-L552), which never calls back while socket.io connections are open; I tested with socket.io 4.8.1 and DocService's socket.io configuration. uncaughtException used to exit with code 1 immediately. Now the process stops accepting connections but keeps running until the last editor leaves. This hits every deployment.
Suggested fix: close the socket.io connections and/or cap the wait with a timeout, and keep the immediate exit for uncaughtException.

3. Sentinel mode runs commands one at a time (reproduced)
node-redis 6.2.1's Sentinel client leases its single master client (masterPoolSize 1) for each operation, so concurrent commands run serially (redisConfig.js#L98-L139). With 50 ms of latency, 200 concurrent PINGs took 10.2 s, against 52 ms with reserveClient: true. Through the adapter, 1,415 of 2,000 concurrent pings hit the 30 s timeout; standalone finished all 2,000 in 0.4 s.
Suggested fix: reserveClient: true (or a larger pool), plus a concurrency test.

4. Lost presence entries are never rebuilt, so live editors can vanish
UPDATE_PRESENCE_SCRIPT only refreshes an entry that still exists (scripts.js#L38-L47), so a connected user whose entry disappears is never re-added. An entry can disappear through:

  • a Redis restart or data loss;
  • a refresh gap longer than the TTL;
  • the cross-node reconnect race in REMOVE_PRESENCE_SCRIPT, which deletes by user id with no connectionId check (#L60-L64).

The next editor who leaves then triggers the final save and deleteChanges while others are still editing. GC also treats the document as expired, and publish() can skip pub/sub for remote users. Traced in the code, not reproduced.
Suggested fix: re-add missing entries on refresh (expireDoc has the connection info), and only remove an entry when its connectionId matches.

5. The Helm chart can't configure Sentinel for this PR
Adding to Blocker 2 in my earlier review, whose entrypoint half #386 addresses. The Kubernetes-Docs chart only emits REDIS_SENTINEL_* when redisConnectorName is ioredis, and this PR rejects that name with Unsupported Redis connector (redisConnection.js#L130-L133). The chart also joins the node lists with spaces, which #386 now splits on commas only (details there). This needs a companion Kubernetes-Docs change.

6. Two additions to @moodyjmz's startup-latch point

  • server.listen() only runs after that connect succeeds (DocsCoServer.js#L4316-L4322), so the process stays up but never listens. Configuration errors, such as setting both Cluster and Sentinel or a name other than redis, end the same way.
  • Correction to my earlier review, where I called maxCommandRediscovers: 0 correct: node-redis uses the same limit for its Sentinel topology-discovery loop (redisConfig.js#L136). A single transient answer, such as s_down mid-failover, therefore fails connect() with no retry. I reproduced this with a fake Sentinel; node-redis's default setting connected after 1.2 s. Once the latch is fixed, the initial connect still needs its own retry.
Should also be fixed (medium)
  • documentsCron vs fixed TTLs. Presence and shard counters are only refreshed once per documentsCron step (DocsCoServer.js#L167), which admins can change in AdminPanel. expire.presence and expire.shard are fixed at 300 s and blocked by the config schema, so a cron of 5 minutes or more makes live documents look expired. Validate the pair.
  • One slow command now fails everything in the process (reproduced). Destroying the client on timeout is the right cure for half-open sockets, as discussed above. But since 6b81708 all stores share one client, so a 30 s timeout on any single command (redisConnection.js#L388-L397) fails every in-flight operation across editorData, editorStat, info and notifications.
  • GC starvation. checkDocumentExpire handles the claimed batch inside one try (gc.js#L146-L174). With the at-least-once claims, a document that always throws comes back in the same position and blocks the documents behind it on every reclaim. Catch errors per item.
  • Abandoned saved-state claims. The claim added in 2579a95 has no TTL (scripts.js#L181-L208) and relies on the same operation being retried. As noted in my earlier review, though, receiveTask acks in finally (canvasservice.js#L2113), so a task that throws is never redelivered. An abandoned claim then makes every later final save of that document throw SavedStateUnknownError. getdelSaved() without an operation id also creates a claim nobody can acknowledge.
  • CI follow-ups.
    • The workflow's path filter (redisEditorDataTests.yml#L7-L29) misses the files changed since the last round: savedState.js, canvasservice.js, DocsCoServer.js, server.js, editorDataMemory.js and routes/info.js.
    • The unit workflow never runs tests/redis; it only selects tests/unit.
    • The Cluster failover test is a graceful CLUSTER FAILOVER, and the Sentinel test waits for promotion before sending anything.
    • The pkg smoke test covers standalone only.
Low / cleanup
  • The force-save "no reset while active" guard exists only in the Redis backend, so the memory backend still allows the duplicate start. On Redis, a rejected request returns UnknownError rather than "in progress".
  • The 30 s command and 15 s connect timeouts can't be tuned: they apply regardless of commandOptions.timeout and socket.connectTimeout.
  • Sentinel node clients inherit redis.options credentials, but Cluster defaults don't. redis.options.url is silently dropped under Sentinel.
  • The proxy store with a non-zero database on Cluster only fails when first used, at pre-stop.
  • Lock and unlock errors are swallowed without logging. Logs print host/port for Cluster and Sentinel, and every reconnect attempt logs a full stack at error level.
  • AdminPanel creates a Redis-backed EditorStat through routes/info but doesn't declare redis.
  • server.js now requires the Redis connection manager at the top, so memory-only deployments load node-redis too.
  • editorStat.js ends with an Ascensio copyright block below the Euro-Office SPDX header, and the editorDataRedis.js shim has no SPDX header.
  • README:
    • it says GC cleanup runs "every 2 seconds (up to 50 members/second)", but the default is every 120 s (about 0.8 per second);
    • the key list omits saved:claim, and the claim and lease keys have no TTL, contrary to "refreshed or expired by the backend";
    • the optionsSentinel structure from your configuration comment isn't in the README yet.

@moodyjmz

Copy link
Copy Markdown
Member

Thanks @MonaAghili. Three additions, and a question:

#4 presence: reproduced. Against node-redis 6.2.1 and a real redis-server at fc7c2201 (the presence scripts are unchanged at 85c8fb3c): two users present, FLUSHALL, a heartbeat for both. getPresence then returns 0 entries, and the document's index score is not re-created.

The latch feeds #4. A latched replica's heartbeats fail, so after expire.presence (300 s) its users' presence entries expire. Other replicas' checkDocumentExpire (gc.js:125-184) can then treat those documents as abandoned while users are still connected to the latched replica. That's traced in the code, not reproduced. Fixing the latch doesn't close #4; it needs its own fix.

#2: unhandled rejections take the same path. server.js has no unhandledRejection handler, and Node's default mode raises an unhandled rejection "as an uncaught exception" (docs). So any unhandled promise rejection now ends in the same hang, not exit(1).

On the stack proposal: your findings map onto it. Sentinel (#3, #5) goes in the connection core, presence (#4) in the presence slice, shutdown (#2) in the top slice, and the saved-state regression (#1) in the separate PR that lands first. @MonaAghili, does that cut work for how you'd review this and for the work on your side? @nvanlaerebeke, you know the code best: would you split it this way, or draw the lines differently?

@MonaAghili

Copy link
Copy Markdown

Thanks @moodyjmz. Both additions check out on my side: the presence scripts are identical between fc7c220 and 85c8fb3, and an unhandled rejection does land in the uncaughtException handler (checked on Node 20; it's Node's default since version 15), so the fix for item 2 has to cover that path as well.

The stack works for me, your placement of items 1 to 5 matches how I'd review it, and I'm happy to go slice by slice. For the rest of my list:

  • Separate PR, first: a test that runs commandSfcCallback against the real in-memory backend rather than a mocked getdelSaved, since that's the default configuration and the path item 1 breaks.
  • Connection core: my item 6 additions, which build on the startup latch, and the shared-client timeout.
  • Document state: the saved-state claim's missing TTL.
  • Presence and the shared sweep: the documentsCron vs TTL check, and per-item error handling in checkDocumentExpire.
  • Top slice: the README fixes; item 2 there should also cover unhandled rejections.

Item 5 also needs a companion Kubernetes-Docs change. Each slice should extend the Redis workflow's path filter to the files it touches, and tests/redis should run in the unit workflow too.

@j-base64

Copy link
Copy Markdown

@nvanlaerebeke @moodyjmz @MonaAghili, regarding the "Packaging check" point above:
I'm looking into whether it's already being handled somewhere, or whether to take it as its own new issue/PR, since it was raised as general rather than specific to this PR.

The end goal would be that the build fails when a pkg.scripts entry matches no file. That is exactly what confused us here: the module was missing but the build never warned us, because an unmatched pkg.scripts glob passes silently.

Does that sound right to you?

@moodyjmz

moodyjmz commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Thanks for the Redis test suite; all 124 existing tests pass against the sketch below unchanged.

Suggestion: restructure #46 along logical boundaries, so each piece can be reviewed and reused on its own. Much of it sits in a few large units: presence, locks, saved state, force-save and the expiry queues all live on EditorData, and the lock logic alone is spread across editorCommon.js, editorData.js and scripts.js. Pulling out pieces with one job each (per-concern classes, a shared transaction primitive, a reusable expiry-claim queue) keeps EditorData's public interface identical, gives each piece its own tests, and turns review into small separate reads. Worth doing early, before any split: each piece then maps onto its own PR, whereas doing it later means reshaping code that's already spread across PRs. Worked example for locks below. What do you think, @nvanlaerebeke @MonaAghili?

Details

Worked example: LockStore, on 85c8fb3c:

class LockStore {
  // ...

  constructor(ops, docBase) {
    this.ops = ops; // {eval, command, transaction}
    this.docBase = docBase;
  }

  lockSave(ctx, docId, userId, ttl) {
    return this.#lock(this.#key(ctx, docId, 'savelock'), userId, ttl);
  }

  async addLocks(ctx, docId, locks) {
    const fields = argsFromObject(locks);
    if (fields.length === 0) return;
    const key = this.#key(ctx, docId, 'locks');
    await this.ops.transaction(key, [['HSET', key, ...fields], ['EXPIRE', key, ttlSeconds(ctx, LOCKS_TTL_PATH, cfgExpLocks)]]);
  }

  async addLocksNX(ctx, docId, locks) {
    // ...HSETNX per field, EXPIRE, HGETALL in one routed MULTI;
    // a field whose HSETNX returned 0 is a conflict.
  }

  async removeLocks(ctx, docId, locks) {
    const lockIds = Object.keys(locks);
    if (lockIds.length > 0) await this.ops.command(['HDEL', this.#key(ctx, docId, 'locks'), ...lockIds]);
  }

  #key(ctx, docId, name) {
    return `${this.docBase(ctx, docId)}${name}`;
  }

  async #lock(key, fencingToken, ttl) {
    // unchanged LOCK_SCRIPT compare-and-set, fail-closed on error
  }

  // ...
}

EditorData constructs it with new LockStore(this._redisOps(), (ctx, docId) => this._docBase(ctx, docId)) and delegates the nine lock methods.

85c8fb3c Sketch
Lua scripts 27 24
Existing Redis tests, standalone and single-node Valkey cluster 124 pass 124 pass
MULTI/EXEC reaching the cluster during the suite (INFO commandstats) 0 13
New memory-vs-Redis lock behaviour test passes passes

Difference found by the behaviour test: the memory backend's getLocks/addLocksNX return its live internal object; Redis returns a copy.

Reusable pieces:

  • A transaction(routingKey, commands) primitive, atomic on every topology (below). Every group that batches writes can use it.
  • The expiry-claim queue: _popExpired/_ackExpired already take their keys as parameters and serve both documents and force-save timers, so they're close to a class of their own.
  • The narrow {eval, command, transaction} surface itself: one seam every unit depends on, testable with a fake.

Units with one job each:

  • Presence: add/update/remove/get, removal preparation, document index.
  • Saved state: setSaved/getdelSaved/ackSaved and the claim.
  • Force-save record and timer.
  • EditorStat: unique users, monthly tallies, connection samples, shard counts.
  • redisConnection.js: connection lifecycle separate from per-topology command routing.

How they map onto the proposed stack:

  • Connection core: the {eval, command, transaction} surface, transaction(), lifecycle vs routing in redisConnection.js.
  • Locks: LockStore.
  • Document state: saved state.
  • Force-save: the force-save record and timer.
  • Presence and the shared sweep: presence plus the expiry-claim queue.
  • EditorStat groups: their own slice.

Lua that could be plain commands. Ten of the 27 scripts are unconditional batches of writes, or map onto an existing command (HSETNX, multi-field HDEL). What blocks them today: commands() uses MULTI on standalone and Sentinel, but on Cluster it's a Promise.all of separate commands, so it isn't atomic there. A transaction(routingKey, commands) built on node-redis's routed multi(routingKey) is atomic on all three topologies (the sketch adds one, about 20 lines). Besides the three lock scripts: ADD_PRESENCE, REMOVE_PRESENCE, ADD_MESSAGE, STORE_FORCE_SAVE, ADD_UNIQUE_USER, SET_CONNECTION_SAMPLE, SET_SHARD_COUNT.

The 17 that need Lua (compare-and-set, read-modify-write, prune-then-read) could be registered with the client via defineScript, so node-redis sends EVALSHA and falls back on NOSCRIPT, instead of sending the full body on every call (eval()). It also gives each script a name and typed arguments.

Syntax. ES classes with #private helpers suit new modules and the repo already uses classes, but that's optional; the separation is the point.

Repro: sketch applied on 85c8fb3c; Valkey 8.1 standalone and single-node cluster; npm run tests -- --runInBand ../tests/redis.

@moodyjmz

moodyjmz commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

@j-base64 sounds right as its own issue; pkg does drop unmatched patterns silently, so a check that fails on an entry matching nothing would have caught it.

@nvanlaerebeke

Copy link
Copy Markdown
Author

@moodyjmz, that sounds like a good idea.

Splitting this PR into multiple smaller units should make it easier to review, it started out much smaller and more localized that it is now.

I’d like to finish a few remaining open items and address the points raised by @MonaAghili before cleaning up the code and splitting it up a bit further.

Once that is done and it's in a state I'm comfortable with, I’ll split the implementation into smaller, focused PRs. Your suggested boundaries make sense and I’ll work toward that.

Distribute presence-expiry and force-save indexes across 16 deterministic
shards derived from tenant and document ID, while keeping all document-local
index operations single-slot routable. Fan out expiration claims across shards
with a bounded aggregate limit of 96 entries per queue per GC pass, preserving
lease, acknowledgement, retry, and response-loss recovery behavior.

Centralize index key construction, update Redis documentation, and add
standalone/Cluster coverage for shard distribution, key slots, expiration,
force-save processing, and batch limits. No migration is included because
Redis support is not currently deployed.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Prevent a timed-out Redis command from poisoning replacement clients or
unrelated Redis-backed operations.

- separate editor data and statistics connection groups
- preserve transport aborts and fail-closed coordination behavior
- allow transient connection failures to reconnect
- guard aborts against stale client generations
- add lifecycle, concurrency, notification, and topology coverage
- require Sentinel topology settings in failover CI
- prevent Redis tests from accumulating log listeners

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Generate canonical optionsSentinel configuration for standalone and
orchestrated deployments, preserving standalone and Cluster behavior.
Validate Sentinel topology input strictly and keep Redis and Sentinel
credentials separate.

Add entrypoint coverage for all Redis topologies and reduce duplication in
the server Redis connection and expiration handling.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Enable Redis offline-queue disabling across supported topologies, fail fast
when clients are reconnecting, and expand standalone, Cluster, and Sentinel
failure/recovery coverage.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>

@MonaAghili MonaAghili left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I re-checked 9050fa9 against the discussion above. The first section lists only what hasn't been raised yet; the second is the status of earlier items at this head. Items marked reproduced were run against Redis 7 with this PR's own modules and node-redis 6.2.1 (the pinned version); the rest were traced in the code.

Not raised yet

1. CI is red at the head (reproduced)
The fork's run for 9050fa9 fails all 16 Redis/Valkey jobs and the Prettier check.

  • All 16 jobs: editorDataRedis.edge.tests.js asserts that the global force-save queue is empty (edge.tests.js#L56), but editorDataRedis.presence.tests.js leaves a timer behind (presence.tests.js#L19). Every job ran presence before edge, and the edge test received [["presence-index-shard-write","document"]]. Locally, edge passes on a fresh prefix and fails the same way after presence.
  • The 5 Sentinel jobs: these also fail rejects commands issued before initial readiness for every configured topology in editorDataRedis.topology.tests.js. With this PR's Sentinel options, a command sent before readiness is still queued after the test's 2 s bound; standalone rejects it in about 130 ms. This reproduces locally, so the fail-fast guarantee the test checks doesn't hold for Sentinel.
  • Prettier: it flags editorDataRedis.connection.tests.js, editorDataRedis.expiration.tests.js, editorDataRedis.failure.integration.tests.js and editorDataRedis.topology.tests.js. ESLint passes.

Suggested fix: clean up the timer in the presence test and have the edge test check its own document instead of the global queue; make pre-ready Sentinel commands fail fast.

2. One failing shard stops expiry and auto-save for every document (reproduced, new in 9b39145)
_popExpiredAcrossShards() pops the 16 shards with Promise.all (editorData.js#L197-L209). If one shard's EVAL rejects, for example while a Cluster master is down, the whole GC pass fails. The members the other 15 shards already moved into their lease are dropped, and they stay invisible until the 5-minute lease expires. With one shard failing, 90 members were stuck. Before this commit the index used a single {editor:index} tag (slot 1073), so GC depended on one master. It now depends on every master that owns one of the 16 shard slots, which in a default three-master cluster is all three.
Suggested fix: Promise.allSettled, returning the shards that succeeded and logging the failures.

3. A live viewer is enough to reach the cleanup orphan path (reproduced)
Earlier in the thread, the orphan path was put down to index-write ordering or clock skew. A viewer alone is enough:

  • The last editor leaves while a viewer is still connected. hasEditors() ignores viewers, so the save or no-changes cleanup runs, and CLEAN_DOCUMENT_SCRIPT deletes nothing because the viewer's presence remains (scripts.js#L376-L401).
  • When the viewer leaves, closeDocument() takes the view branch, which never calls cleanDocumentOnExit() (DocsCoServer.js#L2189-L2192). removePresenceDocument() also drops the documents index entry, so GC never returns to the document.
  • The next session on the same key gets the previous session's chat in its auth response (DocsCoServer.js#L3402) for up to 24 h; the in-memory backend clears it.
  • The force-save record stays for up to 7 days. With autoAssembly on, the leftover force-save timer later starts a conversion for a document that has already been saved.

I reproduced this by replaying closeDocument()'s storage calls in order.
Suggested fix: keep the HLEN guard, but run the deferred cleanup when the last presence entry goes, for example in the view branch or in removePresenceDocument().

4. A saved-state claim can leak without a crash (adds to the abandoned-claim item)
Suppose the integrator sent c=saved with a status other than '1', and the save was encrypted or ended with isError set: a non-corrupted conversion error, or no file URL or users. The callback is still sent; if the integrator answers {error: 0}, the claim is taken. The condition at canvasservice.js#L1324 then skips the whole block containing both the cleanup and ackSaved(). The claim stays with no TTL, and unless a no-changes cleanup clears it first, the next final save of that document throws SavedStateUnknownError. The path is traced; the stuck claim and the throw are reproduced.

Low
  • Capacity: one GC pass claims at most 96 entries per queue per instance, and nothing loops to drain the rest (reproduced: 96 of 500). With autoAssembly on (5-minute interval, 1-minute step), about 480 actively edited documents per instance is enough to grow the auto-save backlog without bound.
  • Clock skew: presence expiry uses each node's own Date.now(), and GET_PRESENCE_SCRIPT deletes expired entries on read. Since presence is never rebuilt (item 4 in my earlier comment), a node whose clock runs about 3 minutes ahead removes other nodes' live users for good. Traced only. Using Redis TIME inside the scripts would avoid it.
  • ackSaved() treats an already-resolved claim as an error (editorData.js#L318-L327). If the same task is delivered twice (same saveKey, so the same claim id), the second ack throws and commandSfcCallback() ends with an error before its updateVersion publish and removeShutdown() (reproduced). The first delivery already ran both, so this is noise rather than a stuck state; logging would be enough.

Status of earlier items at 9050fa9

None of savedState.js, server.js, canvasservice.js, gc.js, DocsCoServer.js or editorDataMemory.js changed after 85c8fb3.

Still open:

  • The in-memory backend treats every final save as failed (re-reproduced: {success: false, claimed: true} where the old check gave true).
  • SIGTERM, SIGINT, uncaughtException and unhandled rejections hang while editors are connected. Re-checked: server.close() never calls back while an upgraded socket is open, and on Node 20 an unhandled rejection reaches the uncaughtException handler.
  • Lost presence entries are never rebuilt (re-reproduced after FLUSHALL: presence 0, the index entry gone, and the remaining editor's locks deleted at cleanup).
  • The saved-state claim has no TTL, while receiveTask acks in finally (claim TTL -1; the next save throws).
  • Sentinel runs commands one at a time: reserveClient is still not set, and node-redis 6.2.1 still defaults to masterPoolSize: 1.
  • maxCommandRediscovers: 0 also limits Sentinel topology discovery: the node-redis #connect() loop throws on its first failure.
  • The entrypoint writes iooptions while the PR reads optionsSentinel, and the PR rejects the ioredis connector name (DocumentServer main unchanged, #386 still open).
  • documentsCron vs the presence and shard TTLs, and GC per-item error handling.
  • The force-save guard returns UnknownError for Form/Internal requests (re-reproduced).
  • The CI path filter still misses the caller files, and the unit workflow still runs only jest unit.
  • The README omits saved:claim, editorStat.js still ends with the Ascensio block, and the editorDataRedis.js shim has no SPDX header.

Fixed:

  • A failed connect no longer disables Redis for the life of the process (b2c1b80).
  • editorData and editorStat now have separate clients. Checked: a timed-out editor-data command no longer affects editor-stat, but it still fails every other in-flight editor-data command.
  • The README's GC rate wording.

The PR description says Prettier and the Redis suite pass, which no longer holds at this head; ESLint does pass.

@nvanlaerebeke

Copy link
Copy Markdown
Author

Thank you @MonaAghili for the summary, it'll make it much easier for me to work starting from that.

I can do only a couple of things per evening, I'll try to get those addressed asap.

@j-base64

j-base64 commented Oct 1, 2026

Copy link
Copy Markdown

Opened #48 with the packaging check you suggested tracking separately from this PR.

  • Flagging it here only as a follow-up since ci: ❌ fail the build when a pkg.scripts entry matches no file #48 suggest a merge-timing option (not decided yet) to remove the stale ./sources/editorDataRedis.js entry from pkg.scripts until the file actually exists on main. If that happens, this PR (or whichever PR ends up adding the module) would need to re-add that entry alongside the module (it only belongs in package.json once the file exists, otherwise pkg would just drop it again).

Clean up presence and force-save test state, scope queue assertions to the
test document, and fail Sentinel commands fast before initial readiness.
Format the affected Redis tests.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Return null for absent saved state to match the Redis contract and prevent
successful normal final saves from being reported as failures.

Add callback and editor-data regression coverage.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Await Socket.IO cleanup before Redis teardown, terminate upgraded connections during shutdown, clean runtime watchers and timers, and add regression coverage for signals, fatal errors, reconnects, and timeout fallback.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Rebuild missing presence entries during valid heartbeats, guard removals by
connection ID, and keep the document presence index synchronized across
replicas, TTL gaps, and cleanup.

Add Redis-backed regression coverage for recovery, reconnect races, replica
coordination, and live-user cleanup/GC behavior.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Add dedicated TTLs and legacy migration for saved-state claims, make
acknowledgements idempotent, and prevent stale workers from deleting newer
saved state. Expand Redis and callback regression coverage.

Assisted-by: Codex:GPT-5
Signed-off-by: Nico Van Laerebeke <1520618+nvanlaerebeke@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants