Repository navigation
docs(redis): correct the cross-datacenter replication guides against a live cluster (cherry-pick of #64) - #65
Merged
chideat merged 1 commit intoSep 10, 2026
Conversation
…a live cluster (#64) * docs(redis): correct the Active-Active setup guide's ports, seeds, and credential Three corrections to functions/95-disaster-recovery/30-active-active/20-setup.mdx, verified against redis-operator and redis-modules release-5.1. announcePort is an identity value, not a transport. The guide described it as "the RESP port for the gossip control plane" and 7379 as the data plane. Both mesh planes actually dial the peer port: replication through connectPeerPort (redis-modules active-redis src/upeer.zig) and gossip through gossipOnePeerPort, which is the only gossip transport left since the plaintext redis-command fallback was removed with the legacy as.peersync path (src/mesh.zig, gossipOne). What announcePort is read for is self-recognition — a node matches a seed against its own announced addr+port (seedMatchesSelfLocked) — which is what lets every datacenter carry one uniform seed list. The guide was the outlier here; 10-architecture.mdx, 10-intro.mdx and 90-limitations.mdx already had it right. The 7379 requirement is unchanged, now stated with its mechanism: the proxy runs ONE listener, and the 7379 entry on the proxy Service is an alias onto it added in mesh mode only (internal/builder/activeredis/service.go, targetPort 6379). The module advertises its own local peer-port verbatim with no announce override, so an external load balancer must forward both numbers and must not remap 7379. Seeds are the remote member's announceAddress:announcePort, not its in-cluster proxy Service. When a load balancer or VIP fronts a member's proxy, that endpoint is what its peers must seed, written verbatim — a mismatch defeats the self-skip that makes a uniform seed list work. Step 1 leads with a dedicated custom RedisUser again, with the default account kept as the documented fallback rather than the opening path. This revisits the framing of 386e527, which made the default account primary after aapeerrepl was removed from the product; the account itself was never the problem, and ResolvePeerAuthCredential still records the dedicated account as the stronger choice because anything holding the default password can otherwise act as a mesh peer. The example creates its own Secret (password Secrets are 1:1 with RedisUser objects) and uses a plainly user-chosen username rather than the retired aapeerrepl. Requirements common to either account are stated once, from ResolvePeerAuthCredential and the RedisUser webhook: local instance, custom or default only, a password Secret, phase Success, and the same username and password value in every datacenter. Also added, from a redis-operator review of the same fields: - announceAddress is now a prerequisite. Nothing defaults it to a cross-datacenter address and admission only warns, so a member left without one advertises an address its peers cannot reach. - The empty-announceAddress fallback is two-tier, not a straight pod-IP fallback: with a managed DNS zone configured the operator derives <instance>.<namespace>.<zone> and publishes it via ExternalDNS, and only with no zone does the module fall back to the pod IP. - announceAddress and announcePort must be set on the Redis instance. ActiveRedis exposes the same fields, but the operator rewrites that child's spec.proxy from the Redis CR on every reconcile, so an edit made there is overwritten. Known gap, tracked separately: docs/shared/crds/ is published as the site's API reference, and the AnnouncePort description generated into middleware.alauda.io_redis.yaml and redis.middleware.alauda.io_activeredis.yaml still carries the old "mesh CONTROL plane (gossip)" wording. It is corrected at source in redis-operator (!803 for release-5.1, !804 for master); those two files should be re-synced here once !803 merges. They are byte-identical to release-5.1 tip today, so that commit is the entire delta. * docs(redis): stop presenting the seed @peer-port suffix as a user knob The Seed syntax section described `@peer-port` as an optional suffix that "names the port to dial", carried over from the CRD description's framing that it is for "when an external LB remaps the module peer-port". That remap does not work, so the sentence advertised a knob the product does not support. DeriveMeshSeeds does accept and pass an explicit `@peer-port` through (internal/controller/middleware/activeredis/meshconfig.go), but a seed is only ever a gossip target: seeds are emitted as bare addr/port records for the gossip round (redis-modules active-redis src/mesh.zig, snapshotGossipTargetsRotated) and never enter the member table. Edges are built from LEARNED records instead — ensurePeer takes np.options.peer_port from node.peer_port (src/mesh_command.zig) — and every member advertises its own local peer-port verbatim, which the operator hardcodes to 7379. So a remapped seed suffix would steer the bootstrap gossip dial and nothing else, leaving replication on 7379. That is the same conclusion 90-limitations.mdx already states: the peer port must not be remapped. The suffix is also not optional in the sense the guide implied. gossipOne fails a round outright for a target with no peer_port ("A record without an advertised peer_port is undialable ... mesh-seeds require @peer-port"), because the peer-port one-shot is the only gossip transport left. The operator appending 7379 is what makes a plain host:port seed dialable at all. The section now tells the reader to write plain `host:port` and explains the completed `host:port@7379` form they will see back from `CONFIG GET activeredis.mesh-seeds`, without offering an override. Not fixed here: the same "when an external LB remaps the module peer-port" wording is generated into the ActiveRedisMesh CRD description and published through docs/shared/crds/, so it needs the same treatment at source in redis-operator as the AnnouncePort description in !803. * docs(redis): drop the module peer-port number from the setup guide The peer port is an internal module mechanism. Everything a datacenter reaches is the redis-tools proxy, which has a single listener and no notion of that number — it appears only as a backend dial target inside the pod. Naming it in a setup guide sent readers looking for a port they never configure. Removed from all four places it appeared: the prerequisite, the rejected announceAddress example, the announcePort directive, and the seed grammar. The announcePort correction is unchanged in substance — it is an identity value, and neither plane is steered by it — but is now stated without naming the port that does carry the traffic. The reachability REQUIREMENT is kept, stated by property instead of by number: forward every port the proxy Service publishes, each under its own number, and do not remap them. It still has to be said, because the operator's default Service type is NodePort and only Ports[0] carries a pinned node port (internal/builder/activeredis/service.go) — the second publishes on an arbitrary one, so an external LB or VIP fronting a NodePort member has to be told what to forward. Under type: LoadBalancer, which is what the guide's example uses, the LB is provisioned from the Service and publishes both ports on its own, so nothing is needed and the reader never meets the number. Seven mentions remain in sibling pages (release_notes.mdx, 90-limitations.mdx, 10-intro.mdx, 20-disaster-recovery/20-setup.mdx, 30-active-active/10-architecture.mdx); whether to extend this is a separate decision, since the architecture page is arguably where an internal mechanism does belong. * docs(redis): drop the module peer-port number from the remaining pages Completes the previous commit across the rest of the doc set, so the number appears nowhere: release notes, the mode comparison and port table in the Cross-Datacenter Replication intro, the limitations entry on address HA, the Disaster Recovery setup note, and the Active-Active architecture page. Everything a datacenter reaches is the redis-tools proxy, which has one listener; the peer port is a backend dial target inside the pod. A reader configuring an instance never sets it, so the docs describe the port by what it is rather than by its number, and point at the proxy Service for the value. Each page keeps its operational content: - The intro's port table still carries the row, so it stays visible that Active-Active publishes a second Service port and Disaster Recovery does not. The Port cell now says to read the number off the Service. - The "expose the peer port unchanged" directive becomes "publish every proxy port unchanged" — same requirement for an external LB or firewall, stated by property. The RESP port remains the documented exception, remappable through announcePort. - The architecture page still says which traffic uses which port and that the module fixes and advertises that number itself; only the literal is gone. The release-note section was included for the same reason 386e527 applied the product rename to historical sections: the docs should read consistently rather than preserve a superseded description of shipped behaviour. * docs(redis): refresh the published CRDs from redis-operator release-5.1 Syncs docs/shared/crds/ with redis-operator release-5.1 at 59c729d5, which merged the announcePort correction (!803, 222781b3). Two files change, 31 lines each; the other four CRDs are already identical. This retires the last copy of the claim that announcePort carries mesh gossip. docs/shared/crds/ is published as the site's API reference and generated into kubectl explain, so until now both were still telling users that gossip rides that port — the same error the prose fixes in this branch corrected two pages over. The regenerated description states it correctly: announcePort is the mesh IDENTITY value, it carries no traffic, and both planes dial the peer port. Generated files, copied whole from config/crd/bases/ — not hand-edited. Not resolved by this: the peer-port number still appears in four CRD field descriptions (redis, activeredis, activeredisconnections, activeredismeshes), which is the one place the prose policy from the previous two commits cannot reach, since these are generated. Fixing it means editing the Go field comments in redis-operator and regenerating. The activeredismeshes description additionally still frames the seed @peer-port suffix as being for "when an external LB remaps the module peer-port", which is the framing removed from the prose in 645925f. * docs(redis): correct the Active-Active setup guide against a live cluster Verified the guide line by line on business-1, running operator v5.1.0-rc.50.g59c729d5 (release-5.1 tip 59c729d5) with redis-modules v5.1.2 and Redis 7.2.16 — three sentinel instances meshed in namespace aa-doc-verify, all three ActiveRedisMesh CRs Healthy 3/3. Four defects, all reproduced: The default-account fallback did not work as written. The guide said every instance has one "so nothing has to be provisioned". True of the RedisUser, not of the credential: the default account holds a password Secret only when the instance was created with spec.passwordSecret, and the create-instance guide's own example sets none. Binding it there leaves ActiveRedis in Failed with "has no password secret: the module rejects credential-less peers" — reproduced on s72-dc1. The fallback now states the precondition and shows that message. The peer-auth password is policy-checked and the guide never said so. A placeholder value is refused at admission with "password should consists of letters, number and special characters"; the rules (8-32 chars, mixed classes) are now stated where the Secret is created. Enabling the mode restarts the data pods, which the guide did not mention, and that interacts with step order: the module's init container is added, so the instance rolls back through Initializing, and while it is not Ready a custom RedisUser cannot be created against it ("redis failover s72-dc1 is not ready"). Doing Step 2 before Step 1 therefore strands you until the roll finishes. Step 1 already came first; the guide now says why that matters. The Step 4 status sample did not match a real mesh: - epoch was shown as 7 at status level and on every member. It is absent from a converged mesh — zero occurrences across all three CRs, and the module reports epoch:0 behind them. Removed rather than left as a number nobody will see. - suspectCount and deadCount were printed as 0; they are omitted at zero. - address was shown as a bare host. It carries the port: activeredis-proxy-rfr-s72-dc1.aa-doc-verify.svc:6379. - every member was shown fully populated. The LOCAL member's entry is bare — uid, serviceID, shardID, state, version and nothing else. Gossip reports the other members, so the operator synthesizes the self record and has no source for address, clockOffsetMs or moduleVersion. The sample now shows three members with the local one short, and a directive says why, so a reader does not file a bug against the member that looks incomplete. moduleVersion 10 was correct for v5.1.2 and is unchanged. Confirmed correct and left alone: mesh rejected on Redis 6.0; serviceID bounds [0-15]; all four announceAddress rejections and the bare-IP warning; the default-account RedisUser naming and both printer-column samples; the proxy Service publishing its second port onto the proxy listener; the ActiveRedis child being overwritten from the Redis CR (patched it to a bogus announce address and watched the operator revert it); seeds completed and pushed identically to masters and replicas; and the Pending, Paused and Failed phases with their status messages. Not testable on one cluster: ExternalDNS binding of the announce host, and genuine cross-datacenter routing. Neither claim was changed. * docs(redis): correct the epoch claim and the operations-page mesh reads Follow-on from the live verification: two more claims outside the setup guide turned out to be wrong in the same ways. epoch does not track membership. The architecture page said the status carries "an epoch that advances as membership changes". It is not a change counter: ms.self.epoch is incremented in exactly two places in redis-modules active-redis src/mesh.zig at d224ac1 — forget (:1039) and forgetPermanent (:1066), the TTL'd-forget and permanent-decommission paths. Join, suspect, dead and recover never touch it, and ordinary per-round freshness is carried by heartbeat_seq instead. That matches the measurement: 0 across all three converged meshes, and epoch:0 from `as.mesh info` on the masters. The page now says it advances only on forget or decommission, so 0 on a group that has only gained members reads as correct rather than as a bug to file. The clock-skew jsonpath sample had the same defect as the Step 4 status sample: it showed the local member as a populated row. Running the doc's own command against a real mesh produces a blank first field and a blank offset, because the local record is synthesized and carries neither: alive activeredis-proxy-rfr-s72-dc2.aa-doc-verify.svc:6379 alive -9 activeredis-proxy-rfr-s72-dc3.aa-doc-verify.svc:6379 alive -12 The sample now shows that row and says the local member is the reference the other offsets are measured against. Remote addresses carry their port, as in the Step 4 fix. Both `redis-cli` reads on that page now say to run against the shard master. Only the master participates in gossip — a replica loads the module but must not advertise itself or initiate gossip (src/mesh.zig:286-292), and the operator agrees by construction, skipping any node that is not a ready master before reading as.mesh (activeredismesh_controller.go:150-152). Measured: on the same converged mesh, the dc2 replica reported alive:0 dead:2 while its master reported alive:2 dead:0. Someone debugging a skew alarm will otherwise run these on whichever pod they reach first and conclude the mesh is broken. Provenance worth recording: the numbers the Step 4 sample used — epoch 7 with clock_offset_ms -5 — are the fixture values in the module's own unit test for the node-line format (src/mesh_command.zig:1157-1168, which also carries hb=42, ver=8.4, mver=2). The sample appears to have been lifted from a test expectation rather than from a deployment, which is why it did not survive contact with one. * docs(redis): warn that Step 3 must wait for the Step 2 restart Found by re-walking the corrected guide end to end on business-1 with a fourth instance (s72-dc4, serviceID 3), following Step 1 -> 2 -> 3 verbatim. Creating the ActiveRedisMesh immediately after Step 2 puts it in Failed, not the Pending the phase table would lead you to expect, and the message says the module's own commands do not exist: set mesh config on pod rfr-s72-dc4-0: ERR Unknown option or number of arguments for CONFIG SET - 'activeredis.mesh-seeds'; shard 0 mesh info: ERR unknown command 'as.mesh', with args beginning with: 'info' The cause is the rolling restart Step 2 triggers. The module arrives with the new pods, and the roll takes them one at a time, so there is a window where the current master is still the pod that has NOT been rolled and carries no module. Observed directly at that moment: rfr-s72-dc4-1 had the install-activeredis-module init container and rfr-s72-dc4-0 did not, and the controller had picked the latter. It does clear on its own — Failed -> Pending once the master carries the module -> Healthy — but while it lasts it is indistinguishable from a real failure, and a first-time reader following the guide in order will hit it. Step 3 now says to wait for the instance to report Ready, shows the message, and says it clears. Also verified in the same walk, nothing to change: - the not-Ready refusal added in the previous commit, message for message: "redis failover s72-dc4 is not ready" - the Secret must exist before the RedisUser that references it; the guide's single create with a --- separator already handles that ordering - the architecture page's "adding a member does not require touching the existing ones": mesh-dc4 seeded ONE existing member (dc1) and reached Healthy 4/4; the other three learned dc4 through gossip and reached 4/4 with their seed lists untouched, confirmed by diffing them before and after - the local member's entry stays bare at four members, as the Step 4 directive now describes One cosmetic fix: the Step 1 sample's column spacing did not match what kubectl prints. * docs(redis): correct the Disaster Recovery setup guide against a live cluster Verified every claim in "Set Up Disaster Recovery Replication" against a Redis 7.2 peerof pair and a Redis 6.0 pair on a business cluster (operator v5.1.0-rc.50.g59c729d5, module image v5.1.2). Most of the guide held; these are the claims that did not. Step 1 now leads with a dedicated custom account, mirroring the Active-Active setup guide. The default account stays as a documented fallback, with the warning that it carries no credential at all unless the instance was created with spec.passwordSecret — binding it there leaves ActiveRedis in Failed. The old text said default passwords are "generated per instance", which they are not. Other corrections, each observed live: * Enabling replication rolls the data pods. Neither the 7.2 nor the 6.0 section said so, and the steps that follow fail against a half-rolled instance. Warned in Step 2, and Step 4 points at it for the downstream. * The Step 6 status sample showed service_metadata, which a 7.2 peerof link never populates — the module reports it empty. It also omitted versionState and was not in the order -o yaml prints. The 6.0 sample is correct as written and keeps service_metadata. * The ActiveRedis printer sample showed DPEERS 0; a zero count renders blank. Replaced with both sides of a live pair. * "Rotating an instance's own password Secret is independent and does not affect peer links" is only true when a custom account is bound. With the default account bound the two are the same Secret — rotating the upstream password made a new link fail WRONGPASS. * Decommission does not force a fresh full synchronization: a connection re-created straight after one resumed PartialSync. What it does is stop retaining the state that makes resuming possible. * A Secret shared by several instances is contended by their RedisUser controllers and never settles, so "the same default password" needs the same value in a Secret per member. * Rotation and the pre-flight inspection: an established link keeps replicating across a one-sided rotation, but a new link cannot be wired while the values differ. * One "master-slave" left in the 6.0 section, next to "master-replica". In the upgrade guide, peer admission matches on the whole credential, not the username alone, and a link wired while the two ends straddle the upgrade parks in Paused rather than failing — captured verbatim. * docs(redis): correct the 6.0-to-7.2 upgrade guide against a live rehearsal Ran the whole procedure on a business cluster: a Redis 6.0 Sentinel pair wired as a Disaster Recovery group, seeded with strings, a list and a hash, then taken through Steps 1-6 to Redis 7.2. The procedure works. Both members kept their data across the version change — verified independently on each member's own master, before any link was re-created — and the re-created link resumed as PartialSync from the offset the Redis 6.0 link had left off at, so Step 2's reason for choosing Detach holds across the generation change too. Corrections: * "Downgrading is not supported" read as though something enforced it. Nothing does: setting spec.version back to 6.0 is accepted and starts rolling the pods onto the older image straight away. Said so plainly. * The prerequisite about the proxy endpoint now states the fact it depends on — the proxy Service is not recreated, so its address, port and node port all survive the upgrade, and the recorded addresses stay valid. * Step 6 gains syncStatus to the checklist, with what a FullSync there would mean, and asks for the pre-upgrade keys as well as a new one. * A link wired while the two ends straddle the upgrade parks in Paused rather than failing or replicating — captured verbatim, since a reader who mistimes Step 5 will see it and it looks alarming. * docs(redis): tighten three claims from the Disaster Recovery pass Follow-ups on the previous two commits, each brought back to what was actually observed. The rolling-restart warning went into the Redis 7.2 section only, but the 6.0 section rolls its pods the same way and its next step — creating the connection — is rejected while the instance is still coming back. Both 6.0 patch steps now say so, with the rejection text. The Decommission directive said the retention that makes resuming possible is no longer kept, so a returning peer needs a full synchronization. The teardown runs on the instance that owns the connection and releases what *it* holds for that peer; the upstream's retention is untouched, which is why a link re-created straight afterwards resumed PartialSync. The directive now describes what the teardown does and says plainly that re-creating the link does not undo it — detach a peer you expect back. versionState was described as where a version-held link shows up. The module reports it as unknown | verified | violating, and a held shard is marked by a non-zero versionHoldSince instead. Also walked the newly-recommended custom-account path on the peerof pair rather than relying on the Active-Active run: both RedisUsers reach Success, switching spec.activeRedis.redisUserRef is a live change with no pod restart, and the link comes up Healthy and replicates. * docs(redis): document the custom peer-auth account on Redis 6.0 The Redis 6.0 section described only the automatic default-account binding, leaving open whether a dedicated account is usable there at all. It is — verified on a purpose-built 6.0 pair. The test was built so the result could not be ambiguous: the two instances were given different default passwords, and the custom account a third one. A link that comes up can therefore only be authenticating as the custom account. It came up Healthy and replicated, and the peer session on the upstream master shows `cmd=as.peerconf user=dr6-peer` — the custom ACL user, not default. Binding it in the same patch that enables replication is enough; the automatic binding leaves an existing one alone. What differs from Redis 7.2 is rotation, and it differs badly enough to warrant a warning of its own. A 6.0 link carries the credential inline in the module's peer record, fixed at wiring time, and the operator's self-healing re-wire is new-module-only. After rotating the password the running link kept working until its session was dropped, then sat at `shard 0 status is Disconnected` and never recovered on its own. Deleting and re-creating the ActiveRedisConnection restored it immediately. Both facts are now in the 6.0 section, and the credential-rotation section carries the 6.0 exception next to the 7.2 procedure it does not share. * docs(redis): record where the Redis 6.0 link credential used to live The docs stated that ActiveRedisConnection.spec.secretName "has been removed" without saying what it was, what replaced it, or what an existing Disaster Recovery group has to do about it. Anyone arriving from a release before v5.1 has that field in their manifests. The Redis 6.0 section now carries the history: until v5.1, Disaster Recovery on Redis 6.0 was the only cross-datacenter replication this product offered, and the link credential lived on the connection, in a Secret holding a password and no username. That is why the account has always been the instance's default one — the legacy module sends a one-argument AUTH and the proxy resolves it to default — and why both datacenters have always needed the same default password. v5.1 removes the field from ActiveRedisConnection and ActiveRedisInspection and takes the credential from spec.activeRedis.redisUserRef alone. Existing groups are migrated with no intervention: the operator binds each replicating instance to its own default-account RedisUser on its first reconcile, which changes nothing on the wire. Stored resources keep working; a manifest that still sets secretName is rejected on its next apply. Also in the release notes, which described the automatic binding as "deriving it from the credential the replication group already shares". Nothing is derived or provisioned — the account already exists and the binding only names it. Added the field removal as an entry of its own. The intro now dates the two generations, so "the legacy module" reads as the product's original Disaster Recovery (v4.1.0) rather than an unexplained alternative. * docs(redis): cover replicating instances in the upgrade to v5.1 The upgrade guide told a reader with a Disaster Recovery group what to do about their Redis version, but nothing about what the operator upgrade itself does to them — which is the question someone planning a 5.0 to 5.1 upgrade actually has. Confirmed by manual test of the upgrade: replication continues across it and the data on both sides is unaffected. The instances keep their Redis version, so they keep the module generation they were already running. What moves is where the link credential is declared: onto the instance's spec.activeRedis.redisUserRef, written by the operator on its first reconcile and naming the default account the link already authenticated as. No action required for the running group; the one thing that does need attention is manifests still setting the removed spec.secretName. * docs(redis): stop using the forbidden word "master" in the pages this pass edits The shared glossary maps master to control plane, so doom lint reports `Forbidden word: "master" (control plane)`. It is a Kubernetes node-naming rule and it fires on Redis's own replication role, but the rule is the rule and CI is red on it. Two things had hidden this. The local lint builds its flag list from the shared terms file, so with that fetch blocked it reports unknown words only and nothing else. And CI reports issues on changed lines, so the occurrences this pass did not touch have never failed a build. Checked against cspell-lib with the mapping applied rather than guessed at: the hyphenated forms are flagged too, because master-replica and master-slave tokenize to master. Only mymaster survives, as one token. That made one more line in this branch a latent failure, the Sentinel shard-0 note whose master-slave I had already changed to master-replica. Every line this branch touches now reads primary node or primary-replica pair and is clean. The remaining occurrences are in lines this branch never edited and are left alone; they are listed in TERMINOLOGY_CANDIDATES.md along with the reason a mechanical sweep would be wrong — the same files carry master as an identifier, in spec.replicas.sentinel.master and in mymaster / MasterName / setMasterName, where changing it breaks the examples. Whether primary node is the accepted replacement, and whether the glossary should carry an exemption for the Redis role sense, are noted there as open questions rather than decided here. (cherry picked from commit 80ee214)
Deploying alauda-redis with
|
| Latest commit: |
f65a52c
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://5c951a8a.alauda-redis.pages.dev |
| Branch Preview URL: | https://cherry-pick-active-active-se.alauda-redis.pages.dev |
chideat
deleted the
cherry-pick/active-active-setup-corrections-to-master
branch
September 10, 2026 09:53
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cherry-pick of #64 (
80ee214) ontomaster. Applied withgit cherry-pick -x, no conflicts, and the result is byte-identical to the merged branch for every documentation file..tekton/doc-build.yamlandsites.yamlkeep master's own versions — they are the only files where the two branches legitimately differ.The original change corrects the three cross-datacenter replication guides against a live cluster (business-1, operator
v5.1.0-rc.50.g59c729d5, redis-modulesv5.1.2, Redis 6.0.22 / 7.2.16). Every claim was run rather than read: positive paths, negative paths, and the exact error text.Active-Active —
announcePortis an identity value rather than a transport; the module peer-port number is out of the doc set; seeds are the peer's configuredannounceAddress:announcePort; Step 1 leads with a dedicatedcustomaccount; enabling the mode rolls the data pods and Step 3 has to wait forReady; the Step 4 status sample was wrong four ways.Disaster Recovery — Step 1 restructured to match Active-Active, with the default account kept as a documented fallback that carries no credential at all unless the instance was created with
spec.passwordSecret. The 7.2 status sample showed a field that a 7.2peeroflink never populates;DPEERS 0renders blank; rotating the instance password is only independent of the link when a custom account is bound;Decommissiondoes not force a resync. Acustomaccount also works on Redis 6.0 — verified with the peer session showinguser=dr6-peeron the upstream — but rotating it there takes the link down until the connection is re-created.Upgrade 6.0 → 7.2 — the draft procedure was rehearsed end to end. Data survived on both members, checked independently before any link was re-created, and the re-created link resumed
PartialSyncfrom the offset the 6.0 link had left off at. "Downgrading is not supported" read as though something enforced it; nothing does.Also documents where the Redis 6.0 link credential used to live, since
ActiveRedisConnection.spec.secretNameis removed in v5.1 and anyone arriving from an earlier release has it in their manifests, and covers what the operator upgrade to v5.1 does to a replicating group.The draft banner on the upgrade guide stays: this was a lab rehearsal, not production validation.