Skip to content

docs(redis): correct the cross-datacenter replication guides against a live cluster (cherry-pick of #64) - #65

Merged
chideat merged 1 commit into
masterfrom
cherry-pick/active-active-setup-corrections-to-master
Sep 10, 2026
Merged

chideat merged 1 commit into
masterfrom
cherry-pick/active-active-setup-corrections-to-master

Conversation

@chideat

@chideat chideat commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Cherry-pick of #64 (80ee214) onto master. Applied with git cherry-pick -x, no conflicts, and the result is byte-identical to the merged branch for every documentation file. .tekton/doc-build.yaml and sites.yaml keep master's own versions — they are the only files where the two branches legitimately differ.

The original change corrects the three cross-datacenter replication guides against a live cluster (business-1, operator v5.1.0-rc.50.g59c729d5, redis-modules v5.1.2, Redis 6.0.22 / 7.2.16). Every claim was run rather than read: positive paths, negative paths, and the exact error text.

Active-Active — announcePort is an identity value rather than a transport; the module peer-port number is out of the doc set; seeds are the peer's configured announceAddress:announcePort; Step 1 leads with a dedicated custom account; enabling the mode rolls the data pods and Step 3 has to wait for Ready; the Step 4 status sample was wrong four ways.

Disaster Recovery — Step 1 restructured to match Active-Active, with the default account kept as a documented fallback that carries no credential at all unless the instance was created with spec.passwordSecret. The 7.2 status sample showed a field that a 7.2 peerof link never populates; DPEERS 0 renders blank; rotating the instance password is only independent of the link when a custom account is bound; Decommission does not force a resync. A custom account also works on Redis 6.0 — verified with the peer session showing user=dr6-peer on the upstream — but rotating it there takes the link down until the connection is re-created.

Upgrade 6.0 → 7.2 — the draft procedure was rehearsed end to end. Data survived on both members, checked independently before any link was re-created, and the re-created link resumed PartialSync from the offset the 6.0 link had left off at. "Downgrading is not supported" read as though something enforced it; nothing does.

Also documents where the Redis 6.0 link credential used to live, since ActiveRedisConnection.spec.secretName is removed in v5.1 and anyone arriving from an earlier release has it in their manifests, and covers what the operator upgrade to v5.1 does to a replicating group.

The draft banner on the upgrade guide stays: this was a lab rehearsal, not production validation.

…a live cluster (#64)

* docs(redis): correct the Active-Active setup guide's ports, seeds, and credential

Three corrections to functions/95-disaster-recovery/30-active-active/20-setup.mdx,
verified against redis-operator and redis-modules release-5.1.

announcePort is an identity value, not a transport. The guide described it as
"the RESP port for the gossip control plane" and 7379 as the data plane. Both
mesh planes actually dial the peer port: replication through connectPeerPort
(redis-modules active-redis src/upeer.zig) and gossip through gossipOnePeerPort,
which is the only gossip transport left since the plaintext redis-command
fallback was removed with the legacy as.peersync path (src/mesh.zig, gossipOne).
What announcePort is read for is self-recognition — a node matches a seed
against its own announced addr+port (seedMatchesSelfLocked) — which is what lets
every datacenter carry one uniform seed list. The guide was the outlier here;
10-architecture.mdx, 10-intro.mdx and 90-limitations.mdx already had it right.

The 7379 requirement is unchanged, now stated with its mechanism: the proxy runs
ONE listener, and the 7379 entry on the proxy Service is an alias onto it added
in mesh mode only (internal/builder/activeredis/service.go, targetPort 6379).
The module advertises its own local peer-port verbatim with no announce
override, so an external load balancer must forward both numbers and must not
remap 7379.

Seeds are the remote member's announceAddress:announcePort, not its in-cluster
proxy Service. When a load balancer or VIP fronts a member's proxy, that
endpoint is what its peers must seed, written verbatim — a mismatch defeats the
self-skip that makes a uniform seed list work.

Step 1 leads with a dedicated custom RedisUser again, with the default account
kept as the documented fallback rather than the opening path. This revisits the
framing of 386e527, which made the default account primary after aapeerrepl was
removed from the product; the account itself was never the problem, and
ResolvePeerAuthCredential still records the dedicated account as the stronger
choice because anything holding the default password can otherwise act as a
mesh peer. The example creates its own Secret (password Secrets are 1:1 with
RedisUser objects) and uses a plainly user-chosen username rather than the
retired aapeerrepl. Requirements common to either account are stated once, from
ResolvePeerAuthCredential and the RedisUser webhook: local instance, custom or
default only, a password Secret, phase Success, and the same username and
password value in every datacenter.

Also added, from a redis-operator review of the same fields:

- announceAddress is now a prerequisite. Nothing defaults it to a
  cross-datacenter address and admission only warns, so a member left without
  one advertises an address its peers cannot reach.
- The empty-announceAddress fallback is two-tier, not a straight pod-IP
  fallback: with a managed DNS zone configured the operator derives
  <instance>.<namespace>.<zone> and publishes it via ExternalDNS, and only with
  no zone does the module fall back to the pod IP.
- announceAddress and announcePort must be set on the Redis instance. ActiveRedis
  exposes the same fields, but the operator rewrites that child's spec.proxy
  from the Redis CR on every reconcile, so an edit made there is overwritten.

Known gap, tracked separately: docs/shared/crds/ is published as the site's API
reference, and the AnnouncePort description generated into
middleware.alauda.io_redis.yaml and redis.middleware.alauda.io_activeredis.yaml
still carries the old "mesh CONTROL plane (gossip)" wording. It is corrected at
source in redis-operator (!803 for release-5.1, !804 for master); those two
files should be re-synced here once !803 merges. They are byte-identical to
release-5.1 tip today, so that commit is the entire delta.

* docs(redis): stop presenting the seed @peer-port suffix as a user knob

The Seed syntax section described `@peer-port` as an optional suffix that
"names the port to dial", carried over from the CRD description's framing that
it is for "when an external LB remaps the module peer-port". That remap does
not work, so the sentence advertised a knob the product does not support.

DeriveMeshSeeds does accept and pass an explicit `@peer-port` through
(internal/controller/middleware/activeredis/meshconfig.go), but a seed is only
ever a gossip target: seeds are emitted as bare addr/port records for the
gossip round (redis-modules active-redis src/mesh.zig,
snapshotGossipTargetsRotated) and never enter the member table. Edges are built
from LEARNED records instead — ensurePeer takes np.options.peer_port from
node.peer_port (src/mesh_command.zig) — and every member advertises its own
local peer-port verbatim, which the operator hardcodes to 7379. So a remapped
seed suffix would steer the bootstrap gossip dial and nothing else, leaving
replication on 7379. That is the same conclusion 90-limitations.mdx already
states: the peer port must not be remapped.

The suffix is also not optional in the sense the guide implied. gossipOne fails
a round outright for a target with no peer_port ("A record without an advertised
peer_port is undialable ... mesh-seeds require @peer-port"), because the
peer-port one-shot is the only gossip transport left. The operator appending
7379 is what makes a plain host:port seed dialable at all.

The section now tells the reader to write plain `host:port` and explains the
completed `host:port@7379` form they will see back from
`CONFIG GET activeredis.mesh-seeds`, without offering an override.

Not fixed here: the same "when an external LB remaps the module peer-port"
wording is generated into the ActiveRedisMesh CRD description and published
through docs/shared/crds/, so it needs the same treatment at source in
redis-operator as the AnnouncePort description in !803.

* docs(redis): drop the module peer-port number from the setup guide

The peer port is an internal module mechanism. Everything a datacenter reaches
is the redis-tools proxy, which has a single listener and no notion of that
number — it appears only as a backend dial target inside the pod. Naming it in
a setup guide sent readers looking for a port they never configure.

Removed from all four places it appeared: the prerequisite, the rejected
announceAddress example, the announcePort directive, and the seed grammar.

The announcePort correction is unchanged in substance — it is an identity
value, and neither plane is steered by it — but is now stated without naming
the port that does carry the traffic.

The reachability REQUIREMENT is kept, stated by property instead of by number:
forward every port the proxy Service publishes, each under its own number, and
do not remap them. It still has to be said, because the operator's default
Service type is NodePort and only Ports[0] carries a pinned node port
(internal/builder/activeredis/service.go) — the second publishes on an
arbitrary one, so an external LB or VIP fronting a NodePort member has to be
told what to forward. Under type: LoadBalancer, which is what the guide's
example uses, the LB is provisioned from the Service and publishes both ports
on its own, so nothing is needed and the reader never meets the number.

Seven mentions remain in sibling pages (release_notes.mdx, 90-limitations.mdx,
10-intro.mdx, 20-disaster-recovery/20-setup.mdx, 30-active-active/10-architecture.mdx);
whether to extend this is a separate decision, since the architecture page is
arguably where an internal mechanism does belong.

* docs(redis): drop the module peer-port number from the remaining pages

Completes the previous commit across the rest of the doc set, so the number
appears nowhere: release notes, the mode comparison and port table in the
Cross-Datacenter Replication intro, the limitations entry on address HA, the
Disaster Recovery setup note, and the Active-Active architecture page.

Everything a datacenter reaches is the redis-tools proxy, which has one
listener; the peer port is a backend dial target inside the pod. A reader
configuring an instance never sets it, so the docs describe the port by what it
is rather than by its number, and point at the proxy Service for the value.

Each page keeps its operational content:

- The intro's port table still carries the row, so it stays visible that
  Active-Active publishes a second Service port and Disaster Recovery does not.
  The Port cell now says to read the number off the Service.
- The "expose the peer port unchanged" directive becomes "publish every proxy
  port unchanged" — same requirement for an external LB or firewall, stated by
  property. The RESP port remains the documented exception, remappable through
  announcePort.
- The architecture page still says which traffic uses which port and that the
  module fixes and advertises that number itself; only the literal is gone.

The release-note section was included for the same reason 386e527 applied the
product rename to historical sections: the docs should read consistently rather
than preserve a superseded description of shipped behaviour.

* docs(redis): refresh the published CRDs from redis-operator release-5.1

Syncs docs/shared/crds/ with redis-operator release-5.1 at 59c729d5, which
merged the announcePort correction (!803, 222781b3). Two files change, 31 lines
each; the other four CRDs are already identical.

This retires the last copy of the claim that announcePort carries mesh gossip.
docs/shared/crds/ is published as the site's API reference and generated into
kubectl explain, so until now both were still telling users that gossip rides
that port — the same error the prose fixes in this branch corrected two pages
over. The regenerated description states it correctly: announcePort is the mesh
IDENTITY value, it carries no traffic, and both planes dial the peer port.

Generated files, copied whole from config/crd/bases/ — not hand-edited.

Not resolved by this: the peer-port number still appears in four CRD field
descriptions (redis, activeredis, activeredisconnections, activeredismeshes),
which is the one place the prose policy from the previous two commits cannot
reach, since these are generated. Fixing it means editing the Go field comments
in redis-operator and regenerating. The activeredismeshes description
additionally still frames the seed @peer-port suffix as being for "when an
external LB remaps the module peer-port", which is the framing removed from the
prose in 645925f.

* docs(redis): correct the Active-Active setup guide against a live cluster

Verified the guide line by line on business-1, running operator
v5.1.0-rc.50.g59c729d5 (release-5.1 tip 59c729d5) with redis-modules v5.1.2 and
Redis 7.2.16 — three sentinel instances meshed in namespace aa-doc-verify, all
three ActiveRedisMesh CRs Healthy 3/3. Four defects, all reproduced:

The default-account fallback did not work as written. The guide said every
instance has one "so nothing has to be provisioned". True of the RedisUser, not
of the credential: the default account holds a password Secret only when the
instance was created with spec.passwordSecret, and the create-instance guide's
own example sets none. Binding it there leaves ActiveRedis in Failed with
"has no password secret: the module rejects credential-less peers" — reproduced
on s72-dc1. The fallback now states the precondition and shows that message.

The peer-auth password is policy-checked and the guide never said so. A
placeholder value is refused at admission with "password should consists of
letters, number and special characters"; the rules (8-32 chars, mixed classes)
are now stated where the Secret is created.

Enabling the mode restarts the data pods, which the guide did not mention, and
that interacts with step order: the module's init container is added, so the
instance rolls back through Initializing, and while it is not Ready a custom
RedisUser cannot be created against it ("redis failover s72-dc1 is not ready").
Doing Step 2 before Step 1 therefore strands you until the roll finishes. Step 1
already came first; the guide now says why that matters.

The Step 4 status sample did not match a real mesh:

- epoch was shown as 7 at status level and on every member. It is absent from a
  converged mesh — zero occurrences across all three CRs, and the module reports
  epoch:0 behind them. Removed rather than left as a number nobody will see.
- suspectCount and deadCount were printed as 0; they are omitted at zero.
- address was shown as a bare host. It carries the port:
  activeredis-proxy-rfr-s72-dc1.aa-doc-verify.svc:6379.
- every member was shown fully populated. The LOCAL member's entry is bare —
  uid, serviceID, shardID, state, version and nothing else. Gossip reports the
  other members, so the operator synthesizes the self record and has no source
  for address, clockOffsetMs or moduleVersion. The sample now shows three
  members with the local one short, and a directive says why, so a reader does
  not file a bug against the member that looks incomplete.

moduleVersion 10 was correct for v5.1.2 and is unchanged.

Confirmed correct and left alone: mesh rejected on Redis 6.0; serviceID bounds
[0-15]; all four announceAddress rejections and the bare-IP warning; the
default-account RedisUser naming and both printer-column samples; the proxy
Service publishing its second port onto the proxy listener; the ActiveRedis
child being overwritten from the Redis CR (patched it to a bogus announce
address and watched the operator revert it); seeds completed and pushed
identically to masters and replicas; and the Pending, Paused and Failed phases
with their status messages.

Not testable on one cluster: ExternalDNS binding of the announce host, and
genuine cross-datacenter routing. Neither claim was changed.

* docs(redis): correct the epoch claim and the operations-page mesh reads

Follow-on from the live verification: two more claims outside the setup guide
turned out to be wrong in the same ways.

epoch does not track membership. The architecture page said the status carries
"an epoch that advances as membership changes". It is not a change counter:
ms.self.epoch is incremented in exactly two places in redis-modules active-redis
src/mesh.zig at d224ac1 — forget (:1039) and forgetPermanent (:1066), the
TTL'd-forget and permanent-decommission paths. Join, suspect, dead and recover
never touch it, and ordinary per-round freshness is carried by heartbeat_seq
instead. That matches the measurement: 0 across all three converged meshes, and
epoch:0 from `as.mesh info` on the masters. The page now says it advances only
on forget or decommission, so 0 on a group that has only gained members reads as
correct rather than as a bug to file.

The clock-skew jsonpath sample had the same defect as the Step 4 status sample:
it showed the local member as a populated row. Running the doc's own command
against a real mesh produces a blank first field and a blank offset, because the
local record is synthesized and carries neither:

	alive
  activeredis-proxy-rfr-s72-dc2.aa-doc-verify.svc:6379  alive  -9
  activeredis-proxy-rfr-s72-dc3.aa-doc-verify.svc:6379  alive  -12

The sample now shows that row and says the local member is the reference the
other offsets are measured against. Remote addresses carry their port, as in the
Step 4 fix.

Both `redis-cli` reads on that page now say to run against the shard master.
Only the master participates in gossip — a replica loads the module but must not
advertise itself or initiate gossip (src/mesh.zig:286-292), and the operator
agrees by construction, skipping any node that is not a ready master before
reading as.mesh (activeredismesh_controller.go:150-152). Measured: on the same
converged mesh, the dc2 replica reported alive:0 dead:2 while its master
reported alive:2 dead:0. Someone debugging a skew alarm will otherwise run these
on whichever pod they reach first and conclude the mesh is broken.

Provenance worth recording: the numbers the Step 4 sample used — epoch 7 with
clock_offset_ms -5 — are the fixture values in the module's own unit test for
the node-line format (src/mesh_command.zig:1157-1168, which also carries hb=42,
ver=8.4, mver=2). The sample appears to have been lifted from a test expectation
rather than from a deployment, which is why it did not survive contact with one.

* docs(redis): warn that Step 3 must wait for the Step 2 restart

Found by re-walking the corrected guide end to end on business-1 with a fourth
instance (s72-dc4, serviceID 3), following Step 1 -> 2 -> 3 verbatim.

Creating the ActiveRedisMesh immediately after Step 2 puts it in Failed, not the
Pending the phase table would lead you to expect, and the message says the
module's own commands do not exist:

  set mesh config on pod rfr-s72-dc4-0: ERR Unknown option or number of arguments
  for CONFIG SET - 'activeredis.mesh-seeds'; shard 0 mesh info: ERR unknown
  command 'as.mesh', with args beginning with: 'info'

The cause is the rolling restart Step 2 triggers. The module arrives with the
new pods, and the roll takes them one at a time, so there is a window where the
current master is still the pod that has NOT been rolled and carries no module.
Observed directly at that moment: rfr-s72-dc4-1 had the install-activeredis-module
init container and rfr-s72-dc4-0 did not, and the controller had picked the
latter.

It does clear on its own — Failed -> Pending once the master carries the module
-> Healthy — but while it lasts it is indistinguishable from a real failure, and
a first-time reader following the guide in order will hit it. Step 3 now says to
wait for the instance to report Ready, shows the message, and says it clears.

Also verified in the same walk, nothing to change:

- the not-Ready refusal added in the previous commit, message for message:
  "redis failover s72-dc4 is not ready"
- the Secret must exist before the RedisUser that references it; the guide's
  single create with a --- separator already handles that ordering
- the architecture page's "adding a member does not require touching the
  existing ones": mesh-dc4 seeded ONE existing member (dc1) and reached Healthy
  4/4; the other three learned dc4 through gossip and reached 4/4 with their
  seed lists untouched, confirmed by diffing them before and after
- the local member's entry stays bare at four members, as the Step 4 directive
  now describes

One cosmetic fix: the Step 1 sample's column spacing did not match what kubectl
prints.

* docs(redis): correct the Disaster Recovery setup guide against a live cluster

Verified every claim in "Set Up Disaster Recovery Replication" against a
Redis 7.2 peerof pair and a Redis 6.0 pair on a business cluster
(operator v5.1.0-rc.50.g59c729d5, module image v5.1.2). Most of the guide
held; these are the claims that did not.

Step 1 now leads with a dedicated custom account, mirroring the
Active-Active setup guide. The default account stays as a documented
fallback, with the warning that it carries no credential at all unless the
instance was created with spec.passwordSecret — binding it there leaves
ActiveRedis in Failed. The old text said default passwords are "generated
per instance", which they are not.

Other corrections, each observed live:

* Enabling replication rolls the data pods. Neither the 7.2 nor the 6.0
  section said so, and the steps that follow fail against a half-rolled
  instance. Warned in Step 2, and Step 4 points at it for the downstream.
* The Step 6 status sample showed service_metadata, which a 7.2 peerof link
  never populates — the module reports it empty. It also omitted
  versionState and was not in the order -o yaml prints. The 6.0 sample is
  correct as written and keeps service_metadata.
* The ActiveRedis printer sample showed DPEERS 0; a zero count renders
  blank. Replaced with both sides of a live pair.
* "Rotating an instance's own password Secret is independent and does not
  affect peer links" is only true when a custom account is bound. With the
  default account bound the two are the same Secret — rotating the upstream
  password made a new link fail WRONGPASS.
* Decommission does not force a fresh full synchronization: a connection
  re-created straight after one resumed PartialSync. What it does is stop
  retaining the state that makes resuming possible.
* A Secret shared by several instances is contended by their RedisUser
  controllers and never settles, so "the same default password" needs the
  same value in a Secret per member.
* Rotation and the pre-flight inspection: an established link keeps
  replicating across a one-sided rotation, but a new link cannot be wired
  while the values differ.
* One "master-slave" left in the 6.0 section, next to "master-replica".

In the upgrade guide, peer admission matches on the whole credential, not
the username alone, and a link wired while the two ends straddle the
upgrade parks in Paused rather than failing — captured verbatim.

* docs(redis): correct the 6.0-to-7.2 upgrade guide against a live rehearsal

Ran the whole procedure on a business cluster: a Redis 6.0 Sentinel pair
wired as a Disaster Recovery group, seeded with strings, a list and a hash,
then taken through Steps 1-6 to Redis 7.2.

The procedure works. Both members kept their data across the version
change — verified independently on each member's own master, before any
link was re-created — and the re-created link resumed as PartialSync from
the offset the Redis 6.0 link had left off at, so Step 2's reason for
choosing Detach holds across the generation change too.

Corrections:

* "Downgrading is not supported" read as though something enforced it.
  Nothing does: setting spec.version back to 6.0 is accepted and starts
  rolling the pods onto the older image straight away. Said so plainly.
* The prerequisite about the proxy endpoint now states the fact it depends
  on — the proxy Service is not recreated, so its address, port and node
  port all survive the upgrade, and the recorded addresses stay valid.
* Step 6 gains syncStatus to the checklist, with what a FullSync there
  would mean, and asks for the pre-upgrade keys as well as a new one.
* A link wired while the two ends straddle the upgrade parks in Paused
  rather than failing or replicating — captured verbatim, since a reader
  who mistimes Step 5 will see it and it looks alarming.

* docs(redis): tighten three claims from the Disaster Recovery pass

Follow-ups on the previous two commits, each brought back to what was
actually observed.

The rolling-restart warning went into the Redis 7.2 section only, but the
6.0 section rolls its pods the same way and its next step — creating the
connection — is rejected while the instance is still coming back. Both 6.0
patch steps now say so, with the rejection text.

The Decommission directive said the retention that makes resuming possible
is no longer kept, so a returning peer needs a full synchronization. The
teardown runs on the instance that owns the connection and releases what
*it* holds for that peer; the upstream's retention is untouched, which is
why a link re-created straight afterwards resumed PartialSync. The
directive now describes what the teardown does and says plainly that
re-creating the link does not undo it — detach a peer you expect back.

versionState was described as where a version-held link shows up. The
module reports it as unknown | verified | violating, and a held shard is
marked by a non-zero versionHoldSince instead.

Also walked the newly-recommended custom-account path on the peerof pair
rather than relying on the Active-Active run: both RedisUsers reach
Success, switching spec.activeRedis.redisUserRef is a live change with no
pod restart, and the link comes up Healthy and replicates.

* docs(redis): document the custom peer-auth account on Redis 6.0

The Redis 6.0 section described only the automatic default-account binding,
leaving open whether a dedicated account is usable there at all. It is —
verified on a purpose-built 6.0 pair.

The test was built so the result could not be ambiguous: the two instances
were given different default passwords, and the custom account a third one.
A link that comes up can therefore only be authenticating as the custom
account. It came up Healthy and replicated, and the peer session on the
upstream master shows `cmd=as.peerconf user=dr6-peer` — the custom ACL user,
not default. Binding it in the same patch that enables replication is enough;
the automatic binding leaves an existing one alone.

What differs from Redis 7.2 is rotation, and it differs badly enough to
warrant a warning of its own. A 6.0 link carries the credential inline in
the module's peer record, fixed at wiring time, and the operator's
self-healing re-wire is new-module-only. After rotating the password the
running link kept working until its session was dropped, then sat at
`shard 0 status is Disconnected` and never recovered on its own. Deleting
and re-creating the ActiveRedisConnection restored it immediately.

Both facts are now in the 6.0 section, and the credential-rotation section
carries the 6.0 exception next to the 7.2 procedure it does not share.

* docs(redis): record where the Redis 6.0 link credential used to live

The docs stated that ActiveRedisConnection.spec.secretName "has been
removed" without saying what it was, what replaced it, or what an existing
Disaster Recovery group has to do about it. Anyone arriving from a release
before v5.1 has that field in their manifests.

The Redis 6.0 section now carries the history: until v5.1, Disaster
Recovery on Redis 6.0 was the only cross-datacenter replication this
product offered, and the link credential lived on the connection, in a
Secret holding a password and no username. That is why the account has
always been the instance's default one — the legacy module sends a
one-argument AUTH and the proxy resolves it to default — and why both
datacenters have always needed the same default password.

v5.1 removes the field from ActiveRedisConnection and
ActiveRedisInspection and takes the credential from
spec.activeRedis.redisUserRef alone. Existing groups are migrated with no
intervention: the operator binds each replicating instance to its own
default-account RedisUser on its first reconcile, which changes nothing on
the wire. Stored resources keep working; a manifest that still sets
secretName is rejected on its next apply.

Also in the release notes, which described the automatic binding as
"deriving it from the credential the replication group already shares".
Nothing is derived or provisioned — the account already exists and the
binding only names it. Added the field removal as an entry of its own.

The intro now dates the two generations, so "the legacy module" reads as
the product's original Disaster Recovery (v4.1.0) rather than an
unexplained alternative.

* docs(redis): cover replicating instances in the upgrade to v5.1

The upgrade guide told a reader with a Disaster Recovery group what to do
about their Redis version, but nothing about what the operator upgrade
itself does to them — which is the question someone planning a 5.0 to 5.1
upgrade actually has.

Confirmed by manual test of the upgrade: replication continues across it
and the data on both sides is unaffected. The instances keep their Redis
version, so they keep the module generation they were already running.

What moves is where the link credential is declared: onto the instance's
spec.activeRedis.redisUserRef, written by the operator on its first
reconcile and naming the default account the link already authenticated
as. No action required for the running group; the one thing that does
need attention is manifests still setting the removed spec.secretName.

* docs(redis): stop using the forbidden word "master" in the pages this pass edits

The shared glossary maps master to control plane, so doom lint reports
`Forbidden word: "master" (control plane)`. It is a Kubernetes node-naming
rule and it fires on Redis's own replication role, but the rule is the rule
and CI is red on it.

Two things had hidden this. The local lint builds its flag list from the
shared terms file, so with that fetch blocked it reports unknown words only
and nothing else. And CI reports issues on changed lines, so the
occurrences this pass did not touch have never failed a build.

Checked against cspell-lib with the mapping applied rather than guessed at:
the hyphenated forms are flagged too, because master-replica and
master-slave tokenize to master. Only mymaster survives, as one token. That
made one more line in this branch a latent failure, the Sentinel shard-0
note whose master-slave I had already changed to master-replica.

Every line this branch touches now reads primary node or primary-replica
pair and is clean. The remaining occurrences are in lines this branch never
edited and are left alone; they are listed in TERMINOLOGY_CANDIDATES.md
along with the reason a mechanical sweep would be wrong — the same files
carry master as an identifier, in spec.replicas.sentinel.master and in
mymaster / MasterName / setMasterName, where changing it breaks the
examples. Whether primary node is the accepted replacement, and whether the
glossary should carry an exemption for the Redis role sense, are noted there
as open questions rather than decided here.

(cherry picked from commit 80ee214)
@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying alauda-redis with  Cloudflare Pages  Cloudflare Pages

Latest commit: f65a52c
Status: ✅  Deploy successful!
Preview URL: https://5c951a8a.alauda-redis.pages.dev
Branch Preview URL: https://cherry-pick-active-active-se.alauda-redis.pages.dev

View logs

@chideat
chideat merged commit f73329d into master Sep 10, 2026
3 checks passed
@chideat
chideat deleted the cherry-pick/active-active-setup-corrections-to-master branch September 10, 2026 09:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant