From 1d3341cb166ba5b444f5c2a251d936e856d5e201 Mon Sep 17 00:00:00 2001 From: Cody Chen Date: Mon, 7 Sep 2026 11:39:12 +0800 Subject: [PATCH 1/4] feat: add getting started section that introduces how to install otterscale step-by-step --- astro.config.mjs | 9 + .../docs/getting-started/01-requirements.mdx | 144 ++++++++ .../getting-started/02-prepare-cluster.mdx | 279 +++++++++++++++ .../getting-started/03-install-server.mdx | 197 +++++++++++ .../docs/getting-started/04-join-cluster.mdx | 327 ++++++++++++++++++ .../docs/getting-started/05-first-login.mdx | 165 +++++++++ src/content/docs/introduction.mdx | 30 +- 7 files changed, 1148 insertions(+), 3 deletions(-) create mode 100644 src/content/docs/getting-started/01-requirements.mdx create mode 100644 src/content/docs/getting-started/02-prepare-cluster.mdx create mode 100644 src/content/docs/getting-started/03-install-server.mdx create mode 100644 src/content/docs/getting-started/04-join-cluster.mdx create mode 100644 src/content/docs/getting-started/05-first-login.mdx diff --git a/astro.config.mjs b/astro.config.mjs index cb42916..7b33e3d 100644 --- a/astro.config.mjs +++ b/astro.config.mjs @@ -43,6 +43,15 @@ export default defineConfig({ 'zh-Hant': '簡介' } }, + { + label: 'Getting Started', + translations: { + 'zh-Hant': '開始使用' + }, + autogenerate: { + directory: 'getting-started' + } + }, ...openAPISidebarGroups ], customCss: ['./src/styles/global.css'], diff --git a/src/content/docs/getting-started/01-requirements.mdx b/src/content/docs/getting-started/01-requirements.mdx new file mode 100644 index 0000000..1502daf --- /dev/null +++ b/src/content/docs/getting-started/01-requirements.mdx @@ -0,0 +1,144 @@ +--- +title: Requirements +description: What you need before installing OtterScale, from clusters and cluster APIs to reserved IPs, storage, and client tooling. +slug: getting-started/requirements +sidebar: + order: 1 +--- + +import { Aside, Steps, Card, CardGrid } from '@astrojs/starlight/components'; + +OtterScale is a hub-and-spoke platform: one **server** (the hub) holds the API, the dashboard, the +identity provider and the registry, and one lightweight **agent** (a spoke) runs inside every +cluster you want to manage. Requirements differ between the two roles, so this page separates them. + + + +## Clusters + +| Role | Count | What runs there | +| :------------------ | :---------- | :--------------------------------------------------------------------------------------------------------- | +| **Hub cluster** | 1 | `otterscale` chart: server, dashboard, Keycloak + PostgreSQL, Valkey, Harbor, and the Gateway API routing. | +| **Managed cluster** | 1 or more | `otterscale-agent` chart: agent, Tenant Operator, and Flux. | + +There is nothing to manage until at least one cluster has joined, so plan for both roles from the +start. Every cluster you want to see in the dashboard needs its own agent release, its own join +token, and its own CA Secret. + +## Cluster APIs and add-ons + +These are prerequisites in the strict sense: the charts render resources from these APIs and fail +without them. + +### On the hub cluster + +| Requirement | Why | +| :--------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Gateway API ≥ v1.5** | The chart publishes its listeners as a `ListenerSet` (`gateway.networking.k8s.io/v1`) rather than editing your `Gateway`, so several releases can share one Gateway. `ListenerSet` requires Gateway API v1.5. | +| **Envoy Gateway ≥ v1.8** | The `ListenerSet` has to be reconciled by a controller that understands it. | +| **cert-manager** (`cert-manager.io/v1`) | With the default `expose.listener.tls.certSource: auto`, the chart creates a self-signed `Issuer`, a CA `Certificate`, and the listener `Certificate`. | +| **A default `StorageClass`** (`ReadWriteOnce`) | Keycloak's PostgreSQL claims 20 GiB, and Harbor claims its own volumes for the registry, database, cache, job service and Trivy. | + +### On every managed cluster + +| Requirement | Why | +| :----------------------------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------- | +| **cert-manager** (`cert-manager.io/v1`) | The Tenant Operator ships its own `Issuer` and `Certificate`, and its admission webhooks are wired with `cert-manager.io/inject-ca-from`. | +| **`admissionregistration.k8s.io/v1`** with `ValidatingAdmissionPolicy` | The agent chart installs a policy that pins every `HelmRelease` to a fixed service account, so a workspace cannot deploy with more privilege than its own. | + + + +## Network + +### Reserved IP addresses + +Reserve **two** addresses in the hub cluster's subnet, outside any DHCP range: + +| Address | Used for | +| :--------------------- | :---------------------------------------------------------------------------------------------------------------------------- | +| **Gateway IP** | The `LoadBalancer` address of the Envoy Gateway proxy Service. Every browser-facing URL is built from it: dashboard, API, Keycloak and Harbor. | +| **Control-plane VIP** | The cluster's own API server VIP (for example the address `kube-vip` announces), so the control plane keeps one stable address. | + +Only the Gateway IP is consumed by OtterScale itself. It is fixed at install time, for the reason +in [Hostnames](#hostnames-and-tls) below, so pick one you will not have to move. + +### Ports on the hub cluster + +| Port | Reached at | Serves | +| :------------- | :---------------- | :----------------------------------------------------------------------------------------------------------------------- | +| **443** | Gateway IP | Dashboard, the API under `/api/`, and Keycloak under `/auth/`. | +| **8443** | Gateway IP | Harbor. It needs a listener of its own whenever it shares a hostname with the dashboard. | +| **30300** | A node address | The agent tunnel. A `NodePort` Service, deliberately *not* behind the Gateway: the tunnel is mTLS end to end, and a terminating HTTPS listener would break it. | + +### Ports on managed clusters + +None. Agents dial out to the hub and keep a reverse tunnel open, so a cluster behind NAT, a +corporate firewall, or in an otherwise closed network needs no inbound rule. What each managed +cluster does need is *outbound* reachability to the hub's 443 and 30300. + +## Hostnames and TLS + +The chart serves HTTPS only, and needs two browser-facing base URLs that differ **by hostname or by +port**: + +- `externalURL` covers the dashboard, the API and Keycloak. +- `harbor.externalURL` covers Harbor. + +Either shape works: + + + + `https://192.0.2.1` and `https://192.0.2.1:8443`. cert-manager mints a private CA and a + certificate covering both. Nothing external to arrange, but browsers will warn until you + trust the CA. + + + `https://otterscale.example.com` and `https://harbor.example.com`, with a certificate you + supply in a Secret. No warnings, and Harbor can share port 443. + + + + + +## GPUs + +Not required. The hub, the dashboard, workspaces and multi-cluster management all run on ordinary +CPU nodes. + +GPUs come in when you want to serve models. The inference stack (`gpu-operator`, `hami`, `kserve` +and friends) ships as separate module charts you install per cluster after the platform is up, and +those need at least one node with a supported GPU. If that is your goal, plan for it; if it is not, +skip it entirely. + +## Client tooling + +Run these from wherever you have `kubectl` access to the clusters. + + + +1. **`kubectl`**, with a context for the hub cluster and one for each cluster you will join. + +2. **Helm 3** with OCI support, since some dependencies are pulled from OCI registries. + +3. **`curl`, `jq` and `openssl`**, used by the script that provisions Harbor's robot account in + [Join a cluster](/getting-started/join-cluster/). + + + +## Next + +[Prepare the cluster](/getting-started/prepare-cluster/) covers cert-manager, a LoadBalancer +address, Envoy Gateway, and the Gateway that OtterScale attaches its listeners to. diff --git a/src/content/docs/getting-started/02-prepare-cluster.mdx b/src/content/docs/getting-started/02-prepare-cluster.mdx new file mode 100644 index 0000000..5f41d0d --- /dev/null +++ b/src/content/docs/getting-started/02-prepare-cluster.mdx @@ -0,0 +1,279 @@ +--- +title: Prepare the cluster +description: Install cert-manager, hand the cluster a LoadBalancer address, install Envoy Gateway, and create the Gateway that OtterScale attaches its listeners to. +slug: getting-started/prepare-cluster +sidebar: + order: 2 +--- + +import { Aside, Steps, Tabs, TabItem } from '@astrojs/starlight/components'; + +Four things have to exist on the hub cluster before the OtterScale chart will install: a certificate +issuer, a reachable `LoadBalancer` address, a Gateway API controller, and a `Gateway` for OtterScale +to attach to. This page sets all four up and verifies each one. + + + +## 1. Install cert-manager + +OtterScale's default TLS mode (`expose.listener.tls.certSource: auto`) does not ship +certificates; it asks cert-manager for them. The chart creates a self-signed `Issuer`, uses it to mint a private +CA, and then issues the listener certificate from that CA. That CA is also what agents and the +Tenant Operator are told to trust, so the whole trust chain starts here. + +```bash +helm install cert-manager \ + oci://quay.io/jetstack/charts/cert-manager \ + --version v1.21.1 \ + --namespace cert-manager --create-namespace \ + --set crds.enabled=true +``` + +`crds.enabled=true` matters: without the CRDs, the `Issuer` and `Certificate` resources the +OtterScale chart renders have nowhere to land, and `helm install` fails on an unknown kind. + +Any cert-manager release that serves `cert-manager.io/v1` will do; v1.21.1 is a known-good version. + +```bash +kubectl -n cert-manager rollout status deploy/cert-manager-webhook +``` + + + +## 2. Give the cluster a LoadBalancer address + +Envoy Gateway asks Kubernetes for a `Service` of type `LoadBalancer` (configured in step 4). In a +managed cloud, the cloud controller assigns an address and you can skip this step. On bare metal +nothing does, the Service sits at `` forever, and no OtterScale URL ever resolves. + +The reference OtterScale deployment runs Cilium, so the example below uses Cilium's own +load-balancer IPAM plus L2 announcements. Any implementation that can hand out an address works +equally well: MetalLB, kube-vip in service mode, or a cloud provider. + + + +1. **Declare the pool the address comes from.** + + A single-address block is deliberate: exactly one Service should ever claim this address, and a + wider range invites a second Service to take it. + + ```bash + kubectl apply -f - < + +## 3. Install Envoy Gateway + +OtterScale needs a Gateway API controller that reconciles `ListenerSet`, which means Envoy Gateway +v1.8 or newer. + +How you install it depends on whether you intend to serve models. Envoy AI Gateway needs Envoy +Gateway started with a specific extension-manager configuration, and that cannot be turned on +later without reinstalling, so choose now, even if the inference modules come much later. + + + + +The `otterscale/envoy-gateway` chart bundles Envoy Gateway, the Envoy AI Gateway controller and its +CRDs, already wired together: the extension-manager hooks AI Gateway needs, and `InferencePool` +registered as a backend resource. + +```bash +helm repo add otterscale https://otterscale.github.io/helm-charts +helm repo update + +helm upgrade --install otterscale-envoy-gateway otterscale/envoy-gateway \ + -n envoy-gateway-system --create-namespace +``` + +`InferencePool` itself is defined by the Gateway API Inference Extension, which is a separate +project and is not part of that chart. Install its manifests too: + +```bash +kubectl apply -f https://github.com/kubernetes-sigs/gateway-api-inference-extension/releases/download/v1.0.2/manifests.yaml +``` + + + + + + +If you only want the dashboard, multi-cluster management and Harbor, upstream Envoy Gateway on its +own is enough: + +```bash +helm upgrade --install otterscale-envoy-gateway \ + oci://docker.io/envoyproxy/gateway-helm --version v1.8.3 \ + -n envoy-gateway-system --create-namespace +``` + +Nothing on this path installs the Envoy AI Gateway or the Inference Extension. Adding inference +later means reinstalling Envoy Gateway with the extension-manager configuration, because it is read +at controller startup. + + + + +Wait for the controller, then confirm the API that OtterScale depends on is actually present: + +```bash +kubectl -n envoy-gateway-system rollout status deploy/envoy-gateway +``` + +```bash +kubectl get crd listenersets.gateway.networking.k8s.io +``` + + + +## 4. Create the GatewayClass, EnvoyProxy and Gateway + +The OtterScale chart does not create a `Gateway`. It hangs a `ListenerSet` off one that already +exists, which is what lets the dashboard, Harbor and anything else you route share a single address. +So the `Gateway` is yours to create, and OtterScale adopts it. + +```bash +kubectl apply -f - < + `PROGRAMMED=True` with no HTTPS listener yet is expected at this stage. The dashboard and Harbor + listeners appear in the next step, when the OtterScale chart adds its `ListenerSet`. + + +## Next + +[Install the server](/getting-started/install-server/) covers the hub itself: API, dashboard, +Keycloak and Harbor. diff --git a/src/content/docs/getting-started/03-install-server.mdx b/src/content/docs/getting-started/03-install-server.mdx new file mode 100644 index 0000000..d4bcd97 --- /dev/null +++ b/src/content/docs/getting-started/03-install-server.mdx @@ -0,0 +1,197 @@ +--- +title: Install the server +description: Install the OtterScale hub with Helm, including the API server, dashboard, Keycloak, Valkey and Harbor. +slug: getting-started/install-server +sidebar: + order: 3 +--- + +import { Aside, Steps, Tabs, TabItem } from '@astrojs/starlight/components'; + +The `otterscale` chart installs the whole hub in one release: + +| Component | What it does | +| :----------------------- | :---------------------------------------------------------------------------------------------------- | +| **Server** | The ConnectRPC API, and the tunnel listener agents dial into. | +| **Dashboard** | The SvelteKit web console. | +| **Keycloak + PostgreSQL**| The identity provider, seeded with a ready-made `otterscale` realm and its OIDC clients. | +| **Valkey** | Dashboard session storage. | +| **Harbor** | The container and chart registry, federated to Keycloak over OIDC. | +| **Gateway routing** | A `ListenerSet` on your Gateway, plus the `HTTPRoute`s that split `/`, `/api/` and `/auth/`. | +| **cert-manager objects** | A self-signed `Issuer`, a private CA, and the listener certificate, under the default `certSource: auto`. | + +## 1. Add the Helm repository + +```bash +helm repo add otterscale https://otterscale.github.io/helm-charts +helm repo update +``` + +## 2. Write the values file + +Three values are mandatory and the chart refuses to render without them. Everything else has a +working default. + +| Value | Meaning | +| :-------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `externalURL` | The browser-facing base URL. OIDC redirect URIs, CORS origins and the Keycloak realm URL are all built from it. Must be `https://`. | +| `harbor.externalURL` | Harbor's browser-facing base URL. Must be `https://`, and must differ from `externalURL` by hostname or by port. | +| `tunnel.externalURL` | The address agents dial for the tunnel. The tunnel certificate is issued for this hostname and agents verify it, so it must be exactly what they will connect to. | + +Pick the shape that matches your environment: + + + + +Nothing to arrange in advance: cert-manager mints a private CA and a certificate covering both +addresses. Harbor takes port `8443` because it shares the hostname with the dashboard. + +Replace `192.0.2.1` with the Gateway IP you reserved. Browsers will warn about the certificate +until you trust the CA; [First login](/getting-started/first-login/#trust-the-certificate) covers +that. + +```yaml +# otterscale-values.yaml +externalURL: 'https://192.0.2.1' + +harbor: + externalURL: 'https://192.0.2.1:8443' + +tunnel: + externalURL: 'https://192.0.2.1:30300' +``` + + + + +With two hostnames, Harbor can share port 443 and needs no listener of its own. Supply the +certificate in a Secret with `tls.crt` and `tls.key` keys, and point `trustedCA` at the issuing CA +if it is not publicly trusted. + +```yaml +# otterscale-values.yaml +externalURL: 'https://otterscale.example.com' + +expose: + listener: + tls: + certSource: secret + secret: + secretName: 'otterscale-tls' + +trustedCA: + secretName: '' # set to the CA Secret when otterscale-tls comes from an internal PKI + +harbor: + externalURL: 'https://harbor.example.com' + caBundleSecretName: '' # same value as trustedCA.secretName + +tunnel: + externalURL: 'https://192.0.2.1:30300' +``` + +`tunnel.externalURL` stays an address rather than a name unless you have a DNS record for the nodes: +it is a `NodePort`, reached directly rather than through the Gateway. + + + + + + +### Values you usually do not need to set + +- **`expose.gateway`** already defaults to `otterscale-gateway` in `envoy-gateway-system`, the + Gateway created in [Prepare the cluster](/getting-started/prepare-cluster/). Change it only if + you named yours differently. +- **`joinSecret`** is generated on first install and reused on upgrade. It is the root secret every + agent join token is derived from, so setting it by hand is only useful when you need the same + tokens across two installs. Changing it invalidates every token already handed out. +- **`trustedCA`** is resolved automatically under `certSource: auto`, becoming the CA the chart + just created. Set it explicitly only when you supply a certificate from a private PKI yourself. +- **`releaseVersion`** is cosmetic: it is what the dashboard displays as its version. Left empty, the + dashboard shows the server and dashboard image tags. + +## 3. Install + +```bash +helm upgrade --install otterscale otterscale/otterscale \ + -f otterscale-values.yaml \ + -n otterscale-system --create-namespace +``` + +The chart validates its inputs before rendering anything, so a missing or malformed value fails +immediately with a message naming the value, rather than half-installing and failing at runtime. + +## 4. Verify + + + +1. **Every pod is running.** + + Harbor and Keycloak take the longest; Keycloak imports its realm and + runs a schema migration on first start. + + ```bash + kubectl -n otterscale-system get pods + ``` + +2. **The listeners were accepted by your Gateway.** + + This is where a missing + `allowedListeners.namespaces.from: All` shows up. + + ```bash + kubectl -n otterscale-system get listenerset otterscale -o wide + ``` + +3. **The routes attached.** + + Two `HTTPRoute`s: one for the dashboard, API and Keycloak, one for + Harbor. + + ```bash + kubectl -n otterscale-system get httproute + ``` + +4. **The dashboard answers.** + + Expect a redirect to Keycloak, and a certificate warning under + `certSource: auto`. + + ```bash + curl -kIL https://192.0.2.1/ + ``` + + + +You now have a running hub, but no clusters in it. Do not log in yet: the first login is worth +doing after at least one cluster is joined, so the dashboard has something to show. + +## Operational notes + +Two properties of the server shape how you run it, and both are by design rather than limitations to +work around: + +- **Run a single replica.** Cluster registrations, allocated loopback addresses and live tunnel + sessions are held in the process that accepted them. A second replica would keep its own registry + and its own CA, so agents registered through one replica are unreachable through the other, and + requests routed to the wrong one fail with `cluster not registered`. +- **A restart re-keys every agent.** The tunnel CA is generated at startup and never persisted, so + certificates issued before a restart stop being trusted. Agents notice the dropped session and + re-register automatically with backoff, but their clusters are briefly unreachable. Expect a + short interruption after every restart or upgrade. + + + +## Next + +[Join a cluster](/getting-started/join-cluster/) covers issuing a join token, handing over the CA, +and installing the agent. diff --git a/src/content/docs/getting-started/04-join-cluster.mdx b/src/content/docs/getting-started/04-join-cluster.mdx new file mode 100644 index 0000000..5974a93 --- /dev/null +++ b/src/content/docs/getting-started/04-join-cluster.mdx @@ -0,0 +1,327 @@ +--- +title: Join a cluster +description: Issue a join token, hand over the CA, provision Harbor access, and install the OtterScale agent. +slug: getting-started/join-cluster +sidebar: + order: 4 +--- + +import { Aside, Steps, Tabs, TabItem } from '@astrojs/starlight/components'; + +Joining a cluster means installing the `otterscale-agent` chart in it. The agent dials out to the +hub, registers, and keeps a reverse tunnel open; from then on the hub can reach that cluster's +`kube-apiserver` without the cluster exposing anything. + +Three pieces of material have to travel from the hub to the joining cluster first: + +| Material | Why the agent needs it | +| :---------------------- | :------------------------------------------------------------------------------------------------- | +| **Join token** | Registration is the one endpoint an agent reaches before it has any credentials, so a token authorises it. | +| **CA certificate** | The agent verifies the hub with its image's system roots. A privately signed certificate has to be handed over out of band. | +| **Harbor robot account** | The Tenant Operator gives every workspace its own Harbor project, robot and image-pull secret, and needs an account of its own to do that. | + +Collect all three, then install once. + + + +## 1. Issue a join token + +The hub holds one root secret and derives each cluster's token from that secret plus the cluster +name. Issuing a token needs nothing but access to the secret, which is why it is done inside the +server pod, so the secret never leaves it. + +```bash +kubectl --context $HUB_CONTEXT -n otterscale-system \ + exec deploy/otterscale-server -- /otterscale join token --cluster $CLUSTER_NAME +``` + +The token is the only output, so it pipes straight into a `helm` invocation if you prefer. + + + +What the token does and does not give you: + +- **It authorises one cluster.** An agent holding `devel`'s token cannot register as `staging`, so a + compromised agent cannot take over another cluster's traffic. +- **A rejected token changes nothing.** The check runs before any state is touched, so a bad + registration cannot displace the agent currently serving that cluster. +- **Tokens do not expire and cannot be revoked individually.** Rotating the hub's `joinSecret` + invalidates all of them at once, after which every agent needs a new token. + +## 2. Hand over the CA + +Under the default `certSource: auto`, the hub's certificate is signed by a private CA that nothing +outside the cluster trusts yet. Export it from the server, which already has it mounted, and +create it as a Secret on the joining cluster. + +```bash +kubectl --context $HUB_CONTEXT -n otterscale-system \ + exec deploy/otterscale-server -- /otterscale join ca > ca.crt +``` + +```bash +kubectl --context $MANAGED_CONTEXT create namespace otterscale-system \ + --dry-run=client -o yaml | kubectl --context $MANAGED_CONTEXT apply -f - + +kubectl --context $MANAGED_CONTEXT -n otterscale-system \ + create secret generic otterscale-ca --from-file=ca.crt +``` + + + +If your hub serves a publicly trusted certificate, skip this step and leave `trustedCA.secretName` +empty. `otterscale join ca` will tell you when there is nothing to hand over. + +Keep `ca.crt` around: the next step uses it to talk to Harbor. + +## 3. Provision a Harbor robot account + +Every `Workspace` gets its own Harbor project, a robot scoped to that project, and an image-pull +secret built from it. The Tenant Operator does that work through Harbor's REST API, so it needs +credentials of its own: a **system-level robot account** on the hub's Harbor. + +Naming the robot after the cluster keeps them separable: revoking one cluster's access does not +touch another's. The script below creates it, or rotates its secret if it already exists, and prints +the credentials. Run it from the directory holding the `ca.crt` from step 2. + +```bash title="fetch-robot-secret.sh" +#!/usr/bin/env bash +set -euo pipefail + +HARBOR_URL="https://${GATEWAY_IP}:8443" +MASTER_NS=otterscale-system +ROBOT_NAME="${CLUSTER_NAME}" + +# Harbor's own admin account. Its password was generated at install time and +# kept in a Secret the chart does not delete on uninstall. +ADMIN_PASS=$(kubectl -n "$MASTER_NS" get secret otterscale-harbor-admin \ + -o jsonpath='{.data.HARBOR_ADMIN_PASSWORD}' | base64 -d) + +# Drop --cacert when Harbor serves a publicly trusted certificate. +api() { + curl -sfS --cacert ca.crt -u "admin:${ADMIN_PASS}" \ + -H 'Content-Type: application/json' "$@" +} + +# Robots are addressed cluster-wide even when scoped to a project, so the +# lookup is a filtered list rather than a path under the project. +ROBOT_ID=$(api "${HARBOR_URL}/api/v2.0/robots?q=name%3D${ROBOT_NAME}" \ + | jq -r --arg n "robot\$${ROBOT_NAME}" '.[] | select(.name==$n) | .id' | head -1) + +if [ -n "${ROBOT_ID}" ]; then + # Harbor reveals a robot secret only when it sets one, so re-reading an + # existing robot's credentials is impossible, so rotate instead. + # The A1a prefix satisfies Harbor's complexity rule for robot secrets. + SECRET="A1a$(openssl rand -hex 16)" + api -X PATCH "${HARBOR_URL}/api/v2.0/robots/${ROBOT_ID}" \ + -d "{\"secret\":\"${SECRET}\"}" > /dev/null +else + RESP=$(api -X POST "${HARBOR_URL}/api/v2.0/robots" -d @- < + Rotating a robot's secret is a `PATCH` on the robot, and Harbor defines no `robot:update` + permission at either scope, so no robot account can perform it, only an administrator. That is + also why the operator itself rebuilds a robot by deleting and re-creating it, which stays within + create/read/list/delete. + + + + +## 4. Write the agent values file + +```yaml +# agent-values.yaml +image: + repository: ghcr.io/otterscale/otterscale + tag: v1.5.0-rc.4 + pullPolicy: IfNotPresent + +agent: + serverURL: 'https://192.0.2.1/api/' + tunnelServerURL: 'https://192.0.2.1:30300' + cluster: 'devel' + joinToken: '' # from step 1 + +trustedCA: + secretName: 'otterscale-ca' # the Secret from step 2 + key: 'ca.crt' + +clusterAdmin: + enabled: true + users: [] # OIDC subjects to bind cluster-admin to, see below + +clusterInfo: + enabled: true + externalAddress: '192.0.2.1' + nodePortRange: '30000-32767' + inferenceURL: 'https://inference.example.com' + +tenantOperator: + enabled: true + harbor: + url: https://192.0.2.1:8443 + robot: + name: 'robot$devel' # from step 3 + secret: '' # from step 3 +``` + +What each value controls: + +- **`agent.serverURL`**: keep the `/api/` suffix. That is the path the Gateway rewrites onto the + server; without it the agent talks to the dashboard instead. +- **`agent.tunnelServerURL`**: must equal the hub's `tunnel.externalURL` exactly. The tunnel + certificate is issued for that hostname and the agent verifies it, so a mismatch fails the TLS + handshake at every agent rather than degrading quietly. +- **`agent.cluster`**: the name the join token is bound to. A token issued for another name is + rejected. +- **`agent.joinToken`**: inline, the token lands in the Helm release and is readable by anyone who + can read Secrets in the namespace. To keep it out, create the Secret yourself and point + `agent.existingSecret` at it instead. Either way the chart mounts it as a file rather than putting + it in the agent's environment. +- **`clusterAdmin.users`**: binds `cluster-admin` on *this* cluster to the listed users. Leave it + empty for now; [First login](/getting-started/first-login/) explains how to find your OIDC subject + and fill it in. Without at least one entry, nobody can create a `Workspace` here. +- **`clusterInfo`**: how users reach workloads on this cluster. The Kubernetes API cannot answer + that, so the answers are written to the `otterscale-info` ConfigMap in `kube-public`, which the + dashboard reads. `externalAddress` must be a bare address: no scheme, no port. `inferenceURL` is + optional; drop it if you are not serving models on this cluster. +- **`tenantOperator.harbor`**: the hub's Harbor and the robot from step 3. All three fields are + required when the operator is enabled: it refuses to run half-configured rather than leave + workspaces without registry access. + + + +## 5. Install + +The release namespace must be `otterscale-system` while the Tenant Operator is enabled, because its +install manifest is pre-rendered against that namespace. + +```bash +helm --kube-context $MANAGED_CONTEXT upgrade --install otterscale-agent \ + otterscale/otterscale-agent \ + -f agent-values.yaml \ + -n otterscale-system --create-namespace +``` + +## 6. Verify + + + +1. **The agent and operator are running.** + + ```bash + kubectl --context $MANAGED_CONTEXT -n otterscale-system get pods + ``` + +2. **Registration succeeded.** The agent logs the registration and then the tunnel it established. + + ```bash + kubectl --context $MANAGED_CONTEXT -n otterscale-system logs deploy/otterscale-agent + ``` + +3. **The dashboard's cluster facts landed.** + + ```bash + kubectl --context $MANAGED_CONTEXT -n kube-public get configmap otterscale-info -o yaml + ``` + +4. **The hub sees the cluster.** It appears in the dashboard's cluster switcher, which is the point + of [First login](/getting-started/first-login/). + + + +## Troubleshooting + +| Symptom | Likely cause | +| :----------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------- | +| Agent logs a TLS handshake failure against the tunnel | `agent.tunnelServerURL` does not match the hub's `tunnel.externalURL`. The certificate is issued for that name and the agent verifies it. | +| Agent logs an `x509` error while registering | The CA Secret is missing, or `trustedCA.secretName` does not point at it. | +| Registration rejected | The token does not match `agent.cluster`, or the hub's join secret has been rotated. | +| `cluster not registered` from the dashboard | The agent has not reconnected yet, or requests are reaching a second server replica. Run one replica. | +| Agent warns about plaintext HTTP at startup | `agent.serverURL` is `http://` to a remote host, which exposes the join token in transit. Legitimate only behind a mesh that terminates TLS. | +| Every cluster goes unreachable at once, then recovers | The hub restarted; its tunnel CA is regenerated at startup and agents must re-register. | +| `Workspace` stalls at `Ready=False` | Harbor credentials are wrong or the robot lacks a permission. Check the Tenant Operator's logs. | + +## Next + +[First login](/getting-started/first-login/) covers signing in, granting yourself cluster access, +and creating a workspace. diff --git a/src/content/docs/getting-started/05-first-login.mdx b/src/content/docs/getting-started/05-first-login.mdx new file mode 100644 index 0000000..07268e9 --- /dev/null +++ b/src/content/docs/getting-started/05-first-login.mdx @@ -0,0 +1,165 @@ +--- +title: First login +description: Sign in to the dashboard, grant yourself cluster access, and create your first workspace. +slug: getting-started/first-login +sidebar: + order: 5 +--- + +import { Aside, Steps, Card, CardGrid } from '@astrojs/starlight/components'; + +The hub is running and at least one cluster has joined. What remains is identity: signing in, and +giving that identity enough authority on the managed cluster to create a workspace. + +## Trust the certificate + +Under the default `certSource: auto` the hub serves a certificate from a private CA, so browsers +will warn. Install that CA, the same `ca.crt` you exported in +[Join a cluster](/getting-started/join-cluster/#2-hand-over-the-ca), into your OS or browser trust +store, or read it out of the hub directly: + +```bash +kubectl -n otterscale-system get secret otterscale-ca \ + -o jsonpath='{.data.ca\.crt}' | base64 -d > otterscale-ca.crt +``` + +Clicking through the warning works for a first look, but do trust the CA before real use: Harbor and +the dashboard exchange OIDC redirects, and a browser that distrusts one of the two hosts can fail +the login mid-flow. + +## Sign in + +Open `externalURL` in a browser. You are redirected to Keycloak, in the seeded `otterscale` realm. + +The realm ships with one user: + +| Field | Value | +| :----------- | :--------- | +| **Username** | `admin` | +| **Password** | `password` | + +Keycloak forces a password change on first sign-in, so this credential is only good once. + + + +That `admin` user starts with two things granted: + +- The `admin` client role on the dashboard client, which the API turns into the Kubernetes group + `oidc:admin`. +- Membership of the `harbor_admin` group, which makes it a Harbor administrator. + +### The Keycloak admin console + +For everything identity-related, such as creating users, groups, and turning off +self-registration, Keycloak's own console is at `/auth/`. It uses the **master** +realm's `admin` account, which is a different account from the realm user above: + +```bash +kubectl -n otterscale-system get secret otterscale-keycloak-admin \ + -o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d +``` + +## Grant yourself cluster access + +At this point the dashboard lists your joined cluster but you cannot create anything in it. Every +request reaches the managed cluster's `kube-apiserver` impersonating *you*, so that cluster's own +RBAC decides what happens, and nothing has granted your identity anything yet. + +The agent chart's `clusterAdmin.users` exists for exactly this. It takes **OIDC subjects**, not +usernames: the subject is what the API sends as `Impersonate-User`, and for Keycloak that is the +user's internal ID. + + + +1. **Find your subject.** In the Keycloak admin console, switch to the `otterscale` realm, open + **Users**, select your user, and copy the **ID** field, which is a UUID. + +2. **Add it to the agent's values.** + + ```yaml + # agent-values.yaml + clusterAdmin: + enabled: true + users: + - '3f9c1e02-7a45-4b8e-9d21-6c0f5ab7e134' + ``` + +3. **Upgrade the agent release.** + + ```bash + helm --kube-context $MANAGED_CONTEXT upgrade --install otterscale-agent \ + otterscale/otterscale-agent \ + -f agent-values.yaml \ + -n otterscale-system + ``` + +4. **Confirm the binding.** + + ```bash + kubectl --context $MANAGED_CONTEXT get clusterrolebinding otterscale-agent-cluster-admin -o yaml + ``` + + + + + +## Create a workspace + +A `Workspace` is the multi-tenant unit: the Tenant Operator turns one into a namespace with Pod +Security Standards labels, RBAC bindings per member, resource quotas and limit ranges, default-deny +network policies, a Harbor project with its own robot and image-pull secret, and a Flux +`HelmRepository` pointing at that project. + +In the dashboard, pick your cluster and create a workspace from the workspace switcher. You are +added as its first `admin` member automatically. + +Two authorisation layers apply, and it helps to know which one is complaining: + +- **Kubernetes RBAC** must allow you to create `workspaces.tenant.otterscale.io` at all. That is + what the `clusterAdmin.users` binding above provides. +- **The operator's admission webhook** then checks that you are either listed as an `admin` member + of the workspace being created, or hold cluster-wide access. Creating a workspace you are not an + admin of is rejected with *"workspace creator must be listed as a member with the 'admin' role"*. + +The same rule governs updates and deletes, evaluated against the *stored* spec, so you cannot +grant yourself admin and approve it in the same request. + + + +## What's next + + + + The dashboard's **Modules** page lists charts from a Flux `HelmRepository` named `modules` in + `otterscale-system` on the target cluster. Neither chart creates it, so point one at + `https://otterscale.github.io/helm-charts` to offer GPU Operator, KubeVirt, KServe, + Prometheus and the rest. + + + The agent proxies read-only Prometheus queries for the dashboard's metrics. Its default + target is `http://otterscale-prometheus-kube-prometheus.monitoring.svc:9090`. Install the + `prometheus-stack` module, or point `--proxy-prometheus-url` at a Prometheus you already run. + + + Repeat [Join a cluster](/getting-started/join-cluster/) for each one. Every cluster needs its + own join token, its own CA Secret, and its own Harbor robot. + + + The ConnectRPC services behind the dashboard are documented under **API**, generated from + OtterScale's OpenAPI schema. + + diff --git a/src/content/docs/introduction.mdx b/src/content/docs/introduction.mdx index 80ef17a..11fd815 100644 --- a/src/content/docs/introduction.mdx +++ b/src/content/docs/introduction.mdx @@ -91,7 +91,31 @@ The OtterScale platform is composed of several open-source components: ## Getting Started -Installation, configuration, and operational guides are coming soon as part of this documentation. In the meantime: + + + + + + + +Along the way: -- Run `otterscale server --help` and `otterscale agent --help` to explore the available options. -- Add the Helm repository: `helm repo add otterscale https://otterscale.github.io/charts` +- `helm repo add otterscale https://otterscale.github.io/helm-charts` adds the chart repository. +- `otterscale server --help` and `otterscale agent --help` are the authoritative reference for every + flag and environment variable. From c5b31d13b84dace3477ad09177e4fbae2cc93a81 Mon Sep 17 00:00:00 2001 From: Cody Chen Date: Mon, 7 Sep 2026 14:12:13 +0800 Subject: [PATCH 2/4] fix(doc): split self-managed and multi-cluster installation page --- .../docs/getting-started/01-requirements.mdx | 182 ++++++----- .../getting-started/02-prepare-cluster.mdx | 30 +- .../getting-started/03-install-server.mdx | 38 ++- .../docs/getting-started/04-self-managed.mdx | 297 ++++++++++++++++++ ...-join-cluster.mdx => 05-multi-cluster.mdx} | 153 ++++----- ...{05-first-login.mdx => 06-first-login.mdx} | 44 +-- src/content/docs/introduction.mdx | 15 +- 7 files changed, 568 insertions(+), 191 deletions(-) create mode 100644 src/content/docs/getting-started/04-self-managed.mdx rename src/content/docs/getting-started/{04-join-cluster.mdx => 05-multi-cluster.mdx} (54%) rename src/content/docs/getting-started/{05-first-login.mdx => 06-first-login.mdx} (76%) diff --git a/src/content/docs/getting-started/01-requirements.mdx b/src/content/docs/getting-started/01-requirements.mdx index 1502daf..2b0d357 100644 --- a/src/content/docs/getting-started/01-requirements.mdx +++ b/src/content/docs/getting-started/01-requirements.mdx @@ -1,40 +1,55 @@ --- title: Requirements -description: What you need before installing OtterScale, from clusters and cluster APIs to reserved IPs, storage, and client tooling. +description: Choose a deployment mode, then check the cluster APIs, addresses and ports that mode needs. slug: getting-started/requirements sidebar: order: 1 --- -import { Aside, Steps, Card, CardGrid } from '@astrojs/starlight/components'; +import { Aside, Steps, Card, CardGrid, LinkCard, Badge } from '@astrojs/starlight/components'; -OtterScale is a hub-and-spoke platform: one **server** (the hub) holds the API, the dashboard, the -identity provider and the registry, and one lightweight **agent** (a spoke) runs inside every -cluster you want to manage. Requirements differ between the two roles, so this page separates them. +OtterScale is one binary in two roles. A **server** (the hub) holds the API, the dashboard, the +identity provider and the registry. A lightweight **agent** (a spoke) runs inside every cluster you +want to manage and dials out to the hub. - +Both roles can live on one cluster, or the hub can sit apart from the clusters it manages. That +choice changes what you need to reserve, so make it first. + +## Deployment modes + +| | Self-managed | Multi-cluster | +| :--------------------- | :-------------------------------------------------------------- | :----------------------------------------------------------------------- | +| **Clusters** | 1 | 1 hub, plus 1 or more managed clusters | +| **Server runs on** | The one cluster | The hub cluster | +| **Agent runs on** | The same cluster | Each managed cluster | +| **Kubectl contexts** | 1 | 1 per cluster | +| **Good for** | Evaluation, a single site, and any deployment where one cluster is the whole estate | Managing clusters that sit behind NAT, a firewall, or in another network | -## Clusters +Both are first-class topologies, and neither is a stepping stone to the other. A self-managed +cluster is not a partial multi-cluster install, and a hub that manages only remote clusters never +needs an agent of its own. -| Role | Count | What runs there | -| :------------------ | :---------- | :--------------------------------------------------------------------------------------------------------- | -| **Hub cluster** | 1 | `otterscale` chart: server, dashboard, Keycloak + PostgreSQL, Valkey, Harbor, and the Gateway API routing. | -| **Managed cluster** | 1 or more | `otterscale-agent` chart: agent, Tenant Operator, and Flux. | + + + + -There is nothing to manage until at least one cluster has joined, so plan for both roles from the -start. Every cluster you want to see in the dashboard needs its own agent release, its own join -token, and its own CA Secret. +## What both modes need -## Cluster APIs and add-ons +### Cluster APIs and add-ons These are prerequisites in the strict sense: the charts render resources from these APIs and fail without them. -### On the hub cluster +On the cluster that runs the **server**: | Requirement | Why | | :--------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------- | @@ -43,48 +58,21 @@ without them. | **cert-manager** (`cert-manager.io/v1`) | With the default `expose.listener.tls.certSource: auto`, the chart creates a self-signed `Issuer`, a CA `Certificate`, and the listener `Certificate`. | | **A default `StorageClass`** (`ReadWriteOnce`) | Keycloak's PostgreSQL claims 20 GiB, and Harbor claims its own volumes for the registry, database, cache, job service and Trivy. | -### On every managed cluster +On every cluster that runs an **agent**: | Requirement | Why | | :----------------------------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------- | | **cert-manager** (`cert-manager.io/v1`) | The Tenant Operator ships its own `Issuer` and `Certificate`, and its admission webhooks are wired with `cert-manager.io/inject-ca-from`. | | **`admissionregistration.k8s.io/v1`** with `ValidatingAdmissionPolicy` | The agent chart installs a policy that pins every `HelmRelease` to a fixed service account, so a workspace cannot deploy with more privilege than its own. | - -## GPUs - -Not required. The hub, the dashboard, workspaces and multi-cluster management all run on ordinary -CPU nodes. - -GPUs come in when you want to serve models. The inference stack (`gpu-operator`, `hami`, `kserve` -and friends) ships as separate module charts you install per cluster after the platform is up, and -those need at least one node with a supported GPU. If that is your goal, plan for it; if it is not, -skip it entirely. +### Client tooling -## Client tooling - -Run these from wherever you have `kubectl` access to the clusters. +Run these from wherever you have `kubectl` access. -1. **`kubectl`**, with a context for the hub cluster and one for each cluster you will join. +1. **`kubectl`**, with a context for the server's cluster and, in multi-cluster mode, one for each + cluster you will join. 2. **Helm 3** with OCI support, since some dependencies are pulled from OCI registries. -3. **`curl`, `jq` and `openssl`**, used by the script that provisions Harbor's robot account in - [Join a cluster](/getting-started/join-cluster/). +3. **`curl`, `jq` and `openssl`**, used by the script that provisions Harbor's robot account. +## Self-managed mode: what to reserve + +One cluster, so everything is local. Reserve **two** addresses in its subnet, outside any DHCP +range, and note one existing node address: + +| Address | Reserved? | Used for | +| :-------------------- | :-------- | :--------------------------------------------------------------------------------------------------------------------------------------------- | +| **Gateway IP** | Yes | The `LoadBalancer` address of the Envoy Gateway proxy Service. Every browser-facing URL is built from it: dashboard, API, Keycloak and Harbor. | +| **Control-plane VIP** | Yes | The cluster's own API server VIP, for example the address `kube-vip` announces, so the control plane keeps one stable address. | +| **A node address** | No | The agent tunnel on `NodePort` 30300, and the address the dashboard builds `NodePort` workload URLs from. | + +Only the Gateway IP is consumed by OtterScale itself, and it is fixed at install time, so pick one +you will not have to move. + +Ports, all on the one cluster: + +| Port | Reached at | Serves | +| :--------- | :----------- | :------------------------------------------------------------------------------------------------------------------------- | +| **443** | Gateway IP | Dashboard, the API under `/api/`, and Keycloak under `/auth/`. | +| **8443** | Gateway IP | Harbor. It needs a listener of its own whenever it shares a hostname with the dashboard. | +| **30300** | Node address | The agent tunnel. A `NodePort` Service, deliberately *not* behind the Gateway: the tunnel is mTLS end to end, and a terminating HTTPS listener would break it. | + +The agent is a pod on this same cluster, so it has to reach the cluster's own Gateway IP on 443 and +a node address on 30300 from inside the cluster. Nothing has to be reachable from outside except by +the people using the dashboard. + +## Multi-cluster mode: what to reserve + +### On the hub cluster + +The same two reserved addresses, the same node address, and the same three ports as above. The hub +is configured identically in both modes; what changes is who connects to it. + +### On every managed cluster + +| Requirement | Detail | +| :-------------------- | :-------------------------------------------------------------------------------------------------------------- | +| **Inbound ports** | None. Agents dial out and keep a reverse tunnel open, so a cluster behind NAT or a firewall needs no ingress. | +| **Outbound access** | To the hub's Gateway IP on 443, and to the hub's node address on 30300. | +| **A node address** | Not reserved. The dashboard builds this cluster's `NodePort` workload URLs from it. | +| **cert-manager** | Installed before the agent, per the table above. | + +Each managed cluster also needs its own join token, its own copy of the hub's CA, and its own Harbor +robot account. None of the three is shared between clusters, so revoking one cluster's access leaves +the others untouched. + +## GPUs + + + +**No GPU is required to install or run OtterScale.** The server, the dashboard, workspaces and +multi-cluster management all run on ordinary CPU nodes, in either deployment mode. Nothing in the +`otterscale` or `otterscale-agent` charts asks for a GPU. + +GPUs become a requirement only when you want to serve models. The inference stack (`gpu-operator`, +`hami`, `kserve` and the rest) ships as separate module charts you install per cluster after the +platform is up, and those need at least one node with a supported GPU in the cluster that will run +the workloads. + +If serving models is your goal, plan for it now: it also decides how you install Envoy Gateway in +[Prepare the cluster](/getting-started/prepare-cluster/), which cannot be changed later without +reinstalling. If it is not, skip it entirely. + ## Next [Prepare the cluster](/getting-started/prepare-cluster/) covers cert-manager, a LoadBalancer -address, Envoy Gateway, and the Gateway that OtterScale attaches its listeners to. +address, Envoy Gateway, and the Gateway that OtterScale attaches its listeners to. Both modes need +it, on the cluster that will run the server. diff --git a/src/content/docs/getting-started/02-prepare-cluster.mdx b/src/content/docs/getting-started/02-prepare-cluster.mdx index 5f41d0d..28cd777 100644 --- a/src/content/docs/getting-started/02-prepare-cluster.mdx +++ b/src/content/docs/getting-started/02-prepare-cluster.mdx @@ -8,13 +8,17 @@ sidebar: import { Aside, Steps, Tabs, TabItem } from '@astrojs/starlight/components'; -Four things have to exist on the hub cluster before the OtterScale chart will install: a certificate -issuer, a reachable `LoadBalancer` address, a Gateway API controller, and a `Gateway` for OtterScale -to attach to. This page sets all four up and verifies each one. +Four things have to exist before the OtterScale chart will install: a certificate issuer, a +reachable `LoadBalancer` address, a Gateway API controller, and a `Gateway` for OtterScale to attach +to. This page sets all four up and verifies each one. + +Everything here happens on the cluster that will run the **server**. In self-managed mode that is +your only cluster; in multi-cluster mode it is the hub, and managed clusters need none of it apart +from cert-manager. + +## 3. Write the agent values file + +```yaml +# agent-values.yaml +image: + repository: ghcr.io/otterscale/otterscale + tag: v1.5.0-rc.4 + pullPolicy: IfNotPresent + +agent: + serverURL: 'https://192.0.2.1/api/' # the Gateway IP + tunnelServerURL: 'https://192.0.2.10:30300' # a node address + cluster: 'devel' + joinToken: '' # from step 1 + +trustedCA: + secretName: 'otterscale-ca' # already present, created by the server release + key: 'ca.crt' + +clusterAdmin: + enabled: true + users: [] # OIDC subjects to bind cluster-admin to, see below + +clusterInfo: + enabled: true + externalAddress: '192.0.2.10' # a node address, for NodePort workload URLs + nodePortRange: '30000-32767' + +tenantOperator: + enabled: true + harbor: + url: https://192.0.2.1:8443 + robot: + name: 'robot$devel' # from step 2 + secret: '' # from step 2 +``` + +What each value controls: + +- **`agent.serverURL`**: keep the `/api/` suffix. That is the path the Gateway rewrites onto the + server; without it the agent talks to the dashboard instead. The agent is a pod on this cluster, + so it reaches the cluster's own Gateway address. +- **`agent.tunnelServerURL`**: must equal the server's `tunnel.externalURL` exactly, which is a + node address rather than the Gateway IP. The tunnel certificate is issued for that host and the + agent verifies it, so a mismatch fails the TLS handshake rather than degrading quietly. +- **`agent.cluster`**: the name the join token is bound to. A token issued for another name is + rejected. +- **`agent.joinToken`**: inline, the token lands in the Helm release and is readable by anyone who + can read Secrets in the namespace. To keep it out, create the Secret yourself and point + `agent.existingSecret` at it instead. Either way the chart mounts it as a file rather than putting + it in the agent's environment. +- **`clusterAdmin.users`**: binds `cluster-admin` on this cluster to the listed users. Leave it + empty for now; [First login](/getting-started/first-login/) explains how to find your OIDC subject + and fill it in. Without at least one entry, nobody can create a `Workspace`. +- **`clusterInfo`**: how users reach workloads here. The Kubernetes API cannot answer that, so the + answers are written to the `otterscale-info` ConfigMap in `kube-public`, which the dashboard + reads. `externalAddress` must be a bare address: no scheme, no port. Add `inferenceURL` only if + you are serving models. +- **`tenantOperator.harbor`**: the Harbor this cluster already runs, and the robot from step 2. All + three fields are required when the operator is enabled: it refuses to run half-configured rather + than leave workspaces without registry access. + + + +## 4. Install + +```bash +helm upgrade --install otterscale-agent otterscale/otterscale-agent \ + -f agent-values.yaml \ + -n otterscale-system +``` + +No `--create-namespace`: the server release already created `otterscale-system`. + +## 5. Verify + + + +1. **The agent and operator are running**, alongside the server components. + + ```bash + kubectl -n otterscale-system get pods + ``` + +2. **Registration succeeded.** The agent logs the registration and then the tunnel it established. + + ```bash + kubectl -n otterscale-system logs deploy/otterscale-agent + ``` + +3. **The dashboard's cluster facts landed.** + + ```bash + kubectl -n kube-public get configmap otterscale-info -o yaml + ``` + + + +## Troubleshooting + +| Symptom | Likely cause | +| :---------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------- | +| Agent logs a TLS handshake failure against the tunnel | `agent.tunnelServerURL` does not match `tunnel.externalURL`. The certificate is issued for that host and the agent verifies it. | +| Agent cannot reach the tunnel at all | The node address is wrong, or `NodePort` 30300 is not reachable from inside the cluster. | +| Agent logs an `x509` error while registering | `trustedCA.secretName` is empty or misspelled. It must be `otterscale-ca`, the Secret the server release created. | +| Registration rejected | The token does not match `agent.cluster`, or the server's join secret has been rotated. | +| `cluster not registered` from the dashboard | The agent has not connected yet, or the server is running more than one replica. Run one. | +| The cluster goes unreachable, then recovers | The server restarted; its tunnel CA is regenerated at startup and the agent must re-register. | +| `Workspace` stalls at `Ready=False` | Harbor credentials are wrong or the robot lacks a permission. Check the Tenant Operator's logs. | + +## Next + +[First login](/getting-started/first-login/) covers signing in, granting yourself cluster access, +and creating a workspace. + +To bring further clusters under this same server later, follow +[Multi-cluster](/getting-started/multi-cluster/). Nothing installed here has to change. diff --git a/src/content/docs/getting-started/04-join-cluster.mdx b/src/content/docs/getting-started/05-multi-cluster.mdx similarity index 54% rename from src/content/docs/getting-started/04-join-cluster.mdx rename to src/content/docs/getting-started/05-multi-cluster.mdx index 5974a93..f96e437 100644 --- a/src/content/docs/getting-started/04-join-cluster.mdx +++ b/src/content/docs/getting-started/05-multi-cluster.mdx @@ -1,33 +1,44 @@ --- -title: Join a cluster -description: Issue a join token, hand over the CA, provision Harbor access, and install the OtterScale agent. -slug: getting-started/join-cluster +title: Multi-cluster +description: Join separate clusters to the hub, each with its own join token, CA copy and Harbor robot. +slug: getting-started/multi-cluster sidebar: - order: 4 + order: 5 --- -import { Aside, Steps, Tabs, TabItem } from '@astrojs/starlight/components'; +import { Aside, Steps } from '@astrojs/starlight/components'; -Joining a cluster means installing the `otterscale-agent` chart in it. The agent dials out to the -hub, registers, and keeps a reverse tunnel open; from then on the hub can reach that cluster's -`kube-apiserver` without the cluster exposing anything. +In a multi-cluster deployment the **hub** runs the server release from +[Install the server](/getting-started/install-server/), and every cluster you want to manage runs an +agent release of its own. The agent dials out to the hub and keeps a reverse tunnel open, so a +cluster behind NAT, a corporate firewall, or in an otherwise closed network needs no inbound rule. -Three pieces of material have to travel from the hub to the joining cluster first: +Work through this page once per managed cluster. Nothing is shared between them: each gets its own +join token, its own copy of the hub's CA, and its own Harbor robot, so revoking one cluster's access +leaves the others untouched. -| Material | Why the agent needs it | -| :---------------------- | :------------------------------------------------------------------------------------------------- | -| **Join token** | Registration is the one endpoint an agent reaches before it has any credentials, so a token authorises it. | -| **CA certificate** | The agent verifies the hub with its image's system roots. A privately signed certificate has to be handed over out of band. | +Three pieces of material have to travel from the hub to each joining cluster: + +| Material | Why the agent needs it | +| :----------------------- | :---------------------------------------------------------------------------------------------------------------------------- | +| **Join token** | Registration is the one endpoint an agent reaches before it has any credentials, so a token authorises it. | +| **CA certificate** | The agent verifies the hub with its image's system roots, so a privately signed certificate has to be handed over out of band. | | **Harbor robot account** | The Tenant Operator gives every workspace its own Harbor project, robot and image-pull secret, and needs an account of its own to do that. | -Collect all three, then install once. +