diff --git a/astro.config.mjs b/astro.config.mjs index cb42916..7b33e3d 100644 --- a/astro.config.mjs +++ b/astro.config.mjs @@ -43,6 +43,15 @@ export default defineConfig({ 'zh-Hant': '簡介' } }, + { + label: 'Getting Started', + translations: { + 'zh-Hant': '開始使用' + }, + autogenerate: { + directory: 'getting-started' + } + }, ...openAPISidebarGroups ], customCss: ['./src/styles/global.css'], diff --git a/src/content/docs/getting-started/01-requirements.mdx b/src/content/docs/getting-started/01-requirements.mdx new file mode 100644 index 0000000..682fe26 --- /dev/null +++ b/src/content/docs/getting-started/01-requirements.mdx @@ -0,0 +1,173 @@ +--- +title: Requirements +description: Choose a deployment mode, then check the cluster APIs, addresses and ports that mode needs. +slug: getting-started/requirements +sidebar: + order: 1 +--- + +import { Aside, Steps, Card, CardGrid, LinkCard, Badge } from '@astrojs/starlight/components'; + +OtterScale is one binary in two roles. A **server** (the hub) holds the API, the dashboard, the +identity provider and the registry. A lightweight **agent** (a spoke) runs inside every cluster you +want to manage and dials out to the hub. + +Both roles can live on one cluster, or the hub can sit apart from the clusters it manages. That +choice changes what you need to reserve, so make it first. + +## Deployment modes + +| | Self-managed | Multi-cluster | +| :--------------------- | :-------------------------------------------------------------- | :----------------------------------------------------------------------- | +| **Clusters** | 1 | 1 hub, plus 1 or more managed clusters | +| **Server runs on** | The one cluster | The hub cluster | +| **Agent runs on** | The same cluster | Each managed cluster | +| **Kubectl contexts** | 1 | 1 per cluster | +| **Good for** | Evaluation, a single site, and any deployment where one cluster is the whole estate | Managing clusters that sit behind NAT, a firewall, or in another network | + +Both are first-class topologies, and neither is a stepping stone to the other. A self-managed +cluster is not a partial multi-cluster install, and a hub that manages only remote clusters never +needs an agent of its own. + +## What both modes need + +### Cluster APIs and add-ons + +These are prerequisites in the strict sense: the charts render resources from these APIs and fail +without them. + +On the cluster that runs the **server**: + +| Requirement | Why | +| :--------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Gateway API ≥ v1.5** | The chart publishes its listeners as a `ListenerSet` (`gateway.networking.k8s.io/v1`) rather than editing your `Gateway`, so several releases can share one Gateway. `ListenerSet` requires Gateway API v1.5. | +| **Envoy Gateway ≥ v1.8** | The `ListenerSet` has to be reconciled by a controller that understands it. | +| **cert-manager** (`cert-manager.io/v1`) | With the default `expose.listener.tls.certSource: auto`, the chart creates a self-signed `Issuer`, a CA `Certificate`, and the listener `Certificate`. | +| **A default `StorageClass`** (`ReadWriteOnce`) | Keycloak's PostgreSQL claims 20 GiB, and Harbor claims its own volumes for the registry, database, cache, job service and Trivy. | + +On every cluster that runs an **agent**: + +| Requirement | Why | +| :----------------------------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------- | +| **cert-manager** (`cert-manager.io/v1`) | The Tenant Operator ships its own `Issuer` and `Certificate`, and its admission webhooks are wired with `cert-manager.io/inject-ca-from`. | +| **`admissionregistration.k8s.io/v1`** with `ValidatingAdmissionPolicy` | The agent chart installs a policy that pins every `HelmRelease` to a fixed service account, so a workspace cannot deploy with more privilege than its own. | + + + +### Hostnames and TLS + +The chart serves HTTPS only, and needs two browser-facing base URLs that differ **by hostname or by +port**: + +- `externalURL` covers the dashboard, the API and Keycloak. +- `harbor.externalURL` covers Harbor. + +Either shape works: + + + + `https://192.0.2.1` and `https://192.0.2.1:8443`. cert-manager mints a private CA and a + certificate covering both. Nothing external to arrange, but browsers will warn until you + trust the CA. + + + `https://otterscale.example.com` and `https://harbor.example.com`, with a certificate you + supply in a Secret. No warnings, and Harbor can share port 443. + + + + + +### Client tooling + +Run these from wherever you have `kubectl` access. + + + +1. **`kubectl`**, with a context for the server's cluster and, in multi-cluster mode, one for each + cluster you will join. + +2. **Helm 3** with OCI support, since some dependencies are pulled from OCI registries. + +3. **`curl`, `jq` and `openssl`**, used by the script that provisions Harbor's robot account. + + + +## Self-managed mode: what to reserve + +One cluster, so everything is local. Reserve **two** addresses in its subnet, outside any DHCP +range, and note one existing node address: + +| Address | Reserved? | Used for | +| :-------------------- | :-------- | :--------------------------------------------------------------------------------------------------------------------------------------------- | +| **Gateway IP** | Yes | The `LoadBalancer` address of the Envoy Gateway proxy Service. Every browser-facing URL is built from it: dashboard, API, Keycloak and Harbor. | +| **Control-plane VIP** | Yes | The cluster's own API server VIP, for example the address `kube-vip` announces, so the control plane keeps one stable address. | +| **A node address** | No | The agent tunnel on `NodePort` 30300, and the address the dashboard builds `NodePort` workload URLs from. | + +Only the Gateway IP is consumed by OtterScale itself, and it is fixed at install time, so pick one +you will not have to move. + +Ports, all on the one cluster: + +| Port | Reached at | Serves | +| :--------- | :----------- | :------------------------------------------------------------------------------------------------------------------------- | +| **443** | Gateway IP | Dashboard, the API under `/api/`, and Keycloak under `/auth/`. | +| **8443** | Gateway IP | Harbor. It needs a listener of its own whenever it shares a hostname with the dashboard. | +| **30300** | Node address | The agent tunnel. A `NodePort` Service, deliberately *not* behind the Gateway: the tunnel is mTLS end to end, and a terminating HTTPS listener would break it. | + +The agent is a pod on this same cluster, so it has to reach the cluster's own Gateway IP on 443 and +a node address on 30300 from inside the cluster. Nothing has to be reachable from outside except by +the people using the dashboard. + +## Multi-cluster mode: what to reserve + +### On the hub cluster + +The same two reserved addresses, the same node address, and the same three ports as above. The hub +is configured identically in both modes; what changes is who connects to it. + +### On every managed cluster + +| Requirement | Detail | +| :-------------------- | :-------------------------------------------------------------------------------------------------------------- | +| **Inbound ports** | None. Agents dial out and keep a reverse tunnel open, so a cluster behind NAT or a firewall needs no ingress. | +| **Outbound access** | To the hub's Gateway IP on 443, and to the hub's node address on 30300. | +| **A node address** | Not reserved. The dashboard builds this cluster's `NodePort` workload URLs from it. | +| **cert-manager** | Installed before the agent, per the table above. | + +Each managed cluster also needs its own join token, its own copy of the hub's CA, and its own Harbor +robot account. None of the three is shared between clusters, so revoking one cluster's access leaves +the others untouched. + +## GPUs + + + +**No GPU is required to install or run OtterScale.** The server, the dashboard, workspaces and +multi-cluster management all run on ordinary CPU nodes, in either deployment mode. Nothing in the +`otterscale` or `otterscale-agent` charts asks for a GPU. + +GPUs become a requirement only when you want to serve models. The inference stack (`gpu-operator`, +`hami`, `kserve` and the rest) ships as separate module charts you install per cluster after the +platform is up, and those need at least one node with a supported GPU in the cluster that will run +the workloads. + +If serving models is your goal, plan for it now: it also decides how you install Envoy Gateway in +[Prepare the cluster](/getting-started/prepare-cluster/), which cannot be changed later without +reinstalling. If it is not, skip it entirely. + +## Next + +[Prepare the cluster](/getting-started/prepare-cluster/) covers cert-manager, a LoadBalancer +address, Envoy Gateway, and the Gateway that OtterScale attaches its listeners to. Both modes need +it, on the cluster that will run the server. diff --git a/src/content/docs/getting-started/02-prepare-cluster.mdx b/src/content/docs/getting-started/02-prepare-cluster.mdx new file mode 100644 index 0000000..28cd777 --- /dev/null +++ b/src/content/docs/getting-started/02-prepare-cluster.mdx @@ -0,0 +1,283 @@ +--- +title: Prepare the cluster +description: Install cert-manager, hand the cluster a LoadBalancer address, install Envoy Gateway, and create the Gateway that OtterScale attaches its listeners to. +slug: getting-started/prepare-cluster +sidebar: + order: 2 +--- + +import { Aside, Steps, Tabs, TabItem } from '@astrojs/starlight/components'; + +Four things have to exist before the OtterScale chart will install: a certificate issuer, a +reachable `LoadBalancer` address, a Gateway API controller, and a `Gateway` for OtterScale to attach +to. This page sets all four up and verifies each one. + +Everything here happens on the cluster that will run the **server**. In self-managed mode that is +your only cluster; in multi-cluster mode it is the hub, and managed clusters need none of it apart +from cert-manager. + + + +## 1. Install cert-manager + +OtterScale's default TLS mode (`expose.listener.tls.certSource: auto`) does not ship certificates; +it asks cert-manager for them. The chart creates a self-signed `Issuer`, uses it to mint a private +CA, and then issues the listener certificate from that CA. That CA is also what agents and the +Tenant Operator are told to trust, so the whole trust chain starts here. + +```bash +helm install cert-manager \ + oci://quay.io/jetstack/charts/cert-manager \ + --version v1.21.1 \ + --namespace cert-manager --create-namespace \ + --set crds.enabled=true +``` + +`crds.enabled=true` matters: without the CRDs, the `Issuer` and `Certificate` resources the +OtterScale chart renders have nowhere to land, and `helm install` fails on an unknown kind. + +Any cert-manager release that serves `cert-manager.io/v1` will do; v1.21.1 is a known-good version. + +```bash +kubectl -n cert-manager rollout status deploy/cert-manager-webhook +``` + + + +## 2. Give the cluster a LoadBalancer address + +Envoy Gateway asks Kubernetes for a `Service` of type `LoadBalancer` (configured in step 4). In a +managed cloud, the cloud controller assigns an address and you can skip this step. On bare metal +nothing does, the Service sits at `` forever, and no OtterScale URL ever resolves. + +The reference OtterScale deployment runs Cilium, so the example below uses Cilium's own +load-balancer IPAM plus L2 announcements. Any implementation that can hand out an address works +equally well: MetalLB, kube-vip in service mode, or a cloud provider. + + + +1. **Declare the pool the address comes from.** + + A single-address block is deliberate: exactly one Service should ever claim this address, and a + wider range invites a second Service to take it. + + ```bash + kubectl apply -f - < + +## 3. Install Envoy Gateway + +OtterScale needs a Gateway API controller that reconciles `ListenerSet`, which means Envoy Gateway +v1.8 or newer. + +How you install it depends on whether you intend to serve models. Envoy AI Gateway needs Envoy +Gateway started with a specific extension-manager configuration, and that cannot be turned on +later without reinstalling, so choose now, even if the inference modules come much later. + + + + +The `otterscale/envoy-gateway` chart bundles Envoy Gateway, the Envoy AI Gateway controller and its +CRDs, already wired together: the extension-manager hooks AI Gateway needs, and `InferencePool` +registered as a backend resource. + +```bash +helm repo add otterscale https://otterscale.github.io/helm-charts +helm repo update + +helm upgrade --install otterscale-envoy-gateway otterscale/envoy-gateway \ + -n envoy-gateway-system --create-namespace +``` + +`InferencePool` itself is defined by the Gateway API Inference Extension, which is a separate +project and is not part of that chart. Install its manifests too: + +```bash +kubectl apply -f https://github.com/kubernetes-sigs/gateway-api-inference-extension/releases/download/v1.0.2/manifests.yaml +``` + + + + + + +If you only want the dashboard, multi-cluster management and Harbor, upstream Envoy Gateway on its +own is enough: + +```bash +helm upgrade --install otterscale-envoy-gateway \ + oci://docker.io/envoyproxy/gateway-helm --version v1.8.3 \ + -n envoy-gateway-system --create-namespace +``` + +Nothing on this path installs the Envoy AI Gateway or the Inference Extension. Adding inference +later means reinstalling Envoy Gateway with the extension-manager configuration, because it is read +at controller startup. + + + + +Wait for the controller, then confirm the API that OtterScale depends on is actually present: + +```bash +kubectl -n envoy-gateway-system rollout status deploy/envoy-gateway +``` + +```bash +kubectl get crd listenersets.gateway.networking.k8s.io +``` + + + +## 4. Create the GatewayClass, EnvoyProxy and Gateway + +The OtterScale chart does not create a `Gateway`. It hangs a `ListenerSet` off one that already +exists, which is what lets the dashboard, Harbor and anything else you route share a single address. +So the `Gateway` is yours to create, and OtterScale adopts it. + +```bash +kubectl apply -f - < + `PROGRAMMED=True` with no HTTPS listener yet is expected at this stage. The dashboard and Harbor + listeners appear in the next step, when the OtterScale chart adds its `ListenerSet`. + + +## Next + +[Install the server](/getting-started/install-server/) covers the server release itself: API, +dashboard, Keycloak and Harbor. It is the same in both deployment modes. diff --git a/src/content/docs/getting-started/03-install-server.mdx b/src/content/docs/getting-started/03-install-server.mdx new file mode 100644 index 0000000..cc551bf --- /dev/null +++ b/src/content/docs/getting-started/03-install-server.mdx @@ -0,0 +1,205 @@ +--- +title: Install the server +description: Install the OtterScale hub with Helm, including the API server, dashboard, Keycloak, Valkey and Harbor. +slug: getting-started/install-server +sidebar: + order: 3 +--- + +import { Aside, Steps, Tabs, TabItem } from '@astrojs/starlight/components'; + +The `otterscale` chart installs the server side of OtterScale in one release. This step is +identical in both deployment modes: install it on your single cluster for a self-managed +deployment, or on the hub for a multi-cluster one. + +| Component | What it does | +| :----------------------- | :---------------------------------------------------------------------------------------------------- | +| **Server** | The ConnectRPC API, and the tunnel listener agents dial into. | +| **Dashboard** | The SvelteKit web console. | +| **Keycloak + PostgreSQL**| The identity provider, seeded with a ready-made `otterscale` realm and its OIDC clients. | +| **Valkey** | Dashboard session storage. | +| **Harbor** | The container and chart registry, federated to Keycloak over OIDC. | +| **Gateway routing** | A `ListenerSet` on your Gateway, plus the `HTTPRoute`s that split `/`, `/api/` and `/auth/`. | +| **cert-manager objects** | A self-signed `Issuer`, a private CA, and the listener certificate, under the default `certSource: auto`. | + +## 1. Add the Helm repository + +```bash +helm repo add otterscale https://otterscale.github.io/helm-charts +helm repo update +``` + +## 2. Write the values file + +Three values are mandatory and the chart refuses to render without them. Everything else has a +working default. + +| Value | Meaning | +| :-------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `externalURL` | The browser-facing base URL. OIDC redirect URIs, CORS origins and the Keycloak realm URL are all built from it. Must be `https://`. | +| `harbor.externalURL` | Harbor's browser-facing base URL. Must be `https://`, and must differ from `externalURL` by hostname or by port. | +| `tunnel.externalURL` | The address agents dial for the tunnel, on `NodePort` 30300. The tunnel certificate is issued for this host and agents verify it, so it must be exactly what they will connect to. | + +Pick the shape that matches your environment: + + + + +Nothing to arrange in advance: cert-manager mints a private CA and a certificate covering both +addresses. Harbor takes port `8443` because it shares the hostname with the dashboard. + +Replace `192.0.2.1` with the Gateway IP you reserved, and `192.0.2.10` with one of your nodes' +addresses. Browsers will warn about the certificate until you trust the CA; +[First login](/getting-started/first-login/#trust-the-certificate) covers that. + +```yaml +# otterscale-values.yaml +externalURL: 'https://192.0.2.1' + +harbor: + externalURL: 'https://192.0.2.1:8443' + +tunnel: + externalURL: 'https://192.0.2.10:30300' # a node address, not the Gateway IP +``` + + + + +With two hostnames, Harbor can share port 443 and needs no listener of its own. Supply the +certificate in a Secret with `tls.crt` and `tls.key` keys, and point `trustedCA` at the issuing CA +if it is not publicly trusted. + +```yaml +# otterscale-values.yaml +externalURL: 'https://otterscale.example.com' + +expose: + listener: + tls: + certSource: secret + secret: + secretName: 'otterscale-tls' + +trustedCA: + secretName: '' # set to the CA Secret when otterscale-tls comes from an internal PKI + +harbor: + externalURL: 'https://harbor.example.com' + caBundleSecretName: '' # same value as trustedCA.secretName + +tunnel: + externalURL: 'https://192.0.2.10:30300' # a node address, not the Gateway IP +``` + +`tunnel.externalURL` stays a bare address rather than a name unless you have a DNS record for the +nodes. It is a `NodePort` on the nodes themselves, reached directly rather than through the +Gateway. + + + + + + +### Values you usually do not need to set + +- **`expose.gateway`** already defaults to `otterscale-gateway` in `envoy-gateway-system`, the + Gateway created in [Prepare the cluster](/getting-started/prepare-cluster/). Change it only if + you named yours differently. +- **`joinSecret`** is generated on first install and reused on upgrade. It is the root secret every + agent join token is derived from, so setting it by hand is only useful when you need the same + tokens across two installs. Changing it invalidates every token already handed out. +- **`trustedCA`** is resolved automatically under `certSource: auto`, becoming the CA the chart + just created. Set it explicitly only when you supply a certificate from a private PKI yourself. +- **`releaseVersion`** is cosmetic: it is what the dashboard displays as its version. Left empty, the + dashboard shows the server and dashboard image tags. + +## 3. Install + +```bash +helm upgrade --install otterscale otterscale/otterscale \ + -f otterscale-values.yaml \ + -n otterscale-system --create-namespace +``` + +The chart validates its inputs before rendering anything, so a missing or malformed value fails +immediately with a message naming the value, rather than half-installing and failing at runtime. + +## 4. Verify + + + +1. **Every pod is running.** + + Harbor and Keycloak take the longest; Keycloak imports its realm and + runs a schema migration on first start. + + ```bash + kubectl -n otterscale-system get pods + ``` + +2. **The listeners were accepted by your Gateway.** + + This is where a missing + `allowedListeners.namespaces.from: All` shows up. + + ```bash + kubectl -n otterscale-system get listenerset otterscale -o wide + ``` + +3. **The routes attached.** + + Two `HTTPRoute`s: one for the dashboard, API and Keycloak, one for + Harbor. + + ```bash + kubectl -n otterscale-system get httproute + ``` + +4. **The dashboard answers.** + + Expect a redirect to Keycloak, and a certificate warning under + `certSource: auto`. + + ```bash + curl -kIL https://192.0.2.1/ + ``` + + + +You now have a running server, but no clusters registered against it. Do not log in yet: the first +login is worth doing after at least one cluster has joined, so the dashboard has something to +show. + +## Operational notes + +Two properties of the server shape how you run it, and both are by design rather than limitations to +work around: + +- **Run a single replica.** Cluster registrations, allocated loopback addresses and live tunnel + sessions are held in the process that accepted them. A second replica would keep its own registry + and its own CA, so agents registered through one replica are unreachable through the other, and + requests routed to the wrong one fail with `cluster not registered`. +- **A restart re-keys every agent.** The tunnel CA is generated at startup and never persisted, so + certificates issued before a restart stop being trusted. Agents notice the dropped session and + re-register automatically with backoff, but their clusters are briefly unreachable. Expect a + short interruption after every restart or upgrade. + + + +## Next + +Pick the page for your deployment mode: + +- [Self-managed cluster](/getting-started/self-managed/) joins this same cluster to the server you + just installed. +- [Multi-cluster](/getting-started/multi-cluster/) joins separate clusters to it. diff --git a/src/content/docs/getting-started/04-self-managed.mdx b/src/content/docs/getting-started/04-self-managed.mdx new file mode 100644 index 0000000..ae0ea88 --- /dev/null +++ b/src/content/docs/getting-started/04-self-managed.mdx @@ -0,0 +1,298 @@ +--- +title: Self-managed cluster +description: Join the cluster that runs the server to itself, so one cluster is both hub and spoke. +slug: getting-started/self-managed +sidebar: + order: 4 +--- + +import { Aside, Steps } from '@astrojs/starlight/components'; + +In a self-managed deployment one cluster plays both roles. The server release you installed in +[Install the server](/getting-started/install-server/) stays where it is, and you add an agent +release alongside it in the same namespace. The agent dials the cluster's own tunnel endpoint, +registers, and from then on the server reaches this cluster's `kube-apiserver` through it, exactly +as it would a remote one. + + + +Two pieces of material are needed before you install, one fewer than a remote cluster needs: the +CA already exists here. + +| Material | Where it comes from | +| :----------------------- | :------------------------------------------------------------------------------------------------ | +| **Join token** | Derived inside the server pod from its root secret. | +| **Harbor robot account** | Created in the Harbor this same release installed. | + + + +## 1. Issue a join token + +Registration is the one endpoint an agent reaches before it has any credentials, so a token +authorises it. The server holds one root secret and derives each cluster's token from that secret +plus the cluster name. Issuing a token needs nothing but access to the secret, which is why it is +done inside the server pod, so the secret never leaves it. + +```bash +kubectl -n otterscale-system \ + exec deploy/otterscale-server -- /otterscale join token --cluster $CLUSTER_NAME +``` + +The token is the only output, so it pipes straight into a `helm` invocation if you prefer. + + + +The token authorises **one cluster**: an agent holding `devel`'s token cannot register as anything +else. Tokens do not expire and cannot be revoked individually; rotating the server's `joinSecret` +invalidates all of them at once. + +## 2. Provision a Harbor robot account + +Every `Workspace` gets its own Harbor project, a robot scoped to that project, and an image-pull +secret built from it. The Tenant Operator does that work through Harbor's REST API, so it needs +credentials of its own: a **system-level robot account**. + +First get a local copy of the CA, so `curl` can verify Harbor's certificate. You are not creating a +Secret here, only a file on disk. The Secret the agent needs already exists. + +```bash +kubectl -n otterscale-system \ + exec deploy/otterscale-server -- /otterscale join ca > ca.crt +``` + +Naming the robot after the cluster keeps accounts separable later, when a second cluster joins. The +script creates it, or rotates its secret if it already exists, and prints the credentials. Run it +from the directory holding `ca.crt`. + +```bash title="fetch-robot-secret.sh" +#!/usr/bin/env bash +set -euo pipefail + +HARBOR_URL="https://${GATEWAY_IP}:8443" +ROBOT_NAME="${CLUSTER_NAME}" + +# Harbor's own admin account, because rotating a robot secret is a PATCH on the +# robot and Harbor defines no robot:update permission at either scope. No robot +# account can do it, only an administrator. The password was generated at +# install time into a Secret the chart keeps on uninstall. +ADMIN_PASS=$(kubectl -n otterscale-system get secret otterscale-harbor-admin \ + -o jsonpath='{.data.HARBOR_ADMIN_PASSWORD}' | base64 -d) + +# Drop --cacert when Harbor serves a publicly trusted certificate. +api() { + curl -sfS --cacert ca.crt -u "admin:${ADMIN_PASS}" \ + -H 'Content-Type: application/json' "$@" +} + +# Robots are addressed cluster-wide even when scoped to a project, so the +# lookup is a filtered list rather than a path under the project. +ROBOT_ID=$(api "${HARBOR_URL}/api/v2.0/robots?q=name%3D${ROBOT_NAME}" \ + | jq -r --arg n "robot\$${ROBOT_NAME}" '.[] | select(.name==$n) | .id' | head -1) + +if [ -n "${ROBOT_ID}" ]; then + # Harbor reveals a robot secret only at the moment it sets one, so an existing + # robot's credentials cannot be read back. Rotate instead. + # The A1a prefix satisfies Harbor's complexity rule for robot secrets. + SECRET="A1a$(openssl rand -hex 16)" + api -X PATCH "${HARBOR_URL}/api/v2.0/robots/${ROBOT_ID}" \ + -d "{\"secret\":\"${SECRET}\"}" > /dev/null +else + RESP=$(api -X POST "${HARBOR_URL}/api/v2.0/robots" -d @- < + Harbor answers `400` when a permission list names a resource/action pair it does not define, and + the response names the pair. Drop it and cross-check the equivalent tick-box in Harbor's **New + Robot** dialog, because the permission vocabulary varies a little between Harbor versions. + + +## 3. Write the agent values file + +```yaml +# agent-values.yaml +image: + repository: ghcr.io/otterscale/otterscale + tag: v1.5.0-rc.4 + pullPolicy: IfNotPresent + +agent: + serverURL: 'https://192.0.2.1/api/' # the Gateway IP + tunnelServerURL: 'https://192.0.2.10:30300' # a node address + cluster: 'devel' + joinToken: '' # from step 1 + +trustedCA: + secretName: 'otterscale-ca' # already present, created by the server release + key: 'ca.crt' + +clusterAdmin: + enabled: true + users: [] # OIDC subjects to bind cluster-admin to, see below + +clusterInfo: + enabled: true + externalAddress: '192.0.2.10' # a node address, for NodePort workload URLs + nodePortRange: '30000-32767' + inferenceURL: "https://inference.example.com" + +tenantOperator: + enabled: true + harbor: + url: https://192.0.2.1:8443 + robot: + name: 'robot$devel' # from step 2 + secret: '' # from step 2 +``` + +What each value controls: + +- **`agent.serverURL`**: keep the `/api/` suffix. That is the path the Gateway rewrites onto the + server; without it the agent talks to the dashboard instead. The agent is a pod on this cluster, + so it reaches the cluster's own Gateway address. +- **`agent.tunnelServerURL`**: must equal the server's `tunnel.externalURL` exactly, which is a + node address rather than the Gateway IP. The tunnel certificate is issued for that host and the + agent verifies it, so a mismatch fails the TLS handshake rather than degrading quietly. +- **`agent.cluster`**: the name the join token is bound to. A token issued for another name is + rejected. +- **`agent.joinToken`**: inline, the token lands in the Helm release and is readable by anyone who + can read Secrets in the namespace. To keep it out, create the Secret yourself and point + `agent.existingSecret` at it instead. Either way the chart mounts it as a file rather than putting + it in the agent's environment. +- **`clusterAdmin.users`**: binds `cluster-admin` on this cluster to the listed users. Leave it + empty for now; [First login](/getting-started/first-login/) explains how to find your OIDC subject + and fill it in. Without at least one entry, nobody can create a `Workspace`. +- **`clusterInfo`**: how users reach workloads here. The Kubernetes API cannot answer that, so the + answers are written to the `otterscale-info` ConfigMap in `kube-public`, which the dashboard + reads. `externalAddress` must be a bare address: no scheme, no port. Add `inferenceURL` only if + you are serving models. +- **`tenantOperator.harbor`**: the Harbor this cluster already runs, and the robot from step 2. All + three fields are required when the operator is enabled: it refuses to run half-configured rather + than leave workspaces without registry access. + + + +## 4. Install + +```bash +helm upgrade --install otterscale-agent otterscale/otterscale-agent \ + -f agent-values.yaml \ + -n otterscale-system +``` + +No `--create-namespace`: the server release already created `otterscale-system`. + +## 5. Verify + + + +1. **The agent and operator are running**, alongside the server components. + + ```bash + kubectl -n otterscale-system get pods + ``` + +2. **Registration succeeded.** The agent logs the registration and then the tunnel it established. + + ```bash + kubectl -n otterscale-system logs deploy/otterscale-agent + ``` + +3. **The dashboard's cluster facts landed.** + + ```bash + kubectl -n kube-public get configmap otterscale-info -o yaml + ``` + + + +## Troubleshooting + +| Symptom | Likely cause | +| :---------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------- | +| Agent logs a TLS handshake failure against the tunnel | `agent.tunnelServerURL` does not match `tunnel.externalURL`. The certificate is issued for that host and the agent verifies it. | +| Agent cannot reach the tunnel at all | The node address is wrong, or `NodePort` 30300 is not reachable from inside the cluster. | +| Agent logs an `x509` error while registering | `trustedCA.secretName` is empty or misspelled. It must be `otterscale-ca`, the Secret the server release created. | +| Registration rejected | The token does not match `agent.cluster`, or the server's join secret has been rotated. | +| `cluster not registered` from the dashboard | The agent has not connected yet, or the server is running more than one replica. Run one. | +| The cluster goes unreachable, then recovers | The server restarted; its tunnel CA is regenerated at startup and the agent must re-register. | +| `Workspace` stalls at `Ready=False` | Harbor credentials are wrong or the robot lacks a permission. Check the Tenant Operator's logs. | + +## Next + +[First login](/getting-started/first-login/) covers signing in, granting yourself cluster access, +and creating a workspace. + +To bring further clusters under this same server later, follow +[Multi-cluster](/getting-started/multi-cluster/). Nothing installed here has to change. diff --git a/src/content/docs/getting-started/05-multi-cluster.mdx b/src/content/docs/getting-started/05-multi-cluster.mdx new file mode 100644 index 0000000..f96e437 --- /dev/null +++ b/src/content/docs/getting-started/05-multi-cluster.mdx @@ -0,0 +1,340 @@ +--- +title: Multi-cluster +description: Join separate clusters to the hub, each with its own join token, CA copy and Harbor robot. +slug: getting-started/multi-cluster +sidebar: + order: 5 +--- + +import { Aside, Steps } from '@astrojs/starlight/components'; + +In a multi-cluster deployment the **hub** runs the server release from +[Install the server](/getting-started/install-server/), and every cluster you want to manage runs an +agent release of its own. The agent dials out to the hub and keeps a reverse tunnel open, so a +cluster behind NAT, a corporate firewall, or in an otherwise closed network needs no inbound rule. + +Work through this page once per managed cluster. Nothing is shared between them: each gets its own +join token, its own copy of the hub's CA, and its own Harbor robot, so revoking one cluster's access +leaves the others untouched. + +Three pieces of material have to travel from the hub to each joining cluster: + +| Material | Why the agent needs it | +| :----------------------- | :---------------------------------------------------------------------------------------------------------------------------- | +| **Join token** | Registration is the one endpoint an agent reaches before it has any credentials, so a token authorises it. | +| **CA certificate** | The agent verifies the hub with its image's system roots, so a privately signed certificate has to be handed over out of band. | +| **Harbor robot account** | The Tenant Operator gives every workspace its own Harbor project, robot and image-pull secret, and needs an account of its own to do that. | + + + + + +## 1. Issue a join token + +The hub holds one root secret and derives each cluster's token from that secret plus the cluster +name. Issuing a token needs nothing but access to the secret, which is why it is done inside the +server pod, so the secret never leaves it. + +```bash +kubectl --context $HUB_CONTEXT -n otterscale-system \ + exec deploy/otterscale-server -- /otterscale join token --cluster $CLUSTER_NAME +``` + +The token is the only output, so it pipes straight into a `helm` invocation if you prefer. + + + +What the token does and does not give you: + +- **It authorises one cluster.** An agent holding `devel`'s token cannot register as `staging`, so a + compromised agent cannot take over another cluster's traffic. +- **A rejected token changes nothing.** The check runs before any state is touched, so a bad + registration cannot displace the agent currently serving that cluster. +- **Tokens do not expire and cannot be revoked individually.** Rotating the hub's `joinSecret` + invalidates all of them at once, after which every agent needs a new token. + +## 2. Hand over the CA + +Under the default `certSource: auto`, the hub's certificate is signed by a private CA that nothing +outside the hub cluster trusts yet. Export it from the server, which already has it mounted, and +create it as a Secret on the joining cluster. + +```bash +kubectl --context $HUB_CONTEXT -n otterscale-system \ + exec deploy/otterscale-server -- /otterscale join ca > ca.crt +``` + +```bash +kubectl --context $MANAGED_CONTEXT create namespace otterscale-system \ + --dry-run=client -o yaml | kubectl --context $MANAGED_CONTEXT apply -f - + +kubectl --context $MANAGED_CONTEXT -n otterscale-system \ + create secret generic otterscale-ca --from-file=ca.crt +``` + + + +If your hub serves a publicly trusted certificate, skip this step and leave `trustedCA.secretName` +empty. `otterscale join ca` will tell you when there is nothing to hand over. + +Keep `ca.crt` around: the next step uses it to talk to the hub's Harbor. + +## 3. Provision a Harbor robot account + +Every `Workspace` gets its own Harbor project, a robot scoped to that project, and an image-pull +secret built from it. The Tenant Operator does that work through the hub's Harbor REST API, so it +needs credentials of its own: a **system-level robot account**, created on the hub. + +Naming the robot after the joining cluster is what keeps the clusters separable. The script creates +it, or rotates its secret if it already exists, and prints the credentials. Run it from the +directory holding the `ca.crt` from step 2. + +```bash title="fetch-robot-secret.sh" +#!/usr/bin/env bash +set -euo pipefail + +HARBOR_URL="https://${GATEWAY_IP}:8443" +ROBOT_NAME="${CLUSTER_NAME}" + +# Harbor's own admin account, because rotating a robot secret is a PATCH on the +# robot and Harbor defines no robot:update permission at either scope. No robot +# account can do it, only an administrator. The password was generated at +# install time into a Secret the chart keeps on uninstall. +ADMIN_PASS=$(kubectl --context "${HUB_CONTEXT}" -n otterscale-system \ + get secret otterscale-harbor-admin \ + -o jsonpath='{.data.HARBOR_ADMIN_PASSWORD}' | base64 -d) + +# Drop --cacert when Harbor serves a publicly trusted certificate. +api() { + curl -sfS --cacert ca.crt -u "admin:${ADMIN_PASS}" \ + -H 'Content-Type: application/json' "$@" +} + +# Robots are addressed cluster-wide even when scoped to a project, so the +# lookup is a filtered list rather than a path under the project. +ROBOT_ID=$(api "${HARBOR_URL}/api/v2.0/robots?q=name%3D${ROBOT_NAME}" \ + | jq -r --arg n "robot\$${ROBOT_NAME}" '.[] | select(.name==$n) | .id' | head -1) + +if [ -n "${ROBOT_ID}" ]; then + # Harbor reveals a robot secret only at the moment it sets one, so an existing + # robot's credentials cannot be read back. Rotate instead. + # The A1a prefix satisfies Harbor's complexity rule for robot secrets. + SECRET="A1a$(openssl rand -hex 16)" + api -X PATCH "${HARBOR_URL}/api/v2.0/robots/${ROBOT_ID}" \ + -d "{\"secret\":\"${SECRET}\"}" > /dev/null +else + RESP=$(api -X POST "${HARBOR_URL}/api/v2.0/robots" -d @- < + Harbor answers `400` when a permission list names a resource/action pair it does not define, and + the response names the pair. Drop it and cross-check the equivalent tick-box in Harbor's **New + Robot** dialog, because the permission vocabulary varies a little between Harbor versions. + + +## 4. Write the agent values file + +Every address here belongs to the **hub**, not to the cluster being joined, except +`clusterInfo.externalAddress`, which is how users reach workloads on this cluster. + +```yaml +# agent-values.yaml +image: + repository: ghcr.io/otterscale/otterscale + tag: v1.5.0-rc.4 + pullPolicy: IfNotPresent + +agent: + serverURL: 'https://192.0.2.1/api/' # the hub's Gateway IP + tunnelServerURL: 'https://192.0.2.10:30300' # the hub's node address + cluster: 'devel' + joinToken: '' # from step 1 + +trustedCA: + secretName: 'otterscale-ca' # the Secret from step 2 + key: 'ca.crt' + +clusterAdmin: + enabled: true + users: [] # OIDC subjects to bind cluster-admin to, see below + +clusterInfo: + enabled: true + externalAddress: '198.51.100.10' # a node address on THIS cluster + nodePortRange: '30000-32767' + +tenantOperator: + enabled: true + harbor: + url: https://192.0.2.1:8443 # the hub's Harbor + robot: + name: 'robot$devel' # from step 3 + secret: '' # from step 3 +``` + +What each value controls: + +- **`agent.serverURL`**: keep the `/api/` suffix. That is the path the hub's Gateway rewrites onto + the server; without it the agent talks to the dashboard instead. +- **`agent.tunnelServerURL`**: must equal the hub's `tunnel.externalURL` exactly, which is one of + the hub's node addresses rather than its Gateway IP. The tunnel certificate is issued for that + host and the agent verifies it, so a mismatch fails the TLS handshake at every agent rather than + degrading quietly. +- **`agent.cluster`**: the name the join token is bound to. A token issued for another name is + rejected. +- **`agent.joinToken`**: inline, the token lands in the Helm release and is readable by anyone who + can read Secrets in the namespace. To keep it out, create the Secret yourself and point + `agent.existingSecret` at it instead. Either way the chart mounts it as a file rather than putting + it in the agent's environment. +- **`clusterAdmin.users`**: binds `cluster-admin` on *this* cluster to the listed users, and it is + per-cluster: a user with access to one managed cluster has none on another until listed there too. + Leave it empty for now; [First login](/getting-started/first-login/) explains how to find your + OIDC subject and fill it in. Without at least one entry, nobody can create a `Workspace` here. +- **`clusterInfo`**: how users reach workloads on this cluster, written to the `otterscale-info` + ConfigMap in `kube-public` for the dashboard to read, because the Kubernetes API cannot answer it. + `externalAddress` must be a bare address: no scheme, no port. Add `inferenceURL` only if this + cluster serves models. +- **`tenantOperator.harbor`**: the hub's Harbor and the robot from step 3. All three fields are + required when the operator is enabled: it refuses to run half-configured rather than leave + workspaces without registry access. + + + +## 5. Install + +The release namespace must be `otterscale-system` while the Tenant Operator is enabled, because its +install manifest is pre-rendered against that namespace. + +```bash +helm --kube-context $MANAGED_CONTEXT upgrade --install otterscale-agent \ + otterscale/otterscale-agent \ + -f agent-values.yaml \ + -n otterscale-system --create-namespace +``` + +## 6. Verify + + + +1. **The agent and operator are running.** + + ```bash + kubectl --context $MANAGED_CONTEXT -n otterscale-system get pods + ``` + +2. **Registration succeeded.** The agent logs the registration and then the tunnel it established. + + ```bash + kubectl --context $MANAGED_CONTEXT -n otterscale-system logs deploy/otterscale-agent + ``` + +3. **The dashboard's cluster facts landed.** + + ```bash + kubectl --context $MANAGED_CONTEXT -n kube-public get configmap otterscale-info -o yaml + ``` + +4. **The hub sees the cluster.** It appears in the dashboard's cluster switcher, which is the point + of [First login](/getting-started/first-login/). + + + +Then repeat the whole page for the next cluster, with a new `CLUSTER_NAME`. + +## Troubleshooting + +| Symptom | Likely cause | +| :---------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------- | +| Agent logs a TLS handshake failure against the tunnel | `agent.tunnelServerURL` does not match the hub's `tunnel.externalURL`. The certificate is issued for that host and the agent verifies it. | +| Agent cannot reach the hub at all | Outbound access from this cluster to the hub's 443 and 30300 is blocked. | +| Agent logs an `x509` error while registering | The CA Secret is missing on this cluster, or `trustedCA.secretName` does not point at it. | +| Registration rejected | The token does not match `agent.cluster`, or the hub's join secret has been rotated. | +| `cluster not registered` from the dashboard | The agent has not reconnected yet, or requests are reaching a second server replica. Run one replica. | +| Agent warns about plaintext HTTP at startup | `agent.serverURL` is `http://` to a remote host, which exposes the join token in transit. Legitimate only behind a mesh that terminates TLS. | +| Every cluster goes unreachable at once, then recovers | The hub restarted; its tunnel CA is regenerated at startup and agents must re-register. | +| `Workspace` stalls at `Ready=False` | Harbor credentials are wrong or the robot lacks a permission. Check the Tenant Operator's logs. | + +## Next + +[First login](/getting-started/first-login/) covers signing in, granting yourself cluster access, +and creating a workspace. diff --git a/src/content/docs/getting-started/06-first-login.mdx b/src/content/docs/getting-started/06-first-login.mdx new file mode 100644 index 0000000..6204fcd --- /dev/null +++ b/src/content/docs/getting-started/06-first-login.mdx @@ -0,0 +1,171 @@ +--- +title: First login +description: Sign in to the dashboard, grant yourself cluster access, and create your first workspace. +slug: getting-started/first-login +sidebar: + order: 6 +--- + +import { Aside, Steps, Card, CardGrid } from '@astrojs/starlight/components'; + +The server is running and at least one cluster has joined, whether that is the same cluster or a +separate one. What remains is identity: signing in, and giving that identity enough authority on +the joined cluster to create a workspace. + +## Trust the certificate + +Under the default `certSource: auto` the server serves a certificate from a private CA, so +browsers will warn. Install that CA into your OS or browser trust store. It is the same certificate +the agent trusts, and you can read it straight out of the cluster running the server: + +```bash +kubectl -n otterscale-system get secret otterscale-ca \ + -o jsonpath='{.data.ca\.crt}' | base64 -d > otterscale-ca.crt +``` + +Clicking through the warning works for a first look, but do trust the CA before real use: Harbor +and the dashboard exchange OIDC redirects, and a browser that distrusts one of the two hosts can +fail the login mid-flow. + +## Sign in + +Open `externalURL` in a browser. You are redirected to Keycloak, in the seeded `otterscale` +realm. + +The realm ships with one user: + +| Field | Value | +| :----------- | :--------- | +| **Username** | `admin` | +| **Password** | `password` | + +Keycloak forces a password change on first sign-in, so this credential is only good once. + + + +That `admin` user starts with two things granted: + +- The `admin` client role on the dashboard client, which the API turns into the Kubernetes group + `oidc:admin`. +- Membership of the `harbor_admin` group, which makes it a Harbor administrator. + +### The Keycloak admin console + +For everything identity-related, such as creating users, groups, and turning off +self-registration, Keycloak's own console is at `/auth/`. It uses the **master** +realm's `admin` account, which is a different account from the realm user above: + +```bash +kubectl -n otterscale-system get secret otterscale-keycloak-admin \ + -o jsonpath='{.data.KC_BOOTSTRAP_ADMIN_PASSWORD}' | base64 -d +``` + +## Grant yourself cluster access + +At this point the dashboard lists your joined cluster but you cannot create anything in it. Every +request reaches that cluster's `kube-apiserver` impersonating *you*, so its own RBAC decides what +happens, and nothing has granted your identity anything yet. That holds for a self-managed cluster +as much as a remote one: the server carries no authority of its own. + +The agent chart's `clusterAdmin.users` exists for exactly this. It takes **OIDC subjects**, not +usernames: the subject is what the API sends as `Impersonate-User`, and for Keycloak that is the +user's internal ID. + + + +1. **Find your subject.** In the Keycloak admin console, switch to the `otterscale` realm, open + **Users**, select your user, and copy the **ID** field, which is a UUID. + +2. **Add it to the agent's values.** + + ```yaml + # agent-values.yaml + clusterAdmin: + enabled: true + users: + - '3f9c1e02-7a45-4b8e-9d21-6c0f5ab7e134' + ``` + +3. **Upgrade the agent release** on the cluster you are granting access to. On a self-managed + cluster that is your only cluster; in multi-cluster mode add `--kube-context` for the target. + + ```bash + helm upgrade --install otterscale-agent otterscale/otterscale-agent \ + -f agent-values.yaml \ + -n otterscale-system + ``` + +4. **Confirm the binding.** + + ```bash + kubectl get clusterrolebinding otterscale-agent-cluster-admin -o yaml + ``` + + `clusterAdmin.users` grants access on the cluster whose agent release carries it, and nowhere + else, so in multi-cluster mode repeat this for every cluster the user needs. + + + + + +## Create a workspace + +A `Workspace` is the multi-tenant unit: the Tenant Operator turns one into a namespace with Pod +Security Standards labels, RBAC bindings per member, resource quotas and limit ranges, default-deny +network policies, a Harbor project with its own robot and image-pull secret, and a Flux +`HelmRepository` pointing at that project. + +In the dashboard, pick your cluster and create a workspace from the workspace switcher. You are +added as its first `admin` member automatically. + +Two authorisation layers apply, and it helps to know which one is complaining: + +- **Kubernetes RBAC** must allow you to create `workspaces.tenant.otterscale.io` at all. That is + what the `clusterAdmin.users` binding above provides. +- **The operator's admission webhook** then checks that you are either listed as an `admin` member + of the workspace being created, or hold cluster-wide access. Creating a workspace you are not an + admin of is rejected with *"workspace creator must be listed as a member with the 'admin' role"*. + +The same rule governs updates and deletes, evaluated against the *stored* spec, so you cannot +grant yourself admin and approve it in the same request. + + + +## What's next + + + + The dashboard's **Modules** page lists charts from a Flux `HelmRepository` named `modules` in + `otterscale-system` on the target cluster. Neither chart creates it, so point one at + `https://otterscale.github.io/helm-charts` to offer GPU Operator, KubeVirt, KServe, + Prometheus and the rest. + + + The agent proxies read-only Prometheus queries for the dashboard's metrics. Its default + target is `http://otterscale-prometheus-kube-prometheus.monitoring.svc:9090`. Install the + `prometheus-stack` module, or point `--proxy-prometheus-url` at a Prometheus you already run. + + + Follow [Multi-cluster](/getting-started/multi-cluster/) for each additional cluster, whether + you started self-managed or not. Every cluster needs its own join token, its own CA Secret, + and its own Harbor robot. + + + The ConnectRPC services behind the dashboard are documented under **API**, generated from + OtterScale's OpenAPI schema. + + diff --git a/src/content/docs/introduction.mdx b/src/content/docs/introduction.mdx index 80ef17a..32ab9dd 100644 --- a/src/content/docs/introduction.mdx +++ b/src/content/docs/introduction.mdx @@ -91,7 +91,38 @@ The OtterScale platform is composed of several open-source components: ## Getting Started -Installation, configuration, and operational guides are coming soon as part of this documentation. In the meantime: + + + + + + + + +The two deployment modes are alternatives, not stages. Requirements explains which to pick. + +Along the way: -- Run `otterscale server --help` and `otterscale agent --help` to explore the available options. -- Add the Helm repository: `helm repo add otterscale https://otterscale.github.io/charts` +- `helm repo add otterscale https://otterscale.github.io/helm-charts` adds the chart repository. +- `otterscale server --help` and `otterscale agent --help` are the authoritative reference for every + flag and environment variable.