Skip to content

Latest commit

 

History

History
201 lines (166 loc) · 10.3 KB

File metadata and controls

201 lines (166 loc) · 10.3 KB

Security model

This page states what drukbox protects, what it deliberately does not, and the tradeoffs behind each default. For how sandboxes are reached read Networking; for the configuration knobs named here read Deploy.

Trust model: trusted callers, untrusted sandboxes

An admin key or service account token can create, list, read, and delete every host. Only an admin key can create or remove a service account. There are no host filters per service account and no tenant isolation. Treat each key and token as an operator credential.

The asymmetry that is part of the model: callers are trusted, the sandboxes they provision are not. Drukbox hands back SSH coordinates and stops — it never runs a sandbox's code and owns no runtime inside the VM. The hardening below is about keeping an untrusted workload on a sandbox from reaching back into drukbox's credentials or its cloud account, not about isolating one token holder from another.

This is the right model for a single team standing up sandboxes behind their own API. It is not a multi-tenant boundary: do not hand drukbox tokens to mutually distrusting users.

Authentication

SERVICE_TOKENS holds the admin keys. Drukbox compares them in constant time and never reads the database for them. A service account token is stored only as its SHA-256 fingerprint and checked against the database on every request, so removal needs no restart. No route returns a token or a fingerprint after creation. Only GET /healthz and the OpenAPI pages skip authentication. GET /doctor is the only route that reports dependency state.

To rotate an admin key, add the new key, restart, move callers, remove the old key, and restart again. To rotate a service account, create a new one, move the caller, and remove the old one.

Control-plane network exposure

The API holds provider credentials and mints cloud resources, so the process is a high-value target. It binds 0.0.0.0 by default for container friendliness. When only a co-located or host-networked caller reaches it, set UVICORN_HOST=127.0.0.1 to keep the control plane off other interfaces. When it must be remote, front it with TLS termination and treat the token as the only thing standing between the internet and your cloud account — drukbox does no TLS itself and adds no rate limiting (see Resource exhaustion).

Sandbox reachability and SSH auth

How a caller reaches a sandbox, and the tradeoffs of each path, are covered in Networking. The security-relevant summary:

  • Per-VM keys. On AWS (Tailscale off), Hetzner, docker, and docker-sbx, drukbox mints a fresh ed25519 keypair per VM, returns the private half once in the create response, and never persists it — a later GET /hosts/{id} returns private_key: null. The key is the auth boundary; password auth is never enabled.
  • AWS ingress fail-open. The managed drukbox-managed security group opens SSH to the detected egress /32, or to whatever AWS_SSH_CIDRS specifies. If egress detection fails and no CIDRs are set, ingress falls back to 0.0.0.0/0 with a warning log. The per-VM key remains the boundary in that case, but set AWS_SSH_CIDRS explicitly in any environment where world-open port 22 is unacceptable.
  • Hetzner has no firewall. A fresh server exposes port 22 to the internet; the per-VM key is the only boundary. There is no ingress configuration to manage.
  • Docker shares the daemon host's identity. A sandbox published on a non-loopback DOCKER_SSH_HOST is reachable by everything that reaches that address. On a tailnet it has no device of its own, so ACL tags cannot scope it. Each sandbox's key is the only boundary.
  • First-keyscan MITM window. With Tailscale off, the known_hosts material is scanned over the public network and carries the usual trust-on-first-use window. Enable Tailscale to run the scan over the authenticated overlay.

Secrets: which field to use

POST /hosts takes two fields that reach the box. env is ordinary configuration. It is plaintext, delivered into the box on purpose, and readable by whatever runs inside. secrets is for a credential. The box gets a placeholder, never the value. Put a token in secrets, never in env.

A secrets entry names a service from the catalog, or a custom host of its own. It holds a static value, or an issuer that mints the value on demand. Drukbox encrypts the entry in the database with AES-256-GCM under SECRETS_KEY. A database dump holds ciphertext, and a process with the key can decrypt it. Rotate the key by prepending a new one. Remove an old key only after no stored row needs it.

Provider tokens (EXE_API_TOKEN, EXE_REGISTRY_PASSWORD, HETZNER_API_TOKEN, Tailscale OAuth) and AWS credentials are read from the environment or the AWS SDK default chain. They are never written to the database and never returned by the API.

What the proxy protects

The placeholder, drk.<host id>.<service>.<random>, works only at the secrets proxy, and only for the host and the service it names. The entry stores a fingerprint of it, so a database read cannot replay it. The proxy swaps every header that carries a placeholder, on HTTPS to a registered host, and touches nothing else. One placeholder it cannot resolve refuses the whole request. Plain HTTP is forwarded unchanged. The proxy refuses a loopback, private, link-local, or metadata destination, so a box cannot reach the exchange or the API through it. It logs no credential.

The real value is encrypted in the database. It passes through the exchange and the proxy for one request, and the exchange keeps an issuer's value in memory. An issuer that ends a value before its expiry orders a refresh through the API at POST /hosts/{host_id}/secrets/{service}/refresh with an admin key or service account token. The exchange listens on loopback beside the API and takes the order from the API only. The order carries no value. It makes the exchange ask the issuer again, so a stray order costs one fetch and nothing else. On docker-sbx the value lives in sbx's own store, scoped to that sandbox, and drukbox runs no proxy there. Host deletion removes the sandbox's secrets and value files before the VM goes. The lease in expires_at schedules that deletion and does not revoke the credential. Revoke it at its source when a box must lose it at once.

A box with secrets trusts the proxy's CA for the registered hosts. Whoever holds the CA key can impersonate any host to that box. The key lives in the proxy's volume. Guard it like SECRETS_KEY. The API reads only the public certificate, from SECRETS_PROXY_CA_FILE.

POST /hosts never returns a secret. A validation response omits the rejected input, so a bad value or a bad issuer header does not reach the caller. An issuer URL must not carry user credentials or a fragment. It can use plain HTTP inside the deployment, where the exchange already answers the proxy in the clear. An issuer outside the deployment uses HTTPS. Put credentials only in the issuer headers, which Drukbox encrypts. The value an issuer returns is never stored.

What env is and is not

Caller env stays plaintext by design. It is ordinary configuration. Each provider writes it to /etc/environment on the VM, and PAM hands it to every session at login. No response echoes it. The schema rejects the reserved key TAILSCALE_AUTHKEY. It also rejects a value that PAM would change. PAM cuts a value at #, treats a quote as the start of a quoted value, and joins the next line after a trailing backslash. It also stops at a line of 8192 bytes and loses every entry after it. So a value must be printable ASCII without #, quotes, or backslashes, and without a space at either end. The whole KEY=VALUE line must stay under 8191 bytes. No secrets in env, ever.

Provider material in the VM

Two pieces of material reach the VM through its provider's user-data / setup-script mechanism, and that channel is the relevant exposure:

  • Tailscale auth key. Minted per host, ephemeral, tag-scoped, and short-lived; it is not persisted in drukbox's database. It is delivered to the VM via user-data, so a process on the box can read it — acceptable given its single-use, ephemeral nature.
  • AWS IMDS. Sandboxes run untrusted code, so launched EC2 instances require IMDSv2 (HttpTokens: required) with a put-response hop limit of 1. This stops an in-VM SSRF or stray process from reading instance metadata over the legacy unauthenticated IMDSv1 path — which would otherwise expose the user-data auth key and, if AWS_INSTANCE_PROFILE is set, live IAM role credentials. The hop limit keeps a containerized workload one network hop from the endpoint; if you run a sandbox payload in a container that genuinely needs IMDS, raise it deliberately. Residual: a local root on the box can still read its own user-data, so scope AWS_INSTANCE_PROFILE tightly (or leave it unset) and keep nothing in caller env that the sandbox workload should not see.

Information disclosure

Provisioning failures are stored on the host as a concrete summary (exception type and message), not a raw Python traceback. That summary is what HostOut.last_error and the POST /hosts 502 detail return to callers; the full traceback stays in the server log only. Keep log sinks access-controlled — they hold the detail the API withholds.

Resource exhaustion

Drukbox adds no quota or rate limit of its own: a valid token can provision paid VMs without bound, so a leaked token is a cost-DoS as well as a control-plane compromise. The controls that exist are operational — expires_at plus the janitor reap idle hosts, PROVISIONING_GRACE_SECONDS bounds strands, and POOL_MAX_CREATES_PER_TICK caps pool over-provision. Put per-caller quotas and rate limiting in the layer that issues and fronts tokens.

Not vulnerabilities by design

  • An admin key or service account token can delete any host. There is no second factor for destructive calls — the token is the boundary.
  • Drukbox never opens an SSH session, runs sandbox code, or creates Linux users. Everything past the returned SSH coordinates is the caller's responsibility.
  • private_key appearing once in the create response is intentional; callers must capture it then, because it is never recoverable later.

Reporting a vulnerability

See SECURITY.md for private reporting.