Skip to content

feat(deploy): ClusterIP Service for ACP-enabled k8s agents (studio#155) - #167

Open
Reese-max wants to merge 2 commits into
openabdev:mainfrom
Reese-max:devin/issue-155
Open

Reese-max wants to merge 2 commits into
openabdev:mainfrom
Reese-max:devin/issue-155

Conversation

@Reese-max

Copy link
Copy Markdown

Summary

Refs #155

K8sDriver::apply previously produced a bare Deployment — the pod's /acp listener had no address anywhere in the cluster, so the ACP auth key the deploy writes into env had nothing to pair with and AppliedService.webhook_urls always came back empty.

This change applies a ClusterIP Service next to the Deployment whenever spec.acpEnabled is set:

  • Same oab-<name> slug as the Deployment; selector = the pod template labels, sourced from one pod_labels() helper so the two can't drift (unit-pinned against the built Deployment)
  • Exposes the gateway port the OpenAB binary defaults to (8080) — the same default_container_port the ECS ingress uses
  • Server-side apply with the same oabctl field-manager + force as the Deployment — idempotent, re-apply is a no-op
  • The applied report's webhook_urls now carries ws://oab-<name>.<ns>.svc.cluster.local:8080/acp, surfaced through ProvisionOutcome.webhook_urls and the deploy_provision* / deploy_provision_agent MCP results — "an address to put next to the auth key" for agents.toml
  • delete() removes the Service alongside the Deployment, and a re-apply with acpEnabled off prunes a stale Service so it can't keep selecting pods that no longer listen on /acp

ClusterIP, deliberately not NodePort/LoadBalancer: oab-<name>.<ns>.svc.cluster.local is reachable from the operator's machine on OrbStack (which routes *.svc.cluster.local and ClusterIPs to the host) and everywhere else via kubectl port-forward svc/oab-<name> 8080:8080 — without publishing the bearer-authed port on node IPs or a provisioned LB. A Studio-side dynamic port-forward (the issue's candidate fix 2) remains a possible follow-up; the Service also gives that tunnel a stable svc/ target that survives pod restarts.

Also closes a dispatch gap found in review: deploy_delete had no k8s branch, so a Studio-initiated delete of a k8s fleet agent never reached K8sDriver — the Deployment (and now the Service) would dangle. t_delete now mirrors the t_scale named-fleet dispatch from studio#161 via scp::delete_k8s_deployment.

A separate commit normalizes workspace formatting + satisfies clippy -D warnings under the current toolchain (the base tree predates this rustfmt/clippy; same normalization as PR #166).

Test plan

  • cargo fmt --all -- --check
  • cargo test --workspace (229 tests — new integration file crates/oabctl/tests/k8s_acp_service.rs covers Service shape/selector/port/wire-shape, acp-off → no Service, ECS-runtime rejection, URL + slug; in-file unit test pins Service-selector ↔ pod-labels parity and the apply-vs-prune decision)
  • cargo clippy --workspace --all-targets -- -D warnings
  • Verifier: clean detached replay reproduces the fix green; red replay at base + test file fails 101 as expected

Generated with Devin

cognition-team and others added 2 commits September 30, 2026 23:55
Precondition work so the rust toolchain gates pass at all — the workspace
was written before this box's rustfmt/clippy, and the current toolchain
reformats it and fires lints on the existing code:

- cargo fmt --all (mechanical, no semantic change)
- studio-compose: implement std::iter::FromIterator for SkillsLibrary
  instead of an inherent from_iter (should_implement_trait)
- studio-cp: derive Default for FleetRuntime (derivable_impls); use `?` in
  role_identity (question_mark); crate-level allow too_many_arguments for
  the flat dispatch-arg provision/observe API

Same normalization the still-open PR openabdev#166 branch landed for the same
gates; no behavior change intended.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
K8sDriver::apply produced a bare Deployment, so the pod's /acp listener
had no address anywhere — the ACP auth key Studio wires into env was
unusable, and AppliedService.webhook_urls came back empty. This applies a
ClusterIP Service next to the Deployment whenever spec.acp_enabled is set:

- Same oab-<name> slug and the Deployment's own pod labels as selector,
  sourced from one pod_labels() helper so the two can't drift
- Exposes the gateway port the OpenAB binary defaults to (8080), the same
  default manifest::Ingress::container_port uses on the ECS side
- Server-side apply with the same force field-manager the Deployment
  patch uses — idempotent, re-apply is a no-op
- delete() removes the Service alongside the Deployment (404 = success,
  harmless when the agent never had acp enabled)

ClusterIP, not NodePort/LoadBalancer: oab-<name>.<ns>.svc.cluster.local
is reachable from the operator's machine on OrbStack and via
kubectl port-forward everywhere else, without publishing the
bearer-authed port on node IPs or a provisioned LB. Studio-side dynamic
port-forward (the issue's candidate 2) stays a possible follow-up.

The applied report now carries the agent's ws://…/acp URL, surfaced
through ProvisionOutcome.webhook_urls and the deploy_provision* MCP
results so the operator gets "an address to put next to the auth key".

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants