Skip to content

feat: stream Kubernetes pod logs to the dashboard - #361

Merged
Priyanshu-u07 merged 5 commits into
mainfrom
k8s-log-streaming
Sep 24, 2026
Merged

Priyanshu-u07 merged 5 commits into
mainfrom
k8s-log-streaming

Conversation

@Priyanshu-u07

@Priyanshu-u07 Priyanshu-u07 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Terminal Logs showed a static snapshot on Kubernetes because the adapter had no streaming implementation get_log_streaming_info returned supported: False.

Closes #356.

The adapter now returns a real ws_url and a subscription, and stream_logs follows the pod's logs through the Kubernetes API. The WebSocket endpoint gains a k8s branch alongside skypilot and nosana.

The client's follow stream blocks, so one dedicated thread opens and reads it, handing lines back through a bounded queue. Nothing touches asyncio's shared executor and nothing is scheduled on the loop per line, so a stream cannot
starve the rest of the service however long it runs.

Closing the response is what unblocks a reader waiting on the next chunk, but close() blocks on the socket that reader is inside — on the event loop that deadlocks the whole orchestration service, so it runs on its own thread. A
reader slower than the engine loses lines rather than stalling it.

The subscription names the StatefulSet, not a pod: a pod name captured at subscribe time goes stale when the pod is replaced, so it is resolved when the socket opens and again on reconnect. Reconnecting replays the last 200 lines, so a refresh keeps recent history.

No dashboard changes: It already opens whatever ws_url it is given and sends the subscription. The gateway already proxies /v1/deployment/ws and rejects connections without a valid token, so this rides an authenticated path that skypilot and nosana already use.

Testing
16 new tests. 10 on the adapter: ordering, lines split across chunks, the response being closed when the caller stops, following the right pod, a missing pod raising rather than hanging, undecodable bytes, that the follow
stream never goes through the shared executor, and that the close never runs on the event loop. 6 on the endpoint: the subscription reaching the adapter with the right arguments, the message framing the dashboard parses, a
subscription without an instance being refused, and adapter failures being reported rather than swallowed.

Known limits
Only the first pod matching the label is followed, so a deployment scaled past one replica streams one of them with no indication.

Verified live
Tested on a GPU box against a running Qwen2.5-14B deployment. Opened and closed Terminal Logs, then chatted through the Sandbox: the orchestration service stayed healthy, new log lines appeared in the browser while the model
was serving, and a thread dump after teardown showed no reader threads left behind.

Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
@Priyanshu-u07
Priyanshu-u07 marked this pull request as ready for review September 24, 2026 10:26
@Priyanshu-u07
Priyanshu-u07 merged commit c8c3d8e into main Sep 24, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Terminal Logs is a snapshot on Kubernetes, not a live stream

1 participant