Skip to content

fix: capture a deployment logs before termination removes them - #366

Merged
Priyanshu-u07 merged 1 commit into
mainfrom
persist-terminal-logs
Sep 28, 2026
Merged

Priyanshu-u07 merged 1 commit into
mainfrom
persist-terminal-logs

Conversation

@Priyanshu-u07

Copy link
Copy Markdown
Collaborator

TerminalLogRepository.save() had no callers, so GET /logs/{deployment_id}/persisted returned nothing for every deployment on every provider.

The capture now runs in terminate_deployment_core, before unload_model stops the engine — afterwards the pod is gone and the logs with it. It is best-effort and bounded at 10 seconds, so a slow or unreachable provider cannot stall or fail a termination. The tail is capped at 500 lines because the rows are written inside the termination transaction.

FAILED and STOPPED are deliberately not wired, and the docstring now says so. STOPPED is set in a finally block after cleanup has already recycled the nodes, so it needs its own pre-teardown capture point inside worker.py. Most FAILED transitions are provisioning failures where no engine ever started.

Also removes the dangling TerminalLogRepository construction in worker_main.py that the issue names.

Closes #360

Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
@Priyanshu-u07
Priyanshu-u07 merged commit 6ff2ac4 into main Sep 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Persisted terminal logs are never written

1 participant