Conversation
The rowwise API executor now sets APIResult.usage to the task's summed token and call accounting, the same numbers the api summary line already logs. APIUsage is a typed model (prompt/completion/reasoning tokens, calls, failures, retries, truncated_calls, wall_sec) replacing the loose dict; a call whose response lacks usage counts in calls and adds no tokens. Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
GET /workflows/{id} now returns a usage object summed over the workflow's finished tasks, computed on read from each task result's usage: API tasks contribute their APIUsage, vLLM inference tasks map their GenerationUsage token counts (reasoning 0, calls from num_requests), and tasks with no usage contribute nothing. The SDK Workflow model gains a typed usage field.
Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Compute each task's usage once at result ingest and store it in the workflow registry's Redis, so GET /workflows/{id} sums stored values without opening result files on every poll. Record the no-usage marker for cloned (merged) children, which make no calls of their own. A model-calling task whose usage cannot be mapped, or a completed task with no recorded usage, makes the workflow's usage null with one log line; a task type that makes no model calls contributes nothing. Return the (possibly all-zero) sum when every completed task is recorded. Unregistering a workflow now deletes its tasks' usage keys too.
Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu>
Co-Authored-By: Claude Code <noreply@anthropic.com>
A merged vLLM dispatch reports the whole batch total on the parent and each child's share in result.children, and every child is ingested separately. The parent now records its own share (total minus the sum over its children's usages) so the workflow sum counts each child's tokens and calls once. wall_sec stays per task as reported. A mirrored child with no result of its own keeps the no-usage marker, since the parent's total already contains its calls. EventMonitor now requires a WorkflowRegistry instead of an optional one. Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
timzsu
added this pull request to stack #156
September 28, 2026 08:44
…into zsu/api-usage Bring in PR 141's review round 2: spec.api is a typed ApiConfig / ApiResponseConfig with range bounds enforced at submission, the request body field is json_body (alias json), _MAX_RETRIES and _MAX_CONCURRENCY live in misc.py, and early stop on a row failure is back. Resolved the one textual conflict in tests/worker/test_api_executor.py's import block to keep both the APIUsage import and the misc constants import. Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
Resolve the api executor test import conflict: keep both 157's python-stage reading (PythonResult) and 161's token usage (APIUsage). Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude <noreply@anthropic.com>
… zsu/api-usage Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
PythonResult was missing from _NO_USAGE_RESULT_TYPES, so every python task mapped to UNKNOWN_USAGE and failed the workflow usage sum closed to null. Add it to the no-usage tuple so python tasks contribute nothing, and cover it with a workflow-sum test and an envelope test. Signed-off-by: Zhengyuan Su <su.zhengyuan@u.nus.edu> Co-Authored-By: Claude Code <noreply@anthropic.com>
Collaborator
Author
|
This PR is closed because the features are folded to #157. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
A caller that runs a workflow had no way to learn what it cost. The API executor already logged per-call token counts, but nothing reached the result or the workflow. This PR reports token and call usage on each API task's result and sums it per workflow on
GET /workflows/{id}, so Lumilake can show per-job usage.Changes
src/shared/schemas/result,api_executor.py).APIUsageis a typed model withprompt_tokens,completion_tokens,reasoning_tokens,calls,failures,retries,truncated_calls(calls that stopped on the length limit) andwall_sec.APIResult.usagecarries the task's sum, the same numbers the executor's summary line logs. A call whose response has no usage counts incallsand adds no tokens.routers/v1/results.py,registries/workflow.py,clients/redis.py). Each task's usage is taken from its result envelope when the result arrives and stored in the workflow registry's Redis undertask:{id}:usage.GET /workflows/{id}sums the stored values and never opens a result file. vLLM inference tasks map theirGenerationUsage:callsfromnum_requests,wall_secfromlatency_sec, reasoning 0.result.children, and each child is ingested separately. The parent records its own share (total minus its children's), so the workflow sum counts every call once. A merged child whose result is copied from the parent (no result of its own) records no usage, because the parent's total already holds its calls.usagenull, with one log line naming the task. Task types that make no model calls (echo, SSH, data profiling, data retrieval) contribute nothing. When every completed task is recorded, the sum is returned, even when it is all zero.EventMonitortakes the workflow registry as a required argument.Workflowmodel gains a typedusagefield.Design
Usage is stored per task at ingest rather than summed on read. Lumilake polls
GET /workflows/{id}for progress, and summing on read reopened and re-parsed every completed task's result file on each poll.Test Plan
The repo's pre-commit gate on the changed files (gitleaks, isort, black, ruff, codespell, mypy), then
tests/server,tests/workerandtests/shared. The new tests cover:read_result.Each rule was also neutralised in src on its own to confirm its test fails.
Test Result
tests/server tests/worker tests/shared: 1905 passed, 1 failed. The failure istest_mp_executor_cleans_up_vllm, a GPU test whose vLLM engine does not start in this environment. This PR changes no vLLM worker code.Pre-submission Checklist
pre-commit run --all-filesand fixed any issues.uv run pytest tests/passes locally.uv sync --all-packages --group ci --frozen).[BREAKING]and described migration steps above.