Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
a693bed
docs(core): every investigation lives in an issue (spec 0001 M5)
bordumb Sep 28, 2026
bf2f461
fix(temporal): LLM failures fail the investigation with the reason
bordumb Sep 28, 2026
e9f99c5
feat(api): every investigation starts in an issue
bordumb Sep 28, 2026
1eb1564
feat(api): check the Anthropic key at startup and say what to fix
bordumb Sep 28, 2026
50abdaf
docs(api): record the M5 API shapes in openapi.json
bordumb Sep 28, 2026
3a71fb9
feat(frontend): Investigate… starts a run from every page (fn-70.22)
bordumb Sep 28, 2026
8192ad9
feat(frontend): the issue page follows the mockup (fn-70.23)
bordumb Sep 28, 2026
049049d
feat(frontend): the run's details page and the LLM status banner (fn-…
bordumb Sep 28, 2026
98ab1bf
fix(frontend): "In progress" in sentence case, as in the mockup
bordumb Sep 28, 2026
7248b65
refactor(core): remove the dormant ways to create an investigation
bordumb Sep 28, 2026
28e3ef7
fix(core): name Claude models in one place and run investigations on …
bordumb Sep 28, 2026
bdd043c
fix(api): send people to sign in when their session is rejected
bordumb Sep 28, 2026
cf09d88
fix(frontend): default an issue's observed date to today where you are
bordumb Sep 28, 2026
1f2e724
fix(core): run issue chat and brief drafting on Claude Sonnet 5.5 too
bordumb Sep 28, 2026
2c8e951
fix(frontend): pick an investigation's tables and days with the old p…
bordumb Sep 29, 2026
ea03b22
feat(core): a confirmed root cause can become a check at any confidence
bordumb Sep 29, 2026
6e12ad2
test(api): stand in for the starter in the end-to-end create test
bordumb Sep 29, 2026
73d1566
feat(frontend): a page to save your own login for a datasource
bordumb Sep 29, 2026
bf4d220
Merge origin/main into claude/issue-first-investigations
bordumb Sep 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -50,9 +50,11 @@ TEMPORAL_TASK_QUEUE=investigations
# API Configuration
# -----------------------------------------------------------------------------

# LLM model to use for investigations
# Options: claude-sonnet-4-20250514, claude-opus-4-20250514
LLM_MODEL=claude-sonnet-4-20250514
# Claude models
# Model overrides. Unset, the defaults in python-packages/dataing/src/dataing/config.py
# apply (claude-sonnet-5-5 for investigations and issue chat).
# LLM_MODEL=claude-sonnet-5-5
# CHAT_AGENT_MODEL=claude-sonnet-5-5

# -----------------------------------------------------------------------------
# Optional: Local Data Sources
Expand Down
2 changes: 1 addition & 1 deletion .flow/tasks/fn-69.12.json
Original file line number Diff line number Diff line change
Expand Up @@ -13,5 +13,5 @@
"spec_path": ".flow/tasks/fn-69.12.md",
"status": "todo",
"title": "[M1] Frontend: per-user datasource credentials page",
"updated_at": "2026-09-27T15:11:28.435439Z"
"updated_at": "2026-09-29T01:05:48.953573Z"
}
13 changes: 10 additions & 3 deletions .flow/tasks/fn-69.12.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,11 +2,18 @@

## Description
### Goal
Build the page that the gateway's 403 links to (`/settings/datasources/{id}/credentials`). It doesn't exist yet (design §5.2).
Extend the per-user credentials page that the gateway's 403 and the issue agent's `credentials_missing` link go to (`/settings/datasources/{id}/credentials`, design §5.2).

A first cut already exists from the issue hub (spec 0001 §8.5): `features/settings/datasource-credentials-page.tsx`, with hooks in `lib/api/credentials.ts`. It has:
- username and password, plus role and warehouse where the source type's config takes them
- a test of the login before it is saved
- Remove
- a message for source types with no login
- a return to the page that linked to it

### Implementation
- **Page:** a route and page under `features/settings/`, with forms generated from fn-69.31's schema for each adapter: username/password, key pair, or service-account JSON.
- Test and Delete buttons.
- **Page:** replace the fixed fields with forms generated from fn-69.31's schema for each adapter: username/password, key pair, or service-account JSON.
- Keep test-before-save and Remove.
- Never echo secrets back.
- **403 handling:** catch 403 `credentials_not_configured` globally in `lib/api/client.ts`, next to the existing `feature_not_available` handling at `:58-65`. Show a toast that links to the page.
## Acceptance
Expand Down
20 changes: 16 additions & 4 deletions .flow/tasks/fn-70.17.json
Original file line number Diff line number Diff line change
@@ -1,14 +1,26 @@
{
"assignee": null,
"assignee": "bordumbb@gmail.com",
"claim_note": "",
"claimed_at": null,
"claimed_at": "2026-09-28T17:52:29.670737Z",
"created_at": "2026-09-28T15:16:44.759472Z",
"depends_on": [],
"epic": "fn-70",
"evidence": {
"commits": [
"1a31082c"
],
"prs": [
"https://github.com/bordumb/dataing/pull/211"
],
"tests": [
"tests/unit/test_config.py::test_chat_agent_defaults_to_claude_opus_5_5",
"tests/unit/agents/test_chat_agent.py::TestBriefDrafting::test_draft_is_parsed_from_a_json_reply_without_forcing_a_tool"
]
},
"id": "fn-70.17",
"priority": null,
"spec_path": ".flow/tasks/fn-70.17.md",
"status": "todo",
"status": "done",
"title": "Run the issue chat agent on Claude Opus 5.5",
"updated_at": "2026-09-28T15:16:44.762271Z"
"updated_at": "2026-09-28T17:52:29.952602Z"
}
9 changes: 4 additions & 5 deletions .flow/tasks/fn-70.17.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,8 @@ TBD


## Done summary
TBD

CHAT_AGENT_MODEL defaults to claude-opus-5-5; brief drafting returns JSON text (PromptedOutput) instead of a forced output tool, which Opus 5.5 rejects. Merged as PR #211 (1a31082c).
## Evidence
- Commits:
- Tests:
- PRs:
- Commits: 1a31082c
- Tests: tests/unit/test_config.py::test_chat_agent_defaults_to_claude_opus_5_5, tests/unit/agents/test_chat_agent.py::TestBriefDrafting::test_draft_is_parsed_from_a_json_reply_without_forcing_a_tool
- PRs: https://github.com/bordumb/dataing/pull/211
21 changes: 21 additions & 0 deletions .flow/tasks/fn-70.18.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
{
"assignee": "bordumbb@gmail.com",
"claim_note": "",
"claimed_at": "2026-09-28T17:54:28.759022Z",
"created_at": "2026-09-28T17:51:58.959592Z",
"depends_on": [],
"epic": "fn-70",
"evidence": {
"commits": [
"f292971a"
],
"prs": [],
"tests": []
},
"id": "fn-70.18",
"priority": null,
"spec_path": ".flow/tasks/fn-70.18.md",
"status": "done",
"title": "M5: spec revision: every run in the hub, one-step start, details page, LLM failures",
"updated_at": "2026-09-28T17:54:29.035901Z"
}
16 changes: 16 additions & 0 deletions .flow/tasks/fn-70.18.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# fn-70.18 M5: spec revision: every run in the hub, one-step start, details page, LLM failures

## Description
TBD

## Acceptance
- docs/specs/0001_issue_chat.md records D12–D15, the §2 "found after M1–M4" table, §7.11 (one starter, open_issue, every path), §7.12 (LLM error table, failing the run, key check), §8.1–8.4 (mockup fidelity, Investigate…, details page, banner), M5 in §10 and its tests in §11
- Owner decisions (2026-09-28): one-click Investigate…, details page off the card, always open an issue, fail the run + key check


## Done summary
Spec revision committed: D12-D15, §7.11, §7.12, §8.1-8.4, M5.
## Evidence
- Commits: f292971a
- Tests:
- PRs:
30 changes: 30 additions & 0 deletions .flow/tasks/fn-70.19.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
{
"assignee": "bordumbb@gmail.com",
"claim_note": "",
"claimed_at": "2026-09-28T17:55:54.224028Z",
"created_at": "2026-09-28T17:51:59.221401Z",
"depends_on": [
"fn-70.18"
],
"epic": "fn-70",
"evidence": {
"commits": [
"b54b8725"
],
"prs": [],
"tests": [
"tests/unit/agents/test_llm_errors.py",
"tests/unit/temporal/test_llm_activity_errors.py",
"tests/unit/temporal/test_investigation_failures.py",
"tests/unit/temporal/test_investigation_replay.py",
"tests/integration/test_publish_outcome.py",
"CE unit suite: 3169 passed"
]
},
"id": "fn-70.19",
"priority": null,
"spec_path": ".flow/tasks/fn-70.19.md",
"status": "done",
"title": "M5: LLM failures fail the run (classify, activities raise, workflow fails)",
"updated_at": "2026-09-28T18:22:06.119415Z"
}
20 changes: 20 additions & 0 deletions .flow/tasks/fn-70.19.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# fn-70.19 M5: LLM failures fail the run (classify, activities raise, workflow fails)

## Description
TBD

## Acceptance
- `agents/errors.py` `classify_llm_error(exc)` maps the real exception chain (LLMError → ModelHTTPError/ModelAPIError → anthropic errors, UserError for a missing key) to the §7.12 codes, retryability and messages; unit-tested per row
- generate_hypotheses, generate_query, interpret_evidence, synthesize and counter_analyze raise ApplicationError `LLMRejected` (non-retryable) or `LLMUnavailable` with `{code, message}` details; AgentClient.interpret_evidence no longer swallows errors
- Every LLM activity call has an explicit RetryPolicy (4 attempts, 5 s, ×2, max 60 s, LLMRejected non-retryable)
- Behind `workflow.patched("llm-failures-v1")`, the workflow fails the run on: generation failure or no hypotheses, any subagent LLM error (others cancelled), all hypotheses untested from errors, synthesis failure. Counter-analysis failure keeps the synthesis and records counter_analysis.error
- A failed run publishes `{"status":"failed","error":{code,message,step}}` via publish_investigation_outcome (outcome, run row completed_at, thread card), then raises a non-retryable ApplicationError
- Workflow tests with fake activities cover each case; Replayer tests pass on the recorded histories


## Done summary
classify_llm_error (agents/errors.py) maps anthropic → pydantic-ai → LLMError chains to the §7.12 codes. LLM activities raise LLMRejected/LLMUnavailable with an explicit retry policy; AgentClient.interpret_evidence no longer swallows errors. Behind llm-failures-v1 the workflow fails the run (generation failure/no hypotheses, subagent LLM error cancelling the others, nothing testable, synthesis failure), publishes {"status":"failed","error":{code,message,step}} and fails the Temporal execution as InvestigationFailed. Counter-analysis failure keeps the conclusion.
## Evidence
- Commits: b54b8725
- Tests: tests/unit/agents/test_llm_errors.py, tests/unit/temporal/test_llm_activity_errors.py, tests/unit/temporal/test_investigation_failures.py, tests/unit/temporal/test_investigation_replay.py, tests/integration/test_publish_outcome.py, CE unit suite: 3169 passed
- PRs:
29 changes: 29 additions & 0 deletions .flow/tasks/fn-70.20.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
{
"assignee": "bordumbb@gmail.com",
"claim_note": "",
"claimed_at": "2026-09-28T19:34:45.513313Z",
"created_at": "2026-09-28T17:51:59.475477Z",
"depends_on": [
"fn-70.19"
],
"epic": "fn-70",
"evidence": {
"commits": [
"784d88d6"
],
"prs": [],
"tests": [
"tests/unit/services/test_llm_status.py",
"tests/unit/entrypoints/api/routes/test_system.py",
"tests/unit/temporal/test_issue_thread_workflow.py",
"tests/integration/test_agent_turn_activity.py",
"CE unit 3186, CE integration 145, EE unit 714"
]
},
"id": "fn-70.20",
"priority": null,
"spec_path": ".flow/tasks/fn-70.20.md",
"status": "done",
"title": "M5: LLM key check at startup, GET /system/llm, readable chat errors",
"updated_at": "2026-09-28T20:10:32.733100Z"
}
18 changes: 18 additions & 0 deletions .flow/tasks/fn-70.20.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# fn-70.20 M5: LLM key check at startup, GET /system/llm, readable chat errors

## Description
TBD

## Acceptance
- API startup checks `GET /v1/models/{id}` for LLM_MODEL and CHAT_AGENT_MODEL in the background (10 s timeout, no SDK retries); an empty key is reported without a request
- `GET /api/v1/system/llm` (ANY_USER, POLICY entry) returns {state, message, models, checked_at}; stale `unreachable` re-checks on read
- run_agent_turn classifies model errors: the reply/brief shows the §7.12 message, not the raw ModelHTTPError; non-retryable errors aren't retried
- Tests fake the Anthropic client for missing, rejected (401), unknown model (404) and ok


## Done summary
LLMStatusChecker checks GET /v1/models/{id} for LLM_MODEL and CHAT_AGENT_MODEL in the background at API startup (10 s timeout, no SDK retries; empty key reported without a request). GET /api/v1/system/llm returns {state, message, models, checked_at}; transient problems re-check after 60 s. Chat turns and brief drafts raise the classified LLMRejected/LLMUnavailable error, so replies show what to fix and rejected keys aren't retried. Verified live: a bogus key → invalid_key.
## Evidence
- Commits: 784d88d6
- Tests: tests/unit/services/test_llm_status.py, tests/unit/entrypoints/api/routes/test_system.py, tests/unit/temporal/test_issue_thread_workflow.py, tests/integration/test_agent_turn_activity.py, CE unit 3186, CE integration 145, EE unit 714
- PRs:
29 changes: 29 additions & 0 deletions .flow/tasks/fn-70.21.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
{
"assignee": "bordumbb@gmail.com",
"claim_note": "",
"claimed_at": "2026-09-28T18:22:06.382549Z",
"created_at": "2026-09-28T17:51:59.729817Z",
"depends_on": [
"fn-70.18"
],
"epic": "fn-70",
"evidence": {
"commits": [
"d7522281"
],
"prs": [],
"tests": [
"tests/integration/test_open_issue.py",
"tests/integration/api/test_start_investigation.py",
"tests/integration/api/test_issue_investigation_runs.py",
"dataing-ee tests/integration/core/automation/test_spawn_investigation.py",
"CE unit 3173, EE unit 714, CE integration 143, EE integration 16, SDK+CLI 272"
]
},
"id": "fn-70.21",
"priority": null,
"spec_path": ".flow/tasks/fn-70.21.md",
"status": "done",
"title": "M5: one starter and open_issue() for every start path",
"updated_at": "2026-09-28T19:34:45.173067Z"
}
22 changes: 22 additions & 0 deletions .flow/tasks/fn-70.21.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# fn-70.21 M5: one starter and open_issue() for every start path

## Description
TBD

## Acceptance
- `adapters/db/issues.py` `open_issue()` is the only issue insert (POST /issues, CE + EE webhooks, the starter) and posts the thread's first entry via the `created` event
- `InvestigationStarterService.start()` resolves or opens the issue, builds the missing brief/alert, writes the investigation row, run row (trigger_type human/api/webhook/rule) and start card, starts the workflow with alert.issue_id, and records a failed outcome if the start fails
- POST /investigations takes exactly one of brief/alert (+ datasource_id, execution_profile, issue_id); bad alert → 422; response adds run_id, issue_id, issue_number
- POST /issues/{id}/investigation-runs, CE webhook-generic AUTO, EE provider webhook AUTO and the EE rule action all use the starter
- InvestigationRunResponse gains number, status, error; InvestigationStateResponse gains issue_id, issue_number, issue_title, run_number, brief, execution_profile, error
- ToolCallRecord stores duration_ms and row_count for run_query
- SDK Investigation gains issue_id/issue_number; `dataing run start` prints the issue URL
- Integration tests on migrated_db for each path in the §7.11 table


## Done summary
open_issue() is the only issue insert (API, CE + EE webhooks, starter) and posts the thread's opening event. InvestigationStarterService.start() resolves/opens the issue, builds the missing brief/alert, writes investigation + run row + card, starts the workflow with alert.issue_id, and records a failed start on the card. POST /investigations takes brief|alert (+issue_id, profile) and returns issue_id/issue_number/run_id; spawn route, webhooks and EE rule action use the starter. Runs report number/status/error; GET /investigations/{id} adds issue, run number, brief, error, hypotheses. Tool calls record duration_ms/row_count. SDK Investigation gains issue fields; CLI links the issue.
## Evidence
- Commits: d7522281
- Tests: tests/integration/test_open_issue.py, tests/integration/api/test_start_investigation.py, tests/integration/api/test_issue_investigation_runs.py, dataing-ee tests/integration/core/automation/test_spawn_investigation.py, CE unit 3173, EE unit 714, CE integration 143, EE integration 16, SDK+CLI 272
- PRs:
26 changes: 26 additions & 0 deletions .flow/tasks/fn-70.22.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
{
"assignee": "bordumbb@gmail.com",
"claim_note": "",
"claimed_at": "2026-09-28T20:41:22.721926Z",
"created_at": "2026-09-28T17:51:59.979116Z",
"depends_on": [
"fn-70.21"
],
"epic": "fn-70",
"evidence": {
"commits": [
"3a71fb9b"
],
"prs": [],
"tests": [
"StartInvestigation.test.tsx",
"frontend vitest 172 passed"
]
},
"id": "fn-70.22",
"priority": null,
"spec_path": ".flow/tasks/fn-70.22.md",
"status": "done",
"title": "M5: frontend: Investigate\u2026 everywhere, remove /investigations/new",
"updated_at": "2026-09-28T20:41:23.797636Z"
}
18 changes: 18 additions & 0 deletions .flow/tasks/fn-70.22.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# fn-70.22 M5: frontend: Investigate… everywhere, remove /investigations/new

## Description
TBD

## Acceptance
- /investigations/new route, NewInvestigation.tsx and the dead components/Layout.tsx are gone; no link points at /investigations/new
- Investigate… (sidebar quick action, dashboard header + empty state, investigations list header + empty state, dataset page header "Investigate this dataset") opens the brief editor in new mode, pre-filled from the page
- New mode: "Start an investigation", symptom + ≥1 scope table required, datasource required only with >1 datasource; Start → POST /investigations → navigate to /issues/{issue_id}
- vitest covers each entry point and the new-mode submit + navigation


## Done summary
Investigate… everywhere: StartInvestigationDialog/InvestigateButton (brief editor in new mode) on the sidebar, dashboard, investigations list and dataset page; POST /investigations → /issues/{id}; /investigations/new, NewInvestigation.tsx and Layout.tsx removed.
## Evidence
- Commits: 3a71fb9b
- Tests: StartInvestigation.test.tsx, frontend vitest 172 passed
- PRs:
28 changes: 28 additions & 0 deletions .flow/tasks/fn-70.23.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
{
"assignee": "bordumbb@gmail.com",
"claim_note": "",
"claimed_at": "2026-09-28T20:41:23.059873Z",
"created_at": "2026-09-28T17:52:00.239444Z",
"depends_on": [
"fn-70.21"
],
"epic": "fn-70",
"evidence": {
"commits": [
"8192ad97"
],
"prs": [],
"tests": [
"IssueWorkspace.test.tsx",
"IssueSidebar.test.tsx",
"IssueThread.test.tsx",
"InvestigationCards.test.tsx"
]
},
"id": "fn-70.23",
"priority": null,
"spec_path": ".flow/tasks/fn-70.23.md",
"status": "done",
"title": "M5: frontend: issue page matches the mockup",
"updated_at": "2026-09-28T20:41:24.133203Z"
}
19 changes: 19 additions & 0 deletions .flow/tasks/fn-70.23.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# fn-70.23 M5: frontend: issue page matches the mockup

## Description
TBD

## Acceptance
- Issue page matches 0001_issue_chat_mockup.html: one-row top bar; description as the thread's first entry; tabs Shared thread / My scratch chats (N) + "N watching · live"; one sidebar panel (Status with note, Details, Dataset + open dataset page, Investigations, Watchers, Your scratch chats); mockup pill colours
- Tool calls collapse to "Ran N queries · X ms · Y rows"; footer "Snapshot saved with this message · copy SQL"
- Investigation card: "Investigation #N", details → link, failed pill + reason + Retry (reopens the editor with the same brief); headless runs attributed to dataing
- Sidebar runs show number, depth and status (failed no longer "running")
- vitest updated/added; screenshots compared against the mockup


## Done summary
Issue page follows the mockup: one-row top bar, IssueOpened first entry (description, dataing-opened line), Shared thread / My scratch chats tabs with watching · live, one sidebar panel (status note, details, dataset link, runs #N with status, watchers, scratch chats), mockup pills, tool-call totals, Investigation #N card with details → and failed state + Retry.
## Evidence
- Commits: 8192ad97
- Tests: IssueWorkspace.test.tsx, IssueSidebar.test.tsx, IssueThread.test.tsx, InvestigationCards.test.tsx
- PRs:
Loading
Loading