Skip to content

feat(core): every investigation starts in an issue (spec 0001 M5) - #233

Merged
bordumb merged 19 commits into
mainfrom
claude/issue-first-investigations
Sep 29, 2026
Merged

bordumb merged 19 commits into
mainfrom
claude/issue-first-investigations

Conversation

@bordumb

@bordumb bordumb commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Spec 0001's M5: the fixes for what testing the issue hub turned up (spec §2 "Found after M1–M4", decisions D12–D15), plus what the demo walkthrough found after that.

Investigations

  • Every investigation starts in an issue (D12). InvestigationStarterService.start() is the only way a run starts: the UI, the API and SDK, webhooks and rules all go through it. A run started without an issue opens one titled with its symptom. The dormant creation paths (queue adapter, old investigation service and repository) are removed.
  • One-click Investigate… (D13). Every start button opens the brief editor in new mode and lands on the new issue. /investigations/new is gone. The editor picks tables with the old page's autocomplete (DatasetEntry) and days with its DatePicker. Leads are hand-off only.
  • The run's details page (D14). /investigations/:id links back to its issue and shows the brief, the hypotheses and how each ended, the evidence, and the outcome.
  • LLM failures fail the run (D15). The run fails with a readable reason instead of "finishing" with nothing tested. Retries are bounded, gated behind workflow.patched("llm-failures-v1") with replay tests. The API checks the Anthropic key at startup, and a banner says what to fix.
  • Add as check follows the review. A confirmed root cause can become a check at any confidence, and a rejected one never can. An unreviewed one still needs 60%, and the card says why when it's below that.

Issue chat

  • A page for your own database login at /settings/datasources/{id}/credentials, laid out like the sign-in page (spec §8.5). The agent linked there on credentials_missing, but the page didn't exist, so chat couldn't query anything. It tests the login before saving it and never shows the password back. The agent now writes that link as Markdown, and app links in the thread open in place.

Models

  • dataing.config is the only place that names a Claude model, and a test enforces that. Investigations, chat and brief drafting run on claude-sonnet-5-5. The old default, claude-sonnet-4-20250514, was retired and made every investigation fail with a 404.

Other fixes

  • A rejected session sends people to sign in instead of every page showing "Missing API key".
  • A new issue's observed date defaults to today where the person is, not in UTC.
  • The end-to-end create test stands in for the starter, which has its own tests.

Test plan

  • On this branch merged with main: CE pytest (3,129 passed), EE pytest (715 passed), mypy, and ruff check
  • Frontend: 192 vitest tests, tsc, eslint, prettier
  • Workflow replay tests for the llm-failures-v1 patch
  • The demo stack run from this branch: starting runs, the details page, chat on Sonnet 5.5, the key check

🤖 Generated with Claude Code

bordumb and others added 19 commits September 28, 2026 21:27
Revise spec 0001 with the owner's decisions from 2026-09-28:
- D12: every run belongs to an issue; API, SDK, webhook and rule runs
  open or reuse one
- D13: Investigate… opens the brief editor where you are and lands on
  the new issue's thread; /investigations/new goes
- D14: /investigations/:id becomes the run's details page
- D15: LLM failures fail the run with the reason; the API checks the key
  and models at startup and every page shows a banner

Adds §7.11 (one starter, open_issue(), every start path), §7.12 (LLM
error table, failing the run, key check), §8.1-8.4 (the mockup as the
acceptance reference) and milestone M5, filed as fn-70.18-25.
fn-70.17 is closed: PR #211 shipped it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
With a rejected key, every LLM activity caught the error and returned
empty results. The workflow logged warnings, synthesized from nothing
and published "completed" at confidence 0. AgentClient also turned
interpretation errors into evidence that read as refuted.

Now (spec 0001 §7.12):
- classify_llm_error maps the real exception chain (anthropic →
  pydantic-ai → LLMError) to a code, retryability and what to fix
- LLM activities raise LLMRejected (non-retryable) or LLMUnavailable
  and retry with LLM_RETRY_POLICY (4 attempts, 5-60 s backoff)
- behind the llm-failures-v1 patch, the run fails when:
  - hypothesis generation fails or proposes nothing
  - a subagent hits an LLM error (the others are cancelled)
  - every hypothesis is left untested by errors
  - synthesis fails
- a failed counter-analysis keeps the conclusion and records the error
- a failed run publishes {"status": "failed", "error": {code, message,
  step}} to the investigation, run row and thread, then fails its
  Temporal execution as InvestigationFailed

Recorded histories still replay. The datasource-down test now expects a
failed run instead of a conclusion drawn from no evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Of the 8 ways a run was created, only starting it from an issue page
linked it fully. POST /investigations (UI, SDK, CLI, notebook) and the
AUTO webhooks inserted runs inline, with no run row or start card, and
the EE rule action never wrote its outcome back.

Now (spec 0001 §7.11):
- open_issue() (adapters/db/issues.py) is the only issue insert. The
  API, both webhooks and the starter use it, and every thread starts
  with the event that opened it
- InvestigationStarterService.start():
  - resolves the given issue or opens one
  - builds the missing brief or alert
  - writes the investigation, the run row (trigger_type human/api/
    webhook/rule) and the thread card
  - starts the workflow with alert.issue_id
  - on a failed start, records the failure on the card
- POST /investigations takes a brief or an alert (a bad alert is a 422,
  not a 500) and returns run_id, issue_id and issue_number
- the issue spawn route, both webhooks and the EE rule action go through
  the starter
- runs report their number, status and error; GET /investigations/{id}
  adds the issue, run number, brief and failure for the details page
- tool calls record duration_ms and row_count for "Ran 1 query · 212 ms"
- the SDK's Investigation gains issue_id and issue_number;
  `dataing run start` links the issue

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Nothing checked the key before, so a rejected key first showed up as a
run that did nothing, and a chat reply or brief draft that failed showed
the raw ModelHTTPError text.

Now (spec 0001 §7.12):
- at startup the API calls GET /v1/models/{id} for LLM_MODEL and
  CHAT_AGENT_MODEL (no tokens). It runs in the background with a 10 s
  timeout and no retries, so startup never waits on it
- GET /api/v1/system/llm returns {state, message, models, checked_at}
  for the app's banner. A transient problem is checked again once it is
  over a minute old
- chat turns and brief drafts raise the classified LLMRejected or
  LLMUnavailable error. The reply shows what to fix, and a rejected key
  isn't retried

Checked against the real API with a bogus key: invalid_key, "Anthropic
rejected the API key (401). Set a valid ANTHROPIC_API_KEY and restart
the API and the worker."

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Copies only the changed operations from the app's schema:
- POST /investigations (brief or alert, issue link in the response)
- GET /investigations/{id} (issue, run number, brief, error, hypotheses)
- the issue's investigation runs and outcome review (number, status,
  error)
- the new GET /system/llm

Also adds the schemas these reference. The rest of the committed spec
is untouched.

The TypeScript client is not regenerated: orval 8 (#221) would rewrite
all 293 generated files (+31k lines) and turn GET hooks into mutations.
The frontend reaches these routes through its hand-written wrappers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Every start button now opens the brief editor in new mode where the
person is, instead of the /investigations/new page (spec 0001 D13, §8.2):
the sidebar quick action, the dashboard header and its empty
recent-investigations card, the investigations list header and empty
state, and "Investigate this dataset" in the dataset page header, which
is shown whether or not the dataset has runs and pre-fills the table's
native path and datasource.

New mode reuses BriefFormView, titled "Start an investigation":
- symptom and at least one scope table are required
- the datasource select lists the tenant's datasources, preselects the
  only one and requires a choice when there are several; it has no
  "The issue's datasource" option
- findings, ruled out and leads start empty
- Start posts the brief to POST /api/v1/investigations and lands on
  /issues/{issue_id}; a 409 ambiguous_datasource asks the person to pick
  a datasource, and other errors toast the server's message

The API client now throws ApiError, which keeps the status and the
detail's error code. BriefFormView takes a mode and an onSubmit; the
hand-off submit moved into HandoffForm unchanged.

Removed: the investigations/new route, NewInvestigation.tsx, the
alert-based useStartInvestigation and the dead components/Layout.tsx.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
0001_issue_chat_mockup.html is the acceptance reference (spec 0001 §8.1),
and the page now follows it:

- A one-row top bar: the #N pill, the title, the status and priority
  pills and "dataset x" on the right, with the back link.
- The thread (1fr) beside one sidebar panel (300px), stacked on narrow
  screens.
- The thread head has the Shared thread / My scratch chats (N) tabs (the
  second opens the scratch drawer) and "N watching · live".
- No description card: the thread's first entry is the `created` event,
  "Maya opened the issue" with the issue's description and Edit for its
  author (PATCHes the description), or "⚑ Issue opened by dataing from x"
  when nobody opened it by hand. Threads without that event still show
  the description, from the issue.
- One sidebar panel: Status (pill, "Change ▾", the menu's note "Only
  moves that will succeed are shown" and the Resolved hint when a cause
  was confirmed), Details (compact editors; labels as pills with +),
  Dataset with "open dataset page →" when a datasource has synced the
  table, Investigations ("#N · depth" with running / done · 0.91 /
  confirmed / rejected / failed), Watchers, Your scratch chats and
  + New scratch chat. The Timeline and the separate cards are gone.
- The investigation card is "Investigation #N" (the run's number from
  the runs list), with the mockup's one-line brief, "details →" in place
  of "Open", and + Add context / + Add hypothesis. A ruled-out
  hypothesis reads "ruled out by Maya · stopped".
- Failed runs: the failed outcome renders the reason, the step and
  Retry (members), which reopens the brief editor with the run's brief.
  The start card turns red with the reason, and offers Retry itself only
  when no failed outcome was posted. The sidebar row says failed.
- A run nobody started by hand reads "dataing · started an
  investigation".
- Tool calls collapse to "▸ Ran 1 query · 212 ms · 4 rows" from the
  record's duration_ms and row_count (older messages fall back to "Ran 1
  query"); the expanded footer adds "Snapshot saved with this message ·
  Copy SQL".
- The outcome card notes when the check of its conclusion didn't run,
  and proposals name their run "#N".

The runs list is hand-typed with number, status and error, refreshes
while a run is running and when an outcome lands in the thread.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
…70.24)

/investigations/:id becomes the run's details page (spec 0001 D14, §8.3):

- Header: "← #42 {issue title}" back to the issue's thread (omitted for
  imported snapshots, which have no issue), "Investigation #N" with its
  status and depth pills, Share (copies the link; the mocked user picker
  is gone), Export snapshot (GET /investigations/{id}/snapshot through
  the API client, saved as snapshot-{id}.tar.gz), Add as check (renamed
  from Codify Test, same gating) and Cancel run while it runs.
- Brief: the brief the run was given.
- Hypotheses: each with its status pill ("ruled out by a person",
  "untested" in amber) and its evidence, grouped by hypothesis_id.
  Queries collapse like the thread's tool calls and expand to the SQL,
  the result summary and the interpretation, with feedback.
- Outcome: the root cause in the thread card's style, with the
  hypothesis pills, causal chain, onset, affected scope, recommendations,
  supporting evidence and the existing feedback buttons.
- A failed run shows the reason, the step and what to do next, with a
  link back to the issue to retry, in place of the outcome. It never
  shows "Unable to determine a definitive root cause".
- The step-timeline card is gone: the header's "running · phase" pill
  and live dot are the compact phase indicator.

A banner under the app header on every page polls GET /system/llm once
a minute and, for any state but ok and checking, shows the server's
message and what it stops: destructive for key and model problems,
amber for unreachable, rate_limited, overloaded and server_error. It
can't be dismissed.

The API client can return a Blob; InvestigationState gains the issue,
run number, brief, profile, error and hypotheses. The unused
StepTimeline, EvidenceCard and EvidenceList are removed; the evidence
field readers moved to components/evidence.tsx.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
The issue page's status pill, menu and filter now read "In progress"
like the other statuses and the mockup.

Also:
- spec §8.2: the hand-off datasource option stays "The issue's
  datasource", because the server resolves it
- close fn-70.22-24 with their commits

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
The starter is the only code that creates an investigation (spec 0001
§7.11). Two paths could still create one, though nothing called them:
- the Redis queue worker (adapters/queue). create_worker was never
  called and nothing enqueued jobs. The package's queue and rate limiter
  had no other users.
- InvestigationService (core/investigation/service.py). The API built it
  into app.state, but nothing read it.

Removed with them, because nothing else uses these:
- the branch/collaboration domain: collaboration, repository, entities,
  values, pattern_extraction, and adapters/db/investigation_repository.py
- the API process's AgentClient and ContextEngine
- get_context_engine_for_tenant, which read the ContextEngine
- their tests

Also fixed:
- the Snowflake docs example, which imported the removed service (and
  was already broken). It now starts a run with the SDK
- CLAUDE.md's description of core/investigation/

ContextEngine (used by the worker), InMemoryPatternRepository, usage
tracking and feedback stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
…Sonnet 5.5

Investigations defaulted to claude-sonnet-4-20250514, which Anthropic
has retired (404), so every run failed. The retired id was hard-coded
in 8 places: config.py, AgentClient, the four legacy LLM adapters, the
quality judge, and docker-compose.yml's LLM_MODEL fallback.

Now dataing.config is the only place that names a model:
- INVESTIGATION_MODEL = "claude-sonnet-5-5", for the manager and its
  subagents
- CHAT_MODEL = "claude-opus-5-5", for issue chat and brief drafts

LLM_MODEL and CHAT_AGENT_MODEL still override them per deployment. An
empty variable, which is what docker compose passes when one is unset,
keeps the default. The compose file no longer hard-codes a model, and
passes CHAT_AGENT_MODEL through as well. Usage pricing is keyed by the
same constants. A test fails if any other source file names a model.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
A login token the API couldn't verify fell through to the API-key check
and failed as "Missing API key". This happens with an expired token, or
one issued by another deployment, such as a demo restarted with a new
JWT secret. The frontend didn't handle the 401, so every page showed
"Failed to load …: Missing API key", which read like a missing
Anthropic key.

Now:
- the API answers a rejected token with "Your session has expired. Sign
  in again."
- a signed-in request that gets a 401 tells the auth provider, which
  signs out (so the route guard shows the login page) and shows one
  toast saying why
- sign-in routes don't count, because a wrong password is a 401 too

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
The new-issue form defaulted to toISOString()'s UTC day, which is the
wrong day near midnight away from UTC. It now defaults to the local date
the date picker shows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
The owner's call: chat moves from claude-opus-5-5 to claude-sonnet-5-5,
the model investigations already use. Checked against the API first,
with both settings the chat agent sends:
- chat turns (effort low, prompt caching) answer
- brief drafts (effort medium, JSON text output) parse

Per-model prices now live in dataing.config next to the model names, so
two roles on one model don't collide in the usage pricing table. Spec
D11 and §12 record the change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
…age's components

The brief editor's scope used a free-text table field and two
datetime-local inputs. Free text let people type names the datasource
can't resolve ("analytics.public.orders" failed a run with "cross-database
references are not implemented"), and the datetime inputs were hard to
use. Both modes now use the removed investigations/new page's components
(spec 0001 §8.2):

- Table(s) is a list of DatasetEntry rows that look tables up in the
  datasource's schema as the person types; "Add another table" adds one.
  A run investigates one datasource, so every row shows the same one. It
  starts as the brief's, else the tenant's default, else its first, which
  replaces the hand-off's "The issue's datasource" option.
- Time window (optional) is the DatePicker, one day or a range; days map
  to whole UTC days.
- New mode has no leads: there is no thread to draft them from.

DatasetEntry's datasource select, table field and remove button get
accessible names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
"Add as check" needed a root cause with 60% confidence, even after a
person confirmed it, so a run that finished at 35% could never become a
check. A person's review now outranks the model's confidence (spec 0001
§7.10):

- a confirmed root cause can become a check at any confidence
- a rejected one never can, so the outcome card hides the button
- an unreviewed one still needs 60%; below that the button is disabled
  and the card says why and what to do: "The root cause's confidence
  (35%) is below 60%. Confirm it to add it as a check."

core/codify.codify_refusal holds the rule, and the codify endpoint looks
up the run's review to apply it; extraction no longer checks confidence
itself. GET /investigations/{id} returns outcome_verdict, so the run's
details page applies the same rule through codifyBlocker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Since every investigation starts in an issue, POST /investigations opens
the issue and records its first thread message, and the end-to-end test's
all-AsyncMock database can't open the thread's transaction. The test now
stands in for InvestigationStarterService, which has its own tests, and
checks what the API does: it parses the SDK alert and answers with the
run, its issue and "queued".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
The issue agent runs every query with the asker's own login (spec 0001
D3), and on credentials_missing it links to
/settings/datasources/{id}/credentials. That page didn't exist, so nobody
could add a login and chat couldn't query anything.

The page (spec 0001 §8.5) is laid out like the sign-in page: a centered
card, "Connect to prod", username and password with icons, plus role and
warehouse where the source type's config takes them (Snowflake).
- Connect tests the login against the database first and saves it only
  if it connects; otherwise the database's reason shows and nothing is
  saved. Then it goes back to where the person came from, usually the
  issue thread, so they can ask again.
- A saved login shows as "You're connected as demo" with the password
  empty (the API never returns it), and Remove my login deletes it.
- A source type with no login says so instead of showing a form.

To make the agent's link work, the agent now writes it as a Markdown link
(for credentials_invalid too), and links to app pages in the thread open
in place instead of a new tab.

It's the first cut of Checks as Code's fn-69.12, whose task now says to
extend this page with key pairs and service-account JSON.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Brings in the orval 8 client regeneration (#232) and just worktrees (#231).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
@vercel

vercel Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
dataing Ready Ready Preview Sep 29, 2026 1:54am UTC
dataing-app Ready Ready Preview Sep 29, 2026 1:54am UTC
dataing-docs Ready Ready Preview Sep 29, 2026 1:54am UTC

@bordumb
bordumb merged commit 0002152 into main Sep 29, 2026
13 checks passed
github-actions Bot pushed a commit that referenced this pull request Sep 29, 2026
# [1.26.0](v1.25.5...v1.26.0) (2026-09-29)

### Features

* **core:** every investigation starts in an issue (spec 0001 M5) ([#233](#233)) ([0002152](0002152)), closes [#211](#211) [#221](#221) [#N](https://github.com/bordumb/dataing/issues/N) [#N](https://github.com/bordumb/dataing/issues/N) [#42](#42)

This branch was successfully deployed

3 active deployments
Preview – dataing-app — bf4d220f Deployed Sep 29, 2026 by vercel[bot]
Preview – dataing — bf4d220f Deployed Sep 29, 2026 by vercel[bot]
Preview – dataing-docs — bf4d220f Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant