feat(core): every investigation starts in an issue (spec 0001 M5) - #233
Merged
Merged
Conversation
Revise spec 0001 with the owner's decisions from 2026-09-28: - D12: every run belongs to an issue; API, SDK, webhook and rule runs open or reuse one - D13: Investigate… opens the brief editor where you are and lands on the new issue's thread; /investigations/new goes - D14: /investigations/:id becomes the run's details page - D15: LLM failures fail the run with the reason; the API checks the key and models at startup and every page shows a banner Adds §7.11 (one starter, open_issue(), every start path), §7.12 (LLM error table, failing the run, key check), §8.1-8.4 (the mockup as the acceptance reference) and milestone M5, filed as fn-70.18-25. fn-70.17 is closed: PR #211 shipped it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
With a rejected key, every LLM activity caught the error and returned
empty results. The workflow logged warnings, synthesized from nothing
and published "completed" at confidence 0. AgentClient also turned
interpretation errors into evidence that read as refuted.
Now (spec 0001 §7.12):
- classify_llm_error maps the real exception chain (anthropic →
pydantic-ai → LLMError) to a code, retryability and what to fix
- LLM activities raise LLMRejected (non-retryable) or LLMUnavailable
and retry with LLM_RETRY_POLICY (4 attempts, 5-60 s backoff)
- behind the llm-failures-v1 patch, the run fails when:
- hypothesis generation fails or proposes nothing
- a subagent hits an LLM error (the others are cancelled)
- every hypothesis is left untested by errors
- synthesis fails
- a failed counter-analysis keeps the conclusion and records the error
- a failed run publishes {"status": "failed", "error": {code, message,
step}} to the investigation, run row and thread, then fails its
Temporal execution as InvestigationFailed
Recorded histories still replay. The datasource-down test now expects a
failed run instead of a conclusion drawn from no evidence.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Of the 8 ways a run was created, only starting it from an issue page
linked it fully. POST /investigations (UI, SDK, CLI, notebook) and the
AUTO webhooks inserted runs inline, with no run row or start card, and
the EE rule action never wrote its outcome back.
Now (spec 0001 §7.11):
- open_issue() (adapters/db/issues.py) is the only issue insert. The
API, both webhooks and the starter use it, and every thread starts
with the event that opened it
- InvestigationStarterService.start():
- resolves the given issue or opens one
- builds the missing brief or alert
- writes the investigation, the run row (trigger_type human/api/
webhook/rule) and the thread card
- starts the workflow with alert.issue_id
- on a failed start, records the failure on the card
- POST /investigations takes a brief or an alert (a bad alert is a 422,
not a 500) and returns run_id, issue_id and issue_number
- the issue spawn route, both webhooks and the EE rule action go through
the starter
- runs report their number, status and error; GET /investigations/{id}
adds the issue, run number, brief and failure for the details page
- tool calls record duration_ms and row_count for "Ran 1 query · 212 ms"
- the SDK's Investigation gains issue_id and issue_number;
`dataing run start` links the issue
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Nothing checked the key before, so a rejected key first showed up as a
run that did nothing, and a chat reply or brief draft that failed showed
the raw ModelHTTPError text.
Now (spec 0001 §7.12):
- at startup the API calls GET /v1/models/{id} for LLM_MODEL and
CHAT_AGENT_MODEL (no tokens). It runs in the background with a 10 s
timeout and no retries, so startup never waits on it
- GET /api/v1/system/llm returns {state, message, models, checked_at}
for the app's banner. A transient problem is checked again once it is
over a minute old
- chat turns and brief drafts raise the classified LLMRejected or
LLMUnavailable error. The reply shows what to fix, and a rejected key
isn't retried
Checked against the real API with a bogus key: invalid_key, "Anthropic
rejected the API key (401). Set a valid ANTHROPIC_API_KEY and restart
the API and the worker."
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Copies only the changed operations from the app's schema:
- POST /investigations (brief or alert, issue link in the response)
- GET /investigations/{id} (issue, run number, brief, error, hypotheses)
- the issue's investigation runs and outcome review (number, status,
error)
- the new GET /system/llm
Also adds the schemas these reference. The rest of the committed spec
is untouched.
The TypeScript client is not regenerated: orval 8 (#221) would rewrite
all 293 generated files (+31k lines) and turn GET hooks into mutations.
The frontend reaches these routes through its hand-written wrappers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Every start button now opens the brief editor in new mode where the
person is, instead of the /investigations/new page (spec 0001 D13, §8.2):
the sidebar quick action, the dashboard header and its empty
recent-investigations card, the investigations list header and empty
state, and "Investigate this dataset" in the dataset page header, which
is shown whether or not the dataset has runs and pre-fills the table's
native path and datasource.
New mode reuses BriefFormView, titled "Start an investigation":
- symptom and at least one scope table are required
- the datasource select lists the tenant's datasources, preselects the
only one and requires a choice when there are several; it has no
"The issue's datasource" option
- findings, ruled out and leads start empty
- Start posts the brief to POST /api/v1/investigations and lands on
/issues/{issue_id}; a 409 ambiguous_datasource asks the person to pick
a datasource, and other errors toast the server's message
The API client now throws ApiError, which keeps the status and the
detail's error code. BriefFormView takes a mode and an onSubmit; the
hand-off submit moved into HandoffForm unchanged.
Removed: the investigations/new route, NewInvestigation.tsx, the
alert-based useStartInvestigation and the dead components/Layout.tsx.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
0001_issue_chat_mockup.html is the acceptance reference (spec 0001 §8.1),
and the page now follows it:
- A one-row top bar: the #N pill, the title, the status and priority
pills and "dataset x" on the right, with the back link.
- The thread (1fr) beside one sidebar panel (300px), stacked on narrow
screens.
- The thread head has the Shared thread / My scratch chats (N) tabs (the
second opens the scratch drawer) and "N watching · live".
- No description card: the thread's first entry is the `created` event,
"Maya opened the issue" with the issue's description and Edit for its
author (PATCHes the description), or "⚑ Issue opened by dataing from x"
when nobody opened it by hand. Threads without that event still show
the description, from the issue.
- One sidebar panel: Status (pill, "Change ▾", the menu's note "Only
moves that will succeed are shown" and the Resolved hint when a cause
was confirmed), Details (compact editors; labels as pills with +),
Dataset with "open dataset page →" when a datasource has synced the
table, Investigations ("#N · depth" with running / done · 0.91 /
confirmed / rejected / failed), Watchers, Your scratch chats and
+ New scratch chat. The Timeline and the separate cards are gone.
- The investigation card is "Investigation #N" (the run's number from
the runs list), with the mockup's one-line brief, "details →" in place
of "Open", and + Add context / + Add hypothesis. A ruled-out
hypothesis reads "ruled out by Maya · stopped".
- Failed runs: the failed outcome renders the reason, the step and
Retry (members), which reopens the brief editor with the run's brief.
The start card turns red with the reason, and offers Retry itself only
when no failed outcome was posted. The sidebar row says failed.
- A run nobody started by hand reads "dataing · started an
investigation".
- Tool calls collapse to "▸ Ran 1 query · 212 ms · 4 rows" from the
record's duration_ms and row_count (older messages fall back to "Ran 1
query"); the expanded footer adds "Snapshot saved with this message ·
Copy SQL".
- The outcome card notes when the check of its conclusion didn't run,
and proposals name their run "#N".
The runs list is hand-typed with number, status and error, refreshes
while a run is running and when an outcome lands in the thread.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
…70.24) /investigations/:id becomes the run's details page (spec 0001 D14, §8.3): - Header: "← #42 {issue title}" back to the issue's thread (omitted for imported snapshots, which have no issue), "Investigation #N" with its status and depth pills, Share (copies the link; the mocked user picker is gone), Export snapshot (GET /investigations/{id}/snapshot through the API client, saved as snapshot-{id}.tar.gz), Add as check (renamed from Codify Test, same gating) and Cancel run while it runs. - Brief: the brief the run was given. - Hypotheses: each with its status pill ("ruled out by a person", "untested" in amber) and its evidence, grouped by hypothesis_id. Queries collapse like the thread's tool calls and expand to the SQL, the result summary and the interpretation, with feedback. - Outcome: the root cause in the thread card's style, with the hypothesis pills, causal chain, onset, affected scope, recommendations, supporting evidence and the existing feedback buttons. - A failed run shows the reason, the step and what to do next, with a link back to the issue to retry, in place of the outcome. It never shows "Unable to determine a definitive root cause". - The step-timeline card is gone: the header's "running · phase" pill and live dot are the compact phase indicator. A banner under the app header on every page polls GET /system/llm once a minute and, for any state but ok and checking, shows the server's message and what it stops: destructive for key and model problems, amber for unreachable, rate_limited, overloaded and server_error. It can't be dismissed. The API client can return a Blob; InvestigationState gains the issue, run number, brief, profile, error and hypotheses. The unused StepTimeline, EvidenceCard and EvidenceList are removed; the evidence field readers moved to components/evidence.tsx. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
The issue page's status pill, menu and filter now read "In progress" like the other statuses and the mockup. Also: - spec §8.2: the hand-off datasource option stays "The issue's datasource", because the server resolves it - close fn-70.22-24 with their commits Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
The starter is the only code that creates an investigation (spec 0001 §7.11). Two paths could still create one, though nothing called them: - the Redis queue worker (adapters/queue). create_worker was never called and nothing enqueued jobs. The package's queue and rate limiter had no other users. - InvestigationService (core/investigation/service.py). The API built it into app.state, but nothing read it. Removed with them, because nothing else uses these: - the branch/collaboration domain: collaboration, repository, entities, values, pattern_extraction, and adapters/db/investigation_repository.py - the API process's AgentClient and ContextEngine - get_context_engine_for_tenant, which read the ContextEngine - their tests Also fixed: - the Snowflake docs example, which imported the removed service (and was already broken). It now starts a run with the SDK - CLAUDE.md's description of core/investigation/ ContextEngine (used by the worker), InMemoryPatternRepository, usage tracking and feedback stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
…Sonnet 5.5 Investigations defaulted to claude-sonnet-4-20250514, which Anthropic has retired (404), so every run failed. The retired id was hard-coded in 8 places: config.py, AgentClient, the four legacy LLM adapters, the quality judge, and docker-compose.yml's LLM_MODEL fallback. Now dataing.config is the only place that names a model: - INVESTIGATION_MODEL = "claude-sonnet-5-5", for the manager and its subagents - CHAT_MODEL = "claude-opus-5-5", for issue chat and brief drafts LLM_MODEL and CHAT_AGENT_MODEL still override them per deployment. An empty variable, which is what docker compose passes when one is unset, keeps the default. The compose file no longer hard-codes a model, and passes CHAT_AGENT_MODEL through as well. Usage pricing is keyed by the same constants. A test fails if any other source file names a model. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
A login token the API couldn't verify fell through to the API-key check and failed as "Missing API key". This happens with an expired token, or one issued by another deployment, such as a demo restarted with a new JWT secret. The frontend didn't handle the 401, so every page showed "Failed to load …: Missing API key", which read like a missing Anthropic key. Now: - the API answers a rejected token with "Your session has expired. Sign in again." - a signed-in request that gets a 401 tells the auth provider, which signs out (so the route guard shows the login page) and shows one toast saying why - sign-in routes don't count, because a wrong password is a 401 too Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
The new-issue form defaulted to toISOString()'s UTC day, which is the wrong day near midnight away from UTC. It now defaults to the local date the date picker shows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
The owner's call: chat moves from claude-opus-5-5 to claude-sonnet-5-5, the model investigations already use. Checked against the API first, with both settings the chat agent sends: - chat turns (effort low, prompt caching) answer - brief drafts (effort medium, JSON text output) parse Per-model prices now live in dataing.config next to the model names, so two roles on one model don't collide in the usage pricing table. Spec D11 and §12 record the change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
…age's components
The brief editor's scope used a free-text table field and two
datetime-local inputs. Free text let people type names the datasource
can't resolve ("analytics.public.orders" failed a run with "cross-database
references are not implemented"), and the datetime inputs were hard to
use. Both modes now use the removed investigations/new page's components
(spec 0001 §8.2):
- Table(s) is a list of DatasetEntry rows that look tables up in the
datasource's schema as the person types; "Add another table" adds one.
A run investigates one datasource, so every row shows the same one. It
starts as the brief's, else the tenant's default, else its first, which
replaces the hand-off's "The issue's datasource" option.
- Time window (optional) is the DatePicker, one day or a range; days map
to whole UTC days.
- New mode has no leads: there is no thread to draft them from.
DatasetEntry's datasource select, table field and remove button get
accessible names.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
"Add as check" needed a root cause with 60% confidence, even after a
person confirmed it, so a run that finished at 35% could never become a
check. A person's review now outranks the model's confidence (spec 0001
§7.10):
- a confirmed root cause can become a check at any confidence
- a rejected one never can, so the outcome card hides the button
- an unreviewed one still needs 60%; below that the button is disabled
and the card says why and what to do: "The root cause's confidence
(35%) is below 60%. Confirm it to add it as a check."
core/codify.codify_refusal holds the rule, and the codify endpoint looks
up the run's review to apply it; extraction no longer checks confidence
itself. GET /investigations/{id} returns outcome_verdict, so the run's
details page applies the same rule through codifyBlocker.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
Since every investigation starts in an issue, POST /investigations opens the issue and records its first thread message, and the end-to-end test's all-AsyncMock database can't open the thread's transaction. The test now stands in for InvestigationStarterService, which has its own tests, and checks what the API does: it parses the SDK alert and answers with the run, its issue and "queued". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Claude <noreply@anthropic.com>
The issue agent runs every query with the asker's own login (spec 0001
D3), and on credentials_missing it links to
/settings/datasources/{id}/credentials. That page didn't exist, so nobody
could add a login and chat couldn't query anything.
The page (spec 0001 §8.5) is laid out like the sign-in page: a centered
card, "Connect to prod", username and password with icons, plus role and
warehouse where the source type's config takes them (Snowflake).
- Connect tests the login against the database first and saves it only
if it connects; otherwise the database's reason shows and nothing is
saved. Then it goes back to where the person came from, usually the
issue thread, so they can ask again.
- A saved login shows as "You're connected as demo" with the password
empty (the API never returns it), and Remove my login deletes it.
- A source type with no login says so instead of showing a form.
To make the agent's link work, the agent now writes it as a Markdown link
(for credentials_invalid too), and links to app pages in the thread open
in place instead of a new tab.
It's the first cut of Checks as Code's fn-69.12, whose task now says to
extend this page with key pairs and service-account JSON.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Claude <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
github-actions Bot
pushed a commit
that referenced
this pull request
Sep 29, 2026
# [1.26.0](v1.25.5...v1.26.0) (2026-09-29) ### Features * **core:** every investigation starts in an issue (spec 0001 M5) ([#233](#233)) ([0002152](0002152)), closes [#211](#211) [#221](#221) [#N](https://github.com/bordumb/dataing/issues/N) [#N](https://github.com/bordumb/dataing/issues/N) [#42](#42)
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Spec 0001's M5: the fixes for what testing the issue hub turned up (spec §2 "Found after M1–M4", decisions D12–D15), plus what the demo walkthrough found after that.
Investigations
InvestigationStarterService.start()is the only way a run starts: the UI, the API and SDK, webhooks and rules all go through it. A run started without an issue opens one titled with its symptom. The dormant creation paths (queue adapter, old investigation service and repository) are removed./investigations/newis gone. The editor picks tables with the old page's autocomplete (DatasetEntry) and days with itsDatePicker. Leads are hand-off only./investigations/:idlinks back to its issue and shows the brief, the hypotheses and how each ended, the evidence, and the outcome.workflow.patched("llm-failures-v1")with replay tests. The API checks the Anthropic key at startup, and a banner says what to fix.Issue chat
/settings/datasources/{id}/credentials, laid out like the sign-in page (spec §8.5). The agent linked there oncredentials_missing, but the page didn't exist, so chat couldn't query anything. It tests the login before saving it and never shows the password back. The agent now writes that link as Markdown, and app links in the thread open in place.Models
dataing.configis the only place that names a Claude model, and a test enforces that. Investigations, chat and brief drafting run onclaude-sonnet-5-5. The old default,claude-sonnet-4-20250514, was retired and made every investigation fail with a 404.Other fixes
Test plan
pytest(3,129 passed), EEpytest(715 passed), mypy, andruff checktsc, eslint, prettierllm-failures-v1patch🤖 Generated with Claude Code