Skip to content

feat(config): minimum-confidence policy and config version stamp - #178

Merged
alxxjohn merged 4 commits into
mainfrom
feat/confidence-policy
Sep 4, 2026
Merged

alxxjohn merged 4 commits into
mainfrom
feat/confidence-policy

Conversation

@alxxjohn

@alxxjohn alxxjohn commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Summary

Two config-level features plus the design work queued behind them.

Minimum-confidence policy turns the confidence every finding already carries into something a team can act on. Confidence has existed since findings gained the field — high, medium, low, set by 177 call sites, normalized in NewFinding, emitted as a SARIF property — but nothing consumed it. Structural analyses (source-to-sink taint, the tree-sitter rule paths) set high while regex line scans leave it unspecified, so a team wanting only the trustworthy half had to disable rules outright or baseline the noise. Now checks.min_confidence filters, checks.confidence_demotion softens, and both are fully accounted for.

Config version stamp records which codeguard wrote a config file, so a config checked into a repository carries its own provenance.

Planning docs for the next three accuracy/performance features: tree-sitter scan scale, cross-file taint, and this confidence policy.

Defaults are unchanged in every respect — an existing config produces byte-identical findings, counts, statuses, and exit codes.

Minimum-confidence policy

checks:
  min_confidence:
    default: medium
    sections:
      security: high
      quality: low
  confidence_demotion: true
  • min_confidence.default drops findings below a level; min_confidence.sections.<section> overrides it for one section. Omitting the block, or setting low, admits everything.
  • confidence_demotion: true reports a low-confidence finding on a failing rule as a warning instead. It never promotes and never touches medium or high. Off by default.

Where it runs, and why that matters

Filtering happens in FinalizeSectionWithDiagnostics — after diff scoping, before waiver auditing, suppression matching, and section status. That is after the per-file findings cache, which gives the property that makes thresholds cheap to tune: changing a threshold re-renders a scan rather than re-running one.

Holding that property took one non-obvious fix. SectionConfigHashes builds a fingerprint per config family, and its conservative "" all-checks fallback hashes cfg.Checks wholesale — so a new field under Checks would have invalidated cached findings for every section outside the named families on every threshold change. Both settings are now stripped through findingRelevantChecks before fingerprinting, and a test asserts every family hash is unchanged while a genuinely finding-relevant setting (parsers.treesitter) still moves it.

Nothing disappears quietly

A filtered finding is counted three ways: per rule as confidence_filtered in the rule_stats artifact, per section as confidence_filtered_count, and listed with reason confidence under --include-suppressed. It is deliberately not folded into Suppressed() or suppression_ratio — a suppression records a team accepting a finding, a confidence filter records the scanner not being sure enough to show it, and collapsing the two would corrupt a metric that means "findings teams work around". A test asserts emitted + confidence_filtered + suppressed == total.

Identity is untouched

Demotion rewrites Level/Severity after NewFinding, so exact, context, and content fingerprints are all unchanged and a demoted finding still matches its baseline entry. Tests pin this in both directions.

Section keys

The policy keys on the ids that appear in scan output and in checks.disabled. Those diverge from the runner's internal registry ids — supply chain finalizes as supply_chain but the registry calls it supply-chain — so keying on the registry id would have made that section silently unconfigurable. core.SectionKeys() holds the report-facing list, and TestReportedSectionIDsAreConfigurableKeys asserts every section a real scan reports is a configurable key, so the divergence cannot reappear unnoticed.

Ordering

Findings today have no explicit sort; they arrive in file-scan order. Sorting section.Findings would therefore have silently reordered JSON and SARIF. Instead the confidence sort is a stable sort within each text finding group: group identity and order are unchanged, equal-confidence findings keep scan order, and machine surfaces keep exact scan order. Pinned by a test that asserts JSON ordering is untouched.

Config version stamp

codeguard_version: v1.9.0
name: my-repo
  • Stamped by config.WriteFile, so codeguard init and every SDK write path record the writing release. Verified end to end against the built binary.
  • Stamped at the write boundary, not in ApplyDefaults: a loaded config reports what its file recorded rather than the running binary, which is what makes the field useful for spotting a config produced by a different release. Rewriting refreshes it.
  • Never validated. Absent, older, newer, and unparseable values all load and validate. Decoding was already non-strict, so older binaries ignore the field.
  • Absent from SectionConfigHashes, so a release upgrade cannot discard cached findings (test included). It does reach ConfigHash, which only feeds waiver-audit history dedupe, where a new observation after a config change is correct.
  • The repository config and the shipped JSON example use the v-prefixed form that release builds write via the GoReleaser ldflag.

Verification

  • Defaults unchanged: an omitted policy and an explicit default: low produce reflect.DeepEqual sections through a real scan.
  • Self-scan: make codeguard-ci is 33 findings / 0 fail — exactly the count at main, confirmed by diffing finding locations against a clean worktree at HEAD. (The repo's own quality.ai.narrative-comment rule caught one of my doc comments during this; it is rewritten.)
  • Race: clean across tests/support, tests/core, tests/report, tests/codeguard, internal/codeguard/runner/....
  • gofmt and go vet clean.
  • New external test packages tests/core and tests/report; the latter drives rendering through codeguard.WriteReport.

Pre-existing failure, not from this branch

TestSecuritySemanticAnalyzerScansNodeModulesWithinTarget fails on main at 1205484 with Security status = "pass", want "fail", reproduced in a clean worktree at HEAD. Its requireTypeScriptSemanticRuntime guard does not skip, so this is a real failure rather than a missing-runtime skip. It plausibly relates to the vendored/node_modules scan bounding in 24cb12c but has not been diagnosed. Worth triaging separately — it means go test ./... is red before this branch.

Local lint caveat

golangci-lint run locally reports 2 typecheck issues inside crypto/internal/randutil / math/rand/v2 because its bundled Go is older than the Homebrew toolchain (1.27.1). Both are stdlib paths; no repo code is implicated. CI's pinned version is the real gate.

Docs

docs/checks.md gains a confidence-policy reference, docs/features.md gains YAML and JSON config examples for both features, and README.md mentions the policy. Knowledge captured in .claude/knowledge/ covers three things that would otherwise be re-learned: finding identity excludes presentation, post-cache policy must be stripped from cache fingerprints, and the section-id divergence.

Follow-ups, deliberately not in scope

  • Nothing surfaces a version mismatch. A codeguard doctor note when the stamp differs from the running binary is the natural next step.
  • Confidence is not mapped onto SARIF rank; it stays a result property so external consumers see no schema churn.
  • The 177 emission sites keep the documented empty-means-medium default; no mass audit.

🤖 Generated with Claude Code

alxxjohn and others added 4 commits September 3, 2026 19:15
Three design specs and their task-by-task implementation plans, covering the
accuracy and performance work queued after 1.9.0:

- tree-sitter scan scale: the tree path is capped at 64 files / 256 KiB per
  scan because trees are retained for the whole scan, so on any real
  repository most files silently fall back to the regex path the spike
  measured at 60%/54.5% precision. Replaces retention with a bounded resident
  cache plus heap-aware parse admission, and counts every fallback.
- cross-file taint: Go taint analysis resolves calls only within one file, so
  a flow through a helper in another file is a false negative. Adds a
  package-qualified summary index iterated to a fixed point in reverse
  topological package order, with a package-closure cache fingerprint so
  cross-file findings cannot go stale.
- confidence policy: findings carry a confidence nothing consumes. Makes it a
  configurable minimum threshold with full accounting, optional demotion, and
  confidence-aware ordering.

Each plan is red-first per task and carries its own acceptance criteria.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Findings have always carried a confidence of high, medium, or low, set by 177
call sites, but nothing consumed it: no filtering, no ordering, no effect on
status. Structural analyses (source-to-sink taint, the tree-sitter rule paths)
set high while regex line scans leave it unspecified, so a team could not ask
for only the trustworthy half without disabling rules outright or baselining
the noise.

checks.min_confidence.default drops findings below a level, and
checks.min_confidence.sections.<section> overrides it per section. Omitting
the block, or setting low, admits everything, so the default configuration is
behavior-preserving. checks.confidence_demotion additionally reports a
low-confidence failing finding as a warning; it never promotes.

Filtering happens in FinalizeSectionWithDiagnostics, before waiver auditing,
suppression matching, and status computation, and therefore after the per-file
findings cache: changing a threshold re-renders a scan rather than re-running
one. SectionConfigHashes strips both settings through findingRelevantChecks so
they stay out of every fingerprint family, including the conservative
all-checks fallback.

Removed findings are never silent. Each is counted per rule as
confidence_filtered, per section as confidence_filtered_count, and listed with
reason "confidence" under --include-suppressed. Confidence is deliberately not
folded into Suppressed() or suppression_ratio, which keep meaning "findings
teams work around". Demotion leaves all three fingerprints untouched, so a
demoted finding still matches its baseline entry.

Section keys are the ids that appear in scan output and in checks.disabled.
Those diverge from the runner's registry ids (supply_chain vs supply-chain), so
core.SectionKeys() holds the report-facing list and a test asserts every
section a real scan reports is a configurable key.

Text output sorts by confidence within each rule group with a stable sort;
section finding order is untouched, so JSON and SARIF keep exact scan order.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A config checked into a repository gave no indication of which codeguard
produced it. WriteFile now records the writing release as a top-level
codeguard_version, so codeguard init and every SDK write path stamp it.

Stamping happens at the write boundary rather than in ApplyDefaults, so a
loaded config reports what its file recorded instead of the running binary;
that is what makes the field useful for spotting a config produced by a
different release. Rewriting refreshes it.

The stamp is provenance, not a compatibility gate. It is never validated, so
an absent, older, newer, or unparseable value always loads, and decoding was
already non-strict, so older binaries ignore the field. It is also absent from
SectionConfigHashes, so a release upgrade cannot discard cached findings; a
test pins that. It does reach ConfigHash, which only feeds waiver-audit history
dedupe, where recording a new observation after a config change is correct.

Release builds write v{{.Version}} via the GoReleaser ldflag, so the repository
config and the shipped JSON example use the v-prefixed form.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CI lint caught four staticcheck findings in the new test files: two
WriteString(fmt.Sprintf(...)) calls that should be fmt.Fprintf (QF1012), and
two negated conjunctions that read better under De Morgan (QF1001). Also drops
a hand-rolled integer formatter in favor of strconv.Itoa.

These were missed locally because golangci-lint suppresses staticcheck findings
in repository code when its typecheck pass fails, and it fails here on the
Homebrew Go 1.27.1 stdlib since the binary is built with go1.26.3. The local
run therefore exited 0 while CI reported four issues. Pinning GOTOOLCHAIN to
the linter's build version reproduces CI exactly and now reports 0 issues; the
knowledge note records the corrected command.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@alxxjohn
alxxjohn merged commit d9def6e into main Sep 4, 2026
16 checks passed
@alxxjohn
alxxjohn deleted the feat/confidence-policy branch September 4, 2026 00:54
alxxjohn added a commit that referenced this pull request Sep 4, 2026
🤖 I have created a release *beep* *boop*
---


##
[1.10.0](v1.9.0...v1.10.0)
(2026-09-04)


### Features

* **config:** add a minimum-confidence policy for findings
([53c7e5d](53c7e5d))
* **config:** minimum-confidence policy and config version stamp
([#178](#178))
([d9def6e](d9def6e))
* **config:** stamp the codeguard version into written configs
([7d8ccd4](7d8ccd4))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant