Skip to content

research(mir): preregister structure feature noninferiority evidence - #1228

Draft
seonghobae wants to merge 217 commits into
bolt-performance-chart-export-13223013812255847379from
research/structure-noninferiority-1225
Draft

seonghobae wants to merge 217 commits into
bolt-performance-chart-export-13223013812255847379from
research/structure-noninferiority-1225

Conversation

@seonghobae

@seonghobae seonghobae commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator

Signal-MIR scientific evidence owner

#1223 withdrew the chroma_cqtchroma_stft production switch because synthetic/mock timing did not establish structure accuracy on rights-cleared decoded audio. This Draft owns preregistered real-audio structure noninferiority evidence before any production default change. Resource Admission remains #866; Project Persistence remains #970; distribution/rights stay with their canonical owners.

Exact identity / stack

  • Protected product truth: develop@314ddeae7b775a4957594b599358c8255617eb2e.
  • Canonical repository Ruff prerequisite: repair(ci): format consolidated supply-chain policy test #1176 exact 8fe6b6d99c009527ef0bcba419e6f6debdb23c23, Open / Draft / mergeable.
  • PR base: bolt-performance-chart-export-13223013812255847379 (repair(ci): format consolidated supply-chain policy test #1176).
  • Branch: research/structure-noninferiority-1225.
  • Exact head: 7eb2b9d1f08254d7ce7300a6f9a5b72ea06b3977.
  • Open / Draft; merge acceptance remains blocked by fresh exact-head gates, central CodeQL lifecycle and scientific acceptance.
  • Production structure extraction still defaults to chroma_feature="cqt"; this PR does not switch production to STFT.

The stack adopts #1176 as real ancestry/content rather than copying a second formatter implementation. Commit 715dc1a6250095352cb5a9b250ddd4271a4656af took test_supply_chain_policy.py from canonical #1176 while preserving the scientific lane. Current 7eb2b9d... is an ordinary one-file descendant.

Frozen scientific semantics retained

Functional-label ACC remains pinned to ismir-mirex/mirex-evaluation@b9fa0b0b32e2145af31f35830f78fc9d09a4301b:music_structure_analysis.eval_script.calculate_accuracy with the 200 ms frame-grid contract. Segmentation metrics remain bound to content-addressed mir_eval==0.8.2, including detection at 0.5/3.0 s (beta=1.0, trim=True), deviation (trim=True) and pairwise grouping (frame_size=0.1, beta=1.0).

macro-track-v1 / paired-track-bootstrap-v1 retain equal-track aggregation, harmonic F reconstruction, macro deviation/ACC/latency means, maximum peak RSS, paired resampling with replacement, numpy.random.Generator(PCG64), and NumPy linear 0.025/0.975 percentile bounds. isolated-single-shot-v1 retains macOS/Windows per-track performance evidence with zero same-process warm-ups, 20 fresh-process observations per lane/track, alternating CQT→STFT / STFT→CQT order, time.perf_counter_ns(), linear p50/p95, and process-lifetime peak RSS.

Synthetic fixtures remain regression evidence only. Rights-cleared real decoded audio is required for production scientific acceptance.

Hosted Ruff RED → minimal causal repair

The merge-tested exact predecessor 715dc1a6250095352cb5a9b250ddd4271a4656af now has terminal repository evidence:

  • build-baseline 35787785367: SUCCESS
  • Security Scan 35787785260: SUCCESS
  • SAST Semgrep 35787785107: SUCCESS
  • sbom 35787785063: SUCCESS
  • repository ci 35787785166: FAILURE
  • CodeQL PR 35787785203: FAILURE in the separately tracked central publication/reconciliation lifecycle

The repository CI failure is local and exact. npm-lock-validation and macOS rust-check passed. Ubuntu ci / build-and-test reached ruff:check, where Ruff 0.15.5 reported one I001 in services/analysis-engine/tests/test_structure_functional_label_contract_policy.py. The previous attempted repair had inserted a blank line between import pytest and import test_structure_noninferiority_policy as noninferiority_policy; hosted Ruff still classified the block as unsorted/unformatted. The neighboring canonical policy test already groups pytest and sibling test-module imports in the same import section.

Minimal causal repair 7eb2b9d1f08254d7ce7300a6f9a5b72ea06b3977 removes exactly that one blank line. Fresh compare from 715dc1a... is one commit, one file, +0/-1. No assertion, metric, threshold, corpus, runtime, dependency, production structure feature, ignore/noqa or scientific claim changes.

Predecessor terminal evidence is RCA/semantic evidence only and does not transfer to 7eb2b9d.... This new exact head must reacquire repository/security/SAST/SBOM/CodeQL and review evidence; absent/queued is not GREEN.

Review state / scientific claim boundary

No qualifying independent non-author current-head APPROVED is claimed. This branch still does not establish approved numeric production margins/latency thresholds, adequacy or representativeness of a concrete rights-cleared corpus, adequacy of bootstrap count/seed, a reviewed concrete host-profile payload, scientific noninferiority/latency superiority on real audio, MIREX leaderboard equivalence, or a production feature switch.

Before candidate results are inspected, reviewers must freeze concrete corpus membership/representativeness, quality noninferiority margins, maximum candidate latency ratio, exact bootstrap count/seed, actual macOS/Windows structure-performance-host-v1 payload/content address, and the population/runtime claim. Only then may the complete rights-cleared corpus run through the canonical experiment runner.

Normal integration order is #1176 authentic central CodeQL/review settlement → protected integration → ordinary/non-force reconciliation of this lane to protected develop → fresh exact-head/base scientific/repository/security evidence → qualifying independent review. No self-approval, force-push, destructive rebase, gate weakening, synthetic status, blind rerun, no-op freshness commit or predecessor-evidence transfer.

UI Delivery Gate: FAIL — this scientific lane does not establish buyer-visible Active Player/UI interaction, accessibility, responsive or 8-locale acceptance.

Commercial Release Gate: FAIL — the exact hosted Ruff finding now has a one-line causal repair, but 7eb2b9d... has not yet reacquired fresh terminal gates, #1176 is not protected truth, independent approval is absent, and rights-cleared real-audio corpus execution plus frozen scientific acceptance thresholds remain missing.

@coderabbitai

coderabbitai Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Copy link
Copy Markdown
Collaborator Author

@opencode-agent review exact head 47d09320957097b5347f81396a57e422dad5813a only. Review the Signal-MIR scientific evidence-admission boundary, especially preregistration/result identity, failed-track fail-closed behavior, aggregate/per-track consistency assumptions, uncertainty claim boundaries, and evidence parsing/security. Review only; do not mutate source.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 47d09320957097b5347f81396a57e422dad5813a. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

Copy link
Copy Markdown
Collaborator Author

@opencode-agent review exact head d7c65fdfbc40a68b066eb092e76fe5e87a5ad0cd only. Review the Signal-MIR scientific evidence-admission boundary after the closed-world measurement-receipt repair, especially post-hoc field admission, aggregate/per-track consistency assumptions, preregistered aggregation/uncertainty claim boundaries, failed-track fail-closed behavior, and evidence parsing/security. Review only; do not mutate source.

@seonghobae seonghobae added the priority: medium Normal-priority or P2 work label Sep 19, 2026 — with ChatGPT Codex Connector
@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 275cadacf672391951e410c1e3c9292e75bc4c99. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 275cadacf672391951e410c1e3c9292e75bc4c99. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 275cadacf672391951e410c1e3c9292e75bc4c99. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

@seonghobae
seonghobae changed the base branch from develop to bolt-performance-chart-export-13223013812255847379 September 22, 2026 21:37

Copy link
Copy Markdown
Collaborator Author

Fresh exact-head verification receipt for 715dc1a6250095352cb5a9b250ddd4271a4656af: build-baseline 35787785367 has completed all four platform jobs successfully (macOS amd64 106949098423, macOS arm64 106949098793, Windows arm64 106949098886, Windows amd64 106949098928). The required final Ubuntu aggregate jobs remain queued without execution: gate / build / macos 106951833521 and gate / build / windows 106952076931, both steps=[], runner_id=0. Repository CI remains blocked before execution at gate / ci / npm-lock-validation 106948699623, also steps=[], runner_id=0.

This is not a scientific GREEN verdict and does not satisfy the real-audio acceptance contract. It only establishes that current-head platform builds completed while later aggregate/CI admission did not receive a runner. The cross-repository queue/admission receipt was routed to canonical owner ContextualWisdomLab/.github#712 (comment 5786151126). No no-op commit, cancellation of sole current-head evidence, or gate bypass was used. Keep this PR Draft pending #1176 settlement, current-head terminal gates/review, and the pre-registered rights-cleared real decoded audio noninferiority run.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 715dc1a6250095352cb5a9b250ddd4271a4656af. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 715dc1a6250095352cb5a9b250ddd4271a4656af. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 715dc1a6250095352cb5a9b250ddd4271a4656af. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 7eb2b9d1f08254d7ce7300a6f9a5b72ea06b3977. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 7eb2b9d1f08254d7ce7300a6f9a5b72ea06b3977. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

@opencode-agent

Copy link
Copy Markdown
Contributor

Queued @opencode-agent for PR #1228 at head 7eb2b9d1f08254d7ce7300a6f9a5b72ea06b3977. Central exact-name Actions artifacts are the durable dispatch ledger; existing review workflows remain authoritative for the final verdict and failure evidence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request priority: medium Normal-priority or P2 work

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant