Skip to content

Prove contract-driven agent workflow composition in a reference app #288

Description

@enricopiovesan

Problem

The MCP real-agent exercise proves governed discovery and invocation, but not an agent proposing a non-trivial multi-step workflow from contracts with clear review and failure/replan behavior.

Scope

Build a reference-app integration where an external or browser-local proposer creates a candidate workflow. Traverse remains the validator/executor; model hosting and prompt retention are not runtime authority.

Definition of done

  • The product decision records the allowed proposer mode (browser-local, external, or both), data handling, and model/provider boundary.
  • A reference app accepts a bounded goal and has the proposer discover contract metadata through public surfaces only.
  • The proposer produces a multi-step candidate; ambiguous mappings/candidates require explicit human selection or review.
  • The runtime validates artifact identity, schema, policy, placement, risk, and approval requirements before execution.
  • The UI exposes redacted trace evidence for each executed step.
  • A deterministic unavailable/failed-step scenario shows the candidate failure and a new reviewed replan; no hidden auto-retry occurs.
  • E2E evidence proves that the final outputs come from invoked governed capabilities, not a visualization or model-generated stand-in.
  • Docs state required additional prompt context and all known limitations.

Governance

Specs 108, 109, and 113 plus ADR-0043 and ADR-0050 govern the runtime boundary. Product selection of model mode requires owner input.

Activity

  1. babyblueviper1 commented on Sep 6, 2026

    @babyblueviper1

    Real question before this gets built, not an assertion: is the "validator" here meant to be one authority or two?

    The DoD lists two genuinely different kinds of check under the same role. "The runtime validates artifact identity, schema, policy, placement, risk, and approval requirements" is deterministic — same inputs, same answer, every time, and Traverse itself is the right party to own it (it's checking conformance to its own contracts). But "ambiguous mappings/candidates require explicit human selection or review" is a judgment call about whether a specific proposed action is prudent given context a schema can't encode — and that's a different kind of check with a different correctness property: it's not reproducible from the contract alone, and a validator judging its own proposer's output has an obvious incentive problem the deterministic checks don't.

    Concretely: if Traverse (validator/executor) is also implicitly the one resolving "is this ambiguous candidate okay to run," that's the same party checking its own downstream execution risk, not an independent second opinion. Worth naming explicitly in the product decision (proposer mode section) whether that judgment call stays inside Traverse's own runtime authority, or whether there's a slot for a genuinely separate reviewer (human or an independent service) for the ambiguous/high-stakes case specifically — the redacted trace evidence you're already planning to expose would be exactly what a real external reviewer needs to do that check without trusting Traverse's own say-so.

  2. enricopiovesan commented on Sep 6, 2026

    @enricopiovesan
    ContributorAuthor

    Decision inventory

    D1 — proposer deployment and data boundary

    Question: May a workflow proposer run only in the browser/local device, or may it call an external model service?

    • Browser/local only: strongest privacy and offline story; limited model capability and device coverage.
    • External service allowed with explicit consent and redaction: broader model choice; introduces data egress, availability, cost, and retention obligations.
    • Both behind an explicit mode: most flexible; requires two auditable data/consent paths.

    Recommendation: support both modes, defaulting to browser/local. External mode must be opt-in, declare the exact egress destination, redact traces, and remain an untrusted proposer.

    D2 — execution authority after proposal

    • Automatic execution of one unambiguous candidate: smoother demo; weakens the reviewed-authority boundary.
    • Always require explicit review/approval: safest and consistent with Specs 108/109/113; adds a user step.

    Recommendation: always require explicit review initially.

    Governing consequence

    Existing Specs 108, 109, and 113 plus ADR-0043/0050 govern D2. D1 requires a narrow reference-app data-handling policy; add a spec/ADR only if external mode becomes a supported product surface.

  3. babyblueviper1 commented on Sep 6, 2026

    @babyblueviper1

    D2's "always require explicit review initially" is the right default, and it's the same conservative choice as the deterministic-vs-judgment split I raised — good to see it land as a real decision rather than left implicit.

    One follow-up question, genuinely open rather than pushing a particular answer: does "explicit review" here mean specifically a human in the loop, or does the design leave room for the reviewer to be an independent automated check (a separate service/party, not the proposer and not Traverse's own runtime) for the cases where human review doesn't scale? Both satisfy "not self-validated" — a human reviewer and an independent third-party reviewer are both distinct from the proposer — but they have very different cost/throughput/consistency tradeoffs, and Specs 108/109/113 sound like they might already take a position on this. If it's scoped human-only for D2's initial version, that's a completely reasonable place to start; just want to understand whether "review" is a role (satisfiable by any independent party) or specifically a person, since that shapes whether an automated reviewer could ever slot into that same seam later without a spec change.

  4. enricopiovesan commented on Sep 6, 2026

    @enricopiovesan
    ContributorAuthor

    Decision record

    Decision

    Support both proposer modes. Browser/local is the default. External proposer mode is opt-in only and must declare its egress destination, minimize/redact transmitted input, and remain an untrusted proposer with no runtime execution authority.

    Rationale

    This preserves the offline/privacy-first default while allowing stronger external models for users who deliberately choose that trade-off.

    Alternatives considered

    • Browser/local only: rejected because it unnecessarily limits model capability and product reach.
    • External service only: rejected because it weakens privacy, offline use, and user control.

    Consequences

    The reference app needs an explicit per-mode data-handling policy and consent UX. Existing Specs 108, 109, and 113 plus ADR-0043/0050 continue to govern proposal/execution authority. A narrow policy/spec addition is required only if external mode becomes a supported reusable product surface.

  5. enricopiovesan commented on Sep 14, 2026

    @enricopiovesan
    ContributorAuthor

    @babyblueviper1 — catching the follow-up that sat under the D1 decision record.

    Honest status: we have not locked a product decision yet on whether “explicit review” for this reference-app is human-only or a role (human or independent third-party checker). Specs 108/109/113 + ADR-0043/0050 keep proposal untrusted and require an explicit review/approval step before execution; they do not, by themselves, answer “person vs independent automated reviewer” for this prototype.

    So the safe read for now:

    • D2 initial default stays: no self-validated auto-execution.
    • Whether an independent automated reviewer can occupy that same review seam later is still open and needs an explicit decision (likely a narrow policy note if we want non-human reviewers in v1).

    I will not invent that answer here. Leaving this marked for owner decision rather than guessing.

  6. babyblueviper1 commented on Sep 14, 2026

    @babyblueviper1

    Fair not to invent it — that's the right call given specs 108/109/113 don't settle it.

    If it's useful for whoever makes that owner decision, here's the concrete shape the "independent automated reviewer can occupy the seat" question probably needs answered, since "person vs. independent checker" collapses into three separate sub-questions once you try to write the policy note:

    1. Scope: which proposed-action classes is an automated reviewer competent to judge at all? Deterministic policy/schema/placement checks are already handled by the runtime per the DoD (my first comment above) — the remaining "explicit review" seam is for judgment calls a schema can't encode. An automated reviewer only occupies that seam credibly for the subset of judgment calls where the failure mode is checkable from available context (e.g. "does this action match the stated intent / stay inside declared bounds"), not ones needing information outside the system (a human's actual preference in an ambiguous case). Worth drawing that line explicitly rather than treating "automated reviewer" as all-or-nothing.
    2. Independence property: same bar Traverse's own runtime fails for self-validation — the reviewer can't be the proposer, and its verdict needs to be checkable by a third party without trusting the reviewer's say-so (a recomputable signature/hash over the same inputs, not just an API response you take on faith).
    3. Accountability over time: a human reviewer's track record is implicit (their org, their job). An automated reviewer's has to be explicit and public — a running record of past verdicts, wrong ones included — or "independent automated review" just means "a black box you're trusting instead of a person you're trusting," which doesn't actually satisfy D2's own "not self-validated" spirit any better.

    None of that answers the human-vs-role question for you — that's genuinely a product call — but it's the minimum a policy note would need to pin down if the answer turns out to be "role."

  7. enricopiovesan commented on Sep 14, 2026

    @enricopiovesan
    ContributorAuthor

    @babyblueviper1 — thanks, that breakdown is useful.

    Scope / independence / accountability is a cleaner split than “human vs bot” as a binary, and it matches why we did not invent an answer this morning: specs 108/109/113 + ADR-0043/0050 require an explicit review step, but they do not pin whether that seat is human-only or a role with an independent checker.

    Parking this for the owner decision (will not invent a product call mid-thread). Your three sub-questions are exactly the shape a policy note would need if the answer lands on “role.” Leaving them attached here so that decision has concrete criteria instead of vibes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestneeds-enricoWaiting on a product or governance decisionruntime-clientTraverse runtime client boundary work

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions