Skip to content

[Portability benchmark] Run Doubt on Claude Code #11

Description

@alsoleg89

Part of #9. This issue covers one client only: Claude Code.

Contribution

Run the frozen Agent Skills portability protocol with the unmodified Doubt v0.5.0 skill and synthetic fixture.

Client ID for the result file: claude-code

Run all three prompts in fresh sessions:

  • direct.txt
  • implicit.txt
  • negative.txt

Deliverables

  • Exact Claude Code version
  • Exact model, operating system, permission mode, install path, and relevant configuration
  • Sanitized raw output for all three prompts
  • Generated artifact and receipt when one exists
  • benchmarks/skill-portability/results/claude-code-<version>.json
  • npm run benchmark:portability passes
  • Skill, fixture, and prompts are unchanged

Start from result.example.json. The validator prints the canonical skill digest.

A fail or blocked run is a valid contribution when the limitation is explicit. Do not turn missing behavior into a success and do not claim cross-client equivalence.

Comment before starting so two people do not duplicate the same client/version. Do not include credentials, private repository data, personal paths, or user names in transcripts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    good first issueGood for newcomershelp wantedExtra attention is neededneeds-evidenceClaim or fixture needs stronger source support

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions