Repository navigation
Evaluate tool rules on catalogue, search and every call - #353
Conversation
Vendors the tinytools package so the build no longer depends on fetching it at build time. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce a new crate for shared agent definition types so downstream crates can depend on a single source of truth. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Moved the rule matching logic out of the tool module into its own rules submodule so the matching behaviour can be tested and reused on its own. No behaviour change. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Renamed the rule type definitions to better reflect their purpose and improve readability across the harness. No behaviour change. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The agent runtime now handles tool calls emitted by the model, dispatching them through the tool registry and feeding results back into the conversation. This lets agents invoke registered tools during a run instead of only producing text responses. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add tests for the runtime module to verify its wiring and behaviour. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
… sections are all empty. Coul Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Move the tool surface and tool definitions into their own modules so the run loop only handles orchestration. This keeps the loop easier to read and gives the tool wiring a place to grow without bloating the loop. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Moved the nested agent loop logic out of run_loop into a dedicated nested module and extracted tool surface handling into tool_surface. This separates the nested execution path from the top-level loop, making both easier to follow without changing behaviour. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Nested tool calls now consult the rule's approval directive instead of relying solely on the tool policy, so rules that waive approval no longer trigger an approval error and rules that require it always do. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Record tinytools as a dependency in Cargo.lock so the resolved dependency graph stays in sync. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds unit tests for the tool rules used by the agent loop, exercising the allow and deny paths so future changes to rule evaluation are caught by the suite. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds test coverage for the tool rule handling in the agent loop, verifying that rules are applied correctly during tool execution. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat long expressions and import lists across the tool rule modules to satisfy rustfmt, and sort the rules and progress module declarations and re-exports alphabetically. No behaviour changes. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add the Tool trait to the tinytools imports in the rules test module so the tests can reference it. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The permissive definition layer test now checks that listing a catalog surface returns an empty result instead of asserting the internal rules field is unset, so the test exercises observable behaviour rather than implementation detail. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a reference to the new tool-rules document in the harness README so the allow, deny, hide, and approval pattern documentation is discoverable from the module index. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Tiny Sweeper review
Last completed reportTiny Sweeper reviewTiny Sweeper reviewed this change across 6 lane(s) and found 10 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Changes requested Review snapshot
Completeness: Complete What changedThe review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred. Findings
Before merge
How this fits togetherflowchart LR
n0["run_loop_body"]:::impacted
n1["Result"]:::impacted
n2["admit_tool_call"]:::impacted
n3["HarnessRunStatus"]:::impacted
n4["run_id"]:::impacted
n5["host_invocation_binding"]:::impacted
n0 -->|uses| n3
n0 -->|calls| n4
n0 -->|uses| n5
n2 -->|uses| n3
n2 -->|uses| n5
n5 -->|uses| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
Requesting changes: 2 lane(s) blocking, worst finding is critical.
Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.
$0.0363 · 740,939 in / 39,889 out · 75,896 cached (10%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0194 · 381,079 in / 23,142 out · 42,350 cached (11%) · gpt-5.6-luna, glm-5.3-flash
security: $0.0164 · 297,690 in / 13,893 out · 30,474 cached (10%) · gpt-5.6-luna
tests: $0.0002 · 20,433 in / 862 out · 1,536 cached (8%) · glm-5.3-flash
description: $0.0002 · 20,495 in / 267 out · 1,408 cached (7%) · glm-5.3-flash
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4eb24b3b13
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Summary
This makes the agent loop evaluate
tinytoolstool rules (tinyhumansai/tinytools#57) on every surface a tool can reach the model on. One rule set now decides the up-front catalogue,tool_search/ the deferred catalogue, and every call, including nested ones.Before this change, the only gate the loop applied to search was an exact-name
host_allowsclosure. A host that denied a tool throughbefore_toolmiddleware still had that tool listed and searchable.RunPolicy::tool_rules: ToolRulePolicy { rules: Arc<ToolRuleSet>, context: RuleContext }: harness-wide rules, plus the context (channel,agent, …) thatwhenconditions match. The default is empty and admits everything, so behaviour is unchanged.AgentDefinition::tool_rules: Option<ToolRules>: a hosted definition's own rules. They are stacked as a further layer and can only narrow.A crate-private
ToolGatereplaces thehost_allows: Fn(&str) -> boolclosure. It combines:The gate is used by:
direct_tool_schemas,build_tool_surface,declare_toolset_changes(catalogue)deferred_catalog, so also thetool_searchmanifest, its answers and replayed promotions (search)admit_tool_callandadmit_nested_tool(call)Call admission:
before_toolmiddleware, so no approval prompt is raised for a call the rules refuse.Tool::indirect_target.require_approvaldefers exactly like a declaredapproval_required;auto_approvewaives it.Docs:
docs/modules/harness/tool-rules.md.API Or Behavior Changes
RunPolicy::tool_rules,tool::ToolRulePolicy,AgentDefinition::tool_rules/with_tool_rules.tinyagents-definitionnow depends ontinytools.RunPolicy { .. }literal without..RunPolicy::default()must now set the new field. Every in-repo site already uses the default.vendor/tinytoolsis pinned to the Add declarative tool rules (allow/deny/hide/approval) tinytools#57 merge commit (bd60b9b).Tests
cargo fmt --checkcargo clippy --workspace --all-targets -- -D warningscargo clippy --workspace --all-targets --all-features -- -D warningscargo test --workspace: 131 suites, 0 failurescargo test --workspace --all-features: 0 failuresNew tests:
tool/rules/mod_tests.rs: how the gate combines its inputs:agent_loop/tool_rules_tests.rs: end-to-end withScriptedModel:tool_searchmanifest and answer omit denied deferred toolswhencontext appliesrequire_approvaldefers andauto_approvewaivesSummary by CodeRabbit
Since review began
family,tagsandindirect_target: the shared-tool adapter and the rename/prefix override tool. Without that, rule metadata was dropped for every tool a host registers through them.IndirectCall(an indirect target is checked against its own arguments).