Skip to content

Evaluate tool rules on catalogue, search and every call - #353

Merged
senamakel merged 17 commits into
mainfrom
tool-rules
Oct 9, 2026
Merged

senamakel merged 17 commits into
mainfrom
tool-rules

Conversation

@senamakel

@senamakel senamakel commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

Summary

This makes the agent loop evaluate tinytools tool rules (tinyhumansai/tinytools#57) on every surface a tool can reach the model on. One rule set now decides the up-front catalogue, tool_search / the deferred catalogue, and every call, including nested ones.

Before this change, the only gate the loop applied to search was an exact-name host_allows closure. A host that denied a tool through before_tool middleware still had that tool listed and searchable.

  • RunPolicy::tool_rules: ToolRulePolicy { rules: Arc<ToolRuleSet>, context: RuleContext }: harness-wide rules, plus the context (channel, agent, …) that when conditions match. The default is empty and admits everything, so behaviour is unchanged.

  • AgentDefinition::tool_rules: Option<ToolRules>: a hosted definition's own rules. They are stacked as a further layer and can only narrow.

  • A crate-private ToolGate replaces the host_allows: Fn(&str) -> bool closure. It combines:

    • the existing exact allowlist (I-9 fail-closed semantics unchanged)
    • the policy's rules
    • the definition's rules

    The gate is used by:

    • direct_tool_schemas, build_tool_surface, declare_toolset_changes (catalogue)
    • deferred_catalog, so also the tool_search manifest, its answers and replayed promotions (search)
    • admit_tool_call and admit_nested_tool (call)
    • the unknown-tool "available" list and rewrite target
  • Call admission:

    • Rules run before before_tool middleware, so no approval prompt is raised for a call the rules refuse.
    • They are evaluated on the raw provider arguments and include Tool::indirect_target.
    • A refusal answers the model with the rule id, layer and reason, and releases the budget slot.
    • require_approval defers exactly like a declared approval_required; auto_approve waives it.

Docs: docs/modules/harness/tool-rules.md.

API Or Behavior Changes

  • Additive, public: RunPolicy::tool_rules, tool::ToolRulePolicy, AgentDefinition::tool_rules / with_tool_rules.
  • tinyagents-definition now depends on tinytools.
  • Breaking for struct literals: a RunPolicy { .. } literal without ..RunPolicy::default() must now set the new field. Every in-repo site already uses the default.
  • With no rules configured, the loop behaves exactly as before.
  • vendor/tinytools is pinned to the Add declarative tool rules (allow/deny/hide/approval) tinytools#57 merge commit (bd60b9b).

Tests

  • cargo fmt --check
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo clippy --workspace --all-targets --all-features -- -D warnings
  • cargo test --workspace: 131 suites, 0 failures
  • cargo test --workspace --all-features: 0 failures

New tests:

  • tool/rules/mod_tests.rs: how the gate combines its inputs:
    • allowlist first
    • policy and definition layers
    • a permissive layer is dropped
    • context
    • the surface follows the tool's exposure
    • call approval and refusal
  • agent_loop/tool_rules_tests.rs: end-to-end with ScriptedModel:
    • denied and hidden tools are absent from the first request
    • the tool_search manifest and answer omit denied deferred tools
    • a denied call is refused with its rule, while a hidden tool still runs
    • when context applies
    • indirect targets are checked
    • require_approval defers and auto_approve waives
    • a hosted definition's rules stack with the policy and the allowlist

Summary by CodeRabbit

  • New Features
    • Added configurable tool rules that control which tools appear in catalogues and search results, and which tools can be called.
    • Rules can account for call context and indirect targets. When a call is denied, the harness returns a refusal.
    • Rule settings can require or waive approval, or use each tool’s existing approval policy by default. With no rules configured, behavior remains permissive.
  • Documentation
    • Added guidance and configuration examples for tool rules, including how rules from different sources combine.

Since review began

  • Wrappers forward family, tags and indirect_target: the shared-tool adapter and the rename/prefix override tool. Without that, rule metadata was dropped for every tool a host registers through them.
  • Adopts tinytools' IndirectCall (an indirect target is checked against its own arguments).

senamakel and others added 17 commits October 9, 2026 09:35
Vendors the tinytools package so the build no longer depends on
fetching it at build time.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce a new crate for shared agent definition types so downstream
crates can depend on a single source of truth.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Moved the rule matching logic out of the tool module into its own rules
submodule so the matching behaviour can be tested and reused on its own.
No behaviour change.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Renamed the rule type definitions to better reflect their purpose and
improve readability across the harness. No behaviour change.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The agent runtime now handles tool calls emitted by the model, dispatching them through the tool registry and feeding results back into the conversation. This lets agents invoke registered tools during a run instead of only producing text responses.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add tests for the runtime module to verify its wiring and behaviour.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
… sections are all empty. Coul

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Move the tool surface and tool definitions into their own modules so the
run loop only handles orchestration. This keeps the loop easier to read
and gives the tool wiring a place to grow without bloating the loop.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Moved the nested agent loop logic out of run_loop into a dedicated nested module and extracted tool surface handling into tool_surface. This separates the nested execution path from the top-level loop, making both easier to follow without changing behaviour.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Nested tool calls now consult the rule's approval directive instead of
relying solely on the tool policy, so rules that waive approval no longer
trigger an approval error and rules that require it always do.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Record tinytools as a dependency in Cargo.lock so the resolved dependency graph stays in sync.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds unit tests for the tool rules used by the agent loop, exercising
the allow and deny paths so future changes to rule evaluation are
caught by the suite.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds test coverage for the tool rule handling in the agent loop, verifying that rules are applied correctly during tool execution.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat long expressions and import lists across the tool rule modules to satisfy rustfmt, and sort the rules and progress module declarations and re-exports alphabetically. No behaviour changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add the Tool trait to the tinytools imports in the rules test module so the
tests can reference it.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The permissive definition layer test now checks that listing a catalog
surface returns an empty result instead of asserting the internal rules
field is unset, so the test exercises observable behaviour rather than
implementation detail.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a reference to the new tool-rules document in the harness README so the allow, deny, hide, and approval pattern documentation is discoverable from the module index.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for 4eb24b3b1361. forge: GitHub: API rate limit exceeded for installation ID 152184043. If you reach out to GitHub Support for help, please include the request ID CC8E:1AA72:1B1EDDF:5969FE8:6AC8C017 and timestamp 2026-10-09 10:21:12 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs\.github\.com/en/site\-policy/github\-terms/github\-terms\-of\-service\)
Documentation URL: https://docs\.github\.com/en/rest/using\-the\-rest\-api/getting\-started\-with\-the\-rest\-api\#rate\-limiting

Last completed report

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 10 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Changes requested
Priority: critical
Reviewed head: 4eb24b3b1361
Updated: 1791528717 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 11 Active findings 10
Tests 3 Noted findings 0
Documentation 2 Resolved findings 0
Configuration 1 Pending checks/questions 0

Completeness: Complete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

  • critical · critique · Add the rules module before declaring it — The reviewed commit contains no `crates/tinyagents-harness/src/tool/rules.rs`, so this module declaration produces a compile error. The added `ToolRulePolicy`, `CallGate`, and `Too (crates/tinyagents\-harness/src/tool/mod\.rs:14)
  • high · critique · Preserve approval directives for rewritten calls — When `UnknownToolPolicy::Rewrite` selects a target, this filter checks whether `admit_call` returns `Admit` but discards the contained `ApprovalDirective`. `rule_approval` remains (crates/tinyagents\-harness/src/agent\_loop/tools\.rs:749)
  • high · critique · Preserve compatibility for RunPolicy struct literals — `RunPolicy` is public, and adding this non-optional field means existing callers using `RunPolicy { ... }` without `..RunPolicy::default()` no longer compile. This is an API-breaki (crates/tinyagents\-harness/src/runtime/types\.rs:492)
  • medium · critique · Evaluate nested-call rules after preparing arguments — `admit_call` receives `call.arguments` before `prepare_tool_arguments` injects authoritative host values and before normalization. Since the gate evaluates rules using the supplied (crates/tinyagents\-harness/src/agent\_loop/nested\.rs:870)
  • medium · critique · Reserve registered names before filtering schemas — `existing` is built from the already-filtered registry schemas. Because `gate.lists` evaluates rules against the concrete tool, a registered tool can be denied while a toolset tool (crates/tinyagents\-harness/src/agent\_loop/tool\_surface\.rs:36)
  • medium · critique · Gate the intrinsic discovery call before answering it — This applies the gate only while building the deferred catalog; it never checks whether the `tool_search` call itself is allowed on the bridge path. For example, with discovery ena (crates/tinyagents\-harness/src/agent\_loop/tools\.rs:360)
  • high · security · Add the declared test module or remove its registration — The declared `tool_rules_tests.rs` file is not present in the repository, so compiling the test target will fail with a missing module file error. Add the sibling test file contain (crates/tinyagents\-harness/src/agent\_loop/mod\.rs:182)
  • high · security · Preserve approval directives for rewritten tool calls — This checks whether the rewrite target is callable, but discards the `ApprovalDirective` returned by `gate.admit_call` by matching it as `_`. For an unknown model-supplied name rew (crates/tinyagents\-harness/src/agent\_loop/tools\.rs:749)
  • medium · security · Start the test module with the required super import — The repository rules require sibling unit-test files to start with `use super::*;`. This file begins with documentation and a standard-library import instead, so it violates the pr (crates/tinyagents\-harness/src/agent\_loop/tool\_rules\_tests\.rs:1)
  • medium · tests · Test the tool-rule branches in the nested-call path — The nested-call path gained two new behaviours — a rule refusal surfacing as `TinyAgentsError::ToolFailed`, and an approval directive (Required/Waived) mapping onto the nested appr (crates/tinyagents\-harness/src/agent\_loop/nested\.rs:870)

Before merge

  • Address Add the rules module before declaring it (crates/tinyagents\-harness/src/tool/mod\.rs).
  • Address Preserve approval directives for rewritten calls (crates/tinyagents\-harness/src/agent\_loop/tools\.rs).
  • Address Preserve compatibility for RunPolicy struct literals (crates/tinyagents\-harness/src/runtime/types\.rs).
  • Address Add the declared test module or remove its registration (crates/tinyagents\-harness/src/agent\_loop/mod\.rs).
  • Address Preserve approval directives for rewritten tool calls (crates/tinyagents\-harness/src/agent\_loop/tools\.rs).

How this fits together

flowchart LR
  n0["run_loop_body"]:::impacted
  n1["Result"]:::impacted
  n2["admit_tool_call"]:::impacted
  n3["HarnessRunStatus"]:::impacted
  n4["run_id"]:::impacted
  n5["host_invocation_binding"]:::impacted
  n0 -->|uses| n3
  n0 -->|calls| n4
  n0 -->|uses| n5
  n2 -->|uses| n3
  n2 -->|uses| n5
  n5 -->|uses| n1
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 17 files; 7 findings. (1 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyagents\-harness/src/tool/mod\.rs — Add the rules module before declaring it
  • Evidence: crates/tinyagents\-harness/src/agent\_loop/tools\.rs — Preserve approval directives for rewritten calls
  • Evidence: crates/tinyagents\-harness/src/runtime/types\.rs — Preserve compatibility for RunPolicy struct literals
  • Evidence: crates/tinyagents\-harness/src/agent\_loop/nested\.rs — Evaluate nested-call rules after preparing arguments
  • Evidence: crates/tinyagents\-harness/src/agent\_loop/tool\_surface\.rs — Reserve registered names before filtering schemas
  • Evidence: crates/tinyagents\-harness/src/agent\_loop/tools\.rs — Gate the intrinsic discovery call before answering it

security

  • Conclusion: Failure
  • Scope reviewed: all assigned evidence
  • Lane summary: Reviewed 15 files; 3 findings. 2 files were not security-reviewed: docs/modules/harness/README.md (prose or tabular data), docs/modules/harness/tool-rules.md (prose or tabular data). _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyagents\-harness/src/agent\_loop/mod\.rs — Add the declared test module or remove its registration
  • Evidence: crates/tinyagents\-harness/src/agent\_loop/tools\.rs — Preserve approval directives for rewritten tool calls
  • Evidence: crates/tinyagents\-harness/src/agent\_loop/tool\_rules\_tests\.rs — Start the test module with the required super import

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The change wires a new tool-rules gate (allow/deny/hide/approval) through the catalogue, tool_search, model calls and nested calls, with two substantial new test files (tool_rules_tests.rs and tool/rules/mod_tests.rs) that exercise refusal, hidden-but-callable, indirect targets, approval directives and hosted-definition stacking — this is well-tested behaviour, not decorative tests. Two gaps remain: the new nested-call branches in nested.rs are never exercised, and the new test module's declaration violates the repository's naming rule. (1 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyagents\-harness/src/agent\_loop/nested\.rs — Test the tool-rule branches in the nested-call path

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The description accurately matches the diff: a new crate-private ToolGate replaces the host_allows closure on every listing and admission site (catalogue, search, calls, nested calls), RunPolicy::tool_rules and AgentDefinition::tool_rules are additive, the breaking struct-literal change is called out, tests follow the repository's sibling test-file convention, and docs are added. The change looks sound and ready to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.036300
  • Tokens: 740939 input · 39889 output · 75896 cached · 0 embedding
Head State Pass summary
4eb24b3b1361 changes requested 10 active finding(s), 0 resolved finding(s) (at 1791528717)

tinysweeper 0.1.0

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-09T06:52:59.874193Z 4eb24b3 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 88d93f93-62e4-4c59-8b35-5c7ea25a3d97
📥 Commits

Reviewing files that changed from the base of the PR and between 37185fd and 4eb24b3.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (18)
  • crates/tinyagents-definition/Cargo.toml
  • crates/tinyagents-definition/src/lib.rs
  • crates/tinyagents-harness/src/agent_loop/mod.rs
  • crates/tinyagents-harness/src/agent_loop/nested.rs
  • crates/tinyagents-harness/src/agent_loop/run_loop.rs
  • crates/tinyagents-harness/src/agent_loop/tool_rules_tests.rs
  • crates/tinyagents-harness/src/agent_loop/tool_surface.rs
  • crates/tinyagents-harness/src/agent_loop/tools.rs
  • crates/tinyagents-harness/src/runtime/agent.rs
  • crates/tinyagents-harness/src/runtime/mod_tests.rs
  • crates/tinyagents-harness/src/runtime/types.rs
  • crates/tinyagents-harness/src/tool/mod.rs
  • crates/tinyagents-harness/src/tool/rules/mod.rs
  • crates/tinyagents-harness/src/tool/rules/mod_tests.rs
  • crates/tinyagents-harness/src/tool/rules/types.rs
  • docs/modules/harness/README.md
  • docs/modules/harness/tool-rules.md
  • vendor/tinytools

Included review availability: This review used your included allowance. Your plan provides up to 2 included reviews per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The harness now supports tool rules from run policies and hosted agent definitions. It applies those rules to tool listings, call admission, and approval decisions. The change also adds tests and harness documentation for these behaviors.

Changes

Tool Rules

Layer / File(s) Summary
Rule inputs and gate evaluation
crates/tinyagents-definition/Cargo.toml, crates/tinyagents-definition/src/lib.rs, crates/tinyagents-harness/src/runtime/types.rs, crates/tinyagents-harness/src/tool/rules/*, crates/tinyagents-harness/src/tool/mod.rs, vendor/tinytools
AgentDefinition gains optional tool_rules. RunPolicy gains a ToolRulePolicy with a context. ToolGate combines policy rules, hosted-definition rules, and the tool allowlist, then evaluates listing and call rules.
Gate enforcement in tool execution
crates/tinyagents-harness/src/runtime/agent.rs, crates/tinyagents-harness/src/agent_loop/*
The runtime passes hosted-definition rules into the invocation binding. Tool surfaces, deferred discovery, direct and nested calls, unknown-tool handling, and approval decisions now use the resolved gate.
Behavior tests and documentation
crates/tinyagents-harness/src/agent_loop/tool_rules_tests.rs, crates/tinyagents-harness/src/agent_loop/mod.rs, crates/tinyagents-harness/src/runtime/mod_tests.rs, crates/tinyagents-harness/src/tool/rules/mod_tests.rs, docs/modules/harness/*
Tests cover listing, call admission, context conditions, indirect targets, approval directives, and combined rule sources. Harness documentation describes rule configuration and behavior.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant AgentDefinition
  participant prepare_agent_turn
  participant run_loop_body
  participant resolve_tool_gate
  participant tools
  AgentDefinition->>prepare_agent_turn: supply tool_rules
  prepare_agent_turn->>run_loop_body: provide HostInvocationBinding
  run_loop_body->>resolve_tool_gate: provide allowlist and rules
  resolve_tool_gate-->>run_loop_body: return ToolGate
  run_loop_body->>tools: dispatch tool call with gate
  tools->>tools: evaluate admission and approval
Loading

Merge Risk

Merge Risk: ⚪ Minimal · up to 4eb24

This change adds optional tool rules that restrict which tools are listed and callable. Behavior is unchanged when no rules are configured. No concrete merge-blocking issue was identified in the supplied material.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to 4eb24

Calls redirected to another tool can lose required approval, and changed calls are not consistently checked against the final restrictions. Existing permission checks limit exposure, but the new rules do not yet provide a complete authorization boundary.

Retained concerns

  • High · security · observed: Unknown-tool recovery checks whether the configured rewrite target is admitted but discards its approval directive. For an initially unknown name, rule_approval remains Default. A model-controlled unknown call can therefore execute the compatibility target without a rule-required approval when the target's declaration does not independently require approval and other authorization checks permit execution. Direct registered calls retain the directive; this gap is introduced with the new rule-based approval contract.
  • Medium · security · inferred: Rules authorize the original tool and provider arguments before mutable before_tool middleware. Final dispatch checks only the exact allowlist and retains the original approval directive. A permitted middleware retargeting an admitted call to an allowlisted but rule-denied tool can consequently reach execution without applying the new rules to that target, provided schema validation and independent authorization allow it. Argument recovery also changes inputs after rule evaluation, but its precise indirect-target exposure remains unresolved because the dependency implementation is unavailable. No in-tree name-rewriting middleware was identified.

Security review details

Security Blast Radius

  • inferred — The evidenced attack scope is execution under a run's available tool authority. Unknown-name recovery reaches a fixed configured compatibility target, while middleware retargeting remains constrained by final allowlist membership and independent authorization. The resources, credentials, or tenants reachable through those tools are not established by the available source.

Security Findings and Attack Paths

  • observed — With unknown-tool rewriting enabled, an attacker-influenced response can supply an unknown name and valid compatibility-tool arguments. The target's Required directive passes the admission filter but is discarded, allowing execution without that approval when the tool declaration and remaining controls allow it.
  • inferred — An installed middleware that changes call identity can move an initially admitted call to a rule-denied, allowlisted target without a second rule decision. This is a conditional API-level bypass, not an established production exploit; an applicable rewriting middleware and permissive remaining controls are required.

Trust Boundaries and Controls

  • observed — Definition rules are carried in invocation-local host binding rather than reusable harness identity. The existing hosted empty-allowlist behavior remains fail-closed by default, and final host authorization remains separate from the new rule decision.

Resilience and Maintainability Implications

  • observed — Ordinary approval denial produces a terminal tool response without execution. Approval with edited arguments updates the call and re-enters serial admission, which reevaluates rules and hosted authorization. Approval tracking remains keyed by call ID; its global uniqueness guarantee was not established, and this mechanism is unchanged by the PR.

Hardening Proposals

  • proposed — Preserve the early rule check and immutable provider payload, but also authorize the final execution target and effective arguments after transformations. Carry the rewrite target's approval directive into the terminal admission decision, combining applicable decisions without weakening a required approval.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 59.09% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 66 functions across 14 files. (4 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check Passed The title clearly and concisely describes the main change: applying tool-rule evaluation to the catalogue, search results, and tool calls.
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 59.09% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 66 functions across 14 files. (4 skipped: 4 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit read the rules at dawn
And checked each tool before moving on
Some stayed hidden from the view
Some needed approval too
The call gate kept its choices clear
Then hopped away with carrots near

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is critical.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0363 · 740,939 in / 39,889 out · 75,896 cached (10%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0194 · 381,079 in / 23,142 out · 42,350 cached (11%) · gpt-5.6-luna, glm-5.3-flash
security:    $0.0164 · 297,690 in / 13,893 out · 30,474 cached (10%) · gpt-5.6-luna
tests:       $0.0002 · 20,433 in  / 862 out    · 1,536 cached (8%)   · glm-5.3-flash
description: $0.0002 · 20,495 in  / 267 out    · 1,408 cached (7%)   · glm-5.3-flash

Comment thread crates/tinyagents-harness/src/tool/mod.rs
Comment thread crates/tinyagents-harness/src/runtime/types.rs
Comment thread crates/tinyagents-harness/src/agent_loop/nested.rs
Comment thread crates/tinyagents-harness/src/agent_loop/tool_surface.rs
Comment thread crates/tinyagents-harness/src/agent_loop/tools.rs
Comment thread crates/tinyagents-harness/src/agent_loop/mod.rs
Comment thread crates/tinyagents-harness/src/agent_loop/tools.rs
Comment thread crates/tinyagents-harness/src/agent_loop/tool_rules_tests.rs
@tinysweeper tinysweeper Bot added the priority: p0 Drop what you are doing. Data loss, a live break, or an exploitable hole. label Oct 9, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4eb24b3b13

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread crates/tinyagents-harness/src/agent_loop/tools.rs
Comment thread crates/tinyagents-harness/src/agent_loop/tools.rs
Comment thread crates/tinyagents-harness/src/agent_loop/tool_rules_tests.rs
@senamakel
senamakel merged commit cf526e3 into main Oct 9, 2026
16 of 18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p0 Drop what you are doing. Data loss, a live break, or an exploitable hole.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant