Skip to content

feat: unify provider factory, protocol surface and credential boundary - #1

Merged
MisonL merged 35 commits into
mainfrom
feat/api-channel-integration
Sep 23, 2026
Merged

MisonL merged 35 commits into
mainfrom
feat/api-channel-integration

Conversation

@MisonL

@MisonL MisonL commented Sep 21, 2026

Copy link
Copy Markdown
Owner

目的

把 provider 层从「各渠道各自为政的调用逻辑」重构为统一工厂 + 抽象请求模型 + 显式协议,并同步收紧凭证边界与失败语义。分支覆盖 feat/api-channel-integration 的 29 个提交,对应 CHANGELOG.md 的 1.4.0。

主要解决的问题:

  1. 协议能力不可见 —— 渠道支不支持 /responses、能不能发 Ark 专属块,此前只能读实现。现在由 src/providers/factory.py 声明能力、抽象接口校验,protocol_status() 可查询。
  2. 本地 SDK 资源被误发 —— 非官方 /responses 端点会被当成官方 OpenAI 放行 SDK 专属资源。现在要求 options.server_verified_protocols 显式登记,未登记时工厂、Provider 和资源 Facade 均在请求前拒绝。
  3. 凭证可能落到日志 —— 模型级 options 与请求级扩展此前可携带密钥、请求头、query 或 Base URL 覆盖。现在由 src/utils/security.py 在配置、请求、日志三个边界校验与脱敏。
  4. 失败被静默掩盖 —— 凭证/SDK/远端调用失败时返回空 embedding、零分 rerank 或占位回答。现在统一显式报错。

影响范围

86 个文件,+34715 / -4312(其中 uv.lock 占 +3583)。

范围 内容
src/providers/ (16) 统一工厂、抽象请求模型、四类线协议适配、原生资源 Facade、SDK 深度适配
src/utils/ (4) 凭证边界与日志脱敏、base_url 升级提示、内置默认值修正
src/services/ (5) src/runtime/ (2) 服务分层与入口重构
src/retrieval/ src/etl/ (7) 注解现代化、死代码清理、embedding 兼容性硬校验
tests/ (15) 协议适配、SDK 能力、凭证边界、Rerank 契约、失败语义、快照兼容性回归
docs/ (4) CHANGELOG.md README.md AGENTS.md 协议选择、options 边界、原生资源入口、Vertex ADC、模型生命周期

破坏性变更(详见 CHANGELOG「移除与调整」):

  • 移除 Dify 上游子模块;src/etl/、src/retrieval/ 移植代码继续遵守 DIFY_LICENSE 并保留上游版权声明。
  • 移除 iFlow 渠道及其配置示例;移除 qwen rerank 示例(该渠道不提供 rerank,调用时显式报错)。
  • Qwen Base URL 改为 compatible-mode/v1,Volcengine Base URL 改为 ark.cn-beijing.volces.com/api/v3。旧值会在配置加载时给出显式升级警告。
  • 发布脚本新增目标环境校验,--target 与主机不匹配时在清理构建产物前失败。

配置与生成数据影响

  • 需要重建知识库快照:gemini-embedding-2 与退役的 text-embedding-004 向量空间不兼容。加载快照时会核对 manifest.toml 记录的 embedding_provider/embedding_model 与当前运行配置,不一致直接报错而非静默复用旧索引。
  • 示例配置与内置兜底默认值已按各渠道当前有效的模型 ID 刷新(Google gemini-embedding-2、OpenAI gpt-5.6-*、Anthropic claude-sonnet-4-6、DeepSeek deepseek-v4-*、SiliconFlow Qwen/Qwen3.5-27B、Grok grok-4.6、Ollama llama3.1 等);修正了 4 个此前无效的兜底 ID。
  • .env 新增 OPENAI_API_BASE 覆盖说明:provider = "openai" 的条目都会走该 Base URL,指向网关时 model_name 必须是网关实际提供的 ID。
  • 火山 doubao-embedding-text-240715 已于 2025-12-26 EOM,官方迁移目标走 /embeddings/multimodal 路径,与本项目文本 embeddings.create 不同,示例配置暂未改动,迁移前需实测。
  • 无新增运行时配置项;config.toml 仍为本地文件,不纳入版本控制。

验证命令

uv sync --group dev
uv run pytest                                    # 823 passed
uv run ruff check main.py src scripts tests      # All checks passed
uv run bandit -r main.py src scripts -ll         # 0 issues
uv run mypy --cache-dir /tmp/pyrag-kit-mypy main.py src   # 49 files, no issues
uv run python -m compileall -q main.py src scripts tests
uv lock --check && git diff --check
uv run main.py --smoke-test

额外验证:

  • CI 可移植性:删除本地 config.toml 与 .env 后完整跑测,823 passed,不依赖开发者本机配置。
  • 逐提交可导入:29/29 提交的 src/ 树均通过导入检查(用 git archive 抽出到临时目录,未触碰工作区)。
  • 端到端链路:快照加载 → 混合检索 → provider 生成,走通完整 RAG 流程。
  • OpenAI 兼容渠道实测:对本地 new-api 网关启用的全部 LLM 条目逐个发起真实请求,均正常返回。
  • 原生 SDK 静态核对:Google GenAI / Anthropic / Volcengine Ark / Jina / SiliconFlow 对照各自官方文档核对调用面(未发真实请求)。

备注

  • HANDOFF.md(本地交接记录)不在本次提交内。
  • 已知未覆盖项:火山 doubao-embedding-vision-251215 迁移路径需真实凭证实测后另行处理。

MisonL and others added 29 commits September 20, 2026 10:39
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Move chat orchestration, retrieval and knowledge-build flows behind
runtime service objects and slim main.py down to wiring. Also harden
the embedding path while touching it: validate the row count returned
by embed_documents against the input, route query embeddings through
the provider-specific task_type, and close the embedding model on
shutdown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add type annotations across retrieval, ETL and release scripts, and
enforce the release target's OS/CPU architecture in
build_binary_release.py so a mismatched --target fails explicitly
before any build artifacts are cleaned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CompletionResult.tool_calls uses the flat shape
{"id","type","name","arguments"}, while OpenAI Chat Completions requires
assistant history tool_calls in the nested shape
{"id","type","function":{"name","arguments"}}. Replaying our own output
into a follow-up request therefore sent a malformed payload and the
server rejected the turn with HTTP 400, breaking multi-turn tool calls
for every OpenAI-compatible channel.

Normalize both shapes at the request boundary so callers can feed the
previous turn's tool_calls back verbatim. Anthropic, Google and Ark
already read the nested form via a flat fallback, so they keep working;
the normalization also fixes the Ark Chat path, which shares
normalize_messages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
153bb07 made normalize_messages require an id/call_id on every
assistant tool_call. That helper runs for every provider, but Gemini's
FunctionCall.id is optional and the SDK leaves it None, so replaying a
Google tool call into the next turn was rejected locally before the
request ever reached the SDK. Google's own _contents accepts id-less
calls, so the regression lived purely in the normalization layer.

Keep the public helper shape-only: it preserves an id when present and
drops the field otherwise. Protocols that genuinely need an id
(Chat Completions, Responses) now pass require_tool_call_ids=True
explicitly, so their stricter validation is unchanged.

Also fix two sibling gaps in the same helper: read the ``input`` alias
used by Responses custom_tool_call and Anthropic tool_use instead of
silently normalizing arguments to an empty string, and map the
function_call/custom_tool_call type values back to ``function``, which
is the only value Chat Completions accepts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two gaps in the credential boundary:

``ak`` and ``sk`` are the key names VolcengineProvider uses to send its
own VOLC_ACCESS_KEY/VOLC_SECRET_KEY, and ``account_key`` is Azure
Storage's, but none were recognized as credential keys, so a request
override carrying them passed validation while ``api_key`` was
rejected. Add them with an exact normalized-key match; ordinary fields
such as task, risk, ask, disk and sketch do not collide.

``RedactingFormatter`` formatted first and redacted second, but
``logging.Formatter.format`` caches the formatted traceback on
``record.exc_text`` for later handlers to reuse. A non-redacting
handler on the same logger therefore still emitted the raw credential.
Clear the cache before formatting and write the redacted text back, so
what later handlers reuse is already scrubbed. This covers handlers
that run after ours; one ordered before ours has already emitted.

Also widen redaction to ``ak=``/``sk=``/``account_key=`` and the URL
query form, using a word-boundary guard so ordinary text is untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1.4.0 moved the Qwen default endpoint to compatible-mode/v1 and the
Volcengine default domain to ark.cn-beijing.volces.com/api/v3, but an
existing config.toml keeps overriding the new defaults with endpoints
that now answer 404/502. Warn at settings load so the fix is obvious
instead of being reverse-engineered from a 404. Only warn: the old
value is a repairable misconfiguration, and failing hard would stop
deployments that still carry it from starting at all.

Jina accepted a return_documents override that SiliconFlow rejects and
that both providers otherwise hardcode to True; pin it so the two
rerank providers behave the same.

Drop the unused splitter_structure_mode parameter from _process_file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add the tool-call normalization regression, the ak/sk and cached
traceback credential fixes, the retired base-url warning and the Jina
return_documents alignment to the 1.4.0 section, and update the test
count to 800.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Record how multi-turn tool history is normalized at the request
boundary and which protocols require an id on assistant tool_calls:
Chat Completions and Responses reject a missing id before the request
is sent, while Gemini's FunctionCall.id is optional and id-less calls
replay fine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The model version regex only matched opus/sonnet/haiku, so the Claude 5
families introduced later (claude-fable-5, claude-mythos-5, and the
unversioned claude-mythos-preview) were classified as pre-4.5 models.
That let temperature/top_p/top_k through to the request, but the
Anthropic Python SDK v1.0+ no longer accepts those parameters at all,
so the call failed with a TypeError instead of the intended omission.

Match the family segment as any word and add a pattern for the
unversioned modern names. Versioned legacy models such as
claude-sonnet-4 and claude-3-5-sonnet-20240620 still report as
non-deprecated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Checked each channel's official model list and replaced IDs that are
retired or no longer offered. Verified against a live gateway: every
replacement answers 200 and every retired ID answers 503.

- Google embedding: text-embedding-004 retired 2026-01-14
- OpenAI: gpt-3.5-turbo retires 2026-10-23, gpt-4o snapshots 2026-10-23
- Anthropic: claude-3-5-sonnet-20240620 retired 2025-10-28
- DeepSeek: deepseek-chat retired 2026-07-24, deepseek-v2 was never a
  real API id
- Volcengine: doubao-pro-32k retired, bge-large-zh absent from the Ark
  model list
- Qwen: qwen-max/qwen-plus are legacy aliases whose snapshots are being
  retired
- SiliconFlow: Qwen2-7B-Instruct and all meta-llama models are gone
- Grok: llama3-70b-8192 is a Groq id, not an xAI one

Also update the built-in llm_configurations defaults and drop the
get_settings() dependency from test_settings_defaults, which made those
assertions depend on whatever the developer's local config.toml held.

Embedding-space changes are not drop-in: gemini-embedding-2 and
text-embedding-004 are incompatible, so an existing snapshot must be
rebuilt rather than reused.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to the catalog refresh: the example section keys still named
retired models (sf-qwen2-7b holding Qwen/Qwen3-8B, grok-llama3-70b
holding grok-4.6), which is misleading for a copy-paste config. Rename
them to match the model they configure and refresh the Ollama entries
to llama3.1/gemma3.

Correct the Volcengine embedding model. The previous value pointed at
doubao-embedding-vision, a multimodal model, but this provider's
embedding path calls the text embeddings.create endpoint; use
doubao-embedding-text-240715 instead.

Also update the provider list examples in the user guide, which still
cited GPT-4o, GPT-3.5-Turbo and Claude 3.5 Sonnet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The version regex matched "claude" anywhere in the string, so an
unrelated name such as "notclaude-sonnet-4-6" was parsed as a modern
Claude model and had its sampling fields silently dropped. Require a
non-alphanumeric boundary before the prefix in both patterns.

Also finish the example alignment: the ollama keys still named llama3
and gemma while holding llama3.1 and gemma3, and the "add a new model"
walkthrough in the user guide used Qwen/Qwen2-57B-A14B-Instruct, which
is no longer offered.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
旧 Qwen 端点的警告文案称其「已失效、会返回 404」,但 DashScope 原生
`/api/v1` 仍在服务,只是路径为 `/services/aigc/text-generation/generation`,
不提供本仓库按 OpenAI 兼容协议请求的 `{base_url}/chat/completions`。
火山旧域名 `maas-api.ml-platform-cn-beijing.volces.com` 才确实已下线。
两者失效方式不同,文案改为分别说明,避免误导用户去排查服务端故障。

同时补充文档:百炼的 `/responses` 只挂在业务空间专属域名
`{WorkspaceId}.{region}.maas.aliyuncs.com/compatible-mode/v1` 下,默认共享
域名未提供该路径,旧版 `/api/v2/apps/protocols/compatible-mode/v1/responses`
已停止维护;在默认域名上启用 `protocol = "responses"` 会收到 HTTP 404。
内置兜底配置里存在几处无效 ID,新用户缺少对应配置段时会照抄到这些值:

- `google` embedding 写的是裸名 `embedding-001`,Gemini API 的有效 ID 带
  `gemini-` 前缀,正确值为 `gemini-embedding-2`(`text-embedding-004` 已于
  2026-01-14 退役,不能作为兜底值)。
- `siliconflow` embedding/rerank 使用 `alibaba/` 命名空间,但官方模型广场
  只有 `BAAI/` 前缀,不存在 `alibaba/`;rerank 现行 ID 为
  `BAAI/bge-reranker-v2-m3`,没有 `bge-reranker-large`。
- `ollama` 的 `llama3` 改为官方现行 tag `llama3.1`。

同步把 config.toml.example 中火山文本向量化模型的注释改为待核实口径:
`doubao-embedding-text-240715` 无法从官方来源确证(文档站为 JS 渲染,
签名接口需账号凭证才能校验模型 ID),已在注释中说明请以控制台开通管理为准,
并区分文本 embeddings.create 与多模态 multimodal_embeddings.create 两条路径。
用 firecrawl 抓到方舟官方文档后,上一轮写在 config.toml.example 的
「doubao-embedding-text-240715 未经官方来源确证」是错的:该 ID 确实存在,
方舟《文本向量化 API》文档给出的请求示例就是它。真实情况是它已进入下线流程:

- 它属于 2025-12-19 启动的那一批,EOM(停止新购)2025-12-26,官方下线公告
  已把它列在迁移表中,建议迁移到 doubao-embedding-vision-251215。
- 方舟当前模型列表的「向量化能力」章节只列 doubao-embedding-vision-251215
  与 -250615,已无任何纯文本 embedding 模型;《文本向量化 API》文档本身位于
  文档站的「下线文档归档」下。
- embedding 模型只有 EOM 阶段、不涉及 EOS,存量接入点不受影响,所以存量部署
  不会被切断;但新接入点已无法创建。

因此示例配置暂不改值,改为在注释中说明真实状态与迁移目标。另需注意迁移目标
doubao-embedding-vision-251215 的官方示例走 /api/v3/embeddings/multimodal,
与本项目火山 embedding 使用的文本 embeddings.create 不是同一条路径,启用前
必须实测。
用 firecrawl 复核各渠道官方页面后,SiliconFlow 的 `Qwen/Qwen3-8B` 已不在该
平台在售列表中——官方模型广场与定价页两个来源均无此 ID,其 Qwen 对话模型现
从 `Qwen/Qwen3.5-27B` 起步。示例配置与键名同步改为 `sf-qwen3-27b`。

同批复核确认无需改动的项:
- Anthropic:`claude-sonnet-4-6` 在官方退役页为 Active(退役不早于 2027-02-17),
  且是 Sonnet 3.5/3.7 与 Sonnet 4 的官方推荐替代;`claude-opus-5`、
  `claude-sonnet-5`、`claude-fable-5-1` 均为 Active。采样参数弃用判定对 5 代与
  4.x 全部生效、对 3.x 正确放行,无需调整正则。
- Google:`gemini-2.5-flash`/`gemini-2.5-pro` 均为 "No shutdown date announced";
  `gemini-embedding-2` 正确,`text-embedding-004` 已于 2026-01-14 退役。
- OpenAI:`gpt-5.6-sol`/`gpt-5.6-terra` 是官方给出的推荐替代目标(现行),
  `text-embedding-3-large` 无退役公告。
- SiliconFlow embedding/rerank:`BAAI/bge-large-zh-v1.5`、`BAAI/bge-m3`、
  `BAAI/bge-reranker-v2-m3` 均在售。
- Jina:`jina-reranker-v2-base-multilingual` 仍在列(官方旗舰已转为 v3.5,
  但旧版未下线,示例保留)。
- Ollama:`llama3.1` 官方 tag 有效(8b/70b/405b,128K 上下文)。
- DeepSeek:`deepseek-v4-pro`/`deepseek-v4-flash` 现行。
- xAI:`grok-4.6` 现行。
`VectorStoreFactory._load_existing_state` 会在活动快照记录的 embedding
provider/模型名与当前运行配置不一致时抛 RuntimeError,但此前没有任何测试覆盖
这条防护。它守的是数据正确性:向量空间不兼容时静默加载旧索引会产出错误检索
结果,而不是显式失败。

新增三个用例,分别覆盖 provider 不同、provider 相同但模型名不同(如
text-embedding-004 与 gemini-embedding-2 同属 google)、以及配置一致时应正常
加载。已做红绿验证:临时禁用检测逻辑后两个不匹配用例失败,还原后全部通过。

同时让 MockFaissStore 实现 load_snapshot,使匹配场景能走到实际加载路径。
服务端返回的凭证错误常带「首尾可见、中间掩码」的形态,例如
`sk-abc***...***xyz`。`_OPENAI_KEY_RE` 要求 `sk-` 之后连续 8 个以上字母数字,
星号会中断匹配,于是整串原样落进日志。

新增 `_MASKED_CREDENTIAL_RE` 单独覆盖这一形态,并在 `redact_sensitive_text`
中于 `_OPENAI_KEY_RE` 之前应用。已做红绿验证:移除该规则后两个掩码用例失败,
还原后通过。
官方 Ark Responses 文档明确 `instructions` 不可与缓存能力一起使用,且
`caching` 配置为 `{"type": "enabled"}` 时请求会直接报错。SDK 不做本地校验,
会原样发到服务端,用户只能从远端错误反推原因。

`_build_responses_request` 此前允许两者同时出现:`instructions` 来自
`system_prompt`,而 `CompletionRequest.system_prompt` 带兼容默认值
`"You are a helpful assistant."`,因此只要用 `prompt=` 且传了 `caching`
就会命中。现在在构造阶段显式拒绝,错误信息点明默认值来源并给出修复方式
(显式传 `system_prompt=None` 并把系统提示放进 `messages`,或移除 `caching`)。

已做红绿验证:移除校验后对应用例失败,还原后通过。
用 AST 扫描全部 52 个源文件的 1519 个函数/方法定义,逐个核对引用后删除以下
死代码。它们的共同特征是同一行为已有另一处实现:

- `has_system_message`:逻辑已在 `normalize_messages` 内联
  (`any(message.get("role") == "system" ...)`)。
- `_configured_extra_body_keys`:`set(configured_extra_body)` 在
  `_build_chat_request` 与 `_build_responses_request` 各内联一次。
- `SnapshotRepository.write_stats` / `load_stats`:`stats.json` 实际由
  `FaissStore.save_snapshot` 内联写入,且全项目从未读取该文件。
- `SnapshotRepository.load_active_manifest`:`get_active_snapshot_dir` +
  `validate_snapshot_dir` + `load_manifest` 的组合封装,无调用点。
- `recursive_text_splitter` 的 `DEFAULT_CHILD_CHUNK_SIZE` /
  `DEFAULT_CHILD_CHUNK_OVERLAP`:`__init__` 用 `None` 默认值并回退到
  `Settings`,这两个常量未被引用,且与 `config.py` 的默认值重复。

保留了由框架调用的函数(pydantic `@field_validator`、自定义
`PydanticBaseSettingsSource.get_field_value`)和公开迁移入口
(`FaissStore.import_legacy_snapshot`),它们无显式调用点但并非死代码。

净删除 37 行,无新增;823 项测试保持通过。
补充三类此前只在代码里存在的约束:

- `.env` 的 `OPENAI_API_BASE` 会覆盖 `config.toml` 的 `openai_api_base`,指向网关时
  `model_name` 必须是网关实际提供的 ID。
- 旧端点升级警告的实际失效方式:DashScope 原生 `/api/v1` 仍在服务但不提供
  `{base_url}/chat/completions`;火山旧域名已下线。
- 活动快照的 embedding 兼容性硬检查,以及切换模型后需重建快照的原因。
- 模型 ID 退役表、`kimi-k3` 的 temperature=1 限制、Ark `instructions` 与 `caching` 互斥,
  以及火山 embedding 与百炼 Responses 域名的已知问题。
@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 25ddec55-b27d-4855-b2da-c4fba89bd6f5


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @MisonL, your pull request is larger than the review limit of 150,000 diff characters

MisonL and others added 5 commits September 22, 2026 01:53
1. 脱敏正则与 is_sensitive_option_key 用了两套键名词表:client_secret、
   private_key、credentials、clientsecret 在配置边界被判为凭证,却因文本
   正则词表更小而明文落进日志与错误信息。改为由单一 _TEXT_CREDENTIAL_KEYS
   派生全部文本正则,消除「同一概念两处实现」的根因。

2. 掩码凭证规则排在 Bearer 规则之后:``_BEARER_TEXT_RE`` 先吃掉
   ``Bearer sk-abc`` 的可见前缀,使 ``sk-`` 锚点失效,``***...***xyz``
   原样残留。改为掩码优先,并放宽字符类以覆盖 ``.`` 与 ``…`` 占位符。

3. ``_url_credential_paths`` 在 urlsplit 抛 ValueError 时直接 return []。
   该分支是叶子值的唯一检查入口,于是任何凭证只要拼在一个畸形 URL 片段
   (如 ``https://[::1``)后面就能整体绕过边界校验。改为退回文本脱敏判定。

4. 新增共用的 safe_exception_text:终端与日志是凭证最易泄漏的出口,
   SDK 异常常回显请求头与 URL,各调用点此前各自处理且不一致。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1. Ark Responses 非流式入口读取 ``response.output_text``,但 Ark SDK 的
   ``Response`` 没有该字段(OpenAI SDK 才把它实现为聚合 property),正文位于
   ``output[].content[].text``。于是 complete / acomplete / invoke(stream=False)
   / ainvoke(stream=False) 全部返回空文本且不报错,ChatService.identify_intent
   随后抛出「意图识别返回空结果」。改为从 output 内容块提取正文,并在标记
   completed 却无正文、工具调用、拒答与推理时显式失败,不再用空结果掩盖。

2. server_verified_protocols 门禁只在显式别名上实现,动态资源树
   (``resources.responses.create(...)``)只做凭证校验即可发出请求。改为在
   ``_NativeResourceProxy.__call__`` 复用 provider 的统一钩子,两条路径共用
   同一不变量;传完整路径而非末段方法名,避免误拦 ``files.create``。

3. instructions 与 caching={"type":"enabled"} 的互斥检查只看顶层 ``caching``,
   漏掉 ``extra_body``(Ark SDK 会把它合并进请求体)与原生 kwargs 两条路径。
   抽出 ``_reject_instructions_with_enabled_caching`` 并覆盖全部来源。

4. 凭证校验提到能力门禁之前:否则携带凭证的调用会报出能力错误,把安全问题
   掩盖成配置问题。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1. Ark 专属内容块(input_audio / input_video / image_pixel_limit)写成 content
   会被 variant 守卫拒绝,写成顶层 item 却因不带 role/content 而直通请求体,
   静默发给不支持它的 OpenAI 端点。补一次等价检查,让两种书写形态一致。

2. Responses custom tool 的 ``input`` 按 SDK 契约是任意文本,但回填时被
   统一走 _responses_json_text 强制 json.loads,导致多轮 custom tool 在请求
   边界直接失败。改为按类型分流:custom 保留原始文本,function 才做 JSON
   规范化。由于 Chat Completions 只接受 ``type: "function"``,归一化时会丢掉
   custom 语义,因此新增 keep_responses_type 开关,仅在 Responses 路径保留
   原始类型,Chat 路径不受影响(内部字段不会进入请求体)。

3. 流式事件同时给出 ``id``(输出项 ID)与 ``call_id``(调用 ID)两个不同的值,
   而 Chat 归一化只保留 ``id``,导致回填 Responses 时 function_call 与
   function_call_output 的 call_id 配不上。Responses 路径一并保留 call_id。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
chat 路径已有 _safe_exception_text,但 retrieval_test 与顶层 CLI 边界仍把裸
异常交给 console.print 与 logger.error,凭证会原样落到终端与日志。SDK 异常
常回显请求头与 URL,这是最现实的泄漏出口。

把该助手提到 utils/security.py 作为共用实现(safe_exception_text),三处调用
点统一使用,消除「同一场景在一条路径做了脱敏、在另一条没做」的不一致。
测试同步改为从共用位置导入。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
为每条修复补回归测试,并逐条做过红绿验证(先回退修复确认测试失败,再恢复
确认通过),避免写出无论代码对错都会通过的测试:

- 脱敏键名遍历 SENSITIVE_OPTION_KEYS,同时断言普通文本(``token usage``、
  ``task = value``)不被误伤。
- 掩码凭证在 Bearer 前缀下完整脱敏。
- 畸形 URL 与可解析 URL 同等严格。
- Ark 非流式正文提取改用**真实** SDK 模型构造响应;此前夹具用
  ``SimpleNamespace(output_text=...)`` 伪造了 Ark 并不存在的字段,恰好掩盖
  了空正文缺陷。
- 动态资源树受 server_verified_protocols 门禁约束,且不误拦 files 等资源。
- instructions × caching 互斥覆盖 extra_body 与原生 kwargs。
- Ark 专属块在顶层 item 形态下同样被拒,Ark 渠道自身仍放行。
- custom_tool_call 往返保留任意文本与空 input,function_call 仍做 JSON 规范化,
  Chat 路径不泄漏 Responses 专属字段。
- function_call 与 function_call_output 回填后共享同一 call_id。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@MisonL

MisonL commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

追加:对抗式审查修复(5 个提交)

在 PR 提交后对全分支做了一轮多维度对抗式审查(9 个维度 × 每条 finding 两路独立反驳验证),确认并修复了 11 个真实缺陷。每条都独立复现过,并做了红绿验证(先回退修复确认测试失败,再恢复确认通过)。

修复清单

# 严重度 位置 缺陷
1 critical volcengine.py Ark Responses 非流式入口读取 response.output_text,但 Ark SDK 的 Response 没有该字段(OpenAI SDK 才把它实现为聚合 property),正文在 output[].content[].text。于是 complete/acomplete/invoke(stream=False)/ainvoke(stream=False) 全部返回空文本且不报错,ChatService.identify_intent 随后抛「意图识别返回空结果」
2 major security.py _url_credential_paths 在 urlsplit 抛 ValueError 时 return []。该分支是叶子值的唯一检查入口,任何凭证只要拼在一个畸形 URL 片段(如 https://[::1)后就能整体绕过边界校验
3 major security.py 掩码凭证正则排在 Bearer 规则之后:_BEARER_TEXT_RE 先吃掉 Bearer sk-abc 的可见前缀,使 sk- 锚点失效,***...***xyz 原样残留
4 major security.py 脱敏正则与 is_sensitive_option_key 用了两套键名词表:client_secret、private_key、credentials 在配置边界被判为凭证,却因文本词表更小而明文落进日志
5 major resources.py server_verified_protocols 门禁只在显式别名上实现,动态资源树 resources.responses.create(...) 只做凭证校验即可发出请求
6 major volcengine.py Ark instructions × caching={"type":"enabled"} 互斥检查只看顶层 caching,漏掉 extra_body(Ark SDK 会把它合并进请求体)与原生 kwargs 两条路径
7 major model_provider.py Ark 专属内容块写成 content 会被 variant 守卫拒绝,写成顶层 item 却因不带 role/content 而直通请求体,静默发给 OpenAI 端点
8 major model_provider.py Responses custom tool 的 input 按 SDK 契约是任意文本,回填时被强制 json.loads,多轮 custom tool 完全不可用
9 major retrieval_test/core.py 裸异常交给 console.print 与 logger.error,凭证原样落到终端与日志
10 major main.py 同上(顶层 CLI 边界)
11 major model_provider.py 流式事件同时给出 id 与 call_id 两个不同的值,Chat 归一化只保留 id,导致回填 Responses 时 function_call 与 function_call_output 的 call_id 配不上

值得注意的两点

测试夹具掩盖了 critical 缺陷。 #1 之所以长期未被发现,是因为既有测试普遍用 SimpleNamespace(output_text=...) 构造响应——而 Ark 的 Response 根本没有这个字段。夹具伪造了线上不存在的形状,让测试永远通过。新测试改用真实 SDK 模型(Response.model_validate)构造。

#11 曾被两路验证判定为误报。 完整性批评复核后指出它是「真实但更窄」的缺陷,验证者关于「该代码路径不可达」的结论不成立。实测确认:流式产出确实携带两个不同值,回填时 call_id 丢失。

验证

uv run pytest -q                                    # 880 passed(新增 57 个回归测试)
uv run ruff check main.py src scripts tests         # All checks passed!
uv run mypy --cache-dir /tmp/pyrag-kit-mypy main.py src   # 49 files, no issues
uv run bandit -r main.py src scripts -ll            # 0 issues
uv run python -m compileall -q main.py src scripts tests
uv lock --check && git diff --check

回归测试均经红绿验证:回退修复后对应测试失败,恢复后通过。

已知未处理

完整性批评还指出几项未纳入本次修复的事项,它们不属于本次改动引入的缺陷,但值得后续跟进:

  • requirements.txt 在依赖升级提交中从「仅生产依赖」变为「含 dev 工具链」,PR 描述未申报。
  • 仓库内不存在 Apache 2.0 全文(DIFY_LICENSE 是摘要),而 CHANGELOG 声称「保留上游版权声明」。
  • 新增 823 项测试与质量工具链,但无 CI 强制执行(release.yml 只在 v* tag 触发)。
  • PyRAG-Kit.spec 是构建产物但未进 .gitignore。
  • src/retrieval_test/core.py 创建 RetrievalService/EmbeddingService 后未释放。

🤖 Generated with Claude Code

对抗式审查的完整性复核指出三项遗留问题,逐条核实后处理:

1. 仓库内没有 Apache 2.0 全文。Dify 衍生代码(src/etl/、src/retrieval/ 等)
   遵循修改后的 Apache 2.0,而 DIFY_LICENSE 只是「引用 Apache 2.0 + 附加
   条款」的摘要。该许可证第 4(a) 条要求随分发提供许可证副本,此前移除
   Dify 子模块时把唯一的副本一并删掉了。补入 licenses/APACHE-2.0.txt,
   纳入发布包 PACKAGE_FILES,并在 DIFY_LICENSE 中指向它。
   PACKAGE_FILES 现在含子目录路径,copy2 不会创建父目录,一并补上。

2. retrieval_test 创建 RetrievalService/EmbeddingService 后从不释放,
   而 chat 与 retriever 路径都做了。Embedding provider 与向量存储持有
   SDK 客户端与连接池,退出时不释放会留下未关闭资源。用 try/finally 包裹
   交互循环,释放失败只记录警告,不掩盖主流程结果。

3. PyRAG-Kit.spec 是构建产物(生成后即 unlink)却未进 .gitignore,工作区
   容易误提交。补 *.spec。

以下两项经核实**不成立**,未改动:
- 称 requirements.txt 由「仅生产依赖」变为「含 dev 工具链」属未申报变更。
  实测 main 版本的 requirements.txt 同样含 pyinstaller 与 pytest——它是
  uv export 的默认全量导出,口径未变,本次只是 dev 分组新增了三个工具。
- 称 recursive_text_splitter.py 与 snapshot_repository.py 属移植代码却缺
  Dify 归属头。两者分别完全建立在 LangChain 的 RecursiveCharacterTextSplitter
  与本项目的 src.runtime.contracts 之上,不含 Dify 特有逻辑;补一个不实的
  归属头反而是错误。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@MisonL

MisonL commented Sep 22, 2026

Copy link
Copy Markdown
Owner Author

追加:遗留问题核实与处理

上一轮完整性复核提出的 5 项遗留问题,逐条核实后处理。3 项成立已修,2 项经实测不成立。

已修(提交 34b8319)

1. 仓库内缺少 Apache 2.0 全文(许可合规)

src/etl/、src/retrieval/ 等 Dify 衍生代码遵循修改后的 Apache 2.0,而 DIFY_LICENSE 只是「引用 Apache 2.0 + 附加条款」的摘要。该许可证第 4(a) 条要求随分发提供许可证副本——移除 Dify 子模块时把唯一的副本一并删掉了,发布包此前只分发摘要。

处理:补入 licenses/APACHE-2.0.txt(202 行官方全文),纳入发布包 PACKAGE_FILES,并在 DIFY_LICENSE 中指向它。PACKAGE_FILES 现在含子目录路径,而 shutil.copy2 不创建父目录,一并补上 mkdir。

2. retrieval_test 未释放服务

该路径创建 RetrievalService/EmbeddingService 后从不释放,而 chat 与 retriever 路径都做了。Embedding provider 与向量存储持有 SDK 客户端与连接池,退出时留下未关闭资源。

处理:用 try/finally 包裹交互循环,新增 _aclose_quietly 助手(优先 aclose,回退 close,失败只记警告,不掩盖主流程结果)。

3. PyRAG-Kit.spec 未进 .gitignore

该文件由 build_binary_release.py 生成后即 unlink,但工作区容易误提交。已补 *.spec。

经核实不成立,未改动

4. 称 requirements.txt 从「仅生产依赖」变为「含 dev 工具链」,属未申报变更

实测 main 版本的 requirements.txt 同样含 pyinstaller 与 pytest——它本来就是 uv export 的默认全量导出(文件头自述的生成命令不含 --no-dev)。口径未变,本次只是 dev 分组新增了 ruff/mypy/bandit 三个工具。该文件也不被任何流程引用(发布流程用 uv sync --group dev)。

5. 称 recursive_text_splitter.py 与 snapshot_repository.py 属移植代码却缺 Dify 归属头

查证两者来源:前者完全建立在 LangChain 的 RecursiveCharacterTextSplitter 上,依赖 src.models.document/src.utils.config;后者依赖 src.runtime.contracts,处理本项目自己的快照目录格式。两者都不含 Dify 特有逻辑(同目录下确实移植的 10 个文件都带归属头,格式一致)。补一个不实的归属头反而是错误,故不加。

另:CI 门禁放在 PR #2

新增 .github/workflows/quality.yml,在 PR 与 main 推送时执行锁文件校验、格式化检查、lint、类型检查、安全扫描、字节码编译与完整测试——此前这些工具「加了但无人运行」(release.yml 只在 v* 标签触发)。

它含 ruff format --check,而本 PR 尚未格式化,放这里会必然失败,因此归入 PR #2(格式化分支)。PR #2 合并后 main 上才有门禁;若先合本 PR,CI 不会生效。

已逐条本地模拟全部 8 步,均通过:

uv lock --check                  Resolved 107 packages
ruff format --check              81 files already formatted
ruff check                       All checks passed!
mypy                             49 source files, no issues
bandit                           0 issues
compileall                       OK
pytest                           880 passed
git diff --check                 OK

🤖 Generated with Claude Code

@MisonL
MisonL merged commit 5899d61 into main Sep 23, 2026
3 checks passed
@MisonL
MisonL deleted the feat/api-channel-integration branch September 23, 2026 15:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant