Skip to content

fix(feishu): 发文件传绝对路径被 lark-cli 拒收,飞书发文件 100% 失败 - #51

Open
AmberCXX wants to merge 12 commits into
LinekForge:mainfrom
AmberCXX:fix/feishu-send-file-relative-path
Open

AmberCXX wants to merge 12 commits into
LinekForge:mainfrom
AmberCXX:fix/feishu-send-file-relative-path

Conversation

@AmberCXX

Copy link
Copy Markdown
Contributor

动机

hub_send_file 走飞书必然失败,不是偶发。

send() 的 file 分支把绝对路径传给 lark-cli im +messages-send --file,而 lark-cli 只接受当前工作目录下的相对路径。两个版本都拒,措辞不同:

1.0.24: --file must be a relative path within the current directory,
        got "/abs/path" (hint: cd to the target directory first)

1.0.95: --file "/abs/path" is outside the built-in allowlist;
        allowed roots are the current working directory ...

升级解决不了 —— 1.0.95 把规则收紧成显式 allowlist,不是放宽。

排查时容易被带偏的一点:失败时 lark-cli 先打一行 proxy 警告,真正的 validation 错误在下一行。

改动概要

file 分支改成 basename(filePath) + cwd: dirname(filePath)。

同一个文件里这个约束已被正确处理过两次,只有 file 分支漏了:

  • voice 分支:"--audio", basename(filePath) + cwd: dirname(filePath),并带注释「lark-cli 的 --audio 接 basename,实际文件要在 cwd 里找」
  • downloadFeishuMedia():cwd: FEISHU_MEDIA_DIR + 相对 --output,注释写着 // lark-cli only accepts relative paths

所以这是漏改,不是设计分歧——修法与既有两处保持一致,未引入新约定。

影响范围

  • 仅 hub-server/channels/feishu-lark-cli.ts 的 send() file 分支,8 增 2 删
  • text / voice / 入站下载 三条路径未动
  • 不改公共契约、不改配置、不加依赖

self-test 结果

(cd hub-server && bun install && bunx tsc --noEmit && bun test)
  tsc  → 0 errors
  test → 356 pass / 0 fail / 2410 expect() · 34 files · 15.16s

bun hub-test-harness/harness.ts
  → 8/8 通过

端到端实测(lark-cli 1.0.24 与 1.0.95 各跑一遍):

写法 结果
绝对路径(改前) exit 2,validation error
basename + cwd(改后) exit 0,返回 message_id,文件实际送达

部署后用 MCP hub_send_file 传绝对路径发送 PDF,两个版本下均成功投递。


⚠️ 顺带一个本 PR 范围外的观察,供参考、不在此处修:

forge-hub sync 经由 ~/bin/forge-hub(symlink → ~/.forge-hub/package/cli.ts)调用时,PKG_ROOT 由 import.meta.url 反推,落在已部署的快照目录上,于是把运行时跟自己同步了一遍并打印「✓ 已更新」。从仓库目录跑 bun cli.ts sync 才会带上源码改动。这与 #50 描述的「sync 报告成功但实际没生效」是不同的成因、同一个症状。

🤖 Generated with Claude Code

https://claude.ai/code/session_01FxWdmB6ACyBEKA8wBAbT1K

AmberCXX and others added 12 commits August 22, 2026 00:51
scheduleOrigin 建定时器时只过 canScheduleToday(仅查 start_date/end_date),
未过 shouldFire,导致两个问题:

1. weekdays/days/months 不参与判定——「每周日」的任务会在周六被
   报成「今天错过」
2. delay <= 0 一律落进 missed 分支——午夜重排本身发生在 00:00,
   重排那一刻就把 00:00 的格子算成已过去,该时段任务永远跑不到

修法:
- 新增 FIRE_GRACE_MS(90s)宽限支,刚过点的任务立即补触发而非报错过
- missed 与宽限两支都补 shouldFire(entry)

shouldFire 改为 export 以便测试,并补 3 个回归测试。

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fd3qWDATPyaLxwM8s1Ag9
错过判定原本只看两样:delay 落在 MISSED_WINDOW_MS(2h) 内,以及 shouldFire
的排期规则。它从不查该条目今天是否已经触发过——fire() 只写 global.json 的
last_fire / last_sender / today_count 这类全局标量,回答不了「这一条今天跑没跑」。

于是任务时刻之后 2 小时内的任何一次重排,都会把已经跑完的任务重新报一遍:
启动、配置热加载的部分重排,以及午夜重排在边界上提前几毫秒跑到前一天时
(此时 22:00 的 delay 约为 -1h59m,正落在窗口里)。2026-08-22 / 08-23 /
08-28 / 08-29 连续复发,每次都要人工判真假,判错的代价是重复推送或漏跑。

改动:新增 per-entry 的已触发记录(forge-state/fires.json)
- fireKey(entry)  来源文件 + 时刻 + 名字(缺省回退 sender)作稳定标识
- hasFiredToday(state, entry, today)  纯函数,state 由调用方注入
- markFired(state, entry, today)  返回新 state,并清掉非当天的键,文件不增长
- fire() 成功推送后落一条记录;scheduleOrigin 每次重排只读一次状态文件
- 错过分支追加 !hasFiredToday(...)

真实的漏跑不受影响:engine 当时没跑就没有记录,照常上报。

验证:forge-engine 28 tests 全过(新增 6 条覆盖键稳定性、label 缺省回退、
跨日失效、清理旧键、同日幂等);bunx tsc --noEmit 无错;
bun hub-test-harness/harness.ts 8/8;fh hub self-test 8/8。

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EWhZ8wD8Ne5yiYaCdL6cjf
engine 的 log() 此前只写 stderr,而 MCP 宿主通常只在连接建立那一刻捕获
stderr——启动之后打印的任何东西都不再被记录。进程一旦异常退出,线索为零。

实测(2026-09-03):engine 在 00:00-03:00 之间消失,宿主日志无断开记录、
系统无崩溃报告、无内存压力事件,调度静默停摆 10 小时 31 分、漏跑 12 个任务,
事后无从判断死因。

改动:
- log()/logError() 同时追加到 <DATA_DIR>/engine.log,超 2MB 转存 .log.1
- 新增 logFatal(),带完整堆栈落盘
- 注册 uncaughtException / unhandledRejection:记录后 stopScheduler()
  释放 PID 锁并 exit(1)

两处取舍写在注释里:
- 抓到异常后退出而非续跑——状态可能已损坏,错误的推送比没有推送更坏
- 不做自动重起——多实例 + 自动重起 + 抢 PID 锁是本项目已踩过的坑;
  恢复由外部检测触发人工重连,engine 自己不做进程管理

验证:tsc --noEmit 通过 · bun test 19 pass · 临时 DATA_DIR 实测日志落盘、
FATAL 带堆栈、2MB 轮转生效

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EWhZ8wD8Ne5yiYaCdL6cjf
本地 main 作为唯一运行版本——engine 的 MCP 入口直接指向工作区,
checkout 在哪个分支跑的就是哪个版本。分散在功能分支上等于生产环境
随 checkout 漂移(2026-09-04 实测:切到修复分支时 crash-log 当场失效)。

PR LinekForge#46 仍挂在上游等合并,本地先用上,不冲突。
同上:本地 main 收敛为唯一运行版本。PR LinekForge#47 仍挂上游等合并。
issue LinekForge#27 的 PID lock(273bf3a)让第二个实例进入 passive 模式以避免重复触发,
但 acquirePidLock() 只在启动时调用一次。active 实例退出后,已经处于 passive 的
实例不会重新竞选——即使它们还活着、MCP 还连着,调度也静默停摆到下一个新 session
启动为止。

原实现处理了 stale 锁回收,但只在启动路径上:新实例发现旧持锁者已死时接管。
没覆盖的是反过来的顺序——已经在场的 passive 实例发现 active 死了该怎么办。

实测(2026-09-08):三个实例,active 被关闭后剩下两个 passive。此时 `/mcp` 显示
连接正常、进程在、MCP 工具可调用,但 PID 锁文件消失、engine.log 停止输出,全部
定时任务不再触发。故障唯一的可见信号是一个没人会主动去看的锁文件。

改动:
- passive 实例启动 30s 间隔的重新竞选定时器,夺锁成功即接管调度
- 抽出 activate(),首次启动与事后接管共用同一段排程逻辑
- stopScheduler() 清理竞选定时器
- 导出 isPassiveMode() / retryPassivePromotion() 供测试直接驱动,不必等 interval

acquirePidLock() 本身未改:持锁者仍存活时照旧返回 false。所以接管只可能发生在锁
被释放(正常退出)或变成 stale(崩溃)之后,不会退回 LinekForge#27 的重复触发。

测试 forge-engine/scheduler-pid-lock.test.ts(4 个,含一条防止修过头的反向测试):
- 持锁者存活时进入 passive
- 持锁者正常退出后自我提升
- 持锁者崩溃留下 stale 锁时自我提升
- 持锁者仍存活时保持 passive,且不夺锁

验证:
- (cd forge-engine && bunx tsc --noEmit && bun test) → 23 pass / 0 fail
- bun hub-test-harness/harness.ts → 8/8

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017pGyNVmhbA1Z1vRoxLZbFe
飞书把图片以 "[图片: img_xxx]" 的形式塞在 content 文本里。入站处理用
content.match() 取 key——没有 g 标志,只返回第一个匹配——随后把**整个
content 覆盖**成那一张的路径:

    const imageKeyMatch = content.match(/\[(?:Image|图片):\s*(img_[^\]]+)\]/);
    if (imageKeyMatch && messageId) {
      const imageKey = imageKeyMatch[1];
      const filePath = await downloadFeishuMedia(messageId, "image", imageKey);
      content = filePath ? `[图片] ${filePath}` : `[图片: ${imageKey}]`;
    }

于是一条「文字 + 6 张图」的富文本到达 agent 时只剩 "[图片] <第一张路径>",
文字和其余 5 张静默消失。用户侧的表现是「一次发六张图,对面只收到第一张」,
而且没有任何错误或日志——后 5 张根本没被下载过。

改动:
- 抽出 resolveImagePlaceholders(content, download):用 matchAll 遍历全部占位符,
  逐个**就地**替换成 "[图片] <路径>",占位符以外的文字原样保留
- 单张下载失败只保留该张的原占位符,不牵连同条消息里其它图片
- download 以参数注入,函数无 IO 依赖,可直接测试
- handleMessage 改用它;检测用不带 g 的常量正则,遍历时另建带 g 的副本,
  避免共享 lastIndex 导致漏匹配

测试 hub-server/feishu-media-placeholders.test.ts(6 个):
- 单张占位符替换为路径
- 多张全部替换,且下载按序对每个 key 各调一次
- 图文混排时占位符以外的文字原样保留
- 兼容英文 [Image: ...] 形式
- 某张下载失败时保留该张占位符,其余照常替换
- 不含占位符的纯文字原样返回,且一次下载都不发起

验证:
- (cd hub-server && bun test) → 356 pass / 0 fail
- bun hub-test-harness/harness.ts → 8/8
- bunx tsc --noEmit:改动前后均只有既有的 channels/feishu.ts 缺
  @larksuiteoapi/node-sdk 一条,未新增

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017pGyNVmhbA1Z1vRoxLZbFe
两套机制都在"确保 Hub 在运行",互相不知道对方存在:

1. launchd(com.forge-hub.plist,RunAtLoad + KeepAlive)
2. hub-client 的 ensureHubRunning():探测不可达就 spawn(..., {detached:true}) + unref()

第 2 条造出的进程不是 launchd 的子进程。它抢到端口后:

- launchd 每 ThrottleInterval(30s) 试着拉起自己的实例,撞端口崩掉,无限循环,
  hub-stderr.log 里刷满 "Failed to start server. Is port 9900 in use?"
- `forge-hub sync` 的 `launchctl bootout` 杀不到它(不归 launchd 管),随后的
  bootstrap 同样撞端口失败,而 `log("✓ Hub 已重启")` 是无条件打印的

净效果:**sync 报告重启成功,实际跑的仍是旧代码**。实测一台机器上 Hub 进程连续
存活 4 天 6 小时,跨越多次 sync 从未被替换;`ps -o ppid` 显示它的父进程是某个
Claude Code session 的 hub-channel.ts,不是 launchd。

修复分两处:

**hub-client(根因)**:新增 hub-starter.ts,装了 plist 时走
`launchctl kickstart` 让 launchd 启动,保证 Hub 始终是 launchd 的子进程;
kickstart 失败(service 未 bootstrap)则回退到原来的 detached spawn。
非 macOS 或没装 plist 时行为完全不变。

**cli.ts sync(止损)**:新增 waitUntil(),把"命令被接受"和"状态真的变了"分开。
bootout 之后轮询确认 Hub 真的下线,等不到就明确报告"它多半不是 launchd 启动的、
bootout 管不到"并给出手动步骤,**不再打印已重启**;bootstrap 之后轮询确认 Hub
恢复响应,成功消息改为有条件打印。

测试:
- hub-client/hub-starter.test.ts(4):darwin+plist→launchd;darwin 无 plist→spawn;
  linux 有/无 plist 均→spawn
- cli.test.ts 新增 waitUntil(4):立即成立不轮询;中途翻转;超时返回 false;
  支持 async predicate

验证:
- (cd hub-client && bunx tsc --noEmit && bun test) → 7 pass / 0 fail
- bun test cli.test.ts → 7 pass / 0 fail
- (cd hub-server && bun test) → 356 pass / 0 fail
- (cd forge-engine && bun test) → 32 pass / 0 fail
- bun hub-test-harness/harness.ts → 8/8

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017pGyNVmhbA1Z1vRoxLZbFe
`hub_send_file` to Feishu always failed. The file branch passed an absolute
path to `lark-cli im +messages-send --file`, but lark-cli only accepts a path
relative to the current working directory:

  1.0.24: "--file must be a relative path within the current directory,
           got "/abs/path" (hint: cd to the target directory first)"
  1.0.95: "--file "/abs/path" is outside the built-in allowlist;
           allowed roots are the current working directory ..."

Verified against both versions — 1.0.95 tightened the rule rather than
relaxing it, so an upgrade does not remove the need for this fix.

The same constraint is already handled twice elsewhere in this file: the
voice branch uses `basename(filePath)` + `cwd: dirname(filePath)` (with a
comment explaining why), and `downloadFeishuMedia` sets `cwd` and passes a
relative `--output`. Only the file branch was missed.

Note for anyone debugging this: on failure lark-cli prints a proxy warning
line first; the real validation error is on the next line.

Verified: absolute path -> exit 2 (validation error); basename + cwd ->
exit 0 with a message_id, and an end-to-end `hub_send_file` with an absolute
path now delivers on both lark-cli 1.0.24 and 1.0.95.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FxWdmB6ACyBEKA8wBAbT1K
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant