Conversation
scheduleOrigin 建定时器时只过 canScheduleToday(仅查 start_date/end_date), 未过 shouldFire,导致两个问题: 1. weekdays/days/months 不参与判定——「每周日」的任务会在周六被 报成「今天错过」 2. delay <= 0 一律落进 missed 分支——午夜重排本身发生在 00:00, 重排那一刻就把 00:00 的格子算成已过去,该时段任务永远跑不到 修法: - 新增 FIRE_GRACE_MS(90s)宽限支,刚过点的任务立即补触发而非报错过 - missed 与宽限两支都补 shouldFire(entry) shouldFire 改为 export 以便测试,并补 3 个回归测试。 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016fd3qWDATPyaLxwM8s1Ag9
错过判定原本只看两样:delay 落在 MISSED_WINDOW_MS(2h) 内,以及 shouldFire 的排期规则。它从不查该条目今天是否已经触发过——fire() 只写 global.json 的 last_fire / last_sender / today_count 这类全局标量,回答不了「这一条今天跑没跑」。 于是任务时刻之后 2 小时内的任何一次重排,都会把已经跑完的任务重新报一遍: 启动、配置热加载的部分重排,以及午夜重排在边界上提前几毫秒跑到前一天时 (此时 22:00 的 delay 约为 -1h59m,正落在窗口里)。2026-08-22 / 08-23 / 08-28 / 08-29 连续复发,每次都要人工判真假,判错的代价是重复推送或漏跑。 改动:新增 per-entry 的已触发记录(forge-state/fires.json) - fireKey(entry) 来源文件 + 时刻 + 名字(缺省回退 sender)作稳定标识 - hasFiredToday(state, entry, today) 纯函数,state 由调用方注入 - markFired(state, entry, today) 返回新 state,并清掉非当天的键,文件不增长 - fire() 成功推送后落一条记录;scheduleOrigin 每次重排只读一次状态文件 - 错过分支追加 !hasFiredToday(...) 真实的漏跑不受影响:engine 当时没跑就没有记录,照常上报。 验证:forge-engine 28 tests 全过(新增 6 条覆盖键稳定性、label 缺省回退、 跨日失效、清理旧键、同日幂等);bunx tsc --noEmit 无错; bun hub-test-harness/harness.ts 8/8;fh hub self-test 8/8。 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EWhZ8wD8Ne5yiYaCdL6cjf
engine 的 log() 此前只写 stderr,而 MCP 宿主通常只在连接建立那一刻捕获 stderr——启动之后打印的任何东西都不再被记录。进程一旦异常退出,线索为零。 实测(2026-09-03):engine 在 00:00-03:00 之间消失,宿主日志无断开记录、 系统无崩溃报告、无内存压力事件,调度静默停摆 10 小时 31 分、漏跑 12 个任务, 事后无从判断死因。 改动: - log()/logError() 同时追加到 <DATA_DIR>/engine.log,超 2MB 转存 .log.1 - 新增 logFatal(),带完整堆栈落盘 - 注册 uncaughtException / unhandledRejection:记录后 stopScheduler() 释放 PID 锁并 exit(1) 两处取舍写在注释里: - 抓到异常后退出而非续跑——状态可能已损坏,错误的推送比没有推送更坏 - 不做自动重起——多实例 + 自动重起 + 抢 PID 锁是本项目已踩过的坑; 恢复由外部检测触发人工重连,engine 自己不做进程管理 验证:tsc --noEmit 通过 · bun test 19 pass · 临时 DATA_DIR 实测日志落盘、 FATAL 带堆栈、2MB 轮转生效 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EWhZ8wD8Ne5yiYaCdL6cjf
本地 main 作为唯一运行版本——engine 的 MCP 入口直接指向工作区, checkout 在哪个分支跑的就是哪个版本。分散在功能分支上等于生产环境 随 checkout 漂移(2026-09-04 实测:切到修复分支时 crash-log 当场失效)。 PR LinekForge#46 仍挂在上游等合并,本地先用上,不冲突。
同上:本地 main 收敛为唯一运行版本。PR LinekForge#47 仍挂上游等合并。
issue LinekForge#27 的 PID lock(273bf3a)让第二个实例进入 passive 模式以避免重复触发, 但 acquirePidLock() 只在启动时调用一次。active 实例退出后,已经处于 passive 的 实例不会重新竞选——即使它们还活着、MCP 还连着,调度也静默停摆到下一个新 session 启动为止。 原实现处理了 stale 锁回收,但只在启动路径上:新实例发现旧持锁者已死时接管。 没覆盖的是反过来的顺序——已经在场的 passive 实例发现 active 死了该怎么办。 实测(2026-09-08):三个实例,active 被关闭后剩下两个 passive。此时 `/mcp` 显示 连接正常、进程在、MCP 工具可调用,但 PID 锁文件消失、engine.log 停止输出,全部 定时任务不再触发。故障唯一的可见信号是一个没人会主动去看的锁文件。 改动: - passive 实例启动 30s 间隔的重新竞选定时器,夺锁成功即接管调度 - 抽出 activate(),首次启动与事后接管共用同一段排程逻辑 - stopScheduler() 清理竞选定时器 - 导出 isPassiveMode() / retryPassivePromotion() 供测试直接驱动,不必等 interval acquirePidLock() 本身未改:持锁者仍存活时照旧返回 false。所以接管只可能发生在锁 被释放(正常退出)或变成 stale(崩溃)之后,不会退回 LinekForge#27 的重复触发。 测试 forge-engine/scheduler-pid-lock.test.ts(4 个,含一条防止修过头的反向测试): - 持锁者存活时进入 passive - 持锁者正常退出后自我提升 - 持锁者崩溃留下 stale 锁时自我提升 - 持锁者仍存活时保持 passive,且不夺锁 验证: - (cd forge-engine && bunx tsc --noEmit && bun test) → 23 pass / 0 fail - bun hub-test-harness/harness.ts → 8/8 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017pGyNVmhbA1Z1vRoxLZbFe
飞书把图片以 "[图片: img_xxx]" 的形式塞在 content 文本里。入站处理用
content.match() 取 key——没有 g 标志,只返回第一个匹配——随后把**整个
content 覆盖**成那一张的路径:
const imageKeyMatch = content.match(/\[(?:Image|图片):\s*(img_[^\]]+)\]/);
if (imageKeyMatch && messageId) {
const imageKey = imageKeyMatch[1];
const filePath = await downloadFeishuMedia(messageId, "image", imageKey);
content = filePath ? `[图片] ${filePath}` : `[图片: ${imageKey}]`;
}
于是一条「文字 + 6 张图」的富文本到达 agent 时只剩 "[图片] <第一张路径>",
文字和其余 5 张静默消失。用户侧的表现是「一次发六张图,对面只收到第一张」,
而且没有任何错误或日志——后 5 张根本没被下载过。
改动:
- 抽出 resolveImagePlaceholders(content, download):用 matchAll 遍历全部占位符,
逐个**就地**替换成 "[图片] <路径>",占位符以外的文字原样保留
- 单张下载失败只保留该张的原占位符,不牵连同条消息里其它图片
- download 以参数注入,函数无 IO 依赖,可直接测试
- handleMessage 改用它;检测用不带 g 的常量正则,遍历时另建带 g 的副本,
避免共享 lastIndex 导致漏匹配
测试 hub-server/feishu-media-placeholders.test.ts(6 个):
- 单张占位符替换为路径
- 多张全部替换,且下载按序对每个 key 各调一次
- 图文混排时占位符以外的文字原样保留
- 兼容英文 [Image: ...] 形式
- 某张下载失败时保留该张占位符,其余照常替换
- 不含占位符的纯文字原样返回,且一次下载都不发起
验证:
- (cd hub-server && bun test) → 356 pass / 0 fail
- bun hub-test-harness/harness.ts → 8/8
- bunx tsc --noEmit:改动前后均只有既有的 channels/feishu.ts 缺
@larksuiteoapi/node-sdk 一条,未新增
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017pGyNVmhbA1Z1vRoxLZbFe
两套机制都在"确保 Hub 在运行",互相不知道对方存在:
1. launchd(com.forge-hub.plist,RunAtLoad + KeepAlive)
2. hub-client 的 ensureHubRunning():探测不可达就 spawn(..., {detached:true}) + unref()
第 2 条造出的进程不是 launchd 的子进程。它抢到端口后:
- launchd 每 ThrottleInterval(30s) 试着拉起自己的实例,撞端口崩掉,无限循环,
hub-stderr.log 里刷满 "Failed to start server. Is port 9900 in use?"
- `forge-hub sync` 的 `launchctl bootout` 杀不到它(不归 launchd 管),随后的
bootstrap 同样撞端口失败,而 `log("✓ Hub 已重启")` 是无条件打印的
净效果:**sync 报告重启成功,实际跑的仍是旧代码**。实测一台机器上 Hub 进程连续
存活 4 天 6 小时,跨越多次 sync 从未被替换;`ps -o ppid` 显示它的父进程是某个
Claude Code session 的 hub-channel.ts,不是 launchd。
修复分两处:
**hub-client(根因)**:新增 hub-starter.ts,装了 plist 时走
`launchctl kickstart` 让 launchd 启动,保证 Hub 始终是 launchd 的子进程;
kickstart 失败(service 未 bootstrap)则回退到原来的 detached spawn。
非 macOS 或没装 plist 时行为完全不变。
**cli.ts sync(止损)**:新增 waitUntil(),把"命令被接受"和"状态真的变了"分开。
bootout 之后轮询确认 Hub 真的下线,等不到就明确报告"它多半不是 launchd 启动的、
bootout 管不到"并给出手动步骤,**不再打印已重启**;bootstrap 之后轮询确认 Hub
恢复响应,成功消息改为有条件打印。
测试:
- hub-client/hub-starter.test.ts(4):darwin+plist→launchd;darwin 无 plist→spawn;
linux 有/无 plist 均→spawn
- cli.test.ts 新增 waitUntil(4):立即成立不轮询;中途翻转;超时返回 false;
支持 async predicate
验证:
- (cd hub-client && bunx tsc --noEmit && bun test) → 7 pass / 0 fail
- bun test cli.test.ts → 7 pass / 0 fail
- (cd hub-server && bun test) → 356 pass / 0 fail
- (cd forge-engine && bun test) → 32 pass / 0 fail
- bun hub-test-harness/harness.ts → 8/8
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017pGyNVmhbA1Z1vRoxLZbFe
用户在飞书里选中一条消息点「回复」,lark-cli 1.0.95 的 im.message.receive_v1 事件带 reply_to(父消息 id),但 feishu 通道只取 sender_id / chat_id / message_type / content 四个字段,reply_to 落地即丢。agent 收到的是一条没有 上下文的裸文本,用户只能把被引用的原文再复制一遍。 改法:事件带 reply_to 时用 `lark-cli im +messages-mget` 取回父消息的已渲染正文, 拼成 `[回复 ▸ 前 200 字]\n<用户的话>` 交给 agent;取不回来退化为 `[回复 om_xxx]\n<用户的话>`,引用关系至少可追。raw 里同时带 reply_to。 实测(bot 身份):mget 读自己发的和用户发的消息都能拿到 content。 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HSo9HCw7JBqQUqZhwnqYee
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HSo9HCw7JBqQUqZhwnqYee
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
动机
用户在飞书里选中一条消息点「回复」再打字,agent 收到的只有打的那几个字——被引用的那条没有一起过来。用户要让 agent 知道自己在回什么,只能把原文再复制一遍。
根因:lark-cli(1.0.95)的
im.message.receive_v1事件已经带reply_to(被引用消息的message_id)和root_id(lark-cli event schema im.message.receive_v1可查)。hub-server/channels/feishu-lark-cli.ts的handleMessage只读sender_id/chat_id/message_type/content四个字段,reply_to落地即丢。改动概要
事件带
reply_to时,用lark-cli im +messages-mget --message-ids <id> --as bot取回被引用消息的已渲染正文,拼成交给 agent。
取不回来(超时 / 非 JSON / 无权限)退化为
[回复 om_xxx]\n<用户打的字>,引用关系至少可追,不阻断消息;失败走hub.logError,不静默。raw里同时带reply_to,下游要用原 id 的还能拿到。纯函数
formatReplyContext抽出来,6 个单测(hub-server/feishu-reply-context.test.ts):正常前缀 / 换行压平 / 200 字截断 / 恰好 200 不加省略号 / 取不到退化 id / 用户自己的换行不动。影响范围
hub-server/channels/feishu-lark-cli.ts的入站路径(新增两个函数、handleMessage里三行)和一个新测试文件。reply_to时多一次lark-cli im +messages-mget调用(10 s 超时);不带reply_to的消息行为完全不变。message_id写进 chat-history 是appendHistory签名的事,另开;root_id(话题串)本 PR 不处理。reply_to字段(1.0.95 有;更老版本没有这个字段时代码路径不触发,等价于改动前)。self-test 结果
cd hub-server && bunx tsc --noEmit && bun test→ 0 错、362 pass(含新增 6 例)。bun hub-test-harness/harness.ts→ 8/8。fh hub self-test→ 8/8。grep -rE '@im\.wechat|ou_[a-f0-9]{16,}|sk-ant-…|sk-…|[0-9]{9,10}:[A-Za-z0-9_-]{35}'对本 PR 改动文件 → 无命中;测试里的 message_id 用的是假值。+messages-mget读自己发的和用户发的消息都能拿到content;部署到本机后由用户在飞书引用回复一次,agent 侧收到[回复 ▸ …]前缀。🤖 Generated with Claude Code
https://claude.ai/code/session_01HSo9HCw7JBqQUqZhwnqYee