Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .coderabbit.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
# yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json
reviews:
auto_review:
# Review stacked pull requests too, not only those that target main.
base_branches:
- ".*"
1 change: 1 addition & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@ OMMS (npm `om-memory-system`) is a memory plugin for AI coding agents. One share
Keep these boundaries:

- `src/core/` and `src/services/` must not import `@opencode-ai/*`, `@earendil-works/*`, or `src/adapters/*`.
- `src/importer/` must not import `src/adapters/*`. Each host's history reader lives in `src/importer/`, and adapters import from it.
- `tests/host-neutral-capture-boundary.test.ts` and `tests/pi-adapter-boundary.test.ts` enforce that rule.
- An adapter must not import the other host's adapter modules.
- Load host SDKs and heavy modules with dynamic `import()`. `tests/plugin-bundle-boundary.test.ts` checks the plugin bundle.
Expand Down
55 changes: 55 additions & 0 deletions docs/tdr/012-tag-migration-touches-only-untagged-memories.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,55 @@
# TDR-012: Tag migration touches only untagged memories

**Date:** 2026-09-28
**Status:** Proposed
**Deciders:** OMMS maintainers
**Tags:** web-ui, embeddings, migration

## Context

The web UI's **Memory Tagging Migration** dialog reports how many memories have no tags ("Found 1 memories needing technical tags") and offers **Start Migration**. On a store with 2,236 memories and one untagged memory, the run showed `/2236` and worked through every memory.

### Root Cause Analysis

`handleDetectTagMigration` counted only rows with empty `tags`, but `handleRunTagMigrationBatch` loaded `SELECT * FROM memories` from every project shard and re-embedded each row's content and tags, calling the model only for the untagged ones. Tagged memories got identical vectors back, so no data was lost, but the run took far longer than the dialog implied and ran the local embedding model thousands of times. A second defect: a memory that threw an error did not advance `processed`, so every later batch restarted at the same memory and the run could never finish.

## Decision

The run builds its work list once, when it starts, from `SELECT id FROM memories WHERE tags IS NULL OR tags = ''` in each project shard, so the total matches the dialog's count. Each memory is re-read by id; one that gained tags or was deleted since is passed over. Only a memory that receives tags is re-embedded. Its tags and both vectors are saved in one `UPDATE`, so a failed embedding leaves it untagged and a later run tries it again. A failure is recorded in `errors` and the run always moves on. When a run completes, the next run starts from a fresh list of whatever is still untagged.

## Consequences

### Positive

- The dialog's count and the run's total agree, and tagged memories are never rewritten.
- A failing memory cannot stall the run; it stays untagged for a later run.

### Negative

- A memory whose model call fails stays untagged, so the dialog reappears until a run succeeds.

### Neutral

- Migration state is still held in the web server's memory; restarting the server restarts the run.

## Alternatives Considered

| Option | Rejected Because |
| ----------------------------------- | ---------------------------------------------------------------- |
| Keep re-embedding everything | Wastes time and CPU, and contradicts the dialog's count |
| Select untagged rows on every batch | Rows that fail stay untagged and would be selected again forever |

## How to Recognise / Handle This Again

1. A migration or backfill total is much larger than the count shown before it starts.
2. Compare the detect query with the query the run iterates over.
3. Build the run's list from the same filter as the detect step, and always advance past failures.

## Revisit Triggers

- Tags or tag vectors change format and existing tagged memories need re-embedding on purpose.

## References

- `src/services/api-handlers.ts` (`handleDetectTagMigration`, `handleRunTagMigrationBatch`)
- `tests/tag-migration.test.ts`
1 change: 1 addition & 0 deletions docs/tdr/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ TDRs capture **implementation-level technical decisions** such as platform-speci
| [009](./009-protect-capture-traces-on-windows.md) | Protect capture traces with Windows access-control lists | Accepted | 2026-09-27 |
| [010](./010-match-windows-native-import-source-paths.md) | Match Windows import-source tests to native canonical paths | Proposed | 2026-09-27 |
| [011](./011-use-execfilesync-in-windows-git-wrapper-test.md) | Use execFileSync in the Windows Git wrapper test | Proposed | 2026-09-27 |
| [012](./012-tag-migration-touches-only-untagged-memories.md) | Tag migration touches only untagged memories | Proposed | 2026-09-28 |

## Status values

Expand Down
11 changes: 2 additions & 9 deletions src/adapters/opencode/user-prompt.ts
Original file line number Diff line number Diff line change
@@ -1,15 +1,8 @@
import { isStructuredSummaryPromptMessage } from "../../core/internal-prompt.js";
import { isInternalStructuredSession } from "../../services/ai/opencode-provider.js";
import { userPromptManager } from "../../services/user-prompt/user-prompt-manager.js";

export function isStructuredSummaryPromptMessage(userMessage: string): boolean {
// This is the plugin's own structured-summary or profile-analysis request.
// OpenCode echoes it through chat.message like a normal user message, but
// capturing it would create self-referential memories / an infinite learning loop.
if (userMessage.includes("# User Profile Analysis")) {
return true;
}
return userMessage.includes("Analyze this conversation.") && userMessage.includes('type="skip"');
}
export { isStructuredSummaryPromptMessage };

/** True when a prompt is omms's own internal traffic and must not be recorded or searched. */
export function isInternalPrompt(sessionID: string, userMessage: string): boolean {
Expand Down
2 changes: 1 addition & 1 deletion src/adapters/pi/capture.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ import { captureConversation } from "../../core/capture.js";
import type { CaptureSummaryProvider } from "../../core/host.js";
import { log } from "../../services/logger.js";
import { memoryClient } from "../../services/client.js";
import { extractPiConversation, type PiSessionEntry } from "./conversation.js";
import { extractPiConversation, type PiSessionEntry } from "../../importer/pi-conversation.js";

export interface PiCaptureState {
/** User entry IDs terminally handled by a settled capture for this session. */
Expand Down
2 changes: 1 addition & 1 deletion src/adapters/pi/extension.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ import { getLanguageName } from "../../services/language-detector.js";
import { log } from "../../services/logger.js";
import { memoryClient } from "../../services/client.js";
import { capturePiSettledWorkUnit, createPiCaptureState } from "./capture.js";
import type { PiSessionEntry } from "./conversation.js";
import type { PiSessionEntry } from "../../importer/pi-conversation.js";
import { createPiLiveModels } from "./live-model.js";
import { registerPiHistoryImportCommand } from "./import-command.js";
import { performPiProfileLearning } from "./profile.js";
Expand Down
11 changes: 11 additions & 0 deletions src/core/internal-prompt.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
/**
* True when a prompt is omms's own structured-summary or profile-analysis
* request. Both hosts echo it back like a user message; recording or learning
* from it would create self-referential memories and a learning loop.
*/
export function isStructuredSummaryPromptMessage(userMessage: string): boolean {
if (userMessage.includes("# User Profile Analysis")) {
return true;
}
return userMessage.includes("Analyze this conversation.") && userMessage.includes('type="skip"');
}
2 changes: 1 addition & 1 deletion src/importer/importer.ts
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ import {
extractPiConversationWindows,
type PiConversationWindow,
type PiSessionEntry,
} from "../adapters/pi/conversation.js";
} from "./pi-conversation.js";
import { discoverPiSessions } from "./discovery.js";
import { resolveImportProject } from "./import-project.js";
import { PiImportLedger, importLedgerDbPath, type ImportLedgerRow } from "./ledger.js";
Expand Down
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
import type { CaptureConversation, CaptureToolCall } from "../../core/host.js";
import type { CaptureConversation, CaptureToolCall } from "../core/host.js";

/**
* Minimal structural view of Pi session entries (see Pi session-format docs).
Expand Down
5 changes: 3 additions & 2 deletions src/importer/profile-import.ts
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
import { isInternalPrompt } from "../adapters/opencode/user-prompt.js";
import { isStructuredSummaryPromptMessage } from "../core/internal-prompt.js";
import { analyzeProfile, type ModelPort } from "../core/profile-analysis.js";
import { isFullyPrivate, stripPrivateContent } from "../services/privacy.js";
import { getTags } from "../services/tags.js";
Expand Down Expand Up @@ -77,7 +77,8 @@ export async function importProfileFromHistory(
if (
!prompt ||
isFullyPrivate(unit.userPrompt) ||
isInternalPrompt(session.sessionId, prompt)
// OpenCode history already leaves out omms's own capture sessions by title.
isStructuredSummaryPromptMessage(prompt)
Comment thread
coderabbitai[bot] marked this conversation as resolved.
) {
continue;
}
Expand Down
2 changes: 1 addition & 1 deletion src/importer/session-loader.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
import { SessionManager } from "@earendil-works/pi-coding-agent";
import type { PiSessionEntry } from "../adapters/pi/conversation.js";
import type { PiSessionEntry } from "./pi-conversation.js";

/**
* Loads a Pi session file for import through Pi's own exported session model
Expand Down
154 changes: 80 additions & 74 deletions src/services/api-handlers.ts
Original file line number Diff line number Diff line change
Expand Up @@ -151,7 +151,8 @@ export async function handleListMemories(
tag?: string,
page: number = 1,
pageSize: number = 20,
includePrompts: boolean = true
includePrompts: boolean = true,
keyword?: string
Comment thread
coderabbitai[bot] marked this conversation as resolved.
): Promise<ApiResponse<PaginatedResponse<Memory | any>>> {
try {
await ensureTursoReady();
Expand Down Expand Up @@ -211,11 +212,22 @@ export async function handleListMemories(
};
});

let timeline: any[] = memoriesWithType;
// A keyword filter keeps memories carrying that keyword (case-insensitive)
// and only the prompts linked to them.
const wanted = keyword?.trim().toLowerCase();
const keptMemories = wanted
? memoriesWithType.filter((m) => m.tags.some((t: string) => t.toLowerCase() === wanted))
: memoriesWithType;
const keptIds = new Set(keptMemories.map((m) => m.id));

let timeline: any[] = keptMemories;
if (includePrompts) {
const projectPath = tag ? await getProjectPathFromTag(tag) : undefined;
const prompts = await userPromptManager.getCapturedPrompts(projectPath);
const promptsWithType = prompts.map((p) => ({
const visiblePrompts = wanted
? prompts.filter((p) => p.linkedMemoryId && keptIds.has(p.linkedMemoryId))
: prompts;
const promptsWithType = visiblePrompts.map((p) => ({
type: "prompt",
id: p.id,
sessionId: p.sessionId,
Expand All @@ -224,7 +236,7 @@ export async function handleListMemories(
projectPath: p.projectPath,
linkedMemoryId: p.linkedMemoryId,
}));
timeline = [...memoriesWithType, ...promptsWithType];
timeline = [...keptMemories, ...promptsWithType];
}

const linkedPairs = new Map<string, { memory: any; prompt: any }>();
Expand Down Expand Up @@ -1434,19 +1446,28 @@ interface MigrationProgress {
errors: string[];
}

const migrationProgress: MigrationProgress = {
const idleMigration = (): MigrationProgress => ({
processed: 0,
total: 0,
currentBatch: 0,
totalBatches: 0,
isComplete: true,
errors: [],
};
});

let migrationProgress: MigrationProgress = idleMigration();
/** The untagged memories this migration run covers, fixed when the run starts. */
let migrationQueue: Array<{ id: string; dbPath: string }> = [];

export async function handleGetTagMigrationProgress(): Promise<ApiResponse<MigrationProgress>> {
return { success: true, data: migrationProgress };
}

/**
* Tag the memories that have no tags, a few per request. Only untagged
* memories are touched: tagged ones keep their vectors. A memory that fails
* is recorded and passed over, so one bad row cannot stall the run.
*/
export async function handleRunTagMigrationBatch(
batchSize: number = 5
): Promise<ApiResponse<{ processed: number; total: number; hasMore: boolean }>> {
Expand All @@ -1459,91 +1480,76 @@ export async function handleRunTagMigrationBatch(
iterationTimeout: 30000,
});
const provider = AIProviderFactory.createProvider(CONFIG.memoryProvider, providerConfig);
const projectShards = await tursoShardManager.getAllShards("project", "");

const allMemories: { memory: any; shard: any }[] = [];

for (const shard of projectShards) {
const db = await tursoConnectionManager.getConnection(shard.dbPath);
const memories = await db.all("SELECT * FROM memories");
for (const m of memories) {
allMemories.push({ memory: m, shard });
if (migrationProgress.isComplete) {
migrationQueue = [];
for (const shard of await tursoShardManager.getAllShards("project", "")) {
const db = await tursoConnectionManager.getConnection(shard.dbPath);
const rows = await db.all(
"SELECT id FROM memories WHERE tags IS NULL OR tags = '' ORDER BY id"
);
for (const row of rows) migrationQueue.push({ id: String(row.id), dbPath: shard.dbPath });
}
migrationProgress = {
...idleMigration(),
total: migrationQueue.length,
totalBatches: Math.ceil(migrationQueue.length / batchSize),
isComplete: migrationQueue.length === 0,
};
}

if (migrationProgress.total === 0) {
migrationProgress.total = allMemories.length;
migrationProgress.totalBatches = Math.ceil(allMemories.length / batchSize);
migrationProgress.isComplete = false;
}

const startIdx = migrationProgress.processed;
const endIdx = Math.min(startIdx + batchSize, allMemories.length);

for (let i = startIdx; i < endIdx; i++) {
const item = allMemories[i];
if (!item) continue;
const { memory: m, shard } = item;
const db = await tursoConnectionManager.getConnection(shard.dbPath);

const batch = migrationQueue.slice(
migrationProgress.processed,
migrationProgress.processed + batchSize
);
for (const item of batch) {
try {
let currentTags = m.tags
? m.tags
.split(",")
.map((t: string) => t.trim().toLowerCase())
.filter((t: string) => t)
: [];

if (currentTags.length === 0) {
const prompt = `Generate 2-4 short technical tags for this memory content:\n\n${m.content}\n\nReturn ONLY a comma-separated list of tags.`;
const result = await provider.executeToolCall(
"You are a technical tagger.",
prompt,
{
type: "function",
function: {
name: "save_tags",
description: "Save generated tags",
parameters: {
type: "object",
properties: { tags: { type: "array", items: { type: "string" } } },
required: ["tags"],
},
const db = await tursoConnectionManager.getConnection(item.dbPath);
const m: any = await db.get("SELECT * FROM memories WHERE id = ?", [item.id]);
// Tagged or deleted since the run started: nothing to do.
if (!m || (m.tags && String(m.tags).trim())) continue;
const prompt = `Generate 2-4 short technical tags for this memory content:\n\n${m.content}\n\nReturn ONLY a comma-separated list of tags.`;
const result = await provider.executeToolCall(
"You are a technical tagger.",
prompt,
{
type: "function",
function: {
name: "save_tags",
description: "Save generated tags",
parameters: {
type: "object",
properties: { tags: { type: "array", items: { type: "string" } } },
required: ["tags"],
},
},
`migration_${m.id}`
);
if (result.success && result.data?.tags) {
currentTags = result.data.tags;
await db.run("UPDATE memories SET tags = ? WHERE id = ?", [
currentTags.join(","),
m.id,
]);
}
},
`migration_${m.id}`
);
const tags: string[] = Array.isArray(result.data?.tags)
? result.data.tags.map((t: unknown) => String(t).trim().toLowerCase()).filter(Boolean)
: [];
if (!result.success || tags.length === 0) {
throw new Error(result.error ?? "The model returned no tags");
}

const vector = await embeddingService.embedWithTimeout(m.content, { task: "document" });
const tagsVector = currentTags.length
? await embeddingService.embedWithTimeout(formatTagsForEmbedding(currentTags), {
task: "document",
})
: undefined;
await tursoVectorSearch.updateVector(db, m.id, vector, tagsVector);

migrationProgress.processed++;
const tagsVector = await embeddingService.embedWithTimeout(formatTagsForEmbedding(tags), {
task: "document",
});
// One statement, so a memory never ends up tagged with stale vectors.
await tursoVectorSearch.updateVector(db, m.id, vector, tagsVector, tags.join(","));
} catch (e) {
const errorMsg = String(e);
migrationProgress.errors.push(errorMsg);
log("Migration error for memory", { id: m.id, error: errorMsg });
log("Migration error for memory", { id: item.id, error: errorMsg });
} finally {
migrationProgress.processed++;
}
}

migrationProgress.currentBatch++;
const hasMore = migrationProgress.processed < migrationProgress.total;

if (!hasMore) {
migrationProgress.isComplete = true;
}
if (!hasMore) migrationProgress.isComplete = true;

return {
success: true,
Expand Down
Loading
Loading