Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 0 additions & 27 deletions .github/workflows/campaign-cron.yml

This file was deleted.

45 changes: 0 additions & 45 deletions .github/workflows/full-cron.yml

This file was deleted.

29 changes: 0 additions & 29 deletions .github/workflows/rss-cron.yml

This file was deleted.

11 changes: 6 additions & 5 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,15 +109,16 @@ Object storage holds generated/imported image assets. Database rows retain owner

Production heavy work runs in the private Cloud repository's persistent worker process. It polls every five seconds and drains at most one campaign item, two SEO jobs, and one deferred-image job per cycle using the existing PostgreSQL claims, retries, stale recovery, and terminal states. The worker runs `server/src/worker.ts` with `BACKGROUND_EXECUTION_MODE=worker`; the API runs `inline`. No Redis or external queue is involved.

Thin scheduled triggers remain for feed processing, daily fallback work, and manual recovery:
The worker also owns periodic work; one thin external trigger remains as a fallback:

| Trigger | Schedule | Work |
| --- | --- | --- |
| Persistent worker | Every 5 seconds | Campaigns, SEO metadata, deferred images |
| Persistent worker | Every 6 hours and on start | Due RSS feeds |
| Persistent worker | Daily and on start | Google indexing, Search Console, expired operation events |
| Cloudflare Worker | Every 6 hours | Protected bounded fallback drain |
| GitHub `rss-cron.yml` | Every 6 hours | Due RSS feeds |
| GitHub `full-cron.yml` | Daily | Campaigns, indexing, feeds, images |
| GitHub `campaign-cron.yml` | Manual | Bounded campaign drain |

Self-hosted Compose and Dokploy installs keep their own `scheduler` container calling the all-task drain.

The backend decides eligibility and claims work. Schedulers must stay thin; do not create a second queue or duplicate job classification in a Worker or workflow.

Expand All @@ -129,7 +130,7 @@ The backend decides eligibility and claims work. Schedulers must stay thin; do n
| Private `BlogFactoryHQ/blogfactory-cloud` | Cloud-only Compose, Caddy, image-build, and deployment ownership |
| Hetzner Nuremberg | PostgreSQL 18, API, worker, backup, and web container runtime |
| Cloudflare R2 EU | Private production image/object storage and encrypted portable backup |
| GitHub Actions + GHCR | Validation and immutable API/web image builds by commit SHA |
| GitHub Actions + GHCR | Validation and immutable API/web image builds by commit SHA; no scheduled production drains |
| Vercel project `editorial-flow-main` | Last clean serverless deployment retained as the DNS rollback target |

The pre-cutover Neon database and manual snapshot remain available only as a
Expand Down
4 changes: 2 additions & 2 deletions docs/operations.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,11 +51,11 @@ Application rollback is redeploying the previous approved image digests. Full-ho

## Background work

The persistent worker runs bounded campaign, SEO, and deferred-image drains every five seconds. GitHub Actions runs RSS every six hours, the full background matrix (including Search Console) daily, and a campaign drain on manual dispatch; the Cloudflare Worker remains a six-hour protected fallback trigger. Every external trigger calls the existing protected cron endpoint and shares `CRON_SECRET`; see the [RSS scheduler guide](rss-scheduler.md).
The persistent worker runs bounded campaign, SEO, and deferred-image drains every five seconds. The same worker starts the due-RSS-feed tick every six hours and the daily indexing, Search Console, and expired operation-event drain; both also run once when the worker starts, and the backend's due checks and claims keep that safe. The Cloudflare Worker remains a six-hour protected fallback trigger for campaign, SEO, and image drains through the existing cron endpoint and `CRON_SECRET`. The public repository has no scheduled production workflows. See the [RSS scheduler guide](rss-scheduler.md).

The API runs `BACKGROUND_EXECUTION_MODE=inline`; the persistent worker runs `BACKGROUND_EXECUTION_MODE=worker` (`server/src/worker.ts`). `BACKGROUND_WORKER_POLL_MS` and the `BACKGROUND_WORKER_CAMPAIGN_ITEMS`/`SEO_JOBS`/`IMAGE_JOBS` bounds default to the 5-second 1/2/1 cycle above. PostgreSQL atomic claims, stale recovery, retries, and feed leases prevent duplicate ownership. Do not run more than one persistent worker until claim and stale-recovery checks pass for that topology.

The existing all-task drain also removes expired `operation_events`. Do not create a separate retention cron. Operation events expire after 30 days.
The worker's daily drain and the existing all-task drain remove expired `operation_events`. Do not create a separate retention cron. Operation events expire after 30 days.

Do not disable a failing scheduled workflow to make Actions appear clean. Confirm the affected task, timeout, and backend behavior before a narrow fix.

Expand Down
32 changes: 12 additions & 20 deletions docs/rss-scheduler.md
Original file line number Diff line number Diff line change
@@ -1,18 +1,20 @@
# RSS Scheduler

BlogFactory uses one protected cron endpoint:
BlogFactory Cloud's persistent worker (`server/src/worker.ts`) starts the RSS
tick every 6 hours and on worker start, plus a daily indexing, Search Console,
and operation-event retention drain. The public repository has no scheduled
production workflows.

Self-hosted installs call one protected cron endpoint from their `scheduler`
container (Compose, Dokploy) or cron service (Railway):

```text
GET /api/cron/drain?task=feeds
Authorization: Bearer $CRON_SECRET
```

Vercel cron is not used.

```text
Cloudflare Worker Cron: campaign, SEO, and image fallback every 6 hours
GitHub Actions: RSS every 6 hours, full background drain daily
```
Vercel cron is not used. Cloudflare Worker Cron remains a campaign, SEO, and
image fallback every 6 hours.

Cloudflare files:

Expand All @@ -21,26 +23,16 @@ wrangler.cron.jsonc
cloudflare/cron-worker.ts
```

GitHub files:

```text
.github/workflows/rss-cron.yml
.github/workflows/full-cron.yml
.github/workflows/campaign-cron.yml
```

Required secrets:
Required secret:

```text
Cloudflare CRON_SECRET = same value as backend CRON_SECRET
GitHub BLOGFACTORY_CRON_SECRET = same value as backend CRON_SECRET
```

Optional variables:
Optional variable:

```text
Cloudflare CRON_BASE_URL = https://app.blogfactory.io
GitHub BLOGFACTORY_BASE_URL = https://app.blogfactory.io
```

The app still decides which feeds are due from
Expand All @@ -54,7 +46,7 @@ RSS_CRON_MAX_POSTS_PER_FEED=1
RSS_FEED_RUN_LEASE_MINUTES=15
```

Raise those only if cron runs finish comfortably under the backend function limit.
The worker and the cron endpoint read the same knobs. Raise them only if runs finish comfortably.

Each feed run is claimed atomically in PostgreSQL before generation starts. Manual
and scheduled runs share the same claim, so only one batch can run for a feed at a
Expand Down
41 changes: 40 additions & 1 deletion server/src/services/background-worker.self-test.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
import assert from "node:assert/strict";
import { readBackgroundWorkerConfig, runBackgroundWorkerCycle } from "./background-worker.js";
import {
createPeriodicScheduler,
DAILY_INTERVAL_MS,
FEEDS_INTERVAL_MS,
readBackgroundWorkerConfig,
runBackgroundWorkerCycle,
} from "./background-worker.js";

assert.deepEqual(readBackgroundWorkerConfig({}), {
pollMs: 5_000,
Expand Down Expand Up @@ -40,4 +46,37 @@ assert.deepEqual(await runBackgroundWorkerCycle(config, {
images: async () => {},
}), { ok: false, failed: ["seo"] });

const settle = () => new Promise((resolve) => setTimeout(resolve, 0));
let clock = 0;
let releaseDaily = () => {};
const periodicRuns: string[] = [];
const periodic = createPeriodicScheduler([
{ name: "feeds", intervalMs: FEEDS_INTERVAL_MS, run: async () => { periodicRuns.push("feeds"); } },
{
name: "daily",
intervalMs: DAILY_INTERVAL_MS,
run: () => new Promise<void>((resolve) => { periodicRuns.push("daily"); releaseDaily = resolve; }),
},
], () => clock);
assert.deepEqual(periodic.tick(), ["feeds", "daily"], "periodic tasks are due on worker start");
await settle();
clock = 5_000;
assert.deepEqual(periodic.tick(), [], "periodic tasks wait for their interval");
clock = FEEDS_INTERVAL_MS;
assert.deepEqual(periodic.tick(), ["feeds"], "feeds run every six hours");
clock = DAILY_INTERVAL_MS;
await settle();
assert.deepEqual(periodic.tick(), ["feeds"], "an unfinished daily run is never started twice");
releaseDaily();
await settle();
assert.deepEqual(periodic.tick(), ["daily"]);
assert.deepEqual(periodicRuns, ["feeds", "daily", "feeds", "feeds", "daily"]);

const failing = createPeriodicScheduler([
{ name: "feeds", intervalMs: FEEDS_INTERVAL_MS, run: async () => { throw new Error("hidden provider detail"); } },
], () => 0);
assert.deepEqual(failing.tick(), ["feeds"]);
await settle();
assert.deepEqual(failing.tick(), [], "a failed task waits for its next interval");

console.log("background worker self-test passed");
81 changes: 81 additions & 0 deletions server/src/services/background-worker.ts
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,16 @@ export type BackgroundWorkerConfig = {
heartbeatUrl?: string;
};

// Feed ticks stay on the fixed six-hour cadence; retention and Search Console work is daily.
export const FEEDS_INTERVAL_MS = 6 * 60 * 60 * 1000;
export const DAILY_INTERVAL_MS = 24 * 60 * 60 * 1000;

export type PeriodicTask = {
name: string;
intervalMs: number;
run(): Promise<unknown>;
};

type WorkerDrains = {
campaigns(maxCampaigns: number, maxItemsPerCampaign: number): Promise<unknown>;
seo(userId: undefined, limit: number): Promise<unknown>;
Expand Down Expand Up @@ -54,6 +64,73 @@ export async function runBackgroundWorkerCycle(config: BackgroundWorkerConfig, d
};
}

// Starts due periodic tasks without blocking the five-second drain cycle.
// Tasks are due on worker start; the backend's own due checks and claims keep that safe.
export function createPeriodicScheduler(tasks: PeriodicTask[], now: () => number = Date.now) {
const state = tasks.map((task) => ({ task, lastStartedAt: Number.NEGATIVE_INFINITY, running: false }));
return {
tick() {
const started: string[] = [];
for (const entry of state) {
if (entry.running || now() - entry.lastStartedAt < entry.task.intervalMs) continue;
entry.running = true;
entry.lastStartedAt = now();
started.push(entry.task.name);
void entry.task.run()
.catch(() => console.error("[worker] Scheduled task failed", { task: entry.task.name }))
.finally(() => { entry.running = false; });
}
return started;
},
};
}

async function loadPeriodicTasks(env: Record<string, string | undefined>): Promise<PeriodicTask[]> {
const [
{ runScheduler },
{ drainQueuedGoogleIndexing },
{ drainSearchConsoleSync },
{ purgeExpiredOperationEvents },
{ readCronDrainConfig },
] = await Promise.all([
import("./scheduler.js"),
import("./indexing.js"),
import("./search-console.js"),
import("./operation-events.js"),
import("../routes/cron.js"),
]);
const config = readCronDrainConfig(() => undefined, env);
const daily = async (name: string, run: () => Promise<unknown>) => {
await run().catch((error) => {
console.error("[worker] Daily drain failed", { task: name });
throw error;
});
};
return [
{
name: "feeds",
intervalMs: FEEDS_INTERVAL_MS,
run: () => runScheduler(undefined, {
awaitGeneration: true,
maxFeeds: config.feeds.maxFeeds,
maxPostsPerFeed: config.feeds.maxPostsPerFeed,
}),
},
{
name: "daily",
intervalMs: DAILY_INTERVAL_MS,
run: async () => {
const results = await Promise.allSettled([
daily("indexing", () => drainQueuedGoogleIndexing(config.indexing.limit)),
daily("search-console", () => drainSearchConsoleSync(config.searchConsole.limit)),
daily("operation-events", () => purgeExpiredOperationEvents()),
]);
if (results.some((result) => result.status === "rejected")) throw new Error("Daily drain failed");
},
},
];
}

async function pingHeartbeat(url: string) {
const response = await fetch(url, { signal: AbortSignal.timeout(5_000) });
if (!response.ok) throw new Error(`HTTP ${response.status}`);
Expand All @@ -71,7 +148,11 @@ export async function runBackgroundWorker(
imageJobs: config.imageJobs,
});

const periodic = createPeriodicScheduler(await loadPeriodicTasks(env));

while (!signal?.aborted) {
const started = periodic.tick();
if (started.length) console.info("[worker] Scheduled tasks started", { tasks: started });
const cycle = await runBackgroundWorkerCycle(config);
await writeFile(config.heartbeatFile, new Date().toISOString());
if (!cycle.ok) console.error("[worker] Drain failed", { tasks: cycle.failed });
Expand Down
Loading