CI/CD runner for TWD (Test while developing). It executes your in-browser TWD tests in a headless environment. Puppeteer is only used to open the page; all tests run inside the real browser context against real DOM.
- Installation
- Usage: running tests, filtering, configuration
- Run report: the
.twd/report/folder every run writes - Recording: capture a run to video, paced so it is watchable
- Contract Validation: check your mocks against OpenAPI specs
- CI/CD Integration: GitHub Action and custom setups
- Beta features: sharding and layout snapshots
- How It Works
- Requirements
npm install twd-cliOr use directly with npx:
npx twd-cli runRun tests with default configuration:
npx twd-cli runnpx twd-cli --help # the commands
npx twd-cli run --help # every run option
npx twd-cli report --help # render a saved reportHelp prints and exits 0 without launching a browser or reading your config.
A flag the CLI does not know is refused before anything runs, with the closest
match suggested, so a typo cannot quietly run the whole suite:
$ npx twd-cli run --tests "Login"
twd-cli run: unknown option --tests
Did you mean --test?
Run `twd-cli run --help` to see every option.
Run only a subset of tests with the repeatable --test flag. Matching is
case-insensitive and matches a substring of each test's full
"Suite > test name" path:
# Run every test whose name contains "shows error"
npx twd-cli run --test "shows error"
# Because matching uses the full "suite > test" path, passing a describe
# name runs every test inside that describe block:
npx twd-cli run --test "Login"
# Multiple --test flags are combined with OR (a test runs if it matches any):
npx twd-cli run --test "Login" --test "Signup"Notes:
- If no test matches any filter, the run exits with code
1and printsNo tests matched filter(s): …, so a typo won't silently look like a pass. - Code coverage collection is skipped while a
--testfilter is active, since a filtered run is a partial (debug) run.
--changed-since <ref> works out which tests the current branch added or
changed and runs only those. It replaces the "diff, grep for it() titles,
build a --test loop" script that every consumer was writing:
npx twd-cli run --changed-since origin/main
# Most useful with --record: a reviewer watches what the PR built,
# not the whole suite.
npx twd-cli run --record --changed-since origin/mainHow the set is worked out:
git merge-base <ref> HEADfor the base, falling back to<ref>itself, because a branch is not always a descendant of wherever the base has moved to.it()titles on lines the branch added, in*.twd.test.*files only. The tests that already lived in the same file are noise, and pacing makes them expensive to record.- If the diff added no
it()at all, every title in the changed files: a body can change without its title line moving, and recording nothing would be worse than recording a little too much.
it() and it.only() are selected; it.skip(), it.todo() and xit() never
are, since a test that does not run cannot be recorded. Uncommitted and
untracked test files count too, so the test you just wrote is picked up without
committing first.
Notes:
- A branch that changed no tests prints one line and exits
0. An empty result is a normal CI outcome, not a failure, unlike--test, which is an assertion you typed and still exits1when it matches nothing. This is decided before the browser launches, so such a run needs no dev server at all. - It unions with
--testrather than overriding it, so you can add one extra test to a branch's own. - The base branch has to be in the clone.
actions/checkoutdefaults tofetch-depth: 1, which fetches no history; setfetch-depth: 0. The error says so if you forget. - It is a filter, not a recording feature.
--recordis optional.
Create a twd.config.json file in your project root:
{
"url": "http://localhost:5173",
"timeout": 10000,
"coverage": true,
"coverageDir": "./coverage",
"nycOutputDir": "./.nyc_output",
"headless": true,
"puppeteerArgs": ["--no-sandbox", "--disable-setuid-sandbox"],
"retryCount": 2,
"protocolTimeout": 300000,
"maxFailures": 10,
"chunkSize": 10
}| Option | Type | Default | Description |
|---|---|---|---|
url |
string | "http://localhost:5173" |
The URL of your development server |
timeout |
number | 10000 |
Timeout in milliseconds for page load |
coverage |
boolean | true |
Enable/disable code coverage collection |
coverageDir |
string | "./coverage" |
Directory to store coverage reports |
nycOutputDir |
string | "./.nyc_output" |
Directory for NYC output |
headless |
boolean | true |
Run browser in headless mode |
puppeteerArgs |
string[] | ["--no-sandbox", "--disable-setuid-sandbox"] |
Additional Puppeteer launch arguments |
retryCount |
number | 2 |
Number of attempts per test before reporting failure. Set to 1 to disable retries |
protocolTimeout |
number | 300000 |
Puppeteer CDP protocolTimeout in ms (5 min). Tests run in chunks via runByIds, so this bounds a single chunk's browser call (not the entire run). Raise it (e.g. 600000) for slow CI or if individual chunks hang; 0 means no timeout. Defaults above Puppeteer's implicit 180000ms ceiling |
maxFailures |
number | 10 |
Stop the run once this many tests have failed in total; the CLI prints the results gathered so far and exits non-zero. Set 0 to disable and always run every test |
chunkSize |
number | 10 |
How many tests run per browser call. Smaller values make the failure limit and timeouts more granular (less work lost if one chunk hangs); larger values reduce overhead. 0 runs everything in one call |
contracts |
array | none | OpenAPI contract validation specs (see Contract Validation) |
contractReportPath |
string | none | Path to write a markdown report for CI/PR integration |
viewport |
object | { "width": 1280, "height": 800 } |
Browser viewport for every run. Layout snapshots are only reproducible when this is fixed and explicit. While recording, record.viewport wins |
snapshotDir |
string | "__twd_snapshots__" |
Where layout snapshot references and failure captures live. Must match the dir given to the twdSnapshot Vite plugin. See layout snapshots |
record |
object | see below | Video recording settings (see Recording) |
report |
object | { "dir": ".twd/report", "formats": ["html", "markdown"] } |
Run report folder and views; false disables it |
Partial Results on Timeout or Crash: Tests run in chunks (controlled by chunkSize), so on a protocolTimeout or unexpected crash mid-run, results from completed chunks are printed instead of being lost entirely.
Every npx twd-cli run writes a report folder, .twd/report/ by default:
.twd/report/
run.json # machine-readable result
index.html # open in a browser: results, failures, recordings, layout snapshots
summary.md # for PRs, GitHub Step Summaries, and CI logs
recordings/ # video clips, when --record is set
snapshots/ # layout snapshot captures, for a run with a failure
The last line of a run points at it:
Report: .twd/report/index.html
An AI agent or script should read run.json rather than parse console output:
outcome ("passed", "failed", or "interrupted"), summary (pass/fail/skip
counts), and tests[].error for what broke.
Print a saved report to stdout, useful for a CI job summary:
npx twd-cli report --format markdown >> "$GITHUB_STEP_SUMMARY"--format accepts markdown (default), html, or json.
{
"report": {
"dir": ".twd/report",
"formats": ["html", "markdown"]
}
}Set "report": false to disable it entirely. Flags override the config for a
single run:
npx twd-cli run --report-dir ./ci-report # write it elsewhere
npx twd-cli run --no-report # skip it for this runAdd .twd/ to your project's .gitignore. The folder is rewritten on every run.
Record a run to a video file, for a PR attachment, a docs clip, or a demo:
npx twd-cli run --record --test "checkout flow"Requires ffmpeg 8 or newer on your PATH, or record.ffmpegPath set. Older builds are checked and rejected before the browser launches, with the reason. See Requirements.
Runs are paced at 300ms by default, so --record on its own produces something watchable rather than a one second blur. Pacing slows the run itself rather than stretching the video, so unlike --record-speed it costs no frame rate. It needs twd-js 1.9.0 or newer; on an older version the run still records, unpaced, with a warning.
npx twd-cli run --record --record-pace 500 --test "checkout flow" # slower
npx twd-cli run --record --record-pace 0 --test "checkout flow" # no pacingWhen several tests match, each one is recorded to its own clip, named after its
suite > test path. A reviewer watches the criterion they doubt instead of
scrubbing a single file for it.
One clip for the whole run is still what you get from a single matched test, from
record.filename (one name cannot address several clips), and from more matched
tests than record.maxClips (default 20, set 0 to disable). The run says which
of those applied.
--test matches a substring of the full "suite > test" path, so one filter can match several tests; re-running overwrites existing clips.
mp4 recordings are converted to H.264 / yuv420p once the run ends, so they open in QuickTime, Preview and every browser, and land at roughly a quarter of the size. If your ffmpeg has no libx264 the original is kept and you get a warning; that file is VP9 and plays only in Chrome or VLC.
The recording viewport is 1280x1600 by default, deliberately taller than a screen. Puppeteer captures exactly the viewport, with no scrolling and no letterboxing, so anything below the fold is simply absent from the video and nothing in the run says so. A short default silently cropped the very content the tests asserted on. Set record.viewport if your app is shorter and you would rather not record empty space.
A recorded run is a demo artifact, not a substitute for a CI run. It sets its own viewport (1280x1600, versus the 1280x800 a normal run uses), reflows the app to full width, and pacing inserts real delays that can mask race conditions. Run CI unrecorded and record separately.
Flags: --record, --record-dir <path>, --record-speed <n>, --record-pace <ms>. Everything else lives under record in twd.config.json.
| Option | Type | Default | Description |
|---|---|---|---|
enabled |
boolean | false |
Turn recording on. Same as --record |
dir |
string | <report dir>/recordings |
Where the video is written |
filename |
string | null | null |
Explicit name. When null, derived from the recorded tests. Setting it also records the whole run to one clip, since one name cannot address several |
maxClips |
number | 20 |
Most clips one run splits into. Past it the whole run goes to a single file. 0 disables the bound |
format |
string | "mp4" |
"mp4" (converted to H.264 after the run), "webm" or "gif" |
viewport |
object | 1280x1600 |
Applied only when recording. width and height set the video dimensions. Tall on purpose: what is below the fold is not in the video. Keep both even: the H.264 conversion needs it |
fps |
number | 30 |
Capture frame rate |
speed |
number | 1 |
Post-hoc playback speed. Costs frame rate, prefer pace |
pace |
number | 300 |
Milliseconds held after each command. 0 disables |
preRoll |
number | 0 |
Milliseconds held on the opening state |
postRoll |
number | 500 |
Milliseconds held on the final state. Without it the last thing your test did never appears in the video |
hideSidebar |
boolean | true |
Hide the TWD sidebar so the frame is just your app |
ffmpegPath |
string | "ffmpeg" |
Path to the binary if it is not on your PATH |
Full explanations, including why postRoll is on by default and the measured frame rate cost of speed, are in the Recording Runs docs.
Important: Puppeteer is not used as a testing framework here. It simply provides a headless browser to load your application, the same way a user would open Chrome. Once the page loads, all test execution happens inside the real browser context through the TWD runner. Your tests interact with real DOM, real components, and real browser APIs. Puppeteer just opens the door and gets out of the way.
Contract Validation: Mock overlaps are automatically handled: if multiple tests or calls use the same alias but with different HTTP methods/URLs/statuses, all are validated separately (no silent drops).
- Launches a headless browser via Puppeteer (the only thing Puppeteer does)
- Navigates to your dev server URL
- Waits for the app and TWD sidebar to be ready
- TWD's in-browser test runner executes all tests against the real DOM
- Collects and reports test results
- Validates collected mocks against OpenAPI contracts (if configured)
- Optionally collects code coverage data
- Exits with appropriate code (0 for success, 1 for failures)
The easiest way to run TWD tests in CI. Handles Puppeteer caching, Chrome installation, and optional contract report posting in a single step:
name: TWD Tests
on:
push:
branches: [main]
pull_request:
branches: [main]
permissions:
pull-requests: write # only needed if using contract-report
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- uses: actions/setup-node@v5
with:
node-version: 24
cache: npm
- name: Install dependencies
run: npm ci
- name: Install mock service worker
run: npx twd-js init public --save
- name: Start dev server
run: |
nohup npm run dev > /dev/null 2>&1 &
npx wait-on http://localhost:5173
- name: Run TWD tests
uses: BRIKEV/twd-cli/.github/actions/run@main
with:
contract-report: 'true'| Input | Default | Description |
|---|---|---|
working-directory |
. |
Directory where twd.config.json lives |
contract-report |
false |
Post contract validation summary as a PR comment |
shard |
(empty) | Run one shard of the suite, as <index>/<total> (e.g. 2/4). Leave empty to run everything in one job. Beta, see docs/sharding.md |
report-dir |
(empty) | Where the run report folder is written. Empty uses report.dir from twd.config.json, or .twd/report if that isn't set either |
upload-report |
true |
Upload the report folder as an artifact named twd-report (twd-report-<index> for a shard) |
The action runs in the same job, so coverage data is available for subsequent steps:
- name: Run TWD tests
uses: BRIKEV/twd-cli/.github/actions/run@main
- name: Display coverage
run: npm run collect:coverage:textThe sibling of the run action, for clips rather than results. It installs a
known-good ffmpeg, records, and uploads the result:
- uses: BRIKEV/twd-cli/.github/actions/record@main
with:
changed-since: ${{ github.event.pull_request.base.sha }}changed-since is what keeps the clip watchable: it records only the tests the
branch touched, rather than the whole suite. See
Running only what this branch changed.
| Input | Default | Description |
|---|---|---|
working-directory |
. |
Directory where twd.config.json lives |
changed-since |
(empty) | A ref. Records only the tests changed since it. Needs fetch-depth: 0 on checkout. Mutually exclusive with tests |
tests |
(empty) | Newline-separated test titles, one --test each. Mutually exclusive with changed-since |
pace |
(empty) | Passed to --record-pace. Empty uses the CLI default of 300; 0 disables pacing |
install-ffmpeg |
true |
Install ffmpeg 8.x. Set false to use whatever is on PATH |
upload-artifact |
true |
Upload the clips as an artifact |
artifact-name |
twd-recording |
Name of the artifact |
retention-days |
14 |
How long to keep it |
| Output | Description |
|---|---|
clip-count |
Number of clips written. One per test when several tests match, one for the whole run when they do not (a single test, an explicit record.filename, or more tests than record.maxClips). 0 is a valid, non-failing result: a branch that changed no tests has nothing to record |
dir |
Where the clips are, for a caller that wants to do its own upload |
artifact-url |
URL of the artifact, when the action uploaded it |
Because the distro build is not good enough, and finding that out the hard way is
expensive. Puppeteer's screencast passes -movflags hybrid_fragmented, which
arrived after ffmpeg 7, and apt-get install ffmpeg on ubuntu-24.04 gets you
6.1.1, which rejects it. The action installs an 8.1.x build whose gpl variant
also carries libx264, which the H.264 conversion needs. Set
install-ffmpeg: false if you manage your own; twd-cli checks the binary can
actually do the job before it launches a browser either way.
Only Linux runners get the bundled build. On macOS or Windows the step warns and skips, so install ffmpeg 8+ yourself there.
Recording is triggered by a label here, but that part is policy: record every PR
to main if you prefer. The trigger, the PR comment and the dev server stay in
your workflow rather than the action, exactly as they do for run:
name: Record a PR's tests
on:
pull_request:
types: [labeled]
jobs:
record:
if: github.event.label.name == 'record'
runs-on: ubuntu-latest
timeout-minutes: 15 # a hung recording must not cost the whole job
permissions: { contents: read, pull-requests: write }
steps:
- uses: actions/checkout@v5
with:
ref: ${{ github.event.pull_request.head.sha }}
fetch-depth: 0 # --changed-since needs history
- uses: actions/setup-node@v5
with: { node-version: 24, cache: npm }
- run: npm ci
- run: |
nohup npm run dev > vite.log 2>&1 &
npx wait-on http://localhost:5173
- uses: BRIKEV/twd-cli/.github/actions/record@main
id: rec
continue-on-error: true # a clip is optional; the PR it describes is not
with:
changed-since: ${{ github.event.pull_request.base.sha }}
- if: steps.rec.outputs.clip-count != '0'
run: gh pr comment "$PR" --body "${{ steps.rec.outputs.clip-count }} clip(s): ${{ steps.rec.outputs.artifact-url }}"
env:
GH_TOKEN: ${{ github.token }}
PR: ${{ github.event.pull_request.number }}The timeout-minutes and continue-on-error are belt, not workaround. A
recording is always optional; the pull request it describes is not.
If you prefer full control, set up each step manually. Puppeteer 24+ no longer auto-downloads Chrome, so you need to install it explicitly:
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- uses: actions/setup-node@v5
with:
node-version: 24
cache: npm
- name: Install dependencies
run: npm ci
- name: Install mock service worker
run: npx twd-js init public --save
- name: Start dev server
run: |
nohup npm run dev > /dev/null 2>&1 &
npx wait-on http://localhost:5173
- name: Cache Puppeteer browsers
uses: actions/cache@v4
with:
path: ~/.cache/puppeteer
key: ${{ runner.os }}-puppeteer-${{ hashFiles('package-lock.json') }}
restore-keys: |
${{ runner.os }}-puppeteer-
- name: Install Chrome for Puppeteer
run: npx puppeteer browsers install chrome
- name: Run TWD tests
run: npx twd-cli run
- name: Display coverage
run: npm run collect:coverage:textValidate your test mocks against OpenAPI specs to catch drift between your mocks and the real API. When a mock response doesn't match the spec, you'll see errors like:
Source: ./contracts/users-3.0.json ERROR
✓ GET /users (200) — mock "getUsers" — in "UserList > should display all users"
✗ GET /users/{userId} (200) — mock "getUserBadAddress" — in "UserDetails > should fetch user details"
→ response.address.city: missing required property
→ response.address.country: missing required property
⚠ GET /users/{userId} (404) — mock "getUserNotFound" 2nd time — in "UserDetails > should show not found"
Status 404 not documented for GET /users/{userId}
- Add your OpenAPI specs to the project (JSON format, 3.0 or 3.1):
contracts/
users-3.0.json
posts-3.1.json
- Configure contracts in
twd.config.json:
{
"url": "http://localhost:5173",
"contracts": [
{
"source": "./contracts/users-3.0.json",
"baseUrl": "/api",
"mode": "error",
"strict": true
},
{
"source": "./contracts/posts-3.1.json",
"baseUrl": "/api",
"mode": "warn",
"strict": true
}
]
}contractReportPath is deprecated and will be removed. Contract results
now appear in the run report's summary.md automatically.
| Option | Type | Default | Description |
|---|---|---|---|
source |
string | none | Path to the OpenAPI spec file (JSON) |
baseUrl |
string | "/" |
Base URL prefix to strip when matching mock URLs to spec paths |
mode |
"error" | "warn" |
"warn" |
error fails the test run, warn reports but doesn't fail |
strict |
boolean | true |
When true, rejects unexpected properties not defined in the spec |
The validator checks all standard OpenAPI/JSON Schema constraints:
- Types:
string,number,integer,boolean,array,object - String:
minLength,maxLength,pattern,format(date, date-time, email, uuid, uri, hostname, ipv4, ipv6) - Number/Integer:
minimum,maximum,exclusiveMinimum,exclusiveMaximum,multipleOf - Array:
minItems,maxItems,uniqueItems - Object:
required,additionalProperties - Composition:
oneOf,anyOf,allOf - Enum: validates against allowed values
- Nullable: supports both OpenAPI 3.0 (
nullable: true) and 3.1 (type: ["string", "null"])
When contractReportPath is set and you use the action with contract-report: 'true', a summary table is posted as a PR comment:
| Spec | Passed | Failed | Warnings | Mode |
|---|---|---|---|---|
users-3.0.json |
2 | 3 | 1 | error |
posts-3.1.json |
2 | 2 | 0 | warn |
Failed validations are included in a collapsible details section with a link to the full CI log.
These work, but their behaviour may still change between minor versions.
- Sharding across CI jobs: split a long run across parallel jobs and merge the reports back into one.
- Layout snapshots: fail a test when the page geometry moves, with
--update-snapshotsand--ci.
-
Node.js >= 20.19.x
-
A running development server with TWD tests
-
ffmpeg 8 or newer, only for
--record. Install withbrew install ffmpeg(macOS),sudo apt-get install ffmpeg(Linux), orwinget install ffmpeg(Windows). Setrecord.ffmpegPathif it is not on yourPATH.The version matters, and "not the distro build" is not enough. Puppeteer's screencast passes
-movflags hybrid_fragmented, which arrived after ffmpeg 7:ffmpeg works 6.1.1 (Ubuntu 24.04) no 7.0.2 (johnvansickle static) no 8.1.2 yes On Ubuntu CI runners, install a build of 8.x rather than the packaged one.
twd-cliprobes the capability, not the version number, so a future ffmpeg that drops the flag is caught too.