Skip to content

Add vocabulary system for domain-specific transcription priming and correction - #8

Merged
reecemiao merged 3 commits into
mainfrom
claude/google-drive-access-check-0rsfo5
Aug 24, 2026
Merged

reecemiao merged 3 commits into
mainfrom
claude/google-drive-access-check-0rsfo5

Conversation

@reecemiao

Copy link
Copy Markdown
Owner

This PR introduces a comprehensive vocabulary system that improves transcription accuracy for domain-specific content by priming Whisper with expected terms and automatically correcting known misrecognitions.

Summary

The changes add three main components:

  1. Vocabulary module (vocabulary.py): Manages domain-specific terms and corrections

    • Parses vocabulary files with two line types: plain terms (for priming) and wrong => right corrections
    • Builds optimized prompts combining a language seed, video metadata (title/description), and ranked terms
    • Includes a built-in vocabulary for Mandarin finance channels (zh-finance.txt)
    • Respects Whisper's 224-token prompt limit with smart term selection
  2. Polish module (polish.py): Post-processing for transcription cleanup

    • Normalizes punctuation (ASCII → fullwidth for Chinese text)
    • Detects and collapses repeated phrases (common Whisper artifact)
    • Filters boilerplate subtitle credits
    • Optionally converts traditional → simplified characters (requires OpenCC)
    • Applies vocabulary corrections to final text
  3. CLI enhancements:

    • --vocabulary flag on run command to specify terms/corrections file
    • New polish subcommand to re-process existing scripts with updated vocabularies
    • Supports both built-in names ("zh-finance") and file paths

Key Changes

  • Config: Added vocabulary, prompt_from_metadata, whisper_condition_on_previous_text, polish, and convert_to_simplified settings
  • Pipeline: Integrates vocabulary loading, prompt generation per-video, and segment polishing
  • Transcribers: Updated base protocol and implementations to accept optional prompt parameter
  • Tests: Comprehensive test coverage for vocabulary parsing, prompt building, correction application, and polishing operations

Implementation Details

  • Corrections use regex patterns with word-boundary awareness for ASCII terms (prevents "CDS" matching inside "CDSX") while allowing substring matches for Chinese
  • Prompt building prioritizes terms mentioned in video metadata to stay within token budget
  • Loop detection requires 3+ identical consecutive segments to avoid false positives
  • Boilerplate detection uses regex patterns for common subtitle credit formats
  • OpenCC integration is optional; gracefully degrades with a warning if not installed

https://claude.ai/code/session_01GARmHk71niEh1TPunhH3Z8

claude added 3 commits August 24, 2026 14:27
…ears

Reading the 83 scripts already in Drive, the errors are not random: they are the
words a general model has no reason to expect. 费半 comes out as 飞班 or 肺瓣,
对冲基金 as 对中基金 (10+ scripts), 资产负债表 as 自然负债表 (6), 每股收益 as
美股收益, SNDK as S&DK, ISRG as ISIG, and Chinese sentences are punctuated with
ASCII commas. Three layers, all of them cheap:

- prompt_from_metadata (on) puts the video's title and the first line of its
  description in front of the model. The title already names the day's tickers.
- vocabulary adds the terms a channel says every episode and rewrites what still
  comes out wrong. zh-finance ships for a Mandarin US-market channel; ASCII
  rewrites only match whole words, so CTS => CDS leaves CTSX alone.
- polish (on) gives Chinese sentences fullwidth punctuation while leaving the
  commas in 1,250 alone, keeps one copy of a looped phrase, and drops subtitle
  boilerplate.

whisper_condition_on_previous_text now defaults to false. Whisper's own default
feeds each clip the previous clip's text, which is what makes it loop; batched
decoding ignores the flag anyway, so off also keeps the batched path and the
out-of-memory retry sounding the same.

convert_to_simplified runs OpenCC over the text when the zh extra is installed,
and warns rather than failing when it is not.

`ytscript polish PATH...` re-applies the clean-up to scripts already written, so
adding a term does not mean re-transcribing a backlog.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GARmHk71niEh1TPunhH3Z8
The zh extra went into pyproject.toml without a matching uv.lock, which `uv
lock --check` catches in the lint job — local runs pass `--frozen` and skip it.
Add the extra to the extras matrix too, so its lazily imported dependency is
checked the way local, openai and drive are.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GARmHk71niEh1TPunhH3Z8
load_vocabulary told a built-in name from a file path by looking for a forward
slash, which a Windows path does not contain: vocabulary = "C:\scripts\mine.txt"
resolved to the built-in directory and failed with "no vocabulary". A built-in
is now a bare stem — no directory separator of either kind, no suffix — and a
file of that name is tried either way, so a built-in never shadows one.

Both separators are checked on every platform, so the unit test covering it
fails everywhere rather than only on the Windows runner.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GARmHk71niEh1TPunhH3Z8
@reecemiao
reecemiao merged commit f3bfb4d into main Aug 24, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants