Add vocabulary system for domain-specific transcription priming and correction - #8
Merged
Merged
Conversation
…ears Reading the 83 scripts already in Drive, the errors are not random: they are the words a general model has no reason to expect. 费半 comes out as 飞班 or 肺瓣, 对冲基金 as 对中基金 (10+ scripts), 资产负债表 as 自然负债表 (6), 每股收益 as 美股收益, SNDK as S&DK, ISRG as ISIG, and Chinese sentences are punctuated with ASCII commas. Three layers, all of them cheap: - prompt_from_metadata (on) puts the video's title and the first line of its description in front of the model. The title already names the day's tickers. - vocabulary adds the terms a channel says every episode and rewrites what still comes out wrong. zh-finance ships for a Mandarin US-market channel; ASCII rewrites only match whole words, so CTS => CDS leaves CTSX alone. - polish (on) gives Chinese sentences fullwidth punctuation while leaving the commas in 1,250 alone, keeps one copy of a looped phrase, and drops subtitle boilerplate. whisper_condition_on_previous_text now defaults to false. Whisper's own default feeds each clip the previous clip's text, which is what makes it loop; batched decoding ignores the flag anyway, so off also keeps the batched path and the out-of-memory retry sounding the same. convert_to_simplified runs OpenCC over the text when the zh extra is installed, and warns rather than failing when it is not. `ytscript polish PATH...` re-applies the clean-up to scripts already written, so adding a term does not mean re-transcribing a backlog. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GARmHk71niEh1TPunhH3Z8
The zh extra went into pyproject.toml without a matching uv.lock, which `uv lock --check` catches in the lint job — local runs pass `--frozen` and skip it. Add the extra to the extras matrix too, so its lazily imported dependency is checked the way local, openai and drive are. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GARmHk71niEh1TPunhH3Z8
load_vocabulary told a built-in name from a file path by looking for a forward slash, which a Windows path does not contain: vocabulary = "C:\scripts\mine.txt" resolved to the built-in directory and failed with "no vocabulary". A built-in is now a bare stem — no directory separator of either kind, no suffix — and a file of that name is tried either way, so a built-in never shadows one. Both separators are checked on every platform, so the unit test covering it fails everywhere rather than only on the Windows runner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GARmHk71niEh1TPunhH3Z8
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR introduces a comprehensive vocabulary system that improves transcription accuracy for domain-specific content by priming Whisper with expected terms and automatically correcting known misrecognitions.
Summary
The changes add three main components:
Vocabulary module (
vocabulary.py): Manages domain-specific terms and correctionswrong => rightcorrectionszh-finance.txt)Polish module (
polish.py): Post-processing for transcription cleanupCLI enhancements:
--vocabularyflag onruncommand to specify terms/corrections filepolishsubcommand to re-process existing scripts with updated vocabularies"zh-finance") and file pathsKey Changes
vocabulary,prompt_from_metadata,whisper_condition_on_previous_text,polish, andconvert_to_simplifiedsettingspromptparameterImplementation Details
https://claude.ai/code/session_01GARmHk71niEh1TPunhH3Z8