Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -91,6 +91,9 @@ jobs:
- extra: drive
# The connector imports the client lazily, so name it too.
import: "ytscript.drive, googleapiclient.discovery, google_auth_oauthlib.flow, google_auth_httplib2, socks"
- extra: zh
# Imported lazily as well, and only when convert_to_simplified is on.
import: "ytscript.polish, opencc"
steps:
- uses: actions/checkout@v5

Expand Down
84 changes: 79 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,8 @@ machine. The `openai` extra needs `OPENAI_API_KEY` in the environment and charge
minute of audio, but needs no local model. They are declared as conflicting extras — two
transcription stacks, nothing needs both — so sync one at a time. The `drive` extra
is independent of both and adds the Google client libraries for [saving the scripts to
Google Drive](#google-drive).
Google Drive](#google-drive), as is the `zh` extra, which adds [OpenCC] for rewriting
traditional characters as simplified.

**On an NVIDIA GPU, install the CUDA libraries too.** faster-whisper runs on
[CTranslate2], which needs cuBLAS and cuDNN 9 and does not bundle them. In a checkout:
Expand Down Expand Up @@ -77,6 +78,8 @@ ytscript drive-auth # optional: sign in to Google Drive once
ytscript run --dry-run # what would be transcribed
ytscript run # first run: the latest 30 videos
ytscript run # later runs: only what is new

ytscript polish scripts # re-clean scripts already written
```

Scripts land in `output_dir` as `2024-05-01_Video-title_VIDEOID.txt`, and every finished
Expand All @@ -91,6 +94,7 @@ Useful flags on `run`:
| `--channel @handle` | Override the configured channel |
| `--language zh` | Override the main language (`auto` detects it) |
| `--batch-size N` | Clips decoded at once; `1` turns batching off |
| `--vocabulary NAME` | Terms and rewrites for the channel: a built-in name or a file |
| `--limit N` | Check the newest N videos, whatever the state file says |
| `--backfill` | Check `initial_backfill` videos again (default 30) |
| `--format txt,md,json` | Write more than one rendering |
Expand All @@ -106,6 +110,11 @@ Useful flags on `run`:

The cookie and members-only flags work on `list` too.

`ytscript polish` runs the same clean-up a run does over scripts that already exist —
useful after adding a term to the vocabulary, or on a backlog transcribed before it had
one. It takes files or directories (`.txt` and `.md`), rewrites them in place, and has
`--vocabulary NAME`, `--simplified` and `--dry-run`.

## Configuration

Settings are read from `ytscript.toml` (searched for in the working directory and its
Expand All @@ -125,11 +134,17 @@ whisper_device = "cuda" # "cpu", "cuda", ...
whisper_compute_type = "float16" # "int8" on CPU, "float16" on GPU
whisper_initial_prompt = "以下是普通话的句子。" # seeds simplified characters
whisper_batch_size = 4 # clips decoded at once; 1 turns batching off
whisper_condition_on_previous_text = false # false stops the model looping a phrase

prompt_from_metadata = true # prime each video with its own title and description
vocabulary = "zh-finance" # terms and rewrites; a built-in name or a file path

output_dir = "scripts"
output_formats = ["txt"] # any of txt, md, json
timestamps = false
paragraph_gap = 2.0 # silence in seconds that starts a new paragraph
polish = true # tidy punctuation, looped phrases and known bad spellings
convert_to_simplified = false # traditional -> simplified; needs --extra zh

state_file = ".ytscript-state.json"
keep_audio = false
Expand Down Expand Up @@ -224,14 +239,73 @@ whisper_initial_prompt = "以下是普通话的句子。" # simplified
whisper_initial_prompt = "以下是普通話的句子。" # traditional
```

It is a nudge, not a guarantee — a few characters can still come out the other way. If
the output has to be uniform, run a converter such as [OpenCC] over the scripts
afterwards. The prompt also matters if you change `language`: a Chinese seed on English
audio makes the transcription worse, so change or remove it along with the language.
It is a nudge, not a guarantee — a few characters can still come out the other way. For
uniform output, `uv sync --extra zh` and set `convert_to_simplified = true`, which runs
[OpenCC] over the text before it is written; without the extra installed the setting
warns once and leaves the characters alone. The prompt also matters if you change
`language`: a Chinese seed on English audio makes the transcription worse, so change or
remove it along with the language. Leaving `whisper_initial_prompt` out of the config
entirely gets the stock sentence for `language`; `whisper_initial_prompt = ""` gets no
seed at all.

Output is written with no spaces between Chinese segments, the way the language is
written; a Latin word inside a sentence still keeps the spaces on either side of it.

### Getting the words right

Whisper decodes a word it is expecting far more readily than one it is not, and the
words a niche channel leans on — tickers, index nicknames, the host's own name — are
exactly the ones a general model has no reason to expect. Left alone it guesses at the
sounds: 费半 (a Mandarin nickname for the Philadelphia semiconductor index) comes out as
飞班 or 肺瓣, 对冲基金 as 对中基金, 资产负债表 as 自然负债表, SNDK as `S&DK`.

Three settings work on this, and they stack:

- **`prompt_from_metadata`** (on by default) puts the video's own title and the first
line of its description in front of the model as the transcript so far. A title like
`半导体、ASML、TSM、NFLX、ISRG、MU、SNDK` names most of the day's tickers before a word
of audio is decoded, and it costs nothing — the metadata is already downloaded.
- **`vocabulary`** adds the terms the channel says every episode, and rewrites what the
model gets wrong anyway. `"zh-finance"` ships with ytscript for a Mandarin US-market
channel; the file is small, commented and meant to be copied:

```bash
cp "$(uv run python -c 'import ytscript.vocabulary as v; print(v.DATA_DIR)')/zh-finance.txt" mine.txt
```

```
# a term is seeded into the prompt
杰克逊霍尔
# a rewrite is seeded and applied to the output
对中基金 => 对冲基金
CTS => CDS
```

Rewrites are literal, not regular expressions. An ASCII left-hand side only matches as
a whole word, so `CTS` never fires inside `CTSX`; a Chinese one matches anywhere, since
the language has no word boundaries to key on. Point `vocabulary` at your copy and add
to it as you read the scripts — the mistakes a model makes are specific to the voice.

- **`polish`** (on by default) cleans the text after recognition: Chinese sentences get
the fullwidth punctuation they are written with (`大家好,` → `大家好,`), while the
commas in `1,250` and the colon in a URL are left alone; a phrase the model looped on
is kept once instead of a dozen times; and the vocabulary's rewrites are applied.

The prompt is capped at 200 characters, because Whisper conditions on about 224 tokens
and silently drops the rest. The seed sentence and the video's own subject go in first,
then terms — the ones named in the title before the rest, since those are the words the
episode actually says.

`whisper_condition_on_previous_text` is the other accuracy setting, and ytscript
defaults it to `false` where Whisper's own default is `true`. Feeding each clip the text
of the one before it is what makes the model repeat a phrase for a minute once the audio
goes quiet. Batched decoding has no previous clip to condition on and behaves as if the
flag were off regardless, so `false` also keeps batched and sequential runs — including
the one-clip-at-a-time retry after an out-of-memory error — sounding the same.

None of this touches the audio, so a script can be re-cleaned without re-transcribing:
add the term, then `ytscript polish scripts`.

### Choosing a model

`large-v3` at `float16` is the default because it is the best Chinese accuracy an 8 GB
Expand Down
3 changes: 3 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,9 @@ dependencies = [
local = ["faster-whisper>=1.1.0"]
# Hosted speech-to-text via the OpenAI audio transcription endpoint.
openai = ["openai>=1.30.0"]
# Traditional -> simplified conversion, for `convert_to_simplified`. Pure Python,
# and independent of which speech-to-text backend is in use.
zh = ["opencc-python-reimplemented>=0.1.7"]
# Copy the finished scripts into Google Drive. Independent of the two backends.
# PySocks is not optional in practice: without it httplib2, which the Google client
# is built on, quietly ignores every proxy setting instead of using one.
Expand Down
85 changes: 84 additions & 1 deletion src/ytscript/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@
from .drive import DriveError, DriveUploader
from .models import RunReport
from .pipeline import Pipeline
from .polish import polish_text
from .vocabulary import VocabularyError, load_vocabulary
from .youtube import YouTubeError

log = logging.getLogger("ytscript")
Expand Down Expand Up @@ -76,6 +78,10 @@ def add_common(target: argparse.ArgumentParser) -> None:
type=int,
help="clips decoded at once; 1 turns batching off",
)
run.add_argument(
"--vocabulary",
help="terms and rewrites for this channel: a built-in name or a file path",
)
run.add_argument("--output-dir", dest="output_dir", type=Path)
run.add_argument(
"--format",
Expand Down Expand Up @@ -112,6 +118,31 @@ def add_common(target: argparse.ArgumentParser) -> None:
help="show what would be transcribed without downloading anything",
)

polish = sub.add_parser(
"polish",
help="re-apply the vocabulary and punctuation clean-up to scripts already written",
)
polish.add_argument(
"paths",
nargs="+",
type=Path,
metavar="PATH",
help="script files, or directories to take the .txt and .md files from",
)
polish.add_argument("--vocabulary", help="a built-in name or a file path")
polish.add_argument(
"--simplified",
dest="convert_to_simplified",
action="store_true",
default=None,
help="also rewrite traditional characters as simplified",
)
polish.add_argument(
"--dry-run",
action="store_true",
help="list the files that would change without writing them",
)

listing = sub.add_parser("list", help="show the newest videos on the channel")
add_common(listing)
listing.add_argument("--limit", type=int, default=10)
Expand All @@ -134,6 +165,7 @@ def add_common(target: argparse.ArgumentParser) -> None:
"backend",
"whisper_model",
"whisper_batch_size",
"vocabulary",
"output_dir",
"output_formats",
"timestamps",
Expand Down Expand Up @@ -211,6 +243,56 @@ def cmd_list(args: argparse.Namespace) -> int:
return 0


SCRIPT_SUFFIXES = (".txt", ".md")


def _script_paths(paths: list[Path]) -> list[Path]:
"""Expand the arguments into the script files to rewrite."""
found: list[Path] = []
for path in paths:
if path.is_dir():
found.extend(
sorted(child for child in path.iterdir() if child.suffix in SCRIPT_SUFFIXES)
)
elif path.is_file():
found.append(path)
else:
raise ConfigError(f"no such file or directory: {path}")
return found


def cmd_polish(args: argparse.Namespace) -> int:
# No channel is needed to rewrite files that already exist.
config = load_config(path=args.config)
name = args.vocabulary if args.vocabulary is not None else config.vocabulary
vocabulary = load_vocabulary(name)
simplified = (
config.convert_to_simplified
if args.convert_to_simplified is None
else args.convert_to_simplified
)
paths = _script_paths(args.paths)
if not paths:
print("no .txt or .md scripts found")
return 0

changed = []
for path in paths:
original = path.read_text(encoding="utf-8")
updated = polish_text(original, vocabulary, simplified=simplified)
if updated == original:
continue
changed.append(path)
if not args.dry_run:
path.write_text(updated, encoding="utf-8")

verb = "would rewrite" if args.dry_run else "rewrote"
print(f"{verb} {len(changed)} of {len(paths)} file(s)")
for path in changed:
print(f" {path}")
return 0


def cmd_drive_auth(args: argparse.Namespace) -> int:
# No channel is needed to authorise, so this skips the usual validation.
config = load_config(path=args.config)
Expand Down Expand Up @@ -251,12 +333,13 @@ def main(argv: list[str] | None = None) -> int:
handlers = {
"run": cmd_run,
"list": cmd_list,
"polish": cmd_polish,
"init": cmd_init,
"drive-auth": cmd_drive_auth,
}
try:
return handlers[args.command](args)
except (ConfigError, YouTubeError, DriveError) as exc:
except (ConfigError, VocabularyError, YouTubeError, DriveError) as exc:
print(f"error: {exc}", file=sys.stderr)
return 1
except KeyboardInterrupt: # pragma: no cover
Expand Down
58 changes: 55 additions & 3 deletions src/ytscript/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,8 @@
from pathlib import Path
from typing import Any

from .vocabulary import VocabularyError, load_vocabulary

DEFAULT_CONFIG_NAMES = ("ytscript.toml", ".ytscript.toml")
ENV_PREFIX = "YTSCRIPT_"

Expand Down Expand Up @@ -45,7 +47,22 @@ class Config:
"""Defaults target an 8 GB NVIDIA card; see the README for the CPU-only settings."""

whisper_initial_prompt: str | None = None
"""Seed text that steers spelling and register — for Chinese, simplified vs traditional."""
"""Seed text that steers spelling and register — for Chinese, simplified vs traditional.

``None`` uses the stock sentence for ``language``; set it to ``""`` for no seed."""

whisper_condition_on_previous_text: bool = False
"""Feed each clip the previous clip's text. Whisper's own default is on, and it is
what makes the model repeat a phrase for a minute once it starts. Batched decoding
turns it off regardless, so leaving it off also keeps the two paths in step."""

prompt_from_metadata: bool = True
"""Prime each video with its own title and description, so the day's tickers and
names are words the model is already expecting."""

vocabulary: str | None = None
"""Terms the channel says every episode, and rewrites for what the model gets wrong.
A built-in name (``"zh-finance"``) or the path to your own file."""

whisper_batch_size: int = 4
"""Clips decoded at once. ``1`` turns batching off; 4 suits large-v3 on an 8 GB card."""
Expand All @@ -62,6 +79,13 @@ class Config:
paragraph_gap: float = 2.0
"""Silence in seconds between segments that starts a new paragraph."""

polish: bool = True
"""Tidy the recognised text before writing it: fullwidth punctuation for Chinese
sentences, one copy of a looped phrase, and the vocabulary's rewrites."""

convert_to_simplified: bool = False
"""Rewrite traditional characters as simplified ones. Needs 'uv sync --extra zh'."""

# --- Google Drive (optional) -----------------------------------------
drive_upload: bool = False
"""Also copy every finished script into Google Drive. Local files are written either way."""
Expand Down Expand Up @@ -152,6 +176,11 @@ def validate(self) -> None:
raise ConfigError("check_limit must be at least 1")
if self.whisper_batch_size < 1:
raise ConfigError("whisper_batch_size must be at least 1 (1 turns batching off)")
# Reading the file now means a typo fails the command, not the first video.
try:
load_vocabulary(self.vocabulary)
except VocabularyError as exc:
raise ConfigError(str(exc)) from exc
if self.include_members_only and not (self.cookies_file or self.cookies_from_browser):
raise ConfigError(
"include_members_only needs a signed-in session: set cookies_file or "
Expand Down Expand Up @@ -286,10 +315,26 @@ def load_config(
whisper_compute_type = "float16"

# Whisper transcribes Mandarin into traditional characters about as readily as
# simplified. A simplified-character seed sentence settles it. Change or clear
# this if you change `language`.
# simplified. A simplified-character seed sentence settles it. Leave it unset to
# get the stock sentence for `language`, or set it to "" for no seed at all.
whisper_initial_prompt = "以下是普通话的句子。"

# The seed is followed by the video's own title and description, so the tickers
# and names that episode is about are words the model already expects.
prompt_from_metadata = true

# Terms the channel says every episode, plus rewrites for the ones the model
# keeps getting wrong ("对中基金 => 对冲基金"). "zh-finance" ships with ytscript
# and suits a Mandarin US-market channel; point this at your own file to extend
# it — `python -c "import ytscript.vocabulary as v; print(v.DATA_DIR)"` finds the
# built-in to copy. Leave it out for no vocabulary.
vocabulary = "zh-finance"

# Whisper's own default feeds each clip the text of the one before it, which is
# what makes it repeat a phrase for a minute when the audio goes quiet. Off also
# matches what batched decoding does, so both paths sound the same.
whisper_condition_on_previous_text = false

# Clips decoded at once — several times faster than one at a time, at the cost
# of VRAM. 4 leaves headroom on an 8 GB card; a 12 GB or larger card can take
# 8 or 16. Drop to 1 to turn batching off.
Expand All @@ -299,6 +344,13 @@ def load_config(
output_formats = ["txt"]
timestamps = false

# Tidy the text before writing: fullwidth punctuation for Chinese sentences, one
# copy of a phrase the model looped on, and the vocabulary's rewrites.
polish = true

# Rewrite traditional characters as simplified. Needs `uv sync --extra zh`.
convert_to_simplified = false

state_file = ".ytscript-state.json"
keep_audio = false

Expand Down
Loading
Loading