Skip to content

Add Google Drive integration for uploading transcripts - #6

Merged
reecemiao merged 1 commit into
mainfrom
claude/google-drive-connector-t0jma5
Aug 24, 2026
Merged

reecemiao merged 1 commit into
mainfrom
claude/google-drive-connector-t0jma5

Conversation

@reecemiao

Copy link
Copy Markdown
Owner

Summary

This PR adds optional Google Drive integration to ytscript, allowing finished transcripts to be automatically uploaded to Google Drive alongside local storage. The feature is entirely optional and disabled by default, with no impact on existing workflows.

Key Changes

  • New DriveUploader class (src/ytscript/drive.py): Handles authentication and uploading to Google Drive

    • Supports both OAuth (interactive sign-in via drive-auth command) and service account credentials
    • Implements folder creation/resolution and file upload with automatic replacement of existing files
    • Provides clear error messages for missing credentials or authorization issues
    • Gracefully handles optional Google client library dependencies
  • Configuration additions (src/ytscript/config.py):

    • drive_upload: Enable/disable Drive uploads
    • drive_folder_id / drive_folder_name: Control where scripts are stored
    • drive_credentials_file: Path to OAuth client secrets JSON
    • drive_token_file: Cached authentication token
    • drive_service_account_file: Service account key for unattended runs
    • drive_scope: Control permission level ("drive.file" or "drive")
    • Validation ensures credentials are configured when uploads are enabled
  • Pipeline integration (src/ytscript/pipeline.py):

    • Uploads each written script to Drive after local writing
    • Handles upload failures gracefully with retry on next run
    • Tracks uploaded files in state for reporting
  • CLI enhancements (src/ytscript/cli.py):

    • New drive-auth command for one-time browser sign-in
    • --drive / --no-drive flags to override config
    • --drive-folder flag to specify target folder
    • Reports uploaded files in run output
  • Testing (tests/test_drive.py):

    • Comprehensive test suite covering authentication flows, folder management, file uploads, and error handling
    • Fake Drive implementation for testing without Google client libraries
  • Documentation (README.md):

    • Setup instructions for Google Cloud console configuration
    • Usage examples and troubleshooting
    • Configuration reference

Implementation Details

  • The Google client libraries are optional dependencies (installed via uv sync --extra drive), so ytscript works without them
  • Authentication is cached after drive-auth, allowing unattended runs
  • File uploads replace earlier copies of the same name rather than creating duplicates
  • Failed uploads don't block the run; they're retried on the next execution
  • Service account support enables CI/CD workflows without browser interaction

https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

Every finished script can now be copied into Google Drive as well. The local
files under output_dir are written either way, so the connector is a copy and
never a destination — with drive_upload off, nothing about a run changes.

The Google client libraries are a new `drive` extra, imported lazily, so a
checkout without them behaves exactly as before and the faked test suite still
runs on a bare `uv sync`.

- `ytscript drive-auth` walks through the browser sign-in once and caches the
  token, which refreshes itself afterwards so a cron run needs no browser. A
  service account key is the alternative for a machine with no browser at all.
- Uploads land in a folder ytscript creates (drive_folder_name) or in one given
  by id or URL (drive_folder_id). The default drive.file scope only ever sees
  what ytscript uploaded; an existing folder needs drive_scope = "drive", and
  the mismatch is reported rather than left to fail at upload time.
- Files are matched by name inside the folder, so re-transcribing a video
  replaces its copy instead of leaving a second one behind.
- Sign-in happens once before the first download, so a stale token costs a
  second rather than a whole backfill. A failed upload fails that video: it is
  not recorded, and the next run picks it up again.

The Drive id and link of each upload are recorded in the state file next to the
local paths, and the run report prints them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62
@reecemiao
reecemiao merged commit f566830 into main Aug 24, 2026
10 checks passed
reecemiao pushed a commit that referenced this pull request Aug 24, 2026
Merging main back into this branch after #6 was squashed conflicted across every
file the squash had rewritten, and the resolution took main's side wholesale.
That reverted four commits: the proxy support, the PySocks dependency without
which httplib2 silently ignores every proxy, the sign-in troubleshooting, and
the account and scope explanations.

The uv.lock and CI entries from those commits survived, which left the tree
contradicting itself — a lockfile naming a dependency pyproject.toml no longer
declared, and an extras job importing modules the extra no longer installed.
Both are why CI is red rather than merely stale.

Restores the files to the branch's state before the merge. No content from main
is lost: everything the merge added back was this branch's own earlier text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62
reecemiao added a commit that referenced this pull request Aug 24, 2026
* Add an optional Google Drive connector for the scripts

Every finished script can now be copied into Google Drive as well. The local
files under output_dir are written either way, so the connector is a copy and
never a destination — with drive_upload off, nothing about a run changes.

The Google client libraries are a new `drive` extra, imported lazily, so a
checkout without them behaves exactly as before and the faked test suite still
runs on a bare `uv sync`.

- `ytscript drive-auth` walks through the browser sign-in once and caches the
  token, which refreshes itself afterwards so a cron run needs no browser. A
  service account key is the alternative for a machine with no browser at all.
- Uploads land in a folder ytscript creates (drive_folder_name) or in one given
  by id or URL (drive_folder_id). The default drive.file scope only ever sees
  what ytscript uploaded; an existing folder needs drive_scope = "drive", and
  the mismatch is reported rather than left to fail at upload time.
- Files are matched by name inside the folder, so re-transcribing a video
  replaces its copy instead of leaving a second one behind.
- Sign-in happens once before the first download, so a stale token costs a
  second rather than a whole backfill. A failed upload fails that video: it is
  not recorded, and the next run picks it up again.

The Drive id and link of each upload are recorded in the state file next to the
local paths, and the run report prints them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

* Say how to get the Drive credentials, and to publish the app

The setup steps skipped past the one thing that breaks an unattended run: while
the Cloud project's publishing status is "Testing", Google's refresh tokens stop
working after 7 days, so a cron run dies a week after it was authorised.
Publishing is free and immediate for the default drive.file scope, which is not
sensitive enough to need Google's verification review.

Also names the file the console hands you, so a service account key or a web
client downloaded by mistake is recognisable before ytscript refuses it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

* Document the "Access blocked" sign-in refusal

Google refuses the sign-in before ytscript sees anything when the app is in
testing and the account is not a test user, when a published app lists a scope
that needs the verification review, or when a Workspace administrator blocks
unverified apps. Name the message and each of the three, since the fix is in the
Cloud console either way.

Also spell out the service account steps, which skip the consent screen, the
verification question and the expiring token altogether — at the cost of the
uploads being owned by the account rather than by you.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

* Note that publishing does not rescue an existing Drive token

The 7 days are stamped on the refresh token when it is issued, so publishing the
app afterwards leaves a token already in hand still expiring — an unpleasant
surprise a week after the setting looks correct. Say to re-authorise, and that
the test-user list stops mattering once the app is published.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

* Say who a published app lets into your Drive, and reorder the section

Publishing the OAuth app sounds like it opens your Drive to strangers, and the
section never said otherwise. It does not: each account that signs in reaches its
own Drive, what reaches yours is the cached token, and the drive.file scope keeps
even that to the files ytscript uploaded. Say so, and say how to revoke.

The console setup had also drifted apart from the sign-in it leads to, with
troubleshooting wedged between them. Setup now runs start to finish, and
publishing, refused sign-ins and access each get a heading of their own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

* Point a connection timeout at the proxy that caused it

A call that never reached googleapis.com surfaced as a bare socket error —
"looking for ytscript failed: [WinError 10060] ..." — with nothing to act on. The
cause is usually that the two halves of a run disagree about where the proxy is:
yt-dlp reads the system proxy settings, while the Google client reads only
HTTPS_PROXY from the environment, so downloads work and uploads time out.

Name that in the error and in the README, on the errors that mean the request
went unanswered. An error Google actually sent back is left alone, since it says
more than the hint would.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

* Use the machine's own proxy for Drive, and make httplib2 able to

Setting HTTPS_PROXY did not help a blocked connection, because httplib2 gates all
proxy support behind PySocks: `ProxyInfo.isgood()` reads `socks and ...`, and
httplib2 no longer bundles that module, so every proxy setting was quietly
skipped and the connection went out direct to time out. PySocks joins the extra,
and a proxy that cannot be used now says so instead of being ignored.

With that fixed, read the proxy the way yt-dlp already does — `getproxies()`,
which prefers the environment and falls back to the Windows registry or the macOS
network settings — and hand it to the client. A machine whose downloads already
work needs nothing configured for its uploads to work too.

A request that goes unanswered now names the proxy it went through, or says it
went direct, rather than leaving a bare socket error to be interpreted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

* Put -v where argparse accepts it

--verbose belongs to the top-level parser, so `ytscript run -v` is rejected as an
unrecognised argument. The proxy section told the reader to type exactly that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

* Restore the work the merge from main dropped

Merging main back into this branch after #6 was squashed conflicted across every
file the squash had rewritten, and the resolution took main's side wholesale.
That reverted four commits: the proxy support, the PySocks dependency without
which httplib2 silently ignores every proxy, the sign-in troubleshooting, and
the account and scope explanations.

The uv.lock and CI entries from those commits survived, which left the tree
contradicting itself — a lockfile naming a dependency pyproject.toml no longer
declared, and an extras job importing modules the extra no longer installed.
Both are why CI is red rather than merely stale.

Restores the files to the branch's state before the merge. No content from main
is lost: everything the merge added back was this branch's own earlier text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MJVQg1LaKgxiVPt2SAjt62

---------

Co-authored-by: reecemiao <reece.miao.jp@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant