You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat: add a nightly copy pipe for repo commit contributors (CM-1824) - #4886
CM-1824 (epic CM-1807, Akrites member contributors) adds a daily sync that writes every commit author crowd.dev knows about into the packages-db repo_contributors table (created in #4885), so a package's repo can answer "who actually commits here" without calling GitHub.
The source has to be Tinybird: public.activities in CDP is deprecated and activityRelations in Postgres is too slow for this. The worker needs the all-time aggregate per (channel, memberId, platform, username) over authored-commit and co-authored-commit activities: commit count, first and last commit date.
That aggregate cannot be run as paged ad-hoc SQL from the worker. Measured on prod (2026-10-02):
query
result
full aggregate, any page
times out at the 15s SQL API limit (218M commit rows, ~7M groups built and sorted before any LIMIT)
one segment at a time (sorting key starts with segmentId)
2.4s each, but 11,437 segments → ~11k requests and 5–9h per daily run
paging by channel or offset
channel is not in the sorting key, every page is a full scan
A nightly COPY pipe runs the aggregation once as a batch job with no API time limit, and the worker then pages the materialized result in sorting-key order (~7M rows, ~140 requests at 50k rows each). This is the same pattern as org_page_contributors_copy_pipe and the other nightly aggregates in this workspace.
What
pipes/repo_commit_contributors_copy.pipe: COPY pipe over activityRelations_deduplicated_cleaned_bucket_union, authored + co-authored commits summed into one count per (channel, memberId, platform, username), min/max commit timestamps, plus max(updatedAt) as lastUpdatedAt. Bots, team members and organization profiles are already excluded by the cleaned source (members_sorted); pre-1971 sentinel timestamps are dropped like the other commit aggregates. COPY_MODE replace, scheduled 03:30 UTC (after repos_channels_copy at 00:00, clear of the health score pipes at 02:00).
datasources/repo_commit_contributors_copy_ds.datasource: ReplacingMergeTree keyed on the same four columns so the consumer can keyset-page, computedAt as version column. lastUpdatedAt moves on ingestion and on member merges (both bump activityRelations.updatedAt), so the worker syncs incrementally: it only pulls rows whose lastUpdatedAt passed the watermark stored from its previous run, instead of rewriting ~7M rows in packages-db every day.
Expected size: ~6.5M rows (authored) + ~0.46M (co-authored) before summing overlaps and removing bots.
The aggregate was validated read-only on prod filtered to one repo (kubernetes/kubernetes) and returns the expected rows. Nothing has been pushed to any Tinybird workspace yet.
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution. You have signed the CLA already but the status is still pending? Let us recheck it.
Medium Risk
A nightly full-replace aggregate feeds packages-db repo contributor sync; incorrect SQL or schema would propagate bad contributor data at scale, though the change is additive and mirrors existing copy pipes.
Overview
Adds Tinybird infrastructure so the packages worker can incrementally sync all-time repo commit authors without hitting the SQL API timeout on a live aggregate over ~218M commit rows.
A nightly COPY pipe (repo_commit_contributors_copy.pipe, 03:30 UTC, full replace) rolls up authored-commit and co-authored-commit rows from activityRelations_deduplicated_cleaned_bucket_union into one row per (channel, memberId, platform, username) with commit count, first/last commit times, and lastUpdatedAt (max(updatedAt)) for watermark-based incremental reads. Filters match other commit aggregates (non-empty channel/username, drop pre-1971 timestamps); member cleanup stays upstream in the cleaned source.
The target datasource (repo_commit_contributors_copy_ds) uses ReplacingMergeTree with a sorting key on those four dimensions so the git-activity contributors sync can keyset-page the ~7M-row result set.
Reviewed by Cursor Bugbot for commit de58994. Bugbot is set up for automated code reviews on this repo. Configure here.
This source has already normalized every channel to lowercase (activityRelations_bucket_clean_enrich_copy_pipe_0.pipe:14). That loses the canonical casing needed to resolve GitLab/Bitbucket repositories: the packages worker deliberately preserves GitLab path casing (member-contributors/__tests__/mapRows.test.ts:17-20), so an exact repo lookup will skip mixed-case repos. Emit the canonical repository URL (for example, by mapping through repos_channels_ds.repoUrl) before grouping.
The reason will be displayed to describe this comment to others. Learn more.
LGTM. The cleaned bucket select keeps channel in its original case (lower() only appears in the repos_to_channels filter), so the mixed-case GitLab path concern does not apply here.
channel is already lowercased by every cleaned-bucket copy (activityRelations_bucket_clean_enrich_copy_pipe_0.pipe:14), while packages-db deliberately preserves GitLab and Bitbucket path casing (services/apps/packages_worker/src/deps-dev/canonicalRepoUrl.ts:18-23). The original case is therefore unrecoverable, so mapping this value to repos.url will skip or misidentify case-sensitive repositories. Materialize the canonical repository URL from repository metadata, retaining segmentId where needed to disambiguate, instead of exporting the cleaned channel.
The reason will be displayed to describe this comment to others. Learn more.
Agreed that updatedAt alone does not see members_sorted or repos_to_channels changes. The consumer (worker PR, same ticket) handles it on its side: the daily run is incremental on lastUpdatedAt, and every two weeks it ignores the watermark, reads the full datasource and reconciles with the same touch-and-delete pattern the governance sync uses, which catches both newly eligible historical contributors and rows that left the source. No pipe change needed; a replace-mode copy cannot cheaply version those dependency changes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
CM-1824 (epic CM-1807, Akrites member contributors) adds a daily sync that writes every commit author crowd.dev knows about into the packages-db
repo_contributorstable (created in #4885), so a package's repo can answer "who actually commits here" without calling GitHub.The source has to be Tinybird:
public.activitiesin CDP is deprecated andactivityRelationsin Postgres is too slow for this. The worker needs the all-time aggregate per(channel, memberId, platform, username)overauthored-commitandco-authored-commitactivities: commit count, first and last commit date.That aggregate cannot be run as paged ad-hoc SQL from the worker. Measured on prod (2026-10-02):
LIMIT)segmentId)channelor offsetchannelis not in the sorting key, every page is a full scanA nightly COPY pipe runs the aggregation once as a batch job with no API time limit, and the worker then pages the materialized result in sorting-key order (~7M rows, ~140 requests at 50k rows each). This is the same pattern as
org_page_contributors_copy_pipeand the other nightly aggregates in this workspace.What
pipes/repo_commit_contributors_copy.pipe: COPY pipe overactivityRelations_deduplicated_cleaned_bucket_union, authored + co-authored commits summed into one count per(channel, memberId, platform, username),min/maxcommit timestamps, plusmax(updatedAt)aslastUpdatedAt. Bots, team members and organization profiles are already excluded by the cleaned source (members_sorted); pre-1971 sentinel timestamps are dropped like the other commit aggregates.COPY_MODE replace, scheduled 03:30 UTC (afterrepos_channels_copyat 00:00, clear of the health score pipes at 02:00).datasources/repo_commit_contributors_copy_ds.datasource: ReplacingMergeTree keyed on the same four columns so the consumer can keyset-page,computedAtas version column.lastUpdatedAtmoves on ingestion and on member merges (both bumpactivityRelations.updatedAt), so the worker syncs incrementally: it only pulls rows whoselastUpdatedAtpassed the watermark stored from its previous run, instead of rewriting ~7M rows in packages-db every day.Expected size: ~6.5M rows (authored) + ~0.46M (co-authored) before summing overlaps and removing bots.
The aggregate was validated read-only on prod filtered to one repo (
kubernetes/kubernetes) and returns the expected rows. Nothing has been pushed to any Tinybird workspace yet.