Repository navigation
Check every pull request, and stop depending on other people's servers - #193
Conversation
actions/checkout was 2 majors behind, actions-hugo and actions-gh-pages 1 each. checkout@v3 runs on the Node 16 runtime and emits deprecation warnings on every build. Also: - ubuntu-22.04 to ubuntu-latest. The pinned image is deprecated. - Drop 'submodules: true'. There are no submodules and no .gitmodules: the theme is vendored in themes/pybcn_theme. - Add an explicit 'permissions: contents: write', which is what the deploy needs. The default token permissions are wider than the job requires.
The GitHub Actions workflow has published the site since 2020. bin/publish is the manual predecessor and it is dangerous to have lying around: it deletes the master branch, checks out an orphan, runs 'git rm -rf' over archetypes, bin, config.toml, content, README.md, resources, static and themes, then force-pushes master. Run by mistake on the wrong branch it destroys the working tree. Nothing calls it, and the README documents pushing to 'edition' as the way to publish.
resources/_gen holds Hugo's SCSS cache and .hugo_build.lock is a scratch file. Neither belongs in the repository: they are regenerated on every build and they produce spurious diffs on unrelated pull requests. Five of the ten cached files are stale, generated from theme paths that no longer exist (scss/scss/, pybcn-site/scss/). Also ignore docs/, which bin/build writes to.
UA-152706496-1 is a Universal Analytics property. Universal Analytics stopped
processing data in July 2023, so it collects nothing.
Scope of this change, stated precisely: no Hugo template calls
_internal/google_analytics.html, so the 246 templated pages load no tracker and
this config line has no effect on them. Removing it closes the door.
It does NOT stop all tracking on pybcn.org. The 43 verbatim snapshots under
static/archives/pybcn.org/ (PyDay 2016, 2018, 2019 and PyData 2017) hardcode
<script async src="https://www.googletagmanager.com/gtag/js?id=UA-152706496-1">
in their own HTML, and they are live. Every visit to one still sends the
visitor's IP and referrer to Google, with no notice and no legal basis, even
though no analytics data is recorded. Those 43 files also load
maxcdn.bootstrapcdn.com, and some of them load platform.twitter.com and Google
Maps iframes.
Stripping the archived snapshots is a separate change, because it edits
archived third-party HTML rather than site configuration and deserves its own
review.
Add enableGitInfo and enableRobotsTXT next to the theme key, not at the end of the file. At the end of the file the keys become members of the last TOML table and have no effect. enableGitInfo adds 248 lastmod entries to sitemap.xml. enableRobotsTXT renders the robots.txt template that the theme has.
The template ended with a bare "User-agent: *" and no rule. This is an empty group. Add "Allow: /" to make the group valid.
The icon link used ".Site.BaseURL" with a literal slash, which gave "https://pybcn.org//favicon.png". Use relURL instead. Add static/favicon.ico with 16, 32 and 48 pixel frames, 4 KB in total, downsampled from static/favicon.png (546x543, 79 KB) with a one-off Python script that uses only zlib and struct. No converter was installed. Browsers that prefer the ico now get 4 KB instead of 79 KB. Point apple-touch-icon at the existing PNG and set theme-color to #0070b0, the blue sampled from the logo.
assets/js/meetup.js injected a JSONP script from api.meetup.com. Meetup retired the REST API and that endpoint returns HTTP 404, so the list never filled. The visible result was a spinner that turned for ever in four lists on the home page and four on /events/monthly_events/, and every page load sent the visitor IP to Meetup. The partial now shows a short static message and a link to the Meetup group. Delete meetup.js, the requiredJS blocks that loaded it, and the li.spinner rule that only the old markup used. A local list of meetups comes later in the project. The partial has a TODO that says so.
The banner promoted PyDay 2024 while the carousel promoted PyDay 2025. The banner now reads the newest page under content/events/pyday_bcn/. The pages are named pyday_bcn_<year>, so the highest file name is the newest edition. The next edition needs no template change. The link points at the event page instead of the #important_dates anchor, because not every edition declares that section.
The .oov::after rule pulled a 489 KB PNG from upload.wikimedia.org on every 404. The site has no control over that file and the request sends the visitor IP to Wikimedia. Remove the rule. The page said "ERROR #404 / Uh ho." and gave no way out. It now says what happened and links to the home page and to the main menu sections. The list is built from site.Menus.main, so it follows the navigation.
static/archives/ holds verbatim mirrors of old sites that pybcn.org still serves. Each page load sent the visitor to Google, Twitter, MaxCDN and Font Awesome. The scripts add nothing to an archived page. Removed: - 43 files: the googletagmanager gtag.js tag and the gtag() block for UA-152706496-1. Universal Analytics stopped recording in July 2023, so this was a beacon with no data behind it. - 43 files: the Bootstrap 3.3.7 script from maxcdn.bootstrapcdn.com. - 5 files: platform.twitter.com/widgets.js. The timeline link stays and now behaves as a normal link to twitter.com/pybcn. - 3 files: the Google Maps iframes. Each one sits next to the venue address in text, so the page still says where the event was. - 2 files: the Google Calendar iframe on the PyDay 2019 page, replaced by one line of text. - 1 file: the unused Source Sans Pro import in pybcn_style.css. No font-family rule in the archive names that family. Vendored, so no page loses its look: - Bootstrap 3.3.7 CSS, checked against the SRI hash that the pages carried: sha384-BVYiiSIFeK1dGmJRAkycuHAHRg32OmUcww7on3RYdg4Va+PmSTsz /K68vbdEjh4u. - Font Awesome 4.7.0 CSS and the woff2/woff face, for the 43 pybcn.org pages that used the use.fontawesome.com/e19927ebb1.js embed. All 16 icon classes used by those pages resolve in the vendored CSS. - Font Awesome 5.3.1 CSS and the brands webfont, for hacktoberfestbarcelona.com. That page uses only "fab" icons. - Roboto 400 and 700, latin and latin-ext, for the same page, in place of the Google Fonts stylesheet. Glyphicon webfonts are not vendored. No archived page uses a glyphicon class, so the browser never requests the face.
partials/head.html loaded the Font Awesome CSS from use.fontawesome.com on all 246 templated pages, so every visitor IP reached a third party before the page could draw an icon. The repo already self-hosts Bootstrap and jQuery under assets/vendor/, so this follows the same pattern. The vendored all.css matches the SRI hash that the external link carried, sha384-oS3vJWv+0UjzBfQzYUhtDYW+Pj2yciDJxpsK1OYPAYjqT085Qq /1cq5FLXAZQ7Ay, so the file is the same one the site used. The stylesheet goes through minify and Fingerprint and keeps the integrity attribute, like partials/js.html. The webfonts go in the theme static directory, because all.css points at ../webfonts/ and that path must stay stable. Only the woff2 and woff faces are vendored. The eot, ttf and svg faces in the CSS are for Internet Explorer and add about 2 MB. Every current browser takes the woff2.
humans.txt credits the work groups and the people pages instead of a name list that goes stale, and states the association facts. The Meetup numbers carry the date they were counted. .well-known/security.txt follows RFC 9116: two Contact lines, the GitHub security advisory form first and the association mailbox second, plus Expires, Preferred-Languages and Canonical. It is not signed. The deploy action writes .nojekyll, so GitHub Pages serves the dot directory. llms.txt is a Hugo template, layouts/index.llms.txt, behind a custom output format. Hugo rebuilds it with the site, so the section lists and the page counts follow the content and cannot drift. The RSS output is kept in the home outputs list.
# Conflicts: # config.toml
Pull requests get no automated checks today. A PR can break the build, or reference a person or sponsor page that does not exist, and nothing reports it. Add three jobs: - build: the site must compile. - content: front matter must parse, a person id must match its filename, ids must be unique, declared photos must exist, and events must not reference unknown person or sponsor ids. - links: a link check that reports but does not block, because external sites rate-limit. The content check finds 6 errors in the current content. Two of them are visible on the live site: a sponsor of PyDay 2021 and 2022 has no sponsor page, so it renders on neither, and a speaker listed on PyDay 2020 and 2021 is dropped from both because their person id does not match the filename the events reference. The fixes go in a separate change. This one only makes the problems visible.
A floating tag is mutable. Whoever controls the action repository can move v5 to new code, which then runs with the workflow's token. A SHA cannot be moved. Each SHA is the latest release of its action, so the pin also advances two of them past the major the previous commit set: actions/checkout v5 -> v7.0.1 actions/setup-python v6 -> v7.0.0 lycheeverse/lychee v2 -> v2.9.0 peaceiris/actions-hugo v3 -> v3.2.1 peaceiris/gh-pages v4 -> v4.1.0 The release each SHA corresponds to is in a comment next to it, so a reader still sees which version is in use. Renovate and Dependabot both understand this format and can keep the pairs in step.
The site uses neither. Hugo enables both by default, so it published two empty index pages and listed them in the sitemap. An empty [taxonomies] table turns them off. The archive work will declare its own taxonomies when it needs them.
The workflow filtered pull_request on the base branch "edition". The four pull requests of this series are stacked on each other, so only the first has "edition" as its base, and that one does not contain the workflow. The result is that the checks never ran on any of them. Remove the branch filter. The workflow now runs on every pull request, whatever its base branch. The permissions block stays at contents: read.
The first Contact pointed at the GitHub private vulnerability report form. Private vulnerability reporting is not enabled on the repository, so the form rejects everyone outside the organization. RFC 9116 reads Contact lines in order of preference, so the dead entry sat in the position a reporter tries first. Remove the line. The mailbox is now the only contact. A comment keeps the URL and says when to put it back: once the repository enables private vulnerability reporting. The rest of the file does not change.
Every action is pinned to a commit SHA, and the commit that pinned them says Dependabot can keep the pins current. There was no Dependabot configuration in the repository, so the pins would never move and would age in silence. Add .github/dependabot.yml with one github-actions entry on a monthly schedule. Dependabot opens a pull request that bumps the SHA and the version comment together, with the "ci" prefix the workflow commits already use.
The Prerequisites section listed pip, but the install script does not use Python. Line 18 of the same file already said so. The script downloads the Hugo binary with curl, verifies it with sha256sum, and unpacks it with tar. Replace pip with those three tools, which matches what bin/install runs.
The links job could not pass. Hugo emitted about 5,000 absolute links of the form "/path/" and lychee ran without --root-dir, so lychee flagged every one of them as an error. The 7,500 "https://pybcn.org/..." self-links went to production instead, so a pull request that adds a page saw 404s for its own new pages. Build the job's site with "--baseURL /", which turns every internal link into "/path/", and pass --root-dir "$GITHUB_WORKSPACE/public" so lychee resolves those paths inside the build. Lychee treats a link to an existing directory as valid, so "/about/" resolves to public/about/ with no extra option. Exclude the frozen snapshots under public/archives (56 pages from 2016 to 2019 that keep their dead links on purpose) and the hosts that answer bots with 400, 403, or 999: twitter.com, x.com, and linkedin.com. The site has about 530 twitter.com, 10 x.com, and 580 linkedin.com anchors. Every flag exists in lychee v0.24.2, the version lychee-action v2.9.0 installs: --root-dir, --exclude-path, --exclude, --exclude-all-private, --max-concurrency, --no-progress. The action runs "eval lychee ${ARGS}", so $GITHUB_WORKSPACE expands at run time. Replace the job-level "continue-on-error: true" with "fail: false" on the lychee step. With continue-on-error the job still shows a red check on the pull request when a remote site rate-limits, which teaches people to ignore red. With "fail: false" the step exits 0, writes the full report to the job summary, and the check stays green. The job still fails on a real problem such as a broken build or an empty link set.
README.md calls .hugo-version the single source of truth, but the version
was also written in gh-pages.yml and twice in pr.yml. The next bump would
miss one of the four copies and the workflows would build with a different
Hugo than bin/install.
Add a step to each job that sets up Hugo. It reads .hugo-version into
HUGO_VERSION through $GITHUB_ENV, and the Setup Hugo step uses
${{ env.HUGO_VERSION }}. The file is now the only place the version lives.
The first run reported 313 link errors out of 15,960 links checked. 251 of them, 80 percent of the report, are one host: pybcn.slack.com answers a bot with 403 by design, and it is linked from the footer of every page. Excluding it leaves about 62 findings, which is a list someone will read. The remaining errors are real: 16 pages link to something that returns 404, and 11 linked domains no longer resolve at all. No internal link failed, which is what the --root-dir change was for.
themes/pybcn_theme/package.json is a leftover from the theme this one was copied from. It does nothing: - its css:build script reads source/scss/style.scss, and source/ does not exist in this repository - no lockfile and no node_modules are tracked, so nobody has ever installed from it - Hugo compiles the SCSS itself through toCSS in partials/head.html It is not inert, though. Dependabot alert 1 reports improper certificate validation in node-sass 4.13.1, which the manifest pins, and that alert stays open as long as the file is in the tree. Deleting the file closes it for good rather than bumping a dependency that nobody installs.
ber2
left a comment
There was a problem hiding this comment.
By looking at these changes I wonder whether we are mixing two artifacts:
- The repo with the static site builder code
- The static build of the site
If this is the case, should these two be separate?
What do we gain from publishing the code that generates the site?
As a developer, what do I gain from cloning a lot of static HTML before I make a contribution?
|
@ber2 they are already separate, just not as two repositories.
That is why the cost is smaller than it looks. The build output is 691 blobs and about 72 MB uncompressed that The weight is not the HTML anyway, it is the images. If we want the build out of the repository entirely, there is a cleaner route than a second repository: switch Pages from "deploy from a branch" to the GitHub Actions source, with |
|
Housekeeping note on the close and reopen in the timeline above. Nothing about this pull request's content changed. After #192 merged I deleted its branch, No commits were lost and the approval stands. The diff is now against For the rest of the stack the order is the other way round: retarget the next pull request first, delete the merged branch after. |
A contributor had no written path from a clone to a merged pull request, and no way to know what the checks do before one of them fails. The section lists the steps, the command for each check, and what each one catches. Asked for in review of #193.
A contributor had no written path from a clone to a merged pull request, and no way to know what the checks do before one of them fails. The section lists the steps, the command for each check, and what each one catches. Asked for in review of #193.
A contributor had no written path from a clone to a merged pull request, and no way to know what the checks do before one of them fails. The section lists the steps, the command for each check, and what each one catches. Asked for in review of #193.
Stacked on #192. Its base is
pr/1-fixes, so the diff shows only this pull request's own changes. GitHub retargets it toeditionautomatically once #192 merges.Nothing checks a pull request in this repository
There is no workflow on
pull_requestat all today, so a change that breaks the build or the content is found after it is merged and deployed. This adds three jobs:bin/check-content, which validates front matter and cross references. Oneditionit reports 6 errors; Fix what is silently broken #192 is what brings it to zero.permissions: contents: read, and every action pinned to a commit SHA with the release named in a comment next to it. A floating tag is mutable: whoever controls the action repository can movev5to new code, which then runs with the workflow's token. Pinning also movesactions/checkoutto v7.0.1 andactions/setup-pythonto v7.0.0, since each SHA is the latest release.The site no longer depends on servers we do not control
What this does not claim. The site still loads a small number of page specific embeds from other servers: six Unsplash hero backgrounds, two Google Forms, one Google Calendar, the PayPal button, and an Eventbrite button. The per page third party that tracked every visitor is gone; those embeds are not. The Unsplash ones are finished off in a later pull request in this stack.
Also
humans.txt,security.txtandllms.txt; a correct favicon set, including an.ico, an apple touch icon and a theme colour, and the old double slash in the URL is gone;robots.txtgeneration and git info enabled; the unusedtagsandcategoriestaxonomies turned off, which removes two empty index pages from the sitemap; Hugo build artifacts untracked;bin/publishretired.What changed after review
A review pass found three defects, all fixed here:
branches: [edition], andpull_request.branchesfilters on the base branch. Because these pull requests are stacked on each other, it would never have fired for this one or the two above it, and Fix what is silently broken #192 does not contain the workflow at all. The filter is gone.--root-dir, which lychee's documentation says makes it flag every absolute link in a local file as an error, and the self links were fetched from production rather than from the build, so a pull request adding a page would see 404s for its own new pages. The job now builds with--baseURL /and points lychee at that build: 7,536 production fetches become 13,295 local checks. The frozen snapshots and the hosts that answer bots with 400 or 403 are excluded.security.txtnamed an unreachable contact first. Private vulnerability reporting is not enabled on this repository, so the GitHub advisory form rejects reports from outside the organization. RFC 9116 treats contact order as precedence, so a dead entry first is the worst position. The line is commented out with a note to restore it once the setting is enabled, which is a repository setting and not a change in this pull request.Plus two loose ends:
.github/dependabot.yml, because a commit here promises that Dependabot can keep the SHA pins in step and nothing configured it; and both workflows now read the Hugo version from.hugo-version, which the README calls the single source of truth and which was in fact pinned in four places.Verified
Build with zero warnings, 317 HTML files,
bin/check-contentreports 0 errors and 8 warnings. Both workflow files parse, all tenuses:lines are 40 character SHAs with an accurate version comment, and the three jobs run what this description says they run.Two things to know before merging
These workflows have never run, because nothing has ever been pushed to this repository from this work. This pull request is their first execution. In particular, lychee could not be run locally, so its flags are verified against the v0.24.2 documentation and source rather than executed.
The checks do not block a merge yet.
editionhas no required status checks configured, so a red pull request can still be merged, andrequire_code_owner_reviewsis off. Both are organization settings worth enabling once this workflow has run green at least once.I will do the merge myself, in order.