Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions .github/workflows/gh-pages.yml
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,23 @@ jobs:
- name: Read the Hugo version
run: echo "HUGO_VERSION=$(tr -d '[:space:]' < .hugo-version)" >> "$GITHUB_ENV"

- name: Restore the generated images
# A cold build resizes more than 400 images, about 90 seconds of CPU
# for output that is the same every time. Hugo names each generated
# file after a hash of the source image and the processing options,
# so a cached file is used only when both match and a stale entry is
# never served. The key follows the image sources; when one changes,
# restore-keys brings the previous cache and only the new image is
# resized. A pull request can restore a cache saved on its base
# branch, and this workflow runs on edition, so the cache it saves
# is the one a new pull request starts from.
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
path: resources/_gen
key: hugo-gen-${{ hashFiles('themes/pybcn_theme/assets/images/**') }}
restore-keys: |
hugo-gen-

- name: Setup Hugo
uses: peaceiris/actions-hugo@2752ce1d29631191ea3f27c23495fa06139a5b78 # v3.2.1
with:
Expand Down
49 changes: 46 additions & 3 deletions .github/workflows/pr.yml
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,21 @@ jobs:
- name: Read the Hugo version
run: echo "HUGO_VERSION=$(tr -d '[:space:]' < .hugo-version)" >> "$GITHUB_ENV"

- name: Restore the generated images
# A cold build resizes more than 400 images, about 90 seconds of CPU
# for output that is the same every time. Hugo names each generated
# file after a hash of the source image and the processing options,
# so a cached file is used only when both match and a stale entry is
# never served. The key follows the image sources; when one changes,
# restore-keys brings the previous cache and only the new image is
# resized.
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
path: resources/_gen
key: hugo-gen-${{ hashFiles('themes/pybcn_theme/assets/images/**') }}
restore-keys: |
hugo-gen-

- name: Setup Hugo
uses: peaceiris/actions-hugo@2752ce1d29631191ea3f27c23495fa06139a5b78 # v3.2.1
with:
Expand Down Expand Up @@ -55,6 +70,21 @@ jobs:
- name: Read the Hugo version
run: echo "HUGO_VERSION=$(tr -d '[:space:]' < .hugo-version)" >> "$GITHUB_ENV"

- name: Restore the generated images
# A cold build resizes more than 400 images, about 90 seconds of CPU
# for output that is the same every time. Hugo names each generated
# file after a hash of the source image and the processing options,
# so a cached file is used only when both match and a stale entry is
# never served. The key follows the image sources; when one changes,
# restore-keys brings the previous cache and only the new image is
# resized.
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
path: resources/_gen
key: hugo-gen-${{ hashFiles('themes/pybcn_theme/assets/images/**') }}
restore-keys: |
hugo-gen-

- name: Setup Hugo
uses: peaceiris/actions-hugo@2752ce1d29631191ea3f27c23495fa06139a5b78 # v3.2.1
with:
Expand All @@ -64,7 +94,16 @@ jobs:
- name: Build
# A root-relative baseURL turns every internal link into "/path/", so
# lychee resolves it inside this build and not on the live site.
run: hugo --minify -d public --baseURL /
# -D builds the drafts: the canary fixture that bin/check-rendered
# reads is three draft pages. The deploy workflow builds without -D,
# so the fixture never reaches the live site.
run: hugo --minify -D -d public --baseURL /

- name: Check the built pages for dangerous markup
# Reads what Hugo wrote, so a template defect is caught here even
# when every content file is clean. Standard library only: the
# runner's python3 is enough.
run: bin/check-rendered public

- name: Check links
uses: lycheeverse/lychee-action@e7477775783ea5526144ba13e8db5eec57747ce8 # v2.9.0
Expand All @@ -73,7 +112,10 @@ jobs:
# every absolute link in a local file as an error.
# --exclude-path: the frozen 2016-2019 snapshots under archives/
# keep their dead links on purpose.
# --exclude: hosts that answer bots with 400, 403, or 999.
# --exclude: hosts that answer bots with 400, 403, or 999, and
# canary.invalid, the host the canary fixture names as text.
# lychee reads URLs out of text nodes too, and .invalid never
# resolves.
args: >-
--no-progress
--exclude-all-private
Expand All @@ -83,7 +125,8 @@ jobs:
--exclude 'twitter\.com'
--exclude '^https?://(www\.)?x\.com/'
--exclude 'linkedin\.com'
--exclude 'pybcn\.slack\.com'
--exclude 'pybcn\.slack\.com'
--exclude 'canary\.invalid'
'public/**/*.html'
# External sites rate-limit. Keep the report in the job summary and
# do not turn the check red.
Expand Down
90 changes: 89 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,47 @@ Once the site was archived, you can create a new link in the navigational menu u

## Information for developers

### Collaborate

A change to this site goes through a pull request. Nothing is pushed to
`edition` directly.

1. Clone the repository. `git clone --single-branch --branch edition` is enough
and skips the built site, which lives on its own branch.
2. Run `bin/install`. It downloads the pinned Hugo binary into `bin/hugo` and
verifies its checksum. You do not need Python, Go, or npm for this step.
3. Run `bin/serve` to see the site at `http://localhost:1313` while you work.
4. Make the change.
5. Run the three checks locally, so you find what the pull request would find:

```
pip install pyyaml
bin/check-content
bin/check-html-safety
bin/hugo --minify -D -d public && bin/check-rendered
```

6. Open the pull request against `edition`.

The `pr-checks` workflow then runs the same checks on your branch, plus a build
and a link check. **All of them have to pass**, and one approving review is
needed before the pull request can be merged. The paths listed in
`.github/CODEOWNERS` also request a review from the web team automatically.

What each check is for:

| Check | Fails when |
|---|---|
| Build the site | Hugo cannot build, or emits a warning |
| `bin/check-content` | Front matter does not parse, a person `id` does not match its filename, an id is duplicated, a declared photo is missing, or an event references a person or sponsor that does not exist |
| `bin/check-html-safety` | Content carries raw HTML that turns a content change into script execution, a redirect, a credential prompt, or a page overlay |
| `bin/check-rendered` | The built pages carry that same markup, which catches a template that produces it even when no content file does |
| Check the links | Reported, never blocking, because external sites rate-limit |

A merge to `edition` deploys the site. The `github-pages` workflow builds it and
publishes the result, so a change is live within a few minutes of the merge.
Comment on lines +85 to +123

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Excellent!



### Content safety check

`bin/check-html-safety` scans `content/` and fails if it finds raw HTML that
Expand Down Expand Up @@ -115,10 +156,52 @@ URL the element loads, so it permits one known embed and not any later one in
the same file, and the file that carries it needs a rule in
`.github/CODEOWNERS`, or any contributor can edit the allowed snippet.

### Rendered page check

`bin/check-rendered` reads the pages Hugo built and fails on dangerous markup
in them: an `on...=` handler on any element, a `javascript:`, `vbscript:` or
`data:` URL in any attribute a browser loads from or navigates to (entities
decoded, the way a browser does), an inline `<script>` or one loaded from off
this site, an `<iframe>`, `<object>`, `<embed>` or `<base>`, a form or a `<meta
refresh>` that sends the visitor off this site, a `style` attribute or `<style>`
block outside a small grammar of sizes, margins and grid properties, a `<link>`,
an image, a media file or a click beacon fetched from off this site, and the
three pieces of markup that `html.parser` and a browser read differently.

It does what the other two checks cannot. `bin/check-content` and
`bin/check-html-safety` read `content/`, which is what an outside contributor
can change, and that is the right place for them to stop. The worst defect
found on this site was not there: a partial assembled the sponsor logo markup
with `printf` and emitted it through `safeHTML`, so a sponsor name reached the
page unescaped while every content file was clean and both content checks
passed. Reading the output catches that class wherever it lives: a template, a
shortcode, a partial, the theme, a Hugo internal template or the minifier.

The check needs a build made with `-D`. Three draft pages are a canary: a
sponsor and a person whose name closes the attribute it sits in and opens a
link to `canary.invalid`, listed on `content/canary.md` through the real
sponsor and people grids. The check fails if either name did not reach that
page as text, or if the page is missing, so the fixture cannot disappear
without notice. The deploy workflow builds without `-D`, so the canary never
reaches the live site.

```
./bin/hugo --minify -D -d public
./bin/check-rendered public
```

The standard library is enough, no `pip install`. The `pr-checks` workflow
runs it in the `links` job, on the same build lychee checks. `ALLOWLIST` in
the script names each accepted finding by page, kind and either the prefix of
the URL the element loads or the SHA-256 of the inline text, with a reason:
today the Google embeds, the PayPal form with its button image, the
Eventbrite button of PyDay 2018, and two inline scripts in the frozen
archives. An entry that no longer matches anything is reported as
stale, so it gets deleted.


This is a work in progress. All the design and implementation decisions are detailed in [Proposta d'estructura](https://docs.google.com/document/d/10YxQeCuGQXUjnN3o9e1oH2HJkrhxJsr31qnxO_aiCNM/edit?usp=sharing) (currently written in catalan). All the tasks are managed through our [private Trello board](https://trello.com/b/cFE8KRTS).

For now, we are not looking for contributors yet.


### Using JS code
Expand Down Expand Up @@ -155,6 +238,11 @@ The template iterates over the defined levels, and for each level:

To add new people to the site, there is a Hugo archetype defined in `themes/pybcn_theme/archetypes/people.md` that can be used by running the command: `hugo new people/my-new-person.md`.

The `photo` field names a file under `themes/pybcn_theme/assets/images/people/`.
A remote `photo_url` is not supported: it would load from a third party on
every page that lists the person and tell that party who viewed it, and
`bin/check-content` reports it as an error.


### Monthly event page implementation details

Expand Down
48 changes: 47 additions & 1 deletion bin/check-content
Original file line number Diff line number Diff line change
Expand Up @@ -3,14 +3,16 @@

Catches the mistakes that a Hugo build does not: a person id that does not
match its filename, a duplicate id, a page that lists a person or sponsor id
with no page behind it, and a declared photo with no file behind it.
with no page behind it, a declared photo with no file behind it, and a
remote photo_url.

Exit code 1 if any error is found. Warnings do not fail the build.

Usage:
bin/check-content
"""

import datetime
import pathlib
import re
import struct
Expand Down Expand Up @@ -158,6 +160,12 @@ def check_photos(people):
small = []
for page_id, path in people.items():
data = front_matter(path) or {}
# A remote photo loads from a third party on every page that lists
# the person and tells that party who viewed it. The template no
# longer reads the field, so a value here would be ignored in silence.
if data.get("photo_url"):
error(path, "photo_url is not supported: put the file under "
"themes/pybcn_theme/assets/images/people/ and use 'photo'")
photo = data.get("photo")
if photo and not bare_file_name(path, "photo", photo):
continue
Expand Down Expand Up @@ -325,6 +333,43 @@ def check_orphans(people, listed_people):
warn(path, f"person '{page_id}' is not listed by any event or page")


SECURITY_TXT = ROOT / "static" / ".well-known" / "security.txt"
SECURITY_TXT_NOTICE_DAYS = 30


def check_security_txt():
"""Report a security.txt that has expired, or is about to.

RFC 9116 makes Expires mandatory and says a reader should treat the file
as stale once that date passes. Nothing else would notice: the file keeps
serving a 200 and keeps looking correct, and the date is a year away when
it is written, which is exactly how long it takes to forget it.
"""
if not SECURITY_TXT.exists():
return
expires = None
for line in SECURITY_TXT.read_text(encoding="utf-8").splitlines():
if line.lower().startswith("expires:"):
expires = line.split(":", 1)[1].strip()
break
if not expires:
error(SECURITY_TXT, "no Expires field, which RFC 9116 requires")
return
try:
when = datetime.datetime.fromisoformat(expires.replace("Z", "+00:00"))
except ValueError:
error(SECURITY_TXT, f"Expires is not a date this can read: '{expires}'")
return
left = (when - datetime.datetime.now(datetime.timezone.utc)).days
if left < 0:
error(SECURITY_TXT, f"expired {-left} days ago, on {when.date()}. "
"A reader treats the whole file as stale. Set a new "
"Expires, at most a year away.")
elif left <= SECURITY_TXT_NOTICE_DAYS:
warn(SECURITY_TXT, f"expires in {left} days, on {when.date()}. "
"Set a new Expires before it does.")


def main():
people = collect_ids("people")
sponsors = collect_ids("sponsors")
Expand All @@ -333,6 +378,7 @@ def main():
listed_people = check_listings(people, sponsors)
check_agenda()
check_orphans(people, listed_people)
check_security_txt()

for line in warnings:
print(f"WARNING {line}")
Expand Down
Loading
Loading