Skip to content

Phase 1: fix what is silently broken #190

Description

@DZPM

Approved by the Permanent Committee on 2026-09-30.

Every item is verified against the live site or the built output. None of them breaks the Hugo build, which is why they went unnoticed, and which is what the phase 2 issue exists to change.

The work is done locally and comes as several small pull requests. This issue closes with the last one.

PR: critical fixes

  • bin/install cannot install Hugo, so a new contributor cannot set the site up. The PyPI package now needs a Go toolchain
  • A speaker is dropped from PyDay 2020 and 2021: a person id does not match the filename the events reference
  • A person's published URL misspells their surname
  • A sponsor of PyDay 2021 and 2022 appears on neither page. The logo has been in the repository the whole time, only the sponsor file was missing

PR: image weight (closes #169 and #171)

  • One sponsor logo is 23,944 x 5,753 px, which needs about 551 MB of RAM to decode. The reader no longer sees it: the pipeline serves it at 560x135 and 2.5 KB. The source is still in the repository and every build decodes it. Cut the image sources down to what the build asks of them #206 cuts it and 25 more, and adds the check that stops the next one. Waiting for review
  • 132 person photos total 38.7 MB. The organizers page ships 8.97 MB of images, roughly 48 s on slow 4G
  • No width, height, loading="lazy" or srcset anywhere, because the files sit in static/ where Hugo cannot process them
  • Thirteen hero images are hotlinked from Unsplash. Seven use source.unsplash.com, which Unsplash retired: they return HTTP 503, so those pages have no hero image today. One more fetches a 3.6 MB original because it carries no size parameter

PR: accessibility

  • The association's email on /contact/ is an icon-only link whose icon is aria-hidden, so it has no accessible name. A screen reader user cannot find how to contact us
  • 2,937 links across the site have no accessible name
  • 46 of 46 images on the organizers page have no alt, including the logo inside the home link
  • 314 modals share a single id, so every dialog announces with the first one's title, and none restores focus on close
  • The agenda conveys difficulty by colour alone, and 13 of 22 cells fail text contrast
  • All 245 pages share one <title>, the home page has nine <h1>, and there is no <main> and no skip link
  • The carousel advances on its own and takes focus from a keyboard user every five seconds
  • The main menu is invalid markup, and a dropdown aria-labelledby points at an id no element carries

PR: plumbing

  • The monthly events page and the home lists spin forever against a Meetup API that returns 404
  • The home page promotes PyDay 2024 under a carousel promoting PyDay 2025
  • enableGitInfo and enableRobotsTXT are unset, so the sitemap has no dates and the robots.txt the theme already contains is never published
  • The favicon URL has a double slash, and the icon is a 79 KB PNG with no .ico
  • Font Awesome loads from a CDN on 289 pages. The repository already self-hosts Bootstrap and jQuery
  • The archived snapshots under /archives/ load third-party scripts, including a tag manager for an analytics property that stopped recording in 2023
  • The 404 page hotlinks a 489 KB image from Wikimedia
  • No humans.txt, no security.txt

PR: housekeeping

  • resources/_gen and .hugo_build.lock are tracked, so a build dirties the tree
  • bin/publish deletes master, runs git rm -rf over the source tree and force-pushes. CI replaced it in 2020
  • A dead analytics property in config.toml
  • A sponsor with no logo renders a broken image, and the build still succeeds
  • /events/ and /people/ return 200 with no listing. Half done: /people/ is a 404 since Fix sitemap, feed, people, sponsors #202. /events/ still answers 200 with the word "Events" and no links in it. It is a real section with real children, so the fix there is a listing, not a removal

PR: content safety

  • Agenda difficulty levels and the legend are raw HTML in front matter, passed through safeHTML. Render them from data instead, which also fixes the colour-only accessibility failure
  • Add CODEOWNERS for the paths where a bad change does most damage
  • Add a check that rejects dangerous raw HTML in content, with an allowlist for the legitimate embeds
  • Two person records claim the same board title at once, and two more share another, because a free-text title carries no term
  • content/pybcn_association/information.md calls the Permanent Committee "Standing Committee" in four places, while the rest of the site calls it Permanent Committee

Verified on production, 2026-10-08, after #202, #204, #207, and #208 went live. Every page of the sitemap was fetched and measured, rather than taken from the pull requests that claim to fix them. The sitemap is 40 URLs now, where it was 247: #202 stopped publishing the 206 person and sponsor pages that nothing linked to.

What was checked Then Now
Images on /pybcn_association/organizers/ 8.97 MB 0.25 MB in 22 files
Person photos in the repository 38.7 MB, 56 of them not square 12.00 MB, all 130 square
width, height, loading, srcset on that page none 46 of 46, 45 lazy, 44 with a srcset
Links with no accessible name 2,937 0
Images with no alt all of them 26, every one of them inside a frozen copy under /archives/
Pages with a repeated modal id every page 0
aria-labelledby pointing at a missing id the three menu sections 0
Hotlinks to Unsplash 13 0
Distinct <title> 1 for 245 pages 40 for 40 pages
<h1> per page 9 on the home page exactly 1 on every page
aria-haspopup and aria-labelledby on a role-less div 7 and 3 per page 0 and 0
<main> and a skip link neither both, on all 247
Agenda difficulty colour only the level in words, at 16.48:1 on all 67 badges
Carousel advances on its own data-interval=false

Two things the W3C validator found that are not on this list, for whoever picks up phase 2:

  • /events/pyday_bcn/pyday_bcn_2025/ has 50 errors, all of them from the raw HTML in the page body: a <style> inside <body> that is not the first child of its parent, a stray </b>, a </p> with no <p>, an </br>, and headings that skip a level.
  • 14 pages skip a heading level, /contact/ among them, which the validator reports as an error. A branch is ready for this one.

Activity

  1. self-assigned this
    on Oct 1, 2026
  2. DZPM commented on Oct 4, 2026

    @DZPM
    MemberAuthor

    Progress update. The first two pull requests are open.

    The grouping changed

    This issue lists the work as six pull requests. It is four, and they are stacked: each one's base is the previous branch, not edition. Two reasons.

    The order turned out to be mandatory, not a preference. Measured with git merge-tree: merged out of order, the third conflicts on .github/workflows/pr.yml and the fourth conflicts on seven files. Only the first is safe alone. Stacking makes GitHub enforce the order, and each diff still shows only its own changes.

    The checks have to come second, not last. There is no workflow on pull_request in this repository at all today, so if the tooling went last, nothing would watch the three pull requests before it. And two fixes exist only because of a later change, so they cannot be pulled forward: the content check has to learn about the image move, and the build needs a longer timeout once it resizes every photo.

    What is in the two that are open

    #192 covers the critical fixes and the housekeeping items: the installer, the dropped speaker and the misspelled URL, the missing sponsor file, the sponsor with no logo, the dead analytics property, the tracked build artifacts, bin/publish, and the missing listings for sections and taxonomies. It also replaces the placeholder text on the How to collaborate page, which is not in the list above and should have been: it had been live since December 2020 and contains racial slurs.

    #193 covers the plumbing: the Meetup API that returns 404, the home page promoting the wrong PyDay, enableGitInfo and enableRobotsTXT, the favicon, Font Awesome off the CDN, the third-party scripts in the archived snapshots, the Wikimedia hotlink, and humans.txt and security.txt. Plus the three check jobs this issue's sibling asks for.

    Still to come, this afternoon: accessibility and image weight in one, content safety in the other.

    Two corrections to this issue

    Numbers in the list above that I measured wrong when I wrote it:

    • "2,937 links have no accessible name" is the reduction, not the total. The site has 3,536. The fix brings the live site to zero; 599 remain inside the frozen snapshots under /archives/.
    • "13 of 22 agenda cells fail text contrast" is right, but it is four colours that fail, not six as I said elsewhere. Green, red, purple and grey already passed.

    What a review changed

    I had five review passes run over the four branches before opening anything. They did not pass. Found and fixed in the two that are open:

    • The new fallback index called the page header a second time, so all six new index pages had two <h1> and two id="main_content".
    • That index listed the two PayPal return pages.
    • bin/install leaked its temporary directory on failure, and would have died on macOS after a successful download.
    • The workflow would have checked nothing: the trigger was filtered to branches: [edition], and that filter matches the base branch, so it would never have fired on a stacked pull request.
    • The links job could not have passed: it checked the self-links against production instead of against the build, so a pull request adding a page would see 404s for its own new pages.
    • security.txt named the GitHub advisory form as its first contact, and private vulnerability reporting is not enabled on this repository, so that channel reaches nobody.

    One open question for reviewers

    content/sponsors/hybridtheory.md existed from November 2021 and was deleted in October 2024, inside a commit that was adding other sponsors and did not mention the removal. The commit message in #192 says the file was never created, which is wrong and which I will correct.

    Was that deletion deliberate? If the sponsorship was intentionally retired, the right fix is to remove the three remaining id references rather than restore the page.

    Two settings that are not code

    Neither check bites at merge time yet. edition has no required status checks, so a red pull request can still be merged, and require_code_owner_reviews is off, so the CODEOWNERS file that the content-safety pull request adds would have no effect. Both are worth enabling now. The workflow has run green on #193: 15,960 links checked, no internal link failed, and 55 external findings, of which 16 are pages linking to something that returns 404 and 11 are domains that no longer resolve. That last list is material for a phase 3 issue, not for these pull requests.

    Unrelated to this issue but relevant to the same cleanup: the wildcard DNS record on pybcn.org is gone as of this morning, so the subdomains that were serving other people's content no longer resolve.

  3. DZPM commented on Oct 4, 2026

    @DZPM
    MemberAuthor

    @rberenguel one for your memory, if you have it.

    content/sponsors/hybridtheory.md was added in November 2021 and removed in October 2024, inside a commit about adding other sponsors that does not mention the removal. The id is still referenced in pyday_bcn_2021.md (Gold), pyday_bcn_2022.md (Supporting) and sponsors/_index.md, and the logo never left the repository. So the grid on both PyDay pages has silently skipped it for two years: an id with no page renders nothing, without a warning.

    Was that removal intentional? If it was an accident, #192 is right and restores the page. If the sponsorship was deliberately retired, I should remove the three references instead and I will swap that commit.

    My reading is accident, because a deliberate retirement would have taken the references and the logo with it. But that is inference from a diff. If nobody remembers, I will restore it and say so in the pull request.

    Either way the web link is already gone: hybridtheory.com now serves a zero second redirect to an Azerion press release.

  4. rberenguel commented on Oct 4, 2026

    @rberenguel
    Contributor

    @DZPM I assume it was not intentional. In 2024 I was not helping with PyDay, so I have no idea why it would have been removed, I don't see a reason. They were sponsors, they should have stayed such. AFAICT that deletion was committed to the branch directly with no PR, so no idea.

  5. DZPM commented on Oct 4, 2026

    @DZPM
    MemberAuthor

    All four pull requests are now open and green.

    #195 in short

    It is the largest of the four and the one a visitor notices.

    Accessibility. Links with no accessible name go from 3,536 to 599, and the 599 that remain are all inside the frozen snapshots under /archives/: the live site is at zero. Four of the agenda colours failed text contrast, which is 13 of the 22 cells on PyDay BCN 2025, and every cell now passes 4.5:1. Plus landmarks and a skip link, one h1 per page, a real <title> per page where all of them used to share one, a unique label on each dialog with focus restored on close, alt text on the people cards, a focus ring that goes from 1.39:1 to 19.59:1, and a carousel that no longer advances on its own or steals focus every five seconds.

    Images. Every photo and logo now goes through the Hugo image pipeline, which needs the files under assets/ rather than static/, and comes out as WebP with a srcset and explicit dimensions. The built site goes from 78.8 MB to 25.9 MB. The organizers page drops from 9.12 MiB of images to 0.39. Thirteen hero images were hotlinked to Unsplash, seven of them to a host Unsplash retired that answers 503, so those pages had no hero image at all; all thirteen now use the association's own photos.

    This issue's checklist

    Closed by #195: #169, the weight of the PyDay participant images. Hugo does have the auto resize this issue guessed at, and that is what the branch uses.

    Partly closed: #171, the carousel. The weight half is done: canodrom_header.png is 945,496 bytes in the repository and now reaches a phone as a 20,178 byte WebP and a desktop as 51,490. The other half is not done. That issue also asks for newer photos, and choosing which photos is a content decision, not a code one. Somebody has to pick them. I would keep #171 open for that after #195 merges.

    Two things worth saying out loud

    A review found a regression this branch itself had introduced. The first version derived a talk title's language from the event's language field, which describes the language the talk is given in, not the language of its title, so 13 English titles ended up tagged as Spanish or Catalan and were read aloud with the wrong voice. They had been read correctly before. Each title now carries its own explicit value. It is the clearest example of why these went through a review pass before being opened.

    #195 needs a squash merge, and the pull request says so. Six commits in its history carry an HTML injection that two of its own changes create between them and that a later commit fixes. Squashing means no state of edition ever holds it.

    Merge order

    The four are stacked, each on the one below, so they have to merge in order: #192, #193, #194, #195. Measured out of order, #194 alone conflicts on the workflow file and #195 alone conflicts on seven files. Only #192 is safe on its own. GitHub retargets each one automatically as the one below it merges.

    One caveat for #195: squash it after the other three have merged. Squashed earlier it would swallow their commits too.

  6. DZPM commented on Oct 4, 2026

    @DZPM
    MemberAuthor

    One more piece of repository archaeology, and @ifosch this one is yours if you happen to remember it, which I would not expect.

    The membership and the cancel pages each load a 1x1 GIF from paypalobjects.com on page load, before anyone clicks anything:

    <img alt="" border="0" src="https://www.paypalobjects.com/es_ES/i/scr/pixel.gif" width="1" height="1">

    It arrived in e77b443, "Publish new site with PyDay 2020", 2020-12-03, which changed 434 files, inside the button snippet that PayPal's own generator produces. A later commit only moved the block from membership.md into a shortcode, so git blame points somewhere misleading. Nobody here ever decided anything about it, which is why I am asking rather than assuming.

    The question: do you know of anything that reads it? If yes I put it back, no argument.

    What I could establish on my own, for the record:

    • 43 bytes, a 1x1 GIF, no query string and no per merchant or per page parameter, cached publicly for a year, and it sets no cookie. The same file for every merchant that uses PayPal, so it cannot distinguish a merchant, a page or a session.
    • A plain <img> with no name, so it is never submitted. Removing it cannot affect a payment: cmd, hosted_button_id and submit are untouched, and the button is the <input type="image">, which stays.
    • PayPal ships it in every generator sample from the 2008 Website Payments Standard guide to the 2019 classic docs, and documents no purpose for it anywhere. It is absent from the HTML variables reference. In their Form Basics example every other line of the snippet carries an explaining HTML comment and this one does not.
    • The only statement of purpose that exists anywhere is a 2009 post on PayPal's community forum by a volunteer, not staff, saying it is for internal traffic monitoring and is optional. Every later summary on the web traces back to that one sentence.
    • PayPal's current button generator emits a JavaScript SDK snippet with no form and no pixel.

    So the honest claim is "PayPal has never documented a function for it", not "PayPal confirms it is unused". That is why this is a question and not a statement.

    And I should be straight about how little removing it buys: the button image is fetched from the same host on the same page load, with the same headers and the same absence of a cookie, so PayPal already learns of the page view either way. The gain is one request and the removal of something nobody can explain, not privacy.

    A separate and much bigger thing, found while looking

    Third parties report that PayPal is retiring Website Payments Standard, which is the flow this button uses: the Buy Now and Add to Cart buttons deprecated in January 2026 and WPS ceasing in January 2027, with Donations and Subscribe buttons claimed to be unaffected. PayPal's own Upgrade Hub calls WPS "legacy" but publishes no dates, so I could not confirm any of it from a PayPal primary source.

    Our button is a legacy-flow Subscribe button. If that exemption is not real, or stops being real, membership payments stop working and nobody finds out until somebody tries to join.

    That deserves its own issue and someone asking PayPal directly whether a WPS Subscribe button is still supported and until when. I will open it separately rather than bury it here.

  7. DZPM commented on Oct 5, 2026

    @DZPM
    MemberAuthor

    #192 merged today and the site is deployed. This is the first deploy since 25 April 2026.

    Verified on https://pybcn.org:

    • /pybcn_association/ serves text/html. It was serving XML, which is what sent me looking at this in the first place.
    • Hybrid Theory is back on PyDay BCN 2021, PyDay BCN 2022, and the sponsors page.
    • Christian Adell is back on PyDay BCN 2020 and PyDay BCN 2021.
    • The two renamed person files keep their old URLs through aliases, so no external link breaks.

    #193 merged as well, so every pull request now builds, checks its content, and checks its links. The remaining three are #194, #195 and #196, all green and waiting for a review.

  8. DZPM commented on Oct 5, 2026

    @DZPM
    MemberAuthor

    Seventeen of the thirty-four boxes above are now ticked. Each one was checked against edition or against the live site rather than against the pull request that claimed it, and the seventeen left are honest: they are in #194, #195 and #196, which are not merged.

    Two notes on items that are not ticked, because the reason is not "waiting on a pull request".

    /events/ and /people/ still have no listing. They return 200 with a correct <title> and a single <h1>, which they did not have before, but neither page links to anything: zero internal links on both. The fallback index gave them a header, not a directory. For /people/ that may be the right answer rather than an oversight, since a browsable directory of named people is a decision with its own consequences, so it is left for someone to decide rather than quietly closed.

    The two person records claiming the same board title is a content problem, not a template one, and nothing in the open pull requests addresses it.

  9. DZPM commented on Oct 5, 2026

    @DZPM
    MemberAuthor

    Status, and what is blocking what.

    17 of 34 done. All verified against edition or the live site, not against the pull request that claimed them.

    The 17 left:

    The blockers

    #194 and #195 have had no review and have been green since this morning. Between them they close 15 of the 17 remaining boxes.

    They also block everything else: #196, #199 and #200 are approved and cannot merge until these two do, because each is based on the one below it.

    #194 first. It is the one about raw HTML reaching the page, so it is the one worth reading carefully.

  10. DZPM commented on Oct 5, 2026

    @DZPM
    MemberAuthor

    Three things found while working through this list are not tracked here, because they are one subject and they have their own issue: the sitemap and the feeds render as raw XML, every feed item is dated 0001-01-01, and 206 of the 247 URLs in the sitemap are pages nothing links to. They are all in #201, which one pull request will close.

  11. ifosch commented on Oct 5, 2026

    @ifosch
    Member

    One more piece of repository archaeology, and @ifosch this one is yours if you happen to remember it, which I would not expect.

    The membership and the cancel pages each load a 1x1 GIF from paypalobjects.com on page load, before anyone clicks anything:

    This blame shows it was @lpmayos. I also guess she will unlikely remember why, but still mentioning here just in case. Thanks!

  12. ifosch commented on Oct 5, 2026

    @ifosch
    Member

    Git blame shows the most recent change.

    According to this: https://github.com/pybcn/pybcn.github.io/commits/edition/content/pybcn_association/membership.md

    This was the initial commmit: https://github.com/pybcn/pybcn.github.io/blame/e77b44366d388934131f5ff16e033ebac61876d7/content/pybcn_association/membership.md

    My bad, sorry! As you expected, I don't remember where this came from... Sorry @DZPM @lpmayos 🙏

  13. added a commit that references this issue on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions