Skip to content

Fix sitemap, feed, people, sponsors - #202

Open
DZPM wants to merge 9 commits into
editionfrom
pr/9-fix-xml-surface
Open

DZPM wants to merge 9 commits into
editionfrom
pr/9-fix-xml-surface

Conversation

@DZPM

@DZPM DZPM commented Oct 5, 2026

Copy link
Copy Markdown
Member

Closes #201.

The machine-readable surface of the site: what it publishes for crawlers and feed readers, and what a person sees if they open one of those pages.

What a person sees

/sitemap.xml and /index.xml rendered as raw XML. They are the only machine-facing pages anyone opens by hand, usually while debugging, and they were the ones with no presentation at all. Two XSLT 1.0 stylesheets now render each as a table, headed by the association logo, linking back to the site. A search engine and a feed reader ignore the stylesheet and read the XML, which is unchanged apart from the xml-stylesheet instruction.

What the feed says

Every item carried the date Mon, 01 Jan 0001 00:00:00 +0000, because no content file has a date in its front matter. Items are now dated and ordered by .Lastmod, which comes from git through the enableGitInfo that landed with #193, the same source the sitemap's lastmod already used.

That is a modification date, not a publication date, so the feed answers "what changed recently" rather than "what is new". A first-publication date is not recoverable here: 17 person files share 2020-12-03 and 19 share 2023-05-25 from bulk imports, and a renamed file reports the date of the rename. A [frontmatter] fallback to :git was tried and rejected, because it stamps article:published_time with that same date onto every HTML page.

How many feeds the site has

One, at /index.xml, advertised from every page and linked from the footer.

It had ten, because Hugo emits one per section by default and nobody chose them: /contact/ with 0 items, /people/ with 138 all lacking a title, /sponsors/ with 68 of the same, and six more. The home feed already carried every usable item they had.

What it stops publishing

The 138 person pages, the 68 sponsor pages, and /people/. No content file is deleted: this is _build.render = "never", and the grids and the modals read the same front matter as before.

Nothing on the site linked to any of them, and they were 206 of the 247 URLs in the sitemap. 73 of the person files carry a bio that rendered, and those stop too, because publishing only the people who happen to have a paragraph written about them would sort the organizers and speakers into two classes decided by whether a volunteer once had time to write it up. They come back with the Archive, which gives each person a page built from the talks they gave and the roles they held.

/sponsors/ stays: it is a real page with links into it.

Two URLs that now return 404

/people/christian-adell-querol/
/people/kemalkan-bora/

Both are aliases added in #192 this week, when two person files were renamed, so that an external link to the old URL kept working. They redirected to the person page, which this pull request stops publishing. A redirect to a page that is not built is worse than a 404, and there is nowhere honest to send them: neither person has a page any more, and the organizers page lists neither, both being speakers.

Verified

On a minified and an unminified build against pr/8-photo-names:

Before After
Sitemap URLs 247 40
Feed items 237 31
Feed items with no title 206 0
Feed items dated 0001 237 0
XML files in the build 11 2

The 207 URLs that left are /people/, the 138 pages under it, and the 68 under /sponsors/. Nothing was added.

/pybcn_association/organizers/, /pyladies_bcn/organizers-and-speakers/, /events/pyday_bcn/pyday_bcn_2025/ and /events/pyday_bcn/pyday_bcn_2021/ are byte-identical before the commit that adds the footer links; /sponsors/ differs only by its removed feed link. After that commit every page differs by the head link and the two footer anchors, and by nothing else.

bin/check-content, bin/check-rendered and bin/check-html-safety pass on both builds. Chromium and Firefox both apply the stylesheets.

Review notes

Nine commits, each one its own argument, written to be read in order.

Two things a reviewer may want to check by hand: open /index.xml and /sitemap.xml in a browser, and confirm the organizers page is unchanged.

One known limitation: if sitemap.xsl or rss.xsl ever failed to deploy, a browser would show a blank page rather than the raw XML. That is how browsers treat a missing xml-stylesheet target. Feed readers and crawlers are unaffected. A build check that every referenced stylesheet exists would be worth having, and belongs in its own change.

@ber2 ber2 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Quick confirm here.

There will be a bunch of /people/<id>/ now returning 404 and we are happy with this: they are low value pages, GDPR, etc.

@DZPM

DZPM commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Quick confirm here.

There will be a bunch of /people/<id>/ now returning 404 and we are happy with this: they are low value pages, GDPR, etc.

Nope! The cards for "people" only render inside modals, they don't have their own page, so no url to return a 404.

If we add more "people" when adding meetup events, then I'm for adding those. That's why I standardised the format, name_surname, fields, etc.

@DZPM
DZPM force-pushed the pr/8-photo-names branch from 52be25a to 97382f6 Compare October 7, 2026 09:47
DZPM added 9 commits October 7, 2026 11:52
/sitemap.xml and the RSS feeds are the machine-facing pages a person
opens, usually while debugging, and a browser showed them as raw XML.

Add static/sitemap.xsl and static/rss.xsl, two XSLT 1.0 stylesheets (the
version browsers implement) that render the XML as a heading, a count,
and a table, newest first. They are written as one pair: the same heading
shape ("PyBCN sitemap", "PyBCN feed", "PyBCN feed: Events" for a section
feed, taken from the part of the channel title before " on "), the same
intro sentence, the same CSS block, the same column widths, and the same
date form. The sitemap carries W3C datetimes and the feed carries RFC 822
dates, as each standard requires; both are shown as "2026-10-05 18:27
+02:00", which is string work with substring and a twelve-way choose for
the month name, since XSLT 1.0 has no date function. The second column
differs on purpose: a sitemap entry has only a URL, a feed item has a
title. The CSS is inline and small, uses the site colours and font stack,
and loads nothing from a third party. Both pages carry robots noindex.

Add layouts/sitemap.xml and layouts/_default/rss.xml to override the Hugo
internal templates. Each is a copy of the template that ships with Hugo
0.139.4 plus one line, the xml-stylesheet processing instruction after
the XML declaration. Search engines and feed readers ignore the
instruction and read the XML as before.

Verified on the minified build against the previous commit:
- sitemap.xml and the ten index.xml feeds are byte-identical apart from
  the processing instruction (11 files compared after stripping it).
- sitemap.xml: 247 URLs before and after.
- Chromium applies both stylesheets and shows the table instead of the
  XML tree, for /sitemap.xml, /index.xml, and /events/index.xml.
Every item of every feed was dated Mon, 01 Jan 0001, and the items came
out in title order. No content file carries a date in its front matter,
so .Date and .PublishDate are zero on every page.

.Lastmod is not zero: enableGitInfo is on, so it resolves to the date of
the last commit that touched the file, which is what the sitemap already
prints as lastmod. The feed template now sorts the items by .Lastmod,
newest first, and prints .Lastmod as the pubDate.

It is a modification date, not a publication date, and the feed should
be read as "what changed recently" rather than "what is new". The real
publication date is not in git either: 17 person files share 2020-12-03
and 19 share 2023-05-25, because those were bulk imports, and a renamed
file reports the date of the rename unless the history is followed. A
date worth the name comes from the events a person spoke at, which is
Archive work.

A [frontmatter] block that lets .Date fall back to :git was tried first
and rejected: it makes .PublishDate non-zero on every page, so the Open
Graph template adds article:published_time with a last-commit date to
every HTML page, and the sitemap reorders because the default page sort
uses .Date. Reading .Lastmod in the feed template changes nothing else.

Verified on the minified build against the previous commit:
- 0 pubDate elements dated 0001 across the ten feeds (was 237 in
  /index.xml alone).
- The first four items of /index.xml are dated 05 Oct, 04 Oct, 04 Oct,
  and 02 Oct 2026, in that order.
- diff -rq between the two builds lists only the ten index.xml files.
Hugo builds an RSS feed for every section by default, so the site had
ten: /index.xml and one per section. Nine of them are a by-product that
nobody asked for and nothing reaches: /contact/index.xml had zero items,
/people/index.xml had 138 items with an empty title (those files carry
name, not title), /sponsors/index.xml had 68 of the same, and the other
six repeated a subset of the home feed. The only link to each one was
the rel=alternate in the head of its own section page.

One feed is a simple way to see which pages have changed, and the home
feed already carries every regular page of the site, so it is the one
that stays. section = ["HTML"] under [outputs] drops the other nine.

Verified on the minified build against the previous commit:
- find <build> -name '*.xml' lists two files: index.xml and sitemap.xml
  (was eleven).
- grep -rlo 'rel=alternate' --include=index.html lists only the home
  page (was ten pages).
- diff -rq between the two builds lists the nine removed index.xml files
  and the nine section index.html files, each of which differs only by
  the removed link rel=alternate element. Nothing else moved.
- sitemap.xml: 247 URLs before and after (feeds are not in it).
Every file under content/people/ and content/sponsors/ exists to feed the
grids and the modals through its front matter. Hugo also built a page from
each one. No page on the site linked to any of them, and they were 206 of
the 247 URLs in the sitemap (#201).

The sponsor pages are empty: 67 of 68 have no body, the one exception has
125 characters. The person pages are not all empty: 73 of the 138 files
carry a bio in the body, up to 1,647 characters, and it did render. They
all stop, bio or not. Publishing only the people who happen to have a
paragraph written about them would sort the organizers and speakers of
this association into two classes, decided by whether a volunteer once had
time to write it up. A site that lists everyone equally in the grids, and
then gives a page to some of them and not to others, is worse than one
that gives a page to nobody yet. The pages come back with the Archive,
when every person has one worth linking to. Every content file stays
exactly as it is: this is _build.render = "never", not a deletion.

Two cascade blocks in config.toml set render = "never" on /people/, on
every page under it, and on every page under /sponsors/. The default
list = "always" is kept on purpose: people-grid.html and
sponsors-grid.html look the pages up in site.Pages by id, and
list = "local" takes them out of site.Pages, which is what emptied the
organizers grid in an earlier attempt. A page that is not rendered has no
permalink, so the feed template now skips pages without one; without that
the home feed listed 206 items with neither a title nor a link. /people/
had nothing but a heading and goes with them. /sponsors/ stays: it is a
real page, linked from the navigation of every page: 253 links on the
live site when #201 was written, 44 anchors on 42 pages in this build.

llms.txt and humans.txt no longer point at /people/ or claim one page per
person or sponsor.

Verified on the minified build against pr/8-photo-names:
- sitemap.xml: 247 URLs before, 40 after. The 207 dropped are /people/,
  the 138 pages under it, and the 68 pages under /sponsors/. Nothing was
  added.
- /index.xml: 237 items before, 31 after, none with an empty title or an
  empty link (was 206 of each).
- <build>/people/ is not built. <build>/sponsors/ holds only index.html.
- /pybcn_association/organizers/, /pyladies_bcn/organizers-and-speakers/,
  /events/pyday_bcn/pyday_bcn_2025/, and /events/pyday_bcn/pyday_bcn_2021/
  are byte for byte identical to the pr/8-photo-names build.
- /sponsors/ differs only by the removed link rel=alternate to its feed.
- bin/check-content (0 errors), bin/check-rendered --without-canary, and
  bin/check-html-safety pass.
Both were added earlier today, when two person files were renamed, so that
an external link to the old URL kept working. They redirected to the person
page, which the previous commit stops publishing, so they now redirect to a
page that is not built.

The two old URLs return 404 from here on:

  /people/christian-adell-querol/
  /people/kemalkan-bora/

A redirect to a page that does not exist is worse than a 404, and there is
nowhere honest to send them instead: neither person has a page any more, and
the organizers page does not list either of them, both being speakers rather
than organizers.

Verified: neither alias directory is built, and the content check still
reports no errors.
The feed and the sitemap are readable in a browser now, and nothing on
the site pointed a person at either. The only reference to the feed was
the link rel=alternate in the head of the home page.

Footer: two icons after the social ones, fas fa-rss and fas fa-sitemap
from the vendored Font Awesome 5.8.2, at fa-2x like their neighbours.
They are not in params.social_items: those are accounts elsewhere,
labelled "PyBCN on X", and they open a new tab. The feed and the sitemap
are pages of this site, so each one has its own label ("RSS feed of
PyBCN", "Sitemap of PyBCN", as aria-label and title) and opens in the
same tab. They share the icon row so they get the same size, colour, and
alignment as the icons next to them, with no new CSS.

Head: every page advertises the site feed, site.Home.OutputFormats.Get
"rss", instead of its own. Since the previous commits the home feed is
the only one, so a section page had nothing to advertise, and a feed
reader should find the feed from wherever the visitor is.

Verified on the minified build against the previous commit:
- Every page with a head carries one link rel=alternate, to
  https://pybcn.org/index.xml: 42 of the 77 index.html files, plus
  404.html. The other 35 have no head partial: 9 alias redirects and 26
  pages of the two static archive mirrors (was 1 page).
- Chromium: the two icons render in the footer at 32px high, the same as
  the six social icons, on the same baseline; /index.xml and
  /sitemap.xml answer 200 with application/xml.
- /pybcn_association/organizers/ differs from the previous build only by
  the head link and the two footer anchors.
- bin/check-content, bin/check-rendered --without-canary, and
  bin/check-html-safety pass.
The header of the stylesheet still said it serves /index.xml and the
index.xml of each section. The section feeds are gone since the commit
that set section = ["HTML"], so the sentence was no longer true. The
section heading logic stays, for a feed that may come back.
The feed table wrote the title of each item as the text of its link. An
item with an empty title gave an anchor with no text: nothing to see and
nothing for a screen reader to announce. No item of the site feed is in
that state since the feed filters out the pages that are not rendered,
but a content file that carries name instead of title would put one
there again. Such a row now shows the link itself.
Both pages carried a title and nothing else, so a person who landed on one
had no way back to the site and nothing saying whose site it was. They now
open with the logo and the title, and the whole block links to the home page.

The logo is the one in assets, processed by Hugo like every other image on
the site, which is why the two stylesheets move from static/ to assets/xsl/
and are run through ExecuteAsTemplate: a static file cannot know the
fingerprinted name of a processed image. The alternative was a second copy of
the logo in static/ with a fixed name, which would go stale the day someone
changes the first one.

The header is white. The logo carries the word "python" in the same blue as
the band it used to sit on, so on blue the word disappeared and only
"barcelona" was legible. White is also what the site's own navigation bar
uses behind it.

resources.Minify is not in the pipeline: Hugo has no media type for .xsl and
the step fails silently, leaving the resource unpublished.

Verified on both the minified and the unminified build: 31 feed items, 40
sitemap URLs, both documents still start with the XML declaration, and
Chromium loads the logo at 58px in both headers.
@DZPM
DZPM force-pushed the pr/9-fix-xml-surface branch from e127979 to 71ed7e4 Compare October 7, 2026 09:53
@DZPM
DZPM changed the base branch from pr/8-photo-names to edition October 7, 2026 09:53
@DZPM
DZPM requested a review from ber2 October 7, 2026 10:16
@DZPM DZPM mentioned this pull request Oct 7, 2026
30 of 34 tasks

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Stop publishing the person and sponsor pages until the Archive gives them content

2 participants