Skip to content

Stop publishing the person and sponsor pages until the Archive gives them content #201

Description

@DZPM

Every file under content/people/ and content/sponsors/ becomes a page, because that is what Hugo does with a content file. Those files exist to feed the grids and the modals: the card a visitor reads is the modal on the organizers page or on an event page.

Nothing on the site links to any of those pages. 206 of the 247 URLs in https://pybcn.org/sitemap.xml are them, 138 people and 68 sponsors, so 83% of what we offer search engines cannot be reached by anyone browsing.

HTTP Visible text Inbound links
https://pybcn.org/people/david/ 200 0 characters 0
https://pybcn.org/sponsors/21buttons/ 200 0 characters 0
https://pybcn.org/people/ 200 6 characters, the word "People" 0

Correcting the first version of this issue

It said these pages carry no content. That is true of the sponsors, where 67 of 68 are empty and the one exception is 125 characters. It is not true of the people: 73 of the 138 have a bio in the body of the file, up to 1,647 characters, median 388. I measured two pages, found both blank, and wrote the sentence as if it covered all 206.

That bio does render, which is why /people/christian-adell/ has text and /people/david/ does not. It is unreachable either way: no page links to it.

What to do

Stop publishing all of them, the 73 with a bio included, and keep every content file exactly as it is. This is _build.render: never, not a deletion.

Publishing only the people who happen to have a paragraph written about them would sort the organizers and speakers of this association into two classes, decided by whether a volunteer once had time to write it up. A site that lists everyone equally in the grids, and then gives a page to some of them and not to others, is worse than one that gives a page to nobody yet.

/people/ goes with them: it returns 200 with a heading and nothing else. /sponsors/ stays: it is a real page with 253 links into it.

The sitemap and the feeds are unreadable, and the feeds have no dates

Two more by-products of the same surface, found while measuring the above.

A browser shows /sitemap.xml and /index.xml as raw XML. They are the only machine-facing pages a person ever opens, usually while debugging, and they are the ones with no presentation at all. An XSL stylesheet fixes that for a reader and changes nothing for a search engine or a feed reader, which read the XML and ignore the instruction.

Every item in every feed is dated Mon, 01 Jan 0001 00:00:00 +0000. No content file carries a date in its front matter, so .Date has no source. The sitemap is unaffected because its lastmod comes from git, through the enableGitInfo that landed with #193. A [frontmatter] block giving date the same :git fallback fixes all 237 items.

That is a modification date, not a publication date. The real one is not in git either: 17 person files share 2020-12-03 and 19 share 2023-05-25, because those were bulk imports, and a renamed file reports the date of the rename unless the history is followed. A publication date worth the name comes from the events a person spoke at, which is Archive work.

There are ten feeds, and no visible link to any of them. Hugo emits one per section by default, so nobody chose them: /contact/index.xml has zero items, /people/index.xml has 138 with no title and no link, /sponsors/index.xml has 68 of the same. Each page does advertise its own section's feed with <link rel="alternate"> in the head, which a browser or a reader can follow, but no anchor anywhere on the site points at a feed, so a person cannot find one. The site keeps one, at /index.xml, advertised from every page and linked from the footer, and drops the other nine.

They come back with the Archive

The Archive planned for this repository gives each person a page built by crossing the talks they gave, the events they organised, and the roles they held, with dates. The 73 bios are part of what it will show. When every person has a page worth linking to, these URLs come back, and that is also the moment to decide what they should be called.

Activity

  1. self-assigned this
    on Oct 5, 2026
  2. changed the title [-]The site publishes 209 empty pages, and submits them to search engines[/-] [+]The site publishes 206 empty pages, and submits them to search engines[/+] on Oct 5, 2026
  3. DZPM commented on Oct 5, 2026

    @DZPM
    MemberAuthor

    Correcting my own figures above. I first counted these with grep -c, which counts matching lines rather than matches, against a build where the sitemap happened to carry newlines. Counted properly against the live sitemap, it is 206 of 247, 138 people and 68 sponsors, so 83% rather than 85%. The title and the body are updated.

    The argument does not change, and neither does anything else in the issue: no page links to any of them, and they carry no text.

  4. changed the title [-]The site publishes 206 empty pages, and submits them to search engines[/-] [+]Stop publishing the person and sponsor pages until the Archive gives them content[/+] on Oct 5, 2026
  5. DZPM commented on Oct 5, 2026

    @DZPM
    MemberAuthor

    Correcting a sentence I added above. I wrote that nothing links to any feed. Each page does advertise its own section's feed with a <link rel="alternate"> in the head; what does not exist is a visible link, so a person cannot find a feed but a browser can. My measurement was wrong rather than my conclusion: the minified HTML writes the attribute unquoted, rel=alternate, and I grepped for rel="alternate", which matches nothing. Counted properly it is 8 pages in the current build.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions