Technology · · 8 min read

Letting Robots Update Your Homepage

My homepage shows what I'm listening to and reading, updated every Monday by a GitHub Action that commits to its own repository. Live data on a static site, no server required.

Photograph © Ken Reid

Letting Robots Update Your Homepage

14 August 2026 · technology

My homepage shows my scrobble count, the book I'm reading through, the last song I played, and whatever I last posted on Bluesky. This is a static site, meaning it can't just run dynamic scripts on a whim. I set up a system to dynamically update my HTML weekly, however, and it works for free. Part of the series on how this site is built.

Quick jargon guide

  • GitHub Actions: GitHub's free automation. You describe a job in a text file, and GitHub runs it on their machines, on a schedule if you like.
  • Cron: the venerable syntax for "run this at these times". Mine says 0 6 * * 1: every Monday at 06:00 UTC.
  • API: a website's machine-readable front door. Last.fm's API answers "what did this user play?" with data instead of a web page.
  • API key / secret: the password for that front door, stored in GitHub's encrypted secrets so it never appears in the code.
  • Scrobble: one logged song play. My Last.fm account has been counting since 2011.
  • Bot commit: a change to the repository made by an automation rather than a person.

The idea: if the site can't fetch, fetch into the site

A normal website with live data has a server that queries things when visitors arrive. A static site can't do that, but it has a loophole: the site is a git repository, and anything that can commit to the repository can change the site. So instead of fetching data per visit, a scheduled job fetches it once a week and commits the results as plain JSON files. The site stays static. Visitors' browsers read those JSON files with an ordinary fetch: same origin, no keys, no rate limits, nothing that can fail separately from the site itself.

The whole pipeline is one workflow file and three Python scripts. The workflow checks out the repository, runs the two fetch scripts, checks what they wrote, and commits whatever changed (the comments are trimmed out here):

name: Refresh live data

on:
  schedule:
    - cron: "0 6 * * 1"   # weekly, Mondays 06:00 UTC
  workflow_dispatch:

permissions:
  contents: write

concurrency:
  group: live-data
  cancel-in-progress: false

jobs:
  refresh:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    env:
      FILES: >-
        data/lastfm.json data/now.json data/lastfm-history.json
        data/topalbums.json data/books.json data/reading.json
    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"

      - name: Fetch listening and currently-reading
        id: now
        continue-on-error: true
        env:
          LASTFM_API_KEY: ${{ secrets.LASTFM_API_KEY }}
        run: python .github/scripts/refresh_now.py

      - name: Refresh the read shelf
        id: books
        continue-on-error: true
        run: python .github/scripts/refresh_books.py

      - name: Validate the data files
        id: validate
        run: python .github/scripts/validate_data.py $FILES

      - name: Commit what refreshed
        if: always() && steps.validate.outcome == 'success'
        run: |
          if git diff --quiet -- $FILES; then
            echo "No change."
            exit 0
          fi
          git config user.name "github-actions[bot]"
          git config user.email "github-actions[bot]@users.noreply.github.com"
          git add -- $FILES
          git commit -m "Refresh live data (Last.fm, Goodreads)"
          git pull --rebase origin "$GITHUB_REF_NAME"
          git push origin "HEAD:$GITHUB_REF_NAME"

      - name: Fail if any source did not refresh
        if: always()
        env:
          NOW: ${{ steps.now.outcome }}
          BOOKS: ${{ steps.books.outcome }}
          VALID: ${{ steps.validate.outcome }}
        run: |
          status=0
          if [ "$NOW" != success ]; then
            echo "::error::Last.fm / currently-reading refresh: $NOW (see 'Fetch listening and currently-reading')"
            status=1
          fi
          if [ "$BOOKS" != success ]; then
            echo "::error::Read-shelf refresh: $BOOKS (see 'Refresh the read shelf')"
            status=1
          fi
          if [ "$VALID" != success ]; then
            echo "::error::Data validation: $VALID; nothing was committed"
            status=1
          fi
          exit $status

The workflow_dispatch line adds a manual "run now" button in GitHub's interface (which you will want the first dozen times), while the diff check means runs where nothing changed produce no commit at all.

Most of the rest is there for when something breaks. Both fetch steps are allowed to fail (continue-on-error), so a Goodreads outage doesn't stop the Last.fm numbers from landing, and the last step turns the run red afterwards so I still hear about it. Nothing is committed unless every file passes its schema in data/schema/, because the pages read these files with no fallback for a wrong shape. The robot and I both push to the same branch, hence the git pull --rebase before its push, and the concurrency group makes a manual run wait for a scheduled one instead of both committing the same files at once.

What the script gathers

The first script, refresh_now.py, makes a handful of HTTP requests. From the Last.fm API: my total play count, the most recent track, and my top albums of the past three months, which feed the homepage stat, the "now playing" strip, and the record crate on the music page respectively. From Goodreads, which shut its API years ago but still publishes RSS feeds: my currently-reading shelf, parsed straight out of the XML.

The second, refresh_books.py, reads the feed of my read shelf the same way and merges it into data/books.json, so a book I mark as read on Goodreads turns up on this site's reading pages by itself. The feed only reports the latest time I finished a book, so when I re-read something the script holds on to the earlier dates itself, because they exist nowhere else. It also writes data/reading.json, the books-read and reviews figures on the homepage. Everything goes into JSON files under data/, most of them a few kilobytes, apart from books.json, which holds every book I've marked read and weighs about 83 KB.

Each run appends that week's play count to a rolling log, capped at ninety entries, which at one measurement a week is close to two years of history.

A brindle Staffordshire terrier standing alert in woodland, looking at the camera
© Ken Reid. All rights reserved. A good fetcher.

What the homepage shows

The strip under the hero is a single row of up to four things, and only two of them come from the robot. The book and the track are read straight out of now.json. The latest short story comes from stories.json, which is an empty list until the first story goes up, so for now the strip leaves that card out. The latest blog post isn't in the strip at all, because the homepage has a whole section of its own for those.

Last.fm distinguishes a track that is playing from one that was played, so the label switches between "Now playing" and "Last played" depending on whether I happen to be listening when you open the page, and the first version gets a little animated equaliser beside it. It is not accurate, because it is updated so infrequently, but it gives some flavor to the site and over time it's likely representative of me. Maybe.

The third item, my most recent Bluesky post, is in no JSON file at all. Your browser fetches it from Bluesky's public API while the page loads. I've been tempted to remove it, as recently I've only posted new blogs to social media, but I hope I will return to posting random thoughts there again soon, making this a bit more useful than another place to see a new blog post (which is already on the index page!)

Design decisions

The homepage HTML contains a hardcoded scrobble count, and the JavaScript replaces it with the fresh value from the JSON after load. If the fetch fails, or JavaScript is off, the visitor sees a slightly stale number instead of a blank. The robot improves the site but it fails elegantly.

Goodreads' RSS goes down more often than Last.fm's API. When a source fails, the script keeps whatever the previous run wrote rather than blanking it: better yesterday's truth than today's error. A failure in one source never stops the others from updating, of course.

The Last.fm key lives in the repository's encrypted secrets and reaches the script as an environment variable. It appears in no file and no log for security reasons (not that I'd be massively put-out if someone broke into my goodreads or last.fm).

The bot's commits are labelled as bot commits, touch only the data files, and say what they did. My git history has a weekly heartbeat in it now: a tidy row of "Refresh live data" commits, one every Monday, each one the robot clocking in so I don't have to manually update stuff.

What else this pattern is good for

Anything that changes slowly and comes from somewhere with an API or a feed: your latest posts elsewhere, sports scores, weather, stars on a project, prices you're tracking, a "days since" counter for whatever you are currently ashamed of. The recipe is always the same three steps: a script that fetches and writes JSON, a workflow that runs it on a schedule and commits, and a page that reads the JSON with a fallback. If the data changes faster than a schedule can keep up with (live chat, comment counts), a static site is the wrong tool and no robot can easily fix that. Everything slower works just fine though.

Common questions

Doesn't committing data on a schedule bloat the repository?

Slowly, and acceptably. Each commit stores a few kilobytes of changed JSON, so a year of weekly commits comes to a couple of hundred kilobytes.

Why weekly instead of hourly, or on every visit?

I could but it's hardly vitally important and worth moving out of the free level of GitHub actions.

What happens when the robot itself breaks?

GitHub emails me when a scheduled workflow fails, and the last step makes sure a run where only one source failed still counts as a failure.

Could this update the page instantly when something changes?

Not this pattern. A schedule can only poll. Instant would need the source to push (webhooks) into something always listening, which drags a server back into the picture.