Technology · · 8 min read
Reading This Site Offline
This site now works offline: one hand-written service worker, no framework, no build step. Here's how.
Photograph © Ken Reid
Reading This Site Offline
This site learned a small trick recently: it works without the internet. I implemented it out of curiosity on how it would be done, but I'm glad I did. If someone tries to read one of my blogs while on a plane, now they can, so long as it was preloaded. Read a few posts over breakfast, open the site again, and the pages you visited are still there, styled, functional. If the connection dies mid-browse, it will still work. Nice, right?
This is done via a service worker, and this post walks through mine: no framework, no build step, in keeping with this site's general philosophy of doing things with the smallest possible amount of infrastructure. If you run a static site anywhere (GitHub Pages included), this is a good philosophy, in my opinion.
Quick jargon guide
- Service worker: a small script the browser installs alongside your site and runs in the background. It sits between your pages and the network, and can answer requests itself, including when there is no network.
- Cache (the browser kind): a local store of responses the service worker can save and replay. Lives on the visitor's device, controlled by your code.
- Network-first: a strategy: try the internet, fall back to the cache. Fresh when online, functional when not.
- Stale-while-revalidate: the inverted strategy: answer instantly from cache, then fetch a new copy in the background for next time. Fast always, fresh eventually, annoying when you are refreshing a page for an update, though.
- Precache: the short list of files saved at install time, before you've visited anything, the skeleton the site can't render without.
- Cache invalidation: famously one of the two hard problems in computer science. A service worker is a machine for having this problem on purpose.
- giscus: the comment system under my posts. It stores every comment as a GitHub Discussion and loads in an embedded frame from giscus.app, so the conversation lives on GitHub rather than on this site.
Two strategies
Everything the worker does comes down to one decision, made per request: who do you trust more, the network or the cache?
Pages get network-first. HTML is where mistakes live: a typo fixed, a broken link repaired, a post updated. When you're online I always want you reading the newest deploy, so the worker tries the network and saves a copy of whatever comes back. It reaches for that copy when the fetch fails, and also when the network has had three and a half seconds and still said nothing, because a connection that's barely there would otherwise leave you staring at a blank tab. The fresh page still goes into the cache when it finally turns up:
// HTML: network-first, with a timeout
function navigation(event) {
var req = event.request;
var key = pageKey(req.url);
var network = Promise.resolve(event.preloadResponse).then(function (pre) {
return pre || fetch(req);
});
event.waitUntil(network.then(function (res) {
return isHtml(res) ? store(PAGES, key, res.clone(), PAGES_LIMIT) : null;
}).catch(noop));
return new Promise(function (resolve) {
var done = false;
function answer(res, fromCache) {
if (done) return;
done = true;
clearTimeout(timer);
if (fromCache) rememberCachedClient(event.resultingClientId);
resolve(res);
}
var timer = setTimeout(function () {
savedPage(key).then(function (hit) { if (hit) answer(hit, true); });
}, NAV_TIMEOUT_MS);
network.then(function (res) { answer(res, false); }, function () {
savedPage(key).then(function (hit) {
if (hit) answer(hit, true);
else offlinePage().then(function (res) { answer(res, true); });
});
});
});
}
The chain reads as a 'politeness ranking': fresh page if possible, your cached copy if not, and a dedicated offline page as the final backup. Every page you visit while online becomes a page you own while offline, the cache is your personal reading history (the last 60 pages of it, anyway). Navigation preload starts that request while the worker is still waking up, instead of after. Pages are saved by their path, so blog.html?tag=books and blog.html share one copy rather than one per filter. And if the offline page itself has been cleared out, offlinePage() builds a tiny one on the spot, so the browser's own error page never gets to blame the site for your connection.
Code and data get network-first too. Scripts, stylesheets and the data/*.json files go to the network first and only fall back to the cache when that fails. A page fetched fresh must never run last week's JavaScript against its new markup, and the data files change on a schedule (Last.fm, Goodreads), so served stale they showed last week's shelf. The one exception is a page that was itself served from the cache: its scripts come from the cache too, so an old page runs the code it was written against and doesn't sit blank on a bad connection waiting for a stylesheet.
Images and fonts get stale-while-revalidate. Thumbnails and fonts change rarely, so the priorities invert. The worker answers from cache immediately and refreshes in the background:
// Images and fonts: stale-while-revalidate
function staleWhileRevalidate(event, cacheName, limit) {
var req = event.request;
var network = fetch(req);
event.waitUntil(network.then(function (res) {
return store(cacheName, req, res.clone(), limit);
}).catch(noop));
return lookup(req).then(function (hit) {
return hit || network.catch(function () { return Response.error(); });
});
}
function store(cacheName, req, res, limit) {
if (!cacheable(res)) return Promise.resolve();
return caches.open(cacheName).then(function (cache) {
return cache.put(req, res).then(function () {
if (cacheName === CORE) return dropOtherVersions(cache, req);
if (limit) return trim(cacheName, limit);
});
});
}
Worst case, you see an old thumbnail for one visit. In exchange, repeat visits render instantly, offline or not. That limit is the housekeeping, meaning that the image cache is capped at 200 entries, oldest evicted first, so browsing my whole gallery doesn't fill up your phone! trim() runs one pass per cache at a time, because a gallery scroll stores hundreds of thumbnails in a burst and overlapping trims would all read the same list and delete the same entries.
The worker keeps three caches: core for the precache plus every script, stylesheet, font and data file, img for images, and pages for HTML, capped at 60. Core is never trimmed, because the repo itself puts a ceiling on how big it can get.
What I deliberately don't cache
The worker ignores anything cross-origin; the full-size photographs stay on GitHub's release servers, caching those would mean warehousing megabytes per photo on your device for a click-through you probably won't repeat. Comments are giscus, which lives in an iframe and belongs to GitHub; offline, it simply doesn't appear, which is the correct behaviour. Analytics likewise gets no offline resurrection, if the network can't see you, neither should it.
Offline mode should preserve the reading, not fake the internet. The precache reflects the same idea: it's just the offline page, the stylesheet, the core scripts, and posts.json (so the blog index still works), so twelve files, everything else earns its place in your cache by you actually opening it.
The footguns
I'm not really a website developer or designer, beyond hobbyist endeavors like this website and a couple of landing pages for research labs I worked in, so much of this section is just what I learned while, well, learning this stuff. Service workers have a deserved reputation for one specific misery: the cache that would not die. Ship a worker with a careless cache-first strategy and your visitors can be pinned to an old version of your site for days: including, delightfully, an old version of the service worker itself. The classic developer experience is editing a file, refreshing, seeing no change, and spending an hour debugging to find it's actually working as intended and you just turn off your monitor for a break and see your sad reflection staring back at yourself. Ahem, anyway.
I've hit two of my own. An earlier version kept everything in one cache capped at 200 entries, and the oldest entries in it were always the precache, so a single scroll through the gallery pushed out the offline page and the stylesheet, and the offline page then had nothing to show. It also served scripts stale-while-revalidate, which for one view after a deploy ran the new markup against the old shared-components.js: an uncaught ReferenceError, no footer, no photo strip.
I'll also admit what this isn't: it isn't a full progressive web app. There's no install banner, no background sync. Those are all possible and all, for a site whose job is being read, beside the point.
Try it
Open a few posts, then turn on airplane mode and keep clicking. The pages you visited load; the search on the blog index still filters; the pages you didn't visit hand you a courteous offline page instead of a browser error, with a list of the posts you do have saved. Turn the network back on and the whole arrangement dissolves back into an ordinary website. Much like an IT worker, if it's doing the job correctly, you won't even notice it's there.
Common questions
Why not just cache the entire site up front?
Precaching everything downloads megabytes a first-time visitor never asked for, most of which they'll never read. As much as I'd like all my visitors to read all my website, most just swing by for something that interests them, then off they go.
Does this let the site track me offline?
No, rather the opposite. The caches live in your browser, managed by your browser, and I have no visibility into them whatsoever. Offline visits send me nothing: no analytics, no logs, no signal you exist. It's the most private way to read the site.
Why not use Workbox or a PWA framework?
Workbox is excellent and I'd reach for it on a complex app. For one hand-written worker on a static site, where I can read every rule top to bottom, the abstraction would outweigh the logic: another layer to learn, and one more dependency to keep updated forever.
How do updates reach me if I'm serving from cache?
Pages are network-first, so any online visit gets the latest deploy automatically, the cache only speaks when the network can't, or when it has taken more than three and a half seconds to answer. Even then, the fresh page still lands in the cache when it arrives.