Skip to main content
← Back to Archives More from 2026

Keeping a 23-Year-Old Website Usable and Accessible

Blue-line technical illustration of an octagonal wooden dice tray with a black dice cup and several ten-sided dice.

The dice setup for Dice Bins, one of the old tools I wanted to get working again.

This website is 23 years old, and some of it has been rotting for a while. When I decided to revamp it this year, I wanted to bring back old material and tools that still had some utility. Some documents I host are hard to find elsewhere; one little program is still used by election officials in California. Getting those things back online was important, but I also wanted to make them more accessible to people with disabilities.

My interest in accessibility goes back to my PhD work on voting systems, when I worked with blind people and other people with disabilities. That included my dear friend Noel Runyan, a blind voting-systems security and accessibility expert.

I used Codex, OpenAI's coding assistant, to help with that work. And, oof, it has not been easy, even with that help. As I repaired things, I also added tests so later changes wouldn't quietly undo the repairs. Eventually, getting those tests under control became a substantial project in itself.

Getting Dice Bins working again

Dice Bins helps a few election officials in California use ten-sided dice to randomly select precincts for audits comparing voting-machine totals with paper ballots. It's a specialized use, but I wanted to keep the software working for them.

One problem was that the replacement I had cooked up in JavaScript worked at its new address, but people following the old dicebins.php link couldn't get there. I'd set up a redirect, but a Cloudflare security rule blocking .php scripts blocked the request before the site could redirect it. So, I made a narrow exception for these redirects while keeping the broader protection in place.

Once people reached the tool, they also needed to be able to use the form. I improved the labels and keyboard access, and added messages that screen readers could announce after a calculation or an input error. I checked both valid and invalid submissions rather than stopping at whether the page loaded.

I also tried navigating selected pages with screen readers—VoiceOver and, later, ChromeVox—as well as checking keyboard navigation, enlarged text, and narrow windows. I wanted to try those screen-reader workflows myself rather than rely entirely on automated reports. This was still my own testing, not an independent specialist audit or a formal evaluation with disabled users.

Keeping the repairs working

During the website ressurection process, I was constantly changing layouts and bringing old content back, and I wanted to ensure I had accessibility checks beyond whichever page I happened to be editing... I want the whole site to be continuously usable and accessible. I used Codex to add a browser crawl: an automated browser opened each of ~1,300 pages generated by the site and ran checks for accessibility problems it could detect. Those full crawls were really useful early on, when I was still finding problems I hadn't anticipated a change in one place could cascade to others.

The crawl ran as part of the site's continuous integration, or CI—the automatic build and tests that run when I make changes to the website. But after all that hard work, it was running on nearly every change. In one August run, scanning the pages for accessibility testing took almost 13 minutes of a build-and-test job that lasted a little over 14 minutes. Not good.

With incremental development using a tool like Codex, it was easy—too easy!—to keep adding checks, reports, and workflow rules, each with a plausible reason. These tend to pile up. And increasingly the scans wouldn't find much. I realized that I needed to spend less time on strict usability testing on each website change without losing the checks that were helping me.

We first tried running parts of the crawl in parallel, break it up. An experiment I ran in August with two parallel acessibility crawls finished about 34% sooner, using about 20% more total computing time across the jobs. That cut the wait, but it still meant crawling the whole set of pages on nearly every change.

Next, in order to "predict" where changes might be relevant, we built a program to work out which pages a change might affect and choose tests smartly. That sounds straightforward for an edited post, but changing the layout or styling shared across the site could affect hundreds of pages. I coded this up and it worked ok, but I decided not to ship it once we had created a simpler alternative, which I'll describe. I didn't want another complicated system just to decide which tests to run! Big websites have big problems.

Checking a smaller set first

The site already used shared templates, which meant a layout change could update many pages at once. That also turns out to be a bit of an opportunity; it gave us a way to reduce the browser testing: choose representative examples of the different page types and archive layouts, and test those on ordinary changes. The selection was fixed rather than recalculated for each edit. It is not a full crawl, it's checking a sample.

We kept checks on the HTML of every generated page too—for things like headings and form labels—without opening each one in a browser (that is cheap). The complete browser crawl ran separately, initially every day, and I could also start one manually. In the August rollout of this new "smart" sample test, the smaller set of browser tests checked 20 pages in about half a minute.

Before-and-after diagram: a full browser crawl on nearly every change is replaced by immediate HTML checks and browser tests on representative pages. A separate scheduled crawl checks all generated pages and can find a defect outside the representative set.

Ordinary changes get HTML checks and a smaller set of browser tests. The complete crawl runs separately, so a problem on another page may be found later.

Pages using the same template can still have different problems, though. We tested that by injecting errors: making a paragraph's text too pale against its background on two pages, for example. The routine checks caught it on a page in the chosen representative set but passed when it was on a page outside that set (as it couldn't see the flaw). Manually running the full crawl against the second test version found it.

This is a delay I would need to be comfortable accepting: a problem on a page outside the smaller set could get published before a later full crawl found it. I was comfortable with that for this site, with the HTML checks still running broadly and the complete crawl checking beyond the selected pages. The browser tests used Chromium, and even the full crawl covered only the generated pages—not every historical file I host or every browser someone might use. Nothing is perfect!

Adjusting as the site settled down

As I got through more of the restoration work, the full crawls started finding less new. I was changing less of the shared structure, and rerunning the same checks every day was becoming tedious. I still wanted the complete crawl, but it didn't need to run as often as it had during the website ressurection.

As of late September, I was trying four scheduled crawls a week instead of seven, using the same pages and checks. Running them less often means a problem they can detect may remain unnoticed for longer. The planned gaps are at most 48 hours, but a delayed or failed run can leave a longer gap. This trial runs through October 7; I suspect I'll eventually settle on weekly, but I haven't made that decision yet.

There are still documents and media I want to make more accessible. For some old PDFs, I've added web-page reading copies alongside the originals; other material still needs work. My accessibility statement records the scope of the evaluation and the remaining limitations. I'm continuing those repairs while working out how much checking the site needs now that the initial overhaul is behind me.