// HACKER NEWS — CYBERSECURITY
Where did the old web go? We followed 657,607 links to find out
An old 0.mk database backup held 657,958 links created between 2009 and 2014, along with their click counts. We restored 657,607 of those records as pre-2015 links and followed every destination in August 2026. Of 655,178 safe, crawlable link records, 76.7% no longer returned a loading page.
Most 0.mk users were in Macedonia, so this is not a census of the entire web. It is a large surviving record of what one online community shared during that period, including local news, personal blogs, photo hosts, forums, and the major platforms of the time.
When 0.mk started in 2009, it was a passion project built by a team of three. We worked on it when we could, usually for a few hours a week around our regular jobs. Seventeen years later, one of us found an old database backup on a disk and decided to bring it back.
Here is what those six years of link creation look like, with the long silence after them:
The crawl covers all 657,607 restored link records dated through December 2014. We excluded 2,429 records whose targets were malformed, internal, credentialed, or policy-blocked, leaving 655,178 crawlable historical links:
Even that 23.3% overstates how much survived. A login wall, a parked domain full of ads, or a "this content is no longer available" notice all count as loading. A working page does not mean the original content is still there.
Why 657,607 links but 494,781 URLs? Multiple short links sometimes point to the exact same destination. There are 162,826 such repeat records. Counting each destination once leaves 494,781 distinct URLs, of which 492,620 were crawlable. Only 21.3% of those loaded. The percentage barely moves when repeated destinations are removed: 78.7% still did not load.
At the unique-URL level, 55.0% failed at the network layer after retrying uncertain results from a second network, and 23.7% returned an HTTP error. The most common HTTP result was 404, across 76,403 distinct URLs. Another 29,663 returned 403 or 429; those pages did not load for the crawler, but may be blocking automated requests rather than missing. A 403 or 429 can mean the site blocked our crawler, so "did not load" is more honest than saying every one of those pages is gone.
The same pattern appears at the domain level. Of 133,605 crawlable hostnames, only 34,827 had even one URL load. The other 98,778 had none.
The 2011 split explains the strange annual totals. One account created 83,398 distinct links to pelaphptutorials.com. At URL level, 92.5% of 2011 destinations did not load. Count that host once and the figure is 61.7%, almost identical to 2010 and 2012.