Permavault › Web Archiving for Evidence and Citation

Web Archiving for Evidence and Citation

Courts have accepted Wayback Machine captures, generally with an affidavit from the Internet Archive or by judicial notice. The limit is not the law, it is coverage: the archive only holds what its crawler happened to reach, and for a growing share of the web it now holds nothing.

Published July 11, 2026 · Updated August 29, 2026

Web archiving covers two jobs that look similar and are not. One is preserving the past: a public archive crawls the web continuously so that history exists at all. The other is preserving a specific page, right now, because your brief, story, or paper depends on what it says today.

The Wayback Machine is unmatched at the first job and structurally unsuited to the second. That distinction is the whole of this page, and getting it wrong is how a citation quietly becomes unsupported three years later.

Pages do not stay put

In 2014, Harvard researchers Lawrence Lessig, Jonathan Zittrain, and Kendra Albert checked every web link cited in United States Supreme Court opinions. Roughly half, 49.9 percent, no longer led to the material the Court had cited. Not obscure blog posts in obscure cases: sources relied on in the reasoning of the highest court in the country, gone or changed.

A decade later the problem has compounded. Pew Research Center’s 2024 study found that 38 percent of web pages that existed in 2013 were no longer accessible ten years on, and a quarter of all pages that existed at any point between 2013 and 2023 had already vanished by the study date. On Wikipedia, 54 percent of articles contain at least one dead link in their references. About one in five government and news pages contains a broken link right now.

Two failure modes hide inside those numbers, and the second is worse.

Link rot is the obvious one: the URL returns an error, the page is gone. Annoying, visible, at least honest.

Reference rot is quieter. The URL still works, but the content is no longer what you cited. Among Supreme Court links that still technically resolved, the Harvard team found only 76 percent still contained the originally cited material. A page that changed under your citation is worse than a dead one, because it can now say something different while appearing to support you. A reader checking your source finds a current page, not the page you read, and has no way to know the difference.

For a brief, that is an authority that no longer says what you quoted. For a journalist, a source that now contradicts the story. For a researcher, a methods page or dataset that silently changed after publication.

Is a Wayback Machine capture admissible as evidence?

Wayback captures are used in litigation regularly and courts have accepted them. In Telewizja Polska USA, Inc. v. Echostar Satellite Corp. (N.D. Ill. 2004), archived pages came in with an affidavit from an Internet Archive representative explaining how the archive works and confirming the printouts matched its records. Many courts since have taken judicial notice of Wayback captures or admitted them on a similar foundation.

Notice what carries the weight there. It is not the archive’s reputation; it is a witness declaration about a specific process, obtained for your case, plus a snapshot that happens to exist for the date in dispute. Take either away and there is nothing to certify.

So the honest answer is conditional, and the conditions are where it usually breaks:

A Wayback capture is genuinely useful evidence when it exists and covers the point. It is not a preservation strategy, because you cannot make it exist on demand for the page in front of you.

Where the public archive falls short

The blocking problem got serious in 2026. More than 340 local news outlets now limit the Internet Archive’s access to their journalism, according to Nieman Lab’s reporting in May 2026, up from 241 in January. The publishers include some of the largest chains in the country: USA Today Co., McClatchy, Advance Local, Tribune Publishing, and Alden Global Capital subsidiaries. The New York Times has also limited the Archive’s access.

The motive is not hostility to preservation. Publishers fear AI companies scraping archived content for training data, and archived copies are a side door. But the effect is the same regardless of motive: for a growing share of the news, there is no Wayback snapshot to point to.

You do not control the crawl. The Wayback Machine archives what its crawlers reach, when they reach it. A page can be created, changed, and deleted between snapshots. If the version that matters existed for six hours on a Tuesday, the odds the crawler saw that exact version are not in your favor. On-demand saving helps when you think of it, but it is subject to the same blocks and capture limits.

Captures are often incomplete. Dynamic pages, JavaScript-heavy layouts, images served from third-party hosts, and interactive elements frequently do not archive faithfully. What replays may be a partial rendering of what a visitor actually saw, and for evidence, the missing element is often the one that matters.

Availability can change retroactively. The Archive has historically honored exclusion requests, and pages can become unavailable after you cited them. A snapshot you relied on is not guaranteed to stay reachable at that URL.

It is one institution. The Internet Archive is a single nonprofit that has spent recent years fighting major lawsuits and absorbing a serious 2024 security breach. It has weathered all of it so far, and the web is lucky it did. But a preservation strategy that depends entirely on one organization’s continued existence and policies is a strategy with a single point of failure.

Why saving a PDF is not an archive either

Saving a local copy is better than nothing, and for private working files it is fine. As a citation or evidence practice it has two gaps. Your reader cannot see your local copy, so the citation still points at a URL that will rot. And when the content is contested, a copy on your own machine proves little: there is no independent record of when it was made, or that it matches what was actually online.

The scholarly world’s institutional answer is Perma.cc, built by the Harvard Library Innovation Lab: it archives the cited page and gives you a stable link, held in library custody. For citation permanence in scholarship it is a genuinely good system, and if your institution participates, use it. Its model is trust in the custodian. A reader relies on the library consortium holding the copy faithfully, rather than on a proof they can check themselves.

What a capture you control has to prove

For work that may be challenged rather than merely followed, permanence is only half the answer. The other half is proof. A capture worth citing should be:

Permanent by design. Stored somewhere engineered to outlive companies and subscription payments, not on one vendor’s cloud or one person’s hard drive.

Fixed at a moment in time. The record shows exactly what the page said on the date you cited it, with the date written by the capture process rather than typed in by hand.

Independently verifiable. Cryptographic fingerprints, computed at capture, let anyone confirm the copy is unaltered without trusting you, your firm, or the archive that holds it. This is the same property that anchors a chain of custody for digital evidence, and you can check it yourself.

Complete. A full-fidelity capture of the page as it rendered, not a text scrape that drops the chart your argument depends on.

The right division of labor

Use the Wayback Machine for what it is unmatched at: history you did not know you would need. What did this company’s homepage claim in 2019? What did this page say before you ever got involved? No tool you run today can capture yesterday, and the Internet Archive is a public good worth supporting for exactly that reason.

Make your own capture for what you know matters now. The page in front of you that your story, case, or paper depends on deserves a capture you control: made at the exact moment you need it, complete as rendered, fingerprinted at capture, and stored somewhere that does not depend on any single organization’s crawler, policies, or survival.

The two are complements. Neither substitutes for the other.

Where Permavault fits

Permavault is built for the second job. Paste the URL and a neutral automated system captures the full page as it rendered, fingerprints every file with cryptographic hashes at capture, and stores the result on a permanent decentralized network of roughly 300 independent nodes, funded by a long-term storage endowment. You choose the moment. The capture is complete as rendered, timestamped by the process, and comes with a verifiable record anyone can check, without trusting us or any archive.

That gives you citations designed not to rot: cite the original URL for convention, and add the capture as the record of what it said. When someone checks your source in five years, what you cited is what they see, verifiable byte for byte.

$4.99 per capture. The price is on the page. No subscriptions, no sales calls, no annual contracts.

Blocked outlets are the reminder that archives you do not control can close to you, and half of what the Supreme Court cited is already gone or changed. Whatever you are writing, the pages you rely on are on the same clock. Capture them while they still say what you read.

This article is general information, not legal advice for any specific matter. Admissibility depends on the facts, the jurisdiction, and the judge.

Go deeper

The guides in this cluster, each covering one part of the problem in detail.

Need a web page preserved exactly as it exists right now?

Capture it with Permavault