ArchiveBox Hosting

Preserve webpages, PDFs, media, and bookmarks in a searchable long-term archive.

Order now No setup fees
  • One click deploy
  • 2 GB RAM Memory needed
  • 25 GB Disk Space Needed
  • From 5 € Price

Tech

Docker image
archivebox/archivebox:latest
Default port
8000

How ArchiveBox works

ArchiveBox imports URLs from lists, browser exports, feeds, or supported integrations and runs several capture methods against each target. Depending on the available tools and source response, it can retain page HTML, screenshots, PDFs, text, media metadata, WARC data, favicons, and other representations. The web interface then provides an index for searching, tagging, viewing, and revisiting saved snapshots.

Multiple formats improve the chance that useful evidence remains available when one representation fails, but no capture method reproduces every modern site. Authentication walls, captchas, streaming media, client-side applications, expiring tokens, anti-automation controls, and robots policies can produce partial or failed snapshots. Each URL needs verification when the preserved result matters.

Key ArchiveBox features

The catalogue creates generated administrator credentials and persists /data, which contains the archive index, SQLite data, configuration, logs, and captured files. Archive visibility and user access need deliberate configuration because captured URLs can include private paths, query parameters, personal information, or session-derived material. A public snapshot can disclose more than the original top-level URL suggests.

ArchiveBox is preservation software, not a legal conclusion about authenticity, admissibility, copyright, or retention compliance. Operators remain responsible for permission to capture and retain content, chain-of-custody procedures, access controls, and independent integrity checks. Large captures also need monitoring because video, PDFs, images, and repeated formats can consume storage quickly.

Who uses ArchiveBox

Researchers preserve source pages cited in a project, journalists retain public material before it changes, and organisations keep authorised web references alongside case or policy work. Individuals can import a bookmark export and build a searchable personal archive. Scheduled or repeated capture workflows require extra care to avoid excessive requests and duplicate storage.

ArchiveBox is not a substitute for the Internet Archive, a credentialed backup of every private account, or a managed compliance archive. Sensitive captures should be restricted, and users should retain original files or official records where available. A visible screenshot proves only what the capture process rendered, not the complete server-side state of the source.

Self-hosting ArchiveBox: requirements and cost

ArchiveBox resource use is dominated by URL count, capture methods, browser processes, screenshots, PDFs, media extraction, WARC files, indexing, imports, concurrent jobs, and recapture frequency. PostgreSQL and MariaDB are Not required in the catalogue template; the package persists its own SQLite data and archive files. Text-only pages are modest, while media-rich captures can expand rapidly. ArchiveBox is marked storage-driven in the plan mapping, so 50 GB on Plan 3 is a realistic starting point and 100 GB on Plan 4 provides more room for a large working set.

On AvaHost, ArchiveBox uses Plan 2 at €5. The hosted ArchiveBox package includes one-click deployment, a custom domain with automated HTTPS, automatic application updates, and scheduled backups. The package supplies one generated administrator account but no source-site credentials, browser extension, legal-preservation process, content rights, proxy network, or capture-quality assurance. Verify representative snapshots, restrict sensitive archives, and keep independent copies of material that must survive an application or account failure.

F.A.Q

  • ArchiveBox starts at €5 on Plan 2. Capture processing and stored snapshot formats make storage the practical limit, so 50 GB on Plan 3 is a realistic starting point and 100 GB on Plan 4 suits a larger archive. URL count, screenshots, PDFs, media, WARC files, indexing, and simultaneous captures affect capacity.

  • No capture method can reproduce every website. Authentication, captchas, client-side applications, streaming media, expiring sessions, robots policies, rate limits, and source changes can produce partial or failed snapshots. Review each important capture, compare several formats, retain authorised originals, and do not treat a screenshot or saved HTML file as automatic legal proof.

  • A custom domain can point to ArchiveBox, with automated HTTPS protecting login and archive traffic. Privacy still depends on application users, snapshot visibility, URL parameters, captured personal data, and how links are shared. Test access from a signed-out browser and review sensitive snapshots before assuming the archive is restricted to administrators.

  • AvaHost applies ArchiveBox application updates automatically and includes scheduled backups for persistent `/data`. After a significant release, verify administrator access, imports, tags, search, index pages, screenshots, HTML, PDFs, WARC output, media extraction, existing snapshots, and new capture jobs. Browser or extractor changes can alter results, so compare representative pages before scaling recaptures.