🗃 Open source self-hosted web archiving. Takes URLs/browser history/bookmarks/Pocket/Pinboard/etc., saves HTML, JS, PDFs, media, and more...
-
Updated
Sep 1, 2026 - Python
🗃 Open source self-hosted web archiving. Takes URLs/browser history/bookmarks/Pocket/Pinboard/etc., saves HTML, JS, PDFs, media, and more...
Core Python Web Archiving Toolkit for replay and recording of web archives
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)
A High-Fidelity Web Archiving Extension for Chrome and Chromium based browsers!
Collect and revisit web pages.
Run a high-fidelity browser-based web archiving crawler in a single Docker container
Automatically archive links to videos, images, and social media content from Google Sheets (and more).
Serverless replay of web archives directly in the browser
Selfhost web archiving and sharing service.
Local-first, open-source AI assistant for your data. Unify tasks, notes, docs, photos, and bookmarks. Private, self-hosted, and extensible via APIs.
InterPlanetary Wayback: A distributed and persistent archive replay system using IPFS
复刻网站的 Agent Skill:抓只读镜像、从压缩代码逐行还原、自动比对验收。An agent skill that mirrors a website, rebuilds it from the minified code, and verifies the result with automated diffs.
Wayback Machine API interface & a command-line tool
🖥️ Official ArchiveBox browser extension: automatically/manually preserve your browsing history using ArchiveBox.
Streaming WARC/ARC library for fast web archive IO
Browsertrix is the hosted, high-fidelity, browser-based crawling service from Webrecorder designed to make web archiving easier and more accessible for all!
Archiveror will help you preserve the webpages you love. 💾
Webrecorder Player for Desktop (OSX/Windows/Linux). (Built with Electron + Webrecorder)
A Tool To Push Web Resources Into Web Archives
To associate your repository with the web-archiving topic, visit your repo's landing page and select "manage topics."