HTTrack and the Wayback Machine: Recovering an Old Website

Illustrated infographic summarizing: HTTRACK and wayback machine

By Greg Nowak. Last updated 2026-09-09.

When an old website disappears, the immediate problem is often missing business material: service descriptions, product documents, images and pages customers still visit. HTTrack and the Wayback Machine can help preserve or recover that material, but they solve different problems.

Use HTTrack to copy a website that is still accessible. Use the Wayback Machine to look for earlier versions. Neither gives you the original CMS, database or working checkout. For a business owner or agency inheriting a neglected site, the useful first step is deciding what needs saving and what will need rebuilding.

Choose the recovery route before downloading

First, ask the hosting provider and previous developer whether they have a usable backup. Check your own storage for database exports, uploaded files and source code. Restoring those may preserve functionality that a public website copy cannot.

Your situation Best starting point Expected outcome
The website is still online HTTrack A local copy of reachable pages and assets
The site is gone or important pages were overwritten Wayback Machine Previously archived content, where available
You need logins, orders or CMS editing restored Hosting and application backups A route to restoring the application and its data
You are replacing an old site A live copy plus selected historical captures Reference material for content, design and URL planning
Choose by what you need to recover: public content, historical material or the working application.

Copy a live website with HTTrack

HTTrack downloads linked files and rewrites links for local browsing. After installation, this command provides a limited first pass, based on the official command-line guide:

httrack https://example.com/ --path ./site-copy --depth=2 --sockets=2 --max-rate=100000

Replace https://example.com/ with the final address shown after any redirects. This saves into site-copy, follows one level of links beyond the starting page, uses two connections and limits transfer speed to 100,000 bytes per second. It is a sample crawl, not a complete site copy.

Open the saved index and inspect representative pages. Read hts-log.txt for failed requests or excluded URLs. Check whether images and stylesheets live on another hostname before expanding the crawl. Keep the project cache if you need to resume.

HTTrack does not execute JavaScript. Content that appears only after scripts run may be missing, even when the live page looks complete. For a planned migration, keep a dated, untouched copy before updating the mirror.

Recover missing material with the Wayback Machine

Enter the original domain or page URL into the Wayback Machine and inspect captures from before the loss or unwanted change. Check important pages individually rather than judging coverage from the homepage.

The Internet Archive’s guidance explains that captures can be incomplete, images may be missing, and navigation can move between capture dates. Record the timestamp for each recovered page. A site that looks consistent while browsing may contain material from several periods.

For missing images or PDFs, search their original URLs separately if you can identify them. Try relevant historical address variants, including HTTP versus HTTPS and domains with or without www.

Create a recovery list with the original URL, capture date, business priority and missing assets. Start with the pages customers need most. A usable service page and contact route usually deserve attention before an old news archive.

Can you download an archived website in bulk?

Bulk recovery is possible when suitable captures exist, but treat it as a technical retrieval task. My recommendation is to use an archive-aware downloader rather than starting a general crawl of Wayback’s replay pages.

The Ruby-based Wayback Machine Downloader documents installation and a listing mode:

gem install wayback_machine_downloader
wayback_machine_downloader https://example.com --list

Run this in a suitable Ruby environment. The second command lists archived URLs and timestamps without downloading their contents. Match the address to the historical site.

The project also documents --from and --to date filters. Without date restrictions, it selects the latest available version of each file, which can mix different designs and content revisions. A date range narrows the selection; it does not guarantee a complete snapshot.

There are open reports of retrieval errors, so these documented commands should be tested before you commit to a recovery schedule. If listing fails, investigate compatibility and archive access before attempting a larger job. An error does not prove the content is absent.

Turn recovered files into a usable website

The download is raw material. Before publishing anything, give the recovery a clear acceptance check:

  • Content: confirm that services, prices, staff details and contact information are still accurate.
  • Assets: inspect images, fonts and documents, and confirm permission to reuse them.
  • Functionality: rebuild and test forms, search, booking and checkout against the intended services.
  • URLs: retain useful original paths where practical and map changed pages to their replacements.
  • Handover: document missing material, assign ownership and establish backups with a tested restore process.

For an agency handover, keep the untouched recovery separate from the edited rebuild. Agree which pages must return first and which can wait. This makes the scope easier to estimate and gives the business a clear basis for approving the work.

If you have an old domain, partial files or a stalled migration, contact Greg with the URL and what you need back. I can help assess the recoverable material and plan the next steps.

Related on GrN.dk

Need help with this kind of work?

Discuss your website recovery with Greg Get in touch with Greg.

Sources

Seneste artikler

Et sikkert AI-workflow kan omsætte Meet- og Teams-transskripter til godkendte beslutninger og opgaver i Jira eller Asana – uden at slippe kontrollen.

AI kan finde opsigelsesfrister og prisreguleringer i leverandørkontrakter, sende usikre fund til godkendelse og oprette de rette påmindelser.

Sådan automatiserer danske virksomheder Gmail og Microsoft 365 med hurtig sortering, begrænsede rettigheder og menneskelig godkendelse.

Samme kunde på flere kort i HubSpot? Se, hvordan CVR-match, AI-forslag og menneskelig godkendelse kan bruges til at rydde op med styr på felter, relationer og kundehistorik.

Få en ugentlig marketingrapport fra GA4 og Google Ads med kontrollerede beregninger, tydelige dataforbehold og et kort AI-udkast, der hjælper jer på mandagsmødet.

Brug AI til webshoppens alt-tekster med en overskuelig pilot: kortlæg billederne, få danske forslag, og kontrollér resultatet i WordPress og WooCommerce.

AI-baseret ticketanalyse kan afsløre gentagne klager, produktfejl og huller i dokumentationen – uden at virksomheden behøver endnu en chatbot.

OpenSSH 10 fjerner DSA og advarer om nøgleudveksling, der ikke er post-kvantesikker. Her får du en metode til at afgrænse SFTP-oprydningen uden at svække alle SSH-forbindelser.

Botforespørgsler overstiger nu menneskelig webtrafik. Lær at auditere AI-crawlere, fastsætte regler på stiniveau, håndhæve robots.txt og måle det forretningsmæssige afkast.

Cloudflares Tunnel-opdateringer fra 2026 forbedrer kortlægning, overvågning af replikaer, logstreaming og overdragelse – men synliggør samtidig svagt ejerskab og mangelfuld praksis for failover og logging.