HTTrack and the Wayback Machine: Recovering an Old Website

Illustrated infographic summarizing: HTTRACK and wayback machine

By Greg Nowak. Last updated 2026-09-09.

When an old website disappears, the immediate problem is often missing business material: service descriptions, product documents, images and pages customers still visit. HTTrack and the Wayback Machine can help preserve or recover that material, but they solve different problems.

Use HTTrack to copy a website that is still accessible. Use the Wayback Machine to look for earlier versions. Neither gives you the original CMS, database or working checkout. For a business owner or agency inheriting a neglected site, the useful first step is deciding what needs saving and what will need rebuilding.

Choose the recovery route before downloading

First, ask the hosting provider and previous developer whether they have a usable backup. Check your own storage for database exports, uploaded files and source code. Restoring those may preserve functionality that a public website copy cannot.

Your situation Best starting point Expected outcome
The website is still online HTTrack A local copy of reachable pages and assets
The site is gone or important pages were overwritten Wayback Machine Previously archived content, where available
You need logins, orders or CMS editing restored Hosting and application backups A route to restoring the application and its data
You are replacing an old site A live copy plus selected historical captures Reference material for content, design and URL planning
Choose by what you need to recover: public content, historical material or the working application.

Copy a live website with HTTrack

HTTrack downloads linked files and rewrites links for local browsing. After installation, this command provides a limited first pass, based on the official command-line guide:

httrack https://example.com/ --path ./site-copy --depth=2 --sockets=2 --max-rate=100000

Replace https://example.com/ with the final address shown after any redirects. This saves into site-copy, follows one level of links beyond the starting page, uses two connections and limits transfer speed to 100,000 bytes per second. It is a sample crawl, not a complete site copy.

Open the saved index and inspect representative pages. Read hts-log.txt for failed requests or excluded URLs. Check whether images and stylesheets live on another hostname before expanding the crawl. Keep the project cache if you need to resume.

HTTrack does not execute JavaScript. Content that appears only after scripts run may be missing, even when the live page looks complete. For a planned migration, keep a dated, untouched copy before updating the mirror.

Recover missing material with the Wayback Machine

Enter the original domain or page URL into the Wayback Machine and inspect captures from before the loss or unwanted change. Check important pages individually rather than judging coverage from the homepage.

The Internet Archive’s guidance explains that captures can be incomplete, images may be missing, and navigation can move between capture dates. Record the timestamp for each recovered page. A site that looks consistent while browsing may contain material from several periods.

For missing images or PDFs, search their original URLs separately if you can identify them. Try relevant historical address variants, including HTTP versus HTTPS and domains with or without www.

Create a recovery list with the original URL, capture date, business priority and missing assets. Start with the pages customers need most. A usable service page and contact route usually deserve attention before an old news archive.

Can you download an archived website in bulk?

Bulk recovery is possible when suitable captures exist, but treat it as a technical retrieval task. My recommendation is to use an archive-aware downloader rather than starting a general crawl of Wayback’s replay pages.

The Ruby-based Wayback Machine Downloader documents installation and a listing mode:

gem install wayback_machine_downloader
wayback_machine_downloader https://example.com --list

Run this in a suitable Ruby environment. The second command lists archived URLs and timestamps without downloading their contents. Match the address to the historical site.

The project also documents --from and --to date filters. Without date restrictions, it selects the latest available version of each file, which can mix different designs and content revisions. A date range narrows the selection; it does not guarantee a complete snapshot.

There are open reports of retrieval errors, so these documented commands should be tested before you commit to a recovery schedule. If listing fails, investigate compatibility and archive access before attempting a larger job. An error does not prove the content is absent.

Turn recovered files into a usable website

The download is raw material. Before publishing anything, give the recovery a clear acceptance check:

  • Content: confirm that services, prices, staff details and contact information are still accurate.
  • Assets: inspect images, fonts and documents, and confirm permission to reuse them.
  • Functionality: rebuild and test forms, search, booking and checkout against the intended services.
  • URLs: retain useful original paths where practical and map changed pages to their replacements.
  • Handover: document missing material, assign ownership and establish backups with a tested restore process.

For an agency handover, keep the untouched recovery separate from the edited rebuild. Agree which pages must return first and which can wait. This makes the scope easier to estimate and gives the business a clear basis for approving the work.

If you have an old domain, partial files or a stalled migration, contact Greg with the URL and what you need back. I can help assess the recoverable material and plan the next steps.

Related on GrN.dk

Need help with this kind of work?

Discuss your website recovery with Greg Get in touch with Greg.

Sources

Latest articles

Google and Bing now offer first-party AI search visibility reports. Here’s how to build a useful baseline without inventing a misleading GEO score.

AI crawlers can copy a familiar name. Here’s how to verify signed agents at the edge while keeping legitimate automated traffic moving.

A critical Webform release is a reminder to audit every Drupal codebase, configuration and deployment—not just the main production website.

A secure AI workflow can turn Meet and Teams transcripts into approved decisions and tasks in Jira or Asana—without giving up control.

NGINX 1.31.5 can route on JSON body values. Here’s how to weigh the performance, security, and operational trade-offs before using it.

OpenAI can keep agent sessions running, but reliable workflows still depend on clear failure states, safe retries, validation, limits and human fallback.

AI can identify termination deadlines and price adjustments in supplier contracts, route uncertain findings for approval and create the right reminders.

Why a DNS record can exist in a dashboard yet fail publicly—and how to trace zone cuts, verify glue, and fix the right side of a live delegation.

An Apache version below 2.4.68 may still be patched. Package provenance, vendor advisories, module checks and runtime evidence reveal the real position.

PHP 8.2 security support ends on December 31, 2026. Here is how to audit, test, and migrate a mixed CMS estate without rushing production changes.