Skip to content

WARC to TXZ Converter

FileFlip converts WARC to TXZ entirely inside your browser, using libarchive compiled to WebAssembly. The file is never uploaded to a server: it is read off your disk, converted on your own machine, and saved back by the browser. That means no file-size limit, no queue and no account.

Convert WARC to

How do I convert a WARC file to TXZ?#

  1. Drop your WARC file onto the converter at the top of this page, or click it to pick one from your computer. You can add as many as you like.
  2. Leave the target set to TXZ and the conversion starts on its own. FileFlip downloads 1.0 MB of WebAssembly the first time and nothing after that, then runs libarchive on your own machine.
  3. Save the TXZ file. Your browser writes it straight to your downloads folder, because there was never a server holding it.

Does converting WARC to TXZ lose quality?#

No. Converting WARC to TXZ does not touch the files inside the archive. Each member is read out and written back byte for byte, so only the wrapper around them changes and nothing is recompressed.

Is it safe to convert WARC files online?#

Yes, because nothing is uploaded. FileFlip converts WARC files inside the browser tab you already have open, using WebAssembly builds of the same engines a desktop converter would install. Your file is read from your own disk into the page's memory, and no request carrying it is ever made.

You do not have to take that on trust. Open your browser's network panel, run a conversion, and you will see the engine come down and your file go nowhere. There is no upload, no server-side copy, no retention window, and no account tied to what you converted.

How FileFlip works lists every engine, its version and its download size, and shows how to check for yourself that nothing leaves your machine.

What WARC to TXZ actually does to the file#

Converting WARC to TXZ runs on libarchive, downloads 1.0 MB of WebAssembly the first time and nothing after that, copies the archived files through unchanged, keeps file contents.

Enginelibarchive
Conversion path2 steps: WARC → the unpacked files → TXZ
Downloaded to your browser1.0 MB, once, then cached by your browser
QualityLossless, contents copied unchanged
You get backOne file
Where it runsIn a background worker, so the page stays responsive

What survives the conversion

What is in the fileThis conversionWhy
File contentsKeptCarried through into the converted file.
File timestampsSometimesSurvives for some files and not others.
File permissionsSometimesSurvives for some files and not others.

Worth knowing before you start:

  • the output is a PAX tar archive inside the compressor, not a single compressed stream
  • every entry is extracted into memory before anything is written, so a very large archive can exhaust it

What is a WARC file?#

A WARC file is a sequence of records that each capture one HTTP request and response, or another web-harvesting transaction, exactly as it crossed the wire, headers included, so a crawl can be replayed rather than merely read.

Use WARC when a crawl or website capture needs to preserve exactly what was sent and received, not just the resulting files.

Convert a WARC's contents to plain files when only the captured pages or media are needed and the HTTP-level record of the crawl itself does not matter.

What is a TXZ file?#

A TXZ file is a TAR archive compressed as one xz stream under a single extension, combining tar's metadata-preserving bundling with the strongest general-purpose compression ratio of the tar pairings offered here.

Use TXZ for a source release or package where the smallest possible download size is worth slower compression.

Convert TXZ to TGZ when a tool in the chain only understands gzip, still the most universally supported tar compressor.

WARC to TXZ questions people ask#

What does a WARC file actually store?

A WARC file stores a sequence of records, each one a full HTTP request or response exactly as it crossed the wire, headers and all, along with metadata like the crawl timestamp. That is different from saving a webpage's files directly, because a WARC also preserves how the server responded, including redirects, status codes and content negotiation.

How do I open a WARC file?

A WARC file needs archive-replay software rather than a browser, since it is a log of HTTP traffic rather than a folder of pages. Webrecorder's ReplayWeb.page opens one in a browser tab and replays the captured site as if it were live, and the warcio command-line tool can inspect or extract individual records.

Why does the Internet Archive use WARC instead of just saving web pages?

WARC preserves the complete HTTP exchange the Internet Archive's crawlers captured, not just the final rendered page, which is what lets the Wayback Machine show accurate status codes, redirects and content-type headers years later. Saving files alone would lose that protocol-level context.

Is WARC an official standard?

Yes. WARC was standardized by the International Organization for Standardization as ISO 28500, developed through the International Internet Preservation Consortium, and it is the format essentially every serious web crawler and archive now writes.