WARC to TBZ2 Converter
FileFlip converts WARC to TBZ2 entirely inside your browser, using libarchive compiled to WebAssembly. The file is never uploaded to a server: it is read off your disk, converted on your own machine, and saved back by the browser. That means no file-size limit, no queue and no account.
How do I convert a WARC file to TBZ2?#
- Drop your WARC file onto the converter at the top of this page, or click it to pick one from your computer. You can add as many as you like.
- Leave the target set to TBZ2 and the conversion starts on its own. FileFlip downloads 1.0 MB of WebAssembly the first time and nothing after that, then runs libarchive on your own machine.
- Save the TBZ2 file. Your browser writes it straight to your downloads folder, because there was never a server holding it.
Does converting WARC to TBZ2 lose quality?#
No. Converting WARC to TBZ2 does not touch the files inside the archive. Each member is read out and written back byte for byte, so only the wrapper around them changes and nothing is recompressed.
Is it safe to convert WARC files online?#
Yes, because nothing is uploaded. FileFlip converts WARC files inside the browser tab you already have open, using WebAssembly builds of the same engines a desktop converter would install. Your file is read from your own disk into the page's memory, and no request carrying it is ever made.
You do not have to take that on trust. Open your browser's network panel, run a conversion, and you will see the engine come down and your file go nowhere. There is no upload, no server-side copy, no retention window, and no account tied to what you converted.
How FileFlip works lists every engine, its version and its download size, and shows how to check for yourself that nothing leaves your machine.
What WARC to TBZ2 actually does to the file#
Converting WARC to TBZ2 runs on libarchive, downloads 1.0 MB of WebAssembly the first time and nothing after that, copies the archived files through unchanged, keeps file contents.
| Engine | libarchive |
|---|---|
| Conversion path | 2 steps: WARC → the unpacked files → TBZ2 |
| Downloaded to your browser | 1.0 MB, once, then cached by your browser |
| Quality | Lossless, contents copied unchanged |
| You get back | One file |
| Where it runs | In a background worker, so the page stays responsive |
What survives the conversion
| What is in the file | This conversion | Why |
|---|---|---|
| File contents | Kept | Carried through into the converted file. |
| File timestamps | Sometimes | Survives for some files and not others. |
| File permissions | Sometimes | Survives for some files and not others. |
Worth knowing before you start:
- the output is a PAX tar archive inside the compressor, not a single compressed stream
- every entry is extracted into memory before anything is written, so a very large archive can exhaust it
What is a WARC file?#
A WARC file is a sequence of records that each capture one HTTP request and response, or another web-harvesting transaction, exactly as it crossed the wire, headers included, so a crawl can be replayed rather than merely read.
Use WARC when a crawl or website capture needs to preserve exactly what was sent and received, not just the resulting files.
Convert a WARC's contents to plain files when only the captured pages or media are needed and the HTTP-level record of the crawl itself does not matter.
What is a TBZ2 file?#
A TBZ2 file is a TAR archive compressed as one continuous bzip2 stream under a single extension, so unlike ZIP, extracting one file means decompressing everything that comes before it in the archive.
Use TBZ2 when a tarball needs to be smaller than TGZ makes it and the slower compression time is acceptable.
Convert TBZ2 to TXZ for a smaller result still, or to TGZ when speed matters more than size.
WARC to TBZ2 questions people ask#
What does a WARC file actually store?
A WARC file stores a sequence of records, each one a full HTTP request or response exactly as it crossed the wire, headers and all, along with metadata like the crawl timestamp. That is different from saving a webpage's files directly, because a WARC also preserves how the server responded, including redirects, status codes and content negotiation.
How do I open a WARC file?
A WARC file needs archive-replay software rather than a browser, since it is a log of HTTP traffic rather than a folder of pages. Webrecorder's ReplayWeb.page opens one in a browser tab and replays the captured site as if it were live, and the warcio command-line tool can inspect or extract individual records.
Why does the Internet Archive use WARC instead of just saving web pages?
WARC preserves the complete HTTP exchange the Internet Archive's crawlers captured, not just the final rendered page, which is what lets the Wayback Machine show accurate status codes, redirects and content-type headers years later. Saving files alone would lose that protocol-level context.
Is WARC an official standard?
Yes. WARC was standardized by the International Organization for Standardization as ISO 28500, developed through the International Internet Preservation Consortium, and it is the format essentially every serious web crawler and archive now writes.