Gzip

Paste gzip, zlib or raw deflate as Base64 to read what it holds — the bytes say which — or type a text to compress it into any of the three, in your browser.

Input
or drop one here — it is read locally and never uploaded
Output
[
  {
    "name": "Olivia",
    "city": "London"
  },
  {
    "name": "Noah",
    "city": "Manchester"
  },
  {
    "name": "Amelia",
    "city": "Dublin"
  },
  {
    "name": "Jack",
    "city": "Toronto"
  },
  {
    "name": "Isla",
    "city": "Sydney"
  },
  {
    "name": "Leo",
    "city": "Auckland"
  },
  {
    "name": "Mia",
    "city": "Cardiff"
  },
  {
    "name": "Oscar",
    "city": "Edinburgh"
  }
]

Read as gzip.

Uncompressed bytes
416
Compressed bytes
168
Change
-59.6%
Base64 characters
224

One compression in three Wrappers, and the word deflate

gzip, zlib and raw deflate are one compression, DEFLATE, wrapped three ways. RFC 1951 defines DEFLATE itself: the bytes cut into blocks, each written as codes that stand for a byte or for a run copied from earlier on. gzip’s Wrapper, RFC 1952, puts a header of at least ten bytes in front and eight bytes behind, a CRC-32 of the original bytes and their length; zlib’s, RFC 1950, puts two bytes in front and an Adler-32 behind; raw deflate is the compression with no Wrapper at all. The compressed data inside can be the same in all three, so the Wrapper is the whole of the difference.

The word deflate is where the names tangle. To HTTP, Content-Encoding: deflate means zlib’s Wrapper, and RFC 9110 notes that some servers send raw deflate under that name all the same. The browser’s own compressor follows HTTP and calls zlib’s Wrapper deflate, and raw deflate deflate-raw. .NET’s DeflateStream and PHP’s gzdeflate write raw deflate, while PHP’s gzcompress and Python’s zlib.compress write zlib’s Wrapper. So bytes labelled deflate can be either, and only the bytes can say which.

That is why the page reads the Wrapper from the bytes when it decompresses, and says which it read under what came out: Read as gzip, Read as zlib or Read as raw deflate. The bytes settle it: gzip always opens 1f 8b, zlib’s first two bytes read as one number are a multiple of 31, and no encoder opens raw deflate with either. A file named .gz that holds zlib is read as zlib, with a note that the name and the bytes disagree, and bytes in none of the three are refused, with nothing to decompress. Compressing, you choose the Wrapper, and it is gzip until you pick another.

Where payloads come from, and what H4sI and eJ are

Compressed bytes reach a developer in a few shapes: as the body of an HTTP response whose Content-Encoding names gzip or deflate, as a file, or as text, written as Base64 or as hex. Every gzip stream opens with the same three bytes, 1f 8b 08 — its two signature bytes and the number of its method, DEFLATE — and three bytes are exactly four characters of Base64, so gzip written as Base64 always begins H4sI. That is what a CloudWatch Logs subscription hands a Lambda function or a Kinesis stream in each record’s data, and how Helm stores a release: its JSON gzipped, then written as Base64. Where Helm keeps a release in a Kubernetes Secret, reading it straight from the Secret gives Base64 twice over, since a Secret holds its data as Base64 of its own; the Base64 tool decodes that outer layer into the H4sI… this page reads.

zlib’s two header bytes name its method, its window and the level it was written at, so with zlib’s default window its Base64 begins one of four ways: eJ at the default level, which is what Python’s zlib.compress writes unless asked otherwise, eA or eF at the levels below it, and eN at those above. Raw deflate has no signature at all, and its Base64 begins however its first block happens to.

Compressed as says how the page reads and writes those bytes as text: Base64, as the page opens, or Hex. It holds in both directions and the page never changes it for you, because every hex digit is a Base64 character too, so 1f8b0800 is text in both and different bytes in each. Base64 is read in the standard alphabet or the URL-safe one, with or without its padding, across spaces and line breaks, and written padded, or URL-safe with no padding when the URL-safe switch is on. Hex is written in spaced pairs and read spaced or run together, with 0x in front, or as a dump — and 0x is how T-SQL writes a binary value, such as the gzip SQL Server’s COMPRESS() returns. When hex read as Base64 fails, or Base64 read as hex is refused at a character hex lacks, and in either case the whole text reads the other way, a note says so beside a button that switches; nothing switches until you press it.

Reading a stream: Members, the bytes after it, an early end

RFC 1952 makes a gzip file a series of Members, complete gzip streams one after another, each with its own header and trailer. cat a.gz b.gz makes such a file, and so does adding to one with gzip -c file >> archive.gz, the GNU gzip manual’s own example; gunzip reads every Member and joins what they hold. This page reads them the same way: each Member is read and checked, what they hold is shown joined, and a note says how many were read. Not every reader does this. The standard the browser’s own decompressor follows allows a gzip stream one Member and calls a second an error, and Python’s zlib.decompress stops after the first one without a word.

Bytes after the end that open no Member are read past rather than refused, since everything before them is complete and checked, and a note says how many there are, the offset they begin at and whether every one is zero. Zeros there are padding — the GNU gzip manual meets them on tape, written up to the end of a block — and gunzip passes over them in silence; it passes over other bytes with a warning that trailing garbage was ignored, where Python’s gzip.decompress refuses the file. The same holds after a zlib stream’s Adler-32, and after a raw stream’s last block, though raw deflate carries no checksum to check.

A stream that ends before its end, such as a paste cut short or a download stopped part way, is shown as far as it goes, with a note that the checksum was not checked. A damaged stream is refused, and nothing of it is shown. That is stricter than gunzip, which writes out what it has decompressed before it reaches the checksum that finds it wrong; but a changed bit can decode in silence for some way before anything shows it, so what came out before the damage showed may already be wrong. No position is given either, since where a decoder notices damage is not where the damage is. And a zlib stream asking for a preset dictionary is refused in a sentence of its own: the stream names the bytes its compressor was primed with and does not carry them, so there is nothing to read it with.

What comes out is written as Bytes as says, as Text in UTF-8 or as Hex. Often it is JSON, which the JSON formatter lays out and checks, pointing at the line and column where it breaks. A gzipped picture or archive is not text, so with Text each sequence of its bytes that is not UTF-8 shows as U+FFFD, with a note counting those sequences beside a button that shows the bytes as hex; Download saves the bytes themselves either way. A line under the result shows what the first Member’s header stores, a name, a time in UTC and a comment, and the stored name never names the file Download saves. An output past the most the page decompresses at once stops there, with a note naming that size, and where a gzip’s trailer states a length past it, the page says so while the work runs.

Why a small input grows, and what Base64 adds

Compression wins by finding repetition, and a short text has little of it, while every Wrapper costs bytes of its own: gzip’s header and trailer come to 18 bytes when the header stores nothing more, zlib’s to 6, and raw deflate has none, though DEFLATE spends a few bits marking out its blocks. So Hello, world!, 13 bytes, becomes 33 bytes as gzip, 21 as zlib and 15 as raw deflate. The sizes line under a result counts the bytes on each side and the change between them, +153.8% for that gzip, and a note says why when a result grows. Text with plenty of repetition, such as JSON or a log, wins the cost back many times over: the example the page opens on, a small JSON, comes out at less than half its size.

Written as text, the bytes cost more again. Base64 takes four characters for every three bytes, a third more, as the Base64 tool’s guide explains, so those 33 bytes of gzip are 44 characters; hex takes two digits a byte, and the page writes them in spaced pairs, 98 characters for the same 33 bytes. The sizes line counts those characters too, and the Byte size converter turns any of its counts into KB or KiB.

Compressing: no level, not gzip -c’s bytes, and no Brotli

Compressing is the browser’s own work, through the CompressionStream the Compression standard defines, and that standard offers no compression level, so neither does this page: what you get is the browser’s default. gzip -9 may come out a little smaller, and seldom by much.

Nor are the bytes the ones gzip -c writes, though both decompress to the same text. The headers differ first. GNU gzip stores a file’s name and time unless told -n, and from a pipe a time of zero; it records -9 or -1 in a byte of the header; and in the tenth byte, which RFC 1952 gives to the operating system, it writes 03 on Linux, Unix’s number. The browser is handed the bytes alone, so neither your file’s name nor its time can leave with the result: Chromium stores no name and a time of zero, which is what gzip -n chooses, and in the tenth byte it writes 03 on Linux, where Chrome on Windows writes 0a. Then the compressed data differs, since GNU gzip’s compressor is not the browser’s: on a long text the two pick different bytes at the same level. A mismatch with gzip -c therefore means nothing by itself; what counts is that both decompress to the same bytes.

Brotli is not here, for now. The Compression standard names it, but not every browser writes it yet, and a choice the page could offer only where the browser makes it would be a different page in different browsers. zstd is not in that standard at all. What this page reads and writes is DEFLATE in its three Wrappers, and nothing else.

A file in, a file out, and nothing uploaded

A file can take the text box’s place in either direction: choose one, or drop it on the box. Its bytes are read as they are, by no Compressed as or Bytes as, and a file has no ceiling here where the text box has one, though a very large file draws a warning that the browser may not hold it. Decompressing, Download names the result as gunzip does, the file’s name less .gz, with .tgz becoming .tar, and less .zz or .deflate as well, the endings this page writes for the other two Wrappers; any other name, and anything pasted, comes out as bytes.bin. Compressing, it adds .gz, .zz or .deflate to the file’s name, .zz being how pigz names zlib, or to the word text for what you typed.

Use as input moves a result into the box and turns the direction, so a round trip takes one click, and a result that came from a file, or is too large for the box, moves as a file. Whichever way it goes, the file is read in this tab and the work runs in a worker the page starts on your own device: a payload or a file is not uploaded to do it.

The same with gunzip and Python

At a command line, gunzip and zcat read gzip as this page does, every Member included, once base64 -d has turned a payload back into bytes, though neither reads zlib or raw deflate, which both reject as not gzip; gzip -k compresses a file and keeps the original beside it. In Python, gzip.decompress reads gzip’s Wrapper and every Member in it, and zlib.decompress reads any of the three Wrappers, told which by its wbits:

# A payload: decode it, then decompress it
printf 'H4sIAAAAAAAAA/NIzcnJ11Eozy/KSVEEAObG5usNAAAA' | base64 -d | gunzip   # Hello, world!

# A file: compress it, keeping the original, then print it back
printf 'Hello, world!' > hello.txt
gzip -k hello.txt
zcat hello.txt.gz                                  # Hello, world!

# Two members, the second appended to the first, read back as one
printf 'Hello, ' | gzip > two.gz
printf 'world!' | gzip >> two.gz
zcat two.gz                                        # Hello, world!

# The same in a script: every member, then one wrapper at a time
import base64, gzip, zlib
two = open('two.gz', 'rb').read()
zl = base64.b64decode('eJzzSM3JyddRKM8vyklRBAAgXgSK')
raw = base64.b64decode('80jNycnXUSjPL8pJUQQA')
gzip.decompress(two)                               # b'Hello, world!'
zlib.decompress(two, wbits=31)                     # b'Hello, '
zlib.decompress(zl, wbits=15)                      # b'Hello, world!'
zlib.decompress(raw, wbits=-15)                    # b'Hello, world!'
zlib.decompress(zl, wbits=47)                      # b'Hello, world!'

wbits 31 asks for gzip’s Wrapper, 15, the default, for zlib’s, and -15 for raw deflate, the 15 in each being the largest window, while 47 takes gzip’s or zlib’s, whichever the bytes open with. zlib.decompress reads a single Member, though, and drops the rest without a word, which is why it gives back b'Hello, ' alone from the two Members above; gzip.decompress reads them all. Python is also stricter than this page twice over: gzip.decompress refuses bytes after the end that are not zeros, and both functions refuse a stream that ends early, where the page shows what came out.

Frequently asked questions

Why is my compressed result bigger than what I typed?
Because each Wrapper costs bytes of its own and a short text has little for the compression to remove: gzip adds 18 bytes, zlib 6, and DEFLATE marks out its blocks besides, so a few words grow where a long JSON shrinks, and the page says so in a note. Raw deflate is the smallest of the three. Written as Base64, any result is a third longer again.
Why does the output not match gzip -c?
Nothing promises that it will. A gzip header can hold a name, a time, a mark of the level and a byte for the operating system, which GNU gzip and the browser fill in differently, and the two are different compressors, so on a longer text even the data between header and trailer can differ. Two gzip streams of one text are both right when each decompresses to it, so compare what they decompress to: Use as input turns this page’s result straight back.
What does H4sI at the start of a payload mean?
That the payload is gzip written as Base64: every gzip stream opens with the bytes 1f 8b 08, and those three bytes are H4sI in Base64. Paste it here as it is. A payload opening eJ is most likely zlib at its default level, and one opening with neither may be raw deflate, or not compressed at all, which the page tells you.
Is gzip encryption?
No. Compression takes no key, so anyone holding the bytes can decompress them, here or with gunzip, and every byte comes back. A password inside a gzipped payload is as exposed as one written out in full.
Is my payload or my file uploaded anywhere?
No. The payload you paste and the file you choose or drop are both read inside this tab, where a worker the page starts does the compressing and the decompressing, and Download builds its file from the result right there. Neither is sent to this site or to anyone else.

Related tools

  • JSON formatter

    Validate and beautify JSON, with clear error locations.

  • Byte size converter

    Convert between KB, MB, GB and KiB, MiB, GiB.

  • Base64

    Encode and decode Base64 — full UTF-8 support.

  • Base32

    Encode and decode Base32 and base32hex — text in UTF-8 or bytes in hex.