Base32
Encode text or hex to Base32 or base32hex, padded or not, upper or lower case, and decode it back in your browser, with a stray character’s line and column.
JBSWY3DP
Bytes in thirty-two characters, and the digits left out
Base32 writes bytes as text, in thirty-two characters that stand for five bits each, so every five bytes — forty bits — become eight characters. Hello is five bytes and is written JBSWY3DP. Choose Decode and paste Base32, and the page gives back the bytes it stands for: as text, or as hex if you choose Bytes as Hex.
RFC 4648, which defines it, uses the twenty-six capital letters and the digits 2 to 7. The RFC gives the reason 0 and 1 are missing: anyone handling the text can easily mistake 0 for O, and 1 for l or I, so the alphabet keeps the letters and leaves those two digits out. Beside the letters it needs six digits, and the six it takes are 2 to 7, so 8 and 9 are not in it either. Decoding in this alphabet, the page refuses a 0, 1, 8 or 9 where it stands, rather than read a 0 as O or a 1 as I: the RFC allows a decoder to read them so, and says that by default it should not.
Case is not part of the value. RFC 4648 designed Base32 for text that has to be read without regard to case, so jbswy3dp is Hello just as JBSWY3DP is. The RFC writes its alphabet in capitals, and so does this page unless you switch lowercase on. Reading also passes over white space anywhere, line breaks included, so Base32 split into groups or across lines reads as if it were written in one piece.
Padding, and the lengths Base32 comes in
Bytes seldom come in whole fives. One to four left at the end take two, four, five or seven characters, and RFC 4648 pads that last group to eight with =: six of them, four, three or one. So f is MY======, fo is MZXQ====, foo is MZXW6=== and foob is MZXW6YQ=, while fooba, five bytes, fills its eight exactly as MZXW6YTB.
Counted in eights, then, the characters before any = always leave 0, 2, 4, 5 or 7 over and never 1, 3 or 6, since no number of bytes leaves those; the page refuses such a text at its last character rather than read it as something shorter. The padding says nothing the length does not, so the page reads Base32 with its padding left off — JBUQ is Hi just as JBUQ==== is — and refuses padding only where it is there and wrong: an = in the middle, with more of the text after it, or the wrong number of = at the end, where the refusal says how many belong.
When you encode, the Padding switch decides whether the = are written. It starts on, because RFC 4648 asks for padding unless the specification using it says otherwise, and several do: the Key Uri Format, which describes the otpauth:// URI in a 2FA QR code, says a secret’s padding should be omitted, and NSEC3 records and IPFS write none. The lowercase switch beside it writes the letters small; the digits and = have no case to change. Both switches are there only while encoding, since reading takes every combination of them.
Other tools read less than this page does. GNU coreutils writes Base32 with base32, and base32hex, the second Alphabet, with basenc --base32hex; Python’s base64 module has a function for each. All of them write what this page writes as it opens, padded and in capitals. But base32 -d refuses Base32 whose padding is missing or whose letters are lowercase, and Python’s b32decode refuses both too, reading lowercase only when it is given casefold=True:
# Writing: the default Alphabet, then the other one
printf 'Hi' | base32 # JBUQ====
printf 'Hi' | basenc --base32hex # 91KG====
# Reading back: padded, then without padding, then in lowercase
printf 'JBUQ====' | base32 -d # Hi
printf 'JBUQ' | base32 -d # base32: invalid input
printf 'jbuq====' | base32 -d # base32: invalid input
# The same in a script
import base64
base64.b32encode(b'Hi') # b'JBUQ===='
base64.b32hexencode(b'Hi') # b'91KG===='
base64.b32decode('JBUQ') # binascii.Error: Incorrect padding
base64.b32decode('jbuq====', casefold=True) # b'Hi'Two Alphabets: base32 and base32hex
RFC 4648 defines a second Alphabet, base32hex: the digits 0 to 9, then the letters A to V, so its characters run in the order of the values they stand for, from 0 for zero to V for thirty-one. That gives it one property the first lacks, which the RFC points out: compared character by character, its text sorts in the order of the bytes it stands for. The byte 00 is AA====== in base32 and 00====== in base32hex, and the byte FF is 74====== and VS======; sorted as text, base32 puts FF first, since digits sort before letters, while base32hex keeps the bytes in order.
DNSSEC’s NSEC3 records use it. A record’s name begins with the hash of a domain name, written in base32hex without padding, and RFC 5155 notes that, written that way, the hashed names sort in the same order as the hashes do by value. Its own example zone hashes example to 0p9mhaveqvm6t7vbl5lop2u3t2rp3tom. With base32hex chosen, the page reads those 32 characters as the hash’s 20 bytes; with base32 chosen, it refuses the 0 — and a note says the text reads whole in base32hex, beside a button that switches.
Since the same bytes are different text in each — Hi is JBUQ==== in base32 and 91KG==== in base32hex — the Alphabet you choose holds in both directions, and the page never changes it for you. When a text you decode holds a character the chosen Alphabet does not have, or ends in spare bits that are not zero, which the questions below explain, while the other Alphabet reads all of it as an encoder would write it, a note says so beside a button that switches; the Alphabet changes only when you press that button or choose it yourself. A text both read cleanly, such as ABCDEFGH, is read in the one chosen, since nothing in it says which was meant.
Where Base32 turns up: 2FA secrets, onion addresses and IPFS
Base32 is how the secret behind a 2FA app’s codes is written in the Key Uri Format, which describes the otpauth:// URI a setup QR code can hold: its secret parameter is Base32, and the padding should be left off. Paste such a secret in capitals or small letters, in groups split by spaces or across lines, with or without = — a hyphen between groups is refused where it stands — or paste the whole URI, and the page reads the secret out of it and says so. What the URI’s other parameters mean, and how a secret becomes the codes an app shows, belong to the TOTP / 2FA code debugger: its guide explains the URI and its parameters, and its page computes the codes. The label and the issuer in the URI are percent-encoded, as the Key Uri Format asks, and the URL encoder / decoder reads them.
An onion address of Tor’s version 3 is Base32 as well. Its 56 characters before .onion are the Base32 of 35 bytes: the service’s 32-byte public key, a two-byte checksum and a version byte, 03. Tor’s specification writes its example addresses in lowercase, and 35 bytes fill whole groups, so there is no padding to leave off. Paste the 56 characters without .onion, choose Bytes as Hex, and the last byte reads 03.
IPFS writes its content identifiers, CIDs, in Base32 by default from version 1 on: lowercase and unpadded, after one letter, b, that is not part of the Base32 but names it — multibase’s prefix, which is taken off before the rest is decoded. So bafybeigdyrzt5sfp7udm7hu76uh7y26nf3efuylqabf3oclgtqy55fbzdi, the example in IPFS’s own documentation, is refused here whole, as a length no bytes produce; without its b it reads as 36 bytes, the binary CID, opening with 01 for its version.
A 2FA secret is bytes, not text
The Key Uri Format’s example secret is JBSWY3DPEHPK3PXP, and that document spells out its value as ten bytes: the characters of Hello!, then DE AD BE EF. Its first eight characters, JBSWY3DP, are Hello on their own. Decoded here with Bytes as Text, the page’s default, it comes out as Hello!, then U+07AD — DE AD happen to make one character, a Thaana mark — and then two U+FFFD, each standing for a sequence that is not text. A note under the result counts the sequences replaced, says the first begins at byte 9, whose bits begin in the character at line 1, column 13, the 3, and says the bytes may not be text at all.
A real secret is random bytes, which seldom make text, so decoded as text it is a scatter of stray characters and U+FFFD that tells you nothing about whether the secret is right. That is what Bytes as Hex is for. Choose it, or press the button beside the note, and the same secret is shown whole as 48 65 6C 6C 6F 21 DE AD BE EF: the bytes themselves, which are what a server holds and what every code is computed from. The page never switches to hex on its own, so what is on the screen is always what you chose.
The other way round is the same choice, made while encoding. A key a server keeps as hex is bytes: choose Encode and Bytes as Hex, paste it — spaced or run together, with 0x before the bytes, as an array of numbers or as a dump in hexdump -C’s or xxd’s layout — and the page writes the Base32 an authenticator takes. 48656C6C6F21DEADBEEF comes out as JBSWY3DPEHPK3PXP, the secret above; ten bytes fill whole groups, and a key that does not, such as one of sixteen bytes, comes out with = after it unless you turn padding off, as the Key Uri Format asks. Hex pasted while decoding by mistake — an even number of hex digits and nothing else but white space, in a form that encoding can read as bytes — draws a note that it looks like hex, beside a button that turns to encoding it. And while decoding, Download saves the bytes themselves as a file named bytes.bin.
Base32 or Base64, and why not Crockford’s alphabet
Base64 does the same job with sixty-four characters, four for every three bytes, which makes its text a third longer than the bytes where Base32’s is three fifths longer. What Base64 gives up for that is case: its alphabet holds capitals and small letters as different values, with + and / besides, so SGk= is Hi and sgk= is two other bytes. Base32 suits text that passes through something that ignores case, or that a person reads off one screen and types into another. When Base64 pasted here holds a + or a /, which no Base32 writes, and is shaped like Base64 throughout, it is refused with a note that it looks like Base64 and a link to the Base64 tool, which reads Base64 as text.
Other alphabets of thirty-two characters exist, and this page offers RFC 4648’s two and no other. One is Douglas Crockford’s, which keeps all ten digits and leaves out I, L, O and U, and which the ULID format uses. Crockford defined it for writing numbers, and written over bytes it has two answers: the sixteen bytes 00 to 0F, packed five bits at a time as RFC 4648 packs bytes, are 000G40R40M30E209185GR38E1W, and read as one 128-bit number, as a ULID’s twenty-six characters are read, they are 00041061050R3GG28A1C60T3GF. A page writing bytes in it would be choosing one of the two on your behalf, unseen.
Frequently asked questions
- Why does my decoded secret show U+FFFD?
- Because a 2FA secret is random bytes, and random bytes are seldom UTF-8 text. With Bytes as Text the page writes each sequence that is not text as U+FFFD, the replacement character, and a note says how many there were and where the first begins, as a byte and as the line and column of the character its bits begin in. The bytes themselves are untouched: the button beside the note shows them as hex, and Download saves them. A U+FFFD that really is in a text, the bytes EF BF BD, is read as text and not counted.
- What does “a length no bytes produce” mean?
- Counted in eights, the characters of Base32, any = set aside, leave 0, 2, 4, 5 or 7 over, whatever the bytes, so a text that leaves 1, 3 or 6 cannot stand for any bytes, and the page refuses it at its last character. Such a text was not written that way by an encoder: something was lost or added on its way to you. JBSWY3DPEHPK3P, the Key Uri Format’s secret two characters short, is refused there. A cut can also land on a length bytes do produce, and then it reads as fewer bytes — JBSWY3DPEHPK3PX as nine — which the page can only notice where the last character’s spare bits are not zero, as that X’s are. A cut at a whole group of eight leaves nothing to notice.
- What are spare bits, and why does the page say mine are not zero?
- Five bits to a character and eight to a byte seldom come out even, so unless the bytes come in fives, the last character carries one to four bits that no byte holds, and an encoder writes them as zeros. MZXW6YQ= and MZXW6YR= are both foob: the last three bits of Q are 000 and those of R are 001, and only the first is what an encoder writes. The page reads both, and for the second its note names the R, where it stands and the Q an encoder writes there. Spare bits that are not zero mean the text is not what an encoder wrote: it may have been cut short, edited by hand or written in the other Alphabet, the last of which the page checks for you.
- Is Base32 encryption?
- No. Base32 changes how bytes are written and nothing else: anyone can decode it, and every decoder that follows RFC 4648 gets the same bytes back in the same Alphabet. A 2FA secret written in Base32 is the secret itself, so anyone who sees the setup key holds your second factor.
- Is anything I paste sent to a server?
- No. Both directions run in this page on your own device: nothing you paste — a secret, a text or hex — is uploaded, and the file Download saves is made in your browser from the bytes already there.
Related tools
- Base64
Base64 writes four characters for every three bytes, where Base32 writes eight for every five, and in Base64 a letter’s case is part of its value: abc is YWJj on that page, and YWJJ reads as abI.
- TOTP / 2FA code debugger
Generate and debug 2FA codes, with the full derivation.
- URL encoder / decoder
Percent-encode a value or a whole URL, and back.
- JSON string escape / unescape
Escape text for a JSON string, or read an escaped one back.