The seven-bit world
For thirty years, "how long is this string" had one answer, and it was right. ASCII gave 128 characters 128 numbers, each fitted in one byte, and every layer this page is about collapsed into a single list.
ASCII was standardised in 1963, revised in 1967, and it is seven bits wide — 128 values, 0 through 127. That leaves the eighth bit of a byte spare, which mattered enormously later and we will come back to it. Of the 128, most are letters, digits and punctuation, and a surprising number are not characters at all.
$python3.14 ascii.py 0 0x00 Cc (unicodedata.name raises ValueError) 7 0x07 Cc (unicodedata.name raises ValueError) 9 0x09 Cc (unicodedata.name raises ValueError) 10 0x0A Cc (unicodedata.name raises ValueError) 13 0x0D Cc (unicodedata.name raises ValueError) 27 0x1B Cc (unicodedata.name raises ValueError) 32 0x20 Zs SPACE 48 0x30 Nd 0 DIGIT ZERO 65 0x41 Lu A LATIN CAPITAL LETTER A 97 0x61 Ll a LATIN SMALL LETTER A 127 0x7F Cc (unicodedata.name raises ValueError) controls in 0..127: 33 everything else : 95
DEL at the top are control codes — instructions to a machine rather than marks on a page. That leaves 95, which is the printable set including the space. Look also at the last column: seven of these eleven rows have no Unicode name at all and only four do, which is a fact we will use in §09.The control codes are the oldest thing in modern computing and they are still in your terminal. 0x07 physically rang a bell on a teletype. 0x0D returned the print carriage to the left margin and 0x0A advanced the paper by one line — two separate mechanical motions, which is exactly why Windows still ends a line with both bytes and Unix with only the second. 0x1B, escape, announced that what followed was a command rather than text, and every colour your terminal prints still begins with it.
The one idea to hold from this section. In ASCII, the byte, the code unit, the code point and the grapheme cluster are the same number, and there are exactly as many of each. A generation of programmers learned "a character is a byte" and was not wrong — it was true of the only text they had. Everything difficult about the rest of this page comes from those four layers coming apart.
LATIN SMALL LETTER H and so on.)Before anything else on this page: do not trust the glyphs. This page loads three remote fonts and none of them carries emoji, nor is any guaranteed to carry every combining mark shown below. On some machines the pictures in the panel above are empty boxes, and on a machine with no emoji font they are nothing at all. So no claim here rests on a glyph appearing. Every character shown travels with its code point and its Unicode name, and those are the claim — a page about text that assumed its own text renders would have missed its own point.
One byte, eight letters
ASCII used seven bits of eight, leaving 128 values spare in every byte. Everyone in the world filled those 128 slots with their own alphabet, and nobody wrote down which one they had used.
The result was the codepage: a table mapping the upper 128 byte values to characters. IBM shipped one per market, Microsoft shipped a different one per market, ISO standardised a third family, Apple had its own. They agree completely below 128 and almost nowhere above it. Here is a single byte, 0xE9, decoded through nine of them:
$python3.14 codepages.py latin_1 0xE9 -> U+00E9 LATIN SMALL LETTER E WITH ACUTE cp1252 0xE9 -> U+00E9 LATIN SMALL LETTER E WITH ACUTE cp437 0xE9 -> U+0398 GREEK CAPITAL LETTER THETA cp850 0xE9 -> U+00DA LATIN CAPITAL LETTER U WITH ACUTE cp1251 0xE9 -> U+0439 CYRILLIC SMALL LETTER SHORT I iso8859_7 0xE9 -> U+03B9 GREEK SMALL LETTER IOTA koi8_r 0xE9 -> U+0418 CYRILLIC CAPITAL LETTER I mac_roman 0xE9 -> U+00C8 LATIN CAPITAL LETTER E WITH GRAVE cp866 0xE9 -> U+0449 CYRILLIC SMALL LETTER SHCHA
| Codec | Byte | What it decodes to |
|---|---|---|
| latin_1 / cp1252 | 0xE9 | éU+00E9Latin small letter e with acute |
| cp437 | 0xE9 | ΘU+0398Greek capital letter theta |
| cp850 | 0xE9 | ÚU+00DALatin capital letter u with acute |
| cp1251 | 0xE9 | йU+0439Cyrillic small letter short i |
| iso8859_7 | 0xE9 | ιU+03B9Greek small letter iota |
| koi8_r | 0xE9 | ИU+0418Cyrillic capital letter i |
| mac_roman | 0xE9 | ÈU+00C8Latin capital letter e with grave |
| cp866 | 0xE9 | щU+0449Cyrillic small letter shcha |
Mojibake is not corruption
You have seen café where a word should be, or ’ in place of an apostrophe. Neither is damage. The second one is the three UTF-8 bytes of ’U+2019Right single quotation mark read one at a time through cp1252, which turns them into âU+00E2Latin small letter a with circumflex €U+20ACEuro sign ™U+2122Trade mark sign. Every byte arrived intact; they were simply read through the wrong table. Watch it happen:
$python3.14 mojibake.py text ['U+0063', 'U+0061', 'U+0066', 'U+00E9'] utf-8 bytes b'caf\xc3\xa9' read as cp1252 -> ['U+0063', 'U+0061', 'U+0066', 'U+00C3', 'U+00A9'] read as latin_1-> ['U+0063', 'U+0061', 'U+0066', 'U+00C3', 'U+00A9'] latin-1 bytes b'caf\xe9' read as utf-8 -> UnicodeDecodeError 'utf-8' codec can't decode byte 0xe9 in position 3: unexpected end of data
0xE9 announces a three-byte sequence and no bytes at all followed — it is the last byte in the payload, which is exactly what unexpected end of data means. UTF-8 can tell you it has been handed the wrong thing. An eight-bit codepage never can.The honest caveat about all of this. The codepages did not merely disagree — each one could represent a few hundred characters at most, so most of the world's text simply could not be written down. '€'.encode('latin_1') raises UnicodeEncodeError: ordinal not in range(256). There is no Latin-1 byte for the euro sign, for Japanese, for Devanagari, or for a document that needs Greek and Russian in the same paragraph. Mojibake was the visible symptom; the invisible one was everything you could not type at all.
The abstract character
Unicode's founding move is to stop talking about bytes altogether. Give every character in every script a number, once, globally — and say nothing whatsoever about how that number is stored.
That number is a code point, written U+ followed by at least four hex digits. U+0041 is the capital A. U+00E9 is the accented e. U+20AC is the euro sign. The notation is hex for the reason the previous page gave — hex makes structure visible, and Unicode's space is organised in blocks of hex-round sizes.
The separation is the point and it is easy to slide past. A code point is not a byte, is not a number of bytes, and does not become bytes until you choose an encoding. U+20AC is three bytes in UTF-8, two in UTF-16, four in UTF-32, and nothing at all until you pick one.
$python3.14 space.py unicodedata.unidata_version = 16.0.0 code point space 1114112 (17 planes of 65,536) unassigned (Cn) 819533 assigned 294579 private use (Co) 137468 surrogates (Cs) 2048 controls (Cc) 65 letters (L*) 4551 + Lo 136477 marks (M*) 2501 symbols (S*) 8514
Why 1,114,112 and not a round number
Because the ceiling is U+10FFFF, and that ceiling was chosen to be exactly what UTF-16 can address — 17 planes of 65,536. Unicode originally promised 16 bits would be enough for everyone, ran out, and the extension it bolted on could reach sixteen more planes and no further. The size of the abstract space was fixed by the limits of one encoding of it. The layers are supposed to be independent; this is the one place where the seam shows, and every later section inherits it.
A code point is not a character either, quite. Unicode's own term is abstract character, and the mismatch runs both ways: some things a reader calls one character are several code points (§08), and some code points are not characters anyone would point at — the 2,048 surrogates, the 65 controls, and U+200DZero width joiner, which is invisible and load-bearing.
UTF-8, byte by byte
A code point is a number up to 1,114,111. A byte holds 0 to 255. UTF-8 is the rule for writing the first as a sequence of the second, and it is small enough to fit on one card.
There are four templates. Which one you use depends only on how many bits the code point needs. The fixed prefix bits are structure; the xs are where the code point's own bits go.
10, and no lead byte ever does. That single property is what §05 is about, and it is why UTF-8 beat every other variable-length encoding proposed alongside it.$python3.14 utf8.py U+00041 1 byte(s) 41 01000001 U+000E9 2 byte(s) C3 A9 11000011 10101001 U+020AC 3 byte(s) E2 82 AC 11100010 10000010 10101100 U+1F468 4 byte(s) F0 9F 91 A8 11110000 10011111 10010001 10101000
01000001 starts with a zero — one byte. 11000011 starts with two ones — two bytes. 11110000 starts with four — four bytes. The remaining bytes on every line start 10 without exception. Nothing here needs a lookup table; the length is written on the front of the character.
Type any hex value up to 10FFFF. The bytes come from the browser's own TextEncoder, so this is the same encoder your network stack uses — not a re-implementation of it.
Why it won
UTF-8 was designed in 1992 on a placemat in a New Jersey diner, and it beat a field of committee proposals on properties rather than politics. Four of them are checkable, so here they are checked:
$python3.14 utf8scan.py scanned every encodable scalar value: 1112064 byte values that never occur: C0 C1 F5 F6 F7 F8 F9 FA FB FC FD FE FF lowest lead byte 0x00 highest lead byte 0xF4 ascii unchanged: True
- ASCII is a subset, unchanged. Verified above.
- No ASCII byte ever appears inside a multi-byte character. Every byte of a multi-byte sequence has its high bit set, so a C string searching for
/or\0cannot get a false hit in the middle of a Japanese character. Legacy encodings did not have this and it caused security bugs for twenty years. - Byte order does not exist. The unit is one byte, so there is no big-endian or little-endian UTF-8, and no byte-order mark is needed. §06 shows what the alternative looks like.
- Sorting byte strings sorts code points. UTF-8's byte order and Unicode's code point order are the same, so a binary sort of UTF-8 gives you code point order for free.
Entering mid-stream
Point at a random byte in the middle of a UTF-8 file. You can tell immediately whether it starts a character, and if it does not, the start is at most three bytes behind you.
This falls straight out of the templates. A lead byte is 0xxxxxxx, 110xxxxx, 1110xxxx or 11110xxx. A continuation byte is 10xxxxxx. Those two sets do not overlap, so one bit test — b & 0xC0 == 0x80 — answers the question for any byte, with no context at all.
The bytes of héllo wörld, encoded as UTF-8 — click any one of them
What a decoder produces from that byte onwards
Dashed cells are continuation bytes. The output comes from the browser's TextDecoder in its default non-fatal mode — the same thing Python's errors="replace" does, and the same thing your text editor does when it opens a truncated file.
$python3.14 sync.py bytes: 68 C3 A9 6C 6C 6F 20 77 C3 B6 72 6C 64 (13) start at 0 (0x68, start of a character): 'héllo wörld' start at 1 (0xC3, start of a character): 'éllo wörld' start at 2 (0xA9, continuation): '�llo wörld' start at 3 (0x6C, start of a character): 'llo wörld'
U+FFFD followed by nine characters that are all correct. And look at the fourth line — byte 3 is 0x6C, an ordinary ASCII lead byte, so from one byte later there is not even a replacement character. Landing mid-sequence costs one character, never the rest of the stream.What this buys, concretely. grep can search a UTF-8 file without decoding it. A video subtitle stream can be joined halfway through. A corrupt disk sector costs you one character rather than the rest of the file. A parallel program can split a huge file at an arbitrary byte offset and have each worker seek forward a byte or two to the nearest lead byte. None of that is possible in an encoding where the meaning of a byte depends on bytes you have not seen — and the shift-state encodings UTF-8 competed against, like ISO 2022, worked exactly that way.
UTF-16 and the surrogates
.length liesBefore UTF-8 won, the industry had already committed. JavaScript, Java, C#, and the Windows API all decided a character was sixteen bits, shipped that decision into the world, and then Unicode outgrew sixteen bits.
The original plan was clean: 65,536 code points, two bytes each, fixed width. That range is now called the Basic Multilingual Plane, the BMP, and it holds most of the world's living scripts. Then Unicode needed more room. The extension had to be backwards compatible with a fixed two-byte format, and the answer was to carve 2,048 code points out of the BMP and declare them to be halves of pairs.
| Range | Code points | Called | Meaning |
|---|---|---|---|
| U+0000–U+D7FF | 55,296 | BMP, lower part | One UTF-16 code unit each |
| U+D800–U+DBFF | 1,024 | High surrogates | Never a character. First half of a pair. |
| U+DC00–U+DFFF | 1,024 | Low surrogates | Never a character. Second half of a pair. |
| U+E000–U+FFFF | 8,192 | BMP, upper part | One UTF-16 code unit each |
| U+10000–U+10FFFF | 1,048,576 | Supplementary planes | Two UTF-16 code units — a surrogate pair |
The arithmetic for a supplementary code point is four lines, and it is worth doing once by hand because it explains every strange thing JavaScript does with emoji:
$python3.14 utf16.py U+1F468 - 0x10000 = 0x0F468 high = 0xD800 + (v >> 10) = 0xD83D low = 0xDC00 + (v & 0x3FF) = 0xDC68 what the codec emits = D83D DC68 '€'.encode('utf-16' ) = FF FE AC 20 '€'.encode('utf-16-be') = 20 AC '€'.encode('utf-16-le') = AC 20 chr(0xD800).encode('utf-8') -> 'utf-8' codec can't encode character '\ud800' in position 0: surrogates not allowed
utf-16 with no suffix emitted four bytes for a two-byte character, because FF FE is a byte-order mark announcing which way round the pairs are. That problem does not exist in UTF-8. Third: a surrogate is not a character, so UTF-8 refuses to encode one at all.What this does to your programs
JavaScript strings are sequences of UTF-16 code units, and .length counts those units. Not characters. Not code points. Units.
$node jslen.js f.length 11 [...f].length 7 f.charCodeAt(0) 0xd83d f.codePointAt(0) 0x1f468 JSON.stringify(f[0]) "\ud83d"
f is the four-person family emoji from the top of the page, and JavaScript reports its length as eleven. Spreading it with [...f] iterates code points instead and gives seven. And f[0] — indexing a string, the most ordinary operation there is — returns half of a character: a lone high surrogate that is not text and cannot be encoded. Every "my emoji turned into two question marks when I truncated the field" bug in the last twenty years is this line.This is not a JavaScript flaw, it is an age. Java's String.length(), C#'s string.Length, and the Windows W APIs all count UTF-16 code units, and only the first two of those three were designed while 16 bits was the whole plan — surrogates arrived in Unicode 2.0 in 1996. C# came a decade later and inherited UTF-16 anyway, from the CLR, from Windows and from the Java lineage it was answering: by then the cost of not matching the platform you run on was higher than the cost of a lying .Length. Languages that arrived after UTF-8 won made the other choice — Go and Rust store UTF-8 and count bytes, Python 3 counts code points. None of the three is wrong. They are answering three different questions, and §10 puts all of them in one table.
Normalisation
Unicode can write the accented e two ways: as one code point that already has the accent, or as a plain e followed by a combining acute. Both are correct, both render identically, and they are not equal.
This is the puzzle from the top of the page, and it exists for a boring historical reason. Unicode had to round-trip cleanly with the legacy codepages of §02, and those codepages had a single byte for éU+00E9Latin small letter e with acute. So that code point had to exist. But Unicode also needs combining marks in general, because no committee could ever pre-compose every base-plus-accent combination in every script. So both spellings exist, permanently.
$python3.14 norm.py NFC 4 code points 5 utf-8 bytes U+0063 U+0061 U+0066 U+00E9 NFD 5 code points 6 utf-8 bytes U+0063 U+0061 U+0066 U+0065 U+0301 nfc == nfd False ud.normalize('NFC', nfd) == nfc True ud.decomposition('é') 0065 0301 ud.decomposition('fi') <compat> 0066 0069 ud.decomposition('①') <circle> 0031 NFKC('fi') 'fi' NFKC('①') '1' NFC('Å' U+212B) U+00C5
== compares code points, so it says False. The fix is one call: normalise both to the same form first. Nothing about the comparison is broken — it is answering the question you asked, which was about code points, and not the one you meant, which was about words.| Form | What it does | Use it for |
|---|---|---|
| NFD | Decompose: replace every pre-composed character with its base plus marks | Stripping accents, analysing marks |
| NFC | Decompose, then re-compose wherever a single code point exists | The default. Storage, transmission, comparison |
| NFKD | Decompose, and also collapse compatibility differences | Search indexes only |
| NFKC | The same collapse, then re-compose | Search indexes only |
| Form | Code points | Length | UTF-8 |
|---|
Normalisation here is the browser's own String.prototype.normalize, which reads the Unicode data your browser was built with. The listing above is CPython's, reading Unicode 16.0.0. They agree on these five inputs; they are not guaranteed to agree on a character added between their two versions.
NFKC is lossy and people reach for it by accident. Look at the two <compat> lines in the listing: fiU+FB01Latin small ligature fi becomes the two letters f and i, and ①U+2460Circled digit one becomes a plain 1. That is the right behaviour for a search index — you want a search for "fi" to find the ligature. It is the wrong behaviour for anything you will store and hand back, because the circle is gone and you cannot get it back. The distinction is not round-tripping — NFC does not always give you back the code points you started with either, as the note below shows — it is that NFC preserves canonical equivalence: whatever it hands you back is a string Unicode considers to be the same text. NFKC does not even promise that; it deliberately throws distinctions away.
One more that catches people. ÅU+212BAngstrom sign and ÅU+00C5Latin capital letter a with ring above are different code points, and NFC turns the first into the second — a compositional normalisation that changes which character you have, not just how it is spelled. Unicode contains a small number of these deliberate duplicates, inherited from the standards it had to absorb, and NFC is where they get merged.
Grapheme clusters
Ask a person how many characters are in the family emoji and they will say one. It is seven code points, twenty-five bytes and eleven UTF-16 units, and the person is right — there is a fourth layer above all three, and it is the one humans actually mean.
Unicode calls it an extended grapheme cluster, and it is defined by a set of rules in annex UAX #29 for where a break may occur between two code points. Four mechanisms cover almost everything you will meet:
| Mechanism | Example, by code point | Points | Bytes | Clusters |
|---|---|---|---|---|
| Combining marks a mark attaches to the base before it |
U+0065 U+0301 Latin small letter e · Combining acute accent |
2 | 3 | 1 |
| Any number of them there is no limit, and the rule does not care |
U+0065 U+0301 U+0327 U+030A e · acute · cedilla · ring above |
4 | 7 | 1 |
| Zero-width joiner U+200D welds two pictures into one |
U+1F468 U+200D U+1F469 U+200D U+1F467 U+200D U+1F466 Man · ZWJ · Woman · ZWJ · Girl · ZWJ · Boy |
7 | 25 | 1 |
| Emoji modifier a skin-tone code point follows its base |
U+1F44D U+1F3FD Thumbs up sign · Emoji modifier Fitzpatrick type-4 |
2 | 8 | 1 |
| Regional indicator pair two letter-symbols make a flag |
U+1F1EF U+1F1F5 Regional indicator symbol letter J · letter P |
2 | 8 | 1 |
Notice what is not in that table: a flag has no code point of its own. There is no U+FLAG-OF-JAPAN. There are twenty-six regional indicator letters, A to Z, and a flag is a pair of them — which is how Unicode encodes every country without having to adjudicate what a country is.
$node grapheme.js node v26.5.1 · ICU 78.3 · Unicode 17.0 cafe NFC bytes = 5 length = 4 code points = 4 clusters = 4 cafe NFD bytes = 6 length = 5 code points = 5 clusters = 4 family bytes = 25 length = 11 code points = 7 clusters = 1 thumbs up bytes = 8 length = 4 code points = 2 clusters = 1 flag JP bytes = 8 length = 4 code points = 2 clusters = 1 e + 3 marks bytes = 7 length = 4 code points = 4 clusters = 1 hello bytes = 5 length = 5 code points = 5 clusters = 5
Where these cluster counts came from, and why it is not the same tool as everything else on this page. Python's standard library has no grapheme segmentation at all — dir(unicodedata) contains nothing that does it, and there is no len_graphemes to call. So the cluster column above was measured with Intl.Segmenter in Node v26.5.1, which links ICU 78.3 and implements Unicode 17.0, while every name, category and normalisation on this page comes from CPython 3.14.6 at Unicode 16.0.0. Two correct tools on one machine, one version apart. That is not sloppiness in the capture — it is the fifth layer showing through, and §09 is about why it is the layer that moves.
The rule that follows from all of this. Never take a substring, a truncation or a character count from user-facing text at the code point level. "Trim to 20 characters" implemented as s[:20] will cut a family emoji in half in Python, cut a surrogate pair in half in JavaScript, and separate an accent from its letter in both. If it is going in front of a person, segment it; if it is a database key, normalise it; if it is a byte budget, count bytes. Those are three different jobs and one of them is not len().
The name is data
Everything below this line is computation: bits into bytes, bytes into code points, code points into clusters. This layer is a lookup in a table that a committee edits, and it is the only layer that changes when you upgrade anything.
Most assigned code points have a name, and the name is not decoration — it is a stable, permanent, machine-readable identifier that Unicode has promised never to change.
Most, not all. Of the 294,579 assigned code points, 148,853 return a name from unicodedata.name() and 145,726 raise instead: the 137,468 private-use code points, whose meaning is yours to decide and so cannot be named centrally; the 2,048 surrogates and 65 controls, which §01 and §06 have already shown are not characters; and 6,145 ideographs in two runs, U+17000–U+187F7 and U+18D00–U+18D08, for which this Python does not generate one. (Those are the runs of nameless code points, measured — not the block boundaries, which are wider.)
Where a name does exist it is usable as a key, and you can look a character up by one:
$python3.14 aliases.py NUL -> U+0000 BEL -> U+0007 LF -> U+000A CR -> U+000D ESC -> U+001B DEL -> U+007F BELL -> U+1F514
unicodedata.name(chr(7)) raises. What they have instead is a set of formal aliases, and lookup() honours them, which is why BEL resolves to the ASCII bell. But BELL, with two Ls, is the real name of an actual assigned character: the bell emoji, at U+1F514. One letter apart, and 128,269 code points apart. The names are data, and data has edge cases.The name is only the first field. Every code point carries a row of properties, and they are what let a program ask real questions about text without knowing anything about the script:
$python3.14 props.py U+00E9 LATIN SMALL LETTER E WITH ACUTE category=Ll bidirectional=L combining=0 east_asian_width=A numeric=None U+0037 DIGIT SEVEN category=Nd bidirectional=EN combining=0 east_asian_width=Na numeric=7.0 U+0663 ARABIC-INDIC DIGIT THREE category=Nd bidirectional=AN combining=0 east_asian_width=N numeric=3.0 U+4E00 CJK UNIFIED IDEOGRAPH-4E00 category=Lo bidirectional=L combining=0 east_asian_width=W numeric=1.0 U+00BD VULGAR FRACTION ONE HALF category=No bidirectional=ON combining=0 east_asian_width=A numeric=0.5
U+0663 is an Arabic-Indic three with numeric value 3.0 and bidirectional class AN, meaning it lays out right-to-left with the text around it. U+4E00 is a CJK ideograph whose numeric value is 1.0 and whose east-Asian width is W — it occupies two terminal columns. And U+00BD has numeric value 0.5. "Is this a digit" is a property lookup, never a range check against '0' and '9'.Case is a property too, and it is not symmetric
Upper-casing is not "subtract 32", and it is not even a one-to-one map:
$python3.14 case.py 'ß' upper='SS' lower='ß' U+00DF 'fi' upper='FI' lower='fi' U+FB01 'İ' upper='İ' lower='i̇' U+0069 U+0307 'ΟΔΟΣ' upper='ΟΔΟΣ' lower='οδος' U+03BF U+03B4 U+03BF U+03C2
U+0130 lower-cases to two code points. And the Greek word ends in ςU+03C2Greek small letter final sigma rather than the ordinary sigma — the correct lower-case form depends on the letter's position in the word. A case conversion that is a per-character table lookup gets that one wrong.The layer that moves. Every number in this section is true for Unicode 16.0.0 and only for Unicode 16.0.0. A new version arrives roughly yearly; it assigns new code points, and it can change a character's properties. Two programs on the same machine, linked against two different Unicode versions, will disagree — and this page has already shown you exactly that, in §08, where the cluster counts came from a library at Unicode 17.0 while everything else came from a library at 16.0.0. This is why the version is in the eyebrow at the top of the page and in the footer at the bottom. A Unicode claim without a version attached is not a claim.
Four answers to one question
The five layers are now all on the table, and the reason len() means something different in every language you use is that each language picked a different layer and none of them told you which.
| # | Code point | Name | UTF-8 | UTF-16 | Glyph |
|---|
The Glyph column is the only thing on this page allowed to be blank. If it is, the row is still complete — the code point and the name are the character; the picture is your operating system's opinion of it.
What five languages answer, on this machine
Same three strings, five runtimes, all measured on the box this page was built on:
| Language | Expression | Layer it counts | hello | String A | String B | Family |
|---|---|---|---|---|---|---|
| C · gcc 16.1.1 | strlen(s) | byte | 5 | 5 | 6 | 25 |
| Go · 1.26.5 | len(s) | byte | 5 | 5 | 6 | 25 |
| Rust · 1.95.0 | s.len() | byte | 5 | 5 | 6 | 25 |
| JavaScript · node v26.5.1 | s.length | code unit | 5 | 4 | 5 | 11 |
| Rust · 1.95.0 | s.encode_utf16().count() | code unit | 5 | 4 | 5 | 11 |
| Python · 3.14.6 | len(s) | code point | 5 | 4 | 5 | 7 |
| Go · 1.26.5 | utf8.RuneCountInString(s) | code point | 5 | 4 | 5 | 7 |
| Rust · 1.95.0 | s.chars().count() | code point | 5 | 4 | 5 | 7 |
| JavaScript · node v26.5.1 | [...s].length | code point | 5 | 4 | 5 | 7 |
| JavaScript · node v26.5.1 | Intl.Segmenter | grapheme cluster | 5 | 4 | 4 | 1 |
Read the hello column, then read any other column. Ten expressions in five languages, and for ASCII they all say 5. That agreement is why the question never comes up until it does, and why it comes up as a bug rather than as a question. The moment the text has an accent in it the same ten expressions give four different answers, and the only way to pick the right one is to know which layer your problem is on: a byte budget is a byte count, a database key is a normalised code point sequence, and anything a person will read or edit is clusters.
Now read them
Three panels opened this page and none of them could be read. Here they are again, and there is nothing left in them you have not met.
String A · five bytes
| Bytes | Code point | Name | Why those bytes |
|---|---|---|---|
| 63 | U+0063 | Latin small letter c | Below U+0080, so the one-byte template 0xxxxxxx — and identical to its ASCII byte from 1963. |
| 61 | U+0061 | Latin small letter a | The same. |
| 66 | U+0066 | Latin small letter f | The same. |
| C3 A9 | U+00E9 | Latin small letter e with acute | U+00E9 needs 8 bits, so the two-byte template: 11000011 10101001. Concatenate the amber bits and you get 00011101001, which is 0xE9. |
4 code points, 5 UTF-8 bytes, 4 UTF-16 code units, 4 grapheme clusters. This is the NFC spelling — the accented e as a single pre-composed code point, inherited from the codepages of §02.
String B · six bytes
| Bytes | Code point | Name | Why those bytes |
|---|---|---|---|
| 63 | U+0063 | Latin small letter c | One-byte template. |
| 61 | U+0061 | Latin small letter a | One-byte template. |
| 66 | U+0066 | Latin small letter f | One-byte template. |
| 65 | U+0065 | Latin small letter e | A plain e. No accent anywhere in this byte. |
| CC 81 | U+0301 | Combining acute accent | A mark, category Mn, canonical combining class 230. It attaches to the code point before it, which is why the two together are one grapheme cluster. |
5 code points, 6 UTF-8 bytes, 5 UTF-16 code units, 4 grapheme clusters. This is the NFD spelling. It has one more code point and one more byte than String A, it draws the same four marks on the screen, and A == B is False because == compares the third row of the stack and the two strings differ there.
The whole puzzle, in one line. unicodedata.normalize("NFC", B) == A returns True — captured in the listing in §07. The two strings were never ambiguous and nothing was broken; they are two legal spellings of one word, and equality was answering a question about code points while you were asking one about language. Choose a normalisation form at the boundary of your system, apply it to everything coming in, and this class of bug disappears.
The picture · twenty-five bytes
| Offset | Bytes | Code point | Name | UTF-16 |
|---|---|---|---|---|
| 0–3 | F0 9F 91 A8 | U+1F468 | Man | D83D DC68 |
| 4–6 | E2 80 8D | U+200D | Zero width joiner | 200D |
| 7–10 | F0 9F 91 A9 | U+1F469 | Woman | D83D DC69 |
| 11–13 | E2 80 8D | U+200D | Zero width joiner | 200D |
| 14–17 | F0 9F 91 A7 | U+1F467 | Girl | D83D DC67 |
| 18–20 | E2 80 8D | U+200D | Zero width joiner | 200D |
| 21–24 | F0 9F 91 A6 | U+1F466 | Boy | D83D DC66 |
Four people and three joiners. Every person is a supplementary code point above U+FFFF, so each takes the four-byte UTF-8 template and a UTF-16 surrogate pair — D83D DC68 is the pair §06 computed by hand. Every joiner is U+200DZero width joiner, three bytes, invisible, and the only reason the four people are one picture rather than four.
7 code points, 25 UTF-8 bytes, 11 UTF-16 code units, 1 grapheme cluster. Four numbers for one thing a reader would call a single character, and every one of them is the right answer to a question somebody's code is actually asking.
And if your screen showed four separate people, or four empty boxes, nothing above is affected. The rendering is a font's job and this page never depended on it. The bytes are the bytes, the code points are the code points, and the names came out of a database that says the same thing on a machine with no fonts installed at all. That was the point of putting the name next to every character from the first section onward.