skylib index
Hashing & content addressing · verified on git 2.55.0 / CPython 3.14.6

The Name computed from the Thing

A name usually gets assigned. This one gets calculated — from the bytes it names, by anybody, anywhere, forever. Follow that single idea up the stack and you arrive inside git. Every number, listing and digest below was produced by running it.

BytesThe content itself. What is actually there, before anything has been said about it.
Hash functionThe machine. Any bytes in, a fixed number of bits out, the same way every time.
DigestWhat comes out. Never stored as the thing — always recomputed, always checkable.
Content addressThe same digest, used as a name. The turn this page is built around.
Object graphA store whose addresses are identities. Where the guarantee holds — or does not.
Two things you cannot read yet · 01
ce013625030ba8dba906f756967f9e9ca394464a

Git's name for a six-byte file. Forty hexadecimal digits, and not one of them was chosen.

Two things you cannot read yet · 02
78 01 4b ca c9 4f 52 30 63 c8 48 cd c9 c9 e7 02 00 1d c5 04 14

All twenty-one bytes of the file git wrote to disk to hold those six. They are not the six.

01

Forty digits and a six-byte file

The hero · nobody chose this name

Make a file with six bytes in it and ask git what it is called. The answer is forty hexadecimal digits long, and it is not stored anywhere — it is worked out from the six bytes, every time.

bash · actually run
$ printf 'hello\n' > hello.txt
$ wc -c hello.txt
6 hello.txt
$ od -An -tx1 hello.txt
 68 65 6c 6c 6f 0a
$ git hash-object hello.txt
ce013625030ba8dba906f756967f9e9ca394464a
Six bytes went in and forty hex digits came out, and the interesting word is hash-object, not name-object. Git did not look the file up in a table of names it keeps. There is no table. It read the bytes, ran a calculation over them, and printed the result. Run it on a different machine, in a different repository, in ten years: same six bytes, same forty digits.

That is the whole idea of this page, and it is worth naming the reframe before any mechanism arrives. A name is normally assigned. You choose hello.txt; the filesystem writes it down; the name and the file are two separate facts held together by bookkeeping. Break the bookkeeping and the name points at the wrong bytes, or at nothing, and nothing anywhere notices.

The name above was not assigned. It was computed, and it can be computed again by anyone holding the bytes. That single change — a name derived from the thing rather than attached to it — is what the rest of this page follows up the stack, one layer at a time, until it turns into git.

The five layers, and the rail. The page widens: bytes → hash function → digest → content address → object graph. Those are the five colours in the legend, and the chip beside each entry in the contents rail says which one that section is standing on. When the colour of the chip changes, the ground has changed.

Two things in the hero are deliberately unreadable, and the last section reads both. The forty digits are the first. The twenty-one bytes beside them are the second — that is the actual file git wrote into .git/objects/ to hold your six bytes, and every one of the six has vanished from it.

An honest warning about this page's own subject. A forty-digit hex string is the single most convincing thing an author can fabricate: it looks correct because it looks random. Every digest here came out of a command, and the test module beside this file re-derives all of them from hashlib and zlib rather than storing copies — so a wrong digit on the page fails a test rather than sitting here looking plausible.

02

What a hash function actually is

Any input · always the same size out

A hash function eats any number of bytes and produces a fixed number of bits. That sentence contains the whole of its power and the whole of its central problem.

python3 · actually run
>>> import hashlib
>>> # nothing at all, six bytes, and a 412 KiB PDF
0 bytes        -> e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855  (64 hex digits)
6 bytes        -> 5891b5b522d5df086d0ff0b110fbd9d21bb4fc7163af34d08286a2e846f6be03  (64 hex digits)
422,435 bytes  -> 2bb787a73e37352f92383abe7e2902936d1059ad9f1ba6daaa9c1e58ee6970d0  (64 hex digits)
The empty file has a digest. That is not a curiosity, it is the definition working: the function is total, so every byte string has exactly one answer, including the one with no bytes in it. And the 412 KiB PDF gets the same sixty-four digits' worth of answer as the six-byte file. Size in, no size out.

Avalanche: one bit in, half the bits out

A useful hash function does not merely produce a number. It produces a number that has no visible relationship to the input, so that similar inputs do not get similar digests. Flip a single bit of hello\n — the low bit of the h, turning it into an i — and look at what happens.

avalanche.py · actually run
$ python3 avalanche.py
a = b'hello\n' 68656c6c6f0a
b = b'iello\n' 69656c6c6f0a
sha256(a) = 5891b5b522d5df086d0ff0b110fbd9d21bb4fc7163af34d08286a2e846f6be03
sha256(b) = a272663cfa53513ba86ba810d7e3d774f3275cf129954d03d37be025f95cd469
bits differing: 125 of 256
One bit of difference in, 125 bits of difference out — call it half. That is the target, and it is what stops a digest from leaking anything about its input. Two files that differ in one byte are exactly as unrelated, digest-wise, as two files that share nothing at all. There is no "close" in this space.
Avalanche · SHA-1 · flip one bit and watch

The pigeonhole, and how far away it is

A SHA-256 digest is 256 bits, so there are exactly 2256 of them. Now count the possible inputs of just 33 bytes: 33 bytes is 264 bits, so there are 2264 — already 256 times more inputs than there are digests, at one fixed length, before you consider any other length at all. So collisions — two different inputs, one digest — do not merely exist. Averaged across the digests, every single one is already the answer for 256 distinct 33-byte inputs. No hash function can avoid this and none claims to; the claim is only that you cannot find one.

How hard is finding one? Not as hard as guessing a specific digest, which is the thing most people expect. If you only need some pair to match, the birthday problem applies and the cost falls to about the square root of the space. That is measurable, so here it is measured, on truncated SHA-256 so the search finishes in under a second.

birthday.py · actually run
$ python3 birthday.py
# how many inputs before two of them share the first N bits of sha256
bits          tries      sqrt(2**N)   the colliding pair
 8 bits  tries =        10   sqrt(2^8) =        16   n-0 and n-10 share 75
16 bits  tries =       466   sqrt(2^16) =       256   n-57 and n-466 share a6f7
24 bits  tries =      6583   sqrt(2^24) =      4096   n-5524 and n-6583 share a4be93
32 bits  tries =     52500   sqrt(2^32) =     65536   n-48716 and n-52500 share 1a3a72b6
40 bits  tries =    899932   sqrt(2^40) =   1048576   n-448373 and n-899932 share 815244b8fb
Every row lands within a small factor of the square root, which is the birthday bound behaving exactly as advertised. Read the column the other way and it tells you why digests are as long as they are: each 16 extra bits multiply the work by 256. Thirty-two bits collides on a laptop in a tenth of a second. Extend the same curve to 160 bits and the square root is 280 — the most a SHA-1 collision was ever supposed to cost anybody. §05 is about the people who found one for a great deal less than that.
03

The humble checksum

Detecting damage · and which direction is the useful one

The oldest use of a digest is the least glamorous: write it down next to the data, and recompute it later to find out whether the data survived.

python3 · actually run
>>> import zlib
>>> # the file, and the same file with one byte changed
hello.txt                   crc32 = 363a3020
hello.txt (1 byte changed)  crc32 = fb603ebe
One byte moved and the checksum is unrecognisable, which is the whole service being rendered. A CRC is cheap enough to run on every disk block and every network frame, which is exactly where it lives.

Now the part that people get backwards. A checksum gives you certainty in one direction only, and it is not the direction you would guess.

What you observeWhat you may concludeHow sure
digests differthe bytes differCertain. Same bytes always give the same digest, so different digests cannot have come from the same bytes.
digests matchthe bytes matchNot certain. Only overwhelmingly likely — and only if nobody was trying.

That asymmetry is why the useful sentence is negative: a different digest proves different bytes. Everything built on hashing, this page's git included, is ultimately leaning on that one implication, plus an engineering bet that the other direction is close enough to certain for the sizes involved.

How close depends entirely on how many bits you are prepared to carry. Every one of these is a digest of the same six-byte file.

python3 · actually run
>>> hashlib.new(name, b'hello\n')
crc32    32 bits  363a3020
md5     128 bits  b1946ac92492d2347c6235b4d2611184
sha1    160 bits  f572d396fae9206628714fb2ce00f72e94f2258f
sha256  256 bits  5891b5b522d5df086d0ff0b110fbd9d21bb4fc7163af34d08286a2e846f6be03
sha512  512 bits  e7c22b994c59d9cf2b48e549b1e24666636045930d3da7c1acb299d1c3b7f931f94aae41edda2c2b207a36e10f8bcb8d45223e54878f5b316e7ce3b6bc019629
Look hard at the sha1 line, because it is going to matter. That is the SHA-1 digest of the six bytes in hello.txt — and it is not the forty digits the hero showed you, which came from the same file and the same function. Two different answers, no mistake in either. §07 is about the difference.
04

The other use: a digest as a place to put things

Hash tables · why a dict lookup does not search

Long before anybody used a digest as a name, they were using it as an index — and that is why looking a key up in a Python dictionary does not get slower as the dictionary grows.

The trick is small. Keep an array of slots. To store a key, hash it, take the remainder modulo the number of slots, and put it there. To find it again, do the identical calculation and look in the one slot the answer names. You never search; you compute where to look.

Python hands you the ingredient directly. For an integer, hash(n) is n.

python3 · actually run
>>> hash(n), and the bucket it names in an 8-slot table
hash(0         ) = 0           & 7 = 0
hash(1         ) = 1           & 7 = 1
hash(7         ) = 7           & 7 = 7
hash(8         ) = 8           & 7 = 0
hash(42        ) = 42          & 7 = 2
hash(1000003   ) = 1000003     & 7 = 3
hash(-1        ) = -2          & 7 = 6
hash(-2        ) = -2          & 7 = 6
hash(2305843009213693951) = 0           & 7 = 0
hash(2305843009213693952) = 1           & 7 = 1
Three separate facts are hiding in that listing. hash(n) == n is true only while n is small — the last two lines are 261−1 and 261, and the identity has folded, because CPython's integer hash is taken modulo the Mersenne prime 261−1. hash(-1) is −2, not −1, because −1 is CPython's C-level "an error happened" return and no hash may collide with it. And hash(-1) and hash(-2) are therefore equal — a real, permanent, shipped collision in the standard library, which the dictionary handles without drama because handling collisions is its job.

Collisions are the normal case, not the failure case

Eight slots and five keys guarantees crowding. Here is a table doing the entire job in twelve lines, with every probe counted.

Hash table · five inserts into eight slots
minitable.py
The eight slots
minitable.py · actually run
$ python3 minitable.py
insert 42  hash=42  bucket=2 probes=1
insert 7   hash=7   bucket=7 probes=1
insert 15  hash=15  bucket=7 probes=2
insert 8   hash=8   bucket=0 probes=2
insert 23  hash=23  bucket=7 probes=5
table: [15, 8, 42, 23, None, None, None, 7]
lookup 15  bucket=7 probes=2 found=True
lookup 99  bucket=3 probes=2 found=False
Watch 23: five probes to store one key in a table holding four. That is clustering — each collision lengthens the run it lands in, and the next collision has a longer run to walk. It is why real hash tables grow the array well before it fills, and why CPython's probe sequence is not the plain "try the next slot" this script uses but a perturbed jump designed to break exactly this pattern up. The mechanism above is the principle, not CPython's implementation of it.

Even so, the payoff is the point. A list has to compare against every element it holds before it can say no. The table computes one index.

ItemsList: comparisons to reach the last oneTable: probes
10101
1,0001,0001
100,000100,0001

The trap this subject has, and it is a real one

Everything above used integer keys, deliberately. Hash a string in Python and you get a number that is different in every process:

Python randomises string and bytes hashing, so a listing of it would be a lie. Run python3 -c "print(hash('alpha'))" twice and you will get two different numbers. There is no listing of unseeded hash() output anywhere on this page, because it could not reproduce — not for you, and not for the tests that check this file. Every string hash shown below pins PYTHONHASHSEED in the command that produced it.

bash · actually run, seeds pinned
$ for s in 0 1 2; do PYTHONHASHSEED=$s python3 -c "print([hash(k) for k in ('alpha','beta','gamma')])"; done
PYTHONHASHSEED=0  [6408890135650488130, 9063898771175018756, -6346723285187719881]
PYTHONHASHSEED=1  [1489347309953176093, -2695237826086174294, 1405729298976596317]
PYTHONHASHSEED=2  [6606819119385553035, -6467206898529584204, -7777297206166926039]
Three seeds, three completely different sets of answers, and each one perfectly reproducible on its own terms. Pinning the seed is what makes this listing exist at all. Notice what it also proves: the seed is not decoration, it goes into the hash.

The randomisation is a security feature, not an inconvenience. If an attacker knows the seed, they know your buckets — and they can send you a request whose keys all land in one.

seedcollide.py · actually run, seed pinned
$ PYTHONHASHSEED=0 python3 seedcollide.py
PYTHONHASHSEED = 0
key-12    hash % 8 = 3
key-18    hash % 8 = 3
key-19    hash % 8 = 3
key-42    hash % 8 = 3
key-52    hash % 8 = 3
key-55    hash % 8 = 3
scanned: 56
Fifty-six candidates tried, six keys found that all land in bucket 3 — on a laptop, in no time at all. Scale that to a few thousand keys in one POST body and every insert walks the whole cluster: an O(1) data structure quietly degraded to O(n2), which is a denial-of-service delivered as ordinary-looking form data. Randomising the seed per process is what makes the search above impossible to do in advance. It is the same trade the rest of this page makes in reverse — here unpredictability is the goal, and from §05 onward reproducibility is.
05

Cryptographic hashes, and the day SHA-1 broke

Collision resistance · the property that can be lost

A hash function used for bucket indexing only has to spread things out. A hash function used as a name has to survive somebody deliberately attacking it — and that is a different, losable property.

PropertyWhat it forbidsStatus of SHA-1
Preimage resistanceGiven a digest, find any input producing itIntact. No practical attack.
Second-preimage resistanceGiven this file, find a different file with the same digestIntact. No practical attack.
Collision resistanceFind any two files sharing a digest, both of your own choosingBroken since 2017. Demonstrated below.

Those three are usually taught as a list. They are better understood as a ladder of decreasing difficulty, and SHA-1 has lost exactly the bottom rung. The 2017 SHAttered result produced two PDF files — both valid, both displaying different content, both crafted by the researchers — with the same SHA-1 digest. They are still published, so this is not a claim, it is a download.

bash + python3 · actually run
$ curl -sO https://shattered.io/static/shattered-1.pdf
$ curl -sO https://shattered.io/static/shattered-2.pdf
$ python3 -c "..."   # sha1 and sha256 of each
shattered-1.pdf 422435
  sha1   38762cf7f55934b34d179ae6a4c80cadccbb7f0a
  sha256 2bb787a73e37352f92383abe7e2902936d1059ad9f1ba6daaa9c1e58ee6970d0
shattered-2.pdf 422435
  sha1   38762cf7f55934b34d179ae6a4c80cadccbb7f0a
  sha256 d4488775d29bdef7993367d541064dbdda50d383f89f0aa13a6ff2e0894ba5ff
Two different files — SHA-256 says so, loudly — and one SHA-1 digest between them. This is the collision the whole page has been building toward as a threat: the moment where "same digest" stops implying "same bytes". Note what has not happened. Nobody found a second file matching a digest somebody else chose. The attack produced both files together, which is why the ladder above matters more than the headline did.

So why is git still running SHA-1?

Two reasons, and both of them are checkable on this machine. The first is that git does not run plain SHA-1.

bash · actually run
$ git version --build-options
git version 2.55.0
…
zlib-ng: 2.3.3
SHA-1: SHA1_DC
SHA-256: SHA256_BLK
default-hash: sha1
SHA1_DC is the answer: collision-detecting SHA-1. It computes the ordinary digest, while watching the message for the internal disturbance patterns that every known collision attack has to produce, and refuses the input if it sees them. Git has shipped it since 2.13. It does not make SHA-1 collision-resistant again; it makes the specific published attacks fail loudly instead of silently.

The second reason is subtler, and it falls out of the way git hashes things — which is §07's subject, arriving early. Ask git to name those two colliding PDFs:

bash · actually run
$ git hash-object shattered-1.pdf
ba9aaa145ccd24ef760cf31c74d8f7ca1a2e47b0
$ git hash-object shattered-2.pdf
b621eeccd5c7edac9b7dcba35a8d5afd075e24f2
Two files with identical SHA-1 digests, and git gives them different names. Not luck, and not the collision detector either — both commands exited cleanly. Git does not hash the file; it hashes a small header followed by the file (§07). SHAttered is an identical-prefix collision: it works from one specific starting state, and prepending the twelve bytes of blob 422435\0 both changes the chaining state those blocks are fed and knocks them out of alignment with SHA-1's 64-byte block boundaries. The collision does not survive the header. That is a genuine property of this construction and it is also nowhere near a security argument — a chosen-prefix collision, of which SHA-1 has had practical ones since 2020, is not stopped by a known prefix.

The honest position. SHA-1 is unsafe to rely on against an adversary who can supply both sides of a comparison, and git's use of it as a name is being retired — not fixed. The transition target is SHA-256, and the recipe does not change at all, only the function.

bash · actually run
$ git init --object-format=sha256 sha256demo
$ printf 'hello\n' > hello.txt
$ git hash-object hello.txt
2cf8d83d9ee29543b34a87727421fdecb7e3f3a183d337639025de576db9ebb4
$ git config extensions.objectformat
sha256
Sixty-four digits instead of forty, and every idea on this page unchanged. The same six bytes, the same header, the same recipe, a different function bolted into the same slot. That is the practical test of whether a design was about hashing or about SHA-1 — and §07 will show you that this digest, too, is one line of Python.
06

Use the digest as the name

The turn · content addressing

Here is the move the whole page exists for. Stop writing the digest down next to the data. Use it as the data's address.

A normal store maps a name you invented to some bytes: you say hello.txt, it hands you six bytes, and the connection between them is a record someone maintains. A content-addressed store inverts it. You hand it bytes, it computes their digest, and that is where they live. There is no name to invent, and no record to maintain.

Two stores · which direction the arrow runs
NAME-ADDRESSED · THE RECORD IS THE WEAK LINK hello.txt a name you chose a lookup table somebody maintains this 68 65 6c 6c 6f 0a six bytes Corrupt the bytes and the name still points at them. Nothing knows. CONTENT-ADDRESSED · THERE IS NO RECORD TO CORRUPT 68 65 6c 6c 6f 0a six bytes sha1(header + bytes) anyone can run this ce013625…464a the digest IS the address Corrupt the bytes and the address stops matching. Anyone sees it.
Read the arrows, not the boxes. In the top row the arrow crosses a box somebody has to keep honest, so the store's integrity is a promise. In the bottom row the arrow is a calculation, so the store's integrity is arithmetic: the address and the content are the same fact written twice, and any disagreement between them is detectable by whoever is holding them.

Two properties fall out for free

Deduplication. If two files hold the same bytes, they compute the same address, so they are the same object. Not "deduplicated by a background job" — never stored twice in the first place.

bash · actually run
$ printf 'hello\n' > copy.txt        # same six bytes, different filename
$ git add copy.txt && git commit -m second
$ git cat-file -p HEAD^{tree}
100644 blob ce013625030ba8dba906f756967f9e9ca394464a	copy.txt
100644 blob ce013625030ba8dba906f756967f9e9ca394464a	hello.txt
$ git count-objects -v
count: 5
size: 20
in-pack: 0
packs: 0
size-pack: 0
prune-packable: 0
garbage: 0
size-garbage: 0
Two filenames, one address, one object. The repository now holds two commits over two files and contains five objects in total — two commits, two directory listings, and a single blob that both filenames point at. Nothing deduplicated anything; there was never a second copy to remove. (size: 20 is kibibytes of disk as the filesystem reports it and will vary; count: 5 will not.)

Tamper-evidence. If the address is computed from the content, then altering the content breaks the relationship, and anything that recomputes will see it. Here is that happening — and here also is the honest limit of it.

bash · actually run
$ ls -l .git/objects/ce/013625030ba8dba906f756967f9e9ca394464a
-r--r--r-- 1 skydude skydude 21 .git/objects/ce/013625030ba8dba906f756967f9e9ca394464a
$ git fsck && echo clean
clean
$ chmod u+w .git/objects/ce/… && python3 -c "…"   # rewrite 'hello' as 'hellp' inside the object
one byte changed inside the object file
$ git cat-file -p ce013625030ba8dba906f756967f9e9ca394464a
hellp
$ git fsck
error: d7a963a648c4564f03a0952546d2800681628048: hash-path mismatch, found at: .git/objects/ce/013625030ba8dba906f756967f9e9ca394464a
missing blob ce013625030ba8dba906f756967f9e9ca394464a
Read the two commands after the tamper in the order they ran, because they disagree and both are right. cat-file printed the corrupted content without complaint: it trusted the path. fsck recomputed, got d7a963a6…, and reported the object as sitting at an address that is not its own — then said the blob the repository actually needs is missing, which it now is. Content addressing does not prevent tampering. It makes tampering detectable by anyone holding the bytes, which is a different and much cheaper guarantee: no signature, no authority, no trusted third party, just arithmetic anyone can rerun. Note too that git had made the file read-only, which is why the tamper needed a chmod first.
07

The git blob, completely

Seven bytes of header · and the mystery from §03 resolves

Git's object id is a SHA-1 digest, but not of your file. It is a digest of your file with a small header glued to the front — and once you know the header, the forty digits stop being magic.

The recipe is complete in one line. Take the object's type, a space, its length in decimal ASCII, a zero byte, then the content. Hash that.

python3 · actually run
>>> content = b'hello\n'
>>> header  = b'blob %d\x00' % len(content)
>>> store   = header + content
content         b'hello\n' 6 bytes
header          b'blob 6\x00' 7 bytes
store           b'blob 6\x00hello\n' 13 bytes
store hex       62 6c 6f 62 20 36 00 68 65 6c 6c 6f 0a
store ascii     b l o b   6 . h e l l o .
sha1(store)     ce013625030ba8dba906f756967f9e9ca394464a

# and the thing §03 left hanging:
sha1(content)   f572d396fae9206628714fb2ce00f72e94f2258f
The last two lines are the same function over the same file, and they disagree, because only one of them was given the header. f572d396… is what sha1sum hello.txt would print. ce013625… is what git calls the file. Anybody comparing a git object id against a sha1sum and concluding that git is doing something exotic has met exactly these seven bytes.

Why have a header at all? Because git stores four kinds of object in one address space, and without a type in the hash, a blob and a tag holding identical bytes would be the same object. The length is in there for the same class of reason — it makes the framing explicit rather than implied by where the reader chose to stop.

The rule, in one line. object id = sha1(type + " " + length + "\0" + content). Every id in git — blob, tree, commit, tag — is that, with a different word in front and different bytes behind. There is no fifth ingredient and no secret.

bash · actually run
$ git cat-file -t ce013625030ba8dba906f756967f9e9ca394464a
blob
$ git cat-file -s ce013625030ba8dba906f756967f9e9ca394464a
6
$ git cat-file -p ce013625030ba8dba906f756967f9e9ca394464a
hello
Type and size come back out of the store because they went into the hash. cat-file -s is not measuring the file on disk — that file is 21 bytes, as §09 will show. It is reading the number out of the header the id was computed over.
08

Trees and commits: naming things that name things

The object graph · one id pins all of it

A blob is content with no name and no context. A tree gives blobs filenames by listing their ids; a commit gives a tree a time, an author and a history by naming its id and its parent's. The whole structure is objects naming objects by digest.

A tree is not a text file, whatever cat-file -p makes it look like. It is a sequence of entries, each one mode, space, filename, a zero byte, and then twenty raw bytes — the object id in binary, not hex.

bash + python3 · actually run
$ git cat-file tree HEAD^{tree} | wc -c
73
$ git cat-file tree HEAD^{tree} | od -An -tx1
 31 30 30 36 34 34 20 63 6f 70 79 2e 74 78 74 00
 ce 01 36 25 03 0b a8 db a9 06 f7 56 96 7f 9e 9c
 a3 94 46 4a 31 30 30 36 34 34 20 68 65 6c 6c 6f
 2e 74 78 74 00 ce 01 36 25 03 0b a8 db a9 06 f7
 56 96 7f 9e 9c a3 94 46 4a
$ # od wraps at 16; the same 73 bytes regrouped by tree entry instead
 31 30 30 36 34 34 20 63 6f 70 79 2e 74 78 74 00 ce 01 36 25 03 0b a8 db a9 06 f7 56 96 7f 9e 9c a3 94 46 4a
 31 30 30 36 34 34 20 68 65 6c 6c 6f 2e 74 78 74 00 ce 01 36 25 03 0b a8 db a9 06 f7 56 96 7f 9e 9c a3 94 46 4a
$ # and as printable characters, dots for the rest
100644 copy.txt...6%.......V......FJ100644 hello.txt...6%.......V......FJ
$ # and rehashed with a 'tree' header instead of a 'blob' one
sha1(b'tree 73\x00' + entries) = 60595ed1e5f5f1f2f414a66105b50067ef61f943
Find ce 01 36 25 in the second grouping — it is there twice. od wraps at sixteen bytes, which cuts an entry in half; the regrouped rows below it are the identical 73 bytes cut at the entry boundary instead, and are the only line here that is not literal output. Those are the first four of the twenty raw bytes of ce013625…464a, once per filename, because both files are the same content. The tree is not describing the blob; it is holding the blob's address, and its own address is a hash of that holding. Change which blob a filename points at and the tree's id changes, necessarily.

A commit does the same thing one level up, and it is plain text.

bash · actually run
$ git cat-file commit HEAD
tree 60595ed1e5f5f1f2f414a66105b50067ef61f943
parent 44e27d33d4885a2c07a97b46bdaefe1d5a1da81e
author skylib <skylib@example.com> 1785628801 +0000
committer skylib <skylib@example.com> 1785628801 +0000

second
$ git cat-file commit HEAD | wc -c
209
$ # rehashed with a 'commit' header
sha1(b'commit 209\x00' + body) = eca56247a5d114cc1cd6e8eddc0088646dd47ec7
Two of those lines are addresses and the rest is metadata, and all of it is inside the hash. The commit names one tree and one parent. That is the entire mechanism by which git has history: not a diff, not a version number — a link, stored as a digest, from a thing to the thing before it.
The object graph · every arrow is a digest, every id is real
COMMITS commit · first 44e27d3 NO PARENT commit · second eca5624 PARENT 44e27d3 PARENT TREES tree aaa96ce HELLO.TXT tree 60595ed COPY.TXT + HELLO.TXT TREE TREE BLOB · ONE COPY blob ce01362 68 65 6C 6C 6F 0A HELLO.TXT COPY.TXT Both filenames hold the same six bytes, so both trees name the same object. There was never a second copy to remove. Every arrow is a 20-byte digest stored inside the box it leaves. Nothing here holds a pointer, a path, or a version number.
Follow any arrow backwards and you are computing, not looking up. Both trees name the same blob because both filenames hold the same six bytes — the deduplication in §06 is visible here as two arrows into one box. And the second commit's id was computed over text containing its parent's id, which was computed over text containing its tree's id, which was computed over bytes containing the blob's id.

That is the entire repository, and it is five objects. Here they are, all of them, with the ids the diagram abbreviates written out in full — every one of which you can reproduce by following the recipe in the footer.

bash · actually run
$ git cat-file --batch-all-objects --batch-check='%(objectname) %(objecttype) %(objectsize)'
44e27d33d4885a2c07a97b46bdaefe1d5a1da81e commit 160
60595ed1e5f5f1f2f414a66105b50067ef61f943 tree 73
aaa96ced2d9a1c8e72c56b253a0e2fe78393feb7 tree 37
ce013625030ba8dba906f756967f9e9ca394464a blob 6
eca56247a5d114cc1cd6e8eddc0088646dd47ec7 commit 209
The listing arrives sorted by address, which is the only order a content-addressed store has. Not by date, not by name, not by the order they were written — there is no such field to sort on. Note the sizes: 6 bytes of your content, 37 and 73 bytes of directory listing, 160 and 209 bytes of commit. The history costs more than the file, which is exactly what you would expect of two commits over six bytes and stops being true immediately.

Why one id pins everything

That chain has a consequence people find surprising the first time. Since each id is computed over content that contains the ids below it, you cannot change anything anywhere in the history without changing every id above it, all the way to the tip. Change one hex digit of the tree line in a commit and watch:

python3 · actually run
>>> raw = the 209 bytes of commit HEAD, exactly as git stores them
>>> sha1(b'commit 209\x00' + raw)                       = eca56247a5d114cc1cd6e8eddc0088646dd47ec7
>>> bad = raw.replace(b'tree 60595ed', b'tree 60595ee')   # one digit
>>> sha1(b'commit 209\x00' + bad)                       = ec7a604ef00d8ab1d93c4f3d336caf8c1dcba350
One character of the 209 changed, and the commit is a different commit. Not a modified version of the old one — a different object, at a different address, which nothing in the repository refers to. This is why quoting a single forty-digit commit id is enough to pin an entire tree of files and the whole history behind it: to hand you different bytes under that id, somebody would need a SHA-1 collision against an id you already have, which is the second-preimage problem from §05's table and is not something anyone can do.
09

What is actually sitting on the disk

Back down to bytes · the file is not the content

The address decides the path: first two hex digits are a directory, the remaining thirty-eight are the filename. But the file at that path does not contain your six bytes, and it does not contain the thirteen either.

bash · actually run
$ find .git/objects -type f | sort
.git/objects/44/e27d33d4885a2c07a97b46bdaefe1d5a1da81e
.git/objects/60/595ed1e5f5f1f2f414a66105b50067ef61f943
.git/objects/aa/a96ced2d9a1c8e72c56b253a0e2fe78393feb7
.git/objects/ce/013625030ba8dba906f756967f9e9ca394464a
.git/objects/ec/a56247a5d114cc1cd6e8eddc0088646dd47ec7
Five files, five addresses, and the split after two digits is pure filesystem pragmatism. A directory holding a million entries is slow on most filesystems; 256 directories holding a few thousand each is not. The 256 comes from those two hex digits and nothing else.

Open the one belonging to hello.txt and you get the twenty-one bytes from the hero.

python3 · actually run
>>> raw = open('.git/objects/ce/013625030ba8dba906f756967f9e9ca394464a','rb').read()
>>> len(raw)
21
>>> raw.hex()
78014bcac94f523063c848cdc9c9e702001dc50414
>>> zlib.decompress(raw)
b'blob 6\x00hello\n'
>>> zlib.compress(b'blob 6\x00hello\n', 1) == raw
True
Twenty-one bytes on disk to hold thirteen, and the six you wrote are nowhere in them. The file is a zlib stream: 78 01 is the zlib header for the fastest compression setting, and the four bytes at the end are an Adler-32 checksum of what was compressed. Git writes loose objects at compression level 1 by default, which is why one line of Python reproduces the file byte for byte — and why the last line above is True rather than "close enough".

The address is not a hash of the file. It is a hash of what the file decompresses to — and this is exactly why the object's own bytes are allowed to change. Recompress it at a different level and every one of those 21 bytes can move; the object id does not, because the id was never about them. That is the layer this section stands on: the bytes are the bottom of the stack again, and they are not the content.

Loose, and then packed

One file per object is fine for five objects and hopeless for five hundred thousand. Git eventually rewrites them into a single pack, and the whole scheme survives the move intact.

bash · actually run
$ git gc -q
$ git count-objects -v
count: 0
size: 0
in-pack: 5
packs: 1
…
$ find .git/objects -type f | sort
.git/objects/info/commit-graph
.git/objects/info/packs
.git/objects/pack/pack-af6d4703fd7ce4ab456add77560922684fe31552.idx
.git/objects/pack/pack-af6d4703fd7ce4ab456add77560922684fe31552.pack
.git/objects/pack/pack-af6d4703fd7ce4ab456add77560922684fe31552.rev
$ git cat-file -p ce013625030ba8dba906f756967f9e9ca394464a
hello
Every loose file is gone, the last command still works, and the pack's own filename is a digest too. Inside a pack, objects that resemble each other are stored as deltas against one another rather than in full, and the .idx beside it is the map from object id to offset — so the lookup is still "compute the address, go straight there". Nothing above this layer noticed the change: not the tree, not the commit, not the forty digits.
——

Derive it yourself

Both unreadable things from the hero · read

You now have every piece. Two things at the top of this page were deliberately opaque; here they are, byte by byte, using nothing that was not on the page above.

The first thing: forty digits

Thirteen bytes go into SHA-1. Seven of them are a header git added; six of them you typed.

#bytehexwhere it came from
0b62the object's type
1l6cthe object's type
2o6fthe object's type
3b62the object's type
4 20a space
5636the content length, in decimal ASCII
6NUL00end of header
7h68your file
8e65your file
9l6cyour file
10l6cyour file
11o6fyour file
12LF0ayour file
Object id builder · the whole recipe, live

The SHA-1 in that tool is twenty-five lines of JavaScript in this file, not a browser API — a page about what a hash function does should be running one. Type hello and a newline into it and it prints the hero's forty digits, because there is nothing else in the recipe.

python3 · actually run · the same thing, four lines
>>> import hashlib
>>> content = b'hello\n'
>>> hashlib.sha1(b'blob %d\x00' % len(content) + content).hexdigest()
'ce013625030ba8dba906f756967f9e9ca394464a'
$ git hash-object hello.txt
ce013625030ba8dba906f756967f9e9ca394464a
Four lines of standard library reproduce git's naming scheme exactly. Not an approximation of it, and not a reimplementation — the same arithmetic, because there is only one thing being done here and no part of it is hidden.

The second thing: twenty-one bytes

And the other unreadable panel from the hero, which is what git wrote to disk under that address.

python3 · actually run
>>> zlib.decompress(bytes.fromhex('78014bcac94f523063c848cdc9c9e702001dc50414'))
b'blob 6\x00hello\n'
>>> hashlib.sha1(_).hexdigest()
'ce013625030ba8dba906f756967f9e9ca394464a'
>>> # which is the path it is stored at:
.git/objects/ce/013625030ba8dba906f756967f9e9ca394464a
The loop closes on itself. Inflate the file, hash what comes out, and you get the name of the file you inflated. That is the whole of content addressing in three lines: the address is a claim about the content, the content is right there, and checking the claim costs one function call that anybody can make.

What you can now say about a repository you have never seen. Hand someone forty hex digits and they can tell whether the bytes you gave them are the bytes you meant — without trusting you, the network, the disk, or the server it came from. That is the same property in every layer above: a hash function makes it computable, a digest makes it small, a content address makes it a name, and an object graph makes it cover an entire history rather than one file. It is one idea, followed until it changed shape.