Forty digits and a six-byte file
Make a file with six bytes in it and ask git what it is called. The answer is forty hexadecimal digits long, and it is not stored anywhere — it is worked out from the six bytes, every time.
$printf 'hello\n' > hello.txt$wc -c hello.txt 6 hello.txt$od -An -tx1 hello.txt 68 65 6c 6c 6f 0a$git hash-object hello.txt ce013625030ba8dba906f756967f9e9ca394464a
That is the whole idea of this page, and it is worth naming the reframe before any mechanism arrives. A name is normally assigned. You choose hello.txt; the filesystem writes it down; the name and the file are two separate facts held together by bookkeeping. Break the bookkeeping and the name points at the wrong bytes, or at nothing, and nothing anywhere notices.
The name above was not assigned. It was computed, and it can be computed again by anyone holding the bytes. That single change — a name derived from the thing rather than attached to it — is what the rest of this page follows up the stack, one layer at a time, until it turns into git.
The five layers, and the rail. The page widens: bytes → hash function → digest → content address → object graph. Those are the five colours in the legend, and the chip beside each entry in the contents rail says which one that section is standing on. When the colour of the chip changes, the ground has changed.
Two things in the hero are deliberately unreadable, and the last section reads both. The forty digits are the first. The twenty-one bytes beside them are the second — that is the actual file git wrote into .git/objects/ to hold your six bytes, and every one of the six has vanished from it.
An honest warning about this page's own subject. A forty-digit hex string is the single most convincing thing an author can fabricate: it looks correct because it looks random. Every digest here came out of a command, and the test module beside this file re-derives all of them from hashlib and zlib rather than storing copies — so a wrong digit on the page fails a test rather than sitting here looking plausible.
What a hash function actually is
A hash function eats any number of bytes and produces a fixed number of bits. That sentence contains the whole of its power and the whole of its central problem.
>>>import hashlib>>> # nothing at all, six bytes, and a 412 KiB PDF0 bytes -> e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855(64 hex digits)6 bytes -> 5891b5b522d5df086d0ff0b110fbd9d21bb4fc7163af34d08286a2e846f6be03(64 hex digits)422,435 bytes -> 2bb787a73e37352f92383abe7e2902936d1059ad9f1ba6daaa9c1e58ee6970d0(64 hex digits)
Avalanche: one bit in, half the bits out
A useful hash function does not merely produce a number. It produces a number that has no visible relationship to the input, so that similar inputs do not get similar digests. Flip a single bit of hello\n — the low bit of the h, turning it into an i — and look at what happens.
$python3 avalanche.py a = b'hello\n' 68656c6c6f0a b = b'iello\n' 69656c6c6f0a sha256(a) = 5891b5b522d5df086d0ff0b110fbd9d21bb4fc7163af34d08286a2e846f6be03 sha256(b) = a272663cfa53513ba86ba810d7e3d774f3275cf129954d03d37be025f95cd469 bits differing: 125 of 256
The pigeonhole, and how far away it is
A SHA-256 digest is 256 bits, so there are exactly 2256 of them. Now count the possible inputs of just 33 bytes: 33 bytes is 264 bits, so there are 2264 — already 256 times more inputs than there are digests, at one fixed length, before you consider any other length at all. So collisions — two different inputs, one digest — do not merely exist. Averaged across the digests, every single one is already the answer for 256 distinct 33-byte inputs. No hash function can avoid this and none claims to; the claim is only that you cannot find one.
How hard is finding one? Not as hard as guessing a specific digest, which is the thing most people expect. If you only need some pair to match, the birthday problem applies and the cost falls to about the square root of the space. That is measurable, so here it is measured, on truncated SHA-256 so the search finishes in under a second.
$python3 birthday.py# how many inputs before two of them share the first N bits of sha256bits tries sqrt(2**N) the colliding pair8 bits tries = 10 sqrt(2^8) = 16 n-0 and n-10 share 75 16 bits tries = 466 sqrt(2^16) = 256 n-57 and n-466 share a6f7 24 bits tries = 6583 sqrt(2^24) = 4096 n-5524 and n-6583 share a4be93 32 bits tries = 52500 sqrt(2^32) = 65536 n-48716 and n-52500 share 1a3a72b6 40 bits tries = 899932 sqrt(2^40) = 1048576 n-448373 and n-899932 share 815244b8fb
The humble checksum
The oldest use of a digest is the least glamorous: write it down next to the data, and recompute it later to find out whether the data survived.
>>>import zlib>>> # the file, and the same file with one byte changedhello.txt crc32 = 363a3020 hello.txt (1 byte changed) crc32 = fb603ebe
Now the part that people get backwards. A checksum gives you certainty in one direction only, and it is not the direction you would guess.
| What you observe | What you may conclude | How sure |
|---|---|---|
| digests differ | the bytes differ | Certain. Same bytes always give the same digest, so different digests cannot have come from the same bytes. |
| digests match | the bytes match | Not certain. Only overwhelmingly likely — and only if nobody was trying. |
That asymmetry is why the useful sentence is negative: a different digest proves different bytes. Everything built on hashing, this page's git included, is ultimately leaning on that one implication, plus an engineering bet that the other direction is close enough to certain for the sizes involved.
How close depends entirely on how many bits you are prepared to carry. Every one of these is a digest of the same six-byte file.
>>>hashlib.new(name, b'hello\n')crc32 32 bits 363a3020 md5 128 bits b1946ac92492d2347c6235b4d2611184 sha1 160 bits f572d396fae9206628714fb2ce00f72e94f2258f sha256 256 bits 5891b5b522d5df086d0ff0b110fbd9d21bb4fc7163af34d08286a2e846f6be03 sha512 512 bits e7c22b994c59d9cf2b48e549b1e24666636045930d3da7c1acb299d1c3b7f931f94aae41edda2c2b207a36e10f8bcb8d45223e54878f5b316e7ce3b6bc019629
hello.txt — and it is not the forty digits the hero showed you, which came from the same file and the same function. Two different answers, no mistake in either. §07 is about the difference.The other use: a digest as a place to put things
Long before anybody used a digest as a name, they were using it as an index — and that is why looking a key up in a Python dictionary does not get slower as the dictionary grows.
The trick is small. Keep an array of slots. To store a key, hash it, take the remainder modulo the number of slots, and put it there. To find it again, do the identical calculation and look in the one slot the answer names. You never search; you compute where to look.
Python hands you the ingredient directly. For an integer, hash(n) is n.
>>>hash(n), and the bucket it names in an 8-slot tablehash(0 ) = 0 & 7 = 0 hash(1 ) = 1 & 7 = 1 hash(7 ) = 7 & 7 = 7 hash(8 ) = 8 & 7 = 0 hash(42 ) = 42 & 7 = 2 hash(1000003 ) = 1000003 & 7 = 3 hash(-1 ) = -2 & 7 = 6 hash(-2 ) = -2 & 7 = 6 hash(2305843009213693951) = 0 & 7 = 0 hash(2305843009213693952) = 1 & 7 = 1
hash(n) == n is true only while n is small — the last two lines are 261−1 and 261, and the identity has folded, because CPython's integer hash is taken modulo the Mersenne prime 261−1. hash(-1) is −2, not −1, because −1 is CPython's C-level "an error happened" return and no hash may collide with it. And hash(-1) and hash(-2) are therefore equal — a real, permanent, shipped collision in the standard library, which the dictionary handles without drama because handling collisions is its job.Collisions are the normal case, not the failure case
Eight slots and five keys guarantees crowding. Here is a table doing the entire job in twelve lines, with every probe counted.
minitable.py
The eight slots
$python3 minitable.py insert 42 hash=42 bucket=2 probes=1 insert 7 hash=7 bucket=7 probes=1 insert 15 hash=15 bucket=7 probes=2 insert 8 hash=8 bucket=0 probes=2 insert 23 hash=23 bucket=7 probes=5 table: [15, 8, 42, 23, None, None, None, 7] lookup 15 bucket=7 probes=2 found=True lookup 99 bucket=3 probes=2 found=False
23: five probes to store one key in a table holding four. That is clustering — each collision lengthens the run it lands in, and the next collision has a longer run to walk. It is why real hash tables grow the array well before it fills, and why CPython's probe sequence is not the plain "try the next slot" this script uses but a perturbed jump designed to break exactly this pattern up. The mechanism above is the principle, not CPython's implementation of it.Even so, the payoff is the point. A list has to compare against every element it holds before it can say no. The table computes one index.
| Items | List: comparisons to reach the last one | Table: probes |
|---|---|---|
| 10 | 10 | 1 |
| 1,000 | 1,000 | 1 |
| 100,000 | 100,000 | 1 |
The trap this subject has, and it is a real one
Everything above used integer keys, deliberately. Hash a string in Python and you get a number that is different in every process:
Python randomises string and bytes hashing, so a listing of it would be a lie. Run python3 -c "print(hash('alpha'))" twice and you will get two different numbers. There is no listing of unseeded hash() output anywhere on this page, because it could not reproduce — not for you, and not for the tests that check this file. Every string hash shown below pins PYTHONHASHSEED in the command that produced it.
$for s in 0 1 2; do PYTHONHASHSEED=$s python3 -c "print([hash(k) for k in ('alpha','beta','gamma')])"; done PYTHONHASHSEED=0 [6408890135650488130, 9063898771175018756, -6346723285187719881] PYTHONHASHSEED=1 [1489347309953176093, -2695237826086174294, 1405729298976596317] PYTHONHASHSEED=2 [6606819119385553035, -6467206898529584204, -7777297206166926039]
The randomisation is a security feature, not an inconvenience. If an attacker knows the seed, they know your buckets — and they can send you a request whose keys all land in one.
$PYTHONHASHSEED=0 python3 seedcollide.py PYTHONHASHSEED = 0 key-12 hash % 8 = 3 key-18 hash % 8 = 3 key-19 hash % 8 = 3 key-42 hash % 8 = 3 key-52 hash % 8 = 3 key-55 hash % 8 = 3 scanned: 56
Cryptographic hashes, and the day SHA-1 broke
A hash function used for bucket indexing only has to spread things out. A hash function used as a name has to survive somebody deliberately attacking it — and that is a different, losable property.
| Property | What it forbids | Status of SHA-1 |
|---|---|---|
| Preimage resistance | Given a digest, find any input producing it | Intact. No practical attack. |
| Second-preimage resistance | Given this file, find a different file with the same digest | Intact. No practical attack. |
| Collision resistance | Find any two files sharing a digest, both of your own choosing | Broken since 2017. Demonstrated below. |
Those three are usually taught as a list. They are better understood as a ladder of decreasing difficulty, and SHA-1 has lost exactly the bottom rung. The 2017 SHAttered result produced two PDF files — both valid, both displaying different content, both crafted by the researchers — with the same SHA-1 digest. They are still published, so this is not a claim, it is a download.
$curl -sO https://shattered.io/static/shattered-1.pdf$curl -sO https://shattered.io/static/shattered-2.pdf$python3 -c "..."# sha1 and sha256 of eachshattered-1.pdf 422435 sha1 38762cf7f55934b34d179ae6a4c80cadccbb7f0a sha256 2bb787a73e37352f92383abe7e2902936d1059ad9f1ba6daaa9c1e58ee6970d0 shattered-2.pdf 422435 sha1 38762cf7f55934b34d179ae6a4c80cadccbb7f0a sha256 d4488775d29bdef7993367d541064dbdda50d383f89f0aa13a6ff2e0894ba5ff
So why is git still running SHA-1?
Two reasons, and both of them are checkable on this machine. The first is that git does not run plain SHA-1.
$git version --build-options git version 2.55.0…zlib-ng: 2.3.3 SHA-1: SHA1_DC SHA-256: SHA256_BLK default-hash: sha1
SHA1_DC is the answer: collision-detecting SHA-1. It computes the ordinary digest, while watching the message for the internal disturbance patterns that every known collision attack has to produce, and refuses the input if it sees them. Git has shipped it since 2.13. It does not make SHA-1 collision-resistant again; it makes the specific published attacks fail loudly instead of silently.The second reason is subtler, and it falls out of the way git hashes things — which is §07's subject, arriving early. Ask git to name those two colliding PDFs:
$git hash-object shattered-1.pdf ba9aaa145ccd24ef760cf31c74d8f7ca1a2e47b0$git hash-object shattered-2.pdf b621eeccd5c7edac9b7dcba35a8d5afd075e24f2
blob 422435\0 both changes the chaining state those blocks are fed and knocks them out of alignment with SHA-1's 64-byte block boundaries. The collision does not survive the header. That is a genuine property of this construction and it is also nowhere near a security argument — a chosen-prefix collision, of which SHA-1 has had practical ones since 2020, is not stopped by a known prefix.The honest position. SHA-1 is unsafe to rely on against an adversary who can supply both sides of a comparison, and git's use of it as a name is being retired — not fixed. The transition target is SHA-256, and the recipe does not change at all, only the function.
$git init --object-format=sha256 sha256demo$printf 'hello\n' > hello.txt$git hash-object hello.txt 2cf8d83d9ee29543b34a87727421fdecb7e3f3a183d337639025de576db9ebb4$git config extensions.objectformat sha256
Use the digest as the name
Here is the move the whole page exists for. Stop writing the digest down next to the data. Use it as the data's address.
A normal store maps a name you invented to some bytes: you say hello.txt, it hands you six bytes, and the connection between them is a record someone maintains. A content-addressed store inverts it. You hand it bytes, it computes their digest, and that is where they live. There is no name to invent, and no record to maintain.
Two properties fall out for free
Deduplication. If two files hold the same bytes, they compute the same address, so they are the same object. Not "deduplicated by a background job" — never stored twice in the first place.
$printf 'hello\n' > copy.txt# same six bytes, different filename$git add copy.txt && git commit -m second$git cat-file -p HEAD^{tree} 100644 blob ce013625030ba8dba906f756967f9e9ca394464a copy.txt 100644 blob ce013625030ba8dba906f756967f9e9ca394464a hello.txt$git count-objects -v count: 5 size: 20 in-pack: 0 packs: 0 size-pack: 0 prune-packable: 0 garbage: 0 size-garbage: 0
size: 20 is kibibytes of disk as the filesystem reports it and will vary; count: 5 will not.)Tamper-evidence. If the address is computed from the content, then altering the content breaks the relationship, and anything that recomputes will see it. Here is that happening — and here also is the honest limit of it.
$ls -l .git/objects/ce/013625030ba8dba906f756967f9e9ca394464a -r--r--r-- 1 skydude skydude 21 .git/objects/ce/013625030ba8dba906f756967f9e9ca394464a$git fsck && echo clean clean$chmod u+w .git/objects/ce/… && python3 -c "…"# rewrite 'hello' as 'hellp' inside the objectone byte changed inside the object file$git cat-file -p ce013625030ba8dba906f756967f9e9ca394464a hellp$git fsck error: d7a963a648c4564f03a0952546d2800681628048: hash-path mismatch, found at: .git/objects/ce/013625030ba8dba906f756967f9e9ca394464a missing blob ce013625030ba8dba906f756967f9e9ca394464a
cat-file printed the corrupted content without complaint: it trusted the path. fsck recomputed, got d7a963a6…, and reported the object as sitting at an address that is not its own — then said the blob the repository actually needs is missing, which it now is. Content addressing does not prevent tampering. It makes tampering detectable by anyone holding the bytes, which is a different and much cheaper guarantee: no signature, no authority, no trusted third party, just arithmetic anyone can rerun. Note too that git had made the file read-only, which is why the tamper needed a chmod first.The git blob, completely
Git's object id is a SHA-1 digest, but not of your file. It is a digest of your file with a small header glued to the front — and once you know the header, the forty digits stop being magic.
The recipe is complete in one line. Take the object's type, a space, its length in decimal ASCII, a zero byte, then the content. Hash that.
>>>content = b'hello\n'>>>header = b'blob %d\x00' % len(content)>>>store = header + content content b'hello\n' 6 bytes header b'blob 6\x00' 7 bytes store b'blob 6\x00hello\n' 13 bytes store hex 62 6c 6f 62 20 36 00 68 65 6c 6c 6f 0a store ascii b l o b 6 . h e l l o . sha1(store) ce013625030ba8dba906f756967f9e9ca394464a# and the thing §03 left hanging:sha1(content) f572d396fae9206628714fb2ce00f72e94f2258f
f572d396… is what sha1sum hello.txt would print. ce013625… is what git calls the file. Anybody comparing a git object id against a sha1sum and concluding that git is doing something exotic has met exactly these seven bytes.Why have a header at all? Because git stores four kinds of object in one address space, and without a type in the hash, a blob and a tag holding identical bytes would be the same object. The length is in there for the same class of reason — it makes the framing explicit rather than implied by where the reader chose to stop.
The rule, in one line. object id = sha1(type + " " + length + "\0" + content). Every id in git — blob, tree, commit, tag — is that, with a different word in front and different bytes behind. There is no fifth ingredient and no secret.
$git cat-file -t ce013625030ba8dba906f756967f9e9ca394464a blob$git cat-file -s ce013625030ba8dba906f756967f9e9ca394464a 6$git cat-file -p ce013625030ba8dba906f756967f9e9ca394464a hello
cat-file -s is not measuring the file on disk — that file is 21 bytes, as §09 will show. It is reading the number out of the header the id was computed over.Trees and commits: naming things that name things
A blob is content with no name and no context. A tree gives blobs filenames by listing their ids; a commit gives a tree a time, an author and a history by naming its id and its parent's. The whole structure is objects naming objects by digest.
A tree is not a text file, whatever cat-file -p makes it look like. It is a sequence of entries, each one mode, space, filename, a zero byte, and then twenty raw bytes — the object id in binary, not hex.
$git cat-file tree HEAD^{tree} | wc -c 73$git cat-file tree HEAD^{tree} | od -An -tx1 31 30 30 36 34 34 20 63 6f 70 79 2e 74 78 74 00 ce 01 36 25 03 0b a8 db a9 06 f7 56 96 7f 9e 9c a3 94 46 4a 31 30 30 36 34 34 20 68 65 6c 6c 6f 2e 74 78 74 00 ce 01 36 25 03 0b a8 db a9 06 f7 56 96 7f 9e 9c a3 94 46 4a$# od wraps at 16; the same 73 bytes regrouped by tree entry instead31 30 30 36 34 34 20 63 6f 70 79 2e 74 78 74 00 ce 01 36 25 03 0b a8 db a9 06 f7 56 96 7f 9e 9c a3 94 46 4a 31 30 30 36 34 34 20 68 65 6c 6c 6f 2e 74 78 74 00 ce 01 36 25 03 0b a8 db a9 06 f7 56 96 7f 9e 9c a3 94 46 4a$# and as printable characters, dots for the rest100644 copy.txt...6%.......V......FJ100644 hello.txt...6%.......V......FJ$# and rehashed with a 'tree' header instead of a 'blob' onesha1(b'tree 73\x00' + entries) = 60595ed1e5f5f1f2f414a66105b50067ef61f943
ce 01 36 25 in the second grouping — it is there twice. od wraps at sixteen bytes, which cuts an entry in half; the regrouped rows below it are the identical 73 bytes cut at the entry boundary instead, and are the only line here that is not literal output. Those are the first four of the twenty raw bytes of ce013625…464a, once per filename, because both files are the same content. The tree is not describing the blob; it is holding the blob's address, and its own address is a hash of that holding. Change which blob a filename points at and the tree's id changes, necessarily.A commit does the same thing one level up, and it is plain text.
$git cat-file commit HEAD tree 60595ed1e5f5f1f2f414a66105b50067ef61f943 parent 44e27d33d4885a2c07a97b46bdaefe1d5a1da81e author skylib <skylib@example.com> 1785628801 +0000 committer skylib <skylib@example.com> 1785628801 +0000 second$git cat-file commit HEAD | wc -c 209$# rehashed with a 'commit' headersha1(b'commit 209\x00' + body) = eca56247a5d114cc1cd6e8eddc0088646dd47ec7
That is the entire repository, and it is five objects. Here they are, all of them, with the ids the diagram abbreviates written out in full — every one of which you can reproduce by following the recipe in the footer.
$git cat-file --batch-all-objects --batch-check='%(objectname) %(objecttype) %(objectsize)' 44e27d33d4885a2c07a97b46bdaefe1d5a1da81e commit 160 60595ed1e5f5f1f2f414a66105b50067ef61f943 tree 73 aaa96ced2d9a1c8e72c56b253a0e2fe78393feb7 tree 37 ce013625030ba8dba906f756967f9e9ca394464a blob 6 eca56247a5d114cc1cd6e8eddc0088646dd47ec7 commit 209
Why one id pins everything
That chain has a consequence people find surprising the first time. Since each id is computed over content that contains the ids below it, you cannot change anything anywhere in the history without changing every id above it, all the way to the tip. Change one hex digit of the tree line in a commit and watch:
>>>raw =the 209 bytes of commit HEAD, exactly as git stores them>>>sha1(b'commit 209\x00' + raw) = eca56247a5d114cc1cd6e8eddc0088646dd47ec7>>>bad = raw.replace(b'tree 60595ed', b'tree 60595ee')# one digit>>>sha1(b'commit 209\x00' + bad) = ec7a604ef00d8ab1d93c4f3d336caf8c1dcba350
What is actually sitting on the disk
The address decides the path: first two hex digits are a directory, the remaining thirty-eight are the filename. But the file at that path does not contain your six bytes, and it does not contain the thirteen either.
$find .git/objects -type f | sort .git/objects/44/e27d33d4885a2c07a97b46bdaefe1d5a1da81e .git/objects/60/595ed1e5f5f1f2f414a66105b50067ef61f943 .git/objects/aa/a96ced2d9a1c8e72c56b253a0e2fe78393feb7 .git/objects/ce/013625030ba8dba906f756967f9e9ca394464a .git/objects/ec/a56247a5d114cc1cd6e8eddc0088646dd47ec7
Open the one belonging to hello.txt and you get the twenty-one bytes from the hero.
>>>raw = open('.git/objects/ce/013625030ba8dba906f756967f9e9ca394464a','rb').read()>>>len(raw) 21>>>raw.hex() 78014bcac94f523063c848cdc9c9e702001dc50414>>>zlib.decompress(raw) b'blob 6\x00hello\n'>>>zlib.compress(b'blob 6\x00hello\n', 1) == raw True
78 01 is the zlib header for the fastest compression setting, and the four bytes at the end are an Adler-32 checksum of what was compressed. Git writes loose objects at compression level 1 by default, which is why one line of Python reproduces the file byte for byte — and why the last line above is True rather than "close enough".The address is not a hash of the file. It is a hash of what the file decompresses to — and this is exactly why the object's own bytes are allowed to change. Recompress it at a different level and every one of those 21 bytes can move; the object id does not, because the id was never about them. That is the layer this section stands on: the bytes are the bottom of the stack again, and they are not the content.
Loose, and then packed
One file per object is fine for five objects and hopeless for five hundred thousand. Git eventually rewrites them into a single pack, and the whole scheme survives the move intact.
$git gc -q$git count-objects -v count: 0 size: 0 in-pack: 5 packs: 1…$find .git/objects -type f | sort .git/objects/info/commit-graph .git/objects/info/packs .git/objects/pack/pack-af6d4703fd7ce4ab456add77560922684fe31552.idx .git/objects/pack/pack-af6d4703fd7ce4ab456add77560922684fe31552.pack .git/objects/pack/pack-af6d4703fd7ce4ab456add77560922684fe31552.rev$git cat-file -p ce013625030ba8dba906f756967f9e9ca394464a hello
.idx beside it is the map from object id to offset — so the lookup is still "compute the address, go straight there". Nothing above this layer noticed the change: not the tree, not the commit, not the forty digits.Derive it yourself
You now have every piece. Two things at the top of this page were deliberately opaque; here they are, byte by byte, using nothing that was not on the page above.
The first thing: forty digits
Thirteen bytes go into SHA-1. Seven of them are a header git added; six of them you typed.
| # | byte | hex | where it came from |
|---|---|---|---|
| 0 | b | 62 | the object's type |
| 1 | l | 6c | the object's type |
| 2 | o | 6f | the object's type |
| 3 | b | 62 | the object's type |
| 4 | 20 | a space | |
| 5 | 6 | 36 | the content length, in decimal ASCII |
| 6 | NUL | 00 | end of header |
| 7 | h | 68 | your file |
| 8 | e | 65 | your file |
| 9 | l | 6c | your file |
| 10 | l | 6c | your file |
| 11 | o | 6f | your file |
| 12 | LF | 0a | your file |
The SHA-1 in that tool is twenty-five lines of JavaScript in this file, not a browser API — a page about what a hash function does should be running one. Type hello and a newline into it and it prints the hero's forty digits, because there is nothing else in the recipe.
>>>import hashlib>>>content = b'hello\n'>>>hashlib.sha1(b'blob %d\x00' % len(content) + content).hexdigest() 'ce013625030ba8dba906f756967f9e9ca394464a'$git hash-object hello.txt ce013625030ba8dba906f756967f9e9ca394464a
The second thing: twenty-one bytes
And the other unreadable panel from the hero, which is what git wrote to disk under that address.
>>>zlib.decompress(bytes.fromhex('78014bcac94f523063c848cdc9c9e702001dc50414')) b'blob 6\x00hello\n'>>>hashlib.sha1(_).hexdigest() 'ce013625030ba8dba906f756967f9e9ca394464a'>>># which is the path it is stored at:.git/objects/ce/013625030ba8dba906f756967f9e9ca394464a
What you can now say about a repository you have never seen. Hand someone forty hex digits and they can tell whether the bytes you gave them are the bytes you meant — without trusting you, the network, the disk, or the server it came from. That is the same property in every layer above: a hash function makes it computable, a digest makes it small, a content address makes it a name, and an object graph makes it cover an entire history rather than one file. It is one idea, followed until it changed shape.