./ahmedhashim

Git Packfiles: Anatomy

Git is the cornerstone of every developer’s toolkit, and most of what it does happens below the surface. When cloning a repository, you’ll see something like this:

$ git clone https://github.com/atom/atom
Cloning into 'atom'...
remote: Enumerating objects: 204170, done.
remote: Counting objects: 100% (86/86), done.
remote: Compressing objects: 100% (24/24), done.
remote: Total 204170 (delta 64), reused 62 (delta 62), pack-reused 204084 (from 1)
Receiving objects: 100% (204170/204170), 322.85 MiB | 48.84 MiB/s, done.
Resolving deltas: 100% (144707/144707), done.

The server found 204,170 objects but only counted 86 and compressed 24. The other 204,084 came straight out of a pack. Then your machine spent the last line resolving 144,707 deltas. What exactly is going on there?

All of it is one file format at work: the packfile.

The packfile

A packfile is one file that holds many Git objects. Normally Git stores each object as its own compressed file. The three kinds of object that matter here are:

  1. Blob: the content of one file.
  2. Tree: the listing of one directory, naming the blobs and subtrees inside it.
  3. Commit: a pointer to the root tree, plus the parent commits, author and message.

A repository with a long history ends up with hundreds of thousands of these. A pack puts them all in a single file. Similar objects are stored as small differences against each other, called deltas. An index sits beside the pack so Git can still jump straight to any object by its id.

That solves two problems Git creates for itself:

  1. Size: Every version of every file is stored whole, and SHA-1 turns a one-line change into a completely different id, so without packs the object store grows by a full copy on every edit. Deltas store just the change instead.
  2. Transfer: A clone would have to send those objects one file at a time. Instead, the pack is the thing Git sends: the clone above received a single pack of 204,170 objects, and that one file now sits in its .git/objects/pack.

Let’s take a look inside one.

Loose objects

When Git writes an object, it prefixes the content with a short header like blob 9982\0 and hashes the result. That hash becomes the object’s id. Git then compresses the object with zlib and writes it to .git/objects/, under a directory named for the first two hex characters of the id. That’s a loose object. It’s the simplest write Git can do, because it never has to touch anything already on disk.

Here’s a repository where I committed a 10 KB header file, appended a comment and committed again:

$ du -b .git/objects/*/*
58    .git/objects/a7/732906e8812408a49c16cb765b85300721fb16
59    .git/objects/20/3b24db4881c92152fbd87d0aeacd7c74be79c5
129   .git/objects/14/80ca830fac79e4e47a514654a7cf0e0428cc9c
174   .git/objects/d7/9cd4c394340fdbf839ecbc71e24b7c1384cf3c
4809  .git/objects/83/299d47324f1ed351db0f6f277b97811cb288a4
4825  .git/objects/3f/20c233eae9d1047b393d111f906c76bfd772c0

The four small files are the commits and trees. The two big ones are the blobs: 83299d is the original file at 9,959 bytes, and 3f20c2 is the file with my line at 9,982 bytes. Each compresses to about 4.8 KB on its own, so the line I added cost the full 4.8 KB, because the loose format has no way to say “the same as 83299d, plus a line.”

The other cost is the filesystem. Every object is a file and every read is an open. Two hex characters only allow for 256 directory names, 00 through ff, so a repository with a few hundred thousand objects is a few hundred thousand small files spread across them. That’s a poor layout for a clone, which needs all of them, or git log -p, which needs a lot of them in a row.

Inside a pack

Git fixes both issues by packing. git gc does it on demand. Git also does it on its own after common commands once there are about 6,700 loose objects, a count it estimates by looking in just one of those directories. After a gc, the six loose files are gone and one pack has replaced them:

$ git gc
$ ls .git/objects/pack
pack-3bf68430aae7e0388ad9fc7a7a353d45e1a9e726.idx
pack-3bf68430aae7e0388ad9fc7a7a353d45e1a9e726.pack
pack-3bf68430aae7e0388ad9fc7a7a353d45e1a9e726.rev

git verify-pack -v lists what went into the packfile. The columns are the object id, type, uncompressed size, size in the pack and offset. Delta entries add a chain depth and the base they point at:

$ git verify-pack -v .git/objects/pack/pack-3bf68430aae7e0388ad9fc7a7a353d45e1a9e726.idx
d79cd4c394340fdbf839ecbc71e24b7c1384cf3c commit 190 138 12
1480ca830fac79e4e47a514654a7cf0e0428cc9c commit 145 112 150
3f20c233eae9d1047b393d111f906c76bfd772c0 blob   9982 3491 262
203b24db4881c92152fbd87d0aeacd7c74be79c5 tree   42 53 3753
a7732906e8812408a49c16cb765b85300721fb16 tree   42 52 3806
83299d47324f1ed351db0f6f277b97811cb288a4 blob   7 18 3858 1 3f20c233eae9d1047b393d111f906c76bfd772c0
non delta: 5 objects
chain length = 1: 1 object

The newer blob, 3f20c2, is stored whole in 3,491 bytes. The older one, 83299d, takes 18 bytes as a delta against it. Both versions of the file now cost less than either did loose. Even the whole blob shrank, because loose objects are compressed at zlib level 1 to keep writes fast, while packs use the default level 6.

The delta points backward on purpose. The newest version is what gets read most, so it’s the one stored whole, and reading it only takes a single decompression.

Laid out on disk, the pack is a header, a run of object entries and a checksum. Here’s the demo pack with the offsets from the listing above:

  flowchart LR
    H["<strong>PACK</strong><br/><span class='mermaid-detail'>v2, 6 objects</span>"]
    C["<strong>2 commits</strong><br/><span class='mermaid-detail'>at 12, 150</span>"]
    B1["<strong style='color:#16181a'>blob</strong><br/><span class='mermaid-detail' style='color:#16181a'>3f20c2 at 262</span>"]
    T["<strong>2 trees</strong><br/><span class='mermaid-detail'>at 3753, 3806</span>"]
    B2["<strong style='color:#16181a'>delta</strong><br/><span class='mermaid-detail' style='color:#16181a'>83299d at 3858</span>"]
    X["<strong>SHA-1</strong><br/><span class='mermaid-detail'>trailer</span>"]
    H --> C --> B1 --> T --> B2 --> X
    B2 -.->|"3,596 bytes back"| B1
    classDef whole fill:#5eff6c,stroke:#5eff6c,color:#16181a
    classDef delta fill:#5ea1ff,stroke:#5ea1ff,color:#16181a
    class B1 whole
    class B2 delta

The header is the four bytes PACK, a version (Git only writes version 2) and a 32-bit object count, so a single pack tops out at 2^32 objects. The trailer is a SHA-1 over everything before it, and the pack takes its filename from that hash. Each entry in between starts with a variable-length header that packs the type into three bits alongside the size, then the zlib stream for that one object. Here are the two blob entries from the demo pack, bit by bit:

3f20c2, whole, at offset 262

         +---+-----+------+
  be     | 1 | 011 | 1110 |  more bytes follow, type 3 = blob, size bits 0-3
         +---+-----+------+
  ef     | 1 | 1101111    |  more bytes follow, size bits 4-10
         +---+------------+
  04     | 0 | 0000100    |  last header byte, size bits 11-17
         +---+------------+
  78 9c  | zlib stream    |  3,488 bytes
  ...    +----------------+
                              size = 9,982

83299d, delta, at offset 3858

         +---+-----+------+
  67     | 0 | 110 | 0111 |  last header byte, type 6 = offset delta, size 7
         +---+-----+------+
  9b     | 1 | 0011011    |  more bytes follow, offset bits
         +---+------------+
  0c     | 0 | 0001100    |  last offset byte, offset bits
         +---+------------+
  78 9c  | zlib stream    |  15 bytes
  ...    +----------------+
                              base is 3,596 bytes back, at offset 262

The top bit of each byte says whether another one follows, so a small object needs one header byte and a big one grows as needed. The remaining bits hold the object’s uncompressed size, with the first byte holding the lowest bits and each byte after it holding higher ones, the same order as little-endian. For a delta, the size is the length of the delta data, not the object it rebuilds, and the base’s location follows the header before the zlib stream begins. The format document has the rest of the layouts.

Notice that each entry has its own zlib stream. Git compresses every object separately rather than compressing the whole pack as one stream. Compressing everything together would be smaller, because zlib finds repeated bytes and two objects often share some. But then reading the last object would mean decompressing everything before it. With one stream per object, Git can seek to any entry and decompress only that one.

Deltas

The delta entry above was only 18 bytes, and 15 of those were the zlib stream. Here’s what’s in it. A delta is a list of instructions for rebuilding one object from another, and there are two kinds:

  1. Copy: take a range of bytes from the base.
  2. Insert: add new bytes that aren’t in the base.

The delta starts with the sizes of the base and the target, each as a variable-length integer. Here’s the entire seven-byte delta for the older blob:

fe 4d      base size    9,982
e7 4d      target size  9,959
b0 e7 26   copy 0x26e7 bytes from offset 0

b0 is 1011 0000 in binary. The top bit means copy, and the low seven bits say which of the following offset and size bytes are present. Only the two size bytes are, so the offset is zero and the size is 0x26e7, which is 9,959. The whole delta says “copy the first 9,959 bytes of the base and stop”. The base is the newer file, so stopping 23 bytes short of its end leaves out the line I appended, and what’s left is the original file.

An insert instruction is more expensive. A copy only has to say where to start and how many bytes to take, but an insert has to carry the new bytes themselves, and it can hold at most 127 of them. So a delta that removes content is small while a delta that adds content isn’t. That’s why Git prefers to store the bigger object whole and delta the smaller one from it: the delta only has to leave bytes out.

Deltas work on bytes and have nothing to do with git diff, which Git computes on demand from whole objects. Aditya Mukerjee’s write-up on parsing packs by hand walks every one of these bytes in more detail.

Every delta needs to say which object its base is, and the pack entry can do that in two ways:

  1. Offset delta: store how many bytes back in the same pack the base starts. That’s what the two offset bytes 9b 0c in the entry header earlier were. They encode 3,596, and 3,858 minus 3,596 is 262, the whole blob.
  2. Ref delta: store the base’s 20-byte id instead. Git only does this when the base isn’t in the same pack.

A base can itself be a delta, so deltas form chains. Here’s a file with four versions in one pack. The newest is stored whole, and each older version is a delta against the one after it:

  flowchart LR
    V1["<strong style='color:#16181a'>oldest</strong><br/><span class='mermaid-detail' style='color:#16181a'>delta, depth 3</span>"]
    V2["<strong>v2</strong><br/><span class='mermaid-detail'>delta, depth 2</span>"]
    V3["<strong>v3</strong><br/><span class='mermaid-detail'>delta, depth 1</span>"]
    V4["<strong style='color:#16181a'>newest</strong><br/><span class='mermaid-detail' style='color:#16181a'>whole</span>"]
    V1 -.->|"base"| V2 -.->|"base"| V3 -.->|"base"| V4
    classDef whole fill:#5eff6c,stroke:#5eff6c,color:#16181a
    classDef delta fill:#5ea1ff,stroke:#5ea1ff,color:#16181a
    class V4 whole
    class V1 delta

Reading the newest version is one decompression. Reading the oldest means Git walks back along the arrows until it finds the whole object, decompresses that, then applies each delta on the way out, three of them here. Two settings limit how much work that takes:

  1. pack.depth: caps a chain at 50 deltas.
  2. core.deltaBaseCacheLimit: gives each thread 96 MiB to hold bases it has already decompressed, so reading several objects that share a base only decompresses it once.

Old versions sit at the far end of their chains, so reading old history costs more than reading recent history. Raising pack.depth allows longer chains, which makes the pack smaller and those old reads slower still.

Two things happened above without explanation. Git chose 3f20c2 as the base for 83299d on its own, and verify-pack knew where every object started without reading through the file. How Git picks a base, and how the index finds it come next.