btay.io/wiki

How git stores history

Blobs, trees, commits and refs — why branches are free, why rebasing rewrites, and why nothing is really deleted.

Updated 9 h ago

Git is a content-addressed key-value store with a thin layer of commands on top. Once you know the four object types and what a branch actually is, the commands stop being a list to memorize: reset, rebase, cherry-pick and revert are all the same few moves against the same small data structure.

Four objects

Everything git stores is one of four things, each named by the SHA-1 of its own contents.

ObjectHoldsNamed by
bloba file's bytes, with no name and no historyhash of the contents
treea directory listing: names pointing at blobs and other treeshash of the listing
commitone tree, zero or more parents, author, messagehash of all of that
taga pointer to an object, with a messagehash of that

A commit does not store a diff. It stores a complete snapshot, as one tree. The diffs you read are computed on demand by comparing two snapshots.

Content addressing pays for itself

Because a name is a hash of contents, identical content is stored once. Two branches with the same unchanged file share one blob. A commit that touches one file in a large tree reuses every other tree and blob, and stores only what actually differs.

This is also why history is tamper-evident. A commit's hash covers its tree and its parent's hash, so changing anything in the past changes every hash after it. You cannot quietly edit an old commit. You can only build a new chain and point at that instead.

A branch is a file with a hash in it

This is the part that makes the rest obvious.

cat .git/refs/heads/main       # one line: a 40-character hash

A branch is a ref: a mutable pointer to one commit. Creating a branch writes 41 bytes. That is the whole reason branching in git is instant and branching in older version control was an event.

RefPoints at
refs/heads/mainthe tip of your local main
refs/remotes/origin/mainwhere origin/main was at your last fetch
refs/tags/v1.0a fixed commit
HEADusually the name of the current branch, not a commit

HEAD holding a name rather than a hash is what makes committing advance the branch. Detached HEAD is simply HEAD holding a hash instead, which is why commits made there belong to no branch and are easy to lose.

The three places a change lives

PlaceIsMoved by
Working treeactual files on diskediting
Index (staging)a proposed next treegit add, git restore --staged
HEAD committhe last committed treegit commit, git reset

The index is a real, complete tree, not a list of changed files. git commit writes it out and points the branch at the result. The three modes of reset are just how far down that list the change is pushed: --soft moves the branch only, --mixed also rewrites the index, --hard also rewrites your files.

Why rebase rewrites and merge does not

Merge creates one new commit with two parents. Both original chains survive untouched, and the graph records that they came together.

Rebase replays your commits onto a new base. Replaying means building new commits: same content, different parent, therefore different hashes. The originals are still in the database, but no branch points at them anymore.

That is the whole meaning of "rebasing rewrites history". It does not modify commits, since commits cannot be modified. It abandons them and builds replacements.

commit --amend and cherry-pick do the same thing for the same reason. So the rule about not rebasing shared branches is not superstition: your collaborator's branch points at commits you just orphaned, and git has no way to know the new ones are meant to be the same work.

Nothing is deleted for a while

A commit with no ref pointing at it is unreachable, not gone. Two things save you:

  • The reflog records every position HEAD and each branch has held, for about 90 days. A bad reset, a lost rebase, a deleted branch: all recoverable from git reflog.
  • Garbage collection only removes objects that are both unreachable and older than the reflog's grace period.

This is why almost every "I destroyed my work" moment in git is recoverable, and why the genuine exceptions are the operations that touch files git never saw: git restore over uncommitted edits, or a --hard reset before anything was committed.

The one place this safety net does not extend is someone else's copy. A force push replaces a remote ref, and the old commits stay reachable by hash on that server until its garbage collector runs. Deleting a repository is the only reliable way to remove content already pushed.

Gotchas

  • A commit is a snapshot, not a diff. Every mental model built on "commits contain changes" eventually produces a wrong prediction about rebase or cherry-pick.
  • Renames are not stored. Git infers them by comparing content between two trees, which is why a rename plus a heavy edit shows up as a delete and an add.
  • The index is a tree, so a partial git add commits a state that never existed on disk. Useful deliberately, surprising accidentally.
  • origin/main is a cache, not a live view. It is where the remote was at your last fetch, and it goes stale silently.
  • Empty directories cannot be committed, because trees only contain blobs and other trees, and there is nothing to point at.

Related pages

On this page