Ctrl + K
Git27 min read

How Git Stores History

A practical explanation of how Git stores project history using commits, trees, blobs, references, object IDs, and parent relationships.

Published: 2026-10-05

Git is often described as a version control system that stores snapshots of a project over time. That description is useful, but it hides an important detail: Git does not primarily store a list of file changes in the way many people imagine. Instead, Git builds a database of objects that describe file contents, directory structures, commits, and relationships between different points in history.

Understanding this internal model makes many Git features easier to understand. Branches, tags, merges, rebases, cherry-picks, detached HEAD states, commit hashes, and even commands such as git reset become much less mysterious once you know what Git is actually storing.

This guide explains Git's storage model from the basic objects upward, including blobs, trees, commits, references, the index, object IDs, parent commits, and the .git directory.

Git Stores Objects, Not Just Diffs

A common mental model is that Git stores a sequence of differences between versions of files. Git can calculate and display diffs, but its fundamental storage model is based on objects representing project data and relationships.

The most important Git object types are blobs, trees, commits, and annotated tags. Together, these objects form the database from which Git reconstructs project history.

ObjectPurpose
BlobStores the contents of a file
TreeStores directory contents and associates names with blobs or other trees
CommitRecords a project snapshot and points to a tree and one or more parent commits
TagStores an annotated reference to another Git object, commonly a commit

The .git Directory

A Git repository normally contains a hidden .git directory at its root. This directory contains the data Git needs to manage the repository's history, references, configuration, index, and other internal information.

ls -la .git

The exact layout of the .git directory can vary between repository configurations and Git versions, but common components include objects, refs, HEAD, index, config, and logs.

PathRole
.git/objectsStores Git objects
.git/refsStores references such as branches and tags
.git/HEADIdentifies the current HEAD reference or commit
.git/indexStores the staging area's current state
.git/configStores repository-specific configuration
.git/logsStores reflog information when reflogs are enabled

What Is a Git Object?

A Git object is a piece of data stored in the repository's object database. Each object is identified by an object ID derived from its contents and type.

Traditionally, Git repositories use SHA-1 object IDs, although Git also supports repositories using SHA-256 object IDs. When developers refer to a commit hash such as a7f3c21, they are referring to an abbreviated representation of an object's identifier.

git rev-parse HEAD

This command prints the object ID of the commit currently referenced by HEAD.

How Git Object IDs Work

A Git object ID is content-derived. Conceptually, Git hashes information that includes the object's type and size together with the object's contents. The resulting identifier is used to locate and reference that object.

This content-addressable design means that the identity of an object depends on what it contains. If the content of an object changes, its object ID changes as well.

💡 This is one reason Git commit IDs change after operations such as rebase. A commit contains references to its parent and tree, so changing the parent relationship creates a different commit object and therefore a different object ID.

What Is a Blob?

A blob is a Git object that stores the contents of a file. The blob does not normally store the file's name or directory location.

For example, two files with exactly the same contents can reference the same blob object. The filename and location are stored by tree objects rather than by the blob itself.

echo "Hello Git" | git hash-object --stdin

The hash-object command can calculate an object ID for data. With the appropriate options, it can also write an object into the repository's object database.

Blobs Do Not Store Filenames

This detail is easy to overlook. A blob represents file content, not the complete file entry you see in a directory.

The association between a filename and its contents is represented by a tree object. This separation allows Git to reuse identical content objects in different locations when appropriate.

What Is a Tree?

A tree object represents a directory at a particular point in history. It contains entries that associate names with other Git objects and include information such as file modes.

A tree can point to blobs for files and to other trees for subdirectories. This allows a complete directory structure to be represented using a hierarchy of Git objects.

git ls-tree HEAD

This command shows the tree entries associated with the root tree of the commit referenced by HEAD.

Trees Store Structure

Suppose a project contains files such as package.json, src/app.ts, and src/utils.ts. Git can represent the root directory with a tree that points to the package.json blob and another tree representing the src directory. The src tree then points to the blobs for app.ts and utils.ts.

The result is a complete representation of the project's directory structure without requiring the commit object itself to contain every file's contents.

What Is a Commit Object?

A commit object records a particular state of the project and connects that state to previous history. A commit points to a tree representing the project's root directory and normally points to one parent commit.

git cat-file -p HEAD

A typical commit contains information such as the tree object, parent commit, author, committer, timestamps, and commit message.

tree 4c8f...
parent a91e...
author Developer <[email protected]> 1750000000 +0000
committer Developer <[email protected]> 1750000000 +0000

Add authentication validation

The exact object IDs and metadata differ for every repository. The important relationship is that the commit points to a tree and its parent or parents.

A Commit Does Not Directly Contain Every File

A commit object is relatively small compared with the complete project snapshot. It points to a tree, and the tree hierarchy points to the blobs containing file contents.

This separation is one of the key ideas behind Git's object model. The commit describes the history event and identifies the project tree, while trees and blobs represent the actual snapshot.

How Git Represents a Commit's Parent

A normal commit has one parent. The parent is the previous commit in that branch's history at the moment the commit was created.

This parent relationship is what allows Git to walk backward through history. Starting from a branch reference, Git can find the commit, then its parent, then that parent's parent, and so on.

git log --oneline

The log command follows these relationships to display the history of the current branch.

Merge Commits Have Multiple Parents

A merge commit differs from an ordinary commit because it can have two or more parents. The multiple parent relationships record that separate lines of development were brought together.

git show --no-patch --pretty=raw <merge-commit>

A merge commit commonly has two parents, although Git's object model supports commits with more than two parents as well.

Git History Is a Graph

Because commits point to their parents, Git history naturally forms a directed acyclic graph. A branch is not a separate copy of the repository. It is simply a reference to a particular commit in that graph.

Branches can therefore diverge when two references point into different parts of the commit graph, and they can converge through merge commits that have multiple parents.

git log --graph --oneline --decorate --all

This command provides a convenient visual representation of the commit relationships, including multiple branches and merge points.

What Is a Git Branch Internally?

A Git branch is essentially a movable reference to a commit. The branch does not contain a separate copy of the project's files or history.

git branch
git rev-parse main

The second command resolves main to the commit currently referenced by that branch.

When you create a new commit while on a branch, Git creates the new commit object and then moves the branch reference forward to point to it.

Why Creating a Branch Is Cheap

Creating a branch does not duplicate the project's entire history. Git only needs another reference pointing to an existing commit.

git switch -c feature/login

At the moment of creation, feature/login points to the same commit as the branch from which it was created. The references can later move independently as new commits are made.

What Is HEAD?

HEAD is a special reference that identifies the currently checked-out position in the repository. In the normal case, HEAD points to a branch, and that branch points to a commit.

cat .git/HEAD

You may see a value similar to a symbolic reference pointing to a branch. This means HEAD follows that branch as the branch advances.

Detached HEAD

A detached HEAD occurs when HEAD points directly to a commit rather than to a branch reference.

git switch --detach <commit>

The commit and its history still exist normally. What changes is the reference through which HEAD identifies the current position. If you create commits in this state without creating another reference, they can become difficult to find after moving elsewhere.

How Git Stores Tags

A lightweight tag is essentially another reference to a Git object, commonly a commit. An annotated tag is different: it is a Git tag object containing metadata and a reference to another object.

git tag v1.0.0
git tag -a v1.0.0 -m "Release 1.0.0"

Tags are useful for naming important points in the commit graph, such as releases.

What Is the Git Index?

The Git index is the staging area. It sits between the working tree and the next commit and records which file contents and metadata should be included when the next commit is created.

git add src/app.ts
git status
git ls-files --stage

The index is not simply a list of filenames. It contains information that allows Git to construct the tree for the next commit.

Working Tree, Index and Repository

Git commonly works with three important states: the working tree, the index, and the committed repository history.

StateWhat it represents
Working treeFiles currently present in the checked-out project directory
IndexThe content selected for the next commit
RepositoryCommitted objects and references stored by Git

The git add command moves selected changes from the working tree into the index. The git commit command then uses the index to construct the tree associated with the new commit.

How a Commit Is Created

When you run git commit, Git takes the state represented by the index and creates the necessary tree objects. It then creates a commit object that points to the resulting root tree and the current parent commit.

git add .
git commit -m "Add user settings"

After the commit is created, the current branch reference moves to the new commit. The previous commit remains in the repository because the new commit points to it as its parent.

Why Old Commits Remain Available

Creating a new commit does not overwrite the previous commit. The new commit simply points to it as a parent. This immutable-object approach is a major part of how Git preserves history.

History can appear to be rewritten when references are moved, for example after reset or rebase. However, the old objects may still exist in the object database for some time even after they are no longer reachable from a normal branch or tag.

Reachable and Unreachable Objects

An object is reachable when Git can find it by following references such as branches and tags and then following the relationships between objects. Objects that are no longer reachable may eventually be removed by Git's garbage collection process.

git fsck --unreachable

Commands such as git fsck can inspect repository objects and help diagnose unusual or damaged repository states. Unreachable objects are not necessarily errors; they can be remnants of previous history after operations such as rebasing or resetting.

How Git Finds Objects

Git's object database traditionally stores loose objects in a directory structure derived from their object IDs. For example, an object whose ID begins with 7a can be associated with a path under .git/objects/7a followed by the remainder of the identifier.

Git also uses packfiles to store many objects efficiently. Packfiles reduce storage overhead and can use delta compression to represent related objects compactly.

Loose Objects and Packfiles

Storage formPurpose
Loose objectIndividual compressed Git object stored separately
PackfileEfficiently stores many objects together
Index fileHelps Git efficiently locate objects inside a pack

You usually do not need to manage these storage details manually. Git handles packing and object storage as part of normal repository maintenance.

What Is Delta Compression?

Many versions of a source file are similar. Storing every version independently can require more space, so Git can use delta compression in packfiles to represent one object's data relative to another related object.

This is an implementation detail of Git's storage system and should not be confused with Git's conceptual history model. Git still exposes commits, trees, and blobs as objects even when their physical storage is optimized internally.

Does Git Store Every Full File on Every Commit?

Conceptually, each commit identifies a complete project snapshot through its tree. However, Git's physical storage does not necessarily keep a completely independent full copy of every file for every commit.

Identical blobs can be reused, and packfiles can store related objects using compression. This means Git can provide snapshot-style semantics without requiring every snapshot to consume the full size of the entire project.

Why Identical File Content Can Be Reused

Because blob identity is based on content, the same file contents can correspond to the same blob object even when those contents appear in different commits.

If a file remains unchanged between two commits, the corresponding tree entries can continue to reference the same blob. Git therefore does not need to create a new blob merely because a new commit exists.

How Git Detects Renames

Git's core object model does not require a special rename object. A rename can be represented by removing one path from a tree and adding another path that points to the same or similar content.

When Git displays history or diffs, it can analyze content similarity and report that a file was renamed. Rename detection is therefore largely a higher-level interpretation of the underlying snapshots and changes.

How Git Calculates a Diff

A diff compares the content represented by two different states. Git can follow the commit and tree relationships to identify the files in each snapshot and then compare their contents.

git diff HEAD^ HEAD

The displayed patch is therefore derived from the underlying snapshots. The patch itself is not the fundamental representation of the commit in Git's object database.

How Merge Uses the Commit Graph

When Git merges two branches, it examines their histories and snapshots to determine a suitable merge base and calculate the changes that need to be combined.

git merge feature

If the histories have diverged and the merge succeeds, Git can create a new commit with both branch tips as parents. This is why merge commits have multiple parent entries.

How Rebase Changes History

Rebase works differently. It takes commits from one part of the graph and reapplies their changes onto another base. The resulting commits are new objects because their parent relationships change.

git switch feature
git rebase main

This explains why rebasing changes commit IDs. The commit message and source changes may look similar, but the commit object itself is different because its ancestry is different.

How Cherry-pick Fits the Object Model

Cherry-pick also creates a new commit. Git takes the changes introduced by an existing commit and applies them to the current branch. The resulting commit has a different parent and therefore a different identity.

git cherry-pick <commit>

This is why cherry-picking does not literally move the original commit from one branch to another. It creates a new commit containing the selected change.

Why Commit Hashes Change After Rebase

A commit's identity depends on the data contained in the commit object, including references to its tree and parent. If a rebase changes the parent, the resulting commit object is different.

This is the underlying reason that rebasing a branch that has already been pushed can require a force push. The remote branch points to the old commit IDs, while the rebased local branch points to newly created commit IDs.

How Git Tags and Releases Relate to History

A release tag gives a human-readable name to an important object in the commit graph. For example, v2.0.0 can identify the commit representing a particular release.

git show v2.0.0
git log v1.0.0..v2.0.0 --oneline

Because tags point to objects, release tooling can use them as stable boundaries for determining which commits belong to a particular release interval.

Git Refs Are Separate From Objects

Branches and tags are references, while commits, trees, and blobs are objects. This distinction is important because moving a branch does not modify the commit object it previously referenced.

For example, after creating a new commit, main moves from the old commit to the new commit. The old commit remains unchanged and still points to its own parent.

Remote Branches Are Also References

Remote-tracking branches such as origin/main are references stored locally that record the state of a corresponding remote branch as last observed by your repository.

git branch -r
git rev-parse origin/main

When you fetch from a remote, Git downloads objects that are needed and updates the appropriate remote-tracking references. The remote repository itself has its own objects and references; your local repository maintains its own copy of the relevant data.

What Happens During Git Fetch?

Git fetch retrieves new objects and updates remote-tracking references without changing your current working branch or working tree.

git fetch origin

After fetching, your local repository may contain commits that were not previously present. The reference origin/main then identifies the latest commit reported by the remote.

What Happens During Git Push?

A push transfers objects and reference updates to a remote repository. Git determines which objects the remote needs and sends the necessary data before updating the remote reference when the operation is allowed.

git push origin main

This means a push is not simply uploading a folder. It is transferring Git objects and requesting an update to a reference.

Why Git Can Be Distributed

A Git repository contains the objects and references necessary to represent its history locally. Developers can therefore inspect commits, branches, and many other aspects of history without continuously contacting a central server.

Remote repositories provide synchronization and collaboration, but the fundamental commit graph and object database can exist locally as well.

Git Object Inspection Commands

Git provides several low-level commands that make the internal storage model visible. They are especially useful when learning how objects are connected.

CommandPurpose
git cat-file -p <object>Print an object's contents in a readable form
git cat-file -t <object>Show the object's type
git ls-tree <tree>Inspect the entries of a tree
git rev-parse <revision>Resolve a revision to an object ID
git show <object>Inspect an object or revision in a convenient form
git fsckCheck repository object connectivity and validity

Inspecting a Commit Step by Step

You can explore the object model starting with HEAD. First, resolve HEAD to its commit ID.

git rev-parse HEAD

Next, inspect the commit itself.

git cat-file -p HEAD

The output reveals the tree and parent references. You can then inspect the tree.

git ls-tree HEAD

A file entry in the tree points to a blob. You can inspect that blob by resolving its object ID and passing it to git cat-file.

git cat-file -p <blob-id>

This sequence demonstrates the core structure: a reference identifies a commit, the commit identifies a tree, the tree identifies file content, and the parent relationship connects the commit to previous history.

Git History Is Mostly Immutable Objects Plus Moving References

One of the most useful mental models for Git is that commits and other objects are generally immutable, while references such as branches are designed to move.

When you create a commit, Git creates new objects and advances a reference. When you create a branch, Git creates another reference. When you reset a branch, Git moves the reference to another commit rather than editing the commits themselves.

How Git Reset Relates to Storage

A reset can move the current branch reference to another commit and optionally change the index and working tree as well.

git reset --soft HEAD^

With a soft reset, the branch reference moves to the parent while the index and working tree remain unchanged. This demonstrates that moving a reference does not destroy the commit object immediately.

Garbage Collection and Repository Cleanup

Git periodically performs maintenance that can pack objects and eventually remove objects that are no longer reachable and have become eligible for cleanup.

git gc

Modern Git versions can perform maintenance automatically in many situations. Manual garbage collection is therefore not something most developers need to run routinely.

⚠️ Do not rely on unreachable objects as a permanent backup. Objects that are no longer referenced can eventually be pruned, so important history should be preserved through appropriate branches, tags, or backups.

Why Git History Can Be Recovered After Mistakes

Operations such as reset, rebase, or deleting a branch can move references away from commits without immediately deleting the underlying objects. Local reflogs can record previous reference positions, which can make recovery possible before unreachable objects are pruned.

git reflog

This is one reason a commit that appears to have disappeared after a history-changing operation may still be recoverable locally.

Why Git History Is Efficient

Git combines several ideas to make repository history efficient: content-addressable objects, reusable blobs and trees, immutable commit relationships, references that move instead of copying history, and packfile compression.

The result is a system that can represent a large number of project snapshots while avoiding unnecessary duplication of identical content and maintaining strong relationships between versions.

A Simple Mental Model of Git Storage

You do not need to memorize every internal detail to work effectively with Git. A compact mental model is enough for most everyday tasks.

  • Blobs represent file contents.
  • Trees represent directory structures.
  • Commits identify project snapshots and connect them to parents.
  • Branches are movable references to commits.
  • Tags give names to important objects.
  • HEAD identifies the current checkout position.
  • The index represents what will be included in the next commit.
  • Object IDs identify Git objects based on their contents.
  • Packfiles optimize the physical storage of many objects.

Common Misconceptions About Git Storage

MisconceptionWhat actually happens
Git stores only diffsGit's object model represents snapshots through trees and blobs; diffs are calculated when needed
A branch contains a copy of the projectA branch is primarily a reference to a commit
A commit contains every file directlyA commit points to a tree that represents the project snapshot
A blob stores a filenameA blob stores file contents; trees associate names with objects
Rebase edits old commitsRebase creates new commits with changed ancestry
Deleting a branch immediately deletes its commitsThe commits may remain reachable through other references or temporarily through reflogs and unreachable-object storage
Push uploads the whole project folderGit transfers required objects and reference updates

How This Knowledge Helps in Everyday Git

Understanding Git's storage model is not only useful for learning Git internals. It explains practical behavior you encounter while working with branches and commits.

  • Branch creation is fast because it creates a reference rather than copying history.
  • Rebase changes commit IDs because it creates commits with different ancestry.
  • Cherry-pick creates new commits because it applies changes to a different history.
  • Merge commits have multiple parents because they connect multiple lines of development.
  • Tags can identify stable points in the commit graph.
  • Reflog can help recover previous reference positions.
  • Git can reuse unchanged file content between snapshots.
  • Fetch can download history without modifying your current branch.

Git Storage and Good Commit Practices

Git's object model also reinforces the value of focused commits. A commit is a named point in the history graph, so its message and contents become part of the long-term record of the project.

Small, logically focused commits are easier to inspect, review, revert, cherry-pick, and understand later. Clear commit messages make those historical objects more useful to both developers and automated release tooling.

Git Storage and Release Workflows

Release workflows often build on the same underlying model. A version tag identifies a commit, commit history describes the changes since an earlier release, and release-note or changelog tooling can derive information from that history.

git log v1.0.0..v1.1.0 --oneline
git diff v1.0.0 v1.1.0

This makes Git history more than a record of who changed a file. It becomes a structured graph that can support development, review, debugging, release management, and automation.

Helpful Git Workflow Tools

Developer tools can simplify several tasks around Git history. A Git commit generator can help create clear commit messages, while a branch name generator can help maintain consistent names for feature, release, and hotfix branches.

Semantic version calculators can help determine version changes in projects following Semantic Versioning. Release notes and changelog generators can turn structured commit history into human-readable release documentation.

Frequently Asked Questions

How does Git store history?

Git stores history as a collection of objects and references. Commits point to tree objects representing project snapshots and to parent commits representing previous history. Trees point to blobs containing file contents, while branches and tags provide references to objects in the repository.

Does Git store full copies of files for every commit?

Git provides snapshot-style semantics, but its physical storage does not require a completely independent full copy of every file for every commit. Unchanged blobs can be reused, and packfiles can compress related objects using deltas.

What is a Git blob?

A blob is a Git object containing file contents. It does not normally contain the filename or directory path. Trees store the relationship between names, file modes, and blobs.

What is a Git tree?

A tree represents a directory at a particular point in history. It contains entries that point to blobs for files and other trees for subdirectories, allowing Git to represent the complete project structure.

What does a Git commit contain?

A commit identifies a tree representing the project snapshot and one or more parent commits. It also contains metadata such as author, committer, timestamps, and the commit message.

Why does rebase change commit hashes?

Rebase creates new commits with different parent relationships. Since a commit's identity depends on information including its parent and tree, changing the ancestry results in different commit IDs.

Is a Git branch a copy of the repository?

No. A branch is primarily a movable reference to a commit. Creating a branch normally requires only another reference to an existing commit rather than copying the repository's history.

Can deleted Git commits be recovered?

Sometimes. Moving or deleting a reference does not necessarily remove the underlying objects immediately. Local reflogs may retain previous reference positions, and unreachable objects can remain temporarily before Git's cleanup process removes them.

Conclusion

Git stores project history using a combination of content-addressable objects and movable references. Blobs contain file contents, trees describe directory structures, and commits connect project snapshots into a history graph. Branches and tags then provide convenient names for important points in that graph.

This model explains many of Git's most important behaviors. Creating a branch is cheap because it creates a reference. Merge commits have multiple parents because they join histories. Rebase and cherry-pick create new commits because they change where existing changes belong in the graph. Tags identify stable points without copying the underlying history.

You do not need to work directly with Git's object database every day, but understanding it gives you a much stronger mental model of the tool. Once commits, trees, blobs, and references make sense, Git's branches, merges, rebases, resets, tags, and recovery mechanisms become much easier to reason about.

Found an issue?

Found an error, outdated information, or something missing from this article? Let me know through the Contact page.

Your feedback helps improve our articles and keep them accurate and useful.