How Git Stores History
A practical explanation of how Git stores project history using commits, trees, blobs, references, object IDs, and parent relationships.
Git is often described as a version control system that stores snapshots of a project over time. That description is useful, but it hides an important detail: Git does not primarily store a list of file changes in the way many people imagine. Instead, Git builds a database of objects that describe file contents, directory structures, commits, and relationships between different points in history.
Understanding this internal model makes many Git features easier to understand. Branches, tags, merges, rebases, cherry-picks, detached HEAD states, commit hashes, and even commands such as git reset become much less mysterious once you know what Git is actually storing.
This guide explains Git's storage model from the basic objects upward, including blobs, trees, commits, references, the index, object IDs, parent commits, and the .git directory.
Git Stores Objects, Not Just Diffs
A common mental model is that Git stores a sequence of differences between versions of files. Git can calculate and display diffs, but its fundamental storage model is based on objects representing project data and relationships.
The most important Git object types are blobs, trees, commits, and annotated tags. Together, these objects form the database from which Git reconstructs project history.
| Object | Purpose |
|---|---|
| Blob | Stores the contents of a file |
| Tree | Stores directory contents and associates names with blobs or other trees |
| Commit | Records a project snapshot and points to a tree and one or more parent commits |
| Tag | Stores an annotated reference to another Git object, commonly a commit |
The .git Directory
A Git repository normally contains a hidden .git directory at its root. This directory contains the data Git needs to manage the repository's history, references, configuration, index, and other internal information.
ls -la .gitThe exact layout of the .git directory can vary between repository configurations and Git versions, but common components include objects, refs, HEAD, index, config, and logs.
| Path | Role |
|---|---|
| .git/objects | Stores Git objects |
| .git/refs | Stores references such as branches and tags |
| .git/HEAD | Identifies the current HEAD reference or commit |
| .git/index | Stores the staging area's current state |
| .git/config | Stores repository-specific configuration |
| .git/logs | Stores reflog information when reflogs are enabled |
What Is a Git Object?
A Git object is a piece of data stored in the repository's object database. Each object is identified by an object ID derived from its contents and type.
Traditionally, Git repositories use SHA-1 object IDs, although Git also supports repositories using SHA-256 object IDs. When developers refer to a commit hash such as a7f3c21, they are referring to an abbreviated representation of an object's identifier.
git rev-parse HEADThis command prints the object ID of the commit currently referenced by HEAD.
How Git Object IDs Work
A Git object ID is content-derived. Conceptually, Git hashes information that includes the object's type and size together with the object's contents. The resulting identifier is used to locate and reference that object.
This content-addressable design means that the identity of an object depends on what it contains. If the content of an object changes, its object ID changes as well.
What Is a Blob?
A blob is a Git object that stores the contents of a file. The blob does not normally store the file's name or directory location.
For example, two files with exactly the same contents can reference the same blob object. The filename and location are stored by tree objects rather than by the blob itself.
echo "Hello Git" | git hash-object --stdinThe hash-object command can calculate an object ID for data. With the appropriate options, it can also write an object into the repository's object database.
Blobs Do Not Store Filenames
This detail is easy to overlook. A blob represents file content, not the complete file entry you see in a directory.
The association between a filename and its contents is represented by a tree object. This separation allows Git to reuse identical content objects in different locations when appropriate.
What Is a Tree?
A tree object represents a directory at a particular point in history. It contains entries that associate names with other Git objects and include information such as file modes.
A tree can point to blobs for files and to other trees for subdirectories. This allows a complete directory structure to be represented using a hierarchy of Git objects.
git ls-tree HEADThis command shows the tree entries associated with the root tree of the commit referenced by HEAD.
Trees Store Structure
Suppose a project contains files such as package.json, src/app.ts, and src/utils.ts. Git can represent the root directory with a tree that points to the package.json blob and another tree representing the src directory. The src tree then points to the blobs for app.ts and utils.ts.
The result is a complete representation of the project's directory structure without requiring the commit object itself to contain every file's contents.
What Is a Commit Object?
A commit object records a particular state of the project and connects that state to previous history. A commit points to a tree representing the project's root directory and normally points to one parent commit.
git cat-file -p HEADA typical commit contains information such as the tree object, parent commit, author, committer, timestamps, and commit message.
tree 4c8f...
parent a91e...
author Developer <[email protected]> 1750000000 +0000
committer Developer <[email protected]> 1750000000 +0000
Add authentication validationThe exact object IDs and metadata differ for every repository. The important relationship is that the commit points to a tree and its parent or parents.
A Commit Does Not Directly Contain Every File
A commit object is relatively small compared with the complete project snapshot. It points to a tree, and the tree hierarchy points to the blobs containing file contents.
This separation is one of the key ideas behind Git's object model. The commit describes the history event and identifies the project tree, while trees and blobs represent the actual snapshot.
How Git Represents a Commit's Parent
A normal commit has one parent. The parent is the previous commit in that branch's history at the moment the commit was created.
This parent relationship is what allows Git to walk backward through history. Starting from a branch reference, Git can find the commit, then its parent, then that parent's parent, and so on.
git log --onelineThe log command follows these relationships to display the history of the current branch.
Merge Commits Have Multiple Parents
A merge commit differs from an ordinary commit because it can have two or more parents. The multiple parent relationships record that separate lines of development were brought together.
git show --no-patch --pretty=raw <merge-commit>A merge commit commonly has two parents, although Git's object model supports commits with more than two parents as well.
Git History Is a Graph
Because commits point to their parents, Git history naturally forms a directed acyclic graph. A branch is not a separate copy of the repository. It is simply a reference to a particular commit in that graph.
Branches can therefore diverge when two references point into different parts of the commit graph, and they can converge through merge commits that have multiple parents.
git log --graph --oneline --decorate --allThis command provides a convenient visual representation of the commit relationships, including multiple branches and merge points.
What Is a Git Branch Internally?
A Git branch is essentially a movable reference to a commit. The branch does not contain a separate copy of the project's files or history.
git branch
git rev-parse mainThe second command resolves main to the commit currently referenced by that branch.
When you create a new commit while on a branch, Git creates the new commit object and then moves the branch reference forward to point to it.
Why Creating a Branch Is Cheap
Creating a branch does not duplicate the project's entire history. Git only needs another reference pointing to an existing commit.
git switch -c feature/loginAt the moment of creation, feature/login points to the same commit as the branch from which it was created. The references can later move independently as new commits are made.
What Is HEAD?
HEAD is a special reference that identifies the currently checked-out position in the repository. In the normal case, HEAD points to a branch, and that branch points to a commit.
cat .git/HEADYou may see a value similar to a symbolic reference pointing to a branch. This means HEAD follows that branch as the branch advances.
Detached HEAD
A detached HEAD occurs when HEAD points directly to a commit rather than to a branch reference.
git switch --detach <commit>The commit and its history still exist normally. What changes is the reference through which HEAD identifies the current position. If you create commits in this state without creating another reference, they can become difficult to find after moving elsewhere.
How Git Stores Tags
A lightweight tag is essentially another reference to a Git object, commonly a commit. An annotated tag is different: it is a Git tag object containing metadata and a reference to another object.
git tag v1.0.0
git tag -a v1.0.0 -m "Release 1.0.0"Tags are useful for naming important points in the commit graph, such as releases.
What Is the Git Index?
The Git index is the staging area. It sits between the working tree and the next commit and records which file contents and metadata should be included when the next commit is created.
git add src/app.ts
git status
git ls-files --stageThe index is not simply a list of filenames. It contains information that allows Git to construct the tree for the next commit.
Working Tree, Index and Repository
Git commonly works with three important states: the working tree, the index, and the committed repository history.
| State | What it represents |
|---|---|
| Working tree | Files currently present in the checked-out project directory |
| Index | The content selected for the next commit |
| Repository | Committed objects and references stored by Git |
The git add command moves selected changes from the working tree into the index. The git commit command then uses the index to construct the tree associated with the new commit.
How a Commit Is Created
When you run git commit, Git takes the state represented by the index and creates the necessary tree objects. It then creates a commit object that points to the resulting root tree and the current parent commit.
git add .
git commit -m "Add user settings"After the commit is created, the current branch reference moves to the new commit. The previous commit remains in the repository because the new commit points to it as its parent.
Why Old Commits Remain Available
Creating a new commit does not overwrite the previous commit. The new commit simply points to it as a parent. This immutable-object approach is a major part of how Git preserves history.
History can appear to be rewritten when references are moved, for example after reset or rebase. However, the old objects may still exist in the object database for some time even after they are no longer reachable from a normal branch or tag.
Reachable and Unreachable Objects
An object is reachable when Git can find it by following references such as branches and tags and then following the relationships between objects. Objects that are no longer reachable may eventually be removed by Git's garbage collection process.
git fsck --unreachableCommands such as git fsck can inspect repository objects and help diagnose unusual or damaged repository states. Unreachable objects are not necessarily errors; they can be remnants of previous history after operations such as rebasing or resetting.
How Git Finds Objects
Git's object database traditionally stores loose objects in a directory structure derived from their object IDs. For example, an object whose ID begins with 7a can be associated with a path under .git/objects/7a followed by the remainder of the identifier.
Git also uses packfiles to store many objects efficiently. Packfiles reduce storage overhead and can use delta compression to represent related objects compactly.
Loose Objects and Packfiles
| Storage form | Purpose |
|---|---|
| Loose object | Individual compressed Git object stored separately |
| Packfile | Efficiently stores many objects together |
| Index file | Helps Git efficiently locate objects inside a pack |
You usually do not need to manage these storage details manually. Git handles packing and object storage as part of normal repository maintenance.
What Is Delta Compression?
Many versions of a source file are similar. Storing every version independently can require more space, so Git can use delta compression in packfiles to represent one object's data relative to another related object.
This is an implementation detail of Git's storage system and should not be confused with Git's conceptual history model. Git still exposes commits, trees, and blobs as objects even when their physical storage is optimized internally.
Does Git Store Every Full File on Every Commit?
Conceptually, each commit identifies a complete project snapshot through its tree. However, Git's physical storage does not necessarily keep a completely independent full copy of every file for every commit.
Identical blobs can be reused, and packfiles can store related objects using compression. This means Git can provide snapshot-style semantics without requiring every snapshot to consume the full size of the entire project.
Why Identical File Content Can Be Reused
Because blob identity is based on content, the same file contents can correspond to the same blob object even when those contents appear in different commits.
If a file remains unchanged between two commits, the corresponding tree entries can continue to reference the same blob. Git therefore does not need to create a new blob merely because a new commit exists.
How Git Detects Renames
Git's core object model does not require a special rename object. A rename can be represented by removing one path from a tree and adding another path that points to the same or similar content.
When Git displays history or diffs, it can analyze content similarity and report that a file was renamed. Rename detection is therefore largely a higher-level interpretation of the underlying snapshots and changes.
How Git Calculates a Diff
A diff compares the content represented by two different states. Git can follow the commit and tree relationships to identify the files in each snapshot and then compare their contents.
git diff HEAD^ HEADThe displayed patch is therefore derived from the underlying snapshots. The patch itself is not the fundamental representation of the commit in Git's object database.
How Merge Uses the Commit Graph
When Git merges two branches, it examines their histories and snapshots to determine a suitable merge base and calculate the changes that need to be combined.
git merge featureIf the histories have diverged and the merge succeeds, Git can create a new commit with both branch tips as parents. This is why merge commits have multiple parent entries.
How Rebase Changes History
Rebase works differently. It takes commits from one part of the graph and reapplies their changes onto another base. The resulting commits are new objects because their parent relationships change.
git switch feature
git rebase mainThis explains why rebasing changes commit IDs. The commit message and source changes may look similar, but the commit object itself is different because its ancestry is different.
How Cherry-pick Fits the Object Model
Cherry-pick also creates a new commit. Git takes the changes introduced by an existing commit and applies them to the current branch. The resulting commit has a different parent and therefore a different identity.
git cherry-pick <commit>This is why cherry-picking does not literally move the original commit from one branch to another. It creates a new commit containing the selected change.
Why Commit Hashes Change After Rebase
A commit's identity depends on the data contained in the commit object, including references to its tree and parent. If a rebase changes the parent, the resulting commit object is different.
This is the underlying reason that rebasing a branch that has already been pushed can require a force push. The remote branch points to the old commit IDs, while the rebased local branch points to newly created commit IDs.
How Git Tags and Releases Relate to History
A release tag gives a human-readable name to an important object in the commit graph. For example, v2.0.0 can identify the commit representing a particular release.
git show v2.0.0
git log v1.0.0..v2.0.0 --onelineBecause tags point to objects, release tooling can use them as stable boundaries for determining which commits belong to a particular release interval.
Git Refs Are Separate From Objects
Branches and tags are references, while commits, trees, and blobs are objects. This distinction is important because moving a branch does not modify the commit object it previously referenced.
For example, after creating a new commit, main moves from the old commit to the new commit. The old commit remains unchanged and still points to its own parent.
Remote Branches Are Also References
Remote-tracking branches such as origin/main are references stored locally that record the state of a corresponding remote branch as last observed by your repository.
git branch -r
git rev-parse origin/mainWhen you fetch from a remote, Git downloads objects that are needed and updates the appropriate remote-tracking references. The remote repository itself has its own objects and references; your local repository maintains its own copy of the relevant data.
What Happens During Git Fetch?
Git fetch retrieves new objects and updates remote-tracking references without changing your current working branch or working tree.
git fetch originAfter fetching, your local repository may contain commits that were not previously present. The reference origin/main then identifies the latest commit reported by the remote.
What Happens During Git Push?
A push transfers objects and reference updates to a remote repository. Git determines which objects the remote needs and sends the necessary data before updating the remote reference when the operation is allowed.
git push origin mainThis means a push is not simply uploading a folder. It is transferring Git objects and requesting an update to a reference.
Why Git Can Be Distributed
A Git repository contains the objects and references necessary to represent its history locally. Developers can therefore inspect commits, branches, and many other aspects of history without continuously contacting a central server.
Remote repositories provide synchronization and collaboration, but the fundamental commit graph and object database can exist locally as well.
Git Object Inspection Commands
Git provides several low-level commands that make the internal storage model visible. They are especially useful when learning how objects are connected.
| Command | Purpose |
|---|---|
| git cat-file -p <object> | Print an object's contents in a readable form |
| git cat-file -t <object> | Show the object's type |
| git ls-tree <tree> | Inspect the entries of a tree |
| git rev-parse <revision> | Resolve a revision to an object ID |
| git show <object> | Inspect an object or revision in a convenient form |
| git fsck | Check repository object connectivity and validity |
Inspecting a Commit Step by Step
You can explore the object model starting with HEAD. First, resolve HEAD to its commit ID.
git rev-parse HEADNext, inspect the commit itself.
git cat-file -p HEADThe output reveals the tree and parent references. You can then inspect the tree.
git ls-tree HEADA file entry in the tree points to a blob. You can inspect that blob by resolving its object ID and passing it to git cat-file.
git cat-file -p <blob-id>This sequence demonstrates the core structure: a reference identifies a commit, the commit identifies a tree, the tree identifies file content, and the parent relationship connects the commit to previous history.
Git History Is Mostly Immutable Objects Plus Moving References
One of the most useful mental models for Git is that commits and other objects are generally immutable, while references such as branches are designed to move.
When you create a commit, Git creates new objects and advances a reference. When you create a branch, Git creates another reference. When you reset a branch, Git moves the reference to another commit rather than editing the commits themselves.
How Git Reset Relates to Storage
A reset can move the current branch reference to another commit and optionally change the index and working tree as well.
git reset --soft HEAD^With a soft reset, the branch reference moves to the parent while the index and working tree remain unchanged. This demonstrates that moving a reference does not destroy the commit object immediately.
Garbage Collection and Repository Cleanup
Git periodically performs maintenance that can pack objects and eventually remove objects that are no longer reachable and have become eligible for cleanup.
git gcModern Git versions can perform maintenance automatically in many situations. Manual garbage collection is therefore not something most developers need to run routinely.
Why Git History Can Be Recovered After Mistakes
Operations such as reset, rebase, or deleting a branch can move references away from commits without immediately deleting the underlying objects. Local reflogs can record previous reference positions, which can make recovery possible before unreachable objects are pruned.
git reflogThis is one reason a commit that appears to have disappeared after a history-changing operation may still be recoverable locally.
Why Git History Is Efficient
Git combines several ideas to make repository history efficient: content-addressable objects, reusable blobs and trees, immutable commit relationships, references that move instead of copying history, and packfile compression.
The result is a system that can represent a large number of project snapshots while avoiding unnecessary duplication of identical content and maintaining strong relationships between versions.
A Simple Mental Model of Git Storage
You do not need to memorize every internal detail to work effectively with Git. A compact mental model is enough for most everyday tasks.
- Blobs represent file contents.
- Trees represent directory structures.
- Commits identify project snapshots and connect them to parents.
- Branches are movable references to commits.
- Tags give names to important objects.
- HEAD identifies the current checkout position.
- The index represents what will be included in the next commit.
- Object IDs identify Git objects based on their contents.
- Packfiles optimize the physical storage of many objects.
Common Misconceptions About Git Storage
| Misconception | What actually happens |
|---|---|
| Git stores only diffs | Git's object model represents snapshots through trees and blobs; diffs are calculated when needed |
| A branch contains a copy of the project | A branch is primarily a reference to a commit |
| A commit contains every file directly | A commit points to a tree that represents the project snapshot |
| A blob stores a filename | A blob stores file contents; trees associate names with objects |
| Rebase edits old commits | Rebase creates new commits with changed ancestry |
| Deleting a branch immediately deletes its commits | The commits may remain reachable through other references or temporarily through reflogs and unreachable-object storage |
| Push uploads the whole project folder | Git transfers required objects and reference updates |
How This Knowledge Helps in Everyday Git
Understanding Git's storage model is not only useful for learning Git internals. It explains practical behavior you encounter while working with branches and commits.
- Branch creation is fast because it creates a reference rather than copying history.
- Rebase changes commit IDs because it creates commits with different ancestry.
- Cherry-pick creates new commits because it applies changes to a different history.
- Merge commits have multiple parents because they connect multiple lines of development.
- Tags can identify stable points in the commit graph.
- Reflog can help recover previous reference positions.
- Git can reuse unchanged file content between snapshots.
- Fetch can download history without modifying your current branch.
Git Storage and Good Commit Practices
Git's object model also reinforces the value of focused commits. A commit is a named point in the history graph, so its message and contents become part of the long-term record of the project.
Small, logically focused commits are easier to inspect, review, revert, cherry-pick, and understand later. Clear commit messages make those historical objects more useful to both developers and automated release tooling.
Git Storage and Release Workflows
Release workflows often build on the same underlying model. A version tag identifies a commit, commit history describes the changes since an earlier release, and release-note or changelog tooling can derive information from that history.
git log v1.0.0..v1.1.0 --oneline
git diff v1.0.0 v1.1.0This makes Git history more than a record of who changed a file. It becomes a structured graph that can support development, review, debugging, release management, and automation.
Helpful Git Workflow Tools
Developer tools can simplify several tasks around Git history. A Git commit generator can help create clear commit messages, while a branch name generator can help maintain consistent names for feature, release, and hotfix branches.
Semantic version calculators can help determine version changes in projects following Semantic Versioning. Release notes and changelog generators can turn structured commit history into human-readable release documentation.
Frequently Asked Questions
How does Git store history?
Git stores history as a collection of objects and references. Commits point to tree objects representing project snapshots and to parent commits representing previous history. Trees point to blobs containing file contents, while branches and tags provide references to objects in the repository.
Does Git store full copies of files for every commit?
Git provides snapshot-style semantics, but its physical storage does not require a completely independent full copy of every file for every commit. Unchanged blobs can be reused, and packfiles can compress related objects using deltas.
What is a Git blob?
A blob is a Git object containing file contents. It does not normally contain the filename or directory path. Trees store the relationship between names, file modes, and blobs.
What is a Git tree?
A tree represents a directory at a particular point in history. It contains entries that point to blobs for files and other trees for subdirectories, allowing Git to represent the complete project structure.
What does a Git commit contain?
A commit identifies a tree representing the project snapshot and one or more parent commits. It also contains metadata such as author, committer, timestamps, and the commit message.
Why does rebase change commit hashes?
Rebase creates new commits with different parent relationships. Since a commit's identity depends on information including its parent and tree, changing the ancestry results in different commit IDs.
Is a Git branch a copy of the repository?
No. A branch is primarily a movable reference to a commit. Creating a branch normally requires only another reference to an existing commit rather than copying the repository's history.
Can deleted Git commits be recovered?
Sometimes. Moving or deleting a reference does not necessarily remove the underlying objects immediately. Local reflogs may retain previous reference positions, and unreachable objects can remain temporarily before Git's cleanup process removes them.
Conclusion
Git stores project history using a combination of content-addressable objects and movable references. Blobs contain file contents, trees describe directory structures, and commits connect project snapshots into a history graph. Branches and tags then provide convenient names for important points in that graph.
This model explains many of Git's most important behaviors. Creating a branch is cheap because it creates a reference. Merge commits have multiple parents because they join histories. Rebase and cherry-pick create new commits because they change where existing changes belong in the graph. Tags identify stable points without copying the underlying history.
You do not need to work directly with Git's object database every day, but understanding it gives you a much stronger mental model of the tool. Once commits, trees, blobs, and references make sense, Git's branches, merges, rebases, resets, tags, and recovery mechanisms become much easier to reason about.