I read this article:
https://cursor.com/blog/git-at-any-scale
It points to the first version of git:
https://github.com/git/git/commit/e83c5163316#diff-505e8cbcfb86ace34be6631e963814dc15014c8d5df255d323c155d11cf8c2e2
The cursor article mentions performance problems with GIT pack files/protocol and how the came out with solutions to try and make git more high performant.
But instead of inventing all kinds of new/akward/work around solutions, how about tackling GIT as it's core.
So I fed the initial version/idea of git/source code into Chatgpt 5.6 web/free chat, probably luna model and here is what it came up with:
I post this here just for inspiration and a starting point, but running this multiple times through multiple AIs it can be re-fined/looked at, improved and ultimately developed into a new specification.
(I might also build an AI agent specially/specifically for design/building/writing/develops specifications... this work is still in progess)
I also have another idea I might like to work on: VirtualGit, it would present every commit of git to a windows explorer like interface/file system driver/whatever so that each commit looks like a checkout to the windows explorer, without actually having to check it out... so that's why I like to call to "virtual" git... It's simply computed... this might be possible with/for systems with lots of memory and compute/cpu... though for 1 gigabyte repositories and up I can imagine that it would quickly run out of memory, however, by sharing objects wisely, keeping it compressed in memory it might still function pretty well, so it would have to be "smart"... so perhaps a better name for it would be: "smart virtual git" :) At least for super large repos, it could at least show some checkouts/last few commits ;)
Anyway, here is ChatGPT re-design ideas for GIT to solve the packfile/protocol performance problem:
Yes. After looking at the original Git commit, the current protocol-v2 design, `upload-pack`/`fetch-pack`, pack heuristics, bitmaps, MIDX, partial clone, and the current pack format, I think the best redesign is **not “make Git negotiation a little smarter.”** It is to separate Git's excellent object model from the transport/storage decisions that accumulated around packfiles.
The key idea would be:
> **Keep Git's content-addressed DAG. Replace negotiated, request-specific pack streams with immutable, indexed, independently transferable object segments.**
The original 2005 design already had the right fundamental abstraction: objects are content-addressed, and objects reference other objects by hash. ([GitHub][1]) Modern Git's performance problems largely come from optimizations layered on top of that abstraction rather than from the DAG itself.
## 1. What I would keep
I would retain almost all of Git's semantic model:
```text blob tree commit tag
object-id = Hash(canonical-object) ```
Commits remain immutable DAG nodes.
Trees remain immutable snapshots.
Refs remain mutable names pointing into the DAG.
You could continue using SHA-1 repositories and SHA-256 repositories exactly as today, although new repositories should obviously use SHA-256.
So:
```text Git object semantics KEEP commit/tree/blob model KEEP content addressing KEEP distributed operation KEEP branch/ref semantics KEEP ```
What I would replace is:
```text fetch negotiation REPLACE pack stream REPLACE pack indexing REPLACE delta representation REDESIGN clone/fetch protocol REDESIGN server-side pack building MOSTLY ELIMINATE ```
---
# 2. The fundamental problem with current Git
Modern Git fetch conceptually looks like this:
```text CLIENT SERVER
capabilities <----------------------------------------
ls-refs ---------------------------------------->
refs <----------------------------------------
want A have X have Y have Z ---------------------------------------->
ACK Y <----------------------------------------
more have... ---------------------------------------->
ACK ... <----------------------------------------
done ---------------------------------------->
enumerate objects
construct pack
select/reuse deltas
compress packPACK stream <======================================== ```
Protocol v2 improved many aspects of the original protocol, but commands are still sequential: only one command can be requested at a time, and the server waits for the complete request before responding. ([Git][2])
The fetch protocol explicitly has `want`, `have`, acknowledgements, `ready`, and `done` machinery for determining the common history. ([GitHub][3])
Git even provides:
```text fetch.negotiationAlgorithm=consecutive fetch.negotiationAlgorithm=skipping fetch.negotiationAlgorithm=noop ```
The existence of `noop` is revealing: Git can eliminate negotiation, but the price is potentially transferring substantially more data. ([Git][4])
So Git is making a trade:
```text
extra RTT / computation
versus
extra bytes
```I don't think that should be the fundamental choice anymore.
---
# 3. Packfiles are the bigger architectural problem
The current packfile has this conceptual form:
```text PACK version object-count
object object object object ...
checksum ```
An object can be:
```text FULL_OBJECT
or
DELTA(base) ```
and a delta can refer to its base by:
```text REF_DELTA -> object ID
OFS_DELTA -> relative position inside this pack ```
That latter mechanism means physical placement becomes part of the dependency structure. ([Git][5])
Git gets extremely good compression from this.
But look at what we lose.
A pack is fundamentally a sequential artifact:
```text +-------------------------------------------------------+ | header | obj | obj | delta | delta | ... | checksum | +-------------------------------------------------------+ ```
Object 8 may require object 3.
Object 3 may require object 1.
And their physical offsets matter.
The `.idx` then exists separately to recover random-access properties from this sequential representation. `index-pack` must construct that index after receiving a pack. ([Git][6])
That design made enormous sense for:
```text 2005 CPUs 2005 disks 2005 network connections 2005 filesystems ```
It is much less attractive for:
```text HTTP/3 CDNs NVMe many-core CPUs parallel decompression range requests object caches edge caches cloud object stores interrupted transfers large monorepositories ```
---
# 4. My replacement: Git Segment Store
I'd call the format something like:
```text Git Object Segment Format GOSF/1 ```
or, if you want to retain the terminology:
```text PACK/4 ```
A repository becomes:
```text
Repository Manifest
|
+--- Segment A
+--- Segment B
+--- Segment C
+--- Segment D
+--- Segment E
```Every segment is:
* immutable * content-addressed * independently verifiable * independently downloadable * independently decompressible * internally indexed * resumable * cacheable forever
For example:
```text
segments/
03/03b8....seg
17/178a....seg
62/62df....seg
cc/cc12....seg
```The segment itself is hashed:
```text segment_id = SHA256(segment_bytes) ```
That ID is **not** a Git object ID.
It identifies the storage artifact.
This distinction is important:
```text Git OID = identity of logical object
Segment ID = identity of physical representation ```
---
# 5. PACK/4 format
Something like:
```text +--------------------------------------+ | PACK4 HEADER | +--------------------------------------+ | OBJECT DIRECTORY | +--------------------------------------+ | FRAME DIRECTORY | +--------------------------------------+ | DEPENDENCY TABLE | +--------------------------------------+ | COMPRESSED FRAME 0 | +--------------------------------------+ | COMPRESSED FRAME 1 | +--------------------------------------+ | COMPRESSED FRAME 2 | +--------------------------------------+ | ... | +--------------------------------------+ | MERKLE / SEGMENT CHECKSUM | +--------------------------------------+ ```
The object directory could conceptually contain:
```text object_oid object_type canonical_size
representation:
FULL
DELTA
CHUNKEDframe_number offset encoded_length
base_object_index // DELTA only ```
Notice what is gone:
```text absolute pack offset as semantic dependency relative pack offset as semantic dependency ```
---
# 6. Delta bases must be local to a segment
This is one of the strongest rules I would introduce.
A transferable segment must be:
```text SELF CONTAINED ```
Therefore:
```text delta object | +-- base object in SAME SEGMENT ```
Never:
```text
Segment B delta
|
+------> object buried in Segment A
```Yes, restricting delta choice can make compression slightly worse.
But it buys something enormously valuable:
```text Segment X can be downloaded, verified, decoded, cached, copied, mirrored, or deleted
without examining any other segment. ```
Git already accepts similar compression-versus-operability compromises with delta islands: Git may accept somewhat larger packs so that server fetches can reuse stored delta relationships instead of recomputing deltas when the receiver lacks their bases. ([Git][7])
PACK/4 would make that idea architectural rather than heuristic.
---
# 7. Make segments moderately sized
Not:
```text one repository = one 80 GB pack ```
and not:
```text one Git object = one network file ```
Instead something around:
```text target compressed size:
4–32 MiB ```
with perhaps:
```text 16 MiB default ```
Large enough for good compression.
Small enough for:
```text parallel downloads parallel decoding CDN distribution retry range requests incremental GC cache eviction ```
So a 2 GB transfer might look like:
```text Worker 1 -> Segment 037 Worker 2 -> Segment 038 Worker 3 -> Segment 039 Worker 4 -> Segment 040 Worker 5 -> Segment 041 Worker 6 -> Segment 042 Worker 7 -> Segment 043 Worker 8 -> Segment 044 ```
rather than:
```text 1 giant pack stream =========================================> ```
---
# 8. Frames inside segments
Even 16 MiB segments should have independently compressed frames.
For example:
```text
Segment
frame 0 256 KB
frame 1 512 KB
frame 2 384 KB
frame 3 640 KB
...
```I would probably use something in the Zstandard family rather than per-object zlib, though the exact codec should be negotiated/versioned.
The important property isn't specifically Zstd.
It is:
```text independent frames ```
so four cores can do:
```text CPU0 -> frame 0 CPU1 -> frame 1 CPU2 -> frame 2 CPU3 -> frame 3 ```
simultaneously.
---
# 9. Large blobs should become chunk-addressable representations
This is another major departure from present packfiles.
Suppose Git has:
```text huge.iso 4 GB huge-v2.iso 4 GB, 5 MB modified ```
Git's logical object model should still see:
```text blob A blob B ```
But PACK/4 doesn't need to represent either blob as one monolithic compressed stream.
Internally:
```text blob B = chunk A chunk B chunk C chunk D chunk E ... ```
using content-defined chunk boundaries.
Something like:
```text 64 KiB minimum 256 KiB target 1 MiB maximum ```
could be a starting point.
Then:
```text Version 1
A B C D E F G H
Version 2
A B C X Y F G H ```
requires transmitting:
```text X Y ```
instead of computing an enormous blob delta.
Important:
The Git object ID still hashes the reconstructed canonical blob.
Therefore existing Git semantics survive.
---
# 10. The server should stop building bespoke packs
Today `pack-objects` can exploit existing deltas, but there are circumstances where it must break or recompute delta relationships because the receiver doesn't have the required base. The Git documentation explicitly discusses this as a significant server CPU issue. ([Git][7])
PACK/4 changes the serving model from:
```text
request arrives
↓
inspect client's history
↓
walk DAG
↓
choose objects
↓
choose deltas
↓
generate pack
↓
compress
↓
stream
```to:
```text
request arrives
↓
identify required immutable segments
↓
return segment IDs
```The expensive work is amortized when segments are created.
Think:
```text repository write-time:
expensive packing once
repository read-time:
cheap immutable serving millions of times ```
That's enormously CDN-friendly.
---
# 11. Repository state should have a manifest
Introduce an immutable repository manifest.
Conceptually:
```text
RepositoryManifest {
repository_idepoch
refs_root
commit_graph_root
segment_set_root
segments[]
previous_manifest
signatures... } ```
and:
```text manifest_id = SHA256(canonical_manifest) ```
A ref snapshot therefore has a cheap immutable identifier:
```text
repo-state:
7c8512...
```---
# 12. Introduce sync receipts
This is how I'd eliminate most negotiation.
When a server sends data to a client it also sends:
```text SYNC_RECEIPT ```
For example:
```text repository = github.com/git/git manifest = abc123 coverage = reachable(refs snapshot X) filters = blob:none epoch = 483918 ```
The server cryptographically authenticates or otherwise validates this opaque token.
The client does **not** have to tell the server:
```text have a have b have c have d have e ... ```
next time.
It just says:
```text I previously synchronized through RECEIPT XYZ. Give me refs/heads/master now. ```
The server already knows exactly what the receipt means.
That turns:
```text "Which objects do you possess?" ```
into:
```text "I know what I previously gave you." ```
Huge difference.
---
# 13. Common fetch becomes one request
The normal fetch would now be:
```text CLIENT SERVER
FETCH {
repo-state-token = ABC
sync-receipt = XYZwant-ref = refs/heads/master
filter = full } --------------------------------------->
resolve ref
compare immutable epochs
determine new segments RESULT {
new-ref = ...
manifest = ...
segments = [
S123,
S124,
S131
]
receipt = NEWXYZ
} SEGMENT S123
SEGMENT S124
SEGMENT S131
<=======================================
```That is **one application-level request/response**.
No:
```text ls-refs RTT have RTT have RTT ACK RTT done RTT pack-generation delay ```
---
# 14. And don't wait for the plan before transmitting
Even better, the response should be multiplexable:
```text RESPONSE HEADER
REF refs/heads/main = A829...
SEGMENT S124 START <data>
MANIFEST ... SEGMENT S125 START <data>
SEGMENT S124 END
SEGMENT S126 START ... ```
In HTTP/3 terms, each segment can be a separate stream.
So head-of-line blocking disappears:
```text Stream 0 metadata Stream 4 segment A Stream 8 segment B Stream 12 segment C Stream 16 segment D ```
A slow segment doesn't stop everything behind it.
---
# 15. Fresh clone becomes one request too
Current Git often conceptually requires discovering refs before choosing what to fetch.
PACK/4 should permit:
```text
CLONE {
ref = refs/heads/main
filter = full
}
```without knowing its OID.
Server replies:
```text
REF:
refs/heads/main = 93824ab...REPOSITORY:
manifest = ...REQUIRED_SEGMENTS:
...SEGMENT... SEGMENT... SEGMENT... ```
So:
```text TCP/TLS/QUIC establishment + one Git request ```
and the data starts flowing.
With HTTP/3 session resumption, the transport itself can sometimes reduce handshake costs too, though that is separate from the Git protocol design.
---
# 16. Historical clones become even better
Static segment IDs mean Git hosting providers can put virtually all object payloads on CDNs:
```text
git.example.com
|
+-- metadata / refs
|
+-- returns segment list
|
V
CDN / object store
```The client downloads:
```text https://cdn/.../segment/abc https://cdn/.../segment/def https://cdn/.../segment/ghi ```
Current Git already has concepts pointing this way: protocol v2 includes `packfile-uris`, and bundle URI support allows large static data to be downloaded while a smaller incremental negotiation proceeds concurrently. ([Git][8])
I would make that **the primary architecture**, rather than an optional acceleration path.
---
# 17. What about clients with arbitrary local history?
That's the difficult case.
For instance:
```text Alice cloned from GitHub. Bob cloned Alice's laptop. Bob added commits. Bob now fetches GitHub. ```
GitHub cannot rely entirely on its old receipt.
I would support two mechanisms.
### First: commit frontier
Client sends selected local graph tips:
```text
local-frontier:
A
B
C
```The server's commit graph/bitmap tells it whether those are known.
One request still suffices.
For most cases:
```text one known ancestor found ```
and no iterative negotiation is needed.
### Second: object-set reconciliation
For pathological cases, use a set-reconciliation structure rather than thousands of `have` messages.
Candidates include techniques in the family of:
```text IBLT Golomb-coded sets XOR filters Merkle set reconciliation ```
But I'd be extremely careful here.
A probabilistic "client has this object" result must **never** silently create an incomplete repository.
So probabilistic structures may only optimize transfer selection.
Correctness still comes from:
```text object identity verification + DAG connectivity verification + missing-object recovery ```
---
# 18. Missing-object recovery becomes normal and cheap
Suppose optimization produces:
```text 99.999% correct object set ```
but misses three objects.
That should not mean:
```text fatal clone corruption start over ```
Client performs:
```text
GET_OBJECTS {
oid1
oid2
oid3
}
```Server returns the corresponding segments or object frames.
This makes aggressive low-latency negotiation possible because correctness no longer requires perfect prediction before transmission begins.
---
# 19. Stop treating a fetch as one atomic pack
This is probably my largest criticism of the current pack abstraction.
Current conceptual model:
```text Fetch result = PACK ```
New model:
```text
Fetch result =
repository metadata
+
set of independently verifiable segments
```Consequently interrupted fetch:
```text Downloaded:
S1 ✓ S2 ✓ S3 ✓ S4 61% S5 pending S6 pending ```
Resume:
```text keep S1 keep S2 keep S3
restart/range-resume S4 get S5 get S6 ```
No need to treat a partially obtained giant pack as a special temporary state.
---
# 20. Garbage collection becomes LSM-like
Git already has MIDX and geometric repacking ideas. A multi-pack-index lets Git address objects across multiple packs, and geometric repack attempts to maintain a progression of pack sizes instead of constantly rewriting everything. ([Git][9])
I'd take that much further.
Think:
```text
Level 0:
newest small segmentsLevel 1:
compacted segmentsLevel 2:
larger stable segmentsArchive:
very stable historical segments
```Almost like an LSM tree:
```text
new objects
↓
small immutable segment
↓
small immutable segment
↓
compaction
↓
larger immutable segment
```Old segment IDs remain valid for some retention period, so CDN caches and clients do not suddenly lose usefulness after server GC.
---
# 21. Segments should be organized by expected transfer locality
Current Git's delta-selection heuristics try to group likely-similar objects. ([GitHub][10])
PACK/4 could explicitly optimize segments for two dimensions:
```text compression similarity + fetch locality ```
For example, objects reachable from nearby commits could tend to live together.
Instead of maximizing:
```text smallest possible repository ```
the objective becomes:
```text Minimize:
storage_bytes + network_bytes + server_CPU + client_CPU + random_IO + RTT_cost + cache_miss_cost ```
That's a much better objective function for 2026.
---
# 22. Push should use the exact same representation
Push currently has a related problem:
```text client computes pack client uploads pack server indexes/verifies pack ```
I'd make push:
```text
PUSH_BEGIN {
old-ref = ABC
new-ref = XYZ
}OBJECT_SEGMENTS:
S1
S2
S3
```Server says:
```text HAVE S1 NEED S2 HAVE S3 ```
If segment IDs are globally content-addressed, existing segments never need retransmission.
Then:
```text UPDATE_REF ```
subject to atomic ref validation.
---
# 23. Better yet: advertise segment presence before uploading bytes
A client could send:
```text
PUSH {
ref-update... offered-segments:
A
B
C
D
}
```Server immediately responds:
```text
need:
C
D
```Client starts them in independent streams.
With HTTP/3 you could even optimistically begin sending likely-new segments before the metadata response comes back.
If the server already has one, it cancels that stream.
The cost is a few redundant bytes instead of another RTT.
That is usually the correct trade on high-latency networks.
---
# 24. Commit graph and object data should be separate planes
I'd explicitly separate:
```text
CONTROL PLANE
refs
commit graph
reachability summaries
manifests
receipts
capabilitiesDATA PLANE
immutable object segments
```Control messages might be kilobytes.
Data plane might be gigabytes.
That means:
```text Control: git.example.com
Data: cdn1.example.net cdn2.example.net cdn3.example.net ```
The data plane requires almost no Git intelligence.
Just:
```text GET segment SHA256:X ```
---
# 25. You could eventually make the central Git server almost stateless
A hosting request might become:
```text FETCH(repo, receipt, ref) ```
The smart front-end looks up:
```text ref -> commit receipt -> known manifest epoch ```
and produces:
```text difference = segment-set(new) - segment-set(old) ```
using bitmaps/set indexes.
Then every byte-heavy operation is handled by blob storage/CDNs.
The Git server itself doesn't run a giant bespoke compression job.
---
# 26. Why reachability bitmaps still matter
Git's modern bitmap machinery is actually something I'd preserve.
Git already uses bitmap indexes to accelerate reachability operations; the source describes a bitmap index associated with a pack or MIDX object order. ([GitHub][11])
But I'd change what the bitmap indexes.
Today roughly:
```text commit ↓ objects ```
PACK/4 could maintain:
```text commit ↓ object bitmap ↓ segment bitmap ```
Thus:
```text reachable-segments(commit A) - reachable-segments(receipt B) ```
becomes mostly bitmap arithmetic.
Very fast.
---
# 27. Segment manifests can form a Merkle tree
For enormous repositories, don't send a million segment identifiers.
Use:
```text
ROOT
/ \
A B
/ \ / \
C D E F
... ... ... ...
```Client knows:
```text old root = X ```
Server has:
```text new root = Y ```
Diff only differing branches.
This allows repository-state comparison proportional to the changed area, rather than repository size.
---
# 28. One important change to Git's concept of a "pack"
I'd make this explicit:
> **A pack should no longer be an archival snapshot. It should be a disposable physical representation of immutable Git objects.**
Git already conceptually separates object identity from pack representation, but implementation details make the representation unusually important.
PACK/4 should enforce the abstraction much more strongly.
Logical world:
```text OID -> canonical object ```
Physical world:
```text OID -> SegmentID -> Frame -> Representation ```
And physical representation can change without changing:
```text commit IDs tree IDs blob IDs ```
---
# 29. Compatibility with existing Git could actually be good
You wouldn't need a flag day.
Architecture:
```text
+----------------------+
old client ---> | Git compatibility |
| upload-pack gateway |
+----------+-----------+
|
v
Segment Store
^
|
+----------+-----------+
new client ---> | Git/3 / PACK4 |
+----------------------+
```For an old client:
```text
segment store
↓
generate legacy PACK2 stream
```For a new client:
```text native segments ```
So the repository can have one modern backend while supporting existing Git wire protocols.
---
# 30. I would also eliminate pkt-line framing for bulk payloads
Git's current pack transmission is multiplexed using pkt-line framing and sideband codes. ([GitHub][12])
That's historical layering we no longer need for the bulk-data path.
Don't do:
```text length channel data length channel data length channel data ... ```
for gigabytes of payload.
Modern transports already provide framing.
Use:
```text QUIC stream HTTP body raw binary segment ```
Progress belongs in another stream.
For example:
```text
Stream 0:
RPC metadataStream 4:
segment A bytesStream 8:
segment B bytesStream 12:
progress/events
```Transport multiplexing should do the multiplexing.
---
# 31. The new protocol
I'd call it something like:
```text Git Transfer Protocol 3 GTP/3 ```
Basic request:
```text
FETCH {
repository: <id> want {
ref: "refs/heads/main"
} base {
receipt: <opaque-token>
} filter {
blobs: all
} capabilities {
segment-format: PACK4
compression: zstd
parallel-streams: 16
}
}
```Response begins immediately:
```text
FETCH_RESULT {
ref {
refs/heads/main:
3c817...
} manifest:
a741... required-segments:
[
12a...,
348...,
9fe...
] new-receipt:
...
}
```followed concurrently by:
```text SEGMENT 12a... SEGMENT 348... SEGMENT 9fe... ```
---
# 32. New clone
```text CLIENT
CLONE {
branch = "main"
filter = full
max_parallel = 16
}
```One request.
Server:
```text metadata + segment streams ```
No explicit:
```text discover choose negotiate ACK negotiate ACK done build transfer ```
pipeline.
---
# 33. Normal fetch
After the initial clone:
```text
FETCH {
ref = main
receipt = R329
}
```Server:
```text R329 corresponds to manifest 713
main now corresponds to manifest 719
needed:
S91
S92
S95
```Transmit three segments.
Done.
---
# 34. No-op fetch becomes extremely cheap
Current conceptual fetch still involves protocol work.
GTP/3:
```text
FETCH {
ref = main
receipt = ABC
}
```response:
```text
UNCHANGED {
receipt = ABC
}
```Potentially a few hundred bytes.
---
# 35. Force pushes become straightforward too
A receipt shouldn't mean:
```text history must be ancestor ```
It means:
```text server knows the exact synchronization state associated with this client token. ```
If:
```text A-B-C-D
becomes
A-B-X-Y ```
server computes required segments for `Y`.
Objects for C/D don't somehow become invalid locally.
Git's immutable-object semantics continue working naturally.
---
# 36. Partial clones improve dramatically
Current partial clone tracks promisor objects and may dynamically fetch missing objects; Git documents the fallback mechanism for retrieving missing promisor objects. ([Git][13])
PACK/4 can make lazy object acquisition first-class.
For example:
```text clone --metadata-only ```
gets:
```text commits trees small blobs perhaps ```
Opening:
```text src/video/test-data.bin ```
causes:
```text GET_OBJECT <OID> ```
which maps immediately:
```text OID ↓ segment ↓ CDN ```
No `upload-pack` traversal is necessary.
---
# 37. Checkout could fetch data in priority order
This creates a very useful capability.
Suppose checkout requires 20 GB.
Instead of waiting for everything:
```text metadata first source files second small blobs third large assets fourth old history later ```
The client can begin constructing the worktree almost immediately.
For instance:
```text Priority 0: commits + trees
Priority 1: README build files source files
Priority 2: normal assets
Priority 3: huge binaries
Priority 4: historical objects ```
Current pack streaming makes this substantially harder because the pack's compression layout heavily influences transfer order.
---
# 38. Clone should become progressive
Eventually:
```text git clone URL repo ```
could behave like:
```text 0.1 sec refs known
0.2 sec trees available
0.4 sec worktree construction begins
0.7 sec most source code available
... historical/archive segments continue in parallel ```
instead of thinking of cloning as:
```text receive opaque giant pack index pack resolve deltas checkout ```
---
# 39. There is a cost
This architecture will usually sacrifice some maximum compression efficiency.
Current Git packing is extraordinarily good at exploiting arbitrary delta relationships.
PACK/4 says:
```text do not optimize only for minimum bytes ```
because another 3% compression improvement isn't necessarily worth:
```text server CPU RTT serialization dependency chains loss of CDN reuse loss of parallelism poor resumability repacking costs ```
The new metric should be:
```text time-to-usable-repository ```
not:
```text smallest .pack possible ```
---
# 40. The distinction is something like this
Classic Git optimized for:
```text
storage
▲
|
compression
|
delta chains
```I'd optimize Git/3 for:
```text
latency
▲
|
parallel transfer
|
CDN reuse -----+----- CPU
|
resumability
|
storage
```It's a different optimum.
---
# 41. My PACK/4 rules
If I were freezing a specification, my strongest invariants would be:
1. **Git object identity remains independent of physical representation.** 2. **Every segment is immutable and byte-content-addressed.** 3. **Every segment is independently verifiable.** 4. **Every transport segment is independently decodable.** 5. **No delta base outside its segment.** 6. **Segment indexes occur before or independently from bulk payloads.** 7. **Compressed payload consists of independently decodable frames.** 8. **Large Git blobs may have chunked physical representations.** 9. **Partial segment download cannot corrupt repository state.** 10. **Repository state is described by immutable manifests.** 11. **Clients receive reusable synchronization receipts.** 12. **Normal fetch requires no iterative HAVE/ACK negotiation.** 13. **Object/segment transfer is inherently parallel.** 14. **The server need not construct a bespoke pack for ordinary fetches.** 15. **Legacy PACK2 can be generated as a compatibility view.**
Those fifteen rules are the heart of my redesign.
---
# 42. Architecture diagram
I would turn Git:
```text
CURRENT GIT refs
|
commit graph
|
objects
|
pack-objects
|
request-specific PACK
|
network
|
index-pack
|
local packs
```into:
```text
GIT/3 refs
|
commit graph
|
objects
|
Segment Builder
|
+-------------+-------------+
| | |
Segment A Segment B Segment C
| | |
+-------------+-------------+
|
Merkle Manifest
|
+-----------+-----------+
| |
Metadata/API CDN
| |
ref/receipt immutable segments
| |
+-----------+-----------+
|
client
```---
# 43. The interesting irony
The initial 2005 Git repository you linked starts from an exceptionally clean idea:
> a content-addressed collection of objects which can refer to other objects. ([GitHub][1])
That abstraction aged extremely well.
I'd argue that **packfiles aged less well than Git objects did**.
Modern Git has spent years successfully mitigating packfile consequences with:
```text bitmaps MIDX reverse indexes delta islands partial clones promisor packs bundle URIs packfile URIs geometric repacking skipping negotiation ```
Those are all useful engineering solutions. ([Git][7])
But taken together, they also suggest the cleaner next step:
> **Stop making the pack the fundamental transfer unit. Make the immutable object segment the transfer unit.**
That would give Git something much closer to **zero sequential negotiation RTT for the normal case, static CDN-compatible payloads, massively parallel decoding and transfer, trivial resume, and far less server-side pack generation**, while preserving the part of Git that was brilliant from day one: its immutable content-addressed object graph.
[1]: https://github.com/git/git/commit/e83c5163316 "Initial revision of \"git\", the information manager from hell · git/git@e83c516 · GitHub" [2]: https://git-scm.com/docs/gitprotocol-v2?utm_source=chatgpt.com "Git - gitprotocol-v2 Documentation" [3]: https://github.com/git/git/blob/master/Documentation/gitprotocol-v2.adoc?utm_source=chatgpt.com "git/Documentation/gitprotocol-v2.adoc at master · git/git · GitHub" [4]: https://git-scm.com/docs/git-fetch?utm_source=chatgpt.com "Git - git-fetch Documentation" [5]: https://git-scm.com/docs/pack-format?utm_source=chatgpt.com "Git - pack-format Documentation" [6]: https://git-scm.com/docs/git-index-pack?utm_source=chatgpt.com "Git - git-index-pack Documentation" [7]: https://git-scm.com/docs/git-pack-objects?utm_source=chatgpt.com "Git - git-pack-objects Documentation" [8]: https://git-scm.com/docs/protocol-v2?utm_source=chatgpt.com "Git - protocol-v2 Documentation" [9]: https://git-scm.com/docs/git-multi-pack-index?utm_source=chatgpt.com "Git - git-multi-pack-index Documentation" [10]: https://github.com/git/git/blob/master/Documentation/technical/pack-heuristics.adoc?utm_source=chatgpt.com "git/Documentation/technical/pack-heuristics.adoc at master · git/git · GitHub" [11]: https://github.com/git/git/blob/master/pack-bitmap.c?utm_source=chatgpt.com "git/pack-bitmap.c at master · git/git · GitHub" [12]: https://github.com/git/git/blob/master/Documentation/gitprotocol-pack.adoc?utm_source=chatgpt.com "git/Documentation/gitprotocol-pack.adoc at master · git/git · GitHub" [13]: https://git-scm.com/docs/partial-clone?utm_source=chatgpt.com "Git - partial-clone Documentation"
Bye for now, Skybuck Flying.