git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Question about your git habits

From
Charles Bailey <charles@hashpling.org>
Date
Feb 23, 2008, 14:01 UTC
Message-ID
<20080223140153.GB5811@hashpling.org>
In-Reply-To
<998d0e4a0802230536w74e93ec3s40c77d52b183a419@mail.gmail.com>
On Sat, Feb 23, 2008 at 02:36:59PM +0100, J.C. Pizarro wrote:
Show 21 quoted lines
> On 2008/2/23, Charles Bailey <charles@hashpling.org> wrote:
> >
> > It shouldn't matter how aggressively the repositories are packed or what
> >  the binary differences are between the pack files are. git clone
> >  should (with the --reference option) generate a new pack for you with
> >  only the missing objects. If these objects are ~52 MiB then a lot has
> >  been committed to the repository, but you're not going to be able to
> >  get around a big download any other way.
> 
> You're wrong, nothing has to be commited ~52 MiB to the repository.
> 
> I'm not saying "commit", i'm saying
> 
> "Assume A & B binary git repos and delta_B-A another binary file, i
> request built
> B' = A + delta_B-A where is verified SHA1(B') = SHA1(B) for avoiding
> corrupting".
> 
> Assume B is the higher repacked version of "A + minor commits of the day"
> as if B was optimizing 24 hours more the minimum spanning tree. Wow!!!
> 

I'm not sure that I understand where you are going with this. Originally, you stated that if you clone a 775 MiB repository on day one, and then you clone it again on day two when it was 777 MiB, then you currently have to download 775 + 777 MiB of data, whereas you could download a 52 MiB binary diff. I have no idea where that value of 52 MiB comes from, and I've no idea how many objects were committed between day one and day two. If we're going to talk about details, then you need to provide more details about your scenario.

Having said that, here is my original point in some more detail. git repositories are not binary blobs, they are object databases. Better than this, they are databases of immutable objects. This means that to get the difference between one database and another, you only need to add the objects that are missing from the other database. If the two databases are actually a database and the same database at short time interval later, then almost all the objects are going to be common and the difference will be a small set of objects. Using git:// this set of objects can be efficiently transfered as a pack file. You may have a corner case scenario where the following isn't true, but in my experience an incremental pack file will be a more compact representation of this difference than a binary difference of two aggressively repacked git repositories as generated by a generic binary difference engine.

I'm sorry if I've misunderstood your last point. Perhaps you could expand in the exact issue that are having if I have, as I'm not sure that I've really answered your last message.

Previous: J.C. PizarroNext: J.C. Pizarro
Message 23 of 29 in “Question about your git habits”
  1. Chase VentersFeb 23, 2008
  2. Tommy ThornFeb 23, 2008
  3. Steven WalterFeb 23, 2008
  4. Jan EngelhardtFeb 23, 2008
  5. Al ViroFeb 23, 2008
  6. Junio C HamanoFeb 23, 2008
  7. Al ViroFeb 23, 2008
  8. Junio C HamanoFeb 23, 2008
  9. Samuel TardieuFeb 23, 2008
  10. Daniel BarkalowFeb 23, 2008
  11. Jeff GarzikFeb 23, 2008
  12. Mike HommeyFeb 23, 2008
  13. Rene HermanFeb 23, 2008
  14. Willy TarreauFeb 23, 2008
  15. Sam RavnborgFeb 23, 2008
  16. Jakub NarebskiFeb 23, 2008
  17. J.C. PizarroFeb 23, 2008
  18. J.C. PizarroFeb 23, 2008
  19. Charles BaileyFeb 23, 2008
  20. J.C. PizarroFeb 23, 2008
  21. Charles BaileyFeb 23, 2008
  22. J.C. PizarroFeb 23, 2008
  23. Charles BaileyFeb 23, 2008
  24. J.C. PizarroFeb 23, 2008
  25. Charles BaileyFeb 23, 2008
  26. J.C. PizarroFeb 23, 2008
  27. Charles BaileyFeb 23, 2008
  28. J.C. PizarroFeb 23, 2008
  29. Mike HommeyFeb 23, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.