git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: two questions about the format of loose object

From
LYLiu Yubao <yubao.liu@gmail.com>
Date
Dec 2, 2008, 03:05 UTC
Message-ID
<4934A5EC.2090708@gmail.com>
In-Reply-To
<20081201153211.GH23984@spearce.org>
Shawn O. Pearce wrote:
Show 25 quoted lines
> Liu Yubao <yubao.liu@gmail.com> wrote:
>> In current implementation the loose objects are compressed:
>>
>>      loose object = deflate(typename + <space> + size + '\0' + data)
> ...
>> * Question 1:
>>
>> Why not use the format below for loose object?
>>     loose object = typename + <space> + size + '\0' + deflate(data)
> 
> Historical accident.  We really should have used a format more
> like what you are asking here, because it makes inflation easier.
> The pack file format uses a header structure sort of like this,
> for exactly that reason.  IOW we did learn our mistakes and fix them.
> 
> If you look up the new style loose object code you'll see that it
> has a format like this (sort of), the header is actually the same
> format that is used in the pack files, making it smaller than what
> you propose but also easier to unpack as the code can be reused
> with the pack reading code.
> 
> Unfortunately the new style loose object was phased out; it never
> really took off and it made the code much more complex.  So it was
> pulled in commit 726f852b0ed7e03e88c419a9996c3815911c9db1:
> 

In fact the format I proposed in my patches is uncompressed loose object, not uncompressed loose object header, that's to say I proposed format 2 in my question 2, I am just curious why the loose object header is compressed in question 1.

I did a test to add all files of git-1.6.1-rc1 with git-add, the time spent decreased by half. Other commands like git diff, git diff --cached, git diff HEAD~ HEAD should be faster now although the change may be not noticable for small and medium project.

Show 18 quoted lines
>  Author: Nicolas Pitre <nico@cam.org>:
>  >  deprecate the new loose object header format
>  >
>  >  Now that we encourage and actively preserve objects in a packed form
>  >  more agressively than we did at the time the new loose object format and
>  >  core.legacyheaders were introduced, that extra loose object format
>  >  doesn't appear to be worth it anymore.
>  >
>  >  Because the packing of loose objects has to go through the delta match
>  >  loop anyway, and since most of them should end up being deltified in
>  >  most cases, there is really little advantage to have this parallel loose
>  >  object format as the CPU savings it might provide is rather lost in the
>  >  noise in the end.
>  >
>  >  This patch gets rid of core.legacyheaders, preserve the legacy format as
>  >  the only writable loose object format and deprecate the other one to
>  >  keep things simpler.
> 
Thank you for dig it out for me!
Best regards,
Liu Yubao
Previous: Shawn O. PearceNext: Nicolas Pitre
Message 26 of 27 in “two questions about the format of loose object”
  1. Liu YubaoDec 1, 2008
  2. Junio C HamanoDec 1, 2008
  3. Liu YubaoDec 1, 2008
  4. Jakub NarebskiDec 1, 2008
  5. Liu YubaoDec 2, 2008
  6. Shawn O. PearceDec 1, 2008
  7. Liu YubaoDec 2, 2008
  8. 0/5 support reading and writing uncompressed loose objectLiu Yubao, Dec 2, 2008
  9. 1/5 avoid parse_sha1_header() accessing memory out of boundLiu Yubao, Dec 2, 2008
  10. Shawn O. PearceDec 2, 2008
  11. Liu YubaoDec 3, 2008
  12. 2/5 don't die immediately when convert an invalid type nameLiu Yubao, Dec 2, 2008
  13. 3/5 optimize parse_sha1_header() a little by detecting object typeLiu Yubao, Dec 2, 2008
  14. Shawn O. PearceDec 2, 2008
  15. Liu YubaoDec 3, 2008
  16. 4/5 support reading uncompressed loose objectLiu Yubao, Dec 2, 2008
  17. Shawn O. PearceDec 2, 2008
  18. Liu YubaoDec 3, 2008
  19. 5/5 support writing uncompressed loose objectLiu Yubao, Dec 2, 2008
  20. Shawn O. PearceDec 2, 2008
  21. Liu YubaoDec 3, 2008
  22. Liu YubaoDec 2, 2008
  23. Nick AndrewDec 1, 2008
  24. Liu YubaoDec 2, 2008
  25. Shawn O. PearceDec 1, 2008
  26. Liu YubaoDec 2, 2008
  27. Nicolas PitreDec 4, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.