git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Figured out how to get Mozilla into git

From
RDRogan Dawes <discard@dawes.za.net>
Date
Jun 10, 2006, 14:47 UTC
Message-ID
<448ADB8A.3070506@dawes.za.net>
In-Reply-To
<7vzmglgyz0.fsf@assigned-by-dhcp.cox.net>
Junio C Hamano wrote:
Show 8 quoted lines
> Rogan Dawes <lists@dawes.za.net> writes:
> 
>> Here's an idea. How about separating trees and commits from the actual
>> blobs (e.g. in separate packs)?
> 
> If I remember my numbers correctly, trees for any project with a
> size that matters contribute nonnegligible amount of the total
> pack weight.  Perhaps 10-25%.

Out of curiosity, do you think that it may be possible for tree objects to compress more/better if they are packed together? Or does the existing pack compression logic already do the diff against similar tree objects?

Show 11 quoted lines
>> In this way, the user has a history that will show all of the commit
>> messages, and would be able to see _which_ files have changed over
>> time e.g. gitk would still work - except for the actual file level
>> diff, "git log" should also still work, etc
> 
> I suspect it would make a very unpleasant system to use.
> Sometimes "git diff -p" would show diffs, and other times it
> mysteriously complain saying that it lacks necessary blobs to do
> its job.  You cannot even run fsck and tell from its output
> which missing objects are OK (because you chose to create such a
> sparse repository) and which are real corruption.

The fsck problem could be worked around by maintaining a list of objects that are explicitly not expected to be present. As the list gets shorter (perhaps as diffs are performed, other parts of the blob history are retrieved, etc), the list will get shorter until we have a complete clone of the original tree.

Of course diffs against a version further back in the history would fail. But if you start with a checkout of a complete tree, any changes made since that point would at least have one version to compare against.

In effect, what we would have is a caching repository (or as Jakub said, a lazy clone). An initial checkout would effectively be pre-seeding the cache. One does not necessarily even need to get the complete set of commit and tree objects, either. The bare minimum would probably be to get the HEAD commit, and the tree objects that correspond to that commit.

At that point, one could populate the "uncached objects" list with the parent commits. One would not be in a position to get any history at all, of course.

As the user performs various operations, e.g. git log, git could either go and fetch the necessary objects (updating the uncached list as it goes), or fail with a message such as "Cannot perform the requested operation - required objects are not available". (We may require another utility that would list the objects required for an operation, and compare it against the list of "uncached objects", printing out a list of which are not yet available locally. I realise that this may be expensive. Maybe a repo configuration option "cached" to enable or disable this.)

As Jakub suggested, it would be necessary to configure the location of the source for any missing objects, but that is probably in the repo config anyway.

Show 6 quoted lines
> A shallow clone with explicit cauterization in grafts file at
> least would not have that problem. Although the user will still
> not see the exact same result as what would happen in a full
> repository, at least we can say "your git log ends at that
> commit because your copy of the history does not go back beyond
> that" and the user would understand.

Or, we could say, perform the operation while you are online, and can access the necessary objects. If the user has explicitly chosen to make a lazy clone, then they should expect that at some point, whatever they do may require them to be online to access items that they have not yet cloned.

Rogan
Previous: Junio C HamanoNext: Jakub Narebski
Message 42 of 67 in “Figured out how to get Mozilla into git”
  1. Jon SmirlJun 9, 2006
  2. Nicolas PitreJun 9, 2006
  3. Martin LanghoffJun 9, 2006
  4. Jon SmirlJun 9, 2006
  5. Jakub NarebskiJun 9, 2006
  6. Linus TorvaldsJun 9, 2006
  7. Nicolas PitreJun 9, 2006
  8. Linus TorvaldsJun 9, 2006
  9. Nicolas PitreJun 9, 2006
  10. Linus TorvaldsJun 9, 2006
  11. Jakub NarebskiJun 9, 2006
  12. Jon SmirlJun 9, 2006
  13. Linus TorvaldsJun 9, 2006
  14. Jon SmirlJun 9, 2006
  15. Linus TorvaldsJun 9, 2006
  16. Jon SmirlJun 9, 2006
  17. Linus TorvaldsJun 9, 2006
  18. Linus TorvaldsJun 9, 2006
  19. Greg KHJun 9, 2006
  20. Martin LanghoffJun 9, 2006
  21. Linus TorvaldsJun 9, 2006
  22. Jon SmirlJun 10, 2006
  23. Linus TorvaldsJun 10, 2006
  24. Jon SmirlJun 10, 2006
  25. Jon SmirlJun 10, 2006
  26. Jakub NarebskiJun 9, 2006
  27. Nicolas PitreJun 9, 2006
  28. Jon SmirlJun 9, 2006
  29. Martin LanghoffJun 10, 2006
  30. Martin LanghoffJun 10, 2006
  31. Linus TorvaldsJun 10, 2006
  32. Linus TorvaldsJun 10, 2006
  33. Jon SmirlJun 10, 2006
  34. Linus TorvaldsJun 10, 2006
  35. Jon SmirlJun 10, 2006
  36. Carl WorthJun 10, 2006
  37. Linus TorvaldsJun 10, 2006
  38. Jakub NarebskiJun 10, 2006
  39. Junio C HamanoJun 10, 2006
  40. Rogan DawesJun 10, 2006
  41. Junio C HamanoJun 10, 2006
  42. Rogan DawesJun 10, 2006
  43. Jakub NarebskiJun 10, 2006
  44. Nicolas PitreJun 10, 2006
  45. Linus TorvaldsJun 10, 2006
  46. Jon SmirlJun 10, 2006
  47. Rogan DawesJun 10, 2006
  48. Linus TorvaldsJun 10, 2006
  49. Jon SmirlJun 10, 2006
  50. Martin LanghoffJun 10, 2006
  51. Junio C HamanoJun 10, 2006
  52. Linus TorvaldsJun 10, 2006
  53. Linus TorvaldsJun 10, 2006
  54. Jon SmirlJun 10, 2006
  55. Junio C HamanoJun 10, 2006
  56. Jon SmirlJun 10, 2006
  57. Timo HirvonenJun 10, 2006
  58. Petr BaudisJun 10, 2006
  59. Lars JohannsenJun 10, 2006
  60. Nicolas PitreJun 11, 2006
  61. Linus TorvaldsJun 18, 2006
  62. Martin LanghoffJun 18, 2006
  63. Linus TorvaldsJun 18, 2006
  64. Broken PPC sha1.. (Re: Figured out how to get Mozilla into git)Linus Torvalds, Jun 18, 2006
  65. Fix PPC SHA1 routine for large input buffersPaul Mackerras, Jun 18, 2006
  66. Linus TorvaldsJun 19, 2006
  67. Pavel RoskinJun 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.