git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Figured out how to get Mozilla into git

From
RDRogan Dawes <lists@dawes.za.net>
Date
Jun 10, 2006, 08:36 UTC
Message-ID
<448A847C.20105@dawes.za.net>
In-Reply-To
<Pine.LNX.4.64.0606092001590.5498@g5.osdl.org>
Linus Torvalds wrote:
Show 37 quoted lines
> 
> On Fri, 9 Jun 2006, Carl Worth wrote:
> 
>> On Fri, 9 Jun 2006 22:21:17 -0400, "Jon Smirl" wrote:
>>> Could you clone the repo and delete changesets earlier than 2004? Then
>>> I would clone the small repo and work with it. Later I decide I want
>>> full history, can I pull from a full repository at that point and get
>>> updated? That would need a flag to trigger it since I don't want full
>>> history to come over if I am just getting updates from someone else's
>>> tree that has a full history.
>> This is clearly a desirable feature, and has been requested by several
>> people (including myself) looking to switch some large-ish histories
>> from an existing system to git.
> 
> The thing is, to some degree it's really fundamentally hard.
> 
> It's easy for a linear history. What you do for a linear history is to 
> just get the top commit, and the tree associated with it, and then you 
> cauterize the parent by just grafting it to go away. Boom. You're done.
> 
> The problems are that if the preceding history _wasn't_ linear (or, in 
> fact, _subsequent_ development refers to it by having branched off at an 
> earlier point), and you try to pull your updates, the other end (that 
> knows about all the history) will assume you have all the history that you 
> don't have, and will send you a pack assuming that.
> 
> Which won't even necessarily have all the tree/blob objects (it assumed 
> you already had them), but more annoyingly, the history won't be 
> cauterized, and you'll have dangling commits. Which you can cauterize by 
> hand, of course, but you literally _will_ have to get the objects and 
> cauterize the thing by hand.
> 
> You're right that it's not "fundamentally impossible" to do: the git 
> format certainly _allows_ it. But the git protocol handshake really does 
> end up optimizing away all the unnecessary work by knowing that the other 
> side will have all the shared history, so lacking the shared history will 
> mean that you're a bit screwed.

Here's an idea. How about separating trees and commits from the actual blobs (e.g. in separate packs)? My reasoning is that the commits and trees should only be a small portion of the overall repository size, and should not be that expensive to transfer. (Of course, this is only a guess, and needs some numbers to back it up.)

So, a shallow clone would receive all of the tree objects, and all of the commit objects, and could then request a pack containing the blobs represented by the current HEAD.

In this way, the user has a history that will show all of the commit messages, and would be able to see _which_ files have changed over time e.g. gitk would still work - except for the actual file level diff, "git log" should also still work, etc

This would also enable other optimisations.

For example, documentation people would only need to get the objects under the doc/ tree, and would not need to actually check out the source. Git could detect any actual changes by checking whether it has the previous blob in its local repository, and whether the file exists locally. Creating a patch would obviously require that the person checks out the previous version, but one could theoretically commit a new blob to a repo without having the previous one (not saying that this would be a good idea, of course)

This would probably require Eric Biederman's "direct access to blob" patches, I guess, in order to be feasible.

Regards,
Rogan
Previous: Junio C HamanoNext: Junio C Hamano
Message 40 of 67 in “Figured out how to get Mozilla into git”
  1. Jon SmirlJun 9, 2006
  2. Nicolas PitreJun 9, 2006
  3. Martin LanghoffJun 9, 2006
  4. Jon SmirlJun 9, 2006
  5. Jakub NarebskiJun 9, 2006
  6. Linus TorvaldsJun 9, 2006
  7. Nicolas PitreJun 9, 2006
  8. Linus TorvaldsJun 9, 2006
  9. Nicolas PitreJun 9, 2006
  10. Linus TorvaldsJun 9, 2006
  11. Jakub NarebskiJun 9, 2006
  12. Jon SmirlJun 9, 2006
  13. Linus TorvaldsJun 9, 2006
  14. Jon SmirlJun 9, 2006
  15. Linus TorvaldsJun 9, 2006
  16. Jon SmirlJun 9, 2006
  17. Linus TorvaldsJun 9, 2006
  18. Linus TorvaldsJun 9, 2006
  19. Greg KHJun 9, 2006
  20. Martin LanghoffJun 9, 2006
  21. Linus TorvaldsJun 9, 2006
  22. Jon SmirlJun 10, 2006
  23. Linus TorvaldsJun 10, 2006
  24. Jon SmirlJun 10, 2006
  25. Jon SmirlJun 10, 2006
  26. Jakub NarebskiJun 9, 2006
  27. Nicolas PitreJun 9, 2006
  28. Jon SmirlJun 9, 2006
  29. Martin LanghoffJun 10, 2006
  30. Martin LanghoffJun 10, 2006
  31. Linus TorvaldsJun 10, 2006
  32. Linus TorvaldsJun 10, 2006
  33. Jon SmirlJun 10, 2006
  34. Linus TorvaldsJun 10, 2006
  35. Jon SmirlJun 10, 2006
  36. Carl WorthJun 10, 2006
  37. Linus TorvaldsJun 10, 2006
  38. Jakub NarebskiJun 10, 2006
  39. Junio C HamanoJun 10, 2006
  40. Rogan DawesJun 10, 2006
  41. Junio C HamanoJun 10, 2006
  42. Rogan DawesJun 10, 2006
  43. Jakub NarebskiJun 10, 2006
  44. Nicolas PitreJun 10, 2006
  45. Linus TorvaldsJun 10, 2006
  46. Jon SmirlJun 10, 2006
  47. Rogan DawesJun 10, 2006
  48. Linus TorvaldsJun 10, 2006
  49. Jon SmirlJun 10, 2006
  50. Martin LanghoffJun 10, 2006
  51. Junio C HamanoJun 10, 2006
  52. Linus TorvaldsJun 10, 2006
  53. Linus TorvaldsJun 10, 2006
  54. Jon SmirlJun 10, 2006
  55. Junio C HamanoJun 10, 2006
  56. Jon SmirlJun 10, 2006
  57. Timo HirvonenJun 10, 2006
  58. Petr BaudisJun 10, 2006
  59. Lars JohannsenJun 10, 2006
  60. Nicolas PitreJun 11, 2006
  61. Linus TorvaldsJun 18, 2006
  62. Martin LanghoffJun 18, 2006
  63. Linus TorvaldsJun 18, 2006
  64. Broken PPC sha1.. (Re: Figured out how to get Mozilla into git)Linus Torvalds, Jun 18, 2006
  65. Fix PPC SHA1 routine for large input buffersPaul Mackerras, Jun 18, 2006
  66. Linus TorvaldsJun 19, 2006
  67. Pavel RoskinJun 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.