git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Recent unresolved issues: shallow clone

From
Junio C Hamano <junkio@cox.net>
Date
Apr 15, 2006, 00:25 UTC
Message-ID
<7vr73zn0rb.fsf@assigned-by-dhcp.cox.net>
In-Reply-To
<87irpb7oma.wl%cworth@cworth.org>
Carl Worth <cworth@cworth.org> writes:
Show 38 quoted lines
> On Fri, 14 Apr 2006 02:31:36 -0700, Junio C Hamano wrote:
>>   I am beginning to think using "graft" to cauterize history
>>   for this, while it technically would work, would not be so
>>   helpful to users, so the design needs to be worked out again.
>
> As context, here is some of what you mentioned in IRC:
>
>>>	Suppose you have this:
>>>
>>>	A---B---C
>>>	 \       \ 
>>>	  D---E---F---G
>>>	 
>>>	and you made a shallow clone of C (because that is where the
>>>	upstream master was when you made that clone).  Then the
>>>	upstream updated the master branch tip to G.
>>>
>>>	The next update from upstream to your shallow clone would break.
>>>	The upstream says: I have G at master.
>>>	You say: I want G then.  By the way, I have C.
>>>
>>>	What it means to tell the other end "I have X" is to promise
>>>	that you have X and _everything_ behind it.  So the upstream
>>>	would send objects necessary to complete D, E, F and G for
>>>	"somebody who already have A and B".  As a consequence, you
>>>	would not see A nor B.
>>>
>>>	Even if the only thing you are interested in is to be in sync
>>>	with the tip of the upstream, you can end up with an
>>>	incomplete tree for G, if some of the blobs or trees contained
>>>	in G already exist in A or B.  They are not sent -- because
>>>	you told the upstream that you have everything necessary to
>>>	get to C.
>
> So that's an argument against using a cauterizing graft for the
> shallow clone of C. It definitely confuses the existing protocol to
> say "I have C" if I have only a cauterized C, (its tree only, but none
> of the commits that should be backing C).

That's what I meant by "graft technically works but is inconvenient".

Maybe after the update to G happens (which means you now have C, F, G but not A B D E commits), the client side could enumerate commits on "rev-list ^C G" and cauterize the ones with missing parents (in this case, F does not have one of its parents). While doing this would help keeping the resulting commit ancestry sane, it does not solve the problem of missing blobs and trees. See below.

Show 7 quoted lines
> So, in the scenario above, the original shallow clone of C would be:
>
> 	Want C->tree, have nothing.
>
> and the later shallow update to G would be:
>
> 	Want G->tree, have C->tree

When you ask for G, you do not know what G^{tree} is, so that is fantasy without a protocol extention. To solve the missing blobs/trees problem we would probably need a protocol extention that says it wants to receive enough data to complete trees and blobs associated with the commits being sent _without_ assuming the recipient has any trees or blobs other than what are contained in "have" commits. Then after such a successful transfer, missing parents of commits listed in "rev-list ^C G" are the ones from the side branch, so the client can cauterize them (F in the above example) appropriately without bothering the server.

However, I think this "do not assume I have any trees behind the commits I explicitly say I have" must be an option, because it makes the resulting transfer unnecessarily more expensive for normal uses. A fetch of the Linux kernel once a day would update about a couple of hundered commits, each of which touches only 3 paths on average (so that would be 600 files out of 18,000 file tree. When side-branch merges are involved, usually many things in G (and F) are unchanged since either A or C, but the extention we are discussing forbids reusing what are found in A (it still allows reusing what are found in C).

> A final step of a shallow clone would then require creating a new
> parent-less commit object so that there's something to point refs/head
> at, (or maybe rather than being parentless, they could be chained
> together with each update?).

Rewriting commit objects transferred to the cloner is something you would _not_ want to do (e.g. rewriting F commits to say it has only one parent C). The history based on that would diverge from parents and would become unmergeable. It is cleaner to just make a new graft entry to say "As far as this repository is concerned, F has one parent C". Shallowness of the repository and its slightly different view of history is a local matter.

Previous: Johannes SchindelinNext: Junio C Hamano
Message 7 of 81 in “Recent unresolved issues”
  1. Junio C HamanoApr 14, 2006
  2. Petr BaudisApr 14, 2006
  3. seanApr 14, 2006
  4. Petr BaudisApr 14, 2006
  5. Carl WorthApr 14, 2006
  6. Johannes SchindelinApr 15, 2006
  7. Junio C HamanoApr 15, 2006
  8. Junio C HamanoApr 15, 2006
  9. Linus TorvaldsApr 14, 2006
  10. Linus TorvaldsApr 15, 2006
  11. Linus TorvaldsApr 15, 2006
  12. Junio C HamanoApr 15, 2006
  13. Linus TorvaldsApr 15, 2006
  14. Linus TorvaldsApr 15, 2006
  15. Linus TorvaldsApr 15, 2006
  16. Junio C HamanoApr 15, 2006
  17. Junio C HamanoApr 15, 2006
  18. Junio C HamanoApr 15, 2006
  19. Johannes SchindelinApr 15, 2006
  20. Linus TorvaldsApr 15, 2006
  21. Linus TorvaldsApr 15, 2006
  22. Junio C HamanoApr 16, 2006
  23. Junio C HamanoApr 15, 2006
  24. Linus TorvaldsApr 15, 2006
  25. Junio C HamanoApr 15, 2006
  26. Unresolved issues #2Junio C Hamano, May 4, 2006
  27. Jakub NarebskiMay 4, 2006
  28. Junio C HamanoMay 4, 2006
  29. Jakub NarebskiMay 4, 2006
  30. Petr BaudisMay 4, 2006
  31. Pavel RoskinMay 4, 2006
  32. Carl WorthMay 4, 2006
  33. Junio C HamanoMay 5, 2006
  34. Martin LanghoffMay 5, 2006
  35. Carl WorthMay 5, 2006
  36. Jakub NarebskiMay 5, 2006
  37. Linus TorvaldsMay 5, 2006
  38. Jakub NarebskiMay 5, 2006
  39. Linus TorvaldsMay 5, 2006
  40. Martin LanghoffMay 6, 2006
  41. Junio C HamanoMay 6, 2006
  42. Martin LanghoffMay 7, 2006
  43. Jeff KingMay 7, 2006
  44. Linus TorvaldsMay 7, 2006
  45. Theodore TsoMay 8, 2006
  46. Linus TorvaldsMay 8, 2006
  47. Theodore TsoMay 8, 2006
  48. Linus TorvaldsMay 8, 2006
  49. Theodore TsoMay 8, 2006
  50. Linus TorvaldsMay 8, 2006
  51. Jeff KingMay 8, 2006
  52. Linus TorvaldsMay 8, 2006
  53. Sergey VlasovMay 7, 2006
  54. Martin LanghoffMay 7, 2006
  55. Junio C HamanoMay 7, 2006
  56. Martin LanghoffMay 7, 2006
  57. Carl WorthMay 5, 2006
  58. Jakub NarebskiMay 7, 2006
  59. Junio C HamanoMay 8, 2006
  60. Jakub NarebskiMay 8, 2006
  61. Jakub NarebskiMay 8, 2006
  62. Daniel BarkalowMay 4, 2006
  63. Linus TorvaldsMay 4, 2006
  64. Junio C HamanoMay 6, 2006
  65. Linus TorvaldsMay 6, 2006
  66. seanMay 6, 2006
  67. Linus TorvaldsMay 6, 2006
  68. seanMay 6, 2006
  69. Linus TorvaldsMay 6, 2006
  70. Junio C HamanoMay 6, 2006
  71. Johannes SchindelinMay 6, 2006
  72. Linus TorvaldsMay 6, 2006
  73. Junio C HamanoMay 7, 2006
  74. Junio C HamanoMay 7, 2006
  75. Johannes SchindelinMay 7, 2006
  76. Jakub NarebskiMay 7, 2006
  77. Junio C HamanoMay 8, 2006
  78. Jakub NarebskiMay 7, 2006
  79. David WoodhouseMay 9, 2006
  80. Bertrand JacquinMay 9, 2006
  81. Nicolas PitreMay 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.