{"thread":{"id":"39611","subject":"RFC: reverse history tree, for faster & background clones","startedAt":"2015-06-12T11:26:42Z","lastAt":"2015-06-14T14:14:53Z","messageCount":5,"participants":["Andres G. Aragoneses","Dennis Kaarsemaker"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"263656","messageId":"mlefli$h6v$1@ger.gmane.org","threadId":"39611","inReplyTo":null,"subject":"RFC: reverse history tree, for faster & background clones","fromName":"Andres G. Aragoneses","fromEmail":"knocte@gmail.com","sentAt":"2015-06-12T11:26:42Z","receivedAt":"2015-06-12T11:26:42Z","isPatch":false,"sender":{"key":"knocte@gmail.com","avatar":"https://gravatar.com/avatar/020c5605dcd7456d87a5f29a2af53e7bdf135fe9ccda514b9f34ab3a7e46db27?d=mp&s=160"},"body":"Hello git devs,\n\nI'm toying with an idea of an improvement I would like to work on, but \nnot sure if it would be desirable enough to be considered good to merge \nin the end, so I'm requesting your opinions before I work on it.\n\nAFAIU git stores the contents of a repo as a sequence of patches in the \n.git metadata folder. So then let's look at an example to illustrate my \npoint more easily.\n\nRepo foo contains the following 2 commits:\n\n1 file, first commit, with the content:\n+First Line\n+Second Line\n+Third Line\n\n2nd and last commit:\n  First Line\n  Second Line\n-Third Line\n+Last Line\n\nSimple enough, right?\n\nBut, what if we decided to store it backwards in the metadata?\n\nSo first commit would be:\n1 file, first commit, with the content:\n+First Line\n+Second Line\n+Last Line\n\n2nd commit:\n  First Line\n  Second Line\n-Last Line\n+Third Line\n\n\nThis would bring some advantages, as far as I understand:\n\n1. `git clone --depth 1` would be way faster, and without the need of \non-demand compressing of packfiles in the server side, correct me if I'm \nwrong?\n2. `git clone` would be able to allow a \"fast operation, complete in the \nbackground\" mode that would allow you to download the first snapshot of \nthe repo very quickly, so that the user would be able to start working \non his working directory very quickly, while a \"background job\" keeps \nretreiving the history data in the background.\n3. Any more advantages you see?\n\n\nI'm aware that this would have also downsides, but IMHO the benefits \nwould outweigh them. The ones I see:\n1. Everytime a commit is made, a big change of the history-metadata tree \nwould need to happen. (Well but this is essentially equivalent to \nenabling an INDEX in a DB, you make WRITES more expensive in order to \nimprove the speed of READS.)\n2. Locking issues? I imagine that rewriting the indexes would open \nlonger time windows to have locking issues, but I'm not an expert in \nthis, please expand.\n3. Any more downsides you see?\n\n\nI would be glad for any feedback you have. Thanks, and have a great day!\n\n   Andrés\n\n-- \n"},{"id":"263657","messageId":"1434108815.5381.3.camel@kaarsemaker.net","threadId":"39611","inReplyTo":"mlefli$h6v$1@ger.gmane.org","subject":"Re: RFC: reverse history tree, for faster & background clones","fromName":"Dennis Kaarsemaker","fromEmail":"dennis@kaarsemaker.net","sentAt":"2015-06-12T11:33:35Z","receivedAt":"2015-06-12T11:33:35Z","isPatch":false,"sender":{"key":"dennis@kaarsemaker.net","avatar":"https://avatars.githubusercontent.com/u/200649?v=4"},"body":"On vr, 2015-06-12 at 13:26 +0200, Andres G. Aragoneses wrote:\n\n> AFAIU git stores the contents of a repo as a sequence of patches in the \n> .git metadata folder. \n\nIt does not, it stores full snapshots of files.\n\n[I've cut the example, as it's not how git works]\n\n> 1. `git clone --depth 1` would be way faster, and without the need of \n> on-demand compressing of packfiles in the server side, correct me if I'm \n> wrong?\n\nYou're wrong due to the misunderstanding of how git works :)\n\n> 2. `git clone` would be able to allow a \"fast operation, complete in the \n> background\" mode that would allow you to download the first snapshot of \n> the repo very quickly, so that the user would be able to start working \n> on his working directory very quickly, while a \"background job\" keeps \n> retreiving the history data in the background.\n\nThis could actually be a good thing, and can be emulated now with git\nclone --depth=1 and subsequent fetches in the background to deepen the\nhistory. I can see some value in clone doing this by itself, first doing\na depth=1 fetch, then launching itself into the background, giving you a\nworktree to play with earlier.\n\n-- \nDennis Kaarsemaker\nhttp://www.kaarsemaker.net\n"},{"id":"263658","messageId":"mlegd8$t71$1@ger.gmane.org","threadId":"39611","inReplyTo":"1434108815.5381.3.camel@kaarsemaker.net","subject":"Re: RFC: reverse history tree, for faster & background clones","fromName":"Andres G. Aragoneses","fromEmail":"knocte@gmail.com","sentAt":"2015-06-12T11:39:19Z","receivedAt":"2015-06-12T11:39:19Z","isPatch":false,"sender":{"key":"knocte@gmail.com","avatar":"https://gravatar.com/avatar/020c5605dcd7456d87a5f29a2af53e7bdf135fe9ccda514b9f34ab3a7e46db27?d=mp&s=160"},"body":"On 12/06/15 13:33, Dennis Kaarsemaker wrote:\n> On vr, 2015-06-12 at 13:26 +0200, Andres G. Aragoneses wrote:\n>\n>> AFAIU git stores the contents of a repo as a sequence of patches in the\n>> .git metadata folder.\n>\n> It does not, it stores full snapshots of files.\n\nIn bare repos too?\n\n\n>> 1. `git clone --depth 1` would be way faster, and without the need of\n>> on-demand compressing of packfiles in the server side, correct me if I'm\n>> wrong?\n>\n> You're wrong due to the misunderstanding of how git works :)\n\nThanks for pointing this out, do you mind giving me a link of some docs \nwhere I can correct my knowledge about this?\n\n\n>> 2. `git clone` would be able to allow a \"fast operation, complete in the\n>> background\" mode that would allow you to download the first snapshot of\n>> the repo very quickly, so that the user would be able to start working\n>> on his working directory very quickly, while a \"background job\" keeps\n>> retreiving the history data in the background.\n>\n> This could actually be a good thing, and can be emulated now with git\n> clone --depth=1 and subsequent fetches in the background to deepen the\n> history. I can see some value in clone doing this by itself, first doing\n> a depth=1 fetch, then launching itself into the background, giving you a\n> worktree to play with earlier.\n\nYou're right, didn't think about the feature that converts a --depth=1 \nrepo to a normal one. Then a patch that would create a --progressive \nflag (for instance, didn't think of a better name yet) for the `clone` \ncommand would actually be trivial to create, I assume, because it would \njust use `depth=1` and then retrieve the rest of the history in the \nbackground, right?\n\nThanks\n"},{"id":"263659","messageId":"1434112410.5381.8.camel@kaarsemaker.net","threadId":"39611","inReplyTo":"mlegd8$t71$1@ger.gmane.org","subject":"Re: RFC: reverse history tree, for faster & background clones","fromName":"Dennis Kaarsemaker","fromEmail":"dennis@kaarsemaker.net","sentAt":"2015-06-12T12:33:30Z","receivedAt":"2015-06-12T12:33:30Z","isPatch":false,"sender":{"key":"dennis@kaarsemaker.net","avatar":"https://avatars.githubusercontent.com/u/200649?v=4"},"body":"On vr, 2015-06-12 at 13:39 +0200, Andres G. Aragoneses wrote:\n> On 12/06/15 13:33, Dennis Kaarsemaker wrote:\n> > On vr, 2015-06-12 at 13:26 +0200, Andres G. Aragoneses wrote:\n> >\n> >> AFAIU git stores the contents of a repo as a sequence of patches in the\n> >> .git metadata folder.\n> >\n> > It does not, it stores full snapshots of files.\n> \n> In bare repos too?\n\nYes. A bare repo is nothing more than the .git dir of a non-bare repo\nwith the core.bare variable set to True :)\n\n> >> 1. `git clone --depth 1` would be way faster, and without the need of\n> >> on-demand compressing of packfiles in the server side, correct me if I'm\n> >> wrong?\n> >\n> > You're wrong due to the misunderstanding of how git works :)\n> \n> Thanks for pointing this out, do you mind giving me a link of some docs \n> where I can correct my knowledge about this?\n\nhttp://git-scm.com/book/en/v2/Git-Internals-Git-Objects should help.\n\n> >> 2. `git clone` would be able to allow a \"fast operation, complete in the\n> >> background\" mode that would allow you to download the first snapshot of\n> >> the repo very quickly, so that the user would be able to start working\n> >> on his working directory very quickly, while a \"background job\" keeps\n> >> retreiving the history data in the background.\n> >\n> > This could actually be a good thing, and can be emulated now with git\n> > clone --depth=1 and subsequent fetches in the background to deepen the\n> > history. I can see some value in clone doing this by itself, first doing\n> > a depth=1 fetch, then launching itself into the background, giving you a\n> > worktree to play with earlier.\n> \n> You're right, didn't think about the feature that converts a --depth=1 \n> repo to a normal one. Then a patch that would create a --progressive \n> flag (for instance, didn't think of a better name yet) for the `clone` \n> command would actually be trivial to create, I assume, because it would \n> just use `depth=1` and then retrieve the rest of the history in the \n> background, right?\n\nA naive implementation that does just clone --depth=1 and then fetch\n--unshallow would probably not be too hard, no. But whether that would\nbe the 'right' way of implementing it, I wouldn't know.\n\n-- \nDennis Kaarsemaker\nhttp://www.kaarsemaker.net\n"},{"id":"263816","messageId":"mlk28u$mm4$1@ger.gmane.org","threadId":"39611","inReplyTo":"1434112410.5381.8.camel@kaarsemaker.net","subject":"Re: RFC: reverse history tree, for faster & background clones","fromName":"Andres G. Aragoneses","fromEmail":"knocte@gmail.com","sentAt":"2015-06-14T14:14:53Z","receivedAt":"2015-06-14T14:14:53Z","isPatch":false,"sender":{"key":"knocte@gmail.com","avatar":"https://gravatar.com/avatar/020c5605dcd7456d87a5f29a2af53e7bdf135fe9ccda514b9f34ab3a7e46db27?d=mp&s=160"},"body":"On 12/06/15 14:33, Dennis Kaarsemaker wrote:\n> On vr, 2015-06-12 at 13:39 +0200, Andres G. Aragoneses wrote:\n>> On 12/06/15 13:33, Dennis Kaarsemaker wrote:\n>>> On vr, 2015-06-12 at 13:26 +0200, Andres G. Aragoneses wrote:\n>>>\n>>>> AFAIU git stores the contents of a repo as a sequence of patches in the\n>>>> .git metadata folder.\n>>>\n>>> It does not, it stores full snapshots of files.\n>>\n>> In bare repos too?\n>\n> Yes. A bare repo is nothing more than the .git dir of a non-bare repo\n> with the core.bare variable set to True :)\n>\n>>>> 1. `git clone --depth 1` would be way faster, and without the need of\n>>>> on-demand compressing of packfiles in the server side, correct me if I'm\n>>>> wrong?\n>>>\n>>> You're wrong due to the misunderstanding of how git works :)\n>>\n>> Thanks for pointing this out, do you mind giving me a link of some docs\n>> where I can correct my knowledge about this?\n>\n> http://git-scm.com/book/en/v2/Git-Internals-Git-Objects should help.\n\nWow, now I wonder if I should also propose a change to make git \noptionally not store the full snapshots, so save disk space. Thanks for \npointing this out to me.\n\n\n>>>> 2. `git clone` would be able to allow a \"fast operation, complete in the\n>>>> background\" mode that would allow you to download the first snapshot of\n>>>> the repo very quickly, so that the user would be able to start working\n>>>> on his working directory very quickly, while a \"background job\" keeps\n>>>> retreiving the history data in the background.\n>>>\n>>> This could actually be a good thing, and can be emulated now with git\n>>> clone --depth=1 and subsequent fetches in the background to deepen the\n>>> history. I can see some value in clone doing this by itself, first doing\n>>> a depth=1 fetch, then launching itself into the background, giving you a\n>>> worktree to play with earlier.\n>>\n>> You're right, didn't think about the feature that converts a --depth=1\n>> repo to a normal one. Then a patch that would create a --progressive\n>> flag (for instance, didn't think of a better name yet) for the `clone`\n>> command would actually be trivial to create, I assume, because it would\n>> just use `depth=1` and then retrieve the rest of the history in the\n>> background, right?\n>\n> A naive implementation that does just clone --depth=1 and then fetch\n> --unshallow would probably not be too hard, no. But whether that would\n> be the 'right' way of implementing it, I wouldn't know.\n\nOk, anyone else that can give an insight here?\n\nI imagine that I would not get real feedback until I send a [PATCH]...\n\nI guess I would use a user-facing message like this one:\n\nFinished cloning the last snapshot of the repository.\nAuto downloading the rest of the history in background.\n\n(Since there's already a similar \"background\" feature already wrt \nauto-packing the repository: `Auto packing the repository in background \nfor optimum performance. See \"git help gc\" for manual housekeeping.`.)\n\nThanks\n"}]}