{"thread":{"id":"50696","subject":"New Ft. for Git : Allow resumable cloning of repositories.","startedAt":"2019-03-08T15:43:50Z","lastAt":"2019-03-10T15:59:47Z","messageCount":3,"participants":["Kapil Jain","Jonathan Tan"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"370993","messageId":"CAMknYENWOW0mj6Bn9OooqKg-sZi9bZUO461Gv1F00=phNwLFQQ@mail.gmail.com","threadId":"50696","inReplyTo":null,"subject":"New Ft. for Git : Allow resumable cloning of repositories.","fromName":"Kapil Jain","fromEmail":"jkapil.cs@gmail.com","sentAt":"2019-03-08T15:43:36Z","receivedAt":"2019-03-08T15:43:50Z","isPatch":false,"sender":{"key":"jkapil.cs@gmail.com","avatar":null},"body":"Objective: Allow pause and resume functionality while cloning repositories.\n\nBelow is a rough idea on how this may be achieved.\n\n1) Create a repository_name.json file.\n2) repository_name.json will be an index file containing list of all\nthe files in the repository with default status being \"False\".\n   \"False\" status of a file signifies that this file is not yet fully\ndownloaded.\n\nSomething like this:\n\n{\n  'file1.ext' : \"False\",\n  'file2.ext' : \"False\",\n  'file3.ext' : \"False\"\n}\n\n3) As a file finishes downloading, say 'file1.ext' and 'file2.ext'\nhave finished downloading, their status will change to:\n\nSomething like this:\n\n{\n  'file1.ext' : \"True\",\n  'file2.ext' : \"True\",\n  'file3.ext' : \"False\"\n}\n\n4) Suppose due to some reason, before 'file3.ext' could finish\ndownload; cloning is interrupted.\n5) After the interruption the repository_name.json and downloaded\nfiles are preserved.\n6) Now, when cloning of the same repository begins next time, files\nwould be downloaded based on information taken from\nrepository_name.json file.\n\nNote 1: Doing this for cloning would be the main objective, further\nthis may be extended for fetching, pulling, and pushing too.\n\nNote 2: Since this is gsoc time, please don't take this to be a\nproject idea for gsoc, as it was pointed out on irc that this would be\na time intensive functionality.\n\nI want to work on building this functionality.\nPlease discuss thoughts on this, so as to make a technically sound to-do list.\n"},{"id":"371001","messageId":"20190308174314.129611-1-jonathantanmy@google.com","threadId":"50696","inReplyTo":"CAMknYENWOW0mj6Bn9OooqKg-sZi9bZUO461Gv1F00=phNwLFQQ@mail.gmail.com","subject":"Re: New Ft. for Git : Allow resumable cloning of repositories.","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2019-03-08T17:43:14Z","receivedAt":"2019-03-08T17:43:19Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> Objective: Allow pause and resume functionality while cloning repositories.\n> \n> Below is a rough idea on how this may be achieved.\n\nThis is indeed a nice feature to have, and thanks for details of how\nthis would be accomplished.\n\n> 1) Create a repository_name.json file.\n> 2) repository_name.json will be an index file containing list of all\n> the files in the repository with default status being \"False\".\n>    \"False\" status of a file signifies that this file is not yet fully\n> downloaded.\n> \n> Something like this:\n> \n> {\n>   'file1.ext' : \"False\",\n>   'file2.ext' : \"False\",\n>   'file3.ext' : \"False\"\n> }\n\nOne issue is that when cloning a repository, we do not download many\nfiles - we only download one dynamically generated packfile containing\nall the objects we want.\n\nYou might be interested in some work I'm doing to offload part of the\npackfile response to CDNs:\n\nhttps://public-inbox.org/git/cover.1550963965.git.jonathantanmy@google.com/\n\nThis means that when cloning/fetching, multiple files could be\ndownloaded, meaning that a scheme like you suggest would be more\nworthwhile. (In fact, I allude to such a scheme in the design document\nin patch 5.)\n"},{"id":"371080","messageId":"CAMknYEMP73D=LSKKvYKpmTdR3LAxc5UMgT3gxiQDZBghkLFo_g@mail.gmail.com","threadId":"50696","inReplyTo":"20190308174314.129611-1-jonathantanmy@google.com","subject":"Re: New Ft. for Git : Allow resumable cloning of repositories.","fromName":"Kapil Jain","fromEmail":"jkapil.cs@gmail.com","sentAt":"2019-03-10T15:59:33Z","receivedAt":"2019-03-10T15:59:47Z","isPatch":false,"sender":{"key":"jkapil.cs@gmail.com","avatar":null},"body":"On Fri, Mar 8, 2019 at 11:13 PM Jonathan Tan <jonathantanmy@google.com> wrote:\n> This is indeed a nice feature to have, and thanks for details of how\n> this would be accomplished.\n>\n> One issue is that when cloning a repository, we do not download many\n> files - we only download one dynamically generated packfile containing\n> all the objects we want.\n\nSince the packfile is dynamically generated specifically for a client\nrequest, and is destroyed from the server as soon as the connection\nbetween them closes.\nIs this the reason why we cannot pause it in between like we can do\nwith download managers ?\n\nI read through the progit ebook 'git internels' chapter and the\nfollowing thought came to me:\n\nAssume a pack file as follows:\n---\n$ git verify-pack -v .git/objects/pack/pack-\n978e03944f5c581011e6998cd0e9e30000905586.idx\nb042a60ef7dff760008df33cee372b945b6e884e blob   22054 5799 1463\n033b4468fa6b2a9547a70d88d1bbe8bf3f9ed0d5 blob   9 20 7262 1 \\\n  b042a60ef7dff760008df33cee372b945b6e884e\n.git/objects/pack/pack-978e03944f5c581011e6998cd0e9e30000905586.pack: ok\n---\n\nHere 033b blob refers b042 blob, and both blobs are different versions\nof the same file.\n\nBefore this pack was made, both of these blobs were stored separately\nand thus were taking more space.\nPackfile is made to save space, by only storing latest version and its\ndelta with earlier version. Both delta and latest version are stored\nin compressed form right ?\n\nNow, here is another approach to save space without needing to create pack:\n\nEarlier both the blobs had their object files as:\n\n.git/objects/03/3b4468fa6b2a9547a70d88d1bbe8bf3f9ed0d5\n.git/objects/b0/42a60ef7dff760008df33cee372b945b6e884e\n\nLets say b042 is latest and 033b is its earlier version.\n\nwhat git does in packfile can be done right here by:\n\nstoring latest version in\n.git/objects/b0/42a60ef7dff760008df33cee372b945b6e884e and its delta\nin .git/objects/03/3b4468fa6b2a9547a70d88d1bbe8bf3f9ed0d5, with the\ndelta version we can add a header that tells it to check for\n.git/objects/b0/42a60ef7dff760008df33cee372b945b6e884e and apply delta\non it to get the earlier version.\n\nDoing this, eliminates the big packfile, and all the objects are\nspread into folders. We can now make this resume-able right ?\n\nPlease point out what i missed here.\nIs it possible to do the above ? if yes then what was the reason to\nintroduce concept of packfile ?\n\n> You might be interested in some work I'm doing to offload part of the\n> packfile response to CDNs:\n>\n> https://public-inbox.org/git/cover.1550963965.git.jonathantanmy@google.com/\n>\n> This means that when cloning/fetching, multiple files could be\n> downloaded, meaning that a scheme like you suggest would be more\n> worthwhile. (In fact, I allude to such a scheme in the design document\n> in patch 5.)\n\ncurrently reading through all the discussion on this strategy.\n"}]}