{"thread":{"id":"3576","subject":"git clone downloads objects that are in GIT_OBJECT_DIRECTORY","startedAt":"2006-03-06T01:08:25Z","lastAt":"2006-03-06T09:20:18Z","messageCount":6,"participants":["Benjamin LaHaise","Shawn Pearce","Junio C Hamano","sean","Johannes Schindelin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"17216","messageId":"20060306010825.GF20768@kvack.org","threadId":"3576","inReplyTo":null,"subject":"git clone downloads objects that are in GIT_OBJECT_DIRECTORY","fromName":"Benjamin LaHaise","fromEmail":"bcrl@kvack.org","sentAt":"2006-03-06T01:08:25Z","receivedAt":"2006-03-06T01:08:25Z","isPatch":false,"sender":{"key":"bcrl@kvack.org","avatar":null},"body":"Hi folks,\n\nDoing a fresh git clone git://some.git.url/ foo seems to download the \nentire remote repository even if all the objects are already stored in \nGIT_OBJECT_DIRECTORY=/home/bcrl/.git/object .  Is this a known bug?  \nAt 100MB for a kernel, this takes a *long* time.\n\n\t\t-ben (who needed to free up disk space)\n-- \n\"Time is of no importance, Mr. President, only life is important.\"\nDon't Email: <dont@kvack.org>.\n"},{"id":"17221","messageId":"20060306014253.GD25790@spearce.org","threadId":"3576","inReplyTo":"20060306010825.GF20768@kvack.org","subject":"Re: git clone downloads objects that are in GIT_OBJECT_DIRECTORY","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-03-06T01:42:53Z","receivedAt":"2006-03-06T01:42:53Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Benjamin LaHaise <bcrl@kvack.org> wrote:\n> Hi folks,\n> \n> Doing a fresh git clone git://some.git.url/ foo seems to download the \n> entire remote repository even if all the objects are already stored in \n> GIT_OBJECT_DIRECTORY=/home/bcrl/.git/object .  Is this a known bug?  \n> At 100MB for a kernel, this takes a *long* time.\n\nI believe it is a known missing feature.  :-) git-clone doesn't\nprep HEAD to have some sort of starting point so the pull it uses\nto download everything literally downloads everything as nothing\nis in common.\n\nOne could work around it by running git-init-db to create the new\nclone locally, git-update-ref HEAD to some commit which you have in\ncommon with the remote, create a origin file, then perform a git-pull.\nThis would only download the objects between the commit you put into\nHEAD and the current master of the remote...  But that is actually\nsome work.\n\nI think Cogito's clone is capable of restarting a failed clone; I\nwonder if that logic would benefit you here?\n\nIs using a common GIT_OBJECT_DIRECTORY across many clones actually\npretty common?  Maybe its time that git-clone gets some more smarts\nwith regards to what it yanks from the origin.\n\n-- \nShawn.\n"},{"id":"17224","messageId":"7vfylwcncn.fsf@assigned-by-dhcp.cox.net","threadId":"3576","inReplyTo":"20060306014253.GD25790@spearce.org","subject":"Re: git clone downloads objects that are in GIT_OBJECT_DIRECTORY","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-06T02:34:48Z","receivedAt":"2006-03-06T02:34:48Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Shawn Pearce <spearce@spearce.org> writes:\n\n> Benjamin LaHaise <bcrl@kvack.org> wrote:\n>> Hi folks,\n>> \n>> Doing a fresh git clone git://some.git.url/ foo seems to download the \n>> entire remote repository even if all the objects are already stored in \n>> GIT_OBJECT_DIRECTORY=/home/bcrl/.git/object .  Is this a known bug?  \n>> At 100MB for a kernel, this takes a *long* time.\n>\n> I believe it is a known missing feature.  :-) git-clone doesn't\n> prep HEAD to have some sort of starting point so the pull it uses\n> to download everything literally downloads everything as nothing\n> is in common.\n\nYou would first 'clone -l -s' from your local repository and\nthen clone into that from whatever remote, perhaps.\n"},{"id":"17227","messageId":"20060306025702.GH25790@spearce.org","threadId":"3576","inReplyTo":"7vfylwcncn.fsf@assigned-by-dhcp.cox.net","subject":"Re: git clone downloads objects that are in GIT_OBJECT_DIRECTORY","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-03-06T02:57:02Z","receivedAt":"2006-03-06T02:57:02Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <junkio@cox.net> wrote:\n> Shawn Pearce <spearce@spearce.org> writes:\n> \n> > Benjamin LaHaise <bcrl@kvack.org> wrote:\n> >> Hi folks,\n> >> \n> >> Doing a fresh git clone git://some.git.url/ foo seems to download the \n> >> entire remote repository even if all the objects are already stored in \n> >> GIT_OBJECT_DIRECTORY=/home/bcrl/.git/object .  Is this a known bug?  \n> >> At 100MB for a kernel, this takes a *long* time.\n> >\n> > I believe it is a known missing feature.  :-) git-clone doesn't\n> > prep HEAD to have some sort of starting point so the pull it uses\n> > to download everything literally downloads everything as nothing\n> > is in common.\n> \n> You would first 'clone -l -s' from your local repository and\n> then clone into that from whatever remote, perhaps.\n\nYea but that's about as much fun as creating a bare repository\nby hand.  (Which I've been doing up until this thread prompted me\nto read git-clone.sh and learn the existance of --bare.)\n\nIt might be nicer if the user could place a list of locally (here\nlocally being possibly remote but closer network-wise) available\nrepositories which should be considered as sources for faster\ncloning.  When cloning a remote repository git-clone would try to\nexamine each of the designated repositories to see if any of them\nhave commits in common with the remote; if so clone off that and\nthen pull from the remote, but designating the remote as `origin'.\n\nThis could be ugly if you have a large number of locally available\ncandidates or if the candidates are many (e.g. 1000s) commits\nbehind the remote being cloned.  But it would save the user from\npulling down 100+MB of objects they already have while making it\nvery convient to establish a new repository+working directory based\non someone else's publically available repository.\n\n\nOr we could just tell the user to create their own clone script,\ne.g. kernel-clone:\n\n\t#!/bin/sh\n\tgit-clone -l -n -s ~/kernel-base \"$2\" &&\n\tcd \"$2\" &&\n\techo \"URL: $1\" >.git/remotes/origin &&\n\techo \"Pull: master:origin\" >>.git/remotes/origin &&\n\tgit-pull\n\n\nBut it would be better if it was more integrated, and somehow\nslightly more automatic...\n\n-- \nShawn.\n"},{"id":"17228","messageId":"BAYC1-PASMTP05F90E7F1807F05A274507AEE90@CEZ.ICE","threadId":"3576","inReplyTo":"20060306025702.GH25790@spearce.org","subject":"Re: git clone downloads objects that are in GIT_OBJECT_DIRECTORY","fromName":"sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2006-03-06T03:31:15Z","receivedAt":"2006-03-06T03:31:15Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Sun, 5 Mar 2006 21:57:02 -0500\nShawn Pearce <spearce@spearce.org> wrote:\n\n> It might be nicer if the user could place a list of locally (here\n> locally being possibly remote but closer network-wise) available\n> repositories which should be considered as sources for faster\n> cloning.  When cloning a remote repository git-clone would try to\n> examine each of the designated repositories to see if any of them\n> have commits in common with the remote; if so clone off that and\n> then pull from the remote, but designating the remote as `origin'.\n\nIt is already easy to start from a similar repo (eg. locally cloned)\nif you wish to conserve bandwidth.\n\nHowever, it might be nice to have a command that allows you to \nchange origin information for a repo without needing to know git\ninternals; maybe something like:\n\n$ git set-origin <URL>\n\nOr maybe better:\n\n$ git set-remote --pull master:origin origin <URL>\n\nSean\n"},{"id":"17249","messageId":"Pine.LNX.4.63.0603061019070.1422@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3576","inReplyTo":"BAYC1-PASMTP05F90E7F1807F05A274507AEE90@CEZ.ICE","subject":"Re: git clone downloads objects that are in GIT_OBJECT_DIRECTORY","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-03-06T09:20:18Z","receivedAt":"2006-03-06T09:20:18Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 5 Mar 2006, sean wrote:\n\n> However, it might be nice to have a command that allows you to \n> change origin information for a repo without needing to know git\n> internals; maybe something like:\n> \n> $ git set-origin <URL>\n> \n> Or maybe better:\n> \n> $ git set-remote --pull master:origin origin <URL>\n\nFWIW, I once sent patches to make this easier by placing this information \ninto the config file, but for reasons I did not understand, they were \nrejected. Sigh!\n\nCiao,\nDscho\n"}]}