{"thread":{"id":"2095","subject":"auto-packing on kernel.org? please?","startedAt":"2005-10-13T18:44:30Z","lastAt":"2005-11-23T14:18:13Z","messageCount":33,"participants":["Linus Torvalds","Dirk Behme","Daniel Barkalow","Nick Hengeveld","Brian Gerst","Junio C Hamano","Johannes Schindelin","Carl Baldwin","Chuck Lever","Catalin Marinas"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"10079","messageId":"Pine.LNX.4.64.0510131113490.15297@g5.osdl.org","threadId":"2095","inReplyTo":null,"subject":"auto-packing on kernel.org? please?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-13T18:44:30Z","receivedAt":"2005-10-13T18:44:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nI know we tried this once earlier, and it caused problems, but that was \nwhen pack-files were new, and not everybody could handle them. These days, \nif you can't handle pack-files, kernel.org is already pretty useless, \nbecause all the major packages use them anyway, because people have \npacked their repositories by hand.\n\nSo I'm suggesting we try to do an automatic repack every once in a while. \n\nIn my suggestion, there would be two levels of repacking: \"incremental\" \nand \"full\", and both of them would count the number of files before they \nrun, so that you'd only do it when it seems worthwhile.\n\nThis is a _really_ simple heuristic:\n\n - incremental repacking run every day:\n\n\t#\n\t# Check if we have more than a couple of hundred\n\t# unpacked objects - approximated by whether we\n\t# have any \"00\" directory with more than one \n\t#\n\t# This means that we don't repack projects that\n\t# that don't have a lot of work going on.\n\t#\n\t# Note: with really new versions of git, the \"00\"\n\t# directory may not exist if it has been pruned\n\t# away, so handle that gracefully.\n\t#\n\texport GIT_DIR=${1:-.}\n\tobjs=$(find \"$GIT_DIR/objects/00\" -type f 2> /dev/null | wc -l)\n\tif [ \"$obj\" -gt 0 ]; then\n\t\tgit repack &&\n\t\t\tgit prune-packed\n\tfi\n\n - \"full repack\" every week if the number of packs has grown to be bigger \n   than say 10 (ie even a very active projects will never have a full \n   repack more than every other week)\n\n\t#\n\t# Check if we have lots of packs, where \"lots\" is defined as 10.\n\t#\n\t# Note: with something that was generated with an old version\n\t# of git, the \"pack\" directory may not exist, so handle that\n\t# gracefully.\n\t#\n\texport GIT_DIR=${1:-.}\n\tpacks=$(find \"$GIT_DIR/objects/pack\" -name '*.idx' 2> /dev/null | wc -l)\n\tif [ \"$packs\" -gt 10 ]; then\n\t\tgit repack -a -d &&\n\t\t\tgit prune-packed\n\tfi\n\n - do a full repack of everything once to start with.\n\n\texport GIT_DIR=${1:-.}\n\tgit repack -a -d &&\n\t\tgit prune-packed\n\nthe above three trivial scripts just take a single argument, which becomes \nthe GIT_DIR (and if no argument exists, it would default to \".\")\n\nIs there any reason not to do this? Right now mirroring is slow, and \nwebgit is also getting to be very slow sometimes. I bet we'd be _much_ \nbetter off with this kind of setup.\n\nNOTE! The above is the \"stupid\" approach, which totally ignores alternate \ndirectories, and isn't able to take advantage of the fact that many \nprojects could share objects. But it's simple, and it's efficient (eg it \nwon't spend time on things like the large historic archives which don't \nchange, but that would be expensive to repack if you didn't check for the \nneed).\n\nSo we could try to come up with a better approach eventually, which would \nautomatically notice alternate directories and not repack stuff that \nexists there, but I'm pretty sure that the above would already help a \n_lot_, and while pack-files have been been around forever, the \n\"alternates\" support is still pretty new, so the above is also the \"safer\" \nthing to do.\n\nWe'd only do the automatic thing on stuff under /pub/scm, of course: not \nstuff in peoples home directories etc..\n\nPeter?\n\n\t\t\tLinus\n"},{"id":"10081","messageId":"Pine.LNX.4.64.0510131422161.23590@g5.osdl.org","threadId":"2095","inReplyTo":"434EC07C.30505@pobox.com","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-13T21:23:56Z","receivedAt":"2005-10-13T21:23:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 13 Oct 2005, Jeff Garzik wrote:\n> \n> Right now, things go through an expand-contract cycle:\n> \n> * people base repos off of Marcelo or Linus's git repo, including using those\n> pack files (saves download bandwidth, disk space through hardlinks).\n> \n> * as 3rd parties and Marcelo/Linus merge stuff, .git/objects/* grows with\n> individual files.\n> \n> * once a month/release/whatever, Linus packs his repo, allowing all the repos\n> following his to use those pack files, pruning a ton of objects off of\n> kernel.org.\n> \n> I have real users of my git repos who can't just download a 100MB pack file in\n> an hour, it takes them many hours.\n\nArgh.\n\nOk, I'm going to follow this up with three small patches that add a \"-l\" \nflag to \"git repack\", which does only a \"local repack\" (ie it will pack \nonly objects that are _not_ in packs in alternate object directories).\n\nThat will hopefully mean that this usage case is supported too.\n\n\t\tLinus\n"},{"id":"10150","messageId":"435264B1.2010204@de.bosch.com","threadId":"2095","inReplyTo":"Pine.LNX.4.64.0510131422161.23590@g5.osdl.org","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Dirk Behme","fromEmail":"dirk.behme@de.bosch.com","sentAt":"2005-10-16T14:33:21Z","receivedAt":"2005-10-16T14:33:21Z","isPatch":false,"sender":{"key":"dirk.behme@de.bosch.com","avatar":null},"body":"\n> On Thu, 13 Oct 2005, Jeff Garzik wrote:\n>\n>>I have real users of my git repos who can't just download a 100MB pack file in\n>>an hour, it takes them many hours.\n\nSeems that I'm one of these users (but using an other repo).\n\nPack files are very nice saving bandwith and disk space. But what I \ndislike is that I often have to download same information twice: Remote \n.git/objects/* repo grows and I update my local repo daily against this. \nThen once a month/release/whatever .git/objects/* are packed into one \nfile. This new pack file then is downloaded as well, but most/all of the \ninformation in this file is already in my local repo and downloaded \nagain. Something like\n\n- detect that there is new pack file in remote repo\n- check what is in this remote pack file\n- if in local repo no or only few .git/objects/* are missing, download \nthe missing ones and create an identical copy of remote pack file using \nlocal .git/objects/*. Don't download remote pack file.\n- remove all local .git/objects/* now in pack file\n\nwould be nice.\n\nOr is this already possible? Or do I misunderstand anything?\n\nDirk\n"},{"id":"10151","messageId":"Pine.LNX.4.63.0510161122570.23242@iabervon.org","threadId":"2095","inReplyTo":"435264B1.2010204@de.bosch.com","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-10-16T15:44:46Z","receivedAt":"2005-10-16T15:44:46Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 16 Oct 2005, Dirk Behme wrote:\n\n> > On Thu, 13 Oct 2005, Jeff Garzik wrote:\n> >\n> > >I have real users of my git repos who can't just download a 100MB pack file\n> > >in\n> > >an hour, it takes them many hours.\n> \n> Seems that I'm one of these users (but using an other repo).\n> \n> Pack files are very nice saving bandwith and disk space. But what I dislike is\n> that I often have to download same information twice: Remote .git/objects/*\n> repo grows and I update my local repo daily against this. Then once a\n> month/release/whatever .git/objects/* are packed into one file. This new pack\n> file then is downloaded as well, but most/all of the information in this file\n> is already in my local repo and downloaded again. Something like\n> \n> - detect that there is new pack file in remote repo\n> - check what is in this remote pack file\n> - if in local repo no or only few .git/objects/* are missing, download the\n> missing ones and create an identical copy of remote pack file using local\n> .git/objects/*. Don't download remote pack file.\n\nThis is the problem: it's impossible to download only a few objects from a \npack file from an HTTP server, because those don't exist on the server as \nseparate files.\n\nThe current HTTP code actually never downloads a pack file unless a needed \nobject is not anywhere else, at which point it has no choice but to \ndownload the pack.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"10152","messageId":"20051016161244.GE5509@reactrix.com","threadId":"2095","inReplyTo":"Pine.LNX.4.63.0510161122570.23242@iabervon.org","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2005-10-16T16:12:44Z","receivedAt":"2005-10-16T16:12:44Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Sun, Oct 16, 2005 at 11:44:46AM -0400, Daniel Barkalow wrote:\n\n> This is the problem: it's impossible to download only a few objects from a \n> pack file from an HTTP server, because those don't exist on the server as \n> separate files.\n\nIs it possible to determine the object locations inside the remote pack\nfile?  If so, it would be possible to use Range: headers to download\nselected objects from a pack.\n\n-- \nFor a successful technology, reality must take precedence over public\nrelations, for nature cannot be fooled.\n"},{"id":"10153","messageId":"43527E86.8000907@didntduck.org","threadId":"2095","inReplyTo":"20051016161244.GE5509@reactrix.com","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Brian Gerst","fromEmail":"bgerst@didntduck.org","sentAt":"2005-10-16T16:23:34Z","receivedAt":"2005-10-16T16:23:34Z","isPatch":false,"sender":{"key":"bgerst@didntduck.org","avatar":null},"body":"Nick Hengeveld wrote:\n> On Sun, Oct 16, 2005 at 11:44:46AM -0400, Daniel Barkalow wrote:\n> \n>> This is the problem: it's impossible to download only a few objects from a \n>> pack file from an HTTP server, because those don't exist on the server as \n>> separate files.\n> \n> Is it possible to determine the object locations inside the remote pack\n> file?  If so, it would be possible to use Range: headers to download\n> selected objects from a pack.\n> \n\nNot possible because the entire pack is compressed.\n\n--\n\t\t\t\tBrian Gerst\n"},{"id":"10154","messageId":"7vzmp9xuwe.fsf@assigned-by-dhcp.cox.net","threadId":"2095","inReplyTo":"43527E86.8000907@didntduck.org","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-16T16:56:49Z","receivedAt":"2005-10-16T16:56:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Brian Gerst <bgerst@didntduck.org> writes:\n\n>> Is it possible to determine the object locations inside the remote\n>> pack\n>> file?  If so, it would be possible to use Range: headers to download\n>> selected objects from a pack.\n\nThat's what the .idx file is for, except that after you fetch\nthe range, you may find you would need something else that the\nobject is delta against.\n"},{"id":"10155","messageId":"Pine.LNX.4.63.0510161906220.25573@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"2095","inReplyTo":"43527E86.8000907@didntduck.org","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2005-10-16T17:10:11Z","receivedAt":"2005-10-16T17:10:11Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 16 Oct 2005, Brian Gerst wrote:\n\n> Nick Hengeveld wrote:\n> > On Sun, Oct 16, 2005 at 11:44:46AM -0400, Daniel Barkalow wrote:\n> > \n> > > This is the problem: it's impossible to download only a few objects from a\n> > > pack file from an HTTP server, because those don't exist on the server as\n> > > separate files.\n> > \n> > Is it possible to determine the object locations inside the remote pack\n> > file?  If so, it would be possible to use Range: headers to download\n> > selected objects from a pack.\n> > \n> \n> Not possible because the entire pack is compressed.\n\nMaybe we should introduce an option which only packs objects of a minimal \nage (something like \"pack only objects 2 days and older\")? This could be \nused to autopackage as long as HTTP is the preferred protocol, so that if \nyou update daily, you already have those objects.\n\nAlternatively, git-prune-packed could have an option to prune only those \nobjects older than 2 days.\n\nCiao,\nDscho\n"},{"id":"10156","messageId":"43528AC1.2060904@didntduck.org","threadId":"2095","inReplyTo":"43527E86.8000907@didntduck.org","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Brian Gerst","fromEmail":"bgerst@didntduck.org","sentAt":"2005-10-16T17:15:45Z","receivedAt":"2005-10-16T17:15:45Z","isPatch":false,"sender":{"key":"bgerst@didntduck.org","avatar":null},"body":"Brian Gerst wrote:\n> Nick Hengeveld wrote:\n> \n>> On Sun, Oct 16, 2005 at 11:44:46AM -0400, Daniel Barkalow wrote:\n>>\n>>> This is the problem: it's impossible to download only a few objects \n>>> from a pack file from an HTTP server, because those don't exist on \n>>> the server as separate files.\n>>\n>>\n>> Is it possible to determine the object locations inside the remote pack\n>> file?  If so, it would be possible to use Range: headers to download\n>> selected objects from a pack.\n>>\n> \n> Not possible because the entire pack is compressed.\n\nI should have looked at the source more closely before stating that. \nEach object gets compressed individually, so this would be possible.\n\n--\n\t\t\t\t\tBrian Gerst\n"},{"id":"10163","messageId":"20051016213341.GF5509@reactrix.com","threadId":"2095","inReplyTo":"7vzmp9xuwe.fsf@assigned-by-dhcp.cox.net","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2005-10-16T21:33:41Z","receivedAt":"2005-10-16T21:33:41Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Sun, Oct 16, 2005 at 09:56:49AM -0700, Junio C Hamano wrote:\n\n> That's what the .idx file is for, except that after you fetch\n> the range, you may find you would need something else that the\n> object is delta against.\n\nWould it make sense to load the pack indexes for each base up front,\nand then fetch individual objects from a pack if they exist in one of\na base's pack indexes?  In such a case, it may not even make sense to\ntry fetching the object directly first.\n\nWhat are the circumstances under which it makes more sense to fetch the\nwhole pack rather than fetching individual objects from it?\n\n-- \nFor a successful technology, reality must take precedence over public\nrelations, for nature cannot be fooled.\n"},{"id":"10166","messageId":"7vwtkd6rik.fsf@assigned-by-dhcp.cox.net","threadId":"2095","inReplyTo":"20051016213341.GF5509@reactrix.com","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-16T22:12:03Z","receivedAt":"2005-10-16T22:12:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nick Hengeveld <nickh@reactrix.com> writes:\n\n> On Sun, Oct 16, 2005 at 09:56:49AM -0700, Junio C Hamano wrote:\n>\n>> That's what the .idx file is for, except that after you fetch\n>> the range, you may find you would need something else that the\n>> object is delta against.\n>\n> Would it make sense to load the pack indexes for each base up front,\n> and then fetch individual objects from a pack if they exist in one of\n> a base's pack indexes?  In such a case, it may not even make sense to\n> try fetching the object directly first.\n>\n> What are the circumstances under which it makes more sense to fetch the\n> whole pack rather than fetching individual objects from it?\n\nIt would make sense if we end up needing most them anyway, I\nthink.\n\nWe are probably far from this, but ideally, we should be able to\nset up something like this.\n\nWe encourage the server side to prepare packs this way [*1*].\n\n -- development --> time --> flows --> this --> way -->\n\n (optional)\n full ------------------------------------------------\n\n base ---------------\n 6mo                 ---------------------------------\n 3mo                               -------------------\n 1mo                                    --------------\n 2wk                                        ----------\n 1wk                                             -----\n                                                      ^\n                                 last pack optimization\n\n\nThat is, a big base pack (say v2.6.12), and multiple packs to\nbring people that were in-sync at various time up-to-date to the\ntime when the set of packs were last optimized.  Any objects\ncreated after the last pack optimization time are left unpacked\nuntil the next pack optimization time.  It might not be a bad\nidea to also have a \"full\" pack.\n\nFor example, if you were in-sync 5-months ago, fetching 3mo pack\nwould not be enough and you would need to get 6mo pack to become\nup-to-date wrt the last pack optimization (say 3 days ago).  You\nwould have obtained the objects not in pack, created within the\nlast 3 days, already as individual objects before realizing that\nyou would need to fetch some pack.\n\nThen, we can teach git-http-fetch to do:\n\n - If an object is unavailable unpacked, get all the indices\n   from that repository (and probably its alternates while we\n   are at it).\n\n - Among the set of packs that contain the object we are\n   currently interested in, try to find the \"best\" pack.  The\n   definition of \"best\" would be a balancing act of finding the\n   one that contains the least number of objects we already\n   have, and the one that contains the most number of objects we\n   do not have yet.\n\nThe commit walker always goes from present to past, so you would\nstart from fetching the latest, presumably unpacked objects, and\nas soon as you hit the last pack optimization boundary, you have\nchoices of multiple packs.  If you are relatively up-to-date,\nyou would find that 1mo pack has more things you already have\nthan 1wk pack, although both of them would fit the bill -- at\nthat point you choose to download 1wk pack.  On the other hand,\nif you are behind, you may find that 3mo pack has more things\nyou do not have than 1wk or 2wk or 1mo pack, and using 3mo pack\nwould become the right choice for you.\n\nI think most repositories have a few related heads and their\nheads almost never rewind, so favoring the pack that contains\nthe most number of objects we do not have would be the right\nstrategy in practice for the downloader.\n\n\n[Footnote]\n\n*1* This is different from a proposal posted on the list earlier\nby somebody (I think it was Pasky but I may be mistaken) which\nlooked like this:\n\n -- development --> time --> flows --> this --> way -->\n\n base ---------------\n 6mo                 --------------\n 3mo                               -----\n 1mo                                    ----\n 2wk                                        -----\n 1wk                                             -----\n\nThe thing is, sum of 3mo+1mo+2wk+1wk packs in the latter scheme\ntends to be a lot bigger than the size of 3mo pack in the former\nscheme.\n"},{"id":"10171","messageId":"20051017060659.GH5509@reactrix.com","threadId":"2095","inReplyTo":"7vwtkd6rik.fsf@assigned-by-dhcp.cox.net","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2005-10-17T06:06:59Z","receivedAt":"2005-10-17T06:06:59Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Sun, Oct 16, 2005 at 03:12:03PM -0700, Junio C Hamano wrote:\n\n>  - Among the set of packs that contain the object we are\n>    currently interested in, try to find the \"best\" pack.  The\n>    definition of \"best\" would be a balancing act of finding the\n>    one that contains the least number of objects we already\n>    have, and the one that contains the most number of objects we\n>    do not have yet.\n\nTo get a complete list of objects we do not have yet, fetch will need\nto walk all the trees first and then make another pass to process\nall the missing objects.  Is it worth considering a case where the\nmissing objects are packed along with objects that don't need to be\ntransferred?  From the use cases you described, it's not clear that\nsituation would ever really happen.\n\nIf the blobs have been packed, it seems likely that the tree objects will\nalso be packed, so fetching them during the first pass will either involve\nfetching a pack without being able to determine which is best or fetching\nthe appropriate ranges from packs to get the tree objects.\n\n-- \nFor a successful technology, reality must take precedence over public\nrelations, for nature cannot be fooled.\n"},{"id":"10172","messageId":"7voe5o366d.fsf@assigned-by-dhcp.cox.net","threadId":"2095","inReplyTo":"20051017060659.GH5509@reactrix.com","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-17T08:21:14Z","receivedAt":"2005-10-17T08:21:14Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nick Hengeveld <nickh@reactrix.com> writes:\n\n> To get a complete list of objects we do not have yet, fetch will need\n> to walk all the trees first and then make another pass to process\n> all the missing objects.\n\nNotice I did not say \"we do not have yet but we will need\" -- I\njust said \"we do not have yet\".\n\nThe assumption, which is the property the suggested packing\nstrategy has, is that older objects that are needed to complete\nthe history leading to the current tip are packed in those\nn-month/n-week packs, so if we do not have them we would likely\nbe needing them, although we might not have walked that far back\nin history yet.\n\nThe previous \"packing strategy\" picture was certainly too\nsimplified.  Obviously we would not want to repack everything\nevery week for different periods all the way back -- we would\nwant to leave old huge pack untouched to help server side (and\nmirroring), so instead of having a single \"pack optimization\nboundary\", we would probably need some staggering as well for\narchived material.\n\nThis is a revised example.\n\n1yr -----\n9mo      --------\n6mo              ----------\n3mo                        ------------------\n1mo                              ------------  \n2wk                                  --------\n1wk                                      ----\n\nWe keep track of \"the current heads and tags\" for each week.\nEvery week, we can do something like this:\n\n - rotate the record, and create a new one:\n   mv .save/wk11 .save/wk12\n   mv .save/wk10 .save/wk11\n   mv .save/wk9 .save/wk10\n   ...\n   mv .save/wk0 .save/wk1\n   find .git/refs -type f -print | xargs cat >.save/wk0\n \n - prepare a pack to allow a single pack fetch to bring a\n   repository that had everything reachable from wk$N refs\n   up-to-date to the current, for selected recent weeks (say N=1,\n   2, 4, 12):\n\n   for N in 1 2 4 12\n   do\n       name=$(git-rev-list --objects \\\n                 $(sed -e 's/^/^/' .save/wk$N) \\\n                 $(cat .save/wk0) |\n              git-pack-object pack-) &&\n       mv pack-$name.* .git/objects/pack/.\n   done\n\n   remove the pack files that we created this way last week from\n   the repository (if the repository did not have any activity\n   during the last week we would have created the same set of\n   packs.  make sure we do not remove them).\n\n - except that, we keep the longest period (i.e. N=12 in this\n   example) one every N weeks (that's how 1yr, 9mo, 6mo packs in\n   the picture are kept).\n\nThis way, really old stuff (say, older than 3mo) will stay\nintact and will not be repacked, so people reasonably up-to-date\n(within 12 weeks in the example) need to fetch only one pack\n(and unpacked objects since the last pack optimization), but\npeople without the ancient history need to go further back.\n"},{"id":"10178","messageId":"20051017174123.GI5509@reactrix.com","threadId":"2095","inReplyTo":"7voe5o366d.fsf@assigned-by-dhcp.cox.net","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2005-10-17T17:41:23Z","receivedAt":"2005-10-17T17:41:23Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Mon, Oct 17, 2005 at 01:21:14AM -0700, Junio C Hamano wrote:\n\n> The assumption, which is the property the suggested packing\n> strategy has, is that older objects that are needed to complete\n> the history leading to the current tip are packed in those\n> n-month/n-week packs, so if we do not have them we would likely\n> be needing them, although we might not have walked that far back\n> in history yet.\n\nGotcha - I'm still thinking in terms of content distribution, where\nyou only need a specific version of a tree to be available locally\nand explicitly don't want to transfer history.  In our case, using\npacks doesn't make sense at the moment.\n\n-- \nFor a successful technology, reality must take precedence over public\nrelations, for nature cannot be fooled.\n"},{"id":"10182","messageId":"Pine.LNX.4.63.0510171348370.23242@iabervon.org","threadId":"2095","inReplyTo":"20051016213341.GF5509@reactrix.com","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-10-17T19:13:21Z","receivedAt":"2005-10-17T19:13:21Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 16 Oct 2005, Nick Hengeveld wrote:\n\n> On Sun, Oct 16, 2005 at 09:56:49AM -0700, Junio C Hamano wrote:\n> \n> > That's what the .idx file is for, except that after you fetch\n> > the range, you may find you would need something else that the\n> > object is delta against.\n> \n> Would it make sense to load the pack indexes for each base up front,\n> and then fetch individual objects from a pack if they exist in one of\n> a base's pack indexes?  In such a case, it may not even make sense to\n> try fetching the object directly first.\n\nAt the start, you have the option of either fetching the list of packs or \nthe object. There are three cases:\n\n 1) the object isn't available separately; we need to fetch the list of \n    packs to find it in a pack.\n 2) there aren't any new packs; we need to fetch the object individually.\n 3) the object is present both individually and in a pack.\n\n(2) is more common than (1), because we don't repack every update. (3) \ndoesn't happen at all, currently, because we prune after packing. So it \nmakes most sense to try the object at once.\n\nOn the other hand, the parallel code should probably do both at the same \ntime, since it can, and it only causes notable latency, not bandwidth. We \nprobably also ought to speculatively get any new index files in parallel \nwith whatever else we're doing, since it is likely that we'll need some \npack at some point, and then we'll need all the index files to decide what \npack to get.\n\n> What are the circumstances under which it makes more sense to fetch the\n> whole pack rather than fetching individual objects from it?\n\nI'm not sure there's a good way of deciding without a plan for what \nconditions cause there to be a choice.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"10183","messageId":"7v3bmzzz30.fsf@assigned-by-dhcp.cox.net","threadId":"2095","inReplyTo":"20051017174123.GI5509@reactrix.com","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-10-17T20:08:03Z","receivedAt":"2005-10-17T20:08:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nick Hengeveld <nickh@reactrix.com> writes:\n\n> Gotcha - I'm still thinking in terms of content distribution, where\n> you only need a specific version of a tree to be available locally\n> and explicitly don't want to transfer history.\n\nIn other words, you'd want to also support CVS-like \"working\ntree has the specific version, and history is not kept here, but\navailable on demand, possibly over the network\" mode of\noperation.  I'd say why not.  We could aim to have \"working tree\nhas the specific version and partial history of recent versions,\nand the ancient history is available on demand, possibly over\nthe network\" mode of operation.\n\nIt is somewhat different from the primary focus of what we have\nbeen doing, but I think it is a natural extension.  The\ninvariant is that once you have a ref pointing at a specific\ncommit, everything reachable from it ought to be available to\nyou.\n\nAnd we have extended the definition of \"available\" over time.\nInitially, you needed to have individual objects, and then we\nmade it so they could live in packs, and now they could even be\nborrowed from another repository via alternates.  We currently\ndo not consider \"lazily fetchable over the network\" as\n\"available\", but I do not object too much to that, as long as it\nis an optional feature.\n\nThis probably is a post 1.0 item, though.  Off the top of my\nhead, we would need:\n\n - a way for the user to say \"unless I ask explicitly otherwise,\n   do not bother me if the commits older than these ones are\n   incomplete\" -- an milder version of cauterizing commit chain\n   via info/grafts.\n\n - a way for the user to say \"this time I am explicitly\n   overriding the above -- I am interested in older history\".\n\n - change to fsck-objects, fetch- and probably upload-pack on\n   the other end, and commit walkers to honor the above two.\n\nMost of these can probably be done by existing info/grafts\nmechanism, but even then definitely would need a nicer user\ninterface.\n\nOnce this is in place, range requests to pick data for\nindividual objects from packs residing on a remote HTTP server\nwould start to make sense.\n"},{"id":"10187","messageId":"Pine.LNX.4.63.0510171830030.23242@iabervon.org","threadId":"2095","inReplyTo":"7v3bmzzz30.fsf@assigned-by-dhcp.cox.net","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-10-17T22:56:38Z","receivedAt":"2005-10-17T22:56:38Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Mon, 17 Oct 2005, Junio C Hamano wrote:\n\n> Nick Hengeveld <nickh@reactrix.com> writes:\n> \n> > Gotcha - I'm still thinking in terms of content distribution, where\n> > you only need a specific version of a tree to be available locally\n> > and explicitly don't want to transfer history.\n> \n> In other words, you'd want to also support CVS-like \"working\n> tree has the specific version, and history is not kept here, but\n> available on demand, possibly over the network\" mode of\n> operation.  I'd say why not.  We could aim to have \"working tree\n> has the specific version and partial history of recent versions,\n> and the ancient history is available on demand, possibly over\n> the network\" mode of operation.\n> \n> It is somewhat different from the primary focus of what we have\n> been doing, but I think it is a natural extension.  The\n> invariant is that once you have a ref pointing at a specific\n> commit, everything reachable from it ought to be available to\n> you.\n\nWouldn't \"git fetch http://.../foo.git/ master^{tree}\" do the right thing?\n\nYou get only the current tree, and write a ref to the tree instead of the \ncommit, maintaining the invariant. Of course, fetch.c needs a bit of work \nso that it can fetch objects in the process of figuring out what the \nrefspec that it's really trying to fetch, but that should be simple \nenough.\n\nOf course, this really isolates you from the history, since you don't even \nremember what the commit was that you've got the tree from, but that may \nnot be an issue in a pure content distribution setup. Also, a pack file of \na single tree isn't going to be terribly efficient, because pack files \nmostly exploit the high similarity between different versions of the same \nfile.\n\nMy other idea is to have a file of things that you expect to be missing, \neven though they are referenced, and where to expect to find them if \nnecessary. Then you could download the latest commit, mark its parents \n(unless you have them) as known-missing, and write the ref.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"10190","messageId":"Pine.LNX.4.64.0510171617460.3369@g5.osdl.org","threadId":"2095","inReplyTo":"Pine.LNX.4.63.0510171830030.23242@iabervon.org","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-10-17T23:19:52Z","receivedAt":"2005-10-17T23:19:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 17 Oct 2005, Daniel Barkalow wrote:\n> \n> Wouldn't \"git fetch http://.../foo.git/ master^{tree}\" do the right thing?\n\nThe pack pullers have trouble with anything that isn't commit-based, \nbecause they do all the \"figure out what we have in common\" logic based on \nthe commit history.\n\nSo if you fetch a tree, it by definition doesn't _have_ any history, and \nthe pack pullers will always pack the whole tree. I think.\n\n\t\tLinus\n"},{"id":"10191","messageId":"20051017235423.GK5509@reactrix.com","threadId":"2095","inReplyTo":"7v3bmzzz30.fsf@assigned-by-dhcp.cox.net","subject":"Re: [kernel.org users] Re: auto-packing on kernel.org? please?","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2005-10-17T23:54:23Z","receivedAt":"2005-10-17T23:54:23Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Mon, Oct 17, 2005 at 01:08:03PM -0700, Junio C Hamano wrote:\n\n>  - a way for the user to say \"unless I ask explicitly otherwise,\n>    do not bother me if the commits older than these ones are\n>    incomplete\" -- an milder version of cauterizing commit chain\n>    via info/grafts.\n> \n>  - a way for the user to say \"this time I am explicitly\n>    overriding the above -- I am interested in older history\".\n> \n>  - change to fsck-objects, fetch- and probably upload-pack on\n>    the other end, and commit walkers to honor the above two.\n\nThat's how I interpreted the -c and -a command-line arguments to the\ncommit walkers.  git-fetch calls them with -a but we've been using -t\nto only follow the tree objects and it's been working great.\n\nPerhaps that would be a good way for the commit walker to decide whether\nto transfer a full pack file - it may not make sense if it wasn't told\nto get history.\n\n-- \nFor a successful technology, reality must take precedence over public\nrelations, for nature cannot be fooled.\n"},{"id":"12457","messageId":"20051121190151.GA2568@hpsvcnb.fc.hp.com","threadId":"2095","inReplyTo":"Pine.LNX.4.64.0510131113490.15297@g5.osdl.org","subject":"Re: auto-packing on kernel.org? please?","fromName":"Carl Baldwin","fromEmail":"cnb@fc.hp.com","sentAt":"2005-11-21T19:01:51Z","receivedAt":"2005-11-21T19:01:51Z","isPatch":false,"sender":{"key":"cnb@fc.hp.com","avatar":null},"body":"I have a question about automatic repacking.\n\nI am thinking of turning something like Linus' repacking heuristic loose\non my repositories.  I just want to make sure it is as safe as possible.\n\nAt the core of the incremental and full repack strategies are these\nstatements.\n\nIncremental...\n> \t\tgit repack &&\n> \t\t\tgit prune-packed\n\nFull...\n> \t\tgit repack -a -d &&\n> \t\t\tgit prune-packed\n\nAre there some built in safety checks in 'git repack' and/or 'git\nprune-packed' to guard against corruption?  In the long run, I would\nfeel more comfortable with somelike like this:\n\ngit repack\ngit verify-pack <new pack>\ngit prune-packed\n\nWould something like this even work with 'git repack -a -d'?  Is there a\nway to do something like the following for a full repack to achieve the\nultimate in paranoia?\n\ngit repack -a\ngit verify-pack <new pack file>\ngit trash-redundant-packs <new pack file>\ngit prune-packed\n\nCarl\n\nOn Thu, Oct 13, 2005 at 11:44:30AM -0700, Linus Torvalds wrote:\n> \n> I know we tried this once earlier, and it caused problems, but that was \n> when pack-files were new, and not everybody could handle them. These days, \n> if you can't handle pack-files, kernel.org is already pretty useless, \n> because all the major packages use them anyway, because people have \n> packed their repositories by hand.\n> \n> So I'm suggesting we try to do an automatic repack every once in a while. \n> \n> In my suggestion, there would be two levels of repacking: \"incremental\" \n> and \"full\", and both of them would count the number of files before they \n> run, so that you'd only do it when it seems worthwhile.\n> \n> This is a _really_ simple heuristic:\n> \n>  - incremental repacking run every day:\n> \n> \t#\n> \t# Check if we have more than a couple of hundred\n> \t# unpacked objects - approximated by whether we\n> \t# have any \"00\" directory with more than one \n> \t#\n> \t# This means that we don't repack projects that\n> \t# that don't have a lot of work going on.\n> \t#\n> \t# Note: with really new versions of git, the \"00\"\n> \t# directory may not exist if it has been pruned\n> \t# away, so handle that gracefully.\n> \t#\n> \texport GIT_DIR=${1:-.}\n> \tobjs=$(find \"$GIT_DIR/objects/00\" -type f 2> /dev/null | wc -l)\n> \tif [ \"$obj\" -gt 0 ]; then\n> \t\tgit repack &&\n> \t\t\tgit prune-packed\n> \tfi\n> \n>  - \"full repack\" every week if the number of packs has grown to be bigger \n>    than say 10 (ie even a very active projects will never have a full \n>    repack more than every other week)\n> \n> \t#\n> \t# Check if we have lots of packs, where \"lots\" is defined as 10.\n> \t#\n> \t# Note: with something that was generated with an old version\n> \t# of git, the \"pack\" directory may not exist, so handle that\n> \t# gracefully.\n> \t#\n> \texport GIT_DIR=${1:-.}\n> \tpacks=$(find \"$GIT_DIR/objects/pack\" -name '*.idx' 2> /dev/null | wc -l)\n> \tif [ \"$packs\" -gt 10 ]; then\n> \t\tgit repack -a -d &&\n> \t\t\tgit prune-packed\n> \tfi\n> \n>  - do a full repack of everything once to start with.\n> \n> \texport GIT_DIR=${1:-.}\n> \tgit repack -a -d &&\n> \t\tgit prune-packed\n> \n> the above three trivial scripts just take a single argument, which becomes \n> the GIT_DIR (and if no argument exists, it would default to \".\")\n> \n> Is there any reason not to do this? Right now mirroring is slow, and \n> webgit is also getting to be very slow sometimes. I bet we'd be _much_ \n> better off with this kind of setup.\n> \n> NOTE! The above is the \"stupid\" approach, which totally ignores alternate \n> directories, and isn't able to take advantage of the fact that many \n> projects could share objects. But it's simple, and it's efficient (eg it \n> won't spend time on things like the large historic archives which don't \n> change, but that would be expensive to repack if you didn't check for the \n> need).\n> \n> So we could try to come up with a better approach eventually, which would \n> automatically notice alternate directories and not repack stuff that \n> exists there, but I'm pretty sure that the above would already help a \n> _lot_, and while pack-files have been been around forever, the \n> \"alternates\" support is still pretty new, so the above is also the \"safer\" \n> thing to do.\n> \n> We'd only do the automatic thing on stuff under /pub/scm, of course: not \n> stuff in peoples home directories etc..\n> \n> Peter?\n> \n> \t\t\tLinus\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n\n-- \n- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -\n Carl Baldwin                        Systems VLSI Laboratory\n Hewlett Packard Company\n MS 88                               work: 970 898-1523\n 3404 E. Harmony Rd.                 work: Carl.N.Baldwin@hp.com\n Fort Collins, CO 80525              home: Carl@ecBaldwin.net\n- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -\n"},{"id":"12461","messageId":"Pine.LNX.4.64.0511211110480.13959@g5.osdl.org","threadId":"2095","inReplyTo":"20051121190151.GA2568@hpsvcnb.fc.hp.com","subject":"Re: auto-packing on kernel.org? please?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-21T19:24:11Z","receivedAt":"2005-11-21T19:24:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Nov 2005, Carl Baldwin wrote:\n>\n> I have a question about automatic repacking.\n> \n> I am thinking of turning something like Linus' repacking heuristic loose\n> on my repositories.  I just want to make sure it is as safe as possible.\n> \n> At the core of the incremental and full repack strategies are these\n> statements.\n> \n> Incremental...\n> > \t\tgit repack &&\n> > \t\t\tgit prune-packed\n> \n> Full...\n> > \t\tgit repack -a -d &&\n> > \t\t\tgit prune-packed\n\nNOTE! Since that email, \"git repack\" has gotten a \"local\" option (-l), \nwhich is very useful if the repositories have pointers to alternates.\n\nSo do\n\n\tgit repack -l\n\ninstead, to get much better packs (and \"-a -d\" for the full case, of \ncourse).\n\nOther that than, the old email suggestion should still be fine.\n\n> Are there some built in safety checks in 'git repack' and/or 'git\n> prune-packed' to guard against corruption?  In the long run, I would\n> feel more comfortable with somelike like this:\n> \n> git repack\n> git verify-pack <new pack>\n> git prune-packed\n\nYou can certainly do that if you are nervous. It might even be a good \nidea: just for fun, I just did\n\n\tgit clone -l git git-clone\n\tcd git-clone\n\n\t# pick an object at random\n\trm .git/objects/f7/c3d39fe3db6da3a307da385a7a1cb563ed15f7\n\n\tgit repack -a -d\n\nand it said:\n\n\terror: Could not read f7c3d39fe3db6da3a307da385a7a1cb563ed15f7\n\tfatal: bad tree object f7c3d39fe3db6da3a307da385a7a1cb563ed15f7\n\nbut then it created the pack _anyway_, and said:\n\n\tPacking 27 objects\n\tPack pack-13bfca704078175c1c1c59964553b14f7b952651 created.\n\nand happily removed all the old ones.\n\nSo right now, repacking a broken archive can actually break it even more.\n\nNOTE! Your \"git verify-pack\" wouldn't even catch this: the _pack_ is fine, \nit's just incomplete.\n\nOf course, this only happens if the repository was broken to begin with, \nso arguably it's not that bad. But it does show that git-repack should be \nmore careful and return an error more aggressively.\n\nCan anybody tell me how to do that sanely? Right now we do\n\n\t..\n\tname=$(git-rev-list --objects $rev_list $(git-rev-parse $rev_parse) |\n\t        git-pack-objects --non-empty $pack_objects .tmp-pack) ||\n\t        exit 1\n\t..\n\nand the thing is, the \"git-pack-objects\" thing is happy, it's the \n\"git-rev-list\" that fails. So because the last command in the pipeline \nreturns ok, we think it all is ok..\n\n(This is one of the reasons I much prefer working in C over working in \nshell: it may be twenty times more lines, but when you have a problem, the \nfix is always obvious..)\n\nAnyway, with that fixed, a \"git repack\" in many ways would be a mini-fsck, \nso it should be very safe in general. Modulo any other bugs like the \nabove.\n\n\t\tLinus\n"},{"id":"12465","messageId":"7v3blprcwk.fsf@assigned-by-dhcp.cox.net","threadId":"2095","inReplyTo":"Pine.LNX.4.64.0511211110480.13959@g5.osdl.org","subject":"Re: auto-packing on kernel.org? please?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-21T19:58:35Z","receivedAt":"2005-11-21T19:58:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Can anybody tell me how to do that sanely? Right now we do\n>\n> \t..\n> \tname=$(git-rev-list --objects $rev_list $(git-rev-parse $rev_parse) |\n> \t        git-pack-objects --non-empty $pack_objects .tmp-pack) ||\n> \t        exit 1\n> \t..\n>\n> and the thing is, the \"git-pack-objects\" thing is happy, it's the \n> \"git-rev-list\" that fails. So because the last command in the pipeline \n> returns ok, we think it all is ok..\n\nOne cop-out: do fsck-objects upfront before making a pack.  This\nwould populate your buffer cache so it might not be a bad thing.\n\nAlternatively:\n\n        name=$( {\n                git-rev-list --objects $rev_list $(git-rev-parse $rev_parse) ||\n                echo Gaaahhh\n        } | git-pack-objects --non-empty $pack_objects .tmp-pack)\n"},{"id":"12468","messageId":"Pine.LNX.4.64.0511211211130.13959@g5.osdl.org","threadId":"2095","inReplyTo":"7v3blprcwk.fsf@assigned-by-dhcp.cox.net","subject":"Re: auto-packing on kernel.org? please?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-21T20:38:31Z","receivedAt":"2005-11-21T20:38:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Nov 2005, Junio C Hamano wrote:\n> \n> One cop-out: do fsck-objects upfront before making a pack.  This\n> would populate your buffer cache so it might not be a bad thing.\n\nWell, it's extremely expensive most of the time. It's often as expensive \nas the packing itself. So I don't like that option very much.\n\n> Alternatively:\n> \n>         name=$( {\n>                 git-rev-list --objects $rev_list $(git-rev-parse $rev_parse) ||\n>                 echo Gaaahhh\n>         } | git-pack-objects --non-empty $pack_objects .tmp-pack)\n\nActually, some dim memories prodded me to some man-page digging, and the \n\"pipefail\" option in particular.\n\nIt seems to be a common option to both ksh and bash, so\n\n\tset -o pipefail\n\nseems like it should fix this. Sadly, I think it's pretty recent in bash \n(ksh apparently got it in -93, bash seems to have gotten it only as of \nversion 3.0, which is definitely recent enough that we can't just assume \nit).\n\n[ Also, bash seems to have a variable called $PIPESTATUS, but that's \n  bash-specific (I don't know when it was enabled). ]\n\nAnyway, doing a\n\n\tset -o pipefail\n\nshould never be the wrong thing to do, but the problem is figuring out \nwhether the option is available or not, since if it isn't available, it's \nconsidered an error ;/\n\nSo with all that, how about we take your \"Gaah\" idea, and simplify it: \njust pipe stderr too. That, together with making git-pack-objects tell \nwhat garbage it got, actually does the rigth thing:\n\n\t[torvalds@g5 git-clone]$ git repack -a -d\n\tfatal: expected sha1, got garbage:\n\t error: Could not read 7f59dbbb8f8d479c1d31453eac06ec765436a780\n\nwith this pretty simple patch.\n\nWhaddaya think?\n\n\t\t\tLinus\n\n---\ndiff --git a/git-repack.sh b/git-repack.sh\nindex 4e16d34..c0f271d 100755\n--- a/git-repack.sh\n+++ b/git-repack.sh\n@@ -41,7 +41,7 @@ esac\n if [ \"$local\" ]; then\n \tpack_objects=\"$pack_objects --local\"\n fi\n-name=$(git-rev-list --objects $rev_list $(git-rev-parse $rev_parse) |\n+name=$(git-rev-list --objects $rev_list $(git-rev-parse $rev_parse) 2>&1 |\n \tgit-pack-objects --non-empty $pack_objects .tmp-pack) ||\n \texit 1\n if [ -z \"$name\" ]; then\ndiff --git a/pack-objects.c b/pack-objects.c\nindex 4e941e7..8864a31 100644\n--- a/pack-objects.c\n+++ b/pack-objects.c\n@@ -524,7 +524,7 @@ int main(int argc, char **argv)\n \t\tunsigned char sha1[20];\n \n \t\tif (get_sha1_hex(line, sha1))\n-\t\t\tdie(\"expected sha1, got garbage\");\n+\t\t\tdie(\"expected sha1, got garbage:\\n %s\", line);\n \t\thash = 0;\n \t\tp = line+40;\n \t\twhile (*p) {\n"},{"id":"12478","messageId":"7vlkzhof9y.fsf@assigned-by-dhcp.cox.net","threadId":"2095","inReplyTo":"Pine.LNX.4.64.0511211211130.13959@g5.osdl.org","subject":"Re: auto-packing on kernel.org? please?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-21T21:35:37Z","receivedAt":"2005-11-21T21:35:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> ...just pipe stderr too. That, together with making git-pack-objects tell \n> what garbage it got, actually does the rigth thing:\n>\n> \t[torvalds@g5 git-clone]$ git repack -a -d\n> \tfatal: expected sha1, got garbage:\n> \t error: Could not read 7f59dbbb8f8d479c1d31453eac06ec765436a780\n>\n> with this pretty simple patch.\n>\n> Whaddaya think?\n\nObviously the right thing to do ;-).  I like it.\n"},{"id":"12516","messageId":"4382AC11.5090209@citi.umich.edu","threadId":"2095","inReplyTo":"Pine.LNX.4.64.0511211110480.13959@g5.osdl.org","subject":"Re: auto-packing on kernel.org? please?","fromName":"Chuck Lever","fromEmail":"cel@citi.umich.edu","sentAt":"2005-11-22T05:26:41Z","receivedAt":"2005-11-22T05:26:41Z","isPatch":false,"sender":{"key":"cel@citi.umich.edu","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Mon, 21 Nov 2005, Carl Baldwin wrote:\n> \n>>I have a question about automatic repacking.\n>>\n>>I am thinking of turning something like Linus' repacking heuristic loose\n>>on my repositories.  I just want to make sure it is as safe as possible.\n>>\n>>At the core of the incremental and full repack strategies are these\n>>statements.\n>>\n>>Incremental...\n>>\n>>>\t\tgit repack &&\n>>>\t\t\tgit prune-packed\n>>\n>>Full...\n>>\n>>>\t\tgit repack -a -d &&\n>>>\t\t\tgit prune-packed\n> \n> \n> NOTE! Since that email, \"git repack\" has gotten a \"local\" option (-l), \n> which is very useful if the repositories have pointers to alternates.\n> \n> So do\n> \n> \tgit repack -l\n> \n> instead, to get much better packs (and \"-a -d\" for the full case, of \n> course).\n> \n> Other that than, the old email suggestion should still be fine.\n\ni've been playing with \"git repack\" on StGIT-managed repositories.\n\non NFS, using packs instead of individual objects is quite a bit faster, \nbecause a single NFS GETATTR will tell you if your NFS client's cached \npack file is still valid, whereas a whole bunch of GETATTRs are required \nfor validating individual object files.\n\nthere are some things repacking does that breaks StGIT, though.\n\ngit repack -d\n\nseems to remove old commits that StGIT was still depending on.\n\ngit repack -a -n\n\nseems to work fine with StGIT, as does\n\ngit prune-packed\n\ni'm really interested in trying out the new command to remove redundant \nobjects and packs, but haven't gotten around to it yet.\n\n\nbegin:vcard\nfn:Chuck Lever\nn:Lever;Charles\norg:Network Appliance, Incorporated;Linux NFS Client Development\nadr:535 West William Street, Suite 3100;;Center for Information Technology Integration;Ann Arbor;MI;48103-4943;USA\nemail;internet:cel@citi.umich.edu\ntitle:Member of Technical Staff\ntel;work:+1 734 763-4415\ntel;fax:+1 734 763 4434\ntel;home:+1 734 668-1089\nx-mozilla-html:FALSE\nurl:http://www.monkey.org/~cel/\nversion:2.1\nend:vcard\n\n"},{"id":"12519","messageId":"Pine.LNX.4.64.0511212134330.13959@g5.osdl.org","threadId":"2095","inReplyTo":"4382AC11.5090209@citi.umich.edu","subject":"Re: auto-packing on kernel.org? please?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-22T05:41:17Z","receivedAt":"2005-11-22T05:41:17Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 22 Nov 2005, Chuck Lever wrote:\n>\n> there are some things repacking does that breaks StGIT, though.\n> \n> git repack -d\n> \n> seems to remove old commits that StGIT was still depending on.\n\nIf that is true, then \"git-fsck-cache\" probably also reports errors on a \nStGIT repository. No? Basically, it implies that the tool doesn't know how \nto find all the \"heads\".\n\nCould somebody (Catalin?) perhaps tell how tools like git-fsck-cache and \ngit-repack could figure out which objects are still in use by stgit?\n\nPreferably with some generic mechanism that _other_ projects (not just \nstgit) might want to use?\n\nThe preferred way would be to just list the references somewhere under \n.git/refs/stgit, in which case fsck and repack should pick them up \nautomatically (so clearly stgit doesn't do that right now ;).\n\nIt also implies that doing a \"git prune\" will do horribly bad things to a \nstgit repo, since it would remove all the objects that it thinks aren't \nreachable..\n\n> git repack -a -n\n> \n> seems to work fine with StGIT,\n\nWell, it \"works\", but not \"fine\". Since it doesn't know about the stgit \nobjects, it won't ever pack them.\n\nBut maybe that's what stgit wants (since they are \"temporary\"), but it \ndoes mean that if you see a big advantage from packing, you might be \nlosing some of it.\n\n\t\tLinus\n"},{"id":"12539","messageId":"b0943d9e0511220613h5978a600l@mail.gmail.com","threadId":"2095","inReplyTo":"Pine.LNX.4.64.0511212134330.13959@g5.osdl.org","subject":"Re: auto-packing on kernel.org? please?","fromName":"Catalin Marinas","fromEmail":"catalin.marinas@gmail.com","sentAt":"2005-11-22T14:13:40Z","receivedAt":"2005-11-22T14:13:40Z","isPatch":false,"sender":{"key":"catalin.marinas@gmail.com","avatar":null},"body":"On 22/11/05, Linus Torvalds <torvalds@osdl.org> wrote:\n> On Tue, 22 Nov 2005, Chuck Lever wrote:\n> > there are some things repacking does that breaks StGIT, though.\n> >\n> > git repack -d\n> >\n> > seems to remove old commits that StGIT was still depending on.\n>\n> If that is true, then \"git-fsck-cache\" probably also reports errors on a\n> StGIT repository. No? Basically, it implies that the tool doesn't know how\n> to find all the \"heads\".\n\nIndeed, 'git repack -d'  or 'git prune' might remove the patches which\nare not applied since there is no link to them from .git/refs/.\n\n> Could somebody (Catalin?) perhaps tell how tools like git-fsck-cache and\n> git-repack could figure out which objects are still in use by stgit?\n\nThey don't figure this out at the moment. I initially thought about\nimplementing these commands in StGIT so that they would pass the\nproper references.\n\n> Preferably with some generic mechanism that _other_ projects (not just\n> stgit) might want to use?\n>\n> The preferred way would be to just list the references somewhere under\n> .git/refs/stgit, in which case fsck and repack should pick them up\n> automatically (so clearly stgit doesn't do that right now ;).\n\nI thought about adding .git/refs/patches/<branch>/* files\ncorresponding to the every StGIT patch. Are the above git commands\nlooking at all depths in the .git/refs/ directory?\n\n> > git repack -a -n\n> >\n> > seems to work fine with StGIT,\n>\n> Well, it \"works\", but not \"fine\". Since it doesn't know about the stgit\n> objects, it won't ever pack them.\n>\n> But maybe that's what stgit wants (since they are \"temporary\"), but it\n> does mean that if you see a big advantage from packing, you might be\n> losing some of it.\n\nThe 'git repack -a' command would include the applied patches in the\nnewly created pack but leave out the unapplied ones. It would be even\nbetter to leave all of them out since the StGIT patches are frequently\nchanged but an independent mechanism for this would complicate GIT -\n'git repack' shouldn't pack any of the objects found in\n.git/refs/patches/, even if they are reachable via .git/refs/heads/*\n(and maybe call the patches directory something like\n.git/refs/unpackable or volatile).\n\n--\nCatalin\n"},{"id":"12546","messageId":"Pine.LNX.4.64.0511220904040.13959@g5.osdl.org","threadId":"2095","inReplyTo":"b0943d9e0511220613h5978a600l@mail.gmail.com","subject":"Re: auto-packing on kernel.org? please?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-22T17:05:44Z","receivedAt":"2005-11-22T17:05:44Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 22 Nov 2005, Catalin Marinas wrote:\n>\n> > The preferred way would be to just list the references somewhere under\n> > .git/refs/stgit, in which case fsck and repack should pick them up\n> > automatically (so clearly stgit doesn't do that right now ;).\n> \n> I thought about adding .git/refs/patches/<branch>/* files\n> corresponding to the every StGIT patch. Are the above git commands\n> looking at all depths in the .git/refs/ directory?\n\nYes. Or at least they're supposed to. If they are not, it's a bug \nregardless, and we'll fix it.\n\n> The 'git repack -a' command would include the applied patches in the\n> newly created pack but leave out the unapplied ones. It would be even\n> better to leave all of them out since the StGIT patches are frequently\n> changed but an independent mechanism for this would complicate GIT -\n> 'git repack' shouldn't pack any of the objects found in\n> .git/refs/patches/, even if they are reachable via .git/refs/heads/*\n> (and maybe call the patches directory something like\n> .git/refs/unpackable or volatile).\n\nIf we have some default location (and .git/refs/patches/ sounds good), we \ncan make git do the right thing - find them for git-fsck-objects, and \nignore them for git-repack.\n\n\t\tLinus\n"},{"id":"12547","messageId":"20051122172558.GA1935@hpsvcnb.fc.hp.com","threadId":"2095","inReplyTo":"Pine.LNX.4.64.0511211110480.13959@g5.osdl.org","subject":"Re: auto-packing on kernel.org? please?","fromName":"Carl Baldwin","fromEmail":"cnb@fc.hp.com","sentAt":"2005-11-22T17:25:58Z","receivedAt":"2005-11-22T17:25:58Z","isPatch":false,"sender":{"key":"cnb@fc.hp.com","avatar":null},"body":"On Mon, Nov 21, 2005 at 11:24:11AM -0800, Linus Torvalds wrote:\n> NOTE! Since that email, \"git repack\" has gotten a \"local\" option (-l), \n> which is very useful if the repositories have pointers to alternates.\n> \n> So do\n> \n> \tgit repack -l\n> \n> instead, to get much better packs (and \"-a -d\" for the full case, of \n> course).\n\nI'm assuming that this option will have no effect on a repository with\nno alternates file.\n\n> Other that than, the old email suggestion should still be fine.\n\n[snip]\n\n> You can certainly do that if you are nervous. It might even be a good \n> idea: just for fun, I just did\n> \n> \tgit clone -l git git-clone\n> \tcd git-clone\n> \n> \t# pick an object at random\n> \trm .git/objects/f7/c3d39fe3db6da3a307da385a7a1cb563ed15f7\n> \n> \tgit repack -a -d\n> \n> and it said:\n> \n> \terror: Could not read f7c3d39fe3db6da3a307da385a7a1cb563ed15f7\n> \tfatal: bad tree object f7c3d39fe3db6da3a307da385a7a1cb563ed15f7\n> \n> but then it created the pack _anyway_, and said:\n> \n> \tPacking 27 objects\n> \tPack pack-13bfca704078175c1c1c59964553b14f7b952651 created.\n> \n> and happily removed all the old ones.\n> \n> So right now, repacking a broken archive can actually break it even more.\n\nInteresting.\n\n> NOTE! Your \"git verify-pack\" wouldn't even catch this: the _pack_ is fine, \n> it's just incomplete.\n\nIn my opinion, git repack did the right thing in creating the pack even\nif it is more broken.  Starting with a broken repository was the real\nproblem.  git repack shouldn't need to worry too much about it.\n\nLooking at it from the nervous repository admin's point of view I think\nhe would want to make sure that the repository is good to begin with.  I\nthink this should be left up to the repository owner and maybe not git\nrepack.  Although, the check that you do following this is probably a\ngood idea.\n\n> Of course, this only happens if the repository was broken to begin with, \n> so arguably it's not that bad. But it does show that git-repack should be \n> more careful and return an error more aggressively.\n> \n> Can anybody tell me how to do that sanely? Right now we do\n> \n> \t..\n> \tname=$(git-rev-list --objects $rev_list $(git-rev-parse $rev_parse) |\n> \t        git-pack-objects --non-empty $pack_objects .tmp-pack) ||\n> \t        exit 1\n> \t..\n> \n> and the thing is, the \"git-pack-objects\" thing is happy, it's the \n> \"git-rev-list\" that fails. So because the last command in the pipeline \n> returns ok, we think it all is ok..\n> \n> (This is one of the reasons I much prefer working in C over working in \n> shell: it may be twenty times more lines, but when you have a problem, the \n> fix is always obvious..)\n> \n> Anyway, with that fixed, a \"git repack\" in many ways would be a mini-fsck, \n> so it should be very safe in general. Modulo any other bugs like the \n> above.\n> \n> \t\tLinus\n\n*NOTE*  There is one question that I feel remains unanswered.  Is it\npossible to split up the repack -a and repack -d so that the nervous\nrepository owner can insert a git verify-pack in the middle.\n\nI'm not nearly this nervous about repositories that I keep for myself\nbut I have ownership of some repositories on which many people may\ndepend.  I will feel better if I can verify the pack separately from\ngit-repack before I do the (potentially destructive) -d to remove old\npacks.\n\nI don't mean to say that I don't trust git repack to do the right thing.\nFundamentally, I just think that I shouldn't depend on it to do the\nright thing in order to avoid corruption in my repository.\n\nCarl\n\nPS  I love that the git object store is designed so that object files\nnever *need* to be removed, renamed, modified or otherwise touched in\nany way after being written to disk.  I think this makes git inherently\nextremely safe from corruption unlike many other older repository\ndesigns.  The only thing that breaks this inherent safety is the desire\nto pack repositories to avoid bloat.\n\nThat is why I want to be a little paranoid when I do the repacking.  I\nwant to maintain some inherent safety in the process that I use to pack\nthem.  This kind of inherent safety is much more valuable then even the\nhighest quality code written to actually do the packing.\n\n-- \n- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -\n Carl Baldwin                        Systems VLSI Laboratory\n Hewlett Packard Company\n MS 88                               work: 970 898-1523\n 3404 E. Harmony Rd.                 work: Carl.N.Baldwin@hp.com\n Fort Collins, CO 80525              home: Carl@ecBaldwin.net\n- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -\n"},{"id":"12553","messageId":"Pine.LNX.4.64.0511220939540.13959@g5.osdl.org","threadId":"2095","inReplyTo":"20051122172558.GA1935@hpsvcnb.fc.hp.com","subject":"Re: auto-packing on kernel.org? please?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-11-22T17:58:45Z","receivedAt":"2005-11-22T17:58:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 22 Nov 2005, Carl Baldwin wrote:\n\n> On Mon, Nov 21, 2005 at 11:24:11AM -0800, Linus Torvalds wrote:\n> > NOTE! Since that email, \"git repack\" has gotten a \"local\" option (-l), \n> > which is very useful if the repositories have pointers to alternates.\n> > \n> > So do\n> > \n> > \tgit repack -l\n> > \n> > instead, to get much better packs (and \"-a -d\" for the full case, of \n> > course).\n> \n> I'm assuming that this option will have no effect on a repository with\n> no alternates file.\n\nCorrect.\n\nThe only thing it does is that when it looks up an object, if it's not in \nour _own_ \".git/objects/\" dir, it won't pack it.\n\nActually, that's not entirely true. It isn't smart enough to know where \nevery object exists, so it only knows about remote _packs_. So what \nhappens is that if you do\n\n\tgit repack -l -a -d\n\nit will create a pack-file that contains _all_ unpacked objects (whether \nlocal or not) and all objects that are in local packs (because of the \n\"-a\"), but not any objects that are in \"alternate packs\".\n\nWhich is actually exactly what you want, if you are in the situation that \nkernel.org is, and you have people who point their alternates to mine: \nwhen I repack my objects, they'll use my packs, but other than that, \nthey'll prefer to use their own packs over any unpacked objects.\n\n> > So right now, repacking a broken archive can actually break it even more.\n> \n> Interesting.\n\nWell, with the latest git repack script, that should no longer be true.\n\n> > NOTE! Your \"git verify-pack\" wouldn't even catch this: the _pack_ is fine, \n> > it's just incomplete.\n> \n> In my opinion, git repack did the right thing in creating the pack even\n> if it is more broken.  Starting with a broken repository was the real\n> problem.  git repack shouldn't need to worry too much about it.\n\nWell, \"git repack\" did the wrong thing in that it never _noticed_, and it \nthen removed all old packs - even though those old packs contained objects \nthat we hadn't repacked because of the broken repository.\n\nOf course, _usually_ a broken repository is just that - broken. The way \nyou fix a broken repo is to find a non-broken one, and clone that. \nHowever, sometimes what you can do (if you literally just lost a few \nobjects) is to find a non-broken repo, and make that the _alternates_, in \nwhich case you may be able to save any work you had in the broken one \n(assuming you only lost objects that were available somewhere else).\n\n> Looking at it from the nervous repository admin's point of view I think\n> he would want to make sure that the repository is good to begin with.\n\nDoing an fsck is certainly always a good idea. I do a \"shallow\" fsck \nusually several times a day (\"shallow\" means that it doesn't fsck packs, \nonly new objects that I have aquired since the last repacking), and I do a \nfull fsck a couple of times a week.\n\nI don't actually know why I do that, though. I don't think I've really \n_ever_ had a broken repo since some very early days, except for the cases \nwhere I break things on purpose (like remove an object to check whether \n\"git repack\" does the right thing or not). I'm just used to it, and the \nshallow fsck takes a fraction of a second, so I tend to do it after each \npull.\n\nSo I really think that an admin has to be more than \"nervous\" to worry \nabout it. He has to be really anal.\n\n(Now, doing a repack and a fsck every week or so might be good, and \nautomatic shallow fsck's daily is probably a great idea too. After all, it \n_is_ checking checksums, so if you worry about security and want to make \nsure that nobody is trying to break in and do bad things to your repo, a \nregular fsck is a good thing even if you're not otherwise worried about \ncorruption).\n\n> *NOTE*  There is one question that I feel remains unanswered.  Is it\n> possible to split up the repack -a and repack -d so that the nervous\n> repository owner can insert a git verify-pack in the middle.\n\nThey are already split up inside \"git-repack\", so we could add a hook \nthere, I guess. See the git-repack.sh file, and notice how it does the \n\"remove_redundant\" part only after it has created the new pack-file and \ndone a \"sync\".\n\n> I don't mean to say that I don't trust git repack to do the right thing.\n> Fundamentally, I just think that I shouldn't depend on it to do the\n> right thing in order to avoid corruption in my repository.\n\nThat's good. However, as the previous failure of git repack showed, to \nsome degree the more likely failure mode is actually that the pack \ngenerated by \"git repack\" is perfectly fine, but it's not _complete_. Say \nwe have a bug in git repack, for example.\n\nAnother case where it's not complete is when you have deleted a branch. \n\"git repack -a -d\" will effectively do a \"git prune\" wrt objects that are \nno longer reachable, and that were in the old packs.\n\nSo I'd actually suggest a slightly different approach. When-ever you \nremove old objects (whether it's \"git prune\" or \"git prune-packed\" or \"git \nrepack -a -d\"), you might want to have an option that doesn't actually \n_remove_ them, but just moves them into \".git/attic\" or something like \nthat.\n\nThen you can clean up the attic after doing your weekly full fsck or \nsomething. And it has the advantage that if somebody has deleted a branch, \nand notices later that maybe he wanted that branch back, you can \"unprune\" \nall the objects, run \"git-fsck-objects --full\" to find any dangling \ncommits, and you'll have all your branches back.\n\nSo in many ways it would perhaps be nicer to have that kind of \"safe \nremove\" option to the pruning commands?\n\n\t\t\tLinus\n"},{"id":"12554","messageId":"4383610D.7080100@citi.umich.edu","threadId":"2095","inReplyTo":"Pine.LNX.4.64.0511212134330.13959@g5.osdl.org","subject":"Re: auto-packing on kernel.org? please?","fromName":"Chuck Lever","fromEmail":"cel@citi.umich.edu","sentAt":"2005-11-22T18:18:53Z","receivedAt":"2005-11-22T18:18:53Z","isPatch":false,"sender":{"key":"cel@citi.umich.edu","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Tue, 22 Nov 2005, Chuck Lever wrote:\n> \n>>there are some things repacking does that breaks StGIT, though.\n>>\n>>git repack -d\n>>\n>>seems to remove old commits that StGIT was still depending on.\n> \n> \n> If that is true, then \"git-fsck-cache\" probably also reports errors on a \n> StGIT repository. No? Basically, it implies that the tool doesn't know how \n> to find all the \"heads\".\n\nindeed.  this is one area where StGIT is \"not safe\" to use with other \nporcelains.  these raw GIT commands can show a bunch of confusing \n\"dangling references\" type errors, or actually modify the index in ways \nthat eliminate StGIT-related commits that aren't currently attached to \nany ancestry.  (i think Catalin mentioned these are related to the \nunapplied patches in a stack, but there could be others; see below).\n\n> The preferred way would be to just list the references somewhere under \n> .git/refs/stgit, in which case fsck and repack should pick them up \n> automatically (so clearly stgit doesn't do that right now ;).\n\nthat could be an extremely large number of commits on a large repository \nwith a lot of patches that have been worked on over a long period.  so \nwhatever mechanism is created to do this needs to scale well in the \nnumber of commits.\n\n> It also implies that doing a \"git prune\" will do horribly bad things to a \n> stgit repo, since it would remove all the objects that it thinks aren't \n> reachable..\n\nyup.  been there, done that.  lucky for me i have an excellent hourly \nbackup scheme.\n\n>>git repack -a -n\n>>\n>>seems to work fine with StGIT,\n> \n> \n> Well, it \"works\", but not \"fine\". Since it doesn't know about the stgit \n> objects, it won't ever pack them.\n\nah!\n\n> But maybe that's what stgit wants (since they are \"temporary\"), but it \n> does mean that if you see a big advantage from packing, you might be \n> losing some of it.\n\nactually, those commits aren't all that \"temporary\".  the \nhistory/revision feature i'm working on would like to maintain all the \ncommits ever done to an StGIT patch.\n\nthe only time you can throw away such commits is when the patch is \ndeleted or when it is finally committed to the repository via \"stg \ncommit\".  otherwise, keeping these commits in a pack would be quite a \ngood thing.\n\nmaybe the first thing to do is to get a basic understanding of an StGIT \ncommit's lifetime.\n\n\nbegin:vcard\nfn:Chuck Lever\nn:Lever;Charles\norg:Network Appliance, Incorporated;Linux NFS Client Development\nadr:535 West William Street, Suite 3100;;Center for Information Technology Integration;Ann Arbor;MI;48103-4943;USA\nemail;internet:cel@citi.umich.edu\ntitle:Member of Technical Staff\ntel;work:+1 734 763-4415\ntel;fax:+1 734 763 4434\ntel;home:+1 734 668-1089\nx-mozilla-html:FALSE\nurl:http://www.monkey.org/~cel/\nversion:2.1\nend:vcard\n\n"},{"id":"12612","messageId":"b0943d9e0511230610x3bdd288ej@mail.gmail.com","threadId":"2095","inReplyTo":"7v1x18eddp.fsf@assigned-by-dhcp.cox.net","subject":"Re: auto-packing on kernel.org? please?","fromName":"Catalin Marinas","fromEmail":"catalin.marinas@gmail.com","sentAt":"2005-11-23T14:10:47Z","receivedAt":"2005-11-23T14:10:47Z","isPatch":false,"sender":{"key":"catalin.marinas@gmail.com","avatar":null},"body":"On 22/11/05, Junio C Hamano <junkio@cox.net> wrote:\n> Catalin Marinas <catalin.marinas@gmail.com> writes:\n>\n> > What I meant is any object whose exact reference is found in\n> > refs/patches (not reachable via refs/patches), even if it is reachable\n> > from refs/heads.\n>\n> do you mean you\n> keep blobs and trees in refs/patches, or \"exactly found in\n> refs/patches\" imply \"commits in refs/patches and trees and blobs\n> reachable from it\"?  If the latter I think it amounts to the\n> same thing.  If some of the blobs are shared with what is\n> reachable from refs/heads or refs/tags I would presume you would\n> want to pack them.\n\nEach patch needs to have 2 commit and 2 tree objects (with the\ncorresponding blobs). I now understand where the problem appears. Most\nof the blobs should actually be packed since they are part of the base\nof the stack.\n\nSince refs/heads files always point to the top of the stack, the\napplied patches (the corresponding objects) would be automatically\npacked. The alternative would be to only pack the objects reachable\nfrom refs/bases but that's really StGIT-specific.\n\nOther algorithm would be to avoid packing objects reachable from\nrefs/patches but not reachable from refs/bases but this would probably\ncomplicate GIT.\n\n> And the \"volatile\" idea may be a good way of doing this.\n> Perhaps \"git repack --volatile <glob>\" to name paths under\n> .git/refs to mark things not to be packed, with a per-repository\n> configuration item to give default 'volatile' patterns?  I could\n> use it when packing my repository to exclude things that are\n> only reachable from \"pu\" branch.\n\nAfter I eventually understood what you meant, the above would still\ninclude the already applied StGIT patches since they are reachable via\nHEAD. Maybe StGIT could avoid modifying refs/heads but I think it\nwould lose some benefits.\n\n--\nCatalin\n"},{"id":"12613","messageId":"b0943d9e0511230618u31d80e57v@mail.gmail.com","threadId":"2095","inReplyTo":"4383610D.7080100@citi.umich.edu","subject":"Re: auto-packing on kernel.org? please?","fromName":"Catalin Marinas","fromEmail":"catalin.marinas@gmail.com","sentAt":"2005-11-23T14:18:13Z","receivedAt":"2005-11-23T14:18:13Z","isPatch":false,"sender":{"key":"catalin.marinas@gmail.com","avatar":null},"body":"On 22/11/05, Chuck Lever <cel@citi.umich.edu> wrote:\n> Linus Torvalds wrote:\n> > But maybe that's what stgit wants (since they are \"temporary\"), but it\n> > does mean that if you see a big advantage from packing, you might be\n> > losing some of it.\n>\n> actually, those commits aren't all that \"temporary\".  the\n> history/revision feature i'm working on would like to maintain all the\n> commits ever done to an StGIT patch.\n\nThat's to avoid pruning them but you might not always want to add them\nto a pack.\n\n> the only time you can throw away such commits is when the patch is\n> deleted or when it is finally committed to the repository via \"stg\n> commit\".  otherwise, keeping these commits in a pack would be quite a\n> good thing.\n>\n> maybe the first thing to do is to get a basic understanding of an StGIT\n> commit's lifetime.\n\nMy initial idea was to throw the old commit away once a patch is\nrefreshed. Even if you want to preserve the history, it would be only\npreserved until you send the patch to be merged upstream and you would\ndelete it locally. If all the patches are meant to be sent upstream at\nsome point, you can avoid packing them.\n\n--\nCatalin\n"}]}