{"thread":{"id":"1192","subject":"[RFC] Design for http-pull on repo with packs","startedAt":"2005-07-10T18:42:45Z","lastAt":"2005-07-12T17:21:40Z","messageCount":9,"participants":["Daniel Barkalow","Dan Holmsand","Junio C Hamano","Tony Luck"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"5902","messageId":"Pine.LNX.4.21.0507101226011.30848-100000@iabervon.org","threadId":"1192","inReplyTo":null,"subject":"[RFC] Design for http-pull on repo with packs","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-07-10T18:42:45Z","receivedAt":"2005-07-10T18:42:45Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"I have a design for using http-pull on a packed repository, and it only\nrequires one extra file in the repository: an append-only list of the pack\nfiles (because getting the directory listing is very painful and\nfailure-prone).\n\nThe first thing to note is that fetch() is allowed to get more than just\nthe requested object. This means that we can get the pack file with the\nrequested object, and this will fulfill the contract of fetch(), and,\nhopefully, be extra-helpful (since we expect the repository owner to have\npacked stuff together usefully). So I do this:\n\n Try to get individual files. So long as this works, everything is as\n  before.\n\n If an individual file is not available, figure out what packs are\n  available:\n\n   Get the list of pack files the repository has\n    (currently, I just use \"e3117bbaf6a59cb53c3f6f0d9b17b9433f0e4135\")\n   For any packs we don't have, get the index files.\n   Keep a list of the struct packed_gits for the packs the server has\n    (these are not used as places to look for objects)\n\n Each time we need an object, check the list for it. If it is in there,\n  download the corresponding pack and report success.\n\nI've nearly got an implementation ready, except for not having a way of\ngetting a list of available packs. It seems to work for getting\ne3117bbaf6a59cb53c3f6f0d9b17b9433f0e4135 when necessary, although I'm\nstill debugging the last few things.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5905","messageId":"42D17D89.9080808@innehallsbolaget.se","threadId":"1192","inReplyTo":"Pine.LNX.4.21.0507101226011.30848-100000@iabervon.org","subject":"Re: [RFC] Design for http-pull on repo with packs","fromName":"Dan Holmsand","fromEmail":"dan@innehallsbolaget.se","sentAt":"2005-07-10T19:56:57Z","receivedAt":"2005-07-10T19:56:57Z","isPatch":false,"sender":{"key":"dan@innehallsbolaget.se","avatar":null},"body":"Daniel Barkalow wrote:\n> I have a design for using http-pull on a packed repository, and it only\n> requires one extra file in the repository: an append-only list of the pack\n> files (because getting the directory listing is very painful and\n> failure-prone).\n\nA few comments (as I've been tinkering with a way to solve the problem \nmyself).\n\nAs long as the pack files are named sensibly (i.e. if they are created \nby git-repack-script), it's not very error-prone to just get the \ndirectory listing, and look for matches for pack-<sha1>.idx. It seems to \nwork quite well (see below). It isn't beautiful in any way, but it works...\n\n[snip]\n\n>  If an individual file is not available, figure out what packs are\n>   available:\n> \n>    Get the list of pack files the repository has\n>     (currently, I just use \"e3117bbaf6a59cb53c3f6f0d9b17b9433f0e4135\")\n>    For any packs we don't have, get the index files.\n\nThis part might be slightly expensive, for large repositories. If one \nassumes that packs are named as by git-repack-script, however, one might \ncache indexes we've already seen (again, see below). Or, if you go for \nthe mandatory \"pack-index-file\", require that it has a reliable order, \nso that you can get the last added index first.\n\n>    Keep a list of the struct packed_gits for the packs the server has\n>     (these are not used as places to look for objects)\n> \n>  Each time we need an object, check the list for it. If it is in there,\n>   download the corresponding pack and report success.\n\nHere you will need some strategy to deal with packs that overlap with \nwhat we've already got. Basically, small and overlapping packs should be \nunpacked, big and non-overlapping ones saved as is (since \ngit-unpack-objects is painfully slow and memory-hungry...).\n\nOne could also optimize the pack-download bit, by figuring out the last \nobject in the pack that we need (easy enough to do from the index file), \n  and just get the part of the pack file leading up to that object. That \ncould be a huge win for independently packed repositories (I don't do \nthat in my code below, though).\n\nAnyway, here's my attempt at the same thing. It introduces \n\"git-dumb-fetch\", with usage like git-fetch-pack (except that it works \nwith http and rsync). And it adds some uglyness to git-cat-file, for \nfiguring out which objects we already have.\n\nI'm sort of using the same basic strategy as you, except that I check \nthe pack files first (I didn't want to mess with http-pull.c, and I \nwanted something that would work with rsync as well).\n\nThe strategy is this:\n\n    o Check if the repository has some pack files we haven't seen\n      already\n\n    o If there are new pack files, download indexes, and see if\n      they contain anything new. If so, download pack file and\n      store or unpack. In either case, note that we have seen the\n      pack file in question (I've used $GIT_DIR/checked_packs).\n\n    o Then\n        o if http: do the git-http-pull stuff, and we're done\n\n        o if rsync: get a list of all object files in the\n          repository, and download the ones we're still missing.\n\nFeel free to take a look, and use anything that might be useful (if \nanything...)\n\nI'm not claiming that this method is better than your way; the only main \ndifferences are the caching of seen index files, and that I download \npacks first.\n\nMy way is faster if the repository contains overlapping object files and \npacks. And doesn't require any new infrastructure.\n\nOn the other hand, my method risks fetching too many objects, if a pack \nfile solely contains stuff from a branch we don't want. And it requires \nthe git-repack-script naming convention to be used on the remote side.\n\n/dan\n\n\ndiff --git a/cat-file.c b/cat-file.c\n--- a/cat-file.c\n+++ b/cat-file.c\n@@ -11,6 +11,42 @@ int main(int argc, char **argv)\n \tchar type[20];\n \tvoid *buf;\n \tunsigned long size;\n+\tint obj_count = 0;\n+\tint missing_count = 0;\n+\tchar line[1000];\n+\n+\tif (argc == 2 && !strcmp(\"--count\", argv[1])) {\n+\t\twhile (fgets(line, sizeof(line), stdin)) {\n+\t\t\tif (get_sha1(line, sha1))\n+\t\t\t\tdie(\"invalid id %s\", line);\n+\t\t\tif (has_sha1_file(sha1))\n+\t\t\t\t++obj_count;\n+\t\t\telse\n+\t\t\t\t++missing_count;\n+\t\t}\n+\t\tprintf(\"%i %i\\n\", obj_count, missing_count);\n+\t\treturn 0;\n+\t}\n+\n+\tif (argc == 2 && !strcmp(\"--existing\", argv[1])) {\n+\t\twhile (fgets(line, sizeof(line), stdin)) {\n+\t\t\tif (get_sha1(line, sha1))\n+\t\t\t\tdie(\"invalid id %s\", line);\n+\t\t\tif (has_sha1_file(sha1))\n+\t\t\t\tprintf (\"%s\", line);\n+\t\t}\n+\t\treturn 0;\n+\t}\n+\n+\tif (argc == 2 && !strcmp(\"--missing\", argv[1])) {\n+\t\twhile (fgets(line, sizeof(line), stdin)) {\n+\t\t\tif (get_sha1(line, sha1))\n+\t\t\t\tdie(\"invalid id %s\", line);\n+\t\t\tif (!has_sha1_file(sha1))\n+\t\t\t\tprintf (\"%s\", line);\n+\t\t}\n+\t\treturn 0;\n+\t}\n \n \tif (argc != 3 || get_sha1(argv[2], sha1))\n \t\tusage(\"git-cat-file [-t | -s | tagname] <sha1>\");\ndiff --git a/git-dumb-fetch b/git-dumb-fetch\nnew file mode 100755\n--- /dev/null\n+++ b/git-dumb-fetch\n@@ -0,0 +1,182 @@\n+#! /bin/sh\n+\n+# git-dumb-fetch pulls objects from (optionally) packed remote\n+# git repositories. \n+\n+. git-sh-setup-script || die \"Not a git archive\"\n+\n+checked_packs=$GIT_DIR/checked_packs\n+\n+usage() {\n+\tdie \"usage: git-dumb-fetch [-w ref] commit-id url\"\n+}\n+\n+http_download() {\n+\ttmpf=$(basename \"$1\")\n+\twget -O \"$tmpd/$tmpf\" \"$1\"\n+}\n+\n+http_cat() {\n+\twget -q -O - \"$1\"\n+}\n+\n+http_list_packs() {\n+\t# XXX: It would be nice to be able to differentiate between failed\n+\t# connections and missing pack dir. For now, assume the latter.\n+\tpindex=$(http_cat \"$1/objects/pack/\") || return 0 \n+\t\t# die \"error getting $1\"\n+\techo \"$pindex\" | \n+\t\tsed -n 's,.*pack-\\([0-9a-f]\\{40\\}\\)\\.idx.*,\\1\\n,gp' |\n+\t\tsed '/^$/d' | sort | uniq\n+}\n+\n+http_pull() {\n+\tgit-http-pull -v -a \"$1\" \"$2/\"\n+}\n+\n+rsync_download() {\n+\trsync \"$1\" \"$tmpd/\" > /dev/null\n+}\n+\n+rsync_cat() {\n+\ttmpf=$(basename \"$1\")\n+\trsync_download \"$1\" && cat \"$tmpd/$tmpf\"\n+}\n+\n+rsync_list_packs() {\n+\t# list every file on the remote side. we'll use that later\n+\techo \"Listing remote objects\" >&2\n+\trsync -zr \"$1/objects/\" > \"$tmpd/files\" &&\n+\tLANG=C sed -n 's,.*pack/pack-\\([0-9a-f]\\{40\\}\\)\\.idx.*,\\1,p' < \\\n+\t\t\"$tmpd/files\" \n+}\n+\n+rsync_pull() {\n+\tLANG=C sed -n 's,.*\\([0-9a-f][0-9a-f]\\)/\\([0-9a-f]\\{38\\}\\).*,\\1\\2,p' \\\n+\t\t< \"$tmpd/files\" | \n+\t\tgit-cat-file --missing > \"$tmpd/missing\" &&\n+\tLANG=C sed 's,^..,\\0/,' < \"$tmpd/missing\" > \"$tmpd/tofetch\" || exit 1\n+\n+\t[ -s \"$tmpd/tofetch\" ] || { echo \"Nothing new to fetch\" >&2; return; }\n+\n+\tif rsync --help 2>&1 | grep -q files-from; then\n+\t\trsync -avz --ignore-existing --whole-file \\\n+\t\t\t--files-from=\"$tmpd/tofetch\" \\\n+\t\t\t\"$2/objects/\" \"$GIT_OBJECT_DIRECTORY/\" >&2\n+\telse\n+\t\tLANG=C sed -n \\\n+\t\t\t's,.*\\([0-9a-f][0-9a-f]\\)/\\([0-9a-f]\\{38\\}\\).*,\\1\\2,p' \\\n+\t\t\t< \"$tmpd/files\" | \n+\t\t\tgit-cat-file --existing > \"$tmpd/got\" && \n+\t\tLANG=C sed 's,^..,\\0/,' < \"$tmpd/got\" > \"$tmpd/excl\" || exit 1\n+\t\tif [ -f \"$checked_packs\" ]; then\n+\t\t\tsed 's,^.*,pack/pack-&.idx,' < \"$checked_packs\"\n+\t\t\tsed 's,^.*,pack/pack-&.pack,' < \"$checked_packs\"\n+\t\tfi >> \"$tmpd/excl\"\n+\t\trsync -avz --ignore-existing --whole-file \\\n+\t\t\t--exclude-from=\"$tmpd/excl\" \\\n+\t\t\t\"$2/objects/\" \"$GIT_OBJECT_DIRECTORY/\" >&2\n+\tfi\n+}\n+\n+idx_policy() {\n+\t# existing=$1 missing=$2\n+\tif [ $1 -eq 0 -a $2 -eq 0 ]; then\n+\t\techo empty\n+\telif [ $2 -eq 0 ]; then\n+\t\techo all\n+\telif [ $1 -eq 0 ]; then\n+\t\techo none\n+\telse\n+\t\tif [ $2 -gt 5000 -a $1 -lt $2 ]; then\n+\t\t\t# It's a really big pack. Don't unpack\n+\t\t\techo all\n+\t\telse\n+\t\t\techo partial\n+\t\tfi\n+\tfi\n+}\n+\n+check_idx() {\n+\tcounts=$(git-show-index | cut -d' ' -f2 | git-cat-file --count) || \n+\t\texit 1\n+\tidx_policy $counts\n+}\n+\n+has_pack() {\n+\t[ -f \"$GIT_OBJECT_DIRECTORY/pack/pack-$1.idx\" -a \\\n+\t\t\"$GIT_OBJECT_DIRECTORY/pack/pack-$1.pack\" ] && return 0\n+\t[ -f \"$checked_packs\" ] && grep -q $1 < \"$checked_packs\"\n+}\n+\n+fetch_packs() {\n+\tidx=$($list_packs \"$1\") || exit 1\n+\t[ \"$idx\" ] || return 0\n+\techo \"Examining remote packs: $idx\" >&2\n+\tfor i in $idx; do\n+\t\thas_pack $i && continue\n+\t\techo \"Downloading pack $i\" >&2\n+\t\t$download \"$1/objects/pack/pack-$i.idx\" &&\n+\t\tgotit=$(check_idx < \"$tmpd/pack-$i.idx\") || exit 1\n+\n+\t\tcase $gotit in \n+\t\t\tpartial | none)\n+\t\t\t$download \"$1/objects/pack/pack-$i.pack\" &&\n+\t\t\tgit-verify-pack \"$tmpd/pack-$i\" || die \"invalid pack\" ;;\n+\t\t\t*)\n+\t\t\techo \"Already got all objects in pack $i\" >&2 ;;\n+\t\tesac\n+\n+\t\tcase $gotit in \n+\t\tpartial)\n+\t\t\tgit-unpack-objects < \"$tmpd/pack-$i.pack\" || exit 1 ;;\n+\t\tnone)\n+\t\t\tmv \"$tmpd/pack-$i.idx\" \"$tmpd/pack-$i.pack\" \\\n+\t\t\t\t\"$GIT_OBJECT_DIRECTORY/pack/\" 2>/dev/null ||\n+\t\t\tcp \"$tmpd/pack-$i.idx\" \"$tmpd/pack-$i.pack\" \\\n+\t\t\t\t\"$GIT_OBJECT_DIRECTORY/pack/\" || exit 1 ;;\n+\t\tesac\n+\t\techo $i >> \"$checked_packs\"\n+\tdone\n+}\n+\n+\n+while true; do\n+\tcase $1 in\n+\t\t--) shift; break ;;\n+\t\t-*) die \"unknown option: $1\" ;;\n+\t\t*) break ;;\n+\tesac\n+\tshift\n+done\n+\n+url=$1 srchead=$2 \n+[ -n \"$srchead\" -a -n \"$url\" ] || usage\n+\n+case $url in\n+\thttp://*) proto=http ;;\n+\trsync://*) proto=rsync ;;\n+\t*) die \"don't know how to fetch from $url\" ;;\n+esac\n+\n+download=${proto}_download\n+cat=${proto}_cat\n+list_packs=${proto}_list_packs\n+pull=${proto}_pull\n+\n+tmpd=$(mktemp -d \"${TMPDIR:-/tmp}/dumbfetch.XXXXXX\") || exit 1\n+trap \"rm -rf '$tmpd'\" 0 1 2 3 15\n+\n+echo \"Fetching from $url\" >&2\n+remoteid=$($cat \"$url/refs/$srchead\") || die \"error reading $srchead\"\n+\n+if [ \"$previd\" = \"$remoteid\" ]; then\n+\techo \"Up to date\" >&2\n+\texit 0\n+fi\n+\n+fetch_packs \"$url\" &&\n+$pull \"$remoteid\" \"$url\" || die \"fetch failed\"\n+\n+echo $remoteid\n+\n"},{"id":"5909","messageId":"Pine.LNX.4.21.0507101557510.30848-100000@iabervon.org","threadId":"1192","inReplyTo":"42D17D89.9080808@innehallsbolaget.se","subject":"Re: [RFC] Design for http-pull on repo with packs","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-07-10T20:29:06Z","receivedAt":"2005-07-10T20:29:06Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 10 Jul 2005, Dan Holmsand wrote:\n\n> Daniel Barkalow wrote:\n> > I have a design for using http-pull on a packed repository, and it only\n> > requires one extra file in the repository: an append-only list of the pack\n> > files (because getting the directory listing is very painful and\n> > failure-prone).\n> \n> A few comments (as I've been tinkering with a way to solve the problem \n> myself).\n> \n> As long as the pack files are named sensibly (i.e. if they are created \n> by git-repack-script), it's not very error-prone to just get the \n> directory listing, and look for matches for pack-<sha1>.idx. It seems to \n> work quite well (see below). It isn't beautiful in any way, but it works...\n\nI may grab your code for that; the version I just sent seems to be working\nexcept for that.\n\n> >  If an individual file is not available, figure out what packs are\n> >   available:\n> > \n> >    Get the list of pack files the repository has\n> >     (currently, I just use \"e3117bbaf6a59cb53c3f6f0d9b17b9433f0e4135\")\n> >    For any packs we don't have, get the index files.\n> \n> This part might be slightly expensive, for large repositories. If one \n> assumes that packs are named as by git-repack-script, however, one might \n> cache indexes we've already seen (again, see below). Or, if you go for \n> the mandatory \"pack-index-file\", require that it has a reliable order, \n> so that you can get the last added index first.\n\nNothing bad happens if you have index files for pack files you don't have,\nas it turns out; the library ignores them. So we can keep the index files\naround so we can quickly check if they have the objects we want. That way,\nwe don't have to worry about skipping something now (because it's not\nneeded) and then ignoring it when the branch gets merged in.\n\nSo what I actually do is make a list of the pack files that aren't already\ndownloaded that are available from the server, and download the index\nfiles for any where the index file isn't downloaded, either.\n\n> >    Keep a list of the struct packed_gits for the packs the server has\n> >     (these are not used as places to look for objects)\n> > \n> >  Each time we need an object, check the list for it. If it is in there,\n> >   download the corresponding pack and report success.\n> \n> Here you will need some strategy to deal with packs that overlap with \n> what we've already got. Basically, small and overlapping packs should be \n> unpacked, big and non-overlapping ones saved as is (since \n> git-unpack-objects is painfully slow and memory-hungry...).\n\nI don't think there's an issue to having overlapping packs, either with\neach other or with separate objects. If the user wants, stuff can be\nrepacked outside of the pull operation (note, though, that the index files\nshould be truncated rather than removed, so that the program doesn't fetch\nthem again next time some object can't be found easily).\n\n> One could also optimize the pack-download bit, by figuring out the last \n> object in the pack that we need (easy enough to do from the index file), \n>   and just get the part of the pack file leading up to that object. That \n> could be a huge win for independently packed repositories (I don't do \n> that in my code below, though).\n\nThat's only possible if you can figure out what you want to have before\nyou get it. My code is walking the reachability graph on the client; it\ncan only figure out what other objects it needs after it's mapped the pack\nfile.\n\n> Anyway, here's my attempt at the same thing. It introduces \n> \"git-dumb-fetch\", with usage like git-fetch-pack (except that it works \n> with http and rsync). And it adds some uglyness to git-cat-file, for \n> figuring out which objects we already have.\n\nI might use that method for listing the available packs, although I'd sort\nof like to encourage a clean solution first.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"5913","messageId":"42D1957F.1050609@gmail.com","threadId":"1192","inReplyTo":"Pine.LNX.4.21.0507101557510.30848-100000@iabervon.org","subject":"Re: [RFC] Design for http-pull on repo with packs","fromName":"Dan Holmsand","fromEmail":"holmsand@gmail.com","sentAt":"2005-07-10T21:39:11Z","receivedAt":"2005-07-10T21:39:11Z","isPatch":false,"sender":{"key":"holmsand@gmail.com","avatar":"https://gravatar.com/avatar/5c722084bafd85e754a02efad01fe69107eb6f393253c49232c5c9f7faa974df?d=mp&s=160"},"body":"Daniel Barkalow wrote:\n> On Sun, 10 Jul 2005, Dan Holmsand wrote:\n>>Daniel Barkalow wrote:\n>>> If an individual file is not available, figure out what packs are\n>>>  available:\n>>>\n>>>   Get the list of pack files the repository has\n>>>    (currently, I just use \"e3117bbaf6a59cb53c3f6f0d9b17b9433f0e4135\")\n>>>   For any packs we don't have, get the index files.\n>>\n>>This part might be slightly expensive, for large repositories. If one \n>>assumes that packs are named as by git-repack-script, however, one might \n>>cache indexes we've already seen (again, see below). Or, if you go for \n>>the mandatory \"pack-index-file\", require that it has a reliable order, \n>>so that you can get the last added index first.\n> \n> \n> Nothing bad happens if you have index files for pack files you don't have,\n> as it turns out; the library ignores them. So we can keep the index files\n> around so we can quickly check if they have the objects we want. That way,\n> we don't have to worry about skipping something now (because it's not\n> needed) and then ignoring it when the branch gets merged in.\n> \n> So what I actually do is make a list of the pack files that aren't already\n> downloaded that are available from the server, and download the index\n> files for any where the index file isn't downloaded, either.\n\nAah. In other words, you do the caching thing as well. It seems a little \nugly, though, to store the index-only index files with the rest of the \npack. It might be preferable to introduce something like \n$GIT_DIR/index-cache or something, so than it can be easily cleaned (and \ndon't follow us around forever when \ncloning-by-hardlinking-the-entire-object-directory).\n\nYou might end up with quite a large number of index files, after a while \nthough, if you pull from several repositories that are regularly repacked.\n\n>>>   Keep a list of the struct packed_gits for the packs the server has\n>>>    (these are not used as places to look for objects)\n>>>\n>>> Each time we need an object, check the list for it. If it is in there,\n>>>  download the corresponding pack and report success.\n>>\n>>Here you will need some strategy to deal with packs that overlap with \n>>what we've already got. Basically, small and overlapping packs should be \n>>unpacked, big and non-overlapping ones saved as is (since \n>>git-unpack-objects is painfully slow and memory-hungry...).\n> \n> \n> I don't think there's an issue to having overlapping packs, either with\n> each other or with separate objects. If the user wants, stuff can be\n> repacked outside of the pull operation (note, though, that the index files\n> should be truncated rather than removed, so that the program doesn't fetch\n> them again next time some object can't be found easily).\n\nWell, the only issue is obviously waste of space. If you fetch a lot of \nbranches from independently packed repos, it might mean a lot of waste, \nthough.\n\nAbout truncating index files: this seems a bit ugly. You get a file that \ndoesn't contain what it says it contains, which may cause trouble if for \nexample the git prune thing is used.\n\nYou might be better off with a simple list of index files we know we \nhave all the objects of (and make sure that git-prune-script deletes \nthis file, since it possibly breaks the contract).\n\n>>One could also optimize the pack-download bit, by figuring out the last \n>>object in the pack that we need (easy enough to do from the index file), \n>>  and just get the part of the pack file leading up to that object. That \n>>could be a huge win for independently packed repositories (I don't do \n>>that in my code below, though).\n> \n> \n> That's only possible if you can figure out what you want to have before\n> you get it. My code is walking the reachability graph on the client; it\n> can only figure out what other objects it needs after it's mapped the pack\n> file.\n\nNo, but we can find out which objects we *don't* want (i.e. the ones we \nhave). And that may be a lot, e.g. if a repository is fully repacked, or \nif we track branches on several similar but independently packed \nrepositories. And as far as I understand git-pack-objects, it tries to \nput recent objects in the front.\n\nI don't have any numbers to back this up with, though. Some testing may \nbe needed, but since the population of packed public repositories is 1, \nthis is tricky...\n\n> I might use that method for listing the available packs, although I'd sort\n> of like to encourage a clean solution first.\n\nEncouraging cleanliness is obviously a good thing :-)\n\n/dan\n"},{"id":"5926","messageId":"7v4qb2ni73.fsf@assigned-by-dhcp.cox.net","threadId":"1192","inReplyTo":"42D17D89.9080808@innehallsbolaget.se","subject":"Re: [RFC] Design for http-pull on repo with packs","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-07-11T03:18:56Z","receivedAt":"2005-07-11T03:18:56Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"One very minor problem I have with Holmsand approach [*1*] is\nthat the original Barkalow puller allowed a really dumb http\nserver by not requiring directory index at all.  For somebody\nlike me with a cheap ISP account [*2*], it was great that I did\nnot have to update 256 index.html files for objects/??/\ndirectories.  Admittedly, it would be just one directory\nobject/pack/, but still...\n\nOn the other hand, picking an optimum set of packs from\noverlapping set of packs is indeed a very interesting (and hard\ncombinatorial) problem to solve.  I am hoping that in practice\npeople would not force clients to do it with \"interesting\" set\nof packs.  I would hope them to have just a full pack and\nincrementals, never having ovelaps, like Linus plans to do on\nhis kernel repo.\n\nOn the other hand, for somebody like Jeff Garzik with 50 heads,\nit might make some sense to have a handful different overlapping\npacks, optimized for different sets of people wanting to pull\nsome but not all of his heads.\n\nHaving said that, even if we want to support such a repository,\nwe should remember that the server side optimization needs to be\ndone only once per push to support many pulls by different\ndownstream clients.  Maybe preparing more than \"list of pack\nfile names\" to help clients decide which packs to pull is\ndesirable anyway.  Say, \"here are the list of packs.  If you want\nto sync with this and that head, I would suggest starting by\ngetting this pack.\"\n\n\n[Footnotes]\n\n*1* I was about to type Dan's, but both of you are ;-).\n\n*2* Not having a public, rsync-reachable repository gave me a\nlot of incentive to think about issues to support small/cheap\nprojects well ;-).\n"},{"id":"5949","messageId":"42D2960E.3050008@gmail.com","threadId":"1192","inReplyTo":"7v4qb2ni73.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] Design for http-pull on repo with packs","fromName":"Dan Holmsand","fromEmail":"holmsand@gmail.com","sentAt":"2005-07-11T15:53:50Z","receivedAt":"2005-07-11T15:53:50Z","isPatch":false,"sender":{"key":"holmsand@gmail.com","avatar":"https://gravatar.com/avatar/5c722084bafd85e754a02efad01fe69107eb6f393253c49232c5c9f7faa974df?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> One very minor problem I have with Holmsand approach [*1*] is\n> that the original Barkalow puller allowed a really dumb http\n> server by not requiring directory index at all.  For somebody\n> like me with a cheap ISP account [*2*], it was great that I did\n> not have to update 256 index.html files for objects/??/\n> directories.  Admittedly, it would be just one directory\n> object/pack/, but still...\n\nI totally agree that you shouldn't have to do any special kind of \nprepping to serve a repository thru http. Which was why I thought it was \na good thing to use the default directory listing of the web-server, \nassuming that this feature would be available on most servers... \nApparently not yours, though :-(\n\nAnd Cogito already relies on directory listings (to find tags to download).\n\nBut if git-repack-script generates a \"pack index file\" automagically, \nthen of course everything is fine.\n\n> On the other hand, picking an optimum set of packs from\n> overlapping set of packs is indeed a very interesting (and hard\n> combinatorial) problem to solve.  I am hoping that in practice\n> people would not force clients to do it with \"interesting\" set\n> of packs.  I would hope them to have just a full pack and\n> incrementals, never having ovelaps, like Linus plans to do on\n> his kernel repo.\n> \n> On the other hand, for somebody like Jeff Garzik with 50 heads,\n> it might make some sense to have a handful different overlapping\n> packs, optimized for different sets of people wanting to pull\n> some but not all of his heads.\n\nWell, it is an interresting problem... But I don't think that the \nsolution is to create more pack files. In fact, you'd want as few pack \nfiles as possible, for maximum overall efficiency.\n\nI did a little experiment. I cloned Linus' current tree, and git \nrepacked everything (that's 63M + 3.3M worth of pack files). Then I got \nsomething like 25 or so of Jeff's branches. That's 6.9M of object files, \nand 1.4M packed. Total size: 70M for the entire .git/objects/pack directory.\n\nRepacking all of that to a single pack file gives, somewhat \nsurprisingly, a pack size of 62M (+ 1.3M index). In other words, the \ncost of getting all those branches, and all of the new stuff from Linus, \nturns out to be *negative* (probably due to some strange deltification \ncoincidence).\n\nI think that this shows that (at least in this case), having many \nbranches isn't particularly wasteful (1.4M in this case with one \nincremental pack).\n\nAnd that fewer packs beats many packs quite handily.\n\nThe big problem, however, comes when Jeff (or anyone else) decides to \nrepack. Then, if you fetch both his repo and Linus', you might end up \nwith several really big pack files, that mostly overlap. That could \neasily mean storing most objects many times, if you don't do some smart \nselective un/repacking when fetching.\n\n/dan\n"},{"id":"5951","messageId":"12c511ca050711100840946891@mail.gmail.com","threadId":"1192","inReplyTo":"42D2960E.3050008@gmail.com","subject":"Re: [RFC] Design for http-pull on repo with packs","fromName":"Tony Luck","fromEmail":"tony.luck@gmail.com","sentAt":"2005-07-11T17:08:15Z","receivedAt":"2005-07-11T17:08:15Z","isPatch":false,"sender":{"key":"tony.luck@gmail.com","avatar":null},"body":"> The big problem, however, comes when Jeff (or anyone else) decides to\n> repack. Then, if you fetch both his repo and Linus', you might end up\n> with several really big pack files, that mostly overlap. That could\n> easily mean storing most objects many times, if you don't do some smart\n> selective un/repacking when fetching.\n\nSo although it is possible to pack and re-pack at any time, perhaps we\nneed some guidelines?  Maybe Linus should just do a re-pack as each\n2.6.x release is made (or perhaps just every 2.6.even release if that is\ntoo often).  It has already been noted offlist that repositories hosted on\nkernel.org can just copy pack files from Linus (or even better hardlink them).\n\n-Tony\n"},{"id":"5979","messageId":"7vu0j0ncnr.fsf@assigned-by-dhcp.cox.net","threadId":"1192","inReplyTo":"42D2960E.3050008@gmail.com","subject":"Re: [RFC] Design for http-pull on repo with packs","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-07-11T23:30:48Z","receivedAt":"2005-07-11T23:30:48Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Dan Holmsand <holmsand@gmail.com> writes:\n\n> I did a little experiment. I cloned Linus' current tree, and git\n> repacked everything (that's 63M + 3.3M worth of pack files). Then I\n> got something like 25 or so of Jeff's branches. That's 6.9M of object\n> files, and 1.4M packed. Total size: 70M for the entire\n> .git/objects/pack directory.\n>\n> Repacking all of that to a single pack file gives, somewhat\n> surprisingly, a pack size of 62M (+ 1.3M index). In other words, the\n> cost of getting all those branches, and all of the new stuff from\n> Linus, turns out to be *negative* (probably due to some strange\n> deltification coincidence).\n\nWe do _not_ want to optimize for initial slurps into empty\nrepositories.  Quite the opposite.  We want to optimize for\nallowing quick updates of reasonably up-to-date developer repos.\nIf initial slurps are _also_ efficient then that is an added\nbonus; that is something the baseline big pack (60M Linus pack)\nwould give us already.  So repacking everything into a single\npack nightly is _not_ what we want to do, even though that would\ngive the maximum compression ;-).  I know you understand this,\nbut just stating the second of the above paragraphs would give\ncasual readers a wrong impression.\n\n> I think that this shows that (at least in this case), having many\n> branches isn't particularly wasteful (1.4M in this case with one\n> incremental pack).\n\n> And that fewer packs beats many packs quite handily.\n\nYou are correct.  For somebody like Jeff, having the Linus\nbaseline pack with one pack of all of his head (incremental that\nexcludes what is already in the Linus baseline pack) would help\npullers.\n\n> The big problem, however, comes when Jeff (or anyone else) decides to\n> repack. Then, if you fetch both his repo and Linus', you might end up\n> with several really big pack files, that mostly overlap. That could\n> easily mean storing most objects many times, if you don't do some\n> smart selective un/repacking when fetching.\n\nIndeed.  Overlapping packs is a possibility, but my gut feeling\nis that it would not be too bad, if things are arranged so that\npacks are expanded-and-then-repacked _very_ rarely if ever.\nInstead, at least for your public repository, if you only repack\nincrementally I think you would be OK.\n"},{"id":"6045","messageId":"42D3FC24.7010509@gmail.com","threadId":"1192","inReplyTo":"7vu0j0ncnr.fsf@assigned-by-dhcp.cox.net","subject":"Re: [RFC] Design for http-pull on repo with packs","fromName":"Dan Holmsand","fromEmail":"holmsand@gmail.com","sentAt":"2005-07-12T17:21:40Z","receivedAt":"2005-07-12T17:21:40Z","isPatch":false,"sender":{"key":"holmsand@gmail.com","avatar":"https://gravatar.com/avatar/5c722084bafd85e754a02efad01fe69107eb6f393253c49232c5c9f7faa974df?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> Dan Holmsand <holmsand@gmail.com> writes:\n>>Repacking all of that to a single pack file gives, somewhat\n>>surprisingly, a pack size of 62M (+ 1.3M index). In other words, the\n>>cost of getting all those branches, and all of the new stuff from\n>>Linus, turns out to be *negative* (probably due to some strange\n>>deltification coincidence).\n> \n> \n> We do _not_ want to optimize for initial slurps into empty\n> repositories.  Quite the opposite.  We want to optimize for\n> allowing quick updates of reasonably up-to-date developer repos.\n> If initial slurps are _also_ efficient then that is an added\n> bonus; that is something the baseline big pack (60M Linus pack)\n> would give us already.  So repacking everything into a single\n> pack nightly is _not_ what we want to do, even though that would\n> give the maximum compression ;-).  I know you understand this,\n> but just stating the second of the above paragraphs would give\n> casual readers a wrong impression.\n\nI agree, to a point: I think the bonus is quite nice to have... As it \nis, it's actually faster on my machine to clone a fresh tree of Linus' \nthan it is to \"git clone\" a local tree (without doing the hardlinking \n\"cheating\", that is). And it's kind of nice to have the option to start \ncompletely fresh.\n\nAnyway, my point is this: to make pulling efficient, we should ideally \nhave (1) as few object files to pull as possible, especially when using \nhttp, and (2) have as few packs as possible, to gain some compression \nfor those who pull more seldom. Point 1 is obviously the most important one.\n\nTo make this happen, relatively frequent repacking and re-repacking \n(even if only on parts of the repository) would be necessary. Or at \nleast nice to have...\n\nWhich was why I wanted the \"dumb fetch\" thingies to at least do some \n\"relatively smart un/repacking\" to avoid duplication. And, ideally, that \nthey would avoid downloading entire packs that we just want the \nbeginning of. That would lessen the cost of repacking, which I happen to \nthink is a good thing.\n\nAlso, it's kind of strange when the ssh/local fetching *always* unpacks \neverything, and rsync/http *never* does this...\n\n> You are correct.  For somebody like Jeff, having the Linus\n> baseline pack with one pack of all of his head (incremental that\n> excludes what is already in the Linus baseline pack) would help\n> pullers.\n\nThat would work, of course. It, however, means that Linus becomes the \n\"official repository maintainer\" in a way that doesn't feel very \ndistributed. Perhaps then Linus' packs should be marked \"official\" in \nsome way?\n\n>>The big problem, however, comes when Jeff (or anyone else) decides to\n>>repack. Then, if you fetch both his repo and Linus', you might end up\n>>with several really big pack files, that mostly overlap. That could\n>>easily mean storing most objects many times, if you don't do some\n>>smart selective un/repacking when fetching.\n> \n> \n> Indeed.  Overlapping packs is a possibility, but my gut feeling\n> is that it would not be too bad, if things are arranged so that\n> packs are expanded-and-then-repacked _very_ rarely if ever.\n> Instead, at least for your public repository, if you only repack\n> incrementally I think you would be OK.\n\nTo be exact, you're ok (in the meaning of avoiding duplicates) as long \nas you always rsync in the \"official packs\", and coordinate with others \nyou're merging with, before you do any repacking of your own. Sure, this \nworks. It just feels a bit \"un-distributed\" for my personal taste...\n\n/dan\n"}]}