{"thread":{"id":"3780","subject":"Cygwin can't handle huge packfiles?","startedAt":"2006-04-03T09:46:13Z","lastAt":"2006-04-07T18:46:55Z","messageCount":18,"participants":["Kees-Jan Dijkzeul","Johannes Schindelin","Morten Welinder","Linus Torvalds","Alex Riesen","Christopher Faylor","Rutger Nijlunsing","Junio C Hamano","Jakub Narebski","Nicolas Pitre"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"18274","messageId":"fa0b6e200604030246q21fccb9ar93004ac67d8b28b3@mail.gmail.com","threadId":"3780","inReplyTo":null,"subject":"Cygwin can't handle huge packfiles?","fromName":"Kees-Jan Dijkzeul","fromEmail":"k.j.dijkzeul@gmail.com","sentAt":"2006-04-03T09:46:13Z","receivedAt":"2006-04-03T09:46:13Z","isPatch":false,"sender":{"key":"k.j.dijkzeul@gmail.com","avatar":"https://gravatar.com/avatar/7093ca9ddb281f666728c9863ed26183dc72ff01d7fd509b053b7e089a93ffcb?d=mp&s=160"},"body":"Hi,\n\nI'm trying to get Git to manage a 5Gb source tree. Under linux, this\nworks like a charm. Under cygwin, however, I run in to difficulties.\nFor example:\n\n$ git-clone sgp-wa/ sgp-wa.clone\nfatal: packfile\n./objects/pack/pack-56aa013a0234e198467ed37ae5db925764a6ee98.pack\ncannot be mapped.\nfatal: unexpected EOF\nfetch-pack from '/cygdrive/e/Projects/sgp-wa/.git' failed.\n\nTo figure out what is happening, I printed the value of errno, which\nturns out to be 12 (Cannot allocate memory). I'm not sure how mmap is\nimplemented in cygwin, but if they allocate memory and load the file\ninto it, then this error is not surprising, as the pack file in\nquestion is 1.5Gb in size.\n\nI'm not sure how to approach this problem. Any tips would be greatly\nappreciated.\n\nThanks a lot!\n\nKees-Jan\n"},{"id":"18276","messageId":"Pine.LNX.4.63.0604031521170.4011@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3780","inReplyTo":"fa0b6e200604030246q21fccb9ar93004ac67d8b28b3@mail.gmail.com","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-04-03T13:23:55Z","receivedAt":"2006-04-03T13:23:55Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 3 Apr 2006, Kees-Jan Dijkzeul wrote:\n\n> I'm trying to get Git to manage a 5Gb source tree. Under linux, this\n> works like a charm. Under cygwin, however, I run in to difficulties.\n> For example:\n> \n> $ git-clone sgp-wa/ sgp-wa.clone\n> fatal: packfile\n> ./objects/pack/pack-56aa013a0234e198467ed37ae5db925764a6ee98.pack\n> cannot be mapped.\n> fatal: unexpected EOF\n> fetch-pack from '/cygdrive/e/Projects/sgp-wa/.git' failed.\n> \n> To figure out what is happening, I printed the value of errno, which\n> turns out to be 12 (Cannot allocate memory). I'm not sure how mmap is\n> implemented in cygwin, but if they allocate memory and load the file\n> into it, then this error is not surprising, as the pack file in\n> question is 1.5Gb in size.\n\nThe problem is not mmap() on cygwin, but that a fork() has to jump through \nloops to reinstall the open file descriptors on cygwin. If the \ncorresponding file was deleted, that fails. Therefore, we work around that \non cygwin by actually reading the file into memory, *not* mmap()ing it.\n\nHth,\nDscho\n"},{"id":"18281","messageId":"118833cc0604030726r44b0682etec3349f62986e3c0@mail.gmail.com","threadId":"3780","inReplyTo":"Pine.LNX.4.63.0604031521170.4011@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2006-04-03T14:26:27Z","receivedAt":"2006-04-03T14:26:27Z","isPatch":false,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"> The problem is not mmap() on cygwin, but that a fork() has to jump through\n> loops to reinstall the open file descriptors on cygwin. If the\n> corresponding file was deleted, that fails. Therefore, we work around that\n> on cygwin by actually reading the file into memory, *not* mmap()ing it.\n\nMaybe, but you aren't going to be able to handler much bigger packs\neven on *nix.  Unless you go 64-bit, that is.\n\nM.\n"},{"id":"18283","messageId":"Pine.LNX.4.64.0604030730040.3781@g5.osdl.org","threadId":"3780","inReplyTo":"Pine.LNX.4.63.0604031521170.4011@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-03T14:33:51Z","receivedAt":"2006-04-03T14:33:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Apr 2006, Johannes Schindelin wrote:\n> \n> The problem is not mmap() on cygwin, but that a fork() has to jump through \n> loops to reinstall the open file descriptors on cygwin. If the \n> corresponding file was deleted, that fails. Therefore, we work around that \n> on cygwin by actually reading the file into memory, *not* mmap()ing it.\n\nWell, we could actually do a _real_ mmap on pack-files. The pack-files are \nmuch better mmap'ed - there we don't _want_ them to be removed while we're \nusing them. It was the index file etc that was problematic.\n\nMaybe the cygwin fake mmap should be triggered only for the index (and \npossibly the individual objects - if only because there doing a \nmalloc+read may actually be faster).\n\nUsing malloc+read on pack-files is pretty wasteful, since we usually only \nuse a very small part of them (ie if we have a 1.5GB pack-file, it's sad \nto read all of it, when we'd usually actually access just a small small \nfraction of it).\n\nThat said, I think git _does_ have problems with large pack-files. We have \nsome 32-bit issues etc, and just virtual address space things. So for now, \nit's probably best to limit pack-files to the few-hundred-meg size, and \ncreate serveral smaller ones rather than one huge one.\n\n\t\tLinus\n"},{"id":"18284","messageId":"Pine.LNX.4.64.0604030734440.3781@g5.osdl.org","threadId":"3780","inReplyTo":"Pine.LNX.4.64.0604030730040.3781@g5.osdl.org","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-04-03T14:36:53Z","receivedAt":"2006-04-03T14:36:53Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Apr 2006, Linus Torvalds wrote:\n> \n> That said, I think git _does_ have problems with large pack-files. We have \n> some 32-bit issues etc\n\nI should clarify that. git _itself_ shouldn't have any 32-bit issues, but \nthe packfile data structure does. The index has 32-bit offsets into \nindividual pack-files. \n\nThat's not hugely fundamental, but I didn't expect people to hit it this \nquickly. What kind of project has a 1.5GB pack-file _already_? I hope it's \nfifteen years of history (so that we'll have another fifteen years before \nwe'll have to worry about 4GB pack-files ;)\n\n\t\t\tLinus\n"},{"id":"18286","messageId":"81b0412b0604030738w3f61f34u8a56b1a6b0b5ef88@mail.gmail.com","threadId":"3780","inReplyTo":"fa0b6e200604030246q21fccb9ar93004ac67d8b28b3@mail.gmail.com","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-04-03T14:38:32Z","receivedAt":"2006-04-03T14:38:32Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 4/3/06, Kees-Jan Dijkzeul <k.j.dijkzeul@gmail.com> wrote:\n> I'm trying to get Git to manage a 5Gb source tree. Under linux, this\n> works like a charm. Under cygwin, however, I run in to difficulties.\n> For example:\n>\n> $ git-clone sgp-wa/ sgp-wa.clone\n> fatal: packfile\n> ./objects/pack/pack-56aa013a0234e198467ed37ae5db925764a6ee98.pack\n> cannot be mapped.\n> fatal: unexpected EOF\n> fetch-pack from '/cygdrive/e/Projects/sgp-wa/.git' failed.\n>\n> To figure out what is happening, I printed the value of errno, which\n> turns out to be 12 (Cannot allocate memory). I'm not sure how mmap is\n\nmmap in git on cygwin does not mmaps anything,\nbut just reads the whole file in memory.\n\n> I'm not sure how to approach this problem. Any tips would be greatly\n> appreciated.\n\nI ended up hacking gitfakemmap like in the attached patches (sorry for mime).\nIt's very ugly and unsafe hack, and it's actually exactly the reason why it was\nnever submitted. Still, it helps me (it speedups revlist, for\ninstance), and maybe\nit'll help you.\nIt is a really good example what stupid windows restrictions can do to\na program.\n\nThe patch is against git as of 3-Apr-2005, ~10 CET\n\n\ndiff --git a/Makefile b/Makefile\nindex c79d646..8a46436\n--- a/Makefile\n+++ b/Makefile\n@@ -389,7 +389,7 @@ ifdef NO_SETENV\n endif\n ifdef NO_MMAP\n \tCOMPAT_CFLAGS += -DNO_MMAP\n-\tCOMPAT_OBJS += compat/mmap.o\n+\tCOMPAT_OBJS += compat/mmap.o compat/realmmap.o\n endif\n ifdef NO_IPV6\n \tALL_CFLAGS += -DNO_IPV6\ndiff --git a/compat/realmmap.c b/compat/realmmap.c\nnew file mode 100644\nindex 0000000..8f26641\n--- /dev/null\n+++ b/compat/realmmap.c\n@@ -0,0 +1,26 @@\n+#include <stdio.h>\n+#include <stdlib.h>\n+#include <unistd.h>\n+#include <errno.h>\n+#include <sys/mman.h>\n+#include \"../git-compat-util.h\"\n+\n+#undef mmap\n+#undef munmap\n+\n+void *realmmap(void *start, size_t length, int prot , int flags, int fd, off_t offset)\n+{\n+\tif (start != NULL || !(flags & MAP_PRIVATE)) {\n+\t\terrno = ENOTSUP;\n+\t\treturn MAP_FAILED;\n+\t}\n+\tstart = mmap(start, length, prot, flags, fd, offset);\n+\treturn start;\n+}\n+\n+int realmunmap(void *start, size_t length)\n+{\n+\treturn munmap(start, length);\n+}\n+\n+\ndiff --git a/diff.c b/diff.c\nindex e496905..f1a2cf0 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -450,7 +450,7 @@ int diff_populate_filespec(struct diff_f\n \t\tfd = open(s->path, O_RDONLY);\n \t\tif (fd < 0)\n \t\t\tgoto err_empty;\n-\t\ts->data = mmap(NULL, s->size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\t\ts->data = realmmap(NULL, s->size, PROT_READ, MAP_PRIVATE, fd, 0);\n \t\tclose(fd);\n \t\tif (s->data == MAP_FAILED)\n \t\t\tgoto err_empty;\n@@ -482,7 +482,7 @@ void diff_free_filespec_data(struct diff\n \tif (s->should_free)\n \t\tfree(s->data);\n \telse if (s->should_munmap)\n-\t\tmunmap(s->data, s->size);\n+\t\trealmunmap(s->data, s->size);\n \ts->should_free = s->should_munmap = 0;\n \ts->data = NULL;\n \tfree(s->cnt_data);\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 5d543d2..85150f8 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -42,22 +42,28 @@ extern int error(const char *err, ...) _\n \n #ifdef NO_MMAP\n \n-#ifndef PROT_READ\n+#include <sys/mman.h>\n+/*#ifndef PROT_READ\n #define PROT_READ 1\n #define PROT_WRITE 2\n #define MAP_PRIVATE 1\n #define MAP_FAILED ((void*)-1)\n-#endif\n+#endif*/\n \n #define mmap gitfakemmap\n #define munmap gitfakemunmap\n extern void *gitfakemmap(void *start, size_t length, int prot , int flags, int fd, off_t offset);\n extern int gitfakemunmap(void *start, size_t length);\n \n+extern void *realmmap(void *start, size_t length, int prot , int flags, int fd, off_t offset);\n+extern int realmunmap(void *start, size_t length);\n+\n #else /* NO_MMAP */\n \n #include <sys/mman.h>\n \n+#define realmmap mmap\n+#define realmunmap munmap\n #endif /* NO_MMAP */\n \n #ifdef NO_SETENV\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 58edec0..712a068 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -330,14 +330,14 @@ void prepare_alt_odb(void)\n \t\tclose(fd);\n \t\treturn;\n \t}\n-\tmap = mmap(NULL, st.st_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\tmap = realmmap(NULL, st.st_size, PROT_READ, MAP_PRIVATE, fd, 0);\n \tclose(fd);\n \tif (map == MAP_FAILED)\n \t\treturn;\n \n \tlink_alt_odb_entries(map, map + st.st_size, '\\n',\n \t\t\t     get_object_directory());\n-\tmunmap(map, st.st_size);\n+\trealmunmap(map, st.st_size);\n }\n \n static char *find_sha1_file(const unsigned char *sha1, struct stat *st)\n@@ -378,7 +378,7 @@ static int check_packed_git_idx(const ch\n \t\treturn -1;\n \t}\n \tidx_size = st.st_size;\n-\tidx_map = mmap(NULL, idx_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\tidx_map = realmmap(NULL, idx_size, PROT_READ, MAP_PRIVATE, fd, 0);\n \tclose(fd);\n \tif (idx_map == MAP_FAILED)\n \t\treturn -1;\n@@ -423,7 +423,7 @@ static int unuse_one_packed_git(void)\n \t}\n \tif (!lru)\n \t\treturn 0;\n-\tmunmap(lru->pack_base, lru->pack_size);\n+\trealmunmap(lru->pack_base, lru->pack_size);\n \tlru->pack_base = NULL;\n \treturn 1;\n }\n@@ -460,7 +460,7 @@ int use_packed_git(struct packed_git *p)\n \t\t}\n \t\tif (st.st_size != p->pack_size)\n \t\t\tdie(\"packfile %s size mismatch.\", p->pack_name);\n-\t\tmap = mmap(NULL, p->pack_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\t\tmap = realmmap(NULL, p->pack_size, PROT_READ, MAP_PRIVATE, fd, 0);\n \t\tclose(fd);\n \t\tif (map == MAP_FAILED)\n \t\t\tdie(\"packfile %s cannot be mapped.\", p->pack_name);\n@@ -494,7 +494,7 @@ struct packed_git *add_packed_git(char *\n \t/* do we have a corresponding .pack file? */\n \tstrcpy(path + path_len - 4, \".pack\");\n \tif (stat(path, &st) || !S_ISREG(st.st_mode)) {\n-\t\tmunmap(idx_map, idx_size);\n+\t\trealmunmap(idx_map, idx_size);\n \t\treturn NULL;\n \t}\n \t/* ok, it looks sane as far as we can check without\n@@ -647,7 +647,7 @@ static void *map_sha1_file_internal(cons\n \t\t */\n \t\tsha1_file_open_flag = 0;\n \t}\n-\tmap = mmap(NULL, st.st_size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\tmap = realmmap(NULL, st.st_size, PROT_READ, MAP_PRIVATE, fd, 0);\n \tclose(fd);\n \tif (map == MAP_FAILED)\n \t\treturn NULL;\n@@ -1184,7 +1184,7 @@ int sha1_object_info(const unsigned char\n \t\t\t*sizep = size;\n \t}\n \tinflateEnd(&stream);\n-\tmunmap(map, mapsize);\n+\trealmunmap(map, mapsize);\n \treturn status;\n }\n \n@@ -1210,7 +1210,7 @@ void * read_sha1_file(const unsigned cha\n \tmap = map_sha1_file_internal(sha1, &mapsize);\n \tif (map) {\n \t\tbuf = unpack_sha1_file(map, mapsize, type, size);\n-\t\tmunmap(map, mapsize);\n+\t\trealmunmap(map, mapsize);\n \t\treturn buf;\n \t}\n \treturn NULL;\n@@ -1493,7 +1493,7 @@ int write_sha1_to_fd(int fd, const unsig\n \t} while (posn < objsize);\n \n \tif (map)\n-\t\tmunmap(map, objsize);\n+\t\trealmunmap(map, objsize);\n \tif (temp_obj)\n \t\tfree(temp_obj);\n \n@@ -1646,7 +1646,7 @@ int index_fd(unsigned char *sha1, int fd\n \n \tbuf = \"\";\n \tif (size)\n-\t\tbuf = mmap(NULL, size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\t\tbuf = realmmap(NULL, size, PROT_READ, MAP_PRIVATE, fd, 0);\n \tclose(fd);\n \tif (buf == MAP_FAILED)\n \t\treturn -1;\n@@ -1660,7 +1660,7 @@ int index_fd(unsigned char *sha1, int fd\n \t\tret = 0;\n \t}\n \tif (size)\n-\t\tmunmap(buf, size);\n+\t\trealmunmap(buf, size);\n \treturn ret;\n }\n \n\n"},{"id":"18288","messageId":"Pine.LNX.4.63.0604031710440.9360@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3780","inReplyTo":"Pine.LNX.4.64.0604030730040.3781@g5.osdl.org","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-04-03T15:12:55Z","receivedAt":"2006-04-03T15:12:55Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 3 Apr 2006, Linus Torvalds wrote:\n\n> On Mon, 3 Apr 2006, Johannes Schindelin wrote:\n> > \n> > The problem is not mmap() on cygwin, but that a fork() has to jump through \n> > loops to reinstall the open file descriptors on cygwin. If the \n> > corresponding file was deleted, that fails. Therefore, we work around that \n> > on cygwin by actually reading the file into memory, *not* mmap()ing it.\n> \n> Well, we could actually do a _real_ mmap on pack-files. The pack-files are \n> much better mmap'ed - there we don't _want_ them to be removed while we're \n> using them. It was the index file etc that was problematic.\n> \n> Maybe the cygwin fake mmap should be triggered only for the index (and \n> possibly the individual objects - if only because there doing a \n> malloc+read may actually be faster).\n\nI hit the problem *only* with \"git-whatchanged -p\". Which means that the \nupcoming we-no-longer-write-temp-files-for-diff version should make that \ngitfakemmap() hack obsolete. (I have not checked whether there are other \nplaces where a file is mmap()ed and then used by a fork()ed process.)\n\nCiao,\nDscho\n"},{"id":"18381","messageId":"fa0b6e200604050624h13ebd8deg241ae98cef1f5a74@mail.gmail.com","threadId":"3780","inReplyTo":"Pine.LNX.4.64.0604030734440.3781@g5.osdl.org","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Kees-Jan Dijkzeul","fromEmail":"k.j.dijkzeul@gmail.com","sentAt":"2006-04-05T13:24:05Z","receivedAt":"2006-04-05T13:24:05Z","isPatch":false,"sender":{"key":"k.j.dijkzeul@gmail.com","avatar":"https://gravatar.com/avatar/7093ca9ddb281f666728c9863ed26183dc72ff01d7fd509b053b7e089a93ffcb?d=mp&s=160"},"body":"On 4/3/06, Linus Torvalds <torvalds@osdl.org> wrote:\n[...]\n> That's not hugely fundamental, but I didn't expect people to hit it this\n> quickly. What kind of project has a 1.5GB pack-file _already_? I hope it's\n> fifteen years of history (so that we'll have another fifteen years before\n> we'll have to worry about 4GB pack-files ;)\n\nI'm trying to get Git to manage my companies source tree. We're\nwriting software for digital TV sets. Anyway, the archive is about 5Gb\nin size and contains binaries, zip files, excel sheets meeting minutes\nand whatnot. So it doesn't compress very well. The 1.5Gb pack file\nhardly contains any history at all (five commits or so). On the flip\nside, for now I'll be the only one adding to the archive, so at least\nit will not grow that fast ;-)\n\nAnyway, to reconstitute the tree, I need very nearly the entire pack,\nso limiting the pack size won't do much good, as git will still try to\nallocate a total of 1.5Gb memory (which, unfortunately, isn't there\n:-)\n\nInspired by a patch of Alex Riesen (thanks, Alex), I tried to use the\nregular mmap for mapping pack files, only to discover that I compile\nwithout defining \"NO_MMAP\", so I've been using the stock mmap all\nalong. So now I'm thinking that the cygwin mmap also does a\nmalloc-and-read, just like git does with NO_MMAP. So I'll continue to\ninvestigate in that direction.\n\nTo be continued...\n\nGroetjes,\n\nKees-Jan\n"},{"id":"18384","messageId":"Pine.LNX.4.63.0604051612200.25304@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3780","inReplyTo":"fa0b6e200604050624h13ebd8deg241ae98cef1f5a74@mail.gmail.com","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-04-05T14:14:20Z","receivedAt":"2006-04-05T14:14:20Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 5 Apr 2006, Kees-Jan Dijkzeul wrote:\n\n> On 4/3/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> [...]\n> > That's not hugely fundamental, but I didn't expect people to hit it this\n> > quickly. What kind of project has a 1.5GB pack-file _already_? I hope it's\n> > fifteen years of history (so that we'll have another fifteen years before\n> > we'll have to worry about 4GB pack-files ;)\n> \n> I'm trying to get Git to manage my companies source tree. We're\n> writing software for digital TV sets. Anyway, the archive is about 5Gb\n> in size and contains binaries, zip files, excel sheets meeting minutes\n> and whatnot. So it doesn't compress very well. The 1.5Gb pack file\n> hardly contains any history at all (five commits or so). On the flip\n> side, for now I'll be the only one adding to the archive, so at least\n> it will not grow that fast ;-)\n> \n> Anyway, to reconstitute the tree, I need very nearly the entire pack,\n> so limiting the pack size won't do much good, as git will still try to\n> allocate a total of 1.5Gb memory (which, unfortunately, isn't there\n> :-)\n> \n> Inspired by a patch of Alex Riesen (thanks, Alex), I tried to use the\n> regular mmap for mapping pack files, only to discover that I compile\n> without defining \"NO_MMAP\", so I've been using the stock mmap all\n> along. So now I'm thinking that the cygwin mmap also does a\n> malloc-and-read, just like git does with NO_MMAP. So I'll continue to\n> investigate in that direction.\n\nI think cygwin's mmap() is based on the Win32 API equivalent, which could \nmean that it *is* memory mapped, but in a special area (which is smaller \nthan 1.5 gigabyte). In this case, it would make sense to limit the pack \nsize, thereby having several packs, and mmap() them as they are needed.\n\nHth,\nDscho\n"},{"id":"18405","messageId":"20060405210844.GN26780@trixie.casa.cgf.cx","threadId":"3780","inReplyTo":"Pine.LNX.4.63.0604051612200.25304@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Christopher Faylor","fromEmail":"me@cgf.cx","sentAt":"2006-04-05T21:08:44Z","receivedAt":"2006-04-05T21:08:44Z","isPatch":false,"sender":{"key":"me@cgf.cx","avatar":null},"body":"On Wed, Apr 05, 2006 at 04:14:20PM +0200, Johannes Schindelin wrote:\n>> Inspired by a patch of Alex Riesen (thanks, Alex), I tried to use the\n>> regular mmap for mapping pack files, only to discover that I compile\n>> without defining \"NO_MMAP\", so I've been using the stock mmap all\n>> along. So now I'm thinking that the cygwin mmap also does a\n>> malloc-and-read, just like git does with NO_MMAP. So I'll continue to\n>> investigate in that direction.\n>\n>I think cygwin's mmap() is based on the Win32 API equivalent, which could \n>mean that it *is* memory mapped, but in a special area (which is smaller \n>than 1.5 gigabyte). In this case, it would make sense to limit the pack \n>size, thereby having several packs, and mmap() them as they are needed.\n\nYes, cygwin's mmap uses CreateFileMapping and MapViewOfFile.  IIRC,\nWindows might have a 2G limitation lurking under the hood somewhere but\nI think that might be tweakable with some registry setting.\n\ncgf\n"},{"id":"18411","messageId":"20060405232739.GA18121@nospam.com","threadId":"3780","inReplyTo":"20060405210844.GN26780@trixie.casa.cgf.cx","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Rutger Nijlunsing","fromEmail":"rutger@nospam.com","sentAt":"2006-04-05T23:27:39Z","receivedAt":"2006-04-05T23:27:39Z","isPatch":false,"sender":{"key":"rutger.nijlunsing@gmail.com","avatar":null},"body":"On Wed, Apr 05, 2006 at 05:08:44PM -0400, Christopher Faylor wrote:\n> On Wed, Apr 05, 2006 at 04:14:20PM +0200, Johannes Schindelin wrote:\n> >> Inspired by a patch of Alex Riesen (thanks, Alex), I tried to use the\n> >> regular mmap for mapping pack files, only to discover that I compile\n> >> without defining \"NO_MMAP\", so I've been using the stock mmap all\n> >> along. So now I'm thinking that the cygwin mmap also does a\n> >> malloc-and-read, just like git does with NO_MMAP. So I'll continue to\n> >> investigate in that direction.\n> >\n> >I think cygwin's mmap() is based on the Win32 API equivalent, which could \n> >mean that it *is* memory mapped, but in a special area (which is smaller \n> >than 1.5 gigabyte). In this case, it would make sense to limit the pack \n> >size, thereby having several packs, and mmap() them as they are needed.\n> \n> Yes, cygwin's mmap uses CreateFileMapping and MapViewOfFile.  IIRC,\n> Windows might have a 2G limitation lurking under the hood somewhere but\n> I think that might be tweakable with some registry setting.\n\nWindows places its DLLs criss-cross through the memory space because\nevery DLL on the system has its own preferred place to be loaded (the\nbase address). This severely limits the amount of largest contiguous\nmemory block available, which is needed for one mmap() I think.\n\nSeveral solutions exist:\n  - enlarge the address space with the /3GB boot flag in boot.ini\n  - rebase all DLLs with REBASE.EXE (part of platform sdk) .\n    Just make them the same and fix them to a low address.\n    Problem is rebasing system dlls since those are locked by the system.\n  - at start of program before other DLLs are loaded,\n    reserve an as large part of the memory as possible with\n    VirtualAlloc()\n\n-- \nRutger Nijlunsing ---------------------------------- eludias ed dse.nl\nnever attribute to a conspiracy which can be explained by incompetence\n----------------------------------------------------------------------\n"},{"id":"18413","messageId":"20060406003449.GA7174@trixie.casa.cgf.cx","threadId":"3780","inReplyTo":"20060405232739.GA18121@nospam.com","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Christopher Faylor","fromEmail":"me@cgf.cx","sentAt":"2006-04-06T00:34:50Z","receivedAt":"2006-04-06T00:34:50Z","isPatch":false,"sender":{"key":"me@cgf.cx","avatar":null},"body":"On Thu, Apr 06, 2006 at 01:27:39AM +0200, Rutger Nijlunsing wrote:\n>On Wed, Apr 05, 2006 at 05:08:44PM -0400, Christopher Faylor wrote:\n>> On Wed, Apr 05, 2006 at 04:14:20PM +0200, Johannes Schindelin wrote:\n>> >> Inspired by a patch of Alex Riesen (thanks, Alex), I tried to use the\n>> >> regular mmap for mapping pack files, only to discover that I compile\n>> >> without defining \"NO_MMAP\", so I've been using the stock mmap all\n>> >> along. So now I'm thinking that the cygwin mmap also does a\n>> >> malloc-and-read, just like git does with NO_MMAP. So I'll continue to\n>> >> investigate in that direction.\n>> >\n>> >I think cygwin's mmap() is based on the Win32 API equivalent, which could \n>> >mean that it *is* memory mapped, but in a special area (which is smaller \n>> >than 1.5 gigabyte). In this case, it would make sense to limit the pack \n>> >size, thereby having several packs, and mmap() them as they are needed.\n>> \n>> Yes, cygwin's mmap uses CreateFileMapping and MapViewOfFile.  IIRC,\n>> Windows might have a 2G limitation lurking under the hood somewhere but\n>> I think that might be tweakable with some registry setting.\n>\n>Windows places its DLLs criss-cross through the memory space because\n>every DLL on the system has its own preferred place to be loaded (the\n>base address). This severely limits the amount of largest contiguous\n>memory block available, which is needed for one mmap() I think.\n>\n>Several solutions exist:\n>  - enlarge the address space with the /3GB boot flag in boot.ini\n\nThanks.  The 3GB boot flag is what I was trying to remember.\n\n>  - rebase all DLLs with REBASE.EXE (part of platform sdk) .\n>    Just make them the same and fix them to a low address.\n>    Problem is rebasing system dlls since those are locked by the system.\n\nCygwin has its own version of rebase and a method for rebasing all of the\ndlls in the distribution.  Using that may help squeeze out a little bit\nof memory.\n\n>  - at start of program before other DLLs are loaded,\n>    reserve an as large part of the memory as possible with\n>    VirtualAlloc()\n\nCygwin actually uses this trick to try to push DLLs into their right\nlocations after a fork.  It sort of works but sometimes, in a child\nproccess, Windows puts \"stuff\" in locations previously occupied by a\nDLL.  I could swear that it does that just to be annoying...\n\nThere is a chicken/egg problem here in that Cygwin uses Doug Lea's malloc\nand that version of malloc will use mmap when sbrk() fails -- as it is\napt to do when allocating gigabytes of memory.  So, using malloc is\nnot a way to avoid mmap.\n\ncgf\n"},{"id":"18416","messageId":"7vfykrwdcf.fsf@assigned-by-dhcp.cox.net","threadId":"3780","inReplyTo":"fa0b6e200604050624h13ebd8deg241ae98cef1f5a74@mail.gmail.com","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-06T04:13:36Z","receivedAt":"2006-04-06T04:13:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Kees-Jan Dijkzeul\" <k.j.dijkzeul@gmail.com> writes:\n\n> I'm trying to get Git to manage my companies source tree. We're\n> writing software for digital TV sets. Anyway, the archive is about 5Gb\n> in size and contains binaries, zip files, excel sheets meeting minutes\n> and whatnot. So it doesn't compress very well. The 1.5Gb pack file\n> hardly contains any history at all (five commits or so). On the flip\n> side, for now I'll be the only one adding to the archive, so at least\n> it will not grow that fast ;-)\n>\n> Anyway, to reconstitute the tree, I need very nearly the entire pack,\n> so limiting the pack size won't do much good, as git will still try to\n> allocate a total of 1.5Gb memory (which, unfortunately, isn't there\n> :-)\n\nRight now we LRU the pack files and evict older ones when we\nmmap too many, but the unit of eviction is the whole file, so it\nwould not help the case like yours at all.  It might be possible\nto mmap only part of a packfile, but it would involve fairly\nmajor surgery to sha1_file.c.\n"},{"id":"18445","messageId":"7vhd55ls24.fsf@assigned-by-dhcp.cox.net","threadId":"3780","inReplyTo":"Pine.LNX.4.64.0604030734440.3781@g5.osdl.org","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-07T08:15:47Z","receivedAt":"2006-04-07T08:15:47Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Mon, 3 Apr 2006, Linus Torvalds wrote:\n>> \n>> That said, I think git _does_ have problems with large pack-files. We have \n>> some 32-bit issues etc\n>\n> I should clarify that. git _itself_ shouldn't have any 32-bit issues, but \n> the packfile data structure does. The index has 32-bit offsets into \n> individual pack-files. \n>\n> That's not hugely fundamental,...\n\nLinus _does_ understand what he means, but let me clarify and\noutline a possible future direction.\n\n * pack-*.pack file has the following format:\n\n   - The header appears at the beginning and consists of the following:\n\n     4-byte signature\n     4-byte version number (network byte order)\n     4-byte number of objects contained in the pack (network byte order)\n\n     Observation: we cannot have more than 4G versions ;-) and\n     more than 4G objects in a pack.\n\n   - The header is followed by number of object entries, each of\n     which looks like this:\n\n     (undeltified representation)\n     n-byte type and length (4-bit type, (n-1)*7+4-bit length)\n     compressed data\n\n     (deltified representation)\n     n-byte type and length (4-bit type, (n-1)*7+4-bit length)\n     20-byte base object name\n     compressed delta data\n\n     Observation: length of each object is encoded in a variable\n     length format and is not constrained to 32-bit or anything.\n\n  - The trailer records 20-byte SHA1 checksum of all of the above.\n\n * pack-*.idx file has the following format:\n\n  - The header consists of 256 4-byte network byte order\n    integers.  N-th entry of this table records the number of\n    objects in the corresponding pack, the first byte of whose\n    object name are smaller than N.\n\n    Observation: we would need to extend this to an array of\n    8-byte integers to go beyond 4G objects per pack, but it is\n    not strictly necessary.\n\n  - The header is followed by sorted 28-byte entries, one entry\n    per object in the pack.  Each entry is:\n\n    4-byte network byte order integer, recording where the\n    object is stored in the packfile as the offset from the\n    beginning.\n\n    20-byte object name.\n\n    Observation: we would definitely need to extend this to\n    8-byte integer plus 20-byte object name to handle a packfile\n    that is larger than 4GB.\n\n  - The file is concluded with a trailer:\n\n    A copy of the 20-byte SHA1 checksum at the end of\n    corresponding packfile.\n\n    20-byte SHA1-checksum of all of the above.\n\nThis is not fundamental, in that pack idx file is something we\ncan regenerate from a packfile.  The push/fetch transfer over\ngit native protocols does not even transfer pack idx file;\ninstead, the recipient uses git-index-pack to generate pack idx.\ngit-index-pack would need to be updated to update the necessary\nfields to 8-byte integers, without breaking existing packfiles.\n\nThe code to read idx file currently has a sanity check logic to\nmake sure that the size of the idx file is consistent with\n24-byte entries (the last entry in the header matches the number\nof objects recorded in the pack).  So we could reliably tell\nbetween the current 24-byte version and 28-byte \"beyond 4GB\"\nversion, and support both formats at the same time.\n\nEven after we start supporting the 28-byte \"beyond 4GB\" format,\nwe can and we should continue writing the current 24-byte\nversion of pack idx file when the packfile offset can be\nexpressed with 32-bit.\n\nHaving said that, I have to warn that this is not for weak of\nheart.  The necessary changes would be somewhat involved.\n\n\n----------------------------------------------------------------\n\nPack idx file\n\n\tidx\n\t    +--------------------------------+\n\t    | fanout[0] = 2                  |-.\n\t    +--------------------------------+ |\n\t    | fanout[1]                      | |\n\t    +--------------------------------+ |\n\t    | fanout[2]                      | |\n\t    ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ |\n\t    | fanout[255]                    | |\n\t    +--------------------------------+ |\nmain\t    | offset                         | |\nindex\t    | object name 00XXXXXXXXXXXXXXXX | |\ntable\t    +--------------------------------+ | \n\t    | offset                         | |\n\t    | object name 00XXXXXXXXXXXXXXXX | |\n\t    +--------------------------------+ |\n\t  .-| offset                         |<+\n\t  | | object name 01XXXXXXXXXXXXXXXX |\n\t  | +--------------------------------+\n\t  | | offset                         |\n\t  | | object name 01XXXXXXXXXXXXXXXX |\n\t  | ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\t  | | offset                         |\n\t  | | object name FFXXXXXXXXXXXXXXXX |\n\t  | +--------------------------------+\ntrailer\t  | | packfile checksum              |\n\t  | +--------------------------------+\n\t  | | idxfile checksum               |\n\t  | +--------------------------------+\n          .-------.      \n                  |\nPack file entry: <+\n\n     packed object header:\n\t1-byte type (bit 4-6)\n\t       size0 (bit 0-3)\n               end-of-length (bit 7)\n        n-byte sizeN (as long as MSB is set, each 7-bit)\n\t\tsize0..sizeN form 4+7+7+..+7 bit integer, size0\n\t\tis the most significant part.\n     packed object data:\n        If it is not DELTA, then deflated bytes (the size above\n\t\tis the size before compression).\n\tIf it is DELTA, then\n\t  20-byte base object name SHA1 (the size above is the\n\t  \tsize of the delta data that follows).\n          delta data, deflated.\n"},{"id":"18446","messageId":"e157ov$nda$1@sea.gmane.org","threadId":"3780","inReplyTo":"7vhd55ls24.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-04-07T08:27:14Z","receivedAt":"2006-04-07T08:27:14Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n>  * pack-*.pack file has the following format:\n[...]\n>  * pack-*.idx file has the following format:\n[...]\nCould you please put the information in parent post somewhere in\nDocumentation, for example Documentation/technical/pack-format.txt\n(perhaps together with putting description of packing heuristic from\nhttp://marc.theaimsgroup.com/?l=git&m=114134881923320 by Jon Loeliger in\nDocumentation/technical/pack-heuristics.txt even if it doesn't conform to\n\"serious documentation\" standards)?\n\nThanks in advance\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"18449","messageId":"Pine.LNX.4.64.0604071002530.2215@localhost.localdomain","threadId":"3780","inReplyTo":"7vhd55ls24.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-04-07T14:11:38Z","receivedAt":"2006-04-07T14:11:38Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Apr 2006, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> > On Mon, 3 Apr 2006, Linus Torvalds wrote:\n> >> \n> >> That said, I think git _does_ have problems with large pack-files. We have \n> >> some 32-bit issues etc\n> >\n> > I should clarify that. git _itself_ shouldn't have any 32-bit issues, but \n> > the packfile data structure does. The index has 32-bit offsets into \n> > individual pack-files. \n> >\n> > That's not hugely fundamental,...\n> \n> Linus _does_ understand what he means, but let me clarify and\n> outline a possible future direction.\n> \n[...]\n\nFor the record, the delta code also has 32-bit limitations of its own \npresently.  It cannot encode a delta against a buffer which is larger \nthan 4GB.\n\nI however made sure the byte 0 could be used as a prefix for future \nencoding extensions, like 64-bit file offsets for example.\n\n\nNicolas\n"},{"id":"18451","messageId":"7vhd55jkz0.fsf@assigned-by-dhcp.cox.net","threadId":"3780","inReplyTo":"Pine.LNX.4.64.0604071002530.2215@localhost.localdomain","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-04-07T18:31:47Z","receivedAt":"2006-04-07T18:31:47Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> On Fri, 7 Apr 2006, Junio C Hamano wrote:\n>\n>> Linus Torvalds <torvalds@osdl.org> writes:\n>> \n>> > On Mon, 3 Apr 2006, Linus Torvalds wrote:\n>> >> \n>> >> That said, I think git _does_ have problems with large pack-files. We have \n>> >> some 32-bit issues etc\n>> >\n>> > I should clarify that. git _itself_ shouldn't have any 32-bit issues, but \n>> > the packfile data structure does. The index has 32-bit offsets into \n>> > individual pack-files. \n>> >\n>> > That's not hugely fundamental,...\n>> \n>> Linus _does_ understand what he means, but let me clarify and\n>> outline a possible future direction.\n>\n> For the record, the delta code also has 32-bit limitations of its own \n> presently.  It cannot encode a delta against a buffer which is larger \n> than 4GB.\n>\n> I however made sure the byte 0 could be used as a prefix for future \n> encoding extensions, like 64-bit file offsets for example.\n\nTrue the delta data representation, not just the \"delta code\",\nhas that limitation, but I do not think you issue \"insert 0-byte\nliteral data\" command from the deltifier side right now, so we\nshould be OK.\n\nMaybe we would want to check (cmd == 0) case to detect delta\nextension that we do not handle right now?\n"},{"id":"18453","messageId":"Pine.LNX.4.64.0604071446010.2215@localhost.localdomain","threadId":"3780","inReplyTo":"7vhd55jkz0.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cygwin can't handle huge packfiles?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-04-07T18:46:55Z","receivedAt":"2006-04-07T18:46:55Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Apr 2006, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > On Fri, 7 Apr 2006, Junio C Hamano wrote:\n> >\n> >> Linus Torvalds <torvalds@osdl.org> writes:\n> >> \n> >> > On Mon, 3 Apr 2006, Linus Torvalds wrote:\n> >> >> \n> >> >> That said, I think git _does_ have problems with large pack-files. We have \n> >> >> some 32-bit issues etc\n> >> >\n> >> > I should clarify that. git _itself_ shouldn't have any 32-bit issues, but \n> >> > the packfile data structure does. The index has 32-bit offsets into \n> >> > individual pack-files. \n> >> >\n> >> > That's not hugely fundamental,...\n> >> \n> >> Linus _does_ understand what he means, but let me clarify and\n> >> outline a possible future direction.\n> >\n> > For the record, the delta code also has 32-bit limitations of its own \n> > presently.  It cannot encode a delta against a buffer which is larger \n> > than 4GB.\n> >\n> > I however made sure the byte 0 could be used as a prefix for future \n> > encoding extensions, like 64-bit file offsets for example.\n> \n> True the delta data representation, not just the \"delta code\",\n> has that limitation, but I do not think you issue \"insert 0-byte\n> literal data\" command from the deltifier side right now, so we\n> should be OK.\n> \n> Maybe we would want to check (cmd == 0) case to detect delta\n> extension that we do not handle right now?\n\nGood idea.  Will send you a patch.\n\n\nNicolas\n"}]}