{"thread":{"id":"30077","subject":"GSoC - Some questions on the idea of \"Better big-file support\".","startedAt":"2012-03-28T04:38:05Z","lastAt":"2012-05-10T22:39:16Z","messageCount":43,"participants":["Bo Chen","Nguyen Thai Ngoc Duy","Sergio","Jeff King","Sergio Callegari","Neal Kreitzinger","Junio C Hamano","Jonathan Nieder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"187913","messageId":"CA+M5ThS2iS-NMNDosk2oR25N=PMJJVTi1D=zg7MnMCUiRoX4BQ@mail.gmail.com","threadId":"30077","inReplyTo":null,"subject":"GSoC - Some questions on the idea of \"Better big-file support\".","fromName":"Bo Chen","fromEmail":"chen@chenirvine.org","sentAt":"2012-03-28T04:38:05Z","receivedAt":"2012-03-28T04:38:05Z","isPatch":false,"sender":{"key":"chen@chenirvine.org","avatar":null},"body":"Hi, Everyone. This is Bo Chen. I am interested in the idea of \"Better\nbig-file support\".\n\nAs it is described in the idea page,\n\"Many large files (like media) do not delta very well. However, some\ndo (like VM disk images). Git could split large objects into smaller\nchunks, similar to bup, and find deltas between these much more\nmanageable chunks. There are some preliminary patches in this\ndirection, but they are in need of review and expansion.\"\n\nCan anyone elaborate a little bit why many large files do not delta\nvery well? Is it a general problem or a specific problem just for Git?\nI am really new to Git, can anyone give me some hints on which source\ncodes I should read to learn more about the current code on delta\noperation? It is said that \"there are some preliminary patches in this\ndirection\", where can I find these patches?\n\nI will appreciate it if anyone can offer some help.\n\nThanks.\n\nBo Chen\n"},{"id":"187915","messageId":"CACsJy8APtMsMJ=FrZjOP=DbzuFoemSLJTmkjaiK5Wkq9XtA4rg@mail.gmail.com","threadId":"30077","inReplyTo":"CA+M5ThS2iS-NMNDosk2oR25N=PMJJVTi1D=zg7MnMCUiRoX4BQ@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of \"Better big-file support\".","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-03-28T06:19:54Z","receivedAt":"2012-03-28T06:19:54Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Mar 28, 2012 at 11:38 AM, Bo Chen <chen@chenirvine.org> wrote:\n> Hi, Everyone. This is Bo Chen. I am interested in the idea of \"Better\n> big-file support\".\n>\n> As it is described in the idea page,\n> \"Many large files (like media) do not delta very well. However, some\n> do (like VM disk images). Git could split large objects into smaller\n> chunks, similar to bup, and find deltas between these much more\n> manageable chunks. There are some preliminary patches in this\n> direction, but they are in need of review and expansion.\"\n>\n> Can anyone elaborate a little bit why many large files do not delta\n> very well?\n\nLarge files are usually binary. Depends on the type of binary, they\nmay or may not delta well. Those that are compressed/encrypted\nobviously don't delta well because one change can make the final\nresult completely different.\n\nAnother problem with delta-ing large files with git is, current code\nneeds to load two files in memory for delta. Consuming 4G for delta 2\n2GB files does not sound good.\n\n> Is it a general problem or a specific problem just for Git?\n> I am really new to Git, can anyone give me some hints on which source\n> codes I should read to learn more about the current code on delta\n> operation? It is said that \"there are some preliminary patches in this\n> direction\", where can I find these patches?\n\nRead about rsync algorithm [2]. Bup [1] implements the same (I think)\nalgorithm, but on top of git. For preliminary patches, have a look at\njc/split-blob series at commit 4a1242d in git.git.\n\n[1] https://github.com/apenwarr/bup\n[2] http://en.wikipedia.org/wiki/Rsync#Algorithm\n-- \nDuy\n"},{"id":"187927","messageId":"loom.20120328T131530-717@post.gmane.org","threadId":"30077","inReplyTo":"CACsJy8APtMsMJ=FrZjOP=DbzuFoemSLJTmkjaiK5Wkq9XtA4rg@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Sergio","fromEmail":"sergio.callegari@gmail.com","sentAt":"2012-03-28T11:33:29Z","receivedAt":"2012-03-28T11:33:29Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"Nguyen Thai Ngoc Duy <pclouds <at> gmail.com> writes:\n\n> \n> On Wed, Mar 28, 2012 at 11:38 AM, Bo Chen <chen <at> chenirvine.org> wrote:\n> > Hi, Everyone. This is Bo Chen. I am interested in the idea of \"Better\n> > big-file support\".\n> >\n> > As it is described in the idea page,\n> > \"Many large files (like media) do not delta very well. However, some\n> > do (like VM disk images). Git could split large objects into smaller\n> > chunks, similar to bup, and find deltas between these much more\n> > manageable chunks. There are some preliminary patches in this\n> > direction, but they are in need of review and expansion.\"\n> >\n> > Can anyone elaborate a little bit why many large files do not delta\n> > very well?\n> \n> Large files are usually binary. Depends on the type of binary, they\n> may or may not delta well. Those that are compressed/encrypted\n> obviously don't delta well because one change can make the final\n> result completely different.\n\nI would add that the larger a file, the larger the temptation to use a\ncompressed format for it, so that large files are often compressed binaries.\n\nFor these, a trick to obtain good deltas can be to decompress before splitting\nin chunks with the rsync algorithm. Git filters can already be used for this,\nbut it can be tricky to assure that the decompress - recompress roundtrip\nre-creates the original compressed file.\n\nFurhermore, some compressed binaries are internally composed by multiple streams\n(think of a zip archive containing multiple files, but this is by no means\nlimited to zip). In this case, it is frequent to have many possible orderings of\nthe streams. If so, the best deltas can be obtained by sorting the streams in\nsome 'canonical' order and decompressing. Even without decompressing, sorting\nalone can obtain good results as long as changes are only due to changes in a\nsingle stream of the container. Personally, I know no example of git filters\nused to perform this sorting which can be extremely tricky in assuring the\npossibility of recovering the file in the original stream order.\n\nMaybe (but this is just speculation), once the bup-inspired file chunking\nsupport is in place, people will start contributing filters to improve the\nmanagement of many types of standard files (obviously 'improve' in terms of\nspace efficiency as filters can be quite slow).\n\nSergio\n"},{"id":"188185","messageId":"CA+M5ThS1XiaGJWmSvfwXoqebnH6fK3h6cC7OnQQi=LXzcA0GRw@mail.gmail.com","threadId":"30077","inReplyTo":"CACsJy8APtMsMJ=FrZjOP=DbzuFoemSLJTmkjaiK5Wkq9XtA4rg@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of \"Better big-file support\".","fromName":"Bo Chen","fromEmail":"chen@chenirvine.org","sentAt":"2012-03-30T19:11:40Z","receivedAt":"2012-03-30T19:11:40Z","isPatch":false,"sender":{"key":"chen@chenirvine.org","avatar":null},"body":"Sorry for replying late.\n\nMy questions are inline in the following.\n\n\nOn Wed, Mar 28, 2012 at 2:19 AM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n> On Wed, Mar 28, 2012 at 11:38 AM, Bo Chen <chen@chenirvine.org> wrote:\n>> Hi, Everyone. This is Bo Chen. I am interested in the idea of \"Better\n>> big-file support\".\n>>\n>> As it is described in the idea page,\n>> \"Many large files (like media) do not delta very well. However, some\n>> do (like VM disk images). Git could split large objects into smaller\n>> chunks, similar to bup, and find deltas between these much more\n>> manageable chunks. There are some preliminary patches in this\n>> direction, but they are in need of review and expansion.\"\n>>\n>> Can anyone elaborate a little bit why many large files do not delta\n>> very well?\n>\n> Large files are usually binary. Depends on the type of binary, they\n> may or may not delta well. Those that are compressed/encrypted\n> obviously don't delta well because one change can make the final\n> result completely different.\n\nJust make clear one of my confusions. Delta operation is to find out\nthe differences between different versions of the same file, right?\nAs I know, delta encoding is to re-encode a file based on the\ndifferences between neighboring blocks, thus can help compress a file\nsince after delta encoding, we will have more similar data within the\nfile. Can anyone elaborate a little bit what is the relation between\ndelta operation in git and delta encoding listed above? Thanks.\n\n>\n> Another problem with delta-ing large files with git is, current code\n> needs to load two files in memory for delta. Consuming 4G for delta 2\n> 2GB files does not sound good.\n\n\nI am wondering why we cannot divide the 2  2GB files into chunks and\ndelta chunks by chunks. Is that any difference, except a little more\nIOs?\n\n>\n>> Is it a general problem or a specific problem just for Git?\n>> I am really new to Git, can anyone give me some hints on which source\n>> codes I should read to learn more about the current code on delta\n>> operation? It is said that \"there are some preliminary patches in this\n>> direction\", where can I find these patches?\n>\n> Read about rsync algorithm [2]. Bup [1] implements the same (I think)\n> algorithm, but on top of git. For preliminary patches, have a look at\n> jc/split-blob series at commit 4a1242d in git.git.\n\nMake clear my another confusion. The file which has been updated\n(added, deleted, and modified) is first delta-compressed, and then\nsynchronize to the remote repo by some mechanism (rsync?). I am\nwondering what is the the relationship between delta operation and\nrsync.\n\n>\n> [1] https://github.com/apenwarr/bup\n> [2] http://en.wikipedia.org/wiki/Rsync#Algorithm\n> --\n> Duy\n\nBo\n"},{"id":"188186","messageId":"CA+M5ThT47twke7xkeAdDbk0c_J_=U6t1swDVexD6WDrQjG9_-w@mail.gmail.com","threadId":"30077","inReplyTo":"loom.20120328T131530-717@post.gmane.org","subject":"Re: GSoC - Some questions on the idea of","fromName":"Bo Chen","fromEmail":"chen@chenirvine.org","sentAt":"2012-03-30T19:44:03Z","receivedAt":"2012-03-30T19:44:03Z","isPatch":false,"sender":{"key":"chen@chenirvine.org","avatar":null},"body":"The following is the list of sub-problems according to my\nunderstanding of the \"big file support\" problem. Can anyone give some\nfeed back and help refine it. Thanks.\n\n            ---- text file (always delta well? need to be confirmed)\n             |\n\n                                               --- delta well (ok)\nlarge file-|                    ----    general binary file (without\nencryption, compression. Other cases which definitely can not delta\nwell)  -|\n             |                     |\n\n                                               --- does not delta well\n(improvement?)\n            ---- binary file   -|---   encrypted file (improvement?\none straightforward method is to decrypt the file before delta-ing it,\nhowever, we don't always have the key for decryption. Other?)\n                                   |\n                                  ---    compressed file (improvement?\nDecompress before delta-ing it? Other?)\n\n\n\nBo\n\nOn Wed, Mar 28, 2012 at 7:33 AM, Sergio <sergio.callegari@gmail.com> wrote:\n> Nguyen Thai Ngoc Duy <pclouds <at> gmail.com> writes:\n>\n>>\n>> On Wed, Mar 28, 2012 at 11:38 AM, Bo Chen <chen <at> chenirvine.org> wrote:\n>> > Hi, Everyone. This is Bo Chen. I am interested in the idea of \"Better\n>> > big-file support\".\n>> >\n>> > As it is described in the idea page,\n>> > \"Many large files (like media) do not delta very well. However, some\n>> > do (like VM disk images). Git could split large objects into smaller\n>> > chunks, similar to bup, and find deltas between these much more\n>> > manageable chunks. There are some preliminary patches in this\n>> > direction, but they are in need of review and expansion.\"\n>> >\n>> > Can anyone elaborate a little bit why many large files do not delta\n>> > very well?\n>>\n>> Large files are usually binary. Depends on the type of binary, they\n>> may or may not delta well. Those that are compressed/encrypted\n>> obviously don't delta well because one change can make the final\n>> result completely different.\n>\n> I would add that the larger a file, the larger the temptation to use a\n> compressed format for it, so that large files are often compressed binaries.\n>\n> For these, a trick to obtain good deltas can be to decompress before splitting\n> in chunks with the rsync algorithm. Git filters can already be used for this,\n> but it can be tricky to assure that the decompress - recompress roundtrip\n> re-creates the original compressed file.\n>\n> Furhermore, some compressed binaries are internally composed by multiple streams\n> (think of a zip archive containing multiple files, but this is by no means\n> limited to zip). In this case, it is frequent to have many possible orderings of\n> the streams. If so, the best deltas can be obtained by sorting the streams in\n> some 'canonical' order and decompressing. Even without decompressing, sorting\n> alone can obtain good results as long as changes are only due to changes in a\n> single stream of the container. Personally, I know no example of git filters\n> used to perform this sorting which can be extremely tricky in assuring the\n> possibility of recovering the file in the original stream order.\n>\n> Maybe (but this is just speculation), once the bup-inspired file chunking\n> support is in place, people will start contributing filters to improve the\n> management of many types of standard files (obviously 'improve' in terms of\n> space efficiency as filters can be quite slow).\n>\n> Sergio\n>\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"188196","messageId":"CA+M5ThTPyic=RhFL2SvuNB0xBWOHxNTaUZrYMB144UjpjCiLoQ@mail.gmail.com","threadId":"30077","inReplyTo":"loom.20120328T131530-717@post.gmane.org","subject":"Re: GSoC - Some questions on the idea of","fromName":"Bo Chen","fromEmail":"chen@chenirvine.org","sentAt":"2012-03-30T19:51:20Z","receivedAt":"2012-03-30T19:51:20Z","isPatch":false,"sender":{"key":"chen@chenirvine.org","avatar":null},"body":"Please neglect my last email.\nFollowing is the version more readable.\nThe sub-problems of \"delta for large file\" problem.\n\n1 large file\n\n1.1 text file (always delta well? need to be confirmed)\n\n1.2 binary file\n\n1.2.1  general binary file (without encryption, compression. Other\ncases which definitely can not delta well)\n\n1.2.1.1 delta well (ok)\n1.2.1.2 does not delta well (improvement?)\n\n1.2.2  encrypted file (improvement? one straightforward method is to\ndecrypt the file before delta-ing it, however, we don't always have\nthe key for decryption. Other?)\n\n1.2.3 compressed file (improvement? Decompress before delta-ing it? Other?)\n\nCan anyone give me any feed back for further refining the problem. Thanks.\n\nBo\n\nOn Wed, Mar 28, 2012 at 7:33 AM, Sergio <sergio.callegari@gmail.com> wrote:\n> Nguyen Thai Ngoc Duy <pclouds <at> gmail.com> writes:\n>\n>>\n>> On Wed, Mar 28, 2012 at 11:38 AM, Bo Chen <chen <at> chenirvine.org> wrote:\n>> > Hi, Everyone. This is Bo Chen. I am interested in the idea of \"Better\n>> > big-file support\".\n>> >\n>> > As it is described in the idea page,\n>> > \"Many large files (like media) do not delta very well. However, some\n>> > do (like VM disk images). Git could split large objects into smaller\n>> > chunks, similar to bup, and find deltas between these much more\n>> > manageable chunks. There are some preliminary patches in this\n>> > direction, but they are in need of review and expansion.\"\n>> >\n>> > Can anyone elaborate a little bit why many large files do not delta\n>> > very well?\n>>\n>> Large files are usually binary. Depends on the type of binary, they\n>> may or may not delta well. Those that are compressed/encrypted\n>> obviously don't delta well because one change can make the final\n>> result completely different.\n>\n> I would add that the larger a file, the larger the temptation to use a\n> compressed format for it, so that large files are often compressed binaries.\n>\n> For these, a trick to obtain good deltas can be to decompress before splitting\n> in chunks with the rsync algorithm. Git filters can already be used for this,\n> but it can be tricky to assure that the decompress - recompress roundtrip\n> re-creates the original compressed file.\n>\n> Furhermore, some compressed binaries are internally composed by multiple streams\n> (think of a zip archive containing multiple files, but this is by no means\n> limited to zip). In this case, it is frequent to have many possible orderings of\n> the streams. If so, the best deltas can be obtained by sorting the streams in\n> some 'canonical' order and decompressing. Even without decompressing, sorting\n> alone can obtain good results as long as changes are only due to changes in a\n> single stream of the container. Personally, I know no example of git filters\n> used to perform this sorting which can be extremely tricky in assuring the\n> possibility of recovering the file in the original stream order.\n>\n> Maybe (but this is just speculation), once the bup-inspired file chunking\n> support is in place, people will start contributing filters to improve the\n> management of many types of standard files (obviously 'improve' in terms of\n> space efficiency as filters can be quite slow).\n>\n> Sergio\n>\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"188197","messageId":"20120330195404.GA20189@sigill.intra.peff.net","threadId":"30077","inReplyTo":"CA+M5ThS1XiaGJWmSvfwXoqebnH6fK3h6cC7OnQQi=LXzcA0GRw@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of \"Better big-file support\".","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-03-30T19:54:04Z","receivedAt":"2012-03-30T19:54:04Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Mar 30, 2012 at 03:11:40PM -0400, Bo Chen wrote:\n\n> Just make clear one of my confusions. Delta operation is to find out\n> the differences between different versions of the same file, right?\n> As I know, delta encoding is to re-encode a file based on the\n> differences between neighboring blocks, thus can help compress a file\n> since after delta encoding, we will have more similar data within the\n> file. Can anyone elaborate a little bit what is the relation between\n> delta operation in git and delta encoding listed above? Thanks.\n\nSort of. Git is snapshot based. So each version of a file is its own\n\"object\", and from a high-level view, we store all objects. But we store\nthe logical objects themselves in packfiles, in which the actual\nrepresentation of the object may be stored as a difference to another\nobject (which is likely to be a different version of the same file, but\ndoes not have to be).\n\nHere's some background reading:\n\n  http://progit.org/book/ch1-3.html\n\n  http://progit.org/book/ch9-4.html\n\n> I am wondering why we cannot divide the 2  2GB files into chunks and\n> delta chunks by chunks. Is that any difference, except a little more\n> IOs?\n\nIt's more complicated than that. What if the file is re-ordered? You\nwould want to compare early chunks in one version against later chunks\nin the other. So yes, you can reduce memory pressure by doing more I/O,\nbut doing too much I/O will be very slow. Coming up with a solution is\npart of what this project is about. And chunking is part of that\nsolution.\n\n> > Read about rsync algorithm [2]. Bup [1] implements the same (I think)\n> > algorithm, but on top of git. For preliminary patches, have a look at\n> > jc/split-blob series at commit 4a1242d in git.git.\n> \n> Make clear my another confusion. The file which has been updated\n> (added, deleted, and modified) is first delta-compressed, and then\n> synchronize to the remote repo by some mechanism (rsync?). I am\n> wondering what is the the relationship between delta operation and\n> rsync.\n\nNo, the updated file is delta compressed into a packfile, and the\npackfile is transmitted. Rsync comes into play because it uses a novel\nchunking algorithm, which was copied by bup (and is referred to as the\n\"bupsplit\" algorithm). Read up on how bup works and why it was invented.\n\n-Peff\n"},{"id":"188202","messageId":"20120330203430.GB20376@sigill.intra.peff.net","threadId":"30077","inReplyTo":"CA+M5ThTPyic=RhFL2SvuNB0xBWOHxNTaUZrYMB144UjpjCiLoQ@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-03-30T20:34:30Z","receivedAt":"2012-03-30T20:34:30Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Mar 30, 2012 at 03:51:20PM -0400, Bo Chen wrote:\n\n> The sub-problems of \"delta for large file\" problem.\n> \n> 1 large file\n> \n> 1.1 text file (always delta well? need to be confirmed)\n\nThey often do, but text files don't tend to be large. There are some\nexceptions (e.g., genetic data is often kept in line-oriented text\nfiles, but is very large).\n\nBut let's take a step back for a moment. Forget about whether a file is\nbinary or not. Imagine you want to store a very large file in git.\n\nWhat are the operations that will perform badly? How can we make them\nperform acceptably, and what tradeoffs must we make? E.g., the way the\ndiff code is written, it would be very difficult to run \"git diff\" on a\n2 gigabyte file. But is that actually a problem? Answering that means\ntalking about the characteristics of 2 gigabyte files, and what we\nexpect to see, and to what degree our tradeoffs will impact them.\n\nHere's a more concrete example. At first, even storing a 2 gigabyte file\nwith \"git add\" was painful, because we would load the whole thing in\nmemory. Repacking the repository was painful, because we had to rewrite\nthe whole 2G file into a packfile. Nowadays, we stream large files\ndirectly into their own packfiles, and we have to pay the I/O only once\n(and the memory cost never). As a tradeoff, we no longer get delta\ncompression of large objects. That's OK for some large objects, like\nmovie files (which don't tend to delta well, anyway). But it's not for\nother objects, like virtual machine images, which do tend to delta well.\n\nSo can we devise a solution which efficiently stores these\ndelta-friendly objects, without losing the performance improvements we\ngot with the stream-directly-to-packfile approach?\n\nOne possible solution is breaking large files into smaller chunks using\nsomething like the bupsplit algorithm (and I won't go into the details\nhere, as links to bup have already been mentioned elsewhere, and Junio's\npatches make a start at this sort of splitting).\n\nNote that there are other problem areas with big files that can be\nworked on, too. For example, some people want to store 100 gigabytes in\na repository. Because git is distributed, that means 100G in the repo\ndatabase, and 100G in the working directory, for a total of 200G. People\nin this situation may want to be able to store part of the repository\ndatabase in a network-accessible location, trading some of the\nconvenience of being fully distributed for the space savings. So another\nproject could be designing a network-based alternate object storage\nsystem.\n\n-Peff\n"},{"id":"188222","messageId":"CA+M5ThR6jtxqs0-Kz-8fcRuOFRbLr-GvsJcTmrOQ7_geNspDLg@mail.gmail.com","threadId":"30077","inReplyTo":"20120330203430.GB20376@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Bo Chen","fromEmail":"chen@chenirvine.org","sentAt":"2012-03-30T23:08:42Z","receivedAt":"2012-03-30T23:08:42Z","isPatch":false,"sender":{"key":"chen@chenirvine.org","avatar":null},"body":"I appreciate for the instant reply.\n\nMy comments are inline below.\n\nOn Fri, Mar 30, 2012 at 4:34 PM, Jeff King <peff@peff.net> wrote:\n> On Fri, Mar 30, 2012 at 03:51:20PM -0400, Bo Chen wrote:\n>\n>> The sub-problems of \"delta for large file\" problem.\n>>\n>> 1 large file\n>>\n>> 1.1 text file (always delta well? need to be confirmed)\n>\n> They often do, but text files don't tend to be large. There are some\n> exceptions (e.g., genetic data is often kept in line-oriented text\n> files, but is very large).\n>\n> But let's take a step back for a moment. Forget about whether a file is\n> binary or not. Imagine you want to store a very large file in git.\n>\n> What are the operations that will perform badly? How can we make them\n> perform acceptably, and what tradeoffs must we make? E.g., the way the\n> diff code is written, it would be very difficult to run \"git diff\" on a\n> 2 gigabyte file. But is that actually a problem? Answering that means\n> talking about the characteristics of 2 gigabyte files, and what we\n> expect to see, and to what degree our tradeoffs will impact them.\n>\n> Here's a more concrete example. At first, even storing a 2 gigabyte file\n> with \"git add\" was painful, because we would load the whole thing in\n> memory. Repacking the repository was painful, because we had to rewrite\n> the whole 2G file into a packfile. Nowadays, we stream large files\n> directly into their own packfiles, and we have to pay the I/O only once\n> (and the memory cost never). As a tradeoff, we no longer get delta\n> compression of large objects. That's OK for some large objects, like\n> movie files (which don't tend to delta well, anyway). But it's not for\n> other objects, like virtual machine images, which do tend to delta well.\n\nIt seems that we should first provide some kind of mechanism which can\ndistinguish the delta-friendly objects and non delta-friendly objects.\nI am wondering whether this algorithm is available now or will be\ndevised.\n\n\n\n>\n> So can we devise a solution which efficiently stores these\n> delta-friendly objects, without losing the performance improvements we\n> got with the stream-directly-to-packfile approach?\n\nAh, I see. Design efficient solution for storing the delta-friendly\nobjects is the main concern. Thank you for helping me clarify this\npoint.\n\n>\n> One possible solution is breaking large files into smaller chunks using\n> something like the bupsplit algorithm (and I won't go into the details\n> here, as links to bup have already been mentioned elsewhere, and Junio's\n> patches make a start at this sort of splitting).\n>\n> Note that there are other problem areas with big files that can be\n> worked on, too. For example, some people want to store 100 gigabytes in\n> a repository. Because git is distributed, that means 100G in the repo\n> database, and 100G in the working directory, for a total of 200G. People\n> in this situation may want to be able to store part of the repository\n> database in a network-accessible location, trading some of the\n> convenience of being fully distributed for the space savings. So another\n> project could be designing a network-based alternate object storage\n> system.\n\n>From the architecture point of view, CVS is fully centralized, and Git\nis fully distributed. It seems that for big repo, the architecture\ndescribed above is in the middle now ^-^.\n\n>\n> -Peff\n\nBo\n"},{"id":"188247","messageId":"4F76E430.6020605@gmail.com","threadId":"30077","inReplyTo":"CA+M5ThR6jtxqs0-Kz-8fcRuOFRbLr-GvsJcTmrOQ7_geNspDLg@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2012-03-31T11:02:08Z","receivedAt":"2012-03-31T11:02:08Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"I wonder if it could make sense to have some pluggable mechanism for file \nsplitting. Something under the lines of filters, so to say.\nBupsplit can be a rather general mechanism, but large binaries that are \ncontainers (zip, jar, docx, tgz, pdf - seen as a collection of streams) may \npossibly be\nmore conveniently split by their inherent components.\n"},{"id":"188253","messageId":"4F77209A.8050607@gmail.com","threadId":"30077","inReplyTo":"20120330203430.GB20376@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-03-31T15:19:54Z","receivedAt":"2012-03-31T15:19:54Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 3/30/2012 3:34 PM, Jeff King wrote:\n> On Fri, Mar 30, 2012 at 03:51:20PM -0400, Bo Chen wrote:\n>\n>> The sub-problems of \"delta for large file\" problem.\n>>\n>> 1 large file\n>>\n> Note that there are other problem areas with big files that can be\n> worked on, too. For example, some people want to store 100 gigabytes\n> in a repository.\n\nI take it that you have in mind a 100G set of files comprised entirely\nof big-files that cannot be logically separated into smaller submodules?\n\nMy understanding is that a main strategy for \"big files\" is to separate\nyour big-files logically into their own submodule(s) to keep them from\nbogging down the not-big-file repo(s).\n\nIs one of the goals of big-file-support to make submodule strategizing \nunconcerned about big-file groupings and only concerned about \nlogical-file groupings?  Big-file groupings are not necessarily logical \nfile groupings, but perhaps a technical file grouping subset of a \nlogical file grouping that is necessitated by big-file performance \nconsiderations.  IOW, is the goal of big-file-support to make big-files \n\"just work\" so that users don't have to think about graphics files, \nbinaries, etc, and just treat them like everything else?  Obviously, a \n100G database file will always be a 'big-file' for the foreseeable \nfuture, but a 0.5G graphics file is not a \"big file\" generally speaking \n(as opposed to git-speaking).\n\n> Because git is distributed, that means 100G in the repo database,\n> and 100G in the working directory, for a total of 200G.\n\nI take it that you are implying that the 100G object-store size is due\nto the notion that binary files cannot-be/are-not compressed well?\n\n> People in this situation may want to be able to store part of the\n> repository database in a network-accessible location, trading some\n> of the convenience of being fully distributed for the space savings.\n> So another project could be designing a network-based alternate\n> object storage system.\n>\nI take it you are implying a local area network with users git repos on \nworkstations?\n\nIn regards to \"network-based alternate objects\" that are in fact on the \ninternet they would need to first be cloned onto the local area network. \n  Or are you imagining this would work for internet \"network-based \nalternate objects\"?\n\nSome setups login to a linux server and have all their repos there.  The \n\"alternate objects\" does not need to network-based in that case.  It is \n\"local\", but local does not mean 20 people cloning the alternate objects \nto their workstations.  It means one copy of alternate objects, and \ntwenty repos referencing that one copy.\n\nv/r,\nneal\n"},{"id":"188254","messageId":"4F772E48.3030708@gmail.com","threadId":"30077","inReplyTo":"4F76E430.6020605@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-03-31T16:18:16Z","receivedAt":"2012-03-31T16:18:16Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 3/31/2012 6:02 AM, Sergio Callegari wrote:\n> I wonder if it could make sense to have some pluggable mechanism for\n>  file splitting. Something under the lines of filters, so to say.\n> Bupsplit can be a rather general mechanism, but large binaries that\n> are containers (zip, jar, docx, tgz, pdf - seen as a collection of\n> streams) may possibly be more conveniently split by their inherent\n> components.\n>\n\ngitattributes or gitconfig could configure the big-file handler for \nspecified files.  Known/supported filetypes like gif, png, zip, pdf, \netc., could be auto-configured by git.  Any yet-unknown/yet-unsupported \nfiletypes could be configured manually by the user, e.g.\n*.zgp=bigcontainer\n\nv/r,\nneal\n"},{"id":"188255","messageId":"4F7735B7.1050707@gmail.com","threadId":"30077","inReplyTo":"20120330203430.GB20376@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-03-31T16:49:59Z","receivedAt":"2012-03-31T16:49:59Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 3/30/2012 3:34 PM, Jeff King wrote:\n> On Fri, Mar 30, 2012 at 03:51:20PM -0400, Bo Chen wrote:\n>\n>> The sub-problems of \"delta for large file\" problem.\n>>\n>> 1 large file\n>>\n>> 1.1 text file (always delta well? need to be confirmed)\n>\n> ...But let's take a step back for a moment. Forget about whether a file\n> is binary or not. Imagine you want to store a very large file in\n> git.\n>\n> ...Nowadays, we stream large files directly into their own packfiles,\n> and we have to pay the I/O only once (and the memory cost never). As\n> a tradeoff, we no longer get delta compression of large objects.\n> That's OK for some large objects, like movie files (which don't tend\n> to delta well, anyway). But it's not for other objects, like virtual\n> machine images, which do tend to delta well.\n>\n> So can we devise a solution which efficiently stores these\n> delta-friendly objects, without losing the performance improvements\n> we got with the stream-directly-to-packfile approach?\n>\n\ngitconfig or gitattributes could specify big-file handlers for \nfiletypes.  It seems a bit ridiculous to expect git to autoconfigure \nbig-file handlers for everything from gif's to vm-images.  In the case \nof vm-images you would need to read the \"big-files\" man-page and then \nconfigure your git for the \"vm image handler\" for whatever your vm-image \nwildcards are for those files.  For movie files you would also read the \nbig-file man-page and configure \"movie file 'x' big file handler' for \nwhatever your movie file wildcards are.  Movie files and vm-images are \nvery expectable (version control) but not very normative (source code \nmanagement) so you need to configure those as needed.  More \nwidely-tracked-by-the-public-at-large files like gif, png, etc, could be \nautoconfigured by git to used the correct big-file handler.\n\nv/r,\nneal\n"},{"id":"188262","messageId":"4F7768D6.3010400@gmail.com","threadId":"30077","inReplyTo":"20120330203430.GB20376@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-03-31T20:28:06Z","receivedAt":"2012-03-31T20:28:06Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 3/30/2012 3:34 PM, Jeff King wrote:\n> On Fri, Mar 30, 2012 at 03:51:20PM -0400, Bo Chen wrote:\n>\n>> The sub-problems of \"delta for large file\" problem.\n>>\n>> 1 large file\n>>\n> But let's take a step back for a moment. Forget about whether a file is\n> binary or not. Imagine you want to store a very large file in git.\n>\n> What are the operations that will perform badly? How can we make them\n> perform acceptably, and what tradeoffs must we make? E.g., the way the\n> diff code is written, it would be very difficult to run \"git diff\" on a\n> 2 gigabyte file. But is that actually a problem? Answering that means\n> talking about the characteristics of 2 gigabyte files, and what we\n> expect to see, and to what degree our tradeoffs will impact them.\n>\n> Here's a more concrete example. At first, even storing a 2 gigabyte file\n> with \"git add\" was painful, because we would load the whole thing in\n> memory. Repacking the repository was painful, because we had to rewrite\n> the whole 2G file into a packfile. Nowadays, we stream large files\n> directly into their own packfiles, and we have to pay the I/O only once\n> (and the memory cost never). As a tradeoff, we no longer get delta\n> compression of large objects. That's OK for some large objects, like\n> movie files (which don't tend to delta well, anyway). But it's not for\n> other objects, like virtual machine images, which do tend to delta well.\n>\n> So can we devise a solution which efficiently stores these\n> delta-friendly objects, without losing the performance improvements we\n> got with the stream-directly-to-packfile approach?\n>\n> One possible solution is breaking large files into smaller chunks using\n> something like the bupsplit algorithm (and I won't go into the details\n> here, as links to bup have already been mentioned elsewhere, and Junio's\n> patches make a start at this sort of splitting).\n>\n(I'm no expert on \"big-files\" in git or elsewhere, but this thread is \nimmensely interesting to me as a git user who wants to track all sorts \nof binary files and possibly large text files in the very near future, \nie. all components tied to a server build and upgrades beyond the \nlinux-distro/rpms and perhaps including them also.)\n\nLet's take an even bigger step back for a moment.  Who determines if a \nfile shall be a big-file or not?  Git or the user?  How is it determined \nif a file shall be a \"big-file\" or not?\n\nWho decides bigness:\nBigness seems to be relative to system resources.  Does the user crunch \nthe numbers to determine if a file is big-file, or does git?  If the \nnumbers are relative then should git query the system and make the \ndetermination?  Either way, once the system-resources are upgraded and \nformerly \"big-files\" are no longer considered \"big\" how is the previous \nhistory refactored to behave \"non-big-file-like\"?  Conversely, if the \nsystem-resources are re-distributed so that formerly non-big files are \nnow relatively big (ie, moved from powerful central server login to \nlaptops), how is the history refactored to accommodate the \nnewly-relative-bigness?\n\nHow bigness is decided:\nThere seems to be two basic types of big-files:  big-worktree-files, and \nbig-history-files.  A big-worktree-file that is delta-friendly is not a \nbig-history-file.  A non-big-worktree-file that is delta-unfriendly is a \nbig-file-history problem.  If you are working alone on an old computer \nyou are probably more concerned about big-worktree-files (memory).  If \nyou are working in a large group making lots of changes to the same \nfiles on a powerful server then you are probably more concerned about \nbig-history-file-size (diskspace).  Of course, all are concerned about \nbig-worktree-files that are delta-unfriendly.\n\nAt what point is a delta-friendly file considered a \"big-file\"?  I \nassume that may depend on the degree delta-friendliness.  I imagine that \na text file and vm-image differ in delta-friendliness by several degrees.\n\nAt what point(s) is a delta-unfriendly file considered a \"big-file\"?  I \nassume that may depend on the degree(s) of delta-unfriendliness.  I \nimagine a compiled program and compressed-container differ in \ndelta-unfriendliness by several degrees.\n\nMy understanding is that git does not ever delta-compress binary files. \n  That would mean even a small-worktree-binary-file becomes a \nbig-history-file over time.\n\nv/r,\nneal\n"},{"id":"188264","messageId":"CA+M5ThTKtSFPq8A3oc1wvc9i0vG1NMyHCRE+poYaq+65FQWOxw@mail.gmail.com","threadId":"30077","inReplyTo":"4F7768D6.3010400@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Bo Chen","fromEmail":"chen@chenirvine.org","sentAt":"2012-03-31T21:27:49Z","receivedAt":"2012-03-31T21:27:49Z","isPatch":false,"sender":{"key":"chen@chenirvine.org","avatar":null},"body":"On Sat, Mar 31, 2012 at 4:28 PM, Neal Kreitzinger\n<nkreitzinger@gmail.com> wrote:\n> On 3/30/2012 3:34 PM, Jeff King wrote:\n>>\n>> On Fri, Mar 30, 2012 at 03:51:20PM -0400, Bo Chen wrote:\n>>\n>>> The sub-problems of \"delta for large file\" problem.\n>>>\n>>> 1 large file\n>>>\n>> But let's take a step back for a moment. Forget about whether a file is\n>> binary or not. Imagine you want to store a very large file in git.\n>>\n>> What are the operations that will perform badly? How can we make them\n>> perform acceptably, and what tradeoffs must we make? E.g., the way the\n>> diff code is written, it would be very difficult to run \"git diff\" on a\n>> 2 gigabyte file. But is that actually a problem? Answering that means\n>> talking about the characteristics of 2 gigabyte files, and what we\n>> expect to see, and to what degree our tradeoffs will impact them.\n>>\n>> Here's a more concrete example. At first, even storing a 2 gigabyte file\n>> with \"git add\" was painful, because we would load the whole thing in\n>> memory. Repacking the repository was painful, because we had to rewrite\n>> the whole 2G file into a packfile. Nowadays, we stream large files\n>> directly into their own packfiles, and we have to pay the I/O only once\n>> (and the memory cost never). As a tradeoff, we no longer get delta\n>> compression of large objects. That's OK for some large objects, like\n>> movie files (which don't tend to delta well, anyway). But it's not for\n>> other objects, like virtual machine images, which do tend to delta well.\n>>\n>> So can we devise a solution which efficiently stores these\n>> delta-friendly objects, without losing the performance improvements we\n>> got with the stream-directly-to-packfile approach?\n>>\n>> One possible solution is breaking large files into smaller chunks using\n>> something like the bupsplit algorithm (and I won't go into the details\n>> here, as links to bup have already been mentioned elsewhere, and Junio's\n>> patches make a start at this sort of splitting).\n>>\n> (I'm no expert on \"big-files\" in git or elsewhere, but this thread is\n> immensely interesting to me as a git user who wants to track all sorts of\n> binary files and possibly large text files in the very near future, ie. all\n> components tied to a server build and upgrades beyond the linux-distro/rpms\n> and perhaps including them also.)\n>\n> Let's take an even bigger step back for a moment.  Who determines if a file\n> shall be a big-file or not?  Git or the user?  How is it determined if a\n> file shall be a \"big-file\" or not?\n>\n> Who decides bigness:\n> Bigness seems to be relative to system resources.  Does the user crunch the\n> numbers to determine if a file is big-file, or does git?  If the numbers are\n> relative then should git query the system and make the determination?\n>  Either way, once the system-resources are upgraded and formerly \"big-files\"\n> are no longer considered \"big\" how is the previous history refactored tot\n> behave \"non-big-file-like\"?  Conversely, if the system-resources are\n> re-distributed so that formerly non-big files are now relatively big (ie,\n> moved from powerful central server login to laptops), how is the history\n> refactored to accommodate the newly-relative-bigness?\n>\n\nIn common sense, a file of tens of MBs should not be considered as a\nbig file, but a file of tens of GBs should definitely be considered as\na big file. I think one simple workable solution is to let the user\nset the threshold of the big file. One complicate but intelligent\nsolution is to let git auto-config the threshold by evaluating current\ncomputing resources in the running platform (a physical machine or\njust a VM). As to the problem of migrating git in different platforms\nwhich equip with different computing power, the git repo should also\nkeep tract of under what big file threshold a specific file is\nhandled.\n\n\n> How bigness is decided:\n> There seems to be two basic types of big-files:  big-worktree-files, and\n> big-history-files.  A big-worktree-file that is delta-friendly is not a\n> big-history-file.  A non-big-worktree-file that is delta-unfriendly is a\n> big-file-history problem.  If you are working alone on an old computer you\n> are probably more concerned about big-worktree-files (memory).  If you are\n> working in a large group making lots of changes to the same files on a\n> powerful server then you are probably more concerned about\n> big-history-file-size (diskspace).  Of course, all are concerned about\n> big-worktree-files that are delta-unfriendly.\n>\n> At what point is a delta-friendly file considered a \"big-file\"?  I assume\n> that may depend on the degree delta-friendliness.  I imagine that a text\n> file and vm-image differ in delta-friendliness by several degrees.\n>\n> At what point(s) is a delta-unfriendly file considered a \"big-file\"?  I\n> assume that may depend on the degree(s) of delta-unfriendliness.  I imagine\n> a compiled program and compressed-container differ in delta-unfriendliness\n> by several degrees.\n>\n> My understanding is that git does not ever delta-compress binary files.\n>  That would mean even a small-worktree-binary-file becomes a\n> big-history-file over time.\n>\n> v/r,\n> neal\n"},{"id":"188274","messageId":"CACsJy8DTegW78Qw7-T6uK_oZj2CELv57bbH6sU=bScHDesGYPQ@mail.gmail.com","threadId":"30077","inReplyTo":"CA+M5ThTKtSFPq8A3oc1wvc9i0vG1NMyHCRE+poYaq+65FQWOxw@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-04-01T04:22:44Z","receivedAt":"2012-04-01T04:22:44Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Apr 1, 2012 at 4:27 AM, Bo Chen <chen@chenirvine.org> wrote:\n>> Who decides bigness:\n>> Bigness seems to be relative to system resources.  Does the user crunch the\n>> numbers to determine if a file is big-file, or does git?  If the numbers are\n>> relative then should git query the system and make the determination?\n>>  Either way, once the system-resources are upgraded and formerly \"big-files\"\n>> are no longer considered \"big\" how is the previous history refactored tot\n>> behave \"non-big-file-like\"?  Conversely, if the system-resources are\n>> re-distributed so that formerly non-big files are now relatively big (ie,\n>> moved from powerful central server login to laptops), how is the history\n>> refactored to accommodate the newly-relative-bigness?\n>>\n>\n> In common sense, a file of tens of MBs should not be considered as a\n> big file, but a file of tens of GBs should definitely be considered as\n> a big file. I think one simple workable solution is to let the user\n> set the threshold of the big file.\n\nWe currently have core.bigFileThreshold = 512MB.\n\n> One complicate but intelligent\n> solution is to let git auto-config the threshold by evaluating current\n> computing resources in the running platform (a physical machine or\n> just a VM). As to the problem of migrating git in different platforms\n> which equip with different computing power, the git repo should also\n> keep tract of under what big file threshold a specific file is\n> handled.\n-- \nDuy\n"},{"id":"188291","messageId":"CA+M5ThTnd+TST6WsAn-Jd=Gb=1EWaJ+QbLMxXgtAVFNVqnRcMw@mail.gmail.com","threadId":"30077","inReplyTo":"CACsJy8DTegW78Qw7-T6uK_oZj2CELv57bbH6sU=bScHDesGYPQ@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Bo Chen","fromEmail":"chen@chenirvine.org","sentAt":"2012-04-01T23:30:51Z","receivedAt":"2012-04-01T23:30:51Z","isPatch":false,"sender":{"key":"chen@chenirvine.org","avatar":null},"body":"One question,  can anyone help me clear?\n\nMy .git/objects has 3 blobs, a, b, and c. a is a unique file, b and c\ntwo sequential versions of the same file. When I run \"git gc\", what\nexactly happens here, e.g., how exactly git (in the latest version)\ndelta compresses-the blobs here?\n\nAny help will be appreciated.\n\nBo\n\nOn Sun, Apr 1, 2012 at 12:22 AM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n> On Sun, Apr 1, 2012 at 4:27 AM, Bo Chen <chen@chenirvine.org> wrote:\n>>> Who decides bigness:\n>>> Bigness seems to be relative to system resources.  Does the user crunch the\n>>> numbers to determine if a file is big-file, or does git?  If the numbers are\n>>> relative then should git query the system and make the determination?\n>>>  Either way, once the system-resources are upgraded and formerly \"big-files\"\n>>> are no longer considered \"big\" how is the previous history refactored tot\n>>> behave \"non-big-file-like\"?  Conversely, if the system-resources are\n>>> re-distributed so that formerly non-big files are now relatively big (ie,\n>>> moved from powerful central server login to laptops), how is the history\n>>> refactored to accommodate the newly-relative-bigness?\n>>>\n>>\n>> In common sense, a file of tens of MBs should not be considered as a\n>> big file, but a file of tens of GBs should definitely be considered as\n>> a big file. I think one simple workable solution is to let the user\n>> set the threshold of the big file.\n>\n> We currently have core.bigFileThreshold = 512MB.\n>\n>> One complicate but intelligent\n>> solution is to let git auto-config the threshold by evaluating current\n>> computing resources in the running platform (a physical machine or\n>> just a VM). As to the problem of migrating git in different platforms\n>> which equip with different computing power, the git repo should also\n>> keep tract of under what big file threshold a specific file is\n>> handled.\n> --\n> Duy\n"},{"id":"188294","messageId":"CACsJy8CJK4gZwnrqkhq2DUqErS94X=99kvB9z=x9TefG=MrE4A@mail.gmail.com","threadId":"30077","inReplyTo":"CA+M5ThTnd+TST6WsAn-Jd=Gb=1EWaJ+QbLMxXgtAVFNVqnRcMw@mail.gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-04-02T01:00:22Z","receivedAt":"2012-04-02T01:00:22Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Apr 2, 2012 at 6:30 AM, Bo Chen <chen@chenirvine.org> wrote:\n> One question,  can anyone help me clear?\n>\n> My .git/objects has 3 blobs, a, b, and c. a is a unique file, b and c\n> two sequential versions of the same file. When I run \"git gc\", what\n> exactly happens here, e.g., how exactly git (in the latest version)\n> delta compresses-the blobs here?\n\nSee Documentation/technical/pack-heuristics.txt for how pack-objects\n(called by\"git gc\") decides to delta either b or c based on the other\none. Once it chooses, say, b to be delta against c, it generates delta\nusing diff-delta.c, then store the delta in either ref-delta or\nofs-delta format. The former stores sha-1 of c, the latter the offset\nof c in the pack.\n-- \nDuy\n"},{"id":"188354","messageId":"20120402210708.GA28926@sigill.intra.peff.net","threadId":"30077","inReplyTo":"4F772E48.3030708@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-04-02T21:07:08Z","receivedAt":"2012-04-02T21:07:08Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Mar 31, 2012 at 11:18:16AM -0500, Neal Kreitzinger wrote:\n\n> On 3/31/2012 6:02 AM, Sergio Callegari wrote:\n> >I wonder if it could make sense to have some pluggable mechanism for\n> > file splitting. Something under the lines of filters, so to say.\n> >Bupsplit can be a rather general mechanism, but large binaries that\n> >are containers (zip, jar, docx, tgz, pdf - seen as a collection of\n> >streams) may possibly be more conveniently split by their inherent\n> >components.\n> >\n> \n> gitattributes or gitconfig could configure the big-file handler for\n> specified files.  Known/supported filetypes like gif, png, zip, pdf,\n> etc., could be auto-configured by git.  Any\n> yet-unknown/yet-unsupported filetypes could be configured manually by\n> the user, e.g.\n> *.zgp=bigcontainer\n\nThis is a tempting route (and one I've even suggested myself before),\nbut I think ultimately it is a bad way to go. The problem is that\nsplitting is only half of the equation. Once you have split contents,\nyou have to use them intelligently, which means looking at the sha1s of\neach split chunk and discarding whole chunks as \"the same\" without even\nlooking at the contents.\n\nWhich means that it is very important that your chunking algorithm\nremain stable from version to version. A change in the algorithm is\ngoing to completely negate the benefits of chunking in the first place.\nSo something configurable, or something that is not applied consistently\n(because it depends on each user's git config, or even on the specific\nversion of a tool used) can end up being no help at all.\n\nProperly applied, I think a content-aware chunking algorithm could\nout-perform a generic one. But I think we need to first find out exactly\nhow well the generic algorithm can perform. It may be \"good enough\"\ncompared to the hassle that inconsistent application of a content-aware\nalgorithm will cause.  So I wouldn't rule it out, but I'd rather try the\nbup-style splitting first, and see how good (or bad) it is.\n\n-Peff\n"},{"id":"188356","messageId":"20120402214049.GB28926@sigill.intra.peff.net","threadId":"30077","inReplyTo":"4F77209A.8050607@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-04-02T21:40:49Z","receivedAt":"2012-04-02T21:40:49Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Mar 31, 2012 at 10:19:54AM -0500, Neal Kreitzinger wrote:\n\n> >Note that there are other problem areas with big files that can be\n> >worked on, too. For example, some people want to store 100 gigabytes\n> >in a repository.\n> \n> I take it that you have in mind a 100G set of files comprised entirely\n> of big-files that cannot be logically separated into smaller submodules?\n\nNot exactly. Two scenarios I'm thinking of are:\n\n  1. You really have 100G of data in the current version that doesn't\n     compress well (e.g., you are storing your music collection). You\n     can't afford to store two copies on your laptop (because you have a\n     fancy SSD, and 100G is expensive again).  You need the working tree\n     version, but it's OK to stream the repo version of a blob from the\n     network when you actually need it (mostly \"checkout\", assuming you\n     have marked the file as \"-diff\").\n\n  2. You have a 100G repository, but only 10G in the most recent\n     version (e.g., because you are doing game development and storing\n     the media assets). You want your clones to be faster and take less\n     space. You can do a shallow clone, but then you're never allowed to\n     look at old history. Instead, it would be nice to clone all of the\n     commits, trees, and small blobs, and then stream large blobs from\n     the network as-needed (again, mostly \"checkout\").\n\n> My understanding is that a main strategy for \"big files\" is to separate\n> your big-files logically into their own submodule(s) to keep them from\n> bogging down the not-big-file repo(s).\n\nThat helps people who want to work on the not-big parts by not forcing\nthem into the big parts (another solution would be partial clone, but\nmore on that in a minute). But it doesn't help people who actually want\nto work on the big parts; they would still have to fetch the whole\nbig-parts repository.\n\nFor splitting the big-parts people from the non-big-parts people, there\nhave been two suggestions: partial checkout (you have all the objects in\nthe repo, but only checkout some of them) and partial clone (you don't\nhave some of the objects in the repo). Partial checkout is a much easier\nproblem, as it is mostly about marking index entries as \"do not bother\nto check this out, and pretend that it is simply unmodified\". Partial\nclone is much harder, because it violates git's usual reachability\nrules. During a fetch, a client will say \"I have commit X\", which the\nserver can then assume means they have all of the ancestors of X, and\nall of the tree and blobs referenced by X and its ancestors.\n\nBut if a client can say \"yes, I have these objects, but I just don't\nwant to get them because it's expensive\", then partial checkout is\nsufficient. The non-big-parts people will clone, omitting the big\nobjects, and then do a partial checkout (to avoid fetching the objects\neven once).\n\nNote that some protocol extension is still needed for the client to tell\nthe server \"don't bother including objects X, Y, and Z in the packfile;\nI'll get them from my alternate big-object repo\". That can either be a\nlist of objects, or it can simply be \"don't bother with objects bigger\nthan N\".\n\n> >Because git is distributed, that means 100G in the repo database,\n> >and 100G in the working directory, for a total of 200G.\n> \n> I take it that you are implying that the 100G object-store size is due\n> to the notion that binary files cannot-be/are-not compressed well?\n\nIn this case, yes. But you could easily tweak the numbers to be 100G and\n150G. The point is that the data is stored twice, and even the\ncompressed version may be big.\n\n> >People in this situation may want to be able to store part of the\n> >repository database in a network-accessible location, trading some\n> >of the convenience of being fully distributed for the space savings.\n> >So another project could be designing a network-based alternate\n> >object storage system.\n> >\n> I take it you are implying a local area network with users git repos\n> on workstations?\n\nNot necessarily. Obviously if you are doing a lot of active work on the\nbig files, the faster your network, the better. But it could work at the\ninternet scale, too, if you don't actually fetch the big files\nfrequently (so part of a scheme like this would be making sure we avoid\naccessing big objects whenever we can; in practice, this is pretty easy,\nas git already tries to avoid accessing objects unnecessarily, because\nit's expensive even on the local end).\n\nYou can also cache a certain number of fetched objects locally. Assuming\nthere is some locality of the objects you ask about (e.g., because you\nare doing \"git checkout\" back and forth between two branches), this can\nhelp.\n\n> Some setups login to a linux server and have all their repos there.\n> The \"alternate objects\" does not need to network-based in that case.\n> It is \"local\", but local does not mean 20 people cloning the\n> alternate objects to their workstations.  It means one copy of\n> alternate objects, and twenty repos referencing that one copy.\n\nRight. This is the same concept, except over the network. So people's\nworking repositories are on their own workstations instead of a central\nserver. You could even do it today by network-mounting a filesystem and\npointing your alternates file at it. However, I think it's worth making\ngit aware that the objects are on the network for a few reasons:\n\n  1. Git can be more careful about how it handles the objects, including\n     when to fetch, when to stream, and when to cache. For example,\n     you'd want to fetch the manifest of objects and cache it in your\n     local repository, because you want fast lookups of \"do I have this\n     object\".\n\n  2. Providing remote filesystems on an Internet scale is a management\n     pain (and it's a pain for the user, too). My thought was that this\n     would be implemented on top of http (the connection setup cost is\n     negligible, since these objects would generally be large).\n\n  3. Usually alternate repositories are full repositories that meet the\n     connectivity requirements (so you could run \"git fsck\" in them).\n     But this is explicitly about taking just a few disconnected large\n     blobs out of the repository and putting them elsewhere. So it needs\n     a new set of tools for managing the upstream repository.\n\n-Peff\n"},{"id":"188361","messageId":"7vvclhdbew.fsf@alter.siamese.dyndns.org","threadId":"30077","inReplyTo":"20120402214049.GB28926@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-04-02T22:19:35Z","receivedAt":"2012-04-02T22:19:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>   1. You really have 100G of data in the current version that doesn't\n>      compress well (e.g., you are storing your music collection). You\n>      can't afford to store two copies on your laptop (because you have a\n>      fancy SSD, and 100G is expensive again).  You need the working tree\n>      version, but it's OK to stream the repo version of a blob from the\n>      network when you actually need it (mostly \"checkout\", assuming you\n>      have marked the file as \"-diff\").\n\nThis feels like a good candidate for an independent project that allows\nyou fuse-mount from a remote repository to give you an illusion that you\nhave a checkout of a specific version.  Such a remote fuse-server would be\nan application that is built using Git, but I do not think we are in any\nbusiness on the client end in such a setup.\n\nSo I'll write it off as a \"non-Git\" issue for now.\n\nThe other parts of your message is much more interesting.\n\n> Right. This is the same concept, except over the network. So people's\n> working repositories are on their own workstations instead of a central\n> server. You could even do it today by network-mounting a filesystem and\n> pointing your alternates file at it. However, I think it's worth making\n> git aware that the objects are on the network for a few reasons:\n>\n>   1. Git can be more careful about how it handles the objects, including\n>      when to fetch, when to stream, and when to cache. For example,\n>      you'd want to fetch the manifest of objects and cache it in your\n>      local repository, because you want fast lookups of \"do I have this\n>      object\".\n>\n>   2. Providing remote filesystems on an Internet scale is a management\n>      pain (and it's a pain for the user, too). My thought was that this\n>      would be implemented on top of http (the connection setup cost is\n>      negligible, since these objects would generally be large).\n>\n>   3. Usually alternate repositories are full repositories that meet the\n>      connectivity requirements (so you could run \"git fsck\" in them).\n>      But this is explicitly about taking just a few disconnected large\n>      blobs out of the repository and putting them elsewhere. So it needs\n>      a new set of tools for managing the upstream repository.\n\nOr you can split out the really large write-only blobs out of SCM control.\nEvery time you introduce a new blob, throw it verbatim in an append-only\ndirectory on a networked filesystem under some unique ID as its filename,\nand maintain a symlink into that networked filesystem under SCM control.\n\nI think git-annex already does something like that...\n"},{"id":"188400","messageId":"4F7AC9E2.60203@gmail.com","threadId":"30077","inReplyTo":"20120402210708.GA28926@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2012-04-03T09:58:58Z","receivedAt":"2012-04-03T09:58:58Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"On 02/04/2012 23:07, Jeff King wrote:\n>> gitattributes or gitconfig could configure the big-file handler for\n>> specified files.  Known/supported filetypes like gif, png, zip, pdf,\n>> etc., could be auto-configured by git.  Any\n>> yet-unknown/yet-unsupported filetypes could be configured manually by\n>> the user, e.g.\n>> *.zgp=bigcontainer\n> This is a tempting route (and one I've even suggested myself before),\n> but I think ultimately it is a bad way to go. The problem is that\n> splitting is only half of the equation. Once you have split contents,\n> you have to use them intelligently, which means looking at the sha1s of\n> each split chunk and discarding whole chunks as \"the same\" without even\n> looking at the contents.\n>\n> Which means that it is very important that your chunking algorithm\n> remain stable from version to version. A change in the algorithm is\n> going to completely negate the benefits of chunking in the first place.\n> So something configurable, or something that is not applied consistently\n> (because it depends on each user's git config, or even on the specific\n> version of a tool used) can end up being no help at all.\nIsn't this the same with filters? The clean algorithms should remain \nstable from\nversion to version. Filters are often perceived as simpler, so that this \nstability seems easier to achieve, but it is not necessarily the case.\n> Properly applied, I think a content-aware chunking algorithm could\n> out-perform a generic one. But I think we need to first find out exactly\n> how well the generic algorithm can perform. It may be \"good enough\"\n> compared to the hassle that inconsistent application of a content-aware\n> algorithm will cause.\nAbsolutely true, but why not giving freedom to the user to chose? Git \ncould provide the bupsplit mechanism and at the same time have a means \nso that the user can plug in a different machinery for specific file \ntypes.  In this case, it is the user responsibility to do it right.\n\nOne could have a special 'filter' for splitting/unsplitting. Say\n\n[splitfilter \"XXX\"]\n     split = xxx\n     unsplit = uxxx\n\nxxx is given the file to split on stdin and returns on stdout a stream \nmade of an index header and the concatenation of the parts in which the \nfile should be split. For unsplitting uxxx is given on stdin the index \nand the concatenation of parts and returns on stdout the binary file.\n\nbupsplit and bupunsplit could be built in, with other tools being user \nprovided.  If the users gets them wrong it is ultimately his/her \nresponsibility. In the end, the user is given even 'rm' isn't he/she? \nGit could provide a header file defining the index header format to help \nthe coding of the alternative, more specific splitters. If people devise \nsome of them that look promising, they can probably be collected in contrib.\n\nPossibly, the index header could comprise starting positions for the \nvarious parts in the stream, but also 'names' for them. This would let \nreusing blob and tree objects to physically store the various parts. For \nbupsplit, names could be flat (e.g. sequence numbers like 0000, 0001). \nFor files that are container, they could reflect the inner names. \nPerspectively, one could even devise specific diff tools for these \n'special' trees of split-object components. With this, when storing say \na very large zip file in git, these tools could help saying things like \n'from version x to version y, only that specific part in the zip file \nhas changed'.\n\nSergio\n"},{"id":"188401","messageId":"20120403100704.GC14483@sigill.intra.peff.net","threadId":"30077","inReplyTo":"7vvclhdbew.fsf@alter.siamese.dyndns.org","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-04-03T10:07:04Z","receivedAt":"2012-04-03T10:07:04Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Apr 02, 2012 at 03:19:35PM -0700, Junio C Hamano wrote:\n\n> >   1. You really have 100G of data in the current version that doesn't\n> >      compress well (e.g., you are storing your music collection). You\n> >      can't afford to store two copies on your laptop (because you have a\n> >      fancy SSD, and 100G is expensive again).  You need the working tree\n> >      version, but it's OK to stream the repo version of a blob from the\n> >      network when you actually need it (mostly \"checkout\", assuming you\n> >      have marked the file as \"-diff\").\n> \n> This feels like a good candidate for an independent project that allows\n> you fuse-mount from a remote repository to give you an illusion that you\n> have a checkout of a specific version.  Such a remote fuse-server would be\n> an application that is built using Git, but I do not think we are in any\n> business on the client end in such a setup.\n\nI think this is backwards. The primary item you want on the laptop is\nthe working directory, because you will be accessing and manipulating\nthe files. That must always work, whether the network is connected or\nnot. You occasionally will want to perform git operations. Most of these\nshould succeed when disconnected, but it's OK for some operations (like\nchecking out an older version of a large blob) to fail.\n\nBut if you are mounting a remote repository and pretending that you have\na local checkout, then just accessing the files either requires a\nnetwork, or you end up caching most of the remote repository.\n\nIt would make more sense to me to clone a bare repository of what's\nupstream, and then fuse-mount the local bare repository to provide a\nfake working directory. And I believe somebody made such a fuse\nfilesystem in the early days of git. However, I recall that it was\nread-only. I'm not sure how you would handle writing to the git-mounted\ndirectory.\n\n> Or you can split out the really large write-only blobs out of SCM control.\n> Every time you introduce a new blob, throw it verbatim in an append-only\n> directory on a networked filesystem under some unique ID as its filename,\n> and maintain a symlink into that networked filesystem under SCM control.\n> \n> I think git-annex already does something like that...\n\nYes, and git-media basically does this, too. But it's awful to use,\nbecause the user has to be constantly aware of these special links and\nmanaging them. You can't just store a symlink into the networked\nfilesystem. For one thing, the path may be different on each client\nmachine, so a simple symlink doesn't work.  For another, symlinks into a\nblob repository mean that the files must be read-only (since they are\nbasically blob-equivalents). So you don't really get your own copy of\nthe file; you can _replace_ it and update the symlink, but you can't\nactually modify it.\n\nSo what things like git-media end up doing is to try to insert\nthemselves between git and the user, and transparently convert the file\ninto its unique ID on \"git add\" and tweak the working directory to\ncontain the actual file on checkout. And it kind of works, but there are\na lot of rough edges (I don't recall the details, but they came up in\npast discussions; clean and smudge filters almost get you there, but not\nquite).\n\nBasically what I'm proposing to do is to just move that logic into git\nitself, so it can just happen at the blob storage level. I don't think\nit would even be that much code inside git; you'd want the interface to\nbe pluggable, so all of the heavy lifting would happen inside of a\nhelper (so really, this isn't necessarily even \"network alternates\" as\nmuch as \"pluggable alternates\").\n\n-Peff\n"},{"id":"188920","messageId":"4F84DD60.20903@gmail.com","threadId":"30077","inReplyTo":"20120402210708.GA28926@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-04-11T01:24:48Z","receivedAt":"2012-04-11T01:24:48Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/2/2012 4:07 PM, Jeff King wrote:\n\n> ...I think we need to first find out exactly\n> how well the generic algorithm can perform. It may be \"good enough\"\n> compared to the hassle that inconsistent application of a content-aware\n> algorithm will cause.  So I wouldn't rule it out, but I'd rather try the\n> bup-style splitting first, and see how good (or bad) it is.\n>\n(I read bup DESIGN doc to see what bup-style splitting is.) When you use \nbup delta technology in git.git I take it that you will use it for \nbig-worktree-files *and* big-history-files (not-big-worktree-files that \nare not xdelta delta-friendly)?  IOW, all binaries plus \nbig-text-worktree-files.  Otherwise, small binaries will become large \nhistories.\n\nIf small binaries are not going to be bup-delta-compressed, then what \nabout using xxd to convert the binary to text and then xdelta \ncompressing the hex dump to achieve efficient delta compression in the \npack file?  You could convert the hexdump back to binary with xxd for \ncheckout and such.\n\nMaybe small binaries do xdelta well and the above is a moot point.  This \nis all theory to me, but the reality is looming over my head since most \nof the components I should be tracking are binaries small (large \nhistory?) and big (but am not yet because of \"big-file\" concerns -- I \ndon't want to have to refactor my vast git ecosystem with filter branch \nlater because I slammed binaries into the main project or superproject \nwithout proper systems programming (I'm not sure what the c/linux term \nis for 'systems programming', but in the mainframe world it meant making \nsure everything was configured for efficient performance)).\n\nNow that I say that out loud I guess a superproject with binaries in \nseparate repos could be easily refactored by creating new efficient \nrepos and making a new commit that points to them instead of the old \ninefficient repos.  That way, when someone checks out the binary repo \n(submodule) into their worktree they get the new efficiency instead of \nthe old inefficiency.  Over time, as folks are less likely to check out \nold stuff the old inefficiency goes away on its own.  I think. \n(Submodules are mostly theory to me at this point also.)\n\nv/r,\nneal\n"},{"id":"188930","messageId":"20120411060357.GA15805@burratino","threadId":"30077","inReplyTo":"4F84DD60.20903@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-11T06:04:04Z","receivedAt":"2012-04-11T06:04:04Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Neal Kreitzinger wrote:\n\n> Maybe small binaries do xdelta well and the above is a moot point.\n\nIf I am reading it correctly, diff-delta copes fine with smallish\nbinary files that have not changed much.  Converting to hex would only\nhurt.\n\nI would suggest tracking source code instead of binaries if possible,\nthough.\n\nJonathan\n"},{"id":"188985","messageId":"4F85B17E.4080005@gmail.com","threadId":"30077","inReplyTo":"20120411060357.GA15805@burratino","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-04-11T16:29:50Z","receivedAt":"2012-04-11T16:29:50Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/11/2012 1:04 AM, Jonathan Nieder wrote:\n> Neal Kreitzinger wrote:\n>\n>> Maybe small binaries do xdelta well and the above is a moot point.\n>\n> If I am reading it correctly, diff-delta copes fine with smallish\n> binary files that have not changed much.  Converting to hex would\n> only hurt.\n>\nHow do I check the history size of a binary?  IOW, how to I check the\nsize of the sum of all the delta-compressions and root blob of a binary?\n  That way I can sample different binary types to get a symptomatic idea\nof how well they are delta compressing.  I suspect that compiled\nbinaries will compress well (efficient history) and graphics files may\nnot compress well (large history).\n\nv/r,\nneal\n"},{"id":"188986","messageId":"4F85B2CB.3070002@gmail.com","threadId":"30077","inReplyTo":"20120411060357.GA15805@burratino","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-04-11T16:35:23Z","receivedAt":"2012-04-11T16:35:23Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/11/2012 1:04 AM, Jonathan Nieder wrote:\n> Neal Kreitzinger wrote:\n>\n>> Maybe small binaries do xdelta well and the above is a moot point.\n>\n> If I am reading it correctly, diff-delta copes fine with smallish\n> binary files that have not changed much.  Converting to hex would\n> only hurt.\n>\n> I would suggest tracking source code instead of binaries if\n> possible, though.\n>\nIs there some documentation out there that lists the common binary\nformats (e.g., pdf, docx, gif, jpg, png, bmp, mpeg, mp3, zip, \nc-binaries, java stuff, website stuff, etc.) and explains their nature \n(container, compressed, encrypted, etc.), how well they currently delta \nin git.git within specified size boundaries and use cases (pdfs with \nonly plain text vs. pdfs with graphics, tables, etc.) so git users can \nreference that to make their git repo/superproject design decisions?\n\nv/r,\nneal\n"},{"id":"188987","messageId":"4F85B4E7.7090603@gmail.com","threadId":"30077","inReplyTo":"20120411060357.GA15805@burratino","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-04-11T16:44:23Z","receivedAt":"2012-04-11T16:44:23Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/11/2012 1:04 AM, Jonathan Nieder wrote:\n> Neal Kreitzinger wrote:\n>\n>> Maybe small binaries do xdelta well and the above is a moot point.\n>\n> If I am reading it correctly, diff-delta copes fine with smallish\n> binary files that have not changed much.\n>\n> I would suggest tracking source code instead of binaries if\n> possible, though.\n>\nI suppose the original \"source\" in git (linux kernel) was so low level\nthat it had no graphics files.  However, most projects are end-user\nprojects and have graphics so I would think that tracking them is a\nnormal expected use of git to version your software.  If you're going to \ndo that then there shouldn't be a problem tracking other binaries that \nare static constants across all servers (as opposed to user edited \ncontent like databases).  I would consider this subset of \"binaries\" to \nbe the expected domain of git revision control for software, ie, gui \nsoftware.  Graphics files for your app are \"source\".  The binary is all \nyou have.  It's the \"source\" that you edit to make changes.\n\nMaybe I'm missing something here.  Maybe graphics files are \"container\" \nfiles and that makes them a problem.\n\nv/r,\nneal\n"},{"id":"188999","messageId":"20120411172034.GE4248@burratino","threadId":"30077","inReplyTo":"4F85B4E7.7090603@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-11T17:20:34Z","receivedAt":"2012-04-11T17:20:34Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Neal Kreitzinger wrote:\n\n>                              Graphics files for your app are\n> \"source\".  The binary is all you have.\n\nOften there is source in SVG or some other simple editable format that\ngets lossily compiled to PNG or JPEG compressed raster graphics.\n"},{"id":"189009","messageId":"4F85CC0A.1020602@gmail.com","threadId":"30077","inReplyTo":"20120411060357.GA15805@burratino","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-04-11T18:23:06Z","receivedAt":"2012-04-11T18:23:06Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/11/2012 1:04 AM, Jonathan Nieder wrote:\n>\n> I would suggest tracking source code instead of binaries if\n> possible, though.\n>\nReasons why we want to track binaries:\n(1) Standard Targets: Our deployment is assembly line style because our\ntarget servers are under our control.\n(2) Copy vs. Recompile:  We run certain \"supported\" linux distro\nversions on our target servers so we can just put our binaries on them\ninstead of recompiling.\n(3) In-house-Source Compiled Binaries:  For our particular proprietary\n(third-party) source language the binaries run on top of a runtime that\nruns on top of the O/S so that makes the need to recompile on a server a\nnon-issue.  We use xxd and compile listings to \"diff\" our compiled\nbinaries to detect missing copybook and data dictionary dependencies \n(missed recompiles), unnecessary recompiles (you didn't really change \nwhat you thought you changed), and miscompiles.  We do this compiled \nbinary validation in git branches and then diff the branches to detect \nthe discrepancies.\n(4) Proprietary-Third-Party Binaries (no source) Versioning:  For our\nthird party binaries we don't have the source.  The are distributed as\nself-extracting-executables.  Changes to third party binaries are\nrelatively infrequent but frequent enough to cause confusion and\ntherefore need to be tracked.\n(5) Graphics \"Source\" Versioning:  Our graphics files are part of our\nsoftware and changes need to be tracked.\n(6) O/S Versioning:  Our linux distro is tracked in a bazaar repo so\nI'm thinking we should be able to track it in a git repo instead.  The\nassembly line just deploys the payload to a new server instead of doing\nmanual install.\n(7) Superproject tracking of \"Super-release\":  The above subsystems are\nrelated in varying degrees (dependent).  A superproject can associate\nall the versions that comprise a \"super\" release of the various\nsubsystem version dependencies.\n\nWhile some of the reasons above may be non-normative for some git-users,\nI think that a large portion (if not the majority) of git-users will\nfind some subset of the above reasons normative for their use-cases\n(namely reasons 5 and 4) therefore making the necessity for binary\ntracking normative for git-users in general.\n\nv/r,\nneal\n"},{"id":"189016","messageId":"7vty0qje4g.fsf@alter.siamese.dyndns.org","threadId":"30077","inReplyTo":"20120411172034.GE4248@burratino","subject":"Re: GSoC - Some questions on the idea of","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-04-11T18:51:43Z","receivedAt":"2012-04-11T18:51:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> Neal Kreitzinger wrote:\n>\n>>                              Graphics files for your app are\n>> \"source\".  The binary is all you have.\n>\n> Often there is source in SVG or some other simple editable format that\n> gets lossily compiled to PNG or JPEG compressed raster graphics.\n\nYou could have just underlined \"if possible\" part in your message and\nended this thread, that seems to be needlessly continuing.\n"},{"id":"189018","messageId":"20120411190351.GF4248@burratino","threadId":"30077","inReplyTo":"7vty0qje4g.fsf@alter.siamese.dyndns.org","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2012-04-11T19:03:51Z","receivedAt":"2012-04-11T19:03:51Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n\n> You could have just underlined \"if possible\" part in your message\n\nYes, that's a more important point than the one I responded to. :)\n"},{"id":"189046","messageId":"20120411213522.GA28199@sigill.intra.peff.net","threadId":"30077","inReplyTo":"4F84DD60.20903@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-04-11T21:35:22Z","receivedAt":"2012-04-11T21:35:22Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 10, 2012 at 08:24:48PM -0500, Neal Kreitzinger wrote:\n\n> (I read bup DESIGN doc to see what bup-style splitting is.) When you\n> use bup delta technology in git.git I take it that you will use it\n> for big-worktree-files *and* big-history-files\n\nI'm not sure what those terms mean. We are talking about files at the\nblob level. So they are either big or not big. We don't know how they\nwill delta, or what their histories will be like.\n\n> (not-big-worktree-files that are not xdelta delta-friendly)?\n> IOW, all binaries plus big-text-worktree-files.  Otherwise, small\n> binaries will become large histories.\n\nFiles that don't delta won't be helped by splitting, as it is just\nanother form of finding deltas (in fact, it should produce worse results\nthan xdelta, because it works with larger granularity; its advantage is\nthat it is not as memory or CPU-hungry as something like xdelta).\n\nSo you really only want to use this for files that are too big to\npractically run through the regular delta algorithm. And if you can\navoid it on files that will never delta well, you are better off\n(because it adds storage overhead over a straight blob).\n\nThe first part is easy: only do it for files that are so big that you\ncan't run the regular delta algorithm. So since your only alternative is\ndoing nothing, you only have to perform better than nothing. :)\n\nThe second part is harder. We generally don't know that a file doesn't\ndelta well until we have two versions of it to try[1]. And that's where\nsome domain-specific knowledge can come in (e.g., knowing that a file is\ncompressed video, and that future versions are likely to differ in the\nvideo content). But sometimes the results can be surprising. I keep a\nrepository of photos and videos, carefully annotated via exif tags. If\nthe media content changes, the results won't delta well. But if I change\nthe exif tags, they _do_ delta very well. So whether something like\nbupsplit is a win depends on the exact update patterns.\n\n[1] I wonder if you could do some statistical analysis on the randomness\n    of the file content to determine this. That is, things which look\n    very random are probably already heavily compressed, and are not\n    going to compress further. You might guess that to mean that they\n    will not delta well, either. And sometimes that is true. But the\n    example I gave above violates it (most of the file is random, but\n    the _changes_ from version to version will not be random, and that\n    is what the delta is compressing).\n\n> If small binaries are not going to be bup-delta-compressed, then what\n> about using xxd to convert the binary to text and then xdelta\n> compressing the hex dump to achieve efficient delta compression in\n> the pack file?  You could convert the hexdump back to binary with xxd\n> for checkout and such.\n\nThat wouldn't help. You are only trading the binary representation for a\nless efficient one. But the data patterns will not change. The\nredundancy you introduced in the first step may mostly come out via\ncompression, but it will never be a net win. I'm sure if I were a better\ncomputer scientist I could write you some proof involving Shannon\nentropy. But here's a fun experiment:\n\n  # create two files, one very compressible and one not very\n  # compressible\n  dd if=/dev/zero of=boring.bin bs=1M count=1\n  dd if=/dev/urandom of=rand.bin bs=1M count=1\n\n  # now make hex dumps of each, and compress the original and the hex\n  # dump\n  for i in boring rand; do\n    xxd <$i.bin >$i.hex\n    for j in bin hex; do\n      gzip -c <$i.$j >$i.$j.gz\n    done\n  done\n\n  # and look at the results\n  du {boring,rand}.*\n\nI get:\n\n  1024    boring.bin\n  4       boring.bin.gz\n  4288    boring.hex\n  188     boring.hex.gz\n  1024    rand.bin\n  1028    rand.bin.gz\n  4288    rand.hex\n  2324    rand.hex.gz\n\nSo you can see that the thing that compresses well will do so in\neither representation, but the end result is a net loss with the less\nefficient representation. Whereas the thing that does not compress well\nwill achieve a better compression ratio in its text form, but will still\nbe a net loss. The reason is that you are just compressing out all of\nthe redundant bits.\n\nYou might observe that this is using gzip, not xdelta. But I think from\nan information theory standpoint, they are two sides of the same coin\n(e.g., you could consider a delta between two things to be equivalent to\nconcatenating them and compressing the result). You should be able to\ndesign a similar experiment with xdelta.\n\n> Maybe small binaries do xdelta well and the above is a moot point.\n\nSome will and some will not. But it has nothing to do with whether they\nare binary, and everything to do with the type of content they store (or\nif binariness does matter, then our delta algorithms should be\nimproved).\n\n> This is all theory to me, but the reality is looming over my head\n> since most of the components I should be tracking are binaries small\n> (large history?) and big (but am not yet because of \"big-file\"\n> concerns -- I don't want to have to refactor my vast git ecosystem\n> with filter branch later because I slammed binaries into the main\n> project or superproject without proper systems programming (I'm not\n> sure what the c/linux term is for 'systems programming', but in the\n> mainframe world it meant making sure everything was configured for\n> efficient performance)).\n\nOne of the things that makes bup not usable as-is for git is that it\nfundamentally changes the object identities. It would be very easy for\n\"git add\" to bupsplit a file into a tree, and store that tree using git\n(in fact, that is more or less how bup works).  But that means that the\nresulting object sha1 is going to depend on the splitting choices made.\nInstead, we want to consider the split version of an object to be simply\nan on-disk representation detail. Just as it is a representation detail\nthat some objects are stored in delta-encoding inside packs, versus as\nloose objects; the sha1 of the object is the same, and we can\nreconstruct it byte-for-byte when we want to.\n\nSo properly implemented, no, you would not have to ever filter-branch to\ntweak these settings. You might have to do a repack to see the gains\n(because you want to delete the old non-split representation you have in\nyour pack and replace it with a split representation), but that is\ntransparent to git's abstract data model.\n\n-Peff\n"},{"id":"189052","messageId":"20120411220955.GB28199@sigill.intra.peff.net","threadId":"30077","inReplyTo":"4F85B17E.4080005@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-04-11T22:09:55Z","receivedAt":"2012-04-11T22:09:55Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 11, 2012 at 11:29:50AM -0500, Neal Kreitzinger wrote:\n\n> How do I check the history size of a binary?  IOW, how to I check the\n> size of the sum of all the delta-compressions and root blob of a binary?\n>  That way I can sample different binary types to get a symptomatic idea\n> of how well they are delta compressing.  I suspect that compiled\n> binaries will compress well (efficient history) and graphics files may\n> not compress well (large history).\n\nI don't think there is a simple command to do it. You have to correlate\nblobs at a given path with objects in the packs yourself. You can script\nit like:\n\n  # get the delta stats from every pack; you only need to do this part\n  # once for a given history state. And obviously you would want to\n  # repack before doing it.\n  for i in .git/objects/pack/*.pack; do\n    git verify-pack -v $i;\n  done |\n  perl -lne '\n    # format is: sha1 type size size-in-pack offset; pick out only the\n    # thing we care about: size in pack\n    /^([0-9a-f]{40}) \\S+\\s+\\d+ (\\d+)/ and print \"$1 $2\";\n  ' |\n  sort >delta-stats\n\n\n  # then you can do this for every path you are interested in.\n\n  # First, get the list of blobs at that path (and follow renames, too).\n  # The second line is picking the \"after\" sha1 from the --raw output.\n  git log --follow --raw --no-abbrev $path |\n  perl -lne '/:\\S+ \\S+ \\S{40} (\\S{40})/ and print $1' |\n  sort -u >blobs\n\n  # Then find the delta stats for those blobs\n  join blobs delta-stats\n\nwhich should give you the stored size of each version of a file.\n\n-Peff\n"},{"id":"189130","messageId":"4F872D24.8010609@gmail.com","threadId":"30077","inReplyTo":"20120411213522.GA28199@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-04-12T19:29:40Z","receivedAt":"2012-04-12T19:29:40Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/11/2012 4:35 PM, Jeff King wrote:\n> On Tue, Apr 10, 2012 at 08:24:48PM -0500, Neal Kreitzinger wrote:\n>\n>> This is all theory to me, but the reality is looming over my head\n>> since most of the components I should be tracking are binaries small\n>> (large history?) and big (but am not yet because of \"big-file\"\n>> concerns -- I don't want to have to refactor my vast git ecosystem\n>> with filter branch later because I slammed binaries into the main\n>> project or superproject without proper systems programming (I'm not\n>> sure what the c/linux term is for 'systems programming', but in the\n>> mainframe world it meant making sure everything was configured for\n>> efficient performance)).\n> So properly implemented, no, you would not have to ever filter-branch to\n> tweak these settings. You might have to do a repack to see the gains\n> (because you want to delete the old non-split representation you have in\n> your pack and replace it with a split representation), but that is\n> transparent to git's abstract data model.\n>\nI'm likely going to have to slam graphics files into the main repo in \nthe very near future.  It sounds like once git.git is updated for \nbig-file optimization I can just upgrade to that git version and repack \nto get the benefits.  Any idea when that version of git will come out \nrelease number wise and calendar wise?\n\n(Don't read this next part if you just ate or are eating or drinking.  \nYou may throw-up from nausea or choke from laughing.)\n(I am forced to deal with a mandated/micromanaged change control menu \ndesign from the powers-that-be that is based on cvs workflow and to \nwipe-your-nose-for-you.  It can't even cope with branches much less \nsubmodules so in that context there isn't time to implement the graphics \ntracking as a submodule.  This change control menu is designed to \nreplace cvs commands with equivalent-results git-command sequences.  \nWhile there are many git users who import from svn into git, do their \nwork in git, and then export back into svn to get work done, ironically \nI am probably the only git user who has to import from git \n(powers-that-be mandated cvs-style menu controlled git-repo) into git \n(separate normal git repo and commandline), do the work in normal git, \nand then export it back into git (cvs-style menu controlled git-repo).)\n\nv/r,\nneal\n"},{"id":"189143","messageId":"20120412210315.GC21018@sigill.intra.peff.net","threadId":"30077","inReplyTo":"4F872D24.8010609@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-04-12T21:03:15Z","receivedAt":"2012-04-12T21:03:15Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Apr 12, 2012 at 02:29:40PM -0500, Neal Kreitzinger wrote:\n\n> I'm likely going to have to slam graphics files into the main repo in\n> the very near future.  It sounds like once git.git is updated for\n> big-file optimization I can just upgrade to that git version and\n> repack to get the benefits.\n\nDepending on the size and number of the files, git may handle them just\nfine. They don't delta well, which means they will bloat your object db\na bit, but if you are talking about a hundreds of megabytes total, it is\nprobably not that big a deal.\n\n> Any idea when that version of git will come out release number wise\n> and calendar wise?\n\nNo idea. This is still in the discussion and experimenting stage. It may\nnot even happen.\n\n-Peff\n"},{"id":"189145","messageId":"4F87443F.10207@gmail.com","threadId":"30077","inReplyTo":"4F872D24.8010609@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-04-12T21:08:15Z","receivedAt":"2012-04-12T21:08:15Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/12/2012 2:29 PM, Neal Kreitzinger wrote:\n>\n> ...ironically I am probably the only git user who has to import from \n> git (powers-that-be mandated cvs-style menu controlled git-repo) into \n> git (separate normal git repo and commandline), do the work in normal \n> git, and then export it back into git (cvs-style menu controlled \n> git-repo).)\n>\n>\naka, git-cotton-picking  ;-)\n\nv/r,\nneal\n"},{"id":"189219","messageId":"CA+M5ThRJYeVgHtKjuDbpDDMUv+k33cVZkeJ887_TnsB9BHAzVg@mail.gmail.com","threadId":"30077","inReplyTo":"4F872D24.8010609@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Bo Chen","fromEmail":"chen@chenirvine.org","sentAt":"2012-04-13T21:36:28Z","receivedAt":"2012-04-13T21:36:28Z","isPatch":false,"sender":{"key":"chen@chenirvine.org","avatar":null},"body":"On Thu, Apr 12, 2012 at 3:29 PM, Neal Kreitzinger\n<nkreitzinger@gmail.com> wrote:\n> On 4/11/2012 4:35 PM, Jeff King wrote:\n>>\n>> On Tue, Apr 10, 2012 at 08:24:48PM -0500, Neal Kreitzinger wrote:\n>>\n>>> This is all theory to me, but the reality is looming over my head\n>>> since most of the components I should be tracking are binaries small\n>>> (large history?) and big (but am not yet because of \"big-file\"\n>>> concerns -- I don't want to have to refactor my vast git ecosystem\n>>> with filter branch later because I slammed binaries into the main\n>>> project or superproject without proper systems programming (I'm not\n>>> sure what the c/linux term is for 'systems programming', but in the\n>>> mainframe world it meant making sure everything was configured for\n>>> efficient performance)).\n>>\n>> So properly implemented, no, you would not have to ever filter-branch to\n>>\n>> tweak these settings. You might have to do a repack to see the gains\n>> (because you want to delete the old non-split representation you have in\n>> your pack and replace it with a split representation), but that is\n>> transparent to git's abstract data model.\n>>\n> I'm likely going to have to slam graphics files into the main repo in the\n> very near future.  It sounds like once git.git is updated for big-file\n> optimization I can just upgrade to that git version and repack to get the\n> benefits.  Any idea when that version of git will come out release number\n> wise and calendar wise?\n>\n> (Don't read this next part if you just ate or are eating or drinking.  You\n> may throw-up from nausea or choke from laughing.)\n> (I am forced to deal with a mandated/micromanaged change control menu design\n> from the powers-that-be that is based on cvs workflow and to\n> wipe-your-nose-for-you.  It can't even cope with branches much less\n> submodules so in that context there isn't time to implement the graphics\n> tracking as a submodule.  This change control menu is designed to replace\n> cvs commands with equivalent-results git-command sequences.  While there are\n> many git users who import from svn into git, do their work in git, and then\n> export back into svn to get work done, ironically I am probably the only git\n> user who has to import from git (powers-that-be mandated cvs-style menu\n> controlled git-repo) into git (separate normal git repo and commandline), do\n> the work in normal git, and then export it back into git (cvs-style menu\n> controlled git-repo).)\n\nIt seems that this is not directly related to the big-file support\nissue. Maybe it is better to discuss it in a new thread ^-^\n>\n> v/r,\n> neal\n"},{"id":"189312","messageId":"20120415021550.GA24102@sigill.intra.peff.net","threadId":"30077","inReplyTo":"4F8A2EBD.1070407@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-04-15T02:15:50Z","receivedAt":"2012-04-15T02:15:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Apr 14, 2012 at 09:13:17PM -0500, Neal Kreitzinger wrote:\n\n> Does a file's delta-compression efficiency in the pack-file directly\n> correlate to its efficiency of transmission size/bandwidth in a\n> git-fetch and git-push?  IOW, are big-files also a problem for\n> git-fetch and git-push by taking too long in a remote transfer?\n\nYes. The on-the-wire format is a packfile. We create a new packfile on\nthe fly, so we may find new deltas (e.g., between objects that were\nstored on disk in two different packs), but we will mostly be reusing\ndeltas from the existing packs.\n\nSo any time you improve the on-disk representation, you are also\nimproving the network bandwidth utilization.\n\n-Peff\n"},{"id":"189313","messageId":"4F8A3381.803@gmail.com","threadId":"30077","inReplyTo":"20120415021550.GA24102@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-04-15T02:33:37Z","receivedAt":"2012-04-15T02:33:37Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/14/2012 9:15 PM, Jeff King wrote:\n> On Sat, Apr 14, 2012 at 09:13:17PM -0500, Neal Kreitzinger wrote:\n>\n>> Does a file's delta-compression efficiency in the pack-file directly\n>> correlate to its efficiency of transmission size/bandwidth in a\n>> git-fetch and git-push?  IOW, are big-files also a problem for\n>> git-fetch and git-push by taking too long in a remote transfer?\n> Yes. The on-the-wire format is a packfile. We create a new packfile on\n> the fly, so we may find new deltas (e.g., between objects that were\n> stored on disk in two different packs), but we will mostly be reusing\n> deltas from the existing packs.\n>\n> So any time you improve the on-disk representation, you are also\n> improving the network bandwidth utilization.\n>\nWe use git to transfer database files from the dev server to \nqa-servers.  Sometimes these barf for some reason and I get called to \nremediate.  I assumed the user closed their session prematurely because \nit was \"taking too long\".  However, now I'm wondering if the git-pull \n--ff-only is dying on its own due to the big-files.  It could be that on \na qa-server that hasn't updated database files in awhile they are \npulling way more than another qa-server that does their git-pull more \nrequently.  How would I go about troubleshooting this?  Is there some \nlog files I would look at?  (I'm using git 1.7.1 compiled with git \nmakefile on rhel6.)  When I go to remediate do git-reset --hard to clear \nout the barfed worktree/index and then run git-pull --ff-only manually \nand it always works.  I'm not sure if that proves it wasn't git that \nbarfed the first time.  Maybe the first time git brought some stuff over \nand barfed because it bit off more than it could chew, but the second \ntime its really having to chew less food because it already chewed some \nof it the first time and therefore works the second time.\n\nv/r,\nneal\n"},{"id":"189408","messageId":"20120416145450.GA14724@sigill.intra.peff.net","threadId":"30077","inReplyTo":"4F8A3381.803@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-04-16T14:54:50Z","receivedAt":"2012-04-16T14:54:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Apr 14, 2012 at 09:33:37PM -0500, Neal Kreitzinger wrote:\n\n> We use git to transfer database files from the dev server to\n> qa-servers.  Sometimes these barf for some reason and I get called to\n> remediate.  I assumed the user closed their session prematurely\n> because it was \"taking too long\".  However, now I'm wondering if the\n> git-pull --ff-only is dying on its own due to the big-files.  It\n> could be that on a qa-server that hasn't updated database files in\n> awhile they are pulling way more than another qa-server that does\n> their git-pull more requently.  How would I go about troubleshooting\n> this?  Is there some log files I would look at?  (I'm using git 1.7.1\n> compiled with git makefile on rhel6.)\n\nNo, git doesn't keep logfiles. Errors go to stderr. So look wherever the\nstderr for your git sessions is going (if you are doing this via cron\njob or something, then that is outside the scope of git).\n\n> When I go to remediate do git-reset --hard to clear out the barfed\n> worktree/index and then run git-pull --ff-only manually and it always\n> works.  I'm not sure if that proves it wasn't git that barfed the\n> first time.  Maybe the first time git brought some stuff over and\n> barfed because it bit off more than it could chew, but the second time\n> its really having to chew less food because it already chewed some of\n> it the first time and therefore works the second time.\n\nTry \"git pull --no-progress\" and see if it still works. If the server\nhas a very long delta-compression phase, there will be no output\ngenerated for a while, which could cause intermediate servers to hang up\n(git won't do this, but if, for example, you are pulling over\ngit-over-http and there is a reverse proxy in the middle, it may hit a\ntimeout). If the automated pulls are happening from a cron job, then\nthey won't have a terminal and progress-reporting will be off by\ndefault.\n\n-Peff\n"},{"id":"191357","messageId":"4FAC367E.8070006@gmail.com","threadId":"30077","inReplyTo":"20120415021550.GA24102@sigill.intra.peff.net","subject":"Re: GSoC - Some questions on the idea of","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-05-10T21:43:26Z","receivedAt":"2012-05-10T21:43:26Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 4/14/2012 9:15 PM, Jeff King wrote:\n> On Sat, Apr 14, 2012 at 09:13:17PM -0500, Neal Kreitzinger wrote:\n>\n>> Does a file's delta-compression efficiency in the pack-file directly\n>> correlate to its efficiency of transmission size/bandwidth in a\n>> git-fetch and git-push?  IOW, are big-files also a problem for\n>> git-fetch and git-push by taking too long in a remote transfer?\n> Yes. The on-the-wire format is a packfile. We create a new packfile on\n> the fly, so we may find new deltas (e.g., between objects that were\n> stored on disk in two different packs), but we will mostly be reusing\n> deltas from the existing packs.\n>\n> So any time you improve the on-disk representation, you are also\n> improving the network bandwidth utilization.\n>\nThe git-clone manpage says you can use the rsync protocol for the url.  \nIf you use rsync:// as your url for your remote does that get you the \nrsync delta-transfer algorithm efficiency for the network bandwidth \nutilization part (as opposed to the on-disk representation part)?  (I'm \nnew to rsync.)\n\nv/r,\nneal\n"},{"id":"191369","messageId":"20120510223916.GB31116@sigill.intra.peff.net","threadId":"30077","inReplyTo":"4FAC367E.8070006@gmail.com","subject":"Re: GSoC - Some questions on the idea of","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-05-10T22:39:16Z","receivedAt":"2012-05-10T22:39:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, May 10, 2012 at 04:43:26PM -0500, Neal Kreitzinger wrote:\n\n> >Yes. The on-the-wire format is a packfile. We create a new packfile on\n> >the fly, so we may find new deltas (e.g., between objects that were\n> >stored on disk in two different packs), but we will mostly be reusing\n> >deltas from the existing packs.\n> >\n> >So any time you improve the on-disk representation, you are also\n> >improving the network bandwidth utilization.\n> >\n> The git-clone manpage says you can use the rsync protocol for the\n> url.  If you use rsync:// as your url for your remote does that get\n> you the rsync delta-transfer algorithm efficiency for the network\n> bandwidth utilization part (as opposed to the on-disk representation\n> part)?  (I'm new to rsync.)\n\nWell, yes. If you use the rsync transport, it literally runs rsync,\nwhich will use the regular rsync algorithm. But it won't be better than\nthe git protocol (and in fact will be much worse) for a few reasons:\n\n  1. The object db files are all named after the sha1 of their content\n     (the object sha1 for loose objects, and the sha1 of the whole pack\n     for packfiles). Rsync will not run its comparison algorithm between\n     files with different names. It will not re-transfer existing loose\n     objects, but it will delete obsolete packfiles and retransfer new\n     ones in their entirety. So it's like re-cloning over again for any\n     fetch after an upstream repack.\n\n  2. Even if you could use the rsync delta algorithm, it will never be\n     as efficient as git. Git understands the structure of the packfile\n     and can tell the other side \"Hey, I have these objects\". Whereas\n     rsync must guess from the bytes in the packfiles. Which is much\n     less efficient to compute, and can be wrong if the representation\n     has changed (e.g., something used to be a whole object, but is now\n     stored as a delta).\n\n  3. Even if you could get the exact right set of objects to transfer,\n     and then use the rsync delta algorithm on them, git would still do\n     better. Git's job is much easier: one side has both sets of\n     objects (those to be sent and those not), and is generating and\n     sending efficient deltas for the other side to apply to their\n     objects. Rsync assumes a harder job: you have one set, and\n     the remote side has the other set, and you must agree on a delta by\n     comparing checksums. So it will fundamentally never do as well.\n\n-Peff\n"}]}