{"thread":{"id":"23500","subject":"GIT Performance question","startedAt":"2010-04-17T09:55:49Z","lastAt":"2010-04-17T11:55:02Z","messageCount":6,"participants":["santos2010","Jeff King","Dmitry Potapov","Geert Bosch"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"139730","messageId":"1271498149921-4917066.post@n2.nabble.com","threadId":"23500","inReplyTo":null,"subject":"GIT Performance question","fromName":"santos2010","fromEmail":"santos.claudia2009@googlemail.com","sentAt":"2010-04-17T09:55:49Z","receivedAt":"2010-04-17T09:55:49Z","isPatch":false,"sender":{"key":"santos.claudia2009@googlemail.com","avatar":null},"body":"\nHello,\n\nOur company is evaluating SCM solutions, one of our most important\nrequirements is performance as we develop over 3 differents sites across the\nworld.\nI read that GIT doesn't use deltas, it uses snapshots. My question is: how\ncould GIT have high performance (most of the users say that) if for\nsynchronization (pull/push command) with e.g. a shared repository GIT\ntransfers all modified files (and references) instead of the respective\ndeltas? \n\nThanks in advance,\n\nSantos\n-- \nView this message in context: http://n2.nabble.com/GIT-Performance-question-tp4917066p4917066.html\nSent from the git mailing list archive at Nabble.com.\n"},{"id":"139734","messageId":"2FFDF076-F06F-427E-92FE-51C2A6AA02BF@adacore.com","threadId":"23500","inReplyTo":"1271498149921-4917066.post@n2.nabble.com","subject":"Re: GIT Performance question","fromName":"Geert Bosch","fromEmail":"bosch@adacore.com","sentAt":"2010-04-17T10:37:11Z","receivedAt":"2010-04-17T10:37:11Z","isPatch":false,"sender":{"key":"bosch@adacore.com","avatar":null},"body":"\nOn Apr 17, 2010, at 05:55, santos2010 wrote:\n\n> \n> Hello,\n> \n> Our company is evaluating SCM solutions, one of our most important\n> requirements is performance as we develop over 3 differents sites across the\n> world.\n> I read that GIT doesn't use deltas, it uses snapshots.\nGit does use deltas between snapshots.\n\n  -Geert"},{"id":"139732","messageId":"20100417103748.GC23110@coredump.intra.peff.net","threadId":"23500","inReplyTo":"1271498149921-4917066.post@n2.nabble.com","subject":"Re: GIT Performance question","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2010-04-17T10:37:49Z","receivedAt":"2010-04-17T10:37:49Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Apr 17, 2010 at 01:55:49AM -0800, santos2010 wrote:\n\n> Our company is evaluating SCM solutions, one of our most important\n> requirements is performance as we develop over 3 differents sites across the\n> world.\n> I read that GIT doesn't use deltas, it uses snapshots. My question is: how\n> could GIT have high performance (most of the users say that) if for\n> synchronization (pull/push command) with e.g. a shared repository GIT\n> transfers all modified files (and references) instead of the respective\n> deltas? \n\nShort answer: Git does store and transfer deltas. It generally beats any\nother system in terms of repo size.\n\nLonger answer:\n\nGit separates the concept of the history graph and the actual storage\nmechanism. So conceptually the history is a directed graph of snapshots,\neach representing the whole tree. But there are two things that save\nspace:\n\n  1. Git addresses content by its sha1. So each snapshot may refer to a\n     file by the sha1 of its content, meaning we only have to store that\n     content once.\n\n  2. Git packs \"objects\" (where each file's content is in a single\n     object) into \"packfiles\", in which it aggressively deltas objects\n     against each other, including objects which do not come from the\n     same path in your tree.\n\nGit will store \"loose\" objects when performing most operations, but will\noccasionally pack when the number of objects get too high. You can also\ninitiate a full pack by running \"git gc\".\n\nFor transferring between repositories, git will figure out which parts\nof the history each side has, and will only send the objects that the\nother side needs. In addition, it will send them as a packfile using\ndelta compression, including deltas against objects that are not being\nsent but that it knows the other side has.\n\n-Peff\n"},{"id":"139733","messageId":"20100417104037.GA20631@dpotapov.dyndns.org","threadId":"23500","inReplyTo":"1271498149921-4917066.post@n2.nabble.com","subject":"Re: GIT Performance question","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2010-04-17T10:40:37Z","receivedAt":"2010-04-17T10:40:37Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Sat, Apr 17, 2010 at 01:55:49AM -0800, santos2010 wrote:\n> \n> I read that GIT doesn't use deltas, it uses snapshots. My question is: how\n> could GIT have high performance (most of the users say that) if for\n> synchronization (pull/push command) with e.g. a shared repository GIT\n> transfers all modified files (and references) instead of the respective\n> deltas? \n\nWell, Git _does_ use deltas for storage and synchronization, but this\ndeltas are unrelated to history of changes stored in the repository. So,\nconceptually, Git just stores snapshots, but files in those snapshots\nare deltified against some old files based on some heuristic of finding\nsimilar files, which allows Git to create deltas not only to previous\nversion of the same file (which most VCSes do), but potentially to any\nfile stored in the repository if it similar enough. So, typically, Git\nhas the most compact storage comparing to other VCSes, in particular, in\ncase of complex history with a lot of branches and merges.\n\n\nDmitry\n"},{"id":"139737","messageId":"1271503273779-4917251.post@n2.nabble.com","threadId":"23500","inReplyTo":"20100417104037.GA20631@dpotapov.dyndns.org","subject":"Re: GIT Performance question","fromName":"santos2010","fromEmail":"santos.claudia2009@googlemail.com","sentAt":"2010-04-17T11:21:13Z","receivedAt":"2010-04-17T11:21:13Z","isPatch":false,"sender":{"key":"santos.claudia2009@googlemail.com","avatar":null},"body":"\nThanks a lot for the quick answer. Are there some references (on-line books\nor web pages) where i could find details about this approach? I need this to\njustify my evaluation :)\n-- \nView this message in context: http://n2.nabble.com/GIT-Performance-question-tp4917066p4917251.html\nSent from the git mailing list archive at Nabble.com.\n"},{"id":"139741","messageId":"20100417115502.GC28623@coredump.intra.peff.net","threadId":"23500","inReplyTo":"1271503273779-4917251.post@n2.nabble.com","subject":"Re: GIT Performance question","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2010-04-17T11:55:02Z","receivedAt":"2010-04-17T11:55:02Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Apr 17, 2010 at 03:21:13AM -0800, santos2010 wrote:\n\n> Thanks a lot for the quick answer. Are there some references (on-line books\n> or web pages) where i could find details about this approach? I need this to\n> justify my evaluation :)\n\nTry the \"Git Internals\" chapter of Scott Chacon's Pro Git book, which is\navailable online here:\n\n  http://progit.org/book/ch9-0.html\n\nHe has some pretty pictures, too.\n\n-Peff\n"}]}