{"thread":{"id":"43124","subject":"Average size of git bookkeeping data (related to Using git as a general backup mechanism)","startedAt":"2006-12-13T19:31:49Z","lastAt":"2006-12-20T01:20:08Z","messageCount":3,"participants":["David Tweed","David Lang","Johannes Schindelin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"293881","messageId":"20061213193149.43284.qmail@web86909.mail.ukl.yahoo.com","threadId":"43124","inReplyTo":null,"subject":"Average size of git bookkeeping data (related to Using git as a general backup mechanism)","fromName":"David Tweed","fromEmail":"tweed314@yahoo.co.uk","sentAt":"2006-12-13T19:31:49Z","receivedAt":"2006-12-13T19:31:49Z","isPatch":false,"sender":{"key":"tweed314@yahoo.co.uk","avatar":null},"body":"Hi,\nThis is in connection with the \"Using git as a general backup mechanism\" thread,\nbut on a different slant more along automatic versioning a subset of the files on the disc (to avoid\nsucking in things like core files and particularly huge multimedia files). In this case I don't\nwant to discard old \"backups\" but just grow the chronological database. (Think a more selective version of\nPlan 9's venti filesystem.)\n\nHow big is the \"metadata\" or \"bookeeping data\" in git related to a commit? (Eg, \"around x bytes per changed file\"\nor \"around x bytes per file being tracked (whether changed in the commit or not)\" )\n\n[I'm trying to get a feel for, if I switched to git, how much overhead would come from having a cron job automatically doing\na snapshot every hour (if anything has changed), plus manual snapshots at points where I want to feel \"safeguarded\".\nI'm currently using my own simple, hacked together system for combined versioning/backups that does\nthis. Using naive tools that don't account for wastes space due to disk block size effects the data being\ntracked is currently just under 9 months of acitvity on  2016 filenames with\n17457599 bytes of data (ie, compressed version of their contents at various times) and 7838546 bytes\nis \"metadata\", ie, 30 percent of the stored data is metadata. This is in a format using 6 bytes to associate a single blob of\ncontents to a filename (whether changed since last snapshot or not).]\n\nMany thanks for any insight,\n\ncheers, dave tweed\n\n\n\n"},{"id":"296677","messageId":"Pine.LNX.4.63.0612140058000.3635@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43124","inReplyTo":"20061213193149.43284.qmail@web86909.mail.ukl.yahoo.com","subject":"Re: Average size of git bookkeeping data (related to Using git as a general backup mechanism)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-12-14T00:02:55Z","receivedAt":"2006-12-14T00:02:55Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 13 Dec 2006, David Tweed wrote:\n\n> How big is the \"metadata\" or \"bookeeping data\" in git related to a \n> commit?\n\nIn my experience, a single commit (with the trees and blobs being \nreachable from it) takes about the same space in a pack as the \ncorresponding tar ball.\n\nWhen you add revisions, the pack grows approximately with the compressed \npatch size.\n\nSo, if you can estimate the size of the initial tar ball and the size of \nthe subsequent patches, you have a ballpark figure of the metadata size of \nthe git repository. (This is for the bare case; for the regular case you \nhave to add the size of a checked out tree, of course.)\n\nHth,\nDscho\n"},{"id":"294527","messageId":"Pine.LNX.4.63.0612191712440.18007@qynat.qvtvafvgr.pbz","threadId":"43124","inReplyTo":"20061213193149.43284.qmail@web86909.mail.ukl.yahoo.com","subject":"Re: Average size of git bookkeeping data (related to Using git as a general backup mechanism)","fromName":"David Lang","fromEmail":"dlang@digitalinsight.com","sentAt":"2006-12-20T01:20:08Z","receivedAt":"2006-12-20T01:20:08Z","isPatch":false,"sender":{"key":"dlang@digitalinsight.com","avatar":null},"body":"On Wed, 13 Dec 2006, David Tweed wrote:\n\n> How big is the \"metadata\" or \"bookeeping data\" in git related to a commit? (Eg, \"around x bytes per changed file\"\n> or \"around x bytes per file being tracked (whether changed in the commit or not)\" )\n>\n> [I'm trying to get a feel for, if I switched to git, how much overhead would come from having a cron job automatically doing\n> a snapshot every hour (if anything has changed), plus manual snapshots at points where I want to feel \"safeguarded\".\n> I'm currently using my own simple, hacked together system for combined versioning/backups that does\n> this. Using naive tools that don't account for wastes space due to disk block size effects the data being\n> tracked is currently just under 9 months of acitvity on  2016 filenames with\n> 17457599 bytes of data (ie, compressed version of their contents at various times) and 7838546 bytes\n> is \"metadata\", ie, 30 percent of the stored data is metadata. This is in a format using 6 bytes to associate a single blob of\n> contents to a filename (whether changed since last snapshot or not).]\n\nif nothing has changed it will take the space of the commit tag, as the tree \nwill remain the same (and you should be able to script detection of this case \nand make it zero overhead)\n\nif something has changed you will have the new tree and the changed object\n\nin a tree each object is ~28 bytes (IIRC from what Linus mentioned in the last \nweek or two)\n\na loose object is compressed, and if you repack it will delta against prior \nversions for even more space savings\n\nlook at the size of the mozilla tree and the kernel tree and you will see that \nwhen packed git is about as efficiant as any other option you have (and more \nefficiant than most)\n\n"}]}