{"thread":{"id":"46426","subject":"Re: Binary files","startedAt":"2017-07-20T07:42:34Z","lastAt":"2017-07-21T17:47:03Z","messageCount":8,"participants":["Volodymyr Sendetskyi","Bryan Turner","Konstantin Khomoutov","Lars Schneider","Stefan Beller","Igor Djordjevic","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"324790","messageId":"CAFc9kS_xYVyPsW7qogDxLugxBb1p2vEFAoP=W9Rdnfqs6XtWKQ@mail.gmail.com","threadId":"46426","inReplyTo":"CAFc9kS8L-JJoJqKi7bB90qwKVW8gB=EFk9D8c=4YShqnamwa2w@mail.gmail.com","subject":"Re: Binary files","fromName":"Volodymyr Sendetskyi","fromEmail":"volodymyrse@devcom.com","sentAt":"2017-07-20T07:41:48Z","receivedAt":"2017-07-20T07:42:34Z","isPatch":false,"sender":{"key":"volodymyrse@devcom.com","avatar":null},"body":"It is known, that git handles badly storing binary files in its\nrepositories at all.\nThis is especially about large files: even without any changes to\nthese files, their copies are snapshotted on each commit. So even\nrepositories with a small amount of code can grove very fast in size\nif they contain some great binary files. Alongside this, the SVN is\nmuch better about that, because it make changes to the server version\nof file only if some changes were done.\n\nSo the question is: why not implementing some feature, that would\nsomehow handle this problem?\n\nOf course, I don't know the internal git structure and the way of\nworking + some nuances (likely about the snapshots at all and the way\nthey are done), so handling this may be a great problem. But the\neasiest feature for me as an end user will be something like\n'.gitbinary', where I can list binary files, that would behave like on\nSVN, or even more optimal, if you can implement it. Maybe there will\nbe a need for separate kinds of repositories, or even servers. But\nthat would be a great change and a logical way of next git's\nevolution.\n"},{"id":"324791","messageId":"CAGyf7-FFQir7tnyW3RZp4vk7Y1qD3m34iTb4xUeKtm8rOrdL0Q@mail.gmail.com","threadId":"46426","inReplyTo":"CAFc9kS_xYVyPsW7qogDxLugxBb1p2vEFAoP=W9Rdnfqs6XtWKQ@mail.gmail.com","subject":"Re: Binary files","fromName":"Bryan Turner","fromEmail":"bturner@atlassian.com","sentAt":"2017-07-20T07:58:47Z","receivedAt":"2017-07-20T07:58:54Z","isPatch":false,"sender":{"key":"bturner@atlassian.com","avatar":"https://gravatar.com/avatar/16bcf3167981c1ef7c804e502642366d888a35b0d0b0a4ca01fdc442aa1acb1e?d=mp&s=160"},"body":"On Thu, Jul 20, 2017 at 12:41 AM, Volodymyr Sendetskyi\n<volodymyrse@devcom.com> wrote:\n> It is known, that git handles badly storing binary files in its\n> repositories at all.\n> This is especially about large files: even without any changes to\n> these files, their copies are snapshotted on each commit. So even\n> repositories with a small amount of code can grove very fast in size\n> if they contain some great binary files. Alongside this, the SVN is\n> much better about that, because it make changes to the server version\n> of file only if some changes were done.\n>\n> So the question is: why not implementing some feature, that would\n> somehow handle this problem?\n\nLike Git LFS or git annex? Features have been implemented to better\nhandle large files; they're just not necessarily part of core Git.\nHave you checked whether one of those solutions might work for your\nuse case?\n\nBest regards,\nBryan Turner\n"},{"id":"324792","messageId":"20170720080116.7os2bxytm25qdt25@tigra","threadId":"46426","inReplyTo":"CAFc9kS_xYVyPsW7qogDxLugxBb1p2vEFAoP=W9Rdnfqs6XtWKQ@mail.gmail.com","subject":"Re: Binary files","fromName":"Konstantin Khomoutov","fromEmail":"kostix+git@007spb.ru","sentAt":"2017-07-20T08:01:16Z","receivedAt":"2017-07-20T08:01:22Z","isPatch":false,"sender":{"key":"kostix+git@007spb.ru","avatar":null},"body":"On Thu, Jul 20, 2017 at 10:41:48AM +0300, Volodymyr Sendetskyi wrote:\n\n> It is known, that git handles badly storing binary files in its\n> repositories at all.\n[...]\n> So the question is: why not implementing some feature, that would\n> somehow handle this problem?\n[...]\n\nHave you examined git-lfs and git-annex?\n(Actually, there are/were more solutions [1] but these two appear to be\nthe most used novadays.)\n\nSuch solutions allow one to use Git for what it does best and defer\nhandling of big files (or files for which lock-modify-unlock works better\nthan the usual modify-merge) to a specialized solution.\n\n1. http://blog.deveo.com/storing-large-binary-files-in-git-repositories/\n\n"},{"id":"324795","messageId":"6C4E3445-7268-47CA-A6D4-514A8EA348AA@gmail.com","threadId":"46426","inReplyTo":"CAFc9kS_xYVyPsW7qogDxLugxBb1p2vEFAoP=W9Rdnfqs6XtWKQ@mail.gmail.com","subject":"Re: Binary files","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-07-20T08:32:03Z","receivedAt":"2017-07-20T08:32:27Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 20 Jul 2017, at 09:41, Volodymyr Sendetskyi <volodymyrse@devcom.com> wrote:\n> \n> It is known, that git handles badly storing binary files in its\n> repositories at all.\n> This is especially about large files: even without any changes to\n> these files, their copies are snapshotted on each commit. So even\n> repositories with a small amount of code can grove very fast in size\n> if they contain some great binary files. Alongside this, the SVN is\n> much better about that, because it make changes to the server version\n> of file only if some changes were done.\n> \n> So the question is: why not implementing some feature, that would\n> somehow handle this problem?\n> \n> Of course, I don't know the internal git structure and the way of\n> working + some nuances (likely about the snapshots at all and the way\n> they are done), so handling this may be a great problem. But the\n> easiest feature for me as an end user will be something like\n> '.gitbinary', where I can list binary files, that would behave like on\n> SVN, or even more optimal, if you can implement it. Maybe there will\n> be a need for separate kinds of repositories, or even servers. But\n> that would be a great change and a logical way of next git's\n> evolution.\n\nGitLFS [1] might be the workaround you want. There are efforts to bring \nlarge file support natively to Git [2].\n\nI tried to explain GitLFS in more detail here: \nhttps://www.youtube.com/watch?v=YQzNfb4IwEY\n\n- Lars\n\n\n[1] https://git-lfs.github.com/\n[2] https://public-inbox.org/git/20170620075523.26961-1-chriscool@tuxfamily.org/\n\n"},{"id":"324814","messageId":"CAGZ79kZ+5G0iOZ5ZpEYpB_Hsc9mftNNWNvixx4k37SzamFvrJw@mail.gmail.com","threadId":"46426","inReplyTo":"CAFc9kS_xYVyPsW7qogDxLugxBb1p2vEFAoP=W9Rdnfqs6XtWKQ@mail.gmail.com","subject":"Re: Binary files","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2017-07-20T17:22:13Z","receivedAt":"2017-07-20T17:22:20Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Thu, Jul 20, 2017 at 12:41 AM, Volodymyr Sendetskyi\n<volodymyrse@devcom.com> wrote:\n> It is known, that git handles badly storing binary files in its\n> repositories at all.\n> This is especially about large files: even without any changes to\n> these files, their copies are snapshotted on each commit. So even\n> repositories with a small amount of code can grove very fast in size\n> if they contain some great binary files. Alongside this, the SVN is\n> much better about that, because it make changes to the server version\n> of file only if some changes were done.\n>\n> So the question is: why not implementing some feature, that would\n> somehow handle this problem?\n\nThere are 'external' solutions such as git LFS and git-annex, mentioned\nin replies nearby.\n\nBut note there are also efforts to handle large binary files internally\nhttps://public-inbox.org/git/3420d9ae9ef86b78af1abe721891233e3f5865a2.1500508695.git.jonathantanmy@google.com/\nhttps://public-inbox.org/git/20170713173459.3559-1-git@jeffhostetler.com/\nhttps://public-inbox.org/git/20170620075523.26961-1-chriscool@tuxfamily.org/\n"},{"id":"324822","messageId":"d4b1b92d-6ab1-7e6f-4afd-6194a5ba8e40@gmail.com","threadId":"46426","inReplyTo":"CAFc9kS_xYVyPsW7qogDxLugxBb1p2vEFAoP=W9Rdnfqs6XtWKQ@mail.gmail.com","subject":"Re: Binary files","fromName":"Igor Djordjevic","fromEmail":"igor.d.djordjevic@gmail.com","sentAt":"2017-07-20T18:49:35Z","receivedAt":"2017-07-20T18:49:51Z","isPatch":false,"sender":{"key":"igor.d.djordjevic@gmail.com","avatar":null},"body":"Hi Volodymyr,\n\nOn 20/07/2017 09:41, Volodymyr Sendetskyi wrote:\n> It is known, that git handles badly storing binary files in its\n> repositories at all.\n> This is especially about large files: even without any changes to\n> these files, their copies are snapshotted on each commit. So even\n> repositories with a small amount of code can grove very fast in size\n> if they contain some great binary files. Alongside this, the SVN is\n> much better about that, because it make changes to the server version\n> of file only if some changes were done.\n\nYou already got some proposals on what you could try for making large \nbinary files handling easier, but I just wanted to comment on this \npart of your message, as it doesn`t seem to be correct.\n\nEven though each repository file is included in each commit (being a \nfull repository state snapshot), meaning big binary files as well, \nthat`s just from an end-user`s perspective.\n\nActual implementation side is smarter than that - if file hasn`t \nchanged between commits, it won`t get copied/written to Git object \ndatabase again.\n\nUnder the hood, many different commits can point to the same \n(unchanged) file, thus repository size _does not_ grow very fast with \neach commit if large binary file is without any changes.\n\nUsually, the biggest concern with Git and large files[1], in \ncomparison to SVN, for example, is something else - Git model \nassuming each repository clone holding the complete repository \nhistory with all the different file versions included, so you can`t \nget just some of them, or the last snapshot only, keeping your local \nrepository small in size.\n\nIf the repository you`re cloning from is a big one, your locally \ncloned repository will be as well, even if you may not really be \ninterested in the big files at all... but you got some suggestions \nfor handling that already, as pointed out :)\n\nJust note that it`s not really Git vs SVN here, but more distributed \nvs centralized approach in general, as you can`t both have everything \nand yet skip something at the same time. Different systems may have \ndifferent workarounds for a specific workflow, though.\n\n[1] Besides taking each file version as a full-sized snapshot (at the \nbeginning, at least, until the delta compression packing occurs).\n\nRegards,\nBuga\n"},{"id":"324836","messageId":"xmqqo9sekfjm.fsf@gitster.mtv.corp.google.com","threadId":"46426","inReplyTo":"d4b1b92d-6ab1-7e6f-4afd-6194a5ba8e40@gmail.com","subject":"Re: Binary files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-07-20T20:40:45Z","receivedAt":"2017-07-20T20:40:58Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Igor Djordjevic <igor.d.djordjevic@gmail.com> writes:\n\n> On 20/07/2017 09:41, Volodymyr Sendetskyi wrote:\n>> It is known, that git handles badly storing binary files in its\n>> repositories at all.\n>> This is especially about large files: even without any changes to\n>> these files, their copies are snapshotted on each commit. So even\n>> repositories with a small amount of code can grove very fast in size\n>> if they contain some great binary files. Alongside this, the SVN is\n>> much better about that, because it make changes to the server version\n>> of file only if some changes were done.\n>\n> You already got some proposals on what you could try for making large \n> binary files handling easier, but I just wanted to comment on this \n> part of your message, as it doesn`t seem to be correct.\n\nAll correct.  Thanks.\n"},{"id":"324880","messageId":"0d2dcfc0-81e0-ff16-3d21-7148978f8100@gmail.com","threadId":"46426","inReplyTo":"xmqqo9sekfjm.fsf@gitster.mtv.corp.google.com","subject":"Re: Binary files","fromName":"Igor Djordjevic","fromEmail":"igor.d.djordjevic@gmail.com","sentAt":"2017-07-21T17:46:51Z","receivedAt":"2017-07-21T17:47:03Z","isPatch":false,"sender":{"key":"igor.d.djordjevic@gmail.com","avatar":null},"body":"On 20/07/2017 22:40, Junio C Hamano wrote:\n> Igor Djordjevic <igor.d.djordjevic@gmail.com> writes:\n>> On 20/07/2017 09:41, Volodymyr Sendetskyi wrote:\n>>> It is known, that git handles badly storing binary files in its\n>>> repositories at all.\n>>> This is especially about large files: even without any changes to\n>>> these files, their copies are snapshotted on each commit. So even\n>>> repositories with a small amount of code can grove very fast in size\n>>> if they contain some great binary files. Alongside this, the SVN is\n>>> much better about that, because it make changes to the server version\n>>> of file only if some changes were done.\n>>\n>> You already got some proposals on what you could try for making large \n>> binary files handling easier, but I just wanted to comment on this \n>> part of your message, as it doesn`t seem to be correct.\n> \n> All correct.  Thanks.\n\nNo problem, thanks for confirmation, being relatively new around it`s \nappreciated, at least knowing that I got it correct myself :)\n"}]}