{"thread":{"id":"30061","subject":"Merge-friendly text-based data storage","startedAt":"2012-03-26T14:19:39Z","lastAt":"2012-03-27T15:46:25Z","messageCount":8,"participants":["Richard Hartmann","Junio C Hamano","Holger Hellmuth"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"187722","messageId":"CAD77+gRTv4Aq_4FLAQcD9E0p7VBD7h6hQq3CJ9Wo5DU9Zjt+Hg@mail.gmail.com","threadId":"30061","inReplyTo":null,"subject":"Merge-friendly text-based data storage","fromName":"Richard Hartmann","fromEmail":"richih.mailinglist@gmail.com","sentAt":"2012-03-26T14:19:39Z","receivedAt":"2012-03-26T14:19:39Z","isPatch":false,"sender":{"key":"richih.mailinglist@gmail.com","avatar":"https://avatars.githubusercontent.com/u/754723?v=4"},"body":"Hi all,\n\nI am looking for information on how to design a merge-friendly data\nlayout. Oddly enough, there does not seem to be much online other than\nthe obvious \"use text-based lines, one per data point\".\n\nMy current plan looks like:\n\n  metamonger\\tversion: 0\n  filename\\towner_name\\tgroup_name\\tetc\\tpp\n  ##########\n  file1\\trichih\\trichih\\tfoo\\tbar\n  relative/path/to/file2\\troot\\troot\\tfoo\\tbar\n\nthe two upper lines are designed to fail a merge if the version of the\nfile layout changes. Anything starting with a hash-pound is a comment\nand will be ignored.\n\nAll other lines are data about random files, relative paths being\nallowed, absolute paths and upper paths being forbidden for security\nreasons. Values are tab-separated as the format is expressively meant\nto be edited by hand. Hex, if needed, would be ASCII-armoured.\n\nAs long as there are no lines that start with the same file name, this\nfile format would allow for efficient merging _if_ git has an internal\nconcept of line identifiers.\n\n\nAre there any considerations I missed? Are there any design\nguides/best practices to follow?\n\n\n\nThanks,\nRichard\n"},{"id":"187756","messageId":"7vfwcvp6pi.fsf@alter.siamese.dyndns.org","threadId":"30061","inReplyTo":"CAD77+gRTv4Aq_4FLAQcD9E0p7VBD7h6hQq3CJ9Wo5DU9Zjt+Hg@mail.gmail.com","subject":"Re: Merge-friendly text-based data storage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-03-26T18:17:45Z","receivedAt":"2012-03-26T18:17:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Richard Hartmann <richih.mailinglist@gmail.com> writes:\n\n> As long as there are no lines that start with the same file name, this\n> file format would allow for efficient merging _if_ git has an internal\n> concept of line identifiers.\n\nYou can write a custom low-level merge driver, and use the attribute\nsystem to mark that your file is meant to be handled by that ll-merge\ndriver.  There is no need for Git to have \"an internal concept of line\nidentifiers\".\n\nIt may be of interest to run \"git help attributes\" and read up on\n\"Defining a custom merge driver\" section.\n"},{"id":"187765","messageId":"CAD77+gQVDtoK0vJnSgX-+i9EeJo6QErUGuzd25_cfBmUfPvW4g@mail.gmail.com","threadId":"30061","inReplyTo":"7vfwcvp6pi.fsf@alter.siamese.dyndns.org","subject":"Re: Merge-friendly text-based data storage","fromName":"Richard Hartmann","fromEmail":"richih.mailinglist@gmail.com","sentAt":"2012-03-26T19:06:30Z","receivedAt":"2012-03-26T19:06:30Z","isPatch":false,"sender":{"key":"richih.mailinglist@gmail.com","avatar":"https://avatars.githubusercontent.com/u/754723?v=4"},"body":"On Mon, Mar 26, 2012 at 20:17, Junio C Hamano <gitster@pobox.com> wrote:\n\n> It may be of interest to run \"git help attributes\" and read up on\n> \"Defining a custom merge driver\" section.\n\nSounds good, thanks.\n\nMy file layout looks fine?\n\n\n-- \nRichard\n"},{"id":"187773","messageId":"7vobrjnnt4.fsf@alter.siamese.dyndns.org","threadId":"30061","inReplyTo":"CAD77+gQVDtoK0vJnSgX-+i9EeJo6QErUGuzd25_cfBmUfPvW4g@mail.gmail.com","subject":"Re: Merge-friendly text-based data storage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-03-26T19:51:19Z","receivedAt":"2012-03-26T19:51:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Richard Hartmann <richih.mailinglist@gmail.com> writes:\n\n> On Mon, Mar 26, 2012 at 20:17, Junio C Hamano <gitster@pobox.com> wrote:\n>\n>> It may be of interest to run \"git help attributes\" and read up on\n>> \"Defining a custom merge driver\" section.\n>\n> Sounds good, thanks.\n>\n> My file layout looks fine?\n\nI have no opinion on it. It is for the consumers of your datafile (the\nones that read it and find these databasy items in it, and your custom\nmerge driver) to decide.\n"},{"id":"187841","messageId":"4F718496.4030808@ira.uka.de","threadId":"30061","inReplyTo":"CAD77+gRTv4Aq_4FLAQcD9E0p7VBD7h6hQq3CJ9Wo5DU9Zjt+Hg@mail.gmail.com","subject":"Re: Merge-friendly text-based data storage","fromName":"Holger Hellmuth","fromEmail":"hellmuth@ira.uka.de","sentAt":"2012-03-27T09:12:54Z","receivedAt":"2012-03-27T09:12:54Z","isPatch":false,"sender":{"key":"hellmuth@ira.uka.de","avatar":null},"body":"On 26.03.2012 16:19, Richard Hartmann wrote:\n> Hi all,\n>\n> I am looking for information on how to design a merge-friendly data\n> layout. Oddly enough, there does not seem to be much online other than\n> the obvious \"use text-based lines, one per data point\".\n>\n> My current plan looks like:\n>\n>    metamonger\\tversion: 0\n>    filename\\towner_name\\tgroup_name\\tetc\\tpp\n>    ##########\n>    file1\\trichih\\trichih\\tfoo\\tbar\n>    relative/path/to/file2\\troot\\troot\\tfoo\\tbar\n>\n> the two upper lines are designed to fail a merge if the version of the\n> file layout changes. Anything starting with a hash-pound is a comment\n> and will be ignored.\n\nI may be misunderstanding something, but lets assume you want to merge a \nfile that has \"version: 0\" with one that has \"version: 1\" and their last \ncommon ancestor would have \"version: 0\" naturally. So the merge would \nnot fail even though the file layout changes.\n\nAnd there would be random merge failures with lines added at the same \nline number even if different.\n\nThe normal merging in git isn't suited for this task, it has different \nobjectives. Without a custom merge driver as Junio suggested the only \nway would be to store each data line in its own file. As you store file \npaths that would even fit, but I doubt it is what you had in mind\n"},{"id":"187851","messageId":"CAD77+gR=p+jhN5qNoRgjtQPHqgqrdtcSmqAy_4d0NUaqE6ZkVg@mail.gmail.com","threadId":"30061","inReplyTo":"4F718496.4030808@ira.uka.de","subject":"Re: Merge-friendly text-based data storage","fromName":"Richard Hartmann","fromEmail":"richih.mailinglist@gmail.com","sentAt":"2012-03-27T13:01:37Z","receivedAt":"2012-03-27T13:01:37Z","isPatch":false,"sender":{"key":"richih.mailinglist@gmail.com","avatar":"https://avatars.githubusercontent.com/u/754723?v=4"},"body":"On Tue, Mar 27, 2012 at 11:12, Holger Hellmuth <hellmuth@ira.uka.de> wrote:\n\n> I may be misunderstanding something, but lets assume you want to merge a\n> file that has \"version: 0\" with one that has \"version: 1\" and their last\n> common ancestor would have \"version: 0\" naturally. So the merge would not\n> fail even though the file layout changes.\n\nUgh, I did not consider that. I can't come up with a way, other than a\ncustom merge driver, to prevent this. Am I correct?\n\n\n> And there would be random merge failures with lines added at the same line\n> number even if different.\n\nYes, I know. That was the main reason why I asked for merge-friendly\ndesigns. I briefly considered union merges, but that's not a good idea\nfor obvious reasons.\n\n\n> The normal merging in git isn't suited for this task, it has different\n> objectives. Without a custom merge driver as Junio suggested\n\nI more or less accepted that I will have to write one, eventually.\n\n\n> the only way\n> would be to store each data line in its own file. As you store file paths\n> that would even fit, but I doubt it is what you had in mind\n\nI considered this as well, but that's extremely expensive and wasteful.\n\n\n-- \nRichard\n"},{"id":"187858","messageId":"7vzkb2jchu.fsf@alter.siamese.dyndns.org","threadId":"30061","inReplyTo":"CAD77+gR=p+jhN5qNoRgjtQPHqgqrdtcSmqAy_4d0NUaqE6ZkVg@mail.gmail.com","subject":"Re: Merge-friendly text-based data storage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-03-27T15:21:33Z","receivedAt":"2012-03-27T15:21:33Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Richard Hartmann <richih.mailinglist@gmail.com> writes:\n\n> On Tue, Mar 27, 2012 at 11:12, Holger Hellmuth <hellmuth@ira.uka.de> wrote:\n>\n>> I may be misunderstanding something, but lets assume you want to merge a\n>> file that has \"version: 0\" with one that has \"version: 1\" and their last\n>> common ancestor would have \"version: 0\" naturally. So the merge would not\n>> fail even though the file layout changes.\n>\n> Ugh, I did not consider that. I can't come up with a way, other than a\n> custom merge driver, to prevent this. Am I correct?\n\nYou are the only judge to that statement: \"I can't come up with...\".\n\nI can't either, but I know a custom ll-merge driver would work.  It is\ndesigned for this kind of thing.  It will know both version 0 and version\n1 format, read from each and writes out the merged result in whatever\nformat it wants to use.\n\n>> the only way\n>> would be to store each data line in its own file. As you store file paths\n>> that would even fit, but I doubt it is what you had in mind\n>\n> I considered this as well, but that's extremely expensive and wasteful.\n\nAnd it does not solve anything.  The \"version\" file may cleanly merge to a\nnew version, and there is no way for the merge result of \"version\" file to\naffect the outcome of merges in other files.\n"},{"id":"187860","messageId":"4F71E0D1.2040600@ira.uka.de","threadId":"30061","inReplyTo":"7vzkb2jchu.fsf@alter.siamese.dyndns.org","subject":"Re: Merge-friendly text-based data storage","fromName":"Holger Hellmuth","fromEmail":"hellmuth@ira.uka.de","sentAt":"2012-03-27T15:46:25Z","receivedAt":"2012-03-27T15:46:25Z","isPatch":false,"sender":{"key":"hellmuth@ira.uka.de","avatar":null},"body":"On 27.03.2012 17:21, Junio C Hamano wrote:\n> Richard Hartmann<richih.mailinglist@gmail.com>  writes:\n>\n>>> the only way\n>>> would be to store each data line in its own file. As you store file paths\n>>> that would even fit, but I doubt it is what you had in mind\n>>\n>> I considered this as well, but that's extremely expensive and wasteful.\n>\n> And it does not solve anything.  The \"version\" file may cleanly merge to a\n> new version, and there is no way for the merge result of \"version\" file to\n> affect the outcome of merges in other files.\n\nIt solves the data merging.\n\nAnd since a version change is presumably a very scarce event, this could \nbe solved with a merge hook that simply aborts the merge with a message \nhow to update the older version, then commit and merge.\n"}]}