{"thread":{"id":"51690","subject":"Only track built files for final output?","startedAt":"2019-08-20T12:21:23Z","lastAt":"2019-08-20T19:42:58Z","messageCount":6,"participants":["Leam Hall","Pratyush Yadav","Randall S. Becker","Phil Hord"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"380797","messageId":"477295c5-f817-e32b-04fd-a41ddfbbac0a@gmail.com","threadId":"51690","inReplyTo":null,"subject":"Only track built files for final output?","fromName":"Leam Hall","fromEmail":"leamhall@gmail.com","sentAt":"2019-08-20T12:21:41Z","receivedAt":"2019-08-20T12:21:23Z","isPatch":false,"sender":{"key":"leamhall@gmail.com","avatar":"https://gravatar.com/avatar/0c8d059b271ba6bf39bd832f0b95c954385219cf3a2906d26a5f35a94be63753?d=mp&s=160"},"body":"Hey all, a newbie could use some help.\n\nWe have some code that generates data files, and as a part of our build \nprocess those files are rebuilt to ensure things work. This causes an \nissue with branches and merging, as the data files change slightly and \ndealing with half a dozen merge conflicts, for files that are in an \ninterim state, is frustrating. The catch is that when the code goes to \nthe production state, those files must be in place and current.\n\nWe use a release branch, and then fork off that for each issue. Testing, \nand file creation, is a part of the pre-merge process. This is what \ncauses the merge conflicts.\n\nRight now my thought is to put the \"final\" versions of the files in some \nother directory, and put the interim file storage directory in \n.gitignore. Is there a better way to do this?\n\nThanks!\n\nLeam\n"},{"id":"380819","messageId":"20190820174640.n3elekpi6l4vwamp@localhost.localdomain","threadId":"51690","inReplyTo":"477295c5-f817-e32b-04fd-a41ddfbbac0a@gmail.com","subject":"Re: Only track built files for final output?","fromName":"Pratyush Yadav","fromEmail":"me@yadavpratyush.com","sentAt":"2019-08-20T17:46:40Z","receivedAt":"2019-08-20T17:46:57Z","isPatch":false,"sender":{"key":"me@yadavpratyush.com","avatar":"https://avatars.githubusercontent.com/u/8817931?v=4"},"body":"On 20/08/19 08:21AM, Leam Hall wrote:\n> Hey all, a newbie could use some help.\n> \n> We have some code that generates data files, and as a part of our build\n> process those files are rebuilt to ensure things work. This causes an issue\n> with branches and merging, as the data files change slightly and dealing\n> with half a dozen merge conflicts, for files that are in an interim state,\n> is frustrating. The catch is that when the code goes to the production\n> state, those files must be in place and current.\n> \n> We use a release branch, and then fork off that for each issue. Testing, and\n> file creation, is a part of the pre-merge process. This is what causes the\n> merge conflicts.\n> \n> Right now my thought is to put the \"final\" versions of the files in some\n> other directory, and put the interim file storage directory in .gitignore.\n> Is there a better way to do this?\n> \n\nMy philosophy with Git is to only track files that I need to generate \nthe final product. I never track the generated files, because I can \nalways get to them via the tracked \"source\" files.\n\nSo for example, I was working on a simple parser in Flex and Bison. Flex \nand Bison take source files in their syntax, and generate a C file each \nthat is then compiled and linked to get to the final binary. So instead \nof tracking the generated C files, I only tracked the source Flex and \nBison files. My build system can always get me the generated files.\n\nSo in your case, what's wrong with just tracking the source files needed \nto generate the other files, and then when you want a release binary, \njust clone the repo, run your build system, and get the generated files?  \nWhat benefit do you get by tracking the generated files?\n\n-- \nRegards,\nPratyush Yadav\n"},{"id":"380822","messageId":"f899594c-4f57-b941-f4f1-fd3b8f81136a@gmail.com","threadId":"51690","inReplyTo":"20190820174640.n3elekpi6l4vwamp@localhost.localdomain","subject":"Re: Only track built files for final output?","fromName":"Leam Hall","fromEmail":"leamhall@gmail.com","sentAt":"2019-08-20T18:01:14Z","receivedAt":"2019-08-20T18:00:57Z","isPatch":false,"sender":{"key":"leamhall@gmail.com","avatar":"https://gravatar.com/avatar/0c8d059b271ba6bf39bd832f0b95c954385219cf3a2906d26a5f35a94be63753?d=mp&s=160"},"body":"On 8/20/19 1:46 PM, Pratyush Yadav wrote:\n> On 20/08/19 08:21AM, Leam Hall wrote:\n>> Hey all, a newbie could use some help.\n>>\n>> We have some code that generates data files, and as a part of our build\n>> process those files are rebuilt to ensure things work. This causes an issue\n>> with branches and merging, as the data files change slightly and dealing\n>> with half a dozen merge conflicts, for files that are in an interim state,\n>> is frustrating. The catch is that when the code goes to the production\n>> state, those files must be in place and current.\n>>\n>> We use a release branch, and then fork off that for each issue. Testing, and\n>> file creation, is a part of the pre-merge process. This is what causes the\n>> merge conflicts.\n>>\n>> Right now my thought is to put the \"final\" versions of the files in some\n>> other directory, and put the interim file storage directory in .gitignore.\n>> Is there a better way to do this?\n>>\n> \n> My philosophy with Git is to only track files that I need to generate\n> the final product. I never track the generated files, because I can\n> always get to them via the tracked \"source\" files.\n> \n> So for example, I was working on a simple parser in Flex and Bison. Flex\n> and Bison take source files in their syntax, and generate a C file each\n> that is then compiled and linked to get to the final binary. So instead\n> of tracking the generated C files, I only tracked the source Flex and\n> Bison files. My build system can always get me the generated files.\n> \n> So in your case, what's wrong with just tracking the source files needed\n> to generate the other files, and then when you want a release binary,\n> just clone the repo, run your build system, and get the generated files?\n> What benefit do you get by tracking the generated files?\n\nFor internal use I agree with you. However, there's an issue.\n\nThe generated files are used by another program's build system, and I \ncan't guarantee the other build system's build system is built like \nours. It seems easier to provide them the generated files and decouple \ntheir build system layout from ours.\n\n\n"},{"id":"380824","messageId":"01b601d55782$c5e19d00$51a4d700$@nexbridge.com","threadId":"51690","inReplyTo":"20190820174640.n3elekpi6l4vwamp@localhost.localdomain","subject":"RE: Only track built files for final output?","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-08-20T18:11:59Z","receivedAt":"2019-08-20T18:12:10Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On August 20, 2019 1:47 PM, Pratyush Yadav\n> On 20/08/19 08:21AM, Leam Hall wrote:\n> > Hey all, a newbie could use some help.\n> >\n> > We have some code that generates data files, and as a part of our\n> > build process those files are rebuilt to ensure things work. This\n> > causes an issue with branches and merging, as the data files change\n> > slightly and dealing with half a dozen merge conflicts, for files that\n> > are in an interim state, is frustrating. The catch is that when the\n> > code goes to the production state, those files must be in place and\ncurrent.\n> >\n> > We use a release branch, and then fork off that for each issue.\n> > Testing, and file creation, is a part of the pre-merge process. This\n> > is what causes the merge conflicts.\n> >\n> > Right now my thought is to put the \"final\" versions of the files in\n> > some other directory, and put the interim file storage directory in\n> .gitignore.\n> > Is there a better way to do this?\n> >\n> \n> My philosophy with Git is to only track files that I need to generate the\nfinal\n> product. I never track the generated files, because I can always get to\nthem\n> via the tracked \"source\" files.\n> \n> So for example, I was working on a simple parser in Flex and Bison. Flex\nand\n> Bison take source files in their syntax, and generate a C file each that\nis then\n> compiled and linked to get to the final binary. So instead of tracking the\n> generated C files, I only tracked the source Flex and Bison files. My\nbuild\n> system can always get me the generated files.\n> \n> So in your case, what's wrong with just tracking the source files needed\nto\n> generate the other files, and then when you want a release binary, just\nclone\n> the repo, run your build system, and get the generated files?\n> What benefit do you get by tracking the generated files?\n\nThe benefit of putting final release packages into git is based on the\nfollowing set of requirements in highly regulated industries:\n\n1. The release artifacts can never change from the point in time at which\nthey are certified as working (a.k.a. passed tests) to the point when they\nare replaced with other artifacts (a subsequent release). Recompiling is not\nsufficient as the compilers themselves may change or be compromised. This is\nan audit requirement.\n2. The source commit(s) used to create the release artifacts must be\nimmutable so that the origins of the release artifacts are always known.\nThis is also an audit requirement in regulated industries.\n3. Disconnecting the source from the object (as is common in artifact\nrepositories) breaks #2 and allows malicious code injection in\nafter-the-test code reproduction. Variant of #2 but from the security\nperspective.\n4. Metadata on the origin of the release artifacts (the clone URL, the\nparent commit, the branch, signed commits), are required for forensic\nanalysis of code in a compliance environment.\n\nThere are other related variants of the above, but those are the essential\nones that are generally accepted in financial, insurance, medical device,\nand industrial applications. Increasingly, food production and distribution\nsectors are realizing that they are also subject to the above. I sadly\ncannot cite specific internal regulations or policies for NDA reasons, but\nhope that others are able to do that.\n\nRegards,\nRandall\n\n-- Brief whoami:\n NonStop developer since approximately 211288444200000000\n UNIX developer since approximately 421664400\n-- In my real life, I talk too much.\n\n\n\n"},{"id":"380838","messageId":"20190820185654.fhelqfub2on67mre@localhost.localdomain","threadId":"51690","inReplyTo":"f899594c-4f57-b941-f4f1-fd3b8f81136a@gmail.com","subject":"Re: Only track built files for final output?","fromName":"Pratyush Yadav","fromEmail":"me@yadavpratyush.com","sentAt":"2019-08-20T18:56:54Z","receivedAt":"2019-08-20T18:57:00Z","isPatch":false,"sender":{"key":"me@yadavpratyush.com","avatar":"https://avatars.githubusercontent.com/u/8817931?v=4"},"body":"On 20/08/19 02:01PM, Leam Hall wrote:\n> On 8/20/19 1:46 PM, Pratyush Yadav wrote:\n> > On 20/08/19 08:21AM, Leam Hall wrote:\n> > > Hey all, a newbie could use some help.\n> > > \n> > > We have some code that generates data files, and as a part of our build\n> > > process those files are rebuilt to ensure things work. This causes an issue\n> > > with branches and merging, as the data files change slightly and dealing\n> > > with half a dozen merge conflicts, for files that are in an interim state,\n> > > is frustrating. The catch is that when the code goes to the production\n> > > state, those files must be in place and current.\n> > > \n> > > We use a release branch, and then fork off that for each issue. Testing, and\n> > > file creation, is a part of the pre-merge process. This is what causes the\n> > > merge conflicts.\n> > > \n> > > Right now my thought is to put the \"final\" versions of the files in some\n> > > other directory, and put the interim file storage directory in .gitignore.\n> > > Is there a better way to do this?\n> > > \n> > \n> > My philosophy with Git is to only track files that I need to generate\n> > the final product. I never track the generated files, because I can\n> > always get to them via the tracked \"source\" files.\n> > \n> > So for example, I was working on a simple parser in Flex and Bison. Flex\n> > and Bison take source files in their syntax, and generate a C file each\n> > that is then compiled and linked to get to the final binary. So instead\n> > of tracking the generated C files, I only tracked the source Flex and\n> > Bison files. My build system can always get me the generated files.\n> > \n> > So in your case, what's wrong with just tracking the source files needed\n> > to generate the other files, and then when you want a release binary,\n> > just clone the repo, run your build system, and get the generated files?\n> > What benefit do you get by tracking the generated files?\n> \n> For internal use I agree with you. However, there's an issue.\n> \n> The generated files are used by another program's build system, and I can't\n> guarantee the other build system's build system is built like ours. It seems\n> easier to provide them the generated files and decouple their build system\n> layout from ours.\n\nMaybe I don't completely understand your use case, but you can still \npass off the generated files to the external build system without having \nto track them. Unless the external build system exclusively relies on \ngit clones/fetches, how about packaging your release with your files \ngenerated from your build system in a tarball (or anything else that \nworks for you) and pushing them to the external build system?\n\nAssuming you just _have_ to track those files, will always resolving the \nmerge conflicts as 'theirs' work?\n\nMy guess about your process works is you branch off, make a new feature \nor fix, and then merge those changes to your master. In that case, the \nchanges that the feature branch made to your generated files should \nalways be the ones that get committed, correct? master's version of the \ngenerated files should be stale. So your merge conflicts always need to \nbe resolved as 'theirs', at least on the generated files. I don't know \nif git-merge supports file-specific merge strategies though, please \ncheck once. Otherwise, maybe you can write a script that resolves \nconflicts as 'theirs' for the generated files, and lets you figure it \nout manually for the rest. \n\nI'm just thinking out loud. I don't know how well this will scale. Maybe \nthe more experienced folks here will have better ideas.\n\n-- \nRegards,\nPratyush Yadav\n"},{"id":"380848","messageId":"CABURp0oCvrVxdvkEyBy_Oe-xy7mEM9Yek-Qe9vnxO9MPfR3Vqg@mail.gmail.com","threadId":"51690","inReplyTo":"f899594c-4f57-b941-f4f1-fd3b8f81136a@gmail.com","subject":"Re: Only track built files for final output?","fromName":"Phil Hord","fromEmail":"phil.hord@gmail.com","sentAt":"2019-08-20T19:42:42Z","receivedAt":"2019-08-20T19:42:58Z","isPatch":false,"sender":{"key":"phil.hord@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123908?v=4"},"body":"On Tue, Aug 20, 2019 at 11:01 AM Leam Hall <leamhall@gmail.com> wrote:\n> On 8/20/19 1:46 PM, Pratyush Yadav wrote:\n\n> > So in your case, what's wrong with just tracking the source files needed\n> > to generate the other files, and then when you want a release binary,\n> > just clone the repo, run your build system, and get the generated files?\n> > What benefit do you get by tracking the generated files?\n>\n> For internal use I agree with you. However, there's an issue.\n>\n> The generated files are used by another program's build system, and I\n> can't guarantee the other build system's build system is built like\n> ours. It seems easier to provide them the generated files and decouple\n> their build system layout from ours.\n\nIt becomes a burden to keep build products in the repo over time, for\nthe reasons you already mentioned (they don't merge and you shouldn't\ntry), but also because those build products never go away, leading to\nrepo-bloat.  Once you realize the cost is too great, it's often too\nlate to do something about it cheaply.  My advice is to keep your\nsource repository clean from the beginning, so it contains only source\ncode.\n\nThis means you still have a problem because you want to distribute\ncertified build artifacts.  I recommend you use some other tool to\nhandle that, like Artifactory.\n\nI recognize it seems easy to use Git for this because Git already acts\nlike a reliable, portable, trackable file distribution system. But\nthat's secondary to Git's purpose; there are better tools for that. If\nyou must lean on Git for this, I like to isolate the binaries into a\nsubmodule so developers who don't want or need them aren't bothered by\nthem, and they can stay out of the way of merges.  But submodules\npresent new workflow challenges and will require some study and\neducation.  If you want to keep them out of the way of developers, you\ncan keep your source code repo and your artifact repo completely\nseparate and make some \"superproject\" which contains both of those\nrepos as submodules.  The nice feature about this setup is you can\npositively associate the set of build products with the set of source\ncode that produced them.\n"}]}