{"thread":{"id":"29276","subject":"How to deal with historic tar-balls","startedAt":"2011-12-31T19:04:58Z","lastAt":"2012-01-07T19:18:45Z","messageCount":12,"participants":["nn6eumtr","Tomas Carnecky","Philip Oakley","Dirk Süsserott","Neal Kreitzinger","Thomas Rast"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"181815","messageId":"4EFF5CDA.5050809@gmail.com","threadId":"29276","inReplyTo":null,"subject":"How to deal with historic tar-balls","fromName":"nn6eumtr","fromEmail":"nn6eumtr@gmail.com","sentAt":"2011-12-31T19:04:58Z","receivedAt":"2011-12-31T19:04:58Z","isPatch":false,"sender":{"key":"nn6eumtr@gmail.com","avatar":null},"body":"I have a number of older projects that I want to bring into a git \nrepository. They predate a lot of the popular scm systems, so they are \nprimarily a collection of tarballs today.\n\nI'm fairly new to git so I have a couple questions related to this:\n\n- What is the best approach for bringing them in? Do I just create a \nrepository, then unpack the files, commit them, clean out the directory \nunpack the next tarball, and repeat until everything is loaded?\n\n- Do I need to pay special attention to files that are renamed/removed \nfrom version to version?\n\n- If the timestamps change on a file but the actual content does not, \nwill git treat it as a non-change once it realizes the content hasn't \nchanged?\n\n- Last, if after loading the repository I find another version of the \nfiles that predates those I've loaded, or are intermediate between two \ncommits I've already loaded, is there a way to go say that commit B is \nactually the ancestor of commit C? (i.e. a->c becomes a->b->c if you \nwere to visualize the commit timeline or do diffs) Or do I just reload \nthe tarballs in order to achieve this?\n\nAll replies appreciated!\n"},{"id":"181820","messageId":"4EFFA868.50605@dbservice.com","threadId":"29276","inReplyTo":"4EFF5CDA.5050809@gmail.com","subject":"Re: How to deal with historic tar-balls","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2012-01-01T00:27:20Z","receivedAt":"2012-01-01T00:27:20Z","isPatch":false,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"On 12/31/11 8:04 PM, nn6eumtr wrote:\n> I have a number of older projects that I want to bring into a git \n> repository. They predate a lot of the popular scm systems, so they are \n> primarily a collection of tarballs today.\n>\n> I'm fairly new to git so I have a couple questions related to this:\n>\n> - What is the best approach for bringing them in? Do I just create a \n> repository, then unpack the files, commit them, clean out the \n> directory unpack the next tarball, and repeat until everything is loaded?\n>\n> - Do I need to pay special attention to files that are renamed/removed \n> from version to version?\n>\n> - If the timestamps change on a file but the actual content does not, \n> will git treat it as a non-change once it realizes the content hasn't \n> changed?\n>\n> - Last, if after loading the repository I find another version of the \n> files that predates those I've loaded, or are intermediate between two \n> commits I've already loaded, is there a way to go say that commit B is \n> actually the ancestor of commit C? (i.e. a->c becomes a->b->c if you \n> were to visualize the commit timeline or do diffs) Or do I just reload \n> the tarballs in order to achieve this?\n\nThere is a script which will import sources from multiple tarballs, \ncreating a commit with the contents of each tarball. It's in the git \nrepository under contrib/fast-import/import-tars.perl.\n\ntom\n"},{"id":"181827","messageId":"B375E525C4704EA8807B5A59257B690B@PhilipOakley","threadId":"29276","inReplyTo":"4EFFA868.50605@dbservice.com","subject":"Re: How to deal with historic tar-balls","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":null,"receivedAt":"2012-01-01T16:27:31Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"From: \"Tomas Carnecky\" <tom@dbservice.com> Sent: Sunday, January 01, 2012\n12:27 AM\n>On 12/31/11 8:04 PM, nn6eumtr wrote:\n>> I have a number of older projects that I want to bring into a git\n>> repository. They predate a lot of the popular scm systems, so they are\n>> primarily a collection of tarballs today.\nI'm doing a similar thing with a set of zip files. I grouped mine into\nbatches for easier checking and putting on to separate branches. Planning\nyour branch requirements is probably the biggest task, and will depend on\nhow you hope to use the new repo.\n\n>> I'm fairly new to git so I have a couple questions related to this:\n>>\n>> - What is the best approach for bringing them in? Do I just create a\n>> repository, then unpack the files, commit them, clean out the\n>> directory unpack the next tarball, and repeat until everything is loaded?\nEssentially yes; Obviously if you have an organisation in mind then you can\nintroduce maintenance branches etc as you develop the import.\n\nIf it is simply to create a nice history that isn't really looked at, then a\nsimple linear model is OK. If you need to keep old mintenance versions and\nobtain diffs with newer versions then look at the various branching models\nand populate appropriately.\n\n>>\n>> - Do I need to pay special attention to files that are renamed/removed\n>> from version to version?\nNo; It is only if you use those file dates to determine the implied date for\nthe commit.\n\n>>\n>> - If the timestamps change on a file but the actual content does not,\n>> will git treat it as a non-change once it realizes the content hasn't\n>> changed?\nCorrect. In fact git doesn't record the time stamps anyway. It simply\nrecords the content, and structure, of the snapshot.\n\n>>\n>> - Last, if after loading the repository I find another version of the\n>> files that predates those I've loaded, or are intermediate between two\n>> commits I've already loaded, is there a way to go say that commit B is\n>> actually the ancestor of commit C? (i.e. a->>c becomes a->>b->>c if you\n>> were to visualize the commit timeline or do diffs) Or do I just reload\n>> the tarballs in order to achieve this?\nYou can use 'grafts' as a mechanism to re-arrange the commit order, and/or\njoin partial repos, and then use git filter-branch to re-write the lot as a\nsingle cohesive repo. But this 'hack' does re-write all the commit SHA1\nvalues, so you should minimise the number of times that happens...\n\nIt is worth capturing your import sequence as a script so that you can wash \n/ rinse / repeat as often as needed to get a result you like.\n\n> There is a script which will import sources from multiple tarballs,\n> creating a commit with the contents of each tarball. It's in the git\n> repository under contrib/fast-import/import-tars.perl.\nI wasn't aware of those scripts. I'll be having a look at the zip import\nscript for my needs.\n\nMy extra problem is that almost all my zips have an extra top level\ndirectory that changes its name for every zip (but some don't..). The TLD\nchanges confuses the git rename detection if I don't remove them before\ncommitting. Fortunately it's an internal development project with no formal\nreleases so creating the history is a bit of a personal project which\ndoesn't affect ongoing development (which is the crunch question for\nfidelity of the repo you create).\n\n> tom\nPhilip\n"},{"id":"181828","messageId":"4F00AE3D.9050102@dirk.my1.cc","threadId":"29276","inReplyTo":"4EFFA868.50605@dbservice.com","subject":"Re: How to deal with historic tar-balls","fromName":"Dirk Süsserott","fromEmail":"newsletter@dirk.my1.cc","sentAt":"2012-01-01T19:04:29Z","receivedAt":"2012-01-01T19:04:29Z","isPatch":false,"sender":{"key":"newsletter@dirk.my1.cc","avatar":null},"body":"Am 01.01.2012 01:27 schrieb Tomas Carnecky:\n> On 12/31/11 8:04 PM, nn6eumtr wrote:\n>> I have a number of older projects that I want to bring into a git\n>> repository. They predate a lot of the popular scm systems, so they are\n>> primarily a collection of tarballs today.\n>>\n>> I'm fairly new to git so I have a couple questions related to this:\n>>\n>> - What is the best approach for bringing them in? Do I just create a\n>> repository, then unpack the files, commit them, clean out the\n>> directory unpack the next tarball, and repeat until everything is loaded?\n>>\n>> - Do I need to pay special attention to files that are renamed/removed\n>> from version to version?\n>>\n>> - If the timestamps change on a file but the actual content does not,\n>> will git treat it as a non-change once it realizes the content hasn't\n>> changed?\n>>\n>> - Last, if after loading the repository I find another version of the\n>> files that predates those I've loaded, or are intermediate between two\n>> commits I've already loaded, is there a way to go say that commit B is\n>> actually the ancestor of commit C? (i.e. a->c becomes a->b->c if you\n>> were to visualize the commit timeline or do diffs) Or do I just reload\n>> the tarballs in order to achieve this?\n> \n> There is a script which will import sources from multiple tarballs,\n> creating a commit with the contents of each tarball. It's in the git\n> repository under contrib/fast-import/import-tars.perl.\n> \n> tom\n\n@tom: True. I didn't know about that script, but it should work.\n\n@nn6eumtr: Basically your workflow is perfect. But let me give you some\nexplanation:\n\ngit init\nforeach archive in *.tar; do\n    tar xf $archive\n    git add --all .\n    git commit -m \"Added $archive\"\n    # now remove everything except for the .git directory\n    # with regular shell commands (rm -rf *). Also remove\n    # any dot-files (and the tarball itself, if it's in the\n    # current directory).\ndone\n\nNotice the '--all' switch to 'git add': Normally, 'git add .' adds all\nfiles that match the given pattern '.', i.e. all files in the current\ndirectory (and below, it's recursive). The '--all' switch together with\nthe pattern '.' adds or updates all files already known to git *AND*\nadds the files not yet known *AND* removes the files that are no longer\nin the working tree. That's exactly what you want.\n\nConsider archive1.tar with files A, B, C:\n\n  git add --all . # will add A, B, and C\n\nNow remove A, B, C, and unpack archive2.tar. Assume it has files B, C,\nD. A was deleted, B was changed, C is unchanged, D is new.\n\n  git add --all . # will remove A, add B, leave C, add D.\n\ngit will notice that C hasn't changed its content (timestamp doesn't\nmatter).\n\nWithout the '--all' switch, git would simply add B and D.\n\nThere is no problem re-arranging the history after your import (see \"git\nrebase --help\", especially the --interactive section), but then you\nprobably will have conflicts and have to resolve them. I'd suggest to\nre-start the import instead.\n\nPlease note that \"for archive in *.tar\" will pick the tarballs in\nlexicographical order. That might not be your intention.\n\nHTH,\n    Dirk\n"},{"id":"181831","messageId":"70916F7E9F934AD3A0DB00C8D6DB6751@PhilipOakley","threadId":"29276","inReplyTo":"B375E525C4704EA8807B5A59257B690B@PhilipOakley","subject":"Re: How to deal with historic tar-balls","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":null,"receivedAt":"2012-01-01T19:57:26Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"From: \"Philip Oakley\" <philipoakley@iee.org> Sent: Sunday, January 01, 2012 \n6:30 PM\n> From: \"Tomas Carnecky\" <tom@dbservice.com> Sent: Sunday, January 01, 2012\n> 12:27 AM\n>>On 12/31/11 8:04 PM, nn6eumtr wrote:\n>>> I have a number of older projects that I want to bring into a git\n>>> repository. They predate a lot of the popular scm systems, so they are\n>>> primarily a collection of tarballs today.\n> I'm doing a similar thing with a set of zip files. I grouped mine into\n> batches for easier checking and putting on to separate branches. Planning\n> your branch requirements is probably the biggest task, and will depend on\n> how you hope to use the new repo.\n>\n<snip>\n>> There is a script which will import sources from multiple tarballs,\n>> creating a commit with the contents of each tarball. It's in the git\n>> repository under contrib/fast-import/import-tars.perl.\n> I wasn't aware of those scripts. I'll be having a look at the zip import\n> script for my needs.\n\nIs there a mechanism for either having fast-import respect a .gitignore,\nor determining if a given file/path should be ignored?\nMy zips contain a lot of compile by-products that should be excluded from \nthe repo.\n\n> My extra problem is that almost all my zips have an extra top level\n> directory that changes its name for every zip (but some don't..). The TLD\n> changes confuses the git rename detection if I don't remove them before\n> committing. Fortunately it's an internal development project with no \n> formal\n> releases so creating the history is a bit of a personal project which\n> doesn't affect ongoing development (which is the crunch question for\n> fidelity of the repo you create).\n>\n>> tom\n> Philip\n"},{"id":"181838","messageId":"4C50794C7EED42A0B1A25ABD77CE7DB0@PhilipOakley","threadId":"29276","inReplyTo":"B375E525C4704EA8807B5A59257B690B@PhilipOakley","subject":"Re: How to deal with historic tar-balls","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":null,"receivedAt":"2012-01-02T09:25:08Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"From: \"Philip Oakley\" <philipoakley@iee.org>: Sunday, January 01, 2012 6:30 \nPM\n> From: \"Tomas Carnecky\" <tom@dbservice.com> : Sunday, January 01, 2012 \n> 12:27 AM\n>>On 12/31/11 8:04 PM, nn6eumtr wrote:\n>>> I have a number of older projects that I want to bring into a git\n>>> repository. They predate a lot of the popular scm systems, so they are\n>>> primarily a collection of tarballs today.\n>> There is a script which will import sources from multiple tarballs,\n>> creating a commit with the contents of each tarball. It's in the git\n>> repository under contrib/fast-import/import-tars.perl.\n> I wasn't aware of those scripts. I'll be having a look at the zip import\n> script for my needs.\n>\n>> tom\n> Philip\n>\nI had a look at the script but Python isn't part of the Msysgit install, so \nthe example wouldn't run.\n\nAlso I couldn't see how the \"fast_import.write(\" method was being created - \nmy ignorance of Python? Otherwise I could look at scripting it.\n\nPhilip \n"},{"id":"181840","messageId":"4F01F6D2.8020005@dirk.my1.cc","threadId":"29276","inReplyTo":"4C50794C7EED42A0B1A25ABD77CE7DB0@PhilipOakley","subject":"Re: How to deal with historic tar-balls","fromName":"Dirk Süsserott","fromEmail":"newsletter@dirk.my1.cc","sentAt":"2012-01-02T18:26:26Z","receivedAt":"2012-01-02T18:26:26Z","isPatch":false,"sender":{"key":"newsletter@dirk.my1.cc","avatar":null},"body":"Am 02.01.2012 11:07 schrieb Philip Oakley:\n> From: \"Philip Oakley\" <philipoakley@iee.org>: Sunday, January 01, 2012\n> 6:30 PM\n>> From: \"Tomas Carnecky\" <tom@dbservice.com> : Sunday, January 01, 2012\n>> 12:27 AM\n>>> On 12/31/11 8:04 PM, nn6eumtr wrote:\n>>>> I have a number of older projects that I want to bring into a git\n>>>> repository. They predate a lot of the popular scm systems, so they are\n>>>> primarily a collection of tarballs today.\n>>> There is a script which will import sources from multiple tarballs,\n>>> creating a commit with the contents of each tarball. It's in the git\n>>> repository under contrib/fast-import/import-tars.perl.\n>> I wasn't aware of those scripts. I'll be having a look at the zip import\n>> script for my needs.\n>>\n>>> tom\n>> Philip\n>>\n> I had a look at the script but Python isn't part of the Msysgit install,\n> so the example wouldn't run.\n> \n> Also I couldn't see how the \"fast_import.write(\" method was being\n> created - my ignorance of Python? Otherwise I could look at scripting it.\n> \n> Philip\n\nPhilip,\n\nI'm not a Python guy, but I think fast_import.write() writes sth. to\nwhatever the popen() call in line 24 returned:\n\n  fast_import = popen('git fast-import --quiet', 'w')\n\nI guess it returns a filehandle and 'git fast-import' reads its data\nfrom stdin. My guess is, that -- instead of writing to that pipe -- you\ncould as well write everything to a temporary file and finally call\n\n  git fast-import < $tempfile\n\nBut that's only a guess.\n\nDirk\n"},{"id":"181938","messageId":"E6EC04F6AE4545F8923507F7CC4A6B07@PhilipOakley","threadId":"29276","inReplyTo":"4F01F6D2.8020005@dirk.my1.cc","subject":"Re: How to deal with historic tar-balls","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":null,"receivedAt":"2012-01-04T19:20:00Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Many thanks - That explanation works for me. I just hadn't seen the \nassociation.\nPhilip\n\nFrom: \"Dirk Süsserott\" <newsletter@dirk.my1.cc>: Monday, January 02, 2012 \n6:26 PM\n> Am 02.01.2012 11:07 schrieb Philip Oakley:\n>>> From: \"Tomas Carnecky\" <tom@dbservice.com> : Sunday, January 01, 2012\n>>> 12:27 AM\n>>>> On 12/31/11 8:04 PM, nn6eumtr wrote:\n>>>>> I have a number of older projects that I want to bring into a git\n>>>>> repository. They predate a lot of the popular scm systems, so they are\n>>>>> primarily a collection of tarballs today.\n>>>> There is a script which will import sources from multiple tarballs,\n>>>> creating a commit with the contents of each tarball. It's in the git\n>>>> repository under contrib/fast-import/import-tars.perl.\n>>> I wasn't aware of those scripts. I'll be having a look at the zip import\n>>> script for my needs.\n>>>\n>>>> tom\n>>> Philip\n>>>\n>> I had a look at the script but Python isn't part of the Msysgit install,\n>> so the example wouldn't run.\n>>\n>> Also I couldn't see how the \"fast_import.write(\" method was being\n>> created - my ignorance of Python? Otherwise I could look at scripting it.\n>>\n>> Philip\n>\n> Philip,\n>\n> I'm not a Python guy, but I think fast_import.write() writes sth. to\n> whatever the popen() call in line 24 returned:\n>\n>  fast_import = popen('git fast-import --quiet', 'w')\n>\n> I guess it returns a filehandle and 'git fast-import' reads its data\n> from stdin. My guess is, that -- instead of writing to that pipe -- you\n> could as well write everything to a temporary file and finally call\n>\n>  git fast-import < $tempfile\n>\n> But that's only a guess.\n>\n> Dirk\n>> \n"},{"id":"181978","messageId":"4F05C0E2.4050101@gmail.com","threadId":"29276","inReplyTo":"4EFF5CDA.5050809@gmail.com","subject":"Re: How to deal with historic tar-balls","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-01-05T15:25:22Z","receivedAt":"2012-01-05T15:25:22Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 12/31/2011 1:04 PM, nn6eumtr wrote:\n> I have a number of older projects that I want to bring into a git\n> repository. They predate a lot of the popular scm systems, so they\n> are primarily a collection of tarballs today.\n>\n> I'm fairly new to git so I have a couple questions related to this:\n>\n> - What is the best approach for bringing them in? Do I just create a\n>  repository, then unpack the files, commit them, clean out the\n> directory unpack the next tarball, and repeat until everything is\n> loaded?\n>\n> - Do I need to pay special attention to files that are\n> renamed/removed from version to version?\n>\n> - If the timestamps change on a file but the actual content does not,\n>  will git treat it as a non-change once it realizes the content\n> hasn't changed?\n>\n> - Last, if after loading the repository I find another version of the\n>  files that predates those I've loaded, or are intermediate between\n> two commits I've already loaded, is there a way to go say that commit\n> B is actually the ancestor of commit C? (i.e. a->c becomes a->b->c if\n> you were to visualize the commit timeline or do diffs) Or do I just\n> reload the tarballs in order to achieve this?\n>\nThe git-rm manpage contains instructions under the \"vendor code drop\"\nsection on how to do this.  I imagine you will want to do each one\nmanually instead of queueing them up in a script because you are likely \ngoing to want to do appropriate clean up of the working tree in each \niteration before committing.  This is where you would review \nrenames/removes with git-status before you git-add and git-commit. \nAlso, if you are tracking permissions in git (the executable bit) then \nyou will want to filter out any noise generated by frivolous permissions \nchanges between the tarball contents.\n\nIn regard to inserting tarballs into the history that depends on when \nyou think you plan on doing that.  You are only going to be able to do \nthat before the history is published (made \"public\" for other repos to \npull down).  Otherwise you will be rewriting published history which is \na big no-no (see git-rebase manpage).  I suggest you do your homework \nand order them properly before you start because that will be less work. \n  If you still find that you missed something then you can use \ninteractive git-rebase to insert.  I'm assuming a single \"master\" branch \nwith linear history is your desired end result.  If you want to create \nmaintenance branches showing release history then you will definitely \nneed to do your homework first (see gitworkflow manpage).\n\nIf you venture into rebase territory by rewriting history (inserting \nmissed tarballs in between older commits) you will need to be sure to \nreview your automatic merge resolutions.  Git only generates \nmerge-conflicts on same-file-same-line conflicts.  It will auto-merge \nsame-file-different-line changes.\n\nYou also need to ask yourself if you really need a history of all those \nversions.  To exaggerate, if all you really need is the current state \nthen you need to ask yourself if it's worth the effort to record the \nprevious states.  Maybe what you want is something in-between (a happy \nmedium).\n\nIn regard to the 'start-over' method of inserting missed tarballs you \nwould just git-reset --hard to the commit you want to insert on-top-of, \nadd the tarball, and then re-apply the subsequent tarballs.  If you are \ndoing cleanup between commits then the rebase or cherry-pick of the \nalready cleaned-up subsequent commits from the \"old-branch\" (previous \nattempt) onto the 'do-over' branch will likely be easier.  (You can just \ndo 'git branch old-branch' on your branch before the git-reset --hard \n(do-over) and that will give you a \"backup copy\" of the \"previous \nattempt\" called \"old-branch\" that you can salvage already-done-work from \nby using rebase or cherry-pick.)\n\nHope this helps.\n\nv/r,\nneal\n"},{"id":"182052","messageId":"4F079BA1.3060907@gmail.com","threadId":"29276","inReplyTo":"4F05C0E2.4050101@gmail.com","subject":"Re: How to deal with historic tar-balls","fromName":"nn6eumtr","fromEmail":"nn6eumtr@gmail.com","sentAt":"2012-01-07T01:10:57Z","receivedAt":"2012-01-07T01:10:57Z","isPatch":false,"sender":{"key":"nn6eumtr@gmail.com","avatar":null},"body":"Thanks for the response, there is lots of good information there.\n\nOne clarification - can you track renames in git? I tried using git mv \nbut from the status output it looks like it deleted the old file  and \nadded the new file. I was expecting it to record some sort of indicator \nof the name change, instead it looks like a short-cut for delete & add, \nthe docs aren't clear if that is the case.\n\nOn 1/5/2012 10:25 AM, Neal Kreitzinger wrote:\n...\n> going to want to do appropriate clean up of the working tree in each\n> iteration before committing. This is where you would review\n> renames/removes with git-status before you git-add and git-commit. Also,\n> if you are tracking permissions in git (the executable bit) then you\n> will want to filter out any noise generated by frivolous permissions\n> changes between the tarball contents.\n...\n"},{"id":"182054","messageId":"87lipk46y2.fsf@thomas.inf.ethz.ch","threadId":"29276","inReplyTo":"4F079BA1.3060907@gmail.com","subject":"Re: How to deal with historic tar-balls","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2012-01-07T01:50:13Z","receivedAt":"2012-01-07T01:50:13Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"nn6eumtr <nn6eumtr@gmail.com> writes:\n\n> Thanks for the response, there is lots of good information there.\n>\n> One clarification - can you track renames in git? I tried using git mv\n> but from the status output it looks like it deleted the old file  and\n> added the new file. I was expecting it to record some sort of\n> indicator of the name change, instead it looks like a short-cut for\n> delete & add, the docs aren't clear if that is the case.\n\nGit only stores snapshots; so for an ordinary (non-merge, non-root)\ncommit, you have the \"before\" (parent) and \"after\" (commit's) snapshot.\nEverything is generated on the fly from that, including diffs, heuristic\nrename detection, pickaxe, ...\n\nTo apply rename detection when diffing (e.g. in diff, log, show,\nformat-patch), use the -M flag.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"182075","messageId":"4F089A95.9050300@gmail.com","threadId":"29276","inReplyTo":"4F079BA1.3060907@gmail.com","subject":"Re: How to deal with historic tar-balls","fromName":"Neal Kreitzinger","fromEmail":"nkreitzinger@gmail.com","sentAt":"2012-01-07T19:18:45Z","receivedAt":"2012-01-07T19:18:45Z","isPatch":false,"sender":{"key":"nkreitzinger@gmail.com","avatar":null},"body":"On 1/6/2012 7:10 PM, nn6eumtr wrote:\n> Thanks for the response, there is lots of good information there.\n>\n> One clarification - can you track renames in git? I tried using git mv \n> but from the status output it looks like it deleted the old file  and \n> added the new file. I was expecting it to record some sort of \n> indicator of the name change, instead it looks like a short-cut for \n> delete & add, the docs aren't clear if that is the case.\n>\n(note: top-posting is not advised.)\nYou are exactly right in your observation:  git-mv is only a shortcut \nfor 'remove old then add new'.  Git does not explicitly track \n\"renames\".  It can detect renames easily in the cases where you really \njust renamed the file and left the contents the same.  Git tracks \ncontent (and trees) as opposed to files (and file names).  Git stores \nthe 'blob', aka 'contents' of files in the object store.  So if you have \n30 files with different names and the exact same contents in your work \ntree they are stored as a single blob in the .git/objects \"object \nstore\".  If some of your \"renames\" are really \"I renamed it and then I \nmodified it\" then git will have a harder time detecting the \"rename\" \ndepending on how much you modified it.  In such cases what you really \ndid is arguably not a \"rename\" anyway.  You can record your \"renames\" \nmanually in your commit message if appropriate.  If you have 5 minutes \nyou can watch this video from the 15:00 min to 20:59 min marks to get an \nexplanation of git-mv and rename-detection: \nhttp://www.youtube.com/watch?v=j45cs5_nY2k (youtube searchstring: 'git \ngoogle tech talks', result: 'contributing with git'.)\n\nv/r,\nneal\n"}]}