{"thread":{"id":"21991","subject":"Git as electronic lab notebook","startedAt":"2009-12-19T12:23:04Z","lastAt":"2009-12-20T04:55:02Z","messageCount":6,"participants":["Thomas Johnson","Ciprian Dorin, Craciun","Johan 't Hart","Nicolas Pitre","Bill Lear"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"130136","messageId":"loom.20091219T130946-844@post.gmane.org","threadId":"21991","inReplyTo":null,"subject":"Git as electronic lab notebook","fromName":"Thomas Johnson","fromEmail":"thomas.j.johnson@gmail.com","sentAt":"2009-12-19T12:23:04Z","receivedAt":"2009-12-19T12:23:04Z","isPatch":false,"sender":{"key":"thomas.j.johnson@gmail.com","avatar":null},"body":"Hello group,\n\nI've been using git on a few different projects over the last couple of months,\nand as a former svn user I really like it. Recently, I've been using it as an\n'electronic lab notebook' for an empirical project. My workflow looks like this:\n1. Start with the stable code base on head\n2. Create  and change to branch 'Experiment123'\n3. Make some changes\n4. Run the program, which generates a giant (10MB-4G) output text file,\nExperiment123.log. Update my LabNotebook.txt file.\n5. Were the new changes helpful?\n5.yes: Bzip Experiment123.log, and commit it on the branch. Merge the\nExperiment123 branch to head and goto 1.\n5.no: Bzip Experiment123.log, and commit it on the branch. Merge LabNotebook.txt\nand Experiment123.log back to head. Switch back to head and goto 1.\n\nThe thing is, Experiment123.log is going to be very similar to Experiment122.log\nand Experiment124.log except for a few details. My understanding is that git is\ngreat at compressing groups of files like this, is that correct? Should I not be\nbzipping them myself? On the other hand, I don't want HEAD to contain hundreds\nof gigs of uncompressed files that bzip down to only a few hundred megs.\n\nAny thoughts on the workflow itself would also be very welcome.\n"},{"id":"130140","messageId":"8e04b5820912190538v2e9ef109me3a1515040127b39@mail.gmail.com","threadId":"21991","inReplyTo":"loom.20091219T130946-844@post.gmane.org","subject":"Re: Git as electronic lab notebook","fromName":"Ciprian Dorin, Craciun","fromEmail":"ciprian.craciun@gmail.com","sentAt":"2009-12-19T13:38:32Z","receivedAt":"2009-12-19T13:38:32Z","isPatch":false,"sender":{"key":"ciprian.craciun@gmail.com","avatar":"https://gravatar.com/avatar/9685eb13288a2c28bdde10acdcda9f15d495efad7b12bf49e3525eb7b8acf944?d=mp&s=160"},"body":"On Sat, Dec 19, 2009 at 2:23 PM, Thomas Johnson\n<thomas.j.johnson@gmail.com> wrote:\n> Hello group,\n>\n> I've been using git on a few different projects over the last couple of months,\n> and as a former svn user I really like it. Recently, I've been using it as an\n> 'electronic lab notebook' for an empirical project. My workflow looks like this:\n> 1. Start with the stable code base on head\n> 2. Create  and change to branch 'Experiment123'\n> 3. Make some changes\n> 4. Run the program, which generates a giant (10MB-4G) output text file,\n> Experiment123.log. Update my LabNotebook.txt file.\n> 5. Were the new changes helpful?\n> 5.yes: Bzip Experiment123.log, and commit it on the branch. Merge the\n> Experiment123 branch to head and goto 1.\n> 5.no: Bzip Experiment123.log, and commit it on the branch. Merge LabNotebook.txt\n> and Experiment123.log back to head. Switch back to head and goto 1.\n>\n> The thing is, Experiment123.log is going to be very similar to Experiment122.log\n> and Experiment124.log except for a few details. My understanding is that git is\n> great at compressing groups of files like this, is that correct? Should I not be\n> bzipping them myself? On the other hand, I don't want HEAD to contain hundreds\n> of gigs of uncompressed files that bzip down to only a few hundred megs.\n>\n> Any thoughts on the workflow itself would also be very welcome.\n\n\n    I have used myself such a similar workflow for parametric studies\non some genetic algorithms, and below are my observations related to\nyour question:\n    * saving the entire log file (either zipped or not) in the\nrepository has some drawbacks with repository clonning; (in my setup\nI've runned the tests in parallel on a different machine, and used Git\nto synchronize between the development machine and the test machine;)\nthe problem lies in the fact that when I wanted to \"clean\" the test\nmachine and start over I had to clone the repository, which also held\nall the unneeded log files;\n    * (actually I've used two Git repositories -- one for the actual\nsource code where I make the commits by hand, and another one which I\nuse for the synchronization;)\n    * even if you prefer having the logs, it's best to let Git handle\nthe compression; because even if only some small parts change from the\noriginal txt file, I would guess that the BZip-ped file looks quite\ndifferent;\n    * maybe it would be better than instead of holding the experiment\nlog, you just keep a sumarization of it (only the important stuff);\nand even if you do need the entire log, you could always recreate it\nby running the code again; (this was the road I took in the end, by\nkeeping a small SQLite database of each experiment;)\n    * (and of course there is also another little trick I've used:\njust put the logs file in a `log` directory which is \"git-ignored\",\nthat way you can switch between branches, but Git won't touch the\n`log` directory, unless you force it by issuing `git clean -f -d -x`;)\n\n    Hope I've been useful,\n    Ciprian.\n"},{"id":"130168","messageId":"4B2D6CA5.3070304@gmail.com","threadId":"21991","inReplyTo":"8e04b5820912190538v2e9ef109me3a1515040127b39@mail.gmail.com","subject":"Re: Git as electronic lab notebook","fromName":"Johan 't Hart","fromEmail":"johanthart@gmail.com","sentAt":"2009-12-20T00:15:33Z","receivedAt":"2009-12-20T00:15:33Z","isPatch":false,"sender":{"key":"johanthart@gmail.com","avatar":"https://gravatar.com/avatar/f7bd2928a1a23e971f4ab68a6c7b35d7ff0efb1c225c570430badf90e404e50b?d=mp&s=160"},"body":"Ciprian Dorin, Craciun schreef:\n> On Sat, Dec 19, 2009 at 2:23 PM, Thomas Johnson\n> <thomas.j.johnson@gmail.com> wrote:\n\n>> 4. Run the program, which generates a giant (10MB-4G) output text file,\n>> Experiment123.log. Update my LabNotebook.txt file.\n\n>     * even if you prefer having the logs, it's best to let Git handle\n> the compression; because even if only some small parts change from the\n> original txt file, I would guess that the BZip-ped file looks quite\n> different;\n>\n\nIs git able to handle 4Gig files? I've heard git loads every file \ncompletely in memory before handling it...\n"},{"id":"130175","messageId":"alpine.LFD.2.00.0912192212400.28241@xanadu.home","threadId":"21991","inReplyTo":"4B2D6CA5.3070304@gmail.com","subject":"Re: Git as electronic lab notebook","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2009-12-20T03:15:00Z","receivedAt":"2009-12-20T03:15:00Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 20 Dec 2009, Johan 't Hart wrote:\n\n> Is git able to handle 4Gig files? I've heard git loads every file completely\n> in memory before handling it...\n\nRight.  Sowith current Git you will be able to deal with 4GB files only \nif you have a 64-bit machine and more than 4GB of RAM.\n\n\nNicolas\n"},{"id":"130176","messageId":"19245.43871.588697.532035@lisa.zopyra.com","threadId":"21991","inReplyTo":"alpine.LFD.2.00.0912192212400.28241@xanadu.home","subject":"Re: Git as electronic lab notebook","fromName":"Bill Lear","fromEmail":"rael@zopyra.com","sentAt":"2009-12-20T04:43:11Z","receivedAt":"2009-12-20T04:43:11Z","isPatch":false,"sender":{"key":"rael@zopyra.com","avatar":"https://gravatar.com/avatar/c4f2d2790ca3828d3b4e7dfebabf61d2fe94fd82fa49cdac2a5295dd2d46a874?d=mp&s=160"},"body":"On Saturday, December 19, 2009 at 22:15:00 (-0500) Nicolas Pitre writes:\n>On Sun, 20 Dec 2009, Johan 't Hart wrote:\n>\n>> Is git able to handle 4Gig files? I've heard git loads every file completely\n>> in memory before handling it...\n>\n>Right.  Sowith current Git you will be able to deal with 4GB files only \n>if you have a 64-bit machine and more than 4GB of RAM.\n\n??\n\n% uname -a\nLinux pppp 2.6.31.6-166.fc12.i686 #1 SMP Wed Dec 9 11:14:59 EST 2009 i686 i686 i386 GNU/Linux\n% cat /proc/meminfo  | grep MemTotal\nMemTotal:        3095296 kB\n% mkdir gogle\n% cd gogle\n% git init\n% dd if=/dev/zero of=zerofile.tst bs=1k count=4700000\n% git add *\n% git commit -a -m new\n[master (root-commit) 35a25be] new\n 1 files changed, 0 insertions(+), 0 deletions(-)\n create mode 100644 zerofile.tst\n% git --version\ngit version 1.6.5.7\n\nSeems ok to me...\n\nThough, I find this interesting:\n\n% git log -p\ncommit 35a25be3fff2f8bbd6ec22c94b9a5c0d66053d21\nAuthor: Bill Lear <rael@zopyra.com>\nDate:   Sat Dec 19 22:38:48 2009 -0600\n\n    new\n\ndiff --git a/zerofile.tst b/zerofile.tst\nnew file mode 100644\nindex 0000000..e5bd39d\nBinary files /dev/null and b/zerofile.tst differ\n\n\nBill\n"},{"id":"130177","messageId":"alpine.LFD.2.00.0912192347250.28241@xanadu.home","threadId":"21991","inReplyTo":"19245.43871.588697.532035@lisa.zopyra.com","subject":"Re: Git as electronic lab notebook","fromName":"Nicolas Pitre","fromEmail":"nico@fluxnic.net","sentAt":"2009-12-20T04:55:02Z","receivedAt":"2009-12-20T04:55:02Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 19 Dec 2009, Bill Lear wrote:\n\n> On Saturday, December 19, 2009 at 22:15:00 (-0500) Nicolas Pitre writes:\n> >On Sun, 20 Dec 2009, Johan 't Hart wrote:\n> >\n> >> Is git able to handle 4Gig files? I've heard git loads every file completely\n> >> in memory before handling it...\n> >\n> >Right.  Sowith current Git you will be able to deal with 4GB files only \n> >if you have a 64-bit machine and more than 4GB of RAM.\n> \n> ??\n> \n> % uname -a\n> Linux pppp 2.6.31.6-166.fc12.i686 #1 SMP Wed Dec 9 11:14:59 EST 2009 i686 i686 i386 GNU/Linux\n> % cat /proc/meminfo  | grep MemTotal\n> MemTotal:        3095296 kB\n> % mkdir gogle\n> % cd gogle\n> % git init\n> % dd if=/dev/zero of=zerofile.tst bs=1k count=4700000\n> % git add *\n> % git commit -a -m new\n> [master (root-commit) 35a25be] new\n>  1 files changed, 0 insertions(+), 0 deletions(-)\n>  create mode 100644 zerofile.tst\n> % git --version\n> git version 1.6.5.7\n> \n> Seems ok to me...\n\nThat's the easy part.  Diffing such files and delta compressing them, or \neven checking them out especially when delta compressed, just won't work \nif you don't have the RAM.  Fixing this limitation would introduce \nsignificant complexity in the code that no one felt was worth it.\n\nI had some thoughts about supporting the addition of really huge files \nin a Git repository where only add/commit/checkout/fetch/push would work \nwith no delta compression.  That didn't materialized yet though.\n\n\nNicolas\n"}]}