{"thread":{"id":"43099","subject":"Re: Fast access git-rev-list output: some OS knowledge required","startedAt":"2006-12-06T19:24:42Z","lastAt":"2006-12-09T12:15:12Z","messageCount":17,"participants":["Andreas Ericsson","Marco Costalba","Shawn Pearce","Johannes Schindelin","Linus Torvalds","Michael K. Edwards"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"298436","messageId":"e5bfff550612061124jcd0d94em47793710866776e7@mail.gmail.com","threadId":"43099","inReplyTo":null,"subject":"Fast access git-rev-list output: some OS knowledge required","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-12-06T19:24:42Z","receivedAt":"2006-12-06T19:24:42Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"I ask help to the list because my knowledge on this is not enough.\n\nCurrently qgit uses, socket based, QProcess class to read data from\n'git rev-list'  when loading the repository at startup.\n\nThe time it takes to read, without processing, the whole Linux tree\nwith this approach it's almost _double_ of the time it takes 'git\nrev-list' to write to a file:\n\n$git rev-list --header --boundary --parents --topo-order HEAD >> tmp.txt\n\nWe are talking of about 7s against less then 4s, on my box (warm cache).\n\nSo I have a patch to make 'git rev-list' writing into a temporary file\nand then read it in memory, perhaps it's not the cleaner way, but it's\nfaster, about 1s less.\n\nI have browsed Qt sources and found that QProcess uses internal\nbuffers that are then copied again before to be used by the\napplication. File approach uses a call to read() /fread() buired\ninside the Qt's QFile class, and no intermediate buffers, so perhaps\nthis could be the reason the second way it's faster.\n\n\nAnyway there are some issues:\n\n1) File tmp.txt is deleted as soon as read, but this is not enough\nsometimes to avoid a costly and wasteful write access to disk by the\nOS. What is the easiest, portable way to create a temporary 'in memory\nonly' file, with no disk access? Or at least delay the HD write access\nenough to be able to read and delete the file before the fist block of\ntmp.txt is flushed to disk?\n\n2) There is a faster/cleaner (and *safe* ) way to access directly 'git\nrev-list' output, something like (just as an example):\n\n$git rev-list --header --boundary --parents --topo-order HEAD >> /dev/mem\n\nOr something similar, possibly _simple_ and _portable_ , so to be able\nto copy the big amount of 'git rev-list' output just once (about 30MB\nwith current tree).\n\n\n3) Other suggestions?  ;-)\n\n\nThanks\n"},{"id":"296399","messageId":"20061206192800.GC20320@spearce.org","threadId":"43099","inReplyTo":"e5bfff550612061124jcd0d94em47793710866776e7@mail.gmail.com","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-06T19:28:01Z","receivedAt":"2006-12-06T19:28:01Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Marco Costalba <mcostalba@gmail.com> wrote:\n> The time it takes to read, without processing, the whole Linux tree\n> with this approach it's almost _double_ of the time it takes 'git\n> rev-list' to write to a file:\n> \n> 3) Other suggestions?  ;-)\n\nThe revision listing machinery is fairly well isolated behind some\npretty clean APIs in Git.  Why not link qgit against libgit.a and\njust do the revision listing in process?\n\n-- \n"},{"id":"297782","messageId":"e5bfff550612061134r3725dcbu2ff2dd6284fcd651@mail.gmail.com","threadId":"43099","inReplyTo":"20061206192800.GC20320@spearce.org","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-12-06T19:34:31Z","receivedAt":"2006-12-06T19:34:31Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 12/6/06, Shawn Pearce <spearce@spearce.org> wrote:\n> Marco Costalba <mcostalba@gmail.com> wrote:\n> > The time it takes to read, without processing, the whole Linux tree\n> > with this approach it's almost _double_ of the time it takes 'git\n> > rev-list' to write to a file:\n> >\n> > 3) Other suggestions?  ;-)\n>\n> The revision listing machinery is fairly well isolated behind some\n> pretty clean APIs in Git.  Why not link qgit against libgit.a and\n> just do the revision listing in process?\n>\n\nWhere can I found some documentation (yes I know RTFS, but...) or,\nbetter, an example of using the API to read git-rev-list output?\n\nif it is possible I also would like to avoid to mess with internal git\nAPI's, of course *if it is possible*  ;-)\n\nThanks\n"},{"id":"297601","messageId":"20061206194258.GD20320@spearce.org","threadId":"43099","inReplyTo":"e5bfff550612061134r3725dcbu2ff2dd6284fcd651@mail.gmail.com","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-06T19:42:58Z","receivedAt":"2006-12-06T19:42:58Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Marco Costalba <mcostalba@gmail.com> wrote:\n> On 12/6/06, Shawn Pearce <spearce@spearce.org> wrote:\n> >Marco Costalba <mcostalba@gmail.com> wrote:\n> >> The time it takes to read, without processing, the whole Linux tree\n> >> with this approach it's almost _double_ of the time it takes 'git\n> >> rev-list' to write to a file:\n> >>\n> >> 3) Other suggestions?  ;-)\n> >\n> >The revision listing machinery is fairly well isolated behind some\n> >pretty clean APIs in Git.  Why not link qgit against libgit.a and\n> >just do the revision listing in process?\n> >\n> \n> Where can I found some documentation (yes I know RTFS, but...) or,\n> better, an example of using the API to read git-rev-list output?\n\nbuiltin-rev-list.c.  :-)\n \nI think all you may need is:\n\n\t#include \"revision.h\"\n\t...\n\tstruct rev_info revs;\n\tinit_revisions(&revs, prefix);\n\trevs.abbrev = 0;\n\trevs.commit_format = CMIT_FMT_UNSPECIFIED;\n\targc = setup_revisions(argc, argv, &revs, NULL);\n\nwhere argv just a char** of the arguments you were going to hand\nto rev-list on the command line.\n\nthen get the data back:\n\n\tstatic void show_commit(struct commit *commit)\n\t{\n\t\tconst char * hex = sha1_to_hex(commit->object.sha1);\n\t\t... copy from hex to your own structures ...\n\t}\n\n\tstatic void show_object(struct object_array_entry *p)\n\t{\n\t\t/* do nothing */\n\t}\n\n\tprepare_revision_walk(&revs);\n\ttraverse_commit_list(&revs, show_commit, show_object);\n\n-- \n"},{"id":"294303","messageId":"20061206195142.GE20320@spearce.org","threadId":"43099","inReplyTo":"20061206194258.GD20320@spearce.org","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-06T19:51:42Z","receivedAt":"2006-12-06T19:51:42Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Shawn Pearce <spearce@spearce.org> wrote:\n> I think all you may need is:\n> \n> \t#include \"revision.h\"\n> \t...\n\nYou'll also need to call:\n\n\tsetup_git_directory();\n\nbefore any of the below; but that should be done once per process.\n\n> \tstruct rev_info revs;\n> \tinit_revisions(&revs, prefix);\n> \trevs.abbrev = 0;\n> \trevs.commit_format = CMIT_FMT_UNSPECIFIED;\n> \targc = setup_revisions(argc, argv, &revs, NULL);\n\nAlthough now that I think about it the library may not be enough\nof a library.  Some data (e.g. commits) will stay in memory forever\nonce loaded.  Pack files won't be released once read; a pack recently\nmade available while the application is running may not get noticed.\n\nPerhaps there is some fast IPC API supported by Qt that you could\nuse to run the revision listing outside of the main UI process,\nto eliminate the bottlenecks you are seeing and remove the problems\nnoted above?  One that doesn't involve reading from a pipe I mean...\n\n-- \n"},{"id":"294212","messageId":"e5bfff550612061208g6e4003e7ifa7dbd5ed69180c9@mail.gmail.com","threadId":"43099","inReplyTo":"20061206195142.GE20320@spearce.org","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-12-06T20:08:55Z","receivedAt":"2006-12-06T20:08:55Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 12/6/06, Shawn Pearce <spearce@spearce.org> wrote:\n> Shawn Pearce <spearce@spearce.org> wrote:\n>\n> Perhaps there is some fast IPC API supported by Qt that you could\n> use to run the revision listing outside of the main UI process,\n> to eliminate the bottlenecks you are seeing and remove the problems\n> noted above?  One that doesn't involve reading from a pipe I mean...\n>\n\nQt it's very fast in reading from files, also git-rev-list is fast in\nwrite to a file...the problem is I would not want the file to be saved\non disk, but stay cached in the OS memory for the few seconds needed\nto be written and read back, and then deleted. It's a kind of shared\nmemory at the end. But I don't know how to realize it.\n\nAlso let git-rev-list to write directly in qgit process address space\nwould be nice, indeed very nice.\n\n\n"},{"id":"297297","messageId":"20061206201816.GF20320@spearce.org","threadId":"43099","inReplyTo":"e5bfff550612061208g6e4003e7ifa7dbd5ed69180c9@mail.gmail.com","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-12-06T20:18:16Z","receivedAt":"2006-12-06T20:18:16Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Marco Costalba <mcostalba@gmail.com> wrote:\n> On 12/6/06, Shawn Pearce <spearce@spearce.org> wrote:\n> >Shawn Pearce <spearce@spearce.org> wrote:\n> >\n> >Perhaps there is some fast IPC API supported by Qt that you could\n> >use to run the revision listing outside of the main UI process,\n> >to eliminate the bottlenecks you are seeing and remove the problems\n> >noted above?  One that doesn't involve reading from a pipe I mean...\n> >\n> \n> Qt it's very fast in reading from files, also git-rev-list is fast in\n> write to a file...the problem is I would not want the file to be saved\n> on disk, but stay cached in the OS memory for the few seconds needed\n> to be written and read back, and then deleted. It's a kind of shared\n> memory at the end. But I don't know how to realize it.\n\nOn a modern Linux (probably your largest target audience) a small\nfile which has a very short lifespan (few seconds) is unlikey to\nhit the platter.  Most filesystems will put the data into buffer\ncache and delay writing to disk because temporary files are so\ncommon on UNIX.\n\nThough our resident Linux experts may chime in with more details...\n \n> Also let git-rev-list to write directly in qgit process address space\n> would be nice, indeed very nice.\n\nAnd ugly.  :-)\n\nSysV IPC (shared memory, semaphores) are messy and difficult to\nget right.  mmap against a random file in the filesystem tends\nto work better on those systems which support it well, provided\nthat the file isn't on a network mount.  But again you still need\nsemaphores or something like them to control access to the data in\nthe mmap'd region.\n\nI was thinking that maybe if Qt had a bounded buffer available for\nuse between a process and its child, that you could use that to run\nyour own \"qgit-rev-list\" child and get the data back more quickly,\nwithout the need for a temporary file.  But it doesn't look like\nthey have one.  Oh well.\n\n\nYour current temporary file approach is probably the best you can\nget, and has the simplest possible implementation.  Doing better\nwould require linking against libgit.a, and getting the core Git\nhackers to make at least the revision machinery more useful in a\nlibrary setting.\n\n-- \n"},{"id":"294484","messageId":"Pine.LNX.4.63.0612070025450.28348@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43099","inReplyTo":"20061206192800.GC20320@spearce.org","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-12-06T23:27:08Z","receivedAt":"2006-12-06T23:27:08Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 6 Dec 2006, Shawn Pearce wrote:\n\n> Marco Costalba <mcostalba@gmail.com> wrote:\n> > The time it takes to read, without processing, the whole Linux tree\n> > with this approach it's almost _double_ of the time it takes 'git\n> > rev-list' to write to a file:\n> > \n> > 3) Other suggestions?  ;-)\n> \n> The revision listing machinery is fairly well isolated behind some \n> pretty clean APIs in Git.  Why not link qgit against libgit.a and just \n> do the revision listing in process?\n\nBecause, depending on what you do, the revision machinery is not \nreentrable. For example, if you filter by filename, the history is \nrewritten in-memory to simulate a history where just that filename was \ntracked, and nothing else. These changes are not cleaned up after calling \nthe internal revision machinery.\n\nHth,\nDscho\n"},{"id":"296349","messageId":"Pine.LNX.4.64.0612061642440.3542@woody.osdl.org","threadId":"43099","inReplyTo":"Pine.LNX.4.63.0612070025450.28348@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-12-07T00:47:54Z","receivedAt":"2006-12-07T00:47:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 7 Dec 2006, Johannes Schindelin wrote:\n> \n> Because, depending on what you do, the revision machinery is not \n> reentrable. For example, if you filter by filename, the history is \n> rewritten in-memory to simulate a history where just that filename was \n> tracked, and nothing else. These changes are not cleaned up after calling \n> the internal revision machinery.\n\nWell, it really wouldn't be that hard to add a new library interface to \n\"reset object state\". We could fairly trivially either:\n\n - walk all objects in the object hashes and clear all the flags.\n - just clear all objects _and_ the hashes.\n\nYes, it implies a small amount of manual \"management\", but considering \nthat the reason it needs to be manual is that the functions simply _need_ \nthe state, is that such a big deal?\n\n"},{"id":"295984","messageId":"e5bfff550612062246m194ce235nf97149f8b041f486@mail.gmail.com","threadId":"43099","inReplyTo":"Pine.LNX.4.64.0612061642440.3542@woody.osdl.org","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-12-07T06:46:05Z","receivedAt":"2006-12-07T06:46:05Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 12/7/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>\n>\n> On Thu, 7 Dec 2006, Johannes Schindelin wrote:\n> >\n> > Because, depending on what you do, the revision machinery is not\n> > reentrable. For example, if you filter by filename, the history is\n> > rewritten in-memory to simulate a history where just that filename was\n> > tracked, and nothing else. These changes are not cleaned up after calling\n> > the internal revision machinery.\n>\n> Well, it really wouldn't be that hard to add a new library interface to\n> \"reset object state\". We could fairly trivially either:\n>\n\nSo the library approach sounds like the best?\n\nOf course in this case the producer git-rev-list and the receiver use\nthe same address space.\n\nIn the case of a temporary file data is first copied to OS disk cache\nbuffers and then again to userspace, in qgit address space. But the\nreal pain is that the temporary file is always flushed to disk after\n4-5 seconds from creation, also if under heavy read/write activity.\nThis is a problem for big repos. I really don't know how to workaround\nthis useless disk flush.\n\nFinally, what about using some kind of shared memory at run time,\ninstead of _sharing_ developer libraries ;-) ? is it too messy?\n\nProbably the concurrent reading while writing is possible without\nsyncro if the reader understands that a sequence of _two_ or more \\0\nit means the end of current write stream if producer is still running\nor the end of data if producer is not running anymore. I use a similar\napproach in the 'temporary file' patch where receiver is able to read\nwhile producer writes without explicit synchronization. In that case a\nread() of a block smaller then maximum with producer still running is\nused as the 'break' condition in the receiver while loop.\n\n\n"},{"id":"295972","messageId":"45781639.1050208@op5.se","threadId":"43099","inReplyTo":"20061206195142.GE20320@spearce.org","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-07T13:25:13Z","receivedAt":"2006-12-07T13:25:13Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Shawn Pearce wrote:\n> \n> Perhaps there is some fast IPC API supported by Qt that you could\n> use to run the revision listing outside of the main UI process,\n> to eliminate the bottlenecks you are seeing and remove the problems\n> noted above?  One that doesn't involve reading from a pipe I mean...\n> \n\nWhy not just fork() + exec() and read from the filedescriptor? You can \nup the output buffer of the forked program to something suitable, which \nmeans the OS will cache it for you until you copy it to a buffer in qgit \n(i.e., read from the descriptor).\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"296228","messageId":"Pine.LNX.4.63.0612071553090.28348@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43099","inReplyTo":"45781639.1050208@op5.se","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-12-07T14:53:58Z","receivedAt":"2006-12-07T14:53:58Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 7 Dec 2006, Andreas Ericsson wrote:\n\n> Shawn Pearce wrote:\n> > \n> > Perhaps there is some fast IPC API supported by Qt that you could use \n> > to run the revision listing outside of the main UI process, to \n> > eliminate the bottlenecks you are seeing and remove the problems noted \n> > above?  One that doesn't involve reading from a pipe I mean...\n> > \n> \n> Why not just fork() + exec() and read from the filedescriptor? You can \n> up the output buffer of the forked program to something suitable, which \n> means the OS will cache it for you until you copy it to a buffer in qgit \n> (i.e., read from the descriptor).\n\nCould somebody remind me why different processes are needed? I thought \nthat the revision machinery should be used directly, by linking to \nlibgit.a...\n\nCiao,\nDscho\n"},{"id":"293843","messageId":"4578330C.9070208@op5.se","threadId":"43099","inReplyTo":"Pine.LNX.4.63.0612071553090.28348@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-12-07T15:28:12Z","receivedAt":"2006-12-07T15:28:12Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Johannes Schindelin wrote:\n> Hi,\n> \n> On Thu, 7 Dec 2006, Andreas Ericsson wrote:\n> \n>> Shawn Pearce wrote:\n>>> Perhaps there is some fast IPC API supported by Qt that you could use \n>>> to run the revision listing outside of the main UI process, to \n>>> eliminate the bottlenecks you are seeing and remove the problems noted \n>>> above?  One that doesn't involve reading from a pipe I mean...\n>>>\n>> Why not just fork() + exec() and read from the filedescriptor? You can \n>> up the output buffer of the forked program to something suitable, which \n>> means the OS will cache it for you until you copy it to a buffer in qgit \n>> (i.e., read from the descriptor).\n> \n> Could somebody remind me why different processes are needed? I thought \n> that the revision machinery should be used directly, by linking to \n> libgit.a...\n> \n\nYou wrote:\n--%<--%<--%<--\nBecause, depending on what you do, the revision machinery is not\nreentrable. For example, if you filter by filename, the history is\nrewritten in-memory to simulate a history where just that filename was\ntracked, and nothing else. These changes are not cleaned up after \ncalling the internal revision machinery.\n--%<--%<--%<--\n\nWhen I wrote the above suggestion, I hadn't read the posts following the \nemail where I cut this text from (where Linus said \"we can add a 'reset' \nthingie to the revision walking machinery\" and Marco replied with some \nmore questions).\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\n"},{"id":"298215","messageId":"Pine.LNX.4.63.0612071649590.28348@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43099","inReplyTo":"4578330C.9070208@op5.se","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-12-07T16:01:54Z","receivedAt":"2006-12-07T16:01:54Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 7 Dec 2006, Andreas Ericsson wrote:\n\n> Johannes Schindelin wrote:\n> > Hi,\n> > \n> > On Thu, 7 Dec 2006, Andreas Ericsson wrote:\n> > \n> > > Shawn Pearce wrote:\n> > > > Perhaps there is some fast IPC API supported by Qt that you could use to\n> > > > run the revision listing outside of the main UI process, to eliminate\n> > > > the bottlenecks you are seeing and remove the problems noted above?  One\n> > > > that doesn't involve reading from a pipe I mean...\n> > > > \n> > > Why not just fork() + exec() and read from the filedescriptor? You can up\n> > > the output buffer of the forked program to something suitable, which means\n> > > the OS will cache it for you until you copy it to a buffer in qgit (i.e.,\n> > > read from the descriptor).\n> > \n> > Could somebody remind me why different processes are needed? I thought that\n> > the revision machinery should be used directly, by linking to libgit.a...\n> > \n> \n> You wrote:\n> --%<--%<--%<--\n> Because, depending on what you do, the revision machinery is not\n> reentrable. For example, if you filter by filename, the history is\n> rewritten in-memory to simulate a history where just that filename was\n> tracked, and nothing else. These changes are not cleaned up after calling the\n> internal revision machinery.\n> --%<--%<--%<--\n> \n> When I wrote the above suggestion, I hadn't read the posts following the \n> email where I cut this text from (where Linus said \"we can add a 'reset' \n> thingie to the revision walking machinery\" and Marco replied with some \n> more questions).\n\nYes. The reset thingie is already in place: clear_commit_marks(). It would \nhave to be enhanced a little, though:\n\n1) the function rewrite_parents(), should add another flag, HALFORPHANED, \n   and\n2) clear_commit_marks() should unset the \"parsed\" flag of the commits for \n   which HALFORPHANED is reset.\n\n-- snip --\ndiff --git a/commit.c b/commit.c\nindex d5103cd..fd225c8 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -431,6 +431,10 @@ void clear_commit_marks(struct commit *commit, unsigned int mark)\n {\n \tstruct commit_list *parents;\n \n+\t/* were parents rewritten? */\n+\tif ((mark & commit->object.flags) & HALFORPHANED)\n+\t\tcommit->object.parsed = 0;\n+\n \tcommit->object.flags &= ~mark;\n \tparents = commit->parents;\n \twhile (parents) {\ndiff --git a/revision.c b/revision.c\nindex 993bb66..461ee06 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -1097,6 +1097,7 @@ static void rewrite_parents(struct rev_info *revs, struct commit *commit)\n \t\tstruct commit_list *parent = *pp;\n \t\tif (rewrite_one(revs, &parent->item) < 0) {\n \t\t\t*pp = parent->next;\n+\t\t\tcommit->object.flags |= HALFORPHANED;\n \t\t\tcontinue;\n \t\t}\n \t\tpp = &parent->next;\ndiff --git a/revision.h b/revision.h\nindex 3adab95..544238c 100644\n--- a/revision.h\n+++ b/revision.h\n@@ -9,6 +9,7 @@\n #define BOUNDARY\t(1u<<5)\n #define BOUNDARY_SHOW\t(1u<<6)\n #define ADDED\t\t(1u<<7)\t/* Parents already parsed and added? */\n+#define HALFORPHANED\t(1u<<8) /* parents were rewritten */\n \n struct rev_info;\n struct log_info;\n-- snap --\n\nNote that this is just the idea. This particular implementation opens a \ngaping memory leak, since the buffer of the commit is not free()d, and a \nreparse would probably not pick up on the fact that the parent commits are \nalready in memory.\n\nCiao,\nDscho\n"},{"id":"296864","messageId":"e5bfff550612081034q5e4c0c93s3512fce2f11b1fab@mail.gmail.com","threadId":"43099","inReplyTo":"45781639.1050208@op5.se","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-12-08T18:34:44Z","receivedAt":"2006-12-08T18:34:44Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 12/7/06, Andreas Ericsson <ae@op5.se> wrote:\n> Shawn Pearce wrote:\n> >\n> > Perhaps there is some fast IPC API supported by Qt that you could\n> > use to run the revision listing outside of the main UI process,\n> > to eliminate the bottlenecks you are seeing and remove the problems\n> > noted above?  One that doesn't involve reading from a pipe I mean...\n> >\n>\n> Why not just fork() + exec() and read from the filedescriptor? You can\n> up the output buffer of the forked program to something suitable, which\n> means the OS will cache it for you until you copy it to a buffer in qgit\n> (i.e., read from the descriptor).\n>\n\nPlease, what do you mean with \"something suitable\"? How can I redirect\nthe output to a memory buffer or to a file that the OS will cache\n*until* I've copied it?\n\nIf I redirect to a 'normal' file, this will be flushed by OS after\nsome time, normally few seconds.\n\nCould you please post links with examples/docs about this kind of\nimplementation?\n\n\nThanks\n"},{"id":"297296","messageId":"f2b55d220612081210u6ec3e95ciec6665a6b5e6a827@mail.gmail.com","threadId":"43099","inReplyTo":"e5bfff550612081034q5e4c0c93s3512fce2f11b1fab@mail.gmail.com","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Michael K. Edwards","fromEmail":"medwards.linux@gmail.com","sentAt":"2006-12-08T20:10:00Z","receivedAt":"2006-12-08T20:10:00Z","isPatch":false,"sender":{"key":"medwards.linux@gmail.com","avatar":null},"body":"There is a very handy solution to this problem called \"tmpfs\".  It\nshould already be mounted at /tmp.  Put tmp.txt there and your problem\nwill go away.\n\nCheers,\n"},{"id":"298037","messageId":"e5bfff550612090415i26af8ea5q383889e951659d7e@mail.gmail.com","threadId":"43099","inReplyTo":"f2b55d220612081210u6ec3e95ciec6665a6b5e6a827@mail.gmail.com","subject":"Re: Fast access git-rev-list output: some OS knowledge required","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-12-09T12:15:12Z","receivedAt":"2006-12-09T12:15:12Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 12/8/06, Michael K. Edwards <medwards.linux@gmail.com> wrote:\n> There is a very handy solution to this problem called \"tmpfs\".  It\n> should already be mounted at /tmp.  Put tmp.txt there and your problem\n> will go away.\n>\n\nThanks Michael,\n\nIt seems to work! patch pushed.\n\n Marco\n\nP.S: I've looked again to Shawn idea (and code) of linking qgit\nagainst libgit.a but I found these two difficult points:\n\n- traverse_commit_list(&revs, show_commit, show_object) is blocking,\ni.e. the GUI will stop responding for few seconds while traversing the\nlist. This is easily and transparently solved by the OS scheduler if\nan external process is used for git-rev-list. To solve this in qgit I\nhave two ways: 1) call QEventLoop() once in a while from inside\nshow_commit()/ show_object() to process pending events  2) Use a\nseparate thread (QThread class). The first idea is not nice, the\nsecond opens a whole a new set of problems and it's a big amount of\nnot trivial new code to add.\n\n-  traverse_commit_list() having an internal state it's not\nre-entrant. git-rev-list it's used to load main view data but also\nfile history in another tab, and the two calls _could_ be ran\nconcurrently. With external process I simply run two instances of\nDataLoader class and consequently two external git-rev-list processes,\n"}]}