{"thread":{"id":"9502","subject":"Can I have this, pretty please?","startedAt":"2007-08-12T13:23:47Z","lastAt":"2007-08-13T05:49:11Z","messageCount":29,"participants":["David Kastrup","Steven Grimm","Linus Torvalds","Jon Smirl","Uwe Kleine-König","Jeff King","Govind Salinas","Martin Langhoff","Paul Mackerras"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"50553","messageId":"85ir7kq42k.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":null,"subject":"Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T13:23:47Z","receivedAt":"2007-08-12T13:23:47Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\nHi,\n\nI have more or less brought my system to a stillstand by trying to\nvisualize branches and histories: the graphical tools really suck\nresources.\n\nSo I have been thinking how I could use Emacs, and how to cache what\nefficiently, and put out information just on-demand and so on.\n\nAnd then it struck me: Emacs has a very efficient browser for linked\none-line information that can be expanded into complete changesets\nwith diffs inside.  It is called \"Gnus\".  A newsreader.\n\nMapping a repository into newsgroups (one per branch head?), complete\nwith threads, references, header display, article fetch (by\ngit-format-patch), Message Ids (=commit id) is much more\nstraightforward than creating an HTML server.  And it means that\neverybody can use his favorite newsreader for navigating a repository.\n\nEven when we are talking about readonly access, this would be simply\ngreat and at once make for a whole bunch of existing tools that would\nprovide much better options in many respects than existing\ngit-specific repository browsers for going through commit histories.\n\nAnd the possibilities for write access are at least intriguing.\n\nSo a lightweight nntp server serving git commits as articles would be\nreally cool.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50557","messageId":"46BF1756.5070305@midwinter.com","threadId":"9502","inReplyTo":"85ir7kq42k.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Steven Grimm","fromEmail":"koreth@midwinter.com","sentAt":"2007-08-12T14:21:10Z","receivedAt":"2007-08-12T14:21:10Z","isPatch":false,"sender":{"key":"koreth@midwinter.com","avatar":"https://gravatar.com/avatar/71b4d2e8b62f168bdc9e9205341159e3567003b4f9e2127c617c5fa0a1f5bad2?d=mp&s=160"},"body":"David Kastrup wrote:\n> Mapping a repository into newsgroups (one per branch head?), complete\n> with threads, references, header display, article fetch (by\n> git-format-patch), Message Ids (=commit id) is much more\n> straightforward than creating an HTML server.  And it means that\n> everybody can use his favorite newsreader for navigating a repository.\n>   \n\nThe news data model has one big problem. It is a tree structure (or \nrather, a set of tree structures). But git's ancestry graphs are not \ntrees; a commit can have multiple parents as well as multiple children, \nand branches can join each other multiple times (via merges) as well as \nsplit off indefinitely.\n\nI realize that you can give a list of parent message IDs in a news \nheader, but I'm going to go out on a limb and guess that all existing \nnewsreaders expect that list to be a linear series of messages going \nback toward the root of the thread (since that's all that ever occurs in \nreal netnews), rather than an arbitrary DAG.\n\nNot saying it's a worthless idea, but I bet you will not be able to get \nan accurate display of a repository's history using a news reader \nwithout modifying it to deal with more complex ancestry structures.\n\n-Steve\n"},{"id":"50560","messageId":"85y7ggogdg.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"46BF1756.5070305@midwinter.com","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T16:40:59Z","receivedAt":"2007-08-12T16:40:59Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Steven Grimm <koreth@midwinter.com> writes:\n\n> David Kastrup wrote:\n>> Mapping a repository into newsgroups (one per branch head?), complete\n>> with threads, references, header display, article fetch (by\n>> git-format-patch), Message Ids (=commit id) is much more\n>> straightforward than creating an HTML server.  And it means that\n>> everybody can use his favorite newsreader for navigating a repository.\n>\n> The news data model has one big problem. It is a tree structure (or\n> rather, a set of tree structures). But git's ancestry graphs are not\n> trees; a commit can have multiple parents as well as multiple\n> children, and branches can join each other multiple times (via merges)\n> as well as split off indefinitely.\n>\n> I realize that you can give a list of parent message IDs in a news\n> header, but I'm going to go out on a limb and guess that all\n> existing newsreaders expect that list to be a linear series of\n> messages going back toward the root of the thread (since that's all\n> that ever occurs in real netnews), rather than an arbitrary DAG.\n\nWell, I never claimed that the threading display would necessarily be\ncorrect.  But with most readers, you can turn it off.  And you can\neven tell git to turn off merges in the revision lists.  After all, it\nhas to linearize things like \"whatchanged\", too.\n\n> Not saying it's a worthless idea, but I bet you will not be able to\n> get an accurate display of a repository's history using a news\n> reader without modifying it to deal with more complex ancestry\n> structures.\n\nIt would still be lots more convenient for finding one's way around\npatch series than the current model.  And putting out every branch\ninto a group of its own (which makes merges somewhat close to\ncrosspostings in that the reader will not usually try tracking the\nchanges in the other group) would help keeping the peculiarities in a\nsingle branch display limited.\n\nAt the current point of time, _all_ tools I have available for\nbrowsing the history of a large project like Emacs suck _big_ time\ncompared to what my newsreader can handle.  They render my system\n(256MB, about 1GHz) unusable if they work at all.  Being able to, say,\ncherrypick stuff together by marking articles and piping the bunch\nthrough git-am (which is what I sometimes do with articles on the git\nlist) would be quite nice.\n\nAnd the charm over an http server is that an nntp server can basically\njust serve git data raw, without the necessity to add any dressing.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50568","messageId":"alpine.LFD.0.999.0708121135050.30176@woody.linux-foundation.org","threadId":"9502","inReplyTo":"85ir7kq42k.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-12T18:38:53Z","receivedAt":"2007-08-12T18:38:53Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Aug 2007, David Kastrup wrote:\n>\n> And then it struck me: Emacs has a very efficient browser for linked\n> one-line information that can be expanded into complete changesets\n> with diffs inside.  It is called \"Gnus\".  A newsreader.\n\nA newsreader is mis-designed for all the same reasons SVN is misdesigned: \nit sees the messages (commits) as a _tree_.\n\nAnybody who sees development as a tree is totally bogus by definition. It \nsees things forking off, but it doesn't see them merging. That's a \nfundamnetal and unfixable design bug.\n\nOf course, for news, that's ok (it might be *nice* if you could reply to \ntwo messages and see it as a merge, but that's not how things work), so \nit wasn't a design mistake for _that_.\n\nBut to visualize a history, it's useless. Merges are as important as forks \n(arguably *more* important). \"Forgetting\" about merges is bad.\n\n\t\t\tLinus\n"},{"id":"50571","messageId":"alpine.LFD.0.999.0708121140190.30176@woody.linux-foundation.org","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121135050.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-12T18:48:57Z","receivedAt":"2007-08-12T18:48:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Aug 2007, Linus Torvalds wrote:\n> \n> A newsreader is mis-designed for all the same reasons SVN is misdesigned: \n> it sees the messages (commits) as a _tree_.\n\nSide note: the lack of this bug is what makes showing large histories \ngraphically be expensive in the first place. \n\nIn fact, in git, merges are \"first-class\" entities, and forking is \nsomething you have to infer from the history (by finding two commits with \nthe same parent), and that's why calculating the graph is actually pretty \nexpensive: when you do so, you have to keep all the commit relationships \nin memory, and you basically have to sort it topologically.\n\nSo even if you don't want to show the graph itself (and just add \nreferences to allow the user to walk to parents/children manually), you'd \nstill have to calculate - and keep track of - the commit relationships. \nAnd I suspect that's what makes gitk and other visualizers take time.\n\nI think one solution is to limit the size fo the visualization by date or \nnumber, ie if you want to see history, it's often useful to do things like\n\n\tgitk --since=10.weeks.ago\n\nto see just the \"recent\" commits. That very fundamentally makes the \nproblem much cheaper, because you simply have to generate the graph for a \nmuch smaller set of commits.\n\nI used to think that we should just default to some reasonable value, but \nthen we optimized the hell out of git-rev-list and Paul fixed a number of \nscalability issues in gitk too, so it kind of fell by the wayside because \nit wasn't as important any more. But if you have a huge project with lots \nof history, the right answer may well be to make gitk *default* to using \nsomething like \"show only the last year unless some revision limiting has \nbeen done explicitly\".\n\nIOW, showing the whole history for a big project is simply pretty \nexpensive. If you have a hundred thousand commits, just keeping track of \nthe tree structure *is* going to take megabytes and megabytes of data. \nLimiting the size of the problem is usually a really good solution, \nespecially since most people tend to care about what happened in the last \nfew days, not what happened five months ago.\n\n\t\t\tLinus\n"},{"id":"50574","messageId":"85abswo9gf.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121135050.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T19:10:24Z","receivedAt":"2007-08-12T19:10:24Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 12 Aug 2007, David Kastrup wrote:\n>>\n>> And then it struck me: Emacs has a very efficient browser for linked\n>> one-line information that can be expanded into complete changesets\n>> with diffs inside.  It is called \"Gnus\".  A newsreader.\n>\n> A newsreader is mis-designed for all the same reasons SVN is\n> misdesigned: it sees the messages (commits) as a _tree_.\n\nIn the first place, it sees linked messages.  They usually correspond\nto something treeish, but a newsreader that would barf when they don't\nwould be unusable.  Newsreaders actually have to deal with stupid\nthings like _loops_ in message referals without going into a tizzy.\nThose things happen in Usenet.\n\n> Anybody who sees development as a tree is totally bogus by\n> definition. It sees things forking off, but it doesn't see them\n> merging. That's a fundamnetal and unfixable design bug.\n\nIt is not inherent in NNTP.  It depends on the particular newsreader,\nand for pretty much all of them, you can turn off threaded display if\nit disturbs you.\n\n> But to visualize a history, it's useless.\n\nNot half as useless as existing git-specific tools.  They thrash my\ncomputer to death on serious sized trees.  Putting every branch into a\nnewsgroup of its own, in contrast, together with the usual header\nsearch and refinement options, would be _much_ _much_ faster for\naccessing a particular patch.\n\nI'll probably be able to create a Gnus _backend_ for this sort of\nsetup (there are even backends for directory browsing: most files\nbecome articles written by their owner that either are plain text, or\nthat contain their file contents as an attachment -- quite more crazy\nthan a git commit tree).  But an nntp server would make the idea\nusable for more than just Emacs users, and it would allow a much more\nconvenient \"what happened on the \"next\" branch in the last few days\"\noverview than existing tools.\n\nIt lends itself not well to actually serving trees and blobs (even\nthough one could superficially rely on a rigid tree topology there):\nnewsreaders just don't match the natural way of accessing them (Gnus\noffers that for files, but it plainly is not much use compared to a\ndedicated directory browser).\n\nBut for commits and patches, one group per branch?  That would be\nfine.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50576","messageId":"alpine.LFD.0.999.0708121219540.30176@woody.linux-foundation.org","threadId":"9502","inReplyTo":"85abswo9gf.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-12T19:24:48Z","receivedAt":"2007-08-12T19:24:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Aug 2007, David Kastrup wrote:\n>\n> > But to visualize a history, it's useless.\n> \n> Not half as useless as existing git-specific tools.  They thrash my\n> computer to death on serious sized trees.\n\nSo, use \"git log --pretty=oneline\" instead, which doesn't have the \nexpense.\n\nI don't see why you think that using nntp would help anything. The \n_problem_ is still the same one, of calculating full reachability. It \ndidn't go away just because you changed to another intermediate protocol.\n\nYes, you could perhaps use the nntp caching, but I don't know if you've \nnoticed: the reason news servers tend to expire old messages is that a \nnews reader and the NNTP protocol won't be able to handle huge histories \neither.\n\nAnd if you just want the \"expire\" feature, then you might as well just \nmake git date-limit things for you, ie \"gitk --since=last.week\"\n\n\t\t\tLinus\n"},{"id":"50577","messageId":"9e4733910708121228v2fa8d356ld93efa7d1d5effd6@mail.gmail.com","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121140190.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-12T19:28:18Z","receivedAt":"2007-08-12T19:28:18Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/12/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> IOW, showing the whole history for a big project is simply pretty\n> expensive. If you have a hundred thousand commits, just keeping track of\n> the tree structure *is* going to take megabytes and megabytes of data.\n> Limiting the size of the problem is usually a really good solution,\n> especially since most people tend to care about what happened in the last\n> few days, not what happened five months ago.\n\nCould the topological graph for a packfile be computed at pack time\nand stored in the packfile so that gitk doesn't have to keep\nrecomputing it? Does it work to merge multiple precomputed graphs\nretrieved from the pack files?\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"50578","messageId":"854pj4o8k5.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121140190.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T19:29:46Z","receivedAt":"2007-08-12T19:29:46Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 12 Aug 2007, Linus Torvalds wrote:\n>> \n>> A newsreader is mis-designed for all the same reasons SVN is misdesigned: \n>> it sees the messages (commits) as a _tree_.\n>\n> Side note: the lack of this bug is what makes showing large\n> histories graphically be expensive in the first place.\n\nNot really.\n\ndak@lola:/home/tmp/emacs$ time git-rev-list --parents --topo-order --all>/dev/null\n\nreal    0m9.042s\nuser    0m8.801s\nsys     0m0.168s\n\nThis does not even start to _think_ of swapping.\n\n> So even if you don't want to show the graph itself (and just add\n> references to allow the user to walk to parents/children manually),\n> you'd still have to calculate - and keep track of - the commit\n> relationships.  And I suspect that's what makes gitk and other\n> visualizers take time.\n\nIt does not bother git-rev-list.  What takes them time is that they\nare simply not written with insane amounts of data in mind.\n\nAnd newsreaders are.\n\n> IOW, showing the whole history for a big project is simply pretty\n> expensive. If you have a hundred thousand commits, just keeping\n> track of the tree structure *is* going to take megabytes and\n> megabytes of data.  Limiting the size of the problem is usually a\n> really good solution, especially since most people tend to care\n> about what happened in the last few days, not what happened five\n> months ago.\n\nAnd newsreaders, for that reason, have a set of strategies for\nlimiting the size of the problem (and changing the limits on the fly\nas needed) as well as being efficient with handling it.  They have to\nbe _good_ at dealing with that amount of data, or they would have\nfallen by the wayside.\n\nAs opposed to gitk and other visualization tools, newsreaders usually\nhave fast and convenient keyboard navigation and an article window\nwhere serious amounts of text can be viewed with readable fonts.\n\nIf you try selecting a more readable font for gitk, you are limited to\nselecting between fonts called something starting with the letters \"a\"\nto \"c\" since the font menu runs off the screen after that.\n\nI find that I can't get much use out of gitweb: like webmail, it is\nsimply too little hands-on for getting at the right stuff efficiently:\nit is all too point and clicky instead direct keyboard access.\n\nSo at least for my preferred human-computer interface style, an\nNNTP-browsable repository would come quite handy.  I'll probably fudge\nsomething in Gnus (which has the advantage that I _can_ create more\ndirect links to files and trees), but I doubt that the usefulness of\nthe concept would not stretch to actual servers.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50580","messageId":"alpine.LFD.0.999.0708121243220.30176@woody.linux-foundation.org","threadId":"9502","inReplyTo":"9e4733910708121228v2fa8d356ld93efa7d1d5effd6@mail.gmail.com","subject":"Re: Can I have this, pretty please?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-12T19:45:26Z","receivedAt":"2007-08-12T19:45:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Aug 2007, Jon Smirl wrote:\n> \n> Could the topological graph for a packfile be computed at pack time\n> and stored in the packfile so that gitk doesn't have to keep\n> recomputing it?\n\nFor a single (full) pack, with no loose objects, sure, you could cache it. \nBut then you might as well just cache it all outside git instead.\n\n>\t\t Does it work to merge multiple precomputed graphs\n> retrieved from the pack files?\n\nNo. For multiple packs, there aren't even any \"precomputed graphs\". You \ncould probably do it with some fragment thing, and then be really clever \nputting all the fragments together, but I think it's complex as hell.\n\n\t\tLinus\n"},{"id":"50581","messageId":"85wsw0mt77.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121219540.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T19:46:52Z","receivedAt":"2007-08-12T19:46:52Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 12 Aug 2007, David Kastrup wrote:\n>>\n>> > But to visualize a history, it's useless.\n>> \n>> Not half as useless as existing git-specific tools.  They thrash my\n>> computer to death on serious sized trees.\n>\n> So, use \"git log --pretty=oneline\" instead, which doesn't have the\n> expense.\n\nYes, like managing a manual with grep is all one needs.  git log\n--pretty=oneline provides just the commit headers, but offers no way\nto jump into the commits themselves and back easily.\n\n> I don't see why you think that using nntp would help anything. The\n> _problem_ is still the same one, of calculating full\n> reachability. It didn't go away just because you changed to another\n> intermediate protocol.\n\nNewsreaders are designed _not_ to calculate full reachability.  They\nwould be unusable otherwise.  They have reasonable heuristics for\ndealing with partial information and getting more only when needed.\n\n> Yes, you could perhaps use the nntp caching, but I don't know if\n> you've noticed: the reason news servers tend to expire old messages\n> is that a news reader and the NNTP protocol won't be able to handle\n> huge histories either.\n\nIt's actually more of a storage problem.  A pretty normal general\nnewsspool with about 2 weeks of storage requires several gigabytes of\ndisk space already.\n\n> And if you just want the \"expire\" feature, then you might as well\n> just make git date-limit things for you, ie \"gitk --since=last.week\"\n\nI actually don't want any \"expire feature\".  Expiry happens at the\nserver, and git is quite efficient enough at \"storing\" the articles\nthat expiry appears pointless (unless one puts all of Sourceforge's\nrecent commit histories onto an NNTP spool, probably an interesting\nexperiment).\n\n\"Marked as read\" could conceivably come handy for keeping on top of\nlarge projects, but basically I'd already be suited fine with\nephemeral groups which look the same whenever I visit them again.\n\nThe thing with newsreaders is that it is easy to say \"since last\nweek\", and then just look at a few more earlier articles.  This sort\nof functionality has been honed and improved over decades.  If I can\navoid starting fresh, with a new user interface and the same old\nproblems, that helps.  Nobody wants tools that require to tell them\nwhen you start them just how much information you'll ever want from\nthem.\n\nThat's the thing why pagers are so convenient with real pipes as\ncompared to temporary files: you can cut off the data generating\nprocess when you decide you don't need more, and you don't need to\nwait until the whole data is there.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50582","messageId":"85sl6omt4d.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"9e4733910708121228v2fa8d356ld93efa7d1d5effd6@mail.gmail.com","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T19:48:34Z","receivedAt":"2007-08-12T19:48:34Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Jon Smirl\" <jonsmirl@gmail.com> writes:\n\n> On 8/12/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>> IOW, showing the whole history for a big project is simply pretty\n>> expensive. If you have a hundred thousand commits, just keeping track of\n>> the tree structure *is* going to take megabytes and megabytes of data.\n>> Limiting the size of the problem is usually a really good solution,\n>> especially since most people tend to care about what happened in the last\n>> few days, not what happened five months ago.\n>\n> Could the topological graph for a packfile be computed at pack time\n> and stored in the packfile so that gitk doesn't have to keep\n> recomputing it? Does it work to merge multiple precomputed graphs\n> retrieved from the pack files?\n\nThe parent information basically _is_ a bare-bones specification of\nthe topological graph.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50583","messageId":"20070812195126.GA17914@informatik.uni-freiburg.de","threadId":"9502","inReplyTo":"854pj4o8k5.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Uwe Kleine-König","fromEmail":"ukleinek@informatik.uni-freiburg.de","sentAt":"2007-08-12T19:51:26Z","receivedAt":"2007-08-12T19:51:26Z","isPatch":false,"sender":{"key":"u.kleine-koenig@pengutronix.de","avatar":"https://gravatar.com/avatar/354b5e3ceb2806a2f1e1e382ac29ddbdad18288654da62b61eb13583a857eee7?d=mp&s=160"},"body":"David Kastrup wrote:\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> > On Sun, 12 Aug 2007, Linus Torvalds wrote:\n> >> \n> >> A newsreader is mis-designed for all the same reasons SVN is misdesigned: \n> >> it sees the messages (commits) as a _tree_.\n> >\n> > Side note: the lack of this bug is what makes showing large\n> > histories graphically be expensive in the first place.\n> \n> Not really.\n> \n> dak@lola:/home/tmp/emacs$ time git-rev-list --parents --topo-order --all>/dev/null\n> \n> real    0m9.042s\n> user    0m8.801s\n> sys     0m0.168s\n> \n> This does not even start to _think_ of swapping.\nrev-list doesn't try to draw a line from each commit to its parents.\nThat's the really intensive part.  So when gitk reads\n\n\td56871cb0e6ceeca8e5435ff95409d78bed014f0 a046fe0cb8697bc97993b2e609688ff5e89e3e9\n\nit must remember this line at least until it sees a line starting with\na046fe0cb8697bc97993b2e609688ff5e89e3e9.\n\nBest regards\nUwe\n\n-- \nUwe Kleine-König\n\ndd if=/proc/self/exe bs=1 skip=1 count=3 2>/dev/null\n"},{"id":"50584","messageId":"alpine.LFD.0.999.0708121246020.30176@woody.linux-foundation.org","threadId":"9502","inReplyTo":"854pj4o8k5.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-12T19:53:57Z","receivedAt":"2007-08-12T19:53:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Aug 2007, David Kastrup wrote:\n> \n> dak@lola:/home/tmp/emacs$ time git-rev-list --parents --topo-order --all>/dev/null\n> \n> real    0m9.042s\n> user    0m8.801s\n> sys     0m0.168s\n> \n> This does not even start to _think_ of swapping.\n\nOk, good. That's the part I care about most. Nine seconds is still a long \ntime to wait for the the window to come up, so I'd still suggest at least \nthinking about limiting it, but..\n\n> It does not bother git-rev-list.  What takes them time is that they\n> are simply not written with insane amounts of data in mind.\n\nWell, gitk has certainly had performance problems in the past, they've \nbeen fixable. I think this should just be fixed too. And if the rev-list \nis fast enough, then the gitk fix may well be to just not compute the \n*whole* history - ie the solution may be as simple as stopping the \nbackground job that does all the graph calculations when it is (pick a \npoint at random) something like a thousand commits into the graph, and the \nuser hasn't scrolled down..\n\nGitk is already incremental (ie it shows the top of the graph long before \nit has drawn it all), so that should not be fundamentally hard. Paul has \nbeen pretty good about these things when we've had problems in the past.\n\nPaul added to Cc. Paul?\n\n> And newsreaders, for that reason, have a set of strategies for\n> limiting the size of the problem (and changing the limits on the fly\n> as needed) as well as being efficient with handling it.  They have to\n> be _good_ at dealing with that amount of data, or they would have\n> fallen by the wayside.\n\nThe reason I argue against this is that (a) the graph really is very \nuseful. It tells you things that you reasonably visualize any other way. \nAnd (b) I think what you suggest wouldn't be trivial at all.\n\nBut if you want to make a virtual NNTP server that exposes the \ngit-rev-list output, go right ahead.\n\nI don't think it should be needed (ie I think we should be able to handle \nthis issue other ways), and I don't think it's as good as the alternatives \n(because I don't think any client will ever be able to show the history \nwell), but hey, alternatives are fine.\n\n\t\tLinus\n"},{"id":"50585","messageId":"alpine.LFD.0.999.0708121255230.30176@woody.linux-foundation.org","threadId":"9502","inReplyTo":"85wsw0mt77.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-12T19:59:48Z","receivedAt":"2007-08-12T19:59:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Aug 2007, David Kastrup wrote:\n> >\n> > So, use \"git log --pretty=oneline\" instead, which doesn't have the\n> > expense.\n> \n> Yes, like managing a manual with grep is all one needs.  git log\n> --pretty=oneline provides just the commit headers, but offers no way\n> to jump into the commits themselves and back easily.\n\nYou misunderstand.\n\nI was suggesting you do a *tool* that bases its listing on \n--pretty=oneline, and then goes from there.\n\nIf you don't show the graph anyway, all the complex and expensive things \nthat \"git-rev-list --topo-order\" does is pretty much totally useless. \nYou're going to show the commits as a list anyway, and then when you \n*select* one commit for closer inspection, you can then try to do a better \njob at that point of doing the reachability (ie parenthood is trivial, and \nthe branch reachability is cheap if it's close to the tip of the tree, \nwhich it would almost always be).\n\nThe real problem with the topological sort is that it requires you to have \nthe full history. That not only makes everything pretty big, it also means \nthat the startup cost is bad, since you can't do things incrementally.\n\nBut if you have a client that is incremental anyway, almost all of that \ngoes away.\n\n\t\t\tLinus\n"},{"id":"50586","messageId":"20070812200258.GA13298@sigill.intra.peff.net","threadId":"9502","inReplyTo":"85abswo9gf.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-08-12T20:02:58Z","receivedAt":"2007-08-12T20:02:58Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Aug 12, 2007 at 09:10:24PM +0200, David Kastrup wrote:\n\n> I'll probably be able to create a Gnus _backend_ for this sort of\n> setup (there are even backends for directory browsing: most files\n\nYou can somewhat prototype this by just dumping the commits to an mbox\n(sorry for the long lines):\n\ngit-log \\\n  --pretty=format:'From %H Mon Sep 17 00:00:00 2001%nFrom: %an <%ae>%nDate: %ad%nSubject: %s%nMessage-ID: <%H@none>%nReferences: %P%n%n%b' \\\n  | perl -pe 's/References: (.*)/\"References: \" .  %join(\" \", map { \"<\" . $_ . \"\\@none>\" } split \\/ \\/, $1)/e' \\\n  >mbox\n\nLooking at an appreciably large chunk of history means that you will be\nvery far down in a subthread. mutt, at least, doesn't display this in a\nvery readable way. But my point is that you are probably better to look\nat a couple of different view strategies just by dumping and tweaking\nthe references relationships (which really only takes about a second for\nme on the git.git repository).\n\nAlso, have you tried looking at tig (make sure to try a recent version\nand use the 'g' command to turn on the graph display)? I think it is\nsimilar to what you are looking for, and I have found it to be very fast\n(both in implementation and in usability).\n\n-Peff\n"},{"id":"50587","messageId":"85k5s0msei.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"20070812195126.GA17914@informatik.uni-freiburg.de","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T20:04:05Z","receivedAt":"2007-08-12T20:04:05Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Uwe Kleine-König <ukleinek@informatik.uni-freiburg.de> writes:\n\n> David Kastrup wrote:\n>> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>> > On Sun, 12 Aug 2007, Linus Torvalds wrote:\n>> >> \n>> >> A newsreader is mis-designed for all the same reasons SVN is misdesigned: \n>> >> it sees the messages (commits) as a _tree_.\n>> >\n>> > Side note: the lack of this bug is what makes showing large\n>> > histories graphically be expensive in the first place.\n>> \n>> Not really.\n>> \n>> dak@lola:/home/tmp/emacs$ time git-rev-list --parents --topo-order --all>/dev/null\n>> \n>> real    0m9.042s\n>> user    0m8.801s\n>> sys     0m0.168s\n>> \n>> This does not even start to _think_ of swapping.\n> rev-list doesn't try to draw a line from each commit to its parents.\n\nWell, that's what --topo-order is somewhat about, but it might\nactually not do much together with --all.\n\n> That's the really intensive part.  So when gitk reads\n>\n> \td56871cb0e6ceeca8e5435ff95409d78bed014f0 a046fe0cb8697bc97993b2e609688ff5e89e3e9\n>\n> it must remember this line at least until it sees a line starting with\n> a046fe0cb8697bc97993b2e609688ff5e89e3e9.\n\n20 bytes of payload for a commit number.  Make a usable hashing data\nstructure for it, adds perhaps another 20 bytes.  Links to all parents\nare 4 bytes each.  All in all, we won't need more than 64 bytes per\ncommit.  Take 100000 of them, and you are at 6.4MB.  And that is not\ntaking into account that you can let git-name-rev cut the information\nretrieval down much much more, and just get the rest of the\ninformation when it is actually moved on-screen.  I don't actually\n_want_ to see 50 parallel lines from bottom to top of screen obscuring\nmy branch display and taking away all the screen estate: that is\ncompletely useless information.  Pack the branches away into a cable\npipe and let them come out isolated again only when they are actually\ninvolved on the screen.\n\nThere is no necessity to prerender/layout 50 yards of graphing.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50589","messageId":"20070812200916.GB13298@sigill.intra.peff.net","threadId":"9502","inReplyTo":"20070812200258.GA13298@sigill.intra.peff.net","subject":"Re: Can I have this, pretty please?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-08-12T20:09:16Z","receivedAt":"2007-08-12T20:09:16Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Aug 12, 2007 at 04:02:58PM -0400, Jeff King wrote:\n\n> git-log \\\n>   --pretty=format:'From %H Mon Sep 17 00:00:00 2001%nFrom: %an <%ae>%nDate: %ad%nSubject: %s%nMessage-ID: <%H@none>%nReferences: %P%n%n%b' \\\n\nEr, sorry, that should be '%aD' in the date.\n\n-Peff\n"},{"id":"50590","messageId":"85bqdcms3a.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121246020.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T20:10:49Z","receivedAt":"2007-08-12T20:10:49Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> But if you want to make a virtual NNTP server that exposes the\n> git-rev-list output, go right ahead.\n>\n> I don't think it should be needed (ie I think we should be able to\n> handle this issue other ways),\n\nSure.  But being able to handle it in a way with which I as well as my\ntools are already fluent is an advantage, for me.  \"One separate\nidiosyncratic tool for every job\" does not cut it for me with regard\nto user interfaces.  Which is part of the reason I am an Emacs user.\nAnd even though you may want to see that breed interned, some of them\ndo useful things at times.\n\n> and I don't think it's as good as the alternatives (because I don't\n> think any client will ever be able to show the history well), but\n> hey, alternatives are fine.\n\nYup.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50591","messageId":"alpine.LFD.0.999.0708121315310.30176@woody.linux-foundation.org","threadId":"9502","inReplyTo":"85k5s0msei.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-12T20:21:38Z","receivedAt":"2007-08-12T20:21:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Aug 2007, David Kastrup wrote:\n> >\n> > rev-list doesn't try to draw a line from each commit to its parents.\n> \n> Well, that's what --topo-order is somewhat about, but it might\n> actually not do much together with --all.\n\nNo, --topo-order works with --all too. In fact, to some degree, it's \n*especially* useful with --all, since having multiple tips makes the whole \ntopological sort all the more interesting, and also usually makes the end \nresult more interesting (ie it's often much more interestign to visualize \ntwo or more branches together, just to see the *relationships* between the \nbranches, and see what is shared.\n\nAnd yes, it keeps track of every single commit, and computes the \nrelationships between them. So it does indeed \"draw the line\", except it \ncan do so in a rather dense and optimized set of data structures.\n\n(That's one reason I love coding in C: it may be more effort, but you can \ntune your data structures in ways you seldom can in higher-level \nlanguages, and git-rev-list and the object representation is some of the \nmost tuned code in git).\n\n> 20 bytes of payload for a commit number.  Make a usable hashing data\n> structure for it, adds perhaps another 20 bytes.  Links to all parents\n> are 4 bytes each.  All in all, we won't need more than 64 bytes per\n> commit.\n\nYeah, that's the rough ballpark (except for 64-bit architectures, the \nlinks are all 8 bytes, but we're pretty careful). See \"object.h\" for most \nof the details.\n\n\t\t\tLinus\n"},{"id":"50592","messageId":"857io0mr6v.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121255230.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T20:30:16Z","receivedAt":"2007-08-12T20:30:16Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 12 Aug 2007, David Kastrup wrote:\n>> >\n>> > So, use \"git log --pretty=oneline\" instead, which doesn't have the\n>> > expense.\n>> \n>> Yes, like managing a manual with grep is all one needs.  git log\n>> --pretty=oneline provides just the commit headers, but offers no way\n>> to jump into the commits themselves and back easily.\n>\n> You misunderstand.\n>\n> I was suggesting you do a *tool* that bases its listing on \n> --pretty=oneline, and then goes from there.\n\nFull agreement here.  My tool was going to pass those lines off as\narticle headers.\n\n> If you don't show the graph anyway, all the complex and expensive\n> things that \"git-rev-list --topo-order\" does is pretty much totally\n> useless.  You're going to show the commits as a list anyway, and\n> then when you *select* one commit for closer inspection, you can\n> then try to do a better job at that point of doing the reachability\n> (ie parenthood is trivial, and the branch reachability is cheap if\n> it's close to the tip of the tree, which it would almost always be).\n\nQuite so.  I was thinking of doing such a tool inside of Emacs (after\nall, Emacs is the most extensive junkyard for prototyping editing\nsolutions that can be had) and was weighing options for what kind of\nstuff I would need to be doing to have it work efficiently, offering\naccess to everything without wasting unnecessary time on those things\nthat don't interest me at the moment.\n\nAnd what I came up with had far too many similarities to what a\nnewsreader does...  So my first reaction to that idea was to post to\nthe Gnus Usenet group proposing some sort of virtual server method for\nGnus (it already has more than a dozen for managing news, mail,\nvarious mail and news spools, diaries, Google, Slashdot, files,\ndirectories, virtual groups...).  But it occured to me that the\nmapping of information to NNTP is actually so straightforward that the\nidea seemed exploitable not just by Emacs users.\n\n> But if you have a client that is incremental anyway, almost all of\n> that goes away.\n\nYup.  Omniscience is overrated in computer science.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50595","messageId":"69b0c0350708121358w13d04047s1916d3599c2e040a@mail.gmail.com","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121255230.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"Govind Salinas","fromEmail":"govindsalinas@gmail.com","sentAt":"2007-08-12T20:58:05Z","receivedAt":"2007-08-12T20:58:05Z","isPatch":false,"sender":{"key":"govindsalinas@gmail.com","avatar":null},"body":"Since you all are talking about such things, I thought I would show\nyou a shot of my git UI.  It does what I think Linus is talking about.\n I have a window of x commits which I show in a list and allow the\nuser to look at each one.  You can click on a commit to see full\ndetails.  There are back/next buttons to browse the entire history and\ndate/author/etc filters to narrow your results.  The only thing I am\nmissing is the pretty chart that gitk and others have.  The chart  (in\nmy app) would only show the chart for the current window of commits.\nI'll get to that sometime after work gives me enough time to start\nworking on this again.\n\nIs this something like what you had in mind?\n\nOn 8/12/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Sun, 12 Aug 2007, David Kastrup wrote:\n> > >\n> > > So, use \"git log --pretty=oneline\" instead, which doesn't have the\n> > > expense.\n> >\n> > Yes, like managing a manual with grep is all one needs.  git log\n> > --pretty=oneline provides just the commit headers, but offers no way\n> > to jump into the commits themselves and back easily.\n>\n> You misunderstand.\n>\n> I was suggesting you do a *tool* that bases its listing on\n> --pretty=oneline, and then goes from there.\n>\n> If you don't show the graph anyway, all the complex and expensive things\n> that \"git-rev-list --topo-order\" does is pretty much totally useless.\n> You're going to show the commits as a list anyway, and then when you\n> *select* one commit for closer inspection, you can then try to do a better\n> job at that point of doing the reachability (ie parenthood is trivial, and\n> the branch reachability is cheap if it's close to the tip of the tree,\n> which it would almost always be).\n>\n> The real problem with the topological sort is that it requires you to have\n> the full history. That not only makes everything pretty big, it also means\n> that the startup cost is bad, since you can't do things incrementally.\n>\n> But if you have a client that is incremental anyway, almost all of that\n> goes away.\n>\n>                         Linus\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n"},{"id":"50598","messageId":"85y7gg5tc3.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"69b0c0350708121358w13d04047s1916d3599c2e040a@mail.gmail.com","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T21:35:56Z","receivedAt":"2007-08-12T21:35:56Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Govind Salinas\" <govindsalinas@gmail.com> writes:\n\n> Since you all are talking about such things, I thought I would show\n> you a shot of my git UI.  It does what I think Linus is talking\n> about.  I have a window of x commits which I show in a list and\n> allow the user to look at each one.  You can click on a commit to\n> see full details.  There are back/next buttons to browse the entire\n> history and date/author/etc filters to narrow your results.  The\n> only thing I am missing is the pretty chart that gitk and others\n> have.  The chart (in my app) would only show the chart for the\n> current window of commits.  I'll get to that sometime after work\n> gives me enough time to start working on this again.\n>\n> Is this something like what you had in mind?\n\nWell, what I have in mind boils down to something I can use without\nleaving my editor...  Your tool does not look all too different from\ngitk, git-gui, giggle, giwhatever.  There is a variety of those\naround, and they all don't really blow me away.  Part of the problem\nis that my work flow involves editing a lot and I naturally use Emacs.\nIf those tools used Emacs for all their editing, I'd probably become\nmore friendly with them (for what it's worth: one can talk with Emacs\nthrough sockets if necessary).  However, Linus might have something\ndifferent in mind.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50600","messageId":"85tzr45smb.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"20070812200258.GA13298@sigill.intra.peff.net","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T21:51:24Z","receivedAt":"2007-08-12T21:51:24Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Sun, Aug 12, 2007 at 09:10:24PM +0200, David Kastrup wrote:\n>\n>> I'll probably be able to create a Gnus _backend_ for this sort of\n>> setup (there are even backends for directory browsing: most files\n>\n> You can somewhat prototype this by just dumping the commits to an mbox\n> (sorry for the long lines):\n>\n> git-log \\\n>   --pretty=format:'From %H Mon Sep 17 00:00:00 2001%nFrom: %an <%ae>%nDate: %ad%nSubject: %s%nMessage-ID: <%H@none>%nReferences: %P%n%n%b' \\\n>   | perl -pe 's/References: (.*)/\"References: \" .  %join(\" \", map { \"<\" . $_ . \"\\@none>\" } split \\/ \\/, $1)/e' \\\n>   >mbox\n\nOne percent too many before join, and the order of the articles is\nreversed (--reverse helps here).\n\nIt is also a good idea to set gnus-thread-indent to 0 or 1, and\ngnus-use-trees seems interesting, though not in a reasonably good\nstate (the graph layout tries to avoid crossing links and node names,\nand that's rather useless).\n\nSo actually Gnus would need some kicking into shape before it actually\nwould present a useful tool.  On the positive side, it takes about 15\nseconds sucking up and toposorting the complete group of about 11000\ncommits from an mbox file (which one would not ever do anyway).  And\nthat is Elisp.  However, the git history is still rather harmless\nconsidering the commit amounts.\n\n> Also, have you tried looking at tig (make sure to try a recent\n> version and use the 'g' command to turn on the graph display)?\n\nThere are too many tools around.  Sigh.  Another to try.  Thanks for\nthe tip.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50601","messageId":"46a038f90708121517s3ce137e6x898e3f7a59d55a2f@mail.gmail.com","threadId":"9502","inReplyTo":"85y7gg5tc3.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-08-12T22:17:20Z","receivedAt":"2007-08-12T22:17:20Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 8/13/07, David Kastrup <dak@gnu.org> wrote:\n> Well, what I have in mind boils down to something I can use without\n> leaving my editor... (...) and I naturally use Emacs.\n\nheh! As an emacs user, I have to say this might just be a tad too much :-)\n\nThe main fix for your immediate woes of having gitk work fast is -\nimho - to limit it by time, which I do all the time.\n\nAnd on that track I'd *love* it if gitk could work as follows:\nstart-up as if I had said --since=10.days.ago (unless I pass an\nexplicit --since) and put a \"get more history\" button at the bottom of\nthe commit list. And make the default --since settable via git config\nas gitk.since or somesuch.\n\nThat'd make newcomers to git go -- WOW -- on gitk, and save old hands\nsome typing ;-)\n\nOn the gnus backend - I don't think the nntp backend is good enough,\nas it can't deal with merges. But if you can write up a new backend\nthat can read merges, you'll be golden. You'll definitely want to\nlimit the number of commits you read initially, too.\n\nNow - both your emacs-gnus-git backend and gitk/qgit would benefit\nfrom having a long-lived git process that you can talk to via a socket\nfor the stuff that you are bound to be asking a lot of (cat-file,\ndiff, etc). Something like git-fastimport but for common queries.\n\nI *thought* there was one -- I was just reading gitk to check and not\nlook like a doofus -- but at least my gitk is exec'ing git cat-file\nall over the place. I am sure that it'd speed up gitk and friends\nenormoustly, specially on non-linux environments where IO isn't as\noptimised.\n\ncheers,\n\n\n\nm\n"},{"id":"50604","messageId":"85ir7k5pp2.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"46a038f90708121517s3ce137e6x898e3f7a59d55a2f@mail.gmail.com","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-12T22:54:33Z","receivedAt":"2007-08-12T22:54:33Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> On 8/13/07, David Kastrup <dak@gnu.org> wrote:\n>> Well, what I have in mind boils down to something I can use without\n>> leaving my editor... (...) and I naturally use Emacs.\n>\n> heh! As an emacs user, I have to say this might just be a tad too\n> much :-)\n>\n> The main fix for your immediate woes of having gitk work fast is -\n> imho - to limit it by time, which I do all the time.\n>\n> And on that track I'd *love* it if gitk could work as follows:\n> start-up as if I had said --since=10.days.ago (unless I pass an\n> explicit --since) and put a \"get more history\" button at the bottom\n> of the commit list. And make the default --since settable via git\n> config as gitk.since or somesuch.\n>\n> That'd make newcomers to git go -- WOW -- on gitk, and save old\n> hands some typing ;-)\n\nSigh.  Why does one have to limit _anything_?  gitk can just keep\nasking git-rev-list -20 --stdin enough questions to fill the screen.\nIt can get more history if it _needs_ it.\n\ntig actually sucks up the whole of Emacs history (100000 commits per\nbranch) as fast as git-rev-list can produce it.  Without locking or\nswapping.\n\n> On the gnus backend - I don't think the nntp backend is good enough,\n> as it can't deal with merges. But if you can write up a new backend\n> that can read merges, you'll be golden. You'll definitely want to\n> limit the number of commits you read initially, too.\n>\n> Now - both your emacs-gnus-git backend and gitk/qgit would benefit\n> from having a long-lived git process that you can talk to via a\n> socket for the stuff that you are bound to be asking a lot of\n> (cat-file, diff, etc). Something like git-fastimport but for common\n> queries.\n\nCan be pipes.  Pretty common way of talking to utilities from within\nEmacs.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"50605","messageId":"20070812231028.GA17620@sigill.intra.peff.net","threadId":"9502","inReplyTo":"85tzr45smb.fsf@lola.goethe.zz","subject":"Re: Can I have this, pretty please?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-08-12T23:10:28Z","receivedAt":"2007-08-12T23:10:28Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Aug 12, 2007 at 11:51:24PM +0200, David Kastrup wrote:\n\n> One percent too many before join, and the order of the articles is\n> reversed (--reverse helps here).\n\nSorry, yes, a cut and paste error on the first. For the second, the\norder is largely irrelevant if your reader is going to sort them anyway\n(and since they are generally all in a single thread, the threading will\ndefine the order).\n\n> So actually Gnus would need some kicking into shape before it actually\n> would present a useful tool.  On the positive side, it takes about 15\n> seconds sucking up and toposorting the complete group of about 11000\n> commits from an mbox file (which one would not ever do anyway).  And\n\nMutt is much faster (about 3 seconds to read and sort). Of course,\nthat's in C and the display doesn't look all that useful. :)\n\n-Peff\n"},{"id":"50607","messageId":"18111.42072.605823.932110@cargo.ozlabs.ibm.com","threadId":"9502","inReplyTo":"alpine.LFD.0.999.0708121246020.30176@woody.linux-foundation.org","subject":"Re: Can I have this, pretty please?","fromName":"Paul Mackerras","fromEmail":"paulus@samba.org","sentAt":"2007-08-13T00:22:48Z","receivedAt":"2007-08-13T00:22:48Z","isPatch":false,"sender":{"key":"paulus@samba.org","avatar":"https://avatars.githubusercontent.com/u/1606439?v=4"},"body":"Linus Torvalds writes:\n\n> Well, gitk has certainly had performance problems in the past, they've \n> been fixable. I think this should just be fixed too. And if the rev-list \n> is fast enough, then the gitk fix may well be to just not compute the \n> *whole* history - ie the solution may be as simple as stopping the \n> background job that does all the graph calculations when it is (pick a \n> point at random) something like a thousand commits into the graph, and the \n> user hasn't scrolled down..\n\nI have made a \"dev\" branch in the gitk.git repository that has some\ntweaks to the graph layout algorithm which change the appearance a\nbit; specifically it doesn't continue the graph lines downwards until\nit has to terminate them with an arrow because the graph is getting\ntoo wide.  Instead, it always terminates them if they are going to be\nlonger than a certain length (about 100 rows).  Also I made some\nchanges to reduce the incidence of two lines having a corner at the\nsame point, for visual clarity.\n\nThe point of terminating the graph lines early is that it means gitk\nwon't have to lay out the whole graph, just the visible bits and a\nlimited number of rows around that.  So I'm interested to know if\npeople think it looks OK visually.  (I think it's actually better,\nmyself.)\n\nThe other thing that takes time is reading in the topology for the\nprevious/next tag computations.  I did a patch that wrote out the\ntopology to a cache file but I ran into some problems where the cache\nincludes commits that have gone away since the cache was created.\nWhat I need to do to update the cached information is basically the\nequivalent of\n\n\tgit rev-list --all ^root1 ^root2 ...\n\nwhere root1, root2, etc. are the commits in the cache that had no\nchildren (and of which all the other commits in the cache are\ndescendents).  However, git rev-list will barf if those commits no\nlonger exist.  Currently the only solution I can see is to validate\nthem one by one with separate invocations of git rev-list or something\n(git rev-parse won't do).\n\nWould it be possible to make git rev-list ignore commits that don't\nexist if they have a \"^\" in front of them, i.e. where we're asking for\nthem to be excluded anyway?  If we can do that (or something\nequivalent) then I can make the cache work reliably.  It does speed up\ngitk enormously, and the cache file is only about 3MB for the kernel\ntree, so it seems well worth while.\n\nPaul.\n"},{"id":"50615","messageId":"85wsw03rxk.fsf@lola.goethe.zz","threadId":"9502","inReplyTo":"18111.42072.605823.932110@cargo.ozlabs.ibm.com","subject":"Re: Can I have this, pretty please?","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-08-13T05:49:11Z","receivedAt":"2007-08-13T05:49:11Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Paul Mackerras <paulus@samba.org> writes:\n\n> Linus Torvalds writes:\n>\n>> Well, gitk has certainly had performance problems in the past, they've \n>> been fixable. I think this should just be fixed too. And if the rev-list \n>> is fast enough, then the gitk fix may well be to just not compute the \n>> *whole* history - ie the solution may be as simple as stopping the \n>> background job that does all the graph calculations when it is (pick a \n>> point at random) something like a thousand commits into the graph, and the \n>> user hasn't scrolled down..\n>\n> I have made a \"dev\" branch in the gitk.git repository that has some\n> tweaks to the graph layout algorithm which change the appearance a\n> bit; specifically it doesn't continue the graph lines downwards until\n> it has to terminate them with an arrow because the graph is getting\n> too wide.  Instead, it always terminates them if they are going to be\n> longer than a certain length (about 100 rows).\n\nHow about terminating them when they are going off-screen?  If you\nworry about reformatting when scrolling, you can terminate them if\nthere will be no change for at least one screen more.\n\nMore importantly: you can do your layout without having to look at\nmore than two screen's worth of commit data.\n\n> Also I made some changes to reduce the incidence of two lines having\n> a corner at the same point, for visual clarity.\n>\n> The point of terminating the graph lines early is that it means gitk\n> won't have to lay out the whole graph, just the visible bits and a\n> limited number of rows around that.\n\nOk, that was what you were already thinking.\n\n> So I'm interested to know if people think it looks OK visually.  (I\n> think it's actually better, myself.)\n\nI'd think so, too, but will be able to check only later this days.\n\n> The other thing that takes time is reading in the topology for the\n> previous/next tag computations.\n\nIf you can move that out of the busy loop and do it in the\nbackground...\n\n> I did a patch that wrote out the topology to a cache file but I ran\n> into some problems where the cache includes commits that have gone\n> away since the cache was created.\n\nI think it should be possible to come up with a data structure that\nswallows less memory than the current one.  All the info you need are\nthe SHA1s and their relations: the rest can be asked from git while\none is scrolling, with a LRU buffer of a few hundred commits for\nspeed.\n\n> Would it be possible to make git rev-list ignore commits that don't\n> exist if they have a \"^\" in front of them, i.e. where we're asking\n> for them to be excluded anyway?  If we can do that (or something\n> equivalent) then I can make the cache work reliably.  It does speed\n> up gitk enormously, and the cache file is only about 3MB for the\n> kernel tree, so it seems well worth while.\n\nCough, cough.  If the cache file is only about 3MB, why wouldn't you\nbe able to keep it in memory?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"}]}