git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: suggestions for gitweb

From
Jakub Narebski <jnareb@gmail.com>
Date
May 13, 2007, 11:50 UTC
Message-ID
<200705131350.04916.jnareb@gmail.com>
In-Reply-To
<7vabw9v906.fsf@assigned-by-dhcp.cox.net>
Junio C Hamano wrote:
Show 17 quoted lines
> Jakub Narebski <jnareb@gmail.com> writes:
> 
>> Lines of code and file sizes: file size needs additional invocation
>> per each file for gitweb; it would be easier for cgit. Costly!
>> Counting LOC is even more costly: take note that 1.) gitweb operates
>> directly on repository / object database, and does not use working
>> area, 2.) git is snapshot based and not changeset based.
> 
> We earlier discussed to make --numstat to allow us add this kind
> of information for easier script consumption.
> 
> Perhaps instead of modifying --numstat, we may be better off to
> add another format that can be more easily extended to support
> other things, like we do for the --porcelain format out of
> git-blame?  It does not have to be one line per record, like the
> way --numstat was done, which was primarily in order to make it
> a compact, human readable format.

Even if we extend --numstat or add yet another diff format meant for porcelain[*1*], and optionally add similar extension to git-ls-tree (as I think object size and LOC of file should be placed there), and the cost of additional fork and exec is not an issue, such extra information be still costly in terms of performance: CPU and I/O.

Currently for difftree (whatchanged-like) we need only to compare trees. For lines added / lines removed statistics we need to _generate_ diff.

For file size (object size) we need at least find the object in question and read it's header; for lines of code we need to get blob contents (find object, uncompress, optionally undeltify) and count the lines.

Its not insurmountable: we can use %feature for that, like in the case of other CPU-intensive features like 'blame' or 'pickaxe', or high-bandwidth features like 'snapshot'.

Footnotes: ---------- [*1*] What we should name it? --numstat-extended, --machinestat, --porcelain, --allstat, <insert your own idea here>?

-- 
Jakub Narebski
Poland
Previous: Junio C HamanoNext: Lars Hjemli
Message 6 of 21 in “suggestions for gitweb”
  1. Michael NiedermayerMay 12, 2007
  2. Junio C HamanoMay 12, 2007
  3. Aaron GrayMay 12, 2007
  4. Jakub NarebskiMay 13, 2007
  5. Junio C HamanoMay 13, 2007
  6. Jakub NarebskiMay 13, 2007
  7. Lars HjemliMay 13, 2007
  8. Suggestions for cgit (was: Re: suggestions for gitweb)Jakub Narebski, May 14, 2007
  9. Lars HjemliMay 14, 2007
  10. Lars HjemliMay 15, 2007
  11. Michael NiedermayerMay 13, 2007
  12. Jakub NarebskiMay 13, 2007
  13. Junio C HamanoMay 14, 2007
  14. Petr BaudisMay 14, 2007
  15. Michael NiedermayerMay 14, 2007
  16. Petr BaudisMay 14, 2007
  17. Michael NiedermayerMay 14, 2007
  18. Petr BaudisMay 14, 2007
  19. Jakub NarebskiMay 14, 2007
  20. Michael NiedermayerMay 14, 2007
  21. Jan HudecMay 15, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.