threads / patch / 29812

patch, 3 parts"Improving the `git add -p` interface" project

Subject: [git wiki PATCH 3/3] "Improving the `git add -p` interface" project

## tl;dr

8 messages between Mar 2, 2012 and Mar 3, 2012. Diffs are folded; open one to read it.

replies: 7people: 5as markdown or json

Thomas Rast· Mar 2, 2012, 11:05 UTC · lore

[git wiki PATCH 1/3] "Improving parallelism in various commands" project

---
 SoC-2012-Ideas.md |   31 +++++++++++++++++++++++++++++++
 1 file changed, 31 insertions(+)
Show changes to SoC-2012-Ideas.md +31 −0
diff --git a/SoC-2012-Ideas.md b/SoC-2012-Ideas.md
index 29a374a..145b379 100644
--- a/SoC-2012-Ideas.md
+++ b/SoC-2012-Ideas.md
@@ -68,3 +68,34 @@ work to be done:
    and only accessed on demand.
 
 Proposed mentor: Jeff King
+
+Improving parallelism in various commands
+-----------------------------------------
+
+Git is mostly written single-threaded, with a few commands having
+bolted-on extensions to support parallel operation (notably git-grep,
+git-pack-objects and the core.preloadIndex feature).
+
+We have recently looked into some of these areas and made a few
+optimizations, but a big roadblock is that pack access is entirely
+single-threaded.  The project would consist of the following steps:
+
+ * In preparation (the half-step): identify commands that could
+   benefit from parallelism.  `git grep --cached` and `git grep
+   COMMIT` come to mind, but most likely also `git diff` and `git log
+   -p`.  You can probably find more.
+
+ * Rework the pack access mechanisms to allow the maximum possible
+   parallel access.
+
+ * Rework the commands found in the first step to use parallel pack
+   access if possible.  Along the way, document the improvements with
+   performance tests.
+
+The actual programming must be done in C using pthreads for obvious
+reasons.  At the very least you should not be scared of low-level
+programming.  Prior experience and access to one or more multi-core
+computers is a plus.
+
+Proposed by: Thomas Rast
+Possible mentor(s): Thomas Rast
-- 
1.7.9.2.467.g7fee4
Thomas Rast· Mar 2, 2012, 11:05 UTC · re: Thomas Rast · lore

[git wiki PATCH 2/3] "Designing a faster index format" project

---
 SoC-2012-Ideas.md |   41 +++++++++++++++++++++++++++++++++++++++++
 1 file changed, 41 insertions(+)
Show changes to SoC-2012-Ideas.md +41 −0
diff --git a/SoC-2012-Ideas.md b/SoC-2012-Ideas.md
index 145b379..59d1baf 100644
--- a/SoC-2012-Ideas.md
+++ b/SoC-2012-Ideas.md
@@ -99,3 +99,44 @@ computers is a plus.
 
 Proposed by: Thomas Rast
 Possible mentor(s): Thomas Rast
+
+Designing a faster index format
+-------------------------------
+
+Git is pretty slow when managing huge repositories in terms of files
+in any given tree, as it needs to rewrite the index (in full) on
+pretty much every operation.  For example, even though _logically_
+`git add already_tracked_file` only changes a single blob SHA-1 in the
+index, Git will verify index correctness during loading and recompute
+the new hash during writing _over the whole index_.  It thus ends up
+spending a large amount of time simply on hashing the index.
+
+A carefully designed index format could help in several ways.  (For the
+complexity estimates below, let n be the number of index entries or
+the size of the index, which is roughly the same.)
+
+ * The work needed for something as simple as entering a new blob into
+   the index, which is possibly the most common operation in git
+   (think `git add -p` etc.) should be at most log(n).
+
+ * The work needed for a more complex operation that changes the
+   number of index entries will have to be larger unless we get into
+   database land.  However the amount of data that we SHA-1 over
+   should still be log(n).
+
+ * It may be possible to store the cache-tree data directly as part of
+   the index, always keeping it valid, and using that to validate
+   index consistency throughout.  If so, this would be a big boost to
+   other git operations that currently suffer from frequent cache-tree
+   invalidation.
+
+Note that there are other criteria than speed: the format should also
+be as easy to parse as possible, so as to simplify work for the other
+.git-reading programs (such as jgit and libgit2).  For the same
+reason, you will also have to show a significant speed boost as
+otherwise the break in compatibility is not worth the fallout.
+
+The programming work will be in C, as it replaces a core part of git.
+
+Proposed by: Thomas Rast
+Possible mentor(s): Thomas Rast
-- 
1.7.9.2.467.g7fee4
Jeff King· Mar 2, 2012, 11:08 UTC · re: Thomas Rast · lore

Re: [git wiki PATCH 2/3] "Designing a faster index format" project

On Fri, Mar 02, 2012 at 12:05:46PM +0100, Thomas Rast wrote:
> ---
>  SoC-2012-Ideas.md |   41 +++++++++++++++++++++++++++++++++++++++++
>  1 file changed, 41 insertions(+)

Thanks, I've applied and pushed all three (but the index one is my favorite).

-Peff
Junio C Hamano· Mar 2, 2012, 18:24 UTC · re: Thomas Rast · lore

Re: [git wiki PATCH 2/3] "Designing a faster index format" project

Having harder and more ambitious ones in the mix is OK, but I suspect this one is probably a bit too ambitious to be realistic for a student project that lasts only 3 months.

This proposal is about a change that touches core parts of the system that has the chance of inflicting permanent damage to end users' histories. I'd have a hard time reviewing and convincing myself that the change is good, if such a change were done by somebody new to the system, even if the work were mentored very closely by one of our top 15 committers.

Nguyen Thai Ngoc Duy· Mar 3, 2012, 03:30 UTC · re: Junio C Hamano · lore

Re: [git wiki PATCH 2/3] "Designing a faster index format" project

On Sat, Mar 3, 2012 at 1:24 AM, Junio C Hamano <gitster@pobox.com> wrote:
Show 9 quoted lines
> Having harder and more ambitious ones in the mix is OK, but I suspect this
> one is probably a bit too ambitious to be realistic for a student project
> that lasts only 3 months.
>
> This proposal is about a change that touches core parts of the system that
> has the chance of inflicting permanent damage to end users' histories.
> I'd have a hard time reviewing and convincing myself that the change is
> good, if such a change were done by somebody new to the system, even if
> the work were mentored very closely by one of our top 15 committers.

I was comparing the excitement of seeing this implemented vs packv4, then realised the complexity of work may be more or less the same. But, if the new format is not mmap'd for direct access, we can reconstruct "struct cache_entry*" exactly as it is now from new format. That makes it less sensitive to the core parts and we still hopefully benefit from new format (mostly I/O I guess). Phase 2 could come later to make core parts aware of new format.

-- 
Duy
Thomas Rast· Mar 2, 2012, 11:05 UTC · re: Thomas Rast · lore
---
 SoC-2012-Ideas.md |   41 +++++++++++++++++++++++++++++++++++++++++
 1 file changed, 41 insertions(+)
Show changes to SoC-2012-Ideas.md +41 −0
diff --git a/SoC-2012-Ideas.md b/SoC-2012-Ideas.md
index 59d1baf..b2cc475 100644
--- a/SoC-2012-Ideas.md
+++ b/SoC-2012-Ideas.md
@@ -140,3 +140,44 @@ The programming work will be in C, as it replaces a core part of git.
 
 Proposed by: Thomas Rast
 Possible mentor(s): Thomas Rast
+
+Improving the `git add -p` interface
+------------------------------------
+
+The interface behind `git {add|commit|stash|reset} {-p|-i}` is shared
+and called `git-add--interactive.perl`.    This project would mostly
+focus on the `--patch` side, as that seems to be much more widely
+used; however, improvements to `--interactive` would probably also be
+welcome.
+
+The `--patch` interface suffers from some design flaws caused largely
+by how the script grew:
+
+ * Application is not atomic: hitting Ctrl-C midway through patching
+   may still touch files.
+
+ * The terminal/line-based interface becomes a problem if diff hunks
+   are too long to fit in your terminal.
+
+ * Cannot go back and forth between files.
+
+ * Cannot reverse the direction of the patch.
+
+ * Cannot look at the diff in word-diff mode (and apply it normally).
+
+Due to the current design it is also pretty hard to add these features
+without adding to the mess.  Thus the project consists of:
+
+ * Come up with more ideas for features/improvements and discuss them
+   with users.
+
+ * Cleanly redesigning the main interface loop to allow for the above
+   features.
+
+ * Implement the new features.
+
+As the existing code is written in Perl, that is what you will use for
+this project.
+
+Proposed by: Thomas Rast
+Possible mentor(s): Thomas Rast
-- 
1.7.9.2.467.g7fee4
Nguyen Thai Ngoc Duy· Mar 2, 2012, 14:29 UTC · re: Thomas Rast · lore

Re: [git wiki PATCH 1/3] "Improving parallelism in various commands" project

On Fri, Mar 2, 2012 at 6:05 PM, Thomas Rast <trast@student.ethz.ch> wrote:
> + * In preparation (the half-step): identify commands that could
> +   benefit from parallelism.  `git grep --cached` and `git grep
> +   COMMIT` come to mind, but most likely also `git diff` and `git log
> +   -p`.  You can probably find more.

I just had a thought this afternoon whether "git add" may benefit from parallelism. It's most likely I/O-bound, although I think if we add a bunch of large files, it might become CPU-bound. To generalize, anything that calls hash_sha1_file() might benefit from parallelism.

Another candidate may be git-apply. Actually I just want to speed up git-rebase and think git-apply may be the culprit. Or it could be unpack-trees code..

-- 
Duy
James Pickens· Mar 2, 2012, 17:35 UTC · re: Thomas Rast · lore

Re: [git wiki PATCH 1/3] "Improving parallelism in various commands" project

[Resend since the first try had HTML and the list rejected it]
On Fri, Mar 2, 2012, Thomas Rast <trast@student.ethz.ch> wrote:
+ * In preparation (the half-step): identify commands that could
+   benefit from parallelism.  `git grep --cached` and `git grep
+   COMMIT` come to mind, but most likely also `git diff` and `git log
+   -p`.  You can probably find more.

For those of us who must work on NFS for various reasons, it would help tremendously to write out work tree files in parallel, during 'git clone', 'git reset --hard', and any other command that writes lots of files to the work tree. You can get a huge speedup (benchmarked at ~3.5x) without even unpacking those files in parallel; unpacking them serially and writing them to disk in parallel is sufficient.

I submitted a patch [1] ~2 years ago that added that capability. It was not accepted, but it did demonstrate the huge potential speedup on NFS, and the pitfall of degrading performance on a local drive. The patch itself is probably not useful any more, but it includes some benchmarks, and the discussion may be helpful.

James
[1] http://thread.gmane.org/gmane.comp.version-control.git/103489

← back to recent threads