git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: newbie questions about git design and features (some wrt hg)

From
Shawn O. Pearce <spearce@spearce.org>
Date
Jan 30, 2007, 16:55 UTC
Message-ID
<20070130165548.GF25950@spearce.org>
In-Reply-To
<3c6c07c20701300820l42cfc8dbsb80393fc1469f667@mail.gmail.com>
Mike Coleman <tutufan@gmail.com> wrote:
> 1.  As of today, is there any real safety concern with either tool's
> repo format?  Is either tool significantly better in this regard?
> (Keith Packard's post hints at a problem here, but doesn't really make
> the case.)

I think the Git format is tighter in terms of compression, and simpler in terms of understanding and writing code. I have personally written the code to read and write the Git repository format in both C and Java, and in both cases it falls out in just a few hundred lines of code (assuming you have libz handy to do the compression/decompression for you).

The Git format is completely safe with regards to parallel modification of a repository, which is good for shared repositories that might have multiple people pushing into it at once.

Git's format is also safe with regards to *any* update. You literally cannot destroy the repository during an update. Its impossible. You'd have to physically destroy the storage device. (OK, that's overstating it a bit, but it is really hard.)

The point Keith was making was the Git format is "add-only". Once something has been stored, we NEVER modify it again. This bypasses any sort of possible problems that can occur with partial modifications caused by a process aborting in the middle of a change.

I think hg modifies files as it goes, which could cause some issues
when a writer is aborted.  I'm sure they have thought about the
problem and tried to make it safe, but there isn't anything safer
than just leaving the damn thing alone.  :)
 
> 2.  Does the git packed object format solve the performance problem
> alluded to in posts from a year or two ago?
Yes.  By a huge margin.  Git's *fast*.  Ignore anything from a year
or two ago.
 
> 3.  Someone mentioned that git bisect can work between any two
> commits, not necessarily just one that happens to be an ancestor of
> the other.  This sounds really cool.  Can hg's bisect do this, too?
No clue.
 
> 4.  What is git's index good for?  I find that I like the idea of it,
> but I'm not sure I could justify it's presence to someone else, as
> opposed to having it hidden in the way that hg's dircache (?) is.  Can
> anyone think of a good scenario where it's a pretty obvious benefit?

Its a good way to stage the stuff in your next commit. By that I mean you edit some code. Then you look at what differs between the index and your working directory. You decide "this hunk is good, it passed the tests, I want to commit that, so toss it into the index". Now that hunk isn't different anymore.

When it comes time to commit, all of your already reviewed stuff is staged in the index. You just need to issue a commit and supply the message. But you can leave modified stuff in the working directory, even for files that were alerady updated in the index.

This really helps during a merge. Only the stuff which Git could not merge for you is seen as different between the index and the working directory; all of the stuff that Git merged for you is already staged in the index. So you can focus on the conflicts, and stage their resolutions into the index as you go. This makes it easier to work through larger merges where more than 1 or 2 files contains conflicts.

> 5.  I think I read that there'd been just one incompatible change over
> time in the git repo format.  What was it?

A LONG time ago, like in the very first version Linus offered out to the public, we computed the identity of an object using the SHA-1 hash of the *compressed* data. This is sensitive to the compression settings used, and was not the best idea as a result.

It was very quickly changed to compute the identity of the object using the SHA-1 has of the raw (user) data, removing any dependence on the compression routine to always yield the same result for the same input.

We haven't had a change since then.  We have added some new
compression options which are just that, options.  If you use them
older Git binaries won't necessarily recognize the repository data,
but these are off by default and can be enabled on a per-repository
basis.  E.g. if you are only using newer Git on a given system you
can enable the newer compression features on all of the repositories
on that system.
 
Show 5 quoted lines
> 6.  Does either tool use hard links?  This matters to me because I do
> development on a connected machine and a disconnected machine, using a
> usb drive to rsync between.  (Perhaps there'll be some way to transfer
> changes using git or hg instead of rsync, but I haven't figured that
> out yet.)

Git can use hardlinks if you ask it to. We only use them for the repository files, not for the user's actual source files.

Git has its own native transport (git-push, git-fetch) which can
move data between two Git repositories via local filesystem access,
SSH, HTTP, FTP, and rsync (latter two are read-only transports).
 
Show 5 quoted lines
> 7.  I'm a fan of Python, and I'm really a fan of using high-level
> languages with performance-critical parts in a lower-level language,
> so in that regard, I really like hg's implementation.  If someone
> wanted to do it, is a Python clone of git conceivable?  Is there
> something about it that just requires C?
Yes, a Python clone of Git is conceivable.  Indeed, there is a
pure Java clone in process (jgit) for an Eclipse plugin (egit).
If you wanted to rewrite Git in Python, knock yourself out.
But we've ported all of our Python to C, as its just faster.
 
Show 6 quoted lines
> 8.  It feels like hg is not really comfortable with parallel
> development over time on different heads within a single repo.
> Rather, it seems that multiple repos are supposed to be used for this.
> Does this lead to any problems?  For example, is it harder or
> different to merge two heads if they're in different repo than if
> they're in the same repo?

No clue. I know multiple heads in one Git repository works *awesome*. Especially on large repositories (>10k files) as the time required to start a new branch is only the time needed to update the files in the working directory which don't have the correct version. Usually that's a small percentage (<200) of the files and thus its very fast to switch to a new branch of development, and switch back.

On a decent UNIX system (and my Mac OS X PowerBook doesn't really count) flipping branches in git-gui is almost immediate. You pick the branch in the menu and *wham* its switched. And that's my PowerBook, which as I said, doesn't quite count as good UNIX system...

-- 
Shawn.
Previous: Johannes SchindelinNext: Theodore Tso
Message 3 of 61 in “newbie questions about git design and features (some wrt hg)”
  1. Mike ColemanJan 30, 2007
  2. Johannes SchindelinJan 30, 2007
  3. Shawn O. PearceJan 30, 2007
  4. Theodore TsoJan 31, 2007
  5. Jakub NarebskiJan 31, 2007
  6. Junio C HamanoJan 31, 2007
  7. Matt MackallJan 31, 2007
  8. Jakub NarebskiJan 31, 2007
  9. Matt MackallFeb 1, 2007
  10. Jakub NarebskiFeb 1, 2007
  11. Simon 'corecode' SchubertFeb 1, 2007
  12. Johannes SchindelinFeb 1, 2007
  13. Simon 'corecode' SchubertFeb 1, 2007
  14. Johannes SchindelinFeb 1, 2007
  15. Linus TorvaldsFeb 1, 2007
  16. Eric WongFeb 1, 2007
  17. Linus TorvaldsFeb 1, 2007
  18. Jakub NarebskiFeb 2, 2007
  19. Simon 'corecode' SchubertFeb 2, 2007
  20. Jakub NarebskiFeb 2, 2007
  21. Shawn O. PearceFeb 2, 2007
  22. Mark WoodingFeb 2, 2007
  23. Jakub NarebskiFeb 2, 2007
  24. Linus TorvaldsFeb 2, 2007
  25. Jakub NarebskiFeb 2, 2007
  26. Linus TorvaldsFeb 2, 2007
  27. Brendan CullyFeb 2, 2007
  28. Jakub NarebskiFeb 2, 2007
  29. Brendan CullyFeb 2, 2007
  30. Giorgos KeramidasFeb 2, 2007
  31. Linus TorvaldsFeb 2, 2007
  32. Giorgos KeramidasFeb 3, 2007
  33. Matthias KestenholzFeb 3, 2007
  34. Linus TorvaldsFeb 3, 2007
  35. Jakub NarebskiFeb 3, 2007
  36. Linus TorvaldsFeb 2, 2007
  37. Brendan CullyFeb 2, 2007
  38. Linus TorvaldsFeb 2, 2007
  39. Brendan CullyFeb 2, 2007
  40. Jakub NarebskiFeb 2, 2007
  41. Linus TorvaldsFeb 2, 2007
  42. Matt MackallFeb 2, 2007
  43. Jakub NarebskiFeb 2, 2007
  44. Matt MackallFeb 2, 2007
  45. Jakub NarebskiFeb 2, 2007
  46. Jakub NarebskiFeb 2, 2007
  47. Brendan CullyFeb 3, 2007
  48. Jakub NarebskiFeb 3, 2007
  49. Jakub NarebskiFeb 3, 2007
  50. Jakub NarebskiJan 30, 2007
  51. Linus TorvaldsJan 30, 2007
  52. Linus TorvaldsJan 30, 2007
  53. Junio C HamanoJan 30, 2007
  54. Mike ColemanJan 31, 2007
  55. Linus TorvaldsJan 31, 2007
  56. Junio C HamanoJan 31, 2007
  57. Linus TorvaldsJan 31, 2007
  58. Johannes SchindelinJan 31, 2007
  59. Mike ColemanJan 31, 2007
  60. Nicolas PitreJan 31, 2007
  61. Mike ColemanJan 31, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.