git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: newbie questions about git design and features (some wrt hg)

From
MMMatt Mackall <mpm@selenic.com>
Date
Feb 1, 2007, 00:34 UTC
Message-ID
<20070201003429.GQ10108@waste.org>
In-Reply-To
<200702010058.43431.jnareb@gmail.com>
On Thu, Feb 01, 2007 at 12:58:42AM +0100, Jakub Narebski wrote:
Show 38 quoted lines
> Matt Mackall wrote:
> > On Wed, Jan 31, 2007 at 11:56:01AM +0100, Jakub Narebski wrote:
> >> Theodore Tso wrote:
> >> 
> >>> On Tue, Jan 30, 2007 at 11:55:48AM -0500, Shawn O. Pearce wrote:
> >>>> I think hg modifies files as it goes, which could cause some issues
> >>>> when a writer is aborted.  I'm sure they have thought about the
> >>>> problem and tried to make it safe, but there isn't anything safer
> >>>> than just leaving the damn thing alone.  :)
> >>> 
> >>> To be fair hg modifies files using O_APPEND only.  That isn't quite
> >>> as safe as "only creating new files", but it is relatively safe.
> >> 
> >>>From (libc.info):
> >> 
> >>  -- Macro: int O_APPEND
> [...] 
> >> I don't quote understand how that would help hg (Mercurial) to have
> >> operations like commit, pull/fetch or push atomic, i.e. all or
> >> nothing. 
> > 
> > That's because it's unrelated.
> [...]
> > Mercurial has write-side locks so there can only ever be one writer at
> > a time. There are no locks needed on the read side, so there can be
> > any number of readers, even while commits are happening.
> > 
> >> What happens if operation is interrupted (e.g. lost connection to
> >> network during fetch)?
> > 
> > We keep a simple transaction journal. As Mercurial revlogs are
> > append-only, rolling back a transaction just means truncating all
> > files in a transaction to their original length.
> 
> Thanks a lot for complete answer. So Mercurial uses write-side locks
> for dealing with concurrent operations, and transaction journal for
> dealing with interrupted operations. I guess that incomplete transactions
> are rolled back on next hg command...

They are either automatically rolled back on abort or if that fails for some reason like power failure the user is prompted to run "hg recover" to complete the rollback. We also save the last transaction journal which allows one level of undo for pulls/commits.

> I guess (please correct me if I'm wrong) that git uses "put reference
> after putting data" scheme, and write-side lock in few places when it
> is needed.
Mercurial also uses a "put reference after putting data" which is what
allows us to have no read vs write locking.
  
Show 20 quoted lines
> >> In git both situations result in some prune-able and fsck-visible crud in
> >> repository, but repository stays uncorrupted, and all operations are atomic
> >> (all or nothing).
> > 
> > If a Mercurial transaction is interrupted and not rolled back, the
> > result is prune-able and fsck-visible crud. But this doesn't happen
> > much in practice.
> > 
> > The claim that's been made is that a) truncate is unsafe because Linux
> > has historically had problems in this area and b) git is safer because
> > it doesn't do this sort of thing. 
> > 
> > My response is a) those problems are overstated and Linux has never
> > had difficulty with the sorts of straightforward single writer
> > operations Mercurial uses and b) normal git usage involves regular
> > rewrites of data with packing operations that makes its exposure to
> > filesystem bugs equivalent or greater.
> 
> Rewrites in git perhaps are (or should be) regular, but need not be often.
> And with new idea/feature of kept packs rewrite need not be of full data.

If the set of files in a given commit (say tip) gets spread out across an arbitrary number of packs ordered by last modification time, performance degrades to O(n) lookups and random seeking.

Show 7 quoted lines
> One command which _is_ (a bit) unsafe in git is git-prune. I'm not sure
> if it could be made safe. But not doing prune affects only a bit
> repository size (where git is best I think of all SCMs) and not performance.
> 
> On the other hand hg repository structure (namely log like append changelog
> / revlog to store commits) makes it I think hard to have multiple persistent
> branches.

Not sure why you think that. There are some difficulties here, but they're mostly owing to the fact that we've always emphasized the one branch per repo approach as being the most user-friendly.

> Sidenote 1: it looks like git is optimized for speed of merge and checkout
> (branch switching, or going to given point in history for bisect), and
> probably accidentally for multi-branch repos, while Mercurial is optimized
> for speed of commit and patch.
I think all of these things are comparable.
> Sidenote 2: Mercurial repository structure might make it use "file-ids"
> (perhaps implicitely), with all the disadvantages (different renames
> on different branches) of those.
Nope.
Show 7 quoted lines
> > In either case, both provide strong integrity checks with recursive
> > SHA1 hashing, zlib CRCs, and GPG signatures (as well as distributed
> > "back-up"!) so this is largely a non-issue relative to traditional
> > systems.
> 
> Integrity checks can tell you that repository is corrupted, but it would
> be better if it didn't get corrupted in first place.
Obviously. Hence our append-only design. Data that's written to a repo
is never rewritten, which minimizes exposure to software bugs and I/O
errors.
 
> Besides: zlib CRC for Mercurial? I thought that hg didn't compress the
> data, only delta chain store it?
We use zlib compression of deltas and have since April 6, 2005.
-- 
Mathematics is the supreme nostalgia of our time.
Previous: Jakub NarebskiNext: Jakub Narebski
Message 9 of 61 in “newbie questions about git design and features (some wrt hg)”
  1. Mike ColemanJan 30, 2007
  2. Johannes SchindelinJan 30, 2007
  3. Shawn O. PearceJan 30, 2007
  4. Theodore TsoJan 31, 2007
  5. Jakub NarebskiJan 31, 2007
  6. Junio C HamanoJan 31, 2007
  7. Matt MackallJan 31, 2007
  8. Jakub NarebskiJan 31, 2007
  9. Matt MackallFeb 1, 2007
  10. Jakub NarebskiFeb 1, 2007
  11. Simon 'corecode' SchubertFeb 1, 2007
  12. Johannes SchindelinFeb 1, 2007
  13. Simon 'corecode' SchubertFeb 1, 2007
  14. Johannes SchindelinFeb 1, 2007
  15. Linus TorvaldsFeb 1, 2007
  16. Eric WongFeb 1, 2007
  17. Linus TorvaldsFeb 1, 2007
  18. Jakub NarebskiFeb 2, 2007
  19. Simon 'corecode' SchubertFeb 2, 2007
  20. Jakub NarebskiFeb 2, 2007
  21. Shawn O. PearceFeb 2, 2007
  22. Mark WoodingFeb 2, 2007
  23. Jakub NarebskiFeb 2, 2007
  24. Linus TorvaldsFeb 2, 2007
  25. Jakub NarebskiFeb 2, 2007
  26. Linus TorvaldsFeb 2, 2007
  27. Brendan CullyFeb 2, 2007
  28. Jakub NarebskiFeb 2, 2007
  29. Brendan CullyFeb 2, 2007
  30. Giorgos KeramidasFeb 2, 2007
  31. Linus TorvaldsFeb 2, 2007
  32. Giorgos KeramidasFeb 3, 2007
  33. Matthias KestenholzFeb 3, 2007
  34. Linus TorvaldsFeb 3, 2007
  35. Jakub NarebskiFeb 3, 2007
  36. Linus TorvaldsFeb 2, 2007
  37. Brendan CullyFeb 2, 2007
  38. Linus TorvaldsFeb 2, 2007
  39. Brendan CullyFeb 2, 2007
  40. Jakub NarebskiFeb 2, 2007
  41. Linus TorvaldsFeb 2, 2007
  42. Matt MackallFeb 2, 2007
  43. Jakub NarebskiFeb 2, 2007
  44. Matt MackallFeb 2, 2007
  45. Jakub NarebskiFeb 2, 2007
  46. Jakub NarebskiFeb 2, 2007
  47. Brendan CullyFeb 3, 2007
  48. Jakub NarebskiFeb 3, 2007
  49. Jakub NarebskiFeb 3, 2007
  50. Jakub NarebskiJan 30, 2007
  51. Linus TorvaldsJan 30, 2007
  52. Linus TorvaldsJan 30, 2007
  53. Junio C HamanoJan 30, 2007
  54. Mike ColemanJan 31, 2007
  55. Linus TorvaldsJan 31, 2007
  56. Junio C HamanoJan 31, 2007
  57. Linus TorvaldsJan 31, 2007
  58. Johannes SchindelinJan 31, 2007
  59. Mike ColemanJan 31, 2007
  60. Nicolas PitreJan 31, 2007
  61. Mike ColemanJan 31, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.