git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: WARNING! Object DB conversion (was Re: [PATCH] write-tree performance problems)

From
David Woodhouse <dwmw2@infradead.org>
Date
Apr 20, 2005, 22:29 UTC
Message-ID
<1114036196.5877.70.camel@localhost.localdomain>
In-Reply-To
<Pine.LNX.4.58.0504200731590.6467@ppc970.osdl.org>
On Wed, 2005-04-20 at 07:59 -0700, Linus Torvalds wrote:
Show 9 quoted lines
>         external-parent <commit-hash> <external-parent-ID>
>                 comment for this parent
> 
> and the nice thing about that is that now that information allows you to 
> add external parents at any point. 
> 
> Why do it like this? First off, I think that the "initial import" ends up
> being just one special case of the much more _generic_ issue of having
> patches come in from other source control systems 

This isn't about patches coming in from other systems -- it's about _history_, and the fact that it's imported from another system is just an implementation detail. It's git history now, and what we have here is just a special case of wanting to prune ancient git history to keep the size of our working trees down. You refer to this yourself...

> Secondly, we do need something like this for pruning off history anyway, 
> so that the tools have a better way of saying "history has been pruned 
> off" than just hitting a missing commit. 

Having a more explicit way of saying "history is pruned" than just a reference to a missing commit is a reasonable request -- but I really don't see how we can do that by changing the now-oldest commit object to contain an 'external-parent' field. Doing that would change the sha1 of the commit object in question, and then ripple through all the subsequent commits.

Come this time next year, if I decide I want to prune anything older than 2.6.40 from all the trees on my laptop, it has to happen _without_ changing the commit objects which occur after my arbitrarily-chosen cutoff point.

If we want to have an explicit record of pruning rather than just copying with a missing object, then I think we'd need to do it with an external note to say "It's OK that commit XXXXXXXXXXX is missing".

Show 6 quoted lines
> Thirdly, I don't actually want my new tree to depend on a conversion of
> the old BK tree.
> 
> Two reasons: if it's a really full conversion, there are definitely going
> to be issues with BitMover. They do not want people to try to reverse
> engineer how they do namespace merges

Don't think of it as "a conversion of the old BK tree". It's just an import of Linux's development history. This isn't going to help reverse-engineer how BK does merges; it's just our own revision history. I'm not sure exactly how Thomas is extracting it, but AIUI it's all obtainable from the SCCS files anyway without actually resorting to using BK itself.

There's nothing here for Larry to worry about. It's not as if we're actually using BK to develop git by observing BK's behaviour w.r.t merges and trying to emulate it. Besides -- if we wanted to do that, we'd need to use the _BK_ version of the tree; the git version wouldn't help us much anyway.

And given that BK's merges are based on individual files and we're not going that route with git, it's not clear how much we could lift directly from BK even if we _were_ going to try that.

> The other reason is just the really obvious one: in the last week, I've
> already changed the format _twice_ in ways that change the hash. As long
> as it's 119MB of data, it's not going to be too nasty to do again.

That's fine. But by the time we settle on a format and actually start using it in anger, it'd be good to be sure that it _is_ possible to track development from current trees all the way back -- be that with explicit reference to pruned history as you suggest, or with absent parents as I still prefer.

> it's not that it's necessarily the wrong thing to do, but I think it
> is the wrogn thing to do _now_.

OK, time for us to keep arguing over the implementation details of how we prune history then :)

-- 
dwmw2
Previous: Linus TorvaldsNext: Chris Mason
Message 26 of 54 in “write-tree performance problems”
  1. write-tree performance problemsChris Mason, Apr 19, 2005
  2. Linus TorvaldsApr 19, 2005
  3. Chris MasonApr 19, 2005
  4. Linus TorvaldsApr 19, 2005
  5. Chris MasonApr 19, 2005
  6. Linus TorvaldsApr 19, 2005
  7. Chris MasonApr 20, 2005
  8. Linus TorvaldsApr 20, 2005
  9. Linus TorvaldsApr 20, 2005
  10. H. Peter AnvinApr 20, 2005
  11. WARNING! Object DB conversion (was Re: [PATCH] write-tree performance problems)Linus Torvalds, Apr 20, 2005
  12. Ingo MolnarApr 20, 2005
  13. Jon SeymourApr 20, 2005
  14. Martin UeckerApr 20, 2005
  15. Morten WelinderApr 20, 2005
  16. Jon SeymourApr 20, 2005
  17. C. Scott AnanianApr 20, 2005
  18. Martin UeckerApr 20, 2005
  19. C. Scott AnanianApr 20, 2005
  20. Martin UeckerApr 20, 2005
  21. Martin UeckerApr 20, 2005
  22. Blob chunking code. [First look.]C. Scott Ananian, Apr 20, 2005
  23. Blob chunking code. [Second look]C. Scott Ananian, Apr 20, 2005
  24. David WoodhouseApr 20, 2005
  25. Linus TorvaldsApr 20, 2005
  26. David WoodhouseApr 20, 2005
  27. Chris MasonApr 20, 2005
  28. C. Scott AnanianApr 20, 2005
  29. Linus TorvaldsApr 20, 2005
  30. C. Scott AnanianApr 20, 2005
  31. Linus TorvaldsApr 20, 2005
  32. Linus TorvaldsApr 20, 2005
  33. David WillmoreApr 20, 2005
  34. Linus TorvaldsApr 20, 2005
  35. Linus TorvaldsApr 20, 2005
  36. Chris MasonApr 20, 2005
  37. Linus TorvaldsApr 20, 2005
  38. Chris MasonApr 20, 2005
  39. Linus TorvaldsApr 20, 2005
  40. Chris MasonApr 20, 2005
  41. Linus TorvaldsApr 20, 2005
  42. Linus TorvaldsApr 20, 2005
  43. David S. MillerApr 20, 2005
  44. David LangApr 19, 2005
  45. Linus TorvaldsApr 19, 2005
  46. David LangApr 19, 2005
  47. Linus TorvaldsApr 19, 2005
  48. David LangApr 19, 2005
  49. Linus TorvaldsApr 19, 2005
  50. Christopher LiApr 19, 2005
  51. Olivier GalibertApr 19, 2005
  52. C. Scott AnanianApr 19, 2005
  53. Linus TorvaldsApr 20, 2005
  54. C. Scott AnanianApr 20, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.