git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: WARNING! Object DB conversion (was Re: [PATCH] write-tree performance problems)

From
Linus Torvalds <torvalds@osdl.org>
Date
Apr 20, 2005, 14:59 UTC
Message-ID
<Pine.LNX.4.58.0504200731590.6467@ppc970.osdl.org>
In-Reply-To
<1114006429.5877.42.camel@localhost.localdomain>
On Thu, 21 Apr 2005, David Woodhouse wrote:
Show 5 quoted lines
> 
> The reason for doing this is that without it, we can't ever have a full
> history actually connected to the current trees. There'd always be a
> break at 2.6.12-rc2, at which point you'd have to switch to an entirely
> different git repository.

Quite frankly, I'd _much_ rather have a notion of "external references" than start depending on external hashes.

IOW, I'd be happier with a new line in the header (after the normal "author"/"committer" lines) that just pointed to an external tree, aka

	external linux-2.6.12-rc2-tree

and then people could literally use this to link whatever they wanted, and it would not force one particular version of an external tree on you.

Why? Because we can't keep re-generating trees.

However, the second part of that plan is that once you do that, you might as well make the "external" linkages be external to the repository itself. IOW, you could just make a file that the git tools can parse that say

	external-parent <root-hash> <external-parent-ID>
		comment for this parent
	external-parent <commit-hash> <external-parent-ID>
		comment for this parent

and the nice thing about that is that now that information allows you to add external parents at any point.

Why do it like this? First off, I think that the "initial import" ends up being just one special case of the much more _generic_ issue of having patches come in from other source control systems (ie the above would actually work with the darcs issues too, and allow people to track the dependencies between a tree maintained in git and maintained elsewhere).

Secondly, we do need something like this for pruning off history anyway, so that the tools have a better way of saying "history has been pruned off" than just hitting a missing commit. That's not a big deal right now, since I'm not planning on letting people prune their history (or at least I'm planning on having tools complain loudly), but it _will_ be an issue. I think history pruning is wonderful, but I do want to have some mechanism to say "it was pruned" as opposed to "it was lost".

Thirdly, I don't actually want my new tree to depend on a conversion of the old BK tree.

Two reasons: if it's a really full conversion, there are definitely going to be issues with BitMover. They do not want people to try to reverse engineer how they do namespace merges, which is why they have the "don't look at git and do another SCM at the same time" clause in the first place. Namespace merges (and probably other things too, for that matter) tend to be the thing they tend to do better than anybody else. The kernel probably does not actually have a lot of those so it might be ok by them, but the keyword is _might_, and I don't want to cloud git by another flamewar.

The other reason is just the really obvious one: in the last week, I've already changed the format _twice_ in ways that change the hash. As long as it's 119MB of data, it's not going to be too nasty to do again. If it's 3+GB of data, I'm going to feel really constrained about the kind of conversions I can do. It's one thing to have something that takes a few minutes and that anybody can do. It's another thing entirely to do something that requires the convertee to dedicate tons of diskspace and hours of work on it.

Let's face it, I doubt we did our last conversion ever. I still think that the git data model is the best model _ever_ for an SCM, but it's not all the minute details I'm proud over, it's the general big things. For example, let's see how the "blobs are sequences of smaller hashes" thing works out. I was doubtful, but Scott's first chunking code doesn't make me hurl chunks, and I've been wrong before.

And the thing is, I'm ok with being wrong. Especially if I can fix things up later.

So I've got tons of reasons (that you may not agree with, obviously) for why I don't think it's a good idea to base the kernel on a large conversion. Some (or all) of those reasons may become moot in another week or month, but I'd definitely _not_ that interested in doing it now. If it turns out later that we do want to re-base the kernel, we can do any conversion we want at a later time - it's not that it's necessarily the wrong thing to do, but I think it is the wrogn thing to do _now_.

		Linus
Previous: David WoodhouseNext: David Woodhouse
Message 25 of 54 in “write-tree performance problems”
  1. write-tree performance problemsChris Mason, Apr 19, 2005
  2. Linus TorvaldsApr 19, 2005
  3. Chris MasonApr 19, 2005
  4. Linus TorvaldsApr 19, 2005
  5. Chris MasonApr 19, 2005
  6. Linus TorvaldsApr 19, 2005
  7. Chris MasonApr 20, 2005
  8. Linus TorvaldsApr 20, 2005
  9. Linus TorvaldsApr 20, 2005
  10. H. Peter AnvinApr 20, 2005
  11. WARNING! Object DB conversion (was Re: [PATCH] write-tree performance problems)Linus Torvalds, Apr 20, 2005
  12. Ingo MolnarApr 20, 2005
  13. Jon SeymourApr 20, 2005
  14. Martin UeckerApr 20, 2005
  15. Morten WelinderApr 20, 2005
  16. Jon SeymourApr 20, 2005
  17. C. Scott AnanianApr 20, 2005
  18. Martin UeckerApr 20, 2005
  19. C. Scott AnanianApr 20, 2005
  20. Martin UeckerApr 20, 2005
  21. Martin UeckerApr 20, 2005
  22. Blob chunking code. [First look.]C. Scott Ananian, Apr 20, 2005
  23. Blob chunking code. [Second look]C. Scott Ananian, Apr 20, 2005
  24. David WoodhouseApr 20, 2005
  25. Linus TorvaldsApr 20, 2005
  26. David WoodhouseApr 20, 2005
  27. Chris MasonApr 20, 2005
  28. C. Scott AnanianApr 20, 2005
  29. Linus TorvaldsApr 20, 2005
  30. C. Scott AnanianApr 20, 2005
  31. Linus TorvaldsApr 20, 2005
  32. Linus TorvaldsApr 20, 2005
  33. David WillmoreApr 20, 2005
  34. Linus TorvaldsApr 20, 2005
  35. Linus TorvaldsApr 20, 2005
  36. Chris MasonApr 20, 2005
  37. Linus TorvaldsApr 20, 2005
  38. Chris MasonApr 20, 2005
  39. Linus TorvaldsApr 20, 2005
  40. Chris MasonApr 20, 2005
  41. Linus TorvaldsApr 20, 2005
  42. Linus TorvaldsApr 20, 2005
  43. David S. MillerApr 20, 2005
  44. David LangApr 19, 2005
  45. Linus TorvaldsApr 19, 2005
  46. David LangApr 19, 2005
  47. Linus TorvaldsApr 19, 2005
  48. David LangApr 19, 2005
  49. Linus TorvaldsApr 19, 2005
  50. Christopher LiApr 19, 2005
  51. Olivier GalibertApr 19, 2005
  52. C. Scott AnanianApr 19, 2005
  53. Linus TorvaldsApr 20, 2005
  54. C. Scott AnanianApr 20, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.