git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: cvs import

From
NSNathaniel Smith <njs@pobox.com>
Date
Sep 14, 2006, 00:32 UTC
Message-ID
<20060914003242.GA19228@frances.vorpus.org>
In-Reply-To
<1158190921.29313.175.camel@neko.keithp.com>
On Wed, Sep 13, 2006 at 04:42:01PM -0700, Keith Packard wrote:
> However, this means that parsecvs must hold the entire tree state in
> memory, which turned out to be its downfall with large repositories.
> Worked great for all of X.org, not so good with Mozilla.

Does anyone know how big Mozilla (or other humonguous repos, like KDE) are, in terms of number of files?

A few numbers for repositories I had lying around:
  Linux kernel -- ~21,000
  gcc -- ~42,000
  NetBSD "src" repo -- ~100,000
  uClinux distro -- ~110,000

These don't seem very indimidating... even if it takes an entire kilobyte per CVS revision to store the information about it that we need to make decisions about how to move the frontier... that's only 110 megabytes for the largest of these repos. The frontier sweeping algorithm only _needs_ to have available the current frontier, and the current frontier+1. Storing information on every version of every file in memory might be worse; but since the algorithm accesses this data in a linear way, it'd be easy enough to stick those in a lookaside table on disk if really necessary, like a bdb or sqlite file or something.

(Again, in practice storing all the metadata for the entire 180k revisions of the 100k files in the netbsd repo was possible on a desktop. Monotone's cvs_import does try somewhat to be frugal about memory, though, interning strings and suchlike.)

-- Nathaniel
-- 
When the flush of a new-born sun fell first on Eden's green and gold,
Our father Adam sat under the Tree and scratched with a stick in the mould;
And the first rude sketch that the world had seen was joy to his mighty heart,
Till the Devil whispered behind the leaves, "It's pretty, but is it Art?"
  -- The Conundrum of the Workshops, Rudyard Kipling
Previous: Keith PackardNext: Jon Smirl
Message 31 of 38 in “Re: cvs import”
  1. Jon SmirlSep 13, 2006
  2. Martin LanghoffSep 13, 2006
  3. Markus SchiltknechtSep 13, 2006
  4. Oswald BuddenhagenSep 13, 2006
  5. Martin LanghoffSep 13, 2006
  6. Michael HaggertySep 14, 2006
  7. Jon SmirlSep 14, 2006
  8. Michael HaggertySep 14, 2006
  9. Martin LanghoffSep 14, 2006
  10. Michael HaggertySep 14, 2006
  11. Jon SmirlSep 14, 2006
  12. Martin LanghoffSep 14, 2006
  13. Markus SchiltknechtSep 13, 2006
  14. Jon SmirlSep 13, 2006
  15. Michael HaggertySep 14, 2006
  16. Shawn PearceSep 14, 2006
  17. Jakub NarebskiSep 14, 2006
  18. Shawn PearceSep 14, 2006
  19. Jon SmirlSep 14, 2006
  20. Michael HaggertySep 14, 2006
  21. Jakub NarebskiSep 14, 2006
  22. Jon SmirlSep 14, 2006
  23. Markus SchiltknechtSep 15, 2006
  24. Shawn PearceSep 16, 2006
  25. Oswald BuddenhagenSep 16, 2006
  26. Nathaniel SmithSep 16, 2006
  27. Nathaniel SmithSep 13, 2006
  28. Daniel CarosoneSep 13, 2006
  29. Daniel CarosoneSep 13, 2006
  30. Keith PackardSep 13, 2006
  31. Nathaniel SmithSep 14, 2006
  32. Jon SmirlSep 14, 2006
  33. Daniel CarosoneSep 14, 2006
  34. Shawn PearceSep 14, 2006
  35. Daniel CarosoneSep 14, 2006
  36. Petr BaudisSep 14, 2006
  37. Shawn PearceSep 14, 2006
  38. Shawn PearceSep 14, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.