git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Creating objects manually and repack

From
Linus Torvalds <torvalds@osdl.org>
Date
Aug 4, 2006, 04:46 UTC
Message-ID
<Pine.LNX.4.64.0608032138330.4168@g5.osdl.org>
In-Reply-To
<9e4733910608032124o5b5b69b5hda2eb8cb1e0ac959@mail.gmail.com>
On Fri, 4 Aug 2006, Jon Smirl wrote:
Show 5 quoted lines
>
> I am converting all of the revisions from each CVS file into git
> objects the first time the file is parsed. The plan was to run repack
> after each file is finished. That way it should be easy to figure out
> the deltas since everything will be a variation on the same file.

Sure. In that case, just list the object ID's in the exact same order you created them.

Basically,as you create them, just keep a list of all ID's you've created, and every (say) 50,000 objects, just do a

	echo all objects you've created | git-pack-objects new-pack

and then move the new pack into place, and remove all the loose objects (don't even bother using "git prune" - just basically do something like "rm -rf .git/objects/??" to get rid of them).

> So what's the best way to pack these objects, append them to the
> existing pack and then clean everything up for the next file? I am
> parsing 120K CVS files containing over 1M revs.

You'll want to repack every once in a while just to not ever have _tons_ of those loose objects around, but if you do it every 50,000 objects, you'll have just twenty nice pack-files once you're done, containing all one million objects, and you'll never have had more than ~200 files in any of the loose object subdirectories.

Of course, you might want to make that "every 50,000 object" thing tunable, so that if you don't have a lot of memory for caching, you might want to do it a bit more often just to make each repack go faster and not have tons of IO.

You can then do a _full_ repack to get one big object, by just listing every object you ever created (in creation order) to git-pack-objects, and then you can replace all the twenty (smaller) pack-files with the resulting single bigger one.

In fact, at that point you no longer even need to worry about "creation order", since you've basically created all the deltas in the first phase, and regardless of ordering, when you then repack everything at the end, it will re-use all earlier delta information.

		Linus
Previous: Jon SmirlNext: Linus Torvalds
Message 5 of 36 in “Creating objects manually and repack”
  1. Jon SmirlAug 4, 2006
  2. Jeff KingAug 4, 2006
  3. Linus TorvaldsAug 4, 2006
  4. Jon SmirlAug 4, 2006
  5. Linus TorvaldsAug 4, 2006
  6. Linus TorvaldsAug 4, 2006
  7. Jon SmirlAug 4, 2006
  8. Jon SmirlAug 4, 2006
  9. Jon SmirlAug 4, 2006
  10. Linus TorvaldsAug 4, 2006
  11. Jon SmirlAug 4, 2006
  12. A Large Angry SCMAug 4, 2006
  13. Jon SmirlAug 4, 2006
  14. Linus TorvaldsAug 4, 2006
  15. Linus TorvaldsAug 4, 2006
  16. Rogan DawesAug 4, 2006
  17. Jon SmirlAug 4, 2006
  18. Linus TorvaldsAug 4, 2006
  19. Jon SmirlAug 4, 2006
  20. Linus TorvaldsAug 4, 2006
  21. Linus TorvaldsAug 4, 2006
  22. Junio C HamanoAug 4, 2006
  23. Linus TorvaldsAug 4, 2006
  24. Carl WorthAug 4, 2006
  25. Junio C HamanoAug 4, 2006
  26. Carl WorthAug 4, 2006
  27. Carl WorthAug 4, 2006
  28. Jakub NarebskiAug 4, 2006
  29. Junio C HamanoAug 4, 2006
  30. Jakub NarebskiAug 4, 2006
  31. Martin LanghoffAug 5, 2006
  32. Jon SmirlAug 5, 2006
  33. Shawn PearceAug 5, 2006
  34. Jon SmirlAug 5, 2006
  35. Shawn PearceAug 5, 2006
  36. Shawn PearceAug 5, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.