git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Creating objects manually and repack

From
Jon Smirl <jonsmirl@gmail.com>
Date
Aug 5, 2006, 05:40 UTC
Message-ID
<9e4733910608042240u581dd23q3859ebcfe4268ce2@mail.gmail.com>
In-Reply-To
<20060805052135.GA18679@spearce.org>
On 8/5/06, Shawn Pearce <spearce@spearce.org> wrote:
Show 11 quoted lines
> I'm almost done with what I'm calling `git-fast-import`.  It takes
> a stream of blobs on STDIN and writes the pack to a file, printing
> SHA1s in hex format to STDOUT.  The basic format for STDIN is a 4
> byte length (native format) followed by that many bytes of blob data.
> It prints the SHA1 for that blob to STDOUT, then waits for another
> length.
>
> It naively deltas each object against the prior object, thus it
> would be best to feed it one ,v file at a time working from the most
> recent revision back to the oldest revision.  This works well for
> an RCS file as that's the natural order to process the file in.  :-)
I am already doing this.
> When done you close STDIN and it'll rip through and update the pack
> object count and the trailing checksum.  This should let you pack
> the entire repository in delta format using only two passes over the
> data: one to write out the pack file and one to compute its checksum.

Thinking about this some more, the existing repack code could be made to work with minor changes. I would like to feed repack 1M revisions which are sorted by file and then newest to oldest. The problem is that my expanded revs take up 12GB disk space.

How about adding a flag to repack that simply says delete the objects when done with them? I'd still create all of the objects on disk. Repack would assume that they have at least been sorted by filename. So repack could read in object names until it sees a change in the file name, sort them by size, deltafy, write out the pack and then delete the objects from that batch. Then repeat this process for the next file name on stdin.

I'm making two assumptions, first that blocks from a deleted file don't get written to disk. And that by deleting the file the file system will use the same blocks over and over. If those assumptions are close to being true then the cache shouldn't thrash. They don't have to be totally true, close is good enough.

Of course eliminating the files all together will be even faster.
-- 
Jon Smirl
jonsmirl@gmail.com
Previous: Shawn PearceNext: Shawn Pearce
Message 34 of 36 in “Creating objects manually and repack”
  1. Jon SmirlAug 4, 2006
  2. Jeff KingAug 4, 2006
  3. Linus TorvaldsAug 4, 2006
  4. Jon SmirlAug 4, 2006
  5. Linus TorvaldsAug 4, 2006
  6. Linus TorvaldsAug 4, 2006
  7. Jon SmirlAug 4, 2006
  8. Jon SmirlAug 4, 2006
  9. Jon SmirlAug 4, 2006
  10. Linus TorvaldsAug 4, 2006
  11. Jon SmirlAug 4, 2006
  12. A Large Angry SCMAug 4, 2006
  13. Jon SmirlAug 4, 2006
  14. Linus TorvaldsAug 4, 2006
  15. Linus TorvaldsAug 4, 2006
  16. Rogan DawesAug 4, 2006
  17. Jon SmirlAug 4, 2006
  18. Linus TorvaldsAug 4, 2006
  19. Jon SmirlAug 4, 2006
  20. Linus TorvaldsAug 4, 2006
  21. Linus TorvaldsAug 4, 2006
  22. Junio C HamanoAug 4, 2006
  23. Linus TorvaldsAug 4, 2006
  24. Carl WorthAug 4, 2006
  25. Junio C HamanoAug 4, 2006
  26. Carl WorthAug 4, 2006
  27. Carl WorthAug 4, 2006
  28. Jakub NarebskiAug 4, 2006
  29. Junio C HamanoAug 4, 2006
  30. Jakub NarebskiAug 4, 2006
  31. Martin LanghoffAug 5, 2006
  32. Jon SmirlAug 5, 2006
  33. Shawn PearceAug 5, 2006
  34. Jon SmirlAug 5, 2006
  35. Shawn PearceAug 5, 2006
  36. Shawn PearceAug 5, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.