git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC] Support projects including other projects

From
DLDavid Lang <david.lang@digitalinsight.com>
Date
May 12, 2005, 17:24 UTC
Message-ID
<Pine.LNX.4.62.0505121006150.25177@qynat.qvtvafvgr.pbz>
In-Reply-To
<Pine.LNX.4.21.0505121218280.30848-100000@iabervon.org>

I was thinking about this recently while reading an article on bittorrent and how it works and it occured to me that perhapse the network access model of git should be reexamined.

git produces a large pool of objects, there are two ways that people want to access these objects.

1. pull the current version of a project (either a straight 'ckeckout' 
type pull or a 'merge' to a local project)
2. pull the objects nessasary for past versions of a project (either all 
the way back to the beginning of time or back to some point, that point 
being a number of possibilities (date, version, things you don't have, 
etc)

in either case the important thing that's key are the indexes related to a particular project, the objects themselves could all be in one huge pool for all projects that ever existed (this doesn't make sense if you use rsync to copy repositories as Linux origionally did, but if you have a more git-aware transport it can make sense)

I believe that there are going to be quite a number of cases where the same object is used for multiple projects (either becouse the project is a fork of another project or becouse some functions (or include files) are so trivial that they are basicly boilerplate and get reused or recreated) if you think about a major mirror server distributing a dozen linux distros via git you will realize that in many cases the source files, scripts, and (in many cases) even the binaries are really going to be identical objects for all the distros so a ftp/http server that used a git filesystem could result in a pretty significant saveings in disk space.

In addition, when you are doing a pull you can accept data from non-authoritative sources since each object (and it's index info) includes enough info to validate the object hasn't been tampered with (at least until such time as the hashes are sufficiantly broken, but that's another debate, and we had that one :-). so a bittorrent-like peer sharing system to fetch objects identified by the index files would open the potential for saving significant bandwith on the master servers while not comprimising the trees at all.

Going back (somewhat) to the subject at hand, with something like this you should be able to combine as many projects as you want in one repository, and the only issue would be the work nessasary to go through that repository and all the index files that point at it when you want to prune old data out of the object pool to save disk space.

thoughts? unfortunnatly I don't have the time to even consider codeing something like this up, but hopefully it will spark interest for someone who does.

David Lang
-- 
There are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.
  -- C.A.R. Hoare
Previous: Daniel BarkalowNext: Junio C Hamano
Message 8 of 14 in “[RFC] Support projects including other projects”
  1. Daniel BarkalowMay 12, 2005
  2. Junio C HamanoMay 12, 2005
  3. Daniel BarkalowMay 12, 2005
  4. Junio C HamanoMay 12, 2005
  5. Daniel BarkalowMay 12, 2005
  6. Junio C HamanoMay 12, 2005
  7. Daniel BarkalowMay 12, 2005
  8. David LangMay 12, 2005
  9. Junio C HamanoMay 12, 2005
  10. Daniel BarkalowMay 12, 2005
  11. Junio C HamanoMay 12, 2005
  12. James PurserMay 12, 2005
  13. Daniel BarkalowMay 12, 2005
  14. James PurserMay 12, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.