{"thread":{"id":"229","subject":"[RFC] A suggestion for more versatile naming conventios in the object database","startedAt":"2005-04-21T19:51:31Z","lastAt":"2005-04-21T19:51:31Z","messageCount":1,"participants":["Imre Simon"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"1182","messageId":"42680443.3050604@ime.usp.br","threadId":"229","inReplyTo":null,"subject":"[RFC] A suggestion for more versatile naming conventios in the object database","fromName":"Imre Simon","fromEmail":"is@ime.usp.br","sentAt":"2005-04-21T19:51:31Z","receivedAt":"2005-04-21T19:51:31Z","isPatch":false,"sender":{"key":"is@ime.usp.br","avatar":null},"body":"In the transition from git-0.04 to git-0.5 (Linus' track) the naming \nconvention of files in the object database has been changed: a file's \nname passed from the sha1 of its contents to the sha1 of its contents \n*before* compression.\n\nThis change was preceded by a long discussion hence both conventions\nhave arguments for and against.\n\nI would like to suggest to adopt a more versatile solution:\n\n   preserve the pure sha1 based names for the sha1 sum of the file's\n   contents. I mean,\n\n       (*)  for files with name xy/z{38} their sha1 sum is xyz{38}\n\n   allow other files (or links) with names of the form\n\n       xy/z{38}.EXTENSION\n\n   where for every EXTENSION the file's content would be the EXTENSION\n   representation of the file xy/z{38} . For every representation type\n   EXTENSION there should be procedures to derive the file xy/z{38}\n   from the name xy/z{38}.EXTENSION and vice-versa (assuming that the\n   representation type EXTENSION cares about the contents of file\n   xy/z{38}).\n\nLet me give two examples:\n\n    all the files in the object database of git-0.04 are just fine, they\n    satisfy axiom (*)\n\n    the name of every file xy/z{38}  in the git-0.5 data base should be\n    changed to xy/z{38}.g assuming that we will use EXTENSION g as the\n    git representation type. The conversion algorithms would be:\n\n        cat-file `cat-file -t xyz{38}` xyz{38}  to obtain the contents\n        represented by xy/z{38}.g whose sha1 is xyz{38}\n\n        and a utility program has to be written to check whether a given\n        file F, is a valid contents as far as git is concerned and in\n        case it is compute its sha1 sum xyz{38} and also comute the file\n        the file xy/z{38}.g .\n\nSo, what are the advantages of this further complication? I see these ones:\n\n   git is separated from the idea of sha1 addressable contents, which\n   indeed is an idea larger than git itself. This same or similar\n   addressing schemes can (and most probably will) be applied to other\n   contents besides SCMs. An example would be a digital library of\n   scientific papers in pdf together with its OAI compliant meta data\n   (don't bother if you are not familiar with these terms, it is just\n   an example and I am sure you are able to come up with many other\n   examples where a sha1 addressable data base would be interesting)\n\n   all these uses could share common backup schemes where axiom (*)\n   would be enforced. One could think of a shared p2p database of\n   repositories of sha1 addressed contents of all kinds. This might be\n   important because, in general, the contents of xyz{38} cannot be\n   reconstructed from its name. The way to defend against file system\n   corruption is replication. Why not share these backup databases?\n\n   it would be easier to experiment with other compression schemes or\n   other proposals for meta data in git itself.\n\n   it would be easier to experiment with the factorization of common\n   chunks of contents, an idea very close to the secret of rsync's\n   amazing efficiency.\n\nWell, that's the proposal. I would be happy to hear comments!\n\nCheers,\n\nImre Simon\n\nPS: the way it is, the git-0.5 README file is inconsistent. The naming\nchange is not reflected in the README file which in many places states\nthat the sha1 sum of file xy/z{38} is xyz{38}.\n\n"}]}