{"thread":{"id":"11","subject":"Yet another base64 patch","startedAt":"2005-04-14T02:24:13Z","lastAt":"2005-04-18T16:42:13Z","messageCount":34,"participants":["H. Peter Anvin","Christopher Li","Linus Torvalds","bert hubert","Paul Jackson","Paul Dickson","David A. Wheeler","David Lang","Daniel Barkalow","Petr Baudis","Kevin Smith"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"39","messageId":"20050414022413.GB18655@64m.dyndns.org","threadId":"11","inReplyTo":"425DEF64.60108@zytor.com","subject":"Re: Yet another base64 patch","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T02:24:13Z","receivedAt":"2005-04-14T02:24:13Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Wed, Apr 13, 2005 at 09:19:48PM -0700, H. Peter Anvin wrote:\n> Checking out the total kernel tree (time checkout-cache -a into an empty \n> directory):\n> \n> \tCache cold\tCache hot\n> stock\t3:46.95\t\t19.95\n> base64\t5:56.20\t\t23.74\n> flat\t2:44.13\t\t15.68\n\n\n> It seems that the flat format, at least on ext3 with dircache, is \n> actually a major performance win, and that the second level loses quite \n> a bit.\n\nThat is not surprising due to the directory index in ext3. Htree is pretty\ngood at random access and the hashed file name distribute evenly, that is\nthe best case for htree. \n\nChris\n\n"},{"id":"41","messageId":"20050414024228.GC18655@64m.dyndns.org","threadId":"11","inReplyTo":"425E0174.4080404@zytor.com","subject":"Re: Yet another base64 patch","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-14T02:42:28Z","receivedAt":"2005-04-14T02:42:28Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Wed, Apr 13, 2005 at 10:36:52PM -0700, H. Peter Anvin wrote:\n> Christopher Li wrote:\n> >On Wed, Apr 13, 2005 at 09:19:48PM -0700, H. Peter Anvin wrote:\n> >\n> >That is not surprising due to the directory index in ext3. Htree is pretty\n> >good at random access and the hashed file name distribute evenly, that is\n> >the best case for htree. \n> >\n> \n> Right, so by not trying to do the filesystem's job for it we actually \n> come out ahead.\n>\n\nBut if you write a large number of random files, when htree has three\nlevels index. htree will suffer on the effect that it dirty random block\nvery quickly, most block get dirty only contain one or two new entries.\nExt3 will choke on it due to the limited journal size.\n\nWhile non-index directory, new entry are very compact on the blocks.\nSo it end up dirty a lot less blocks, of course, lookup will suffer.\n\nDepend on you want check out fast or write a big tree fast, you can't\nwin it all.\n\nChris\n\n \n"},{"id":"34","messageId":"425DEF64.60108@zytor.com","threadId":"11","inReplyTo":null,"subject":"Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-14T04:19:48Z","receivedAt":"2005-04-14T04:19:48Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"I am assuming this will be the last one one way or another...\n\nI decided that filenames/tags beginning with - was a really bad thing,\nso I decided that, ugly though it might be, the best was to do a hybrid\nbetween regular base64 (+ /) and filesystem-safe base64 (- _) and use\n+ _ as the nonalpha characters needed.  I have updated the base64\npatches as well as gitcvt, and also put out a flat version of gitcvt.\n\ngitcvt also now converts the HEAD file over.  This requires pointing it\nat the .dircache/.git directory instead of the objects directory inside.\n  I have tested it on both the git and the kernel-test repositories.\n\nChecking out the total kernel tree (time checkout-cache -a into an empty \ndirectory):\n\n\tCache cold\tCache hot\nstock\t3:46.95\t\t19.95\nbase64\t5:56.20\t\t23.74\nflat\t2:44.13\t\t15.68\n\nIt seems that the flat format, at least on ext3 with dircache, is \nactually a major performance win, and that the second level loses quite \na bit.\n\n\t-hpa\n\n"},{"id":"35","messageId":"425DF0BB.8020802@zytor.com","threadId":"11","inReplyTo":"425DEF64.60108@zytor.com","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-14T04:25:31Z","receivedAt":"2005-04-14T04:25:31Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"H. Peter Anvin wrote:\n> \n> It seems that the flat format, at least on ext3 with dircache, is \n> actually a major performance win, and that the second level loses quite \n> a bit.\n> \n\ns/dircache/dir_index/\n\n\t-hpa\n"},{"id":"40","messageId":"425E0174.4080404@zytor.com","threadId":"11","inReplyTo":"20050414022413.GB18655@64m.dyndns.org","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-14T05:36:52Z","receivedAt":"2005-04-14T05:36:52Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Christopher Li wrote:\n> On Wed, Apr 13, 2005 at 09:19:48PM -0700, H. Peter Anvin wrote:\n> \n>>Checking out the total kernel tree (time checkout-cache -a into an empty \n>>directory):\n>>\n>>\tCache cold\tCache hot\n>>stock\t3:46.95\t\t19.95\n>>base64\t5:56.20\t\t23.74\n>>flat\t2:44.13\t\t15.68\n> \n>>It seems that the flat format, at least on ext3 with dircache, is \n>>actually a major performance win, and that the second level loses quite \n>>a bit.\n> \n> That is not surprising due to the directory index in ext3. Htree is pretty\n> good at random access and the hashed file name distribute evenly, that is\n> the best case for htree. \n> \n\nRight, so by not trying to do the filesystem's job for it we actually \ncome out ahead.\n\n\t-hpa\n"},{"id":"43","messageId":"425E0D62.9000401@zytor.com","threadId":"11","inReplyTo":"20050414024228.GC18655@64m.dyndns.org","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-14T06:27:46Z","receivedAt":"2005-04-14T06:27:46Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Christopher Li wrote:\n> \n> But if you write a large number of random files, when htree has three\n> levels index. htree will suffer on the effect that it dirty random block\n> very quickly, most block get dirty only contain one or two new entries.\n> Ext3 will choke on it due to the limited journal size.\n> \n> While non-index directory, new entry are very compact on the blocks.\n> So it end up dirty a lot less blocks, of course, lookup will suffer.\n> \n> Depend on you want check out fast or write a big tree fast, you can't\n> win it all.\n> \n\nActually, the subdirectory hack has the same effect, so you lose \nregardless.  Doesn't mean that you can't construct cases where the \nsubdirectory hack doesn't win, but I maintain that those are likely to \nbe artificial.\n\nIt's probably worth noting that you have to assume htree is on, since \nthat's the typical default for a Linux installation, even if you use the \nsubdirectory hack.\n\n\t-hpa\n"},{"id":"44","messageId":"425E0F2D.4040507@zytor.com","threadId":"11","inReplyTo":"425E0D62.9000401@zytor.com","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-14T06:35:25Z","receivedAt":"2005-04-14T06:35:25Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"H. Peter Anvin wrote:\n\n> \n> Actually, the subdirectory hack has the same effect, so you lose \n> regardless.  Doesn't mean that you can't construct cases where the \n> subdirectory hack doesn't win, but I maintain that those are likely to \n> be artificial.\n> \n\nThat should, of course, be \"... where the subdirectory hack does win ...\"\n\nReally, the subdirectory hack is a workaround for broken filesystems, \nand we don't use those anymore.\n\n\t-hpa\n"},{"id":"49","messageId":"Pine.LNX.4.58.0504140038450.7211@ppc970.osdl.org","threadId":"11","inReplyTo":"425E0D62.9000401@zytor.com","subject":"Re: Yet another base64 patch","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-14T07:40:55Z","receivedAt":"2005-04-14T07:40:55Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 13 Apr 2005, H. Peter Anvin wrote:\n> \n> Actually, the subdirectory hack has the same effect, so you lose \n> regardless.  Doesn't mean that you can't construct cases where the \n> subdirectory hack doesn't win, but I maintain that those are likely to \n> be artificial.\n\nI'll tell you why a flat object directory format simply isn't an option.\n\nHint: maximum directory size. It's limited by n_link, and it's almost\nuniversally a 16-bit number on Linux (and generally artifically limited to\n32000 entries).\n\nIn other words, if you ever expect to have more than 32000 objects, a flat \nspace simply isn't possible.\n\n\t\tLinus\n"},{"id":"51","messageId":"Pine.LNX.4.58.0504140114260.7211@ppc970.osdl.org","threadId":"11","inReplyTo":"425DEF64.60108@zytor.com","subject":"Re: Yet another base64 patch","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-14T08:17:18Z","receivedAt":"2005-04-14T08:17:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 13 Apr 2005, H. Peter Anvin wrote:\n> \n> Checking out the total kernel tree (time checkout-cache -a into an empty \n> directory):\n> \n> \t\tCache cold\tCache hot\n> stock\t\t3:46.95\t\t19.95\n> base64\t5:56.20\t\t23.74\n> flat\t\t2:44.13\t\t15.68\n\nSo why is \"base64\" worse than the stock one?\n\nAs mentioned, the \"flat\" version may be faster, but it really isn't an\noption. 32000 objects is peanuts. Any respectable source tree may hit that\nin a short time, and will break in horrible ways on many Linux\nfilesystems.\n\nSo you need at least a single level of subdirectory. \n\nWhat I don't get is why the stock hex version would be better than base64.\n\nI like the result, I just don't _understand_ it.\n\n\t\tLinus\n"},{"id":"90","messageId":"425EA152.4090506@zytor.com","threadId":"11","inReplyTo":"Pine.LNX.4.58.0504140038450.7211@ppc970.osdl.org","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-14T16:58:58Z","receivedAt":"2005-04-14T16:58:58Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> I'll tell you why a flat object directory format simply isn't an option.\n> \n> Hint: maximum directory size. It's limited by n_link, and it's almost\n> universally a 16-bit number on Linux (and generally artifically limited to\n> 32000 entries).\n> \n> In other words, if you ever expect to have more than 32000 objects, a flat \n> space simply isn't possible.\n> \n\nEh?!  n_link limits the number of *subdirectories* a directory can \ncontain, not the number of *entries*.\n\n\t-hpa\n\n\t\n"},{"id":"91","messageId":"425EA23F.6010900@zytor.com","threadId":"11","inReplyTo":"Pine.LNX.4.58.0504140114260.7211@ppc970.osdl.org","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-14T17:02:55Z","receivedAt":"2005-04-14T17:02:55Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> So why is \"base64\" worse than the stock one?\n> \n> As mentioned, the \"flat\" version may be faster, but it really isn't an\n> option. 32000 objects is peanuts. Any respectable source tree may hit that\n> in a short time, and will break in horrible ways on many Linux\n> filesystems.\n> \n\nIf it does, it's not because of n_link; see previous email.\n\nI have used ext2 filesystems with hundreds of thousands of files per \ndirectory back in 1996.  It was slow but didn't break anything.\n\nThe only filesystem I know of which has a 2^16 entry limit is FAT.\n\n> So you need at least a single level of subdirectory. \n> \n> What I don't get is why the stock hex version would be better than base64.\n> \n> I like the result, I just don't _understand_ it.\n\nThe base64 version has 2^12 subdirectories instead of 2^8 (I just used 2 \ncharacters as the hash key just like the hex version.)  So it ascerbates \nthe performance penalty of subdirectory hashing.\n\n\t-hpa\n"},{"id":"94","messageId":"Pine.LNX.4.58.0504141042450.7211@ppc970.osdl.org","threadId":"11","inReplyTo":"425EA152.4090506@zytor.com","subject":"Re: Yet another base64 patch","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-14T17:42:56Z","receivedAt":"2005-04-14T17:42:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 14 Apr 2005, H. Peter Anvin wrote:\n> \n> Eh?!  n_link limits the number of *subdirectories* a directory can \n> contain, not the number of *entries*.\n\nDuh. I'm a git.\n\n\t\tLinus\n"},{"id":"109","messageId":"20050414191157.GA27696@outpost.ds9a.nl","threadId":"11","inReplyTo":"Pine.LNX.4.58.0504141042450.7211@ppc970.osdl.org","subject":"Re: Yet another base64 patch","fromName":"bert hubert","fromEmail":"ahu@ds9a.nl","sentAt":"2005-04-14T19:11:58Z","receivedAt":"2005-04-14T19:11:58Z","isPatch":false,"sender":{"key":"ahu@ds9a.nl","avatar":null},"body":"On Thu, Apr 14, 2005 at 10:42:56AM -0700, Linus Torvalds wrote:\n> > Eh?!  n_link limits the number of *subdirectories* a directory can \n> > contain, not the number of *entries*.\n> \n> Duh. I'm a git.\n\nThat may be true :-), but from the \"front lines\" I can report that\ndirectories with > 32000 or > 65000 entries is *asking* for trouble. There\nis a whole chain of systems that need to get things right for huge\ndirectories to work well, and it often is not that way.\n\nEven though it should be.\n\nSo the question is, should git be the harbringer of improvements in this\narea, or should it go with the flow.\n\nBert.\n\n-- \nhttp://www.PowerDNS.com      Open source, database driven DNS Software \nhttp://netherlabs.nl              Open and Closed source services\n"},{"id":"115","messageId":"425EC3B4.6090908@zytor.com","threadId":"11","inReplyTo":"20050414191157.GA27696@outpost.ds9a.nl","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-14T19:25:40Z","receivedAt":"2005-04-14T19:25:40Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"bert hubert wrote:\n> \n> That may be true :-), but from the \"front lines\" I can report that\n> directories with > 32000 or > 65000 entries is *asking* for trouble. There\n> is a whole chain of systems that need to get things right for huge\n> directories to work well, and it often is not that way.\n> \n\nSpecifics, please?\n\n\t-hpa\n"},{"id":"132","messageId":"20050414214756.GA31249@outpost.ds9a.nl","threadId":"11","inReplyTo":"425EC3B4.6090908@zytor.com","subject":"Re: Yet another base64 patch","fromName":"bert hubert","fromEmail":"ahu@ds9a.nl","sentAt":"2005-04-14T21:47:56Z","receivedAt":"2005-04-14T21:47:56Z","isPatch":false,"sender":{"key":"ahu@ds9a.nl","avatar":null},"body":"On Thu, Apr 14, 2005 at 12:25:40PM -0700, H. Peter Anvin wrote:\n> >That may be true :-), but from the \"front lines\" I can report that\n> >directories with > 32000 or > 65000 entries is *asking* for trouble. There\n> >is a whole chain of systems that need to get things right for huge\n> >directories to work well, and it often is not that way.\n> >\n> \n> Specifics, please?\n\nWe've seen even Linus assume there is a 65K limit, and it appears more\npeople have been confused.\n\nThe systems I've seen mess this up include backup tools (quite serious ones\ntoo), NetApp NFS servers, Samba shares and archivers.\n\nSome tools just fail visibly, which is good, others become so slow as to\neffectively lock up, which was the case with the backup tools. \n\nI've quite often been able to fix broken systems by hashing directories -\nmany problems just vanish. \n\nIt is too easy to get into a O(N^2) situation. Git may be able to deal with\nit but you may hurt yourself when making backups, or if you ever want to\nshare your tree (possibly with yourself) over the network.\n\nBut if you live in an all Linux world, and use mostly tar and rsync, it\nshould work.\n\nBert.\n\n-- \nhttp://www.PowerDNS.com      Open source, database driven DNS Software \nhttp://netherlabs.nl              Open and Closed source services\n"},{"id":"175","messageId":"Pine.LNX.4.58.0504141743360.7211@ppc970.osdl.org","threadId":"11","inReplyTo":"20050414214756.GA31249@outpost.ds9a.nl","subject":"Re: Yet another base64 patch","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-15T00:44:09Z","receivedAt":"2005-04-15T00:44:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 14 Apr 2005, bert hubert wrote:\n> \n> It is too easy to get into a O(N^2) situation. Git may be able to deal with\n> it but you may hurt yourself when making backups, or if you ever want to\n> share your tree (possibly with yourself) over the network.\n\nEven something as simple as \"ls -l\" has been known to have O(n**2)  \nbehaviour for big directories.\n\n\t\tLinus\n"},{"id":"177","messageId":"425F1394.5020709@zytor.com","threadId":"11","inReplyTo":"Pine.LNX.4.58.0504141743360.7211@ppc970.osdl.org","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-15T01:06:28Z","receivedAt":"2005-04-15T01:06:28Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> Even something as simple as \"ls -l\" has been known to have O(n**2)  \n> behaviour for big directories.\n> \n\nFor filesystems with linear directories, sure.  For sane filesystems, it \nshould have O(n log n).\n\n\t-hpa\n"},{"id":"179","messageId":"425F13C9.5090109@zytor.com","threadId":"11","inReplyTo":"Pine.LNX.4.58.0504141743360.7211@ppc970.osdl.org","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-15T01:07:21Z","receivedAt":"2005-04-15T01:07:21Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Thu, 14 Apr 2005, bert hubert wrote:\n> \n>>It is too easy to get into a O(N^2) situation. Git may be able to deal with\n>>it but you may hurt yourself when making backups, or if you ever want to\n>>share your tree (possibly with yourself) over the network.\n> \n> \n> Even something as simple as \"ls -l\" has been known to have O(n**2)  \n> behaviour for big directories.\n> \n\nUltimately the question is: do we care about old (broken) filesystems?\n\n\t-hpa\n"},{"id":"184","messageId":"20050414205831.01039ee8.pj@engr.sgi.com","threadId":"11","inReplyTo":"425F13C9.5090109@zytor.com","subject":"Re: Yet another base64 patch","fromName":"Paul Jackson","fromEmail":"pj@engr.sgi.com","sentAt":"2005-04-15T03:58:31Z","receivedAt":"2005-04-15T03:58:31Z","isPatch":false,"sender":{"key":"pj@engr.sgi.com","avatar":null},"body":"Earlier, hpa wrote:\n> The base64 version has 2^12 subdirectories instead of 2^8 (I just used 2 \n> characters as the hash key just like the hex version.)\n\nLater, hpa wrote:\n> Ultimately the question is: do we care about old (broken) filesystems?\n\nI'd imagine we care a little - just not alot.\n\nI'd think that going to 2^12 subdirectories, which with 2^12 entries per\nsubdirectory gets us to 16 million files before the leaf directories get\nbigger than the parent, is a good tradeoff.\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"259","messageId":"20050415165532.05ed5dc4.paul@permanentmail.com","threadId":"11","inReplyTo":"425DEF64.60108@zytor.com","subject":"Re: Yet another base64 patch","fromName":"Paul Dickson","fromEmail":"paul@permanentmail.com","sentAt":"2005-04-15T23:55:32Z","receivedAt":"2005-04-15T23:55:32Z","isPatch":false,"sender":{"key":"paul@permanentmail.com","avatar":null},"body":"On Wed, 13 Apr 2005 21:19:48 -0700, H. Peter Anvin wrote:\n\n> Checking out the total kernel tree (time checkout-cache -a into an empty \n> directory):\n> \n>         Cache cold      Cache hot\n> stock   3:46.95         19.95\n> base64  5:56.20         23.74\n> flat    2:44.13         15.68\n> \n> It seems that the flat format, at least on ext3 with dircache, is \n> actually a major performance win, and that the second level loses quite \n> a bit.\n\nSince 160-bits does not go into base64 evenly anyways, what happens if\nyou use 2^10 instead of 2^12 for the subdir names?  That will be 1/4 the\ndirectories of the base64 given above.\n\n\t-Paul\n\n"},{"id":"426","messageId":"4261DDBC.3050706@dwheeler.com","threadId":"11","inReplyTo":"20050414205831.01039ee8.pj@engr.sgi.com","subject":"Re: Yet another base64 patch","fromName":"David A. Wheeler","fromEmail":"dwheeler@dwheeler.com","sentAt":"2005-04-17T03:53:32Z","receivedAt":"2005-04-17T03:53:32Z","isPatch":false,"sender":{"key":"dwheeler@dwheeler.com","avatar":"https://avatars.githubusercontent.com/u/813150?v=4"},"body":"Paul Jackson wrote:\n> Earlier, hpa wrote:\n> \n>>The base64 version has 2^12 subdirectories instead of 2^8 (I just used 2 \n>>characters as the hash key just like the hex version.)\n> \n> Later, hpa wrote:\n> \n>>Ultimately the question is: do we care about old (broken) filesystems?\n> \n> \n> I'd imagine we care a little - just not alot.\n\nSome people (e.g., me) would really like for \"git\"\nto be more forgiving of nasty filesystems,\nso that git can be used very widely.\nI.E., be forgiving about case insensitivity,\npoor performance or problems with a large # of files\nin a directory, etc.  You're already working to make\nsure git handles filenames with spaces & i18n filenames,\na common failing of many other SCM systems.\n\nIf \"git\" is used for Linux kernel development & nothing else,\nit's still a success.  But it'd be even better from\nmy point of view if \"git\" was a useful tool for MANY\nother projects.  I think there are advantages, even if you\nonly plan to use git for the kernel, to making \"git\" easier\nto use for other projects.  By making git less\nsensitive to the filesystem, you'll attract more (non-kernel-dev)\nusers, some of whom will become new git developers who\nadd cool new functionality.\n\nAs noted in my SCM survey (http://www.dwheeler.com/essays/scm.html),\nI think SCM Windows support is really important to a lot of\nOSS projects.  Many OSS projects, even if they start\nUnix/Linux only, spin off a Windows port, and it's\npainful if their SCM can't run on Windows then.\nProblems running on NFS filesystems have caused problems\nwith GNU Arch users (there are workarounds, but now you\nneed to learn about workarounds instead of things\n\"just working\").  If nothing else, look at the history\nof other SCM projects: all too many have undergone radical and\npainful surgeries so that they can be more portable to\nvarious filesystems.\n\nIt's a trade-off, I know.\n\n--- David A. Wheeler\n"},{"id":"427","messageId":"20050416210513.1ba26967.pj@sgi.com","threadId":"11","inReplyTo":"4261DDBC.3050706@dwheeler.com","subject":"Re: Yet another base64 patch","fromName":"Paul Jackson","fromEmail":"pj@sgi.com","sentAt":"2005-04-17T04:05:13Z","receivedAt":"2005-04-17T04:05:13Z","isPatch":false,"sender":{"key":"pj@sgi.com","avatar":null},"body":"David wrote:\n> It's a trade-off, I know.\n\nSo where do you recommend we make that trade-off?\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"430","messageId":"Pine.LNX.4.62.0504162107040.22904@qynat.qvtvafvgr.pbz","threadId":"11","inReplyTo":"425F1394.5020709@zytor.com","subject":"Re: Yet another base64 patch","fromName":"David Lang","fromEmail":"dlang@digitalinsight.com","sentAt":"2005-04-17T04:10:09Z","receivedAt":"2005-04-17T04:10:09Z","isPatch":false,"sender":{"key":"dlang@digitalinsight.com","avatar":null},"body":"On Thu, 14 Apr 2005, H. Peter Anvin wrote:\n\n> Linus Torvalds wrote:\n>> \n>> Even something as simple as \"ls -l\" has been known to have O(n**2) \n>> behaviour for big directories.\n>> \n>\n> For filesystems with linear directories, sure.  For sane filesystems, it \n> should have O(n log n).\n\nnote that default configs of ext2 and ext3 don't qualify as sane \nfilesystems by this definition.\n\next3 does have an extention that you can enable to have it hash the \ndirectory access, but even if you enable that on a filesystem you aren't \nguaranteed that it will be active (if the directory existed before it was \nturned on, or has been accessed by a kernel that didn't understand the \nextention then the htree functionality won't be used until you manually \ntell the system to generate the tree)\n\nDavid Lang\n\n-- \nThere are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.\n  -- C.A.R. Hoare\n"},{"id":"451","messageId":"42620452.4080809@dwheeler.com","threadId":"11","inReplyTo":"20050416210513.1ba26967.pj@sgi.com","subject":"Re: Yet another base64 patch","fromName":"David A. Wheeler","fromEmail":"dwheeler@dwheeler.com","sentAt":"2005-04-17T06:38:10Z","receivedAt":"2005-04-17T06:38:10Z","isPatch":false,"sender":{"key":"dwheeler@dwheeler.com","avatar":"https://avatars.githubusercontent.com/u/813150?v=4"},"body":"Paul Jackson wrote:\n> David wrote:\n> \n>>It's a trade-off, I know.\n> \n> \n> So where do you recommend we make that trade-off?\n\nI'd look at some of the more constraining, yet still\ncommon cases, and make sure it worked reasonably\nwell without requiring magic. My list would be:\next2, ext3, NFS, and Windows' NTFS (stupid short filenames,\ncase-insensitive/case-preserving).  Samba shouldn't be\nmore constraining than NTFS, and I would expect ReiserFS\nwouldn't be a constraining case.  Bonus points if the\nnames lengths are inside POSIX guarantees, but I bet the\nPOSIX limits are so tiny as to be laughable.  Bonus points for\nCD-ROM format with the Rock Ridge extensions (I _think_ DVDs\nand later use that format too, yes?), though if that\ndidn't work tar files are an easy workaround. Imagine a full\nLinux kernel source repository, for 30+ (pick a number) years..\ncan the filesystems handle the number of objects in those cases?\nIf it works, your infrastructure should be sufficiently\nportable to \"just work\" on others too.\n\nAnyway, my two cents.\n\n--- David A. Wheeler\n"},{"id":"458","messageId":"20050417011615.3e7dfb29.pj@sgi.com","threadId":"11","inReplyTo":"42620452.4080809@dwheeler.com","subject":"Re: Yet another base64 patch","fromName":"Paul Jackson","fromEmail":"pj@sgi.com","sentAt":"2005-04-17T08:16:15Z","receivedAt":"2005-04-17T08:16:15Z","isPatch":false,"sender":{"key":"pj@sgi.com","avatar":null},"body":"David wrote:\n> My list would be:\n> ext2, ext3, NFS, and Windows' NTFS (stupid short filenames,\n> case-insensitive/case-preserving).\n\nI'm no mind reader, but I'd bet a pretty penny that what you have in\nmind and what Linus has in mind have no overlaps in their solution sets.\n\nHappy coding ...\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@engr.sgi.com> 1.650.933.1373, 1.925.600.0401\n"},{"id":"473","messageId":"Pine.LNX.4.21.0504171018410.30848-100000@iabervon.org","threadId":"11","inReplyTo":"20050416210513.1ba26967.pj@sgi.com","subject":"Re: Yet another base64 patch","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-17T14:30:37Z","receivedAt":"2005-04-17T14:30:37Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sat, 16 Apr 2005, Paul Jackson wrote:\n\n> David wrote:\n> > It's a trade-off, I know.\n> \n> So where do you recommend we make that trade-off?\n\nSo why do we have to be consistant? It seems like we need a standard\nformat for these reasons:\n\n - We use rsync to interact with remote repositories, and rsync won't\n   understand if they aren't organized the same way. But I'm working on\n   having everything go through git-specific code, which could understand\n   different layouts.\n\n - Everything that shares a local repository needs to understand the\n   format of that repository. But the filesystem constraints on the local\n   repository will be the same regardless of who is looking, so they'd all\n   expect the same format anyway.\n\nSo my idea is, once we're using git-smart transfer code (which can verify\nobjects, etc.), add support for different implementations of \nsha1_file_name suitable for different filesystems, and vary based either\non a compile-time option or on a setting stored in the objects\ndirectory. The only thing that matters is that repositories on\nnon-special web servers have a standard format, because they'll be serving\nobjects by URL, not by sha1.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"494","messageId":"42628EFD.3030509@dwheeler.com","threadId":"11","inReplyTo":"Pine.LNX.4.21.0504171018410.30848-100000@iabervon.org","subject":"Re: Yet another base64 patch","fromName":"David A. Wheeler","fromEmail":"dwheeler@dwheeler.com","sentAt":"2005-04-17T16:29:49Z","receivedAt":"2005-04-17T16:29:49Z","isPatch":false,"sender":{"key":"dwheeler@dwheeler.com","avatar":"https://avatars.githubusercontent.com/u/813150?v=4"},"body":"I wrote:\n>>>It's a trade-off, I know.\n\nPaul Jackson replied:\n>>So where do you recommend we make that trade-off?\n\nDaniel Barkalow wrote:\n> So why do we have to be consistant? It seems like we need a standard\n> format for these reasons:\n> \n>  - We use rsync to interact with remote repositories, and rsync won't\n>    understand if they aren't organized the same way. But I'm working on\n>    having everything go through git-specific code, which could understand\n>    different layouts.\n> \n>  - Everything that shares a local repository needs to understand the\n>    format of that repository. But the filesystem constraints on the local\n>    repository will be the same regardless of who is looking, so they'd all\n>    expect the same format anyway.\n> \n> So my idea is, once we're using git-smart transfer code (which can verify\n> objects, etc.), add support for different implementations of \n> sha1_file_name suitable for different filesystems, and vary based either\n> on a compile-time option or on a setting stored in the objects\n> directory.\n\nI think that's the perfect answer: make it a setting stored\nin the objects directory (presumably set during\ninitialization of the directory), and handled automagically\nby the tools.  I recommend handling them NOT be a compile-time option,\nso that the same set of tools works everywhere automatically\n(who wants to recompile tools just to work on a different file layout?).\n\n\n> The only thing that matters is that repositories on\n> non-special web servers have a standard format, because they'll be serving\n> objects by URL, not by sha1.\n\nIf the \"layout info\" is stored in a standard location for a\ngiven repository, then the rest doesn't matter. The library would just\ndownload that, then know how to find the rest.\n\n--- David A. Wheeler\n"},{"id":"508","messageId":"4262A238.3050207@dwheeler.com","threadId":"11","inReplyTo":"20050417011615.3e7dfb29.pj@sgi.com","subject":"Re: Yet another base64 patch","fromName":"David A. Wheeler","fromEmail":"dwheeler@dwheeler.com","sentAt":"2005-04-17T17:51:52Z","receivedAt":"2005-04-17T17:51:52Z","isPatch":false,"sender":{"key":"dwheeler@dwheeler.com","avatar":"https://avatars.githubusercontent.com/u/813150?v=4"},"body":"Paul Jackson wrote:\n> David wrote:\n> \n>>My list would be:\n>>ext2, ext3, NFS, and Windows' NTFS (stupid short filenames,\n>>case-insensitive/case-preserving).\n> \n> \n> I'm no mind reader, but I'd bet a pretty penny that what you have in\n> mind and what Linus has in mind have no overlaps in their solution sets.\n\nSadly, I lack the mind reading ability as well.\n\nOur goals are, I suspect, somewhat different.\nLinus wants to build a tool that meets his specific needs\n(managing kernel development), and he has particular requirements\n(such as fast simple merging when working at large scales).\nIn contrast, I'm hoping for a more\ngeneral OSS/FS SCM tool that many others can use as well.\n\nBut I think there's heavy overlap in the solution space.\nThe Linux kernel project is, to my knowledge, the largest\nproject using a truly distributed SCM process.\nAnyone else who is considering a distributed SCM process\nwould at _least_ want to think about how the Linux kernel\nproject works, and if they're doing so, they\nmight also want to reuse the development tools.\n\nI'm just taking a peek, and\nlooking for situations where a design decision is irrelevant\nfor his purposes, but a particular direction would be of\nparticular help to other projects.  I'm more worried about the\nstorage format; if the code doesn't support some particular\nfeature but it could be added later without great pain, no big deal.\nIf something would imply a complete rewrite, that's undesirable.\n\n--- David A. Wheeler\n"},{"id":"513","messageId":"20050417181935.GD1461@pasky.ji.cz","threadId":"11","inReplyTo":"42620452.4080809@dwheeler.com","subject":"Re: Yet another base64 patch","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-17T18:19:36Z","receivedAt":"2005-04-17T18:19:36Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sun, Apr 17, 2005 at 08:38:10AM CEST, I got a letter\nwhere \"David A. Wheeler\" <dwheeler@dwheeler.com> told me that...\n> I'd look at some of the more constraining, yet still\n> common cases, and make sure it worked reasonably\n> well without requiring magic. My list would be:\n> ext2, ext3, NFS, and Windows' NTFS (stupid short filenames,\n> case-insensitive/case-preserving).  Samba shouldn't be\n> more constraining than NTFS, and I would expect ReiserFS\n> wouldn't be a constraining case.  Bonus points if the\n> names lengths are inside POSIX guarantees, but I bet the\n> POSIX limits are so tiny as to be laughable.  Bonus points for\n> CD-ROM format with the Rock Ridge extensions (I _think_ DVDs\n> and later use that format too, yes?), though if that\n> didn't work tar files are an easy workaround. Imagine a full\n> Linux kernel source repository, for 30+ (pick a number) years..\n> can the filesystems handle the number of objects in those cases?\n> If it works, your infrastructure should be sufficiently\n> portable to \"just work\" on others too.\n\nI personally don't mind getting it work on more places, if it doesn't\nmake git work (measurably) worse on modern Linux systems, the code will\nnot go to hell, you tell me what needs to be done and preferably give me\nthe patches. ;-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"628","messageId":"426341FC.7090600@dwheeler.com","threadId":"11","inReplyTo":"20050417181935.GD1461@pasky.ji.cz","subject":"Re: Yet another base64 patch","fromName":"David A. Wheeler","fromEmail":"dwheeler@dwheeler.com","sentAt":"2005-04-18T05:13:32Z","receivedAt":"2005-04-18T05:13:32Z","isPatch":false,"sender":{"key":"dwheeler@dwheeler.com","avatar":"https://avatars.githubusercontent.com/u/813150?v=4"},"body":"I said:\n>>I'd look at some of the more constraining, yet still\n>>common cases, and make sure it worked reasonably\n>>well without requiring magic. My list would be:\n>>ext2, ext3, NFS, and Windows' NTFS (stupid short filenames,\n>>case-insensitive/case-preserving).\n\nPetr Baudis replied:\n> I personally don't mind getting it work on more places, if it doesn't\n> make git work (measurably) worse on modern Linux systems, the code will\n> not go to hell, you tell me what needs to be done and preferably give me\n> the patches. ;-)\n\nOkay, that's great.\n\nThe one potential issue I know of (after trying to read from the\nfirehose^Wlist archives) is that some are worried about poor filesystems\nwhen there are a large number of objects in an object directory.\n\nAfter doing some calculations, it seems to me that perhaps this\nisn't really such a big deal, if there's a top directory such as\nthe 16-bit (2-char) top directory currently in git-pasky.\nRemoving the top directory would improve performance for the better\nfilesystems, but would be an absolute KILLER to poorer systems, so\nI'd keep the 2**8 top directory just as it is in git-pasky.\nIt's a compromise that means people can ease into git, and then\nswitch when their projects grow to large sizes.\nMy calculations are below, but I could be mistaken; let me\nknow if I'm all wet.\n\nDoes anyone know of any other issues in how git data is stored that\nmight cause problems for some situations?  Windows' case-insensitive/\ncase-preserving model for NTFS and vfat32 seems to be enough\n(since the case is preserved) so that the format should work,\nand you can just demand that\nspecial git files use Unix formats (\"/\" as dir separator,\nUnix end-of-lines).  The implementation currently would need\nchange to work easily on Windows (dealing with binary opens at least,\nand probably rewriting the shell programs for those unwilling to\ninstall Cygwin), but those can be done later if desired\nwithout interfering with the interface formats.\n\n\n========================= Details =========================\n\nBasically, I'd like \"git\" to work on:\n(1) nearly ANY system on small-to-medium projects,\n     even if their filesystems do linear searches in directories,\n     over a lengthy time.  Ideally possibly (though poorly)\n     on larger systems.\n(2) work well on large projects (e.g., kernel) on _common_\n     development platforms (ext2, ext3, NTFS, NFS).\n\nIt all depends on what you're optimizing for; but humor me\nif those were your requirements...\n\nCase 1:\nThe top (2-char) directory appears likely to make small projects\nperform okay, and large projects possible, on stupid filesystems.\nThe one level extra directory is actually not a bad compromise\nto make things \"just work\" on just about anything for smaller scales.\n* git-paskey (a tiny project) has ~2K objects in 2weeks; at that pace,\n4Kobjects/month for 10 years, you'd have 480K objects.\nThat's absurd for even tiny projects, and it's unlikely that\na participant in a tiny project would be willing to change\nfilesystems just to participate.  But then if you\ndivide it among 256 directories = 1875 files/directory average.\nLinear search is undesirable (about 1000 entry checks on\naverage to find each entry), but it's nowhere near the\n2^16 dir entries that made people afraid.\nSwitching to a 2^12 top directory, you have an average of 117 entries\nin each subdir (and 4096 entries at the top), yielding\nan average of (117+4096)/2 = 2106 entry checks to find an entry.\n* I estimated also for the big end, using the Linux kernel;\nI guesstimated 36,000 objects/month for the kernel**. Over 10 years that\naccumulates 4,320,000 objects, completely insane for a flat file\non a stupid filesystem. If it has a one-level 256dir directory, that's\n16875 objects/directory.  Now THAT'S painful,\nthough nowhere near the 2^16 limit most quoted as bad.\n* For 10K objects/month, and a top dir of 2**8, you have 1,200,000\nobjects; each dir has 4680 entries (average lookup: 2468 entries).\nDividing into 2**12 has 292/directory, average lookup: 2194.\n\nOn 2**12 vs. 2**8, it's not clear-cut. 2**8 works best for small\nprojects, 2**12 for larger.  My guess is that stupid filesystems\nwill tend to be used primarily only on small projects, so 2**8 might\nbe the better choice but that's debatable.\n\nCase 2:\nThankfully adequate systems are finally more common, and they're\ncommon enough that for really large projects (kernel) it seems\nreasonable to demand such filesystems.\nExt2 & ext3 have had htree for a while now, and it's enabled by\ndefault on at least Fedora Core 3.  If it's off, just do:\n  tune2fs -O dir_index /dev/hdFOO; e2fsck -fD /dev/hdFOO\nThis stuff has been around so long that it should just be\na trivial command by any developer today.\nReiserFS has hashing too.  Windows' NTFS does\ntree-balancing (it appears not as good as the hashing htree\nsystem of ext2/ext3, but it should work tolerably since it's no\nlonger a linear search).  One useful factoid: For good NTFS\nperformance with git on large projects,\nyou should disable short name generation on the big directories\n(Microsoft recommends this when >300,000 names are in one dir).\nNTFS (and VFAT32) allow filenames up to 255 chars, and\nfilepaths up to 260 chars, so that seems okay.\nI was primarily concerned about NTFS, and that seems to have\nthe necessities.  This info should in some FAQ or\ndocumentation (\"Using git for large projects\").\n\nIt _seems_ to me that the NFS implementations are likely to\ndo similar things, but I don't know.  And I've not tested\nanything on real systems, which is the real test.\nAnyone know more about the limits of the NFS implementations?\n\nMore directory levels could be created to make\nstupid filesystems happier, but that interferes with smart filesystems.\nYou could try to make filesystem layout a per-user issue,\nbut that makes using rsync more complicated.\nA link farm could be created, though those are a pain to maintain.\nIt DOES turn out there are many alternatives if necesary, e.g.,\nconfigurations per object database, or automatically \"fixing\"\nthings for a local configuration as data comes in or out,\nthough if you can avoid that it'd be better.\n\n\n** Looking at \"linux-2.4.0-to-2.6.12-rc2-patchset\", I count\n28237 patches; \"RCS file:\" occurs 188119 times & I'll claim\nthat that approximates the number of different file objects\nIF there were no intermediate files.  If on average there are\n5 versions of a file before it gets into the mainline,\nand 3 commits before the final mainline patch, I get\napproximately this many objects in a \"real\" object db:\n  (28237*(3+1) trees) *2 (if #commits==#trees) +\n  (188119*(5+1) file objs))\n= 1,354,610 objects from 2002/02/05 to 2005/04/04\n= about 36,000 objects/month.\n\n\nAm I missing anything?\n\n--- David A. Wheeler\n"},{"id":"631","messageId":"42635256.7020701@zytor.com","threadId":"11","inReplyTo":"Pine.LNX.4.62.0504162107040.22904@qynat.qvtvafvgr.pbz","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-18T06:23:18Z","receivedAt":"2005-04-18T06:23:18Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"David Lang wrote:\n> \n> note that default configs of ext2 and ext3 don't qualify as sane \n> filesystems by this definition.\n\nNot using dir_index *IS* insane.\n\n\t-hpa\n"},{"id":"632","messageId":"42635388.90207@zytor.com","threadId":"11","inReplyTo":"20050415165532.05ed5dc4.paul@permanentmail.com","subject":"Re: Yet another base64 patch","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-18T06:28:24Z","receivedAt":"2005-04-18T06:28:24Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Paul Dickson wrote:\n> \n> Since 160-bits does not go into base64 evenly anyways, what happens if\n> you use 2^10 instead of 2^12 for the subdir names?  That will be 1/4 the\n> directories of the base64 given above.\n> \n\nI was going to try one-character subdirs, so 2^6, but I haven't had a \nchance to do that since I'm at LCA.\n\nAnyway, I'm starting to suspect it's too late to change the format, \nespecially since Linus seems highly disinclined.\n\n\t-hpa\n"},{"id":"655","messageId":"4263AF41.1070806@qualitycode.com","threadId":"11","inReplyTo":"426341FC.7090600@dwheeler.com","subject":"Re: Yet another base64 patch","fromName":"Kevin Smith","fromEmail":"yarcs@qualitycode.com","sentAt":"2005-04-18T12:59:45Z","receivedAt":"2005-04-18T12:59:45Z","isPatch":false,"sender":{"key":"yarcs@qualitycode.com","avatar":null},"body":"David A. Wheeler wrote:\n> Does anyone know of any other issues in how git data is stored that\n> might cause problems for some situations?  Windows' case-insensitive/\n> case-preserving model for NTFS and vfat32 seems to be enough\n> (since the case is preserved) so that the format should work,\n\nIf git is retaining hex naming, and not moving to base64, then I don't\nthink what I am about to say is relevant. However, if base64 file naming\nis still being considered, then vfat32 compatibility may be a concern\n(I'm not sure about NTFS). Although it is case-preserving, it actually\nconsiders both cases as being the same name. So AaA would overwrite aAa.\n\nIf I'm doing the math right, we would effectively be ignoring roughly\none out of 6 base64 bits. This would reduce the collision avoidance\ncapability of SHA-1 (on vfat32) from 160 bits to about 133 bits. Still\nstrong, and probably acceptable, but worth noting.\n\nI'll take this opportunity to support David's position that it would be\nfantastic if git could end up being valuable for a wide range of\nprojects, rather than just the kernel. I also fully understand that the\nkernel is the primary target, but when there are opportunities to make\nthe data structures more generally useful without causing problems for\nthe kernel project, I hope they are taken.\n\nThanks,\n\nKevin\n"},{"id":"668","messageId":"E1DNZK9-0003c7-3r@fenris.runbox.com","threadId":"11","inReplyTo":"4263AF41.1070806@qualitycode.com","subject":"Re: Yet another base64 patch","fromName":"David A. Wheeler","fromEmail":"dwheeler@dwheeler.com","sentAt":"2005-04-18T16:42:13Z","receivedAt":"2005-04-18T16:42:13Z","isPatch":false,"sender":{"key":"dwheeler@dwheeler.com","avatar":"https://avatars.githubusercontent.com/u/813150?v=4"},"body":"I asked:\n> > Does anyone know of any other issues in how git data is stored that\n> > might cause problems for some situations? ...\n\nKevin said:\n> If git is retaining hex naming, and not moving to base64, then I don't\n> think what I am about to say is relevant. However, if base64 file naming\n> is still being considered, then vfat32 compatibility may be a concern\n> (I'm not sure about NTFS).\n\nI can't speak for the git developers. However, I think the current\nnaming scheme for the object database as used in git-pasky\nis actually a very good one and should be left as-is\n(SHA-1 hex values, directory of 2-char prefixes,\nfilenames with the rest of the value).\n\nAs far as I can tell from various calculations (& supported by the\nperformance measurements done by others), the hex values\nwith one level of directory turns out to work pretty well!\nIt's easily understood, works with non-massive projects on stupid\nfilesystems, and it has good performance on good filesystems\neven with massive projects with huge histories.  You could\ntune it further, but a single approach that works \"everywhere\"\nis a whole lot simpler.  So I'd recommend keeping that\napproach.\n\nAs far as base64/32 vs. hex names, I think there\nare many reasons to stay with the hex names.\nUsing hex names is a good idea for the simple reason that\nnormally SHA-1 hashes are presented as hex values;\nyou'll work WITH instead of AGAINST other tools, and\nhumans who deal with this stuff will \"see what they expect\".\nIt takes a few more characters, but not many, and it's not\nlike base64 is any more comprehensible to humans.\nAnd the fact that hex values don't allow \"all\" legal values\nmeans that some errors are trivially detectable.\n\nYou're right, base64 eliminates many bits of differentiation,\nand in a very non-obvious way (I _hate_ weird surprises like\nthat, they cause lots of trouble).  I think there's another\nproblem too that's more insideous. Although the _filesystem_\nis case-preserving, I suspect some _tools_ on Windows don't take\ncare to preserve case.  If that's so, it'd be easily possible for a\nWindows user to use some tools that screw up a Unix/Linux user\nonce they were imported, causing all sorts of \"extraneous\" files &\nfiles that mysteriously disappeared (they were only accessible\nfrom Windows). Ugh.\nThis can even happen on Unix/Linux systems if they use\na fileserver with NTFS semantics. In contrast,\nif a hex value has its case changed, it's easy to fix locally.\n\nBy choosing the more traditional hex representation, you\neliminate lots of problems, and it's easier to explain too.\n\nKevin added:\n> I'll take this opportunity to support David's position that it would be\n> fantastic if git could end up being valuable for a wide range of\n> projects, rather than just the kernel. I also fully understand that the\n> kernel is the primary target, but when there are opportunities to make\n> the data structures more generally useful without causing problems for\n> the kernel project, I hope they are taken.\n\nThanks for the vote of confidence!\n\n--- David A. Wheeler\n"}]}