{"thread":{"id":"48","subject":"full kernel history, in patchset format","startedAt":"2005-04-16T13:15:28Z","lastAt":"2005-04-19T08:33:21Z","messageCount":42,"participants":["Ingo Molnar","David Mansfield","Francois Romieu","Linus Torvalds","Petr Baudis","Thomas Gleixner","Junio C Hamano","Mike Taht","Christopher Li","Jan-Benedict Glaw","Daniel Barkalow","David Lang","David Woodhouse","Catalin Marinas"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"299","messageId":"20050416131528.GB19908@elte.hu","threadId":"48","inReplyTo":null,"subject":"full kernel history, in patchset format","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-16T13:15:28Z","receivedAt":"2005-04-16T13:15:28Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\ni've converted the Linux kernel CVS tree into 'flat patchset' format, \nwhich gave a series of 28237 separate patches. (Each patch represents a \nchangeset, in the order they were applied. I've used the cvsps utility.)\n\nthe history data starts at 2.4.0 and ends at 2.6.12-rc2. I've included a \nscript that will apply all the patches in order and will create a \npristine 2.6.12-rc2 tree.\n\nit needed many hours to finish, on a very fast server with tons of RAM, \nand it also needed a fair amount of manual work to extract it and to \nmake it usable, so i guessed others might want to use the end result as \nwell, to try and generate large GIT repositories from them (or to run \nanalysis over the patches, etc.).\n\nthe patches contain all the existing metadata, dates, log messages and \nrevision history. (What i think is missing is the BK tree merge \ninformation, but i'm not sure we want/need to convert them to GIT.)\n\nit's a 136 MB tarball, which can be downloaded from:\n\n   http://kernel.org/pub/linux/kernel/people/mingo/Linux-2.6-patchset/\n\nthe ./generate-2.6.12-rc2 script generates the 2.6.12-rc2 tree into \nlinux/, from scratch. (No pre-existing kernel is needed, as 2.patch \ngenerates the full 2.4.0 kernel tree.) The patching takes a couple of \nminutes to finish, on a fast box.\n\nbelow i've attached a sample patch from the series.\n\nnote: i kept the patches the cvsps utility generated as-is, to have a \nverifiable base to work on. There were a very small amount of deltas \nmissed (about a dozen), probably resulting from CVS related errors, \nthese are included in the diff-CVS-to-real patch. Also, the patch format \ncannot create the Documentation/logo.gif file, so the script does this \ntoo - just to be able to generate a complete 2.6.12-rc2 tree that is \nbyte-for-byte identical to the real thing.\n\n\tIngo\n\n---------------------\nPatchSet 1234 \nDate: 2002/04/11 18:29:07\nAuthor: viro\nBranch: HEAD\nTag: (none) \nLog:\n[PATCH] crapectomy in include/linux/nfsd/syscall.h\n\nRemoves an atavism in declaration of sys_nfsservctl() - sorry, I should've\nremove that junk when cond_syscall() thing was done.\n\nBKrev: 3cb5c7e3phTYgiz1YLsjQ_McTo9pOQ\n\nMembers: \n\tChangeSet:1.1234->1.1235 \n\tinclude/linux/nfsd/syscall.h:1.3->1.4 \n\nIndex: linux/include/linux/nfsd/syscall.h\n===================================================================\nRCS file: /home/mingo/linux-CVS/linux/include/linux/nfsd/syscall.h,v\nretrieving revision 1.3\nretrieving revision 1.4\ndiff -u -r1.3 -r1.4\n--- linux/include/linux/nfsd/syscall.h\t15 Mar 2002 23:06:06 -0000\t1.3\n+++ linux/include/linux/nfsd/syscall.h\t11 Apr 2002 17:29:07 -0000\t1.4\n@@ -132,11 +132,7 @@\n /*\n  * Kernel syscall implementation.\n  */\n-#if defined(CONFIG_NFSD) || defined(CONFIG_NFSD_MODULE)\n extern asmlinkage long\tsys_nfsservctl(int, struct nfsctl_arg *, void *);\n-#else\n-#define sys_nfsservctl\t\tsys_ni_syscall\n-#endif\n extern int\t\texp_addclient(struct nfsctl_client *ncp);\n extern int\t\texp_delclient(struct nfsctl_client *ncp);\n extern int\t\texp_export(struct nfsctl_export *nxp);\n"},{"id":"304","messageId":"20050416133513.GA21678@elte.hu","threadId":"48","inReplyTo":"20050416131528.GB19908@elte.hu","subject":"Re: full kernel history, in patchset format","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-16T13:35:13Z","receivedAt":"2005-04-16T13:35:13Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Ingo Molnar <mingo@elte.hu> wrote:\n\n> the patches contain all the existing metadata, dates, log messages and \n> revision history. (What i think is missing is the BK tree merge \n> information, but i'm not sure we want/need to convert them to GIT.)\n\nauthor names are abbreviated, e.g. 'viro' instead of \nviro@parcelfarce.linux.theplanet.co.uk, and no committer information is \nincluded (albeit commiter ought to be Linus in most cases). These are \nlimitations of the BK->CVS gateway i think.\n\n\tIngo\n"},{"id":"308","messageId":"42612771.5000705@cobite.com","threadId":"48","inReplyTo":"20050416133513.GA21678@elte.hu","subject":"Re: full kernel history, in patchset format","fromName":"David Mansfield","fromEmail":"david@cobite.com","sentAt":"2005-04-16T14:55:45Z","receivedAt":"2005-04-16T14:55:45Z","isPatch":false,"sender":{"key":"david@cobite.com","avatar":null},"body":"Ingo Molnar wrote:\n> * Ingo Molnar <mingo@elte.hu> wrote:\n> \n> \n>>the patches contain all the existing metadata, dates, log messages and \n>>revision history. (What i think is missing is the BK tree merge \n>>information, but i'm not sure we want/need to convert them to GIT.)\n> \n> \n> author names are abbreviated, e.g. 'viro' instead of \n> viro@parcelfarce.linux.theplanet.co.uk, and no committer information is \n> included (albeit commiter ought to be Linus in most cases). These are \n> limitations of the BK->CVS gateway i think.\n> \n\nGlad to hear cvsps made it through!  I'm curious what the manual fixups \nrequired were, except for the binary file issue (logo.gif).\n\nAs to the actual email addresses, for more recent patches, the \nSigned-off should help.  For earlier ones, isn't their some script which \n'knows' a bunch of canonical author->email mappings? (the shortlog \nscript or something)?\n\nIs the full committer email address actually in the changeset in BK?  If \nso, given that we have the unique id (immutable I believe) of the \nchangset, could it be extracted directly from BK?\n\nDavid\n\n"},{"id":"312","messageId":"20050416150816.GA4943@electric-eye.fr.zoreil.com","threadId":"48","inReplyTo":"20050416131528.GB19908@elte.hu","subject":"Re: full kernel history, in patchset format","fromName":"Francois Romieu","fromEmail":"romieu@fr.zoreil.com","sentAt":"2005-04-16T15:08:17Z","receivedAt":"2005-04-16T15:08:17Z","isPatch":false,"sender":{"key":"romieu@fr.zoreil.com","avatar":null},"body":"Ingo Molnar <mingo@elte.hu> :\n[...]\n> the history data starts at 2.4.0 and ends at 2.6.12-rc2. I've included a \n> script that will apply all the patches in order and will create a \n> pristine 2.6.12-rc2 tree.\n\n127 weeks of bk-commit mail for the 2.6 branch alone since october 2002\nprovides more than 44000 messages here. The figures are surprisingly\ndifferent.\n\n> it needed many hours to finish, on a very fast server with tons of RAM, \n> and it also needed a fair amount of manual work to extract it and to \n> make it usable, so i guessed others might want to use the end result as \n> well, to try and generate large GIT repositories from them (or to run \n> analysis over the patches, etc.).\n\nHas anyone already compared the (split/digested) content of the ChangeLog\nfile with the commit messages ? It raises the interesting question of\ninserting the merge messages/patches in the sequence at the right place\nbut I'd like to know if someone met other issues.\n\n--\nUeimor\n"},{"id":"340","messageId":"20050416154441.GA30392@64m.dyndns.org","threadId":"48","inReplyTo":"20050416174327.GG19099@pasky.ji.cz","subject":"Re: Re: full kernel history, in patchset format","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-16T15:44:41Z","receivedAt":"2005-04-16T15:44:41Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"On Sat, Apr 16, 2005 at 07:43:27PM +0200, Petr Baudis wrote:\n> Dear diary, on Sat, Apr 16, 2005 at 07:04:31PM CEST, I got a letter\n> where Linus Torvalds <torvalds@osdl.org> told me that...\n> > So I'd _almost_ suggest just starting from a clean slate after all.  \n> > Keeping the old history around, of course, but not necessarily putting it\n> > into git now. It would just force everybody who is getting used to git in \n> > the first place to work with a 3GB archive from day one, rather than \n> > getting into it a bit more gradually.\n> > \n> > Comments?\n> \n> FWIW, it looks pretty reasonable to me. Perhaps we should have a\n> separate GIT repository with the previous history though, and in the\n> first new commit the parent could point to the last commit from the\n> other repository.\n> \n> Just if it isn't too much work, though. :-)\n\nI think we can make the git using stackable repository. When it fail\nto find an object, it will try it's to read from parent repository.\nIt is useful to slice the history.\n\nI can have local repository that all the new object create by me will\nstore in my tree instead of the official one. Clean up the object in the\nmy local tree will be much easier it only need to work on a much smaller\nrepository. If all my change is merge to official tree, I just simply\nempty my local repository.\n\nAbout the kernel git repository. I think it is much easier just put\nthem in one tree.  So I don't need to worry about \"if I need to see\npre 2.6.12, I need to do this\". And the full repository  need to\nstore in the server some where any way.\n\nHowever I totally agree that people should not deal with unnecessary the history\nwhen they start using the git tools. We should just make the tools\nby default don't download all the histories. Only get it when user specific \nask for it.\n\nWhy 2.6.12-rc2? When kernel grows to 2.6.15, a new user might not even need\npre 2.6.13 most of the time. If we make it very easier for people to get\nhistory if they need, it will make them less motivate to store unnecessary\nhistory locally (just in case I need it).\n\nI think we should not advise using rsync to sync the whole git tree as\nway to get update. We need to get use to only have a slice of the history\nand get more if we needed.\nThe server should should provide some small metadata file like the\nthe rev-tool cache, so the SCM tools can download it to figure out what file\nis needed to download to get to certain revision. Instead of download the\nwhole repository to figure out what is new.\n\nWe can even slice that metadata information to smaller pieces base on major release point.\n\nChris\n \n"},{"id":"349","messageId":"20050416162632.GA3309@64m.dyndns.org","threadId":"48","inReplyTo":"7v8y3i7cn9.fsf@assigned-by-dhcp.cox.net","subject":"Re: full kernel history, in patchset format","fromName":"Christopher Li","fromEmail":"git@chrisli.org","sentAt":"2005-04-16T16:26:32Z","receivedAt":"2005-04-16T16:26:32Z","isPatch":false,"sender":{"key":"git@chrisli.org","avatar":null},"body":"We can just have a baseline file contain all the commit objects.\nThen have the git \"download on demand\". The problem with diff\npackage  is that I it is harder to merge with more than one diff.\n\nI bet 90% of the time people sync to the repository head first\nwant to check out the last bits. And maybe reading some change\nlog to see what is changed.\n\nSo having all the commit object, the user will able to see\nwhat is change and which version he we like to check out.\n\nThen he can issue a command \"download me all the objects is needed\nfor checkout the this commit\". Download of demand should be\neven better.\n\nChris\n\n\nOn Sat, Apr 16, 2005 at 12:19:22PM -0700, Junio C Hamano wrote:\n> >>>>> \"MT\" == Mike Taht <mike.taht@timesys.com> writes:\n> \n> MT> alternatively, \"git-archive-torrent\" to create a list of files for a\n> MT> bittorrent feed....\n> \n> That is certainly good for establishing the baseline, but you\n> still need to leverage the inherent delta-compressibility\n> between related blobs/trees by also doing something like what I\n> described as \"diff package\", don't you?\n> \n> \n> \n> \n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"327","messageId":"Pine.LNX.4.58.0504160953310.7211@ppc970.osdl.org","threadId":"48","inReplyTo":"20050416131528.GB19908@elte.hu","subject":"Re: full kernel history, in patchset format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T17:04:31Z","receivedAt":"2005-04-16T17:04:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Ingo Molnar wrote:\n> \n> i've converted the Linux kernel CVS tree into 'flat patchset' format, \n> which gave a series of 28237 separate patches. (Each patch represents a \n> changeset, in the order they were applied. I've used the cvsps utility.)\n> \n> the history data starts at 2.4.0 and ends at 2.6.12-rc2. I've included a \n> script that will apply all the patches in order and will create a \n> pristine 2.6.12-rc2 tree.\n\nHey, that's great. I got the CVS repo too, and I was looking at it, but \nthe more I looked at it, the more I felt that the main reason I want to \nimport it into git ends up being to validate that my size estimates are at \nall realistic.\n\nI see that Thomas Gleixner seems to have done that already, and come to a \nfigure of 3.2GB for the last three years, which I'm very happy with, \nmainly because it seems to match my estimates to a tee. Which means that I \njust feel that much more confident about git actually being able to handle \nthe kernel long-term, and not just as a stop-gap measure.\n\nBut I wonder if we actually want to actually populate the whole history.. \nNow that my size estimates have been verified, I have little actual real \nreason to put the history into git. There are no visualization tools done \nfor git yet, and no helpers to actually find problems, and by the time \nthere will be, we'll have new history.\n\nSo I'd _almost_ suggest just starting from a clean slate after all.  \nKeeping the old history around, of course, but not necessarily putting it\ninto git now. It would just force everybody who is getting used to git in \nthe first place to work with a 3GB archive from day one, rather than \ngetting into it a bit more gradually.\n\nWhat do people think? I'm not so much worried about the data itself: the\ngit architecture is _so_ damn simple that now that the size estimate has\nbeen confirmed, that I don't think it would be a problem per se to put\n3.2GB into the archive. But it will bog down \"rsync\" horribly, so it will\nactually hurt synchronization untill somebody writes the rev-tree-like\nstuff to communicate changes more efficiently..\n\nIOW, it smells to me like we don't have the infrastructure to really work \nwith 3GB archives, and that if we start from scratch (2.6.12-rc2), we can \nbuild up the infrastructure in parallell with starting to really need it.\n\nBut it's _great_ to have the history in this format, especially since \nlooking at CVS just reminded me how much I hated it.\n\nComments?\n\n\t\tLinus\n"},{"id":"330","messageId":"20050416174327.GG19099@pasky.ji.cz","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504160953310.7211@ppc970.osdl.org","subject":"Re: Re: full kernel history, in patchset format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-16T17:43:27Z","receivedAt":"2005-04-16T17:43:27Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Apr 16, 2005 at 07:04:31PM CEST, I got a letter\nwhere Linus Torvalds <torvalds@osdl.org> told me that...\n> So I'd _almost_ suggest just starting from a clean slate after all.  \n> Keeping the old history around, of course, but not necessarily putting it\n> into git now. It would just force everybody who is getting used to git in \n> the first place to work with a 3GB archive from day one, rather than \n> getting into it a bit more gradually.\n> \n> Comments?\n\nFWIW, it looks pretty reasonable to me. Perhaps we should have a\nseparate GIT repository with the previous history though, and in the\nfirst new commit the parent could point to the last commit from the\nother repository.\n\nJust if it isn't too much work, though. :-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"332","messageId":"7vmzry7ev5.fsf@assigned-by-dhcp.cox.net","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504160953310.7211@ppc970.osdl.org","subject":"Re: full kernel history, in patchset format","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T18:31:26Z","receivedAt":"2005-04-16T18:31:26Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"LT\" == Linus Torvalds <torvalds@osdl.org> writes:\n\nLT> What do people think? I'm not so much worried about the data itself: the\nLT> git architecture is _so_ damn simple that now that the size estimate has\nLT> been confirmed, that I don't think it would be a problem per se to put\nLT> 3.2GB into the archive. But it will bog down \"rsync\" horribly, so it will\nLT> actually hurt synchronization untill somebody writes the rev-tree-like\nLT> stuff to communicate changes more efficiently..\n\nLT> IOW, it smells to me like we don't have the infrastructure to really work \nLT> with 3GB archives, and that if we start from scratch (2.6.12-rc2), we can \nLT> build up the infrastructure in parallell with starting to really need it.\n\nLT> But it's _great_ to have the history in this format, especially since \nLT> looking at CVS just reminded me how much I hated it.\n\nLT> Comments?\n\nI have been cooking this idea before I dove into the merge stuff\nand did not have time to implement it myself (Hint Hint), but I\nthink something along the following lines would work nicely:\n\n * A script git-archive-tar is used to create a \"base tarball\"\n   that roughly corresponds to \"linux-*.tar.gz\".  This works as\n   follows:\n\n    $ git-archive-tar C [B1 B2...]\n\n   This reads the named commit C, grabs the associated tree\n   (i.e.  its sub-tree objects and the blob they refer to), and\n   makes a tarball of ??/??????????????????????????????????????\n   files.  The tarball does not have to contain any extra\n   information to reproduce any ancestor of the named commit.\n\n   When extra parameters, B1 B2..., are given, it also creates\n   \"diff package\" that roughly corresponds to \"patch-*.gz\" for\n   each Bn given.  They must be ancestors of commit.  The\n   intention is to store enough information to ensure that the\n   recipient can recreate all the SHA1 files \"base tarball\" for\n   commits between (Bn, C] would contain, provided if the\n   recipient already has all the SHA1 files \"base tarball\" for\n   Bn.\n\n * A script git-archive-patch is used to read such a \"diff\n   package\".\n\nSo a user needs to:\n\n * First pick some baseline B and download the base tarball for\n   commit B.  It is up to him to make trade-offs between how far\n   back he wants to see the history and how much bandwidth he\n   wants to waste.  Untar it to get the baseline.\n\n * Then periodically pick up \"diff package\" for (C, B] where C\n   is the latest available.  Run git-archive-patch to populate\n   the rest.\n\n * In addition the user can run rsync with timestamp option to\n   pick up SHA1 files created upstream since C after this\n   happens.\n\nWhat git-archive-tar needs to do to produce \"diff package\" for\n(Bn, C] is fairly obvious.\n\n * From rev-tree output, find all the commits that are on path\n   from Bn to C.\n\n * Find all the SHA1 objects that appear on this commit chain;\n   subtract what is in Bn since we assume the recipient has them\n   already.\n\n * Run diff-tree between neighboring commits [*1*] to find out\n   the set of blobs that are \"related\".  Extract those related\n   blobs and run \"diff\" [*2*] between them to see if it produces\n   a patch smaller than the whole thing when compressed.  If\n   diff+patch is a win, then we do not have to transmit the blob\n   that we could reproduce by sending the diff.  Note that fact.\n\n * When you are all done, you have a single patch file that\n   contains small edits on numerous blobs, and set of SHA1 files\n   that are cheaper to transmit than in the patch form.\n   Compress the patch file and package them together to make a\n   tar archive.\n\nGiven the above, the operation of git-archive-patch is also\nquite obvious.  Extract the \"diff package\" tarball into the\nobjects/ directory that has (at least) the full Bn, uncompress\nthe patch file part, and run patch on it. \n\n\n[Footnotes]\n\n*1* Alternatively, this diff-tree can be run between Bn and each\ncommit between (Bn, C].  It is like incremental dump strategy.\nWe should experiment and find a good balance.\n\n*2* This does not have to be \"diff -u\" --- we are assuming the\nexact patch so diff -e or xdelta would do.  We should experiment\nand find a good diff+patch pair.\n\n\n"},{"id":"333","messageId":"20050416183232.GH19099@pasky.ji.cz","threadId":"48","inReplyTo":"1113679421.28612.16.camel@tglx.tec.linutronix.de","subject":"Re: Re: full kernel history, in patchset format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-16T18:32:32Z","receivedAt":"2005-04-16T18:32:32Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Apr 16, 2005 at 09:23:40PM CEST, I got a letter\nwhere Thomas Gleixner <tglx@linutronix.de> told me that...\n> One remark on the tree blob storage format. \n> The binary storage of the sha1sum of the refered object is a PITA for\n> scripting. \n> Converting the ASCII -> binary for the sha1sum comparision should not\n> take much longer than the binary -> ASCII conversion for the file\n> reference. Can this be changed ?\n\nHuh, you aren't supposed to peek into trees directly. What's wrong with\nls-tree?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"334","messageId":"42615B00.6090106@timesys.com","threadId":"48","inReplyTo":"7vmzry7ev5.fsf@assigned-by-dhcp.cox.net","subject":"Re: full kernel history, in patchset format","fromName":"Mike Taht","fromEmail":"mike.taht@timesys.com","sentAt":"2005-04-16T18:35:44Z","receivedAt":"2005-04-16T18:35:44Z","isPatch":false,"sender":{"key":"mike.taht@timesys.com","avatar":null},"body":"\n>  * A script git-archive-tar is used to create a \"base tarball\"\n>    that roughly corresponds to \"linux-*.tar.gz\".  This works as\n>    follows:\n> \n>     $ git-archive-tar C [B1 B2...]\n> \n>    This reads the named commit C, grabs the associated tree\n>    (i.e.  its sub-tree objects and the blob they refer to), and\n>    makes a tarball of ??/??????????????????????????????????????\n>    files.  The tarball does not have to contain any extra\n>    information to reproduce any ancestor of the named commit.\n\nalternatively, \"git-archive-torrent\" to create a list of files for a \nbittorrent feed....\n\n-- \n\nMike Taht\n"},{"id":"335","messageId":"20050416183656.GI19099@pasky.ji.cz","threadId":"48","inReplyTo":"20050416183232.GH19099@pasky.ji.cz","subject":"Re: Re: Re: full kernel history, in patchset format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-16T18:36:56Z","receivedAt":"2005-04-16T18:36:56Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Apr 16, 2005 at 08:32:32PM CEST, I got a letter\nwhere Petr Baudis <pasky@ucw.cz> told me that...\n> Dear diary, on Sat, Apr 16, 2005 at 09:23:40PM CEST, I got a letter\n> where Thomas Gleixner <tglx@linutronix.de> told me that...\n> > One remark on the tree blob storage format. \n> > The binary storage of the sha1sum of the refered object is a PITA for\n> > scripting. \n> > Converting the ASCII -> binary for the sha1sum comparision should not\n> > take much longer than the binary -> ASCII conversion for the file\n> > reference. Can this be changed ?\n> \n> Huh, you aren't supposed to peek into trees directly. What's wrong with\n> ls-tree?\n\n(I meant, you aren't supposed to peek into trees from scripts. Or well,\nnot \"not supposed\", but it does not make much sense when you have\nls-tree.)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"337","messageId":"Pine.LNX.4.58.0504161135480.7211@ppc970.osdl.org","threadId":"48","inReplyTo":"1113679421.28612.16.camel@tglx.tec.linutronix.de","subject":"Re: full kernel history, in patchset format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T18:44:34Z","receivedAt":"2005-04-16T18:44:34Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Thomas Gleixner wrote:\n> \n> One remark on the tree blob storage format. \n> The binary storage of the sha1sum of the refered object is a PITA for\n> scripting. \n> Converting the ASCII -> binary for the sha1sum comparision should not\n> take much longer than the binary -> ASCII conversion for the file\n> reference. Can this be changed ?\n\nI'd really rather not. Why don't you just use \"ls-tree\" for scripting? \nThat's why it exists in the first place. \n\nIt might make sense to have some simple selection capabilities built into \nls-tree (ie \"ls-tree --match drivers/char/ -z <treesha1>\" to get just a \nsubtree out), but that depends entirely on how you end up using it.\n\nThe fact is, there should _never_ any reason to look at the objects\nthemselves directly. \"cat-file\" is a debugging aid, it shouldn't be\nscripted (with the possible exception of \"cat-file blob xxxx\" to just\nextract the blob contents, since that object doesn't have any internal\nstructure).\n\nThat level of abstraction (\"we never look directly at the objects\") is \nwhat allows us to change the object structure later. For example, we \nalready changed the \"commit\" date thing once, and the tree object has \nobviously evolved a bit, and if we ever change the hash, the objects will \nchange too, but if you always just script them using nice helper tools, \nyou won't ever need to _care_. And that's how it should be.\n\nIf there's a tool missing, holler. THAT is the part I've been trying to\nwrite: all the plumbing so that you _can_ script the thing sanely, and not\nworry about how objects are created and worked with. \n\nFor example, that \"index\" file format likely _will_ change. I ended up\ndoing the new \"stage\" flags in a way that kept the index file compatible\nwith old ones, but I did that mainly because it also happened to be the\neasiest way to enforce the rule I wanted to enforce (ie the \"stage\" really\n_is_ a part of the filename from a \"compare filenames\" standpoint, in\norder to make sure that the stages are always ordered).\n\nSo if the index file change hadn't had that property, I'd have just said\n\"I'll change the format\", and anybody who tried to parse the index file\nwould have been _broken_.\n\n\t\tLinus\n"},{"id":"338","messageId":"7vd5su7e5j.fsf@assigned-by-dhcp.cox.net","threadId":"48","inReplyTo":"7vmzry7ev5.fsf@assigned-by-dhcp.cox.net","subject":"Re: full kernel history, in patchset format","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T18:46:48Z","receivedAt":"2005-04-16T18:46:48Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"JCH\" == Junio C Hamano <junkio@cox.net> writes:\n\nJCH> I have been cooking this idea before I dove into the merge stuff\nJCH> and did not have time to implement it myself (Hint Hint), but I\nJCH> think something along the following lines would work nicely:\n\nIt should be fairly obvious from the context what I meant to\nsay, but in case somebody gets confused by my inaccurate\ndescription of small details (or, before somebody nitpicks ;-),\nI'd add some clarifications and corrections.\n\nJCH>  * Run diff-tree between neighboring commits [*1*] to find out\nJCH>    the set of blobs that are \"related\".  Extract those related\nJCH>    blobs and run \"diff\" [*2*] between them to see if it produces\nJCH>    a patch smaller than the whole thing when compressed.  If\nJCH>    diff+patch is a win, then we do not have to transmit the blob\nJCH>    that we could reproduce by sending the diff.  Note that fact.\n\nI talked only about blobs here, but I really mean all types:\ncommits, trees and blobs here.  Nothing prevents us from\nextracting the raw data for trees and commits and run diff\nbetween them.  We can use cat-file to do that today.\n\nWhat we do not have is the reverse of \"$ cat-file type >rawdata\"\n(i.e. \"$ write-file type <rawdata\"), but that is trivial to\nwrite.  The raw data for related tree objects should delta well.\nI do not think it is worth the effort to attempt delta for\ncommit objects.  Anything that git-archive-tar decides not to\nsend in diff+patch form, be it blob or tree or commit, should be\nnoted here, not just blob as my previous message incorrectly\nimplies.\n\nJCH> Given the above, the operation of git-archive-patch is also\nJCH> quite obvious.  Extract the \"diff package\" tarball into the\nJCH> objects/ directory that has (at least) the full Bn, uncompress\nJCH> the patch file part, and run patch on it. \n\nOf course after you ran patch to reproduce the raw data for the\nblob or tree, we need the reverse of cat-file to register such\ndata under object/ hierarchy.\n\n"},{"id":"341","messageId":"20050416185751.GJ19099@pasky.ji.cz","threadId":"48","inReplyTo":"1113681021.28612.29.camel@tglx.tec.linutronix.de","subject":"Re: Re: full kernel history, in patchset format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-16T18:57:51Z","receivedAt":"2005-04-16T18:57:51Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Apr 16, 2005 at 09:50:21PM CEST, I got a letter\nwhere Thomas Gleixner <tglx@linutronix.de> told me that...\n> On Sat, 2005-04-16 at 11:44 -0700, Linus Torvalds wrote:\n> \n> > That level of abstraction (\"we never look directly at the objects\") is \n> > what allows us to change the object structure later. For example, we \n> > already changed the \"commit\" date thing once, and the tree object has \n> > obviously evolved a bit, and if we ever change the hash, the objects will \n> > change too, but if you always just script them using nice helper tools, \n> > you won't ever need to _care_. And that's how it should be.\n> \n> For the export stuff its terrible slow. :(\n\nIt seems to me that you must be doing something wrong then. I can't see\nanything which would not make ls-tree blindingly fast (except for when\nbeing recursive, see below).\n\nBTW, what do you need ls-tree output for, when doing export _to_ git?\n\nP.S.: It seems that Linus applied a patch to ls-tree which will make it\nread_sha1_file() on each item when ls-tree is recursive. Junio, why did\nyou do it? Is there any possible case when the item would not be marked\nas directory but it would be a tree object? I could imagine it bogging\ndown ls-tree on big tree a lot.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"342","messageId":"20050416191312.GT9461@lug-owl.de","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504160953310.7211@ppc970.osdl.org","subject":"Re: full kernel history, in patchset format","fromName":"Jan-Benedict Glaw","fromEmail":"jbglaw@lug-owl.de","sentAt":"2005-04-16T19:13:12Z","receivedAt":"2005-04-16T19:13:12Z","isPatch":false,"sender":{"key":"jbglaw@lug-owl.de","avatar":null},"body":"On Sat, 2005-04-16 10:04:31 -0700, Linus Torvalds <torvalds@osdl.org>\nwrote in message <Pine.LNX.4.58.0504160953310.7211@ppc970.osdl.org>:\n\n> What do people think? I'm not so much worried about the data itself: the\n> git architecture is _so_ damn simple that now that the size estimate has\n> been confirmed, that I don't think it would be a problem per se to put\n> 3.2GB into the archive. But it will bog down \"rsync\" horribly, so it will\n> actually hurt synchronization untill somebody writes the rev-tree-like\n> stuff to communicate changes more efficiently..\n> \n> IOW, it smells to me like we don't have the infrastructure to really work \n> with 3GB archives, and that if we start from scratch (2.6.12-rc2), we can \n> build up the infrastructure in parallell with starting to really need it.\n\n3GB is quite some data, but I'd accept and prefer to download it from\nsomewhere. I think that it's worth it.\n\nI accept that there are people out there which would love to get a\nsmaller archive, but at least most developers that would actually use it\nfor day-to-day work *do* have the bandwidth to download it. Maybe we'd\nalso prepare (from time to time) bzip'ed tarballs, which I expect to be\na tad smaller.\n\nMfG, JBG\n\n-- \nJan-Benedict Glaw       jbglaw@lug-owl.de    . +49-172-7608481             _ O _\n\"Eine Freie Meinung in  einem Freien Kopf    | Gegen Zensur | Gegen Krieg  _ _ O\n fuer einen Freien Staat voll Freier Bürger\" | im Internet! |   im Irak!   O O O\nret = do_actions((curr | FREE_SPEECH) & ~(NEW_COPYRIGHT_LAW | DRM | TCPA));\n"},{"id":"343","messageId":"Pine.LNX.4.58.0504161208410.7211@ppc970.osdl.org","threadId":"48","inReplyTo":"1113681021.28612.29.camel@tglx.tec.linutronix.de","subject":"Re: full kernel history, in patchset format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-16T19:15:35Z","receivedAt":"2005-04-16T19:15:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 16 Apr 2005, Thomas Gleixner wrote:\n> \n> For the export stuff its terrible slow. :(\n\nI don't really see your point.\n\nIf you already know what the tree is like you say, you don't care about\nthe tree object. And if you don't know what the tree is, what _are_ you\ndoing?\n\nIn other words, show us what you're complaining about. If you're looking\ninto the trees yourself, then the binary representation of the sha1 is\nalready what you want. That _is_ the hash. So why do you want it in ASCII?  \nAnd if you're not looking into the tree directly, but using \"cat-file\ntree\" and you were hoping to see ASCII data, then that's certainly not\ngoing to be any faster than just doing \"ls-tree\" instead.\n\nIn other words, I don't see your point. Either you want ascii output for \nscripting, or you don't. First you claimed that you did, and that you \nwould want the tree object to change in order to do so. Now you claim that \nyou can't use \"ls-tree\" because it's too slow. \n\nThat just isn't making any sense. You're mixing two totally different\nlevels, and complaining about performance when scripting things. Yet\nyou're talking about a 20-byte data structure that is trivial to convert\nto any format you want.\n\nWhat kind of _strange_ scripting architecture is so fast that there's a\ndifference between \"cat-file\" and \"ls-tree\" and can handle 17,000 files in\n60,000 revisions, yet so slow that you can't trivially convert 20 bytes of \ndata?\n\n\t\tLinus\n"},{"id":"345","messageId":"20050416191810.GA27748@elte.hu","threadId":"48","inReplyTo":"42612771.5000705@cobite.com","subject":"Re: full kernel history, in patchset format","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-16T19:18:10Z","receivedAt":"2005-04-16T19:18:10Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* David Mansfield <david@cobite.com> wrote:\n\n> Ingo Molnar wrote:\n> >* Ingo Molnar <mingo@elte.hu> wrote:\n> >\n> >\n> >>the patches contain all the existing metadata, dates, log messages and \n> >>revision history. (What i think is missing is the BK tree merge \n> >>information, but i'm not sure we want/need to convert them to GIT.)\n> >\n> >\n> >author names are abbreviated, e.g. 'viro' instead of \n> >viro@parcelfarce.linux.theplanet.co.uk, and no committer information is \n> >included (albeit commiter ought to be Linus in most cases). These are \n> >limitations of the BK->CVS gateway i think.\n> >\n> \n> Glad to hear cvsps made it through!  I'm curious what the manual \n> fixups required were, except for the binary file issue (logo.gif).\n\n--cvs-direct was needed to speed it up from 'several days to finish' to \n'several hours to finish', but it crashed on a handful of patches [i \nused the latest devel snapshot so this isnt a complaint]. (one of the \ncrashes was when generating 1860.patch.) Also, 'cvs rdiff' apparently \nemits an empty patch for diffs that remove a file that end without \nhaving a newline character - but this isnt cvsps's problem.  (grep for \n+++ in the patchset to find those cases.)\n\n> As to the actual email addresses, for more recent patches, the \n> Signed-off should help.  For earlier ones, isn't their some script \n> which 'knows' a bunch of canonical author->email mappings? (the \n> shortlog script or something)?\n\nyeah, that's not that much of a problem, most of the names are unique, \nand the rest can be fixed up too.\n\n> Is the full committer email address actually in the changeset in BK?  \n> If so, given that we have the unique id (immutable I believe) of the \n> changset, could it be extracted directly from BK?\n\ni think it's included in BK.\n\n\tIngo\n"},{"id":"346","messageId":"7v8y3i7cn9.fsf@assigned-by-dhcp.cox.net","threadId":"48","inReplyTo":"42615B00.6090106@timesys.com","subject":"Re: full kernel history, in patchset format","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T19:19:22Z","receivedAt":"2005-04-16T19:19:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"MT\" == Mike Taht <mike.taht@timesys.com> writes:\n\nMT> alternatively, \"git-archive-torrent\" to create a list of files for a\nMT> bittorrent feed....\n\nThat is certainly good for establishing the baseline, but you\nstill need to leverage the inherent delta-compressibility\nbetween related blobs/trees by also doing something like what I\ndescribed as \"diff package\", don't you?\n\n\n\n\n"},{"id":"331","messageId":"1113679421.28612.16.camel@tglx.tec.linutronix.de","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504160953310.7211@ppc970.osdl.org","subject":"Re: full kernel history, in patchset format","fromName":"Thomas Gleixner","fromEmail":"tglx@linutronix.de","sentAt":"2005-04-16T19:23:40Z","receivedAt":"2005-04-16T19:23:40Z","isPatch":false,"sender":{"key":"tglx@linutronix.de","avatar":null},"body":"On Sat, 2005-04-16 at 10:04 -0700, Linus Torvalds wrote:\n\n> So I'd _almost_ suggest just starting from a clean slate after all.  \n> Keeping the old history around, of course, but not necessarily putting it\n> into git now. It would just force everybody who is getting used to git in \n> the first place to work with a 3GB archive from day one, rather than \n> getting into it a bit more gradually.\n\nSure. We can export the 2.6.12-rc2 version of the git'ed history tree\nand start from there. Then the first changeset has a parent, which just\nlives in a different place. \nThats the only difference to your repository, but it would change the\nsha1 sums of all your changesets.\n\n> What do people think? I'm not so much worried about the data itself: the\n> git architecture is _so_ damn simple that now that the size estimate has\n> been confirmed, that I don't think it would be a problem per se to put\n> 3.2GB into the archive. But it will bog down \"rsync\" horribly, so it will\n> actually hurt synchronization untill somebody writes the rev-tree-like\n> stuff to communicate changes more efficiently..\n\nWe have all the tracking information in SQL and we will post the data\nbase dump soon, so people interested in revision tracking can use this\nas an information base.\n\n> But it's _great_ to have the history in this format, especially since \n> looking at CVS just reminded me how much I hated it.\n\n:)\n\nOne remark on the tree blob storage format. \nThe binary storage of the sha1sum of the refered object is a PITA for\nscripting. \nConverting the ASCII -> binary for the sha1sum comparision should not\ntake much longer than the binary -> ASCII conversion for the file\nreference. Can this be changed ?\n\ntglx\n\n\n"},{"id":"347","messageId":"7v3btq7c9n.fsf@assigned-by-dhcp.cox.net","threadId":"48","inReplyTo":"20050416185751.GJ19099@pasky.ji.cz","subject":"Re: full kernel history, in patchset format","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T19:27:32Z","receivedAt":"2005-04-16T19:27:32Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"PB\" == Petr Baudis <pasky@ucw.cz> writes:\n\nPB> P.S.: It seems that Linus applied a patch to ls-tree which will make it\nPB> read_sha1_file() on each item when ls-tree is recursive. Junio, why did\nPB> you do it?\n\nSorry it was my misunderstanding, before I found out exactly how\nS_ISDIR is used.  Thank you for pointing it out.\n\nI was confused by this comment around the area I changed:\n\n    /* XXX: We do some ugly mode heuristics here.\n     * It seems not worth it to read each file just to get this\n     * and the file size. -- pasky@ucw.cz\n\nI mistakenly inferred from that comment that S_ISDIR(mode) is\nnot a guarantee.  So I mistakenly optimized it for non-recursive\ncase by keeping that \"heuristics\".  The logic was: If recursive\nwe will need to run read_sha1_file() to find out if it is really\na tree anyway.\n\nI'll fix it up, now I know S_ISDIR(mode) is a guarantee that it\nis a tree, I'll do the \"heuristics\" first, and do read_sha1_file\nonly when it is a tree and I am recursive.\n\n"},{"id":"336","messageId":"1113680477.28612.24.camel@tglx.tec.linutronix.de","threadId":"48","inReplyTo":"20050416183232.GH19099@pasky.ji.cz","subject":"Re: Re: full kernel history, in patchset format","fromName":"Thomas Gleixner","fromEmail":"tglx@linutronix.de","sentAt":"2005-04-16T19:41:17Z","receivedAt":"2005-04-16T19:41:17Z","isPatch":false,"sender":{"key":"tglx@linutronix.de","avatar":null},"body":"On Sat, 2005-04-16 at 20:32 +0200, Petr Baudis wrote:\n> Dear diary, on Sat, Apr 16, 2005 at 09:23:40PM CEST, I got a letter\n> where Thomas Gleixner <tglx@linutronix.de> told me that...\n> > One remark on the tree blob storage format. \n> > The binary storage of the sha1sum of the refered object is a PITA for\n> > scripting. \n> > Converting the ASCII -> binary for the sha1sum comparision should not\n> > take much longer than the binary -> ASCII conversion for the file\n> > reference. Can this be changed ?\n> \n> Huh, you aren't supposed to peek into trees directly. What's wrong with\n> ls-tree?\n\nWhy I'm not supposed ? Is this evil ?\n\nMy export script has all the data available, so I write the tree refs\ndirectly. The full export runs ~1 hour. Thats long enough :) I tried the\ngit way and it slows me down by factor \"BIG\" (I dont remember the\nnumber)\n\nAlso for reference tracking all the information might be available e.g.\nby a database. Why should the revtool then use some tool to retrieve\ninformation which is already there ?\n\ntglx\n\n\n"},{"id":"350","messageId":"20050416194126.GA28429@elte.hu","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504160953310.7211@ppc970.osdl.org","subject":"Re: full kernel history, in patchset format","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-16T19:41:26Z","receivedAt":"2005-04-16T19:41:26Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Linus Torvalds <torvalds@osdl.org> wrote:\n\n> > the history data starts at 2.4.0 and ends at 2.6.12-rc2. I've included a \n> > script that will apply all the patches in order and will create a \n> > pristine 2.6.12-rc2 tree.\n> \n> Hey, that's great. I got the CVS repo too, and I was looking at it, \n> but the more I looked at it, the more I felt that the main reason I \n> want to import it into git ends up being to validate that my size \n> estimates are at all realistic.\n> \n> I see that Thomas Gleixner seems to have done that already, and come \n> to a figure of 3.2GB for the last three years, which I'm very happy \n> with, mainly because it seems to match my estimates to a tee. [...]\n\n(yeah, we apparently worked in parallel - i only learned about his \nefforts after i sent my mail. He was using BK to extract info, i was \nusing the CVS tree alone and no BK code whatsoever. (I dont think there \nwill be any argument about who owns what, but i wanted to be on the safe \nside, and i also wanted to see how complete and usable the CVS metadata \nis - it's close to perfect i'd say, for the purposes i care about.))\n\n> But I wonder if we actually want to actually populate the whole \n> history..\n\nyeah, it definitely feels a bit brave to import 28,000 changesets into a \nsource-code database project that will be a whopping 2 weeks old in 2 \ndays ;) Even if we felt 100% confident about all the basics (which we do \nof course ;), it's just simply too young to tie things down via a 3.2GB \ndatabase. It feels much more natural to grow it gradually, 28,000 \nchangesets i'm afraid would just suffocate the 'project growth \ndynamics'. Not going too fast is just as important as not going too \nslow.\n\nI didnt generate the patchset to get it added into some central \nrepository right now, i generated it to check that we _do_ have all the \nrevision history in an easy to understand format which does generate \ntoday's kernel tree, so that we can lean back and worry about the full \ndatabase once things get a bit more settled down (in a couple of months \nor so). It's also an easy testbed for GIT itself.\n\nbut the revision history was one of the main reasons i used BK myself, \nso we'll need a merged database eventually. Occasionally i needed to \ncheck who was the one who touched a particular piece of code - was that \nfantastic new line of code written by me, or was that buggy piece of \ncrap written by someone else? ;) Also, looking at a change and then \ngoing to the changeset that did it, and then looking at the full picture \nwas pretty useful too. So that sort of annotation, and generally \nnavigating around _quickly_ and looking at the 'flow' of changes going \ninto a particular file was really useful (for me).\n\n\tIngo\n"},{"id":"351","messageId":"42616A9F.1030302@timesys.com","threadId":"48","inReplyTo":"7v8y3i7cn9.fsf@assigned-by-dhcp.cox.net","subject":"Re: full kernel history, in patchset format","fromName":"Mike Taht","fromEmail":"mike.taht@timesys.com","sentAt":"2005-04-16T19:42:23Z","receivedAt":"2005-04-16T19:42:23Z","isPatch":false,"sender":{"key":"mike.taht@timesys.com","avatar":null},"body":"Junio C Hamano wrote:\n>>>>>>\"MT\" == Mike Taht <mike.taht@timesys.com> writes:\n> \n> \n> MT> alternatively, \"git-archive-torrent\" to create a list of files for a\n> MT> bittorrent feed....\n> \n> That is certainly good for establishing the baseline, but you\n> still need to leverage the inherent delta-compressibility\n> between related blobs/trees by also doing something like what I\n> described as \"diff package\", don't you?\n\nYes... yes you could have files and diffs generated statically...\n\nalthough something like a bittorrent server/client/frontend, call it \n\"gittorrent\" (I hate being the first to make this pun) could walk the \nhashes dynamically (\nIhave: sha,sha,sha,sha... Sendme: shaxxxxxxxxxxxxxxxxxxx\nHereswhatyouneedfromgit: file,file,file,diff,diff,diff,...)\n\n-- \n\nMike Taht\n\n\n   \"It looks like blind screaming hedonism won out.\"\n"},{"id":"352","messageId":"7vy8bi5wwr.fsf@assigned-by-dhcp.cox.net","threadId":"48","inReplyTo":"20050416162632.GA3309@64m.dyndns.org","subject":"Re: full kernel history, in patchset format","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-04-16T19:44:36Z","receivedAt":"2005-04-16T19:44:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":">>>>> \"CL\" == Christopher Li <git@chrisli.org> writes:\n\nCL> I bet 90% of the time people sync to the repository head first\nCL> want to check out the last bits. And maybe reading some change\nCL> log to see what is changed.\n\nCL> So having all the commit object, the user will able to see\nCL> what is change and which version he we like to check out.\n\nMakes sense.\n\n"},{"id":"339","messageId":"1113681021.28612.29.camel@tglx.tec.linutronix.de","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504161135480.7211@ppc970.osdl.org","subject":"Re: full kernel history, in patchset format","fromName":"Thomas Gleixner","fromEmail":"tglx@linutronix.de","sentAt":"2005-04-16T19:50:21Z","receivedAt":"2005-04-16T19:50:21Z","isPatch":false,"sender":{"key":"tglx@linutronix.de","avatar":null},"body":"On Sat, 2005-04-16 at 11:44 -0700, Linus Torvalds wrote:\n\n> That level of abstraction (\"we never look directly at the objects\") is \n> what allows us to change the object structure later. For example, we \n> already changed the \"commit\" date thing once, and the tree object has \n> obviously evolved a bit, and if we ever change the hash, the objects will \n> change too, but if you always just script them using nice helper tools, \n> you won't ever need to _care_. And that's how it should be.\n\nFor the export stuff its terrible slow. :(\n\nI agree that using common tools is good. But we talk also about an open\nformat, so using a script to speed up certain tasks is not bad at all.\n\ntglx\n\n\n\n"},{"id":"353","messageId":"Pine.LNX.4.21.0504161546120.30848-100000@iabervon.org","threadId":"48","inReplyTo":"42616A9F.1030302@timesys.com","subject":"Re: full kernel history, in patchset format","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-16T20:19:01Z","receivedAt":"2005-04-16T20:19:01Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sat, 16 Apr 2005, Mike Taht wrote:\n\n> Junio C Hamano wrote:\n> >>>>>>\"MT\" == Mike Taht <mike.taht@timesys.com> writes:\n> > \n> > \n> > MT> alternatively, \"git-archive-torrent\" to create a list of files for a\n> > MT> bittorrent feed....\n> > \n> > That is certainly good for establishing the baseline, but you\n> > still need to leverage the inherent delta-compressibility\n> > between related blobs/trees by also doing something like what I\n> > described as \"diff package\", don't you?\n> \n> Yes... yes you could have files and diffs generated statically...\n> \n> although something like a bittorrent server/client/frontend, call it \n> \"gittorrent\" (I hate being the first to make this pun) could walk the \n> hashes dynamically (\n> Ihave: sha,sha,sha,sha... Sendme: shaxxxxxxxxxxxxxxxxxxx\n> Hereswhatyouneedfromgit: file,file,file,diff,diff,diff,...)\n\nI'm actually working on a trivial HTTP client to do this. The user says\n\"get <commit-id> from <url>\", and it gets that object, the associated\ntrees, and the associated blobs, skipping any that it already has.\n\nThis should save having a non-standard public-facing server process, and\nbe essentially as effective, at least once I have it using a single\nconnection for everything.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"348","messageId":"1113683730.28612.34.camel@tglx.tec.linutronix.de","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504161208410.7211@ppc970.osdl.org","subject":"Re: full kernel history, in patchset format","fromName":"Thomas Gleixner","fromEmail":"tglx@linutronix.de","sentAt":"2005-04-16T20:35:30Z","receivedAt":"2005-04-16T20:35:30Z","isPatch":false,"sender":{"key":"tglx@linutronix.de","avatar":null},"body":"On Sat, 2005-04-16 at 12:15 -0700, Linus Torvalds wrote:\n> \n> On Sat, 16 Apr 2005, Thomas Gleixner wrote:\n> > \n> > For the export stuff its terrible slow. :(\n>\n> What kind of _strange_ scripting architecture is so fast that there's a\n> difference between \"cat-file\" and \"ls-tree\" and can handle 17,000 files in\n> 60,000 revisions, yet so slow that you can't trivially convert 20 bytes of \n> data?\n\nSorry I was neither talking about \"cat-file ...\" nor about the 20 byte\nconversion. I was talking about the bk export script, which writes the\nobjects itself. Doing this with the git-tools would slow it down, as I\nhave the retrieved data already in memory. It does not slow me down to\ncreate the binary ref, but its annoying.\n\nI just figured, that some revtools might have the need to use direct\npointers into objects and face the same problem the other way round.\n\ntglx\n\n\n"},{"id":"378","messageId":"Pine.LNX.4.62.0504161606590.22652@qynat.qvtvafvgr.pbz","threadId":"48","inReplyTo":"1113679421.28612.16.camel@tglx.tec.linutronix.de","subject":"Re: full kernel history, in patchset format","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2005-04-16T23:08:55Z","receivedAt":"2005-04-16T23:08:55Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"On Sat, 16 Apr 2005, Thomas Gleixner wrote:\n\n> On Sat, 2005-04-16 at 10:04 -0700, Linus Torvalds wrote:\n>\n>> So I'd _almost_ suggest just starting from a clean slate after all.\n>> Keeping the old history around, of course, but not necessarily putting it\n>> into git now. It would just force everybody who is getting used to git in\n>> the first place to work with a 3GB archive from day one, rather than\n>> getting into it a bit more gradually.\n>\n> Sure. We can export the 2.6.12-rc2 version of the git'ed history tree\n> and start from there. Then the first changeset has a parent, which just\n> lives in a different place.\n> Thats the only difference to your repository, but it would change the\n> sha1 sums of all your changesets.\n\nat least start with a full release. say 2.6.11\n\nthe history won't be blank, but it's far more likly that people will care \nabout the details between 2.6.11 and 2.6.12 and will want to go back \nbefore -rc2\n\nDavid Lang\n\n-- \nThere are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.\n  -- C.A.R. Hoare\n"},{"id":"570","messageId":"1113780698.11910.8.camel@localhost.localdomain","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504160953310.7211@ppc970.osdl.org","subject":"Re: full kernel history, in patchset format","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-17T23:31:36Z","receivedAt":"2005-04-17T23:31:36Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Sat, 2005-04-16 at 10:04 -0700, Linus Torvalds wrote:\n> So I'd _almost_ suggest just starting from a clean slate after all.  \n> Keeping the old history around, of course, but not necessarily putting it\n> into git now. It would just force everybody who is getting used to git in \n> the first place to work with a 3GB archive from day one, rather than \n> getting into it a bit more gradually.\n> \n> What do people think? I'm not so much worried about the data itself: the\n> git architecture is _so_ damn simple that now that the size estimate has\n> been confirmed, that I don't think it would be a problem per se to put\n> 3.2GB into the archive. But it will bog down \"rsync\" horribly, so it will\n> actually hurt synchronization untill somebody writes the rev-tree-like\n> stuff to communicate changes more efficiently..\n\nNote that any given copy of a tree doesn't _need_ to keep all the\nhistory back the beginning of time. It's OK if the oldest commit object\nin your tree actually refers back to a parent which doesn't exist\nlocally. I can well imagine that some people will want to keep their\ntrees pruned to keep only a few weeks of history, while other copies of\nthe tree will keep everything.\n\nHowever, if we _don't_ base our current work on an existing import of\nthe kernel, then we don't retain that option. We can't just change the\n'parent' field of your 2.6.12-rc2 import, without changing the sha1 hash\nof _everything_ that happens thereafter. \n\nSo I'd say we should take Thomas' import, and base new work on that --\nbut then possibly leave out the older objects from the 'working'\nrepository which everyone is rsyncing from; just make them available in\na 'linux-history.git' object database elsewhere.\n\n-- \ndwmw2\n\n"},{"id":"574","messageId":"20050417233936.GV1461@pasky.ji.cz","threadId":"48","inReplyTo":"1113780698.11910.8.camel@localhost.localdomain","subject":"Re: full kernel history, in patchset format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-17T23:39:36Z","receivedAt":"2005-04-17T23:39:36Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, Apr 18, 2005 at 01:31:36AM CEST, I got a letter\nwhere David Woodhouse <dwmw2@infradead.org> told me that...\n> Note that any given copy of a tree doesn't _need_ to keep all the\n> history back the beginning of time. It's OK if the oldest commit object\n> in your tree actually refers back to a parent which doesn't exist\n> locally. I can well imagine that some people will want to keep their\n> trees pruned to keep only a few weeks of history, while other copies of\n> the tree will keep everything.\n\nI think this is bad, bad, bad. If you don't keep around all the\n_commits_, you get into all sorts of troubles - when merging, when doing\ngit log, etc. And the commits themselves are probably actually pretty\nsmall portion of the thing. I didn't do any actual measurement but I\nwould be pretty surprised if it would be much more than few megabytes of\ndata for the kernel history.\n\nOf course an entirely different thing are _trees_ associated with those\ncommits. As long as you stay with a simple three-way merge, you\nbasically never want to look at trees which aren't heads and which you\ndon't specifically request to look at. And the trees and what they carry\ninside is the main bulk of data.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"579","messageId":"1113782805.11910.36.camel@localhost.localdomain","threadId":"48","inReplyTo":"20050417233936.GV1461@pasky.ji.cz","subject":"Re: full kernel history, in patchset format","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-18T00:06:43Z","receivedAt":"2005-04-18T00:06:43Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Mon, 2005-04-18 at 01:39 +0200, Petr Baudis wrote:\n> I think this is bad, bad, bad. If you don't keep around all the\n> _commits_, you get into all sorts of troubles - when merging, when doing\n> git log, etc. And the commits themselves are probably actually pretty\n> small portion of the thing. I didn't do any actual measurement but I\n> would be pretty surprised if it would be much more than few megabytes of\n> data for the kernel history.\n\nI'm not sure it's that bad -- and everyone already seems perfectly happy\nnot to have history going back before 2.6.12-rc2. We're not talking\nabout doing this by _default_ -- we're talking about allowing people to\nkeep trees pruned if they _want_ to. So I might want to drop history\nbefore 2.6.0 on my laptop, for example.\n\n> Of course an entirely different thing are _trees_ associated with those\n> commits. As long as you stay with a simple three-way merge, you\n> basically never want to look at trees which aren't heads and which you\n> don't specifically request to look at. And the trees and what they carry\n> inside is the main bulk of data.\n\nIf the trees are absent and you're trying to merge, what do you gain\nfrom having the commit objects? And for the case of 'git log', I\ncertainly think it's acceptable that you lose out on those parts of\nprehistory which you've explicitly removed from your local tree --\nthat's a feature, not a bug. \n\nFor the special case of removing history before 2.6.12-rc2 from the\ntrees, I certainly think we can do it by leaving out all the commits,\nnot just the trees. We can do that easily, but there's no way we can\n_add_ that history retrospectively if we omit it in the first place.\n\nFor history older than 2.6.12-rc2 I'd suggest that it would be available\nin a different place, and absent from the 'main' working tree that\neveryone uses by default. The only difference we'd see in the working\ntree is that the 2.6.12-rc2 commit -- the oldest commit in that tree --\nwould actually have an absentee parent instead of appearing to be an\nimport. And all the sha1 hashes of all subsequent commits would be\ndifferent, of course.\n\nTo allow pruning of older objects in the general case would be a little\nbit harder than that, because as things stand you'd be re-fetching them\nevery time you rsync from elsewhere -- but that wouldn't really be hard\nto fix if we care.\n\nEither way, I think it can probably be done by omitting the commit\nobjects as well as the trees -- but the important point is that we\n_should_ include a 'parent' pointer in the oldest commit of the tree\nwe're working with, pointing back to the imported history.\n\n-- \ndwmw2\n\n"},{"id":"582","messageId":"20050418003526.GD1461@pasky.ji.cz","threadId":"48","inReplyTo":"1113782805.11910.36.camel@localhost.localdomain","subject":"Re: full kernel history, in patchset format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-18T00:35:26Z","receivedAt":"2005-04-18T00:35:26Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, Apr 18, 2005 at 02:06:43AM CEST, I got a letter\nwhere David Woodhouse <dwmw2@infradead.org> told me that...\n> On Mon, 2005-04-18 at 01:39 +0200, Petr Baudis wrote:\n> > Of course an entirely different thing are _trees_ associated with those\n> > commits. As long as you stay with a simple three-way merge, you\n> > basically never want to look at trees which aren't heads and which you\n> > don't specifically request to look at. And the trees and what they carry\n> > inside is the main bulk of data.\n> \n> If the trees are absent and you're trying to merge, what do you gain\n> from having the commit objects?\n\nmerge-base\n\n> For the special case of removing history before 2.6.12-rc2 from the\n> trees, I certainly think we can do it by leaving out all the commits,\n> not just the trees. We can do that easily, but there's no way we can\n> _add_ that history retrospectively if we omit it in the first place.\n\nI'm confused by this paragraph, but that might be my English skills\nfailing somehow.\n\n> For history older than 2.6.12-rc2 I'd suggest that it would be available\n> in a different place, and absent from the 'main' working tree that\n> everyone uses by default. The only difference we'd see in the working\n> tree is that the 2.6.12-rc2 commit -- the oldest commit in that tree --\n> would actually have an absentee parent instead of appearing to be an\n> import. And all the sha1 hashes of all subsequent commits would be\n> different, of course.\n\nYes, that's what I suggested too.\n\n> To allow pruning of older objects in the general case would be a little\n> bit harder than that, because as things stand you'd be re-fetching them\n> every time you rsync from elsewhere -- but that wouldn't really be hard\n> to fix if we care.\n\nI think http-pull is very promising. :-)\n\nIt could be actually much faster than rsync, since you don't need to\nbuild directory listings etc, which actually takes non-trivial amount of\ntime already with the kernel git repository.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"583","messageId":"1113785123.11910.43.camel@localhost.localdomain","threadId":"48","inReplyTo":"20050418003526.GD1461@pasky.ji.cz","subject":"Re: full kernel history, in patchset format","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-18T00:45:22Z","receivedAt":"2005-04-18T00:45:22Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Mon, 2005-04-18 at 02:35 +0200, Petr Baudis wrote:\n> > For the special case of removing history before 2.6.12-rc2 from the\n> > trees, I certainly think we can do it by leaving out all the commits,\n> > not just the trees. We can do that easily, but there's no way we can\n> > _add_ that history retrospectively if we omit it in the first place.\n> \n> I'm confused by this paragraph, but that might be my English skills\n> failing somehow.\n\n\"For the general case of people pruning their own trees, _maybe_ you're\nright that it would be good to keep the commits even if we delete the\nactual trees. But for history older than 2.6.12-rc2, that's a special\ncase -- I think we can happily delete the commits too.\n\n\"We can delete old trees/commits easily, but we can't _add_ them to the\nexisting linux-2.6.git tree, because the oldest commit in that tree\n(b4ceb6e27e4cc3f37d26e04c4535c79b98a9f889) doesn't have a parent.\"\n\n-- \ndwmw2\n\n"},{"id":"585","messageId":"20050418005032.GE1461@pasky.ji.cz","threadId":"48","inReplyTo":"1113785123.11910.43.camel@localhost.localdomain","subject":"Re: full kernel history, in patchset format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-18T00:50:32Z","receivedAt":"2005-04-18T00:50:32Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, Apr 18, 2005 at 02:45:22AM CEST, I got a letter\nwhere David Woodhouse <dwmw2@infradead.org> told me that...\n> On Mon, 2005-04-18 at 02:35 +0200, Petr Baudis wrote:\n> > > For the special case of removing history before 2.6.12-rc2 from the\n> > > trees, I certainly think we can do it by leaving out all the commits,\n> > > not just the trees. We can do that easily, but there's no way we can\n> > > _add_ that history retrospectively if we omit it in the first place.\n> > \n> > I'm confused by this paragraph, but that might be my English skills\n> > failing somehow.\n> \n> \"For the general case of people pruning their own trees, _maybe_ you're\n> right that it would be good to keep the commits even if we delete the\n> actual trees. But for history older than 2.6.12-rc2, that's a special\n> case -- I think we can happily delete the commits too.\n\nAh _so_. Thanks for explanation.\n\nI think I will make git-pasky's default behaviour (when we get\nhttp-pull, that is) to keep the complete commit history but only trees\nyou need/want; togglable to both sides.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"586","messageId":"1113785521.11910.45.camel@localhost.localdomain","threadId":"48","inReplyTo":"20050418005032.GE1461@pasky.ji.cz","subject":"Re: full kernel history, in patchset format","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-18T00:51:59Z","receivedAt":"2005-04-18T00:51:59Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Mon, 2005-04-18 at 02:50 +0200, Petr Baudis wrote:\n> I think I will make git-pasky's default behaviour (when we get\n> http-pull, that is) to keep the complete commit history but only trees\n> you need/want; togglable to both sides.\n\nI think the default behaviour should probably be to fetch everything.\n\n-- \ndwmw2\n\n"},{"id":"588","messageId":"20050418005955.GG1461@pasky.ji.cz","threadId":"48","inReplyTo":"1113785521.11910.45.camel@localhost.localdomain","subject":"Re: full kernel history, in patchset format","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-04-18T00:59:55Z","receivedAt":"2005-04-18T00:59:55Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, Apr 18, 2005 at 02:51:59AM CEST, I got a letter\nwhere David Woodhouse <dwmw2@infradead.org> told me that...\n> On Mon, 2005-04-18 at 02:50 +0200, Petr Baudis wrote:\n> > I think I will make git-pasky's default behaviour (when we get\n> > http-pull, that is) to keep the complete commit history but only trees\n> > you need/want; togglable to both sides.\n> \n> I think the default behaviour should probably be to fetch everything.\n\nI think fetching gigs of data just won't work for many people,\nespecially if they could do with a fraction of that.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"592","messageId":"Pine.LNX.4.58.0504171816040.7211@ppc970.osdl.org","threadId":"48","inReplyTo":"20050418003526.GD1461@pasky.ji.cz","subject":"Re: full kernel history, in patchset format","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-18T01:16:43Z","receivedAt":"2005-04-18T01:16:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 18 Apr 2005, Petr Baudis wrote:\n\n> Dear diary, on Mon, Apr 18, 2005 at 02:06:43AM CEST, I got a letter\n> where David Woodhouse <dwmw2@infradead.org> told me that...\n> > On Mon, 2005-04-18 at 01:39 +0200, Petr Baudis wrote:\n> > > Of course an entirely different thing are _trees_ associated with those\n> > > commits. As long as you stay with a simple three-way merge, you\n> > > basically never want to look at trees which aren't heads and which you\n> > > don't specifically request to look at. And the trees and what they carry\n> > > inside is the main bulk of data.\n> > \n> > If the trees are absent and you're trying to merge, what do you gain\n> > from having the commit objects?\n> \n> merge-base\n\nAlternatively, you can have just the rev-tree cache of them. That's what\nit was designed for (along with avoiding to have to read 60,000 commits).\n\n\t\tLinus\n"},{"id":"597","messageId":"1113788210.11910.52.camel@localhost.localdomain","threadId":"48","inReplyTo":"Pine.LNX.4.58.0504171816040.7211@ppc970.osdl.org","subject":"Re: full kernel history, in patchset format","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2005-04-18T01:36:48Z","receivedAt":"2005-04-18T01:36:48Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Sun, 2005-04-17 at 18:16 -0700, Linus Torvalds wrote:\n> Alternatively, you can have just the rev-tree cache of them. That's what\n> it was designed for (along with avoiding to have to read 60,000 commits).\n\nPurely from a conceptual POV I'd be a little happier with the history\njust ending with a parent pointer to a commit object which is absent,\nrather than having commit objects which point to _trees_ which are\nabsent. But I suppose I can't really justify that, and I'm not overly\nbothered about it either.\n\nThe important thing to get right at this point is that the tree we all\nwork with should refer to the history, regardless of how we choose to\nprune it. The current linux-2.6.git tree has a parentless commit for the\n2.6.12-rc2 import, which is bad. We should start with Thomas' git tree\nrepresenting the real history, and work from that. You don't even need\nto see his tree; you only need the final sha1 hash of the commit in his\ntree which matches 2.6.12-rc2, so you can use that as the 'parent' of\nthe first change you import yourself.\n\n-- \ndwmw2\n\n"},{"id":"638","messageId":"tnx4qe49z4h.fsf@arm.com","threadId":"48","inReplyTo":"20050416131528.GB19908@elte.hu","subject":"Re: full kernel history, in patchset format","fromName":"Catalin Marinas","fromEmail":"catalin.marinas@arm.com","sentAt":"2005-04-18T10:07:42Z","receivedAt":"2005-04-18T10:07:42Z","isPatch":false,"sender":{"key":"catalin.marinas@arm.com","avatar":null},"body":"Ingo Molnar <mingo@elte.hu> wrote:\n> i've converted the Linux kernel CVS tree into 'flat patchset' format, \n> which gave a series of 28237 separate patches. (Each patch represents a \n> changeset, in the order they were applied. I've used the cvsps\n> utility.)\n\nAFAIK, cvsps uses the date/time to create the changesets. There is a\nproblem with the BKCVS export since some files in the same commit can\nhave a different time (by an hour). I posted a mail some time ago\nabout this - \nhttp://marc.theaimsgroup.com/?l=linux-kernel&m=110026570201544&w=2\n\nI read that the old history won't be merged into the new repository\nbut, if you are interested, I have a script that can do this based on\nthe \"(Logical change ...)\" string in the file commit logs and it is\nquite fast at generating the patches.\n\n-- \nCatalin\n\n"},{"id":"715","messageId":"426437A2.7080903@cobite.com","threadId":"48","inReplyTo":"tnx4qe49z4h.fsf@arm.com","subject":"Re: full kernel history, in patchset format","fromName":"David Mansfield","fromEmail":"david@cobite.com","sentAt":"2005-04-18T22:41:38Z","receivedAt":"2005-04-18T22:41:38Z","isPatch":false,"sender":{"key":"david@cobite.com","avatar":null},"body":"Catalin Marinas wrote:\n> Ingo Molnar <mingo@elte.hu> wrote:\n> \n>>i've converted the Linux kernel CVS tree into 'flat patchset' format, \n>>which gave a series of 28237 separate patches. (Each patch represents a \n>>changeset, in the order they were applied. I've used the cvsps\n>>utility.)\n> \n> \n> AFAIK, cvsps uses the date/time to create the changesets. There is a\n> problem with the BKCVS export since some files in the same commit can\n> have a different time (by an hour). I posted a mail some time ago\n> about this - \n> http://marc.theaimsgroup.com/?l=linux-kernel&m=110026570201544&w=2\n> \n> I read that the old history won't be merged into the new repository\n> but, if you are interested, I have a script that can do this based on\n> the \"(Logical change ...)\" string in the file commit logs and it is\n> quite fast at generating the patches.\n> \n\nHmmm.  I read that message just now.  Is it a matter of 'perfection' \nthat is the issue here, or actual correctness when applying the patches \nin order?\n\n(perhaps this has now been fixed).\n\n\n\nDavid\n\n"},{"id":"795","messageId":"tnxfyxngo8e.fsf@arm.com","threadId":"48","inReplyTo":"426437A2.7080903@cobite.com","subject":"Re: full kernel history, in patchset format","fromName":"Catalin Marinas","fromEmail":"catalin.marinas@arm.com","sentAt":"2005-04-19T08:33:21Z","receivedAt":"2005-04-19T08:33:21Z","isPatch":false,"sender":{"key":"catalin.marinas@arm.com","avatar":null},"body":"David Mansfield <david@cobite.com> wrote:\n> Catalin Marinas wrote:\n>> AFAIK, cvsps uses the date/time to create the changesets. There is a\n>> problem with the BKCVS export since some files in the same commit can\n>> have a different time (by an hour). I posted a mail some time ago\n>> about this -\n>> http://marc.theaimsgroup.com/?l=linux-kernel&m=110026570201544&w=2\n>> I read that the old history won't be merged into the new repository\n>> but, if you are interested, I have a script that can do this based on\n>> the \"(Logical change ...)\" string in the file commit logs and it is\n>> quite fast at generating the patches.\n>>\n>\n> Hmmm.  I read that message just now.  Is it a matter of 'perfection'\n> that is the issue here, or actual correctness when applying the\n> patches in order?\n\nI see it as a matter of correctness since in a given BKCVS changeset\n(i.e. revision in the ChangeSet,v file) you may miss files. You would\neventually get them, with the same log, but in a different patch. If\nyou don't care about this, you can call it 'perfection'.\n\nAt that time I thought about modifying cvsps to use the \"(Logical\nchange ...)\" string instead of time/date for grouping the files but I\nrealised it is easier with a shell script.\n\n> (perhaps this has now been fixed).\n\nThere was no reply to this e-mail. It might have been fixed in the\nmeantime but I don't think the history was fixed as well.\n\n-- \nCatalin\n\n"}]}