{"thread":{"id":"52511","subject":"Possible improvement in DB structure","startedAt":"2019-12-23T13:22:02Z","lastAt":"2019-12-23T21:41:30Z","messageCount":4,"participants":["Arnaud Bertrand","brian m. carlson","Jonathan Nieder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"388806","messageId":"CAEW0o+gwbNyDqmiouFzO16LsRUfcAnSwj9K77oGe5hi=EVMB=w@mail.gmail.com","threadId":"52511","inReplyTo":null,"subject":"Possible improvement in DB structure","fromName":"Arnaud Bertrand","fromEmail":"xda@abalgo.com","sentAt":"2019-12-23T13:00:46Z","receivedAt":"2019-12-23T13:22:02Z","isPatch":false,"sender":{"key":"xda@abalgo.com","avatar":null},"body":"Hello,\n\nAccording to my understanding, git has only 3 kinds of objects:\n(excluding the packed version)\n- the blobs\n- the trees\n- the commits\n\nToday to parse all objects of the same type, it is necessary to parse\nall the objects and test them one by one.\n\nIt should be so simple to organize objects in\n.git/objects/blobs\n.git/objects/trees\n.git/object/commits\n\nMay be due to my limited knowledge of git, I don't see any advantage\nto put everything together.\nBy splitting the objects directory, the gain in performance could be\nimportant, the scripts simplified, the representation more clear.\n\nTo be backward compatible, we can imagine a get-object() function that parses\n.git/objects/blobs\n.git/objects/trees\n.git/object/commits\nand, when not found\n.git/objects\n\nA get-tree() function that first parses\ngit/objects/trees\nand when not found\n.git/objects\n\nidem for getblob() and getcommit()\n\nIs there a reason that I don't understand behind the decision to put\neverything together ?\n\nBest regards,\n\nArnaud Bertrand\n"},{"id":"388829","messageId":"20191223190950.GA6240@camp.crustytoothpaste.net","threadId":"52511","inReplyTo":"CAEW0o+gwbNyDqmiouFzO16LsRUfcAnSwj9K77oGe5hi=EVMB=w@mail.gmail.com","subject":"Re: Possible improvement in DB structure","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2019-12-23T19:09:50Z","receivedAt":"2019-12-23T19:09:58Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2019-12-23 at 13:00:46, Arnaud Bertrand wrote:\n> Hello,\n> \n> According to my understanding, git has only 3 kinds of objects:\n> (excluding the packed version)\n> - the blobs\n> - the trees\n> - the commits\n\nThere are also tags.\n\n> Today to parse all objects of the same type, it is necessary to parse\n> all the objects and test them one by one.\n\nThis isn't a behavior we often want.  Can you say more about why you\nwant to do this?\n\n> May be due to my limited knowledge of git, I don't see any advantage\n> to put everything together.\n> By splitting the objects directory, the gain in performance could be\n> important, the scripts simplified, the representation more clear.\n\nOftentimes, we want to look up an item that we would refer to as a\ntree-ish.  That means that any tag, commit, or tree can be used in this\ncase and it will automatically be resolved to the appropriate tree.\n\nCurrently, we can look for any loose object, and then look for any\npacked object, which is a limited number of lookups (at most, the number\nof packs plus one).  Your proposal would have us look up at most the\nnumber of packs plus six.\n\nIn addition, we sometimes know that we need to look up an object, but\ndon't know its type.  We would incur additional costs in this case as\nwell.\n\nI'm not sure that we would gain a lot other than conceptual tidiness,\nbut we would incur additional performance costs.  We can currently\ndistinguish between the type of all of these objects by simply reading\nthe object header, which on a 64-bit system cannot exceed 28 bytes,\nwhich we do in some cases, such as `git cat-file --batch`.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"388830","messageId":"CAEW0o+jRW8LJqfjsDVtUiSNxwM9yBkj0c=Ddy3kEGUdsYM8myQ@mail.gmail.com","threadId":"52511","inReplyTo":"20191223190950.GA6240@camp.crustytoothpaste.net","subject":"Re: Possible improvement in DB structure","fromName":"Arnaud Bertrand","fromEmail":"xda@abalgo.com","sentAt":"2019-12-23T20:46:06Z","receivedAt":"2019-12-23T21:09:09Z","isPatch":false,"sender":{"key":"xda@abalgo.com","avatar":null},"body":"Hello Brian,\n\nToday, I think that tags are not located in objects directory but in\nrefs/tags which is a good idea.;-)\n\nThe origin of my reflection was that I wanted to find an old file.\n\nI knew that in the past of my project, we had started to write a\ndriver for a device and it was abandoned. I wanted to find this file.\nI knew a \"key line\" to search for and I knew the file was a .c file\nbut I didn't know the exact name.\n\nSo, the goal was to parse all the database, find all the different .c\nfiles and grep it to find the the driver.\n\nAnd there began the problems.... I had a huge database and I've\nwritten a script that had to:\n1. Identify all the trees (straight forward if all trees are in objects/trees)\n2. In each trees, identify all different *.c files\n3. grep \"key line\" in them\n\nWell, as I said, I had a huge database and I took a long time to get\nthe information.\n\nIf the objects had been separated directly, it would have been much simpler.\n\nIt is just an example, finally, I've written a cron job that unpacks\neverything and saves all the trees sha in a file that can be parsed by\nscripts.\n\nSo, a small change in the db structure could be very helpful for this\nkind of needs.\n\nAbout the fact that searching for an arbitrary object will consume\nmore time... It's very rare to look for an object without knowing it's\ntype, and parsing 3 subdirs instead of one is not so time consuming by\ncomparison of the operation described above.\n\nArnaud Bertrand, Belgium\n\n\n\nLe lun. 23 déc. 2019 à 20:10, brian m. carlson\n<sandals@crustytoothpaste.net> a écrit :\n>\n> On 2019-12-23 at 13:00:46, Arnaud Bertrand wrote:\n> > Hello,\n> >\n> > According to my understanding, git has only 3 kinds of objects:\n> > (excluding the packed version)\n> > - the blobs\n> > - the trees\n> > - the commits\n>\n> There are also tags.\n>\n> > Today to parse all objects of the same type, it is necessary to parse\n> > all the objects and test them one by one.\n>\n> This isn't a behavior we often want.  Can you say more about why you\n> want to do this?\n>\n> > May be due to my limited knowledge of git, I don't see any advantage\n> > to put everything together.\n> > By splitting the objects directory, the gain in performance could be\n> > important, the scripts simplified, the representation more clear.\n>\n> Oftentimes, we want to look up an item that we would refer to as a\n> tree-ish.  That means that any tag, commit, or tree can be used in this\n> case and it will automatically be resolved to the appropriate tree.\n>\n> Currently, we can look for any loose object, and then look for any\n> packed object, which is a limited number of lookups (at most, the number\n> of packs plus one).  Your proposal would have us look up at most the\n> number of packs plus six.\n>\n> In addition, we sometimes know that we need to look up an object, but\n> don't know its type.  We would incur additional costs in this case as\n> well.\n>\n> I'm not sure that we would gain a lot other than conceptual tidiness,\n> but we would incur additional performance costs.  We can currently\n> distinguish between the type of all of these objects by simply reading\n> the object header, which on a 64-bit system cannot exceed 28 bytes,\n> which we do in some cases, such as `git cat-file --batch`.\n> --\n> brian m. carlson: Houston, Texas, US\n> OpenPGP: https://keybase.io/bk2204\n"},{"id":"388831","messageId":"20191223214125.GA38316@google.com","threadId":"52511","inReplyTo":"CAEW0o+jRW8LJqfjsDVtUiSNxwM9yBkj0c=Ddy3kEGUdsYM8myQ@mail.gmail.com","subject":"Re: Possible improvement in DB structure","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2019-12-23T21:41:25Z","receivedAt":"2019-12-23T21:41:30Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi Arnaud,\n\nArnaud Bertrand wrote:\n\n> Today, I think that tags are not located in objects directory but in\n> refs/tags which is a good idea.;-)\n\nNot precisely.  See \"git help repository-layout\" for more details, or\nhttps://www.kernel.org/pub/software/scm/git/docs/user-manual.html#hacking-git\nor the \"git internals\" chapter of https://git-scm.com/book/.\n\n> The origin of my reflection was that I wanted to find an old file.\n>\n> I knew that in the past of my project, we had started to write a\n> driver for a device and it was abandoned. I wanted to find this file.\n> I knew a \"key line\" to search for and I knew the file was a .c file\n> but I didn't know the exact name.\n\nThanks for this context!  It's very helpful.\n\n> So, the goal was to parse all the database, find all the different .c\n> files and grep it to find the the driver.\n\nGit intends to make this kind of history mining not too difficult.\nYou can run a command like\n\n\tgit log --all -S'the key line' -- '*.c'\n\nand it should do the right thing.  Or you can do something more\ncomplex using something like \"git rev-list --all | git diff-tree\n--stdin --name-only --diff-filter=D\" (to show deleted files).\n\nIs the problem that that command is too slow?\n\nHope that helps,\nJonathan\n"}]}