{"thread":{"id":"14909","subject":"[RFC] Plumbing-only support for storing object metadata","startedAt":"2008-08-09T21:07:33Z","lastAt":"2008-08-18T23:28:00Z","messageCount":22,"participants":["Jamey Sharp","Scott Chacon","Shawn O. Pearce","Jan Hudec","Stephen R. van den Berg","david@lang.hm","Junio C Hamano","Josh Triplett","Derek Fawcus","Marcus Griep"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"86615","messageId":"20080809210733.GA6637@oh.minilop.net","threadId":"14909","inReplyTo":null,"subject":"[RFC] Plumbing-only support for storing object metadata","fromName":"Jamey Sharp","fromEmail":"jamey@minilop.net","sentAt":"2008-08-09T21:07:33Z","receivedAt":"2008-08-09T21:07:33Z","isPatch":false,"sender":{"key":"jamey@minilop.net","avatar":"https://gravatar.com/avatar/979ed4f9190c19c13a84945afe9a6cf141e8884ef6ddaae342db7a74ae28828e?d=mp&s=160"},"body":"The attached test illustrates a proposal for minimal plumbing support\nusable to store permissions, ownership, and other metadata in git\nrepositories. This proposal is fully compatible with existing\nrepositories when the new functionality is not in use. Similar to the\nintroduction of subprojects, we have not yet specified the porcelain. We\nbelieve that the plumbing will provide sufficient functionality for many\nuses, and these uses will help determine the appropriate porcelain.\n\nWe would have included an implementation along with the test, but we\nneed help with a detail of git internals. More on that at the end. We'd\nalso appreciate feedback on the proposal.\n\nWe propose representing objects with metadata using a new \"inode\"\nobject. An inode object contains the hash of the real object and the\nhash of a \"props\" (properties) object. A props object contains a set of\nname-value pairs. Tree objects can reference inode objects in addition\nto the current possibilities of blobs, trees, and subproject commits; we\npropose using the currently invalid type 110000 (S_IFREG | S_IFIFO) for\ninode objects. We primarily see a use case for inodes referencing blobs\nand trees, though as defined they support any object type.\n\nBy separating property objects from inodes, objects with the same\nproperties can share the same property object; we expect, for instance,\nthat repositories reflecting /etc will have many references to the\n\"root:root 644\" and \"root:root 755\" properties.\n\nBoth object types have a unique representation: equivalent inodes and\nprops objects will have the same hash. The exact format of an inode\nlooks like:\n\t<object_type> SP <object_sha1> LF\n\tprops SP <props_sha1> LF\nA property object looks like a sorted list of one or more of:\n\t<key> SP <value> LF\nThe same key is allowed to appear more than once, in which case the\nlines will be sorted by the bytes of the values. Allowing duplicate keys\nwill make it easier to retrieve a set of similar properties such as\nacls.\n\nThis format implies certain constraints on property names and values. We\npropose limiting both names and values to printable ASCII (\\x20-\\x7E),\nand disallowing spaces in keys. If some use case requires property names\nor values with binary data, that property could use a printable encoding\nsuch as base64.\n\nWe believe this proposal provides a sensible approach to storing\nmetadata in Git repositories; however, we're happy with any reasonable\nsolution that provides equivalent functionality. Some alternatives we\nconsidered:\n\n  - We could allow UTF-8 property names or values, rather than strictly\n    ASCII. Our proposal is conservative in this regard, allowing an\n    extension to UTF-8 later while remaining compatible with existing\n    repositories.\n\n  - We could allow arbitrary property names or values, by changing the\n    props format to store lengths rather than using delimiters. This\n    would not be a compatible change, so it needs to be decided early.\n\n  - Tree objects already store mode bits, but we believe that it would\n    prove simpler to store complete modes in properties rather than\n    adjusting Git internals to preserve arbitrary mode bits in trees.\n    Even if new versions of Git preserved the full mode, existing\n    versions of Git might silently give incorrect results. Furthermore,\n    mode bits other than executability seem of limited value without\n    ownership information.\n\n  - inode objects could directly store properties, rather than\n    referencing a separate props object. This would eliminate one\n    indirection needed to access properties. However, it would also\n    reduce sharing of data for objects with the same properties.\n    Furthermore, we expect that the indirection will have negligible\n    cost when accessing objects from packs, given appropriately sorted\n    packs. Shared props objects also suggest caching at various layers.\n\n  - We could have called them \"meta\" objects instead of \"props\", but\n    then we couldn't make \"mad props\" jokes.\n\nWe began trying to implement this proposal, but we found this enum\ndefinition in cache.h, which made us think there's only room for one\nmore kind of object:\n\n\tenum object_type {\n\t\tOBJ_BAD = -1,\n\t\tOBJ_NONE = 0,\n\t\tOBJ_COMMIT = 1,\n\t\tOBJ_TREE = 2,\n\t\tOBJ_BLOB = 3,\n\t\tOBJ_TAG = 4,\n\t\t/* 5 for future expansion */\n\t\tOBJ_OFS_DELTA = 6,\n\t\tOBJ_REF_DELTA = 7,\n\t\tOBJ_ANY,\n\t\tOBJ_MAX,\n\t};\n\nDo these object_type values appear in any on-disk structure, or does any\nother reason exist why this set of values cannot change? Can we add\nadditional object types for inodes and props? If not, what would you\nrecommend instead?\n\n- Jamey Sharp and Josh Triplett\n\n\n#!/bin/sh\n#\n# Copyright (c) 2008 Josh Triplett and Jamey Sharp\n#\n\ntest_description=\"Test inode plumbing\"\n\n. ./test-lib.sh\n\ncat > shadow <<EOF\nroot:*:13943:0:99999:7:::\nEOF\nshadow_sha1=`git hash-object -t blob -w shadow`\n\ncat > props <<EOF\ngroup shadow\nmode 640\nowner root\nEOF\nprops_sha1=FIXME\n\ncat > inode <<EOF\nblob $shadow_sha1\nprops $props_sha1\nEOF\ninode_sha1=FIXME\n\ncat > tree <<EOF\n110644 inode $inode_sha1\tshadow\nEOF\ntree_sha1=FIXME\n\ntest_expect_success 'hash a props' '\n\ttest $props_sha1 = \"`git hash-object -t props -w props`\"\n'\n\ntest_expect_success 'cat-file a props' '\n\tgit cat-file props $props_sha1 | cmp -s - props\n'\n\ntest_expect_success 'hash an inode' '\n\ttest $inode_sha1 = \"`git hash-object -t inode -w inode`\"\n'\n\ntest_expect_success 'cat-file an inode' '\n\tgit cat-file inode $inode_sha1 | cmp -s - inode\n'\n\ntest_expect_success 'tree with inode' '\n\ttest $tree_sha1 = \"`git mktree < tree`\"\n'\n\ntest_expect_success 'ls-tree of tree with inode' '\n\tgit ls-tree $tree_sha1 | cmp -s - tree\n'\n\ntest_expect_success 'check type with cat-file' '\n\ttest inode = \"`git cat-file -t $tree_sha1:shadow`\"\n'\n\ntest_expect_success 'cat-file inode tree:inode' '\n\tgit cat-file inode $tree_sha1:shadow | cmp -s - inode\n'\n\ntest_expect_success 'cat-file blob tree:inode' '\n\tgit cat-file blob $tree_sha1:shadow | cmp -s - shadow\n'\n\ntest_expect_success 'cat-file props tree:inode' '\n\tgit cat-file props $tree_sha1:shadow | cmp -s - props\n'\n\ntest_expect_success 'read-tree' '\n\tgit read-tree $tree_sha1\n'\n\ntest_expect_success 'ls-files shows no modified files' '\n\ttest -z \"`git ls-files -m || echo fail`\"\n'\n\ntest_expect_success 'write-tree' '\n\ttest $tree_sha1 = \"`git write-tree`\"\n'\n\ntest_expect_success 'commit-tree' '\n\tCOMMIT=`echo Commit with an inode | git commit-tree $tree_sha1` &&\n\tgit update-ref HEAD $COMMIT\n'\n\ncat >shadow <<EOF\nroot:*:13943:0:99999:7:::\njamey:*:13943:0:99999:7:::\njosh:*:13943:0:99999:7:::\nEOF\nshadow_sha1=FIXME\n\ntest_expect_success 'ls-files shows modified file' '\n\ttest \"shadow\" = \"`git ls-files -m`\"\n'\n\ntest_expect_success 'add modified file to index' '\n\tgit add shadow\n'\n\ntest_expect_success 'commit modification' '\n\tgit commit -m \"Modify shadow\"\n'\n\ntest_expect_success 'ls-files shows no modified files' '\n\ttest -z \"`git ls-files -m || echo fail`\"\n'\n\ntest_expect_success 'check type with cat-file, after modification' '\n\ttest inode = \"`git cat-file -t HEAD:shadow`\"\n'\n\ncat > inode <<EOF\nblob $shadow_sha1\nprops $props_sha1\nEOF\ninode_sha1=FIXME\n\ntest_expect_success 'cat-file inode HEAD:inode, after modification' '\n\tgit cat-file inode HEAD:shadow | cmp -s - inode\n'\n\ntest_expect_success 'cat-file blob HEAD:inode, after modification' '\n\tgit cat-file blob HEAD:shadow | cmp -s - shadow\n'\n\ntest_expect_success 'cat-file props HEAD:inode, after modification' '\n\tgit cat-file props HEAD:shadow | cmp -s - props\n'\n\ntest_done\n"},{"id":"86617","messageId":"d411cc4a0808091449n7e0c9b7et7980cf668106aead@mail.gmail.com","threadId":"14909","inReplyTo":"20080809210733.GA6637@oh.minilop.net","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Scott Chacon","fromEmail":"schacon@gmail.com","sentAt":"2008-08-09T21:49:20Z","receivedAt":"2008-08-09T21:49:20Z","isPatch":false,"sender":{"key":"schacon@gmail.com","avatar":"https://gravatar.com/avatar/9b13a8a078e1dcf8588c4eea9554445d51ebed6c41b51f56f4d96738130b05c6?d=mp&s=160"},"body":"> We began trying to implement this proposal, but we found this enum\n> definition in cache.h, which made us think there's only room for one\n> more kind of object:\n>\n>        enum object_type {\n>                OBJ_BAD = -1,\n>                OBJ_NONE = 0,\n>                OBJ_COMMIT = 1,\n>                OBJ_TREE = 2,\n>                OBJ_BLOB = 3,\n>                OBJ_TAG = 4,\n>                /* 5 for future expansion */\n>                OBJ_OFS_DELTA = 6,\n>                OBJ_REF_DELTA = 7,\n>                OBJ_ANY,\n>                OBJ_MAX,\n>        };\n>\n> Do these object_type values appear in any on-disk structure, or does any\n> other reason exist why this set of values cannot change? Can we add\n> additional object types for inodes and props? If not, what would you\n> recommend instead?\n\nIf I'm not mistaken, these are the values used to identify data in the\nheader sections of packfile objects.  The first four bits are used to\nidentify the object type, where the first bit is static and the next\nthree are the object type of the data following the header.  Since the\ntype is encoded using those three bits, 0-7 is the valid range.  I\nwould assume that would be difficult to change, since all the\npackfiles depend on that range.\n\nScott\n"},{"id":"86634","messageId":"20080810035101.GA22664@spearce.org","threadId":"14909","inReplyTo":"d411cc4a0808091449n7e0c9b7et7980cf668106aead@mail.gmail.com","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-10T03:51:01Z","receivedAt":"2008-08-10T03:51:01Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Scott Chacon <schacon@gmail.com> wrote:\n> > We began trying to implement this proposal, but we found this enum\n> > definition in cache.h, which made us think there's only room for one\n> > more kind of object:\n> >\n> >        enum object_type {\n> >                OBJ_BAD = -1,\n> >                OBJ_NONE = 0,\n> >                OBJ_COMMIT = 1,\n> >                OBJ_TREE = 2,\n> >                OBJ_BLOB = 3,\n> >                OBJ_TAG = 4,\n> >                /* 5 for future expansion */\n> >                OBJ_OFS_DELTA = 6,\n> >                OBJ_REF_DELTA = 7,\n> >                OBJ_ANY,\n> >                OBJ_MAX,\n> >        };\n> >\n> > Do these object_type values appear in any on-disk structure, or does any\n> > other reason exist why this set of values cannot change? Can we add\n> > additional object types for inodes and props? If not, what would you\n> > recommend instead?\n> \n> If I'm not mistaken, these are the values used to identify data in the\n> header sections of packfile objects.  The first four bits are used to\n> identify the object type, where the first bit is static and the next\n> three are the object type of the data following the header.  Since the\n> type is encoded using those three bits, 0-7 is the valid range.  I\n> would assume that would be difficult to change, since all the\n> packfiles depend on that range.\n\nCorrect.  There is only room in the pack file for 3 bits in the\ntype field, resulting in types 0-7 as being the only valid range.\n\nOnly type 0 and 5 are available for use.\n\nNico and I have (at least in the past) agreed that type 0 is meant\nas an escape indicator.  If the type is set to 0 then the real type\ncode appears in another byte of data which follows the object's\ninflated length.\n\nThat leaves only type 5 available.  Note that because type 5 can be\nencoded into a really small space (3 bits) compared to any other\ntype we may add we really want to use it for something which will\nappear _very_frequently_.  The OBJ_DICT_TREE encoding we were talking\nabout doing for pack v4 fits that bill, as nearly any project (even\nhuge ones like Mozilla or KDE) would probably be using OBJ_DICT_TREE\nthoughout their pack files, and there is a noticable reduction in\ndisk usage (and increased performance due to lower page faults)\nas a result.\n\nThe proposed \"inode\" and \"props\" types sound like they are useful\nfor only less common cases, and would appear very infrequently\ncompared to a tree object.\n\nSo yea, there really aren't any new type bits available.\n\nBut tossing aside the type bit argument, I'm not sure I see the\nvalue in adding limited arbitrary properties to names in a tree.\nHow does one edit these?  How do you inspect them before you get\na checkout, assuming they might actually have an impact on the\ncheckout process?  How the hell do you merge them?\n\nI'm also very concerned about the limited range of values for both\nkeys and values in a \"props\" type.  Even if we did go down this\nroad of supporting such a concept at the plumbing layer (and in the\nstorage modal) everwhere else we are 8-bit clean.  Commit messages,\ntag messages, blob contents, even file names in tree objects.\n(OK, file names cannot contain a NUL byte, but whatever, that is\ntheir only limitation.)\n\nThe proper encoding for both keys and values should permit any data\nto be stored.  Doesn't the extended attributes feature in Linux and\nFreeBSD both support any data to be attached to an inode in the fs?\n\nPlease don't get me wrong.\n\nI think this is a _BAD_ idea.\n\nA bad idea that will only clutter up the core object model, and\nthe core processing code of that object model.  Extended attributes\naren't used that much on local filesystems, because they are hard\nto work with and suck performance wise.  Performance in Git is\na _feature_.  It matters.  Our clean object model really helps to\nmake that possible.\n\n-- \nShawn.\n"},{"id":"86660","messageId":"20080810110925.GB3955@efreet.light.src","threadId":"14909","inReplyTo":"20080809210733.GA6637@oh.minilop.net","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2008-08-10T11:09:25Z","receivedAt":"2008-08-10T11:09:25Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"Hello All,\n\nI am glad you came up with this, as I think this is the only reasonable way\nto support things like etckeeper. The metastore and similar solutions are\na kludge and fall apart in so many cases.\n\nI am not sure your approach is the right one, though. I tend to agree with\nShawn it's not. So here is a couple of alternate proposals (sorry, it's a bit\nlong, as I have several variants with different drawbacks I would like to\ndiscuss).\n\nOn Sat, Aug 09, 2008 at 14:07:33 -0700, Jamey Sharp wrote:\n> The attached test illustrates a proposal for minimal plumbing support\n> usable to store permissions, ownership, and other metadata in git\n> repositories. This proposal is fully compatible with existing\n> repositories when the new functionality is not in use. Similar to the\n> introduction of subprojects, we have not yet specified the porcelain. We\n> believe that the plumbing will provide sufficient functionality for many\n> uses, and these uses will help determine the appropriate porcelain.\n\nI think the main way to use it would be a hook, that would read/write the\nattributes to/from the tree. That will do the right thing for storing\npermissions, owners and other things represented in the worktree. And\nmetadata that are neither part of the tree or directly related to git's\nfunctionality are out of our scope.\n\n> [...]\n> We propose representing objects with metadata using a new \"inode\"\n> object. An inode object contains the hash of the real object and the\n> hash of a \"props\" (properties) object. A props object contains a set of\n> name-value pairs. Tree objects can reference inode objects in addition\n> to the current possibilities of blobs, trees, and subproject commits; we\n> propose using the currently invalid type 110000 (S_IFREG | S_IFIFO) for\n> inode objects. We primarily see a use case for inodes referencing blobs\n> and trees, though as defined they support any object type.\n\nI think this is the overly complex -- and also the needlessly incompatible\npart. By the way, I don't think you need separate type for props -- it can be\na blob too.\n\nI would suggest investigating following options:\n\n 1. It would be possible to use clean/smudge filters to encode the attributes\n    in the blob itself.\n 2. Store the metadata in separate objects, but link them in the parent tree\n    directly. In this case, each attribute could probably get it's own blob,\n    so eg. for a file foo the tree containing it would have entries:\n      foo\n      foo<sep>owner\n      foo<sep>permissions\n      ...\n    Where <sep> would be some sepatator (more on that below).\n\nAdvantages (+), disadvantages (-) and possible (*) extensions of 1:\n \n + It should be possible to get to something useful with very little changes\n   to git. Basically all it needs to be useful for things like etckeeper is\n   to:\n    . Make sure both clean and smudge filter always get filehandle to the\n      disk file in question (I am /not/ suggesting path as the file may be\n      written in a staging area and moved into place later).\n    . Pass the blob id currently in index to the clean filter, so it can\n      maintain the data if they are not representable in this particular\n      checkout (eg. when checking out such repo on windows). Note, that this\n      would also be useful for ignoring insignificant changes, eg. when\n      a in some config file order is not important and the tool using it\n      randomly changes that order when changing that file.\n\n - It does not support metadata for directories, but could be crossed with\n   approach 2 to fix that. Git could special-case entry '.' for storing\n   \"content\" of a directory, which would be wholly created by running the\n   clean filter on a directory (I am not sure directory handles are portable,\n   but running with that directory as current should be). This would not have\n   the problem of approach 2 with the entry names for the metadata.\n\n * Default processing could be added to strip the metadata in smudge and\n   re-add them from index on clean. This would require adding some marker to\n   know which blobs need this treatment. I see two ways:\n    . Using different file type for them. There are already two types\n      pointing to blob (S_IFREG and S_IFLNK) and they are treated differently\n      on read (clean) / write (smudge) from/to tree, so third type should be\n      workable.\n    . Using additional format. Currently a blob is encoded as\n\t\"blob\" <LF> <content>\n      so maybe an extneded blob could be encoded as\n\t\"blob extended\" <LF> <content>\n      without needing a special type for it. But I don't know git internals\n      enough to know how easy, hard or dirty this would be.\n\nAdvantages (+), disadvantages (-) and possible (*) extensions of 2:\n\n + It would work the same way for directories and file, or mostly so.\n + Different metadata would be handled independently, so it would be easier\n   to combine support for multiple attributes (not that I can imagine any\n   sensible use beyond access lists (owner, permission, posix acl)).\n + Checking out without the hooks could easily create special metadata files,\n   providing easy way to work with the attributes where they are not\n   supported by the underlaying filesystem.\n - It would require reserving some names for the metadata entries. I see\n   basically three ways to name the attribues:\n    . Reserving some character for the separator, eg. @ or # or something\n      like that. So with file foo, there would be entries:\n        foo\n\tfoo@owner\n\tfoo@permissions\n      This has following pros and cons:\n       + Minimal changes to the index <-> tree logic (remember, index is\n         flat and has no directory entries, so the tree writer must decide\n\t to which tree each entry goes).\n       + Trivially supports checking the metadata entries out as special\n         files on filesystem without metadata support.\n       - The character is reserved in trees that need the feature (the trees\n\t that don't need it don't need to care).\n      Note, that the metadata entries could have mode either S_ISREG, or\n      a new one. Inclined to say S_ISREG -- we have the special name to\n      distinguish them.\n    . Using something that does not exist in a normalized path, ie either\n      \"//\" or \"/./\". So with file foo, there would be entries:\n        foo\n\tfoo//owner\n\tfoo//permissions\n      This has following pros and cons:\n       + Does not reserve any characters. Every filename is permitted even\n         when the freature is used.\n       - Harder on the index <-> tree logic, as it would have to not consider\n\t such strings as not being directory separators.\n       - Such files could not be checked out, though they could still be\n\t manipulated using cat-file and update-index.\n      The metadata entries could have mode either S_ISREG or a new one again.\n      New mode would be sensible if it would make easier on the index <->\n      tree logic (it's easier to check 3 bits than search string for\n      a substring).\n    . Leave the suffix for metadata entries to the hooks. This would be\n      middle road between the above two:\n       + Reserves as little as possible, while not complicating the index <->\n         tree logic.\n       + Remains easy to check out as special files where you can't run the\n         hooks, though this would require some special-casing similar to\n\t symlinks on Windows.\n       - Would require new mode for these entries, so we know they are\n\t created and consumed by the hooks rather than directly read/written\n\t to the tree.\n\nBest regards,\nJan\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"86662","messageId":"20080810112038.GB30892@cuci.nl","threadId":"14909","inReplyTo":"20080810035101.GA22664@spearce.org","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-08-10T11:20:38Z","receivedAt":"2008-08-10T11:20:38Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Shawn O. Pearce wrote:\n>The proper encoding for both keys and values should permit any data\n>to be stored.  Doesn't the extended attributes feature in Linux and\n>FreeBSD both support any data to be attached to an inode in the fs?\n\nI'd think so yes, so any attempt to store the metadata should support it\nas well.\nThat also would imply that any such metadata storage would have to allow\nfor arbitrary blobs to be stored under tag-names.\nAnd *that* would imply that anything that implements a kludge like\nspecifying a flat-file format to encode name/value pairs doesn't scale.\n\n>I think this is a _BAD_ idea.\n\n>A bad idea that will only clutter up the core object model, and\n>the core processing code of that object model.  Extended attributes\n>aren't used that much on local filesystems, because they are hard\n>to work with and suck performance wise.  Performance in Git is\n>a _feature_.  It matters.  Our clean object model really helps to\n>make that possible.\n\nQuite right.\n\nHowever, pondering the idea a bit more, I could envision something\nsimilar to the following:\n\nIn the git tree the following layout would be used:\n\nplainfile.txt\notherdir/otherplainfile.txt\nprojects/README\nprojects/README/_owner\nprojects/README/_acl\nprojects/README/_icon\nprojects/README/_mimetype\nprojects/something.mpeg\nprojects/something.mpeg/_icon\nprojects/something.mpeg/_mimetype\nprojects/asubdir/thirdplainfile.txt\n\nThat would imply that in the tree storage, the only extension would be\nthat for any given reference to a blob in a tree object, there could be\na reference to a tree object as well.  I.e. something like this in the\ntree object:\n\n100644 blob f7b7414159b8a7159538fac543b2b19ef531968e  README\n000000 tree df6ee415f04d6ccea5dab0de562c2f155583a2c4  README\n100644 blob 0a54f8ec13df03cf6bdb5b973acec6d8141c01cc  something.mpeg\n000000 tree a421448d765abb7bb979dc1d56621d0fc9b41229  soemthing.mpeg\n\nThe extra tree reference for README would actually refer to something like:\n\n100644 blob be3365fdaae0f4ed8c22c4cf38a4b1f88f9069c3  _owner\n100644 blob 739e9e8f3d095931084b54cbf7f90d8f64eb0ac6  _acl\n100644 blob bc1a868bb50644712966a50150d21199c401d6d5  _icon\n100644 blob 6076bde5b3b6b8bed4ec4968d09abdbf015b3b75  _mimetype\n\nWhich would contain the extra attributes.\n\nAnd that would imply that during checkout you can do a rich checkout or a\nflat checkout for any files under the projects directory.\n\nA flat checkout results in the following files in the filesystem:\n\nplainfile.txt\notherdir/otherplainfile.txt\nprojects/README\nprojects/README.attr/_owner\nprojects/README.attr/_acl\nprojects/README.attr/_icon\nprojects/README.attr/_mimetype\nprojects/something.mpeg\nprojects/something.mpeg.attr/_icon\nprojects/something.mpeg.attr/_mimetype\nprojects/asubdir/thirdplainfile.txt\n\nA rich checkout results in the following files in the filesystem:\n\nplainfile.txt\notherdir/otherplainfile.txt\nprojects/README\nprojects/something.mpeg\nprojects/asubdir/thirdplainfile.txt\nprojects/asubdir/fourthplainfile.txt\n\nThe rich checkout also applies the extended attributes/metadata to the\nfilesystem (i.e. it would store all the metadata in the appropriate\nplaces).\n\nThe nice thing about this setup is that:\na. There is *no* change whatsoever to existing repositories or\n   repositoryformat.\nb. It's less filling (i.e. there are no special bits or object types to be\n   used).\nc. Speed for files without attributes is not affected.\nd. It's fully 8-bit-transparent.\ne. It scales, even if you have large or many attributes.\nf. It uses the natural tree storage abstraction already supported in\n   git repositories to store the additional data.\ng. It allows reuse of attribute information at many levels.\nh. It even allows for a hierarchy of attributes attached to a single\n   file (no current filesystem supports that (yet)).\ni. The only change in the fast-path of core-git is that it would have to\n   know how to skip tree objects referenced in a tree object if a\n   same-name blob object is already there.  This can even be optimised\n   by requiring the attribute-tree to have a very specific (e.g. 0)\n   mode to ease detection.\nj. Editing and merging the meta-information could be made an almost\n   natural operation in the flat-checkout mode (the extension to be used\n   to name the attribute subdir should be made configurable).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\nReal programmers don't produce results, they return exit codes.\n"},{"id":"86664","messageId":"alpine.DEB.1.10.0808100502530.32620@asgard.lang.hm","threadId":"14909","inReplyTo":"20080810112038.GB30892@cuci.nl","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-10T12:16:47Z","receivedAt":"2008-08-10T12:16:47Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Sun, 10 Aug 2008, Stephen R. van den Berg wrote:\n\n> Shawn O. Pearce wrote:\n>> The proper encoding for both keys and values should permit any data\n>> to be stored.  Doesn't the extended attributes feature in Linux and\n>> FreeBSD both support any data to be attached to an inode in the fs?\n>\n> I'd think so yes, so any attempt to store the metadata should support it\n> as well.\n> That also would imply that any such metadata storage would have to allow\n> for arbitrary blobs to be stored under tag-names.\n> And *that* would imply that anything that implements a kludge like\n> specifying a flat-file format to encode name/value pairs doesn't scale.\n>\n>> I think this is a _BAD_ idea.\n>\n>> A bad idea that will only clutter up the core object model, and\n>> the core processing code of that object model.  Extended attributes\n>> aren't used that much on local filesystems, because they are hard\n>> to work with and suck performance wise.  Performance in Git is\n>> a _feature_.  It matters.  Our clean object model really helps to\n>> make that possible.\n>\n> Quite right.\n>\n> However, pondering the idea a bit more, I could envision something\n> similar to the following:\n>\n> In the git tree the following layout would be used:\n>\n> plainfile.txt\n> otherdir/otherplainfile.txt\n> projects/README\n> projects/README/_owner\n> projects/README/_acl\n> projects/README/_icon\n> projects/README/_mimetype\n> projects/something.mpeg\n> projects/something.mpeg/_icon\n> projects/something.mpeg/_mimetype\n> projects/asubdir/thirdplainfile.txt\n>\n> That would imply that in the tree storage, the only extension would be\n> that for any given reference to a blob in a tree object, there could be\n> a reference to a tree object as well.  I.e. something like this in the\n> tree object:\n>\n> 100644 blob f7b7414159b8a7159538fac543b2b19ef531968e  README\n> 000000 tree df6ee415f04d6ccea5dab0de562c2f155583a2c4  README\n> 100644 blob 0a54f8ec13df03cf6bdb5b973acec6d8141c01cc  something.mpeg\n> 000000 tree a421448d765abb7bb979dc1d56621d0fc9b41229  soemthing.mpeg\n>\n> The extra tree reference for README would actually refer to something like:\n>\n> 100644 blob be3365fdaae0f4ed8c22c4cf38a4b1f88f9069c3  _owner\n> 100644 blob 739e9e8f3d095931084b54cbf7f90d8f64eb0ac6  _acl\n> 100644 blob bc1a868bb50644712966a50150d21199c401d6d5  _icon\n> 100644 blob 6076bde5b3b6b8bed4ec4968d09abdbf015b3b75  _mimetype\n>\n> Which would contain the extra attributes.\n>\n> And that would imply that during checkout you can do a rich checkout or a\n> flat checkout for any files under the projects directory.\n>\n> A flat checkout results in the following files in the filesystem:\n>\n> plainfile.txt\n> otherdir/otherplainfile.txt\n> projects/README\n> projects/README.attr/_owner\n> projects/README.attr/_acl\n> projects/README.attr/_icon\n> projects/README.attr/_mimetype\n> projects/something.mpeg\n> projects/something.mpeg.attr/_icon\n> projects/something.mpeg.attr/_mimetype\n> projects/asubdir/thirdplainfile.txt\n>\n> A rich checkout results in the following files in the filesystem:\n>\n> plainfile.txt\n> otherdir/otherplainfile.txt\n> projects/README\n> projects/something.mpeg\n> projects/asubdir/thirdplainfile.txt\n> projects/asubdir/fourthplainfile.txt\n>\n> The rich checkout also applies the extended attributes/metadata to the\n> filesystem (i.e. it would store all the metadata in the appropriate\n> places).\n>\n> The nice thing about this setup is that:\n> a. There is *no* change whatsoever to existing repositories or\n>   repositoryformat.\n> b. It's less filling (i.e. there are no special bits or object types to be\n>   used).\n> c. Speed for files without attributes is not affected.\n> d. It's fully 8-bit-transparent.\n> e. It scales, even if you have large or many attributes.\n> f. It uses the natural tree storage abstraction already supported in\n>   git repositories to store the additional data.\n> g. It allows reuse of attribute information at many levels.\n> h. It even allows for a hierarchy of attributes attached to a single\n>   file (no current filesystem supports that (yet)).\n> i. The only change in the fast-path of core-git is that it would have to\n>   know how to skip tree objects referenced in a tree object if a\n>   same-name blob object is already there.  This can even be optimised\n>   by requiring the attribute-tree to have a very specific (e.g. 0)\n>   mode to ease detection.\n> j. Editing and merging the meta-information could be made an almost\n>   natural operation in the flat-checkout mode (the extension to be used\n>   to name the attribute subdir should be made configurable).\n\nyou also need to be able to add something to the attribute tree to \nindicate what type of metadata is being stored in it. you could  have *nix \nperms, windows perms, posix extended attributes, or other things.\n\nI could see this as a great way to deal with editing exif data for images. \nwhen checking in a .jpg, extract the .exif data and store it seperately, \nwhen doing a rich checkout combine it back into the .jpg file. now the \nlarge binary blob doesn't change so you don't have to try and find deltas \nfor it.\n\nall the special case things would be in the helper routines written to do \nthe 'rich checkin/checkout' of each type. people who don't care about \nthis don't enable these helpers in the configs and so don't suffer any \noverhead (other then item (i) above)\n\nthis has the potential to be horribly abused, but it also has the \npotential to open up some very interesting possibilities as well.\n\nDavid Lang\n"},{"id":"86673","messageId":"20080810145019.GC3955@efreet.light.src","threadId":"14909","inReplyTo":"alpine.DEB.1.10.0808100502530.32620@asgard.lang.hm","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2008-08-10T14:50:19Z","receivedAt":"2008-08-10T14:50:19Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Sun, Aug 10, 2008 at 05:16:47 -0700, david@lang.hm wrote:\n> On Sun, 10 Aug 2008, Stephen R. van den Berg wrote:\n>> However, pondering the idea a bit more, I could envision something\n>> similar to the following:\n>>\n>> In the git tree the following layout would be used:\n>>\n>> plainfile.txt\n>> otherdir/otherplainfile.txt\n>> projects/README\n>> projects/README/_owner\n>> projects/README/_acl\n>> projects/README/_icon\n>> projects/README/_mimetype\n>> projects/something.mpeg\n>> projects/something.mpeg/_icon\n>> projects/something.mpeg/_mimetype\n>> projects/asubdir/thirdplainfile.txt\n>>\n>> That would imply that in the tree storage, the only extension would be\n>> that for any given reference to a blob in a tree object, there could be\n>> a reference to a tree object as well.  I.e. something like this in the\n>> tree object:\n>>\n>> 100644 blob f7b7414159b8a7159538fac543b2b19ef531968e  README\n>> 000000 tree df6ee415f04d6ccea5dab0de562c2f155583a2c4  README\n>> 100644 blob 0a54f8ec13df03cf6bdb5b973acec6d8141c01cc  something.mpeg\n>> 000000 tree a421448d765abb7bb979dc1d56621d0fc9b41229  soemthing.mpeg\n>>\n>> The extra tree reference for README would actually refer to something like:\n>>\n>> 100644 blob be3365fdaae0f4ed8c22c4cf38a4b1f88f9069c3  _owner\n>> 100644 blob 739e9e8f3d095931084b54cbf7f90d8f64eb0ac6  _acl\n>> 100644 blob bc1a868bb50644712966a50150d21199c401d6d5  _icon\n>> 100644 blob 6076bde5b3b6b8bed4ec4968d09abdbf015b3b75  _mimetype\n>>\n>> Which would contain the extra attributes.\n\n... provided the two entries under the same name wouldn't drive the internal\nlogic completely mad, I quite like this. Note by the way, that you need to\nallow for two trees too, because you may want to store attributes for\ndirectories too. It's no problem to differentiate them by type 04755\nvs. 00000 or 11000 or whatever, but it is a problem for index, because that\ndoes not store directory entries, so metadata for a directory would conflict\nwith regular entries in it. Can be fixed by using different filetype for the\nmetadata.\n\n>> And that would imply that during checkout you can do a rich checkout or a\n>> flat checkout for any files under the projects directory.\n>>\n>> A flat checkout results in the following files in the filesystem:\n>>\n>> plainfile.txt\n>> otherdir/otherplainfile.txt\n>> projects/README\n>> projects/README.attr/_owner\n>> projects/README.attr/_acl\n>> projects/README.attr/_icon\n>> projects/README.attr/_mimetype\n>> projects/something.mpeg\n>> projects/something.mpeg.attr/_icon\n>> projects/something.mpeg.attr/_mimetype\n>> projects/asubdir/thirdplainfile.txt\n\nStoring like this in index as well would make it even more compatible. Of\ncourse you are reserving the .attr suffix. But it's probably OK to reserve\n/something/ for this functionality (when the functionality is needed only).\nMaybe it could use some special character (@, #, =, $ or something) to\nseparate the suffix instead of normal . to decrease the chance to conflict\nwith other use.\n\n>> A rich checkout results in the following files in the filesystem:\n>>\n>> plainfile.txt\n>> otherdir/otherplainfile.txt\n>> projects/README\n>> projects/something.mpeg\n>> projects/asubdir/thirdplainfile.txt\n>> projects/asubdir/fourthplainfile.txt\n>>\n>> The rich checkout also applies the extended attributes/metadata to the\n>> filesystem (i.e. it would store all the metadata in the appropriate\n>> places).\n>>\n>> The nice thing about this setup is that:\n>> a. There is *no* change whatsoever to existing repositories or\n>>   repositoryformat.\n\nWell, there is a small change -- it needs to support multiple entries with\ndifferent type but same name in the tree object (but could be avoided by\nusing some special reserved suffix). Plus the index functionality needs to be\nmodified to put the metadata entries in the right places. Still of course\nmuch less invasive than the proposal from OP.\n\n>> b. It's less filling (i.e. there are no special bits or object types to be\n>>   used).\n>> c. Speed for files without attributes is not affected.\n>> d. It's fully 8-bit-transparent.\n>> e. It scales, even if you have large or many attributes.\n>> f. It uses the natural tree storage abstraction already supported in\n>>   git repositories to store the additional data.\n>> g. It allows reuse of attribute information at many levels.\n>> h. It even allows for a hierarchy of attributes attached to a single\n>>   file (no current filesystem supports that (yet)).\n>> i. The only change in the fast-path of core-git is that it would have to\n>>   know how to skip tree objects referenced in a tree object if a\n>>   same-name blob object is already there.  This can even be optimised\n>>   by requiring the attribute-tree to have a very specific (e.g. 0)\n>>   mode to ease detection.\n>> j. Editing and merging the meta-information could be made an almost\n>>   natural operation in the flat-checkout mode (the extension to be used\n>>   to name the attribute subdir should be made configurable).\n>\n> you also need to be able to add something to the attribute tree to  \n> indicate what type of metadata is being stored in it. you could  have \n> *nix perms, windows perms, posix extended attributes, or other things.\n\nWell, not really. I think the best way to implement the 'rich' checkout is to\nuse a hook to read/write the metadata. Git-core should just support storing\nattributes, but not actually store any of it's own, since they are nt needed\nfor it's main purpose, which is source code control.\n\n> I could see this as a great way to deal with editing exif data for \n> images. when checking in a .jpg, extract the .exif data and store it \n> seperately, when doing a rich checkout combine it back into the .jpg \n> file. now the large binary blob doesn't change so you don't have to try \n> and find deltas for it.\n>\n> all the special case things would be in the helper routines written to do \n> the 'rich checkin/checkout' of each type. people who don't care about  \n> this don't enable these helpers in the configs and so don't suffer any  \n> overhead (other then item (i) above)\n>\n> this has the potential to be horribly abused, but it also has the  \n> potential to open up some very interesting possibilities as well.\n\nI would say your example above belongs in the categry of abuses. The binary\ndiffer can deal with exif just OK (it's not compressed IIRC), so all you need\nis a custom diff driver for merging -- and that's already supported.\nCompressed stuff can be already handled for the differ with clean & smudge.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"86692","messageId":"20080810175735.GA14237@cuci.nl","threadId":"14909","inReplyTo":"20080810145019.GC3955@efreet.light.src","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-08-10T17:57:35Z","receivedAt":"2008-08-10T17:57:35Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jan Hudec wrote:\n>On Sun, Aug 10, 2008 at 05:16:47 -0700, david@lang.hm wrote:\n>> On Sun, 10 Aug 2008, Stephen R. van den Berg wrote:\n>>> However, pondering the idea a bit more, I could envision something\n>>> similar to the following:\n\n>.... provided the two entries under the same name wouldn't drive the internal\n>logic completely mad, I quite like this. Note by the way, that you need to\n>allow for two trees too, because you may want to store attributes for\n\nWell, in theory yes, but currently git doesn't store directories.\nHow about extending git-core to allow for storage of directories by\nvirtue of the following object in a tree:\n\n040000 blob e69de29bb2d1d6434b8b29ae775ad8c2e48c5391  .\n\nI.e. the hash belongs to the empty blob.\nNormally you don't (have to) store these directory blobs, but if you\ninsist on having them, git will create the empty directory on checkout\n(i.e. you wouldn't need the dummy file trick anymore to force the\ndirectory to be present).\n-- \nSincerely,\n           Stephen R. van den Berg.\n\nReal programmers don't produce results, they return exit codes.\n"},{"id":"86693","messageId":"20080810181115.GA3906@efreet.light.src","threadId":"14909","inReplyTo":"20080810175735.GA14237@cuci.nl","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2008-08-10T18:11:16Z","receivedAt":"2008-08-10T18:11:16Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Sun, Aug 10, 2008 at 19:57:35 +0200, Stephen R. van den Berg wrote:\n> Jan Hudec wrote:\n> >On Sun, Aug 10, 2008 at 05:16:47 -0700, david@lang.hm wrote:\n> >> On Sun, 10 Aug 2008, Stephen R. van den Berg wrote:\n> >>> However, pondering the idea a bit more, I could envision something\n> >>> similar to the following:\n> \n> >.... provided the two entries under the same name wouldn't drive the internal\n> >logic completely mad, I quite like this. Note by the way, that you need to\n> >allow for two trees too, because you may want to store attributes for\n> \n> Well, in theory yes, but currently git doesn't store directories.\n\nIt depends. It does store directories in the tree objects, it just does not\ndo that in index. And we are talking about tree objects, where git does store\ndirectories.\n\nBesides, that is irrelevant to storing attributes for directories -- the\nattribute objects are not themselves directories, so git would store them\njust fine.\n\n> How about extending git-core to allow for storage of directories by\n> virtue of the following object in a tree:\n> \n> 040000 blob e69de29bb2d1d6434b8b29ae775ad8c2e48c5391  .\n> \n> I.e. the hash belongs to the empty blob.\n\nSorry, but this is insane. If git was to store anything for empty\ndirectories, it would be empty tree, not a tree containing empty blob called\n'.'. There was even a prototype patch to do that sent to the list (I believe\nit was from Linus and was part of an argument along the lines \"you could do\nit like this, so stop talking and finish it if you have good enough reason to\nwant it (which you obviously don't)\").\n\n> Normally you don't (have to) store these directory blobs, but if you\n> insist on having them, git will create the empty directory on checkout\n> (i.e. you wouldn't need the dummy file trick anymore to force the\n> directory to be present).\n\nNo, I don't give a damn about directories themselves. I want to store their\nattributes, which is completely different thing.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"86714","messageId":"20080810201651.GB14237@cuci.nl","threadId":"14909","inReplyTo":"20080810181115.GA3906@efreet.light.src","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-08-10T20:16:51Z","receivedAt":"2008-08-10T20:16:51Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"Jan Hudec wrote:\n>On Sun, Aug 10, 2008 at 19:57:35 +0200, Stephen R. van den Berg wrote:\n>> Jan Hudec wrote:\n>If git was to store anything for empty\n>directories, it would be empty tree, not a tree containing empty blob called\n>'.'. There was even a prototype patch to do that sent to the list (I believe\n\nOk, sounds reasonable.\n\nWith respect to the storage inside the tree, using a duplicate name with\nmode 0 or a name with some kind of rare extension...\nIt should probably be investigated how much of the existing core needs\nto be touched/changed to support the duplicate name.\nI agree that using a custom rare extension would allow for almost no\nchange to git-core.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\nReal programmers don't produce results, they return exit codes.\n"},{"id":"86726","messageId":"7v7iao1oua.fsf@gitster.siamese.dyndns.org","threadId":"14909","inReplyTo":"20080810201651.GB14237@cuci.nl","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-08-10T22:34:37Z","receivedAt":"2008-08-10T22:34:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Stephen R. van den Berg\" <srb@cuci.nl> writes:\n\n> I agree that using a custom rare extension would allow for almost no\n> change to git-core.\n\nAnd at that point there is no \"plumbing\" side change necessary.  You just\nhave to teach your Porcelain to notice the associated \"metainfo\" files and\ndeal with them.\n\nFor merging such \"metainfo\", you would need to do your \"flattish/unrich\"\ncheckout anyway, so it might be that an easier approach for such a\nPorcelain might be:\n\n * Define a specific leading path, say \".attrs\" the hierarchy to store the\n   attributes information.  Attributes to a file README and t/Makefile\n   will be stored in .attrs/README and .attrs/t/Makefile.  They are\n   probably just plain text file you can do your merges and parsing easily\n   but with this counterproposal the only requirement is they are simple\n   plain blobs.  The plumbing layer does not care what payload they carry.\n\n * When you want to \"git setattr $path\", the Porcelain mucks with\n   \".attr/$path\".  Probably checkout codepath would give you a hook that\n   lets you reflect what \".attr/$path\" records to \"$path\", and checkin\n   (i.e. not commit but update-index) codepath would have another hook to\n   let you grab attributes for \"$path\" and update \".attr/$path\".\n\n * Merging and handling updates to \".attrs/\" hierarchy are done the usual\n   way we handle blobs.  Your Porcelain would then take the result and do\n   whatever changes to ACL or xattrs to the corresponding path, perhaps\n   from a hook after merge.\n\nSo it will most likely boild down to a \"Porcelain only\" convention that\ndifferent Porcelains would agree on.\n\nMy reaction for the initial proposal was very similar to the one given by\nShawn.  I do not see much point on having plumbing side support (yet).\n"},{"id":"86733","messageId":"alpine.DEB.1.10.0808101550570.32620@asgard.lang.hm","threadId":"14909","inReplyTo":"7v7iao1oua.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-10T23:10:44Z","receivedAt":"2008-08-10T23:10:44Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Sun, 10 Aug 2008, Junio C Hamano wrote:\n\n>> I agree that using a custom rare extension would allow for almost no\n>> change to git-core.\n>\n> And at that point there is no \"plumbing\" side change necessary.  You just\n> have to teach your Porcelain to notice the associated \"metainfo\" files and\n> deal with them.\n>\n> For merging such \"metainfo\", you would need to do your \"flattish/unrich\"\n> checkout anyway, so it might be that an easier approach for such a\n> Porcelain might be:\n>\n> * Define a specific leading path, say \".attrs\" the hierarchy to store the\n>   attributes information.  Attributes to a file README and t/Makefile\n>   will be stored in .attrs/README and .attrs/t/Makefile.  They are\n>   probably just plain text file you can do your merges and parsing easily\n>   but with this counterproposal the only requirement is they are simple\n>   plain blobs.  The plumbing layer does not care what payload they carry.\n>\n> * When you want to \"git setattr $path\", the Porcelain mucks with\n>   \".attr/$path\".  Probably checkout codepath would give you a hook that\n>   lets you reflect what \".attr/$path\" records to \"$path\", and checkin\n>   (i.e. not commit but update-index) codepath would have another hook to\n>   let you grab attributes for \"$path\" and update \".attr/$path\".\n>\n> * Merging and handling updates to \".attrs/\" hierarchy are done the usual\n>   way we handle blobs.  Your Porcelain would then take the result and do\n>   whatever changes to ACL or xattrs to the corresponding path, perhaps\n>   from a hook after merge.\n>\n> So it will most likely boild down to a \"Porcelain only\" convention that\n> different Porcelains would agree on.\n>\n> My reaction for the initial proposal was very similar to the one given by\n> Shawn.  I do not see much point on having plumbing side support (yet).\n\na few items\n\nconvienience\n\n1. tieing the attributes to the file more directly will make it much \neasier to deal with them along with the file in the non-rich checkout \n(it's much easier to say README* then README .attr/README*)\n\n\nconsisntancy\n\n2. putting hooks into the plumbing that can call external programs for the \nrich checkin/checkout will let all porcelains make use of the features \nwithout having to modify all of them independanty.\n\nsafety\n\n3. when doing checkins/checkouts of individual files you need to be sure \nthat you deal with the correct attributes at the same time (or else that \nthe person is explicity requesting only a piece of it) with the attributes \nclosely associated with the file this is much easier to do (this is \nanother aspect of the convienience in #1 above)\n\n4. if the configuration of what helper to use changes from one revision to \nanother the plumbing (which is already looking at the tree object for both \nrevisions) is in a better position to detect and alert then the porcelains\n\nDavid Lang\n"},{"id":"86756","messageId":"20080811101132.GB31686@cuci.nl","threadId":"14909","inReplyTo":"alpine.DEB.1.10.0808101550570.32620@asgard.lang.hm","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Stephen R. van den Berg","fromEmail":"srb@cuci.nl","sentAt":"2008-08-11T10:11:32Z","receivedAt":"2008-08-11T10:11:32Z","isPatch":false,"sender":{"key":"srb@cuci.nl","avatar":"https://gravatar.com/avatar/f75389059e827634d38e9df2a9b6ecbd50028b5a454442efa1c7205b7ff29c6a?d=mp&s=160"},"body":"david@lang.hm wrote:\n>On Sun, 10 Aug 2008, Junio C Hamano wrote:\n>>* Define a specific leading path, say \".attrs\" the hierarchy to store the\n>>  attributes information.  Attributes to a file README and t/Makefile\n\n>1. tieing the attributes to the file more directly will make it much \n>easier to deal with them along with the file in the non-rich checkout \n>(it's much easier to say README* then README .attr/README*)\n\nI have to agree that from a practical standpoint for the user, having\nthe file and the attribute tree right next to each other in the tree is\na lot easier to manage.\n\nSo even though setting up a shadow attribute tree is cleaner because it\ndoesn't need some kind of magic extension, it tends to clutter the\nmanagement in the flat-file checkout case.\n-- \nSincerely,\n           Stephen R. van den Berg.\n\n\"Beware: In C++, your friends can see your privates!\"\n"},{"id":"87367","messageId":"20080816062130.GA4554@oh.minilop.net","threadId":"14909","inReplyTo":"20080811101132.GB31686@cuci.nl","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Josh Triplett","fromEmail":"josh@freedesktop.org","sentAt":"2008-08-16T06:21:30Z","receivedAt":"2008-08-16T06:21:30Z","isPatch":false,"sender":{"key":"josh@joshtriplett.org","avatar":"https://avatars.githubusercontent.com/u/162737?v=4"},"body":"We want to reply to a few of the common points raised in this thread\nfirst, and then we have a few point-by-point replies later in this mail.\nIn particular, we see two common questions: whether Git should include\nsupport for metadata such as permissions and ownership at all, and\nhow Git should store this metadata if so.\n\nWe agree entirely with Jan Hudec's first point:\n\nOn Sun, Aug 10, 2008 at 01:09:25PM +0200, Jan Hudec wrote:\n> I am glad you came up with this, as I think this is the only reasonable way\n> to support things like etckeeper. The metastore and similar solutions are\n> a kludge and fall apart in so many cases.\n\nMetastore, etckeeper, and other existing \"hook-based\" solutions, which\nattempt to handle permissions separately, have several fundamental\nproblems.  They do not integrate well with the normal Git workflow, they\noften have race conditions that can lead to security problems, and they\nstore working-copy permissions separately from the filesystem\npermissions where they can potentially become out of sync.\n\nWe want to emphasize that we really don't have a preference amongst the\nvarious reasonable proposals for storing object metadata or for\npresenting that metadata in porcelain.  We're happy that our proposal\nstimulated discussion on the topic and that we now understand relevant\nGit internals much better.  We made our proposal and test case to\ndemonstrate that we're willing to design and implement a solution, not\njust complain that Git does not support permissions.\n\nAmong the proposals mentioned in this thread, we see some common\nrequirements:\n\n- All of the proposals suggest referencing the properties from the tree\n  containing the object they apply to, rather than creating an\n  extra object to store both hashes together.  We originally thought\n  that having a single object reference in the tree would make it easier\n  to iterate over the tree, construct each object, and apply its\n  permissions.  However, several of the proposals address that in other\n  ways.\n\n- Several proposals suggest storing the metadata as a tree object,\n  rather than a custom \"props\" object.  This makes a lot of sense.  It\n  allows Git to use existing logic for parsing, reachability\n  checking, merging, and checkouts.  On the other hand, we want to\n  optimize for the common cases such as POSIX permissions and ownership\n  rather than the unusual cases like extended attributes, so it might\n  make sense to store all the metadata for a particular object as a\n  single blob.\n\n- Several responses expressed concerns about merges and conflicts.  We\n  propose implementing support for this in plumbing the same way Git\n  does for everything else: put entries into the index with stages\n  marked.  This works whether metadata storage uses a tree or a blob.\n  Porcelains can choose how to resolve these merges and present\n  conflicts to the user for resolution.\n\n- Several proposals suggest using a magic suffix or special mode to\n  distinguish object file entries from their metadata entries.  Either\n  of these approaches seem fine.  In the case of a suffix, we think it\n  makes the most sense to use '/' or \"//\" in this suffix; any other\n  suffix would potentially conflict with legitimate filenames.  \"//\" has\n  the advantage of working unambiguously in the index as well.  Either\n  way, any porcelain on top of this could choose a different naming\n  scheme for non-\"rich\" checkouts, or check out the properties as a\n  separate top-level directory as Junio proposed.\n\n- Several people complained about our initial proposal of printable\n  ASCII for property names and values.  We used this approach solely\n  because it seemed like a reasonable starting place.  Length-prefixed\n  binary would work fine and provide 8-bit cleanness, as would the\n  proposals that store properties as trees of blobs.\n\nOn Sun, Aug 10, 2008 at 01:09:25PM +0200, Jan Hudec wrote:\n> Advantages (+), disadvantages (-) and possible (*) extensions of 1:\n>  \n>  + It should be possible to get to something useful with very little changes\n>    to git. Basically all it needs to be useful for things like etckeeper is\n>    to:\n>     . Make sure both clean and smudge filter always get filehandle to the\n>       disk file in question (I am /not/ suggesting path as the file may be\n>       written in a staging area and moved into place later).\n>     . Pass the blob id currently in index to the clean filter, so it can\n>       maintain the data if they are not representable in this particular\n>       checkout (eg. when checking out such repo on windows). Note, that this\n>       would also be useful for ignoring insignificant changes, eg. when\n>       a in some config file order is not important and the tool using it\n>       randomly changes that order when changing that file.\n\nIt might prove possible to implement a reasonable and secure interface\nfor permissions on top of Git without standardizing the plumbing and\nstorage formats, true.  With enough specialized hooks, some of the\nexisting problems with solutions like etckeeper and metastore go away.\nHowever, we feel that most of these solutions will have to deal with the\nsame problems, such as storage and merging, and the solutions will end\nup re-solving problems already handled by Git plumbing.  Those who do\nnot understand Git's solutions are doomed to re-invent them poorly. :)\n\nOn Sat, Aug 09, 2008 at 08:51:01PM -0700, Shawn O. Pearce wrote:\n> Nico and I have (at least in the past) agreed that type 0 is meant\n> as an escape indicator.  If the type is set to 0 then the real type\n> code appears in another byte of data which follows the object's\n> inflated length.\n>\n> That leaves only type 5 available.\n[...]\n> So yea, there really aren't any new type bits available.\n\nIf consensus opinion was that new object types were a reasonable way to\nsolve this problem, then it sounds as if there's plenty of room to\ncreate new types using this escape mechanism.  As a result we found your\nsubsequent comments a bit confusing since they seem to say only one more\nnew object type can exist.\n\n> But tossing aside the type bit argument, I'm not sure I see the\n> value in adding limited arbitrary properties to names in a tree.\n> How does one edit these?  How do you inspect them before you get\n> a checkout, assuming they might actually have an impact on the\n> checkout process?  How the hell do you merge them?\n\nSeveral of those questions depend on the porcelain.  The plumbing\nwould provide support for adding these properties to the index,\ncommitting them, viewing them, and doing merges in the index.  The\nporcelain would handle friendly editing, application to the working\ntree, and friendly merges.\n\n> A bad idea that will only clutter up the core object model, and\n> the core processing code of that object model.  Extended attributes\n> aren't used that much on local filesystems, because they are hard\n> to work with and suck performance wise.  Performance in Git is\n> a _feature_.  It matters.  Our clean object model really helps to\n> make that possible.\n\nIf you mean that our proposal seems too general, like extended\nattributes, then we can't argue with that. :)  We would have no problem\nwith a solution that only supported the standard POSIX info found in\n\"stat\" (permissions, ownership, times).  We just felt that such a\nspecific proposal would not go over well; if consensus points toward a\nmore specialized solution, that works fine for us too.\n\nWe actually proposed the simple name/value storage for props objects\nbecause we primarily cared about the case of small values like\npermissions, not large values like arbitrary xattrs.\n\nOn Sun, Aug 10, 2008 at 03:34:37PM -0700, Junio C Hamano wrote:\n> For merging such \"metainfo\", you would need to do your \"flattish/unrich\"\n> checkout anyway,\n\nWhy not just put entries into the index for each stage as merging\ncurrently does?  You could then compare the metadata in the index with\nthe filesystem metadata in the \"rich\" checkout, and resolve the conflict\nby adding the desired metadata to the index as stage 0 as usual.  You\nwould just need some sort of interface like \"git add --metadata file\" to\nadd the metadata for file to the index.  Alternatively, you could have\nsome simple wrappers to directly edit the metadata in the index, much\nlike the existing \"git update-index --chmod\" does for the execute bit.\n\n>  * Define a specific leading path, say \".attrs\" the hierarchy to store the\n>    attributes information.  Attributes to a file README and t/Makefile\n>    will be stored in .attrs/README and .attrs/t/Makefile.  They are\n>    probably just plain text file you can do your merges and parsing easily\n>    but with this counterproposal the only requirement is they are simple\n>    plain blobs.  The plumbing layer does not care what payload they carry.\n\nUsing a top-level tree to store all of the permissions makes sub-trees\nnot stand alone; the tree sha1 of a subdirectory doesn't give you enough\ninformation to recreate the metadata for that subdirectory.\n\n>  * When you want to \"git setattr $path\", the Porcelain mucks with\n>    \".attr/$path\".  Probably checkout codepath would give you a hook that\n>    lets you reflect what \".attr/$path\" records to \"$path\", and checkin\n>    (i.e. not commit but update-index) codepath would have another hook to\n>    let you grab attributes for \"$path\" and update \".attr/$path\".\n\nThis hook would need to provide a way to process these updates before\nthe blob or tree contents get put into place.  For example, if you check\nout /etc/shadow, you need to apply the non-world-readable permissions\n*before* you write out the contents.\n\n> So it will most likely boild down to a \"Porcelain only\" convention that\n> different Porcelains would agree on.\n> \n> My reaction for the initial proposal was very similar to the one given by\n> Shawn.  I do not see much point on having plumbing side support (yet).\n\nWe agree in principle that a sufficiently rich set of hooks might make\nit possible to implement metadata outside of the Git plumbing.  However,\nin practice the set of hooks necessary for complete integration seems\nquite large.  Furthermore, implementing these hooks efficiently seems\ndifficult.  We also don't want to force people to use a non-Git\nporcelain just to get support for permissions.  Finally, we think that\nalong with a common storage format, these porcelains will all have a\ncommon set of problems to solve, and it seems better to solve them once\ncorrectly in Git using code that mostly already exists.\n\n- Josh Triplett and Jamey Sharp\n"},{"id":"87368","messageId":"alpine.DEB.1.10.0808160046250.12859@asgard.lang.hm","threadId":"14909","inReplyTo":"20080816062130.GA4554@oh.minilop.net","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-16T07:56:19Z","receivedAt":"2008-08-16T07:56:19Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 15 Aug 2008, Josh Triplett wrote:\n\n>\n> - Several proposals suggest storing the metadata as a tree object,\n>  rather than a custom \"props\" object.  This makes a lot of sense.  It\n>  allows Git to use existing logic for parsing, reachability\n>  checking, merging, and checkouts.  On the other hand, we want to\n>  optimize for the common cases such as POSIX permissions and ownership\n>  rather than the unusual cases like extended attributes, so it might\n>  make sense to store all the metadata for a particular object as a\n>  single blob.\n\nahh, but if the 'tree object' that you are storing is named file.attr and \ncontains just the posix permissions and ownership, there are a very small \nnumber of different permutations that you will see on any one system (let \nalone in any one repository), as such the duplicates will all hash to the \nsame value and be combined in storage. your rich checkout porceleans can \ncache these into a lookup table and gain performance basicly equivalent to \ndefining a custom object.\n\nin fact, I'd be willing to bet that even when extended attributes are in \nuse (say SELinux tags) the number of different tree objects that would be \nused would still be pretty small.\n\n> On Sun, Aug 10, 2008 at 03:34:37PM -0700, Junio C Hamano wrote:\n>> For merging such \"metainfo\", you would need to do your \"flattish/unrich\"\n>> checkout anyway,\n>\n> Why not just put entries into the index for each stage as merging\n> currently does?  You could then compare the metadata in the index with\n> the filesystem metadata in the \"rich\" checkout, and resolve the conflict\n> by adding the desired metadata to the index as stage 0 as usual.  You\n> would just need some sort of interface like \"git add --metadata file\" to\n> add the metadata for file to the index.  Alternatively, you could have\n> some simple wrappers to directly edit the metadata in the index, much\n> like the existing \"git update-index --chmod\" does for the execute bit.\n\nbecouse the tools to work directly on the index are very limited. yes they \ncan be left in the index, but then the index-manipulation tools need to \nunderstand every type of metadata. if it's able to be presented in the \n\"flattish/unrich\" mode it will work anywhere, even on operating systems \nthat can't run your 'rich' tools\n\nDavid Lang\n"},{"id":"87372","messageId":"7vd4k9e120.fsf@gitster.siamese.dyndns.org","threadId":"14909","inReplyTo":"20080816062130.GA4554@oh.minilop.net","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-08-16T09:55:51Z","receivedAt":"2008-08-16T09:55:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Josh Triplett <josh@freedesktop.org>, Jamey Sharp <jamey@minilop.net>\nwrites:\n\n> This hook would need to provide a way to process these updates before\n> the blob or tree contents get put into place.  For example, if you check\n> out /etc/shadow, you need to apply the non-world-readable permissions\n> *before* you write out the contents.\n\nI think such atomicity or \"checkout race problem\" is irrelevant.\n\nI'd like to make a comment on this point, even though at the moment\n(especially before the real release), I am not very interested in where\nthis \"proposal\" is going.\n\nYou mention that you would resolve attribute conflicts just the same way\nyou would resolve contents conflicts, which in turn means that you would\ncheck out a half-merged state with conflict markers to the working tree,\nfix up the filesystem entity (both contents and presumably its attributes\nlike perm bits, ownership, xa and whatnot), and mark the path resolved.\nEven without talking about attributes conflicts, what's your position on\nthe time-window during which the contents of /etc/shadow and /etc/password\nhave conflict markers in them?\n\nLuckily, the markers do not have sufficient number of colons, and that\nwould protect your system from attempts to break into it with a phoney\nusername '=======' with an empty password ;-), but I think you get the\nidea.  Anything that has to be in some consistent state that cannot see\nconflicted state in the middle should not be merged in-place [*1*], [*2*].\n\nSo please simplify your requirements and at least drop atomicity argument.\n\nI am _not_ fundamentally opposed to somebody who wants to use git or any\nother SCM as a cooler representation of snapshots than a sequence of\ntarballs.  I however would be unhappy if your design and implementation\nbecomes more complicated than otherwise only because you try to deal with\nthe atomicity issue.  IOW, if your solution would become much simpler once\nyou pare down the atomicity requirement, then I'd reject the more complex\nvariant with atomicity in any second, even though I might still find the\nsimpler variant that does not care about atomicity worth considering.\n\n\n[Footnotes]\n\n*1* That is why people often frown upon \"using SCM to track changes of a\nlive system in-place\", and suggest tracking source material in SCM, and\nbuild material to deploy from the source and install into the final\ndestination (not limited to /etc but more often so for e.g. web server\nassets) as a better practice.\n\n*2* Also you should realize your \"/etc/shadow must be non-world-readable\nfrom the beginning\" is a very application specific wish.  What if the\nattribute you are trying to enforce is \"this path must always be\nworld-readable\"?  Are you going to limit this \"attribute enhancements\" to\nwhat you can specify at creat(2) time only?  How would you handle \"this\npath must be owned by user 'www-data' (assuming root drives git)\", which\nwould be done by creat(2) followed by chown(2)?\n"},{"id":"87380","messageId":"20080816150715.GA4057@efreet.light.src","threadId":"14909","inReplyTo":"7vd4k9e120.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2008-08-16T15:07:15Z","receivedAt":"2008-08-16T15:07:15Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Sat, Aug 16, 2008 at 02:55:51 -0700, Junio C Hamano wrote:\n> Josh Triplett <josh@freedesktop.org>, Jamey Sharp <jamey@minilop.net>\n> writes:\n> \n> > This hook would need to provide a way to process these updates before\n> > the blob or tree contents get put into place.  For example, if you check\n> > out /etc/shadow, you need to apply the non-world-readable permissions\n> > *before* you write out the contents.\n> \n> I think such atomicity or \"checkout race problem\" is irrelevant.\n> \n> I'd like to make a comment on this point, even though at the moment\n> (especially before the real release), I am not very interested in where\n> this \"proposal\" is going.\n> \n> You mention that you would resolve attribute conflicts just the same way\n> you would resolve contents conflicts, which in turn means that you would\n> check out a half-merged state with conflict markers to the working tree,\n> fix up the filesystem entity (both contents and presumably its attributes\n> like perm bits, ownership, xa and whatnot), and mark the path resolved.\n> Even without talking about attributes conflicts, what's your position on\n> the time-window during which the contents of /etc/shadow and /etc/password\n> have conflict markers in them?\n\nWell, there are situations where conflicts can happen and situations where\nthey can't. So I think the solution is \"don't merge in the live directory\"\n(applicable to other uses of version control in other kind of live copies\ntoo).\n\n> Luckily, the markers do not have sufficient number of colons, and that\n> would protect your system from attempts to break into it with a phoney\n> username '=======' with an empty password ;-), but I think you get the\n> idea.  Anything that has to be in some consistent state that cannot see\n> conflicted state in the middle should not be merged in-place [*1*], [*2*].\n> \n> So please simplify your requirements and at least drop atomicity argument.\n\nThe atomicity requirement is real for some applications, like the etckeeper.\n\nIt should be restated in terms of moving the content to the work tree rather\nthan before writing it out -- the content can be written out to a staging\narea, attributes applied and than moved into the tree. IIUC git already uses\na staging area during checkout, no?\n\n> I am _not_ fundamentally opposed to somebody who wants to use git or any\n> other SCM as a cooler representation of snapshots than a sequence of\n> tarballs.  I however would be unhappy if your design and implementation\n> becomes more complicated than otherwise only because you try to deal with\n> the atomicity issue.  IOW, if your solution would become much simpler once\n> you pare down the atomicity requirement, then I'd reject the more complex\n> variant with atomicity in any second, even though I might still find the\n> simpler variant that does not care about atomicity worth considering.\n\nI don't think the atomicity requirement should make anything more\ncomplicated. It is only a matter of running the hook applying the attributes\n-- I think git should not define meaning of the attributes -- at the right\npoint during the checkout process.\n\n> [Footnotes]\n> \n> *1* That is why people often frown upon \"using SCM to track changes of a\n> live system in-place\", and suggest tracking source material in SCM, and\n> build material to deploy from the source and install into the final\n> destination (not limited to /etc but more often so for e.g. web server\n> assets) as a better practice.\n\nYes, unless you need to track the changes done in the live directory by other\nsoftware, which is the case for /etc. It is also the case for ikiwiki-based\nweb sites.\n\nYou still need to avoid merging in the live tree to avoid breaking it, but\ngit always allows you to create a separate staging tree for such tasks.\n\n> *2* Also you should realize your \"/etc/shadow must be non-world-readable\n> from the beginning\" is a very application specific wish.  What if the\n> attribute you are trying to enforce is \"this path must always be\n> world-readable\"?  Are you going to limit this \"attribute enhancements\" to\n> what you can specify at creat(2) time only?  How would you handle \"this\n> path must be owned by user 'www-data' (assuming root drives git)\", which\n> would be done by creat(2) followed by chown(2)?\n\nYes, that does not make sense. But if you restate the requirement that the\nattributes must be applied when the file becomes accessible in the work tree,\nthan it makes sense and is easily doable by writing the file to a temporary\nlocation -- which is sufficiently protected if it is inside .git -- and\nmoving it into the tree as the last step. (The data is available inside\n.git/objects and .git/packs, so they are only as well protected as the .git\ndir itself is, so no restrictions as long as the file is inside .git).\n\nBest regards,\n\nJan\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"87518","messageId":"20080818061236.GF7376@spearce.org","threadId":"14909","inReplyTo":"20080816062130.GA4554@oh.minilop.net","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-18T06:12:36Z","receivedAt":"2008-08-18T06:12:36Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Josh Triplett <josh@freedesktop.org>, Jamey Sharp <jamey@minilop.net> wrote:\n> On Sat, Aug 09, 2008 at 08:51:01PM -0700, Shawn O. Pearce wrote:\n> > Nico and I have (at least in the past) agreed that type 0 is meant\n> > as an escape indicator.  If the type is set to 0 then the real type\n> > code appears in another byte of data which follows the object's\n> > inflated length.\n> >\n> > That leaves only type 5 available.\n> [...]\n> > So yea, there really aren't any new type bits available.\n> \n> If consensus opinion was that new object types were a reasonable way to\n> solve this problem, then it sounds as if there's plenty of room to\n> create new types using this escape mechanism.\n\nYes, but we'd hate to see the majority of the encodings within a\npack using the escape mechanism.\n\nSo a lot of my argument here was just trying to point out that\ntype bits aren't free, and we need to make sure the limited ones\navailable are applied to the majority of the pack contents.\n\nAdding a new type bit is a lot more than just adding it to the pack\ndata field.  Look at the amount of code that needed to be changed to\nsupport gitlink in trees, and that was \"reusing\" the OBJ_COMMIT type.\nAnytime you start poking at the core object enumeration code with\nnew cases there's a lot of corners that are affected.\n\n-- \nShawn.\n"},{"id":"87583","messageId":"20080818230646.GA11044@cisco.com","threadId":"14909","inReplyTo":"20080818061236.GF7376@spearce.org","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Derek Fawcus","fromEmail":"dfawcus@cisco.com","sentAt":"2008-08-18T23:06:46Z","receivedAt":"2008-08-18T23:06:46Z","isPatch":false,"sender":{"key":"dfawcus@cisco.com","avatar":null},"body":"On Sun, Aug 17, 2008 at 11:12:36PM -0700, Shawn O. Pearce wrote:\n> Adding a new type bit is a lot more than just adding it to the pack\n> data field.  Look at the amount of code that needed to be changed to\n> support gitlink in trees, and that was \"reusing\" the OBJ_COMMIT type.\n> Anytime you start poking at the core object enumeration code with\n> new cases there's a lot of corners that are affected.\n\nActually,  I'd been thinking of how to attach metadata - but more from\nthe perspective of attaching it to commits,  rather than individual\nblobs or trees.\n\nAt the moment,  my workaround is simply to add well known lines to\nthe end of the commit comments,  the downside being that it makes\nthe comments a bit ugly,  and one needs to know the protocol for\nparsing them.\n\nMy other hacky thought was that tag object could be overloaded for\nthis purpose.  It is already sort of an indirect object,  but seems\nto be limited to appearing at the edge of the graph.\n\nIf we could say have:\n\n  commit -> tag -> tree\n\nthen arbitrary data could be stored in the tag,  similarly this\ncould be extended for when a tree or blob object is expected\n(I'm not sure about the blob case).\n\nI guess there'd have to be some rule - like only one indirect\nobject allowed to be inserted (otherwise its awkward to check\nfor loops),  and there would need to be some custom merge rules.\n\nDF\n"},{"id":"87589","messageId":"20080818231844.GC9572@spearce.org","threadId":"14909","inReplyTo":"20080818230646.GA11044@cisco.com","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-18T23:18:44Z","receivedAt":"2008-08-18T23:18:44Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Derek Fawcus <dfawcus@cisco.com> wrote:\n> On Sun, Aug 17, 2008 at 11:12:36PM -0700, Shawn O. Pearce wrote:\n> > Adding a new type bit is a lot more than just adding it to the pack\n> > data field.  Look at the amount of code that needed to be changed to\n> > support gitlink in trees, and that was \"reusing\" the OBJ_COMMIT type.\n> > Anytime you start poking at the core object enumeration code with\n> > new cases there's a lot of corners that are affected.\n> \n> Actually,  I'd been thinking of how to attach metadata - but more from\n> the perspective of attaching it to commits,  rather than individual\n> blobs or trees.\n>\n> At the moment,  my workaround is simply to add well known lines to\n> the end of the commit comments,\n\nWe've talked about adding additional header lines to the commit after\nthe \"committer\" or \"encoding\" line but before the first blank line\nthat ends the headers and starts the message.  Most of the code will\nskip over an unknown header at this position, as we went through\nthat pain when we added the \"encoding\" header to the commit format.\n\nHowever, once you start putting headers into there one has to\nactually understand what they mean.  And it gets really ugly if\nyour tool thinks \"fixed XXX\\n\" means something different from what\nmy tool thinks \"fixed YYY\\n\" means and I use my tool against a\nclone of your repository.  In other words there is no concept of\n\"header namespaces\".\n\nThus far I don't think anyone has really tried adding more headers\nhere because nobody has come up with a concrete example of how it\nis useful.\n \n> I guess there'd have to be some rule - like only one indirect\n> object allowed to be inserted (otherwise its awkward to check\n> for loops),  and there would need to be some custom merge rules.\n\nLoops aren't possible.  If you can create a loop you have a very\nreal and very valid attack against SHA-1.  You will probably be\nable to use that in some way that profits you better than a loop\nwithin some random Git repository.\n\nYou may also want to look into the \"notes\" idea floated on the list\nin the past.  It allowed attaching trees (IIRC) to any commit, and\nfinding that later on in O(1) time during say git-log.  This can\nbe useful to attach a build report or a test report to a commit\nhours after it was created.\n\n-- \nShawn.\n"},{"id":"87591","messageId":"48AA0487.8050009@gmail.com","threadId":"14909","inReplyTo":"20080818230646.GA11044@cisco.com","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Marcus Griep","fromEmail":"neoeinstein@gmail.com","sentAt":"2008-08-18T23:23:51Z","receivedAt":"2008-08-18T23:23:51Z","isPatch":false,"sender":{"key":"neoeinstein@gmail.com","avatar":"https://gravatar.com/avatar/75d467077b37e56699d408fb97545e9a92a2907ff1feea4ba3a4b861f7cb7af4?d=mp&s=160"},"body":"Derek Fawcus wrote:\n> My other hacky thought was that tag object could be overloaded for\n> this purpose.  It is already sort of an indirect object,  but seems\n> to be limited to appearing at the edge of the graph.\n> \n> If we could say have:\n> \n>   commit -> tag -> tree\n> \n> then arbitrary data could be stored in the tag,  similarly this\n> could be extended for when a tree or blob object is expected\n> (I'm not sure about the blob case).\n\nI was under the impression that tags were references to commit objects,\nand they to tree objects:\n\ntag -> commit -> tree\n\nAlso, wouldn't this require large numbers tags, or the ability to multi-\ntarget tags?\n\n-- \nMarcus Griep\nGPG Key ID: 0x5E968152\n——\nhttp://www.boohaunt.net\nאת.ψο´\n"},{"id":"87593","messageId":"20080818232800.GD9572@spearce.org","threadId":"14909","inReplyTo":"48AA0487.8050009@gmail.com","subject":"Re: [RFC] Plumbing-only support for storing object metadata","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-18T23:28:00Z","receivedAt":"2008-08-18T23:28:00Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Marcus Griep <neoeinstein@gmail.com> wrote:\n> I was under the impression that tags were references to commit objects,\n> and they to tree objects:\n> \n> tag -> commit -> tree\n\nNo.  A tag can reference any object.  See for example the\njunio-gpg-pub tag in git.git, it references a blob, not a commit.\nThe linux-2.6.git tree has a tag which references a tree.\n\nTags may also reference other tags.\n \n> Also, wouldn't this require large numbers tags, or the ability to multi-\n> target tags?\n\nTag objects don't have to have names in the repository's ref space,\nbut it helps that they do when you are doing git-lost-found.\nHaving a tag in the database which shouldn't have a ref name in\nrefs/tags is more than a bit funny.\n\n-- \nShawn.\n"}]}