{"thread":{"id":"64841","subject":"Missing and omitted objects","startedAt":"2026-01-21T12:02:09Z","lastAt":"2026-01-26T12:48:21Z","messageCount":3,"participants":["Simon Richter","Philip Oakley"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"534329","messageId":"a612ea8e-a741-436d-8ed2-6ff09ba7945b@hogyros.de","threadId":"64841","inReplyTo":null,"subject":"Missing and omitted objects","fromName":"Simon Richter","fromEmail":"simon.richter@hogyros.de","sentAt":"2026-01-21T11:54:32Z","receivedAt":"2026-01-21T12:02:09Z","isPatch":false,"sender":{"key":"simon.richter@hogyros.de","avatar":"https://gravatar.com/avatar/1192aa9fa5dd19ce258b12b044cc27111dd7cdb58d920dc2123a02a24f55b5c5?d=mp&s=160"},"body":"Hi,\n\nwe're having a bit of a discussion in Debian.\n\nThe goal is to move towards git based storage for source packages, away \nfrom tarballs; ideally we'd like to reuse the upstream git archive as \nfar as possible, so it is easy to check for differences.\n\nHowever, some projects are shipping files that aren't redistributable, \nor that we want to omit for other reasons (such as vendored \ndependencies, when there is a perfectly working common version \navailable, and we really really want to make sure these don't get used \naccidentally).\n\nThe goal here is to allow the recipient of such a bundle to verify that \nany files received are unmodified, and get a list of paths that were \nremoved (which may be an entire subdirectory). Ideally, they could also \ncontinue working on a clone of this and generate commits on top as long \nas the affected paths aren't touched.\n\nThe minimal amount of data we'd want to archive is a single commit and \nits tree and dependencies, plus optionally a signed tag pointing at it \nif it exists (i.e. the same information we get if we use git-archive, \nplus the signature on the tag, plus the option to clone from such a \nsnapshot). For the simple case where nothing is removed, this already \nworks well and covers most of the use cases, but, sadly, not all of them.\n\nAs a side effect, this could make recovery of a broken repository that \nis missing objects more robust.\n\nRight now, I'd like some feedback whether someone has a better idea, and \nif such a feature could ever work or if it violates some fundamental \ndesign principles.\n\n    Simon\n"},{"id":"534614","messageId":"f6cc0420-1be6-4855-8c0f-b79c683203ee@iee.email","threadId":"64841","inReplyTo":"a612ea8e-a741-436d-8ed2-6ff09ba7945b@hogyros.de","subject":"Re: Missing and omitted objects","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2026-01-25T15:42:00Z","receivedAt":"2026-01-25T17:20:32Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 21/01/2026 11:54, Simon Richter wrote:\n> Hi,\n> \n> we're having a bit of a discussion in Debian.\n> \n> The goal is to move towards git based storage for source packages, away\n> from tarballs; ideally we'd like to reuse the upstream git archive as\n> far as possible, so it is easy to check for differences.\n> \n> However, some projects are shipping files that aren't redistributable,\n> or that we want to omit for other reasons (such as vendored\n> dependencies, when there is a perfectly working common version\n> available, and we really really want to make sure these don't get used\n> accidentally).\n\nThere was a discussion about allowing objects to be 'redacted' back at\nthe Git Merge 2020 (https://git-merge.com/),\nunder [TOPIC 3/17] Obliterate.\nhttps://lore.kernel.org/git/5B2FEA46-A12F-4DE7-A184-E8856EF66248@jramsay.com.au/\n\n> \n> The goal here is to allow the recipient of such a bundle to verify that\n> any files received are unmodified, and get a list of paths that were\n> removed (which may be an entire subdirectory). Ideally, they could also\n> continue working on a clone of this and generate commits on top as long\n> as the affected paths aren't touched.\n> \nThat discussion on redacting objects didn't reach any actionable\nconclusion that allows objects to be omitted/redacted, while keeping the\nbranch based directed graph flow. I've continued to consider options for\ndeliberately creating 'counterfeit' objects (old name/oid, but\nnew/limited content) which could then be 'verified' through a facsimile\nobject with the same new/limited content but a properly hashed name/oid.\nI haven't shared any of that with the list.\n\n> The minimal amount of data we'd want to archive is a single commit and\n> its tree and dependencies, plus optionally a signed tag pointing at it\n> if it exists (i.e. the same information we get if we use git-archive,\n> plus the signature on the tag, plus the option to clone from such a\n> snapshot). For the simple case where nothing is removed, this already\n> works well and covers most of the use cases, but, sadly, not all of them.\n\nYou could simply branch that special commit that will have all the\ndeletions, plus a 'deletions' file diff file (assuming you want to\nhighlight those deletions..), and then leave that branch as a stub, with\na tag, and remove that old branch name such that the tag is the thing\nthat retains the special commit in the hierarchy, and it's parent still\nholds within the regular git commit graph.\n> \n> As a side effect, this could make recovery of a broken repository that\n> is missing objects more robust.\n\nBroken repos are scarce, more often than not being compatibility issues\nbetween (*nix) Git and Git-for-Windows (case sensitivity, sizeof(long),\ncharacter limits, etc.). However redaction and overlarge files still fit\ninto the 'Don't do that' category (expect the unexpected..).\n\nThere is also the distinction between the meta-data and content. The\nformer also includes the data that holds together the commit graphs\nintegrity (hash of hashes) and filenames, directory names and commit\ntexts (point 15 of the Git Merge discussion). Being inside the hash\nverified meta data makes it \"hard\" to break and create exceptions.\n\nA mechanism for marking leaf objects as 'removed'/abscissed/absconded\nwould help here. It's tricky to do that safely for a commit, as it also\ncarries parent information which must be retained.\n\nFor a blob (leaf) object, with its free form text, it is possible to\nhave a fixed format, fixed length (hash specific) counterfeit object,\ne.g. \"Git redact abcd01245..\"(*) which would then also exist as a\nfacsimile (i.e.has a true hash oid) object within some authenticated\npart of the graph, and the counterfeit exist in place of the 'broken'\nblob object with that self referential \"abcd01245..\" oid.\n\nFor trees, it becomes necessary to locate a bit of free text in the meta\ndata to provide self reference, and make it appear as either the empty\ntree or empty file(blob). The true oid of such a counterfeit tree\nlikewise would need a way of existing within some authenticated part of\nthe wider graph. Perhaps a step too far at this stage of hand waving.\n\n> \n> Right now, I'd like some feedback whether someone has a better idea, and\n> if such a feature could ever work or if it violates some fundamental\n> design principles.\n\nIt's a big ask. Finding one specific feature (just on) that could\nactually be made to work would provide a toe hold for discussion.\n\nAt least this is a solid desire from within the community's infrastructure..\n\nAt present there is no mechanism for assuming that a piece of blob\n*content* is \"correct\" but that the oid it is stored under is incorrect\n/ does not match. We already have/had the `--literally` option for\ncreating arbitrary content, but not it's corollary `--use-oid=abcd01245..`.\nsee\nhttps://lore.kernel.org/git/20250516045010.GL22242@coredump.intra.peff.net/\nPeff cc'd\n> \n>    Simon\n> \n(*) I more wanted \"Git redact abcd01245.hexoid Base64oid\" to reduce\naccidental creation of such objects and allow double checking of the\noid. But maybe that's too cute.  ;-)\n"},{"id":"534664","messageId":"92cd5477-c2fc-42c5-b678-aa95b0999b2a@iee.email","threadId":"64841","inReplyTo":"f6cc0420-1be6-4855-8c0f-b79c683203ee@iee.email","subject":"Re: Missing and omitted objects","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2026-01-26T12:48:18Z","receivedAt":"2026-01-26T12:48:21Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Had a bit of a think overnight. Some parts probably don't apply to this\napplication, see below.\n\nOn 25/01/2026 15:42, Philip Oakley wrote:\n> On 21/01/2026 11:54, Simon Richter wrote:\n>> Hi,\n>>\n>> we're having a bit of a discussion in Debian.\n>>\n>> The goal is to move towards git based storage for source packages, away\n>> from tarballs; ideally we'd like to reuse the upstream git archive as\n>> far as possible, so it is easy to check for differences.\n>>\n>> However, some projects are shipping files that aren't redistributable,\n>> or that we want to omit for other reasons (such as vendored\n>> dependencies, when there is a perfectly working common version\n>> available, and we really really want to make sure these don't get used\n>> accidentally).\n> \n> There was a discussion about allowing objects to be 'redacted' back at\n> the Git Merge 2020 (https://git-merge.com/),\n> under [TOPIC 3/17] Obliterate.\n> https://lore.kernel.org/git/5B2FEA46-A12F-4DE7-A184-E8856EF66248@jramsay.com.au/\n>\nThe discussion is like still informative.\n\n>>\n>> The goal here is to allow the recipient of such a bundle to verify that\n>> any files received are unmodified, and get a list of paths that were\n>> removed (which may be an entire subdirectory). Ideally, they could also\n>> continue working on a clone of this and generate commits on top as long\n>> as the affected paths aren't touched.\n>>\n> That discussion on redacting objects didn't reach any actionable\n> conclusion that allows objects to be omitted/redacted, while keeping the\n> branch based directed graph flow. I've continued to consider options for\n> deliberately creating 'counterfeit' objects (old name/oid, but\n> new/limited content) which could then be 'verified' through a facsimile\n> object with the same new/limited content but a properly hashed name/oid.\n> I haven't shared any of that with the list.\n> \nThe idea of eliminating objects by OID, totally, from the repo is not\nsuitable for the use case. It would be an all-or-nothing response,\nrather than a tailored response.\n\n\n>> The minimal amount of data we'd want to archive is a single commit and\n>> its tree and dependencies, plus optionally a signed tag pointing at it\n>> if it exists (i.e. the same information we get if we use git-archive,\n>> plus the signature on the tag, plus the option to clone from such a\n>> snapshot). For the simple case where nothing is removed, this already\n>> works well and covers most of the use cases, but, sadly, not all of them.\n> \n> You could simply branch that special commit that will have all the\n> deletions, plus a 'deletions' file diff file (assuming you want to\n> highlight those deletions..), and then leave that branch as a stub, with\n> a tag, and remove that old branch name such that the tag is the thing\n> that retains the special commit in the hierarchy, and it's parent still\n> holds within the regular git commit graph.\n\nThis may still be a useful tailoring where a separate commit is\ngenerated which omits unwanted files/content. This is quite lightweight\nin terms of repo size because of the inherent de-duplication of common\ncontent. It's only the updated trees that need storing.\n\nPhilip\n\n>>\n>> As a side effect, this could make recovery of a broken repository that\n>> is missing objects more robust.\n> \n> Broken repos are scarce, more often than not being compatibility issues\n> between (*nix) Git and Git-for-Windows (case sensitivity, sizeof(long),\n> character limits, etc.). However redaction and overlarge files still fit\n> into the 'Don't do that' category (expect the unexpected..).\n> \n> There is also the distinction between the meta-data and content. The\n> former also includes the data that holds together the commit graphs\n> integrity (hash of hashes) and filenames, directory names and commit\n> texts (point 15 of the Git Merge discussion). Being inside the hash\n> verified meta data makes it \"hard\" to break and create exceptions.\n> \n> A mechanism for marking leaf objects as 'removed'/abscissed/absconded\n> would help here. It's tricky to do that safely for a commit, as it also\n> carries parent information which must be retained.\n> \n> For a blob (leaf) object, with its free form text, it is possible to\n> have a fixed format, fixed length (hash specific) counterfeit object,\n> e.g. \"Git redact abcd01245..\"(*) which would then also exist as a\n> facsimile (i.e.has a true hash oid) object within some authenticated\n> part of the graph, and the counterfeit exist in place of the 'broken'\n> blob object with that self referential \"abcd01245..\" oid.\n> \n> For trees, it becomes necessary to locate a bit of free text in the meta\n> data to provide self reference, and make it appear as either the empty\n> tree or empty file(blob). The true oid of such a counterfeit tree\n> likewise would need a way of existing within some authenticated part of\n> the wider graph. Perhaps a step too far at this stage of hand waving.\n> \n>>\n>> Right now, I'd like some feedback whether someone has a better idea, and\n>> if such a feature could ever work or if it violates some fundamental\n>> design principles.\n> \n> It's a big ask. Finding one specific feature (just on) that could\n> actually be made to work would provide a toe hold for discussion.\n> \n> At least this is a solid desire from within the community's infrastructure..\n> \n> At present there is no mechanism for assuming that a piece of blob\n> *content* is \"correct\" but that the oid it is stored under is incorrect\n> / does not match. We already have/had the `--literally` option for\n> creating arbitrary content, but not it's corollary `--use-oid=abcd01245..`.\n> see\n> https://lore.kernel.org/git/20250516045010.GL22242@coredump.intra.peff.net/\n> Peff cc'd\n>>\n>>    Simon\n>>\n> (*) I more wanted \"Git redact abcd01245.hexoid Base64oid\" to reduce\n> accidental creation of such objects and allow double checking of the\n> oid. But maybe that's too cute.  ;-)\n> \n\n"}]}