{"thread":{"id":"49202","subject":"Questions about the hash function transition","startedAt":"2018-08-23T14:02:58Z","lastAt":"2018-08-29T23:45:28Z","messageCount":33,"participants":["Ævar Arnfjörð Bjarmason","Junio C Hamano","brian m. carlson","Jonathan Nieder","Johannes Schindelin","Derrick Stolee","Edward Thomson","Stefan Beller","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"356361","messageId":"878t4xfaes.fsf@evledraar.gmail.com","threadId":"49202","inReplyTo":null,"subject":"Questions about the hash function transition","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-08-23T14:02:51Z","receivedAt":"2018-08-23T14:02:58Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"I wanted to send another series to clarify things in\nhash-function-transition.txt, but for some of the issues I don't know\nthe answer, and I had some questions after giving this another read.\n\nSo let's discuss that here first. Quoting from the document (available\nat\nhttps://github.com/git/git/blob/v2.19.0-rc0/Documentation/technical/hash-function-transition.txt)\n\n> Git hash function transition\n> ============================\n>\n> Objective\n> ---------\n> Migrate Git from SHA-1 to a stronger hash function.\n\nShould way say \"Migrate Git from SHA-1 to SHA-256\" here instead?\n\nMaybe it's overly specific, i.e. really we're also describnig how /any/\nhash function transition might happen, but having just read this now\nfrom start to finish it takes us a really long time to mention (and at\nfirst, only offhand) that SHA-256 is the new hash.\n\n> [...]\n> Goals\n> -----\n> 1. The transition to SHA-256 can be done one local repository at a time.\n>    a. Requiring no action by any other party.\n>    b. A SHA-256 repository can communicate with SHA-1 Git servers\n>       (push/fetch).\n>    c. Users can use SHA-1 and SHA-256 identifiers for objects\n>       interchangeably (see \"Object names on the command line\", below).\n>    d. New signed objects make use of a stronger hash function than\n>       SHA-1 for their security guarantees.\n> 2. Allow a complete transition away from SHA-1.\n>    a. Local metadata for SHA-1 compatibility can be removed from a\n>       repository if compatibility with SHA-1 is no longer needed.\n> 3. Maintainability throughout the process.\n>    a. The object format is kept simple and consistent.\n>    b. Creation of a generalized repository conversion tool.\n>\n> Non-Goals\n> ---------\n> 1. Add SHA-256 support to Git protocol. This is valuable and the\n>    logical next step but it is out of scope for this initial design.\n\nThis is a non-goal according to the docs, but now that we have protocol\nv2 in git, perhaps we could start specifying or describing how this\nprotocol extension will work?\n\n> [...]\n> 3. Intermixing objects using multiple hash functions in a single\n>    repository.\n\nBut isn't that the goal now per \"Translation table\" & writing both SHA-1\nand SHA-256 versions of objects?\n\n> [...]\n> Pack index\n> ~~~~~~~~~~\n> Pack index (.idx) files use a new v3 format that supports multiple\n> hash functions. They have the following format (all integers are in\n> network byte order):\n>\n> - A header appears at the beginning and consists of the following:\n>   - The 4-byte pack index signature: '\\377t0c'\n>   - 4-byte version number: 3\n>   - 4-byte length of the header section, including the signature and\n>     version number\n>   - 4-byte number of objects contained in the pack\n>   - 4-byte number of object formats in this pack index: 2\n>   - For each object format:\n>     - 4-byte format identifier (e.g., 'sha1' for SHA-1)\n\nSo, given that we have 4-byte limit and have decided on SHA-256 are we\njust going to call this 'sha2'? That might be confusingly ambiguous\nsince SHA2 is a standard with more than just SHA-256, maybe 's256', or\nmaybe we should give this 8 bytes with trailing \\0s so we can have\n\"SHA-1\\0\\0\\0\" and \"SHA-256\\0\"?\n\n> [...]\n> - The trailer consists of the following:\n>   - A copy of the 20-byte SHA-256 checksum at the end of the\n>     corresponding packfile.\n>\n>   - 20-byte SHA-256 checksum of all of the above.\n\nWe need to update both of these to 32 byte, right? Or are we planning to\ntruncate the checksums?\n\nThis seems like just a mistake when we did s/NewHash/SHA-256/g, but then\nagain it was originally \"20-byte NewHash checksum\" ever since 752414ae43\n(\"technical doc: add a design doc for hash function transition\",\n2017-09-27), so what do we mean here?\n\n> Loose object index\n> ~~~~~~~~~~~~~~~~~~\n> A new file $GIT_OBJECT_DIR/loose-object-idx contains information about\n> all loose objects. Its format is\n>\n>   # loose-object-idx\n>   (sha256-name SP sha1-name LF)*\n>\n> where the object names are in hexadecimal format. The file is not\n> sorted.\n>\n> The loose object index is protected against concurrent writes by a\n> lock file $GIT_OBJECT_DIR/loose-object-idx.lock. To add a new loose\n> object:\n>\n> 1. Write the loose object to a temporary file, like today.\n> 2. Open loose-object-idx.lock with O_CREAT | O_EXCL to acquire the lock.\n> 3. Rename the loose object into place.\n> 4. Open loose-object-idx with O_APPEND and write the new object\n> 5. Unlink loose-object-idx.lock to release the lock.\n>\n> To remove entries (e.g. in \"git pack-refs\" or \"git-prune\"):\n>\n> 1. Open loose-object-idx.lock with O_CREAT | O_EXCL to acquire the\n>    lock.\n> 2. Write the new content to loose-object-idx.lock.\n> 3. Unlink any loose objects being removed.\n> 4. Rename to replace loose-object-idx, releasing the lock.\n\nDo we expect multiple concurrent writers to poll the lock if they can't\naquire it right away? I.e. concurrent \"git commit\" would block? Has this\noverall approach been benchmarked somewhere?\n\nI wonder if some lock-less variant of this would perform\nbetter. E.g. that we'd consult not one, but any\nloose-object-{1..Inf}.idx files written in sequence, and we's specify\nthat each file would contain no more than N mappings.\n\nThen writers could stat() the file to figure out if it has more than N\nalready by looking at the size (we have fixed-width records). Once we've\nfilled loose-object-1.idx we start writing to loose-object-2.idx and so\non.\n\nThe advantage of this is that writers wouldn't have to block one\nanother, and could just O_APPEND write to the file(s), although we could\nend up with duplicate entries (which readers would need to tolerate).\n\nThen some GC process could look at the set of loose-object-{1..Inf}.idx\nfiles, find one that was at the max size (or slightly above, due to the\nO_APPEND race condition), and whose mtime was deemed old enough to be\n\"safe\" (to guard against an hours-long git-commit writing to it),\ncompact it, and rename a new one in-place, or better yet get rid of it\nin favor of a pack).\n\nMaybe I've missed some subtlety where that won't work, I'm just\nconcerned that something that's writing a lot of objects in parallel\nwill be slowed down (e.g. the likes of BFG repo cleaner).\n\n> Translation table\n> ~~~~~~~~~~~~~~~~~\n> The index files support a bidirectional mapping between sha1-names\n> and sha256-names. The lookup proceeds similarly to ordinary object\n> lookups. For example, to convert a sha1-name to a sha256-name:\n>\n>  1. Look for the object in idx files. If a match is present in the\n>     idx's sorted list of truncated sha1-names, then:\n>     a. Read the corresponding entry in the sha1-name order to pack\n>        name order mapping.\n>     b. Read the corresponding entry in the full sha1-name table to\n>        verify we found the right object. If it is, then\n>     c. Read the corresponding entry in the full sha256-name table.\n>        That is the object's sha256-name.\n>  2. Check for a loose object. Read lines from loose-object-idx until\n>     we find a match.\n>\n> Step (1) takes the same amount of time as an ordinary object lookup:\n> O(number of packs * log(objects per pack)). Step (2) takes O(number of\n> loose objects) time. To maintain good performance it will be necessary\n> to keep the number of loose objects low. See the \"Loose objects and\n> unreachable objects\" section below for more details.\n>\n> Since all operations that make new objects (e.g., \"git commit\") add\n> the new objects to the corresponding index, this mapping is possible\n> for all objects in the object store.\n\nAre we going to need a midx version of these mapping files? How does\nmidx fit into this picture? Perhaps it's too obscure to worry about...\n\n> Reading an object's sha1-content\n> ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> The sha1-content of an object can be read by converting all sha256-names\n> its sha256-content references to sha1-names using the translation table.\n>\n> Fetch\n> ~~~~~\n> Fetching from a SHA-1 based server requires translating between SHA-1\n> and SHA-256 based representations on the fly.\n>\n> SHA-1s named in the ref advertisement that are present on the client\n> can be translated to SHA-256 and looked up as local objects using the\n> translation table.\n>\n> Negotiation proceeds as today. Any \"have\"s generated locally are\n> converted to SHA-1 before being sent to the server, and SHA-1s\n> mentioned by the server are converted to SHA-256 when looking them up\n> locally.\n>\n> After negotiation, the server sends a packfile containing the\n> requested objects. We convert the packfile to SHA-256 format using\n> the following steps:\n>\n> 1. index-pack: inflate each object in the packfile and compute its\n>    SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against\n>    objects the client has locally. These objects can be looked up\n>    using the translation table and their sha1-content read as\n>    described above to resolve the deltas.\n> 2. topological sort: starting at the \"want\"s from the negotiation\n>    phase, walk through objects in the pack and emit a list of them,\n>    excluding blobs, in reverse topologically sorted order, with each\n>    object coming later in the list than all objects it references.\n>    (This list only contains objects reachable from the \"wants\". If the\n>    pack from the server contained additional extraneous objects, then\n>    they will be discarded.)\n> 3. convert to sha256: open a new (sha256) packfile. Read the topologically\n>    sorted list just generated. For each object, inflate its\n>    sha1-content, convert to sha256-content, and write it to the sha256\n>    pack. Record the new sha1<->sha256 mapping entry for use in the idx.\n> 4. sort: reorder entries in the new pack to match the order of objects\n>    in the pack the server generated and include blobs. Write a sha256 idx\n>    file\n> 5. clean up: remove the SHA-1 based pack file, index, and\n>    topologically sorted list obtained from the server in steps 1\n>    and 2.\n\nDoesn't this process require us to implement a \"fetch quarantine\"? Least\nwe have (e.g. other concurrent fetches) referencing those new SHA-1\nobjects we've fetched in a pack that we'll remove in step #5?\n\n> [...]\n> The user can also explicitly specify which format to use for a\n> particular revision specifier and for output, overriding the mode. For\n> example:\n>\n> git --output-format=sha1 log abac87a^{sha1}..f787cac^{sha256}\n\nHow is this going to interact with other peel syntax? I.e. now we have\n<object>^{commit} <sha>^{tag} etc. It seems to me we'll need not ^{sha1}\nbut ^{sha1:<current_type>}, e.g. ^{sha1:commit} or ^{sha1:tag}, with\ncurrent ^{} being a synonym for ^{sha1:}.\n\nOr is this expected to be chained, as e.g. <object>^{tag}^{sha256} ?\n\n> Transition plan\n> ---------------\n\nOne thing that's not covered in this document at all, which I feel is\nmissing, is how we're going to handle references to old commit IDs in\ncommit messages, bug trackers etc. once we go through the whole\nmigration process.\n\nI.e. are users who expect to be able to read old history and \"git show\n<sha1 I found>\" expected to maintain a repository that has a live\nsha1<->sha256 mapping forever, or could we be smarter about this and\nsupport some sort of marker in the repository saying \"maintain the\nmapping up until this point\".\n\nThen, along with some v2 protocol extension to transfer such a\nhistorical mapping (and perhaps a default user option to request it)\nwe'd be guaranteed to be able to read old log messages and \"git show\"\nthem, and servers could avoid breaking past URLs without maintaining the\nmapping going forward.\n\nOne example of this on the server is that on GitLab (I don't know how\nGitHub does this) when you reference a commit from e.g a bug, a\nrefs/keep-around/<sha1> is created, to make sure it doesn't get GC'd.\n\nThose sorts of hosting providers would like to not break *existing*\nlinks, without needing to forever maintain a bidirectional mapping.\n"},{"id":"356363","messageId":"xmqqy3cxjgz1.fsf@gitster-ct.c.googlers.com","threadId":"49202","inReplyTo":"878t4xfaes.fsf@evledraar.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-23T14:27:30Z","receivedAt":"2018-08-23T14:27:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> - The trailer consists of the following:\n>>   - A copy of the 20-byte SHA-256 checksum at the end of the\n>>     corresponding packfile.\n>>\n>>   - 20-byte SHA-256 checksum of all of the above.\n>\n> We need to update both of these to 32 byte, right? Or are we planning to\n> truncate the checksums?\n\nhttps://public-inbox.org/git/CA+55aFwc7UQ61EbNJ36pFU_aBCXGya4JuT-TvpPJ21hKhRengQ@mail.gmail.com/\n"},{"id":"356370","messageId":"876001f6u3.fsf@evledraar.gmail.com","threadId":"49202","inReplyTo":"xmqqy3cxjgz1.fsf@gitster-ct.c.googlers.com","subject":"Re: Questions about the hash function transition","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-08-23T15:20:04Z","receivedAt":"2018-08-23T15:20:10Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Aug 23 2018, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>\n>>> - The trailer consists of the following:\n>>>   - A copy of the 20-byte SHA-256 checksum at the end of the\n>>>     corresponding packfile.\n>>>\n>>>   - 20-byte SHA-256 checksum of all of the above.\n>>\n>> We need to update both of these to 32 byte, right? Or are we planning to\n>> truncate the checksums?\n>\n> https://public-inbox.org/git/CA+55aFwc7UQ61EbNJ36pFU_aBCXGya4JuT-TvpPJ21hKhRengQ@mail.gmail.com/\n\nThanks.\n\nYeah for this checksum purpose even 10 or 5 characters would do, but\nsince we'll need a new pack format anyway for SHA-256 why not just use\nthe full length of the SHA-256 here? We're using the full length of the\nSHA-1.\n\nI don't see it mattering for security / corruption detection purposes,\nbut just to avoid confusion. We'll have this one place left where\nsomething looks like a SHA-1, but is actually a trunctated SHA-256.\n"},{"id":"356380","messageId":"xmqqa7pdjc1s.fsf@gitster-ct.c.googlers.com","threadId":"49202","inReplyTo":"876001f6u3.fsf@evledraar.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-23T16:13:51Z","receivedAt":"2018-08-23T16:13:57Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> On Thu, Aug 23 2018, Junio C Hamano wrote:\n>\n>> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n>>\n>>>> - The trailer consists of the following:\n>>>>   - A copy of the 20-byte SHA-256 checksum at the end of the\n>>>>     corresponding packfile.\n>>>>\n>>>>   - 20-byte SHA-256 checksum of all of the above.\n>>>\n>>> We need to update both of these to 32 byte, right? Or are we planning to\n>>> truncate the checksums?\n>>\n>> https://public-inbox.org/git/CA+55aFwc7UQ61EbNJ36pFU_aBCXGya4JuT-TvpPJ21hKhRengQ@mail.gmail.com/\n>\n> Thanks.\n>\n> Yeah for this checksum purpose even 10 or 5 characters would do, but\n> since we'll need a new pack format anyway for SHA-256 why not just use\n> the full length of the SHA-256 here? We're using the full length of the\n> SHA-1.\n>\n> I don't see it mattering for security / corruption detection purposes,\n> but just to avoid confusion. We'll have this one place left where\n> something looks like a SHA-1, but is actually a trunctated SHA-256.\n\nI would prefer to see us at least explore if the gain in throughput\nis sufficiently big if we switch to weaker checksum, like crc32.  If\ndoes not give us sufficient gain, I'd agree with you that consistently\nusing full hash everywhere would conceptually be cleaner.\n\n\n"},{"id":"356427","messageId":"20180824014007.GF535143@genre.crustytoothpaste.net","threadId":"49202","inReplyTo":"878t4xfaes.fsf@evledraar.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-08-24T01:40:07Z","receivedAt":"2018-08-24T01:40:18Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Thu, Aug 23, 2018 at 04:02:51PM +0200, Ævar Arnfjörð Bjarmason wrote:\n> > [...]\n> > Goals\n> > -----\n> > 1. The transition to SHA-256 can be done one local repository at a time.\n> >    a. Requiring no action by any other party.\n> >    b. A SHA-256 repository can communicate with SHA-1 Git servers\n> >       (push/fetch).\n> >    c. Users can use SHA-1 and SHA-256 identifiers for objects\n> >       interchangeably (see \"Object names on the command line\", below).\n> >    d. New signed objects make use of a stronger hash function than\n> >       SHA-1 for their security guarantees.\n> > 2. Allow a complete transition away from SHA-1.\n> >    a. Local metadata for SHA-1 compatibility can be removed from a\n> >       repository if compatibility with SHA-1 is no longer needed.\n> > 3. Maintainability throughout the process.\n> >    a. The object format is kept simple and consistent.\n> >    b. Creation of a generalized repository conversion tool.\n> >\n> > Non-Goals\n> > ---------\n> > 1. Add SHA-256 support to Git protocol. This is valuable and the\n> >    logical next step but it is out of scope for this initial design.\n> \n> This is a non-goal according to the docs, but now that we have protocol\n> v2 in git, perhaps we could start specifying or describing how this\n> protocol extension will work?\n\nI have code that does this.  The reason is that the first stage of the\ntransition code is to implement stage 4 of the transition: that is, a\nfull SHA-256 implementation without any SHA-1 support.  Implementing it\nthat way means that we don't have to deal with any of the SHA-1 to\nSHA-256 mapping in the first stage of the code.\n\nIn order to clone an SHA-256 repo (which the testsuite is completely\nbroken without), you need to be able to have basic SHA-256 support in\nthe protocol.  I know this was a non-goal, but the alternative is a an\ninability to run the testsuite using SHA-256 until all the code is\nmerged, which is unsuitable for development.  The transition plan also\nanticipates stage 4 (full SHA-256) support before earlier stages, so\nthis will be required.\n\nI hope to be able to spend some time documenting this in a little bit.\nI have documentation for that code in my branch, but I haven't sent it\nin yet.\n\nI realize I have a lot of code that has not been sent in yet, but I also\ntend to build on my own series a lot, and I probably need to be a bit\nbetter about extracting reusable pieces that can go in independently\nwithout waiting for the previous series to land.\n\n> > [...]\n> > 3. Intermixing objects using multiple hash functions in a single\n> >    repository.\n> \n> But isn't that the goal now per \"Translation table\" & writing both SHA-1\n> and SHA-256 versions of objects?\n\nNo, I think this statement is basically that you have to have the entire\nrepository use all one algorithm under the hood in the .git directory,\ntranslation tables excluded.  I don't think that's controversial.\n\n> > [...]\n> > Pack index\n> > ~~~~~~~~~~\n> > Pack index (.idx) files use a new v3 format that supports multiple\n> > hash functions. They have the following format (all integers are in\n> > network byte order):\n> >\n> > - A header appears at the beginning and consists of the following:\n> >   - The 4-byte pack index signature: '\\377t0c'\n> >   - 4-byte version number: 3\n> >   - 4-byte length of the header section, including the signature and\n> >     version number\n> >   - 4-byte number of objects contained in the pack\n> >   - 4-byte number of object formats in this pack index: 2\n> >   - For each object format:\n> >     - 4-byte format identifier (e.g., 'sha1' for SHA-1)\n> \n> So, given that we have 4-byte limit and have decided on SHA-256 are we\n> just going to call this 'sha2'? That might be confusingly ambiguous\n> since SHA2 is a standard with more than just SHA-256, maybe 's256', or\n> maybe we should give this 8 bytes with trailing \\0s so we can have\n> \"SHA-1\\0\\0\\0\" and \"SHA-256\\0\"?\n\nThis is the format_version field in struct git_hash_algo.\n\nFor SHA-1, I have 0x73686131, which is \"sha1\", big-endian, and for\nSHA-256, I have 0x73323536, which is \"s256\", big-endian.  The former is\nin the codebase already; the latter, in my hash-impl branch.\n\nIf people have objections, we can change this up until we merge the pack\nindex v3 code (which is not yet finished).  It needs to be unique, and\nthat's it.  We could specify 0x00000001 and 0x00000002 if we wanted,\nalthough I feel the values I mentioned above are self-documenting, which\nis desirable.\n\n> > [...]\n> > - The trailer consists of the following:\n> >   - A copy of the 20-byte SHA-256 checksum at the end of the\n> >     corresponding packfile.\n> >\n> >   - 20-byte SHA-256 checksum of all of the above.\n> \n> We need to update both of these to 32 byte, right? Or are we planning to\n> truncate the checksums?\n> \n> This seems like just a mistake when we did s/NewHash/SHA-256/g, but then\n> again it was originally \"20-byte NewHash checksum\" ever since 752414ae43\n> (\"technical doc: add a design doc for hash function transition\",\n> 2017-09-27), so what do we mean here?\n\nYes, this will be 32 bytes.  The code I have uses 32 bytes, because\ntruncating it means that we have to write special code just for that\ncase, which seems silly.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"356428","messageId":"20180824014703.GE99542@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"878t4xfaes.fsf@evledraar.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-24T01:47:03Z","receivedAt":"2018-08-24T01:47:09Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nÆvar Arnfjörð Bjarmason wrote:\n\n> I wanted to send another series to clarify things in\n> hash-function-transition.txt, but for some of the issues I don't know\n> the answer, and I had some questions after giving this another read.\n\nThanks for looking it over!  Let's go. :)\n\n[...]\n>> Objective\n>> ---------\n>> Migrate Git from SHA-1 to a stronger hash function.\n>\n> Should way say \"Migrate Git from SHA-1 to SHA-256\" here instead?\n>\n> Maybe it's overly specific, i.e. really we're also describnig how /any/\n> hash function transition might happen, but having just read this now\n> from start to finish it takes us a really long time to mention (and at\n> first, only offhand) that SHA-256 is the new hash.\n\nWell, the objective really is to migrate to a stronger hash function,\nand that we chose SHA-256 is part of the details of how we chose to do\nthat.  So I think this would be a misleading change.\n\nYou can tell that I'm not just trying to justify after the fact\nbecause the initial version of the design doc at [*] already uses this\nwording, and that version assumed that the hash function was going to\nbe SHA-256.\n\n[*] https://public-inbox.org/git/20170304011251.GA26789@aiede.mtv.corp.google.com/\n\n[...]\n>> Non-Goals\n>> ---------\n>> 1. Add SHA-256 support to Git protocol. This is valuable and the\n>>    logical next step but it is out of scope for this initial design.\n>\n> This is a non-goal according to the docs, but now that we have protocol\n> v2 in git, perhaps we could start specifying or describing how this\n> protocol extension will work?\n\nYes, that would be great!  But I suspect it's cleanest to do so in a\nseparate doc.  That would allow clarifying this part, by pointing to\nthe protocol doc.\n\n[...]\n>> 3. Intermixing objects using multiple hash functions in a single\n>>    repository.\n>\n> But isn't that the goal now per \"Translation table\" & writing both SHA-1\n> and SHA-256 versions of objects?\n\nNo, we don't write both versions of objects.  The translation records\nboth names of an object.\n\n[...]\n>>   - For each object format:\n>>     - 4-byte format identifier (e.g., 'sha1' for SHA-1)\n>\n> So, given that we have 4-byte limit and have decided on SHA-256 are we\n> just going to call this 'sha2'?\n\nGood question.  'sha2' sounds fine to me.  If we want to do\nSHA-512/256 later, say, we'd just have to come up with a name for that\nat that point (and it doesn't have to be ASCII).\n\n>                                 That might be confusingly ambiguous\n\nThis is a binary format.  Are you really worried that people are going\nto misinterpret the magic numbers it contains?\n\n> since SHA2 is a standard with more than just SHA-256, maybe 's256', or\n> maybe we should give this 8 bytes with trailing \\0s so we can have\n> \"SHA-1\\0\\0\\0\" and \"SHA-256\\0\"?\n\nFor what it's worth, if that's the alternative, I'd rather have four\nrandom bytes.\n\n[...]\n>> The loose object index is protected against concurrent writes by a\n>> lock file $GIT_OBJECT_DIR/loose-object-idx.lock. To add a new loose\n>> object:\n>>\n>> 1. Write the loose object to a temporary file, like today.\n>> 2. Open loose-object-idx.lock with O_CREAT | O_EXCL to acquire the lock.\n>> 3. Rename the loose object into place.\n>> 4. Open loose-object-idx with O_APPEND and write the new object\n>> 5. Unlink loose-object-idx.lock to release the lock.\n>>\n>> To remove entries (e.g. in \"git pack-refs\" or \"git-prune\"):\n>>\n>> 1. Open loose-object-idx.lock with O_CREAT | O_EXCL to acquire the\n>>    lock.\n>> 2. Write the new content to loose-object-idx.lock.\n>> 3. Unlink any loose objects being removed.\n>> 4. Rename to replace loose-object-idx, releasing the lock.\n>\n> Do we expect multiple concurrent writers to poll the lock if they can't\n> aquire it right away? I.e. concurrent \"git commit\" would block? Has this\n> overall approach been benchmarked somewhere?\n\nGit doesn't support concurrent \"git commit\" today.\n\nMy feeling is that if loose object writing becomes a performance\nproblem, we should switch to writing packfiles instead (as \"git\nreceive-pack\" already does).  So when there's a choice between better\nperformance of writing loose objects and simplicity, I lean toward\nsimplicity (though that's not absolute, there are definitely tradeoffs\nto be made).\n\nEarlier discussion about this had sharded loose object indices for\neach xy/ subdir.  It was more complicated, for not much gain.\n\n[...]\n> Maybe I've missed some subtlety where that won't work, I'm just\n> concerned that something that's writing a lot of objects in parallel\n> will be slowed down (e.g. the likes of BFG repo cleaner).\n\nBFG repo cleaner is an application like fast-import that is a good fit\nfor writing packs, not loose objects.\n\n[...]\n>> Since all operations that make new objects (e.g., \"git commit\") add\n>> the new objects to the corresponding index, this mapping is possible\n>> for all objects in the object store.\n>\n> Are we going to need a midx version of these mapping files? How does\n> midx fit into this picture? Perhaps it's too obscure to worry about...\n\nThat's a great question!  I think the simplest answer is to have a\nmidx only for the primary object format and fall back to using\nordinary idx files for the others.\n\nThe midx format already has a field for hash function (thanks,\nDerrick!).\n\n[...]\n>> 5. clean up: remove the SHA-1 based pack file, index, and\n>>    topologically sorted list obtained from the server in steps 1\n>>    and 2.\n>\n> Doesn't this process require us to implement a \"fetch quarantine\"? Least\n> we have (e.g. other concurrent fetches) referencing those new SHA-1\n> objects we've fetched in a pack that we'll remove in step #5?\n\nDuring a fetch today, objects aren't accessible until the\ncorresponding .idx file has been put in place.\n\n[...]\n>> The user can also explicitly specify which format to use for a\n>> particular revision specifier and for output, overriding the mode. For\n>> example:\n>>\n>> git --output-format=sha1 log abac87a^{sha1}..f787cac^{sha256}\n>\n> How is this going to interact with other peel syntax? I.e. now we have\n> <object>^{commit} <sha>^{tag} etc. It seems to me we'll need not ^{sha1}\n> but ^{sha1:<current_type>}, e.g. ^{sha1:commit} or ^{sha1:tag}, with\n> current ^{} being a synonym for ^{sha1:}.\n>\n> Or is this expected to be chained, as e.g. <object>^{tag}^{sha256} ?\n\nGreat question.  The latter (well, <hexdigits>^{sha256}^{tag}, not the\nother way around).\n\n>> Transition plan\n>> ---------------\n>\n> One thing that's not covered in this document at all, which I feel is\n> missing, is how we're going to handle references to old commit IDs in\n> commit messages, bug trackers etc. once we go through the whole\n> migration process.\n>\n> I.e. are users who expect to be able to read old history and \"git show\n> <sha1 I found>\" expected to maintain a repository that has a live\n> sha1<->sha256 mapping forever, or could we be smarter about this and\n> support some sort of marker in the repository saying \"maintain the\n> mapping up until this point\".\n\nThat's a good question, too.  My feeling is that such a selective\nmapping could be invented later and would want to work differently\nthan this design.  The important thing with this design is that the\ninformation is not lost, so the door to implementing that is not\nclosed.\n\nAs a brief strawman of what I mean, I wouldn't be surprised if\nprojects want to distribute a simple signed flat sha1<->sha256 mapping\ntable for commits from \"before the SHA-256 era\", and Git could learn\nto consume that.\n\nThanks,\nJonathan\n"},{"id":"356429","messageId":"20180824015438.GF99542@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"20180824014007.GF535143@genre.crustytoothpaste.net","subject":"Re: Questions about the hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-24T01:54:38Z","receivedAt":"2018-08-24T01:54:43Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nbrian m. carlson wrote:\n> On Thu, Aug 23, 2018 at 04:02:51PM +0200, Ævar Arnfjörð Bjarmason wrote:\n\n>>> 1. Add SHA-256 support to Git protocol. This is valuable and the\n>>>    logical next step but it is out of scope for this initial design.\n>>\n>> This is a non-goal according to the docs, but now that we have protocol\n>> v2 in git, perhaps we could start specifying or describing how this\n>> protocol extension will work?\n>\n> I have code that does this.  The reason is that the first stage of the\n[nice explanation snipped]\n> I hope to be able to spend some time documenting this in a little bit.\n> I have documentation for that code in my branch, but I haven't sent it\n> in yet.\n\nYay!\n\n> I realize I have a lot of code that has not been sent in yet, but I also\n> tend to build on my own series a lot, and I probably need to be a bit\n> better about extracting reusable pieces that can go in independently\n> without waiting for the previous series to land.\n\nFor what it's worth, even if it all is in one commit with message\n\"wip\", I think I'd benefit from being able to see this code.  I can\npromise not to critique it, and to only treat it as a rough\npremonition of the future.\n\n[...]\n> For SHA-1, I have 0x73686131, which is \"sha1\", big-endian, and for\n> SHA-256, I have 0x73323536, which is \"s256\", big-endian.  The former is\n> in the codebase already; the latter, in my hash-impl branch.\n\nI mentioned in another reply that \"sha2\" sounds fine.  \"s256\" of\ncourse also sounds fine to me.  Thanks to Ævar for asking so that we\nhave the reminder to pin it down in the doc.\n\nThanks,\nJonathan\n"},{"id":"356431","messageId":"20180824025123.GA186259@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"878t4xfaes.fsf@evledraar.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-24T02:51:23Z","receivedAt":"2018-08-24T02:51:28Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Ævar Arnfjörð Bjarmason wrote:\n\n>> Objective\n>> ---------\n>> Migrate Git from SHA-1 to a stronger hash function.\n>\n> Should way say \"Migrate Git from SHA-1 to SHA-256\" here instead?\n>\n> Maybe it's overly specific, i.e. really we're also describnig how /any/\n> hash function transition might happen, but having just read this now\n> from start to finish it takes us a really long time to mention (and at\n> first, only offhand) that SHA-256 is the new hash.\n\nI answered this question in my other reply, but my answer missed the\npoint.\n\nI think it would be fine for this to say \"Migrate Git from SHA-1 to a\nstronger hash function (SHA-256)\".  More importantly, I think the\nBackground section should say something about SHA-256 --- e.g. how about\nreplacing the sentence\n\n  SHA-1 still possesses the other properties such as fast object\n  lookup and safe error checking, but other hash functions are equally\n  suitable that are believed to be cryptographically secure.\n\nwith something about SHA-256?\n\nRereading the background section, I see some other bits that could be\nclarified, too.  It has a run-on sentence:\n\n  Thus Git has in effect already migrated to a new hash that isn't\n  SHA-1 and doesn't share its vulnerabilities, its new hash function\n  just happens to produce exactly the same output for all known\n  inputs, except two PDFs published by the SHAttered researchers, and\n  the new implementation (written by those researchers) claims to\n  detect future cryptanalytic collision attacks.\n\nThe \",\" after vulnerabilities should be a period, ending the sentence.\nMy understanding is that sha1collisiondetection's safe-hash is meant\nto protect against known attacks and that the code is meant to be\nadaptable for future attacks of the same kind (by updating the list of\ndisturbance vectors), but it doesn't claim to guard against future\nnovel cryptanalysis methods that haven't been published yet.\n\nThanks,\nJonathan\n"},{"id":"356435","messageId":"20180824044720.GG535143@genre.crustytoothpaste.net","threadId":"49202","inReplyTo":"20180824015438.GF99542@aiede.svl.corp.google.com","subject":"Re: Questions about the hash function transition","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2018-08-24T04:47:20Z","receivedAt":"2018-08-24T04:49:31Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Thu, Aug 23, 2018 at 06:54:38PM -0700, Jonathan Nieder wrote:\n> brian m. carlson wrote:\n> > I realize I have a lot of code that has not been sent in yet, but I also\n> > tend to build on my own series a lot, and I probably need to be a bit\n> > better about extracting reusable pieces that can go in independently\n> > without waiting for the previous series to land.\n> \n> For what it's worth, even if it all is in one commit with message\n> \"wip\", I think I'd benefit from being able to see this code.  I can\n> promise not to critique it, and to only treat it as a rough\n> premonition of the future.\n\nIt's in my object-id-partn branch at https://github.com/bk2204/git.\n\nIt doesn't do protocol v2 yet, but it does do protocol v1.  It is, of\ncourse, subject to change (especially naming) depending on what the list\nthinks is most appropriate.\n-- \nbrian m. carlson: Houston, Texas, US\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"356436","messageId":"20180824045254.GA214696@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"20180824044720.GG535143@genre.crustytoothpaste.net","subject":"Re: Questions about the hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-24T04:52:54Z","receivedAt":"2018-08-24T04:52:59Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"brian m. carlson wrote:\n> On Thu, Aug 23, 2018 at 06:54:38PM -0700, Jonathan Nieder wrote:\n\n>> For what it's worth, even if it all is in one commit with message\n>> \"wip\", I think I'd benefit from being able to see this code.  I can\n>> promise not to critique it, and to only treat it as a rough\n>> premonition of the future.\n>\n> It's in my object-id-partn branch at https://github.com/bk2204/git.\n>\n> It doesn't do protocol v2 yet, but it does do protocol v1.  It is, of\n> course, subject to change (especially naming) depending on what the list\n> thinks is most appropriate.\n\n $ git diff --shortstat origin/master...bmc/object-id-partn\n  185 files changed, 2263 insertions(+), 1535 deletions(-)\n\nBeautiful.  Thanks much for this.\n\nSincerely,\nJonathan\n"},{"id":"356660","messageId":"nycvar.QRO.7.76.6.1808281402510.73@tvgsbejvaqbjf.bet","threadId":"49202","inReplyTo":"20180824014703.GE99542@aiede.svl.corp.google.com","subject":"Re: Questions about the hash function transition","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2018-08-28T12:04:02Z","receivedAt":"2018-08-28T12:04:17Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 23 Aug 2018, Jonathan Nieder wrote:\n\n> Ævar Arnfjörð Bjarmason wrote:\n> \n> [...]\n> >> Since all operations that make new objects (e.g., \"git commit\") add\n> >> the new objects to the corresponding index, this mapping is possible\n> >> for all objects in the object store.\n> >\n> > Are we going to need a midx version of these mapping files? How does\n> > midx fit into this picture? Perhaps it's too obscure to worry about...\n> \n> That's a great question!  I think the simplest answer is to have a\n> midx only for the primary object format and fall back to using\n> ordinary idx files for the others.\n> \n> The midx format already has a field for hash function (thanks,\n> Derrick!).\n\nRelated: I wondered whether we could simply leverage the midx code for the\nbidirectional SHA-1 <-> SHA-256 mapping, as it strikes me as very similar\nin concept and challenges.\n\nCiao,\nDscho"},{"id":"356685","messageId":"c098b0c6-1062-6581-81a9-7ce15f3738de@gmail.com","threadId":"49202","inReplyTo":"nycvar.QRO.7.76.6.1808281402510.73@tvgsbejvaqbjf.bet","subject":"Re: Questions about the hash function transition","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2018-08-28T12:49:16Z","receivedAt":"2018-08-28T12:49:20Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/28/2018 8:04 AM, Johannes Schindelin wrote:\n> Hi,\n>\n> On Thu, 23 Aug 2018, Jonathan Nieder wrote:\n>\n>> Ævar Arnfjörð Bjarmason wrote:\n>>\n>> [...]\n>>>> Since all operations that make new objects (e.g., \"git commit\") add\n>>>> the new objects to the corresponding index, this mapping is possible\n>>>> for all objects in the object store.\n>>> Are we going to need a midx version of these mapping files? How does\n>>> midx fit into this picture? Perhaps it's too obscure to worry about...\n>> That's a great question!  I think the simplest answer is to have a\n>> midx only for the primary object format and fall back to using\n>> ordinary idx files for the others.\n>>\n>> The midx format already has a field for hash function (thanks,\n>> Derrick!).\n> Related: I wondered whether we could simply leverage the midx code for the\n> bidirectional SHA-1 <-> SHA-256 mapping, as it strikes me as very similar\n> in concept and challenges.\n\nIf we would like such a mapping, then I would propose the following:\n\n1. The object store has everything in SHA-256, so the HASH_LEN parameter \nof the multi-pack-index is 32.\n\n2. We create an optional chunk to add to the multi-pack-index that \nstores the SHA-1 for each object. This list would be in lex order.\n\n3. We create two optional chunks that store the bijection between \nSHA-256 and SHA-1: the first is a list of integers i_0, i_1, ..., \ni_{N-1} such that i_k is the position in the SHA-1 list corresponding to \nthe kth SHA-256. The second is a list of integers j_0, j_1, ..., j_{N-1} \nsuch that j_k is the position in the SHA-256 list of the kth SHA-1.\n\nI'm not super-familiar with how the transition plan specifically needs \nthis mapping, but it seems like a good place to put it.\n\nThanks,\n\n-Stolee\n\n"},{"id":"356695","messageId":"87h8jeeh2e.fsf@evledraar.gmail.com","threadId":"49202","inReplyTo":"878t4xfaes.fsf@evledraar.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-08-28T13:50:17Z","receivedAt":"2018-08-28T13:50:23Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Aug 23 2018, Ævar Arnfjörð Bjarmason wrote:\n\n>> Transition plan\n>> ---------------\n>\n> One thing that's not covered in this document at all, which I feel is\n> missing, is how we're going to handle references to old commit IDs in\n> commit messages, bug trackers etc. once we go through the whole\n> migration process.\n>\n> I.e. are users who expect to be able to read old history and \"git show\n> <sha1 I found>\" expected to maintain a repository that has a live\n> sha1<->sha256 mapping forever, or could we be smarter about this and\n> support some sort of marker in the repository saying \"maintain the\n> mapping up until this point\".\n>\n> Then, along with some v2 protocol extension to transfer such a\n> historical mapping (and perhaps a default user option to request it)\n> we'd be guaranteed to be able to read old log messages and \"git show\"\n> them, and servers could avoid breaking past URLs without maintaining the\n> mapping going forward.\n>\n> One example of this on the server is that on GitLab (I don't know how\n> GitHub does this) when you reference a commit from e.g a bug, a\n> refs/keep-around/<sha1> is created, to make sure it doesn't get GC'd.\n>\n> Those sorts of hosting providers would like to not break *existing*\n> links, without needing to forever maintain a bidirectional mapping.\n\nConsidering this a bit more, I think this would nicely fall under what I\nsuggested in\nhttps://public-inbox.org/git/874ll3yd75.fsf@evledraar.gmail.com/\n\nI.e. the interface that's now proposed / documented is fairly\ninelastic. I.e.:\n\n    [extensions]\n        objectFormat = sha256\n        compatObjectFormat = sha1\n\nIf we instead had something like clean/smudge filters:\n\n    [extensions]\n        objectFilter = sha256-to-sha1\n        compatObjectFormat = sha1\n    [objectFilter \"sha256-to-sha1\"]\n        clean  = ...\n        smudge = ...\n\nWe could apply arbitrary transformations on objects through filters\nwhich would accept/return some simple format requesting them to\ntranslate such-and-such objects, and would either return object\nnames/types under which to store them, or \"nothing to do\".\n\nSo we could also have filters that would munge the contents of objects\nbetween local & remote (for e.g. this \"use a public remote host for\nstoring an encrypted repo\" that'll fsck on their end) use-case, but also\ne.g. be able to pass arguments to the filters saying that only commits\nolder than so-and-so are to have a reverse mapping (for looking up old\ncommits), or just ones on some branch etc.\n\nIt wouldn't be any slower than the current proposal, since some subset\nof it would be picked up and implemented in C directly via some fast\npath, similar to the proposal that e.g. some encoding filters be\nimplemented as built-ins.\n\nBut by having it be more extendable it'll be easy to e.g. pass options,\nor implement custom transformations.\n\nWe're still far away from reviewing patches to implement this, but in\nanticipation of that I'd like to see what people think about\nfuture-proofing this objectFilter syntax.\n"},{"id":"356696","messageId":"CA+WKDT1k1SpHQmUKunV+vC+VLBfTBjZBgw+n4NeTE=oKxWL-Sg@mail.gmail.com","threadId":"49202","inReplyTo":"87h8jeeh2e.fsf@evledraar.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Edward Thomson","fromEmail":"ethomson@edwardthomson.com","sentAt":"2018-08-28T14:15:35Z","receivedAt":"2018-08-28T14:15:40Z","isPatch":false,"sender":{"key":"ethomson@edwardthomson.com","avatar":"https://avatars.githubusercontent.com/u/1130014?v=4"},"body":"On Tue, Aug 28, 2018 at 2:50 PM, Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n> If we instead had something like clean/smudge filters:\n>\n>     [extensions]\n>         objectFilter = sha256-to-sha1\n>         compatObjectFormat = sha1\n>     [objectFilter \"sha256-to-sha1\"]\n>         clean  = ...\n>         smudge = ...\n>\n> We could apply arbitrary transformations on objects through filters\n> which would accept/return some simple format requesting them to\n> translate such-and-such objects, and would either return object\n> names/types under which to store them, or \"nothing to do\".\n\nIf I'm understanding you correctly, then on the libgit2 side, I'm very much\nopposed to this proposal.  We never execute commands, nor do I want to start\nthinking that we can do so arbitrarily.  We run in environments where that's\na non-starter\n\nAt present, in libgit2, users can provide their own mechanism for running\nclean/smudge filters.  But hash transformation / compatibility is going to\nbe a crucial compatibility component.  So this is not something that we\ncould simply opt out of or require users to implement themselves.\n\n-ed\n"},{"id":"356699","messageId":"87ftyyedqd.fsf@evledraar.gmail.com","threadId":"49202","inReplyTo":"CA+WKDT1k1SpHQmUKunV+vC+VLBfTBjZBgw+n4NeTE=oKxWL-Sg@mail.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-08-28T15:02:18Z","receivedAt":"2018-08-28T15:02:24Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Aug 28 2018, Edward Thomson wrote:\n\n> On Tue, Aug 28, 2018 at 2:50 PM, Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n>> If we instead had something like clean/smudge filters:\n>>\n>>     [extensions]\n>>         objectFilter = sha256-to-sha1\n>>         compatObjectFormat = sha1\n>>     [objectFilter \"sha256-to-sha1\"]\n>>         clean  = ...\n>>         smudge = ...\n>>\n>> We could apply arbitrary transformations on objects through filters\n>> which would accept/return some simple format requesting them to\n>> translate such-and-such objects, and would either return object\n>> names/types under which to store them, or \"nothing to do\".\n>\n> If I'm understanding you correctly, then on the libgit2 side, I'm very much\n> opposed to this proposal.  We never execute commands, nor do I want to start\n> thinking that we can do so arbitrarily.  We run in environments where that's\n> a non-starter\n\nI'm being unclear. I'm suggesting that we slightly amend the syntax of\nwhat we're proposing to put in the .git/config to leave the door open\nfor *optionally* doing arbitrary mappings.\n\nIt would still work exactly the same internally for the common\nsha1<->sha256 case, i.e. neither git, libgit, jgit or anyone else would\nneed to shell out to anything.\n\nThey'd just pick up that common case and handle it internally, similar\nto how e.g. the crlf filter (v.s. full clean/smudge support) works in\ngit & libgit2:\nhttps://github.com/libgit2/libgit2/blob/master/tests/filter/crlf.c\n\nSo the sha256<->sha1 support would be an implicit built-in like crlf, it\nwould just leave the door open to having something like git-lfs.\n\nNow what does that really mean? And I admit I may be missing something\nhere.\n\nUnlike smudge/clean filters we're going to be constrained by having\nhashes of length 20 or 32, locally & remotely, since we wouldn't want to\nsupport arbitrary lengths, but with relatively small changes it'll allow\nfor changing just:\n\n    # local  remote\n    sha256<->sha1\n\nTo also support:\n\n    # local  remote\n    fn(sha1)<->fn(sha1)\n    fn(sha1)<->fn(sha256)\n    fn(sha256)<->fn(sha1)\n    fn(sha256)<->fn(sha256)\n\nWhere fn() is some hook you'd provide to hook into the bits where we're\ne.g. unpacking SHA-1 objects from the remote, and writing them locally\nas SHA-256, except instead of (as we do by default) writing:\n\n    SHA256_map(sha256(content)) = content\n\nYou'd write:\n\n    SHA256_map(sha256(fn(content))) = fn(content)\n\nWhere fn() would need to be idempotent.\n\nNow, why is this useful or worth considering? As noted in the E-Mail I\nlinked to it allows for some novel use cases for doing local to remote\nobject translation.\n\nBut really, I'm not suggesting that *that* is something we should\nconsider. *All* I'm saying is that given the experience of how we\nstarted out with stuff like built-in \"crlf\", and then grew smudge/clean\nfilters, that it's worth considering what sort of .git/config key-value\npairs we'd pick that would yield themselves to such future extensions,\nshould that be something we deem to be a good idea in the future.\n\nBecause if we don't we've lost nothing, but if we do we'd need to\nsupport two sets of config syntaxes to do those two related things.\n\n> At present, in libgit2, users can provide their own mechanism for running\n> clean/smudge filters.  But hash transformation / compatibility is going to\n> be a crucial compatibility component.  So this is not something that we\n> could simply opt out of or require users to implement themselves.\n\nIndeed.\n"},{"id":"356705","messageId":"xmqqzhx6h4ux.fsf@gitster-ct.c.googlers.com","threadId":"49202","inReplyTo":"CA+WKDT1k1SpHQmUKunV+vC+VLBfTBjZBgw+n4NeTE=oKxWL-Sg@mail.gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-28T15:45:42Z","receivedAt":"2018-08-28T15:45:47Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Edward Thomson <ethomson@edwardthomson.com> writes:\n\n> If I'm understanding you correctly, then on the libgit2 side, I'm very much\n> opposed to this proposal.  We never execute commands, nor do I want to start\n> thinking that we can do so arbitrarily.  We run in environments where that's\n> a non-starter\n>\n> At present, in libgit2, users can provide their own mechanism for running\n> clean/smudge filters.  But hash transformation / compatibility is going to\n> be a crucial compatibility component.  So this is not something that we\n> could simply opt out of or require users to implement themselves.\n\nWhile I suspect the \"apparent flexibility\" does not equal to \"we\nmust be able to run arbitrary external programs\" in the proposal, I\ndo agree that hash transformation MUST NOT be configurable like\nthis.  We do not want to add random source of incompatible mappings\nwhen there is no need to introduce confusion.\n\nIf old object names under old hash users find in log messages and\nother places need to be easily looked up in a repository that has\nbeen converted, then:\n\n (1) get_sha1() equivalent in the new world should learn to fall\n     back to use old hash when there is no object with that name\n     under new hash;\n\n (2) in addition to the above fallback, there should be a syntax to\n     explicitly tell that function that it is using the old hash;\n\n (3) get_commit_buffer() should learn to optionally allow converting\n     old hash in log messages to new ones, in a way similar to how\n     textconv filter can be specified by the end-users to make\n     binary blob easier to grok by text-based tools (the important\n     part is that such a filter does not have to be limited to\n     \"upgrade hash algorithm\"---it can be more general \"correct\n     misspelt words automatically\" filter).\n\nWith 1+2, you can say \"git log $sha1\" and also \"git log sha1:$sha1\"\nto disambiguate.  3 would be icing on the cake.\n"},{"id":"356709","messageId":"20180828171113.GA23314@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"nycvar.QRO.7.76.6.1808281402510.73@tvgsbejvaqbjf.bet","subject":"Re: Questions about the hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-28T17:11:13Z","receivedAt":"2018-08-28T17:11:25Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nJohannes Schindelin wrote:\n> On Thu, 23 Aug 2018, Jonathan Nieder wrote:\n> > Ævar Arnfjörð Bjarmason wrote:\n\n>>> Are we going to need a midx version of these mapping files? How does\n>>> midx fit into this picture? Perhaps it's too obscure to worry about...\n>>\n>> That's a great question!  I think the simplest answer is to have a\n>> midx only for the primary object format and fall back to using\n>> ordinary idx files for the others.\n>>\n>> The midx format already has a field for hash function (thanks,\n>> Derrick!).\n>\n> Related: I wondered whether we could simply leverage the midx code for the\n> bidirectional SHA-1 <-> SHA-256 mapping, as it strikes me as very similar\n> in concept and challenges.\n\nInteresting: tell me more.\n\nMy first instinct is to prefer the idx-based design that is already\ndescribed in the design doc.  If we want to change that, we should\nhave a motivating reason.\n\nMidx is designed to be optional and to not necessarily cover all\nobjects, so it doesn't seem like a good fit.\n\nThanks,\nJonathan\n"},{"id":"356710","messageId":"20180828171200.GB23314@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"c098b0c6-1062-6581-81a9-7ce15f3738de@gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-28T17:12:00Z","receivedAt":"2018-08-28T17:12:05Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Derrick Stolee wrote:\n\n> I'm not super-familiar with how the transition plan specifically needs this\n> mapping, but it seems like a good place to put it.\n\nWould you mind reading it through and letting me know your thoughts?\nMore eyes can't hurt.\n\nThanks,\nJonathan\n"},{"id":"356820","messageId":"877ek9edsa.fsf@evledraar.gmail.com","threadId":"49202","inReplyTo":"20180824014703.GE99542@aiede.svl.corp.google.com","subject":"How is the ^{sha256} peel syntax supposed to work?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-08-29T09:13:25Z","receivedAt":"2018-08-29T09:13:31Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Aug 24 2018, Jonathan Nieder wrote:\n\n> Hi,\n>\n> Ævar Arnfjörð Bjarmason wrote:\n>\n>>> git --output-format=sha1 log abac87a^{sha1}..f787cac^{sha256}\n>>\n>> How is this going to interact with other peel syntax? I.e. now we have\n>> <object>^{commit} <sha>^{tag} etc. It seems to me we'll need not ^{sha1}\n>> but ^{sha1:<current_type>}, e.g. ^{sha1:commit} or ^{sha1:tag}, with\n>> current ^{} being a synonym for ^{sha1:}.\n>>\n>> Or is this expected to be chained, as e.g. <object>^{tag}^{sha256} ?\n>\n> Great question.  The latter (well, <hexdigits>^{sha256}^{tag}, not the\n> other way around).\n\nSince nobody's chimed in with an answer, and I suspect many have an\nadversion to that big thread I thought I'd spin out just this small\nquestion into its own thread.\n\nbrian m. carlson did some prep work for this in his just-submitted\nhttps://public-inbox.org/git/20180829005857.980820-2-sandals@crustytoothpaste.net/\n\nI was going to work on some of the peel code soon (digging up the type\ndisambiguation patches I still need to re-submit), so could do this\nwhile I'm at it, i.e. implement ^{sha1}.\n\nBut as noted above it's not clear how it should work. Jonathan's\nchaining suggestion (<hexdigits>^{sha256}^{tag} not\n<hexdigits>^{tag}^{sha256}) makes more sense than mine, but is that what\nwe're going for, or ^{sha256:tag}?\n"},{"id":"356833","messageId":"nycvar.QRO.7.76.6.1808291458480.71@tvgsbejvaqbjf.bet","threadId":"49202","inReplyTo":"20180828171113.GA23314@aiede.svl.corp.google.com","subject":"Re: Questions about the hash function transition","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2018-08-29T13:09:07Z","receivedAt":"2018-08-29T13:09:21Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Jonathan,\n\nOn Tue, 28 Aug 2018, Jonathan Nieder wrote:\n\n> Johannes Schindelin wrote:\n> > On Thu, 23 Aug 2018, Jonathan Nieder wrote:\n> > > Ævar Arnfjörð Bjarmason wrote:\n> \n> >>> Are we going to need a midx version of these mapping files? How does\n> >>> midx fit into this picture? Perhaps it's too obscure to worry\n> >>> about...\n> >>\n> >> That's a great question!  I think the simplest answer is to have a\n> >> midx only for the primary object format and fall back to using\n> >> ordinary idx files for the others.\n> >>\n> >> The midx format already has a field for hash function (thanks,\n> >> Derrick!).\n> >\n> > Related: I wondered whether we could simply leverage the midx code for\n> > the bidirectional SHA-1 <-> SHA-256 mapping, as it strikes me as very\n> > similar in concept and challenges.\n> \n> Interesting: tell me more.\n> \n> My first instinct is to prefer the idx-based design that is already\n> described in the design doc.  If we want to change that, we should\n> have a motivating reason.\n> \n> Midx is designed to be optional and to not necessarily cover all\n> objects, so it doesn't seem like a good fit.\n\nRight.\n\nWhat I meant was to leverage the midx code, not the .midx files.\n\nMy comment was motivated by my realizing that both the SHA-1 <-> SHA-256\nmapping and the MIDX code have to look up (in a *fast* way) information\nwith hash values as keys. *And* this information is immutable. *And* the\namount of information should grow with new objects being added to the\ndatabase.\n\nI know that Stolee performed a bit of performance testing regarding\ndifferent data structures to use in MIDX. We could benefit from that\ntesting by using not only the results from those tests, but also the code.\n\nIIRC one of the insights was that packs are a natural structure that\ncan be used for the MIDX mapping, too (you could, for example, store the\nSHA-1 <-> SHA-256 mapping *only* for objects inside packs, and re-generate\nthem on the fly for loose objects all the time).\n\nStolee can speak with much more competence and confidence about this,\nthough, whereas all of what I said above is me waving my hands quite\nfrantically.\n\nCiao,\nDscho"},{"id":"356836","messageId":"04300dbc-622a-c8cb-172e-985726249a8e@gmail.com","threadId":"49202","inReplyTo":"nycvar.QRO.7.76.6.1808291458480.71@tvgsbejvaqbjf.bet","subject":"Re: Questions about the hash function transition","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2018-08-29T13:27:36Z","receivedAt":"2018-08-29T13:27:40Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/29/2018 9:09 AM, Johannes Schindelin wrote:\n> Hi Jonathan,\n>\n> On Tue, 28 Aug 2018, Jonathan Nieder wrote:\n>\n>> Johannes Schindelin wrote:\n>>> On Thu, 23 Aug 2018, Jonathan Nieder wrote:\n>>>> Ævar Arnfjörð Bjarmason wrote:\n>>>>> Are we going to need a midx version of these mapping files? How does\n>>>>> midx fit into this picture? Perhaps it's too obscure to worry\n>>>>> about...\n>>>> That's a great question!  I think the simplest answer is to have a\n>>>> midx only for the primary object format and fall back to using\n>>>> ordinary idx files for the others.\n>>>>\n>>>> The midx format already has a field for hash function (thanks,\n>>>> Derrick!).\n>>> Related: I wondered whether we could simply leverage the midx code for\n>>> the bidirectional SHA-1 <-> SHA-256 mapping, as it strikes me as very\n>>> similar in concept and challenges.\n>> Interesting: tell me more.\n>>\n>> My first instinct is to prefer the idx-based design that is already\n>> described in the design doc.  If we want to change that, we should\n>> have a motivating reason.\n>>\n>> Midx is designed to be optional and to not necessarily cover all\n>> objects, so it doesn't seem like a good fit.\n\nIt is optional, but shouldn't this mode where a Git repo that needs to \nknow about two different versions of all files be optional? Or at least \ntemporary?\n\nThe multi-pack-index is intended to cover all packed objects, so covers \nthe same number of objects as an IDX-based strategy. If we are \nrebuilding the repo from scratch by translating the hashes, then \"being \ntoo big to repack\" is probably not a problem, so we would expect a \nsingle IDX file anyway.\n\nIn my opinion, whatever we do for the IDX-based approach will need to be \nduplicated in the multi-pack-index. The multi-pack-index does have a \nnatural mechanism (optional chunks) for inserting this data without \nincrementing the version number.\n\n> Right.\n>\n> What I meant was to leverage the midx code, not the .midx files.\n>\n> My comment was motivated by my realizing that both the SHA-1 <-> SHA-256\n> mapping and the MIDX code have to look up (in a *fast* way) information\n> with hash values as keys. *And* this information is immutable. *And* the\n> amount of information should grow with new objects being added to the\n> database.\n\nI'm unsure what this means, as the multi-pack-index simply uses \nbsearch_hash() to find hashes in the list. The same method is used for \nIDX lookups.\n\n> I know that Stolee performed a bit of performance testing regarding\n> different data structures to use in MIDX. We could benefit from that\n> testing by using not only the results from those tests, but also the code.\n\nI did test ways to use something other than bsearch_hash(), such as \nusing a 65,536-entry fanout table for lookups using the first two bytes \nof a hash (tl;dr: it speeds things up a bit, but the super-small \nimprovement is probably not worth the space and complexity). I've also \ntoyed with the idea of using interpolation search inside bsearch_hash(), \nbut I haven't had time to do that.\n\n> IIRC one of the insights was that packs are a natural structure that\n> can be used for the MIDX mapping, too (you could, for example, store the\n> SHA-1 <-> SHA-256 mapping *only* for objects inside packs, and re-generate\n> them on the fly for loose objects all the time).\n>\n> Stolee can speak with much more competence and confidence about this,\n> though, whereas all of what I said above is me waving my hands quite\n> frantically.\n\nI understand the hesitation to pair such an important feature (hash \ntransition) to a feature that hasn't even shipped. We will need to see \nhow things progress on both fronts to see how mature the \nmulti-pack-index is when we need this transition table.\n\nThanks,\n\n-Stolee\n\n"},{"id":"356842","messageId":"74787b14-ea63-de76-4eba-ce322aa7b1d2@gmail.com","threadId":"49202","inReplyTo":"04300dbc-622a-c8cb-172e-985726249a8e@gmail.com","subject":"Re: Questions about the hash function transition","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2018-08-29T14:43:37Z","receivedAt":"2018-08-29T14:43:49Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 8/29/2018 9:27 AM, Derrick Stolee wrote:\n> On 8/29/2018 9:09 AM, Johannes Schindelin wrote:\n>>\n>> What I meant was to leverage the midx code, not the .midx files.\n>>\n>> My comment was motivated by my realizing that both the SHA-1 <-> SHA-256\n>> mapping and the MIDX code have to look up (in a *fast* way) information\n>> with hash values as keys. *And* this information is immutable. *And* the\n>> amount of information should grow with new objects being added to the\n>> database.\n>\n> I'm unsure what this means, as the multi-pack-index simply uses \n> bsearch_hash() to find hashes in the list. The same method is used for \n> IDX lookups.\n>\nI talked with Johannes privately, and we found differences in our \nunderstanding of the current multi-pack-index feature. Johannes thought \nthe feature was farther along than it is, specifically related to how \nmuch we value the data in the multi-pack-index when adding objects to \npack-files or repacking. Some of this misunderstanding is due to how the \nequivalent feature works in VSTS (where there is no IDX-file equivalent, \nevery object in the repo is tracked by a multi-pack-index).\n\nI'd like to point out a few things about how the multi-pack-index works \nnow, and how we hope to extend it in the future.\n\nCurrently:\n\n1. Objects are added to the multi-pack-index by adding a new set of \n.idx/.pack file pairs. We scan the .idx file for the objects and offsets \nto add.\n\n2. We re-use the information in the multi-pack-index only to write the \nnew one without re-reading the .pack files that are already covered.\n\n3. If a 'git repack' command deletes a pack-file, then we delete the \nmulti-pack-index. It must be regenerated by 'git multi-pack-index write' \nlater.\n\nIn the current world, the multi-pack-index is completely secondary to \nthe .idx files.\n\nIn the future, I hope these features exist in the multi-pack-index:\n\n1. A stable object order. As objects are added to the multi-pack-index, \nwe assign a distinct integer value to each. As we add objects, those \nintegers values do not change. We can then pair the reachability bitmap \nto the multi-pack-index instead of a specific pack-file (allowing repack \nand bitmap computations to happen asynchronously). The data required to \nstore this object order is very similar to storing the bijection between \nSHA-1 and SHA-256 hashes.\n\n2. Incremental multi-pack-index: Currently, we have only one \nmulti-pack-index file per object directory. We can use a mechanism \nsimilar to the split-index to keep a small number of multi-pack-index \nfiles (at most 3, probably) such that the \n'.git/objects/pack/multi-pack-index' file is small and easy to rewrite, \nwhile it refers to larger '.git/objects/pack/*.midx' files that change \ninfrequently.\n\n3. Multi-pack-index-aware repack: The repacker only knows about the \nmulti-pack-index enough to delete it. We could instead directly \nmanipulate the multi-pack-index during repack, and we could decide to do \nmore incremental repacks based on data stored in the multi-pack-index.\n\nIn conclusion: please keep the multi-pack-index in mind as we implement \nthe transition plan. I'll continue building the feature as planned (the \nnext thing to do after the current series of cleanups is 'git \nmulti-pack-index verify') but am happy to look into other applications \nas we need it.\n\nThanks,\n\n-Stolee\n\n"},{"id":"356862","messageId":"CAGZ79kaGb_TL7SiR4CFGFzrfy2Lotioy76o6sUK4=vZK5qwqNA@mail.gmail.com","threadId":"49202","inReplyTo":"877ek9edsa.fsf@evledraar.gmail.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-08-29T17:51:22Z","receivedAt":"2018-08-29T17:51:36Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Wed, Aug 29, 2018 at 2:13 AM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Fri, Aug 24 2018, Jonathan Nieder wrote:\n>\n> > Hi,\n> >\n> > Ævar Arnfjörð Bjarmason wrote:\n> >\n> >>> git --output-format=sha1 log abac87a^{sha1}..f787cac^{sha256}\n> >>\n> >> How is this going to interact with other peel syntax? I.e. now we have\n> >> <object>^{commit} <sha>^{tag} etc. It seems to me we'll need not ^{sha1}\n> >> but ^{sha1:<current_type>}, e.g. ^{sha1:commit} or ^{sha1:tag}, with\n> >> current ^{} being a synonym for ^{sha1:}.\n> >>\n> >> Or is this expected to be chained, as e.g. <object>^{tag}^{sha256} ?\n> >\n> > Great question.  The latter (well, <hexdigits>^{sha256}^{tag}, not the\n> > other way around).\n>\n> Since nobody's chimed in with an answer, and I suspect many have an\n> adversion to that big thread I thought I'd spin out just this small\n> question into its own thread.\n>\n> brian m. carlson did some prep work for this in his just-submitted\n> https://public-inbox.org/git/20180829005857.980820-2-sandals@crustytoothpaste.net/\n>\n> I was going to work on some of the peel code soon (digging up the type\n> disambiguation patches I still need to re-submit), so could do this\n> while I'm at it, i.e. implement ^{sha1}.\n>\n> But as noted above it's not clear how it should work. Jonathan's\n> chaining suggestion (<hexdigits>^{sha256}^{tag} not\n> <hexdigits>^{tag}^{sha256}) makes more sense than mine, but is that what\n> we're going for, or ^{sha256:tag}?\n\nThe choice of hash seems position independent to me, so as a user\nI would expect both to work at first. Though when looking at more\nsyntax of these expressions, e.g. b9dfa238d5c34~1^2^^, it is\nread left to right, i.e. you arrive at the destination by evaluating\nthe next part of the expression and then jumping around based on\neach expression. And with that model, <hexdigits>^{sha256}^{tree}\ncould mean to obtain the sha256 value of <hexvalue> and then derive\nthe tree from that object, so it is unclear if the tree object would also come\nin sha256 or if we could just return the tree in sha1 notation (as it would\nbe correctly - though confusingly - described that way. The sha256\nconversion happened at an intermediate step.)\n\nSo with that said, I would expect the hash specifier at the end of the chain.\n\nWould the position of the hash specifier make any difference for\nverifying signed tags/commits ? (subtle asking to verify the sha1\nsignature or the sha256 signature explicitly vs asking to verify an object\nthat is given with <hexval> in sha1 or in sha256)\n\nThanks,\nStefan\n"},{"id":"356863","messageId":"20180829175602.GA7547@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"877ek9edsa.fsf@evledraar.gmail.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-29T17:56:02Z","receivedAt":"2018-08-29T17:56:07Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nÆvar Arnfjörð Bjarmason wrote:\n> On Fri, Aug 24 2018, Jonathan Nieder wrote:\n>> Ævar Arnfjörð Bjarmason wrote:\n\n>>> Or is this expected to be chained, as e.g. <object>^{tag}^{sha256} ?\n>>\n>> Great question.  The latter (well, <hexdigits>^{sha256}^{tag}, not the\n>> other way around).\n>\n> Since nobody's chimed in with an answer, and I suspect many have an\n> adversion to that big thread I thought I'd spin out just this small\n> question into its own thread.\n>\n> brian m. carlson did some prep work for this in his just-submitted\n> https://public-inbox.org/git/20180829005857.980820-2-sandals@crustytoothpaste.net/\n>\n> I was going to work on some of the peel code soon (digging up the type\n> disambiguation patches I still need to re-submit), so could do this\n> while I'm at it, i.e. implement ^{sha1}.\n\nCool!\n\n> But as noted above it's not clear how it should work. Jonathan's\n> chaining suggestion (<hexdigits>^{sha256}^{tag} not\n> <hexdigits>^{tag}^{sha256}) makes more sense than mine, but is that what\n> we're going for, or ^{sha256:tag}?\n\nI don't have a strong opinion about this, but since it affects the\ninterpretation of <hexdigits>, my assumption has been that, in the\nspirit of referential transparency, you would put\n'<hexdigits>^{format}' and could put any additional specifiers after\nthat.\n\nIn other words, ^{format} changes the interpretation of <hexdigits> so\nmy assumption is that people would want it to be close by.\n\nBut if something else is easier to implement, we can start with that\nsomething else and figure out whether we like it in review.\n\nThanks,\nJonathan\n"},{"id":"356864","messageId":"20180829175950.GB7547@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"CAGZ79kaGb_TL7SiR4CFGFzrfy2Lotioy76o6sUK4=vZK5qwqNA@mail.gmail.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-29T17:59:50Z","receivedAt":"2018-08-29T17:59:55Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Stefan Beller wrote:\n\n>                  And with that model, <hexdigits>^{sha256}^{tree}\n> could mean to obtain the sha256 value of <hexvalue> and then derive\n> the tree from that object,\n\nWhat does \"the sha256 value of <hexvalue>\" mean?\n\nFor example, in a repository with two objects:\n\n 1. an object with sha1-name abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd\n    and sha256-name ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n\n 2. an object with sha1-name ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n    and sha256-name abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd...\n\nwhat objects would you expect the following to refer to?\n\n  abcdabcd^{sha1}\n  abcdabcd^{sha256}\n  ef01ef01^{sha1}\n  ef01ef01^{sha256}\n\nThanks,\nJonathan\n"},{"id":"356867","messageId":"CAGZ79kZsAFNE0GgUHociSDejwB+HsPVscZ4jcq__sFew7g85Bw@mail.gmail.com","threadId":"49202","inReplyTo":"20180829175950.GB7547@aiede.svl.corp.google.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-08-29T18:34:40Z","receivedAt":"2018-08-29T18:34:55Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Wed, Aug 29, 2018 at 10:59 AM Jonathan Nieder <jrnieder@gmail.com> wrote:\n>\n> Stefan Beller wrote:\n>\n> >                  And with that model, <hexdigits>^{sha256}^{tree}\n> > could mean to obtain the sha256 value of <hexvalue> and then derive\n> > the tree from that object,\n>\n> What does \"the sha256 value of <hexvalue>\" mean?\n\ns/hexvalue/hexdigits/\n..\n\nAnd with that model, <hexdigits>^{sha256}^{tree}\ncould mean to obtain the object using sha256 descriptors\n(for trees/blobs/commits/tags) of <hexdigits> (as defined by\nthe step of the transition plan, it could mean <hexdigits>\nto be interpreted as SHA1 or SHA256 or DWIM).\n\n>\n> For example, in a repository with two objects:\n>\n>  1. an object with sha1-name abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd\n>     and sha256-name ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n>\n>  2. an object with sha1-name ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n>     and sha256-name abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd...\n>\n\nIt would be super cool to have hash values to match vice versa,\nbut for the sake of the example, let's go with that.\n\n> what objects would you expect the following to refer to?\n\nThat generally depends on the step of the transition plan\nin that specific Git repository.\n\nI thought these format specifiers would only describe how to\noutput the object names correctly for now, so it could be\npossible to have:\n\n$ git show abcdabcd^{sha1}\ncommit abcdabcd...\n....\n\n$ git show abcdabcd^{sha256}\ncommit ef01ef01e...\n....\n\nin one step and\n\n$ git show abcdabcd^{sha1}\ncommit ef01ef01e...\n....\n\n$ git show abcdabcd^{sha256}\ncommit abcdabcd...\n....\n\nin another step, and in yet another step it could mean\n\n$ git show abcdabcd^{sha1}\ncommit abcdabcd[...]^{sha1}\n...\n\nBut my question was more hinting to the point that we should not\noverload the syntax to mean much more than either output formatting\nor hash selection.\n\nThe third meaning could be used for verifying objects as we could\nuse this syntax to mean\n\n  \"please verify the signature of the object (as given by ^{hash}\"\n\nor it could mean\n\n  \"please verify the signature of the object as given and ensure that\n    it was signed in this ^{hash} and not in a weaker hash world\".\n\nAnd I would think all the verification should not be folded into this\nnotation for now, but we only want to ask for the output to be\none or the other hash, or we could ask for an object that is\n<hexdigits> in the specified hash, but these two modes depend\non the step in the transition plan.\n\nStefan\n"},{"id":"356868","messageId":"87zhx5c8wo.fsf@evledraar.gmail.com","threadId":"49202","inReplyTo":"20180829175950.GB7547@aiede.svl.corp.google.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-08-29T18:41:43Z","receivedAt":"2018-08-29T18:41:50Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Aug 29 2018, Jonathan Nieder wrote:\n\n> Stefan Beller wrote:\n>\n>>                  And with that model, <hexdigits>^{sha256}^{tree}\n>> could mean to obtain the sha256 value of <hexvalue> and then derive\n>> the tree from that object,\n>\n> What does \"the sha256 value of <hexvalue>\" mean?\n>\n> For example, in a repository with two objects:\n>\n>  1. an object with sha1-name abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd\n>     and sha256-name ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n>\n>  2. an object with sha1-name ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n>     and sha256-name abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd...\n\nI'm not saying this makes sense, or that it doesn't honestly my head's\nstill spinning a bit from this mail exchange (these are the patches I\nneed to re-submit):\nhttps://public-inbox.org/git/878t8txfyf.fsf@evledraar.gmail.com/#t\n\nBut paraphrasing my understanding of what Junio & Jeff are saying in\nthat thread, basically what the peel syntax means is different in the\ntwo completely different scenarios it's used:\n\n 1. When it's being used as <object>^{<thing>}[...^{<thing2>}] AND\n    <object> is an unambiguous SHA1 it's fairly straightforward, i.e. if\n    <object> is a commit and you say ^{tree} it lists the tree SHA-1,\n    but if <object> is e.g. a tree and you say ^{blob} it produces an\n    error, since there's no one blob.\n\n 2. When it's used in the same way, but <object> is an ambiguous SHA1 we\n    fall back on a completely different sort of behavior.\n\n    Now it's, or well, supposed to be, I haven't worked through the\n    feedback and rewritten the patches, this weird sort of filter syntax\n    where <ambiguous_object>^{<type>} will return SHA1s of starting with\n    a prefix of <ambiguous_object> IF the types of such SHA1s could be\n    contained within that type of object.\n\n    So e.g. abcabc^{tree} is supposed to list all tree and blob objects\n    starting with a prefix of abcabc, even though some of the blobs\n    could not be reachable from those trees.\n\n    It doesn't make sense to me, but there it is.\n\nNow, because of this SHA1 v.s. SHA256 thing we have a third case.\n\n> what objects would you expect the following to refer to?\n>\n>   abcdabcd^{sha1}\n>   abcdabcd^{sha256}\n>   ef01ef01^{sha1}\n>   ef01ef01^{sha256}\n\nI still can't really make any sense of why anyone would even want #2 as\ndescribed above, but for this third case I think we should do this:\n\n    abcdabcd^{sha1}   = abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd\n    abcdabcd^{sha256} = ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n    ef01ef01^{sha1}   = ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n    ef01ef01^{sha256} = abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd...\n\nI.e. a really useful thing about this peel syntax is that it's\nforgiving, and will try to optimistically look up what you want.\n\nSo e.g. <hash>^{commit} is not an error if <hash> is already a commit,\nit could be (why are you trying to peel something already peeled!),\nbecause it's useful to be able to feed it a set of things, some of which\nare commits, some of which are tags, and have it always resolve things\nwithout error handling on the caller side.\n\nSimilarly, I think it would be very useful if we just make this work:\n\n    git rev-parse $some_hash^{sha256}^{commit}\n\nAnd not care whether $some_hash is SHA-1 or SHA-256, if it's the former\nwe'd consult the SHA-1 <-> SHA-256 lookup table and go from there, and\nalways return a useful value.\n"},{"id":"356870","messageId":"20180829191232.GC7547@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"87zhx5c8wo.fsf@evledraar.gmail.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-29T19:12:32Z","receivedAt":"2018-08-29T19:12:37Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nÆvar Arnfjörð Bjarmason wrote:\n> On Wed, Aug 29 2018, Jonathan Nieder wrote:\n\n>> what objects would you expect the following to refer to?\n>>\n>>   abcdabcd^{sha1}\n>>   abcdabcd^{sha256}\n>>   ef01ef01^{sha1}\n>>   ef01ef01^{sha256}\n>\n> I still can't really make any sense of why anyone would even want #2 as\n> described above, but for this third case I think we should do this:\n>\n>     abcdabcd^{sha1}   = abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd\n>     abcdabcd^{sha256} = ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n>     ef01ef01^{sha1}   = ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n>     ef01ef01^{sha256} = abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd...\n>\n> I.e. a really useful thing about this peel syntax is that it's\n> forgiving, and will try to optimistically look up what you want.\n\nSorry, I'm still not understanding.\n\nI am not attached to any particular syntax, but what I really want is\nthe following:\n\n\tSomeone who only uses SHA-256 sent me the commit id\n\tabcdabcdabcdabcdabcdabcdabcdabcdabcdabcd... out of band.\n\tShow me that commit.\n\n\tI don't care what object id you show me when you show that\n\tcommit.  If I pass --output-format=sha1, then that means I\n\tcare, and show me the SHA-1.\n\nIn other words, I want the input format and output format completely\ndecoupled.  If I pass ^{sha1}, I am indicating the input format.  To\nspecify the output format, I'd use --output-format instead.\n\nThat lets me mix both hash functions in my input:\n\n\tgit --output-format=sha256 diff abcdabcd^{sha1} abcdabcd^{sha256}\n\nI learned about these two commits out of band from different users,\none who only uses SHA-1 and the other who only uses SHA-256.\n\nIn other words:\n\n[...]\n> Similarly, I think it would be very useful if we just make this work:\n>\n>     git rev-parse $some_hash^{sha256}^{commit}\n>\n> And not care whether $some_hash is SHA-1 or SHA-256, if it's the former\n> we'd consult the SHA-1 <-> SHA-256 lookup table and go from there, and\n> always return a useful value.\n\nThe opposite of this. :)\n\nThanks,\nJonathan\n"},{"id":"356871","messageId":"87y3cpc6bt.fsf@evledraar.gmail.com","threadId":"49202","inReplyTo":"20180829191232.GC7547@aiede.svl.corp.google.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-08-29T19:37:26Z","receivedAt":"2018-08-29T19:37:32Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Aug 29 2018, Jonathan Nieder wrote:\n\n> Hi,\n>\n> Ævar Arnfjörð Bjarmason wrote:\n>> On Wed, Aug 29 2018, Jonathan Nieder wrote:\n>\n>>> what objects would you expect the following to refer to?\n>>>\n>>>   abcdabcd^{sha1}\n>>>   abcdabcd^{sha256}\n>>>   ef01ef01^{sha1}\n>>>   ef01ef01^{sha256}\n>>\n>> I still can't really make any sense of why anyone would even want #2 as\n>> described above, but for this third case I think we should do this:\n>>\n>>     abcdabcd^{sha1}   = abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd\n>>     abcdabcd^{sha256} = ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n>>     ef01ef01^{sha1}   = ef01ef01ef01ef01ef01ef01ef01ef01ef01ef01\n>>     ef01ef01^{sha256} = abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd...\n>>\n>> I.e. a really useful thing about this peel syntax is that it's\n>> forgiving, and will try to optimistically look up what you want.\n>\n> Sorry, I'm still not understanding.\n>\n> I am not attached to any particular syntax, but what I really want is\n> the following:\n>\n> \tSomeone who only uses SHA-256 sent me the commit id\n> \tabcdabcdabcdabcdabcdabcdabcdabcdabcdabcd... out of band.\n> \tShow me that commit.\n\nThis is reasonable.\n\n> \tI don't care what object id you show me when you show that\n> \tcommit.  If I pass --output-format=sha1, then that means I\n> \tcare, and show me the SHA-1.\n>\n> In other words, I want the input format and output format completely\n> decoupled.  If I pass ^{sha1}, I am indicating the input format.  To\n> specify the output format, I'd use --output-format instead.\n\nThis is also a reasonable thing to want, but I don't see how it can be\nsensibly squared with the existing peel syntax.\n\nThe peel syntax <thing>^{commit} doesn't mean <thing> is a commit, it\nmeans that thing might be some thing (commit, tag), and it should be\n(recursively if needed) *resolved* as the thing on the RHS.\n\nSo to be consistent <thing>^{sha1} shouldn't mean <thing> is SHA-1, but\nthat I want a SHA-1 out of <thing>.\n\n> That lets me mix both hash functions in my input:\n>\n> \tgit --output-format=sha256 diff abcdabcd^{sha1} abcdabcd^{sha256}\n\nPresumably you mean something like:\n\n     git diff-tree --raw -r -p bcdabcd^{sha1} abcdabcd^{sha256}\n\nI.e. we don't show any sort of SHAs in diff output, so what would this\n--output-format=sha256 mean?\n\n> I learned about these two commits out of band from different users,\n> one who only uses SHA-1 and the other who only uses SHA-256.\n\nI think for those cases we would just support:\n\n     git diff-tree --raw -r -p bcdabcd abcdabcd\n\nI.e. there's no need to specify the hash type, unless the two happen to\nbe ambiguous, but yeah, if that's the case we'd need to peel them (or\nsupply more hexdigits).\n\n> In other words:\n>\n> [...]\n>> Similarly, I think it would be very useful if we just make this work:\n>>\n>>     git rev-parse $some_hash^{sha256}^{commit}\n>>\n>> And not care whether $some_hash is SHA-1 or SHA-256, if it's the former\n>> we'd consult the SHA-1 <-> SHA-256 lookup table and go from there, and\n>> always return a useful value.\n>\n> The opposite of this. :)\n\nCan you elaborate on that? What do you think that should do? Return an\nerror if $some_hash is SHA-1, even though we have a $some_hash =\n$some_hash_256 mapping?\n\nI.e. if I'm using this in a script I'd need:\n\n    if x = git rev-parse $some_hash^{sha256}^{commit}\n        hash = x\n    elsif x = git rev-parse $some_hash^{sha1}^{commit}\n        hash = x\n    endif\n\nAs opposed to the thing I'm saying is the redeeming quality of the peel\nsyntax:\n\n    hash = git rev-parse $some_hash^{sha256}^{commit}\n"},{"id":"356877","messageId":"20180829204623.GD7547@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"87y3cpc6bt.fsf@evledraar.gmail.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-29T20:46:23Z","receivedAt":"2018-08-29T20:46:28Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nÆvar Arnfjörð Bjarmason wrote:\n> On Wed, Aug 29 2018, Jonathan Nieder wrote:\n\n>> In other words, I want the input format and output format completely\n>> decoupled.  If I pass ^{sha1}, I am indicating the input format.  To\n>> specify the output format, I'd use --output-format instead.\n>\n> This is also a reasonable thing to want, but I don't see how it can be\n> sensibly squared with the existing peel syntax.\n\nAll the weight here is on the word \"sensibly\".  Currently, ^{thing}\nmeans \"act on the object\" and @{thing} means \"act on the ref\".  This\n^{sha1} syntax is really a new kind of modifier, ~{thing}, meaning\n\"act on the string\".\n\nThat said, we can make it do anything we want.  There is nothing\nforcing us to make it more similar to ^{commit} than to\n^{/searchstring}, say.\n\nIn that context:\n\n[...]\n>> Ævar Arnfjörð Bjarmason wrote:\n\n>>> Similarly, I think it would be very useful if we just make this work:\n>>>\n>>>     git rev-parse $some_hash^{sha256}^{commit}\n>>>\n>>> And not care whether $some_hash is SHA-1 or SHA-256, if it's the former\n>>> we'd consult the SHA-1 <-> SHA-256 lookup table and go from there, and\n>>> always return a useful value.\n>>\n>> The opposite of this. :)\n>\n> Can you elaborate on that?\n\nWhat I'm saying is, regardless of the syntax used, as a user I *need*\na way to look up $some_hash as a sha256-name, with zero risk of Git\ntrying to outsmart me and treating $some_hash as a sha1-name instead.\n\nAny design without that capability is a non-starter.\n\n[...]\n> I.e. if I'm using this in a script I'd need:\n>\n>     if x = git rev-parse $some_hash^{sha256}^{commit}\n>         hash = x\n>     elsif x = git rev-parse $some_hash^{sha1}^{commit}\n>         hash = x\n>     endif\n\nWhy wouldn't you use \"git rev-parse $some_hash^{commit}\" instead?\n\nThanks,\nJonathan\n"},{"id":"356878","messageId":"xmqq36uwc2s8.fsf@gitster-ct.c.googlers.com","threadId":"49202","inReplyTo":"20180829191232.GC7547@aiede.svl.corp.google.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-29T20:53:59Z","receivedAt":"2018-08-29T20:54:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> In other words, I want the input format and output format completely\n> decoupled.\n\nI thought that the original suggestion was to use \"hashname:\" as a\nprefix to specify input format.  In other words\n\n\tsha1:abababab\n\tsha256:abababab\n\nAnd an unadorned abababab is first looked up in sha256 space for\nuniqueness, and if and only if there is only one object whose sha256\nname begins with abababab and there is *no* object whose sha1 name\nbegins with that hexstring (or vice versa), that string will be\nresolved to an object name.\n\nI do not think ^{hashname} mixes well with ^{objecttype} syntax at\nall as an output specifier, either.  It would make sense to be more\nexplicit, I would think, e.g.\n\n\tgit rev-parse --output=sha1 sha256:abababab\n\n(or would that be the job for name-rev?)\n"},{"id":"356882","messageId":"20180829210129.GE7547@aiede.svl.corp.google.com","threadId":"49202","inReplyTo":"xmqq36uwc2s8.fsf@gitster-ct.c.googlers.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2018-08-29T21:01:29Z","receivedAt":"2018-08-29T21:01:35Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>> In other words, I want the input format and output format completely\n>> decoupled.\n>\n> I thought that the original suggestion was to use \"hashname:\" as a\n> prefix to specify input format.  In other words\n>\n> \tsha1:abababab\n> \tsha256:abababab\n\nThat's fine with me too, and it's probably easier to understand than\n^{sha1}.  The disadvantage is that it clashes with existing meaning of\n\"path abababab in branch sha1\".  If we're okay with that change, then\nit's a good syntax.\n\nIf we have a collection of proposed syntaxes, I can get some help from\na UI designer here, too, to help find any ramifications we've missed.\n\n[...]\n> I do not think ^{hashname} mixes well with ^{objecttype} syntax at\n> all as an output specifier, either.  It would make sense to be more\n> explicit, I would think, e.g.\n>\n> \tgit rev-parse --output=sha1 sha256:abababab\n\nAgreed.  I don't think it makes sense to put output specifiers in\nrevision names.  It would create a lot of unnecessary complexity and\nambiguity.\n\nThanks,\nJonathan\n"},{"id":"356917","messageId":"20180829234524.GA15802@sigill.intra.peff.net","threadId":"49202","inReplyTo":"20180829204623.GD7547@aiede.svl.corp.google.com","subject":"Re: How is the ^{sha256} peel syntax supposed to work?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2018-08-29T23:45:24Z","receivedAt":"2018-08-29T23:45:28Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Aug 29, 2018 at 01:46:23PM -0700, Jonathan Nieder wrote:\n\n> > Can you elaborate on that?\n> \n> What I'm saying is, regardless of the syntax used, as a user I *need*\n> a way to look up $some_hash as a sha256-name, with zero risk of Git\n> trying to outsmart me and treating $some_hash as a sha1-name instead.\n> \n> Any design without that capability is a non-starter.\n\nRight, this is IMHO the only thing that makes sense for ^{hash} to do:\nit disambiguates the sha1 that you just gave it. Nothing more, nothing\nless.\n\n> > I.e. if I'm using this in a script I'd need:\n> >\n> >     if x = git rev-parse $some_hash^{sha256}^{commit}\n> >         hash = x\n> >     elsif x = git rev-parse $some_hash^{sha1}^{commit}\n> >         hash = x\n> >     endif\n> \n> Why wouldn't you use \"git rev-parse $some_hash^{commit}\" instead?\n\nYes, the sane rules seem to me to be:\n\n  # try any available hash for $some_hash\n  git rev-parse $some_hash\n\n  # look _only_ for $some_hash as a sha1\n  git rev-parse $some_hash^{sha1}\n\n  # ditto for sha256\n  git rev-parse $some_hash^{sha256}\n\n  # ditto, but then peel the result to a commit\n  git rev-parse $some_hash^{sha256}^{commit}\n\n  # this is nonsense, and should produce an error\n  git rev-parse $some_hash^{commit}^{sha256}\n\nFor convenience of scripts, we may also want:\n\n  git rev-parse --input-hash=sha256 $some_hash\n\nto pretend as if \"^{sha256}\" was appended to each command-line hash we\ntry to resolve (e.g., consider a case where a script is feeding 0 or\nmore hashes).\n\n-Peff\n"}]}