{"thread":{"id":"65201","subject":"[GSoC Proposal] Implement promisor remote fetch ordering","startedAt":"2026-03-10T18:25:08Z","lastAt":"2026-03-18T16:29:11Z","messageCount":3,"participants":["Lorenzo Pegorari","Christian Couder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"538503","messageId":"abBh__zmlWXY-yjI@lorenzo-VM","threadId":"65201","inReplyTo":null,"subject":"[GSoC Proposal] Implement promisor remote fetch ordering","fromName":"Lorenzo Pegorari","fromEmail":"lorenzo.pegorari2002@gmail.com","sentAt":"2026-03-10T18:25:03Z","receivedAt":"2026-03-10T18:25:08Z","isPatch":false,"sender":{"key":"lorenzo.pegorari2002@gmail.com","avatar":"https://avatars.githubusercontent.com/u/132087553?v=4"},"body":"The following is my proposal for the GSoC'26 for the project \"Implement\npromisor remote fetch ordering\".\n\nAs soon as the the contributor application period begins, I will submit\nthe proposal in PDF format to the official GSoC website.\n\nI have dedicated a large section (about 40%) of the proposal to\nexplaining the current situation and the tests that I have done to gain a\nlot of hands-on experience. I consider this section important, but if it\ntoo long-winded, please let me know.\n\nThank you so much to everyone that is going to spend their time reading\nthis proposal and giving me their feedback.\n\n\n==============================\n\n\"Implement promisor remote fetch ordering\"\n\n==============================\n\n\n# Personal information\n\nName: Lorenzo Pegorari\nPronouns: he/him\nLocation: Cremona, Lombardy, Italy\nTimezone: CET (UTC+1)\nEmail: lorenzo.pegorari2002@gmail.com\nGitHub: https://github.com/LorenzoPegorari\nLinkedIn: https://www.linkedin.com/in/lorenzopegorari/\n\n------------------------------\n\n# Background\n\n## General\n\nHi Git team! My name is Lorenzo Pegorari and I am a 23-year-old student\nfrom Italy.\n\nThroughout my undergraduate studies, I constantly tried to differentiate\nmyself by taking part in as many interesting experiences as possible, to\nbetter define my future professional path. And so, at the end of 2024, I\ndecided to join the FOSS world in order to improve my software\nengineering skills and to contribute to projects that I find meaningful.\n\nMy goal is to join the broader Linux community, mostly focusing on the\nLinux kernel and Git, and to prove to myself that I am capable of\nparticipating in prestigious (and, in the context of the GSoC,\ncompetitive) organizations. My dream is to one day become a cornerstone\nin one of these open-source communities.\n\n## Education\n\nI am currently in the final year of my BSc in Computer Science and\nEngineering at Politecnico di Milano (Milan, Italy).\n\n## Previous Open-Source Experience\n\nI am fairly new to contributing to open-source projects, which is why I\nam applying for the Google Summer of Code in the first place.\n\nMy first step in the open-source world happened on October 16th 2024,\nwhen I released my first project Simply Colorful [1], a free theme for\nthe note-taking application Obsidian that has now gathered close to 10000\ndownloads since its release.\n\nLast year, in the summer of 2025, I also had the honor of participating\nin the Google Summer of Code 2025 with the organization BRL-CAD, where I\nsuccessfully completed my proposed project \"Developing a MOOSE-based\nConsole for arbalest: a first step to merge arbalest and qged\" [2].\n\nThe general goal of the project was to take the initial step in merging\nBRL-CAD's two in-development GUIs: \"arbalest\" and \"qged\". The primary\nobjective was to transfer qged's sophisticated GED (Geometry EDiting)\nconsole, which works via low-level calls to BRL-CAD's core libraries,\ninto arbalest, while preserving the application's distinctive clean and\neasy-to-scale architecture. To support this endeavor, I also expanded\nBRL-CAD's new lightweight, modular, object-oriented API, known as\n\"MOOSE\". In addition to these core tasks, I also fixed some compatibility\nissues related to arbalest's Qt6 widgets to ensure proper display across\ndifferent OSs, and resolved various GUI-related bugs.\n\nMore information regarding my previous GSoC participation can be found in\nmy final project report [3] and in my GSoC'25 daily notes [4]. I am also\nextremely happy to say that my work was much appreciated by the BRL-CAD\ncommunity, with the organization admin Christopher Sean Morrison stating\nthat my final project report was \"outstanding\", and my amazing mentor,\nDr. Daniel Rossberg, noting that my performance was \"awesome\" and that I\nwas \"a pleasure to mentor\".\n\nI would also like to state that, even though I have been quite busy\nlately, I still consider myself a part of the BRL-CAD community, having\ndone some small fixes after the GSoC'25 ended, and now having personally\nhelped some new developers to (hopefully) join BRL-CAD for the GSoC'26.\n\n## Git Experience\n\nI joined the Git community at the beginning of 2026, but I have been\ninterested in this project since 2024. In fact, last year, when I was\ndeciding which organization to join the GSoC'25 with, I seriously\nconsidered Git, but in the end I discarded it because I felt not skilled\nenough to take part in such a complex organization. This time, though, I\nfeel much more confident, and so here I am!\n\nSince the start of this year, I have explored the codebase as much as\npossible, focusing on the GSoC project ideas regarding partial clones,\nwhich were the ones that I personally found most interesting and valuable.\n\nSo far, I have made the following contributions to Git:\n\n * [GSoC PATCH v2] diff: improve scaling of filenames in diffstat to handle UTF-8 chars\n   * Link: https://lore.kernel.org/git/cover.1768520441.git.lorenzo.pegorari2002@gmail.com\n   * Description: The computation of column width made by `git diff --stat`\n                  was confused when pathnames contained non-ASCII chars.\n\t\t  This issue was reported by a `NEEDSWORK` comment.\n   * Status: Merged to `master`\n \n * [GSoC PATCH v3] diff: handle ANSI escape codes in prefix when calculating diffstat width\n   * Link: https://lore.kernel.org/git/cover.1772226209.git.lorenzo.pegorari2002@gmail.com\n   * Description: Fixed `git log --graph --stat` not correctly counting\n                  the display width of colored graph part of its own\n\t\t  output. This issue was reported by a `NEEDSWORK` comment.\n   * Status: Merged to `master`.\n    \n * [GSoC PATCH v3] doc: improve gitprotocol-pack\n   * Link: https://lore.kernel.org/git/cover.1772502209.git.lorenzo.pegorari2002@gmail.com\n   * Description: Improved the `gitprotocol-pack` documentation.\n   * Status: Will merge to `master`.\n\n## Experience With C\n\nC is my primary language. I used it throughout my university courses, for\nmost of my personal projects, and during my GSoC'25 project with BRL-CAD,\nwhere, although my main tasks involved using C++ and Qt6, I had to\nconstantly interface with BRL-CAD's core libraries, which are written in C.\n\n------------------------------\n\n# Current Situation & Testing\n\n## Partial Clones\n\nThe \"partial clone\" feature was introduced to better handle extremely\nlarge repositories, particularly those that contain large binary files.\nThe problem is clearly shown with the following example, which\nillustrates how quickly the size of a Git repository can grow when just\none single 1 MB binary file is frequently committed:\n\n```\n#!/bin/bash\ngit init size_check\n# Create 1MB file of random data, to simulate a compressed binary\nhead -c 1M </dev/urandom >size_check/foo \ngit -C size_check add foo\ngit -C size_check commit -m \"foo\"\ndu -hs size_check/.git/objects  # .git/objects size after 1 foo commit = 1.1 MB\nfor i in {1..50}; do\n    head -c 1M </dev/urandom >size_check/foo  # Change the 1MB file\n    git -C size_check commit -a -m \"foo\"\ndone\ndu -hs size_check/.git/objects  # .git/objects size after 51 foo commits = 53 MB\n```\n\nPartial clones avoid this issue during `clone` and `fetch` operations by\npassing all the objects to download through a `--filter=<filter-spec>`\nspecified by the user, which will limit the number of blobs and trees\nthat actually get downloaded. The `<filter-spec>`, can, for example, be:\n * `blob:none`, which will filter out all blobs.\n * `tree:0`, which will filter out all trees.\n * `blob:limit=5k`, which will filter out all blobs whose size is greater\n   than $5$kB.\n\nThe filtered out objects will be lazily downloaded when the user runs a\ncommand that requires those missing data.\n\nThis mechanism works with the following steps:\n * When the client wants to fetch some objects from the server using a\n   filter, the client, after sending a list of capabilities it wants to\n   be in effect, sends the `filter: <filter-spec>` capability, followed\n   by a request for the objects that the client wants to retrieve. The\n   following is an example of a request (extracted using\n   `GIT_TRACE_PACKET=1`) made by a client to a server to fetch 1 object\n   using the `<filter-spec>=blob:none`:\n\n   ```\n   [...]\n   pkt-line.c:85           packet:        fetch< 0000  # \"flush-pkt\"\n   pkt-line.c:85           packet:        fetch> command=fetch  # Execute fetch\n   pkt-line.c:85           packet:        fetch> agent=git/2.43.0\n   pkt-line.c:85           packet:        fetch> object-format=sha1\n   pkt-line.c:85           packet:        fetch> 0001  # \"delim-pkt\"\n   pkt-line.c:85           packet:        fetch> thin-pack  # Capability\n   pkt-line.c:85           packet:        fetch> no-progress  # Capability\n   pkt-line.c:85           packet:        fetch> ofs-delta  # Capability\n   pkt-line.c:85           packet:        fetch> filter blob:none  # Filter capability\n   # OID of the object the client wants to retrieve\n   pkt-line.c:85           packet:        fetch> want 394ca7a7b5e75a57e736040480f685c8b71844eb  \n   pkt-line.c:85           packet:        fetch> done  # End fetch\n   pkt-line.c:85           packet:        fetch> 0000  # \"flush-pkt\"\n   [...]\n   ```\n\n * The server will apply the requested `<filter-spec>` as it creates the\n   \"promisor packfile\" of the requested objects. A packfile is a binary\n   file that is used to compress many \"loose objects\", and it does so by\n   containing the most recent versions of the stored objects and deltas\n   of the previous versions of those objects. A promisor packfile is a\n   filtered packfile, where the unwanted objects are not present. The\n   promisor packfile is sent to the client.\n\n * When the client receives the promisor packfile, it can operate\n   normally by knowing that the \"promisor objects\" (the filtered-out\n   objects) will be dynamically fetched when needed from \"promisor\n   remotes\" (remotes that have \"promised\" that they have the missing\n   objects). The promisor remotes are defined using the\n   `remote.<name>.promisor` and `remote.<name>.partialCloneFilter`\n   configuration variables.\n\n## Multiple Promisor Remotes & Testing\n\nFocusing on promisor remotes, currently multiple of them can be\nconfigured and used. This feature gives users the flexibility to use a\nspecific promisor remote when convenient (e.g., a remote that is\ncloser/faster for some kind of object). This is particularly useful when\nworking with extremely large repositories ($+100$GB) that contain many\nlarge binary files. These types of repositories can greatly benefit from\nhaving multiple promisor remotes: a common example is setting them up so\nthat one promisor remote can act as a \"Large Object Promisor\" (LOP),\nmeaning a remote that is used only to store large blobs, while the other\none will be the main remote, used to store everything else.\n\nI created a minimal example setup, mostly based on the test\n`t/t5710-promisor-remote-capability` added by `4602676` (\"Add\n'promisor-remote' capability to protocol v2\", 2025-02-18), to experiment\nwith multiple promisor remotes, in order to not simply rely on the\ndocumentation, but to actually get hands-on experience. The example setup\ncreates a `server`, a 'lopm' (\"Large Object Promisor medium\") for blobs\nlarger than 5kB, a `lopl` (\"Large Object Promisor large\") for blobs\nlarger than 50kB, and a `client` that interfaces with all of these\nremotes. It is created in the following way:\n\n * Initially, a very simple Git repository `template` is created, which\n   contains just 3 commits, 3 trees, and 3 blobs of different sizes (as\n   shown by using `git verify-pack -v`:\n\n   ```\n   # \"git -C template verify-pack -v *.pack\" output:\n   <OID-commit-large> commit 217 157 12\n   <OID-commit-medium> commit 218 159 296\n   <OID-commit-small> commit 169 127 169\n   <OID-tree-large> tree   102 105 455\n   <OID-tree-medium> tree   69 76 606\n   <OID-tree-small> tree   35 46 560\n   <OID-blob-large> blob   102400 102444 682  # 100kB blob\n   <OID-blob-medium> blob   10240 10254 103126  # 10kB blob\n   <OID-blob-small> blob   6 15 113380  # 6 bytes blob\n   non delta: 9 objects\n   <*.pack>: ok\n   ```\n\n * The bare `server`, based on the `template`, and the bare and empty\n   `lopm` and `lopl`, are generated using, respectively, `git clone\n   --bare --no-local template server` and `git init --bare lop[m|l]`.\n \n * The objects inside the `server` are unpacked, with all blobs larger\n   than 5kB copied inside `lopm` and all blobs larger than 50kB copied\n   inside `lopl`. The `server` is then repacked using the command `git\n   repack -a -d --filter=blob:limit=5k` (to remove blobs larger than\n   5kB), and finally a \".promisor\" file is created with the same name as\n   the \".pack\" file, to tell Git that all missing objects from the pack\n   can be found in the configured promisor remotes.\n\n * The `server` configuration is modified to support `lopm` and `lopl` as\n   promisor remotes. Also, make it so that inside `server`, `lopm`, and\n   `lopl`, `upload-pack` will support partial clone and partial fetch\n   object filtering (using `uploadpack.allowFilter`), and will accept a\n   fetch request that asks for any object at all (using\n   `uploadpack.allowAnySHA1InWant`):\n\n   ```\n   git -C server remote add lopm \"file://$(pwd)/lopm\"  # Add lopm remote to server\n   git -C server config remote.lopm.promisor true  # Make lopm a promisor remote\n   git -C server remote add lopl \"file://$(pwd)/lopl\"  # Add lopl remote to server\n   git -C server config remote.lopl.promisor true  # Make lopl a promisor remote\n   git -C server config uploadpack.allowFilter true\n   git -C server config uploadpack.allowAnySHA1InWant true\n   git -C lopm config uploadpack.allowFilter true\n   git -C lopm config uploadpack.allowAnySHA1InWant true\n   git -C lopl config uploadpack.allowFilter true\n   git -C lopl config uploadpack.allowAnySHA1InWant true\n   ```\n\n * The `client` is created by doing a partial clone of the `server`, and\n   adding `lopl` and `lopm` as promisor remotes: \n\n   ```\n   GIT_TRACE=$(pwd)/trace \\  # Env var to trace general messages\n   GIT_TRACE_PACKET=$(pwd)/packet \\  # Env var to trace messages for in/out packets \n   GIT_NO_LAZY_FETCH=0 \\  # Env var to enable lazily fetch missing objects on demand\n   git clone \\\n       -c remote.lopl.url=\"file://$(pwd)/lopl\" \\  # Add remote lopl\n       -c remote.lopl.fetch=\"+refs/heads/*:refs/remotes/lopl/*\" \\\n       -c remote.lopl.promisor=true \\  # Make lopl a promisor remote\n       -c remote.lopm.url=\"file://$(pwd)/lopm\" \\  # Add remote lopm\n       -c remote.lopm.fetch=\"+refs/heads/*:refs/remotes/lopm/*\" \\\n       -c remote.lopm.promisor=true \\  # Make lopm a promisor remote\n       --no-local --filter=\"blob:limit=5k\" server client\n   ```\n\nNow, with this setup, by slightly tweaking the configurations of each\nrepository, it is possible to deeply test how multiple promisor remotes\nare handled in various situations, and actually see what is described in\nthe documentation.\n\n## Testing Promisor Remotes Advertisement\n\nAn important thing to test is the promisor remotes advertisement feature.\nThis feature is dependent on 2 main configuration options: the\nserver-side option `promisor.advertise`, which enables the server to\nadvertise the promisor remotes it is using to the client, and the\nclient-side option `promisor.acceptFromServer`, which describes how the\nclient should handle the promisor remotes advertised:\n\n * If `promisor.advertise=false`, when the `client` wants to fetch an\n   object that the `server` does not have, the `server` will not\n   advertise the `promisor-remote` capability, and so it has no other\n   choice than to first fetch the object from `lopl` and/or `lopm`, and\n   then give it to the `client`. This can be checked by doing `git -C\n   server rev-list --objects --all --missing=print`, and seeing that the\n   previously missing large blobs are now present inside the `server`, or\n   by directly looking into the `GIT_TRACE_PACKET` output, and seeing\n   that there is no reference to the `promisor-remote` capability.\n\n * If `promisor.advertise=true`, when the `client` wants to fetch an\n   object that the `server` does not have, the `server` will advertise\n   its promisor remotes, as seen by the `GIT_TRACE_PACKET` output, which\n   will contain:\n    \n   ```\n   [...]\n   packet: upload-pack> promisor-remote= \\\n       name=lopl,url=file://$(pwd)/lopl; \\  # Adv lopl\n       name=lopm,url=file://$(pwd)/lopm  # Adv lopm\n   [...]\n   ```\n\n   The `client` can control what advertised promisor remote to accept with\n   the following options:\n\n    * If `promisor.acceptFromServer=All`, the `client` will accept all\n      advertised promisor remotes. This can be seen by looking at the\n      `GIT_TRACE_PACKET` output, which will contain:\n      \n      ```\n      [...]\n      packet: clone> promisor-remote=lopl;lopm  # Accept lopl and lopm\n      [...]\n      ```\n    \n    * If `promisor.acceptFromServer=KnownName`, the `client` will accept\n      promisor remotes which are already configured and have the same\n      name. This can be seen by changing the `lopl` name in the `client`\n      configuration, and looking at the `GIT_TRACE_PACKET` output, which\n      will contain:\n    \n      ```\n      [...]\n      packet: clone> promisor-remote=lopm  # Accept lopm (no reference to lopl!)\n      [...]\n      ```\n        \n    * If `promisor.acceptFromServer=KnownUrl`, the `client` will accept\n      promisor remotes which are already configured and have the same\n      name and URL. This can be seen by changing the `lopl` URL in the\n      `client` configuration, and looking at the `GIT_TRACE_PACKET`\n      output, which will contain:\n    \n      ```\n      [...]\n      packet: clone> promisor-remote=lopm  # Accept lopm (no reference to lopl!)\n      [...]\n      ```\n\n    * If `promisor.acceptFromServer=None`, the `client` won't accept any\n      advertised promisor remotes.\n\nAdditional pieces of information can be sent by the server when\nadvertising its promisor remotes to the client. These pieces of\ninformation are configured in the server-side configuration variable\n`promisor.sendFields`, and currently can be:\n\n * `partialCloneFilter`, which contains the partial clone filter used for\n   the remote.\n\n * `token`, which contains an authentication token for the remote.\n\nThe client-side configuration variable `promisor.checkFields` can be used\nby the client to check if the values transmitted by a server correspond\nto the values in its own configuration, and accept the promisor remote if\nthey are the same.\n\nA simple test can be done by adding the `remote.lopl.partialCloneFilter`,\n`remote.lopl.token`, and the `promisor.sendFields` variables to the\n`server` configuration. The output of `LOG_TRACE_PACKET` will contain:\n\n```\n[...]\npacket:  upload-pack> promisor-remote=name=lopl, \\  # name field will always be sent\n    url=file://$(pwd)/lopl, \\  # url field will always be sent\n    partialCloneFilter=blob:none, \\  # partialCloneFilter field of lopl\n    token=value; \\  # token field of lopl\n    [...]\n```\n\nRecently, with the patch series \"Implement `promisor.storeFields` and\n`--filter=auto`\" [5], the new client-side configuration variable\n`promisor.storeFields` was added. It contains a list of field names\n`partialCloneFilter` and/or `token`), and the values of these fields,\nwhen transmitted by the server, will be stored in the local configuration\non the client.\n\n## Testing Multiple Promisor Remotes Fetch Order\n\nFinally, the last mechanism that is fundamental to understand is the\nfetch order when multiple promisor remotes are defined:\n\n * When multiple remotes are configured, they are tried one after the\n   other in the order in which they appear in the configuration, until\n   all objects are fetched. This can be easily seen from the output of\n   `GIT_TRACE`, which initially tries to fetch the objects from `lopl`,\n   and then from `lopm`:\n\n   ```\n   [...]\n   trace: built-in: git fetch lopl [...] --filter=blob:none [...]\n   [...]\n   trace: built-in: git fetch lopm [...] --filter=blob:none [...]\n   [...]\n   ```\n\n   While, if we make it so that we first define `lopm` in the `client`\n   configuration, then initially `lopm` will be used to fetch the\n   objects, and `lopl` will not be used at all (because `lopm` contains\n   all required objects:\n\n   ```\n   [...]\n   trace: built-in: git fetch lopm [...] --filter=blob:none [...]\n   [...]\n   ```\n\n * If the configuration option `extensions.partialClone` is present, the\n   promisor remote that it specifies will always be the last one tried\n   when fetching objects.\n    \n------------------------------\n\n# \"Implement promisor remote fetch ordering\"\n\n## Project Goal\n\nThis project aims to improve Git by implementing a fetch ordering\nmechanism for multiple promisor remotes, that can be:\n\n * Configured locally by the client.\n * Advertised by servers through the `promisor-remote` protocol.\n\n## Approach\n\nThe bulk of the project will be the creation of a system that allows to\ndefine the order with which the promisor remotes will be tried when\nfetching an object.\n\nThe first goal will be the creation of a `remote.<name>.promisorPriority`\nconfiguration option, which will hold a number between 1 and 'UCHAR_MAX',\nand which defines the priority of that promisor remote in the fetch\norder. This means that the order in which the promisor are tried will be\nthe following:\n\n * All promisor remotes that have a valid `remote.<name>.promisorPriority`,\n   starting from the one with higher priority (the lower `promisorPriority`\n   value). If 2 or more promisor remotes have the same priority, they will be\n   tried following the order in which they appear in the configuration file.\n\n * All promisor remotes that don't have or have an invalid\n   `remote.<name>.promisorPriority` configuration option. If 2 or more\n   promisor remotes don't define any priority, or have an invalid priority,\n   they will be tried following the order in which they appear in the\n   configuration file.\n\n * The promisor remote defined inside the `extensions.partialClone`, no\n   matter their priority (which will be ignored if present). This is\n   necessary for backward compatibility.\n    \nHaving already taken a look at the code, I have a general idea of th\nmajor steps to take to actually introduce the\n`remote.<name>.promisorPriority` configuration option:\n\n * Modify the `promisor_remote` linked list (inside `promisor-remote.h`)\n   to introduce the new member `promisor_priority}` and the\n   `promisor_remote_config()` function (inside `promisor-remote.c`), to\n   correctly fill the `promisor_priority` of all promisor remotes read\n   from the configuration file.\n\n * Modify the `promisor_remote_get_direct()` function (defined inside\n   `promisor-remote.c`), which fetches all requested objects from all\n   promisor remotes, trying them one at a time until all objects are\n   fetched, to make it follow the previously defined promisor remote order.\n\nWhen the first goal is achieved, the client-side-only fetch ordering\nmechanism for multiple promisor remotes, controllable locally from the\nclient configuration, will be complete.\n\nThe second goal will be the introduction of the new `promisorPriority`\nfield for the `promisor.sendFields`, `promisor.checkFields`, and\n`promisor.storeFields` configuration variables. With this new field, the\nserver will be able to tell the priorities of the promisor remotes that\nit advertises to the client, and the client will be able to either check\nor store these suggested priorities.\n\nMy general plan to implement the `promisorPriority` field is the following:\n\n * Create the `static const char promisor_field_priority` variable inside\n   `promisor-remote.c`, and add this variable inside the `known_fields` array.\n\n * Introduce the new member `priority` to the `promisor_info struct`, a\n   structure for promisor remotes involved in the `promisor-remote`\n   protocol capability, and the new member `store_priority` to the\n   `store_info struct`, a structure used in the \"store fields\" mechanism.\n\n * Create the new `valid_priority()` function, which has to parse the\n   value inside the `promisorPriority` field, and check if it is valid.\n\n * Modify many functions inside of the `promisor-remote.c` file to\n   support the new field. Some of these functions are:\n\n    * `promisor_remote_info()`\n    * `set_one_field()`\n    * `match_field_against_config()`\n    * `all_fields_match()\n    * `parse_one_advertised_remote()`\n    * `store_info_new()`\n    * `promisor_store_advertised_fields()`\n\nWhen the second goal is achieved, the mechanism for servers and clients\nto, respectively, advertise and check/store the promisor remote fetch\norder will be complete.\n\n# Possible Issues\n\nFrom my understanding, the project as it is proposed will handle all\npossible cases, except for one. Let's imagine the following situation:\n\n * `server1` and `server2` both use the promisor remotes `lop1` and `lop2`.\n * `client` has both `server1` and `server2` as remotes.\n\nIn this situation, the `client` has no way to specifically say that when\nfetching from `server1`, it wants to first try `lop1` and then `lop2`, while\nwhen fetching from `server2`, it wants to first try `lop2` and then `lop1`.\n\nOne way to solve this very specific (and maybe unusual) issue is to\nintroduce a way to associate a `promisorPriority` to a specific remote. \n\n## Development Schedule\n\nProject size: large (350 hours).\n\nTimeline:\n\n * May 01 - May 24 (Community Bonding Period):\n    * Discuss with the mentor(s) the best plan to implement the new features.\n    * Get familiar with the Git components that are required to implement the new features.\n * May 25 - June 14\n    * Add the `remote.<name>.promisorPriority` configuration option.\n    * Write tests for the new feature.\n    * Update the documentation.\n * June 15 - June 28\n    * Implement all the suggestions made by the mentor(s)/community.\n    * Refine the patch series\n * June 29 - July 10\n    * Complete all remaining work.\n    * Submit the midterm project report for evaluation.\n * July 11 - August 02\n    * Add support for the `promisorPriority` field.\n    * Write tests for the new feature.\n    * Update the documentation.\n * August 03 - August 16\n    * Implement all the suggestions made by the mentor(s)/community.\n    * Refine the patch series\n * August 17 - August 24\n    * Wrap up everything that is still pending.\n    * Submit the final project report for evaluation.\n\nThis development schedule can be subject to changes/corrections during\nthe \"Community Bonding Period\".\n\n## Time Availability\n\nI plan to spend 5-6 hours a day from Monday to Saturday on this projects,\nso roughly around 30-36 hours a week.\n\nI intend to keep a daily log of what I do, similar to what I have done\nduring the GSoC'25.\n\n------------------------------\n\n# Possible questions\n\n## Am I eligible for the GSoC?\n\nYes. It is possible to participate for a second GSoC term as long as the\ncontributor is still a student.\n\n## Will I use AI?\n\nMostly no. Most studies right now show that the use of LLM-assisted\ncoding, for junior developers, is detrimental in many ways: spending more\ntime on tasks, creating worse code, and learning less during the process.\n\nConsidering that my very first goal as a GSoC contributor is to use this\nexperience to learn as much as possible, I will not use AI for coding.\n\nI will exclusively use AI to check for grammatical and/or syntactical\nerrors in sentences I have written. I will never use AI to generate text,\nbut only to double check it.\n\n## What is my reasoning behind proposing a new feature?\n\nAs clearly stated in the \"General Applicant Information\" in the \"Git\nDeveloper Pages\" [6], contributors suggesting new features should\ncarefully consider the many potential issues that may arise, and see if\nthey can be mitigated before the project is submitted.\n\nMy reasoning behind the proposal of this new feature is the following:\n\n * I think that in my proposal I have shown that I have considered\n   thoroughly all possible cases regarding the introduction of the\n   \"promisor remote fetch ordering\" feature, and so I feel that the\n   necessary discussion to define the details of the project will be\n   very quick.\n\n * I think the proposed new feature is not prone to long naming or user\n   interface discussions.\n\n * I think that the \"promisor remote fetch ordering\" feature is a\n   necessary step to fully support multiple promisor remotes, and to\n   fully support the partial clone mechanism.\n\n * I think the proposed project is not too complex or too difficult for\n   me to handle. In fact, although I was interested, I discarded the\n   \"enhance promisor-remote protocol for better-connected remotes\" project\n   idea, precisely because it seemed like a way too big and complex\n   feature to handle for a GSoC project.\n\n## Why Git?\n\nAs I have said already, I have been interested in contributing to Git\nsince 2024.\n\nThe sheer amount of people all across the globe actively using Git and/or\nengaging with software that was produced also thanks to Git, makes this\nFOSS project, to me, one of the most interesting ones in the world.\n\nBeing responsible for the maintenance and development of a software with\nthis amount of users is extremely challenging, but also really rewarding.\nFurthermore, the developers in this community are some of the best in the\nindustry, and working with them is an amazing opportunity that cannot be\nmissed.\n\nFinally, simply put, joining the broader Linux community is a dream of\nmine, particularly to work on Git and the Linux kernel. In fact, during\nFebruary 2026, I didn't work as much on Git, because I was focused on\napplying for the LFX \"Linux kernel Spring 2026\" mentorship to fix bugs in\nthe Linux kernel.\n\n## Why me?\n\nI hope that it's evident the amount of time and effort that I have put\ninto this proposal. I intend to give my absolute best to make this\nproject a success.\n\nAlso, having already participated last year in the Google Summer of Code,\nI am already very familiar with this online program: I know the dos and\ndon'ts, and will apply what I learned last year during this term.\n\nFinally, I also intend to continue contributing to Git, particularly to\ncontinue to expand and improve the partial clone feature, which I find\nparticularly fascinating.\n\n------------------------------\n\n# Links\n\n[1]: https://github.com/LorenzoPegorari/SimplyColorful\n[2]: https://summerofcode.withgoogle.com/archive/2025/projects/25f08iuM\n[3]: https://lorenzopegorari.github.io/GSoC25-report/\n[4]: https://lorenzopegorari.github.io/GSoC25-report/logs\n[5]: https://lore.kernel.org/git/20260216132317.15894-1-christian.couder@gmail.com/\n[6]: https://git.github.io/General-Application-Information/\n\n\n==============================\n\n\nThanks,\n\nLorenzo\n"},{"id":"538983","messageId":"CAP8UFD1=Ow6NNFKK6y5csmneVaS0J+e5z9pGjFmaVoJ2g1OPFg@mail.gmail.com","threadId":"65201","inReplyTo":"abBh__zmlWXY-yjI@lorenzo-VM","subject":"Re: [GSoC Proposal] Implement promisor remote fetch ordering","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2026-03-14T17:30:57Z","receivedAt":"2026-03-14T17:31:09Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Tue, Mar 10, 2026 at 7:25 PM Lorenzo Pegorari\n<lorenzo.pegorari2002@gmail.com> wrote:\n>\n> The following is my proposal for the GSoC'26 for the project \"Implement\n> promisor remote fetch ordering\".\n\nThank you for your interest in Git and this project.\n\n> As soon as the the contributor application period begins, I will submit\n> the proposal in PDF format to the official GSoC website.\n\nGood idea.\n\n> I have dedicated a large section (about 40%) of the proposal to\n> explaining the current situation and the tests that I have done to gain a\n> lot of hands-on experience. I consider this section important, but if it\n> too long-winded, please let me know.\n\n[...]\n\n> So far, I have made the following contributions to Git:\n>\n>  * [GSoC PATCH v2] diff: improve scaling of filenames in diffstat to handle UTF-8 chars\n>    * Link: https://lore.kernel.org/git/cover.1768520441.git.lorenzo.pegorari2002@gmail.com\n>    * Description: The computation of column width made by `git diff --stat`\n>                   was confused when pathnames contained non-ASCII chars.\n>                   This issue was reported by a `NEEDSWORK` comment.\n>    * Status: Merged to `master`\n>\n>  * [GSoC PATCH v3] diff: handle ANSI escape codes in prefix when calculating diffstat width\n>    * Link: https://lore.kernel.org/git/cover.1772226209.git.lorenzo.pegorari2002@gmail.com\n>    * Description: Fixed `git log --graph --stat` not correctly counting\n>                   the display width of colored graph part of its own\n>                   output. This issue was reported by a `NEEDSWORK` comment.\n>    * Status: Merged to `master`.\n\nFor the patches that are merged to master, it could help if you could\ngive the object ID of the merge commit that merged your commits into\nmaster, or alternatively the object ID of all your commits.\n\n>  * [GSoC PATCH v3] doc: improve gitprotocol-pack\n>    * Link: https://lore.kernel.org/git/cover.1772502209.git.lorenzo.pegorari2002@gmail.com\n>    * Description: Improved the `gitprotocol-pack` documentation.\n>    * Status: Will merge to `master`.\n\nYeah, this has been merged to master after your email.\n\n[...]\n\n> Partial clones avoid this issue during `clone` and `fetch` operations by\n> passing all the objects to download through a `--filter=<filter-spec>`\n> specified by the user, which will limit the number of blobs and trees\n> that actually get downloaded. The `<filter-spec>`, can, for example, be:\n>  * `blob:none`, which will filter out all blobs.\n>  * `tree:0`, which will filter out all trees.\n>  * `blob:limit=5k`, which will filter out all blobs whose size is greater\n>    than $5$kB.\n\nWhy are there '$' signs above?\n\n> The filtered out objects will be lazily downloaded when the user runs a\n> command that requires those missing data.\n>\n> This mechanism works with the following steps:\n>  * When the client wants to fetch some objects from the server using a\n>    filter, the client, after sending a list of capabilities it wants to\n>    be in effect, sends the `filter: <filter-spec>` capability, followed\n>    by a request for the objects that the client wants to retrieve. The\n>    following is an example of a request (extracted using\n>    `GIT_TRACE_PACKET=1`) made by a client to a server to fetch 1 object\n>    using the `<filter-spec>=blob:none`:\n>\n>    ```\n>    [...]\n>    pkt-line.c:85           packet:        fetch< 0000  # \"flush-pkt\"\n>    pkt-line.c:85           packet:        fetch> command=fetch  # Execute fetch\n>    pkt-line.c:85           packet:        fetch> agent=git/2.43.0\n>    pkt-line.c:85           packet:        fetch> object-format=sha1\n>    pkt-line.c:85           packet:        fetch> 0001  # \"delim-pkt\"\n>    pkt-line.c:85           packet:        fetch> thin-pack  # Capability\n>    pkt-line.c:85           packet:        fetch> no-progress  # Capability\n>    pkt-line.c:85           packet:        fetch> ofs-delta  # Capability\n>    pkt-line.c:85           packet:        fetch> filter blob:none  # Filter capability\n>    # OID of the object the client wants to retrieve\n>    pkt-line.c:85           packet:        fetch> want 394ca7a7b5e75a57e736040480f685c8b71844eb\n>    pkt-line.c:85           packet:        fetch> done  # End fetch\n>    pkt-line.c:85           packet:        fetch> 0000  # \"flush-pkt\"\n>    [...]\n>    ```\n\nI think when lazy fetching like this, the filter is always blob:none.\nIt's not really used anyway because the objects that the client wants\nare specified explicitly.\n\nThe filter is important when initially cloning or fetching from the\nserver to specify which objects are initially excluded, even if some\nof these  objects will be lazy fetched soon. For example the checkout\npart of a clone might need objects that were initially excluded, so it\nmight lazy fetch some.\n\n>  * The server will apply the requested `<filter-spec>` as it creates the\n>    \"promisor packfile\" of the requested objects.\n\nThis is important during an initial clone or fetch, not when lazy fetching.\n\n> A packfile is a binary\n>    file that is used to compress many \"loose objects\", and it does so by\n>    containing the most recent versions of the stored objects and deltas\n>    of the previous versions of those objects. A promisor packfile is a\n>    filtered packfile, where the unwanted objects are not present. The\n>    promisor packfile is sent to the client.\n\n\n> I created a minimal example setup, mostly based on the test\n> `t/t5710-promisor-remote-capability` added by `4602676` (\"Add\n> 'promisor-remote' capability to protocol v2\", 2025-02-18), to experiment\n> with multiple promisor remotes, in order to not simply rely on the\n> documentation, but to actually get hands-on experience. The example setup\n> creates a `server`, a 'lopm' (\"Large Object Promisor medium\") for blobs\n> larger than 5kB, a `lopl` (\"Large Object Promisor large\") for blobs\n> larger than 50kB, and a `client` that interfaces with all of these\n> remotes. It is created in the following way:\n\n[...]\n\n> Now, with this setup, by slightly tweaking the configurations of each\n> repository, it is possible to deeply test how multiple promisor remotes\n> are handled in various situations, and actually see what is described in\n> the documentation.\n\nYeah, it's quite complex to set up.\n\n> ## Testing Promisor Remotes Advertisement\n>\n> An important thing to test is the promisor remotes advertisement feature.\n> This feature is dependent on 2 main configuration options: the\n> server-side option `promisor.advertise`, which enables the server to\n> advertise the promisor remotes it is using to the client, and the\n> client-side option `promisor.acceptFromServer`, which describes how the\n> client should handle the promisor remotes advertised:\n>\n>  * If `promisor.advertise=false`, when the `client` wants to fetch an\n>    object that the `server` does not have,\n\nI don't think it depends on the client fetching an object the server\ndoes not have. It depends on the client using a filter because the\npromisor-remote capability only makes sense in the case of partial\nclones (or fetches).\n\n> the `server` will not\n>    advertise the `promisor-remote` capability, and so it has no other\n>    choice than to first fetch the object from `lopl` and/or `lopm`, and\n>    then give it to the `client`. This can be checked by doing `git -C\n>    server rev-list --objects --all --missing=print`, and seeing that the\n>    previously missing large blobs are now present inside the `server`, or\n>    by directly looking into the `GIT_TRACE_PACKET` output, and seeing\n>    that there is no reference to the `promisor-remote` capability.\n>\n>  * If `promisor.advertise=true`, when the `client` wants to fetch an\n>    object that the `server` does not have,\n\nSame as above, it doesn't depend on the client fetching an object the\nserver does not have. It depends on the client using a filter because\nthe promisor-remote capability only makes sense in the case of partial\nclones (or fetches).\n\n> the `server` will advertise\n>    its promisor remotes, as seen by the `GIT_TRACE_PACKET` output, which\n>    will contain:\n>\n>    ```\n>    [...]\n>    packet: upload-pack> promisor-remote= \\\n>        name=lopl,url=file://$(pwd)/lopl; \\  # Adv lopl\n>        name=lopm,url=file://$(pwd)/lopm  # Adv lopm\n>    [...]\n>    ```\n\n[...]\n\n> Recently, with the patch series \"Implement `promisor.storeFields` and\n> `--filter=auto`\" [5], the new client-side configuration variable\n> `promisor.storeFields` was added. It contains a list of field names\n> `partialCloneFilter` and/or `token`), and the values of these fields,\n> when transmitted by the server, will be stored in the local configuration\n> on the client.\n>\n> ## Testing Multiple Promisor Remotes Fetch Order\n\nYeah, I think this is the most relevant for the project.\n\n> Finally, the last mechanism that is fundamental to understand is the\n> fetch order when multiple promisor remotes are defined:\n>\n>  * When multiple remotes are configured, they are tried one after the\n>    other in the order in which they appear in the configuration, until\n>    all objects are fetched.\n\nRight, but there is the exception of a remote configured with\n`extensions.partialClone` that will be tried last. You mention it\nlater though.\n\n> This can be easily seen from the output of\n>    `GIT_TRACE`, which initially tries to fetch the objects from `lopl`,\n>    and then from `lopm`:\n>\n>    ```\n>    [...]\n>    trace: built-in: git fetch lopl [...] --filter=blob:none [...]\n>    [...]\n>    trace: built-in: git fetch lopm [...] --filter=blob:none [...]\n>    [...]\n>    ```\n>\n>    While, if we make it so that we first define `lopm` in the `client`\n>    configuration, then initially `lopm` will be used to fetch the\n>    objects, and `lopl` will not be used at all (because `lopm` contains\n>    all required objects:\n>\n>    ```\n>    [...]\n>    trace: built-in: git fetch lopm [...] --filter=blob:none [...]\n>    [...]\n>    ```\n\nYeah, when all the needed objects have been lazy fetched, there is no\npoint in further fetching from any remote.\n\n>  * If the configuration option `extensions.partialClone` is present, the\n>    promisor remote that it specifies will always be the last one tried\n>    when fetching objects.\n>\n> ------------------------------\n>\n> # \"Implement promisor remote fetch ordering\"\n>\n> ## Project Goal\n>\n> This project aims to improve Git by implementing a fetch ordering\n> mechanism for multiple promisor remotes, that can be:\n>\n>  * Configured locally by the client.\n>  * Advertised by servers through the `promisor-remote` protocol.\n>\n> ## Approach\n>\n> The bulk of the project will be the creation of a system that allows to\n> define the order with which the promisor remotes will be tried when\n> fetching an object.\n>\n> The first goal will be the creation of a `remote.<name>.promisorPriority`\n\nYeah, or just `remote.<name>.priority`. The name is to be discussed.\n\n> configuration option, which will hold a number between 1 and 'UCHAR_MAX',\n\nUCHAR_MAX could be system dependent. It might be better to have\nconfigurations work in the same way on all machines though. So perhaps\na fixed range like 1 to 100 would be better. Or are there other ranges\nof values used for similar things in Git or other well known software\nthat could be reused?\n\n> and which defines the priority of that promisor remote in the fetch\n> order. This means that the order in which the promisor are tried will be\n> the following:\n>\n>  * All promisor remotes that have a valid `remote.<name>.promisorPriority`,\n>    starting from the one with higher priority (the lower `promisorPriority`\n>    value). If 2 or more promisor remotes have the same priority, they will be\n>    tried following the order in which they appear in the configuration file.\n>\n>  * All promisor remotes that don't have or have an invalid\n>    `remote.<name>.promisorPriority` configuration option. If 2 or more\n>    promisor remotes don't define any priority, or have an invalid priority,\n>    they will be tried following the order in which they appear in the\n>    configuration file.\n>\n>  * The promisor remote defined inside the `extensions.partialClone`, no\n>    matter their priority (which will be ignored if present). This is\n>    necessary for backward compatibility.\n\nYeah, I think something like what you describe makes sense.\n\n> Having already taken a look at the code, I have a general idea of th\n\ns/of th/of the/\n\n> major steps to take to actually introduce the\n> `remote.<name>.promisorPriority` configuration option:\n\n[...]\n\n> # Possible Issues\n>\n> From my understanding, the project as it is proposed will handle all\n> possible cases, except for one. Let's imagine the following situation:\n>\n>  * `server1` and `server2` both use the promisor remotes `lop1` and `lop2`.\n>  * `client` has both `server1` and `server2` as remotes.\n>\n> In this situation, the `client` has no way to specifically say that when\n> fetching from `server1`, it wants to first try `lop1` and then `lop2`, while\n> when fetching from `server2`, it wants to first try `lop2` and then `lop1`.\n\nRight, but lazy fetching does not only happen as part of a clone or\nfetch from a server. It happens when for some reason (like a git show\nor a git blame for example) the user needs some objects it doesn't\nhave locally, and when that happens, this is not related to a single\nserver.\n\nSo global priorities are likely the most useful ones to have.\n\n> One way to solve this very specific (and maybe unusual) issue is to\n> introduce a way to associate a `promisorPriority` to a specific remote.\n\nYeah, but I don't think it would be used a lot. We can perhaps think\nof some cases where it could be useful, but in practice it is likely\nthat if there is an optimal order for one server, it will be optimal\nfor all other servers too.\n\n[...]\n\nThanks!\n"},{"id":"539295","messageId":"abrS0q_Oc3kn_T3Y@lorenzo-VM","threadId":"65201","inReplyTo":"CAP8UFD1=Ow6NNFKK6y5csmneVaS0J+e5z9pGjFmaVoJ2g1OPFg@mail.gmail.com","subject":"Re: [GSoC Proposal] Implement promisor remote fetch ordering","fromName":"Lorenzo Pegorari","fromEmail":"lorenzo.pegorari2002@gmail.com","sentAt":"2026-03-18T16:29:06Z","receivedAt":"2026-03-18T16:29:11Z","isPatch":false,"sender":{"key":"lorenzo.pegorari2002@gmail.com","avatar":"https://avatars.githubusercontent.com/u/132087553?v=4"},"body":"On Sat, Mar 14, 2026 at 06:30:57PM +0100, Christian Couder wrote:\n> On Tue, Mar 10, 2026 at 7:25 PM Lorenzo Pegorari\n> <lorenzo.pegorari2002@gmail.com> wrote:\n> >\n> > The following is my proposal for the GSoC'26 for the project \"Implement\n> > promisor remote fetch ordering\".\n> \n> Thank you for your interest in Git and this project.\n\nThank you for reading and giving me feedback on my proposal!\n\n> > As soon as the the contributor application period begins, I will submit\n> > the proposal in PDF format to the official GSoC website.\n> \n> Good idea.\n\nI will send v2 and upload it pretty soon.\n\n> For the patches that are merged to master, it could help if you could\n> give the object ID of the merge commit that merged your commits into\n> master, or alternatively the object ID of all your commits.\n\nAck.\n\n> >  * [GSoC PATCH v3] doc: improve gitprotocol-pack\n> >    * Link: https://lore.kernel.org/git/cover.1772502209.git.lorenzo.pegorari2002@gmail.com\n> >    * Description: Improved the `gitprotocol-pack` documentation.\n> >    * Status: Will merge to `master`.\n> \n> Yeah, this has been merged to master after your email.\n\nAck.\n\n> > Partial clones avoid this issue during `clone` and `fetch` operations by\n> > passing all the objects to download through a `--filter=<filter-spec>`\n> > specified by the user, which will limit the number of blobs and trees\n> > that actually get downloaded. The `<filter-spec>`, can, for example, be:\n> >  * `blob:none`, which will filter out all blobs.\n> >  * `tree:0`, which will filter out all trees.\n> >  * `blob:limit=5k`, which will filter out all blobs whose size is greater\n> >    than $5$kB.\n> \n> Why are there '$' signs above?\n\nOps. I wrote the proposal on Markdown with LaTeX support. Text between\n\"$\" is considered LaTeX. Forgot to delete it when sending the email. My\nfault.\n\n> > The filtered out objects will be lazily downloaded when the user runs a\n> > command that requires those missing data.\n> >\n> > This mechanism works with the following steps:\n> >  * When the client wants to fetch some objects from the server using a\n> >    filter, the client, after sending a list of capabilities it wants to\n> >    be in effect, sends the `filter: <filter-spec>` capability, followed\n> >    by a request for the objects that the client wants to retrieve. The\n> >    following is an example of a request (extracted using\n> >    `GIT_TRACE_PACKET=1`) made by a client to a server to fetch 1 object\n> >    using the `<filter-spec>=blob:none`:\n> >\n> >    ```\n> >    [...]\n> >    pkt-line.c:85           packet:        fetch< 0000  # \"flush-pkt\"\n> >    pkt-line.c:85           packet:        fetch> command=fetch  # Execute fetch\n> >    pkt-line.c:85           packet:        fetch> agent=git/2.43.0\n> >    pkt-line.c:85           packet:        fetch> object-format=sha1\n> >    pkt-line.c:85           packet:        fetch> 0001  # \"delim-pkt\"\n> >    pkt-line.c:85           packet:        fetch> thin-pack  # Capability\n> >    pkt-line.c:85           packet:        fetch> no-progress  # Capability\n> >    pkt-line.c:85           packet:        fetch> ofs-delta  # Capability\n> >    pkt-line.c:85           packet:        fetch> filter blob:none  # Filter capability\n> >    # OID of the object the client wants to retrieve\n> >    pkt-line.c:85           packet:        fetch> want 394ca7a7b5e75a57e736040480f685c8b71844eb\n> >    pkt-line.c:85           packet:        fetch> done  # End fetch\n> >    pkt-line.c:85           packet:        fetch> 0000  # \"flush-pkt\"\n> >    [...]\n> >    ```\n> \n> I think when lazy fetching like this, the filter is always blob:none.\n> It's not really used anyway because the objects that the client wants\n> are specified explicitly.\n\nOh, I didn't know that. Makes sense.\n\n> The filter is important when initially cloning or fetching from the\n> server to specify which objects are initially excluded, even if some\n> of these  objects will be lazy fetched soon. For example the checkout\n> part of a clone might need objects that were initially excluded, so it\n> might lazy fetch some.\n\nOoh ok, with this comment I actually fully understand now. Looking back\nat the `GIT_TRACE_PACKET` output, I actually understand almost all of\nit. So the partial clone fetches (usually) the `HEAD`, excluding the\nfiltered out objects, while the lazy fetching directly asks for the\nmissing objects when they are needed, so the filter is not used. Got it!\n\n> >  * The server will apply the requested `<filter-spec>` as it creates the\n> >    \"promisor packfile\" of the requested objects.\n> \n> This is important during an initial clone or fetch, not when lazy fetching.\n\nGot it. I will revisit all the instances where I made some confusion\nbetween lazy fetching and initial cloning/fetching. Thank you so much\nfor your explaination Christian!\n\n> > A packfile is a binary\n> >    file that is used to compress many \"loose objects\", and it does so by\n> >    containing the most recent versions of the stored objects and deltas\n> >    of the previous versions of those objects. A promisor packfile is a\n> >    filtered packfile, where the unwanted objects are not present. The\n> >    promisor packfile is sent to the client.\n> \n> \n> > I created a minimal example setup, mostly based on the test\n> > `t/t5710-promisor-remote-capability` added by `4602676` (\"Add\n> > 'promisor-remote' capability to protocol v2\", 2025-02-18), to experiment\n> > with multiple promisor remotes, in order to not simply rely on the\n> > documentation, but to actually get hands-on experience. The example setup\n> > creates a `server`, a 'lopm' (\"Large Object Promisor medium\") for blobs\n> > larger than 5kB, a `lopl` (\"Large Object Promisor large\") for blobs\n> > larger than 50kB, and a `client` that interfaces with all of these\n> > remotes. It is created in the following way:\n> \n> [...]\n> \n> > Now, with this setup, by slightly tweaking the configurations of each\n> > repository, it is possible to deeply test how multiple promisor remotes\n> > are handled in various situations, and actually see what is described in\n> > the documentation.\n> \n> Yeah, it's quite complex to set up.\n\nYep. The complexity of the tests are the reason behind my decision to\ndeeply describe them in the proposal.\n\n> > ## Testing Promisor Remotes Advertisement\n> >\n> > An important thing to test is the promisor remotes advertisement feature.\n> > This feature is dependent on 2 main configuration options: the\n> > server-side option `promisor.advertise`, which enables the server to\n> > advertise the promisor remotes it is using to the client, and the\n> > client-side option `promisor.acceptFromServer`, which describes how the\n> > client should handle the promisor remotes advertised:\n> >\n> >  * If `promisor.advertise=false`, when the `client` wants to fetch an\n> >    object that the `server` does not have,\n> \n> I don't think it depends on the client fetching an object the server\n> does not have. It depends on the client using a filter because the\n> promisor-remote capability only makes sense in the case of partial\n> clones (or fetches).\n\nOk yeah, I should have explained this better. Of course this depends on\nthe client using a filter. Thanks for the feedback.\n\n> > the `server` will not\n> >    advertise the `promisor-remote` capability, and so it has no other\n> >    choice than to first fetch the object from `lopl` and/or `lopm`, and\n> >    then give it to the `client`. This can be checked by doing `git -C\n> >    server rev-list --objects --all --missing=print`, and seeing that the\n> >    previously missing large blobs are now present inside the `server`, or\n> >    by directly looking into the `GIT_TRACE_PACKET` output, and seeing\n> >    that there is no reference to the `promisor-remote` capability.\n> >\n> >  * If `promisor.advertise=true`, when the `client` wants to fetch an\n> >    object that the `server` does not have,\n> \n> Same as above, it doesn't depend on the client fetching an object the\n> server does not have. It depends on the client using a filter because\n> the promisor-remote capability only makes sense in the case of partial\n> clones (or fetches).\n\nAck. Same as above.\n\n> > the `server` will advertise\n> >    its promisor remotes, as seen by the `GIT_TRACE_PACKET` output, which\n> >    will contain:\n> >\n> >    ```\n> >    [...]\n> >    packet: upload-pack> promisor-remote= \\\n> >        name=lopl,url=file://$(pwd)/lopl; \\  # Adv lopl\n> >        name=lopm,url=file://$(pwd)/lopm  # Adv lopm\n> >    [...]\n> >    ```\n> \n> [...]\n> \n> > Recently, with the patch series \"Implement `promisor.storeFields` and\n> > `--filter=auto`\" [5], the new client-side configuration variable\n> > `promisor.storeFields` was added. It contains a list of field names\n> > `partialCloneFilter` and/or `token`), and the values of these fields,\n> > when transmitted by the server, will be stored in the local configuration\n> > on the client.\n> >\n> > ## Testing Multiple Promisor Remotes Fetch Order\n> \n> Yeah, I think this is the most relevant for the project.\n\nAgreed.\n\n> > Finally, the last mechanism that is fundamental to understand is the\n> > fetch order when multiple promisor remotes are defined:\n> >\n> >  * When multiple remotes are configured, they are tried one after the\n> >    other in the order in which they appear in the configuration, until\n> >    all objects are fetched.\n> \n> Right, but there is the exception of a remote configured with\n> `extensions.partialClone` that will be tried last. You mention it\n> later though.\n\nYep, will mention it also here.\n\n> > This can be easily seen from the output of\n> >    `GIT_TRACE`, which initially tries to fetch the objects from `lopl`,\n> >    and then from `lopm`:\n> >\n> >    ```\n> >    [...]\n> >    trace: built-in: git fetch lopl [...] --filter=blob:none [...]\n> >    [...]\n> >    trace: built-in: git fetch lopm [...] --filter=blob:none [...]\n> >    [...]\n> >    ```\n> >\n> >    While, if we make it so that we first define `lopm` in the `client`\n> >    configuration, then initially `lopm` will be used to fetch the\n> >    objects, and `lopl` will not be used at all (because `lopm` contains\n> >    all required objects:\n> >\n> >    ```\n> >    [...]\n> >    trace: built-in: git fetch lopm [...] --filter=blob:none [...]\n> >    [...]\n> >    ```\n> \n> Yeah, when all the needed objects have been lazy fetched, there is no\n> point in further fetching from any remote.\n\nYeah, and so `lopl` is not tried at all.\n\n> >  * If the configuration option `extensions.partialClone` is present, the\n> >    promisor remote that it specifies will always be the last one tried\n> >    when fetching objects.\n> >\n> > ------------------------------\n> >\n> > # \"Implement promisor remote fetch ordering\"\n> >\n> > ## Project Goal\n> >\n> > This project aims to improve Git by implementing a fetch ordering\n> > mechanism for multiple promisor remotes, that can be:\n> >\n> >  * Configured locally by the client.\n> >  * Advertised by servers through the `promisor-remote` protocol.\n> >\n> > ## Approach\n> >\n> > The bulk of the project will be the creation of a system that allows to\n> > define the order with which the promisor remotes will be tried when\n> > fetching an object.\n> >\n> > The first goal will be the creation of a `remote.<name>.promisorPriority`\n> \n> Yeah, or just `remote.<name>.priority`. The name is to be discussed.\n\nAck.\n\n> > configuration option, which will hold a number between 1 and 'UCHAR_MAX',\n> \n> UCHAR_MAX could be system dependent. It might be better to have\n> configurations work in the same way on all machines though. So perhaps\n> a fixed range like 1 to 100 would be better. Or are there other ranges\n> of values used for similar things in Git or other well known software\n> that could be reused?\n\nMmh true. A fixed range might be better, I agree.\n\n> > and which defines the priority of that promisor remote in the fetch\n> > order. This means that the order in which the promisor are tried will be\n> > the following:\n> >\n> >  * All promisor remotes that have a valid `remote.<name>.promisorPriority`,\n> >    starting from the one with higher priority (the lower `promisorPriority`\n> >    value). If 2 or more promisor remotes have the same priority, they will be\n> >    tried following the order in which they appear in the configuration file.\n> >\n> >  * All promisor remotes that don't have or have an invalid\n> >    `remote.<name>.promisorPriority` configuration option. If 2 or more\n> >    promisor remotes don't define any priority, or have an invalid priority,\n> >    they will be tried following the order in which they appear in the\n> >    configuration file.\n> >\n> >  * The promisor remote defined inside the `extensions.partialClone`, no\n> >    matter their priority (which will be ignored if present). This is\n> >    necessary for backward compatibility.\n> \n> Yeah, I think something like what you describe makes sense.\n\nNice! :-)\n\n> > Having already taken a look at the code, I have a general idea of th\n> \n> s/of th/of the/\n\nAck.\n\n> > major steps to take to actually introduce the\n> > `remote.<name>.promisorPriority` configuration option:\n> \n> [...]\n> \n> > # Possible Issues\n> >\n> > From my understanding, the project as it is proposed will handle all\n> > possible cases, except for one. Let's imagine the following situation:\n> >\n> >  * `server1` and `server2` both use the promisor remotes `lop1` and `lop2`.\n> >  * `client` has both `server1` and `server2` as remotes.\n> >\n> > In this situation, the `client` has no way to specifically say that when\n> > fetching from `server1`, it wants to first try `lop1` and then `lop2`, while\n> > when fetching from `server2`, it wants to first try `lop2` and then `lop1`.\n> \n> Right, but lazy fetching does not only happen as part of a clone or\n> fetch from a server. It happens when for some reason (like a git show\n> or a git blame for example) the user needs some objects it doesn't\n> have locally, and when that happens, this is not related to a single\n> server.\n> \n> So global priorities are likely the most useful ones to have.\n> \n> > One way to solve this very specific (and maybe unusual) issue is to\n> > introduce a way to associate a `promisorPriority` to a specific remote.\n> \n> Yeah, but I don't think it would be used a lot. We can perhaps think\n> of some cases where it could be useful, but in practice it is likely\n> that if there is an optimal order for one server, it will be optimal\n> for all other servers too.\n\nI agree. I should have pointed out clearly that, to me, this unusual\nsituation doesn't seem worth the effort.\n\n> [...]\n> \n> Thanks!\n\nThank you Christian!\n"}]}