Re: [GSoC PATCH v2 2/4] pack-write: add helper to fill promisor file after repack
- From
Junio C Hamano <gitster@pobox.com>
- Date
- Mar 23, 2026, 21:30 UTC
- Message-ID
- <xmqqfr5q44jx.fsf@gitster.g>
- In-Reply-To
- <0bb031e7443bb53abbbb0afaa347285d6d8cf7b8.1774205661.git.lorenzo.pegorari2002@gmail.com>
LorenzoPegorari <lorenzo.pegorari2002@gmail.com> writes:
Show 13 quoted lines
> Create a `copy_all_promisor_files()` helper function used to copy the > contents of all ".promisor" files in a `repository` inside another > ".promisor" file. > > This function can be used to preserve the contents of all ".promisor" > files inside a new ".promisor" file, for example when a repack happens. > > This function is written in such a way so that it will read all the > ".promisor" files inside the given `repository` line by line, and copy > only the lines that are not already present in the destination file. This > is done to avoid copying the same lines multiple times that may come from > multiple (redundant) packfiles. There might be another better/cleaner way > to achieve this.
In the previous step, we extablished that these "back then their ref X used to point at object Y" records are there so that we can identify which refs were fetched at the time the packfile was downloaded to help debugging. When repacking, losing these records certainly would lose information.
But would concatenating all into a single file help preserve the useful information? Don't we need do better than that?
A NEEDSWORK comment, as was discussed in another thread or two in the recent past, is not necessarily a well thought out fully finished specification of an additional piece of work. "We know this has a problem, we may need to do something about it, like concatenating to save the contents, perhaps? We do not know the answer, and we do not bother thinking it through right at this moment. It is left to the future developers to figure it out" is what a NEEDSWORK comment is about.
Your first response to such a comment may be "yeah, I agree that it is bad to lose information we added to help debugging", but the second one should be to wonder if the "like concatenating..." is the best approach going forward.
In other words, we should take a NEEDSWORK comment as a mere starting point, and what NEEDS your work begins at thinking what needs to be done about the problem raised there.
By mixing them up all into a single list, you no longer can tell when their ref X was observed to be pointing at object Y anymore. You may have two packs originally, with a record for "ref X pointing at object Y" in each of them, but by deduping them, you lose the information that you cloned at one time, and made an additional fetch on another day, and the fact the ref X was pointing at the same value at both times. I am not sure if it is a good implementation if the objective of this topic is to preserve information that is useful for debugging.
I wonder if it helps to append to each line the file timestamp of the .promisor file we took the record originally? For the sake of completeness, we could consider adding the filename as well, but we can quickly dismiss it as not so useful ;-)
If repacking already repacked promisor packfile, the records would already contain such a timestamp at the end, so the code to copy existing records must be prepared to see if the records are the <ref, oid> tuple, or <ref, oid, timestamp> tuple, and act accordingly.
I am *not* saying that without such a "preserve timestamp" column in the record, copying existing records to a new .promisor file is useless. But we do not see any explanation why the author thinks that it is sufficient to copy existing records while silently deduping. We can implement only one choice backed by series of decisions like "timestamp might help" and "original filenames would probably not help", and the design should describe what was considered and rejected (as opposed to "we didn't think things through---we just did what the original NEEDSWORK comment suggested doing").
Thanks.