git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC PATCH 7/7] update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flags

From
NSNeeraj Singh <nksingh85@gmail.com>
Date
Mar 23, 2022, 20:19 UTC
Message-ID
<CANQDOddpo+a8r_0yghgy_1bHvfUe5XQaaaWc7D-OLqX6Anhgiw@mail.gmail.com>
In-Reply-To
<220323.86sfr9ndpr.gmgdl@evledraar.gmail.com>

I'm going to respond in more detail to your individual patches, (expect the last mail to contain a comment at the end "LAST MAIL").

On Wed, Mar 23, 2022 at 3:52 AM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:

Show 120 quoted lines
>
>
> On Tue, Mar 22 2022, Neeraj Singh wrote:
>
> > On Tue, Mar 22, 2022 at 8:48 PM Ævar Arnfjörð Bjarmason
> > <avarab@gmail.com> wrote:
> >>
> >> As with unpack-objects in a preceding commit have update-index.c make
> >> use of the HASH_N_OBJECTS{,_{FIRST,LAST}} flags. We now have a "batch"
> >> mode again for "update-index".
> >>
> >> Adding the t/* directory from git.git on a Linux ramdisk is a bit
> >> faster than with the tmp-objdir indirection:
> >>
> >>         git hyperfine -L rev ns/batched-fsync,HEAD -s 'make CFLAGS=-O3' -p 'rm -rf repo && git init repo && cp -R t repo/' 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' --warmup 1 -r 10
> >>         Benchmark 1: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync
> >>           Time (mean ± σ):     289.8 ms ±   4.0 ms    [User: 186.3 ms, System: 103.2 ms]
> >>           Range (min … max):   285.6 ms … 297.0 ms    10 runs
> >>
> >>         Benchmark 2: git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD
> >>           Time (mean ± σ):     273.9 ms ±   7.3 ms    [User: 189.3 ms, System: 84.1 ms]
> >>           Range (min … max):   267.8 ms … 291.3 ms    10 runs
> >>
> >>         Summary
> >>           'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran
> >>             1.06 ± 0.03 times faster than 'git ls-files -- t | ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'
> >>
> >> And as before running that with "strace --summary-only" slows things
> >> down a bit (probably mimicking slower I/O a bit). I then get:
> >>
> >>         Summary
> >>           'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'HEAD' ran
> >>             1.21 ± 0.02 times faster than 'git ls-files -- t | strace --summary-only ./git -c core.fsync=loose-object -c core.fsyncMethod=batch -C repo update-index --add --stdin' in 'ns/batched-fsync'
> >>
> >> We also go from ~51k syscalls to ~39k, with ~2x the number of link()
> >> and unlink() in ns/batched-fsync.
> >>
> >> In the process of doing this conversion we lost the "bulk" mode for
> >> files added on the command-line. I don't think it's useful to optimize
> >> that, but we could if anyone cared.
> >>
> >> We've also converted this to a string_list, we could walk with
> >> getline_fn() and get one line "ahead" to see what we have left, but I
> >> found that state machine a bit painful, and at least in my testing
> >> buffering this doesn't harm things. But we could also change this to
> >> stream again, at the cost of some getline_fn() twiddling.
> >>
> >> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>
> >> ---
> >>  builtin/update-index.c | 31 +++++++++++++++++++++++++++----
> >>  1 file changed, 27 insertions(+), 4 deletions(-)
> >>
> >> diff --git a/builtin/update-index.c b/builtin/update-index.c
> >> index af02ff39756..c7cbfe1123b 100644
> >> --- a/builtin/update-index.c
> >> +++ b/builtin/update-index.c
> >> @@ -1194,15 +1194,38 @@ int cmd_update_index(int argc, const char **argv, const char *prefix)
> >>         }
> >>
> >>         if (read_from_stdin) {
> >> +               struct string_list list = STRING_LIST_INIT_NODUP;
> >>                 struct strbuf line = STRBUF_INIT;
> >>                 struct strbuf unquoted = STRBUF_INIT;
> >> +               size_t i, nr;
> >> +               unsigned oflags;
> >>
> >>                 setup_work_tree();
> >> -               while (getline_fn(&line, stdin) != EOF)
> >> -                       line_from_stdin(&line, &unquoted, prefix, prefix_length,
> >> -                                       nul_term_line, set_executable_bit, 0);
> >> +               while (getline_fn(&line, stdin) != EOF) {
> >> +                       size_t len = line.len;
> >> +                       char *str = strbuf_detach(&line, NULL);
> >> +
> >> +                       string_list_append_nodup(&list, str)->util = (void *)len;
> >> +               }
> >> +
> >> +               nr = list.nr;
> >> +               oflags = nr > 1 ? HASH_N_OBJECTS : 0;
> >> +               for (i = 0; i < nr; i++) {
> >> +                       size_t nth = i + 1;
> >> +                       unsigned f = i == 0 ? HASH_N_OBJECTS_FIRST :
> >> +                                 nr == nth ? HASH_N_OBJECTS_LAST : 0;
> >> +                       struct strbuf buf = STRBUF_INIT;
> >> +                       struct string_list_item *item = list.items + i;
> >> +                       const size_t len = (size_t)item->util;
> >> +
> >> +                       strbuf_attach(&buf, item->string, len, len);
> >> +                       line_from_stdin(&buf, &unquoted, prefix, prefix_length,
> >> +                                       nul_term_line, set_executable_bit,
> >> +                                       oflags | f);
> >> +                       strbuf_release(&buf);
> >> +               }
> >>                 strbuf_release(&unquoted);
> >> -               strbuf_release(&line);
> >> +               string_list_clear(&list, 0);
> >>         }
> >>
> >>         if (split_index > 0) {
> >> --
> >> 2.35.1.1428.g1c1a0152d61
> >>
> >
> > This buffering introduces the same potential risk of the
> > "stdin-feeder" process not being able to see objects right away as my
> > version had. I'm planning to mitigate the issue by unplugging the bulk
> > checkin when issuing a verbose report so that anyone who's using that
> > output to synchronize can still see what they're expecting.
>
> I was rather terse in the commit message, I meant (but forgot some
> words) "doesn't harm thing for performance [in the above test]", but
> converting this to a string_list is clearly & regression that shouldn't
> be kept.
>
> I just wanted to demonstrate method of doing this by passing down the
> HASH_* flags, and found that writing the state-machine to "buffer ahead"
> by one line so that we can eventually know in the loop if we're in the
> "last" line or not was tedious, so I came up with this POC. But we
> clearly shouldn't lose the "streaming" aspect.
>

From my experience working on several state machines in the Windows OS, they are notoriously difficult to understand and extend. I wouldn't want every top-level command that does something interesting to have to deal with that.

Show 76 quoted lines
> But anyway, now that I look at this again the smart thing here (surely?)
> is to keep the simple getline() loop and not ever issue a
> HASH_N_OBJECTS_LAST for the Nth item, instead we should in this case do
> the "checkpoint fsync" at the point that we write the actual index.
>
> Because an existing redundancy in your series is that you'll do the
> fsync() the same way for "git unpack-objects" as for "git
> {update-index,add}".
>
> I.e. in the former case adding the N objects is all we're doing, so the
> "last object" is the point at which we need to flush the previous N to
> disk.
>
> But for "update-index/add" you'll do at least 2 fsync()'s in the bulk
> mode, when it should be one. I.e. the equivalent of (leaving aside the
> tmp-objdir migration part of it), if writing objects A && B:
>
>     ## METHOD ONE
>     # A
>     write(objects/A.tmp)
>     bulk_fsync(objects/A.tmp)
>     rename(objects/A.tmp, objects/A)
>     # B
>     write(objects/B.tmp)
>     bulk_fsync(objects/B.tmp)
>     rename(objects/B.tmp, objects/B)
>     # "cookie"
>     write(bulk_fsync_XXXXXX)
>     fsync(bulk_fsync_XXXXXX)
>     # ref
>     write(INDEX.tmp, $(git rev-parse B))
>     fsync(INDEX.tmp)
>     rename(INDEX.tmp, INDEX)
>
> This series on top changes that so we know that we're doing N, so we
> don't need the seperate "cookie", we can just use the B object as the
> cookie, as we know it comes last;
>
>     ## METHOD TWO
>     # A -- SAME as above
>     write(objects/A.tmp)
>     bulk_fsync(objects/A.tmp)
>     rename(objects/A.tmp, objects/A)
>     # B -- SAME as above, with s/bulk_fsync/fsync/
>     write(objects/B.tmp)
>     fsync(objects/B.tmp)
>     rename(objects/B.tmp, objects/B)
>     # "cookie" -- GONE!
>     # ref -- SAME
>     write(INDEX.tmp, $(git rev-parse B))
>     fsync(INDEX.tmp)
>     rename(INDEX.tmp, INDEX)
>
> But really, we should instead realize that we're not doing
> "unpack-objects", but have a "ref update" at the end (whether that's a
> ref, or an index etc.) and do:
>
>     ## METHOD THREE
>     # A -- SAME as above
>     write(objects/A.tmp)
>     bulk_fsync(objects/A.tmp)
>     rename(objects/A.tmp, objects/A)
>     # B -- SAME as the first
>     write(objects/B.tmp)
>     bulk_fsync(objects/B.tmp)
>     rename(objects/B.tmp, objects/B)
>     # ref -- SAME
>     write(INDEX.tmp, $(git rev-parse B))
>     fsync(INDEX.tmp)
>     rename(INDEX.tmp, INDEX)
>
> Which cuts our number of fsync() operations down from 2 to 1, ina
> addition to removing the need for the "cookie", which is only there
> because we didn't keep track of where we were in the sequence as in my
> 2/7 and 5/7.
>

I agree that this is a great direction to go in as an extension to this work (i.e. a subsequent patch). I saw in one of your mails on v2 of your rfc series that you mentioned a "lightweight transaction-y thing". I've been thinking along the same lines myself, but wanted to treat that as a separable concern. In my ideal world, we'd just use a real database for loose objects, the index, and refs and let that handle the transaction management. But in lieu of that, having a transaction that looks across the ODB, index, and refs would let us batch syncs optimally.

Show 30 quoted lines
> And it would be the same for tmp-objdir, the rename dance is a bit
> different, but we'd do the "full" fsync() while on the INDEX.tmp, then
> migrate() the tmp-objdir, and once that's done do the final:
>
>     rename(INDEX.tmp, INDEX)
>
> I.e. we'd fsync() the content once, and only have the renme() or link()
> operations left. For POSIX we'd need a few more fsync() for the
> metadata, but this (i.e. your) series already makes the hard assumption
> that we don't need to do that for rename().
>
> > I think the code you've presented here is a lot of diff to accomplish
> > the same thing that my series does, where this specific update-index
> > caller has been roto-tilled to provide the needed
> > begin/end-transaction points.
>
> Any caller of these APIs will need the "unsigned oflags" sooner than
> later anyway, as they need to pass down e.g. HASH_WRITE_OBJECT. We just
> do it slightly earlier.
>
> And because of that in the general case it's really not the same, I
> think it's a better approach. You've already got the bug in yours of
> needlessly setting up the bulk checkin for !HASH_WRITE_OBJECT in
> update-index, which this neatly solves by deferring the "bulk" mechanism
> until the codepath that's past that and into the "real" object writing.
>
> We can also die() or error out in the object writing before ever getting
> to writing the object, in which case we'd do some setup that we'd need
> to tear down again, by deferring it until the last moment...
>

I'll be submitting a new version to the list which sets up the tmp objdir lazily on first actual write, so the concern about writing to the ODB needlessly should go away.

Show 25 quoted lines
> > And I think there will be a lot of
> > complexity in supporting the same hints for command-line additions
> > (which is roughly equivalent to the git-add workflow).
>
> I left that out due to Junio's comment in
> https://lore.kernel.org/git/xmqqzgljyz34.fsf@gitster.g/; i.e. I don't
> see why we'd find it worthwhile to optimize that case, but we easily
> could (especially per the "just sync the INDEX.tmp" above).
>
> But even if we don't do "THREE" above I think it's still easy, for "TWO"
> we already have as parse_options() state machine to parse argv as it
> comes in. Doing the fsync() on the last object is just a matter of
> "looking ahead" there).
>
> > Every caller
> > that wants batch treatment will have to either implement a state
> > machine or implement a buffering mechanism in order to figure out the
> > begin-end points. Having a separate plug/unplug call eliminates this
> > complexity on each caller.
>
> This is subjective, but I really think that's rather easy to do, and
> much easier to reason about than the global state on the side via
> singletons that your method of avoiding modifying these callers and
> instead having them all consult global state via bulk-checkin.c and
> cache.h demands.
The nice thing about having the ODB handle the batch stuff internally
is that it can present a nice minimal interface to all of the callers.
Yes, it has a complex implementation internally, but that complexity
backs a rather simple API surface:
1. Begin/end transaction (plug/unplug checkin).
2. Find-object by SHA
3. Add object if it doesn't exist
4. Get the SHA without adding anything.

The ODB work is implemented once and callers can easily adopt the transaction API without having to implement their own stuff on the side. Future series can make the transaction span nicely across the ODB, index, and refs.

Show 7 quoted lines
> That API also currently assumes single-threaded writers, if we start
> writing some of this in parallel in e.g. "unpack-objects" we'd need
> mutexes in bulk-object.[ch]. Isn't that a lot easier when the caller
> would instead know something about the special nature of the transaction
> they're interacting with, and that the 1st and last item are important
> (for a "BEGIN" and "FLUSH").
>

The API as sketched above doesn't deeply assume single-threadedness for the "find object by SHA" or "add object if it doesn't exist". There is a single-threaded assumption for begin/end-transaction. The implementation can use pthread_once to handle anything that needs to be done lazily when adding objects.

Show 9 quoted lines
> > Btw, I'm planning in a future series to reduce the system calls
> > involved in renaming a file by taking advantage of the renameat2
> > system call and equivalents on other platforms.  There's a pretty
> > strong motivation to do that on Windows.
>
> What do you have in mind for renameat2() specifically?  I.e. which of
> the 3x flags it implements will benefit us? RENAME_NOREPLACE to "move"
> the tmp_OBJ to an eventual OBJ?
>

Yes RENAME_NOREPLACE. I'd want to introduce a helper called git_rename_noreplace and use it instead of the link dance.

Show 9 quoted lines
> Generally: There's some low-hanging fruit there. E.g. for tmp-objdir we
> slavishly go through the motion of writing an tmp_OBJ, writing (and
> possibly syncing it), then renaming that tmp_OBJ to OBJ.
>
> We could clearly just avoid that in some/all cases that use
> tmp-objdir. I.e. we're writing to a temporary store anyway, so why the
> tmp_OBJ files? We could just write to the final destinations instead,
> they're not reachable (by ref or OID lookup) from anyone else yet.
>

We were thinking before that there could be some concurrency in the tmp_objdir, though I personally don't believe it's possible for the typical bulk checkin case. Using the final name in the tmp objdir would be a nice optimization, but I think that it's a separable concern that shouldn't block the bigger win from eliminating the cache flushes.

Show 29 quoted lines
> But even then I don't see how you'd get away with reducing some classes
> of syscalls past the 2x increase for some (leading an overall increase,
> but not a ~2x overall increase as noted in:
> https://lore.kernel.org/git/RFC-patch-7.7-481f1d771cb-20220323T033928Z-avarab@gmail.com/)
> as long as you use the tmp-objdir API. It's always going to have to
> write tmpdir/OBJ and link()/rename() that to OBJ.
>
> Now, I do think there's an easy way by extending the API use I've
> introduced in this RFC to do it. I.e. we'd just do:
>
>     ## METHOD FOUR
>     # A -- SAME as THREE, except no rename()
>     write(objects/A.tmp)
>     bulk_fsync(objects/A.tmp)
>     # B -- SAME as THREE, except no rename()
>     write(objects/B.tmp)
>     bulk_fsync(objects/B.tmp)
>     # ref -- SAME
>     write(INDEX.tmp, $(git rev-parse B))
>     fsync(INDEX.tmp)
>     # NEW: do all the renames at the end:
>     rename(objects/A.tmp, objects/A)
>     rename(objects/B.tmp, objects/B)
>     rename(INDEX.tmp, INDEX)
>
> That seems like an obvious win to me in any case. I.e. the tmp-objdir
> API isn't really a close fit for what we *really* want to do in this
> case.
>

I think this is the right place to get to eventually. I believe the best way to get there is to keep the plug/unplug bulk checkin functionality (rebranding it as an 'ODB transaction') and then make that a sub-transaction of a larger 'git repo transaction.'

Show 15 quoted lines
> I.e. the reason it does everything this way is because it was explicitly
> designed for 722ff7f876c (receive-pack: quarantine objects until
> pre-receive accepts, 2016-10-03), where it's the right trade-off,
> because we'd like to cheaply "rm -rf" the whole thing if e.g. the
> "pre-receive" hook rejects the push.
>
> *AND* because it's made for the case of other things concurrently
> needing access to those objects. So pedantically you would need it for
> some modes of "git update-index", but not e.g. "git unpack-objects"
> where we really are expecting to keep all of them.
>
> > Thanks for the concrete code,
>
> ..but no thanks? I.e. it would be useful to explicitly know if you're
> interested or open to running with some of the approach in this RFC.

I'm still at the point of arguing with you about your RFC, but I'm _not_ currently leaning toward adopting your approach. I think from a separation-of-concerns perspective, we shouldn't change top-level git commands to try hard to track first/last object. The ODB should conceptually handle it internally as part of a higher-level transaction. Consider cmd_add, which does its interesting add_file_to_index from the update_callback coming from the diff code: I believe it would be hopelessly complex/impossible to do the tracking required to pass the LAST_OF_N flag to a multiplexed write API.

We have a pretty clear example from the database world that begin/end-transaction is the right way to design the API for the task we want to accomplish. It's also how many filesystems work internally. I don't want to reinvent the bicycle here.

Thanks, Neeraj

Previous: Ævar Arnfjörð BjarmasonNext: Ævar Arnfjörð Bjarmason
Message 42 of 175 in “core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objects”
  1. 0/7 core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objectsNeeraj K. Singh via GitGitGadget, Mar 15, 2022
  2. 1/7 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Mar 15, 2022
  3. Junio C HamanoMar 16, 2022
  4. Neeraj SinghMar 16, 2022
  5. Junio C HamanoMar 16, 2022
  6. Neeraj SinghMar 16, 2022
  7. Junio C HamanoMar 16, 2022
  8. Neeraj SinghMar 16, 2022
  9. 2/7 core.fsyncmethod: batched disk flushes for loose-objectsNeeraj Singh via GitGitGadget, Mar 15, 2022
  10. Patrick SteinhardtMar 16, 2022
  11. Neeraj SinghMar 16, 2022
  12. Patrick SteinhardtMar 17, 2022
  13. Bagas SanjayaMar 16, 2022
  14. Neeraj SinghMar 16, 2022
  15. 4/7 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 15, 2022
  16. 3/7 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 15, 2022
  17. 5/7 core.fsync: use batch mode and sync loose objects by default on WindowsNeeraj Singh via GitGitGadget, Mar 15, 2022
  18. 6/7 core.fsyncmethod: tests for batch modeNeeraj Singh via GitGitGadget, Mar 15, 2022
  19. 7/7 core.fsyncmethod: performance tests for add and stashNeeraj Singh via GitGitGadget, Mar 15, 2022
  20. 0/7 core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objectsNeeraj K. Singh via GitGitGadget, Mar 20, 2022
  21. 1/7 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Mar 20, 2022
  22. 2/7 core.fsyncmethod: batched disk flushes for loose-objectsNeeraj Singh via GitGitGadget, Mar 20, 2022
  23. Ævar Arnfjörð BjarmasonMar 21, 2022
  24. Neeraj SinghMar 21, 2022
  25. Ævar Arnfjörð BjarmasonMar 21, 2022
  26. Neeraj SinghMar 21, 2022
  27. Ævar Arnfjörð BjarmasonMar 21, 2022
  28. Neeraj SinghMar 22, 2022
  29. Ævar Arnfjörð BjarmasonMar 22, 2022
  30. Neeraj SinghMar 22, 2022
  31. 0/7 bottom-up ns/batched-fsync & "plugging" in object-file.cÆvar Arnfjörð Bjarmason, Mar 23, 2022
  32. 2/7 unpack-objects: add skeleton HASH_N_OBJECTS{,_{FIRST,LAST}} flagsÆvar Arnfjörð Bjarmason, Mar 23, 2022
  33. 1/7 write-or-die.c: remove unused fsync_component() functionÆvar Arnfjörð Bjarmason, Mar 23, 2022
  34. Neeraj SinghMar 23, 2022
  35. 4/7 update-index: use a utility function for stdin consumptionÆvar Arnfjörð Bjarmason, Mar 23, 2022
  36. 3/7 object-file: pass down unpack-objects.c flags for "bulk" checkinÆvar Arnfjörð Bjarmason, Mar 23, 2022
  37. 5/7 update-index: pass down an "oflags" argumentÆvar Arnfjörð Bjarmason, Mar 23, 2022
  38. 6/7 update-index: rename "buf" to "line"Ævar Arnfjörð Bjarmason, Mar 23, 2022
  39. 7/7 update-index: make use of HASH_N_OBJECTS{,_{FIRST,LAST}} flagsÆvar Arnfjörð Bjarmason, Mar 23, 2022
  40. Neeraj SinghMar 23, 2022
  41. Ævar Arnfjörð BjarmasonMar 23, 2022
  42. Neeraj SinghMar 23, 2022
  43. 0/7 bottom-up ns/batched-fsync & "plugging" in object-file.cÆvar Arnfjörð Bjarmason, Mar 23, 2022
  44. 1/7 unpack-objects: add skeleton HASH_N_OBJECTS{,_{FIRST,LAST}} flagsÆvar Arnfjörð Bjarmason, Mar 23, 2022
  45. Neeraj SinghMar 23, 2022
  46. 2/7 object-file: pass down unpack-objects.c flags for "bulk" checkinÆvar Arnfjörð Bjarmason, Mar 23, 2022
  47. Neeraj SinghMar 23, 2022
  48. 3/7 update-index: pass down skeleton "oflags" argumentÆvar Arnfjörð Bjarmason, Mar 23, 2022
  49. 4/7 update-index: have the index fsync() flush the loose objectsÆvar Arnfjörð Bjarmason, Mar 23, 2022
  50. Neeraj SinghMar 23, 2022
  51. 5/7 add: use WLI_NEED_LOOSE_FSYNC for new "only the index" bulk fsync()Ævar Arnfjörð Bjarmason, Mar 23, 2022
  52. 7/7 fsync docs: add new fsyncMethod.batch.quarantine, elaborate on oldÆvar Arnfjörð Bjarmason, Mar 23, 2022
  53. Neeraj SinghMar 23, 2022
  54. 6/7 fsync docs: update for new syncing semanticsÆvar Arnfjörð Bjarmason, Mar 23, 2022
  55. Junio C HamanoMar 21, 2022
  56. Neeraj SinghMar 21, 2022
  57. Ævar Arnfjörð BjarmasonMar 23, 2022
  58. Neeraj SinghMar 24, 2022
  59. 3/7 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 20, 2022
  60. Ævar Arnfjörð BjarmasonMar 21, 2022
  61. Neeraj SinghMar 21, 2022
  62. Ævar Arnfjörð BjarmasonMar 21, 2022
  63. Junio C HamanoMar 21, 2022
  64. Neeraj SinghMar 21, 2022
  65. 4/7 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 20, 2022
  66. Junio C HamanoMar 21, 2022
  67. Neeraj SinghMar 21, 2022
  68. Neeraj SinghMar 22, 2022
  69. 7/7 core.fsyncmethod: performance tests for add and stashNeeraj Singh via GitGitGadget, Mar 20, 2022
  70. 5/7 core.fsync: use batch mode and sync loose objects by default on WindowsNeeraj Singh via GitGitGadget, Mar 20, 2022
  71. 6/7 core.fsyncmethod: tests for batch modeNeeraj Singh via GitGitGadget, Mar 20, 2022
  72. Junio C HamanoMar 21, 2022
  73. Neeraj SinghMar 22, 2022
  74. Junio C HamanoMar 21, 2022
  75. Neeraj SinghMar 21, 2022
  76. Junio C HamanoMar 21, 2022
  77. 00/11 core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objectsNeeraj K. Singh via GitGitGadget, Mar 24, 2022
  78. 01/11 bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'Neeraj Singh via GitGitGadget, Mar 24, 2022
  79. Ævar Arnfjörð BjarmasonMar 24, 2022
  80. Neeraj SinghMar 24, 2022
  81. 02/11 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Mar 24, 2022
  82. 03/11 object-file: pass filename to fsync_or_dieNeeraj Singh via GitGitGadget, Mar 24, 2022
  83. 04/11 core.fsyncmethod: batched disk flushes for loose-objectsNeeraj Singh via GitGitGadget, Mar 24, 2022
  84. 05/11 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 24, 2022
  85. Junio C HamanoMar 24, 2022
  86. Neeraj SinghMar 24, 2022
  87. Junio C HamanoMar 24, 2022
  88. Neeraj SinghMar 24, 2022
  89. 06/11 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 24, 2022
  90. 07/11 core.fsync: use batch mode and sync loose objects by default on WindowsNeeraj Singh via GitGitGadget, Mar 24, 2022
  91. 08/11 test-lib-functions: add parsing helpers for ls-files and ls-treeNeeraj Singh via GitGitGadget, Mar 24, 2022
  92. 10/11 core.fsyncmethod: performance tests for add and stashNeeraj Singh via GitGitGadget, Mar 24, 2022
  93. 11/11 core.fsyncmethod: correctly camel-case warning messageNeeraj Singh via GitGitGadget, Mar 24, 2022
  94. 09/11 core.fsyncmethod: tests for batch modeNeeraj Singh via GitGitGadget, Mar 24, 2022
  95. Ævar Arnfjörð BjarmasonMar 24, 2022
  96. Neeraj SinghMar 24, 2022
  97. Ævar Arnfjörð BjarmasonMar 26, 2022
  98. Junio C HamanoMar 24, 2022
  99. Neeraj SinghMar 24, 2022
  100. 00/13 core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objectsNeeraj K. Singh via GitGitGadget, Mar 29, 2022
  101. 01/13 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Mar 29, 2022
  102. 02/13 bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'Neeraj Singh via GitGitGadget, Mar 29, 2022
  103. 03/13 object-file: pass filename to fsync_or_dieNeeraj Singh via GitGitGadget, Mar 29, 2022
  104. 04/13 core.fsyncmethod: batched disk flushes for loose-objectsNeeraj Singh via GitGitGadget, Mar 29, 2022
  105. 05/13 cache-tree: use ODB transaction around writing a treeNeeraj Singh via GitGitGadget, Mar 29, 2022
  106. 06/13 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 29, 2022
  107. 07/13 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 29, 2022
  108. 08/13 core.fsync: use batch mode and sync loose objects by default on WindowsNeeraj Singh via GitGitGadget, Mar 29, 2022
  109. 12/13 core.fsyncmethod: performance tests for add and stashNeeraj Singh via GitGitGadget, Mar 29, 2022
  110. Neeraj SinghMar 29, 2022
  111. 10/13 core.fsyncmethod: tests for batch modeNeeraj Singh via GitGitGadget, Mar 29, 2022
  112. 13/13 core.fsyncmethod: correctly camel-case warning messageNeeraj Singh via GitGitGadget, Mar 29, 2022
  113. 11/13 t/perf: add iteration setup mechanism to perf-libNeeraj Singh via GitGitGadget, Mar 29, 2022
  114. Neeraj SinghMar 29, 2022
  115. Junio C HamanoMar 29, 2022
  116. 09/13 test-lib-functions: add parsing helpers for ls-files and ls-treeNeeraj Singh via GitGitGadget, Mar 29, 2022
  117. Ævar Arnfjörð BjarmasonMar 29, 2022
  118. Neeraj SinghMar 29, 2022
  119. Ævar Arnfjörð BjarmasonMar 29, 2022
  120. Neeraj SinghMar 29, 2022
  121. 00/14 core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objectsNeeraj K. Singh via GitGitGadget, Mar 30, 2022
  122. 02/14 bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'Neeraj Singh via GitGitGadget, Mar 30, 2022
  123. Junio C HamanoMar 30, 2022
  124. Neeraj SinghMar 31, 2022
  125. 01/14 bulk-checkin: rename 'state' variable and separate 'plugged' booleanNeeraj Singh via GitGitGadget, Mar 30, 2022
  126. Junio C HamanoMar 30, 2022
  127. Neeraj SinghMar 30, 2022
  128. Junio C HamanoMar 30, 2022
  129. Neeraj SinghMar 31, 2022
  130. Junio C HamanoMar 31, 2022
  131. Neeraj SinghMar 31, 2022
  132. 03/14 object-file: pass filename to fsync_or_dieNeeraj Singh via GitGitGadget, Mar 30, 2022
  133. Junio C HamanoMar 30, 2022
  134. Neeraj SinghMar 30, 2022
  135. 04/14 core.fsyncmethod: batched disk flushes for loose-objectsNeeraj Singh via GitGitGadget, Mar 30, 2022
  136. Junio C HamanoMar 30, 2022
  137. Neeraj SinghMar 31, 2022
  138. Junio C HamanoMar 31, 2022
  139. Neeraj SinghMar 31, 2022
  140. Junio C HamanoApr 1, 2022
  141. 05/14 cache-tree: use ODB transaction around writing a treeNeeraj Singh via GitGitGadget, Mar 30, 2022
  142. Junio C HamanoMar 30, 2022
  143. Neeraj SinghMar 30, 2022
  144. 07/14 update-index: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 30, 2022
  145. Junio C HamanoMar 30, 2022
  146. Neeraj SinghMar 30, 2022
  147. 09/14 core.fsync: use batch mode and sync loose objects by default on WindowsNeeraj Singh via GitGitGadget, Mar 30, 2022
  148. 10/14 test-lib-functions: add parsing helpers for ls-files and ls-treeNeeraj Singh via GitGitGadget, Mar 30, 2022
  149. 08/14 unpack-objects: use the bulk-checkin infrastructureNeeraj Singh via GitGitGadget, Mar 30, 2022
  150. 06/14 builtin/add: add ODB transaction around add_files_to_cacheNeeraj Singh via GitGitGadget, Mar 30, 2022
  151. Junio C HamanoMar 30, 2022
  152. 11/14 core.fsyncmethod: tests for batch modeNeeraj Singh via GitGitGadget, Mar 30, 2022
  153. Junio C HamanoMar 30, 2022
  154. Neeraj SinghMar 31, 2022
  155. 12/14 t/perf: add iteration setup mechanism to perf-libNeeraj Singh via GitGitGadget, Mar 30, 2022
  156. 13/14 core.fsyncmethod: performance tests for batch modeNeeraj Singh via GitGitGadget, Mar 30, 2022
  157. Neeraj SinghMar 31, 2022
  158. 14/14 core.fsyncmethod: correctly camel-case warning messageNeeraj Singh via GitGitGadget, Mar 30, 2022
  159. 01/12 bulk-checkin: rename 'state' variable and separate 'plugged' booleannksingh85@gmail.com, Apr 5, 2022
  160. 00/12 core.fsyncmethod: add 'batch' mode for faster fsyncing of multiple objectsnksingh85@gmail.com, Apr 5, 2022
  161. Junio C HamanoApr 6, 2022
  162. Junio C HamanoMay 19, 2022
  163. Neeraj SinghMay 19, 2022
  164. Johannes SchindelinMay 24, 2022
  165. 04/12 cache-tree: use ODB transaction around writing a treenksingh85@gmail.com, Apr 5, 2022
  166. 10/12 core.fsyncmethod: tests for batch modenksingh85@gmail.com, Apr 5, 2022
  167. 12/12 core.fsyncmethod: performance tests for batch modenksingh85@gmail.com, Apr 5, 2022
  168. 11/12 t/perf: add iteration setup mechanism to perf-libnksingh85@gmail.com, Apr 5, 2022
  169. 09/12 test-lib-functions: add parsing helpers for ls-files and ls-treenksingh85@gmail.com, Apr 5, 2022
  170. 07/12 unpack-objects: use the bulk-checkin infrastructurenksingh85@gmail.com, Apr 5, 2022
  171. 05/12 builtin/add: add ODB transaction around add_files_to_cachenksingh85@gmail.com, Apr 5, 2022
  172. 08/12 core.fsync: use batch mode and sync loose objects by default on Windowsnksingh85@gmail.com, Apr 5, 2022
  173. 02/12 bulk-checkin: rebrand plug/unplug APIs as 'odb transactions'nksingh85@gmail.com, Apr 5, 2022
  174. 03/12 core.fsyncmethod: batched disk flushes for loose-objectsnksingh85@gmail.com, Apr 5, 2022
  175. 06/12 update-index: use the bulk-checkin infrastructurenksingh85@gmail.com, Apr 5, 2022

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.