git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git in Outreachy December 2019?

From
Jeff King <peff@peff.net>
Date
Sep 23, 2019, 21:28 UTC
Message-ID
<20190923212834.GA19504@sigill.intra.peff.net>
In-Reply-To
<20190923203854.171170-1-jonathantanmy@google.com>
On Mon, Sep 23, 2019 at 01:38:54PM -0700, Jonathan Tan wrote:
> I didn't have any concrete ideas so I didn't include those, but some
> unrefined ideas:

One risk to a mentoring project like this is that the intern does a good job of steps 1-5, and then in step 6 we realize that the whole thing is not useful, and upstream doesn't want it. Which isn't to say the intern didn't learn something, and the project didn't benefit. Negative results can be useful; but it can also be demoralizing.

I'm not arguing that's going to be the case here. But I do think it's worth talking through these things a bit as part of thinking about proposals.

Show 8 quoted lines
>  - index-pack has the CLI option to specify a message to be written into
>    the .promisor file, but in my patch to write fetched refs to
>    .promisor [1], I ended up making fetch-pack.c write the information
>    because I didn't know how many refs were going to be written (and I
>    didn't want to bump into CLI argument length limits). If we had this
>    feature, I might have been able to pass a callback to index-pack that
>    writes the list of refs once we have the fd into .promisor,
>    eliminating some code duplication (but I haven't verified this).

That makes some sense. We could pass the data over a pipe, but obviously stdin is already in use to receive the pack here. Ideally we'd be able to pass multiple streams between the programs, but I think due to Windows support, we can't assume that arbitrary pipe descriptors will make it across the run-command boundary. So I think we'd be left with communicating via temporary files (which really isn't the worst thing in the world, but has its own complications).

Show 8 quoted lines
>  - In your reply [2] to the above [1], you mentioned the possibility of
>    keeping a list of cutoff points. One way of doing this, as I state in
>    [3], is my original suggestion back in 2017 of one such
>    repository-wide list. If we do this, it would be better for
>    fetch-pack to handle this instead of index-pack, and it seems more
>    efficient to me to have index-pack be able to pass objects to
>    fetch-pack as they are inflated instead of fetch-pack rereading the
>    compressed forms on disk (but again, I haven't verified this).

And this is the flip-side problem: we need to get data back, but we have only stdout, which is already in use (so we need some kind of protocol). That leads to things like the horrible NUL-byte added by 83558686ce (receive-pack: send keepalives during quiet periods, 2016-07-15).

> There are also the debuggability improvements of not having to deal with
> 2 processes.

I think it can sometimes be easier to debug with two separate processes, because the input to index-pack is well-defined and can be repeated without hitting the network (though you do have to figure out how to record the network response, which can be non-trivial). I've also done similar things for running performance simulations.

We'll still have the stand-alone index-pack command, so it can be used for those cases. But as we add more features that utilize the in-process interface, that may eventually stop being feasible.

Show 9 quoted lines
> > [dropping unpack-objects]
> >     Maybe that would be worth making part of the project?
> 
> I'm reluctant to do so because I don't want to increase the scope too
> much - although if my project has relatively narrow scope for an
> Outreachy project, we can do so. As for eliminating the utility of
> having richer communication, I don't think so, because in the situations
> where we require richer communication (right now, situations to do with
> partial clone), we specifically run index-pack anyway.

Yeah, we're in kind of a weird situation there, where unpack-objects is used less and less. I wonder how many surprises are lurking where somebody reasoned about index-pack behavior, but unpack-objects may do something slightly differently (I know this came up when we looked at fsck-ing incoming objects for submodule vulnerabilities).

I kind of wonder if it would be reasonable to just always use index-pack for the sake of simplicity, even if it never learns to actually unpack objects. We've been doing that for years on the server side at GitHub without ill effects (I think the unpack route is slightly more efficient for a thin pack, but since it only kicks in when there are few objects anyway, I wonder how big an advantage it is in general).

-Peff
Previous: Jonathan TanNext: Jonathan Tan
Message 61 of 63 in “Git in Outreachy December 2019?”
  1. Jeff KingAug 27, 2019
  2. Christian CouderAug 31, 2019
  3. Olga TelezhnayaAug 31, 2019
  4. Jeff KingSep 4, 2019
  5. Christian CouderSep 5, 2019
  6. Emily ShafferSep 5, 2019
  7. Carlo ArenasSep 6, 2019
  8. Jeff KingSep 7, 2019
  9. Carlo ArenasSep 7, 2019
  10. Jeff KingSep 7, 2019
  11. Pratyush YadavSep 8, 2019
  12. Jeff KingSep 9, 2019
  13. SZEDER GáborSep 23, 2019
  14. SZEDER GáborSep 26, 2019
  15. Johannes SchindelinSep 26, 2019
  16. SZEDER GáborSep 26, 2019
  17. Johannes SchindelinSep 26, 2019
  18. Jonathan TanSep 13, 2019
  19. Jeff KingSep 13, 2019
  20. Emily ShafferSep 16, 2019
  21. Eric WongSep 16, 2019
  22. SZEDER GáborSep 16, 2019
  23. Jonathan NiederSep 16, 2019
  24. Jeff KingSep 17, 2019
  25. Johannes SchindelinSep 17, 2019
  26. SZEDER GáborSep 17, 2019
  27. Johannes SchindelinSep 23, 2019
  28. SZEDER GáborSep 23, 2019
  29. Johannes SchindelinSep 26, 2019
  30. SZEDER GáborSep 26, 2019
  31. Johannes SchindelinSep 26, 2019
  32. SZEDER GáborSep 26, 2019
  33. Jeff KingSep 27, 2019
  34. SZEDER GáborOct 9, 2019
  35. Jeff KingOct 11, 2019
  36. Jeff KingSep 23, 2019
  37. Johannes SchindelinSep 24, 2019
  38. Christian CouderSep 17, 2019
  39. Johannes SchindelinSep 23, 2019
  40. Jeff KingSep 23, 2019
  41. Jeff KingSep 23, 2019
  42. Johannes SchindelinSep 24, 2019
  43. Jeff KingSep 24, 2019
  44. Junio C HamanoSep 28, 2019
  45. Eric WongSep 24, 2019
  46. Johannes SchindelinSep 26, 2019
  47. Eric WongSep 30, 2019
  48. Junio C HamanoSep 28, 2019
  49. Jonathan TanSep 20, 2019
  50. Emily ShafferSep 21, 2019
  51. Christian CouderSep 23, 2019
  52. Jeff KingSep 23, 2019
  53. Philip OakleySep 23, 2019
  54. Emily ShafferOct 22, 2019
  55. Christian CouderSep 23, 2019
  56. Jonathan TanSep 23, 2019
  57. Jeff KingSep 23, 2019
  58. Jonathan TanSep 23, 2019
  59. Jeff KingSep 23, 2019
  60. Jonathan TanSep 23, 2019
  61. Jeff KingSep 23, 2019
  62. Jonathan TanSep 24, 2019
  63. Jeff KingSep 26, 2019

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.