git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] ref-filter: sort numerically when ":size" is used

From
Jeff King <peff@peff.net>
Date
Sep 1, 2023, 19:16 UTC
Message-ID
<20230901191639.GA1955435@coredump.intra.peff.net>
In-Reply-To
<ZPI0e1XzZrDV2fJk@five231003>
On Sat, Sep 02, 2023 at 12:29:07AM +0530, Kousik Sanagavarapu wrote:
Show 12 quoted lines
> > > > I think they are covered implicitly by the "else" block of the
> > > > conditional that checks for FIELD_STR.
> > > 
> > > Ah, OK.  That needs to be future-proofed to force future developers
> > > who want to add different FIELD_FOO type to look at the comparison
> > > logic.  If we want to do so, it should be done as a separate topic
> > > for cleaning-up the mess, not as part of this effort.
> 
> What I also find weird is the fact that we assign a "cmp_type" to the
> whole atom. Like "contents" is FIELD_STR and "objectsize" is "FIELD_ULONG"
> in "valid_atom". This seems wrong because the options of the atoms should be
> the ones deciding the "cmp_type", no?

I think the data structure is a little confusing if you haven't worked with it before. But basically each "atom" corresponds to a single "%()" block in the format. So if you ran:

  git for-each-ref --format="%(contents:size) %(contents:body)"

you'd have two atoms in the used_atom struct: one for the size and one for the body.

IMHO the code would be a lot easier to work with if the atoms were structured as a parse tree with child pointers (especially when you get into things like "if" that have sub-expressions). I think one of the reasons that used_atom is an array is to de-duplicate repeated mentions (so if you formatted "%(foo) %(foo)" it would only have to store the computed value once).

But I think that is the wrong way to optimize it. We shouldn't be storing any strings per-atom, but rather walking the parse tree to produce a single output buffer. And the values should be cheap to fill in, because we should parse the object as necessary up front. This is more or less the way the pretty.c parser does it.

But that is all quite a large tangent from what you're working on, and would probably be a ground-up rewrite of the formatting code. You can safely ignore my rant for the purposes of your patch. ;)

> I wanted to leave the "cmp_type" field of the atom untouched because that
> would mess up this "global" setting of "contents" to be a "FIELD_STR" (or
> even "raw" for that matter). Although that seems like a bad idea, after
> I've read Junio's and your comments.

Yeah, I agree that would be a problem if there were one global "contents". But we are allocating a new atom struct on the fly for the contents:size directive that we parse, and so on.

-Peff
Previous: Kousik SanagavarapuNext: Junio C Hamano
Message 7 of 14 in “ref-filter: sort numerically when ":size" is used”
  1. ref-filter: sort numerically when ":size" is usedKousik Sanagavarapu, Sep 1, 2023
  2. Junio C HamanoSep 1, 2023
  3. Jeff KingSep 1, 2023
  4. Junio C HamanoSep 1, 2023
  5. Jeff KingSep 1, 2023
  6. Kousik SanagavarapuSep 1, 2023
  7. Jeff KingSep 1, 2023
  8. Junio C HamanoSep 1, 2023
  9. Jeff KingSep 1, 2023
  10. Junio C HamanoSep 1, 2023
  11. Jeff KingSep 1, 2023
  12. ref-filter: sort numerically when ":size" is usedKousik Sanagavarapu, Sep 2, 2023
  13. Kousik SanagavarapuSep 2, 2023
  14. Junio C HamanoSep 2, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.