git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 00/22] Refactor to accept NUL in commit messages

From
Nguyen Thai Ngoc Duy <pclouds@gmail.com>
Date
Oct 23, 2011, 01:24 UTC
Message-ID
<CACsJy8B=TsC4A=R6b3jyYBCvorEDBYHQ8uA864WrB0-3pgNyKA@mail.gmail.com>
In-Reply-To
<7vobx863v3.fsf@alter.siamese.dyndns.org>
2011/10/23 Junio C Hamano <gitster@pobox.com>:
Show 21 quoted lines
> I do not think we want to go this route.
>
> There are two possible approaches to attack this.
>
>  - If we want to show everything after a potential and rare NUL in the log
>   message most of the time, then "struct commit" should just store
>   <ptr,len> pair. This grows "struct commit" with one extra ulong.
>
>  - If we want to give us a way to notice and show these "funnily, this
>   commit log message has a NUL in it" case as an exception in only
>   selected codepaths, then "struct commit" should just gain "flags"
>   4-byte int field between "indegree" and "date", and
>   parse_commit_buffer() should set one bit in the flags when the log
>   message has NUL in it. And teach only these selected codepaths to find
>   the length from the object name with sha1_object_info() as needed. This
>   grows "struct commit" with one 4-byte int, with runtime overhead only
>   where it matters.
>
> The approach taken by the patch wastes two malloc() blocks with their own
> allocation overhead, and unused "alloc" field in the strbuf that does not
> have to be there.
We could allocate just one block with length as the first field:
struct commit_buffer {
        unsigned long len;
        char buf[FLEX_ARRAY];
};

The downside is commit_buffer field type in struct commit changes, which impacts many codepaths. If we agree to allow NUL in commit objects, then all codepaths should be aware of that fact, otherwise funny things may happen because string processing in this function stops early due to NUL, but others run fine..

Jeff's low-level normalization approach sounds much simpler, but probably trickier because we need to identify where to normalize and denormalize. At least with type change, the compiler spots all the places for me.

I would not worry about runtime processing overhead. The string end check is basically converted from "if (*msg)" to "if (msg < msg_end)". There will one more pointer (msg_end) in stack for each call. Unless we do deep recursion, we should be fine. Memory overhead is still something to profile.

-- 
Duy
Previous: Junio C HamanoNext: Junio C Hamano
Message 5 of 26 in “Re: [PATCH 00/22] Refactor to accept NUL in commit messages”
  1. Jeff KingOct 22, 2011
  2. Robin RosenbergOct 23, 2011
  3. Jeff KingOct 23, 2011
  4. Junio C HamanoOct 22, 2011
  5. Nguyen Thai Ngoc DuyOct 23, 2011
  6. Junio C HamanoOct 23, 2011
  7. Nguyen Thai Ngoc DuyOct 23, 2011
  8. Junio C HamanoOct 23, 2011
  9. Nguyen Thai Ngoc DuyOct 23, 2011
  10. Jeff KingOct 23, 2011
  11. Junio C HamanoOct 23, 2011
  12. Junio C HamanoOct 24, 2011
  13. Nguyen Thai Ngoc DuyOct 24, 2011
  14. Štěpán NěmecOct 24, 2011
  15. Jeff KingOct 24, 2011
  16. Štěpán NěmecOct 25, 2011
  17. Junio C HamanoOct 25, 2011
  18. Jeff KingOct 27, 2011
  19. Junio C HamanoOct 27, 2011
  20. Jeff KingOct 27, 2011
  21. Junio C HamanoOct 27, 2011
  22. Jeff KingOct 27, 2011
  23. Junio C HamanoOct 28, 2011
  24. Jeff KingOct 28, 2011
  25. Miles BaderOct 28, 2011
  26. Junio C HamanoOct 28, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.