git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v3 4/4] commit: rewrite read_graft_line

From
Jeff King <peff@peff.net>
Date
Aug 18, 2017, 06:43 UTC
Message-ID
<20170818064335.h5sr5iz7mh64axji@sigill.intra.peff.net>
In-Reply-To
<cb98970b3f6c175321f52efb65deb48f9cfeabae.1503020338.git.patryk.obara@gmail.com>
On Fri, Aug 18, 2017 at 03:59:38AM +0200, Patryk Obara wrote:
Show 6 quoted lines
> Determine the number of object_id's to parse in a single graft line by
> counting separators (whitespace characters) instead of dividing by
> length of hash representation.
> 
> This way graft parsing code can support different sizes of hashes
> without any further code adaptations.

Sounds like a reasonable approach, though I wonder what happens if our counting pass differs in its behavior from the actual parse.

E.g., here:
> +	/* count number of blanks to determine size of array to allocate */
> +	for (i = 0, n = 0; i < line->len; i++)
> +		if (isspace(line->buf[i]))
> +			n++;

If we see multiple spaces like "1234abcd 5678abcd" we'll allocate a slot for each. I think that's OK because here:

> +	for (i = 0; i < graft->nr_parent; i++)
> +		if (!isspace(*tail++) || parse_oid_hex(tail, &graft->parent[i], &tail))
>  			goto bad_graft_data;

We'd reject such an input totally (though as an interesting side effect, you can convince the parser to allocate 20x as much RAM as you send it; one oid for each space).

So we're probably fine. The two parsing passes are right next to each other and are sufficiently simple and strict that we don't have to worry about them diverging.

The single-pass alternative would probably be to read into a dynamic structure like an oid_array, and then copy the result into the flex structure.

Or of course to stop using a flex structure, as your original pass did. I agree with Junio that the use of object_id's is orthogonal to using a FLEX_ARRAY. But I could also see an argument that the complexity the flex array adds here isn't worth the savings. The main benefits of a flex array are:

  1. Less memory used. But we don't expect to see a large enough number
     of grafts for this to matter.
  2. A more compact memory representation, which can be faster. But
     accessing the parent list of a graft isn't going to be the hot code
     path. It probably only happens once in a program run when we
     rewrite the parents (versus checking the grafted commit, which is
     looked up once per commit we access).
  3. It's easier to free the struct and its associated resources in a
     single free(). But we never free the graft list.

Reading your original, my thought was "why _not_ keep doing it as a FLEX_ARRAY, as it saves a little memory, which can't hurt". But seeing this pre-counting phase, I think it does make the code a little more complicated. But I'd be OK with doing it either way.

-Peff
Previous: Patryk ObaraNext: Junio C Hamano
Message 42 of 59 in “Modernize read_graft_line implementation”
  1. 0/5 Modernize read_graft_line implementationPatryk Obara, Aug 15, 2017
  2. 1/5 cache: extend object_id size to sha3-256Patryk Obara, Aug 15, 2017
  3. 2/5 sha1_file: fix hardcoded size in null_sha1Patryk Obara, Aug 15, 2017
  4. Junio C HamanoAug 15, 2017
  5. Stefan BellerAug 15, 2017
  6. Junio C HamanoAug 15, 2017
  7. Patryk ObaraAug 16, 2017
  8. Junio C HamanoAug 16, 2017
  9. 4/5 commit: implement free_commit_graftPatryk Obara, Aug 15, 2017
  10. Stefan BellerAug 15, 2017
  11. Junio C HamanoAug 15, 2017
  12. Patryk ObaraAug 16, 2017
  13. 3/5 commit: replace the raw buffer with strbuf in read_graft_linePatryk Obara, Aug 15, 2017
  14. Stefan BellerAug 15, 2017
  15. Patryk ObaraAug 16, 2017
  16. brian m. carlsonAug 16, 2017
  17. Jeff KingAug 17, 2017
  18. Junio C HamanoAug 17, 2017
  19. Patryk ObaraAug 17, 2017
  20. Junio C HamanoAug 15, 2017
  21. 5/5 commit: rewrite read_graft_linePatryk Obara, Aug 15, 2017
  22. Stefan BellerAug 15, 2017
  23. Junio C HamanoAug 15, 2017
  24. Patryk ObaraAug 16, 2017
  25. Junio C HamanoAug 16, 2017
  26. Stefan BellerAug 15, 2017
  27. 0/4 Modernize read_graft_line implementationPatryk Obara, Aug 16, 2017
  28. 1/4 sha1_file: fix hardcoded size in null_sha1Patryk Obara, Aug 16, 2017
  29. Junio C HamanoAug 16, 2017
  30. 3/4 commit: implement free_commit_graftPatryk Obara, Aug 16, 2017
  31. 2/4 commit: replace the raw buffer with strbuf in read_graft_linePatryk Obara, Aug 16, 2017
  32. 4/4 commit: rewrite read_graft_linePatryk Obara, Aug 16, 2017
  33. Junio C HamanoAug 17, 2017
  34. Patryk ObaraAug 17, 2017
  35. 0/4 Modernize read_graft_line implementationPatryk Obara, Aug 18, 2017
  36. 2/4 commit: replace the raw buffer with strbuf in read_graft_linePatryk Obara, Aug 18, 2017
  37. Jeff KingAug 18, 2017
  38. Patryk ObaraAug 18, 2017
  39. Jeff KingAug 18, 2017
  40. 1/4 sha1_file: fix definition of null_sha1Patryk Obara, Aug 18, 2017
  41. 4/4 commit: rewrite read_graft_linePatryk Obara, Aug 18, 2017
  42. Jeff KingAug 18, 2017
  43. Junio C HamanoAug 18, 2017
  44. Patryk ObaraAug 18, 2017
  45. Jeff KingAug 18, 2017
  46. Junio C HamanoAug 18, 2017
  47. Patryk ObaraAug 18, 2017
  48. Junio C HamanoAug 18, 2017
  49. 3/4 commit: allocate array using object_id sizePatryk Obara, Aug 18, 2017
  50. Junio C HamanoAug 18, 2017
  51. 0/4 Modernize read_graft_line implementationPatryk Obara, Aug 18, 2017
  52. 2/4 commit: replace the raw buffer with strbuf in read_graft_linePatryk Obara, Aug 18, 2017
  53. 4/4 commit: rewrite read_graft_linePatryk Obara, Aug 18, 2017
  54. Patryk ObaraAug 18, 2017
  55. Junio C HamanoAug 18, 2017
  56. Patryk ObaraAug 18, 2017
  57. Junio C HamanoAug 18, 2017
  58. 1/4 sha1_file: fix definition of null_sha1Patryk Obara, Aug 18, 2017
  59. 3/4 commit: allocate array using object_id sizePatryk Obara, Aug 18, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.