git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git in Outreachy December 2019?

From
EWEric Wong <e@80x24.org>
Date
Sep 24, 2019, 00:55 UTC
Message-ID
<20190924005529.GA8354@dcvr>
In-Reply-To
<nycvar.QRO.7.76.6.1909171158090.15067@tvgsbejvaqbjf.bet>
Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:
Show 8 quoted lines
> On Mon, 16 Sep 2019, Emily Shaffer wrote:
> >  - try and make progress towards running many tests from a single test
> >    file in parallel - maybe this is too big, I'm not sure if we know how
> >    many of our tests are order-dependent within a file for now...
> 
> Another, potentially more rewarding, project would be to modernize our
> test suite framework, so that it is not based on Unix shell scripting,
> but on C instead.

I worry more C would reduce the amount of contributors (some of the C rewrites already scared me off hacking years ago). I figure more users are familiar with sh than C.

It would also increase the disparity between tests and use of actual users from the command-line.

> The fact that it is based on Unix shell scripting not only costs a lot
> of speed, especially on Windows, it also limits us quite a bit, and I am
> talking about a lot more than just the awkwardness of having to think
> about options of BSD vs GNU variants of common command-line tools.

I agree that it costs a lot of time, and I'm even on Linux using dash as /bin/sh + eatmydata (but ancient laptop)

Show 55 quoted lines
> For example, many, many, if not all, test cases, spend the majority of
> their code on setting up specific scenarios. I don't know about you,
> but personally I have to dive into many of them when things fail (and I
> _dread_ the numbers 0021, 0025 and 3070, let me tell you) and I really
> have to say that most of that code is hard to follow and does not make
> it easy to form a mental model of what the code tries to accomplish.
> 
> To address this, a while ago Thomas Rast started to use `fast-export`ed
> commit histories in test scripts (see e.g. `t/t3206/history.export`). I
> still find that this fails to make it easier for occasional readers to
> understand the ideas underlying the test cases.
> 
> Another approach is to document heavily the ideas first, then use code
> to implement them. For example, t3430 starts with this:
> 
> 	[...]
> 
> 	Initial setup:
> 
> 	    -- B --                   (first)
> 	   /       \
> 	 A - C - D - E - H            (master)
> 	   \    \       /
> 	    \    F - G                (second)
> 	     \
> 	      Conflicting-G
> 
> 	[...]
> 
> 	test_commit A &&
> 	git checkout -b first &&
> 	test_commit B &&
> 	git checkout master &&
> 	test_commit C &&
> 	test_commit D &&
> 	git merge --no-commit B &&
> 	test_tick &&
> 	git commit -m E &&
> 	git tag -m E E &&
> 	git checkout -b second C &&
> 	test_commit F &&
> 	test_commit G &&
> 	git checkout master &&
> 	git merge --no-commit G &&
> 	test_tick &&
> 	git commit -m H &&
> 	git tag -m H H &&
> 	git checkout A &&
> 	test_commit conflicting-G G.t
> 
> 	[...]
> 
> While this is _somewhat_ better than having only the code, I am still
> unhappy about it: this wall of `test_commit` lines interspersed with
> other commands is very hard to follow.
Agreed.  More on the readability part below...

As far as speeding that up, I think moving some parts of test setup to Makefiles + fast-import/fast-export would give us a nice balance of speed + maintainability:

1. initial setup is done using normal commands (or graph drawing tool)
2. the result of setup is "built" with fast-export
3. test uses fast-import
Makefile rules would prevent subsequent test runs from repeating
1. and 2.
Show 8 quoted lines
> If we were to (slowly) convert our test suite framework to C, we could
> change that.
> 
> One idea would be to allow recreating commit history from something that
> looks like the output of `git log`, or even `git log --graph --oneline`,
> much like `git mktree` (which really should have been a test helper
> instead of a Git command, but I digress) takes something that looks like
> the output of `git ls-tree` and creates a tree object from it.

I've been playing with Graph::Easy (Perl5 module) in other projects, and I also think the setup could be more easily expressed with a declarative language (e.g. GNU make)

Show 6 quoted lines
> Another thing that would be much easier if we moved more and more parts
> of the test suite framework to C: we could implement more powerful
> assertions, a lot more easily. For example, the trace output of a failed
> `test_i18ngrep` (or `mingw_test_cmp`!!!) could be made a lot more
> focused on what is going wrong than on cluttering the terminal window
> with almost useless lines which are tedious to sift through.

I fail to see how language choice here matters. But then again, I have plenty of experience writing bad code in ALL languages I know :>

Show 7 quoted lines
> Likewise, having a framework in C would make it a lot easier to improve
> debugging, e.g. by making test scripts "resumable" (guarded by an
> option, it could store a complete state, including a copy of the trash
> directory, before executing commands, which would allow "going back in
> time" and calling a failing command with a debugger, or with valgrind, or
> just seeing whether the command would still fail, i.e. whether the test
> case is flaky).

Resumability sounds like a perfect job for GNU make. (that said, I don't know if you use make or something else to build gfw)

Show 5 quoted lines
> In many ways, our current test suite seems to test Git's functionality
> as much as (core) contributors' abilities to implement test cases in
> Unix shell script, _correctly_, and maybe also contributors' patience.
> You could say that it tests for the wrong thing at least half of the
> time, by design.
Basic (not advanced) sh is already a prerequisite for using git.

Writing correct code and tests in ANY language is still a challenge for me; but I'm least convinced a low-level language such as C is the right language for writing integration tests in.

C is fine for unit tests, and maybe we can use more unit tests and less integration tests.

> It might look like a somewhat less important project, but given that we
> exercise almost 150,000 test cases with every CI build, I think it does
> make sense to grind our axe for a while, so to say.

Something that would benefit both users and regular contributors is the use and adoption of more batch and eval-friendly interfaces. e.g. fast-import/export, cat-file --batch, for-each-ref --perl...

I haven't used hg since 2005, but I know "hg server" exists nowadays to get rid of a lot of startup overhead in Mercurial, and maybe git could steal that idea, too...

Show 7 quoted lines
> Therefore, it might be a really good project to modernize our test
> suite. To take ideas from modern test frameworks such as Jest and try to
> bring them to C. Which means that new contributors would probably be
> better suited to work on this project than Git old-timers!
> 
> And the really neat thing about this project is that it could be done
> incrementally.

I hope to find time to hack some more batch/eval-friendly stuff that can make scripting git more performant; but no idea on my availability :<

Previous: Junio C HamanoNext: Johannes Schindelin
Message 45 of 63 in “Git in Outreachy December 2019?”
  1. Jeff KingAug 27, 2019
  2. Christian CouderAug 31, 2019
  3. Olga TelezhnayaAug 31, 2019
  4. Jeff KingSep 4, 2019
  5. Christian CouderSep 5, 2019
  6. Emily ShafferSep 5, 2019
  7. Carlo ArenasSep 6, 2019
  8. Jeff KingSep 7, 2019
  9. Carlo ArenasSep 7, 2019
  10. Jeff KingSep 7, 2019
  11. Pratyush YadavSep 8, 2019
  12. Jeff KingSep 9, 2019
  13. SZEDER GáborSep 23, 2019
  14. SZEDER GáborSep 26, 2019
  15. Johannes SchindelinSep 26, 2019
  16. SZEDER GáborSep 26, 2019
  17. Johannes SchindelinSep 26, 2019
  18. Jonathan TanSep 13, 2019
  19. Jeff KingSep 13, 2019
  20. Emily ShafferSep 16, 2019
  21. Eric WongSep 16, 2019
  22. SZEDER GáborSep 16, 2019
  23. Jonathan NiederSep 16, 2019
  24. Jeff KingSep 17, 2019
  25. Johannes SchindelinSep 17, 2019
  26. SZEDER GáborSep 17, 2019
  27. Johannes SchindelinSep 23, 2019
  28. SZEDER GáborSep 23, 2019
  29. Johannes SchindelinSep 26, 2019
  30. SZEDER GáborSep 26, 2019
  31. Johannes SchindelinSep 26, 2019
  32. SZEDER GáborSep 26, 2019
  33. Jeff KingSep 27, 2019
  34. SZEDER GáborOct 9, 2019
  35. Jeff KingOct 11, 2019
  36. Jeff KingSep 23, 2019
  37. Johannes SchindelinSep 24, 2019
  38. Christian CouderSep 17, 2019
  39. Johannes SchindelinSep 23, 2019
  40. Jeff KingSep 23, 2019
  41. Jeff KingSep 23, 2019
  42. Johannes SchindelinSep 24, 2019
  43. Jeff KingSep 24, 2019
  44. Junio C HamanoSep 28, 2019
  45. Eric WongSep 24, 2019
  46. Johannes SchindelinSep 26, 2019
  47. Eric WongSep 30, 2019
  48. Junio C HamanoSep 28, 2019
  49. Jonathan TanSep 20, 2019
  50. Emily ShafferSep 21, 2019
  51. Christian CouderSep 23, 2019
  52. Jeff KingSep 23, 2019
  53. Philip OakleySep 23, 2019
  54. Emily ShafferOct 22, 2019
  55. Christian CouderSep 23, 2019
  56. Jonathan TanSep 23, 2019
  57. Jeff KingSep 23, 2019
  58. Jonathan TanSep 23, 2019
  59. Jeff KingSep 23, 2019
  60. Jonathan TanSep 23, 2019
  61. Jeff KingSep 23, 2019
  62. Jonathan TanSep 24, 2019
  63. Jeff KingSep 26, 2019

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.