git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: RFE: "git bisect reverse"

From
EWEaldwulf Wuffinga <ealdwulf@googlemail.com>
Date
May 28, 2009, 21:07 UTC
Message-ID
<efe2b6d70905281407x56bb788aq3dba4b27eb91d7a6@mail.gmail.com>
In-Reply-To
<4A1E00F1.4030709@zytor.com>
On Thu, May 28, 2009 at 4:11 AM, H. Peter Anvin <hpa@zytor.com> wrote:
Show 13 quoted lines
> Again, given a bisection, the information gain by "bisecting" at point x
>  where 0 < x < 1 is:
>
>        -(x log2 x)-((1-x) log2 (1-x))
>
> At x = 0.5 this gives the optimal 1 bit, but the curve is rather flat
> near the top.  You don't drop to 1/2 bit of information until
> x = 0.11 or 0.89, and it doesn't drop to 1/4 bit of information until
> x = 0.04 or 0.96.
>
> Thus, the lack of optimality in searching away from a skip point is much
> smaller than the potential cost of having to having to skip multiple
> nearby points.

I understand that. I didn't mean to imply that there was anything wrong with your proposal, indeed, it makes sense for git-bisect.

What I am interested in is how to extend bisection to the case of intermittent bugs; where a test which observes the fault means that it cannot have been introduced in subsequent commits, but a test which does not observe the fault cannot guarantee that it must have been introduced in a subsequent commit.

The simplest way to deal with this is to try to reduce it to the deterministic case by repeating the test some number of times. It turns out, that this is rather inefficient.

In bbchop, the search algorithm does not assume that the test is deterministic. Therefore, it has to calculate the probabilities in order to know when it has accumulated enough evidence to accuse a particular commit. It turns out that it is not much more expensive to calculate which commit we can expect to gain the most information from by testing it next.

How can I incorporate your skipping feature into this model? The problem is that while (just thinking about the linear case for the moment) there is a fixed boundary at one end - where we actually saw a fault - on the other side there are a bunch of fuzzy probabilities, ultimately bounded by wherever we decided the limit of the search was. So when we get a skip we could hop half way toward the limit. That would be reasonable toward the beginning of the search, but towards the end when most of the probability is concentrated in a small number of commits, it would make no sense.

It would fit a lot better into this algorithm to have some model of the probability that a commit will cause a skip. It doesn't actually have to be a very good one, because if it's poor it will only make the search slightly less efficient, not affect the reliability of the final result.

Ealdwulf
Previous: H. Peter AnvinNext: H. Peter Anvin
Message 13 of 18 in “RFE: "git bisect reverse"”
  1. H. Peter AnvinMay 26, 2009
  2. Sam VilainMay 27, 2009
  3. H. Peter AnvinMay 27, 2009
  4. Christian CouderMay 27, 2009
  5. Ealdwulf WuffingaMay 27, 2009
  6. Clemens BuchacherMay 27, 2009
  7. Ealdwulf WuffingaMay 27, 2009
  8. Sam VilainMay 27, 2009
  9. Ealdwulf WuffingaMay 28, 2009
  10. Sam VilainMay 29, 2009
  11. Ealdwulf WuffingaMay 31, 2009
  12. H. Peter AnvinMay 28, 2009
  13. Ealdwulf WuffingaMay 28, 2009
  14. H. Peter AnvinMay 28, 2009
  15. Ealdwulf WuffingaMay 31, 2009
  16. Christian CouderMay 27, 2009
  17. Nanako ShiraishiMay 27, 2009
  18. Matthieu MoyMay 27, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.