git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: data loss when doing ls-remote and piped to command

From
RBRolf Eike Beer <eb@emlix.com>
Date
Sep 17, 2021, 06:59 UTC
Message-ID
<2722184.bRktqFsmb4@devpool47>
In-Reply-To
<xmqq7dfgtfpt.fsf@gitster.g>
Am Donnerstag, 16. September 2021, 22:42:22 CEST schrieb Junio C Hamano:
Show 19 quoted lines
> Linus Torvalds <torvalds@linux-foundation.org> writes:
> > On Thu, Sep 16, 2021 at 5:17 AM Rolf Eike Beer <eb@emlix.com> wrote:
> >> Am Donnerstag, 16. September 2021, 12:12:48 CEST schrieb Tobias Ulmer:
> >> > > The redirection seems to be an important part of it. I now did:
> >> > > 
> >> > > git ... 2>&1 | sha256sum
> >> > 
> >> > I've tried to reproduce this since yesterday, but couldn't until now:
> >> > 
> >> > 2>&1 made all the difference, took less than a minute.
> > 
> > So if that redirection is what matters, and what causes problems, I
> > can almost guarantee that the reason is very simple:
> > ...
> > Anyway. That was a long email just to tell people it's almost
> > certainly user error, not the kernel.
> 
> Yes, 2>&1 will mix messages from the standard error stream at random
> places in the output, which explains the checksum quite well.

If there would be any errors. The point is: if I run the command with ">/dev/ null" just to the terminals a hundred times there is never any output on stderr at all. If I pipe stderr into a file it's empty after all of this (yes, I did append, not overwrite).

That the particular construct in this case is sort of nonsense is granted, I just hit it because some tool here used some very similar construct and suddenly started failing. "less" isn't the original reproducer, it was just something I started testing with to be able to easily visually inspect the output.

What you need is a _fast_ git server. kernel.org or github.com seem to be too slow for this if you don't sit somewhere in their datacenter. Use something in your local network, a Xeon E5 with lot's of RAM and connected with 1GBit/s Ethernet in my case.

And the reader must be "somewhat" slow. Using sha256sum works reliably for me. Using "wc -l" does not, also md5sum and sha1sum are too fast as it seems.

When I run the whole thing with strace I can't see the effect, which isn't really surprising. But there is a difference between the cases where I run with redirection "2>&1":

ioctl(2, TCGETS, 0x7ffd6f119b10) = -1 ENOTTY (Inappropriate ioctl for device)

and without:
ioctl(2, TCGETS, {B38400 opost isig icanon echo ...}) = 0
AFAICT this is the only place where fd 2 is used at all during the whole time.
Regards,
Eike
-- 
Rolf Eike Beer, emlix GmbH, http://www.emlix.com
Fon +49 551 30664-0, Fax +49 551 30664-11
Gothaer Platz 3, 37083 Göttingen, Germany
Sitz der Gesellschaft: Göttingen, Amtsgericht Göttingen HR B 3160
Geschäftsführung: Heike Jordan, Dr. Uwe Kracke – Ust-IdNr.: DE 205 198 055

emlix - smart embedded open source
Previous: Junio C HamanoNext: Jeff King
Message 10 of 13 in “data loss when doing ls-remote and piped to command”
  1. Rolf Eike BeerSep 15, 2021
  2. Junio C HamanoSep 15, 2021
  3. Rolf Eike BeerSep 16, 2021
  4. Tobias UlmerSep 16, 2021
  5. Rolf Eike BeerSep 16, 2021
  6. Mike GalbraithSep 16, 2021
  7. Mike GalbraithSep 17, 2021
  8. Linus TorvaldsSep 16, 2021
  9. Junio C HamanoSep 16, 2021
  10. Rolf Eike BeerSep 17, 2021
  11. Jeff KingSep 17, 2021
  12. Linus TorvaldsSep 17, 2021
  13. Mike GalbraithSep 18, 2021

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.