{"thread":{"id":"61295","subject":"[RFC] Avaiable disk space when automatic garbage collection kicks in","startedAt":"2024-04-08T16:29:27Z","lastAt":"2024-06-12T17:25:54Z","messageCount":4,"participants":["Dragan Simic","rsbecker@nexbridge.com"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"492560","messageId":"950d4ad3bcee79df1424faee09eb6b00@manjaro.org","threadId":"61295","inReplyTo":null,"subject":"[RFC] Avaiable disk space when automatic garbage collection kicks in","fromName":"Dragan Simic","fromEmail":"dsimic@manjaro.org","sentAt":"2024-04-08T16:29:25Z","receivedAt":"2024-04-08T16:29:27Z","isPatch":false,"sender":{"key":"dsimic@manjaro.org","avatar":null},"body":"Hello all,\n\nA few days ago I've noticed a rather unusual issue, but still\na realistic one.  When automatic garbage collection kicks in,\nas a result of gc.auto >= 0, which is also the default, the\nlocal repository can be left in a rather strange state if there\nisn't enough free space available on the respective filesystem\nfor writing the objects, etc.\n\nIt might be a good idea to estimate the required amount of free\nfilesystem space before starting the garbage collection, be it\nautomatic or manual, and refuse the operation if there isn't\nenough free space available.\n\nAs a note, the need_to_gc() function already does something a bit\nsimilar with the available system RAM.\n\nAny thoughts?\n"},{"id":"497040","messageId":"164fc547afd66caf58019b6c614b5134@manjaro.org","threadId":"61295","inReplyTo":"950d4ad3bcee79df1424faee09eb6b00@manjaro.org","subject":"Re: [RFC] Avaiable disk space when automatic garbage collection kicks in","fromName":"Dragan Simic","fromEmail":"dsimic@manjaro.org","sentAt":"2024-06-12T16:25:11Z","receivedAt":"2024-06-12T16:25:19Z","isPatch":false,"sender":{"key":"dsimic@manjaro.org","avatar":null},"body":"[Maybe this RFC deserves a \"bump\", so let me try.]\n\nOn 2024-04-08 18:29, Dragan Simic wrote:\n> Hello all,\n> \n> A few days ago I've noticed a rather unusual issue, but still\n> a realistic one.  When automatic garbage collection kicks in,\n> as a result of gc.auto >= 0, which is also the default, the\n> local repository can be left in a rather strange state if there\n> isn't enough free space available on the respective filesystem\n> for writing the objects, etc.\n> \n> It might be a good idea to estimate the required amount of free\n> filesystem space before starting the garbage collection, be it\n> automatic or manual, and refuse the operation if there isn't\n> enough free space available.\n> \n> As a note, the need_to_gc() function already does something a bit\n> similar with the available system RAM.\n> \n> Any thoughts?\n"},{"id":"497041","messageId":"123401dabcea$8fe50110$afaf0330$@nexbridge.com","threadId":"61295","inReplyTo":"164fc547afd66caf58019b6c614b5134@manjaro.org","subject":"RE: [RFC] Avaiable disk space when automatic garbage collection kicks in","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2024-06-12T17:04:12Z","receivedAt":"2024-06-12T17:04:34Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On Wednesday, June 12, 2024 12:25 PM, Dragan Simic wrote:\n>[Maybe this RFC deserves a \"bump\", so let me try.]\n>On 2024-04-08 18:29, Dragan Simic wrote:\n>> Hello all,\n>>\n>> A few days ago I've noticed a rather unusual issue, but still a\n>> realistic one.  When automatic garbage collection kicks in, as a\n>> result of gc.auto >= 0, which is also the default, the local\n>> repository can be left in a rather strange state if there isn't enough\n>> free space available on the respective filesystem for writing the\n>> objects, etc.\n>>\n>> It might be a good idea to estimate the required amount of free\n>> filesystem space before starting the garbage collection, be it\n>> automatic or manual, and refuse the operation if there isn't enough\n>> free space available.\n>>\n>> As a note, the need_to_gc() function already does something a bit\n>> similar with the available system RAM.\n>>\n>> Any thoughts?\n\nI am not sure there is a good portable way of reliably doing this using OS\nAPIs, particularly with virtual disks and shared file sets. An edge\ncondition would be setting up a separate file set for content inside .git\nfor massive repositories, so taking an estimate in the working index would\nnot fix the above.\n\nIt might be useful to add a configuration item like: \n\ngc.reserve = size   # possibly with mb, kb, gb, tb, or some other suffix\nindicating how much space must be available to reserve prior to starting the\noperation.\n\nThen creating a file (with real content) inside .git (or .git/objects) with\nthe reserved size. If the file cannot be constructed, gc gets suppressed.\nThis can happen for more than size issue - permissions, for example. Note\nalso that some file systems to not actually allocate the entire space just\nsetting EOF, so that technique, while fast, will also not work portably.\n\nAfter the reserve works, it can be removed (and hopefully NFS will properly\nclose it), providing a lock is put in place, followed by gc running. It\nmight be useful to do this even on a non-auto gc. While this can be\nexpensive (writing a block of stuff twice), it is safer this way.\n\nJust a thought.\n\nRandall.\n\n"},{"id":"497042","messageId":"4951aa9acafcbb723c64ed0df9b5fe41@manjaro.org","threadId":"61295","inReplyTo":"123401dabcea$8fe50110$afaf0330$@nexbridge.com","subject":"Re: [RFC] Avaiable disk space when automatic garbage collection kicks in","fromName":"Dragan Simic","fromEmail":"dsimic@manjaro.org","sentAt":"2024-06-12T17:25:51Z","receivedAt":"2024-06-12T17:25:54Z","isPatch":false,"sender":{"key":"dsimic@manjaro.org","avatar":null},"body":"Hello Randall,\n\nOn 2024-06-12 19:04, rsbecker@nexbridge.com wrote:\n> On Wednesday, June 12, 2024 12:25 PM, Dragan Simic wrote:\n>> [Maybe this RFC deserves a \"bump\", so let me try.]\n>> On 2024-04-08 18:29, Dragan Simic wrote:\n>>> A few days ago I've noticed a rather unusual issue, but still a\n>>> realistic one.  When automatic garbage collection kicks in, as a\n>>> result of gc.auto >= 0, which is also the default, the local\n>>> repository can be left in a rather strange state if there isn't \n>>> enough\n>>> free space available on the respective filesystem for writing the\n>>> objects, etc.\n>>> \n>>> It might be a good idea to estimate the required amount of free\n>>> filesystem space before starting the garbage collection, be it\n>>> automatic or manual, and refuse the operation if there isn't enough\n>>> free space available.\n>>> \n>>> As a note, the need_to_gc() function already does something a bit\n>>> similar with the available system RAM.\n>>> \n>>> Any thoughts?\n> \n> I am not sure there is a good portable way of reliably doing this using \n> OS\n> APIs, particularly with virtual disks and shared file sets. An edge\n> condition would be setting up a separate file set for content inside \n> .git\n> for massive repositories, so taking an estimate in the working index \n> would\n> not fix the above.\n> \n> It might be useful to add a configuration item like:\n> \n> gc.reserve = size   # possibly with mb, kb, gb, tb, or some other \n> suffix\n> indicating how much space must be available to reserve prior to \n> starting the\n> operation.\n> \n> Then creating a file (with real content) inside .git (or .git/objects) \n> with\n> the reserved size. If the file cannot be constructed, gc gets \n> suppressed.\n> This can happen for more than size issue - permissions, for example. \n> Note\n> also that some file systems to not actually allocate the entire space \n> just\n> setting EOF, so that technique, while fast, will also not work \n> portably.\n> \n> After the reserve works, it can be removed (and hopefully NFS will \n> properly\n> close it), providing a lock is put in place, followed by gc running. It\n> might be useful to do this even on a non-auto gc. While this can be\n> expensive (writing a block of stuff twice), it is safer this way.\n\nThanks for your response!\n\nOne of the troubles with the introduction of \"gc.reserve\" is that it \nwould\nbe probably used by advanced users only, which may already turn \nautomatic\ngarbage collection off for their repositories on filesystems without \nenough\nfree space for the garbage collection to succeed.  Another issue is that \nthe\non-disk footprint of large repositories can grow significantly over \ntime,\nso rather frequent updates to the \"gc.reserve\" values would be needed.\n\nThere are aven more issues, which you already mentioned...  One of them \nis\nthe additional time required to create a large file, and another is the\nadditional wear that creating a large temporary file puts on flash-based\nstorage.  Moreover, if the total block usage of an underlying SSD gets \nclose\nto 100% after the large temporary file is created, we'd be putting that \nSSD\nin a rather unfavorable position because no TRIM operation may be \nperformed\non that large file when it gets removed, and we'd then \"hammer\" the SSD\nwith a whole lot of small writes.\n"}]}