{"thread":{"id":"57921","subject":"About GIT Internals","startedAt":"2022-05-25T16:11:14Z","lastAt":"2022-06-06T13:00:52Z","messageCount":17,"participants":["Aman","Emily Shaffer","Erik Cervin Edin","git-vger@eldondev.com","Philip Oakley","Konstantin Khomoutov","Kerry, Richard","Ævar Arnfjörð Bjarmason","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"456104","messageId":"CACMKQb0Mz4zBoSX2CdXkeF51z_mh3had7359J=LmXGzJM1WYLg@mail.gmail.com","threadId":"57921","inReplyTo":null,"subject":"About GIT Internals","fromName":"Aman","fromEmail":"amanmatreja@gmail.com","sentAt":"2022-05-25T16:10:42Z","receivedAt":"2022-05-25T16:11:14Z","isPatch":false,"sender":{"key":"amanmatreja@gmail.com","avatar":null},"body":"Hello there,\n\nI have recently been reading The Architecture for Open Source\nApplications book - and read the chapters dedicated to GIT internals.\nAnd if I am being completely honest, I didn't understand most of it.\n\nCould someone please assist - in sharing some resources - which I\ncould go through, to better understand GIT software internals.\n\n(I am a high school student, and really want to learn more about how\nall the great software and hardware around us work - which so many of\nus take for granted)\n\nRegards,\n"},{"id":"456108","messageId":"CAJoAoZmtzRakZB2Pgez9YW7URXaxTF3RZ2-ZWrRHMhbj81u2yA@mail.gmail.com","threadId":"57921","inReplyTo":"CACMKQb0Mz4zBoSX2CdXkeF51z_mh3had7359J=LmXGzJM1WYLg@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2022-05-25T16:49:29Z","receivedAt":"2022-05-25T16:49:47Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Wed, May 25, 2022 at 9:11 AM Aman <amanmatreja@gmail.com> wrote:\n>\n> Hello there,\n>\n> I have recently been reading The Architecture for Open Source\n> Applications book - and read the chapters dedicated to GIT internals.\n> And if I am being completely honest, I didn't understand most of it.\n>\n> Could someone please assist - in sharing some resources - which I\n> could go through, to better understand GIT software internals.\n\nI am really excited you asked! This puts you firmly on the road to\nbeing the person who can help unstick all your friends when they get\ninto Git messes later on. ;)\n\nhttps://docs.google.com/presentation/d/1IQCRPHEIX-qKo7QFxsD3V62yhyGA9_5YsYXFOiBpgkk/edit?usp=sharing\n<- This is a really great intro to the internals which I love. I\npretty much always recommend it as the place to start for someone\ncurious about learning how Git works.\nhttps://www.youtube.com/watch?v=5Gq3KVvcfDk <- This covers much of the\nsame territory but has a nice video to go through it, in case it's\neasier for you to learn that way instead of reading slides.\n\nIf you have additional questions about the technical design of Git\nfollowing one or both of those presentations above, I think you could\nget far starting with Git's own design documentation:\nhttps://github.com/git/git/tree/master/Documentation/technical\n\nFrom there I think the list will be the best place for specific\nfollowup questions you might have.\n\nHappy learning!\n\n - Emily\n"},{"id":"456121","messageId":"CA+JQ7M_OmfJVJpnYc1h=L226nZqo0A=5USCEJAgrK2EKPRsW4w@mail.gmail.com","threadId":"57921","inReplyTo":"CACMKQb0Mz4zBoSX2CdXkeF51z_mh3had7359J=LmXGzJM1WYLg@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"Erik Cervin Edin","fromEmail":"erik@cervined.in","sentAt":"2022-05-25T21:14:17Z","receivedAt":"2022-05-25T21:15:00Z","isPatch":false,"sender":{"key":"erik@cervined.in","avatar":null},"body":"On Wed, May 25, 2022 at 10:14 PM Aman <amanmatreja@gmail.com> wrote:\n>\n> And if I am being completely honest, I didn't understand most of it.\n\nYou are not alone, there are many that struggle with understanding how\ngit works internally.\n\n> (I am a high school student, and really want to learn more about how\n> all the great software and hardware around us work - which so many of\n> us take for granted)\n\nPerhaps not a good resource, depending on your familiarity with\ncomputer science but\nhttps://eagain.net/articles/git-for-computer-scientists/\nis an article that is often recommended.\n\nI think for me, the hardest part of understanding Git was the\ndifficulty conceptualizing it.\nBut at its core Git is very simple.\n\nYou can think of it as a folder of files that you can \"save\" (commit)\nwhenever you want.\nEach time you \"save\" (commit), all files and folders are \"copied\" to\nanother folder (the local repository).\nThat means that if you ever want to look at a previous version of a\nfile, it's there.\nFor simplicity's sake you can think of this as being unchangeable.\nOnce a file is saved it's saved forever.\n\nJust having a messy pile of every single version of a file is not useful,\nso the rest of git consists of making this manageable.\nFor example by remembering who saved it, when and why (by making them\nwrite a message when they save).\n\nThe main thing however is that Git orders saves.\nThis order is not necessarily one version after another, sorted by\nwhen they were saved.\nInstead, order is manually controlled by saving files in different\nplaces (branches).\nIn its simplest form, a branch is several saves, one after another.\n\nBecause of how Git orders saves, I can work on files, save them and\ngive them to you.\nYou can keep working on those files and make your own saves.\nBut I don't have to wait for you to send your work back to me.\nI can keep working on the same files and making my own saves.\n\nWhen you're done you can put your saves in a \"shared folder\" (a remote\nrepository).\nLater, when I'm done, I can get your saves and Git can help me figure\nout which parts of the files that you changed that I didn't and copy\nboth of our work into new files (merging).\n\nThis is a bit of an oversimplification and Git allows users to do more\nadvanced things but the gist is basically this.\n"},{"id":"456131","messageId":"Yo68+kjAeP6tnduW@invalid","threadId":"57921","inReplyTo":"CACMKQb0Mz4zBoSX2CdXkeF51z_mh3had7359J=LmXGzJM1WYLg@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"","fromEmail":"git-vger@eldondev.com","sentAt":"2022-05-25T23:34:18Z","receivedAt":"2022-05-25T23:50:00Z","isPatch":false,"sender":{"key":"git-vger@eldondev.com","avatar":null},"body":"Hi Aman, responses inline below.\n\nOn Wed, May 25, 2022 at 09:40:42PM +0530, Aman wrote:\n> Could someone please assist - in sharing some resources - which I\n> could go through, to better understand GIT software internals.\n\nThere is an excellent free book at https://git-scm.com/book/en/v2 .\n\nChapter 10 is about git internals. It is important to realize that,\nunlike many other version control systems, git works effectively on\nfiles locally on your computer, without any server or other shared\nresources to manage. Also, one good way to learn may be to form a\nquestion that you want to answer first. \"How do I ....\" or \"what happens\nwhen I ....\". Since git works locally, it is possible to create a git\nrepo, look at the files contained in the .git directory, take action\nwith git, and then look at the files again.\n\nMany people use git from the command line. If you are not familiar with\nthe command line, you may be interesting in learning more about it.\nMozilla, the makers of the Firefox web browser, have a wiki page to\nfamiliarize yourself with the command line here: \nhttps://developer.mozilla.org/en-US/docs/Learn/Tools_and_testing/Understanding_client-side_tools/Command_line\n\nHappy Explorations!\nEldon\n"},{"id":"456154","messageId":"8adba93c-7671-30d8-5a4c-4ad6e1084a22@iee.email","threadId":"57921","inReplyTo":"Yo68+kjAeP6tnduW@invalid","subject":"Re: About GIT Internals","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-05-26T08:47:37Z","receivedAt":"2022-05-26T08:47:45Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 26/05/2022 00:34, git-vger@eldondev.com wrote:\n> Hi Aman, responses inline below.\n>\n> On Wed, May 25, 2022 at 09:40:42PM +0530, Aman wrote:\n>> Could someone please assist - in sharing some resources - which I\n>> could go through, to better understand GIT software internals.\n> There is an excellent free book at https://git-scm.com/book/en/v2 .\n>\n> Chapter 10 is about git internals. It is important to realize that,\n> unlike many other version control systems, git works effectively on\n> files locally on your computer, without any server or other shared\n> resources to manage. Also, one good way to learn may be to form a\n> question that you want to answer first. \"How do I ....\" or \"what happens\n> when I ....\". Since git works locally, it is possible to create a git\n> repo, look at the files contained in the .git directory, take action\n> with git, and then look at the files again.\n>\n>\nAnother Git feature, compared to older version control systems, is that\nit flips the 'control' aspect on its head. (who controls what you can\nstore?)\n\nIt does this by using the hash (sha1, or sha256) values as a way of\nusers _checking_ that they have the right copy of a file or commit,\nrather than needing special permissions to access (write/read) some\nalleged 'master' copy (in the sense of a unique artefact) of the\nparticular version. Maintainers now check and authorise particular\nversions much more easily.\n\nHence Git _Distributes Control_ - you no longer need permission to keep\nversioned copies of your work. This was, in my mind, a core element of\nits success.\n\nThere is other stuff about how Git splits the (file) content from it's\nmeta-data, so if say 10 files contain the same licence text, then it\nonly hold one copy of that text, with its own unique hash. Then has a\nhierarchy (pyramid) of hashes of the meta-data to build up a whole\nproject's hash (the top level 'tree'), and the same hierarchy technique\nis repeated for the project's history of commits.\n\nIf you have a copy of the repository with the latest (same) hash then\nyou have a perfect copy, indistinguishable to the 'original'! Older\nversioning systems did not have those guarantees, many were derived from\nsystems for versioning engineering and architectural drawings such as\nthose that were used for the RMS Titanic or Empire State Building.\n\nPhilip\n\nPS it's worth checking out the distinction between having hash (a magic\nid) of some text, and encrypting (a magic translation of) some text.\n\n\n"},{"id":"456171","messageId":"20220526124547.5txalnlophx7b57l@carbon","threadId":"57921","inReplyTo":"CACMKQb0Mz4zBoSX2CdXkeF51z_mh3had7359J=LmXGzJM1WYLg@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"Konstantin Khomoutov","fromEmail":"kostix@bswap.ru","sentAt":"2022-05-26T12:45:47Z","receivedAt":"2022-05-26T12:46:00Z","isPatch":false,"sender":{"key":"kostix@bswap.ru","avatar":null},"body":"In addition to what others have said, I would recommend to start with \"The Git\nParable\" [1] - which is an ideal gentle, non-technical introduction to the\nconcept of distributed version control systems, - and then read \"Git from the\nBottom Up\" [2] and \"Git for Computer Scientists\" which has already been\nmentioned.\n\n 1. https://tom.preston-werner.com/2009/05/19/the-git-parable.html\n 2. https://jwiegley.github.io/git-from-the-bottom-up/\n\n"},{"id":"456290","messageId":"0201db28-d788-4458-e31d-c6cdedf5c9cf@iee.email","threadId":"57921","inReplyTo":"CACMKQb3exv13sYN5uEP_AG-JYu1rmVj4HDxjdw8_Y-+maJPwGg@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-05-27T14:40:27Z","receivedAt":"2022-05-27T14:40:38Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Aman,\nWe try to keep all the cc's so every one can gain from the learning! \ncomments in-line.\n\nOn 26/05/2022 15:17, Aman wrote:\n> Hey Phillip.\n>\n> Thanks a lot for your email, and for sharing the book! This is great.\nThat was Eldon, thank you.. \n(https://lore.kernel.org/git/Yo68+kjAeP6tnduW@invalid/)\n\nThere is also the Git Magic 'book' from Stanford, with Ch8 covering the \ninternals http://www-cs-students.stanford.edu/~blynn/gitmagic/ch08.html\n\n>\n> Just a  follow up questions- if you don't mind:\n>\n> 1. I haven't had the experience of working with other (perhaps even\n> older) version control systems, like subversion. So when refering to\n> the \"control\" aspect,\n\nThe \"control\" aspect was from whoever was the 'manager' that limited \naccess to the version system (i.e. acting like a museum curator), and \ndeciding if your masterpiece was worthy of inclusion as a significant \nexample of your craft, whether that was an engineering drawing or some \nsoftware code.\n\n>   you mean because with hashes we can verify\n\nIf you have a look at \nhttps://www.makeuk.org/insights/blogs/how-to-read-engineering-drawings-a-simple-guide \nand the part about the Title Block has a drawing (DWG) number \n(EEF-001-AM) that is used to reference it and, while it feels nice, the \nreference is rather arbitrary, (could someone else use that number? \nwhat's the next in the sequence? what happens when we reach EEF-999-AM? \netc.).\n\nSo the computer hash (40 digits of 0-9a-f !) solves all those problems, \nit is unique (>40 card shuffle level), depends only on the content, \ncomputers like it. Yay.\n(Computers are great at perfect replication, so cost of manufacture \ntends to zero! Cost of design wanders in the other direction;-)\n\n> the\n> integrity of the files (like code) in git - there is no need for\n> having a central authority to guarantee that's it's the right content\n> files (which is great)?\n\nAnd it means managers no longer worry about _your_ working copy - \ncomputers have digital storage space to spare. That wasn't the case when \nit was on paper, and we didn't have photocopies - have a look at 'blue \nprints' https://en.wikipedia.org/wiki/Blueprint (see the invention \ndate!, I still remember the smell from the late 1970s)\n>\n> On Thu, May 26, 2022 at 2:17 PM Philip Oakley <philipoakley@iee.email> wrote:\n>> On 26/05/2022 00:34, git-vger@eldondev.com wrote:\n>>> Hi Aman, responses inline below.\n>>>\n>>> On Wed, May 25, 2022 at 09:40:42PM +0530, Aman wrote:\n>>>> Could someone please assist - in sharing some resources - which I\n>>>> could go through, to better understand GIT software internals.\n>>> There is an excellent free book at https://git-scm.com/book/en/v2 .\n>>>\n>>> Chapter 10 is about git internals. It is important to realize that,\n>>> unlike many other version control systems, git works effectively on\n>>> files locally on your computer, without any server or other shared\n>>> resources to manage. Also, one good way to learn may be to form a\n>>> question that you want to answer first. \"How do I ....\" or \"what happens\n>>> when I ....\". Since git works locally, it is possible to create a git\n>>> repo, look at the files contained in the .git directory, take action\n>>> with git, and then look at the files again.\n>>>\n>>>\n>> Another Git feature, compared to older version control systems, is that\n>> it flips the 'control' aspect on its head. (who controls what you can\n>> store?)\n>>\n>> It does this by using the hash (sha1, or sha256) values as a way of\n>> users _checking_ that they have the right copy of a file or commit,\n>> rather than needing special permissions to access (write/read) some\n>> alleged 'master' copy (in the sense of a unique artefact) of the\n>> particular version. Maintainers now check and authorise particular\n>> versions much more easily.\n>>\n>> Hence Git _Distributes Control_ - you no longer need permission to keep\n>> versioned copies of your work. This was, in my mind, a core element of\n>> its success.\n>>\n>> There is other stuff about how Git splits the (file) content from it's\n>> meta-data, so if say 10 files contain the same licence text, then it\n>> only hold one copy of that text, with its own unique hash. Then has a\n>> hierarchy (pyramid) of hashes of the meta-data to build up a whole\n>> project's hash (the top level 'tree'), and the same hierarchy technique\n>> is repeated for the project's history of commits.\n>>\n>> If you have a copy of the repository with the latest (same) hash then\n>> you have a perfect copy, indistinguishable to the 'original'! Older\n>> versioning systems did not have those guarantees, many were derived from\n>> systems for versioning engineering and architectural drawings such as\n>> those that were used for the RMS Titanic or Empire State Building.\n>>\n>> Philip\n>>\n>> PS it's worth checking out the distinction between having hash (a magic\n>> id) of some text, and encrypting (a magic translation of) some text.\n>>\n>>\n\n"},{"id":"456295","messageId":"b269bafc-0a9e-e676-6fb7-10c39212985b@iee.email","threadId":"57921","inReplyTo":"C4B1A93D-800F-4C49-93D5-86FE58B1DDCA@hxcore.ol","subject":"Re: About GIT Internals","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-05-27T15:14:01Z","receivedAt":"2022-05-27T15:14:09Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 27/05/2022 16:01, Aman wrote:\n>\n> Hey, thank you again.\n>\n> I am finding this mailing list format of talking a bit confusing, sorry.\n>\nNo problem, blame Microsoft for following the business ($$$) way of \ndoing stuff.\n\nThe 'plain text, in-line replies, with trimming of unrelated items' \nstyle helps the lurkers, and folks who come to the discussion later.\n\nMain point is that we convert each point of interest into its own \ndiscussion, rather than it being a big challenge-response style between \nlegal negotiators - it's not a win/lose discussion ;-)\n\n\n> Would there be any way address everyone on the mailing list – like in \n> the future – to continue this conversation about git internals?\n>\nKey method is to locate the \"reply All\" option in your mail app. That \nmakes sure every one in the discussion is copied, and the mailing lists \nas well for all the 'lurkers';-)\n\nThe mailing list archive uses both the message titles and any \nin-reply-to hidden headers (see if your mailer has a 'show source' \noption to see those interesting bits) to organise the list archive.\n>\n> I found out the mailing list archive (so confusing)– and saw these \n> personal replies don’t get added in the thread. I would appreciate if \n> you give some advice, thank you\n>\n> Sent from Mail <https://go.microsoft.com/fwlink/?LinkId=550986> for \n> Windows\n>\n> From: Philip Oakley <mailto:philipoakley@iee.email>\n> Sent: 27 May 2022 08:10 PM\n> To: Aman <mailto:amanmatreja@gmail.com>\n> Cc: Git List <mailto:git@vger.kernel.org>; git-vger@eldondev.com\n> Subject: Re: About GIT Internals\n>\n> Hi Aman,\n>\n> We try to keep all the cc's so every one can gain from the learning!\n>\n> comments in-line.\n>\n> On 26/05/2022 15:17, Aman wrote:\n>\n> > Hey Phillip.\n>\n> >\n>\n> > Thanks a lot for your email, and for sharing the book! This is great.\n>\n> That was Eldon, thank you..\n>\n> (https://lore.kernel.org/git/Yo68+kjAeP6tnduW@invalid/)\n>\n> There is also the Git Magic 'book' from Stanford, with Ch8 covering the\n>\n> internals http://www-cs-students.stanford.edu/~blynn/gitmagic/ch08.html\n>\n> >\n>\n> > Just a  follow up questions- if you don't mind:\n>\n> >\n>\n> > 1. I haven't had the experience of working with other (perhaps even\n>\n> > older) version control systems, like subversion. So when refering to\n>\n> > the \"control\" aspect,\n>\n> The \"control\" aspect was from whoever was the 'manager' that limited\n>\n> access to the version system (i.e. acting like a museum curator), and\n>\n> deciding if your masterpiece was worthy of inclusion as a significant\n>\n> example of your craft, whether that was an engineering drawing or some\n>\n> software code.\n>\n> >   you mean because with hashes we can verify\n>\n> If you have a look at\n>\n> https://www.makeuk.org/insights/blogs/how-to-read-engineering-drawings-a-simple-guide \n>\n>\n> and the part about the Title Block has a drawing (DWG) number\n>\n> (EEF-001-AM) that is used to reference it and, while it feels nice, the\n>\n> reference is rather arbitrary, (could someone else use that number?\n>\n> what's the next in the sequence? what happens when we reach EEF-999-AM?\n>\n> etc.).\n>\n> So the computer hash (40 digits of 0-9a-f !) solves all those problems,\n>\n> it is unique (>40 card shuffle level), depends only on the content,\n>\n> computers like it. Yay.\n>\n> (Computers are great at perfect replication, so cost of manufacture\n>\n> tends to zero! Cost of design wanders in the other direction;-)\n>\n> > the\n>\n> > integrity of the files (like code) in git - there is no need for\n>\n> > having a central authority to guarantee that's it's the right content\n>\n> > files (which is great)?\n>\n> And it means managers no longer worry about _your_ working copy -\n>\n> computers have digital storage space to spare. That wasn't the case when\n>\n> it was on paper, and we didn't have photocopies - have a look at 'blue\n>\n> prints' https://en.wikipedia.org/wiki/Blueprint (see the invention\n>\n> date!, I still remember the smell from the late 1970s)\n>\n> >\n>\n> > On Thu, May 26, 2022 at 2:17 PM Philip Oakley \n> <philipoakley@iee.email> wrote:\n>\n> >> On 26/05/2022 00:34, git-vger@eldondev.com wrote:\n>\n> >>> Hi Aman, responses inline below.\n>\n> >>>\n>\n> >>> On Wed, May 25, 2022 at 09:40:42PM +0530, Aman wrote:\n>\n> >>>> Could someone please assist - in sharing some resources - which I\n>\n> >>>> could go through, to better understand GIT software internals.\n>\n> >>> There is an excellent free book at https://git-scm.com/book/en/v2 .\n>\n> >>>\n>\n> >>> Chapter 10 is about git internals. It is important to realize that,\n>\n> >>> unlike many other version control systems, git works effectively on\n>\n> >>> files locally on your computer, without any server or other shared\n>\n> >>> resources to manage. Also, one good way to learn may be to form a\n>\n> >>> question that you want to answer first. \"How do I ....\" or \"what \n> happens\n>\n> >>> when I ...\". Since git works locally, it is possible to create a git\n>\n> >>> repo, look at the files contained in the .git directory, take action\n>\n> >>> with git, and then look at the files again.\n>\n> >>>\n>\n> >>>\n>\n> >> Another Git feature, compared to older version control systems, is that\n>\n> >> it flips the 'control' aspect on its head. (who controls what you can\n>\n> >> store?)\n>\n> >>\n>\n> >> It does this by using the hash (sha1, or sha256) values as a way of\n>\n> >> users _checking_ that they have the right copy of a file or commit,\n>\n> >> rather than needing special permissions to access (write/read) some\n>\n> >> alleged 'master' copy (in the sense of a unique artefact) of the\n>\n> >> particular version. Maintainers now check and authorise particular\n>\n> >> versions much more easily.\n>\n> >>\n>\n> >> Hence Git _Distributes Control_ - you no longer need permission to keep\n>\n> >> versioned copies of your work. This was, in my mind, a core element of\n>\n> >> its success.\n>\n> >>\n>\n> >> There is other stuff about how Git splits the (file) content from it's\n>\n> >> meta-data, so if say 10 files contain the same licence text, then it\n>\n> >> only hold one copy of that text, with its own unique hash. Then has a\n>\n> >> hierarchy (pyramid) of hashes of the meta-data to build up a whole\n>\n> >> project's hash (the top level 'tree'), and the same hierarchy technique\n>\n> >> is repeated for the project's history of commits.\n>\n> >>\n>\n> >> If you have a copy of the repository with the latest (same) hash then\n>\n> >> you have a perfect copy, indistinguishable to the 'original'! Older\n>\n> >> versioning systems did not have those guarantees, many were derived \n> from\n>\n> >> systems for versioning engineering and architectural drawings such as\n>\n> >> those that were used for the RMS Titanic or Empire State Building.\n>\n> >>\n>\n> >> Philip\n>\n> >>\n>\n> >> PS it's worth checking out the distinction between having hash (a magic\n>\n> >> id) of some text, and encrypting (a magic translation of) some text.\n>\n> >>\n>\n> >>\n>\n\n"},{"id":"456365","messageId":"AS8PR02MB730274D473C2BC3846D9FA3F9CDD9@AS8PR02MB7302.eurprd02.prod.outlook.com","threadId":"57921","inReplyTo":"0201db28-d788-4458-e31d-c6cdedf5c9cf@iee.email","subject":"RE: About GIT Internals","fromName":"Kerry, Richard","fromEmail":"richard.kerry@atos.net","sentAt":"2022-05-30T09:49:57Z","receivedAt":"2022-05-30T09:50:36Z","isPatch":false,"sender":{"key":"richard.kerry@atos.net","avatar":null},"body":"\n\n> -----Original Message-----\n> From: Philip Oakley <philipoakley@iee.email>\n> Sent: 27 May 2022 15:40\n> To: Aman <amanmatreja@gmail.com>\n> Cc: Git List <git@vger.kernel.org>; git-vger@eldondev.com\n> Subject: Re: About GIT Internals\n> \n> > Just a  follow up questions- if you don't mind:\n> >\n> > 1. I haven't had the experience of working with other (perhaps even\n> > older) version control systems, like subversion. So when refering to\n> > the \"control\" aspect,\n> \n> The \"control\" aspect was from whoever was the 'manager' that limited\n> access to the version system (i.e. acting like a museum curator), and deciding\n> if your masterpiece was worthy of inclusion as a significant example of your\n> craft, whether that was an engineering drawing or some software code.\n\nI'm not sure I get that idea.  I worked using server-based Version Control systems from the mid 80s until about 5 years ago when the team moved from Subversion to Git.  There was never a \"curator\" who controlled what went into VC.  You did your work, developed files, and committed when you thought it necessary.  When a build was to be done there would then be some consideration of what from VC would go into the build.\nThat is all still there nowadays using a distributed system (ie Git).  Those doing Open source work might operate a bit differently, as there is of necessity distribution of control of what gets into a release. But those of us who are developing proprietary software are still going through the same sort of release process.  And that's even if there isn't actually a separate person actively manipulating the contents of a release, it's just up to you to do what's necessary (actually there are others involved in dividing what will be in, but in our case they don't actively manipulate a repository).\n\n\n> >>> Chapter 10 is about git internals. It is important to realize that,\n> >>> unlike many other version control systems, git works effectively on\n> >>> files locally on your computer, without any server or other shared\n> >>> resources to manage. Also, one good way to learn may be to form a\n> >>> question that you want to answer first. \"How do I ....\" or \"what\n> >>> happens when I ....\". Since git works locally, it is possible to\n> >>> create a git repo, look at the files contained in the .git\n> >>> directory, take action with git, and then look at the files again.\n> >>>\n> >>>\n> >> Another Git feature, compared to older version control systems, is\n> >> that it flips the 'control' aspect on its head. (who controls what\n> >> you can\n> >> store?)\n\nAgain, I don't really recognize that.  You store what you want, probably with some sort of arrangement with the others on the team.  The important bit is determining what will go into the release.  Ie in choosing what, from everything that is stored, will be released.\n\n> >> Hence Git _Distributes Control_ - you no longer need permission to\n> >> keep versioned copies of your work. This was, in my mind, a core\n> >> element of its success.\n\nMaybe you do.  If you're working with others there will probably be \"permission\" in some sense involved.  I can store what I like locally, but then I miss out on some protection of my work, against a technical fault locally that might cause a loss of the whole repository.  If there is a remote server then I am probably only allowed to store company work to the company server.\n\nA lot of this discussion seems to be more about the differences between the nature of Git and its client-server rivals.  I thought the original query was about how its internals worked, which would seem to be a slightly different question.\n\nRegards,\nRichard.\n(Not old enough to remember the smell of blue prints, but old enough to know of the term)\n\n\n\n"},{"id":"456367","messageId":"20220530115339.3torgv5c2zw75okg@carbon","threadId":"57921","inReplyTo":"AS8PR02MB730274D473C2BC3846D9FA3F9CDD9@AS8PR02MB7302.eurprd02.prod.outlook.com","subject":"Re: About GIT Internals","fromName":"Konstantin Khomoutov","fromEmail":"kostix@bswap.ru","sentAt":"2022-05-30T11:53:39Z","receivedAt":"2022-05-30T11:53:56Z","isPatch":false,"sender":{"key":"kostix@bswap.ru","avatar":null},"body":"On Mon, May 30, 2022 at 09:49:57AM +0000, Kerry, Richard wrote:\n\n[...]\n> > > 1. I haven't had the experience of working with other (perhaps even\n> > > older) version control systems, like subversion. So when refering to\n> > > the \"control\" aspect,\n> > \n> > The \"control\" aspect was from whoever was the 'manager' that limited\n> > access to the version system (i.e. acting like a museum curator), and deciding\n> > if your masterpiece was worthy of inclusion as a significant example of your\n> > craft, whether that was an engineering drawing or some software code.\n> \n> I'm not sure I get that idea.  I worked using server-based Version Control\n> systems from the mid 80s until about 5 years ago when the team moved from\n> Subversion to Git.  There was never a \"curator\" who controlled what went\n> into VC.  You did your work, developed files, and committed when you thought\n> it necessary.  When a build was to be done there would then be some\n> consideration of what from VC would go into the build. That is all still\n> there nowadays using a distributed system (ie Git).  Those doing Open source\n> work might operate a bit differently, as there is of necessity distribution\n> of control of what gets into a release. But those of us who are developing\n> proprietary software are still going through the same sort of release\n> process.  And that's even if there isn't actually a separate person actively\n> manipulating the contents of a release, it's just up to you to do what's\n> necessary (actually there are others involved in dividing what will be in,\n> but in our case they don't actively manipulate a repository).\n\nI think, the \"inversion of control\" brought in by DVCS-es about a bit\ndifferet set of things.\n\nI would say it is connected to F/OSS and the way most projects have been\nhosted before the DVCS-es over: usually each project had a single repository\n(say, on Sourceforge or elsewhere), and it was \"truly central\" in the sense\nthat if anyone were to decide to work on that project, they would need to\ncontact whoever were in charge of that project and ask them to set up\npermissions allowing commits - may be not to \"the trunk\", but anyway the\ncommit access was required because in centralized VCS commits are made on the\nserver side.\n(Of course, there were projects where you could mail your patchset to a\nmaintainer, but maintaining such patchset was not convenient: you would either\nneed to host your own fully private VCS or use a tool like Quilt [1].\nAlso note that certain high-profile projects such as Linux and Git use mailing\nlists for submission and review of patch series; this workflow coexists with\nthe concept of DVCS just fine.)\n\nThis approach has been effectively reversed by what was a killer-feature of\nGithub (I honestly am not sure whether Github was the first to implement it\nbut it was, and arguably is, the most popular): a network of \"forks\".\nIf a project is hosted using a DVCS, anyone is free to clone it and push their\nwork _elsewhere._ This point is crucial: you do not need to ask the project\nmaintainers to publish your modifications. Github pushed this concept quite\nfar: creating a fork and pushing your work there is actually a device to create\na pull request - a request to incorporate your changes into the original\nproject. While this approach has obvious upsides, it also has possible\ndownsides; one of a more visible is that when an original project becomes\ndormant for some reason, its users might have hard time understanding which\none of competing forks to switch to, and there are cases when multiple\ncompeting forks implement different features and bugfixes, in parallel.\nOne of the guys behind Subversion expressed his concerns about this back then\nwgen Git was in its relative infancy [2].\n\n 1. https://en.wikipedia.org/wiki/Quilt_(software)\n 2. http://blog.red-bean.com/sussman/?p=20\n\n"},{"id":"456370","messageId":"220530.8635gr2jsh.gmgdl@evledraar.gmail.com","threadId":"57921","inReplyTo":"20220530115339.3torgv5c2zw75okg@carbon","subject":"Re: About GIT Internals","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-05-30T13:50:54Z","receivedAt":"2022-05-30T15:13:26Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, May 30 2022, Konstantin Khomoutov wrote:\n\n> On Mon, May 30, 2022 at 09:49:57AM +0000, Kerry, Richard wrote:\n>\n> [...]\n>> > > 1. I haven't had the experience of working with other (perhaps even\n>> > > older) version control systems, like subversion. So when refering to\n>> > > the \"control\" aspect,\n>> > \n>> > The \"control\" aspect was from whoever was the 'manager' that limited\n>> > access to the version system (i.e. acting like a museum curator), and deciding\n>> > if your masterpiece was worthy of inclusion as a significant example of your\n>> > craft, whether that was an engineering drawing or some software code.\n>> \n>> I'm not sure I get that idea.  I worked using server-based Version Control\n>> systems from the mid 80s until about 5 years ago when the team moved from\n>> Subversion to Git.  There was never a \"curator\" who controlled what went\n>> into VC.  You did your work, developed files, and committed when you thought\n>> it necessary.  When a build was to be done there would then be some\n>> consideration of what from VC would go into the build. That is all still\n>> there nowadays using a distributed system (ie Git).  Those doing Open source\n>> work might operate a bit differently, as there is of necessity distribution\n>> of control of what gets into a release. But those of us who are developing\n>> proprietary software are still going through the same sort of release\n>> process.  And that's even if there isn't actually a separate person actively\n>> manipulating the contents of a release, it's just up to you to do what's\n>> necessary (actually there are others involved in dividing what will be in,\n>> but in our case they don't actively manipulate a repository).\n>\n> I think, the \"inversion of control\" brought in by DVCS-es about a bit\n> differet set of things.\n\nRe the \"I'm not sure I get that idea\" from Richard I think his point\nstands that some of the stories we carry around about the VCS v.s. DVCS\nin free/open source software was more particular to how things were done\nin those online communities, and not really about the implicit\nconstraints of centralized VCS per-se.\n\nPartly those two mix: It was quite common for free software projects not\nto have any public VCS (usually CVS) access at all, some did, but it was\nquite a hassle to set up, and not part of your \"normal\" workflow (as\nopposed setting up a hoster git repository, which everyone uses) that\nmany just didn't do it.\n\n> I would say it is connected to F/OSS and the way most projects have been\n> hosted before the DVCS-es over: usually each project had a single repository\n> (say, on Sourceforge or elsewhere), and it was \"truly central\" in the sense\n> that if anyone were to decide to work on that project, they would need to\n> contact whoever were in charge of that project and ask them to set up\n> permissions allowing commits - may be not to \"the trunk\", but anyway the\n> commit access was required because in centralized VCS commits are made on the\n> server side.\n\nWe may have tried this in different eras, but from what I recall it was\na crapshoot whether there was any public VCS access at all. Some\nprojects were quite good about it, and sourceforge managed to push that\nto more of them early on by making anonymous CVS access something you\ncould get by default.\n\nBut a lot of projects simply didn't have it at all, you'll still find\nsome of them today, i.e. various bits of \"infrastructure\" code that the\nmaintainers are (presumably) still manually managing with zip snapshots\nand manually applied patches.\n\n> (Of course, there were projects where you could mail your patchset to a\n> maintainer, but maintaining such patchset was not convenient: you would either\n> need to host your own fully private VCS or use a tool like Quilt [1].\n> Also note that certain high-profile projects such as Linux and Git use mailing\n> lists for submission and review of patch series; this workflow coexists with\n> the concept of DVCS just fine.)\n\nI'd add though that this isn't really \"co-existing\" with DVSC so much as\nusing patches on a ML as an indirect transport protocol for \"git push\".\n\nI.e. if you contributed to some similar projects \"back in the day\" you\ncould expect to effectively send your patche into a black-hole until the\nnext release, the maintainer would apply them locally, you wouldn't be\nable to pull them back down via the DVCS.\n\nPerhaps there would be development releases, but those could be weeks or\neven months apart, and a \"real\" release might be once every 1-2 years.\n\nWhereas both Junio and Linus (and other linux maintainers) publish their\nversion of the patches they do integrate fairly quickly.\n\n> [...] it also has possible\n> downsides; one of a more visible is that when an original project becomes\n> dormant for some reason, its users might have hard time understanding which\n> one of competing forks to switch to, and there are cases when multiple\n> competing forks implement different features and bugfixes, in parallel.\n> One of the guys behind Subversion expressed his concerns about this back then\n> wgen Git was in its relative infancy [2].\n>\n>  1. https://en.wikipedia.org/wiki/Quilt_(software)\n>  2. http://blog.red-bean.com/sussman/?p=20\n\nIt's interesting that this aspect of what proponents of centralized VCS\nwere fearful of when it came to DVCS turned out to be the exact\nopposite:\n\n    Notice what this user is now able to do: he wants to to crawl off\n    into a cave, work for weeks on a complex feature by himself, then\n    present it as a polished result to the main codebase. And this is\n    exactly the sort of behavior that I think is bad for open source\n    communities.\n\nI.e. lowering the cost to publish early and often has had the effect\nthat people are less likely to \"crawl off into a cave\" and work on\nsomething for a long time without syncing up with other parallel\ndevelopment.\n"},{"id":"456571","messageId":"CACMKQb3_j+iFcf5trZEcWoU7vAsscKv+_sLaEqg_qfazBPTo+Q@mail.gmail.com","threadId":"57921","inReplyTo":"220530.8635gr2jsh.gmgdl@evledraar.gmail.com","subject":"Re: About GIT Internals","fromName":"Aman","fromEmail":"amanmatreja@gmail.com","sentAt":"2022-06-03T12:18:14Z","receivedAt":"2022-06-03T12:18:32Z","isPatch":false,"sender":{"key":"amanmatreja@gmail.com","avatar":null},"body":"Hello everyone. I sent out an email here last week, asking for a list\nof resources, so I could better understand the workings and design of\ngit. I really appreciate everyone, who gave the links and their\nadvice.\n\nI have been reading about GIT for some time now, and have looked at\nalmost all of the resources plus some others. I think I could say, I\nnow have a decent conceptual understanding of how GIT  works\ninternally.\n\n(Also, I understood the chapter about git I read in the book I am\nreading, Architecture of Open Source Applications: Volume 2, which I\ndidn't understand at all, the reason I started this thread). Although\nthere must definitely be a lot of details and subtle things I may not\nunderstand yet (like branches are nothing but pointers to commits,\nwow! btw)\n\nNow, continuing this discussion, and talking about the implementation\nand engineering side of things, I wanted to ask another question and\nhence wanted some advice.\n\nThough I may understand the internal design and high-level\nimplementation of GIT, I really want to know how it's implemented and\nwas made, which means reading the SOURCE CODE.\n\n1. I don't know how absurd of a quest this is, please enlighten me.\n2. How do I do it? Where do I start? It's such a BIG repository - and\nI am not guessing it's going to be easy.\n3. Would someone advise, perhaps, to have a look at an older version\nof the source code? rather than the latest one, for some reason.\n\n\nAgain, I would really appreciate it if someone could give their\nthoughts on this.\n\nThank you,\n\nRegards,\nAman\n\n\nOn Mon, May 30, 2022 at 7:40 PM Ævar Arnfjörð Bjarmason\n<avarab@gmail.com> wrote:\n>\n>\n> On Mon, May 30 2022, Konstantin Khomoutov wrote:\n>\n> > On Mon, May 30, 2022 at 09:49:57AM +0000, Kerry, Richard wrote:\n> >\n> > [...]\n> >> > > 1. I haven't had the experience of working with other (perhaps even\n> >> > > older) version control systems, like subversion. So when refering to\n> >> > > the \"control\" aspect,\n> >> >\n> >> > The \"control\" aspect was from whoever was the 'manager' that limited\n> >> > access to the version system (i.e. acting like a museum curator), and deciding\n> >> > if your masterpiece was worthy of inclusion as a significant example of your\n> >> > craft, whether that was an engineering drawing or some software code.\n> >>\n> >> I'm not sure I get that idea.  I worked using server-based Version Control\n> >> systems from the mid 80s until about 5 years ago when the team moved from\n> >> Subversion to Git.  There was never a \"curator\" who controlled what went\n> >> into VC.  You did your work, developed files, and committed when you thought\n> >> it necessary.  When a build was to be done there would then be some\n> >> consideration of what from VC would go into the build. That is all still\n> >> there nowadays using a distributed system (ie Git).  Those doing Open source\n> >> work might operate a bit differently, as there is of necessity distribution\n> >> of control of what gets into a release. But those of us who are developing\n> >> proprietary software are still going through the same sort of release\n> >> process.  And that's even if there isn't actually a separate person actively\n> >> manipulating the contents of a release, it's just up to you to do what's\n> >> necessary (actually there are others involved in dividing what will be in,\n> >> but in our case they don't actively manipulate a repository).\n> >\n> > I think, the \"inversion of control\" brought in by DVCS-es about a bit\n> > differet set of things.\n>\n> Re the \"I'm not sure I get that idea\" from Richard I think his point\n> stands that some of the stories we carry around about the VCS v.s. DVCS\n> in free/open source software was more particular to how things were done\n> in those online communities, and not really about the implicit\n> constraints of centralized VCS per-se.\n>\n> Partly those two mix: It was quite common for free software projects not\n> to have any public VCS (usually CVS) access at all, some did, but it was\n> quite a hassle to set up, and not part of your \"normal\" workflow (as\n> opposed setting up a hoster git repository, which everyone uses) that\n> many just didn't do it.\n>\n> > I would say it is connected to F/OSS and the way most projects have been\n> > hosted before the DVCS-es over: usually each project had a single repository\n> > (say, on Sourceforge or elsewhere), and it was \"truly central\" in the sense\n> > that if anyone were to decide to work on that project, they would need to\n> > contact whoever were in charge of that project and ask them to set up\n> > permissions allowing commits - may be not to \"the trunk\", but anyway the\n> > commit access was required because in centralized VCS commits are made on the\n> > server side.\n>\n> We may have tried this in different eras, but from what I recall it was\n> a crapshoot whether there was any public VCS access at all. Some\n> projects were quite good about it, and sourceforge managed to push that\n> to more of them early on by making anonymous CVS access something you\n> could get by default.\n>\n> But a lot of projects simply didn't have it at all, you'll still find\n> some of them today, i.e. various bits of \"infrastructure\" code that the\n> maintainers are (presumably) still manually managing with zip snapshots\n> and manually applied patches.\n>\n> > (Of course, there were projects where you could mail your patchset to a\n> > maintainer, but maintaining such patchset was not convenient: you would either\n> > need to host your own fully private VCS or use a tool like Quilt [1].\n> > Also note that certain high-profile projects such as Linux and Git use mailing\n> > lists for submission and review of patch series; this workflow coexists with\n> > the concept of DVCS just fine.)\n>\n> I'd add though that this isn't really \"co-existing\" with DVSC so much as\n> using patches on a ML as an indirect transport protocol for \"git push\".\n>\n> I.e. if you contributed to some similar projects \"back in the day\" you\n> could expect to effectively send your patche into a black-hole until the\n> next release, the maintainer would apply them locally, you wouldn't be\n> able to pull them back down via the DVCS.\n>\n> Perhaps there would be development releases, but those could be weeks or\n> even months apart, and a \"real\" release might be once every 1-2 years.\n>\n> Whereas both Junio and Linus (and other linux maintainers) publish their\n> version of the patches they do integrate fairly quickly.\n>\n> > [...] it also has possible\n> > downsides; one of a more visible is that when an original project becomes\n> > dormant for some reason, its users might have hard time understanding which\n> > one of competing forks to switch to, and there are cases when multiple\n> > competing forks implement different features and bugfixes, in parallel.\n> > One of the guys behind Subversion expressed his concerns about this back then\n> > wgen Git was in its relative infancy [2].\n> >\n> >  1. https://en.wikipedia.org/wiki/Quilt_(software)\n> >  2. http://blog.red-bean.com/sussman/?p=20\n>\n> It's interesting that this aspect of what proponents of centralized VCS\n> were fearful of when it came to DVCS turned out to be the exact\n> opposite:\n>\n>     Notice what this user is now able to do: he wants to to crawl off\n>     into a cave, work for weeks on a complex feature by himself, then\n>     present it as a polished result to the main codebase. And this is\n>     exactly the sort of behavior that I think is bad for open source\n>     communities.\n>\n> I.e. lowering the cost to publish early and often has had the effect\n> that people are less likely to \"crawl off into a cave\" and work on\n> something for a long time without syncing up with other parallel\n> development.\n"},{"id":"456587","messageId":"CAJoAoZmPOOp41KF=V4EhUopu+P8=UW55bkUJm6ZYiKziprtWug@mail.gmail.com","threadId":"57921","inReplyTo":"CACMKQb3_j+iFcf5trZEcWoU7vAsscKv+_sLaEqg_qfazBPTo+Q@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"Emily Shaffer","fromEmail":"emilyshaffer@google.com","sentAt":"2022-06-03T15:25:24Z","receivedAt":"2022-06-03T15:25:52Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Fri, Jun 3, 2022 at 5:21 AM Aman <amanmatreja@gmail.com> wrote:\n>\n> Hello everyone. I sent out an email here last week, asking for a list\n> of resources, so I could better understand the workings and design of\n> git. I really appreciate everyone, who gave the links and their\n> advice.\n>\n> I have been reading about GIT for some time now, and have looked at\n> almost all of the resources plus some others. I think I could say, I\n> now have a decent conceptual understanding of how GIT  works\n> internally.\n>\n> (Also, I understood the chapter about git I read in the book I am\n> reading, Architecture of Open Source Applications: Volume 2, which I\n> didn't understand at all, the reason I started this thread). Although\n> there must definitely be a lot of details and subtle things I may not\n> understand yet (like branches are nothing but pointers to commits,\n> wow! btw)\n>\n> Now, continuing this discussion, and talking about the implementation\n> and engineering side of things, I wanted to ask another question and\n> hence wanted some advice.\n>\n> Though I may understand the internal design and high-level\n> implementation of GIT, I really want to know how it's implemented and\n> was made, which means reading the SOURCE CODE.\n>\n> 1. I don't know how absurd of a quest this is, please enlighten me.\n\nIt's a lot :) But I don't think that should discourage you.\n\n> 2. How do I do it? Where do I start? It's such a BIG repository - and\n> I am not guessing it's going to be easy.\n\nI would start actually with \"Documentation/MyFirstContribution.txt\"\nand \"Documentation/MyFirstRevisionWalk.txt\" - but I am biased towards\nthose documents. ;) The other subtle hint I would give is that the\nentry point for almost every command is at a function called\n\"cmd_cmdname()\", so for example \"git status\" is at \"cmd_status()\",\nusually somewhere in 'builtin/'.\n\n> 3. Would someone advise, perhaps, to have a look at an older version\n> of the source code? rather than the latest one, for some reason.\n\nSome other piece of the developer documentation (maybe\n\"SubmittingPatches\"?) suggests that you start from the initial commit\nand understand that part first. I personally don't find this exercise\nvery useful anymore as Git has grown quite a lot since then (and is\neven primarily in a different language, although we still have some\nbash scripts here and there).\n\n> Again, I would really appreciate it if someone could give their\n> thoughts on this.\n\nIn your journeys, also watch out for some libraries in common, like\ncalls from \"run-command.h\" or \"parse-opt.h\", to help you understand\nhow we make stuff work more or less consistently across the codebase,\nor libraries like \"strbuf.h\" and \"string-list.h\" to understand some of\nthe things that we do to make working with C a little less fraught.\n\n>\n> Thank you,\n>\n> Regards,\n> Aman\n>\n>\n> On Mon, May 30, 2022 at 7:40 PM Ævar Arnfjörð Bjarmason\n> <avarab@gmail.com> wrote:\n> >\n> >\n> > On Mon, May 30 2022, Konstantin Khomoutov wrote:\n> >\n> > > On Mon, May 30, 2022 at 09:49:57AM +0000, Kerry, Richard wrote:\n> > >\n> > > [...]\n> > >> > > 1. I haven't had the experience of working with other (perhaps even\n> > >> > > older) version control systems, like subversion. So when refering to\n> > >> > > the \"control\" aspect,\n> > >> >\n> > >> > The \"control\" aspect was from whoever was the 'manager' that limited\n> > >> > access to the version system (i.e. acting like a museum curator), and deciding\n> > >> > if your masterpiece was worthy of inclusion as a significant example of your\n> > >> > craft, whether that was an engineering drawing or some software code.\n> > >>\n> > >> I'm not sure I get that idea.  I worked using server-based Version Control\n> > >> systems from the mid 80s until about 5 years ago when the team moved from\n> > >> Subversion to Git.  There was never a \"curator\" who controlled what went\n> > >> into VC.  You did your work, developed files, and committed when you thought\n> > >> it necessary.  When a build was to be done there would then be some\n> > >> consideration of what from VC would go into the build. That is all still\n> > >> there nowadays using a distributed system (ie Git).  Those doing Open source\n> > >> work might operate a bit differently, as there is of necessity distribution\n> > >> of control of what gets into a release. But those of us who are developing\n> > >> proprietary software are still going through the same sort of release\n> > >> process.  And that's even if there isn't actually a separate person actively\n> > >> manipulating the contents of a release, it's just up to you to do what's\n> > >> necessary (actually there are others involved in dividing what will be in,\n> > >> but in our case they don't actively manipulate a repository).\n> > >\n> > > I think, the \"inversion of control\" brought in by DVCS-es about a bit\n> > > differet set of things.\n> >\n> > Re the \"I'm not sure I get that idea\" from Richard I think his point\n> > stands that some of the stories we carry around about the VCS v.s. DVCS\n> > in free/open source software was more particular to how things were done\n> > in those online communities, and not really about the implicit\n> > constraints of centralized VCS per-se.\n> >\n> > Partly those two mix: It was quite common for free software projects not\n> > to have any public VCS (usually CVS) access at all, some did, but it was\n> > quite a hassle to set up, and not part of your \"normal\" workflow (as\n> > opposed setting up a hoster git repository, which everyone uses) that\n> > many just didn't do it.\n> >\n> > > I would say it is connected to F/OSS and the way most projects have been\n> > > hosted before the DVCS-es over: usually each project had a single repository\n> > > (say, on Sourceforge or elsewhere), and it was \"truly central\" in the sense\n> > > that if anyone were to decide to work on that project, they would need to\n> > > contact whoever were in charge of that project and ask them to set up\n> > > permissions allowing commits - may be not to \"the trunk\", but anyway the\n> > > commit access was required because in centralized VCS commits are made on the\n> > > server side.\n> >\n> > We may have tried this in different eras, but from what I recall it was\n> > a crapshoot whether there was any public VCS access at all. Some\n> > projects were quite good about it, and sourceforge managed to push that\n> > to more of them early on by making anonymous CVS access something you\n> > could get by default.\n> >\n> > But a lot of projects simply didn't have it at all, you'll still find\n> > some of them today, i.e. various bits of \"infrastructure\" code that the\n> > maintainers are (presumably) still manually managing with zip snapshots\n> > and manually applied patches.\n> >\n> > > (Of course, there were projects where you could mail your patchset to a\n> > > maintainer, but maintaining such patchset was not convenient: you would either\n> > > need to host your own fully private VCS or use a tool like Quilt [1].\n> > > Also note that certain high-profile projects such as Linux and Git use mailing\n> > > lists for submission and review of patch series; this workflow coexists with\n> > > the concept of DVCS just fine.)\n> >\n> > I'd add though that this isn't really \"co-existing\" with DVSC so much as\n> > using patches on a ML as an indirect transport protocol for \"git push\".\n> >\n> > I.e. if you contributed to some similar projects \"back in the day\" you\n> > could expect to effectively send your patche into a black-hole until the\n> > next release, the maintainer would apply them locally, you wouldn't be\n> > able to pull them back down via the DVCS.\n> >\n> > Perhaps there would be development releases, but those could be weeks or\n> > even months apart, and a \"real\" release might be once every 1-2 years.\n> >\n> > Whereas both Junio and Linus (and other linux maintainers) publish their\n> > version of the patches they do integrate fairly quickly.\n> >\n> > > [...] it also has possible\n> > > downsides; one of a more visible is that when an original project becomes\n> > > dormant for some reason, its users might have hard time understanding which\n> > > one of competing forks to switch to, and there are cases when multiple\n> > > competing forks implement different features and bugfixes, in parallel.\n> > > One of the guys behind Subversion expressed his concerns about this back then\n> > > wgen Git was in its relative infancy [2].\n> > >\n> > >  1. https://en.wikipedia.org/wiki/Quilt_(software)\n> > >  2. http://blog.red-bean.com/sussman/?p=20\n> >\n> > It's interesting that this aspect of what proponents of centralized VCS\n> > were fearful of when it came to DVCS turned out to be the exact\n> > opposite:\n> >\n> >     Notice what this user is now able to do: he wants to to crawl off\n> >     into a cave, work for weeks on a complex feature by himself, then\n> >     present it as a polished result to the main codebase. And this is\n> >     exactly the sort of behavior that I think is bad for open source\n> >     communities.\n> >\n> > I.e. lowering the cost to publish early and often has had the effect\n> > that people are less likely to \"crawl off into a cave\" and work on\n> > something for a long time without syncing up with other parallel\n> > development.\n"},{"id":"456590","messageId":"20220603152301.olqbjtt5j2kqyc3e@carbon","threadId":"57921","inReplyTo":"CACMKQb3_j+iFcf5trZEcWoU7vAsscKv+_sLaEqg_qfazBPTo+Q@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"Konstantin Khomoutov","fromEmail":"kostix@bswap.ru","sentAt":"2022-06-03T15:23:01Z","receivedAt":"2022-06-03T15:57:36Z","isPatch":false,"sender":{"key":"kostix@bswap.ru","avatar":null},"body":"On Fri, Jun 03, 2022 at 05:48:14PM +0530, Aman wrote:\n\n[...]\n> Though I may understand the internal design and high-level\n> implementation of GIT, I really want to know how it's implemented and\n> was made, which means reading the SOURCE CODE.\n> \n> 1. I don't know how absurd of a quest this is, please enlighten me.\n> 2. How do I do it? Where do I start? It's such a BIG repository - and\n> I am not guessing it's going to be easy.\n> 3. Would someone advise, perhaps, to have a look at an older version\n> of the source code? rather than the latest one, for some reason.\n\nWell, depends on what you mean when talking about the two mentioned designs.\nI mean, there's the design of the approach to manage data and there's the\ndesign of the software package (which Git is).\n\nIf you do also understand the latter - that is, understanding that Git is an\nassortment of CLI tools combined into two layers called \"plumbing\" and\n\"porcelain\", - then you should have no difficulty starting to read the code:\nbasically locate the source code of the entry point Git binary (which is,\nwell, \"git\", or \"git.exe\" on Windows) and start reading it. You'll find it\nparses its command-line arguments and calls out to other executable modules\nwhich are parts of the Git software package to do heavy lifting.\nYou then read the source code of the packages of interest, and so on and so\non. I'm not sure there could be any other \"guide\" to read the source code.\n\nIf you're not familiar with the design of Git-as-a-software-package, it's\nprobably time to clone the Git repository and explore the contents of the\ndirectory named \"Documentation\" there.\n\n"},{"id":"456595","messageId":"xmqqk09xhdma.fsf@gitster.g","threadId":"57921","inReplyTo":"CAJoAoZmPOOp41KF=V4EhUopu+P8=UW55bkUJm6ZYiKziprtWug@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-06-03T17:15:57Z","receivedAt":"2022-06-03T17:16:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Emily Shaffer <emilyshaffer@google.com> writes:\n\n>> 3. Would someone advise, perhaps, to have a look at an older version\n>> of the source code? rather than the latest one, for some reason.\n\nFor those who want to learn from source files, I would recommend\nreading all the files in the very initial commit, cover to cover.\n\ne83c5163 (Initial revision of \"git\", the information manager from\nhell, 2005-04-07)\n\nWith only 1244 lines spread across 11 files, it is a short-read that\nis completable in a single sitting for those who are reasonably\nfluent in C.  It does not have any frills, but the basic data\nstructures to express the important concepts are already there.\n\n\n"},{"id":"456681","messageId":"CACMKQb3gYwdyfRSLWO4FWb6+Kxrk-WURpLayrgFsszCKMhWONw@mail.gmail.com","threadId":"57921","inReplyTo":"20220603152301.olqbjtt5j2kqyc3e@carbon","subject":"Re: About GIT Internals","fromName":"Aman","fromEmail":"amanmatreja@gmail.com","sentAt":"2022-06-04T15:24:10Z","receivedAt":"2022-06-04T15:24:26Z","isPatch":false,"sender":{"key":"amanmatreja@gmail.com","avatar":null},"body":"On Fri, Jun 3, 2022, at 8:53 PM Konstantin Khomoutov <kostix@bswap.ru> wrote:\n\n> Well, depends on what you mean when talking about the two mentioned designs.\n> I mean, there's the design of the approach to manage data and there's the\n> design of the software package (which Git is).\n\nThat's a good perspective on the distinction between the designs. I am\nnot familiar yet, with the design of GIT as a software package, and I\nam guessing most people who'll be learning about GIT internals won't\nbe.\n\n> If you do also understand the latter - that is, understanding that Git is an\n> assortment of CLI tools combined into two layers called \"plumbing\" and\n> \"porcelain\", - then you should have no difficulty starting to read the code:\n> basically locate the source code of the entry point Git binary (which is,\n> well, \"git\", or \"git.exe\" on Windows) and start reading it.\n\nHow do I do that?  What do you mean by the \"entry point\" of the git binary?\n"},{"id":"456712","messageId":"20220606115215.mxzney54vf6vkzlp@carbon","threadId":"57921","inReplyTo":"CACMKQb3gYwdyfRSLWO4FWb6+Kxrk-WURpLayrgFsszCKMhWONw@mail.gmail.com","subject":"Re: About GIT Internals","fromName":"Konstantin Khomoutov","fromEmail":"kostix@bswap.ru","sentAt":"2022-06-06T11:52:15Z","receivedAt":"2022-06-06T13:00:52Z","isPatch":false,"sender":{"key":"kostix@bswap.ru","avatar":null},"body":"On Sat, Jun 04, 2022 at 08:54:10PM +0530, Aman wrote:\n\n[...]\n> > If you do also understand the latter - that is, understanding that Git is an\n> > assortment of CLI tools combined into two layers called \"plumbing\" and\n> > \"porcelain\", - then you should have no difficulty starting to read the code:\n> > basically locate the source code of the entry point Git binary (which is,\n> > well, \"git\", or \"git.exe\" on Windows) and start reading it.\n\n(I have reversed the order of your questions below so that my comments follow\nlogically one after another.)\n\n> What do you mean by the \"entry point\" of the git binary?\n\nWell, porcelain Git commands (those supposed to be used by users to carry out\ntheir day-to-day tasks) are all implemented as subcommands of a single\nexecutable image file called \"git\" on all supported platforms (except Windows,\nwhere it's called \"git.exe\"): for instance, you run \"git init\" to initialize a\nrepository, and your OS looks up the executable image file named \"git\"\nsomewhere in the list of directories containing such files (it's usually\ncontained in the environment variable named \"PATH\"), executes it and passes it\na single command-line argument - \"init\". The rest of the commands works the\nsame way. Therefore, that binary named \"git\" is an entry point of the Git\nsoftware package: the execution of most Git commands starts there (not *all*\nGit commands, but let's not touch this yet).\n\n> How do I do that?\n\nWell, basically that's out of the scope of this list, but let's try...\n\nGit is a complex software package mostly written in C (and POSIX shell).\nAs many F/OSS projects written in C, it has a top-level Makefile which is a\nfile supposed to be processed by GNU Make; this file contains a set of rules\nfor generating files from other files (compiling C source code into object\nfiles and linking those into libraries and executable image files is exactly\nthis - generating files from other files). So usually you start from reading\nthe Makefile to find where the binary file of interest is generated, and from\nwhich source files.\n\nThe problem is that Git's Makefile is *complex.*\nSo let's save you some headache and cut straight to the point: of the top\ninterest to you are the two files: git.c and common-main.c. The former is\nexactly what implements that top-level entry point program, \"git\", while the\nlatter implements the function called \"main\" which is an entry point to any\nprogram written in C which is supposed to be runnable standalone (as opposed\nto becoming a library); the object file generated when compiling common-main.c\nis linked to every other compiled code implementing Git commands, its main()\ncalls cmd_main() which is supposed to be implemented in the code of those\ncommands.\n\nThe rest is basically just usual C stuff - source files and header files.\nIf you're not familiar with these basics, then, I'm afraid, Git may be not the\nbest project to dive into.\n\nIn any case, I find the idea proposed by Junio elsewhere in this thread to be\nvery smart: it should be quite enlightening to read the \"early\" Git code to\nmake yourself accustomed to its overal architecture before moving on to its\npresent - much more complicated - implementation which nevertheless still\nmaintains the same architecture.\n\n"}]}