{"thread":{"id":"26916","subject":"Git Pack: Improving cache performance (maybe a good GSoC practice)","startedAt":"2011-03-29T22:21:23Z","lastAt":"2011-03-30T12:07:22Z","messageCount":6,"participants":["Sebastian Thiel","Shawn Pearce","Vicent Marti","Erik Faye-Lund"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"164638","messageId":"4D925B63.9040405@googlemail.com","threadId":"26916","inReplyTo":null,"subject":"Git Pack: Improving cache performance (maybe a good GSoC practice)","fromName":"Sebastian Thiel","fromEmail":"byronimo@googlemail.com","sentAt":"2011-03-29T22:21:23Z","receivedAt":"2011-03-29T22:21:23Z","isPatch":false,"sender":{"key":"byronimo@googlemail.com","avatar":null},"body":"Hi,\n\nWhat follows is a summary of how I approached the git cache in order to \nwrite my own improved version. The conclusions can be found further \ndown, in case you want to skip all the extra words for now.\n\nI am currently working on a c++ implementation of the git core, which \nfor now includes reading and writing of loose objects, as well as \nreading and verifying pack files. Actually, this is not the first time I \ndo this, as I made my first experience in that matter with a pure python \nimplementation of the git core \n(https://github.com/gitpython-developers/gitdb). This time though, I \nwanted to see whether I can achieve better performance, and how I can \nmake git more suitable to handle big files.\n\nWhen profiling my initial uncached version of my pack reading \nimplementation, I noticed that most of the time was actually spent in \nzlibs inflate method. Apparently, a cache was a good idea - git already \nhas one. Ignoring its implementation, I wrote my own naive one right \naway which stored only base objects and inflated deltas. Interestingly, \nthis could already double the performance of my test case, which would \njust stream all data of all objects contained in the pack, sha by sha, \nresembling a random access pattern.\n\nTo compare my cache with git, I implemented pack verification, which \nbasically generates sha1 of all uncompressed objects in the pack, and \nruns a crc32 on the compressed streams. As the objects are streamed \nordered by offset, the access pattern can be described as sequential.\n\nIt came at no surprise, that git would verify my test-pack (aggressively \npacked git source repository with 137k objects, one 27mb pack) much \nfaster, i.e. 25%. After some profiling and optimizations, I could bring \nit down to being just 15% faster. Considering that my cache, which only \ngot faster in the course of the optimizations, could speed up random \naccess by a factor of about 2.5, it was hard to understand that I \ncouldn't reach git's performance.\n\nThe major difference turned out to be the way the cache works. Git has a \nsmall delta cache with only 256 [hardcoded] cache entries and a default \nmemory limit of 16mb. There it stores fully decompressed objects. It \nmaps objects to entries by hashing their pack offsets into the available \nrange of entries. When the pack is accessed sequentially, the cache will \nbe filled with related uncompressed objects, which can in turn reduce \nthe time required to apply the next delta by a huge amount, as only a \nsingle delta has to be applied instead of a possibly long delta chain. \nAs git appears to pack deltas of related objects close to each other \n(regarding their offset in the pack), the cache will be hit quite often \nautomatically. As the number of entries is small, and as entries are \nconnected using a doubly linked list, reclaiming of memory is rather \nefficient, as it hits the first used objects first, which are unlikely \nto be needed ever again. Collection doesn't necessarily run too often as \nwell, as most entries will be overwritten with new data during hash \ncollisions.\nThis cache implementation is clearly suitable for sequential access.\n\nMy cache was optimized for random access, hence it stores only base \nobjects and uncompressed delta streams, using many more entries to \nachieve good cache hit ratios. The reason for the performance gain of \nthe random access cache was that it stores full objects. This fills the \ncache memory up much faster, so having a lot of cache entries makes no \nsense.\n\nBoth cache types are optimized for different kinds of access modes, and \nboth are required to efficiently deal with everything git usually has to do.\nHence I changed my cache to support both modes, and rerun the pack \nverification test.\n\nThe result was better than expected, as the my implementation now takes \nthe lead by a tiny amount (25.3s vs. 26.0s) with a 16mb cache size. On \nmy way to make it even faster, I experimented with different cache \nsizes, amounts of entries and of course different packs, which ranged \nfrom 20mb to 600mb, which helped me fine tune the relations of these \nvariables.\nIn the end, with a cache of 27mb, my implementation took 20.6s, whereas \nthe git implementation could only improve slightly to finish after 25.3s.\nI believe the cause of this is the fixed amount of entries. My cache \nadjusts this amount depending on the packs size, the amount of objects, \nas well as the size of the cache. In the this case, my cache would have \nnearly 1000 entries, which helped to spread the amount of available \nmemory. Due to the limited amount of entries, git will not even benefit \nfrom further increasing the size, whereas I could get as low as 13.5s by \nincreasing the cache size to 48mb for instance.\n\nJust for the fun of it, I increased the amount of entries in the git \ncache to the same amount my cache was using, and suddenly git was \nperforming equally well, finishing after just 20.8s with a 27mb cache size.\n\nAs my random access cache performed worse in sequential access mode, I \nran a test to see whether the opposite is true as well: Does the \nsequential cache harm performance in random access mode ? The answer is: \nYes it does ! To show some numbers:  34mb of objects per second could be \nstreamed without cache, which was reduced to 28mb/s with a random access \ncache. The cache in that case just causes overhead (especially when \nreclaiming memory), and is hit just rarely.\n\nTo test my assumptions not only with my code, but also with git itself, \nI used a test written for git-python, which streams blobs from the 27mb \ngit pack. With the default cache, I get 14mb/s. When I removed the \ncache, it was upped to 15mb, which was less than expected, but we must \nnot forget the git-python overhead here. Finally, with the sequential \naccess cache enabled, its entries increased to 1000, and the cache size \nupped to 27mb, suddenly I would get 34.2mb/s ! A new record, for \ngit-python at least ;).\n\nAs a final disclaimer, please let me emphasize that the tests I run are \nneither statistically profound, nor are the pack verification tests \nnecessarily comparable in all details. Additionally, the git-python \nobject throughput tests cannot be directly compared to the c++ test \nwhich has much less overhead. The tests were made to show performance \nrelations and uncover ways to improve performance, and not to claim that \none implementation is 'better' than the other.\n\n-- Conclusions --\n* delta cache needs to deal with random and sequential access.\n* current implementation deals with sequential access only, which is \nonly suitable for pack verification, and in fact hurts performance in \nother cases if the amount of entries (at least) is not dynamically \nadjusted depending on parameters of the actual pack.\n* random access caches work well with plenty of entries, when storing \nonly uncompressed deltas and base objects, as reapplying a delta is very \nfast.\n* Sequential access caches have to dynamically adjust their amount of \nentries according to the amount of available cache memory and the \naverage packed object size, to make best use of the available memory.\n* it should be possible to adjust the caching mode at runtime, or to \nfully disable the cache.\n* it might be useful/necessary to have one cache per pack sharing global \nmemory limits, instead of having one global cache, as caches need to be \nadjusted depending on the actual pack.\n\nIn case anyone is interested in having a look at the way the I determine \nthe cache parameters (which really are the key to optimizing \nperformance), this is the line you would have to focus on: \nhttps://github.com/Byron/gitplusplus/blob/deltastream/src/git/db/pack_file.cpp#L103. \nThe cache is used by the pack stream, whose core is in the \nunpack_object_recursive method (equivalent to unpack_delta_entry in the \ngit source): \nhttps://github.com/Byron/gitplusplus/blob/deltastream/src/git/db/pack_stream.cpp#L247 \n.\n\nKind Regards,\nSebastian\n"},{"id":"164641","messageId":"AANLkTin6z3hM7nyMqOUPdHrY9TmRVAzpchM+4O=S7KKj@mail.gmail.com","threadId":"26916","inReplyTo":"4D925B63.9040405@googlemail.com","subject":"Re: Git Pack: Improving cache performance (maybe a good GSoC practice)","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2011-03-29T22:45:35Z","receivedAt":"2011-03-29T22:45:35Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Mar 29, 2011 at 15:21, Sebastian Thiel <byronimo@googlemail.com> wrote:\n> I am currently working on a c++ implementation of the git core, which for\n> now includes reading and writing of loose objects, as well as reading and\n> verifying pack files.\n\nHave considered wrapping libgit2 with a C++ binding? Just curious.\n\n> Actually, this is not the first time I do this, as I\n> made my first experience in that matter with a pure python implementation of\n> the git core (https://github.com/gitpython-developers/gitdb).\n\nI think I saw this the other week... why this project vs. using Dulwich[1]?\n\n[1] http://samba.org/~jelmer/dulwich/\n\n> This time\n> though, I wanted to see whether I can achieve better performance, and how I\n> can make git more suitable to handle big files.\n\nA noble goal...\n\n> When profiling my initial uncached version of my pack reading\n> implementation, I noticed that most of the time was actually spent in zlibs\n> inflate method.\n\nYes. The profile is somewhere in this ballpark if Git is doing\nrev-list --objects, aka the \"Counting\" phase of a git clone:\n\n- 30% in zlib inflate()\n- 30% in object map lookup/insertion\n- 30% misc. elsewhere\n\n> The major difference turned out to be the way the cache works. Git has a\n> small delta cache with only 256 [hardcoded] cache entries and a default\n> memory limit of 16mb. There it stores fully decompressed objects. It maps\n> objects to entries by hashing their pack offsets into the available range of\n> entries.\n\nRight, a very simple cache. FWIW, I've tried to use more complex cache\nrules inside of JGit, to no avail. A more complex cache implementation\n(e.g. one that supports a limited number of collisions in the hash\nbuckets and uses a full LRU) runs slow enough relative to this simple\ncache that performance actually gets worse.\n\n> When the pack is accessed sequentially, the cache will be filled\n> with related uncompressed objects, which can in turn reduce the time\n> required to apply the next delta by a huge amount, as only a single delta\n> has to be applied instead of a possibly long delta chain.\n\nYes... mostly.\n\n> As git appears to\n> pack deltas of related objects close to each other (regarding their offset\n> in the pack),\n\nThis isn't true. Git packs object by time, *not* delta ordering.\nHowever objects are delta compressed by commonality on tree path *and*\ntime. An example repository I like to play with is the linux-2.6\nrepository; in that repository the pack is around 370 MiB. If you\nbreak the pack up into 1 MiB slices by offset, you will find that an\nobject at the end of a 50 deep delta chain touches about 50 unique 1\nMiB slices in order to build itself up.  :-)\n\nThis is caused by things being clustered by both time and path. If a\npath is heavily modified within a short time period, sure, those will\nbe clustered together in the file. But if a path is rarely modified,\nits objects will be distributed throughout the file.\n\n> the cache will be hit quite often automatically.\n\nThe hit rate happens to work well because most uses access less than\n256 distinct similar things at once. I forget what the stats are for\nthe linux-2.6 repository, but I think there are less than 256 unique\ndirectories. As Git walks through the history sequentially from\nmost-recent to least-recent, its priming the cache with objects that\nhave very short delta chains and are thus more likely to be used as\ndelta bases for objects later in the file. Since each directory or\nfile acts as a delta base for someone else later, its likely to be in\nthis cache as the reader walks backwards through time. As bases\nswitch, the cache is updated at a relatively low penalty, because the\nnew base was itself recently accessed using the base that is already\nin the cache.\n\nThe simple % 256 rule the cache uses is effective because objects are\npretty randomly allocated as far as offsets go in the file. We just\ndamn lucky. :-)\n\n> This cache implementation is clearly suitable for sequential access.\n\nYes.\n\n> Both cache types are optimized for different kinds of access modes, and both\n> are required to efficiently deal with everything git usually has to do.\n> Hence I changed my cache to support both modes, and rerun the pack\n> verification test.\n>\n> The result was better than expected, as the my implementation now takes the\n> lead by a tiny amount (25.3s vs. 26.0s) with a 16mb cache size. On my way to\n\nThis isn't a very significant speed difference given the differences\nin implementation. We're not really looking to shave 3% off the\nrunning time for operation X, we're looking to shave >10%.\n\n> make it even faster, I experimented with different cache sizes, amounts of\n> entries and of course different packs, which ranged from 20mb to 600mb,\n> which helped me fine tune the relations of these variables.\n> In the end, with a cache of 27mb, my implementation took 20.6s, whereas the\n\nOK, this is pretty significant. Saving 21% of the running time, at the\nexpense of an extra 11M of working set.\n\nBut the verify pack workload is pretty useless, nobody accesses data\nby SHA-1 order. Most uses of Git are going backwards through time. log\nand blame are the two notable things that happen *a lot* and that\nusers complain about being slow. These also aren't random accesses,\nthere is a definite pattern and the pattern can be exploited. I'm\nreally only interested in improving these two patterns.\n\nAs far as verify-pack improving, Junio improved it by switching to use\nindex-pack with the new --verify flag. There really isn't a faster way\nto scan through a pack than the way index-pack does it.\n\nSo, all I'm trying to say is, verify-pack isn't the right thing to\ntarget when you are looking at \"how do I make Git faster\".\n\n> -- Conclusions --\n> * delta cache needs to deal with random and sequential access.\n\nI'm not sure where the random access case is coming from. Who is doing\nrandom access except verify-pack?\n\n> * current implementation deals with sequential access only, which is only\n> suitable for pack verification,\n\nNot true. First, pack verification is horrifically random, since its\nby SHA-1 order and not sequential order. Second, every other use of\nthe pack data is generally sequential in time, because every other use\nis starting from the current revisions as found from the refs and\nwalking backwards in time, which is forwards sequentially in the pack.\n\n-- \nShawn.\n"},{"id":"164670","messageId":"4D92EDA0.8030309@googlemail.com","threadId":"26916","inReplyTo":"AANLkTin6z3hM7nyMqOUPdHrY9TmRVAzpchM+4O=S7KKj@mail.gmail.com","subject":"Re: Git Pack: Improving cache performance (maybe a good GSoC practice)","fromName":"Sebastian Thiel","fromEmail":"byronimo@googlemail.com","sentAt":"2011-03-30T08:45:20Z","receivedAt":"2011-03-30T08:45:20Z","isPatch":false,"sender":{"key":"byronimo@googlemail.com","avatar":null},"body":"Hi Shawn,\n\nThank you for your detailed answer, especially about how deltas are \nordered within the pack. First things first: Where does my random access \n(sha1 by sha1) access pattern come from ?\nIts clearly just part of my test, as its easy to just iterate shas in \nthe index and query their data in the pack. The pack verification though \nis not using sha1 order, but offset order, iterating the pack from the \nsmallest to the largest offset. This is true for the git implementation, \nas well as for my one, which is why I would believe the access pattern \nis quite sequential here.\nWhen reading the reply, at first I thought we agreed that pack \nverification is sequential, but what confused me is one of your last \nstatement: \"First, pack verification is horrifically random, since its \nby SHA-1 order and not sequential order\".\n\nNonetheless, you are absolutely right that the sha1 ordered access is \nnothing that would usually happen in real life, but I didn't yet \nimplement commit walking or tree iteration.\n\nWhat stays is my observation that a larger, or lets say, more adaptive \namount of entries, can greatly improve performance. The git-python test \nactually iterates commits, new to old, iterates the respective trees \ndepth first, and streams all blobs. As I understand it, this is a common \naccess pattern, which would greatly benefit from a larger entry cache. \nIt improved performance from 14mb/s to 34mb/s, I used about 1000 entries \nin the cache, and a memory cap of 27mb.\nThe default pack-verify implementation would also benefit from more \ncache entries\nMaybe the default of 256 entries is sufficient if the trees are iterated \nbreadth first, but to my mind depth first would be a valid access \npattern as well.\n\nThe simplicity of the cache to me is the right approach, but I cannot \nagree with its statically allocated amount of entries, as it apparently \ndoesn't suit any but the smallest packs I tried. Even though it might \nnot be statistically relevant, completely disabling the cache boosted \nthe git-python test\ndescribed previously by 1mb/s, reproducibly, which seems to show that \nthe cache can hurt if there aren't enough entries at least.\n\nTo my mind, changing the cache to be per-pack with dynamically allocated \nentries depending on the average size of uncompressed objects will help \nperformance enough to be worth the effort.\n\nPlease see some more comments further down the email.\nKind Regards,\nSebastian\n\nOn 03/30/2011 12:45 AM, Shawn Pearce wrote:\n> On Tue, Mar 29, 2011 at 15:21, Sebastian Thiel<byronimo@googlemail.com>  wrote:\n>> I am currently working on a c++ implementation of the git core, which for\n>> now includes reading and writing of loose objects, as well as reading and\n>> verifying pack files.\n> Have considered wrapping libgit2 with a C++ binding? Just curious.\nThe project appears to be silent for nearly 5 months now, and it is in a \nrather early stage of development. There is no delta cache yet, nor is \nthere a sliding window mmap implementation which would be required on 32 \nbit systems, at least if you want to have big file support.\n>> Actually, this is not the first time I do this, as I\n>> made my first experience in that matter with a pure python implementation of\n>> the git core (https://github.com/gitpython-developers/gitdb).\n> I think I saw this the other week... why this project vs. using Dulwich[1]?\n>\n> [1] http://samba.org/~jelmer/dulwich/\nJelmer and I talked about how both projects could benefit from each \nother, but we dropped the idea once it turned out that the licenses are \nquite incompatible (gpl vs. bsd). Besides, I like big file support, \nwhich also means that the system should internally stream all data, \nusing a stream-like interface. Dulwich currently puts all data into RAM, \nand so does git. Gitdb uses stream interfaces exclusively, but \nadmittedly I still didn't implement a delta decompression that would \nwork without plenty of buffers ... but that's a different topic.\n>> This time\n>> though, I wanted to see whether I can achieve better performance, and how I\n>> can make git more suitable to handle big files.\n> A noble goal...\n... which can be reached :). Git-like databases could greatly improve \nthe performance of existing technologies, like package managers or \nupdate systems (for games, for instance) if people wouldn't have to \nre-download whole packages although only a few bytes/files changed in \nthe new version. Having a customizable git-library for this would allow \nanyone to easily implement his custom git-like database solution to \noptimize these kinds of transfers. This is what drives me.\n>> When profiling my initial uncached version of my pack reading\n>> implementation, I noticed that most of the time was actually spent in zlibs\n>> inflate method.\n> Yes. The profile is somewhere in this ballpark if Git is doing\n> rev-list --objects, aka the \"Counting\" phase of a git clone:\n>\n> - 30% in zlib inflate()\n> - 30% in object map lookup/insertion\n> - 30% misc. elsewhere\n>\n>> The major difference turned out to be the way the cache works. Git has a\n>> small delta cache with only 256 [hardcoded] cache entries and a default\n>> memory limit of 16mb. There it stores fully decompressed objects. It maps\n>> objects to entries by hashing their pack offsets into the available range of\n>> entries.\n> Right, a very simple cache. FWIW, I've tried to use more complex cache\n> rules inside of JGit, to no avail. A more complex cache implementation\n> (e.g. one that supports a limited number of collisions in the hash\n> buckets and uses a full LRU) runs slow enough relative to this simple\n> cache that performance actually gets worse.\n>\n>> When the pack is accessed sequentially, the cache will be filled\n>> with related uncompressed objects, which can in turn reduce the time\n>> required to apply the next delta by a huge amount, as only a single delta\n>> has to be applied instead of a possibly long delta chain.\n> Yes... mostly.\n>\n>> As git appears to\n>> pack deltas of related objects close to each other (regarding their offset\n>> in the pack),\n> This isn't true. Git packs object by time, *not* delta ordering.\n> However objects are delta compressed by commonality on tree path *and*\n> time. An example repository I like to play with is the linux-2.6\n> repository; in that repository the pack is around 370 MiB. If you\n> break the pack up into 1 MiB slices by offset, you will find that an\n> object at the end of a 50 deep delta chain touches about 50 unique 1\n> MiB slices in order to build itself up.  :-)\n>\n> This is caused by things being clustered by both time and path. If a\n> path is heavily modified within a short time period, sure, those will\n> be clustered together in the file. But if a path is rarely modified,\n> its objects will be distributed throughout the file.\n>\n>> the cache will be hit quite often automatically.\n> The hit rate happens to work well because most uses access less than\n> 256 distinct similar things at once. I forget what the stats are for\n> the linux-2.6 repository, but I think there are less than 256 unique\n> directories. As Git walks through the history sequentially from\n> most-recent to least-recent, its priming the cache with objects that\n> have very short delta chains and are thus more likely to be used as\n> delta bases for objects later in the file. Since each directory or\n> file acts as a delta base for someone else later, its likely to be in\n> this cache as the reader walks backwards through time. As bases\n> switch, the cache is updated at a relatively low penalty, because the\n> new base was itself recently accessed using the base that is already\n> in the cache.\n>\n> The simple % 256 rule the cache uses is effective because objects are\n> pretty randomly allocated as far as offsets go in the file. We just\n> damn lucky. :-)\n>\n>> This cache implementation is clearly suitable for sequential access.\n> Yes.\n>\n>> Both cache types are optimized for different kinds of access modes, and both\n>> are required to efficiently deal with everything git usually has to do.\n>> Hence I changed my cache to support both modes, and rerun the pack\n>> verification test.\n>>\n>> The result was better than expected, as the my implementation now takes the\n>> lead by a tiny amount (25.3s vs. 26.0s) with a 16mb cache size. On my way to\n> This isn't a very significant speed difference given the differences\n> in implementation. We're not really looking to shave 3% off the\n> running time for operation X, we're looking to shave>10%.\n>\n>> make it even faster, I experimented with different cache sizes, amounts of\n>> entries and of course different packs, which ranged from 20mb to 600mb,\n>> which helped me fine tune the relations of these variables.\n>> In the end, with a cache of 27mb, my implementation took 20.6s, whereas the\n> OK, this is pretty significant. Saving 21% of the running time, at the\n> expense of an extra 11M of working set.\n>\n> But the verify pack workload is pretty useless, nobody accesses data\n> by SHA-1 order. Most uses of Git are going backwards through time. log\n> and blame are the two notable things that happen *a lot* and that\n> users complain about being slow. These also aren't random accesses,\n> there is a definite pattern and the pattern can be exploited. I'm\n> really only interested in improving these two patterns.\n>\n> As far as verify-pack improving, Junio improved it by switching to use\n> index-pack with the new --verify flag. There really isn't a faster way\n> to scan through a pack than the way index-pack does it.\n>\n> So, all I'm trying to say is, verify-pack isn't the right thing to\n> target when you are looking at \"how do I make Git faster\".\n>\nI couldn't find the index-pack --verify flag in 1.7.4.2 - but maybe it \nis even more bleeding edge, or I am looking in the wrong place.\n>> -- Conclusions --\n>> * delta cache needs to deal with random and sequential access.\n> I'm not sure where the random access case is coming from. Who is doing\n> random access except verify-pack?\n>\nSee top of reply.\n>> * current implementation deals with sequential access only, which is only\n>> suitable for pack verification,\n> Not true. First, pack verification is horrifically random, since its\n> by SHA-1 order and not sequential order. Second, every other use of\n> the pack data is generally sequential in time, because every other use\n> is starting from the current revisions as found from the refs and\n> walking backwards in time, which is forwards sequentially in the pack.\n>\n"},{"id":"164675","messageId":"AANLkTim+Ge2c-i_jUi8YN8g+cQmXuyKYrdHC+jYukjQy@mail.gmail.com","threadId":"26916","inReplyTo":"4D92EDA0.8030309@googlemail.com","subject":"Re: Git Pack: Improving cache performance (maybe a good GSoC practice)","fromName":"Vicent Marti","fromEmail":"vicent@github.com","sentAt":"2011-03-30T09:46:18Z","receivedAt":"2011-03-30T09:46:18Z","isPatch":false,"sender":{"key":"vicent@github.com","avatar":"https://gravatar.com/avatar/9d57a2b1e3137bf84342ac1dfdf1cde409b86e8fda5397d05f40f17fa5b84a63?d=mp&s=160"},"body":"On Wed, Mar 30, 2011 at 11:45 AM, Sebastian Thiel\n<byronimo@googlemail.com> wrote:\n>> Have considered wrapping libgit2 with a C++ binding? Just curious.\n>\n> The project appears to be silent for nearly 5 months now, and it is in a\n> rather early stage of development. There is no delta cache yet, nor is there\n> a sliding window mmap implementation which would be required on 32 bit\n> systems, at least if you want to have big file support.\n\nwat?\n\nhttp://libgit2.github.com\nhttps://github.com/libgit2/libgit2\n\nCheers,\nVicent\n"},{"id":"164677","messageId":"4D92FD9D.8040707@gmail.com","threadId":"26916","inReplyTo":"AANLkTim+Ge2c-i_jUi8YN8g+cQmXuyKYrdHC+jYukjQy@mail.gmail.com","subject":"Re: Git Pack: Improving cache performance (maybe a good GSoC practice)","fromName":"Sebastian Thiel","fromEmail":"byronimo@googlemail.com","sentAt":"2011-03-30T09:53:33Z","receivedAt":"2011-03-30T09:53:33Z","isPatch":false,"sender":{"key":"byronimo@googlemail.com","avatar":null},"body":"Thank you very much for the heads-up - I was using old mirrors it appears:\n\ngit://repo.or.cz/libgit2.git\ngit://repo.or.cz/libgit2/raj.git\n\nIts quite terrible that I was left thinking that the project stalled for\nso long, and it was hard for me to understand why people would continue\nto bring it up :).\nNow it all makes sense !\n\nCheers,\nSebastian\n\nOn 30.03.11 11:46, Vicent Marti wrote:\n> On Wed, Mar 30, 2011 at 11:45 AM, Sebastian Thiel\n> <byronimo@googlemail.com> wrote:\n>>> Have considered wrapping libgit2 with a C++ binding? Just curious.\n>> The project appears to be silent for nearly 5 months now, and it is in a\n>> rather early stage of development. There is no delta cache yet, nor is there\n>> a sliding window mmap implementation which would be required on 32 bit\n>> systems, at least if you want to have big file support.\n> wat?\n>\n> http://libgit2.github.com\n> https://github.com/libgit2/libgit2\n>\n> Cheers,\n> Vicent\n"},{"id":"164681","messageId":"AANLkTikD9OrboQ_0Qi+6vsJz3Bxe5g5GazTCP4LUFynN@mail.gmail.com","threadId":"26916","inReplyTo":"4D92FD9D.8040707@gmail.com","subject":"Re: Git Pack: Improving cache performance (maybe a good GSoC practice)","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@gmail.com","sentAt":"2011-03-30T12:07:22Z","receivedAt":"2011-03-30T12:07:22Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Wed, Mar 30, 2011 at 11:53 AM, Sebastian Thiel\n<byronimo@googlemail.com> wrote:\n> Thank you very much for the heads-up - I was using old mirrors it appears:\n>\n> git://repo.or.cz/libgit2.git\n> git://repo.or.cz/libgit2/raj.git\n>\n> Its quite terrible that I was left thinking that the project stalled for\n> so long, and it was hard for me to understand why people would continue\n> to bring it up :).\n> Now it all makes sense !\n>\n\nI was also confused by this at first. Shawn, would you mind updating\nthe readme-field in the repo.or.cz-mirror to reflect that the project\nhas moved to GitHub?\n"}]}