git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git's database structure

From
Junio C Hamano <gitster@pobox.com>
Date
Sep 4, 2007, 18:06 UTC
Message-ID
<7v1wdenw4n.fsf@gitster.siamese.dyndns.org>
In-Reply-To
<9e4733910709041044r71264346n341d178565dd0521@mail.gmail.com>
"Jon Smirl" <jonsmirl@gmail.com> writes:
Show 22 quoted lines
> On 9/4/07, Junio C Hamano <gitster@pobox.com> wrote:
>> "Jon Smirl" <jonsmirl@gmail.com> writes:
>>
>> > Another way of looking at the problem,
>> >
>> > Let's build a full-text index for git. You put a string into the index
>> > and it returns the SHAs of all the file nodes that contain the string.
>> > How do I recover the path names of these SHAs?
>>
>> That question does not make much sense without specifying "which
>> commit's path you are talking about".
>>
>> If you want to encode such "contextual information" in addition
>> to "contents", you could do so, but you essentially need to
>> record commit + pathname + mode bits + contents as "blob" and
>> hash that to come up with a name.
>
> I left the details out of the full-text example to make it more
> obvious that we can't recover the path names.
>
> Doing this type of analysis may point out that even more fields are
> missing from the blob table such as commit id.

Quite the contrary. You just illustrated why it is wrong to put anything but contents in the blob.

The specialized indexing is a different issue. If you want to have a full text index to answer "what paths in which commits had this string?", then your database table would have columns such as commit (sha-1), path (string) as values, indexed with the search string.

Now the current set of "git" operation does not need to answer that query, so we do not build nor maintain such an index that nobody uses. But your application may benefit from such an index, and as others said, nobody prevents you from building one.

Previous: Reece DunnNext: Theodore Tso
Message 18 of 39 in “Git's database structure”
  1. Jon SmirlSep 4, 2007
  2. Andreas EricssonSep 4, 2007
  3. Mike HommeySep 4, 2007
  4. Andreas EricssonSep 4, 2007
  5. Jon SmirlSep 4, 2007
  6. Andreas EricssonSep 4, 2007
  7. Jeff KingSep 4, 2007
  8. David TweedSep 4, 2007
  9. Junio C HamanoSep 4, 2007
  10. Jon SmirlSep 4, 2007
  11. Andreas EricssonSep 4, 2007
  12. Jon SmirlSep 4, 2007
  13. Andreas EricssonSep 4, 2007
  14. Junio C HamanoSep 4, 2007
  15. Jon SmirlSep 4, 2007
  16. Mike HommeySep 4, 2007
  17. Reece DunnSep 4, 2007
  18. Junio C HamanoSep 4, 2007
  19. Theodore TsoSep 4, 2007
  20. Jon SmirlSep 4, 2007
  21. Andreas EricssonSep 5, 2007
  22. Jon SmirlSep 5, 2007
  23. Andreas EricssonSep 5, 2007
  24. Jon SmirlSep 5, 2007
  25. Julian PhillipsSep 5, 2007
  26. Jon SmirlSep 5, 2007
  27. Julian PhillipsSep 5, 2007
  28. Kyle MoffettSep 6, 2007
  29. Mike HommeySep 5, 2007
  30. Andreas EricssonSep 6, 2007
  31. Junio C HamanoSep 6, 2007
  32. Wincent ColaiutaSep 6, 2007
  33. Johannes SchindelinSep 6, 2007
  34. Steven GrimmSep 6, 2007
  35. Martin LanghoffSep 7, 2007
  36. Andy ParkinsSep 5, 2007
  37. Julian PhillipsSep 4, 2007
  38. Jon SmirlSep 4, 2007
  39. Andreas EricssonSep 4, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.