git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 11/14] rust: add functionality to hash an object

From
Ezekiel Newren <ezekielnewren@gmail.com>
Date
Oct 28, 2025, 18:05 UTC
Message-ID
<CAH=ZcbDCrYuSW7nLerQZnT-R_CoCtN2RNycLqOEEV-T-T7VoZQ@mail.gmail.com>
In-Reply-To
<20251027004404.2152927-12-sandals@crustytoothpaste.net>

On Sun, Oct 26, 2025 at 6:44 PM brian m. carlson <sandals@crustytoothpaste.net> wrote:

Show 37 quoted lines
>
> In a future commit, we'll want to hash some data when dealing with a
> loose object map.  Let's make this easy by creating a structure to hash
> objects and calling into the C functions as necessary to perform the
> hashing.  For now, we only implement safe hashing, but in the future we
> could add unsafe hashing if we want.  Implement Clone and Drop to
> appropriately manage our memory.  Additionally implement Write to make
> it easy to use with other formats that implement this trait.
>
> While we're at it, add some tests for the various cases in this file.
>
> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>
> ---
>  src/hash.rs | 157 ++++++++++++++++++++++++++++++++++++++++++++++++++++
>  1 file changed, 157 insertions(+)
>
> diff --git a/src/hash.rs b/src/hash.rs
> index a5b9493bd8..8798a50aef 100644
> --- a/src/hash.rs
> +++ b/src/hash.rs
> @@ -10,6 +10,7 @@
>  // You should have received a copy of the GNU General Public License along
>  // with this program; if not, see <https://www.gnu.org/licenses/>.
>
> +use std::io::{self, Write};
>  use std::os::raw::c_void;
>
>  pub const GIT_MAX_RAWSZ: usize = 32;
> @@ -39,6 +40,81 @@ impl ObjectID {
>      }
>  }
>
> +pub struct Hasher {
> +    algo: HashAlgorithm,
> +    safe: bool,
> +    ctx: *mut c_void,
> +}

The name _Hasher_ is already used by std::hash::Hasher. It would be preferable to pick a different name to avoid confusion. Perhaps CryptoHasher, SecureHasher?

Show 11 quoted lines
> +impl Hasher {
> +    /// Create a new safe hasher.
> +    pub fn new(algo: HashAlgorithm) -> Hasher {
> +        let ctx = unsafe { c::git_hash_alloc() };
> +        unsafe { c::git_hash_init(ctx, algo.hash_algo_ptr()) };
> +        Hasher {
> +            algo,
> +            safe: true,
> +            ctx,
> +        }
> +    }
-    pub fn new(algo: HashAlgorithm) -> Hasher {
+    pub fn new(algo: HashAlgorithm) -> Self {
         let ctx = unsafe { c::git_hash_alloc() };
         unsafe { c::git_hash_init(ctx, algo.hash_algo_ptr()) };
-        Hasher {
+        Self {
            algo,
            safe: true,
            ctx,
        }
> +    /// Return whether this is a safe hasher.
> +    pub fn is_safe(&self) -> bool {
> +        self.safe
> +    }

I don't understand the point in being able to query whether a given hasher is safe or not. How does that change how this hasher code is used? If the functions are safe then you wouldn't wrap it in an unsafe block. If the functions are declared with unsafe then you'd always need to wrap it in an unsafe block whether it's actually safe or not. Using unsafe in Rust isn't like error handling where you do something different on failure. If something fails in unsafe it's usually unrecoverable e.g. segfault due to invalid memory access. My understanding of unsafe in Rust means "The compiler can't verify that this code is actually safe to run, so I've made sure that it is safe myself and I'll let the compiler know what code to ignore during compilation."

Show 51 quoted lines
> +    /// Update the hasher with the specified data.
> +    pub fn update(&mut self, data: &[u8]) {
> +        unsafe { c::git_hash_update(self.ctx, data.as_ptr() as *const c_void, data.len()) };
> +    }
> +
> +    /// Return an object ID, consuming the hasher.
> +    pub fn into_oid(self) -> ObjectID {
> +        let mut oid = ObjectID {
> +            hash: [0u8; 32],
> +            algo: self.algo as u32,
> +        };
> +        unsafe { c::git_hash_final_oid(&mut oid as *mut ObjectID as *mut c_void, self.ctx) };
> +        oid
> +    }
> +
> +    /// Return a hash as a `Vec`, consuming the hasher.
> +    pub fn into_vec(self) -> Vec<u8> {
> +        let mut v = vec![0u8; self.algo.raw_len()];
> +        unsafe { c::git_hash_final(v.as_mut_ptr(), self.ctx) };
> +        v
> +    }
> +}
> +
> +impl Write for Hasher {
> +    fn write(&mut self, data: &[u8]) -> io::Result<usize> {
> +        self.update(data);
> +        Ok(data.len())
> +    }
> +
> +    fn flush(&mut self) -> io::Result<()> {
> +        Ok(())
> +    }
> +}
> +
> +impl Clone for Hasher {
> +    fn clone(&self) -> Hasher {
> +        let ctx = unsafe { c::git_hash_alloc() };
> +        unsafe { c::git_hash_clone(ctx, self.ctx) };
> +        Hasher {
> +            algo: self.algo,
> +            safe: self.safe,
> +            ctx,
> +        }
> +    }
> +}
> +
> +impl Drop for Hasher {
> +    fn drop(&mut self) {
> +        unsafe { c::git_hash_free(self.ctx) };
> +    }
> +}
Make sense.
Show 97 quoted lines
>  /// A hash algorithm,
>  #[repr(C)]
>  #[derive(Debug, Copy, Clone, Ord, PartialOrd, Eq, PartialEq)]
> @@ -167,6 +243,11 @@ impl HashAlgorithm {
>      pub fn hash_algo_ptr(self) -> *const c_void {
>          unsafe { c::hash_algo_ptr_by_offset(self as u32) }
>      }
> +
> +    /// Create a hasher for this algorithm.
> +    pub fn hasher(self) -> Hasher {
> +        Hasher::new(self)
> +    }
>  }
>
>  pub mod c {
> @@ -174,5 +255,81 @@ pub mod c {
>
>      extern "C" {
>          pub fn hash_algo_ptr_by_offset(n: u32) -> *const c_void;
> +        pub fn unsafe_hash_algo(algop: *const c_void) -> *const c_void;
> +        pub fn git_hash_alloc() -> *mut c_void;
> +        pub fn git_hash_free(ctx: *mut c_void);
> +        pub fn git_hash_init(dst: *mut c_void, algop: *const c_void);
> +        pub fn git_hash_clone(dst: *mut c_void, src: *const c_void);
> +        pub fn git_hash_update(ctx: *mut c_void, inp: *const c_void, len: usize);
> +        pub fn git_hash_final(hash: *mut u8, ctx: *mut c_void);
> +        pub fn git_hash_final_oid(hash: *mut c_void, ctx: *mut c_void);
> +    }
> +}
> +
> +#[cfg(test)]
> +mod tests {
> +    use super::{HashAlgorithm, ObjectID};
> +    use std::io::Write;
> +
> +    fn all_algos() -> &'static [HashAlgorithm] {
> +        &[HashAlgorithm::SHA1, HashAlgorithm::SHA256]
> +    }
> +
> +    #[test]
> +    fn format_id_round_trips() {
> +        for algo in all_algos() {
> +            assert_eq!(
> +                *algo,
> +                HashAlgorithm::from_format_id(algo.format_id()).unwrap()
> +            );
> +        }
> +    }
> +
> +    #[test]
> +    fn offset_round_trips() {
> +        for algo in all_algos() {
> +            assert_eq!(*algo, HashAlgorithm::from_u32(*algo as u32).unwrap());
> +        }
> +    }
> +
> +    #[test]
> +    fn slices_have_correct_length() {
> +        for algo in all_algos() {
> +            for oid in [algo.null_oid(), algo.empty_blob(), algo.empty_tree()] {
> +                assert_eq!(oid.as_slice().len(), algo.raw_len());
> +            }
> +        }
> +    }
> +
> +    #[test]
> +    fn hasher_works_correctly() {
> +        for algo in all_algos() {
> +            let tests: &[(&[u8], &ObjectID)] = &[
> +                (b"blob 0\0", algo.empty_blob()),
> +                (b"tree 0\0", algo.empty_tree()),
> +            ];
> +            for (data, oid) in tests {
> +                let mut h = algo.hasher();
> +                assert_eq!(h.is_safe(), true);
> +                // Test that this works incrementally.
> +                h.update(&data[0..2]);
> +                h.update(&data[2..]);
> +
> +                let h2 = h.clone();
> +
> +                let actual_oid = h.into_oid();
> +                assert_eq!(**oid, actual_oid);
> +
> +                let v = h2.into_vec();
> +                assert_eq!((*oid).as_slice(), &v);
> +
> +                let mut h = algo.hasher();
> +                h.write_all(&data[0..2]).unwrap();
> +                h.write_all(&data[2..]).unwrap();
> +
> +                let actual_oid = h.into_oid();
> +                assert_eq!(**oid, actual_oid);
> +            }
> +        }
>      }
>  }
Looks good.
Previous: Patrick SteinhardtNext: brian m. carlson
Message 14 of 118 in “SHA-1/SHA-256 interoperability, part 2”
  1. 00/14 SHA-1/SHA-256 interoperability, part 2brian m. carlson, Oct 27, 2025
  2. 14/14 object-file-convert: always make sure object ID algo is validbrian m. carlson, Oct 27, 2025
  3. 05/14 rust: add a hash algorithm abstractionbrian m. carlson, Oct 27, 2025
  4. Patrick SteinhardtOct 28, 2025
  5. Ezekiel NewrenOct 28, 2025
  6. Junio C HamanoOct 28, 2025
  7. Ezekiel NewrenOct 28, 2025
  8. Junio C HamanoOct 29, 2025
  9. Junio C HamanoOct 29, 2025
  10. 11/14 rust: add functionality to hash an objectbrian m. carlson, Oct 27, 2025
  11. Patrick SteinhardtOct 28, 2025
  12. brian m. carlsonOct 29, 2025
  13. Patrick SteinhardtOct 29, 2025
  14. Ezekiel NewrenOct 28, 2025
  15. brian m. carlsonOct 29, 2025
  16. Ben KnobleOct 29, 2025
  17. 07/14 csum-file: define hashwrite's count as a uint32_tbrian m. carlson, Oct 27, 2025
  18. Ezekiel NewrenOct 28, 2025
  19. 09/14 hash: expose hash context functions to Rustbrian m. carlson, Oct 27, 2025
  20. Junio C HamanoOct 29, 2025
  21. brian m. carlsonOct 30, 2025
  22. Junio C HamanoOct 30, 2025
  23. 13/14 rust: add a small wrapper around the hashfile codebrian m. carlson, Oct 27, 2025
  24. Ezekiel NewrenOct 28, 2025
  25. brian m. carlsonOct 29, 2025
  26. 06/14 hash: add a function to look up hash algo structsbrian m. carlson, Oct 27, 2025
  27. Patrick SteinhardtOct 28, 2025
  28. Junio C HamanoOct 28, 2025
  29. brian m. carlsonNov 4, 2025
  30. Junio C HamanoNov 4, 2025
  31. 10/14 rust: add a build.rs script for testsbrian m. carlson, Oct 27, 2025
  32. Patrick SteinhardtOct 28, 2025
  33. Ezekiel NewrenOct 28, 2025
  34. Junio C HamanoOct 29, 2025
  35. Ezekiel NewrenOct 29, 2025
  36. Junio C HamanoOct 29, 2025
  37. Patrick SteinhardtOct 30, 2025
  38. Junio C HamanoOct 30, 2025
  39. Ezekiel NewrenOct 31, 2025
  40. Junio C HamanoNov 1, 2025
  41. 12/14 rust: add a new binary loose object map formatbrian m. carlson, Oct 27, 2025
  42. Patrick SteinhardtOct 28, 2025
  43. brian m. carlsonOct 29, 2025
  44. Patrick SteinhardtOct 29, 2025
  45. Junio C HamanoOct 29, 2025
  46. Junio C HamanoOct 29, 2025
  47. 08/14 write-or-die: add an fsync component for the loose object mapbrian m. carlson, Oct 27, 2025
  48. 02/14 conversion: don't crash when no destination algobrian m. carlson, Oct 27, 2025
  49. 03/14 hash: use uint32_t for object_id algorithmbrian m. carlson, Oct 27, 2025
  50. Patrick SteinhardtOct 28, 2025
  51. Ezekiel NewrenOct 28, 2025
  52. Junio C HamanoOct 28, 2025
  53. Ezekiel NewrenOct 28, 2025
  54. Junio C HamanoOct 28, 2025
  55. brian m. carlsonOct 30, 2025
  56. Collin FunkOct 30, 2025
  57. brian m. carlsonNov 3, 2025
  58. brian m. carlsonOct 29, 2025
  59. Patrick SteinhardtOct 29, 2025
  60. 04/14 rust: add a ObjectID structbrian m. carlson, Oct 27, 2025
  61. Patrick SteinhardtOct 28, 2025
  62. Ezekiel NewrenOct 28, 2025
  63. brian m. carlsonOct 29, 2025
  64. Junio C HamanoOct 28, 2025
  65. brian m. carlsonOct 29, 2025
  66. brian m. carlsonOct 29, 2025
  67. Patrick SteinhardtOct 29, 2025
  68. brian m. carlsonOct 30, 2025
  69. 01/14 repository: require Rust support for interoperabilitybrian m. carlson, Oct 27, 2025
  70. Patrick SteinhardtOct 28, 2025
  71. Junio C HamanoOct 29, 2025
  72. Junio C HamanoOct 29, 2025
  73. Ezekiel NewrenNov 11, 2025
  74. Junio C HamanoNov 14, 2025
  75. Junio C HamanoNov 14, 2025
  76. Junio C HamanoNov 17, 2025
  77. brian m. carlsonNov 17, 2025
  78. Junio C HamanoNov 18, 2025
  79. brian m. carlsonNov 19, 2025
  80. Junio C HamanoNov 19, 2025
  81. Ezekiel NewrenNov 19, 2025
  82. Ezekiel NewrenNov 20, 2025
  83. brian m. carlsonNov 20, 2025
  84. Ezekiel NewrenNov 20, 2025
  85. Junio C HamanoNov 20, 2025
  86. 00/15 SHA-1/SHA-256 interoperability, part 2brian m. carlson, Nov 17, 2025
  87. 02/15 conversion: don't crash when no destination algobrian m. carlson, Nov 17, 2025
  88. 03/15 hash: use uint32_t for object_id algorithmbrian m. carlson, Nov 17, 2025
  89. 01/15 repository: require Rust support for interoperabilitybrian m. carlson, Nov 17, 2025
  90. 04/15 rust: add a ObjectID structbrian m. carlson, Nov 17, 2025
  91. 06/15 hash: add a function to look up hash algo structsbrian m. carlson, Nov 17, 2025
  92. 08/15 csum-file: define hashwrite's count as a uint32_tbrian m. carlson, Nov 17, 2025
  93. 05/15 rust: add a hash algorithm abstractionbrian m. carlson, Nov 17, 2025
  94. 09/15 write-or-die: add an fsync component for the object mapbrian m. carlson, Nov 17, 2025
  95. 10/15 hash: expose hash context functions to Rustbrian m. carlson, Nov 17, 2025
  96. 07/15 rust: add additional helpers for ObjectIDbrian m. carlson, Nov 17, 2025
  97. 12/15 rust: add functionality to hash an objectbrian m. carlson, Nov 17, 2025
  98. 11/15 rust: add a build.rs script for testsbrian m. carlson, Nov 17, 2025
  99. 14/15 rust: add a small wrapper around the hashfile codebrian m. carlson, Nov 17, 2025
  100. 15/15 object-file-convert: always make sure object ID algo is validbrian m. carlson, Nov 17, 2025
  101. 13/15 rust: add a new binary object map formatbrian m. carlson, Nov 17, 2025
  102. 00/16 SHA-1/SHA-256 interoperability, part 2brian m. carlson, Feb 7, 2026
  103. 04/16 rust: add a ObjectID structbrian m. carlson, Feb 7, 2026
  104. 02/16 conversion: don't crash when no destination algobrian m. carlson, Feb 7, 2026
  105. 01/16 repository: require Rust support for interoperabilitybrian m. carlson, Feb 7, 2026
  106. 03/16 hash: use uint32_t for object_id algorithmbrian m. carlson, Feb 7, 2026
  107. 07/16 rust: add additional helpers for ObjectIDbrian m. carlson, Feb 7, 2026
  108. 14/16 rust: add a new binary object map formatbrian m. carlson, Feb 7, 2026
  109. 08/16 csum-file: define hashwrite's count as a uint32_tbrian m. carlson, Feb 7, 2026
  110. 06/16 hash: add a function to look up hash algo structsbrian m. carlson, Feb 7, 2026
  111. 11/16 rust: fix linking binaries with cargobrian m. carlson, Feb 7, 2026
  112. 12/16 rust: add a build.rs script for testsbrian m. carlson, Feb 7, 2026
  113. 10/16 hash: expose hash context functions to Rustbrian m. carlson, Feb 7, 2026
  114. 05/16 rust: add a hash algorithm abstractionbrian m. carlson, Feb 7, 2026
  115. 09/16 write-or-die: add an fsync component for the object mapbrian m. carlson, Feb 7, 2026
  116. 13/16 rust: add functionality to hash an objectbrian m. carlson, Feb 7, 2026
  117. 15/16 rust: add a small wrapper around the hashfile codebrian m. carlson, Feb 7, 2026
  118. 16/16 object-file-convert: always make sure object ID algo is validbrian m. carlson, Feb 7, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.