From: Phillip Wood Date: Thu, 06 Nov 2025 09:55:31 GMT Subject: Re: [PATCH v2 01/10] doc: define unambiguous type mappings across C and Rust Message-ID: <995f77a3-b94c-46df-87d3-22c7b2a3c762@gmail.com> In-Reply-To: <88133848d1a317f8a95c19ee5482b828a3f8705f.1761776388.git.gitgitgadget@gmail.com> Hi Ezekiel On 29/10/2025 22:19, Ezekiel Newren via GitGitGadget wrote: > From: Ezekiel Newren > > Document other nuances with crossing the FFI boundary. Other language > mappings may be added in the future. Thanks for adding this, I've left a few comments below. Overall I thought it was very well written. I tried building an html version of this but even after adding it to the list of TECH_DOCS in Documentation/Makefile with diff --git a/Documentation/Makefile b/Documentation/Makefile index 47208269a2e..2699f0b24af 100644 --- a/Documentation/Makefile +++ b/Documentation/Makefile @@ -143,6 +143,7 @@ TECH_DOCS += technical/shallow TECH_DOCS += technical/sparse-checkout TECH_DOCS += technical/sparse-index TECH_DOCS += technical/trivial-merge +TECH_DOCS += technical/unambiguous-types TECH_DOCS += technical/unit-tests SP_ARTICLES += $(TECH_DOCS) SP_ARTICLES += technical/api-index it fails with $ make -C Documentation/ technical/unambiguous-types.html Merge branch 'ps/object-source-loose' into seen make: Entering directory '/home/phil/src/git/Documentation' GEN asciidoc.conf * new asciidoc flags ASCIIDOC technical/unambiguous-types.html asciidoc: ERROR: unambiguous-types.adoc: line 139: undefined filter attribute in command: source-highlight --gen-version -f xhtml -s {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} {args=} asciidoc: ERROR: unambiguous-types.adoc: line 162: undefined filter attribute in command: source-highlight --gen-version -f xhtml -s {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} {args=} asciidoc: ERROR: unambiguous-types.adoc: line 177: undefined filter attribute in command: source-highlight --gen-version -f xhtml -s {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} {args=} asciidoc: ERROR: unambiguous-types.adoc: line 187: undefined filter attribute in command: source-highlight --gen-version -f xhtml -s {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} {args=} asciidoc: ERROR: unambiguous-types.adoc: line 199: undefined filter attribute in command: source-highlight --gen-version -f xhtml -s {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} {args=} asciidoc: ERROR: unambiguous-types.adoc: line 213: undefined filter attribute in command: source-highlight --gen-version -f xhtml -s {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} {args=} asciidoc: ERROR: unambiguous-types.adoc: line 224: undefined filter attribute in command: source-highlight --gen-version -f xhtml -s {language} {src_numbered?--line-number=' '} {src_tab?--tab={src_tab}} {args=} make: *** [Makefile:396: technical/unambiguous-types.html] Error 1 make: *** Deleting file 'technical/unambiguous-types.html' make: Leaving directory '/home/phil/src/git/Documentation' > +== Character types > + > +This is where C and Rust don't have a clean one-to-one mapping. A C `char` is > +an 8-bit type that is signless (neither signed nor unsigned) I found this a bit confusing. Isn't the signedness of "char" implementation defined rather than it being "signless" > which causes > +problems with e.g. `make DEVELOPER=1`. I'm not sure what this is referring to - maybe -Wsign-compare? > Rust's `char` type is an unsigned 32-bit > +integer that is used to describe Unicode code points. Even though a C `char` > +is the same width as `u8`, `char` should be converted to u8 where it is > +describing bytes in memory. I'm dreading the point where we start sharing "struct strbuf" with rust and have to change the "buf" member from "char*" to "uint8_t*". While it is not used in the xdiff code it is ubiquitous everywhere else and there are lots of places where be pass the "buf" member to functions expecting a "char*". git grep -E '(\.|->)buf\W' has over 4000 matches > If a C `char` is not describing bytes, then it > +should be converted to a more accurate unambiguous type. That's a good point. > +While you could specify `char` in the C code and `u8` in Rust code, it's not as > +clear what the appropriate type is, but it would work across the FFI boundary. > +However the bigger problem comes from code generation tools like cbindgen and > +bindgen. When cbindgen see u8 in Rust it will generate uint8_t on the C side > +which will cause differ in signedness warnings/errors. Similaraly if bindgen > +see `char` on the C side it will generate `std::ffi::c_char` which has its own > +problems. Yeah, we definitely don't want to be using "std::ffi::c_char" in our rust implementations. I do wonder if we might want to use it (or CStr) judiciously in function parameters and immediately convert it to u8 in the function body where the function is called from C though. > +=== Notes > +^1^ This is only true if stdbool.h (or equivalent) is used. + > +^2^ C does not enforce IEEE-754 compatibility, but Rust expects it. If the > +platform/arch for C does not follow IEEE-754 then this equivalence does not > +hold. Also, it's assumed that `float` is 32 bits and `double` is 64, but > +there may be a strange platform/arch where even this isn't true. + > +^3^ C also defines uintptr_t, but this should not be used in Git. + > +^4^ C also defines ssize_t and intptr_t, but these should not be used in Git. + [u]intptr_t and ssize_t are used in git already. As Junio has pointed out there are sane uses for these types but we don't want to use them in structs or function parameters where the struct or function is shared with rust. > + > +== Problems with std::ffi::c_* types in Rust > +TL;DR: They're not guaranteed to match C types for all possible C > +compilers/platforms/architectures. Is this official policy of the rust project? Thanks Phillip > +Only a few of Rust's C FFI types are considered safe and semantically clear to > +use: + > + > +* `c_void` > +* `CStr` > +* `CString` > + > +Even then, they should be used sparingly, and only where the semantics match > +exactly. > + > +The std::os::raw::c_* (which is deprecated) directly inherits the problems of > +core::ffi, which changes over time and seems to make a best guess at the > +correct definition for a given platform/target. This probably isn't a problem > +for all platforms that Rust supports currently, but can anyone say that Rust > +got it right for all C compilers of all platforms/targets? > + > +On top of all of that we're targeting an older version of Rust which doesn't > +have the latest mappings. > + > +To give an example: c_long is defined in > +footnote:[https://doc.rust-lang.org/1.63.0/src/core/ffi/mod.rs.html#175-189[c_long in 1.63.0]] > +footnote:[https://doc.rust-lang.org/1.89.0/src/core/ffi/primitives.rs.html#135-151[c_long in 1.89.0]] > + > +=== Rust version 1.63.0 > + > +[source] > +---- > +mod c_long_definition { > + cfg_if! { > + if #[cfg(all(target_pointer_width = "64", not(windows)))] { > + pub type c_long = i64; > + pub type NonZero_c_long = crate::num::NonZeroI64; > + pub type c_ulong = u64; > + pub type NonZero_c_ulong = crate::num::NonZeroU64; > + } else { > + // The minimal size of `long` in the C standard is 32 bits > + pub type c_long = i32; > + pub type NonZero_c_long = crate::num::NonZeroI32; > + pub type c_ulong = u32; > + pub type NonZero_c_ulong = crate::num::NonZeroU32; > + } > + } > +} > +---- > + > +=== Rust version 1.89.0 > + > +[source] > +---- > +mod c_long_definition { > + crate::cfg_select! { > + any( > + all(target_pointer_width = "64", not(windows)), > + // wasm32 Linux ABI uses 64-bit long > + all(target_arch = "wasm32", target_os = "linux") > + ) => { > + pub(super) type c_long = i64; > + pub(super) type c_ulong = u64; > + } > + _ => { > + // The minimal size of `long` in the C standard is 32 bits > + pub(super) type c_long = i32; > + pub(super) type c_ulong = u32; > + } > + } > +} > +---- > + > +Even for the cases where C types are correctly mapped to Rust types via > +std::ffi::c_* there are still problems. Let's take c_char for example. On some > +platforms it's u8 on others it's i8. > + > +=== Subtraction underflow in debug mode > + > +The following code will panic in debug on platforms that define c_char as u8, > +but won't if it's an i8. > + > +[source] > +---- > +let mut x: std::ffi::c_char = 0; > +x -= 1; > +---- > + > +=== Inconsistent shift behavior > + > +`x` will be 0xC0 for platforms that use i8, but will be 0x40 where it's u8. > + > +[source] > +---- > +let mut x: std::ffi::c_char = 0x80; > +x >>= 1; > +---- > + > +=== Equality fails to compile on some platforms > + > +The following will not compile on platforms that define c_char as i8, but will > +if it's u8. You can cast x e.g. `assert_eq!(x as u8, b'a');`, but then you get > +a warning on platforms that use u8 and a clean compilation where i8 is used. > + > +[source] > +---- > +let mut x: std::ffi::c_char = 0x61; > +assert_eq!(x, b'a'); > +---- > + > +== Enum types > +Rust enum types should not be used as FFI types. Rust enum types are more like > +C union types than C enum's. For something like: > + > +[source] > +---- > +#[repr(C, u8)] > +enum Fruit { > + Apple, > + Banana, > + Cherry, > +} > +---- > + > +It's easy enough to make sure the Rust enum matches what C would expect, but a > +more complex type like. > + > +[source] > +---- > +enum HashResult { > + SHA1([u8; 20]), > + SHA256([u8; 32]), > +} > +---- > + > +The Rust compiler has to add a discriminant to the enum to distinguish between > +the variants. The width, location, and values for that discriminant is up to > +the Rust compiler and is not ABI stable.