From: Simon Richter Date: Sun, 07 Dec 2025 05:26:46 GMT Subject: Re: Git for structured data Message-ID: <2ae0a2d5-e909-4c51-9459-83f5c6950d51@hogyros.de> In-Reply-To: Hi, On 12/6/25 01:51, Cedric Sodhi wrote: > Why can't we have structured, version controlled data? You can version control inside a relational database, by adding valid time columns with a range-between-timestamps type and a constraint to disallow overlaps. There are good indexing techniques, the first thing that springs to mind is [1], but I'm fairly sure there are others, and a modern RDBMS should provide constraints on range types. A valid time column can encode either "time at which the data is valid", or "time at which the data was current in the database", with two columns, you can encode both at the same time. If you hide the "data is current within" column behind a view and automatically update it, this creates the historical log of when an entry was updated. Tracking arbitrary data in git is, of course, also possible, but requires diff/merge tools adequate for the data. The built-in tools are adequate for the main use case, text files that usually change on a line-by-line basis and are seldom reorganized as a whole, so we can pretend they are one-dimensional. In KiCad, the files we generate describe a three-dimensional structure. No matter how we normalize the file contents, elements can only be moved on one axis without requiring us to move them to a different position in the file. So if I sort by z,y,x, then moving an object to a different z coordinate likely results in "deletion" of the old object at the existing place, and "creation" of a new object at a different place in the file, the one-dimensional diff algorithm is unable to create a minimal diff here that shows that only the z coordinate changed. Not sorting (i.e. leaving elements in creation order) means that deleting and recreating an object with the same parameters causes it to move within the file. The solution is to treat the serialized representation as just that, a serialization, and not try to interpret order in any meaningful way, but this requires dedicated diff/patch tools and heuristics that guess whether deleting and creating similar objects constitutes a move or if the objects are unrelated, same as git does in its move detection. I think that diff/merge on relational data is more difficult than expressing history inside the relational tables. For other data structures, this may be different, and git might be a viable storage method for history -- but in any case it requires the effort to build an appropriate plug-in. Simon [1] https://link.springer.com/chapter/10.1007/BFb0054512