git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: GDPR compliance best practices?

From
Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Date
Jun 3, 2018, 10:45 UTC
Message-ID
<87vab087y2.fsf@evledraar.gmail.com>
In-Reply-To
<20180603092736.GA5510@helen.PLASMA.Xg8.DE>
On Sun, Jun 03 2018, Peter Backes wrote:
> Unfortunatly this important topic of GDPR compliance has not seen much
> interest.

I don't think you can infer that there's not much interest, but maybe people just don't have anything to say about it.

There's a lot of discussions about this that I've seen, but what they all have in common is that nobody really knows. Just like nobody really knew what the "cookie law" would be like.

So I think all of us are just waiting to see.

I took the bite and tried to paraphrase some stuff I've read about it, but as you pointed out in 20180417232504.GA4626@helen.PLASMA.Xg8.DE I incorrectly surmised some stuff, although I very much suspect that *in practice* the GDPR is going to be more about "consumer protection". I.e. regulators / prosecutors are much likely to go after some advertising company than some project using a Git repo.

Just like nobody's going after some local computer club's internal-only website because it sets cookies without asking, but they might go after Facebook for doing the same.

Show 18 quoted lines
> [...]
> In course of this, anonymization could also be added. My idea would be
> as follows:
>
> Do not hash anything directly to obtain the commit ID. Instead, hash a
> list of hashes of [$random_number, $information] pairs. $information
> could be an author id, a commit date, a comment, or anything else. Then
> store the commit id, the list of hashes, and the list of pairs to form
> the commit.
>
> If someone requests erasure, simply empty the corresponding pair in the
> list. All that would be left would be the hash of the pair, which is
> completely anonymous (not more useful than a random number) and thus
> not covered by the GDPR. The history could still be completely
> verified, and when displaying the log, the erased entry could be
> displayed as "<<ERASED>>".
>
> What do you think about this?

Since the Author is free-form this sort of thing doesn't need to be part of the git data format. You can just generate a UUID like "5c679eda-b4e5-4f35-b691-8e13862d4f79" and then set user.name to "refval:5c679eda-b4e5-4f35-b691-8e13862d4f79" and user.email to "refval:5c679eda-b4e5-4f35-b691-8e13862d4f79".

Then you'd create a ref on the server like refs/refval/5c679eda-b4e5-4f35-b691-8e13862d4f79 containing the real "$user <$email>". If you then wanted to erase that field you'd just delete the ref, and it would be much easier to teach stuff that renders the likes of git-log to lookup these refs than changing the data format.

Sites that are paranoid about the GDPR could have a pre-receive hook rejecting any pushes from EU customers unless their commits were in this format.

Perhaps some variation of this is where the GDPR v2 will go. It'll be an "obligation to be forgotten", and I won't be allowed to use my own name anymore. Instead I'll have a daily UUID issued from a government API to use on various forms, and the only way for anyone to resolve that will be going through a webservice that'll reject UUID lookups older than N months, caching those requests will be met with the death penalty. We'll all be free at last.

Okey, that last paragraph is just trolling, but I think that refval: -> ref convention is something worth considering if things *really* go in this direction.

Previous: Peter BackesNext: Peter Backes
Message 5 of 53 in “GDPR compliance best practices?”
  1. Peter BackesApr 17, 2018
  2. Ævar Arnfjörð BjarmasonApr 17, 2018
  3. Peter BackesApr 17, 2018
  4. Peter BackesJun 3, 2018
  5. Ævar Arnfjörð BjarmasonJun 3, 2018
  6. Peter BackesJun 3, 2018
  7. Ævar Arnfjörð BjarmasonJun 3, 2018
  8. Peter BackesJun 3, 2018
  9. Philip OakleyJun 3, 2018
  10. Peter BackesJun 3, 2018
  11. Theodore Y. Ts'oJun 3, 2018
  12. Peter BackesJun 3, 2018
  13. Peter BackesJun 3, 2018
  14. Theodore Y. Ts'oJun 3, 2018
  15. Peter BackesJun 3, 2018
  16. Theodore Y. Ts'oJun 3, 2018
  17. Peter BackesJun 3, 2018
  18. Theodore Y. Ts'oJun 4, 2018
  19. Peter BackesJun 4, 2018
  20. Philip OakleyJun 3, 2018
  21. Peter BackesJun 3, 2018
  22. Philip OakleyJun 4, 2018
  23. David LangJun 7, 2018
  24. Peter BackesJun 7, 2018
  25. Philip OakleyJun 7, 2018
  26. Peter BackesJun 7, 2018
  27. David LangJun 7, 2018
  28. Peter BackesJun 7, 2018
  29. David LangJun 7, 2018
  30. Peter BackesJun 8, 2018
  31. David LangJun 8, 2018
  32. Peter BackesJun 8, 2018
  33. David LangJun 8, 2018
  34. David LangJun 12, 2018
  35. Peter BackesJun 12, 2018
  36. Martin FickJun 12, 2018
  37. Theodore Y. Ts'oJun 13, 2018
  38. Peter BackesJun 13, 2018
  39. Theodore Y. Ts'oJun 8, 2018
  40. Peter BackesJun 8, 2018
  41. Ævar Arnfjörð BjarmasonJun 8, 2018
  42. Peter BackesJun 8, 2018
  43. Ævar Arnfjörð BjarmasonJun 8, 2018
  44. Theodore Y. Ts'oJun 8, 2018
  45. Peter BackesJun 8, 2018
  46. Johannes SixtJun 8, 2018
  47. Philip OakleyJun 9, 2018
  48. Theodore Y. Ts'oJun 10, 2018
  49. Philip OakleyJun 3, 2018
  50. Ævar Arnfjörð BjarmasonJun 3, 2018
  51. Peter BackesJun 3, 2018
  52. Jonathan NiederJun 8, 2018
  53. Ævar Arnfjörð BjarmasonJun 8, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.