git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: GDPR compliance best practices?

From
Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Date
Jun 3, 2018, 19:48 UTC
Message-ID
<87sh638xdr.fsf@evledraar.gmail.com>
In-Reply-To
<20180603141801.GA8898@helen.PLASMA.Xg8.DE>
On Sun, Jun 03 2018, Peter Backes wrote:
Show 11 quoted lines
> On Sun, Jun 03, 2018 at 02:59:26PM +0200, Ævar Arnfjörð Bjarmason wrote:
>> I'm not trying to be selfish, I'm just trying to counter your literal
>> reading of the law with a comment of "it'll depend".
>>
>> Just like there's a law against public urination in many places, but
>> this is applied very differently to someone taking a piss in front of
>> parliament v.s. someone taking a piss in the forest on a hike, even
>> though the law itself usually makes no distinction about the two.
>
> We have huge companies using git now. This is not the tool used by a
> few kernel hackers anymore.

Sure, but what I'm pointing out is a) you can't focus on git as the technology because it tells you nothing about what's being done with it (e.g. the log file case I mentioned b) nobody who came up with the GDPR was concerned with some free software projects or the SCM used by companies, so this is very unlikely to be enforced.

Show 10 quoted lines
>> In this example once you'd delete the UUID ref you don't have the UUID
>> -> author mapping anymore (and b.t.w. that could be a many to one
>> mapping).
>
> It is not relevant whether you have that mapping or not, it is enough
> that with additional information you could obtain it. For example, say,
> you have 5000 commits with the same UUID. Now your delete the mapping.
> But your friend still has it on his local copy. Now your friendly
> merely needs to tell you who is behind that UUID and instantly you can
> associate all 5000 commits with that person again.

So nobody can be GDPR compliant in the face of archive.org and the like? If the law says that you need to delete information you published in the past, and you do so, how is it your problem that someone mirrored & re-published it? That's their compliance problem at that point.

Show 6 quoted lines
> The GDPR is very explict about this, see recital 26. It says that
> pseudonymization is not enough, you need anonymization if you want to
> be free from regulation.
>
> In addition, and in contrast to my proposal, your solution doesn't
> allow verification of the author field.

It does if you've got the ref. Maybe I just don't get your proposal, quote:

    Do not hash anything directly to obtain the commit ID. Instead, hash a
    list of hashes of [$random_number, $information] pairs. $information
    could be an author id, a commit date, a comment, or anything else. Then
    store the commit id, the list of hashes, and the list of pairs to form
    the commit.

You're just proposing (if I've read this correctly) that the commit object should have some list of headers pointing to other SHA1s, and that fsck and the like be OK with these going away. Right?

How is this intrinsically different from referring to something in the ref namespace that may be deleted in the future?

In both cases you're just trying to solve the problem of trying to somehow encode data into a git repository today, that may go away tomorrow. Similar to how a reference to some LFS object today going away doesn't fail "git fsck".

Show 20 quoted lines
>> I think again that this is taking too much of a literalist view. The
>> intent of that policy is to ensure that companies like Google can't just
>> close down their EU offices weasel out of compliance be saying "we're
>> just doing business from the US, it doesn't apply to us".
>>
>> It will not be used against anyone who's taking every reasonable
>> precaution from doing business with EU customers.
>
> I think you are underestimating the political intention behind the
> GDPR. It has kind of an imperialist goal, to set international
> standards, to enforce them against foreign companies and to pressure
> other nations to establish the same standards.
>
> If I would read the GPDR in a literal sense, I would in fact come to
> the same conclusion as you: It's about companies doing substantial
> business in the EU. But the GDPR is carefully constructed in such a way
> that it is hard not to be affected by the GDPR in one way or another,
> and the obvious way to cope with that risk is to more or less obey the
> GDPR rules even if one does not have substantial business interests in
> the EU.

Okey, so you're not reading the GDPR in some literal sense, but you're coming to a conclusion that's supported by ... what? To echo Theodore Y. Ts'o E-Mail have you consulted with someone who's an actual lawyer on this subject?

I haven't but, I'm not suggesting that the git data format needs to change because of some new EU law. You are, what's your basis for that opinion?

It seems to me that the git project doesn't need to do anything about this. There's plenty of things that are illegal to publish, and some of which may be made illegal after the fact (e.g. national security related information). If those things are incidentally saved in git repositories the parties involved may need to run git-filter-branch.

Of course if they need to do that on a weekly basis because of some overzealous law we may need to have some "native" support for that, but I see zero signs of that so far.

Previous: Philip OakleyNext: Peter Backes
Message 50 of 53 in “GDPR compliance best practices?”
  1. Peter BackesApr 17, 2018
  2. Ævar Arnfjörð BjarmasonApr 17, 2018
  3. Peter BackesApr 17, 2018
  4. Peter BackesJun 3, 2018
  5. Ævar Arnfjörð BjarmasonJun 3, 2018
  6. Peter BackesJun 3, 2018
  7. Ævar Arnfjörð BjarmasonJun 3, 2018
  8. Peter BackesJun 3, 2018
  9. Philip OakleyJun 3, 2018
  10. Peter BackesJun 3, 2018
  11. Theodore Y. Ts'oJun 3, 2018
  12. Peter BackesJun 3, 2018
  13. Peter BackesJun 3, 2018
  14. Theodore Y. Ts'oJun 3, 2018
  15. Peter BackesJun 3, 2018
  16. Theodore Y. Ts'oJun 3, 2018
  17. Peter BackesJun 3, 2018
  18. Theodore Y. Ts'oJun 4, 2018
  19. Peter BackesJun 4, 2018
  20. Philip OakleyJun 3, 2018
  21. Peter BackesJun 3, 2018
  22. Philip OakleyJun 4, 2018
  23. David LangJun 7, 2018
  24. Peter BackesJun 7, 2018
  25. Philip OakleyJun 7, 2018
  26. Peter BackesJun 7, 2018
  27. David LangJun 7, 2018
  28. Peter BackesJun 7, 2018
  29. David LangJun 7, 2018
  30. Peter BackesJun 8, 2018
  31. David LangJun 8, 2018
  32. Peter BackesJun 8, 2018
  33. David LangJun 8, 2018
  34. David LangJun 12, 2018
  35. Peter BackesJun 12, 2018
  36. Martin FickJun 12, 2018
  37. Theodore Y. Ts'oJun 13, 2018
  38. Peter BackesJun 13, 2018
  39. Theodore Y. Ts'oJun 8, 2018
  40. Peter BackesJun 8, 2018
  41. Ævar Arnfjörð BjarmasonJun 8, 2018
  42. Peter BackesJun 8, 2018
  43. Ævar Arnfjörð BjarmasonJun 8, 2018
  44. Theodore Y. Ts'oJun 8, 2018
  45. Peter BackesJun 8, 2018
  46. Johannes SixtJun 8, 2018
  47. Philip OakleyJun 9, 2018
  48. Theodore Y. Ts'oJun 10, 2018
  49. Philip OakleyJun 3, 2018
  50. Ævar Arnfjörð BjarmasonJun 3, 2018
  51. Peter BackesJun 3, 2018
  52. Jonathan NiederJun 8, 2018
  53. Ævar Arnfjörð BjarmasonJun 8, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.