26  Data and privacy

Use the workshop’s small, source-free synthetic data. It resembles survey/panel data so you can practise joins, metadata, flags, weights, missingness, and release checks. It is not SOEP, does not reproduce SOEP weights, and is not substantive evidence.

“Public” does not mean “approved for an external AI service.” Before sending data, check institutional policy, contractual restrictions, provider retention and training settings, and account controls. In this workshop, never send restricted rows, unpublished results, personal or participant data, or credentials.

Check the OpenCode Zen terms, general OpenAI data controls, and separate Codex agent approvals and security guidance. For Google/Antigravity, check the current permissions and CLI usage pages. Telemetry controls are not a confidentiality guarantee. Recheck these policies before publication.

26.1 Further reading

  • EDPB Opinion 28/2024 — The strongest public authority here for why public availability, anonymity, lawful processing, and downstream responsibility are separate questions.
  • NIST’s Generative AI Profile — A broader risk-management frame for data governance, privacy, security, evaluation, and incident response.
  • Security in GitHub Codespaces — Clarifies the boundary between isolating workshop execution and authorizing data disclosure to a model provider.