Home / Insights / Test data

Test data and the GDPR: why copying the production database is the most expensive shortcut

Production data in a test system is convenient, widespread and in most cases unlawful. The German Federal Commissioner for Data Protection has published a clear order of precedence. Knowing it saves more time than it costs.

It is the fastest route to a realistic test system: take a dump of the production database and load it into the test environment. Within two hours you have real customer records, real edge cases and real data volumes. No synthetic data set can match that.

And hardly any other routine in a software project is as exposed under data protection law. Not because anyone has bad intentions, but because at that moment personal data is being processed for a purpose it was not collected for, in an environment that is less well protected than production, accessible to a wider group of people, and as a rule without any deletion deadline.

Why testing is a purpose in its own right

The principles in Article 5 GDPR are uncomfortably clear on this point. Personal data may only be collected for specified, explicit and legitimate purposes and must not be further processed in a manner incompatible with those purposes. It must be adequate and limited to what is necessary. And it may be kept in identifiable form only for as long as the purpose requires.

A customer placing an order provides their data for the fulfilment of that order. They do not provide it for trialling the next release. Software testing is a different purpose, and that purpose needs its own assessment, its own legal basis and its own safeguards.

There is also a practical point that often weighs heavier than the legal one: test systems are structurally less well protected. More people have access, external staff among them. Permissions are granted more generously, because people need to be able to work. Extracts get copied locally, attached to defect tickets and carried along in backups. A data breach in a test system is just as notifiable as one in production, and the data subjects are the same people.

The BfDI's order of precedence

In June 2025 the German Federal Commissioner for Data Protection and Freedom of Information (BfDI) published a short position paper on precisely this question. It sets out a three-stage sequence of assessment:

  • First, non-personal, anonymous or synthetic data. It falls outside the GDPR and can be used for development and testing without restriction.
  • Then pseudonymised data, where the first stage is not sufficient.
  • Last, personal data unaltered, and only once test data without a personal reference is genuinely not an option for the specific test in question.

The decisive word in that formulation is “specific”. The assessment is not a one-off for the project as a whole but has to be made for each individual test. For the great majority of functional test cases the question can be answered with synthetic data. For a load test requiring a realistic data distribution, or a migration check across historically grown records, the answer looks different, and it is exactly those cases that have to be justified and documented.

The real reason for the database dump

Ask honestly why a team wants production data and the answer is rarely “convenience”. The reason is almost always coverage. Production data contains edge cases nobody would invent: the customer with an umlaut in their surname and a hyphen in their house number, the contract from a decommissioned system with a field that has not been populated since 2011, the booking across the turn of the year with a retrospective cancellation.

That is a legitimate business concern, but it is a test design problem, not a data procurement problem. The edge cases do not sit in the data; they sit in the rules that produced the data. They can be derived systematically: from boundary value analysis and equivalence partitioning, from the history of production defects, from the field formats of the legacy systems, and from a statistical analysis of the production data that describes distributions and formats without copying individual records.

The effort is incurred once. After that you have a test data set that covers the edge cases deliberately rather than by chance, that is reproducible, that can be used without a sign-off, and that does not have to be sourced afresh with every extract. A production dump, by contrast, is a matter of luck: it contains the edge cases that happen to be in it, and nobody knows which ones are missing.

Anonymisation is harder than it sounds

Anonymous data falls outside the GDPR; Recital 26 of the Regulation makes that clear. That is precisely why the bar is set high: a data set is anonymous only once attribution to a person is no longer possible, or only with disproportionate effort.

The BfDI points to the place where most anonymisation attempts fail: the quasi-identifiers. Replacing name and customer number is not enough if date of birth, postcode and contract start date remain in the record. A combination of a few individually harmless attributes identifies people reliably. And what counts is not only the means available to whoever receives the data, but also the means that could be brought to bear from outside to identify individuals.

A common mistake in practice: names are replaced with random names but the relationships between records are preserved. The data set then looks anonymous, and is not.

Pseudonymisation is not an exit from the GDPR

Pseudonymised data remains personal data. Article 4(5) GDPR defines pseudonymisation as processing after which the data can no longer be attributed to a person without the use of additional information, that additional information being kept separately. As long as the key exists, the personal reference exists.

What pseudonymisation delivers is risk reduction. Article 32 names it explicitly as a technical measure for the security of processing, and Article 25 cites it as an example of data protection by design. The European Data Protection Board issued dedicated guidelines on the subject in January 2025; the BfDI summarises their practical value by noting that pseudonymisation can make it easier for controllers to rely on legitimate interests as a legal basis.

For test planning that means: pseudonymisation is a good measure and a poor excuse. A pseudonymised production dump in a test system still needs a legal basis, an access control concept and a deletion deadline.

When production data is unavoidable

There are tests for which there is no way around real data, certain migration checks among them. At that point the task shifts from “avoid” to “control”. The BfDI names as conditions a sound legal basis, appropriate technical and organisational measures, and compliance with the general conditions for lawful processing, including data subject rights and transparency obligations.

In practice that means, among other things: a narrow, individually documented group of people with access rather than a blanket permission for the test system; a fixed deletion deadline with a mechanism that enforces it, rather than an intention; encryption at rest and in transit; no attachments containing production data in defect tickets; and a processor agreement under Article 28 as soon as an external service provider processes the data. Whether a data protection impact assessment is additionally required depends on the individual case and should be assessed, not guessed.

What this means for test planning

Test data is not a by-product of test execution but a planning item in its own right. This is why ISO/IEC/IEEE 29119-3 treats test data requirements as a distinct part of the documentation. In practice three things belong in the test plan:

  • Test data requirements per test case or test group: which attribute values are needed, which boundary values, which edge cases. Only this list shows whether synthetic data is sufficient.
  • The sourcing decision with its rationale: synthetic, anonymised, pseudonymised or production data, and why. This doubles as the documentation of the order of precedence the BfDI requires.
  • The deletion concept for the test environment: when are data sets reset, who ensures it happens, how is it evidenced. Test systems are the environments in which data sits unnoticed for the longest.

With these three points in the test plan, the discussion with the data protection officer happens once at the outset rather than again shortly before every test start.

Who does what

Qelivia specifies the test data requirements in the test plan and derives the edge cases needed from the risk picture and the test design techniques. We need no access to your production data for this, and we do not take it. Generating and providing the test data happens in your environment, as does test execution.

Equally clearly: we are not a data protection consultancy. Assessing the legal basis, carrying out a data protection impact assessment and signing off a processing activity belong with your data protection officer or a law firm. What test management contributes is the groundwork that makes such an assessment possible in the first place: a well-founded statement of which data is genuinely needed for the business case, and as a rule the evidence that it is less than assumed.

Conclusion

The production dump in the test system is a shortcut taken at the start of a project and paid for dearly at the end: with a discussion before go-live, with an environment that can no longer be cleaned up, in the worst case with a notification under Article 33. The alternative is less convenient but one-off: derive the edge cases systematically, generate them synthetically and document the sourcing decision. After that the test data set is an asset rather than a liability.

Note: This article reflects the position as at 21 August 2026 and summarises publicly available sources from the German Federal Commissioner for Data Protection, the European Data Protection Board, and the text of the Regulation. This article is not legal advice. Assessing your specific processing activity belongs with your data protection officer or in a legal review.

Sources

  1. BfDI, short position paper “Personenbezogene Daten bei Software-Entwicklung und -Tests”, 10 June 2025: bfdi.bund.de, Kurzposition Testdaten (PDF)
  2. Regulation (EU) 2016/679 (GDPR), in particular Article 4(5), Article 5, Article 25, Article 28, Article 32 and Recital 26: eur-lex.europa.eu/eli/reg/2016/679
  3. European Data Protection Board, Guidelines 01/2025 on Pseudonymisation: edpb.europa.eu, Guidelines 01/2025 (PDF)
  4. BfDI, press release “EDSA schafft mehr Klarheit bei Pseudonymisierung”, 16 January 2025: bfdi.bund.de, Pressemitteilung 1/2025
  5. ISO/IEC/IEEE 29119-3:2021, Software testing — Part 3: Test documentation (test data requirements as a documentation item): iso.org/standard/79429.html
Next step

Test data nobodyhas to justify.

We derive the edge cases you need from the risk picture and record the test data requirements in the test plan, without access to your production data. Your team runs the tests.

Explore Test Management as a Service →