One in five social science datasets shared for replication purposes fails to meet basic privacy standards, according to a new analysis that raises concerns about how researchers protect the identities of study participants. The finding suggests that promises of anonymity made to subjects may be routinely undermined when data is made public.
The study examined datasets that accompany published social science research and are intended to allow other scientists to verify results. Researchers found that a substantial share of these files contained information that could be used to re-identify individuals, even when names and other obvious identifiers had been removed. The failures appeared across multiple disciplines and journals, indicating a systemic problem rather than isolated mistakes.
Replication is a cornerstone of the scientific method, and many funders and journals now require authors to share their data. But the push for openness has outpaced the development of robust anonymization practices. Simple removal of names, addresses, and other direct identifiers is often insufficient. Combinations of seemingly innocuous details — such as age, occupation, geographic location, and income — can uniquely fingerprint individuals, especially in small or specialized samples.
The analysis did not name specific studies or researchers, but its broad scope suggests that the issue affects a wide range of work, from survey research to behavioral experiments. The consequences can be serious: participants who were promised confidentiality could face embarrassment, discrimination, or other harms if their responses become publicly linked to them. In some cases, the data may even contain sensitive information about health, finances, or political views.
The problem is compounded by the fact that many researchers lack formal training in data privacy. Anonymization is technically challenging, and best practices are not always well known. Journals and repositories often provide only general guidance, leaving authors to make judgment calls about what constitutes adequate protection. As a result, even well-intentioned researchers may inadvertently release identifiable data.
The study’s authors call for stronger safeguards, including standardized privacy reviews before data is posted, better training for researchers, and the use of formal disclosure-control methods. They also suggest that journals and funders should treat privacy protection as a requirement on par with data sharing itself. Without such measures, the drive for transparency could erode the trust that makes social science research possible.
The findings add to a growing debate about the ethics of data sharing in the social sciences. While open data can improve reproducibility and accelerate discovery, it must be balanced against the rights and expectations of study participants. The analysis indicates that current practices are falling short, and that the research community needs to act before more participants are exposed.





