"I cannot live without looking over my shoulder," a sexual abuse survivor told US lawmakers in May, describing how her personal information was publicly exposed in January, when the Department of Justice (DOJ) released hundreds of thousands of documents related to the financier Jeffrey Epstein. "I can only imagine the long-term impact this mistake will have on my life."

Three days after the DOJ released the largest of 12 datasets on Epstein on January 30, a department lawyer told a federal court that "human or technical error" had been among the factors that led to the publishing of personally identifying information of people who had been sexually abused by Epstein and his associates. As a result, the DOJ said it would remove, review and redact flagged documents before republishing them.

DW's investigative unit, along with its Innovation and Data teams, spent months reviewing thousands of documents released by the DOJ. In the six months since the release, DW identified dozens of files that still contained information — including names, faces and email addresses — that could be used to identify survivors, witnesses and informants, among other people. DOJ Epstein Library

By mid-February, DW had scraped more than 800,000 files from the DOJ's Epstein Library — a challenging task with the archive constantly changing: New files were being added, others were taken down and modified files were popping up after days of absence.