If you’ve ever tried to extract useful data from SEC filings, you know the experience sits somewhere between reading hieroglyphics and assembling IKEA furniture without the manual. The documents are dense, inconsistently formatted, and designed for human lawyers, not machine learning models.

A team from Stanford’s Advanced Financial Technologies Lab just dropped something that could change that. The Stanford EDGAR Filings Dataset, or SEFD, is a massive reconstruction of US SEC EDGAR filings spanning from 1994 to the present, reformatted into a layout-faithful MultiMarkdown style that machines can actually parse without losing the financial meaning buried in the structure.

What makes this dataset different

The initial public snapshot contains 152 billion tokens covering filings from January 2022 to June 2025. The full dataset, when complete, is estimated to reach roughly 550 billion tokens drawn from approximately 18.5 million filings.

The project was led by Nick Bettencourt, affiliated with UCLA and collaborating with Stanford. It was announced on June 16, 2026.