
Librarians and staff at the UC Berkeley Library work to ensure our students, researchers, and faculty have access to the journals and books they need. But what about data? Our Data and Digital Scholarship Services team purchases and licenses datasets that span the university’s numerous disciplines. Many of these datasets can be found in Dataverse, an open-source repository where UC Berkeley and Lawrence Berkeley National Lab affiliates can search for data, access platforms, and download files.
Getting started is simple. Navigate to UC Berkeley Library’s Dataverse and log in using your Calnet ID and passphrase in the upper right hand corner. You will be asked to review and agree to our general terms of use. Note that individual datasets may have additional terms.
A few datasets to check out:
- Web of Science XML data: metadata from over 63 million scholarly article records spanning multiple disciplines (1900 – 2025)
- Linguistic Data Consortium datasets: hundreds of datasets from the University of Pennsylvania’s Linguistic Data Consortium supporting linguistic education and research.
- RateWatch Scholar: financial data covering banks, credit unions, and savings and loan associations
- Historical Newspapers: San Francisco Chronicle (and other titles!) from 1865 – 1922 in machine readable format
Questions or suggestions?
If you have questions on how you may or may not use the data in Dataverse, especially related to AI, email tdm-access@berkeley.edu. And if you have a dataset that you would like the library to purchase, please fill out our purchase request form or email librarydataservices@berkeley.edu.
Check back for our October spotlight to learn about ICPSR!




