Accessibility Tools

Select your language

Select your language

Vera Clemens, a software developer working on the NFDI4Health project, has become the first female external developer from Europe to be in the Core Team of Harvard University’s Dataverse Project. In an interview, she discusses the background of the Dataverse initiative, her new role, and the opportunities it offers for a FAIR research data landscape.

A young woman with brown hair and glasses against a blurry background.
Her work receives international recognition: Vera Clemens appointed as an external developer for the Dataverse Project. Copyright: ZB MED

By joining the Core Team of the Dataverse Project, Vera Clemens is taking on a key role in the further development of a global open-source platform and strengthening the international networking of NFDI4Health. We invited her to share her insights and asked her five questions

1. What is the Dataverse Project, when and by whom was it founded?
The Dataverse Project is an open-source platform for research data repositories. It serves to share, archive, cite, find, and analyze research data. Institutions can use it to operate their own repositories. The best-known instance is Harvard Dataverse. The Dataverse Project was founded in 2006 at the Institute for Quantitative Social Science (IQSS) at Harvard University and has since been primarily developed by this institution together with the open-source community and the Global Dataverse Community Consortium (GDCC).

2. What is the (scientific) gap that the Dataverse Project addresses, and how does Dataverse help to close it?
Dataverse bridges the gap between scientific publications and the underlying research data. It transforms data into an independently citable, discoverable, long-term archived, and reusable research output. In doing so, it supports Open Science, FAIR principles, reproducibility, and the recognition of data-driven work.

3. What exactly does it mean to be part of the Dataverse Core Team? What tasks are involved?
In addition to the institute's full-time developers, the software is collaboratively developed further by the open-source community. The core team consists of long-standing contributors to the open-source project who are more closely integrated into the institute's development processes. This enables a deep and strategic integration of their joint work.

4.What kind of interface is there with the NFDI (National Research Data Infrastructure in Germany)?
The Health Study Hub, which is being developed within the framework of NFDI4Health, is based on Dataverse software. The same applies to the FAIRagro Dataset Finder, which is being developed within the FAIRAgro project. The ZB MED Data Portals Working Group is actively involved in these developments. Furthermore, the Dataverse software is also used in other consortia and projects. German Dataverse users and developers regularly exchange information within the FDM.NRW Dataverse interest group. Currently, 150 Dataverse installations are known worldwide; 73 of these are located in Europe. This means that approximately half of the documented Dataverse installations are in Europe. There are currently 15 known Dataverse installations in Germany.

5. Is there anything else important to know about this topic?
Dataverse is widely used and well-established internationally as repository software for research data (see map of installations on https://dataverse.org/ as well as https://dataverse.org/metrics). In our federal state of North Rhine-Westphalia, a statewide repository for the publication of scientific research data is currently being developed as part of the DH.NRW initiative; it will also be based on DataPublication.nrw.