RDM Weekly - Issue 053
A weekly roundup of Research Data Management resources.
Welcome to Issue 53 of the RDM Weekly Newsletter!
The content of this newsletter is divided into 4 categories:
✅ What’s New in RDM?
These are resources that have come out within the last year or so
✅ Oldies but Goodies
These are resources that came out over a year ago but continue to be excellent ones to refer to as needed
✅ Research Data Management Job Opportunities
Research data management related job opportunities that I have come across in the past week
✅ Just for Fun
A data management meme or other funny data management content
What’s New in RDM?
Resources from the past year
1. Data Analysis Journeys
This open access book contains data analysis journeys to practice your wrangling, visualisation, and analysis skills (using R) in a less structured format than the main PsyTeachR books. The authors have designed these tasks as a bridge between the structured learning in the core course book chapters and your assessments. The authors present you with a new data set, show you what the end product should look like, and see if you can apply your data wrangling, visualisation, and/or analysis skills to get there. Data analysis is all about seeing the data you have available to you and identifying what the end product needs to be to apply your visualisation and analysis techniques. You can then mentally (or physically) create a checklist of tasks to work backwards to get there. There might be a lot of trial and error as you try one thing, it does not quite work, so you go back and try something else. If you get stuck though, the book provides a range of hints and task lists that you can unhide, then the solution to check your attempts against.
2. Where’s the Evidence that Respondents Understand Your Survey Questions?
Survey researchers try to align respondent question interpretations with their own by applying theoretical rules, such as writing questions simply, concretely, and without biasing words or double-barreled inquiries, or via post hoc plausibility checks. Few, however, conduct ex ante empirical evaluations of their application of these rules. The widely recognized best empirical approach is “cognitive debriefing,” that follows a survey question with a detailed conversation with the respondent. Unfortunately, this time-consuming procedure is rarely used, especially for novel surveys on online platforms. This paradoxically leaves the veracity of the discipline’s most used empirical methodology depending largely on theoretical assumptions. For cognitive debriefing to live up to its intended potential, the authors of this study first formalize the essential interpretation-related assumptions underlying all survey questions. They then automate cognitive debriefing to satisfy these assumptions using a specially-tuned chatbot, apply it to larger numbers than heretofore possible, analyze the collection of transcripts, and summarize the interpretations it reveals. By iteratively rewriting questions and rerunning this procedure, researchers can help ensure respondents understand questions as they do. The authors also apply their procedure to some of the most common survey questions and demonstrate large disconnects between researcher intentions and respondent understandings. The authors make available easy-to-use open source software that implements their suggestions.
3. data-dict.yaml
data-dict.yaml is a data dictionary specification that describes a collection of related tables: their contents, constraints, connections, and the specialised vocabulary you need to understand them. It is designed to be a living document, co-written by humans and agents, that tracks your understanding of a dataset as it evolves. data-dict.yaml is designed to be lightweight. It doesn’t attempt to precisely describe every possible type of metadata in a machine-readable way. Instead it focuses on precisely recording the most important components, leaving the remainder to plain text fields that require a human or agent to interpret. This means that data-dict.yaml doesn’t itself do data cleaning, but it is a useful complement to tools that do.
4. DataCite DOI Re-Curation Watch
DataCite DOI metadata should evolve over time as new related identifiers and resource types emerge, and the authors point out that a robust versioning/provenance system for tracking these changes already exists within DataCite's activities API — repositories don't need to build one themselves. Because that raw API output is hard to browse, Metadata Game Changers built a "Re-Curation Watch" tool that lets users pull recently updated records from a repository and see a simple history of what changed and when. The tool is now open for public testing, with the authors inviting feedback and sharing examples of repositories already using metadata updates as a best practice.
5. Eleven Strategies for Making Reproducible Research and Open Science Training the Norm at Research Institutions
Reproducible research and open science practices have the potential to accelerate scientific progress by allowing others to reuse research outputs, and by promoting rigorous research that is more likely to yield trustworthy results. However, these practices are uncommon in many fields, so there is a clear need for training that helps and encourages researchers to integrate reproducible research and open science practices into their daily work. Here, the authors outline eleven strategies for making training in these practices the norm at research institutions. The strategies, which emerged from a virtual brainstorming event organized in collaboration with the German Reproducibility Network, are concentrated in three areas: (i) adapting research assessment criteria and program requirements; (ii) training; (iii) building communities. The authors provide a brief overview of each strategy, offer tips for implementation, and provide links to resources. They also highlight the importance of allocating resources and monitoring impact. Their goal is to encourage researchers – in their roles as scientists, supervisors, mentors, instructors, and members of curriculum, hiring or evaluation committees – to think creatively about the many ways they can promote reproducible research and open science practices in their institutions.
6. Modern Python Project Management for Researchers with uv
The Python programming language is enjoying great popularity in Scientific Computing due to its ease of use and its vast ecosystem of available software packages. However, due to the explosive growth of the number of available Python packages, complex dependency trees are becoming an increasing concern for the maintainability and reproducibility of Python software. This workshop provides a practical guide to uv, a powerful tool designed to facilitate the whole experience of working with Python code and dependencies. The 90 minute online workshop, happening this Thursday July 16th, will introduce the core functionalities of uv, from the efficient execution of standalone scripts and the management of project environments using lockfiles to structuring, building, and shipping your own project for public distribution. By the end of the course, researchers will have a streamlined toolkit to ensure their computational work is both sustainable and easily shareable with the scientific community.
7. ECR Asks: Dr. Ed Ivimey-Cook
‘ECR Asks’ is a series of Q&A sessions where Oakleigh Wilson speaks with experienced SORTEE members to explore their journey in, and perspectives on, open science and transparent research. With the goal of supporting early career researchers, this series aims to answer big questions and share practical insights on navigating the ever evolving landscape of open science and academia. In this conversation, Ed shares his time, experience, and tips. This conversation was an insightful glimpse into the challenges and future of code review with valuable advice for open science focused ECRs.
Oldies but Goodies
Older resources that are still helpful
1. National Library of Medicine Data Glossary
This resource connects and defines concepts, services, and tools relevant to librarians working in data-driven discovery. A definition, relevant literature, and web resources accompany each term along with links to related terms. Search by term or keyword or browse the terms listed. The Data Glossary is created and maintained by the NCDS.
2. GREI Collaborative Webinar Series on Data Sharing in Generalist Repositories
On this page you can access resources from a series of four presentations and panel discussions by generalist repositories to learn about available repository resources and best practices for sharing NIH-funded research. Presented by the members of the NIH Generalist Repository Ecosystem Initiative (GREI): Dryad, Dataverse, Figshare, Mendeley Data, Open Science Framework, and Vivli.
3. Open Science + Data Management Policies
This resource contains slides from a training webinar on Open Science practices in LD research, put together by the Florida Learning Disabilities Research Center. Topics covered include effective management of research data, federal data policies, and data documentation.
4. Cultural Obstacles to Research Data Management and Sharing at TU Delft
Research data management (RDM) is increasingly important in scholarship. Many researchers are, however, unaware of the benefits of good RDM and unsure about the practical steps they can take to improve their RDM practices. Delft University of Technology (TU Delft) addresses this cultural barrier by appointing Data Stewards at every faculty. By providing expert advice and increasing awareness, the Data Stewardship project focuses on incremental improvements in current data and software management and sharing practices. This cultural change is accelerated by the Data Champions who share best practices in data management with their peers. The Data Stewards and Data Champions build a community that allows a discipline-specific approach to RDM. Nevertheless, cultural change also requires appropriate rewards and incentives. While local initiatives are important, and the authors discuss several examples in this paper, systemic changes to the academic rewards system are needed. This will require collaborative efforts of a broad coalition of stakeholders and the authors will mention several such initiatives. This article demonstrates that community building is essential in changing the code and data management culture at TU Delft.
Research Data Management Job Opportunities
These are data management job opportunities that I have seen posted in the last week. I have no affiliation with these organizations.
Just for Fun
Sponsor
This newsletter is supported in part by the Eunice Kennedy Shriver National Institute Of Child Health & Human Development of the National Institutes of Health under Award Number R25HD114368. The content is solely the responsibility of the author and does not necessarily represent the official views of the National Institutes of Health. Read more about the NIH Data Management for Data Sharing Workshop Project.
Thank you for reading! If you enjoy this content, please like, comment, or share this post! You can also support this work through Buy Me A Coffee.



