How to increase the reusability of research outputs?
Sure, Open Science has gained some momentum during the last decade. The Horizon Europe programme supports a number of open infrastructure projects, one of them being AquaINFRA, and the next call is already out. The German National Research Data Infrastructure (NFDI) has just announced that funding will continue until 2038. And let’s not forget the local initiatives in universities, such as Open Science Communities, reproducibility networks, and research data management centers. We also talked about 52°North’s efforts in our blog post on Open Science Capacity Building last week. Ultimately, however, we haven’t yet tapped into the full potential when it comes to publishing research outputs. Let’s say we collect some data and analyze it in an R script. We then implement functions for data pre-processing, analysis, and visualization. The configuration of the function parameters is buried in the code and the computational environment is specified on our local computer. We write a paper and send it to a journal or conference for peer review. At best, the submission requirements ask us to publish all materials in a repository and paste the DOI into the article. While we should appreciate that this is more than it was a few years ago, we also need to acknowledge that this is not the most efficient way to share high-value and easily reusable research outputs such as data and source code. Others need to download the materials, install the right version of the libraries, understand the code, find out how to change the parameters and so on and so forth. Many people tend to avoid these efforts and instead develop something new from scratch. So the challenge we want to address in this week’s blog post is:
How to increase the reusability of research outputs?
We addressed this issue in the EU-funded project AquaINFRA and finally came up with the so-called Data-to-Knowledge Package (D2K-Package) to tackle the following requirements:
- Verifying and reproducing the research results presented in an article
- Exploring the analysis and evaluating its relevance and reusability
- Reusing parts of the analysis in a custom script
- Understanding the role of each function in an analysis, which input parameters are needed, and how the functions are connected
The main objective of the D2K-Package is to enhance the reusability of computational research outputs by enabling users to seamlessly trace the process from raw data to analysis and knowledge generation. It contains everything that is needed to reproduce the analysis plus additional products that make reuse and interaction with the analysis easier (see Fig. 1). As presented here, the D2K-Package has five components, but it is extensible. Simply speaking, it is a collection of links to digital research assets that together unlock the potential of open reproducible research and open FAIR data.

Let’s start with the first component: Data. Since we want to promote transparency and reusability, data should follow the FAIR principles (findable, accessible, interoperable, reusable) and be available under an open license, e.g., CC0 or alike. Ideally it is made available via a repository providing DOIs and an open data format, though binary formats are also fine if its format is released under an open specification.
Next, the analysis (e.g., written in R or Python) is made available in a Toolbox, which can be also stored in a repository in order to retrieve a DOI. The toolbox contains the entire analysis pipeline split into separate, containerized, and self-containing functions following the input-processing-output mechanism. Data and a parameter configuration (input) go into the function, the data is processed or analyzed (processing), and the result is the processed data, for instance, a table or figure (output). Hence, every function fulfills a certain task thus increasing reusability. We know this mechanism from R libraries or Python packages. The functions are designed such that the output of one function becomes the input of the next function and so forth until we reach the final result of the analysis. We will use this mechanism later to create a workflow. For the containerization of the analysis including the computational environment (versions of the libraries and the runtime), we make use of Docker. Data and Toolbox form the reproducible basis for the next products that aim at supporting the reuse of the materials and the process from data to knowledge generation!
The third component of the D2K-Package is the Computational Workflow using the Galaxy platform. It might be difficult to find out how the functions are connected, how to change the parameter configuration, or maybe you just want to click on run to see the results quickly. For that, we use Galaxy as a platform to make the entire analysis pipeline available as a readily shareable workflow (see Figure 2). Because concrete implementation details are abstracted away, users can utilize the workflow as an initial entry point for understanding the analysis, later reviewing the underlying specifics by accessing the virtual lab or toolbox. They can also run the workflow quickly using different parameter settings and compare the results. Finally, it is also possible to reuse single steps from different workflows and combine them in a new way, provided they are interoperable. The workflow file can be exported from Galaxy and stored in a repository to receive a DOI and increase its findability through metadata. That DOI can be added to the D2K-Package.

We dive one level deeper and take a look at the fourth component, the Web API Service. You might want to reuse the functions in your own script. Copying all relevant code snippets and installing all relevant libraries in the right version might be a daunting task, so we make the toolbox functions available as OGC API Processes – a standardized web service for wrapping computational tasks into executable processes (learn more). Thus, every function can be called via an HTTP-request. This feature doesn’t come for free. We need a server running, for instance, pygeoapi, to make the web processes available. The amount of resources depends mainly on the expected size of the input and output datasets and the computation. A link to that service will also become part of the D2K-Package.
Last but not least, the Virtual Lab is the fifth component of the D2K-Package. Based on the computational environment defined in the toolbox, the virtual lab recreates that environment and provides it as a JupyterLab in the browser. No local installation of the runtime and dependencies is needed, just follow the URL and start working with the source code! We use the MyBinder web application to realize this feature, but please note that MyBinder is just a test instance with limited computational resources. An alternative to MyBinder is Replay, provided by EGI.
As mentioned above, the D2K-Package is a collection of links to the five (or more) components stored in a machine-readable and light-weight JSON-LD file. It can be easily parsed by websites and enriched with metadata. Moreover, the D2K-Package is extensible and not limited to the components described above. It could also contain a link to open educational resources, scientific articles, or webinars.
Of course, we haven’t just developed the theoretical concept of the D2K-Package, we have actively put it into practice. You can explore practical examples on our AquaINFRA Interaction Platform. Additionally, a comprehensive training video created by Sadra Matmir from the Bochum University of Applied Sciences is available online. Sadra also recently earned his Master’s degree; his thesis focused on Evaluating technology adoption by planners using a data-to-knowledge package for flood risk management. Congratulations Sadra! It goes without saying that we’re a little proud that the D2K-Package became an essential component in a thesis ;).

So, how can you get a D2K-Package for your analysis? We wrote a user guide explaining every single step, but we’re also ready to help and become part of your project. We can help you split your analysis script into separate and self-contained functions, create a virtual lab environment, and make the functions available as OGC API Processes. In addition, we’re also keen to help you put your work on the Galaxy platform in order to create readily shareable workflows. Don’t hesitate to get in touch with us!
See you next week.
References
- Konkol M, Labuce A, Domisch S et al. Encouraging reusability of computational research through Data-to-Knowledge Packages – A hydrological use case. Open Research Europe 2025, 5:123 (https://doi.org/10.12688/openreseurope.20221.3)
Leave a Reply