{
  "type": "Article",
  "authors": [
    {
      "type": "Person",
      "familyNames": [
        "Weber"
      ],
      "givenNames": [
        "TobiasBoege,RenéFritze,ChristianeGörgen,JeroenHanselman,DorotheaIglezakis,LarsKastner,",
        "ThomasKoprucki,TabeaH.Krause,ChristophLehrenfeld,SilviaPolla,MarcoReidelbach,ChristianRiedel,JensSaak,",
        "BjörnSchembera,KarstenTabelow,Marcus"
      ]
    }
  ],
  "description": [
    {
      "type": "Paragraph",
      "content": [
        "In this paper we discuss the notion of research data for the field of mathematics and report on the status quo of\nresearch-data management and planning. A number of decentralized approaches are presented and compared to needs and\nchallenges faced in three use cases from different mathematical subdisciplines. We highlight the importance of\ntailoring research-data management plans to mathematicians’ research processes and discuss their usage all along the\ndata life cycle.\n"
      ]
    }
  ],
  "identifiers": [],
  "references": [
    {
      "type": "Article",
      "id": "bib-bib1",
      "authors": [],
      "title": "\nR. Arp,\nB. Smith and A. D. Spear, Building Ontologies with Basic\nFormal Ontology. MIT Press, Cambridge, MA (2015) ",
      "url": "https://doi.org/10.7551/mitpress/9780262527811.001.0001"
    },
    {
      "type": "Article",
      "id": "bib-bib2",
      "authors": [],
      "title": "\nW. Bangerth and T. Heister,\nQuo vadis, scientific software? SIAM News 47, 8–7\n(2014)\n"
    },
    {
      "type": "Article",
      "id": "bib-bib3",
      "authors": [],
      "title": "\nK. Berčič, M. Kohlhase and F. Rabe, (Deep) FAIR mathematics. it –\nInformation Technology 62, 7–17 (2020) ",
      "url": "https://doi.org/10.1515/itit-2019-0028"
    },
    {
      "type": "Article",
      "id": "bib-bib4",
      "authors": [],
      "title": "\nJ. Dierkes, 4.1 Planung, Beschreibung und Dokumentation von Forschungsdaten. In\nPraxishandbuch Forschungsdatenmanagement, pp. 303–325, De Gruyter Saur,\nBerlin, Boston (2021) ",
      "url": "https://doi.org/10.1515/9783110657807-018"
    },
    {
      "type": "Article",
      "id": "bib-bib5",
      "authors": [],
      "title": "\nJ. Fehr, J. Heiland, C. Himpe and J. Saak, Best practices for replicability,\nreproducibility and reusability of computer-based experiments exemplified by\nmodel reduction software. AIMS Mathematics 1, 261–281 (2016) ",
      "url": "https://doi.org/10.3934/math.2016.3.261"
    },
    {
      "type": "Article",
      "id": "bib-bib6",
      "authors": [],
      "title": "\nJ. Fehr, C. Himpe, S. Rave and J. Saak, Sustainable research software\nhand-over. Journal of Open Research Software 9, article no. 5 (2021) ",
      "url": "https://doi.org/10.5334/jors.307"
    },
    {
      "type": "Article",
      "id": "bib-bib7",
      "authors": [],
      "title": "\nP. B. Heidorn, Shedding light on the dark data in the long tail of science.\nLibrary Trends 57, 280–299 (2008) ",
      "url": "https://doi.org/10.1353/lib.0.0036"
    },
    {
      "type": "Article",
      "id": "bib-bib8",
      "authors": [],
      "title": "\nK. Hulek, F. Müller, M. Schubotz and O. Teschke, Mathematical research data –\nan analysis through zbMATH references. EMS Newsl. 113, 54–57\n(2019) ",
      "url": "https://doi.org/10.4171/news/113/14"
    },
    {
      "type": "Article",
      "id": "bib-bib9",
      "authors": [],
      "title": "\nM. Kindling and P. Schirmbacher, ,,Die digitale Forschungswelt“ als\nGegenstand der Forschung / Research on digital research / Recherche dans la domaine de la recherche numérique. Information – Wissenschaft &\nPraxis 64, 127–136 (2013) ",
      "url": "https://doi.org/10.1515/iwp-2013-0017"
    },
    {
      "type": "Article",
      "id": "bib-bib10",
      "authors": [],
      "title": "\nT. Koprucki, K. Tabelow and I. Kleinod, Mathematical research data.\nPAMM. Proc. Appl. Math. Mech. 16, 959–960 (2016) ",
      "url": "https://doi.org/10.1002/pamm.201610458"
    },
    {
      "type": "Article",
      "id": "bib-bib11",
      "authors": [],
      "title": "\nM. Kostre, V. Sunkara, C. Schütte and N. D. Conrad, Understanding the\nromanization spreading on historical interregional networks in Northern\nTunisia. Applied Network Science 7, article no. 53 (2022) ",
      "url": "https://doi.org/10.1007/s41109-022-00492-w"
    },
    {
      "type": "Article",
      "id": "bib-bib12",
      "authors": [],
      "title": "\nM. S. Krafczyk, A. Shi, A. Bhaskar, D. Marinov and V. Stodden, Learning from\nreproducing computational results: introducing three principles and the\nReproduction Package. Philos. Trans. Roy. Soc. A 379, article no. 20200069 (2021) ",
      "url": "https://doi.org/10.1098/rsta.2020.0069"
    },
    {
      "type": "Article",
      "id": "bib-bib13",
      "authors": [],
      "title": "\nK. Lejaeghere, G. Bihlmayer, T. Björkman, P. Blaha, S. Blügel, V. Blum,\nD. Caliste, I. E. Castelli, S. J. Clark, A. Dal Corso, S. de Gironcoli,\nT. Deutsch, J. K. Dewhurst, I. Di Marco, C. Draxl, M. Dułak, O. Eriksson,\nJ. A. Flores-Livas, K. F. Garrity, L. Genovese, P. Giannozzi, M. Giantomassi,\nS. Goedecker, X. Gonze, O. Grånäs, E. K. U. Gross, A. Gulans, F. Gygi,\nD. R. Hamann, P. J. Hasnip, N. A. W. Holzwarth, D. Iuşan, D. B. Jochym,\nF. Jollet, D. Jones, G. Kresse, K. Koepernik, E. Küçükbenli, Y. O.\nKvashnin, I. L. M. Locht, S. Lubeck, M. Marsman, N. Marzari, U. Nitzsche,\nL. Nordström, T. Ozaki, L. Paulatto, C. J. Pickard, W. Poelmans, M. I. J.\nProbert, K. Refson, M. Richter, G.-M. Rignanese, S. Saha, M. Scheffler,\nM. Schlipf, K. Schwarz, S. Sharma, F. Tavazza, P. Thunström, A. Tkatchenko,\nM. Torrent, D. Vanderbilt, M. J. van Setten, V. Van Speybroeck, J. M. Wills,\nJ. R. Yates, G.-X. Zhang and S. Cottenier, Reproducibility in density\nfunctional theory calculations of solids. Science 351, article no. aad3000 (2016) ",
      "url": "https://doi.org/10.1126/science.aad3000"
    },
    {
      "type": "Article",
      "id": "bib-bib14",
      "authors": [],
      "title": "\nF. Matúš,\nConditional independences\namong four random variables. II. Combin. Probab. Comput.\n4, 407–417 (1995) ",
      "url": "https://dx.doi.org/10.1017/S0963548300001747"
    },
    {
      "type": "Article",
      "id": "bib-bib15",
      "authors": [],
      "title": "\nF. Matúš,\nConditional independences\namong four random variables. III. Final conclusion. Combin.\nProbab. Comput. 8, 269–276 (1999) ",
      "url": "https://dx.doi.org/10.1017/S0963548399003740"
    },
    {
      "type": "Article",
      "id": "bib-bib16",
      "authors": [],
      "title": "\nF. Matúš and M. Studený,\nConditional independences\namong four random variables. I. Combin. Probab. Comput. 4,\n269–278 (1995) ",
      "url": "https://dx.doi.org/10.1017/S0963548300001644"
    },
    {
      "type": "Article",
      "id": "bib-bib17",
      "authors": [],
      "title": "\nW. K. Michener, Ten simple rules for creating a good data management plan.\nPLOS Computational Biology 11, article no. e1004525\n(2015) ",
      "url": "https://doi.org/10.1371/journal.pcbi.1004525"
    },
    {
      "type": "Article",
      "id": "bib-bib18",
      "authors": [],
      "title": "\nC. Riedel, H. Geßner, A. Seegebrecht, S. I. Ayon, S. H. Chowdhury, R. Engbert and U. Lucke, Including data management\nin research culture increases the reproducibility of scientific results. In Proceedings of INFORMATIK 2022, Lecture Notes in Informatik P-326, pp. 1341–1352, Gesellschaft für Informatik, Bonn\n(2022) ",
      "url": "https://dx.doi.org/10.18420/inf2022_114"
    },
    {
      "type": "Article",
      "id": "bib-bib19",
      "authors": [],
      "title": "\nM. Schappals, A. Mecklenfeld, L. Kröger, V. Botan, A. Köster, S. Stephan,\nE. J. García, G. Rutkai, G. Raabe, P. Klein, K. Leonhard, C. W. Glass,\nJ. Lenhard, J. Vrabec and H. Hasse, Round robin study: Molecular simulation\nof thermodynamic properties from models with internal degrees of freedom.\nJ. Chem. Theory Comput. 13, 4270–4280 (2017) ",
      "url": "https://doi.org/10.1021/acs.jctc.7b00489"
    },
    {
      "type": "Article",
      "id": "bib-bib20",
      "authors": [],
      "title": "\nD. Schober, B. Smith, S. E. Lewis, W. Kusnierczyk, J. Lomax, C. Mungall, C. F.\nTaylor, P. Rocca-Serra and S.-A. Sansone, Survey-based naming conventions\nfor use in OBO Foundry ontology development. BMC Bioinformatics\n10, article no. 125 (2009) ",
      "url": "https://doi.org/10.1186/1471-2105-10-125"
    },
    {
      "type": "Article",
      "id": "bib-bib21",
      "authors": [],
      "title": "\nP. Šimeček, A short note on discrete representability of independence\nmodels. In Proceedings of the 3rd European Workshop on Probabilistic\nGraphical Models, pp. 287–292, Action M Agency, Prague (2006)\n"
    },
    {
      "type": "Article",
      "id": "bib-bib22",
      "authors": [],
      "title": "\nO. Teschke, Some heuristics about the ecosystem of mathematics research data.\nPAMM. Proc. Appl. Math. Mech. 16, 963–964\n(2016) ",
      "url": "https://doi.org/10.1002/pamm.201610460"
    },
    {
      "type": "Article",
      "id": "bib-bib23",
      "authors": [],
      "title": "\nThe LMFDB Collaboration, The L-functions and modular forms database.\nhttp://www.lmfdb.org (2022) ",
      "url": "http://www.lmfdb.org"
    },
    {
      "type": "Article",
      "id": "bib-bib24",
      "authors": [],
      "title": "\nThe MaRDI consortium, MaRDI: Mathematical Research Data Initiative\nproposal. Zenodo (2022) ",
      "url": "https://doi.org/10.5281/zenodo.6552436"
    },
    {
      "type": "Article",
      "id": "bib-bib25",
      "authors": [],
      "title": "\nM. Wilkinson, M. Dumontier, IJ. J. Aalbersberg, G. Appleton, M. Axton, A. Baak,\nN. Blomberg, J.-W. Boiten, L. O. Bonino da Silva Santos, P. Bourne,\nJ. Bouwman, A. Brookes, T. Clark, M. Crosas, I. Dillo, O. Dumon, S. Edmunds,\nC. Evelo, R. Finkers and B. Mons, The FAIR Guiding Principles for\nscientific data management and stewardship. Scientific Data\n3, article no. 160018 (2016) ",
      "url": "https://doi.org/10.1038/sdata.2016.18"
    }
  ],
  "title": "Research-data management planning in the German mathematical community",
  "meta": {},
  "content": [
    {
      "type": "Heading",
      "id": "S1",
      "depth": 1,
      "content": [
        "1 Introduction"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S1.p1",
      "content": [
        "Scientific progress heavily relies on the reusability of previous results.\nThis in turn is closely linked to reliability and reproducibility of research, and to the question whether another\nresearcher would arrive at the same result with the same material. In mathematics proofs, together with references to\ndefinitions of mathematical objects and already verified theorems, traditionally contained all the information needed\nin order to verify results.\nHowever, the advent of computers has opened up new resources previously deemed impossible, while increasing the need\nfor well-adapted research-data management (RDM). For example, algorithms are now implemented to arrive at new\nconclusions. The size of examples has exploded several orders in magnitude. And some proofs have become too complicated\nfor even the brightest minds, such that software is consulted for thorough understanding and\nverification.",
        {
          "type": "Note",
          "id": "idm17",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote1",
              "content": [
                "See, e.g., the story outlined in\n",
                {
                  "type": "Link",
                  "target": "https://xenaproject.wordpress.com/2020/12/05/liquid-tensor-experiment",
                  "content": [
                    "https://xenaproject.wordpress.com/2020/12/05/liquid-tensor-experiment"
                  ]
                },
                ", solved with the lean project\n",
                {
                  "type": "Link",
                  "target": "https://github.com/leanprover-community/lean-liquid",
                  "content": [
                    "https://github.com/leanprover-community/lean-liquid"
                  ]
                },
                "."
              ]
            }
          ]
        },
        "\nStudies [",
        {
          "type": "Cite",
          "target": "bib-bib13",
          "content": [
            "13"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib19",
          "content": [
            "19"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib12",
          "content": [
            "12"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib18",
          "content": [
            "18"
          ]
        },
        "] from various fields of applied mathematics show that\nnowadays many results cannot be easily reproduced and hence verified.\n"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S1.p2",
      "content": [
        "As\nwe outline in ",
        {
          "type": "Cite",
          "target": "S2",
          "content": [
            "Section 2"
          ]
        },
        ", there are research data in all subdisciplines of mathematics that need responsible\norganization and documentation in order to ensure they are handled according to the FAIR principles [",
        {
          "type": "Cite",
          "target": "bib-bib25",
          "content": [
            "25"
          ]
        },
        "] for\nsustainable, reproducible, and reusable research. One way to achieve this is via a tailored research-data management\nplan (RDMP), describing the data life cycle over the course of a project\n[",
        {
          "type": "Cite",
          "target": "bib-bib17",
          "content": [
            "17"
          ]
        },
        "] and providing guidance to fulfill funding requirements.",
        {
          "type": "Note",
          "id": "idm38",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote2",
              "content": [
                "E.g., at the European level\n",
                {
                  "type": "Link",
                  "target": "https://ec.europa.eu/info/funding-tenders/opportunities/docs/2021-2027/horizon/guidance/programme-guide_horizon_en.pdf",
                  "content": [
                    "https://ec.europa.eu/info/funding-tenders/opportunities/docs/2021-2027/horizon/guidance/programme-guide_horizon_en.pdf"
                  ]
                },
                ",\nand at the German national level\n",
                {
                  "type": "Link",
                  "target": "https://www.dfg.de/download/pdf/foerderung/grundlagen_dfg_foerderung/forschungsdaten/forschungsdaten_checkliste_de.pdf",
                  "content": [
                    "https://www.dfg.de/download/pdf/foerderung/grundlagen_dfg_foerderung/forschungsdaten/forschungsdaten_checkliste_de.pdf"
                  ]
                },
                "."
              ]
            }
          ]
        },
        "\nIn mathematics, it is particularly important to treat the RDMP as a living document [",
        {
          "type": "Cite",
          "target": "bib-bib4",
          "content": [
            "4"
          ]
        },
        "] because the\nmathematical research process is hardly projectable and does usually not follow a standardized\ncollection–analysis–report procedure.\nIn subfields with experience in using such documentation, three-fold reports – at the grant-application stage, as a\nworking document, and as a final report – have proven useful. We discuss this in Sections ",
        {
          "type": "Cite",
          "target": "S3",
          "content": [
            "3"
          ]
        },
        " and ",
        {
          "type": "Cite",
          "target": "S4",
          "content": [
            "4"
          ]
        },
        ", spotlighting\nexamples from different subfields, and conclude this article by listing central topics for RDMPs in all areas of\nmathematics."
      ]
    },
    {
      "type": "Heading",
      "id": "S2",
      "depth": 1,
      "content": [
        "2 Mathematical research data"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S2.p1",
      "content": [
        "Following [",
        {
          "type": "Cite",
          "target": "bib-bib9",
          "content": [
            "9"
          ]
        },
        ", p. 130], we define ",
        {
          "type": "Emphasis",
          "content": [
            "research data"
          ]
        },
        " as all digital and analog objects that\nare generated or handled in the process of doing research.",
        {
          "type": "Note",
          "id": "idm65",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote3",
              "content": [
                "This is in line with the notions employed by the\nDFG ",
                {
                  "type": "Link",
                  "target": "https://www.dfg.de/foerderung/grundlagen_rahmenbedingungen/forschungsdaten/index.html",
                  "content": [
                    "https://www.dfg.de/foerderung/grundlagen_rahmenbedingungen/forschungsdaten/index.html"
                  ]
                },
                ", forschungsdaten.info\n",
                {
                  "type": "Link",
                  "target": "https://www.forschungsdaten.info/themen/informieren-und-planen/was-sind-forschungsdaten",
                  "content": [
                    "https://www.forschungsdaten.info/themen/informieren-und-planen/was-sind-forschungsdaten"
                  ]
                },
                ", and the MPG\n",
                {
                  "type": "Link",
                  "target": "https://rdm.mpdl.mpg.de/introduction/research-data-management",
                  "content": [
                    "https://rdm.mpdl.mpg.de/introduction/research-data-management"
                  ]
                },
                ", e.g."
              ]
            }
          ]
        },
        " In mathematics,\nresearch data thus include paper publications and proofs therein as well as computational results, code, software, and\nlibraries of classifications of mathematical objects. A non-exhaustive list of possible formats and examples is\npresented in ",
        {
          "type": "Cite",
          "target": "S2-T1",
          "content": [
            "Table 1"
          ]
        },
        ", and ",
        {
          "type": "Cite",
          "target": "S4",
          "content": [
            "Section 4"
          ]
        },
        ".\nThe apparent diversity of mathematical research-data formats is also reflected in other characteristics, such as their\nstorage size, longevity, and state of standardization [",
        {
          "type": "Cite",
          "target": "bib-bib10",
          "content": [
            "10"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib22",
          "content": [
            "22"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib8",
          "content": [
            "8"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib3",
          "content": [
            "3"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib24",
          "content": [
            "24"
          ]
        },
        "],\nleading to RDM needs and challenges that are very specific to the discipline of mathematics."
      ]
    },
    {
      "type": "Table",
      "id": "S2-T1",
      "caption": [
        {
          "type": "Paragraph",
          "content": [
            "Mathematical research data come in a variety of data formats. Updated table based on [",
            {
              "type": "Cite",
              "target": "bib-bib24",
              "content": [
                "24"
              ]
            },
            ", pp. 26–27]."
          ]
        }
      ],
      "rows": [
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                {
                  "type": "Strong",
                  "content": [
                    "Research-data type"
                  ]
                }
              ]
            },
            {
              "type": "TableCell",
              "content": [
                {
                  "type": "Strong",
                  "content": [
                    "Examples of data formats"
                  ]
                }
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Mathematical documents"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "PDF, LATEX, XML, MathML"
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Literate programming sources"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "Maple Worksheets, Jupyter/Mathematica/Pluto Notebooks"
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Domain-specific research software packages and libraries"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "R for statistics, Octave, NumPy/SciPy or\nJulia for matrix computations, CPLEX, Gurobi, Mosel and SCIP for integer programming, or DUNE, deal.II and Trilinos for\nnumerical simulation"
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Computer-algebra systems"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "SageMath, SINGULAR, Macaulay2, GAP, polymake, Pari/GP, Linbox, OSCAR, and\ntheir embedded data collections"
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Programs and scripts"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "written in the packages and systems above, in systems not developed within the\nmathematical community, input data for these systems (algorithmic parameters, meshes, mathematical objects stored in\nsome collection, the definition of a deep neural network as a graph in machine learning)"
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Experimental and simulation data"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "usually series of states of representative snapshots of an observed\nsystem, discretized fields, more generally very large but structured datasets as simulation output or experimental\noutput (simulation input and validation), stored in established data formats (i.e., HDF5) or in domain-specific\nformats, e.g., CT scans in neuroscience, material science or hydrology"
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Formalized mathematics"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "Coq, HOL, Isabelle, Lean, Mizar, NASA PVS library"
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Collections of mathematical objects"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "L-Functions and Modular Forms Database (LMFDB), Online Encyclopedia of\nInteger Sequences (OEIS), Class Group Database, ATLAS of Finite Group Representations, Manifold Atlas, GAP Small Groups\nLibrary"
              ]
            }
          ]
        },
        {
          "type": "TableRow",
          "cells": [
            {
              "type": "TableCell",
              "content": [
                "Descriptions of mathematical models in mathematical modeling languages"
              ]
            },
            {
              "type": "TableCell",
              "content": [
                "Modelica for component-oriented\nmodeling of complex systems, Systems Biology Markup Language (SBML) for computational models of biological processes,\nSPICE for modeling of electronic circuits and devices, and AIMMS or LINGO as a modeling language for integer programming"
              ]
            }
          ]
        }
      ]
    },
    {
      "type": "Paragraph",
      "id": "S2.p2",
      "content": [
        "One of the most apparent challenges is the question what metadata are sufficient for reusability. We will answer this\nquestion partially for the mathematical subfields presented in ",
        {
          "type": "Cite",
          "target": "S4",
          "content": [
            "4"
          ]
        },
        ". However,\nas the authors of [",
        {
          "type": "Cite",
          "target": "bib-bib3",
          "content": [
            "3"
          ]
        },
        "] note,\n‘the meaning and provenance of [mathematical research] data must usually be given in the form of complex\nmathematical data themselves.’ It is thus not surprising that there is no common, standardized metadata format yet.\nA search in the RDA Metadata Standards\nCatalog",
        {
          "type": "Note",
          "id": "idm176",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote4",
              "content": [
                {
                  "type": "Link",
                  "target": "https://rdamsc.bath.ac.uk/subject/Mathematics%20and%20statistics",
                  "content": [
                    "https://rdamsc.bath.ac.uk/subject/Mathematics%20and%20statistics"
                  ]
                }
              ]
            }
          ]
        },
        " at the time of writing\nreveals five hits, four from a subfield of statistics and one from economics, none of which could encode information\nabout, say, a computer-algebra experiment. This lack of standardization is in contrast to other disciplines such as the\nlife sciences, where the OBO Foundry",
        {
          "type": "Note",
          "id": "idm183",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote5",
              "content": [
                {
                  "type": "Link",
                  "target": "https://obofoundry.org/principles/fp-000-summary.html",
                  "content": [
                    "https://obofoundry.org/principles/fp-000-summary.html"
                  ]
                }
              ]
            }
          ]
        },
        "\nhosts more than one hundred interoperable ontologies to describe and link research results, including common naming\nconventions [",
        {
          "type": "Cite",
          "target": "bib-bib20",
          "content": [
            "20"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib1",
          "content": [
            "1"
          ]
        },
        "].\n"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S2.p3",
      "content": [
        "Another important aspect of mathematical research data is its particular data life cycle. Again in contrast, for\ninstance, to the life sciences, where older results can be overruled by new evidence, mathematical results that have\nbeen proven true remain true indefinitely. Since they cater for other disciplines such as the physical, social, health\nor life sciences [",
        {
          "type": "Cite",
          "target": "bib-bib24",
          "content": [
            "24"
          ]
        },
        ", Fig. 1 and discussion], mathematics has a particular responsibility to science to\npreserve their results in a sustainable manner. We discuss this aspect and how mathematics can be embedded in\ninterdisciplinary research pipelines in more detail in Sections ",
        {
          "type": "Cite",
          "target": "S4-SS2",
          "content": [
            "4.2"
          ]
        },
        " and ",
        {
          "type": "Cite",
          "target": "S4-SS1",
          "content": [
            "4.1"
          ]
        },
        ". Section ",
        {
          "type": "Cite",
          "target": "S4-SS3",
          "content": [
            "4.3"
          ]
        },
        " stresses the role thorough\ndocumentation plays in this context, using classifications as an example."
      ]
    },
    {
      "type": "Heading",
      "id": "S3",
      "depth": 1,
      "content": [
        "3 Status quo of RDMPs in mathematics"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S3.p1",
      "content": [
        "In the narrower sense of data (rather than research data), it is a common claim in the community at the time of writing\nthat mathematics rarely produces data",
        {
          "type": "Note",
          "id": "idm213",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote6",
              "content": [
                "Usually, only statistics is mentioned as a data-producing\nsubdiscipline, see,\ne.g., ",
                {
                  "type": "Link",
                  "target": "https://wissenschaftliche-integritaet.de/kommentare/software-entwicklung-und-umgang-mit-forschungsdaten-in-der-mathematik",
                  "content": [
                    "https://wissenschaftliche-integritaet.de/kommentare/software-entwicklung-und-umgang-mit-forschungsdaten-in-der-mathematik"
                  ]
                },
                "."
              ]
            }
          ]
        },
        " and that the few data available need no particular management.",
        {
          "type": "Note",
          "id": "idm220",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote7",
              "content": [
                "See, e.g., the unofficial\ndocument ",
                {
                  "type": "Link",
                  "target": "https://www.math.harvard.edu/media/DataManagement.pdf",
                  "content": [
                    "https://www.math.harvard.edu/media/DataManagement.pdf"
                  ]
                },
                "."
              ]
            }
          ]
        },
        " This is often based on an\ninterpretation of data being something computational, and mathematics being a discipline which is very much paper\nrather than computer based. Anecdotal evidence suggests that this view is widely established, that there is little\nknowledge about general RDM, that existing local facilities are hardly used, and that RDMPs are not a standard tool at\nany stage of the research process.",
        {
          "type": "Note",
          "id": "idm227",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote8",
              "content": [
                "In fact, RDMPs became compulsory in DFG-funding applications only in March\n2022, and there are no statistics available on how many mathematics proposals included such a document. See also\n",
                {
                  "type": "Link",
                  "target": "https://www.dfg.de/foerderung/info_wissenschaft/2022/info_wissenschaft_22_25/index.html",
                  "content": [
                    "https://www.dfg.de/foerderung/info_wissenschaft/2022/info_wissenschaft_22_25/index.html"
                  ]
                },
                "."
              ]
            }
          ]
        },
        "\nThe proposal [",
        {
          "type": "Cite",
          "target": "bib-bib24",
          "content": [
            "24"
          ]
        },
        "] has identified the need to build common infrastructures for all subdisciplines of\nmathematics, and mathematics-specific DFG guidelines for FAIR research data will be developed in the foreseeable\nfuture.",
        {
          "type": "Note",
          "id": "idm236",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote9",
              "content": [
                {
                  "type": "Link",
                  "target": "https://www.dfg.de/foerderung/info_wissenschaft/2022/info_wissenschaft_22_25/index.html",
                  "content": [
                    "https://www.dfg.de/foerderung/info_wissenschaft/2022/info_wissenschaft_22_25/index.html"
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "type": "Paragraph",
      "id": "S3.p2",
      "content": [
        "Now, the question of what these guidelines should be is not trivial.\nA large number of questions from a general RDMP catalogue,",
        {
          "type": "Note",
          "id": "idm245",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote10",
              "content": [
                "For instance, the current questionnaire supplied by\nthe DFG-funded research-data management organiser RDMO\n",
                {
                  "type": "Link",
                  "target": "https://github.com/rdmorganiser/rdmo-catalog/releases/tag/1.1.0-rdmo-1.6.0",
                  "content": [
                    "https://github.com/rdmorganiser/rdmo-catalog/releases/tag/1.1.0-rdmo-1.6.0"
                  ]
                },
                "."
              ]
            }
          ]
        },
        " are\nirrelevant for a community which produces foremostly theoretical results. For instance, for mathematicians the cost of\nproducing data is rarely relevant – unlike, e.g., in the life sciences where data might have to be collected in the\nfield. In the same vein, ethical or data-protection questions most often do not play a role, save for, for instance,\nindustry collaborations or studies conducted in didactics. Large parts of the community have little training in legal\naspects as, for example, formulae cannot be assigned proprietary rights. In order to avoid the impression that thus all\ngeneral RDMP questions apply only to sciences different than mathematics, it is imperative to design bespoke catalogues\nof questions. These should (a) use unambiguous language, for instance, using the term ‘research data’ rather\nthan the more specific ‘data’ which many mathematicians do not handle in their research, and (b) avoid\nsuperfluous topics while at the same time including sufficient detail, for instance, for mathematics’ metadata and\npreservation needs identified in the previous section. Now, rather than endeavoring to find a one-size-fits-all\nsolution, in the subsequent section we identify important RDM questions for a number of use cases which are known to\nthe authors – focusing on metadata, software, data formats and size, versioning, and storage – and provide those with\nwhat we consider to be sensible answers.\nWe use the remainder of this section to report on two RDM solutions implemented in DFG-funded Collaborative Research Centers (CRC).\n"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S3.p3",
      "content": [
        "The CRC 1456 ‘Mathematics of Experiment’ includes 17 scientific projects in applied mathematics, computer\nscience, and natural sciences such as biophysics and astronomy, aiming to improve the analysis of experimental data.\nThe research data here are extremely diverse (e.g., mathematical documents, notebooks, programs, simulation data or\nexperimental measurements) and their handling is supported by the CRC’s dedicated infrastructure project.\nIn regular RDM meetings, four themes are recurrent. First, reusage scenarios: especially in interdisciplinary research\nthe same datasets may be processed or used by different groups; documentation, curation, and publication should be\ntailored to those groups’ needs. Second, reproducibility, both computationally and practically in data recreation.\nThird, metadata: finding accurate descriptors to help the user understand cross-scientific research data. And fourth,\nvisibility: receiving recognition for stand-alone research data beyond a journal publication is hard.\nThis last topic is usually not part of a standard set of RDMP questions but aims to provide an incentive to increase\nthe effort in research-data creation, publication, and curation."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S3.p4",
      "content": [
        "The CRC 1294 ‘Data Assimilation’ includes 15 interdisciplinary research projects focusing on the development\nand integration of algorithms, e.g., in earthquake prediction, medication dosing, or cell-shape dynamics.\nResearchers are thus confronted both with diverse research data and varying cultural data-handling habits. A central\nproject supports their RDM, and IT infrastructure to facilitate collaborative work and knowledge perpetuation to\nadvance good scientific practice is provided. In particular, the CRC designed an RDMP template in collaboration with\nthe University of Potsdam’s research-data group. This covers policies and guidelines, legal and ethical considerations,\ndocumentation, and dataset-specific aspects. A vital component of the training is then the classification of the\ndigital objects that are reused and created by the individual researchers. This helps them to develop tailored\nstrategies to improve the quality and reproducibility of published results and to sensitize their research-data\nhandling throughout the data life cycle."
      ]
    },
    {
      "type": "Heading",
      "id": "S4",
      "depth": 1,
      "content": [
        "4 Use cases"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.p1",
      "content": [
        "We now consider four very different mathematical use cases and discuss their particular research-data needs. These use\ncases have been identified by the different mathematical subfields in [",
        {
          "type": "Cite",
          "target": "bib-bib24",
          "content": [
            "24"
          ]
        },
        "] as particularly\nrepresentative for the research community. Central in these expositions for us is to find out how, using RDMPs, we can\nprovide the best, case-specific guidance to make a project reusable."
      ]
    },
    {
      "type": "Heading",
      "id": "S4.SS1",
      "depth": 2,
      "content": [
        "4.1 Applied and interdisciplinary mathematics"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS1.p1",
      "content": [
        "In numerous scientific fields real-world problems are simplified, e.g., to experiments, and subsequently described in\nabstract ways using mathematical models. If a model is combined with input data, it forms a concrete instance of such a\nproblem. With the help of algorithms, the input data are then transformed into output data. Following validations, the\ninterpretation of outputs provides the solution of the initial problem in a so-called Modeling–Simulation–Optimization\nworkflow [",
        {
          "type": "Cite",
          "target": "bib-bib24",
          "content": [
            "24"
          ]
        },
        ", p. 77]. For complete RDM, such workflows should be documented in detail as part of an\nRDMP."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS1.p2",
      "content": [
        "A standard RDMP questionnaire includes some guidance for the documentation of workflows, such as the main research\nquestion, involved disciplines, tools, software, technologies, processes, research-data aspects, and reproducibility.\nUsing this as a template, a tailored questionnaire is currently being developed within the framework of\nMaRDI",
        {
          "type": "Note",
          "id": "idm282",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote11",
              "content": [
                {
                  "type": "Link",
                  "target": "https://www.mardi4nfdi.de",
                  "content": [
                    "https://www.mardi4nfdi.de"
                  ]
                }
              ]
            }
          ]
        },
        " to document workflows in detail. This is divided into four sections\ndealing with the problem statement (object of research, data streams), the model (discretization, variables), the\nprocess information (process steps, applied methods), and reproducibility. It is aimed at all disciplines and differs\nonly slightly in whether a theoretical or experimental workflow is documented. The central element of the questionnaire\nis to establish connections between different steps of the research process in order to improve interoperability of\nresearch data. The description of an individual process step, for example, requires the assignment of the relevant\ninput and output data, the method and the (software) environment. At the same time, the documentation of the methods,\nsoftware, input and output data requires persistent identifiers (e.g., Wikidata, swMATH, DOI) in addition to\ntopic-dependent information."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS1.p3",
      "content": [
        "We consider the documentation of a concrete workflow combining archaeology and mathematics as an example. This is based\non [",
        {
          "type": "Cite",
          "target": "bib-bib11",
          "content": [
            "11"
          ]
        },
        "], was created by Margarita Kostre independently afterwards, and described in personal communication\nas ‘very helpful for the reflection of the own work.’ The author commented that she will use workflow\ndocumentation in the future again, as she believes it facilitates interdisciplinary communication, e.g., about the\nstatus of a project, its goal, and data transfer, it provides better clarity in larger collaborations and allows\ncolleagues to enter a project more easily. The aim of this work is to understand the Romanization of Northern Africa\nusing a susceptible infectious epidemic model. On the process level, the workflow starts with data preparation,\ne.g., collecting, discretizing, and reducing archaeological data. Once a suitable epidemics model is found, the\ninverse problem is solved to determine contact networks and spreading-rate functions. Subsequent analysis allows the\nidentification of three different possibilities of the Romanization of Northern Africa. The detailed documentation can\nbe found on the MaRDI\nPortal.",
        {
          "type": "Note",
          "id": "idm293",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote12",
              "content": [
                {
                  "type": "Link",
                  "target": "https://portal.mardi4nfdi.de/wiki/Romanization_spreading_on_historical_interregional_networks_in_Northern_Tunisia",
                  "content": [
                    "https://portal.mardi4nfdi.de/wiki/Romanization_spreading_on_historical_interregional_networks_in_Northern_Tunisia"
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "type": "Heading",
      "id": "S4.SS2",
      "depth": 2,
      "content": [
        "4.2 Scientific computing"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS2.p1",
      "content": [
        "While research in pure mathematics strives to determine an ultimate truth, applied or computational mathematics in\nmajority need to deal with approximations to reality: models are usually expressed in terms of real or complex numbers\nand only finite subsets of these can actually be implemented on computer hardware.\nConsequently, the result of a computation depends on the format of the finite-precision numbers used and on the\nspecific hardware executing the computations, making a detailed documentation of the computer-based experiment crucial\nand reusability of code a must-have [",
        {
          "type": "Cite",
          "target": "bib-bib5",
          "content": [
            "5"
          ]
        },
        "].\nThus, the input data and results of a computer experiment and also the precise implementation (code, software, and\nhardware) of the algorithms used constitute important research data."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS2.p2",
      "content": [
        "Absence of such details in documentation makes applied mathematics face the same reproducibility issues\n(e.g., [",
        {
          "type": "Cite",
          "target": "bib-bib2",
          "content": [
            "2"
          ]
        },
        "]) as other scientific fields. Still mathematical algorithms make up the foundation of many\ncomputational experiments, for instance as solvers for linear systems of equations, eigenvalue problems, or\noptimization problems, and are thus at the heart of science today. This responsibility calls for rigorous RDM and\ndocumentation in RDMPs.\n"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS2.p3",
      "content": [
        "The main difficulty in establishing RDMPs in scientific computing seems to be in creating incentives to adhere to\ncommon standards. In case of a single multi-author paper within a larger project cluster, there are two levels to this\nquestion: the funding context and local RDM. Regarding the first, incentives should clearly address reporting\nrequirements and incorporate rewards for sustainable RDM, rather than merely counting publications and citations, to\nensure the cluster can ",
        {
          "type": "Emphasis",
          "content": [
            "stand on the shoulders of giants"
          ]
        },
        " instead of ",
        {
          "type": "Emphasis",
          "content": [
            "building on quicksand"
          ]
        },
        ". The\nbeneficiaries here are other researchers in the project and world-wide. Consequently, global RDM needs to answer what\nis reported where and why. Regarding the local context, for the collaborative work of the\nauthors,\nincentives are far\nmore evident. Thorough RDM, documented in a living RDMP, not only accelerates the paper writing, it also improves the\nreusability of information for future endeavors of the individual authors. Questions center around ‘When is the\ncode/data provided? Where in the (local) infrastructure is it stored? By whom? Who is processing it next?’"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS2.p4",
      "content": [
        "Consequently, RDMPs should be modularized to enable the single modules to change at their appropriate pace. While the\nglobal management rules of a project cluster may not change at all, or at best very slowly, the findings in a single\nwork package may alter the RDMP and thus RDM needs high agility to react to changes. For the software pipeline of an\nexample paper that means: a task-based RDMP, updated as the pipeline evolves, needs to fulfill the requirements\nof [",
        {
          "type": "Cite",
          "target": "bib-bib5",
          "content": [
            "5"
          ]
        },
        "], while for the project cluster sustainable handover, following\n(e.g., [",
        {
          "type": "Cite",
          "target": "bib-bib6",
          "content": [
            "6"
          ]
        },
        "]),\nneeds to be addressed in the overarching RDMP."
      ]
    },
    {
      "type": "Heading",
      "id": "S4.SS3",
      "depth": 2,
      "content": [
        "4.3 Computer algebra and theoretical statistics"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS3.p1",
      "content": [
        "Large parts of the German mathematical community consider themselves as not doing applied work. This includes fields\nsuch as geometry, topology, algebra, analysis or number theory, and also mathematical statistics, for instance.\nHowever, these researchers increasingly use computers, too, to explore the viability of proof strategies, test their\nown conjectures or refute established ones.\nAs a consequence, ",
        {
          "type": "Emphasis",
          "content": [
            "classifications"
          ]
        },
        ", the systematic and complete tabulation of all objects with a given property,\ngrow wildly in size and complexity. They give a complete picture for some aspect of a theory and may be used in many\nways, from the search of (counter)examples over building blocks for constructive proofs to benchmark problems."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS3.p2",
      "content": [
        "For\ninstance, the L-Functions and Modular Forms Database (LMFDB) [",
        {
          "type": "Cite",
          "target": "bib-bib23",
          "content": [
            "23"
          ]
        },
        "]\ncontains over 4.8 TB of data relating\nobjects conjectured to have strong connections by the Langlands program: number fields, elliptic curves, modular forms,\nL-functions, Galois representations. It includes tens of millions of individual objects and stores the relations\nbetween these. Entries contain detailed information on reliability, completeness, and several versions of the code\nneeded to compute them. The database has a public reporting system which allows all users to have visibility of any\nissues or\nerrors.",
        {
          "type": "Note",
          "id": "idm341",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote13",
              "content": [
                "Other classification databases targeted at specific audiences are listed at ",
                {
                  "type": "Link",
                  "target": "https://mathdb.mathhub.info",
                  "content": [
                    "https://mathdb.mathhub.info"
                  ]
                },
                "."
              ]
            }
          ]
        }
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS3.p3",
      "content": [
        "Computing mathematical objects for classification can often be algorithmically hard and time-consuming. But once\ncomputed, results are final and independent of the software used.\nWith larger computer clusters and better algorithms, it is unreasonable and unsustainable to expect researchers who\nwant to build on existing research to repeat individual computations. This expected reuse increases the need for\nresponsible RDM and triggers challenges which need to be addressed in an RDMP. In particular, four themes are central\nin this regard. First, how can researchers ensure that their research data are correct and complete? Is the connection\nof mathematical theory and code sound? Second, how can other researchers access, understand, and reuse the research\ndata? Third, how can one ensure longevity of their research data? And fourth, how can researchers report\nerrors/corrections and upload new versions of research data if necessary?"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS3.p4",
      "content": [
        "These questions are neatly addressed in the LMFDB mentioned above. To show how things can go wrong without proper RDM,\nwe discuss a classification of all conditional independence structures on up to four discrete random variables,\noriginally published in a series of papers [",
        {
          "type": "Cite",
          "target": "bib-bib16",
          "content": [
            "16"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib14",
          "content": [
            "14"
          ]
        },
        ", ",
        {
          "type": "Cite",
          "target": "bib-bib15",
          "content": [
            "15"
          ]
        },
        "].\nOf the ",
        {
          "type": "MathFragment",
          "mathLanguage": "mathml",
          "text": "<mml:math xmlns:mml=\"http://www.w3.org/1998/Math/MathML\" id=\"S4.SS3.p4.m1\" alttext=\"2^{24}=16\\,777\\,216\" display=\"inline\"><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mn>24</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>16 777 216</mml:mn></mml:mrow></mml:math>",
          "meta": {
            "altText": "2^{24}=16\\,777\\,216"
          }
        },
        {
          "type": "Emphasis",
          "content": [
            "a priori"
          ]
        },
        " possible patterns of how four random variables can influence each other,\nonly ",
        {
          "type": "MathFragment",
          "mathLanguage": "mathml",
          "text": "<mml:math xmlns:mml=\"http://www.w3.org/1998/Math/MathML\" id=\"S4.SS3.p4.m2\" alttext=\"18\\,478\" display=\"inline\"><mml:mn>18 478</mml:mn></mml:math>",
          "meta": {
            "altText": "18\\,478"
          }
        },
        " (",
        {
          "type": "MathFragment",
          "mathLanguage": "mathml",
          "text": "<mml:math xmlns:mml=\"http://www.w3.org/1998/Math/MathML\" id=\"S4.SS3.p4.m3\" alttext=\"\\approx 0.11\\%\" display=\"inline\"><mml:mrow><mml:mi/><mml:mo>≈</mml:mo><mml:mrow><mml:mn>0.11</mml:mn><mml:mo>%</mml:mo></mml:mrow></mml:mrow></mml:math>",
          "meta": {
            "altText": "\\approx 0.11\\%"
          }
        },
        ") are realizable with a probability distribution.\nŠimeček, the author of [",
        {
          "type": "Cite",
          "target": "bib-bib21",
          "content": [
            "21"
          ]
        },
        "],\ndigitized this result and then left the field after his PhD in 2007. His research data was deleted in 2021 from his\nformer institute’s website – the only public place which ever held the database.",
        {
          "type": "Note",
          "id": "idm378",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote14",
              "content": [
                "A backup is still\navailable on the Internet Archive at\n",
                {
                  "type": "Link",
                  "target": "http://web.archive.org/web/20190516145904/http://atrey.karlin.mff.cuni.cz/~simecek/skola/models/",
                  "content": [
                    "http://web.archive.org/web/20190516145904/http://atrey.karlin.mff.cuni.cz/~simecek/skola/models/"
                  ]
                },
                "."
              ]
            }
          ]
        },
        " It was encoded in a packed binary format which is hard to read, search, and reuse. Some files supporting\nthe correctness of the classification for binary distributions use an unspecified, compiler-specific binary\nserialization format for floating-point data.",
        {
          "type": "Note",
          "id": "idm385",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote15",
              "content": [
                "A set of scripts for reading these files is available at\n",
                {
                  "type": "Link",
                  "target": "https://github.com/taboege/simecek-tools",
                  "content": [
                    "https://github.com/taboege/simecek-tools"
                  ]
                },
                "."
              ]
            }
          ]
        },
        " The programs used for the creation and inspection\nof the database were written in a dialect of the Pascal programming language, which has not been maintained since 2006.\nThe sparse documentation is in Czech."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S4.SS3.p5",
      "content": [
        "This situation can only be fixed by recreating the database from scratch, including proofs.\nAn RDMP for this project should emphasize the need to list and document each step of redoing the computations, the use\nof standard data formats with rich metadata for interoperability and searchability of the database, and ensure future\nreusability of Šimeček’s results."
      ]
    },
    {
      "type": "Heading",
      "id": "S5",
      "depth": 1,
      "content": [
        "5 Discussion and outlook"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S5.p1",
      "content": [
        "The problem of reusability strongly relates to a phenomenon called ",
        {
          "type": "Emphasis",
          "content": [
            "dark data"
          ]
        },
        ", which ‘exists only in the\nbottom left-hand desk drawer of scientists on some media that are quickly aging’ [",
        {
          "type": "Cite",
          "target": "bib-bib7",
          "content": [
            "7"
          ]
        },
        "]. If research data\nare not available, they are of course neither traceable nor reusable or FAIR. This phenomenon extends from lost USB\nsticks and conflicting cloud-based collaboration tools like\nDropbox",
        {
          "type": "Note",
          "id": "idm407",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote16",
              "content": [
                {
                  "type": "Link",
                  "target": "https://www.dropbox.com",
                  "content": [
                    "https://www.dropbox.com"
                  ]
                }
              ]
            }
          ]
        },
        "nd\nOverleaf",
        {
          "type": "Note",
          "id": "idm414",
          "noteType": "Footnote",
          "content": [
            {
              "type": "Paragraph",
              "id": "footnote17",
              "content": [
                {
                  "type": "Link",
                  "target": "https://www.overleaf.com",
                  "content": [
                    "https://www.overleaf.com"
                  ]
                }
              ]
            }
          ]
        },
        "ithout local backup to papers containing very condensed\ncomplicated proofs that can only be taken up in future work if access to handwritten notes of the authors is also\npossible. A prime example of this is presented in Section ",
        {
          "type": "Cite",
          "target": "S4-SS3",
          "content": [
            "4.3"
          ]
        },
        " where unavailable research data is in stark contrast to\nthe everlasting truth of mathematical results.\nRDMPs are a tool of choice against such issues, serving as a basic measure to organize the full data life cycle.\n"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S5.p2",
      "content": [
        "From the three case studies considered in ",
        {
          "type": "Cite",
          "target": "S4",
          "content": [
            "Section 4"
          ]
        },
        ", we derive that RDMPs in mathematics in particular\n(a) stimulate reflection, clarity, and interdisciplinary communication, (b) require flexibility and modularization as\nliving RDMPs, and\n(c) facilitate the documentation of iterative computational processes by fostering research-data\ninteroperability and reusability."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S5.p3",
      "content": [
        "We further conclude that archiving and preservation is key in any mathematical subdiscipline. As a very first step to\nimprove the status quo, all research results necessary for reusability (data, code, notes, …) should be stored in\na sustainable and findable manner, using resources already documented in an RDMP before a project starts. Ideally, in a\nsecond step, citable repositories with persistent identifiers for these research data can be chosen, and, in a third\nstep, these can be annotated with interlinked metadata, implemented via knowledge graphs. Because of the diversity of\nmathematical research data, the choice of metadata should be made carefully with possible reusage scenarios and\ninterest groups in mind, also documented in an RDMP. If code is part of a publication, thoughts should be given to the\ndetail of documentation and again appropriate citable long-term repositories. In addition, an RDMP should be used as a\ntool to identify legal constraints, like the compatibility of software licenses, before any actual work is conducted."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S5.p4",
      "content": [
        {
          "type": "Emphasis",
          "content": [
            "Acknowledgements. "
          ]
        },
        "\nThe authors are grateful to Margarita Kostre for retrospectively compiling an RDMP for her project\n[",
        {
          "type": "Cite",
          "target": "bib-bib11",
          "content": [
            "11"
          ]
        },
        "] and to Tim Hasler for background and discussion regarding the MATH+ research-data management organizer.\n"
      ]
    },
    {
      "type": "Paragraph",
      "id": "S5.p5",
      "content": [
        {
          "type": "Emphasis",
          "content": [
            "Funding."
          ]
        },
        " René Fritze, Christiane Görgen, Jeroen Hanselman, Lars Kastner, Thomas Koprucki, Tabea Krause, Marco Reidelbach,\nJens Saak, Björn Schembera, Karsten Tabelow and Marcus Weber are at the time of writing supported by MaRDI, funded by\nthe Deutsche Forschungsgemeinschaft (DFG), project number 460135501, NFDI 29/1 ‘MaRDI – Mathematische\nForschungsdateninitiative.’ Christian Riedel is supported by the DFG, project-ID 318763901 – SFB1294. Christoph\nLehrenfeld is supported by the DFG, project-ID 432680300 – SFB1456."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S5.p6",
      "content": [
        "The authors have no competing interests to declare."
      ]
    },
    {
      "type": "Paragraph",
      "id": "S5.p7",
      "content": [
        "All authors made significant contributions to the design of this review as well as drafting and revising the\nmanuscript. All have approved this final version, agreed to be accountable and have approved of the inclusion of those\nin the list of authors."
      ]
    },
    {
      "type": "Paragraph",
      "id": "authorinfo",
      "content": [
        "\nTobias Boege got his PhD from the Otto-von-Guericke-Universität Magdeburg working on conditional independence in\nalgebraic statistics. After a postdoc position at the Max Planck Institute for Mathematics in the Sciences, he is\ncurrently at Aalto University.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0001-7284-1827",
          "content": [
            "orcid.org/0000-0001-7284-1827"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:post@taboege.de",
          "content": [
            "post@taboege.de"
          ]
        },
        "\nRené Fritze received his diploma in mathematics at the University of Münster. After working on various research\nprojects, including MaRDI, he has joined the Digital Technology group at Arup as a Senior Software Engineer.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0002-9548-2238",
          "content": [
            "orcid.org/0000-0002-9548-2238"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:rene.fritze@wwu.de",
          "content": [
            "rene.fritze@wwu.de"
          ]
        },
        "\nChristiane Görgen (corresponding author) holds a PhD in statistics from Warwick University and has been doing research in algebraic statistics\nat the Max Planck Institute for Mathematics in the Sciences. Since 2021 she works at the University of Leipzig as\nMaRDI’s mathematical data consultant.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0002-6476-956X",
          "content": [
            "orcid.org/0000-0002-6476-956X"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:goergen@math.uni-leipzig.de",
          "content": [
            "goergen@math.uni-leipzig.de"
          ]
        },
        "\nJeroen Hanselman did his PhD in mathematics at Ulm University and is currently working as a postdoc at the RPTU\nKaiserslautern-Landau concerning himself with improving the software peer reviewing process for MaRDI. His area of\nresearch is computational arithmetic geometry, with a focus on Jacobians of curves.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0002-1298-0961",
          "content": [
            "orcid.org/0000-0002-1298-0961"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:hanselman@mathematik.uni-kl.de",
          "content": [
            "hanselman@mathematik.uni-kl.de"
          ]
        },
        "\nDorothea Iglezakis holds a diploma in psychology and a PhD in Computer Science. She is head of the research-data\nmanagement team of the University of Stuttgart. Dorothea is mainly interested in (semantically enriched) metadata, the\nautomation of research-data management processes and the interlinking of different research outputs, actors, and\nconcepts in a global knowledge graph.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0002-8524-0569",
          "content": [
            "orcid.org/0000-0002-8524-0569"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:dorothea.iglezakis@ub.uni-stuttgart.de",
          "content": [
            "dorothea.iglezakis@ub.uni-stuttgart.de"
          ]
        },
        "\nLars Kastner holds a PhD in mathematics from Freie Universität Berlin. His main research area lies at the intersection\nof algebraic geometry and combinatorics, with a focus on computational aspects. In 2022, he joined the MaRDI task area\non computer algebra at TU Berlin.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0001-9224-7761",
          "content": [
            "orcid.org/0000-0001-9224-7761"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:kastner@math.tu-berlin.de",
          "content": [
            "kastner@math.tu-berlin.de"
          ]
        },
        "\nThomas Koprucki holds a Diploma degree in physics and a PhD degree in mathematics. He works in the field of\nmathematical modeling and numerical simulation in nano- and opto-electronics at WIAS Berlin.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0001-6235-9412",
          "content": [
            "orcid.org/0000-0001-6235-9412"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:thomas.koprucki@wias-berlin.de",
          "content": [
            "thomas.koprucki@wias-berlin.de"
          ]
        },
        "\nTabea H. Krause holds a degree in mathematics and logic from the University of Leipzig. Since 2022 she works as MaRDI’s\nconsortia contact at Leipzig University.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0001-7275-5830",
          "content": [
            "orcid.org/0000-0001-7275-5830"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:tabea.krause@math.uni-leipzig.de",
          "content": [
            "tabea.krause@math.uni-leipzig.de"
          ]
        },
        "\nChristoph Lehrenfeld holds a PhD in mathematics from RWTH Aachen University, and his research focuses on numerical\nmethods for partial differential equations. He has been a professor at the Georg-August-Universität Göttingen\nsince 2016.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0003-0170-8468",
          "content": [
            "orcid.org/0000-0003-0170-8468"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:lehrenfeld@math.uni-goettingen.de",
          "content": [
            "lehrenfeld@math.uni-goettingen.de"
          ]
        },
        "\nSilvia Polla (PhD in archaeology, University of Siena) works since 2021 as a research data steward and library manager\nat the Weierstrass Institute for Applied Analysis and Stochastics (WIAS).\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0002-2395-2448",
          "content": [
            "orcid.org/0000-0002-2395-2448"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:silvia.polla@wias-berlin.de",
          "content": [
            "silvia.polla@wias-berlin.de"
          ]
        },
        "\nMarco Reidelbach holds a PhD in bioinformatics from Freie Universität Berlin and works in the field of molecular\nmodeling and simulation. Since 2021 he works for MaRDI at Zuse Institute Berlin focusing on mathematics in an\ninterdisciplinary context.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0002-1919-1834",
          "content": [
            "orcid.org/0000-0002-1919-1834"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:reidelbach@zib.de",
          "content": [
            "reidelbach@zib.de"
          ]
        },
        "\nChristian Riedel graduated in geoinformation and obtained a PhD in planetary sciences. He currently works at Potsdam\nUniversity in the research-data management of an interdisciplinary Collaborative Research Center on mathematical data\nassimilation. His work involves research on data and software-based procedures in geosciences and the sustainable\nprovision of interdisciplinary research data.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0001-5154-4153",
          "content": [
            "orcid.org/0000-0001-5154-4153"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:christian.riedel@uni-potsdam.de",
          "content": [
            "christian.riedel@uni-potsdam.de"
          ]
        },
        "\nJens Saak holds a PhD in applied mathematics from TU Chemnitz. Since 2010, he has been a team leader at the Max Planck\nInstitute for Dynamics of Complex Technical Systems in Magdeburg. His research covers various aspects of industrial and\napplied mathematics, with an increasing focus on research-software engineering and research-data management.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0001-5567-9637",
          "content": [
            "orcid.org/0000-0001-5567-9637"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:saak@mpi-magdeburg.mpg.de",
          "content": [
            "saak@mpi-magdeburg.mpg.de"
          ]
        },
        "\nBjörn Schembera holds a diploma degree in computer science and a PhD in engineering. His research interests include\ndark data, semantic technology and research-data management and he currently works as a knowledge engineer in MaRDI at\nthe University of Stuttgart.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0003-2860-6621",
          "content": [
            "orcid.org/0000-0003-2860-6621"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:bjoern.schembera@mathematik.uni-stuttgart.de",
          "content": [
            "bjoern.schembera@mathematik.uni-stuttgart.de"
          ]
        },
        "\nKarsten Tabelow holds a PhD in physics. He works in the field of medical image analysis with a focus on neuroimaging,\nquantitative imaging and statistical methods at WIAS Berlin. Since 2016 he has been also working on mathematical\nresearch data and the concepts behind MaRDI, the Mathematical Research Data Initiative.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0003-1274-9951",
          "content": [
            "orcid.org/0000-0003-1274-9951"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:karsten.tabelow@wias-berlin.de",
          "content": [
            "karsten.tabelow@wias-berlin.de"
          ]
        },
        "\nMarcus Weber holds a PhD in mathematics and did a habilitation at FU Berlin. He is head of the research group\n‘Computational Molecular Design’ at Zuse Institute Berlin. Marcus is mainly interested in molecular simulation and\nin the theory of Markov processes.\n",
        {
          "type": "Link",
          "target": "https://orcid.org/0000-0003-3939-410X",
          "content": [
            "orcid.org/0000-0003-3939-410X"
          ]
        },
        ",\n",
        {
          "type": "Link",
          "target": "mailto:weber@zib.de",
          "content": [
            "weber@zib.de"
          ]
        }
      ]
    }
  ]
}