Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reproducible science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES↗

Stewardship of NASA's Earth Science Data and Ensuring Long-Term Active Archives

Program, NASA has followed an open data policy, with non-discriminatory access to data with no period of exclusive access. NASA has well-established processes for assigning and or accepting datasets into one of 12 Distributed Active Archive Centers (DAACs) that are parts of EOSDIS. EOSDIS has been evolving through several information technology cycles, adapting to hardware and software changes in the commercial sector. NASA is responsible for maintaining Earth science data as long as users are interested in using them for research and applications, which is well beyond the life of the data gathering missions. For science data to remain useful over long periods of time, steps must be taken to preserve: (1) Data bits with no corruption, (2) Discoverability and access, (3) Readability, (4) Understandability, (5) Usability' and (6). Reproducibility of results. NASAs Earth Science data and Information System (ESDIS) Project, along with the 12 EOSDIS Distributed Active Archive Centers (DAACs), has made significant progress in each of these areas over the last decade, and continues to evolve its active archive capabilities. Particular attention is being paid in recent years to ensure that the datasets are published in an easily accessible and citable manner through a unified metadata model, a common metadata repository (CMR), a coherent view through the earthdata.gov website, and assignment of Digital Object Identifiers (DOI) with well-designed landing product information pages.

Data Management↗

Rigor and Reproducibility in Electrocatalysis: Best Practices for Operando Studies

Operando measurements have rapidly expanded the scope of electrocatalysis by enabling direct observation of catalytic interfaces under working conditions and by linking structural, compositional, and spectroscopic observables to activity and selectivity. However, the growth of operando methods has outpaced the adoption of broadly shared experimental standards, creating persistent challenges in reproducibility, interpretation, and comparison across laboratories and platforms. This perspective synthesizes discussions from the 2025 National Science Foundation Workshop on Rigor and Reproducibility in Electrocatalysis and outlines a practical framework for the rigorous use of operando measurements in electrocatalysis. We highlight three recurring needs: careful implementation of complex methods to avoid overinterpretation; recognition that (subtle) differences in reactor architecture, hydrodynamics, and electrical boundary conditions can alter apparent kinetics and selectivity; and transparent reporting standards that enable meaningful cross-comparison without constraining measurement-specific cell innovation. Focusing on widely used techniques (including X-ray and vibrational spectroscopies, mass spectrometry, and electron microscopy), we discuss technique-specific pitfalls, cross-validation strategies, and recurring platform-agnostic considerations such as mass transport, current distribution, temporal-resolution mismatches, and catalyst evolution. This Perspective aims to strengthen the mechanistic inference and improve the reproducibility, comparability, and predictive value of operando electrocatalysis research.

X-ray absorption spectroscopy↗

NLSP: NASA Life Sciences Portal

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles. Some of these improvements will at the same time support the twin pillars of Open Science: transparency of methods and reproducibility of results. This video is a high level overview of the NLSP for existing and new users.

Life Sciences data↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

Best practices in catalysis: A perspective

Catalysis, from its roots in petrochemical refining and conversion, has emerged as a transdisciplinary field that now encompasses synthesis of materials and molecules that enable applications in energy conversion and storage, environmental remediation, medicine, plastics, and fertilizer production, among numerous others. A handful of disciplines can claim relevance and success over such an extended period of time and continue to claim a preeminent role in defining the state-of-the-art in science and technology. Syncretic and rapid advancements in formulation and spectroscopic characterization of materials and molecules useful as catalysts, high-level density functional and molecular orbital theory calculations, and a strong foundation in concepts of physical chemistry, thermodynamics, and chemical kinetics offer new and abundant opportunities at the present day for addressing the grand challenge of controlling chemical transformations using catalysis. Here, we examine what we have learned of concepts that underpin heterogeneous catalysis but more importantly, how we learned to archive our knowledge in context of a set of best practices and standards that have emerged in the course of our learnings—ones we seek to highlight herein. Our perspective emphasizes concepts in synthesis, characterization, kinetics, and theory, because these four elements combined enable description of molecular acts that happen on surfaces and how fast they occur. In revisiting best practices in heterogeneous catalysis in these sub-fields and in authorship and peer review we aspire to augment clarity, reproducibility, and rigor in the science and practice of catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

Towards Interactive, Reproducible Analytics at Scale on HPC Systems

The growth in scientific data volumes has resulted in a need to scale up processing and analysis pipelines using High Performance Computing (HPC) systems. These workflows need interactive, reproducible analytics at scale. The Jupyter platform provides core capabilities for interactivity but was not designed for HPC systems. In this paper, we outline our efforts that bring together core technologies based on the Jupyter Platform to create interactive, reproducible analytics at scale on HPC systems. Our work is grounded in a real world science use case-applying geophysical simulations and inversions for imaging the subsurface. Our core platform addresses three key areas of the scientific analysis workflow-reproducibility, scalability, and interactivity. We describe our implemention of a system, using Binder, Science Capsule, and Dask software. We demonstrate the use of this software to run our use case and interactively visualize real-Time streams of HDF5 data.

containers↗

NASA Open Science Data Repository: Open Science for Life in Space

Space biology and health data are critical for the success of deep space missions and sustainable human presence off-world. At the core of effectively managing biomedical risks is the commitment to open science principles, which ensure that data are findable, accessible, interoperable, reusable, reproducible and maximally open. The 2021 integration of the Ames Life Sciences Data Archive with GeneLab to establish the NASA Open Science Data Repository significantly enhanced access to a wide range of life sciences, biomedical-clinical, and mission telemetry data alongside existing ‘omics data from GeneLab. This paper describes the new database, its architecture, and new data streams supporting diverse data types and enhancing data submission, retrieval, and analysis. Features include the Biological Data Management Environment for improved data submission, a new user interface, controlled data access, an enhanced API, and comprehensive public visualization tools for environmental telemetry, radiation dosimetry data, and ‘omics analyses. By fostering global collaboration through its Analysis Working Groups and training programs, the Open Science Data Repository promotes widespread engagement in space biology, ensuring transparency and inclusivity in research. It supports the global scientific community in advancing our understanding of spaceflight's impact on biological systems, ensuring humans will thrive in future deep space missions.

OSDR↗

A variational multiscale immersed meshfree method for heterogeneous materials

Abstract We introduce an immersed meshfree formulation for modeling heterogeneous materials with flexible non-body-fitted discretizations, approximations, and quadrature rules. The interfacial compatibility condition is imposed by a volumetric constraint, which avoids a tedious contour integral for complex material geometry. The proposed immersed approach is formulated under a variational multiscale based formulation, termed the variational multiscale immersed method (VMIM). Under this framework, the solution approximation on either the foreground or the background can be decoupled into coarse-scale and fine-scale in the variational equations, where the fine-scale approximation represents a correction to the residual of the coarse-scale equations. The resulting fine-scale solution leads to a residual-based stabilization in the VMIM discrete equations. The employment of reproducing kernel (RK) approximation for the coarse- and fine-scale variables allows arbitrary order of continuity in the approximation, which is particularly advantageous for modeling heterogeneous materials. The effectiveness of VMIM is demonstrated with several numerical examples, showing accuracy, stability, and discretization efficiency of the proposed method.

36 MATERIALS SCIENCE↗

Shifting institutional culture to develop climate solutions with Open Science

To address our climate emergency, “we must rapidly, radically reshape society”—Johnson & Wilkinson, All We Can Save. In science, reshaping requires formidable technical (cloud, coding, reproducibility) and cultural shifts (mindsets, hybrid collaboration, inclusion). We are a group of cross-government and academic scientists that are exploring better ways of working and not being too entrenched in our bureaucracies to do better science, support colleagues, and change the culture at our organizations. We share much-needed success stories and action for what we can all do to reshape science as part of the Open Science movement and 2023 Year of Open Science.

54 ENVIRONMENTAL SCIENCES↗

Not just for programmers: How GitHub can accelerate collaborative and reproducible research in ecology and evolution

Abstract Researchers in ecology and evolutionary biology are increasingly dependent on computational code to conduct research. Hence, the use of efficient methods to share, reproduce, and collaborate on code as well as document research is fundamental. GitHub is an online, cloud‐based service that can help researchers track, organize, discuss, share, and collaborate on software and other materials related to research production, including data, code for analyses, and protocols. Despite these benefits, the use of GitHub in ecology and evolution is not widespread. To help researchers in ecology and evolution adopt useful features from GitHub to improve their research workflows, we review 12 practical ways to use the platform. We outline features ranging from low to high technical difficulty, including storing code, managing projects, coding collaboratively, conducting peer review, writing a manuscript, and using automated and continuous integration to streamline analyses. Given that members of a research team may have different technical skills and responsibilities, we describe how the optimal use of GitHub features may vary among members of a research collaboration. As more ecologists and evolutionary biologists establish their workflows using GitHub, the field can continue to push the boundaries of collaborative, transparent, and open research.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Curifactory: A research experiment manager

Curifactory is a command line tool and framework for organizing Python experiment code, configuration parameters, and results. It is an opinionated and lightweight approach to workflow management infrastructure and is primarily intended to support researchers conducting experiments on one machine. This software was developed to support the reproducibility of results for several data science projects in the Nuclear Nonproliferation Division at Oak Ridge National Laboratory. Curifactory is intended to be a general framework and is not specific to machine learning or data science. It can aid in any field in which experiments are primarily computation-based studies and can be implemented in Python (e.g., high-energy physics, astronomy, computational chemistry). Here, the design emphasizes the automated caching of intermediate data analysis artifacts to speed up development involving computationally intensive tasks. It also allows for data provenance and experiment reproduction. Individual experiment runs are tracked through logs and their output reports, and entire copies of a run with all cached data and metadata can be exported for others to run using Curifactory on another machine. Curifactory experiments can either be integrated into a project from the beginning or can be written on top of an existing codebase without needing significant modification. A few important views of the Curifactory library can be seen in Figure 1.

97 MATHEMATICS AND COMPUTING↗

Persistent Identifiers Implementation in EOSDIS

This presentation provides the motivation for and status of implementation of persistent identifiers in NASA's Earth Observation System Data and Information System (EOSDIS). The motivation is provided from the point of view of long-term preservation of datasets such that a number of questions raised by current and future users can be answered easily and precisely. A number of artifacts need to be preserved along with datasets to make this possible, especially when the authors of datasets are no longer available to address users questions. The artifacts and datasets need to be uniquely and persistently identified and linked with each other for full traceability, understandability and scientific reproducibility. Current work in the Earth Science Data and Information System (ESDIS) Project and the Distributed Active Archive Centers (DAACs) in assigning Digital Object Identifiers (DOI) is discussed as well as challenges that remain to be addressed in the future.

persistent identifiers↗

Vibration Isolation Technology (VIT) ATD Project

A fundamental advantage for performing material processing and fluid physics experiments in an orbital environment is the reduction in gravity driven phenomena. However, experience with manned spacecraft such as the Space Transportation System (STS) has demonstrated a dynamic acceleration environment far from being characterized as a 'microgravity' platform. Vibrations and transient disturbances from crew motions, thruster firings, rotating machinery etc. can have detrimental effects on many proposed microgravity science experiments. These same disturbances are also to be expected on the future space station. The Microgravity Science and Applications Division (MSAD) of the Office of Life and Microgravity Sciences and Applications (OLMSA), NASA Headquarters recognized the need for addressing this fundamental issue. As a result an Advanced Technology Development (ATD) project was initiated in the area of Vibration Isolation Technology (VIT) to develop methodologies for meeting future microgravity science needs. The objective of the Vibration Isolation Technology ATD project was to provide technology for the isolation of microgravity science experiments by developing methods to maintain a predictable, well defined, well characterized, and reproducible low-gravity environment, consistent with the needs of the microgravity science community. Included implicitly in this objective was the goal of advising the science community and hardware developers of the fundamental need to address the importance of maintaining, and how to maintain, a microgravity environment. This document will summarize the accomplishments of the VIT ATD which is now completed. There were three specific thrusts involved in the ATD effort. An analytical effort was performed at the Marshall Space Flight Center to define the sensitivity of selected experiments to residual and dynamic accelerations. This effort was redirected about half way through the ATD focusing specifically on the sensitivity of protein crystals to a realistic orbital environment. The other two thrusts of the ATD were performed at the Lewis Research Center. The first was to develop technology in the area of reactionless mechanisms and robotics to support the eventual development of robotics for servicing microgravity science experiments. This activity was completed in 1990. The second was to develop vibration isolation and damping technology providing protection for sensitive science experiments. In conjunction with the this activity, two workshops were held. The results of these were summarized and are included in this report.

Lubomski, Joseph F.↗

Breaking the reproducibility barrier with standardized protocols for plant–microbiome research

Inter-laboratory replicability is crucial yet challenging in microbiome research. Leveraging microbiomes to promote soil health and plant growth requires understanding underlying molecular mechanisms using reproducible experimental systems. In a global collaborative effort involving five laboratories, we aimed to help advance reproducibility in microbiome studies by testing our ability to replicate synthetic community assembly experiments. Our study compared fabricated ecosystems constructed using two different synthetic bacterial communities, the model grass Brachypodium distachyon, and sterile EcoFAB 2.0 devices. All participating laboratories observed consistent inoculum-dependent changes in plant phenotype, root exudate composition, and final bacterial community structure, where Paraburkholderia sp. OAS925 could dramatically shift microbiome composition. Comparative genomics and exudate utilization linked the pH-dependent colonization ability of Paraburkholderia, which was further confirmed with motility assays. The study provides detailed protocols, benchmarking datasets, and best practices to help advance replicable science and inform future multi-laboratory reproducibility studies.

Novak, Vlastimil↗

Using AI to Reproduce Neutrino Cross Section Analysis - Prototyping the Neutrino Discovery Platform

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]↗