Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reproducible research”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Materials characterization: Can artificial intelligence be used to address reproducibility challenges?

Material characterization techniques are widely used to characterize the physical and chemical properties of materials at the nanoscale and, thus, play central roles in material scientific discoveries. However, the large and complex datasets generated by these techniques often require significant human effort to interpret and extract meaningful physicochemical insights. Artificial intelligence (AI) techniques such as machine learning (ML) have the potential to improve the efficiency and accuracy of surface analysis by automating data analysis and interpretation. In this perspective paper, we review the current role of AI in surface analysis and discuss its future potential to accelerate discoveries in surface science, materials science, and interface science. We highlight several applications where AI has already been used to analyze surface analysis data, including the identification of crystal structures from XRD data, analysis of XPS spectra for surface composition, and the interpretation of TEM and SEM images for particle morphology and size. We also discuss the challenges and opportunities associated with the integration of AI into surface analysis workflows. These include the need for large and diverse datasets for training ML models, the importance of feature selection and representation, and the potential for ML to enable new insights and discoveries by identifying patterns and relationships in complex datasets. Most importantly, AI analyzed data must not just find the best mathematical description of the data, but it must find the most physical and chemically meaningful results. In addition, the need for reproducibility in scientific research has become increasingly important in recent years. The advancement of AI, including both conventional and the increasing popular deep learning, is showing promise in addressing those challenges by enabling the execution and verification of scientific progress. By training models on large experimental datasets and providing automated analysis and data interpretation, AI can help to ensure that scientific results are reproducible and reliable. Although integration of knowledge and AI models must be considered for the transparency and interpretability of models, the incorporation of AI into the data collection and processing workflow will significantly enhance the efficiency and accuracy of various surface analysis techniques and deepen our understanding at an accelerated pace.

Materials Science↗

Standards, dissemination, and best practices in systems biology

In this study, the reproducibility of scientific research is crucial to the success of the scientific method. Here, we review the current best practices when publishing mechanistic models in systems biology. We recommend, where possible, to use software engineering strategies such as testing, verification, validation, documentation, versioning, iterative development, and continuous integration. In addition, adhering to the Findable, Accessible, Interoperable, and Reusable modeling principles allows other scientists to collaborate and build off of each other’s work. Existing standards such as Systems Biology Markup Language, CellML, or Simulation Experiment Description Markup Language can greatly improve the likelihood that a published model is reproducible, especially if such models are deposited in well-established model repositories. Where models are published in executable programming languages, the source code and their data should be published as open-source in public code repositories together with any documentation and testing code. For complex models, we recommend container-based solutions where any software dependencies and the run-time context can be easily replicated.

59 BASIC BIOLOGICAL SCIENCES↗

Anion-exchange membrane water electrolysis: insights from round-robin testing

As research and industrial interest in anion-exchange membrane water electrolysis (AEMWE) grows, there is an increasing need for reliable baselines and cross-lab validation of results. The wide variety of material sets and operating conditions under consideration for AEMWE has thus far limited efforts for standardization. In this study, round-robin testing was conducted in deionized water and KOH-based supporting electrolyte by 5 institutions from academia, national laboratories, and industry to provide baseline performance data and identify sources of cross-lab variability. Baseline membrane electrode assemblies were fabricated with commercial catalysts, membranes, and transport layers using standard techniques and tested using reagent-grade electrolytes, aiming for accessibility rather than state-of-the-art performance. From all tests, the average voltage at 1 A/cm 2 was 2.72 ± 0.17 V and 1.87 ± 0.03 V in deionized water and 0.1 M KOH, respectively. The maximum in-house and cross-lab variations at this current density were 118 mV and 476 mV in water and 60 and 88 mV in 0.1 M KOH. The KOH purity, station contamination, and temperature control were identified as possible factors affecting performance between labs, with in-house specific variation attributed to sample-to-sample differences in fabrication, cell assembly, and station contamination. This work provides a commercial baseline for the field and highlights the need for improved standardization and reproducibility in AEMWE research.

08 HYDROGEN↗

Why don't we share data and code? Perceived barriers and benefits to public archiving practices

The biological sciences community is increasingly recognizing the value of open, reproducible and transparent research practices for science and society at large. Despite this recognition, many researchers fail to share their data and code publicly. This pattern may arise from knowledge barriers about how to archive data and code, concerns about its reuse, and misaligned career incentives. Here, we define, categorize and discuss barriers to data and code sharing that are relevant to many research fields. We explore how real and perceived barriers might be overcome or reframed in the light of the benefits relative to costs. By elucidating these barriers and the contexts in which they arise, we can take steps to mitigate them and align our actions with the goals of open science, both as individual scientists and as a scientific community.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Polarized Deep-Inelastic Scattering with Spin Correlations in Herwig 7

This repository is the research software and reproducibility companion for the HerwigPol polarized deep-inelastic scattering implementation developed for Herwig 7. It brings together the modified Herwig and ThePEG source snapshots, the curated POLDIS fixed-order reference code, the custom Rivet analyses, the DIS validation workflow, and the paper source in a single formal repository layout. The repository is intended to preserve the source-level ingredients needed to rebuild and re-run the validated DIS studies. It therefore tracks code, input cards, workflow drivers, and technical notes, while intentionally excluding generated artifacts such as build products, campaign outputs, merged YODA files, plots, and rendered paper outputs.

Papaefstathioou, Andreas [Kennesaw State Universit↗

Ontology Engineering in Provenance Enablement for the National Climate Assessment

The National Climate Assessment of the U.S. Global Change Research Program (USGCRP) analyzes and presents the impacts of climate change on the United States. The provenance information in the assessment is important because the assessment findings are of great public and academic concern and are used in policy and decision-making. By applying a use case-driven iterative methodology, we developed information models and ontology to represent the content structure of the recent National Climate Assessment draft report and its associated provenance information. We tested the ontology by using it in pilot systems serving information about instances of chapters, scientific findings, figures, tables, images, datasets, references, people, and organizations, etc. in the draft report, as well as interrelationships among those instances. The results successfully help users trace provenance in the draft report, such as finding all the journal articles from which a figure in the report was derived. The provenance information in our work was maintained in the context of the "Web of Data". In addition to the pilot systems we developed, other tools and services are also able to retrieve and utilize the provenance information. Our work is part of a Global Change Information System coordinated by the USGCRP that will eventually cover provenance information for the entire scope of global change research. Such a system will greatly increase understanding, credibility and trust in the global change research and foster reproducibility of scientific results and conclusions.

Ontology engineering↗

Evolution of NASA's Earth Science Digital Object Identifier Registration System

NASA's Earth Science Data and Information System (ESDIS) Project has implemented a fully automated system for assigning Digital Object Identifiers (DOIs) to Earth Science data products being managed by its network of 12 distributed active archive centers (DAACs). A key factor in the successful evolution of the DOI registration system over last 7 years has been the incorporation of community input from three focus groups under the NASA's Earth Science Data System Working Group (ESDSWG). These groups were largely composed of DOI submitters and data curators from the 12 data centers serving the user communities of various science disciplines. The suggestions from these groups were formulated into recommendations for ESDIS consideration and implementation. The ESDIS DOI registration system has evolved to be fully functional with over 5,000 publicly accessible DOIs and over 200 DOIs being held in reserve status until the information required for registration is obtained. The goal is to assign DOIs to the entire 8000+ data collections under ESDIS management via its network of discipline-oriented data centers. DOIs make it easier for researchers to discover and use earth science data and they enable users to provide valid citations for the data they use in research. Also for the researcher wishing to reproduce the results presented in science publications, the DOI can be used to locate the exact data or data products being cited.

DOI; digital object identifier; mappin↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering↗

Editorial: Towards the rapid and systematic assessment of vaccine technologies

The COVID-19 pandemic highlighted both the extraordinary potential of modern vaccinology and persistent challenges in how vaccine technologies are assessed. While vaccines can be developed and deployed at unprecedented speed, our ability to predict efficacy in a population is constrained by methodological difficulties, underreporting of negative results, and limited comparability across studies. This editorial introduces a Research Topic that brings together an interdisciplinary collection of experimental, computational, and theoretical contributions spanning multiple pathogens and vaccine platforms. Across these contributions, emerging themes emphasize the need for standardized immunogenicity metrics, transparent reporting including negative findings, and harmonized experimental protocols to support meaningful comparisons. This editorial highlights community practices and shared commitments – supported by researchers, funders, and journals – that could strengthen reproducibility, transparency, and cumulative learning in vaccine research.

59 BASIC BIOLOGICAL SCIENCES↗

A comprehensive guide to CAN IDS data and introduction of the ROAD dataset

Although ubiquitous in modern vehicles, Controller Area Networks (CANs) lack basic security properties and are easily exploitable. A rapidly growing field of CAN security research has emerged that seeks to detect intrusions or anomalies on CANs. Producing vehicular CAN data with a variety of intrusions is a difficult task for most researchers as it requires expensive assets and deep expertise. To illuminate this task, we introduce the first comprehensive guide to the existing open CAN intrusion detection system (IDS) datasets. We categorize attacks on CANs including fabrication (adding frames, e.g., flooding or targeting and ID), suspension (removing an ID’s frames), and masquerade attacks (spoofed frames sent in lieu of suspended ones). We provide a quality analysis of each dataset; an enumeration of each datasets’ attacks, benefits, and drawbacks; categorization as real vs. simulated CAN data and real vs. simulated attacks; whether the data is raw CAN data or signal-translated; number of vehicles/CANs; quantity in terms of time; and finally a suggested use case of each dataset. State-of-the-art public CAN IDS datasets are limited to real fabrication (simple message injection) attacks and simulated attacks often in synthetic data, lacking fidelity. In general, the physical effects of attacks on the vehicle are not verified in the available datasets. Only one dataset provides signal-translated data but is missing a corresponding “raw” binary version. This issue pigeon-holes CAN IDS research into testing on limited and often inappropriate data (usually with attacks that are too easily detectable to truly test the method). The scarcity of appropriate data has stymied comparability and reproducibility of results for researchers. As our primary contribution, we present the Real ORNL Automotive Dynamometer (ROAD) CAN IDS dataset, consisting of over 3.5 hours of one vehicle’s CAN data. ROAD contains ambient data recorded during a diverse set of activities, and attacks of increasing stealth with multiple variants and instances of real (i.e. non-simulated) fuzzing, fabrication, unique advanced attacks, and simulated masquerade attacks. To facilitate a benchmark for CAN IDS methods that require signal-translated inputs, we also provide the signal time series format for many of the CAN captures. Our contributions aim to facilitate appropriate benchmarking and needed comparability in the CAN IDS research field.

97 MATHEMATICS AND COMPUTING↗

HamLib: A library of Hamiltonians for benchmarking quantum algorithms and hardware

In order to characterize and benchmark computational hardware, software, and algorithms, it is essential to have many problem instances on-hand. This is no less true for quantum computation, where a large collection of real-world problem instances would allow for benchmarking studies that in turn help to improve both algorithms and hardware designs. To this end, here we present a large dataset of qubit-based quantum Hamiltonians. The dataset, called HamLib (for Hamiltonian Library), is freely available online and contains problem sizes ranging from 2 to 1000 qubits. HamLib includes problem instances of the Heisenberg model, Fermi-Hubbard model, Bose-Hubbard model, molecular electronic structure, molecular vibrational structure, MaxCut, Max- k -SAT, Max- k -Cut, QMaxCut, and the traveling salesperson problem. The goals of this effort are (a) to save researchers time by eliminating the need to prepare problem instances and map them to qubit representations, (b) to allow for more thorough tests of new algorithms and hardware, and (c) to allow for reproducibility and standardization across research studies.

97 MATHEMATICS AND COMPUTING↗

A Novel Dosimetry Method for Small Animal Irradiators Using 3D-printed Mouse Phantoms and Alanine Dosimeters

Abstract Accurate dosimetry is a crucial component of small animal and preclinical irradiation studies. Various dosimetry options are available but fail to characterize complex geometries and variable energy spectra of modern x-ray irradiators accurately. These options also lack national/international standards of recognition. This paper presents a novel dosimetry system, Dosequate, which uses murine phantoms embedded with alanine dosimeters and paired with x-ray energy spectra corrections to deliver consistent and accurate dosimetry measurements. This study compares Dosequate measurements against a Precision X-RAD 320 internal ion chamber and the treatment planning system of an Xstrahl Small Animal Radiation Research Platform. Results demonstrate the accuracy and reproducibility of the Dosequate system, highlighting its potential for standardizing dosimetry in preclinical research throughout the industry. Results demonstrate the accuracy and reproducibility of the Dosequate system, highlighting its potential for standardizing dosimetry in preclinical research throughout the industry. The Dosequate dosimetry method provides a robust and standardized approach for measuring absorbed dose in small animal irradiators. Its accuracy, reproducibility, and ability to account for complex irradiation geometries make it a valuable tool for preclinical research. This system has the potential to significantly improve the intercomparability of studies across different facilities and enhance the reliability of results.

Environmental Sciences & Ecology↗

Photocatalytic water splitting

With the goal of achieving large-scale H 2 production from renewable resources, water splitting into H 2 and O 2 using semiconductor photocatalysts (sometimes called artificial photosynthesis) has been studied for five decades. Unfortunately, the lack of rigour and reproducibility in the data collection and analysis of experimental results has hindered progress in the field. This Primer provides a comprehensive overview of proper characterization and evaluation of photocatalysts for overall water splitting. In particular, the Primer covers various pitfalls in photocatalysis research, best practices for reproducibility and reliable methods for conducting rigorous experiments. As a result, the recommendations are intended to reduce false positives in the literature and to promote progress towards a practical technology for producing H 2 from water by using sunlight.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NASA Earth eXchange (NEX) App Store

NASA Earth Exchange (NEX), and her public cloud version OpenNEX, have become platforms supporting scientific collaboration, knowledge sharing and research for the entire Earth science community. To date, a number of custom tools and capabilities have been integrated into the platforms. However, such integration has to undergo a case-by-case manual process thus lacks scalability. This timely project builds an App Store onto OpenNEX as a building block. Climate data analytics tools/programs can be easily uploaded, shared, organized, searched, and recommended like photos and videos on the YouTube. The foundation of our App Store is a provenance server, which not only records metadata but also execution history of climate data analytics apps including the input data and parameters, output data and products, who runs the app for which purpose, and how apps may be chained into workflows. Researchers can thus understand, reproduce, and repurpose existing apps and workflows. Machine learning approaches are applied to mine provenance to provide recommend-as-you-go services for Earth scientists, such as to recommend suitable apps and workflow snippets. A browser-based workflow tool is also provided for researchers to explore the provenance server and design value-added workflows. Scalability, sustainability, extensibility, usability, adaptability, security and privacy are considered in the App Store.

eXchange↗

rSHUD v2.0: advancing the Simulator for Hydrologic Unstructured Domains and unstructured hydrological modeling in the R environment

Abstract. Hydrological modeling is a crucial component in hydrology research, particularly for projecting future scenarios. However, achieving reproducibility and automation in distributed hydrological modeling research for modeling, simulation, and analysis is challenging. This paper introduces rSHUD v2.0, an innovative, open-source toolkit developed in the R environment to enhance the deployment and analysis of the Simulator for Hydrologic Unstructured Domains (SHUD). The SHUD is an integrated surface–subsurface hydrological model that employs a finite-volume method to simulate hydrological processes at various scales. The rSHUD toolkit includes pre- and post-processing tools, facilitating reproducibility and automation in hydrological modeling. The utility of rSHUD is demonstrated through case studies of the Shale Hills Critical Zone Observatory in the USA and the Waerma watershed in China. The rSHUD toolkit's ability to quickly and automatically deploy models while ensuring reproducibility has facilitated the implementation of the Global Hydrological Data Cloud (https://ghdc.ac.cn, last access: 1 September 2023), a platform for automatic data processing and model deployment. This work represents a significant advancement in hydrological modeling, with implications for future scenario projections and spatial analysis.

Shu, Lele (ORCID:0000000269034466)↗

Material Needs and Measurement Challenges for Advanced Semiconductor Packaging: Understanding the Soft Side of Science

This Perspective builds upon insights from the National Institute of Standards and Technology (NIST)-organized workshop, “Materials and Metrology Needs for Advanced Semiconductor Packaging Strategies,” held at the 35th annual Electronics Packaging Symposium in Binghamton, NY, on September 5, 2024. It outlines critical challenges and opportunities related to polymer-based “soft” materials in advanced semiconductor packaging, with emphasis on polymer science, measurement science (metrology), and the strategic development of Research-Grade Test Materials (RGTMs). These efforts, led by the NIST CHIPS team, aim to advance the fundamental understanding of structure-property-processing relationships, promote standardized guidelines and innovative methods for material characterization, and accelerate the development, qualification, and adoption of next-generation packaging materials. The Perspective also distills key insights from the panel discussion with industry experts, emphasizing the need for close collaboration among materials scientists, process engineers, and metrology experts to enable a holistic strategy, further highlighting the importance of cross-sector partnerships among industry, academia, and government to address pressing challenges in packaging materials and processes.

97 MATHEMATICS AND COMPUTING↗

EDD Basic Stats and Graphs Notebook analysis (EDD BSG Notebook) v1.0

This jupyter notebook calculates basic statistics (e.g., mean, standard deviation, coefficient of variation) and simple graphs (e.g., bar graphs, line plots) for data from the Experiment Data Depot (EDD) to provide rapid and reproducible assessment of data quality to aid research efforts across the JBEI and ABF projects. It rapidly and reproducibly calculates basic statistical values for data stored in the EDD which aids researchers and strengthens comparisons across different experiments and projects.

Petzold, ChristopherJ↗

Machine learning for surrogate process models of bioproduction pathways

Technoeconomic analysis and life-cycle assessment are critical to guiding and prioritizing bench-scale experiments and to evaluating economic and environmental performance of biofuel or biochemical production processes at scale. Traditionally, commercial process simulation tools have been used to develop detailed models for these purposes. However, developing and running such models can be costly and computationally intensive, which limits the degree to which they can be shared and reproduced in the broader research community. This study evaluates the potential of an automated machine learning approach to develop surrogate models based on conventional process simulation models. The analysis focuses on several high-value biofuels and bioproducts for which pathways of production from biomass feedstocks have been well-established. The results demonstrate that surrogate models can be an accurate and effective tool for approximating the cost, mass and energy balance outputs of more complex process simulations at a fraction of the computational expense.

09 BIOMASS FUELS↗