Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reproducible science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The Quick Rise and Fall of LK-99 as a Room Temperature Superconductor [Slides]

A swift determination of superconductivity vs non-superconductivity of LK-99 demonstrates the importance of reproducibility in science. This has benefited from 4 decades study of high-temperature cuprate mechanism. Holy grail of room-temperature ambient-pressure superconductor remains to be pursued. Caution is needed about DFT calculations. Crucial old experimental data should be referenced. With the existing amount of experimentally discovered superconductors, could data science make a stride?

36 MATERIALS SCIENCE↗

From Reads to Function Workshop - Milano 2026

The Bicocca Sampling Days (BSDs) model offers a reproducible “citizen science” framework integrating research, education, and public engagement through large-scale microbiome sampling, followed by a workshop of data analysis on select samples. We identified 9 bacterial and archaeal metagenome-assembled genomes from six soil samples across three separate sampling days in two approaches with indidivual sample and replicate co-assembly spanning three unique classes, providing genomic insights into microbial nutrient cycling in these systems.

59 BASIC BIOLOGICAL SCIENCES↗

Reproducible Crystal Growth Experiments in Microgravity Science Glovebox at the International Space Station (SUBSA Investigation)

Solidification Using a Baffle in Sealed Ampoules (SUBSA) is the first investigation conducted in the Microgravity Science Glovebox (MSG) Facility at the International Space Station (ISS) Alpha. 8 single crystals of InSb, doped with Te and Zn, were directionally solidified in microgravity. The experiments were conducted in a furnace with a transparent gradient section, and a video camera, sending images to the earth. The real time images (i) helped seeding, (ii) allowed a direct measurement of the solidification rate. The post-flight characterization of the crystals includes: computed x-ray tomography, Secondary Ion Mass Spectroscopy (SIMS), Hall measurements, Atomic Absorption (AA), and 4 point probe analysis. For the first time in microgravity, several crystals having nearly identical initial transients were grown. Reproducible initial transients were obtained with Te-doped InSb. Furthermore, the diffusion controlled end-transient was demonstrated experimentally (SUBSA 02). From the initial transients, the diffusivity of Te and Zn in InSb was determined.

Ostrogorsky, A.↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

How should reproducibility be approached in plastic recycling?

With the growing importance of developing new and improved methodologies for plastic recycling, conducting reproducible research and ensuring that results are transferable across labs are increasingly important. This Voices article reflects on how academia and industry view the path forward for strengthening reproducibility to advance science and enable a circular plastics economy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

UVCGAN: UNet Vision Transformer cycle-consistent GAN for unpaired image-to-image translation

Unpaired image-to-image translation has broad applications in art, design, and scientific simulations. One early breakthrough was CycleGAN that emphasizes one-to-one mappings between two unpaired image domains via generative-adversarial networks (GAN) coupled with the cycle-consistency constraint, while more recent works promote one-to-many mapping to boost diversity of the translated images. Motivated by scientific simulation and one-to-one needs, this work revisits the classic CycleGAN framework and boosts its performance to outperform more contemporary models without relaxing the cycle-consistency constraint. To achieve this, we equip the generator with a Vision Transformer (ViT) and employ necessary training and regularization techniques. Compared to previous best-performing models, our model performs better and retains a strong correlation between the original and translated image. An accompanying ablation study shows that both the gradient penalty and self-supervised pre-training are crucial to the improvement. To promote reproducibility and open science, the source code, hyperparameter configurations, and pre-trained model are available at https: //github.com/LS4GAN/uvcgan.

97 MATHEMATICS AND COMPUTING↗

Data Readiness for Scientific AI at Scale

This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, bio/health, and materials—to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework that combines canonical preprocessing patterns with a five-level operational readiness scale, both tailored to high-performance computing (HPC) environments. This framework helps outline key challenges in transforming large-scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

gRASPA

GPU Monte Carlo Simulation Code with a taste of RASPA We present enhancements in Monte Carlo simulation speed and functionality within an open-source code, gRASPA, which uses graphical processing units (GPUs) to achieve significant performance improvements compared to serial, CPU implementations of Monte Carlo. The code supports a wide range of Monte Carlo simulations, including canonical ensemble (NVT), grand canonical, NVT Gibbs, Widom test particle insertions, and continuous-fractional component Monte Carlo. Implementation of grand canonical transition matrix Monte Carlo (GC-TMMC) and a novel feature to allow different moves for the different components of metal-organic framework (MOF) structures exemplify the capabilities of gRASPA for precise free energy calculations and enhanced adsorption studies, respectively. The introduction of a High-Throughput Computing (HTC) mode permits many Monte Carlo simulations on a single GPU device for accelerated materials discovery. The code can incorporate machine learning (ML) potentials. The open-source nature of gRASPA promotes reproducibility and openness in science, and users may add features to the code and optimize it for their own purposes. The code is written in CUDA/C++ and SYCL/C++ to support different GPU vendors. The gRASPA code is publicly available at https://github.com/snurr-group/gRASPA.

Li, Zhao [Purdue/Northwestern/Notre Dame Universit↗

[Re] Drivers of evapotranspiration from boreal wildfires

Computational reproducibility is a difficult challenge across science. I attempted to use R 3.6.1 to reproduce linear model fits, done originally using v2.6.0 for a 2009 paper on the drivers of large-scale forest evapotranspiration after wildfire. Model outputs were largely identical, aside from minor formatting changes, except for one–out of 12 total– regression in which the median residual value changed very slightly (in the sixth decimal place). I suggest that this essentially successful reproducibility is due to the relative simplicity of the script, its use of only base R functions, and R’s historically conservative approach to breaking changes.

54 ENVIRONMENTAL SCIENCES↗

Exploring Blockchain to Support Open Science Practices

Open science aims to foster transparent sharing of scientific processes including open access, incentivization, provenance, open source code and tools, metrics, and resource sharing. However, effective management of these processes remains a challenge. This paper explores the application of blockchain technology to address these key aspects of open science. Blockchain offers a decentralized and secureplatform for information exchange and verification. By leveraging blockchain, open science can enhance transparency and reproducibility. In this paper, we present an implementation of blockchain for Earth science data synchronization across organizations, enabling tracking of data copying, citation, anddownload. The findings highlight the potential of blockchain in supporting open science objectives.

Iksha Gurung↗

Why don't we share data and code? Perceived barriers and benefits to public archiving practices

The biological sciences community is increasingly recognizing the value of open, reproducible and transparent research practices for science and society at large. Despite this recognition, many researchers fail to share their data and code publicly. This pattern may arise from knowledge barriers about how to archive data and code, concerns about its reuse, and misaligned career incentives. Here, we define, categorize and discuss barriers to data and code sharing that are relevant to many research fields. We explore how real and perceived barriers might be overcome or reframed in the light of the benefits relative to costs. By elucidating these barriers and the contexts in which they arise, we can take steps to mitigate them and align our actions with the goals of open science, both as individual scientists and as a scientific community.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source↗

From Reproducible Edge–Cloud Experimentation to Real-World Practice: The E2Clab Experience

Reproducibility is already difficult in distributed systems; on the computing continuum, it becomes substantially harder. Applications that span sensing devices, edge and fog resources, and cloud platforms must be evaluated across heterogeneous hardware, variable network conditions, cross-layer orchestration decisions, and long-running workflow lifecycles. We use E2Clab as a case study to examine these challenges and their implications for experimental methodology. We explain why reproducible experimentation is harder on the continuum, then revisit E2Clab as an initial response based on explicit modeling of infrastructure, workflow lifecycle, and artifacts. Lastly, we discuss how its evolution toward more realistic application settings can be understood through the lens of Translational Computer Science. We argue that reproducible continuum experimentation requires methods that are rigorous enough for research while remaining adaptable to real-world practice.

42 ENGINEERING↗

Good practices for documenting AI-based studies on energy and buildings

Artificial intelligence has transformed building science research over the past decade, with applications spanning energy modeling, energy prediction, HVAC optimization and controls, fault detection, and occupancy modeling. However, many studies lack adequate documentation of datasets, algorithms, training procedures, and validation methods. Building science research faces additional challenges including inconsistent evaluation metrics, limited generalizability across building types, climates, and significant gaps between experimental studies and deployed systems. This communication provides practical guidance for good practices in documenting and publishing AI-based research following established standards from the computer science and machine learning communities. By adopting frameworks such as Datasheets for Datasets, Model Cards, and standardized reproducibility checklists, researchers can ensure their work meets the rigorous documentation standards necessary for reproducible, comparable, and impactful building science research.

Hong, Tianzhen [Lawrence Berkeley National Laborat↗

The fortedata R package: open-science datasets from a manipulative experiment testing forest resilience

The fortedata R package is an open data notebook from the Forest Resilience Threshold Experiment (FoRTE) – a modeling and manipulative field experiment that tests the effects of disturbance severity and disturbance type on carbon cycling dynamics in a temperate forest. Package data consist of measurements of carbon pools and fluxes and ancillary measurements to help analyze and interpret carbon cycling over time. Currently the package includes data and metadata from the first three FoRTE field seasons, serves as a central, updatable resource for the FoRTE project team, and is intended as a resource for external users over the course of the experiment and in perpetuity. Further, it supports all associated FoRTE publications, analyses, and modeling efforts. This increases efficiency, consistency, compatibility, and productivity while minimizing duplicated effort and error propagation that can arise as a function of a large, distributed and collaborative effort. More broadly, fortedata represents an innovative, collaborative way of approaching science that unites and expedites the delivery of complementary datasets to the broader scientific community, increasing transparency and reproducibility of taxpayer-funded science. The fortedata package is available via GitHub: https://github.com/FoRTExperiment/fortedata (last access: 19 February 2021), and detailed documentation on the access, used, and applications of fortedata are available at https://fortexperiment.github.io/fortedata/ (last access: 19 February 2021). The first public release, version 1.0.1 is also archived at https://doi.org/10.5281/zenodo.4399601 (Atkins et al., 2020b). All data products are also available outside of the package as .csv files: https://doi.org/10.6084/m9.figshare.13499148.v1 (Atkins et al., 2020c).

97 MATHEMATICS AND COMPUTING↗

LLaMP v0.1.0

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. LLaMP is a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning and acting (ReAct) agents that can dynamically and recursively interact with Materials Project to ground large language models on high-fidelity materials informatics.

Riebesell, Janosh [Lawrence Berkeley National Labo↗

Evolution of NASA's Earth Science Digital Object Identifier Registration System

NASA's Earth Science Data and Information System (ESDIS) Project has implemented a fully automated system for assigning Digital Object Identifiers (DOIs) to Earth Science data products being managed by its network of 12 distributed active archive centers (DAACs). A key factor in the successful evolution of the DOI registration system over last 7 years has been the incorporation of community input from three focus groups under the NASA's Earth Science Data System Working Group (ESDSWG). These groups were largely composed of DOI submitters and data curators from the 12 data centers serving the user communities of various science disciplines. The suggestions from these groups were formulated into recommendations for ESDIS consideration and implementation. The ESDIS DOI registration system has evolved to be fully functional with over 5,000 publicly accessible DOIs and over 200 DOIs being held in reserve status until the information required for registration is obtained. The goal is to assign DOIs to the entire 8000+ data collections under ESDIS management via its network of discipline-oriented data centers. DOIs make it easier for researchers to discover and use earth science data and they enable users to provide valid citations for the data they use in research. Also for the researcher wishing to reproduce the results presented in science publications, the DOI can be used to locate the exact data or data products being cited.

DOI; digital object identifier; mappin↗

Efforts to enhance reproducibility in a human performance research project

Background: Ensuring the validity of results from funded programs is a critical concern for agencies that sponsor biological research. In recent years, the open science movement has sought to promote reproducibility by encouraging sharing not only of finished manuscripts but also of data and code supporting their findings. While these innovations have lent support to third-party efforts to replicate calculations underlying key results in the scientific literature, fields of inquiry where privacy considerations or other sensitivities preclude the broad distribution of raw data or analysis may require a more targeted approach to promote the quality of research output. Methods: We describe efforts oriented toward this goal that were implemented in one human performance research program, Measuring Biological Aptitude, organized by the Defense Advanced Research Project Agency's Biological Technologies Office. Our team implemented a four-pronged independent verification and validation (IV&V) strategy including 1) a centralized data storage and exchange platform, 2) quality assurance and quality control (QA/QC) of data collection, 3) test and evaluation of performer models, and 4) an archival software and data repository. Results: Our IV&V plan was carried out with assistance from both the funding agency and participating teams of researchers. QA/QC of data acquisition aided in process improvement and the flagging of experimental errors. Holdout validation set tests provided an independent gauge of model performance. Conclusions: In circumstances that do not support a fully open approach to scientific criticism, standing up independent teams to cross-check and validate the results generated by primary investigators can be an important tool to promote reproducibility of results.

59 BASIC BIOLOGICAL SCIENCES↗