Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “reproducible research”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Explaining and predicting human behavior and social dynamics in simulated virtual worlds: reproducibility, generalizability, and robustness of causal discovery methods

Ground Truth program was designed to evaluate social science modeling approaches using simulation test beds with ground truth intentionally and systematically embedded to understand and model complex Human Domain systems and their dynamics Lazer et al. (Science 369:1060–1062, 2020). Our multidisciplinary team of data scientists, statisticians, experts in Artificial Intelligence (AI) and visual analytics had a unique role on the program to investigate accuracy, reproducibility, generalizability, and robustness of the state-of-the-art (SOTA) causal structure learning approaches applied to fully observed and sampled simulated data across virtual worlds. In addition, we analyzed the feasibility of using machine learning models to predict future social behavior with and without causal knowledge explicitly embedded. In this paper, we first present our causal modeling approach to discover the causal structure of four virtual worlds produced by the simulation teams—Urban Life, Financial Governance, Disaster and Geopolitical Conflict. Our approach adapts the state-of-the-art causal discovery (including ensemble models), machine learning, data analytics, and visualization techniques to allow a human-machine team to reverse-engineer the true causal relations from sampled and fully observed data. We next present our reproducibility analysis of two research methods team’s performance using a range of causal discovery models applied to both sampled and fully observed data, and analyze their effectiveness and limitations. We further investigate the generalizability and robustness to sampling of the SOTA causal discovery approaches on additional simulated datasets with known ground truth. Our results reveal the limitations of existing causal modeling approaches when applied to large-scale, noisy, high-dimensional data with unobserved variables and unknown relationships between them. We show that the SOTA causal models explored in our experiments are not designed to take advantage from vasts amounts of data and have difficulty recovering ground truth when latent confounders are present; they do not generalize well across simulation scenarios and are not robust to sampling; they are vulnerable to data and modeling assumptions, and therefore, the results are hard to reproduce. Finally, when we outline lessons learned and provide recommendations to improve models for causal discovery and prediction of human social behavior from observational data, we highlight the importance of learning data to knowledge representations or transformations to improve causal discovery and describe the benefit of causal feature selection for predictive and prescriptive modeling.

97 MATHEMATICS AND COMPUTING↗

Material Needs and Measurement Challenges for Advanced Semiconductor Packaging: Understanding the Soft Side of Science

This Perspective builds upon insights from the National Institute of Standards and Technology (NIST)-organized workshop, “Materials and Metrology Needs for Advanced Semiconductor Packaging Strategies,” held at the 35th annual Electronics Packaging Symposium in Binghamton, NY, on September 5, 2024. It outlines critical challenges and opportunities related to polymer-based “soft” materials in advanced semiconductor packaging, with emphasis on polymer science, measurement science (metrology), and the strategic development of Research-Grade Test Materials (RGTMs). These efforts, led by the NIST CHIPS team, aim to advance the fundamental understanding of structure-property-processing relationships, promote standardized guidelines and innovative methods for material characterization, and accelerate the development, qualification, and adoption of next-generation packaging materials. The Perspective also distills key insights from the panel discussion with industry experts, emphasizing the need for close collaboration among materials scientists, process engineers, and metrology experts to enable a holistic strategy, further highlighting the importance of cross-sector partnerships among industry, academia, and government to address pressing challenges in packaging materials and processes.

97 MATHEMATICS AND COMPUTING↗

EDD Basic Stats and Graphs Notebook analysis (EDD BSG Notebook) v1.0

This jupyter notebook calculates basic statistics (e.g., mean, standard deviation, coefficient of variation) and simple graphs (e.g., bar graphs, line plots) for data from the Experiment Data Depot (EDD) to provide rapid and reproducible assessment of data quality to aid research efforts across the JBEI and ABF projects. It rapidly and reproducibly calculates basic statistical values for data stored in the EDD which aids researchers and strengthens comparisons across different experiments and projects.

Petzold, ChristopherJ↗

Machine learning for surrogate process models of bioproduction pathways

Technoeconomic analysis and life-cycle assessment are critical to guiding and prioritizing bench-scale experiments and to evaluating economic and environmental performance of biofuel or biochemical production processes at scale. Traditionally, commercial process simulation tools have been used to develop detailed models for these purposes. However, developing and running such models can be costly and computationally intensive, which limits the degree to which they can be shared and reproduced in the broader research community. This study evaluates the potential of an automated machine learning approach to develop surrogate models based on conventional process simulation models. The analysis focuses on several high-value biofuels and bioproducts for which pathways of production from biomass feedstocks have been well-established. The results demonstrate that surrogate models can be an accurate and effective tool for approximating the cost, mass and energy balance outputs of more complex process simulations at a fraction of the computational expense.

09 BIOMASS FUELS↗

Abstraction hierarchy to define biofoundry workflows and operations for interoperable synthetic biology research and applications

Lack of standardization in biofoundries limits the scalability and efficiency of synthetic biology research. Here, we propose an abstraction hierarchy that organizes biofoundry activities into four interoperable levels: Project, Service/Capability, Workflow, and Unit Operation, effectively streamlining the Design‑Build‑Test‑Learn (DBTL) cycle. This framework enables more modular, flexible, and automated experimental workflows. It improves communication between researchers and systems, supports reproducibility, and facilitates better integration of software tools and artificial intelligence. Our approach lays the foundation for a globally interoperable biofoundry network, advancing collaborative synthetic biology and accelerating innovation in response to scientific and societal challenges.

Kim, Haseong↗

Exploring the roles of microbes in facilitating plant adaptation to climate change

Plants benefit from their close association with soil microbes which assist in their response to abiotic and biotic stressors. Yet much of what we know about plant stress responses is based on studies where the microbial partners were uncontrolled and unknown. Under climate change, the soil microbial community will also be sensitive to and respond to abiotic and biotic stressors. Thus, facilitating plant adaptation to climate change will require a systems-based approach that accounts for the multi-dimensional nature of plant–microbe–environment interactions. In this perspective, we highlight some of the key factors influencing plant–microbe interactions under stress as well as new tools to facilitate the controlled study of their molecular complexity, such as fabricated ecosystems and synthetic communities. When paired with genomic and biochemical methods, these tools provide researchers with more precision, reproducibility, and manipulability for exploring plant–microbe–environment interactions under a changing climate.

59 BASIC BIOLOGICAL SCIENCES↗

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES↗

JetNet: A Python package for accessing open datasets and benchmarking machine learning methods in high energy physics

JetNet is a Python package that aims to increase accessibility and reproducibility for machinelearning (ML) research in high energy physics (HEP), primarily related to particle jets. Basedon the popular PyTorch ML framework, it provides easy-to-access and standardized interfacesfor multiple heterogeneous HEP datasets and implementations of evaluation metrics, lossfunctions, and more general utilities relevant to HEP.

97 MATHEMATICS AND COMPUTING↗

WINDPROF: Merged Best-Estimate Wind Profile Data – Block Island (WFIP3 Campaign)

WINDPROF provides 10-minute wind and turbulence profiles, integrating Doppler lidars, wind profiling radars, and sonic anemometers across Northeast U.S. coastal/offshore sites during the WFIP3 campaign. Key data include wind speed, direction, vertical velocity, and turbulence parameters, with standardized quality control (e.g., instrument-specific thresholds and inter-instrument validation). Profiles are interpolated to a height grid (20 m spacing below 100 m; 30 m above) and include comprehensive uncertainty estimates. The Block Island dataset covers February 2024–September 2025, offering reproducible methods for atmospheric research, model validation, and wind energy studies.

17 WIND ENERGY↗

WINDPROF: Merged Best-Estimate Wind Profile Data – Nantucket (WFIP3 Campaign)

WINDPROF provides 10-minute wind and turbulence profiles, integrating Doppler lidars, wind profiling radars, and sonic anemometers across Northeast U.S. coastal/offshore sites during the WFIP3 campaign. Key data include wind speed, direction, vertical velocity, and turbulence parameters, with standardized quality control (e.g., instrument-specific thresholds and inter-instrument validation). Profiles are interpolated to a height grid (20 m spacing below 100 m; 30 m above) and include comprehensive uncertainty estimates. The Nantucket dataset covers February 2024–September 2025, offering reproducible methods for atmospheric research, model validation, and wind energy studies.

17 WIND ENERGY↗

WINDPROF: Merged Best-Estimate Wind Profile Data – Site A1 (AWAKEN Campaign)

WINDPROF provides 10-minute wind and turbulence profiles, integrating Doppler lidars and anemometers during the AWAKEN campaign. Key data include wind speed, direction, vertical velocity, and turbulence parameters, with standardized quality control (e.g., instrument-specific thresholds and inter-instrument validation). Profiles are interpolated to a height grid (20 m spacing below 100 m; 30 m above) and include uncertainty estimates, offering reproducible methods for atmospheric research, model validation, and wind energy studies.

17 WIND ENERGY↗

WINDPROF: Merged Best-Estimate Wind Profile Data – Site A2 (AWAKEN Campaign)

WINDPROF provides 10-minute wind and turbulence profiles, integrating Doppler lidars and anemometers during the AWAKEN campaign. Key data include wind speed, direction, vertical velocity, and turbulence parameters, with standardized quality control (e.g., instrument-specific thresholds and inter-instrument validation). Profiles are interpolated to a height grid (20 m spacing below 100 m; 30 m above) and include uncertainty estimates, offering reproducible methods for atmospheric research, model validation, and wind energy studies.

17 WIND ENERGY↗

WINDPROF: Merged Best-Estimate Wind Profile Data – Site H (AWAKEN Campaign)

WINDPROF provides 10-minute wind and turbulence profiles, integrating Doppler lidars and anemometers during the AWAKEN campaign. Key data include wind speed, direction, vertical velocity, and turbulence parameters, with standardized quality control (e.g., instrument-specific thresholds and inter-instrument validation). Profiles are interpolated to a height grid (20 m spacing below 100 m; 30 m above) and include uncertainty estimates, offering reproducible methods for atmospheric research, model validation, and wind energy studies.

17 WIND ENERGY↗

Input files and WRF run directory for a LASSO-CACTI simulation

Tar file containing files necessary for reproducing a given Weather Research and Forecasting (WRF) simulation from the Large-Eddy Simulation (LES) Atmospheric Radiation Measurement (ARM) Symbiotic Simulation and Observation (LASSO) deep-convection scenario for the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign. The LASSO-CACTI simulations span grid spacings from 7.5 km to 100 m for convection near the Sierras de Córdoba mountain range, roughly centered on the ARM Mobile Facility. More information can be found at https://www.arm.gov/capabilities/modeling/lasso. Types of files in this tar include the initial and boundary conditions, WRF namelist, and other files present in a WRF run directory.

54 ENVIRONMENTAL SCIENCES↗

Are Fullerenes Relevant to Cosmochemistry? A New Finding

The abundances of noble gases found in primitive, carbonaceous meteorites are unexpected when compared with our Sun. Known as Q-gases (Q for some unknown carrier dubbed quintessence ), this anomaly has remained a mystery since it was discovered in 1975. Q-gases are characterized by increasing depletions with decreasing atomic number (Z) relative to solar noble gases and normalized to 132Xe (Figure 1). This Q-gas mass fractionation is unexplained, and its investigation is important to understanding the origin of the solar system. However, the subject is fraught with controversy, in part due to the complex nature of Q and in part due to claims of some researchers that cannot be reproduced by other investigators. The topic is discussed in numerous places [e.g., 1-4], with models of Q falling into two basic categories, both involving carbon entrapment of noble gases. First (Group A), there is the conservative two-dimensional view that Q-gases are adsorbed or sorbed onto a "labyrinth" of graphite or carbon grains [5-9], or they undergo active capture onto growing surfaces [6]. Second (Group B), there is the view holding to the remarkable property of carbon discovered in 1985. Carbon can curl up into closed geometries of hexagon- and pentagon-shaped carbon-ring configurations, a property ignored completely by Group A. Group B thinks of Q as a three-dimensional structure of endohedral carbon cages like fullerenes, carbon onions, or some class of carbon nanotubes [3, 4, 10]. Group B does not exclude Group A effects.

Wilson, T. L.↗

An information maximization model of eye movements

We propose a sequential information maximization model as a general strategy for programming eye movements. The model reconstructs high-resolution visual information from a sequence of fixations, taking into account the fall-off in resolution from the fovea to the periphery. From this framework we get a simple rule for predicting fixation sequences: after each fixation, fixate next at the location that minimizes uncertainty (maximizes information) about the stimulus. By comparing our model performance to human eye movement data and to predictions from a saliency and random model, we demonstrate that our model is best at predicting fixation locations. Modeling additional biological constraints will improve the prediction of fixation sequences. Our results suggest that information maximization is a useful principle for programming eye movements.

NASA Discipline Neuroscience↗

DATASET RELEASE AND QUALITY CONTROL REVIEW OF LIVERMORE NEVADA NETWORK (LNN) RECORDINGS OF A SUBSET OF NEVADA NUCLEAR SECURITY SITE NUCLEAR EXPLOSIONS FROM 1979 TO 1992.

Geophysical research on historical nuclear tests is an important aspect of future monitoring capabilities in seismic research. This research is challenging due to the limited number of digital seismic recordings during the peak of nuclear testing (1945-1992). These limited records are unique and non-reproducible data with potential high research impact. Releasing available nuclear explosion seismic records to the explosion monitoring community is thus of high value and is the motivation for this dataset release. The target of this effort was on compilation and quality control of regional seismic records of nuclear explosions recorded on Lawrence Livermore National Laboratory stations ELK, KNB, LAC, and MNV, known collectively as the Livermore National Network (LNN) (Figure 1). LNN was established in the early 1960s for the primary purpose of monitoring underground nuclear testing at the former Nevada Test Site (NTS), now known as the Nevada Nuclear Security Site (NNSS) following the signing of the Limited Test Ban Treaty (LTBT). LNN consisted initially of short-period vertical component Benioff’s recorded on film located at Mina, NV (MNV) and Kanab, Utah (KNB). LNN added two additional stations at Landers, CA (LAC) and Elko, NV (ELK) in 1967 and upgraded equipment to broadband seismometers recorded on frequency modulation (FM) tapes from 1967-1979, followed by digital recordings after 1979 (Jarpe, 1989). The digital recordings were on a variety of now obsolete media, including 9-track, Exabyte, and DAT tapes. Jarpe (1989) describes the seismic station instrumentation details over the period of deployment. LNN recorded valuable non-repeatable unique data of several hundreds of nuclear explosions at NNSS, as well as earthquakes and chemical and mining explosions (Walter, 2020). The details of these nuclear tests are provided in the Department of Energy Report NV-209 Rev 16 (DOE, 2015).

58 GEOSCIENCES↗

Reproducibility of Radiokrypton in Deep Desert Aquifers: Insights from a Decade of Research

Great technical advances have been achieved since the first atom-trap trace analysis (ATTA) -based radiokrypton application in Egypt, where 1 Myr old groundwater was discovered. Beyond advances in ATTA measurement capabilities, including reduction in sample size, analysis duration, and analytical uncertainty, major progress has been achieved over the past two decades in the sample collection and preparation techniques. These advances paved the expansion of ATTA-based noble gas applications to many other aquifers worldwide, illuminating the nature and flow pattern of deep groundwater systems. While the potential of this new analytical technique for old groundwater dating is well recognized, another important aspect yet to be examined is the reproducibility of radiokrypton in aquifers over time, i.e., how representative is a discrete groundwater sample, collected at a specific time and location, for the natural groundwater system? The likelihood of a negative answer is increased by flow-field disturbance in aquifers following massive groundwater abstraction. Here, in this work, we present repeated 81 Kr sampling and measurements in twenty-one sites over Israel, mostly of deep (up to 1 km) wells tapping confined aquifers in the arid to hyperarid Negev desert. The results demonstrate that radiokrypton measurements are indeed reproducible, even in cases where samples were collected as long as nine years apart and from highly productive (∼1 Mm 3 /yr order) pumping wells. Furthermore, many of the repeated measurements in this study (17 out of the 21 sites) were conducted with different ATTA Instruments in two different laboratories using slightly different sampling, preparation, and analysis techniques, yet with an overall good agreement. The consistency in the ATTA-based 81 Kr-dating results over time highlights the robustness of this state-of-the-art technique as a tool to unravel groundwater flow patterns and encourages further applications to many other yet-to-be-explored deep aquifers.

atom-trap trace analysis↗