Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “process analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Simulation process and data flow for a large system dynamics model

This paper documents the workflow and supporting technologies that a large system dynamics model, the biomass scenario model, employs to streamline the data preparation, simulation, quality control, and analysis process at the National Renewable Energy Laboratory. The workflow centers on automation of routine aspects of the flow of data between data stores, simulations, and visualizations. It enforces quality checks on data, reproducibility of computations, and traceability of results, while maintaining complete archives of modeling and analysis artifacts. The resulting frictionless simulation/analysis environment supports large-scale sensitivity analysis, interactive creation of ensembles of simulations, and rapid visualization-based exploration of simulation results.

09 BIOMASS FUELS↗

A graph signal processing‐based multiple model Kalman filter ( GSP‐MMKF ) tool for predictive analytics: An air separation unit process application

Abstract The industrial Air Separations Unit (ASU) is a complicated and tightly operated process. The use of dynamic process analytics is also a key element of safe and economic operation of these processes, with increasing focus on predictive analytics to take preemptive actions. With the availability of real‐time data from hundreds of sensors, the data analysis process should also consider the topology of the data, as seen in sensor networks. In this paper, a novel tool is presented that considers the complex connectivity patterns in the sensor network and uses local adaptive disturbance estimations to predict global network‐scale trends. The paper introduces the emerging field of Graph Signal Processing (GSP) and presents a rigorous derivation of the tool starting from the extraction of the sensor‐network (in a graph theoretical sense) from the data. This network, which is in the form of a matrix, is then used to derive a Kalman‐filter type of state‐space model driven by input disturbances. Multiple disturbance models (e.g., step, ramp, periodic) are included to allow the model to have different kinds of disturbance propagation. Each graph node (representing the sensors used) dynamically adapts to the most recent detected disturbance individually. These estimated disturbances are propagated to the global network using the graph. Modifications to ensure stability are also discussed. The fidelity of the tool is tested on certain downtime events and the paper concludes by discussing the advantages of the method and planned future improvements.

Ghosh, Sambit↗

Used Nuclear Fuel Management Using the Next Generation System Analysis Model

The U.S. Department of Energy (DOE) is leading the National effort to manage the back end of the nuclear fuel cycle, encompassing the safe transportation, storage/staging, and/or eventual disposal of used nuclear fuel (UNF) and high-level radioactive waste. The Next Generation System Analysis Model (NGSAM) is DOE’s discrete-event, agent-based simulation tool designed to model the full life cycle of UNF from reactor discharge to final disposal. NGSAM supports the DOE Office of Spent Fuel and High-Level Waste Disposition by enabling a detailed, scenario-based analysis of logistics, infrastructure, and shipping strategies. NGSAM replaces legacy models with a modern, flexible platform built on Repast Simphony and enhanced by the Process Analysis Tool. NGSAM simulates the movement and interaction of individual fuel assemblies with system components such as canisters, casks, railcars, and facilities. The model integrates with the Java Transportation Operations Model to plan and execute transportation scenarios, supporting both constrained and unconstrained resource allocation. Key features include customizable allocation and acceptance algorithms, detailed facility-level operations, and a Quick Edit tool for rapid scenario adjustments. NGSAM supports multimodal transportation modeling (e.g. rail, road, barge) and provides comprehensive cost, schedule, and infrastructure data. NGSAM utilizes data from sources such as DOE’s STANDARDS UNF database and DOE’s Stakeholder Tool for Assessing Radioactive Transportation, while also allowing user-defined inputs for scenario customization. NGSAM enables stakeholders to evaluate complex UNF management strategies, assess system performance under varying assumptions, and inform decision making for future infrastructure investments. Its modular architecture and integration with other Integrated Waste Management System tools make it a critical asset for planning the safe and efficient disposition of the Nation’s growing UNF inventory.

Craig, Brian [Argonne National Laboratory (ANL)]↗

eagles-project/asediag

Aerosol process analysis in model-native Spectral Element (SE) grid. Aerosol Diagnostics on Model Native Grid is a Python-based tool designed for diagnosing aerosol processes in E3SM. It is particularly designed to analyze simulation data on the model-native spectral element (SE) grid, soit is also known as Aerosol SE Diagnostics or “asediag”

Hassan, Taufiq↗

srlife : a software tool for estimating the life of high temperature concentrating solar receivers. Part I – metallic receivers

Here, this paper introduces srlife, a tool for estimating the structural service life of concentrating solar power (CSP) receivers operating at high temperatures. Supporting both metallic and ceramic receiver designs, srlife is available as open-source software at https://github.com/applied-material-modeling/srlife and can be installed via the PyPi package manager (https://pypi.org). Given basic receiver geometry and incident heat flux, the tool performs thermohydraulic and structural analysis and estimates the life of a receiver. Designed for easy integration into a software stack, including solar field and levelized cost analysis, the tool can be utilized for optimizing receiver designs to meet service life and economic targets. This paper is Part I in a two-part series. Part I discusses the analysis process used to estimate the life of metallic receivers, along with a description of the required input data. Additionally, several heuristics applied within srlife can reduce analysis time significantly while maintaining accurate life estimations for metallic receivers when compared to full analyses. Several examples demonstrating the utility of srlife in receiver design are also discussed. Part II focuses on the life estimation of ceramic receivers, using time-dependent reliability analysis and various ceramic failure models implemented in srlife.

Creep-fatigue analysis↗

EMPHATIC Silicon Strip Detector Efficiencies

EMPHATIC is an experiment at Fermilab which aims to reduce current neutrino flux uncertainties. This report discusses the limitations current neutrino flux uncertainties places on large scale neutrino experiments, provides background on the EMPHATIC experiment, and details the project of determining the efficiency of the Silicon Strip Detectors (SSDs) used in EMPHATIC. As part of the data analysis process and in order to increase the accuracy of EMPHATIC’s simulations a representation of efficiency of each SSD is required. To achieve this a data-driven analysis was performed on EMPHATIC's collected data using the Root and Art frameworks. Visual and numerical representations of efficiency were determined. The average efficiency over all SSDs is 98.58\%, however this number deflated as it includes known bad channels.

Olson, Virginia [Illinois U., Urbana (main)]↗

Determining the Efficiency of EMPHATICs Silicon Strip Detectors (SSDs)

EMPHATIC is an experiment at Fermilab which aims to reduce current neutrino flux uncertainties. This report discusses the limitations current neutrino flux uncertainties places on large scale neutrino experiments, provides background on the EMPHATIC experiment, and details the project of determining the efficiency of the Silicon Strip Detectors (SSDs) used in EMPHATIC. As part of the data analysis process and in order to increase the accuracy of EMPHATIC’s simulations a representation of efficiency of each SSD is required. To achieve this a data-driven analysis was performed on EMPHATIC's collected data using the Root and Art frameworks. Visual and numerical representations of efficiency were determined. The average efficiency over all SSDs is 98.58\%, however this number deflated as it includes known bad channels.

Olson, V. [Illinois U., Urbana (main)]↗

What Is Data Analysis?

A quick guide to understand your data and use it to tell compelling stories. Data analysis helps develop insights for research projects, planning interventions, or systematic information gathering. This guide highlights important aspects of the data analysis process.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

AGR-2 TRISO Layer Thickness Imaging Archive

As a part of fuel quality control characterization, optical microscopy images of particle cross sections near midplane were acquired at Oak Ridge National Laboratory (ORNL). These particles were produced by the Advanced Gas Reactor Fuel Development and Qualification (AGR) Program’s AGR-2 irradiation campaign. These images may be of use for the development of image processing algorithms with the benchmark values measured at ORNL. This report provides those benchmark values, along with the raw images and data generated by the ORNL particle layer thickness analysis process.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Parallel I/O Evaluation Techniques and Emerging HPC Workloads: A Perspective

Emerging workloads such as artificial intelligence, big data analytics and complex multi-step workflows alongside future exascale applications are anticipated future HPC workloads, which will result in a more diverse I/O system workload and even less predictable I/O behavior and access patterns. Along with the ever increasing gap between the compute and storage performance capabilities, the in-depth understanding of extreme-scale I/O behavior and the I/O performance modeling and prediction are essential tools of the large-scale I/O evaluation process for addressing the needs of extreme-scale hybrid workloads. In this survey article, we focus on the state-of-the-art of the I/O behavior and performance analysis process for HPC systems in a 5-year time window and identify future research challenges.

Neuwirth, Sarah↗

Visualizing metagenomic and metatranscriptomic data: A comprehensive review

The fields of Metagenomics and Metatranscriptomics involve the examination of complete nucleotide sequences, gene identification, and analysis of potential biological functions within diverse organisms or environmental samples. Despite the vast opportunities for discovery in metagenomics, the sheer volume and complexity of sequence data often present challenges in processing analysis and visualization. This article highlights the critical role of advanced visualization tools in enabling effective exploration, querying, and analysis of these complex datasets. Emphasizing the importance of accessibility, the article categorizes various visualizers based on their intended applications and highlights their utility in empowering bioinformaticians and non-bioinformaticians to interpret and derive insights from meta-omics data effectively.

59 BASIC BIOLOGICAL SCIENCES↗

Human Liver Epithelium Response to HCoV-229E Infection Epigenomics (ACS-DP4)

The purpose of this experiment was to evaluate how wild-type Human coronavirus strain 229E (HCoV-229E) infection alters chromatin accessibility in infected cells only. Sample data was obtained for mock and infected (standard and UV-inactivated) immortalized human liver cells (HuH-7) and collected 24 hrs. post infection. Samples were processed using assay for transposase-accessible chromatin using high-throughput sequencing (ATAC-Seq) and generated bar coded library samples were evaluated for RNA sequencing (RNA-Seq) expression analysis. Processed ATAC-Seq datasets are openly accessible from the download button and contain secondary processed RNA-Seq results files and supporting metadata materials. Data download includes a sample naming key, infection titer metadata, normalized counts, and relevant computational source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

Chemical characterization of microplastic particles formed in airborne waste discharged from sewer pipe repairs

Microplastic particles are of increasing environmental concern due to the widespread uncontrolled degradation of various commercial products made of plastic and their associated waste disposal. Recently, common technology used to repair sewer pipes was reported as one of the emission sources of airborne microplastics in urban areas. This research presents results of the multi-modal comprehensive chemical characterization of the microplastic particles related to waste discharged in the pipe repair process and compares particle composition with the components of uncured resin and cured plastic composite used in the process. Analysis of these materials employs complementary use of surface-enhanced Raman spectroscopy, scanning transmission X-ray spectro-microscopy, single particle mass spectrometry, and direct analysis in real-time high-resolution mass spectrometry. Further, it is shown that the composition of the relatively large (100 μm) microplastic particles resembles components of plastic material used in the process. In contrast, the composition of the smaller (micrometer and sub-micrometer) particles is significantly different, suggesting their formation from unintended polymerization of water-soluble components occurring in drying droplets of the air-discharged waste. In addition, resin material type influences the composition of released microplastic particles. Results are further discussed to guide the detection and advanced characterization of airborne microplastics in future field and laboratory studies pertaining to sewer pipe repair technology.

54 ENVIRONMENTAL SCIENCES↗

HopPyBar

HopPyBar is a python program to import, analyze, and export split-Hopkinson pressure bar (SHPB, also known as Kolsky bar) data. Traditional analysis offers a black box approach, where input data is converted to analyzed output by performing a series of calculations without user involvement. This program serves as a developmental platform to "white box" the data analysis process. Data streams can be captured (in-situ) to enable advanced or unconventional analyses, statistics, and comparisons. Additionally, the program is geared towards the standardized forms of input and output used at LANL to streamline analysis, but the open nature of the program makes additional input/output schemes straightforward to add. General workflow will import SHPB data in one of a number of formats, identify relevant portions of data signals, and convert to stress-strain-strain rate to show material behavior as a function of dynamic testing.

Morrow, Benjamin↗

Methods for Incorporating Model Uncertainty into Exoplanet Atmospheric Analysis

A key goal of exoplanet spectroscopy is to measure atmospheric properties, such as abundances of chemical species, in order to connect them to our understanding of atmospheric physics and planet formation. In this new era of high-quality JWST data, it is paramount that these measurement methods are robust. When comparing atmospheric models to observations, multiple candidate models may produce reasonable fits to the data. Typically, conclusions are reached by selecting the best-performing model according to some metric. This ignores model uncertainty in favor of specific model assumptions, potentially leading to measured atmospheric properties that are overconfident and/or incorrect. In this paper, we compare three ensemble methods for addressing model uncertainty by combining posterior distributions from multiple analyses: Bayesian model averaging, a variant of Bayesian model averaging using leave-one-out predictive densities, and stacking of predictive distributions. We demonstrate these methods by fitting the Hubble Space Telescope (HST) + Spitzer transmission spectrum of the hot Jupiter HD 209458b using models with different cloud and haze prescriptions. All of our ensemble methods lead to uncertainties on retrieved parameters that are larger but more realistic and consistent with physical and chemical expectations. Since they have not typically accounted for model uncertainty, uncertainties of retrieved parameters from HST spectra have likely been underreported. We recommend stacking as the most robust model combination method. Our methods can be used to combine results from independent retrieval codes and from different models within one code. They are also widely applicable to other exoplanet analysis processes, such as combining results from different data reductions.

79 ASTRONOMY AND ASTROPHYSICS↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Solar PV

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) research platform. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence data centers and other variable loads. This dataset entry describes the behavior of a 1.25-MW proton exchange membrane MC250 electrolyzer system, manufactured by Nel Hydrogen , [1] when fed historical data generated by the 430-kW, fixed-axis solar photovoltaic (PV) array located at NLR’s Flatirons Campus. (While the electrolyzer balance of plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack.) Solar PV power output data for the 2020 calendar year were categorized on a daily basis by total energy generation and standard deviation. Each day was then ranked by these metrics, and the 25th, 50th, and 100th percentiles were selected. The 75th percentile day did not exhibit sufficient variability to make for a valuable experiment. A similar process was used for the related historical wind dataset . [2] The historical days in 2020 that represented these percentiles are Dec. 19, March 29, and May 4, respectively. The entire solar day’s power profile was then fed through the MC250 electrolyzer. Due to its length, the 100th percentile day experiment was split into two parts, and the final 3 hours of the solar day were not captured. These final 3 hours contained no spikes or dips of interest and simply represented a slow decay of input solar power. Also, a single timestamp (13:13:47 on Jan. 14, 2026) was lost in the hydrogen system supervisory control and data acquisition. Finally, during the 25th percentile experiment (solar day Dec. 19, 2020) data recording was lost from 11:00:13 to 11:14:45. The roughly 15 minutes of the solar profile were rerun at the end of the experiment and spliced into this time slot during post-processing. The electrolysis system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operation of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical solar profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. For more details on the statistical analysis process, see the slide deck “Public Reference Data for Megawatt-Scale Hydrogen Electrolysis: NLR Historical Solar PV Analysis and Profile Generation” accessible with this data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single solar PV electrolysis experiment and is formatted as: {technology}_{percentile}_{scaling factor} For instance, “solarPV-430kW_25_2x.zip” reports the experiment using the 25th percentile solar data from the historical 2020 solar PV dataset, scaled to 200%. Scaling factors were applied to the generated solar PV power output files to more closely match the 1.25-MW capacity of the electrolyzer. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and solar power input. A PDF file detailing the historical solar data statistical analysis used to generate the solar profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all experiments combined into one dataset labeled "combined_solarPV_experiments.csv". [1] nelhydrogen.com/product/mc-series-electrolyser . [2] data.nlr.gov/submissions/316 .

08 HYDROGEN↗

Summary of the 5th IAEA technical meeting on fusion data processing, validation and analysis (FDPVA)

The purpose of the 5th International Atomic Energy Agency technical meeting on fusion data processing, validation and analysis (FDPVA) (Ghent University, Ghent, Belgium, 12–15 June 2023) was to provide a platform during which a set of topics relevant to FDPVA were discussed with the view of meeting the needs of next step fusion devices such as ITER. The validation and analysis of experimental data obtained from diagnostics used to characterize fusion plasmas are crucial for a knowledge-based understanding of the physical processes governing the dynamics of these plasmas. This paper presents the recent progress and achievements in the domain of plasma diagnostics data analysis and synthetic diagnostics reported at the meeting, including concept description of new devices; fusion databases; integrated data analysis; inverse problems; uncertainty propagation, verification and validation; probabilistic methods and machine learning. The relevant results underline trends observed in the current major fusion confinement devices.

fusion databases↗