Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Application of Principal Component Analysis to Electrochemical Reprocessing PM and NMAC

In this report, data from an electrorefiner (ER) for nuclear fuel reprocessing is evaluated for process monitoring (PM) conclusions. This data comes from tests performed at the Idaho National Laboratory in 2022. Multivariate approaches utilizing methods of Principal Component Analysis (PCA) is applied. This is based off established work in process monitoring for fault detection in industrial facilities. This report will discuss the background, methods, and results of the application and some of the conclusions and applications that can be drawn from them. PCA is applied to two different electrorefiner (ER) operations that occurred at Idaho National Laboratory between August and October 2022. The first operation occurred with little incident while the second had several noted faults in the equipment in operational logs. The data from the first run was used to train the data for “normal” operations and applied to both sets of data to determine when operations were in an “off-normal” condition and identify where the fault occurs through PCA. PCA was able to identify off-normal events and identify the cause for off-normal operations. These identified off-normal events matched with the events and their causes in the operational logs. However, small amounts of variance in the data led to false detection of “off-normal” events. Thus, careful selection of training data and a-posteriori conclusions based off operator assessments will both be required for application of PCA to PM applications. This work demonstrated that multivariate approaches and latent variables are applicable to pyroprocessing PM applications and can be further expanded in future work as quality variables such as salt concentration from sensors and sampling become available.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Factors associated with treatment limitations in two Swedish intensive care units: Prevalence and patient involvement

Abstract The aim was to study the prevalence, documentation, and patient involvement in treatment limitations (TLs) in two Swedish intensive care units (ICUs). All patients admitted to the ICUs of two Swedish regional hospitals in 2019 were screened for inclusion. Exclusion criteria included postanesthesia care <24 h. Patients were identified using the Swedish Intensive Care Registry (SIR) and data were extracted from SIR and hospital charts. Uni‐ and multivariable logistic analysis was performed to investigate associations with the presence of TLs. A total of 3090 patients were admitted to the two ICUs in 2019. After exclusion, 1019 patients were included in the study. 45.5% were women and the mean age was 62.9 years. 26.5% of the patients had one or several TLs. Age (OR 1.04 per one year increase 95% confidence interval (CI) 1.02–1.05), SAPS3‐score (OR 1.08 per one unit increase 95% CI 1.06–1.09) and ICU length of stay (OR 1.11 per one day increase 95% CI 1.05–1.17) were independently associated with an increased likelihood of receiving a TL. 17% of the patients were involved in the decision‐making process and in >30% of cases neither the patient nor next‐of‐kin were informed. Women were to a larger extent involved in the decision process than men (24.5 vs. 12.5% p < .05). When the intensivist documented why a TL was established, patient autonomy was four times more commonly stated as the motivation for the TL among women compared to men (15.5% vs. 3.8% p < .05). TLs were common in two Swedish ICUs but a substantial number of patients and next‐of‐kin were not involved in the decision‐making process or informed of the decision. Women were more often than men engaged in the decision to establish a TL.

Jönsson, Nino↗

Identification and correction of temporal and spatial distortions in scanning transmission electron microscopy

Scanning transmission electron microscopy (STEM) has become the technique of choice for quantitative characterization of atomic structure of materials, where the minute displacements of atomic columns from high-symmetry positions can be used to map strain, polarization, octahedra tilts, and other physical and chemical order parameter fields. The latter can be used as inputs into mesoscopic and atomistic models, providing insight into the correlative relationships and generative physics of materials on the atomic level. However, these quantitative applications of STEM necessitate understanding the microscope induced image distortions and developing the pathways to compensate them both as part of a rapid calibration procedure for in situ imaging, and the post-experimental data analysis stage. Here, we explore the spatiotemporal structure of the microscopic distortions in STEM using multivariate analysis of the atomic trajectories in the image stacks. Based on the behavior of principal component analysis (PCA), we develop the Gaussian process (GP)-based regression method for quantification of the distortion function. The limitations of such an approach and possible strategies for implementation as a part of in-line data acquisition in STEM are discussed. Here, the analysis workflow is summarized in a Jupyter notebook that can be used to retrace the analysis and analyze the reader's data.

36 MATERIALS SCIENCE↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Accelerated Probabilistic Marching Cubes by Deep Learning for Time-Varying Scalar Ensembles

Visualizing the uncertainty of ensemble simulations is challenging due to the large size and multivariate and temporal features of en-semble data sets. One popular approach to studying the uncertainty of ensembles is analyzing the positional uncertainty of the level sets. Probabilistic marching cubes is a technique that performs Monte Carlo sampling of multivariate Gaussian noise distributions for positional uncertainty visualization of level sets. However, the technique suffers from high computational time, making interactive visualization and analysis impossible to achieve. This paper introduces a deep-learning-based approach to learning the level-set uncertainty for two-dimensional ensemble data with a multivariate Gaussian noise assumption. We train the model using the first few time steps from time-varying ensemble data in our workflow. We demonstrate that our trained model accurately infers uncertainty in level sets for new time steps and is up to 170X faster than that of the original probabilistic model with serial computation and 10X faster than that of the original parallel computation.

Han, Mengjiao↗

Resolving the Evolution of Atomic Layer-Deposited Thin-Film Growth by Continuous In Situ X-Ray Absorption Spectroscopy

In situ synchrotron X-ray absorption near-edge structure characterization of thin-film titania growth by atomic layer deposition (ALD) over ZnO nanowires reveals persistent low-coordinated Ti motifs leading to a new picture of ALD growth. Through the design of growth and measurement cycles, Ti K-edge spectral data are continuously recorded so as to characterize the film evolution as a function of ALD cycle number and the surface changes within the time scale of the ALD cycle. A unified set of analysis tools is developed to interpret the time-series of spectral data. A prenucleation stage of growth, a transition region, and then a steady-state growth stage are observed with distinguishable features. Multivariate curve resolution analysis, that is physically constrained, demonstrates two specific spectral components with associated, time-dependent concentrations. The bulk-film component tracks the stages of growth. The surface and interface components, present throughout the stages of growth, reveal a significant coverage of relatively isolated or loosely networked tetrahedrally coordinated Ti atomic motifs. Lastly, spectral signatures for the intra-cycle growth kinetics are reconstructed at a time resolution of ~1 s and demonstrate that the transient Ti motifs on the growing surface stabilize within a few seconds of the Ti precursor pulse.

36 MATERIALS SCIENCE↗

New methods for trace analysis of gamma-irradiated pentaerythritol tetranitrate

High explosives (HEs) are used in a diverse range of applications in which they could be exposed to various radiation levels that may cause potential chemical changes. This study further evaluated pentaerythritol tetranitrate (PETN) that was previously aged with a low-level 2 kGy dose of gamma irradiation in order to understand chemical changes caused by irradiation. Both unirradiated PETN and gamma-irradiated PETN were analyzed using ultra high-pressure liquid chromatography coupled to quadrupole time of flight mass spectrometry (UHPLC-QTOF). The resulting data were processed in a non-targeted manner using Fisher's ratio analysis and multivariate curve resolution-alternating least squares (MCR-ALS) to aid in discovery and identification of the chemical changes brought about by irradiation without a priori knowledge. The application of using UHPLC-QTOF in combination with chemometric techniques for the analysis of irradiated samples has not previously been performed. Here, in this work, we show how to use this method to provide chemical information that would otherwise not be discernible, such as the discovery of degradation of the various homologues of PETN. Major differences identified with radiolytic aging of the PETN sample included decomposition products that resulted from the degradation of the trigger linkage – the O–NO 2 bonds – resulting in the formation of alcohol and aldehyde groups. Similar degradation was also observed in the PETN homologues as well as interconversion from one homologue to another.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Robust framework and software implementation for fast speciation mapping

One of the greatest benefits of synchrotron radiation is the ability to perform chemical speciation analysis through X-ray absorption spectroscopies (XAS). XAS imaging of large sample areas can be performed with either full-field or raster-scanning modalities. A common practice to reduce acquisition time while decreasing dose and/or increasing spatial resolution is to compare X-ray fluorescence images collected at a few diagnostic energies. In this work, several authors have used different multivariate data processing strategies to establish speciation maps. Furthermore, the theoretical aspects and assumptions that are often made in the analysis of these datasets are focused on. A robust framework is developed to perform speciation mapping in large bulk samples at high spatial resolution by comparison with known references. Two fully operational software implementations are provided: a user-friendly implementation within the MicroAnalysis Toolkit software, and a dedicated script developed under the R environment. The procedure is exemplified through the study of a cross section of a typical fossil specimen. Additionally, the algorithm provides accurate speciation and concentration mapping while decreasing the data collection time by typically two or three orders of magnitude compared with the collection of whole spectra at each pixel. Whereas acquisition of spectral datacubes on large areas leads to very high irradiation times and doses, which can considerably lengthen experiments and generate significant alteration of radiation-sensitive materials, this sparse excitation energy procedure brings the total irradiation dose greatly below radiation damage thresholds identified in previous studies. This approach is particularly adapted to the chemical study of heterogeneous radiation-sensitive samples encountered in environmental, material, and life sciences.

47 OTHER INSTRUMENTATION↗

Chemistry imaging and distribution analysis of rare earth elements in coal using LIBS and LA-ICP-MS instruments

Currently, demand for rare earth elements (REEs) increased significantly. Coal is actively evaluated as potential economic sources for extraction of REEs. Here, in this work, laser-induced breakdown spectroscopy (LIBS) was evaluated for rapid estimation of REEs content and their distribution in the natural coal samples. The results were compared with similar laser ablation–inductively coupled plasma–mass spectrometry (LA-ICP-MS) measurements. Thirteen coal samples (nine standard samples and five natural samples) were used in this study. Powder samples were pressed into pellets while coal chunks were directly ablated for data recording. Pellets of the powder standard samples were used to optimize the data acquisition system and then data recorded with this optimized system was used to identify the proper data acquisition and analysis models. After establishing the proper data acquisition system and analysis model using the standard samples, natural coal samples in powder form and their chunks were utilized to record LIBS and LA-ICP-MS spectra. Multivariate calibration models were developed using four of the natural samples, which were evaluated by predicting the REE content in the fifth sample. Principal component analysis was performed on the LIBS data obtained from the natural samples and it classified all the samples with high accuracy. Two-dimensional (2D) elemental mapping on coal chunk samples was also performed using both LIBS and LA-ICP-MS to study the distribution of REEs in the samples. The resulting elemental images and their correlations can be used to infer mineral distributions.

01 COAL, LIGNITE, AND PEAT↗

Stochastic Simulation of Daily Suspended Sediment Concentration Using Multivariate Copulas

Estimation of daily suspended sediment concentration (SSC) is required for water resources and environment management. In this paper, a copula-based stochastic method was proposed for daily SSC simulation. Here, the multivariate copula function, constructed based on a bivariate copula and two bivariate conditional probability distributions, was used to model the temporal and cross dependence structures in daily SSCs. Then, the daily SSCs were generated by sampling from the multivariate conditional distribution. As a result, synthetic long-term SSCs data beyond the limited observation period can be provided for water resources managers, which plays a critical role in accurately estimating frequency and magnitude of extreme SSCs events. The proposed method was under rigorous examination by applying to a case study at Pingshan station in the Jinsha River Basin, China. Results showed that the generated daily SSC sequences not only had a high degree of accuracy in preserving the statistical characteristics of the daily SSC observations, but also captured both the temporal correlation and the cross-correlation between the daily streamflow and daily SSC. Specifically, the average daily relative error values corresponding to mean, standard deviation, skewness, lag-1 temporal correlation, and cross correlation were 0.87%, 4.24%, 7.52%, 0.51% and 2.02%, respectively. The multivariate copula framework proposed here can accurately and efficiently generate long-term daily SSC data for water resources management such as frequency analysis and risk assessment of extreme SSC events.

54 ENVIRONMENTAL SCIENCES↗

MetaboDirect: an analytical pipeline for the processing of FT-ICR MS-based metabolomic data

Background: Microbiomes are now recognized as the main drivers of ecosystem function ranging from the oceans and soils to humans and bioreactors. However, a grand challenge in microbiome science is to characterize and quantify the chemical currencies of organic matter (i.e., metabolites) that microbes respond to and alter. Critical to this has been the development of Fourier transform ion cyclotron resonance mass spectrometry (FT-ICR MS), which has drastically increased molecular characterization of complex organic matter samples, but challenges users with hundreds of millions of data points where readily available, user-friendly, and customizable software tools are lacking. Results: Here, we build on years of analytical experience with diverse sample types to develop MetaboDirect, an open-source, command-line-based pipeline for the analysis (e.g., chemodiversity analysis, multivariate statistics), visualization (e.g., Van Krevelen diagrams, elemental and molecular class composition plots), and presentation of direct injection high-resolution FT-ICR MS data sets after molecular formula assignment has been performed. When compared to other available FT-ICR MS software, MetaboDirect is superior in that it requires a single line of code to launch a fully automated framework for the generation and visualization of a wide range of plots, with minimal coding experience required. Among the tools evaluated, MetaboDirect is also uniquely able to automatically generate biochemical transformation networks (ab initio) based on mass differences (mass difference network-based approach) that provide an experimental assessment of metabolite connections within a given sample or a complex metabolic system, thereby providing important information about the nature of the samples and the set of microbial reactions or pathways that gave rise to them. Finally, for more experienced users, MetaboDirect allows users to customize plots, outputs, and analyses. Conclusion: Application of MetaboDirect to FT-ICR MS-based metabolomic data sets from a marine phage-bacterial infection experiment and a Sphagnum leachate microbiome incubation experiment showcase the exploration capabilities of the pipeline that will enable the research community to evaluate and interpret their data in greater depth and in less time. It will further advance our knowledge of how microbial communities influence and are influenced by the chemical makeup of the surrounding system. The source code and User’s guide of MetaboDirect are freely available through (https://github.com/Coayala/MetaboDirect) and (https://metabodirect.readthedocs.io/en/latest/), respectively.

54 ENVIRONMENTAL SCIENCES↗

Visual HPC Workflows for the Analysis of System Dynamics Models

Visual analytics supported by high performance computing (HPC) accelerates and enhances the discovery, exploration, and analysis of causal patterns in complex system dynamics (SD) models. We present a suite of visualization-assisted ensemble-based techniques for hypothesis generation and testing, and for sensitivity analysis. By employing HPC to provide parallel, on-demand simulation of SD models, one can “steer” an ensemble of simulated scenarios in real time as one first formulates and then informally tests those hypotheses: this provides rapid feedback for analysts to refine their understanding of the causal relationships emergent from a model. Such understandings can be followed and augmented by rigorous application of statistical methods, namely global variance-based sensitivity analysis, Monte-Carlo filtering, adaptive regional sensitivity analysis, and self-organized maps: here timely computation relies on HPC, while effective presentation emphasizes high-dimensional multivariate data visualization. Immersive visualization in virtual 3D environments provides an excellent adjunct to the traditional 2D graphics typically used for SD models, as it generates an embodied understanding of model behavior and facilitates an active, collaborative critique of model structure and output. Finally, we summarize prospects for HPC-enabled visual analytics applied to SD modeling.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Monitoring Radiochemical Processing Streams for the 238 Pu Supply Program with Process Pulse II

Oak Ridge National Laboratory (ORNL) is developing advanced spectroscopic and real-time monitoring capabilities to improve the timeliness of analytical measurements and process decisions for the 238Pu Supply Program. Reducing the time, resources, and costs associated with each production campaign is critical because overlapping campaigns will be required to meet the production goals of the National Aeronautics and Space Administration. Real-time, in situ analytical measurements in the heavily shielded hot cells at the Radiochemical Engineering Development Center (REDC) will allow for rapid process information feedback and operational benefits that help the 238Pu supply program scale-up production efforts. Noteworthy steps were taken during Campaign 5 to establish the ability to monitor processing streams in real time with spectrophotometry and a commercially available online monitoring software called The Unscrambler X Process Pulse II (PP) multivariate statistical process monitoring system by Camo Analytics (version 5.60). PP automates univariate-type calculations within the software itself and executes multivariate models built using The Unscrambler X (version 10.4 or newer). The Unscrambler is a commercially available data analysis software made by the same company. PP is composed of easy-to-use-tools for all personnel, including data scientists and technicians. The software can be used to plot analyte concentration profiles, spectral data, and other process variables in real time. All process data are represented in a single view with interactive charts useful for viewing how a process evolves over time.

07 ISOTOPE AND RADIATION SOURCES↗

Monitoring the Caustic Dissolution of Aluminum Alloy in a Radiochemical Hot Cell Using Raman Spectroscopy

Chemical processing of highly radioactive materials commonly takes place in heavily shielded hot cells. The remote, real-time monitoring of chemical processing streams via optical spectroscopic techniques in hot cells may be particularly useful. Here in this paper, we describe the implementation of Raman spectroscopy and chemometric analysis to monitor the dissolution of aluminum-clad targets containing irradiated aluminum–neptunium oxide cermet pellets in caustic solutions in a hot cell environment. Partial least squares regression analysis was used to generate calibration models to quantify the concentration of dissolved aluminum, nitrate, and hydroxide in solutions within the radiochemical hot cell. This work explored a systematic approach to optimize a matrix of calibration standards using a D-optimal experimental design. The Design of Experiments-based regression model, in comparison to more traditional analytical approaches, was found to be the more practical method for building calibration models, with fewer samples, to obtain informative analytical data from Raman spectra.

36 MATERIALS SCIENCE↗

PyKrev: A Python Library for the Analysis of Complex Mixture FT-MS Data

In this study, we present PyKrev, a Python library for the analysis of complex mixture Fourier transform mass spectrometry (FT-MS) data. PyKrev is a comprehensive suite of tools for analysis and visualization of FT-MS data after formula assignment has been performed. These comprise formula manipulation and calculation of chemical properties, intersection analysis between multiple lists of formulas, calculation of chemical diversity, assignment of compound classes to formulas, multivariate analysis, and a variety of visualization tools producing van Krevelen diagrams, class histograms, PCA score, and loading plots, biplots, scree plots, and UpSet plots. The library is showcased through analysis of hot water green tea extracts and Scotch whisky FT-ion cyclotron resonance-MS data sets. PyKrev addresses the lack of a single, cohesive toolset for researchers to perform FT-MS analysis in the Python programming environment encompassing the most recent data analysis techniques used in the field.

47 OTHER INSTRUMENTATION↗

Accelerating Multivariate Functional Approximation Computation with Domain Decomposition Techniques⋆

Modeling large datasets through Multivariate Functional Approximations (MFA) provide an elegant way to handle many visualization and scientific analysis workflows. The process necessitates scalable data partitioning methods to compute MFA representations efficiently without compromising the accuracy or continuity of the reconstructed solution. We propose a domain -decomposed method for computing the MFA with B -spline bases, which reduces the total work per task and uses a restricted Additive Schwarz (RAS) method to converge the control point data degrees -of -freedom along subdomain boundaries. We provide an in-depth analysis of the parallel approach with domain decomposition solvers, aiming to minimize local subdomain error residuals and recover high -order continuity at subdomain interfaces with appropriate choices of knot overlaps. The communication cost, determined by the overlap regions in the RAS implementation, is optimized to recover the numerical error profile of the single subdomain case. Our proposed method stands in contrast to previous methods, which typically only recover either C 0 or at best C 1 continuity for arbitrary B -spline degree expansions, or those that require post -processing to blend discontinuities in the reconstructed data. We demonstrate the effectiveness of our approach using analytical and real -world datasets in 1D, 2D, and 3D through both strong and weak scaling studies. The performance results indicate that the overall cost of computing the approximation is directly proportional to the underlying nearest -neighbor communication implementation, and is only weakly dependent on the overlap region size that determines the size of the messages. This finding underscores the efficiency and scalability of our proposed method, making it a promising solution for handling large datasets in scientific workflows.

additive Schwarz solvers↗