Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Predicting the Seawater Chemistry of an Ocean World Using Machine Learning on Isotopic Measurements of Volatile CO2

Introduction: Given the long time intervals required for data transmission to and from ocean worlds targets, low bandwidth for data transmission, time required for data processing and analysis, and potentially extreme radiation environments (e.g., Europa), it is clear that ocean worlds missions will need more autonomous flight instruments and software in order to achieve established science goals. Protracted time intervals for data analysis (e.g., Europa Lander) strongly motivates the development of rapid, consistent and streamlined methods for interpreting data from flight mass spectrometers to e.g., determine how mass spectra from a plume or surface liquid/ice relates to the surface/subsurface. Since mass spectrometry also has the potential to correctly identify biosignatures[1], it is imperative that such methods for interpreting data are consistent and accurate. We used 848 isotope ratio mass spectra from laboratory analyses of CO2 that interacted with ocean worlds-relevant seawaters as a ‘training’ dataset for ‘unsupervised’ machine learning. In unsupervised learning, characteristics of the data are not labeled or linked, and any similarities found only result from the neural network. CO2 isotopologues analyzed for this dataset mimic the remote measurements of CO2 by a flight mass spectrometer, and are detailed in Theiling [2]. From this dataset, we used measured features of the spectra, such as retention time, intensity, and (isotopologue) mass ratios as inputs for our autoencoder neural network. Our neural network was trained to find similarities in these and other spectral features for seawaters of a particular composition and amount of initial CO2. Successful training then created an output of these similarities for various seawaters, which included MgSO4, Na2SO4, NaCl, MgCl2, KCl, and NaHCO3, and combinations of these salts. We then applied dimensionality reduction techniques such as Principal Component Analysis (PCA), T-Distributed Stochastic Neighbor Embedding (TSNE), and Uniform Manifold Approximation and Projection (UMAP) to demonstrate latent data features as a two-dimensional projection in a unitless, high-dimensional space. In this projection, a data point represents the combined effect of spectral features such as intensity, retention time, and isotope ratio. Our initial UMAP demonstrates data clustering (organization of the data by the neural network) based on the amount of CO2 that had initially interacted with each seawater. Further training using more ‘supervised’ learning techniques demonstrate strong clustering of preliminary data based on initial CO2 concentration, seawater chemical composition, and ionic strength (salinity). Our preliminary work therefore suggests that machine learning has the potential to identify compositional variants of an ocean world seawater based on mass spectra from volatile CO2 measurements. Acknowledgments: This work was funded through a Strategic Task Group at NASA Goddard Space Flight Center. The training dataset was collected through funding from the Oklahoma Space Grant Consortium. References: [1] Pappalardo, R. et al. (2013) Astrobiology, 13, 740–773. [2] Theiling (2020) Icarus, 114216.

Europa↗

Surface analysis insight note: Differentiation methods applicable to noisy data for determination of sp2‐ versus sp3‐hybridization of carbon allotropes and AES signal strengths

The derivatives of the spectra are commonly used for quantification in Auger Electron Spectroscopy (AES) spectra, while the derivative of the KLL C Auger line has proven to be valuable in obtaining a measure of the relative proportions of sp 2 ‐ and sp 3 ‐hybridization using the D‐parameter in both AES and X‐ray Photoelectron Spectroscopy (XPS). Differentiation of X‐ray Photoelectron Spectroscopy (XPS) and Auger Electron Spectroscopy (AES) spectra by numerical means is presented and illustrated for polymeric, such as PEEK and Nylon, as well as for graphitic materials including highly ordered pyrolytic graphite and graphene oxide. The most commonly available Savitzky–Golay method is explained mathematically and developed through the case of constructing a 5‐point quadratic polynomial convolution kernel suitable for differentiating spectra of adequate signal to noise. The concept of differentiation of spectra where signal to noise is less than adequate is also developed. Two alternative strategies to Savitzky–Golay differentiation are presented, which fit curves to data that allow derivatives to be obtained where Savitzky–Golay would otherwise fail. These alternative methods involve constructing a parametric curve that fits data over the entire energy interval of interest. Derivatives of spectra are then obtained by differentiating these parametric curves directly. A comparison of results for different materials for which specific sp 2 ‐ vs sp 3 ‐hybridized carbon proportions are of interest is used to emphasize the importance of characterizing methods used to differentiate spectra and understanding the characteristics of instrumentation used to measure spectra. The case for using Principal Component Analysis noise reduction with C KLL spectra is made for spectra collected from a heterogeneous graphene oxide sample.

Fairley, Neal↗

Analysis of Salinity Intrusion in the San Francisco Bay-Delta using a GA- Optimized Neural Net, and Application of the Model to Prediction in the Elkhorn Slough Habitat

The San Francisco Bay Delta is a large hydrodynamic complex that incorporates the Sacramento and San Joaquin Estuaries, the Burman Marsh, and the San Francisco Bay proper. Competition exists for the use of this extensive water system both from the fisheries industry, the agricultural industry, and from the marine and estuarine animal species within the Delta. As tidal fluctuations occur, more saline water pushes upstream allowing fish to migrate beyond the Burman Marsh for breeding and habitat occupation. However, the agriculture industry does not want extensive salinity intrusion to impact water quality for human and plant consumption. The balance is regulated by pumping stations located alone the estuaries and reservoirs whereby flushing of fresh water keeps the saline intrusion at bay. The pumping schedule is driven by data collected at various locations within the Bay Delta and by numerical models that predict the salinity intrusion as part of a larger model of the system. The Interagency Ecological Program (IEP) for the San Francisco Bay/Sacramento-San Joaquin Estuary collects, monitors, and archives the data, and the Department of Water Resources provides a numerical model simulation (DSM2) from which predictions are made that drive the pumping schedule. A problem with this procedure is that the numerical simulation takes roughly 16 hours to complete a C:~ prediction. We have created a neural net, optimized with a genetic algorithm, that takes as input the archived data from multiple stations and predicts stage, salinity, and flow at the Carquinez Straits (at the downstream end of the Burman Marsh). This model seems to be robust in its predictions and operates much faster than the current numerical DSM2 model. Because the system is strongly tidal driven, we used both Principal Component Analysis and Fast Fourier Transforms to discover dominant features within the IEP data. We then filtered out the dominant tidal forcing to discover non-primary tidal effects, and used this to enhance the neural network by mapping input-output relationships in a more efficient manner. Furthermore, the neural network implicitly incorporates both the hydrodynamic and water quality models into a single predictive system. Although our model has not yet been enhanced to demonstrate improve pumping schedules, it has the possibility to support better decision-making procedures that may then be implemented by State agencies if desired. Our intention is now to use this model in the smaller Elkhorn Slough complex near Monterey Bay where no such hydrodynamic model currently exists. At the Elkhorn Slough, we are fusing the neural net model of tidally-driven flow with in situ flow data and airborne and satellite remote sensation data. These further constrain the behavior of the model in predicting the longer-term health and future of this vital estuary.

Thompson, David E.↗

DEEPEN 3D PFA Weights for Exploration Datasets in Magmatic Environments

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This GDR submission includes those weights. The weighting was done using two different approaches: one based on expert opinions, and one based on statistical learning. The weights are intended to describe how useful a particular exploration method is for imaging each component of each play type. They may be adjusted based on the characteristics of the resource under investigation, knowledge of the quality of the dataset, or simply to reduce the impact a single dataset has on the resulting outputs. Within the DEEPEN PFA, separate sets of weights are produced for each component of each play type, since exploration methods hold different levels of importance for detecting each play component, within each play type. The weights for conventional hydrothermal systems were based on the average of the normalized weights used in the DOE-funded PFA projects that were focused on magmatic plays. This decision was made because conventional hydrothermal plays are already well-studied and understood, and therefore it is logical to use existing weights where possible. In contrast, a true PFA has never been applied to superhot EGS or supercritical plays, meaning that exploration methods have never been weighted in terms of their utility in imaging the components of these plays. To produce weights for superhot EGS and supercritical plays, two different approaches were used: one based on expert opinion and the analytical hierarchy process (AHP), and another using a statistical approach based on principal component analysis (PCA). The weights are intended to provide standardized sets of weights for each play type in all magmatic geothermal systems. Two different approaches were used to investigate whether a more data-centric approach might allow new insights into the datasets, and also to analyze how different weighting approaches impact the outcomes. The expert/AHP approach involved using an online tool (https://bpmsg.com/ahp/) with built-in forms to make pairwise comparisons which are used to rank exploration methods against one-another. The inputs are then combined in a quantitative way, ultimately producing a set of consensus-based weights. To minimize the burden on each individual participant, the forms were completed in group discussions. While the group setting means that there is potential for some opinions to outweigh others, it also provides a venue for conversation to take place, in theory leading the group to a more robust consensus then what can be achieved on an individual basis. This exercise was done with two separate groups: one consisting of U.S.-based experts, and one consisting of Iceland-based experts in magmatic geothermal systems. The two sets of weights were then averaged to produce what we will from here on refer to as the "expert opinion-based weights," or "expert weights" for short. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. More information on this approach along with the dataset used to produce the statistical weights may be found in the linked dataset below.

15 GEOTHERMAL ENERGY↗

Sea Ice Motion from Wavelet Analysis of Satellite Data

Wavelet analysis of NASA scatterometer (NSCAT) backscatter and Defense Meteorological Satellite Program (DMSP) Special Sensor Microwave/Imager (SSM/I) radiance data can be used to obtain daily sea ice drift information for the Arctic region. This technique provides improved spatial coverage over the existing array of Arctic Ocean buoys and better temporal resolution over techniques utilizing data from satellite synthetic aperture radars. Comparisons with ice motion derived from ocean buoys give good quantitative agreement. Both comparison results from NSCAT and SSM/I are compatible, and the results from NSCAT can definitely complement that from SSM/I when there are cloud or surface effects. Then three sea-ice drift daily results from NSCAT, SSM/I, and buoy data can be merged as a composite map by some data fusion techniques. The ice flow streamlines are highly correlated with surface air pressure contours. Examples of derived ice-drift maps in December 1996 illustrate large-scale circulation reversals over a period of four days. A method for deriving divergence and shear at the large-scale has been developed and comparison between buoys and satellite results shows a good agreement. These calibrated/validated results indicate that NSCAT, SSM/I merged daily ice motion are suitably accurate to identify and closely locate sea ice processes, and to improve our current knowledge of sea ice drift and related processes through the data assimilation of ocean-ice numerical model. For demonstration purpose, the ice velocities derived from satellite data are compared with the ice velocities derived from a coupled ice-ocean interaction model. The comparison reveals that the general circulation patterns of the two are quite similar but the ice velocity differences between the two are quite significant. In order to quantify the wind effects on ice motion, empirical orthogonal functions (EOF) are used in the principal component analysis for both ice motion and pressure field. Some preliminary results of sea-ice motion from QuikScat will also be presented.

Liu, Antony K.↗

Multivariate Machine Learning Models of Nanoscale Porosity from Ultrafast NMR Relaxometry

Abstract Nanoporous materials are of great interest in many applications, such as catalysis, separation, and energy storage. The performance of these materials is closely related to their pore sizes, which are inefficient to determine through the conventional measurement of gas adsorption isotherms. Nuclear magnetic resonance (NMR) relaxometry has emerged as a technique highly sensitive to porosity in such materials. Nonetheless, streamlined methods to estimate pore size from NMR relaxometry remain elusive. Previous attempts have been hindered by inverting a time domain signal to relaxation rate distribution, and dealing with resulting parameters that vary in number, location, and magnitude. Here we invoke well‐established machine learning techniques to directly correlate time domain signals to BET surface areas for a set of metal‐organic frameworks (MOFs) imbibed with solvent at varied concentrations. We employ this series of MOFs to establish a correlation between NMR signal and surface area via partial least squares (PLS), following screening with principal component analysis, and apply the PLS model to predict surface area of various nanoporous materials. This approach offers a high‐throughput, non‐destructive way to assess porosity in c.a. one minute. We anticipate this work will contribute to the development of new materials with optimized pore sizes for various applications.

Fricke, Sophia N.↗

Multivariate Machine Learning Models of Nanoscale Porosity from Ultrafast NMR Relaxometry

Abstract Nanoporous materials are of great interest in many applications, such as catalysis, separation, and energy storage. The performance of these materials is closely related to their pore sizes, which are inefficient to determine through the conventional measurement of gas adsorption isotherms. Nuclear magnetic resonance (NMR) relaxometry has emerged as a technique highly sensitive to porosity in such materials. Nonetheless, streamlined methods to estimate pore size from NMR relaxometry remain elusive. Previous attempts have been hindered by inverting a time domain signal to relaxation rate distribution, and dealing with resulting parameters that vary in number, location, and magnitude. Here we invoke well‐established machine learning techniques to directly correlate time domain signals to BET surface areas for a set of metal‐organic frameworks (MOFs) imbibed with solvent at varied concentrations. We employ this series of MOFs to establish a correlation between NMR signal and surface area via partial least squares (PLS), following screening with principal component analysis, and apply the PLS model to predict surface area of various nanoporous materials. This approach offers a high‐throughput, non‐destructive way to assess porosity in c.a. one minute. We anticipate this work will contribute to the development of new materials with optimized pore sizes for various applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Decomposition of Irganox 1010 in plastic bonded explosives

Abstract Degradation pathways of Irganox 1010 in aged plastic bonded explosive (PBX) 9501 were investigated using ultrahigh performance liquid chromatography coupled to quadrupole time of flight mass spectrometry (UHPLC‐QTOF). Using a targeted approach, a total of 44 Irganox 1010 decomposition products were discovered. These decomposition products were formed through hydrolysis, scission, and/or oxidation of Irganox 1010. The hydrolytic decomposition of Irganox is a straightforward process resulting in the cleavage of the ester group(s) while oxidation and scission are more complicated and can happen at multiple locations on the Irganox 1010 molecule. Moreover, due to the symmetric nature of Irganox 1010, multiple decomposition reactions can occur. Indeed some decomposition products exhibited hydrolysis, oxidation, and scission. In order to probe any trends in the aged PBX 9501 samples, principal component analysis (PCA) was implemented. The greatest chemical differences between the aged PBX samples was hydrolysis of the ester functional groups on Irganox 1010. Despite the negative connotations of hydrolysis, the Irganox 1010 decomposition products are still able to function as a radical scavenger in PBX 9501 as intended.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A novel methodology to integrate outcomes regarding perioperative pain experience into a composite score: Prediction model development and validation

Abstract Background An integrated score that globally assesses perioperative pain experience and rationally weights each component has not yet been developed. Methods A development dataset specific to adult Chinese patients undergoing orthopaedic surgery was obtained from PAIN OUT (1985 qualified patients of 2244). A more recent validation dataset obeying the same conditions was obtained from the Chinese Anaesthesia Shared‐database Platform (1004 qualified patients of 1032). Outcomes were assessed using the International Pain Outcomes Questionnaire (IPO‐Q), which comprises key patient‐level outcomes of perioperative pain management, including pain experience and perceptions of care. Using principal component analysis and regression models, a composite score (CS) was inferred to integrate pain experience. The discrimination of the CS for dissatisfaction and desire for more pain treatment was compared with that of the worst pain score. Results A CS was developed from the 12 items of the IPO‐Q regarding pain experience. The weight for calculating the CS was worst pain 11, least pain 17, time spent in severe pain 11, interference with activity in bed 9, interference with breathing deeply or coughing 10, interference with sleep 9, anxiety 12, helplessness 12, nausea 0, drowsiness 2, itch 5 and dizziness 2. In external validation, the CS indicated superior discrimination to the worst pain in predicting dissatisfaction ( p < 0.001) and desire for more pain treatment ( p < 0.001). Conclusions This study introduced a methodology to integrate outcomes regarding perioperative pain experience into a CS, which was based on the weight of each item. Significance This novel methodology sheds additional light on the riveting issue of carefully integrating several measures into a composite endpoint, which may be useful for quality improvement purposes when addressing the impact of a change in clinical practice.

Jiang, Bailin↗

Hidden Features: How Subsurface and Landscape Heterogeneity Govern Hydrologic Connectivity and Stream Chemistry in a Montane Watershed

ABSTRACT Hydrologic connectivity is defined as the connection among stores of water within a watershed and controls the flux of water and solutes from the subsurface to the stream. Hydrologic connectivity is difficult to quantify because it is goverened by heterogeniety in subsurface storage and permeability and responds to seasonal changes in precipitation inputs and subsurface moisture conditions. How interannual climate variability impacts hydrologic connectivity, and thus stream flow generation and chemistry, remains unclear. Using a rare, four‐year synoptic stream chemistry dataset, we evaluated shifts in stream chemistry and stream flow source of Coal Creek, a montane, headwater tributary of the Upper Colorado River. We leveraged compositional principal component analysis and end‐member mixing to evaluate how seasonal and interannual variation in subsurface moisture conditions impacts stream chemistry. Overall, three main findings emerged from this work. First, three geochemically distinct end members were identified that constrained stream flow chemistry: reach inflows, and quick and slow flow groundwater contributions. Reach inflows were impacted by historic base and precious metal mine inputs. Bedrock fractures facilitated much of the transport of quick flow groundwater and higher‐storage subsurface features (e.g., alluvial fans) facilitated the transport of slow flow groundwater. Second, the contributions of different end members to the stream changed over the summer. In early summer, stream flow was composed of all three end members, while in late summer, it was composed predominantly of reach inflows and slow flow groundwater. Finally, we observed minimal differences in proportional composition in stream chemistry across all four years, indicating seasonal variability in subsurface moisture and spatial heterogeneity in landscape and geologic features had a greater influence than interannual climate fluctuation on hydrologic connectivity and stream water chemistry. These findings indicate that mechanisms controlling solute transport (e.g., hydrologic connectivity and flow path activation) may be resilient (i.e., able to rebound after perturbations) to predicted increases in climate variability. By establishing a framework for assessing compositional stream chemistry across variable hydrologic and subsurface moisture conditions, our study offers a method to evaluate watershed biogeochemical resilience to variations in hydrometeorological conditions.

Johnson, Keira [College of Earth, Ocean, and Atmos↗

Bayesian Calibration of Stochastic Agent Based Model via Random Forest

Agent-based models (ABM) provide an excellent framework for modeling outbreaks and interventions in epidemiology by explicitly accounting for diverse individual interactions and environments. However, these models are usually stochastic and highly parametrized, requiring precise calibration for predictive performance. When considering realistic numbers of agents and properly accounting for stochasticity, this high-dimensional calibration can be computationally prohibitive. This paper presents a random forest-based surrogate modeling technique to accelerate the evaluation of ABMs and demonstrates its use to calibrate an epidemiological ABM named CityCOVID via Markov chain Monte Carlo (MCMC). The technique is first outlined in the context of CityCOVID's quantities of interest, namely hospitalizations and deaths, by exploring dimensionality reduction via temporal decomposition with principal component analysis (PCA) and via sensitivity analysis. The calibration problem is then presented, and samples are generated to best match COVID-19 hospitalization and death numbers in Chicago from March to June in 2020. Further, these results are compared with previous approximate Bayesian calibration (IMABC) results, and their predictive performance is analyzed, showing improved performance with a reduction in computation.

60 APPLIED LIFE SCIENCES↗

Robust Spectral Anomaly Detection in EELS Spectral Images via 3D Convolutional Variational Autoencoders

Abstract A 3D Convolutional Variational Autoencoder (3D‐CVAE) is introduced for automated anomaly detection in electron energy‐loss spectroscopy spectrum imaging (EELS‐SI) data. This approach leverages the full 3D structure of EELS‐SI data to detect subtle spectral anomalies while preserving both spatial and spectral correlations across the datacube. By employing cross‐entropy loss and training on bulk spectra, the model learns to reconstruct bulk features characteristic of the defect‐free material. In exploring methods for anomaly detection, both the 3D‐CVAE approach and principal component analysis (PCA) are evaluated, testing their performance using FeL‐edge ΔEpeak shifts designed to simulate material defects. These results show that 3D‐CVAE achieves superior anomaly detection and maintains consistent performance across various shift magnitudes. The method demonstrates clear bimodal separation between bulk and anomalous spectra, enabling reliable classification. Further analysis verifies that lower‐dimensional representations are robust to anomalies in the data. While performance advantages over PCA diminish with decreasing anomaly concentration, our method maintains high reconstruction quality even in challenging, noise‐dominated spectral regions. This approach provides a robust framework for unsupervised automated detection of spectral anomalies in EELS‐SI data, particularly valuable for analyzing complex material systems.

Chemistry↗

Reinforcement learning for real-time process control in high-temperature superconductor manufacturing

With high efficiency and low energy loss, high-temperature superconductors (HTS) have demonstrated their profound applications in various fields, such as medical imaging, transportation, accelerators, microwave devices, and power systems. The high-field applications of HTS tapes have raised the demand for producing cost-effective tapes with long lengths in superconductor manufacturing. However, achieving the uniform and enhanced performance of a long HTS tape is challenging due to the unstable growth conditions in the manufacturing process. Although it is confirmed that the process parameters during the advanced metal organic chemical vapor deposition (A-MOCVD) process influence the uniformity of the produced HTS tapes, the high-dimensional process parameter signals and their complicated interactions make it difficult to develop an effective control policy. In this paper, we propose a local measure for the uniformity of HTS tapes to provide instant feedback for our control policy. Then, we model the manufacturing of HTS tapes as a Markov decision process (MDP) with continuous state and action spaces to assess the instant reward in real time in our feedback control model. As our MDP involves continuous and high-dimensional state and action spaces, a neural fitted Q-iteration (NFQ) algorithm is adopted to solve the MDP with artificial neural network (ANN) function approximation. The collinearity of process parameters can restrict our capability of adjusting the process parameters, which is addressed by the principal component analysis (PCA) in our method. The control policy adjusts the PCA of process parameters using the NFQ algorithm. In conclusion, based on our case studies on real A-MOCVD dataset, the obtained control policy increases the average uniformity of tapes by 5.6% and performs especially well on sample HTS tapes with a low uniformity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Spatiotemporal distribution of chemical signatures exhibited by Myxococcus xanthus in response to metabolic conditions

Myxococcus xanthus is a common soil bacterium with a complex life cycle, which is known for production of secondary metabolites. However, little is known about the effects of nutrient availability on M. xanthus metabolite production. In this study, we utilize confocal Raman microscopy (CRM) to examine the spatiotemporal distribution of chemical signatures secreted by M. xanthus and their response to varied nutrient availability. Here, ten distinct spectral features are observed by CRM from M. xanthus grown on nutrient-rich medium. However, when M. xanthus is constrained to grow under nutrient-limited conditions, by starving it of casitone, it develops fruiting bodies, and the accompanying Raman microspectra are dramatically altered. The reduced metabolic state engendered by the absence of casitone in the medium is associated with reduced, or completely eliminated, features at 1140 cm –1 , 1560 cm –1 , and 1648 cm –1 . In their place, a feature at 1537 cm –1 is observed, this feature being tentatively assigned to a transitional phase important for cellular adaptation to varying environmental conditions. In addition, correlating principal component analysis heat maps with optical images illustrates how fruiting bodies in the center co-exist with motile cells at the colony edge. While the metabolites responsible for these Raman features are not completely identified, three M. xanthus peaks at 1004, 1151, and 1510 cm –1 are consistent with the production of lycopene. Thus, a combination of CRM imaging and PCA enables the spatial mapping of spectral signatures of secreted factors from M. xanthus and their correlation with metabolic conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mallat Scattering Transformation based surrogate for Magnetohydrodynamics

Abstract A Machine and Deep Learning (MLDL) methodology is developed and applied to give a high fidelity, fast surrogate for 2D resistive MagnetoHydroDynamic (MHD) simulations of Magnetic Liner Inertial Fusion (MagLIF) implosions. The resistive MHD code is used to generate an ensemble of implosions with different liner aspect ratios, initial gas preheat temperatures (that is, different adiabats), and different liner perturbations. The liner density and magnetic field as functions of x , y , and z were generated. The Mallat Scattering Transformation (MST) is taken of the logarithm of both fields and a Principal Components Analysis (PCA) is done on the logarithm of the MST of both fields. The fields are projected onto the PCA vectors and a small number of these PCA vector components are kept. Singular Value Decompositions of the cross correlation of the input parameters to the output logarithm of the MST of the fields, and of the cross correlation of the SVD vector components to the PCA vector components are done. This allows the identification of the PCA vectors vis-a-vis the input parameters. Finally, a Multi Layer Perceptron (MLP) neural network with ReLU activation and a simple three layer encoder/decoder architecture is trained on this dataset to predict the PCA vector components of the fields as a function of time. Details of the implosion, stagnation, and the disassembly are well captured. Examination of the PCA vectors and a permutation importance analysis of the MLP show definitive evidence of an inverse turbulent cascade into a dipole emergent behavior. The orientation of the dipole is set by the initial liner perturbation. The analysis is repeated with a version of the MST which includes phase, called Wavelet Phase Harmonics (WPH). While WPH do not give the physical insight of the MST, they can and are inverted to give field configurations as a function of time, including field-to-field correlations.

97 MATHEMATICS AND COMPUTING↗

A Gaussian process autoregressive model capturing microstructure evolution paths in a Ni–Mo–Nb alloy

Additive manufacturing is increasingly being employed to produce components of complex geometries in structural alloys because of the expected energy savings associated with the near-net-shape capability and the ability to build in novel internal features that are not possible with many conventional manufacturing approaches. However, because of the extreme thermal conditions encountered, the non-equilibrium microstructures produced during powder bed-based additive manufacturing processes must be subjected to custom post-heat treatment processes to recover the target mechanical properties. Phase-field models and simulation techniques have matured to a state where the microstructure evolution paths, and the morphologies of the resulting precipitate phases can be predicted reasonably accurately, considering alloy-specific thermodynamic and kinetic aspects of the nucleation and growth processes. However, phase-field simulations are computationally intensive, which precludes the ability to apply the simulations directly to the length scale of the entire component. Therefore, it is highly desirable to develop low-computational-cost surrogate models that effectively capture the physics at the microstructural length scale, while facilitating the design of optimized processing conditions resulting in location-specific targeted microstructures at the component scale. The work presented here demonstrates the application of the materials knowledge system framework to develop a surrogate model that effectively captures the microstructural path during annealing of a Ni–Mo–Nb alloy containing different Mo and Nb compositions known to segregate during solidification under additive manufacturing conditions. Specifically, the surrogate model built in this work is based on a Gaussian process autoregressive model informed by statistical representation of simulated microstructures using two-point correlations and dimensionality reduction through principal component analysis. In conclusion, this surrogate model is shown to capture the bifurcation of the microstructural path during precipitation, which yields a microstructure dominated by the $\gamma^{\prime\prime}$ phase at high Nb concentrations and the $\delta$ phase at low Nb concentrations.

36 MATERIALS SCIENCE↗

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka↗

Data Science Techniques, Assumptions, and Challenges in Alloy Clustering and Property Prediction

Data analytics methods have been increasingly applied to understanding materials chemistry, processing due to the manufacturing approach, and uni-axial and cyclic property relationships in the highly complex space of alloy design. There are several benefits to applying data analytics to this space, including the ability to manage non-linearities in the responses of the alloy attributes and the resulting mechanical properties. However, key difficulties in applying and understanding the results of data analytics include the often lack of reported assumptions and data processing steps necessary to improve interpretation and reproducibility in derived results. In this work, the methods used to generate clustering and correlation analyses for experimental 9% Cr ferritic-martensitic steel data were investigated and the resulting implications for mechanical property predictions were assessed. This work uses principal component analysis, partitioning around medoids, t-SNE, and k-means clustering to investigate trends in composition, processing and microstructure information with creep and tensile properties, building on work done previously using a smaller version of the same dataset. The initial assumptions, preprocessing steps and methods are investigated and outlined in order to depict the fine level of detail required to convey the steps taken to process data and produce analytical results. Here, the variations in the resulting analyses are explored due to the influence of new and more varied data.

36 MATERIALS SCIENCE↗