Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine-learning featurization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A machine-learning-aided data recovery approach for predicting multi-material thermal behaviors in advanced test reactor capsules

Instrumented experiments conducted at test reactors are essential to the deployment of new advanced reactor systems. Designing new experiments and generating data on specific reactor conditions require significant investments in terms of both time and cost. Finite element analysis software can be used to create high-fidelity models of experiment environments in order to support the actual experiments, but computation time remains a concern in terms of applying outcomes to real-time usage of data (e.g., a digital twin [DT]). Here, the present research proposes a machine-learning (ML) aided approach to making temperature and displacement predictions based on the thickness of the outer gas gap on the experimental capsule used for in-pile demonstration of a novel new thermal conductivity probe in the Advanced Test Reactor (ATR). This capsule consisted of U10Zr fuel, a rodlet, sodium, and inner and outer capsules. Gas gaps existed between the fuel and the rodlet, and between the inner and the outer capsule. The learning data pertained to an experimental capsule's radial distributions of temperature and displacement, as obtained based on Abaqus and the physical features. For the first step of ML sequence, the temperature was predicted using three positional parameters. Next, the displacement was predicted using seven additional parameters. Each physical feature was normalized in order to be both nondimensional and standardized. The temperature and displacement predictions showed good agreement with the simulation results in all cases involving interpolation and extrapolation. Furthermore, data similarity enhancement increased the similarity between the training and the target data, thereby increasing the predictive accuracy of the ML models. In certain extrapolation cases involving limited original ML model accuracy, data similarity enhancement and data recovery was able to somewhat improve this accuracy.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Machine-learning and first-principles investigation of lightweight medium-entropy alloys for hydrogen-storage applications

The transition to a low-carbon economy demands efficient and sustainable energy-storage solutions, with hydrogen emerging as a promising clean-energy carrier and with metal hydrides recognized for their hydrogen-storage capacity. Here, we leverage machine learning (ML) to predict hydrogen-to-metal (H/M) ratios and solution energy by incorporating thermodynamic parameters and local lattice distortion (LLD) as key features. Our best-performing ML model provides improvements to H/M ratios and solution energies over a broad class of medium-entripy alloys (easily extendable to multi-principal-element alloys), such as Ti–Nb-X (X = Mo, Cr, Hf, Ta, V, Zr) and Co–Ni-X (X = Al, Mg, V). Ti–Nb–Mo alloys reveal compositional effects in H-storage behavior, in particular Ti, Nb, and V enhance H-storage capacity, while Mo reduces H/M and hydrogen weight percent by 40–50 %. We attributed results in molybdenum-rich alloys to slow hydrogen kinetics, as validated by our pressure-composition-temperature (PCT) isotherm experiments on pure Ti and Ti 5 Mo 95 alloys. Density functional theory (DFT) and molecular dynamics (MD) simulations also confirm that Ti and Nb promote H diffusion, whereas Mo hinders it, highlighting the interplay between electronic structure, lattice distortions, and hydrogen uptake. Notably, our Gradient Boosting Regression model identifies LLD as a critical factor in H/M predictions. Here, to aid material selection, we present two periodic tables illustrating elemental effects on (a) H 2 wt% and (b) solution energy, derived from ML, and provide a reference for identifying alloying elements that enhance hydrogen solubility and storage.

08 HYDROGEN↗

Recent advances in rational design of defect-engineered photocatalysts toward sustainable NH 3 synthesis as H 2 carrier: From fundamental and development to machine-learning

In this study, we provide a detailed overview of the fundamental mechanisms underpinning photocatalytic N 2 reduction. We also discuss advances in catalyst design for the synthesis of NH 3 . Particular emphasis is placed on the role of surface defect engineering, which includes the creation of surface defects to enhance the performance of semiconducting photocatalysts for efficient N 2 reduction. In addition, the application of a machine learning-based computational modeling approach is discussed as an important driving force for predicting and regulating NH 3 synthesis efficiency based on catalyst features and reaction conditions. Finally, existing challenges and future perspectives for improving the performance of defect-engineered photocatalysts are outlined to contribute to the ongoing discourse on sustainable ammonia generation. This review aims to clarify recent progress in the rational design of defect-containing photocatalysts for the synthesis of NH 3 and encourages innovative approaches to catalyst optimization rather than solely focusing on new materials.

08 HYDROGEN↗

Predicting Si-Anode Calendar Life Using Machine Learning: Correlating Electrolyte Properties and Electrochemical Signals

This study evaluates novel electrolytes tailored for Si-containing anodes to promote calendar-life. Drawing inspiration from advancements in electrolytes for Li-metal cells, the work investigates correlations between predicted electrolyte properties and measured electrochemical performance using several machine-learning models. By leveraging machine learning and advanced modeling techniques, this study aims to establish predictive frameworks that accelerate calendar-aging experiments and inform rational electrolyte design for Si-containing cells. In the present study, fifteen different electrolytes are evaluated in a Si-containing cell using an accelerated calendar-life protocol. For each electrolyte considered, 87 properties (features) from the Advanced Electrolyte Model were produced to identify key property/performance relationships. In this study, the best performing electrolytes were generally those formulations that included non-coordinating fluoroether solvents, and the most predictive features for long-term calendar-life were features related to salt concentration and electrolyte viscosity as well as early capacity, ionic conductivity, and Coulombic efficiency measurements. The framework developed in this study correlating electrolyte properties to measured electrochemical performance is expected to accelerate electrolyte design for Si-containing anodes and ultimately enable high-energy-density, long-life Li-ion batteries.

25 - ENERGY STORAGE↗

Unveiling X-ray absorption signatures of boron nitride via first-principles simulation and machine learning

Boron nitride (BN) allotropes hold great promise in many advanced applications ranging from optical and photonic devices to energy storage and battery systems to tribological components. The diverse functionalities of this material stem from BN’s highly tunable structural and electronic properties, which are governed by the versatile boron–nitrogen bonding configurations. Exploring the structural landscape of BN can unveil novel structures possessing unique properties suited for specific applications, therefore accelerating the design of next-generation advanced functional materials. In this work, we leverage boron K-edge X-ray absorption spectroscopy (XAS) as an effective probe for local structural features and chemical environments. A total of 210 BN crystal structures are generated via analogies to the extensive array of carbon allotropes, and XAS is simulated for each unique local motif within the resulting collection of structures. A mapping between structural features and spectral signatures was established by synergizing first-principle simulations with data-driven based post-analysis approaches. Specifically, we developed a neural network model that can satisfactorily predict spectra line shapes from local structural descriptors. Toward automatic spectroscopic interpretation of any new BN structures, supervised machine learning models, trained on this structure–spectrum dataset, can accurately infer local coordination environments from simulated XAS, highlighting the strength of this unique approach of combining high-fidelity first-principles simulation and machine-learning to accelerate target design of novel BN materials via rational understanding of local structure-spectrum correlations.

36 MATERIALS SCIENCE↗

Thermodynamically Optimized Machine-Learned Reaction Coordinates for Hydrophobic Ligand Dissociation

Ligand unbinding is mediated by its free energy change, which has intertwined contributions from both energy and entropy. It is important, but not easy, to quantify their individual contributions to the free energy profile. We model hydrophobic ligand unbinding for two systems, a methane particle and a C 60 fullerene, both unbinding from hydrophobic pockets in all-atom water. Using a modified deep learning framework, we learn a thermodynamically optimized reaction coordinate to describe the hydrophobic ligand dissociation for both systems. Interpretation of these reaction coordinates reveals the roles of entropic and enthalpic forces as the ligand and pocket sizes change. In both cases, we observe that the free-energy barrier to unbinding is dominated by entropy considerations. Furthermore, the process of methane unbinding is driven by methane solvation, while fullerene unbinding is driven first by pocket wetting and then fullerene wetting. For both solutes, the direct importance of the distance from the binding pocket to the learned reaction coordinate is present, but low. Furthermore, our framework and subsequent feature important analysis thus give useful thermodynamic insight into hydrophobic ligand dissociation problems that are otherwise difficult to glean.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Combining Deep Learning With Physics Based Features in Explosion-Earthquake Discrimination

This paper combines the power of deep-learning with the generalizability of physics-based features, to present an advanced method for seismic discrimination between earthquakes and explosions. The proposed method contains two branches: a deep learning branch operating directly on seismic waveforms or spectrograms, and a second branch operating on physics-based parametric features. These features are high-frequency P/S amplitude ratios and the difference between local magnitude (ML) and coda duration magnitude (MC). The combination achieves better generalization performance when applied to new regions than models that are developed solely with deep learning. Further, we also examined which parts of the waveform data dominate deep learning decisions (i.e., via Grad-CAM). Such visualization provides a window into the black-box nature of the machine-learning models and offers new insight into how the deep learning derived models use data to make decisions.

58 GEOSCIENCES↗

Ensemble Federated Machine Learning‐Based Cybersecurity Situational Awareness in Microgrid Network

Cyber-physical microgrids are vulnerable to stealthy cybersecurity threats that disguise their actions through the exploitation of system knowledge. Such actions can severely impacts microgrids deployed in defense bases, slowing the response time of military forces during national emergencies. Several machine-learning algorithms have been proposed to detect intrusions in the grid networks; however, these traditional machine-learning algorithms lack data privacy and are subject to several adversarial machine-learning threats. This paper proposes a novel federated machine learning (FML)-based three-model framework to detect and identify stealthy data-integrity attacks while ensuring data privacy in microgrid networks. The proposed architecture uses a variational mode decomposition technique to extract derived features from incoming measurement and control datasets. The extraction of these derived features allows FML models to learn minute variations in data patterns that allow them to perform significantly better than the models trained with generic datasets consisting of raw features. Our experimental results show the efficient performance of the proposed methodology against different types of data integrity attacks while considering primary and secondary controllers in microgrids. Further, the applied FML-integrated random forest ensemble algorithm outperforms the existing generic FML algorithms during noisy and noise-free datasets with prediction latencies of only 91–134 µs per sample within the 0.1 s sampling interval and requires communication bandwidth of around ∼8.25 KB/s at the control center and ∼2.7 KB/s per edge client for communication.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine Learning for Mapping Multipactor Susceptibility in RF Systems: Capabilities and Generalization Constraints

Multipactor is a surface-driven electron avalanche phenomenon that degrades the performance and reliability of radio-frequency (RF) systems in particle accelerator and vacuum electronics applications. Multipactor behavior in a given device structure is conventionally assessed through susceptibility charts, which provide a parameter-space characterization of the instability. In this work, we assess the capabilities of machine-learning (ML) models to learn and predict such susceptibility charts and analyze the constraints governing their generalization across materials. Using a simulation-derived dataset spanning six distinct secondary-electron-yield material profiles in a canonical two-surface planar geometry, we train supervised regression models and artificial neural networks to predict the time-averaged electron growth rate, δavg, across the relevant parameter space. Model performance is evaluated using metrics that explicitly probe the structure of susceptibility charts, including Intersection over Union, Structural Similarity Index, and correlation analysis. Tree-based ensemble models outperform neural-network models in reconstructing susceptibility regions and in generalizing across material domains. Principal-component analysis reveals disjoint material feature distributions, indicating that the piecewise mode structure of multipactor susceptibility is difficult to represent with a single global model and that generalization is constrained by data coverage rather than by model complexity. An exhaustive reduced-coverage study further shows that sparse material-space coverage can yield mean performance in the same general range but producing large variability in the susceptibility-region overlap. These results clarify the capabilities of ML-based surrogate models for parameter-space characterization of multipactor discharge. They also provide guidance for their appropriate use in RF system design.

43 PARTICLE ACCELERATORS↗

Physics-guided dual implicit neural representations for source separation

Significant challenges exist in efficient data analysis of most advanced experimental and observational techniques because the collected signals often include unwanted contributions, such as background and signal distortions, that can obscure the physically relevant information of interest. To address this, we have developed a self-supervised machine-learning approach for source separation using a dual implicit neural representation framework that jointly trains two neural networks: one for approximating distortions of the physical signal of interest and the other for learning the effective background contribution. Our method learns directly from the raw data by minimizing a reconstruction-based loss function without requiring labeled data or pre-defined dictionaries. We demonstrate the effectiveness of our framework by considering a challenging case study involving large-scale simulated, as well as experimental, momentum-energy-dependent inelastic neutron scattering data in a four-dimensional parameter space, characterized by heterogeneous background contributions and unknown distortions to the target signal. The method is found to successfully separate physically meaningful signals from a complex or structured background even when the signal characteristics vary across all four dimensions of the parameter space. An analytical approach that informs the choice of the regularization parameter is presented. Our method offers a versatile framework for addressing source separation problems across diverse domains, ranging from superimposed signals in astronomical measurements to structural features in biomedical image reconstructions.

47 OTHER INSTRUMENTATION↗

Novel machine-learning method for spin classification of neutron resonances

The performance of nuclear reactors and other nuclear systems depends on a precise understanding of the neutron interaction cross sections for materials used in these systems. These cross sections exhibit resonant structure whose shape is determined in part by the angular-momentum quantum numbers of the resonances. The correct assignment of the quantum numbers of neutron resonances is, therefore, paramount. In this project, we apply machine learning to automate the quantum number assignments using only the resonances' energies and widths and not relying on detailed transmission or capture measurements. The classifier used for quantum number assignment is trained using stochastically generated resonance sequences whose distributions mimic those of real data. Here we explore the use of several physics-motivated features for training our classifier. These features amount to out-of-distribution tests of a given resonance's widths and resonance-pair spacings. We pay special attention to situations where either capture widths cannot be trusted for classification purposes or where there is insufficient information to classify resonances by the total spin J. We demonstrate the efficacy of our classification approach using simulated and actual 52 Cr resonance data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Explainable machine learning for hydrogen diffusion in metals and random binary alloys

Hydrogen diffusion in metals and alloys plays an important role in the discovery of new materials for fuel cell and energy storage technology. While analytic models use hand-selected features that have clear physical ties to hydrogen diffusion, they often lack accuracy when making quantitative predictions. Machine learning models are capable of making accurate predictions, but their inner workings are obscured, rendering it unclear which physical features are truly important. To develop interpretable machine learning models to predict the activation energies of hydrogen diffusion in metals and random binary alloys, we create a database for physical and chemical properties of the species and use it to fit six machine learning models. Our models achieve root-mean-squared errors between 98–119 meV on the testing data and accurately predict that elemental Ru has a large activation energy, while elemental Cr and Fe have small activation energies. By analyzing the feature importances of these fitted models, we identify relevant physical properties for predicting hydrogen diffusivity. While metrics for measuring the individual feature importances for machine learning models exist, correlations between the features lead to disagreement between models and limit the conclusions that can be drawn. Instead grouped feature importance, formed by combining the features via their correlations, agree across the six models and reveal that the two groups containing the packing factor and electronic specific heat are particularly significant for predicting hydrogen diffusion in metals and random binary alloys. In conclusion, this framework allows us to interpret machine learning models and enables rapid screening of new materials with the desired rates of hydrogen diffusion.

36 MATERIALS SCIENCE↗

Entropy-driven Optimal Sub-sampling of Fluid Dynamics for Developing Machine-learned Surrogates

Optimal sub-sampling of large datasets from fluid dynamics simulations is essential for training reduced-order machine learned models. A method using Shannon entropy was developed to weight flow features according to their level of information content, such that the most informative features can be extracted and used for training a surrogate model. The method is demonstrated in the canonical flow over a cylinder problem simulated with OpenFOAM. Both time-independent predictions and temporal forecasting were investigated as well as two types of prediction targets: local per-grid-point predictions and global per-time-step predictions. When tested on training a surrogate model, results indicate that our entropy-based sampling method typically outperforms random sampling and yields more reproducible results in less iterations. Finally, the method was used to train a surrogate model for modeling turbulence in magnetohydrodynamic flows, which revealed various challenges and opportunities for future research.

Brewer, Wes↗

STM/S Grid LDOS Data and Analysis Code for Deciphering Majorana Zero Modes in Topological Superconductor

This dataset provides raw millikelvin scanning tunneling microscopy/spectroscopy (STM/S) grid spectroscopy data and Python analysis scripts supporting the manuscript “Deciphering Majorana Zero Modes in Topological Superconductor FeTe0.55Se0.45 with Machine-Learning-Assisted Spectral Deconvolution.” The dataset includes a raw grid spectroscopy file acquired on FeTe0.55Se0.45 at 40 mK under magnetic field, together with Python/Jupytext analysis scripts used for STM/S data processing, visualization, spectral deconvolution, Lorentzian peak fitting, feature extraction, machine-learning-assisted clustering, and figure generation. These files support the analysis of vortex-core local density of states and the identification of zero-bias-peak-related spectral components from complex in-gap states. The dataset is intended to provide a citable archival record of the data and analysis code associated with the published manuscript and to support transparency and reproducibility of the reported STM/S and machine-learning workflow.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

TCR-H: explainable machine learning prediction of T-cell receptor epitope binding on unseen datasets

Artificial-intelligence and machine-learning (AI/ML) approaches to predicting T-cell receptor (TCR)-epitope specificity achieve high performance metrics on test datasets which include sequences that are also part of the training set but fail to generalize to test sets consisting of epitopes and TCRs that are absent from the training set, i.e., are ‘unseen’ during training of the ML model. We present TCR-H, a supervised classification Support Vector Machines model using physicochemical features trained on the largest dataset available to date using only experimentally validated non-binders as negative datapoints. TCR-H exhibits an area under the curve of the receiver-operator characteristic (AUC of ROC) of 0.87 for epitope ‘hard splitting’ (i.e., on test sets with all epitopes unseen during ML training), 0.92 for TCR hard splitting and 0.89 for ‘strict splitting’ in which neither the epitopes nor the TCRs in the test set are seen in the training data. Furthermore, we employ the SHAP (Shapley additive explanations) eXplainable AI (XAI) method for post hoc interrogation to interpret the models trained with different hard splits, shedding light on the key physiochemical features driving model predictions. TCR-H thus represents a significant step towards general applicability and explainability of epitope:TCR specificity prediction.

60 APPLIED LIFE SCIENCES↗

PowerModel-AI: A First On-the-Fly Machine-Learning Predictor for AC Power Flow Solutions

The real-time creation of machine-learning models via active or on-the-fly learning has attracted considerable interest across various scientific and engineering disciplines. These algorithms enable machines to build models autonomously while remaining operational. Through a series of query strategies, the machine can evaluate whether newly encountered data fall outside the scope of the existing training set. In this study, we introduce PowerModel-AI, an end-to-end machine learning software designed to accurately predict AC power flow solutions. We present detailed justifications for our model design choices and demonstrate that selecting the right input features effectively captures load flow decoupling inherent in power flow equations. Our approach incorporates on-the-fly learning, where power flow calculations are initiated only when the machine detects a need to improve the dataset in regions where the model’s suboptimal performance is based on specific criteria. Otherwise, the existing model is used for power flow predictions. This study includes analyses of five Texas A&M synthetic power grid cases, encompassing the 14-, 30-, 37-, 200-, and 500-bus systems. The training and test datasets were generated using PowerModels.jl, an open-source power flow solver/optimizer developed at Los Alamos National Laboratory, NM, USA.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Advances in Computational Approaches for Estimating Passive Permeability in Drug Discovery

Passive permeation of cellular membranes is a key feature of many therapeutics. The relevance of passive permeability spans all biological systems as they all employ biomembranes for compartmentalization. A variety of computational techniques are currently utilized and under active development to facilitate the characterization of passive permeability. These methods include lipophilicity relations, molecular dynamics simulations, and machine learning, which vary in accuracy, complexity, and computational cost. This review briefly introduces the underlying theories, such as the prominent inhomogeneous solubility diffusion model, and covers a number of recent applications. Various machine-learning applications, which have demonstrated good potential for high-volume, data-driven permeability predictions, are also discussed. Due to the confluence of novel computational methods and next-generation exascale computers, we anticipate an exciting future for computationally driven permeability predictions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗