Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Predictive Data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Attack Surface Analysis of the Digital Twins interface with Advanced Sensor and Instrumentation Interfaces: Cyber Threat Assessment and Attack Demonstration for Digital Twins in Advanced Reactor Architectures

A digital twin is a virtual representation of a physical system or object using real-time data that can predict and analyze how the system or object performs. This relatively new technology can be applied to the field of nuclear power generation, to aid in the design and development of new nuclear power plants and reduce operation costs using predictive maintenance and other data analytical methods. While there are already companies utilizing simulation software to train operators and technicians in the nuclear industry, some are now transitioning to utilizing their existing technology, software, and methods to develop digital twin solutions for the next generation of nuclear power plants, offering their services to utilities and government organizations around the world.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Novel Hot Gas Components for Gas Turbine Engines Enabled by Materials and Additive Manufacturing Process Development

Additive Manufacturing (AM), also known as 3D printing, has emerged as a manufacturing method that enables new design freedom for gas turbine engine manufacturers. However, the material selection for AM processable high-temperature super alloys is currently limited. Additionally, the heat transfer performance of AM enabled micro-cooling architectures is not yet well understood. Accordingly, in support of advanced manufacturing and engine performance development, Oak Ridge National Laboratory (ORNL)and Solar Turbines (Solar) conducted a multidisciplinary project to generate both AM super alloy material properties data and micro-channel performance data for two AM super alloys. The data supported the design and analysis of an internally cooled turbine hot section AM tip shoe component. This data was used to analytically predict the reduction in operating temperature of a gas turbine tip shoe. The work concluded that the cooling flow required to cool the tip shoe can be tuned to suit the efficiency improvements desired in an industrial gas turbine.

36 MATERIALS SCIENCE↗

Predictive modeling of Néel temperature in austenitic alloys using CALPHAD and data analytics

The Néel temperature is a crucial yet often overlooked parameter in calculating the stacking fault energy (SFE) of austenitic alloys. Several empirical equations have been proposed to estimate the Néel temperature of austenitic alloys, which are then used to calculate the SFE and explain deformation mechanisms. However, these empirical equations, typically derived using linear regression algorithms, are often simplistic and may fail to capture the complex interactions among multiple alloying elements that influence the Néel temperature. Moreover, their applicability is usually limited to specific compositional ranges. In this study, we propose a CALPHAD based approach and develop a surrogate decision tree based regression model capable of capturing the interactions among multiple alloying elements to predict the Néel temperature. Predictions from both the CALPHAD approach and the regression model show close agreement with experimental measurements reported in the literature. In conclusion, the implications of accurate Néel temperature predictions on the calculated SFE and deformation mechanisms are also discussed.

36 MATERIALS SCIENCE↗

Predicting Sintering Window of Binder Jet Additively Manufactured Parts Using a Coupled Data Analytics and CALPHAD Approach

Batch-to-batch variation in powder compositions for binder jet additive manufacturing (BJAM) can significantly deter defining an “ideal” sintering window for a given alloy. One way to overcome the problem is by running sintering experiments at various temperatures for each batch of the powder. However, such an approach increases the time required to achieve large-scale production of parts. The predictive capabilities of computational thermodynamic tools like CALPHAD can be leveraged to overcome the challenge, especially for binder jet additive manufacturing, since the process occurs under near-equilibrium conditions. However, calculating the sintering window using CALPHAD can be computationally expensive, considering many possible feedstock compositions within “specification”. Here, we generate high throughput CALPHAD data for nickel-based superalloys to develop machine learning models to predict the sintering window rapidly. The predictive capability of the models has been validated using published results on BJAM of Inconel 718 and 625. Further, validated models are lightweight and can be deployed in an industrial setting to get sintering window in an accelerated manner.

36 MATERIALS SCIENCE↗

Accelerated Materials Design for Molten Salt Technologies Using Innovative High-Throughput Methods

The focus of the project is on building an innovative accelerated materials design platform for molten salt technologies using novel high-throughput methods coupled to data analytics. The main objectives is to predict a FeCrMnNi alloy compositional space with better corrosion resistance than stainless steel 316 and identify new molten salt corrosion mechanisms. The project demonstrates the feasibility to use high-throughput methods coupled to data analytics to accelerate alloy design for extreme environments applications. Using a trained and tested machine learning (ML) model, 2000 FeCrMnNi alloy corrosion rate in molten chloride salts were predicted and a compositional field with corrosion rate lower than 316 stainless steel was identified. The ML model interpretability unveiled multiple features of importance in the model prediction. Some features were expected to be of relative significance, such as work function, surface energy and alloy electronegativity, and the ML model interpretability analysis confirmed those. On the other hand, the most important feature is the diffusion coefficient of Ni in the bulk alloy which indicates that a surface diffusion mechanism plays an important role in the overall corrosion mechanism in molten salts.

36 MATERIALS SCIENCE↗

Battery Life Prediction Using Reduced-Order Physics Models and Machine Learning (CRADA Final Report)

Phase 1 (Original CRADA, plus no-cost extension modifications #1-3, 6/1/2017 to 3/13/2021): The Australian Department of Defence (AUDoD) is performing accelerated aging tests of Li-ion batteries to benchmark their reliability and degradation characteristics. Using its previously developed battery lifetime predictive model framework, the National Laboratory of the Rockies (NLR) will develop analytical models based the AUDoD data to predict lifetime of the multiple Li-ion battery chemistries under real-world use scenarios of interest to AUDoD. The NLR model is based on physical degradation mechanisms encountered by Li-ion batteries and has been previously validated. Phase 2 (CRADA modification #4, plus no-cost extension modification #5, 2/22/2021 to 3/30/2025): Train and support Australian Department of Defence personnel to use NLR software for model-based estimation of Li-ion battery lifetime using accelerated battery aging data collected by the Australian Department of Defence. Under separate DOE funding from 2019 to 2021, NLR enhanced its battery life-prediction software using machine learning algorithms to automate portions of the model-fitting process, requiring significantly less labor and expert judgment and also adding uncertainty quantification, increasing statistical rigor. Under Phase 2, NLR will customize NLR Software and provide it to AuDoD. NLR will enhance its NLR Model to capture aging modes of AuDoD's multi-cell modules, including cell-balancing effects. NLR will develop example single-cell and multi-cell models based on one AuDoD battery aging dataset. NLR will train AuDoD personnel on NLR Software. By the conclusion of the project, NLR will have provided AuDoD the training materials, a user manual and software needed to perform their own analysis of additional and/or future battery aging datasets.

33 ADVANCED PROPULSION SYSTEMS↗

Comparing Experimentally-Measured Sand’s Times with Concentrated Solution Theory Predictions in a Polymer Electrolyte

We compare the electrochemically measured Sand’s time, the time required for the cell potential to diverge when the applied current density exceeds the limiting current, with theoretical predictions for a 0.47 M poly(ethylene oxide) (5 kg mol −1 )/LiTFSI electrolyte. The theoretical predictions are made using concentrated solution theory which accounts for both concentration polarization and polymer motion, using independently measured parameters that depend on concentration, c : conductivity ( κ ), salt diffusion coefficient ( D ), cationic transference number with respect to the solvent velocity ( t + 0 ), thermodynamic factor 1 + dln f ± dln c , and partial molar volume of the salt ( V ̅ ); f ± is the mean molar activity coefficient of the salt. We find quantitative agreement between experimental data and theoretical predictions. We derive a generalized analytical expression for Sand’s time for electrolytes based on dilute solution theory. This expression correctly predicts the divergence of the Sand’s time at the limiting current, in agreement with experimental data and concentrated solution theory predictions. When the applied current is large compared to the limiting current, the analytical expression approaches the standard expression for Sand’s time used in the literature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NIOLD Run 4 Report

In this report we discuss the motivation, methods, and analysis of the IOTA Run 4 studies of the Nonlinear Integrable Optics, Landau Damping (NIOLD) experiment. We introduce Landau Damping, an effect that damps collective instabilities in particle accelerators. We also introduce and discuss a new method to measure stability diagrams which quantify the strength of Landau Damping, that employs an antidamper. We also present the data collection process, the data analysis procedure, and the preliminary results. The first results qualitatively agree with the analytical predictions and the simulations. Next steps for the current data and goals for future data are also discussed.

43 PARTICLE ACCELERATORS↗

Constraining Cosmology with Simulation-based inference and Optical Galaxy Cluster Abundance

We test the robustness of simulation-based inference (SBI) in the context of cosmological parameter estimation from galaxy cluster counts and masses in simulated optical datasets. We construct ``simulations'' using analytical models for the galaxy cluster halo mass function (HMF) and for the observed richness (number of observed member galaxies) to train and test the SBI method. We compare the SBI parameter posterior samples to those from an MCMC analysis that uses the same analytical models to construct predictions of the observed data vector. The two methods exhibit comparable performance, with reliable constraints derived for the primary cosmological parameters, ($\Omega_m$ and $\sigma_8$), and richness-mass relation parameters. We also perform out-of-domain tests with observables constructed from galaxy cluster-sized halos in the Quijote simulations. Again, the SBI and MCMC results have comparable posteriors, with similar uncertainties and biases. Unsurprisingly, upon evaluating the SBI method on thousands of simulated data vectors that span the parameter space, SBI exhibits worsened posterior calibration metrics in the out-of-domain application. We note that such calibration tests with MCMC is less computationally feasible and highlight the potential use of SBI to stress-test limitations of analytical models, such as in the use for constructing models for inference with MCMC.

79 ASTRONOMY AND ASTROPHYSICS↗

Artificial Intelligence for Earth System Predictability (AI4ESP) (2021 Workshop Report)

In October 2021, the U.S. Department of Energy (DOE) welcomed participants to the Artificial Intelligence for Earth System Predictability (AI4ESP) Workshop, hosted by the Office of Biological and Environmental Research (BER)—Advanced Scientific Computing Research (ASCR). The workshop is part of BER-ASCR’s ambition to more radically and aggressively advance prediction capabilities in the climate, Earth, and environmental sciences through the use of modern data analytics and artificial intelligence (AI). Advances in these capabilities are needed to improve predictions of climate change and extreme events that provide actionable information for planning and building resilience to their impacts.

54 ENVIRONMENTAL SCIENCES↗

Training data selection for event classification in a highly variable environment

A problem of interest for nuclear nonproliferation is monitoring activities at nuclear facilities, where proliferation events may only take place a few times and often under variable conditions. Machine learning has revolutionized data analytics by enabling the use of measurable signatures to generate predictive models of facility operations. However, traditional methods for training these models require large, reliable data sets with labeled observations, a challenge for nonproliferation. Highly variable conditions further complicate this as events from training data may have occurred in conditions quite different from the event of interest. Our hypothesis is that when events occur in a highly variable environment, careful training data selection for each test event could outperform the standard approach of using all available training data. We developed a method to optimize training data selection for the given test event and applied it to predicting the power level of the High Flux Isotope Reactor (HFIR) at Oak Ridge National Laboratory. In this study, the reactor startup exhibits variability between occurrences due to natural variability in environmental conditions and operational procedures. Using a combination of analysis techniques, a similitude assessment was performed on data collected from HFIR to isolate clusters that were optimal for training a predictive model. Concepts such as dynamic time warping and Jaccard similarity were used in conjunction with clustering analysis. In order to validate this approach, the model was trained on every combination of unique training events and the predictive performance was compared to the performance using a subset of the training data selected by isolated clusters found through the similitude assessment.

Iyer, A↗

Physics-Informed and Data-Driven Prediction of Residual Stress in Three-Dimensional Machining

Efficient and reliable prediction of machining-induced residual stress (RS) is a key requirement for truly integrated computational materials engineering (ICME). Currently available process modeling approaches, including empirical, analytical, and numerical methodologies lack predictive power and require substantial calibration and validation data. Moreover, most model-based approaches consider only two-dimensional (2D) (i.e., orthogonal), cutting processes. Meanwhile, industrial processes such as milling, turning, and drilling are inherently three-dimensional (3D). The present work attempts to bridge the gap between 2D and 3D through careful consideration of the process physics, including geometric, kinematic, and size-effect constraints to realize robust prediction of how RS develops in 3D machining. Using a novel in-situ experimental technique and digital image correlation (DIC) to determine equivalent Hertzian contact widths, contact pressures, and friction coefficients, the proposed methodology leverages a discretized conversion algorithm that includes multi-pass shakedown effects. This paper presents a semi-analytical model to predict machining-induced RS in 3D turning operations, which are used representatively for 3D processes more generally. Rather than follow a ‘brute force’ 3D FEM approach or conduct countless experiments to train a purely data-driven machine learning algorithm, the proposed approach builds on previous 2D modeling work. Through careful consideration of the process physics, including complex geometry/kinematic considerations of 3D turning, the authors demonstrated an experimentally calibrated approach, as well as validation based on published RS data. Model predictions and previously published measurement data of RS depth profiles for turning of Inconel 718 were compared for a range of process parameters. Correlation between the proposed 3D model and validation data was found to be within the margin of experimental error for most conditions. The proposed model appears to capture the overall behavior of 3D RS depth profiles with acceptable accuracy, particularly the key metrics of near-surface stress, peak stress magnitude and location, as well as overall stress profile depth. This report presents a physics-informed, data-driven approach for efficient calibration of a 2D model for machining-induced RS through DIC analysis of in-situ characterized subsurface displacement fields.

42 ENGINEERING↗

Bridging Equipment Reliability Data and Robust Decisions in a Plant Operation Context

In order to reduce operation and maintenance (O&M) costs, nuclear power plants (NPPs) are moving from corrective and periodic maintenance to predictive maintenance strategies. Such transition requires changes on the data that needs to be retrieved and on the type of decision processes to be employed. Advanced monitoring and data analysis technologies are essential to support predictive strategies. They can in fact provide precise information about health of a component, track its degradation trends, and provide information of its expected failure time. With such information, maintenance operations for a component can be performed right before its expected failure time. This dynamic context of O&M operations requires new methods to analyze data, propagate component health information from the component to the system level, and optimize plant resources. In this respect, the risk informed asset management (RIAM) project has been tasked to develop and test this new class of methods into a risk analytics toolset. This toolset consists of data analytics tools coupled with reliability methods designed to manage plant assets and performances in a predictive maintenance context. This report shows the latest improvements on such development and the initial testing of our methods on the three main research areas that the RIAM project is focusing on. These areas are the following: equipment reliability data analytics, system reliability modeling, and plant resources optimization methods. We show how the methods developed in these areas can support predictive maintenance strategies by: 1) analyzing equipment reliability data (either in numeric and textual form), 2) assessing component and system health through an innovative margin-based reliability approach, and 3) identifying the most critical components and set optimal maintenance schedule based on plant economic and operational constraints.

97 MATHEMATICS AND COMPUTING↗

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES↗

Benchmarking magnetised three-wave coupling for laser backscattering: analytic solutions and kinetic simulations

Understanding magnetised laser–plasma interactions is important for controlling magneto-inertial fusion experiments and developing magnetically assisted radiation and particle sources. For nanosecond pulses at non-relativistic intensities, interactions are dominated by coherent three-wave interactions, whose nonlinear coupling coefficients became known only recently when waves propagate at oblique angles with the magnetic field. In this paper, backscattering coupling coefficients predicted by warm-fluid theory are benchmarked using particle-in-cell simulations in one spatial dimension, and excellent agreements are found for a wide range of plasma temperatures, magnetic field strengths and laser propagation angles, when the interactions are mediated by electron-dominant hybrid waves. Systematic comparisons between theory and simulations are made possible by a rigorous protocol. On the theory side, the initial boundary value problem of linearised three-wave equations is solved, and the transient-time solutions allow the effects of growth and damping to be distinguished. On the simulation side, parameters are carefully chosen and calibration runs are performed to ensure that comparisons are well controlled. Fitting simulation data to analytical solutions yields numerical growth rates that match theory predictions within error bars. Although warm-fluid theory is found to be valid for a wide parameter range, genuine kinetic effects have also been observed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗

A structured framework for predicting sustainable aviation fuel properties using liquid-phase FTIR and machine learning

Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.

Fourier transform infrared spectroscopy↗