Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Comparison of Probabilistic Forecasts for Predictive Voltage Control

This paper explores predictive cooperative voltage control in distribution systems with highly variable sources such as photovoltaics (PV). The goal is to maintain the voltage profile within the limits despite the fluctuations due to sudden changes in solar power generation. The predictive voltage control method relies on probabilistic solar power and load forecasts to select the optimum Voltage Regulator (VR) taps appropriately. VR taps are selected to minimize the risk of voltage violation. A modified version of the IEEE 123 system is used as the case study. A 100% penetration of solar power is assumed for the distribution system with profiles for solar generation and loads added to the system. Three different probabilistic forecast models (Quantile Regression (QR), Gaussian distribution and volatility forecasting using Generalized Autoregressive Conditional Heteroskedasticity (GARCH)) are explored in this study. The results for the VR taps and Voltage Deviation Index (VDI) are compared to find the most effective forecast model.

Panamtash, Hossein↗

Data‐Driven Safety Risk Prediction of Lithium‐Ion Battery

Abstract Inevitable safety issues have pushed battery engineers to become more conservative in battery system design; however, battery‐involved accidents still frequently are reported in headlines. Identifying, understanding, and predicting safety risks have become priorities to further accelerate technology and industry development. However, diverse loading scenarios, significantly varied stress‐induced short circuit mechanisms, and highly coupled mechanical–electrochemical safety behaviors have remained grand challenges. Herein, the safety risk is termed as the probability of the mechanical triggering of an internal short circuit, to reflect the safety related behaviors of lithium‐ion batteries. Based on a mechanical model and experimental results, a sufficient dataset is generated consisting of strain states and their corresponding safety risks, covering both cylindrical and pouch cells, various states of charges, and loading conditions. Machine‐learning tools combined with the established finite element mechanical model are applied to predict the safety risks of the cells. The results achieve a high level of accuracy on the test data (the relative error of the average short circuit prediction deviation is less than 6.2%.). This work underpins the safety risk concept and highlights the promise of physics combined with data‐driven modeling methodology to predict the safety behaviors of energy storage systems.

Jia, Yikai↗

Physics-Informed Machine Learning Models for Predicting the Progress of Reactive-Mixing

This paper presents a physics-informed machine learning (ML) framework to construct reduced-order models (ROMs) for reactive-transport quantities of interest (QoIs) based on high-fidelity numerical simu-lations. QoIs include species decay, product yield, and degree of mixing. The ROMs for QoIs are applied to quantify and understand how the chemical species evolve over time. First, high-resolution datasets for constructing ROMs are generated by solving anisotropic reaction-di?usion equations using a non-negative finite element formulation for di?erent input parameters. The reactive-mixing model input parameters are: time-scale associated with flipping of velocity, spatial-scale controlling small/large vortex structures of velocity, perturbation parameter of the vortex-based velocity, anisotropic dispersion strength/contrast, and molecular diffusion. Second, random forests, F-test, and mutual information criterion are used to evaluate the importance of model inputs/features with respect to QoIs. We observed that anisotropic dispersion strength/contrast is the most important feature and time-scale associated with flipping of velocity is the least important feature. Third, Support Vector Machines (SVM) and Support Vector Regression (SVR) are used to construct ROMs based on the model inputs. The constructed SVR-ROMs are then used to predict scaling of QoIs. We also present estimates and inequalities on the QoIs, which inform that the species decay, mix, and produce in an exponential fashion. These inequalities also inform that a radial basis function is the most suitable kernel for the SVM/SVR models for QoIs. It is observed that R2-score for SVR-ROMs on unseen data is greater than 0.9, implying that the SVR-ROMs are able to predict the reaction-diffusion system state reasonably well. Finally, in terms of the computational cost, the proposed SVM-ROMs are O(107) times faster than running a high-fidelity finite element simulation for evaluating QoIs. This makes the proposed ML-based ROMs attractive for reactive-transport sensing and real-time monitoring applications as they are significantly faster yet reasonably accurate.

Mudunuru, Maruti K.↗

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

HydraGNN_OPF_GFM_2026 - Ensemble of predictive graph foundation models for power grid applications

This dataset supports research on graph foundation models for optimal power flow (OPF) on electric grids using HydraGNN. It contains heterogeneous graph representations of PGLib-OPF cases spanning systems from 14 to 13,659 buses, together with packed HDF5 datasets for pretraining, feasibility classification, and N-1 contingency analysis. The release includes OPF solution data, downstream fine-tuning datasets, pretrained HeteroSAGE and HeteroHEAT model checkpoints, hyperparameter-optimization summaries across multiple heterogeneous GNN architectures, and aggregated fine-tuning results for sample-efficiency studies. The dataset is designed to enable scalable training, evaluation, and transfer-learning studies for OPF surrogate modeling, including node-level AC-OPF solution prediction, graph-level prediction, feasibility classification, operating-condition generalization, and contingency-response tasks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗

AI-Enabled Operations at Fermi Complex: Multivariate Time Series Prediction for Outage Prediction and Diagnosis

The Main Control Room of the Fermilab accelerator complex continuously gathers extensive time-series data from thousands of sensors monitoring the beam. However, unplanned events such as trips or voltage fluctuations often result in beam outages, causing operational downtime. This downtime not only consumes operator effort in diagnosing and addressing the issue but also leads to unnecessary energy consumption by idle machines awaiting beam restoration. The current threshold-based alarm system is reactive and faces challenges including frequent false alarms and inconsistent outage-cause labeling. To address these limitations, we propose an AI-enabled framework that leverages predictive analytics and automated labeling. Using data from $2,703$ Linac devices and $80$ operator-labeled outages, we evaluate state-of-the-art deep learning architectures, including recurrent, attention-based, and linear models, for beam outage prediction. Additionally, we assess a Random Forest-based labeling system for providing consistent, confidence-scored outage annotations. Our findings highlight the strengths and weaknesses of these architectures for beam outage prediction and identify critical gaps that must be addressed to fully harness AI for transitioning downtime handling from reactive to predictive, ultimately reducing downtime and improving decision-making in accelerator management.

Jain, Milan [PNL, Richland] (ORCID:000000021676111↗

A probabilistic inverse prediction method for predicting plutonium processing conditions

In the past decade, nuclear chemists and physicists have been conducting studies to investigate the signatures associated with the production of special nuclear material (SNM). In particular, these studies aim to determine how various processing parameters impact the physical, chemical, and morphological properties of the resulting special nuclear material. By better understanding how these properties relate to the processing parameters, scientists can better contribute to nuclear forensics investigations by quantifying their results and ultimately shortening the forensic timeline. This paper aims to statistically analyze and quantify the relationships that exist between the processing conditions used in these experiments and the various properties of the nuclear end-product by invoking inverse methods. In particular, these methods make use of Bayesian Adaptive Spline Surface models in conjunction with Bayesian model calibration techniques to probabilistically determine processing conditions as an inverse function of morphological characteristics. Not only does the model presented in this paper allow for providing point estimates of a sample of special nuclear material, but it also incorporates uncertainty into these predictions. This model proves sufficient for predicting processing conditions within a standard deviation of the observed processing conditions, on average, provides a solid foundation for future work in predicting processing conditions of particles of special nuclear material using only their observed morphological characteristics, and is generalizable to the field of chemometrics for applicability across different materials.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Conotoxin Prediction: New Features to Increase Prediction Accuracy

Conotoxins are toxic, disulfide-bond-rich peptides from cone snail venom that target a wide range of receptors and ion channels with multiple pathophysiological effects. Conotoxins have extraordinary potential for medical therapeutics that include cancer, microbial infections, epilepsy, autoimmune diseases, neurological conditions, and cardiovascular disorders. Despite the potential for these compounds in novel therapeutic treatment development, the process of identifying and characterizing the toxicities of conotoxins is difficult, costly, and time-consuming. This challenge requires a series of diverse, complex, and labor-intensive biological, toxicological, and analytical techniques for effective characterization. While recent attempts, using machine learning based solely on primary amino acid sequences to predict biological toxins (e.g., conotoxins and animal venoms), have improved toxin identification, these methods are limited due to peptide conformational flexibility and the high frequency of cysteines present in toxin sequences. This results in an enumerable set of disulfide-bridged foldamers with different conformations of the same primary amino acid sequence that affect function and toxicity levels. Consequently, a given peptide may be toxic when its cysteine residues form a particular disulfide-bond pattern, while alternative bonding patterns (isoforms) or its reduced form (free cysteines with no disulfide bridges) may have little or no toxicological effects. Similarly, the same disulfide-bond pattern may be possible for other peptide sequences and result in different conformations that all exhibit varying toxicities to the same receptor or to different receptors. We present here new features, when combined with primary sequence features to train machine learning algorithms to predict conotoxins, that significantly increase prediction accuracy.

collisional cross section↗

Improving Prediction of Surface Solar Irradiance Variability by Integrating Observed Cloud Characteristics and Machine Learning

A 5-year, 1-minute resolution observational dataset of clouds and solar radiation was produced that includes two metrics of the variability in surface solar irradiance due to cloud type and fractional sky cover. Multiple regression models were trained to fit observations of surface solar irradiance variability from those two cloud property predictors. We found that ensemble tree-based methods, Random Forest and Gradient Boosting Machine, have the least overfitting issues and showed the best performance with an R2 of 0.42. While the observational data trained in this study was only from one site, the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) site in Oklahoma, initial comparisons of the seasonality of the statistics suggest that these results are relatively weather regime independent; the generality of such a finding across sites will be tested in future work. The observational data and developed machine learning model are being used to create a numerical weather prediction model parameterization to enable day-ahead solar variability prediction in a computationally efficient way. This is a first step towards creating a new paradigm of predicting day-ahead variability with the potential to provide a new tool to improve grid operation, planning, and resilience.

Riihimaki, Laura↗

Machine Learning Self-Diffusion Prediction for Lennard-Jones Fluids in Pores

Predicting the diffusion coefficient of fluids under nanoconfinement is important for many applications including the extraction of shale gas from kerogen and product turnover in porous catalysts. Due to the large number of important variables, including pore shape and size, fluid temperature and density, and the fluid–wall interaction strength, simulating diffusion coefficients using molecular dynamics (MD) in a systematic study could prove to be prohibitively expensive. Here, we use machine learning models trained on a subset of MD data to predict the self-diffusion coefficients of Lennard-Jones fluids in pores. Our MD data set contains 2280 simulations of ideal slit pore, cylindrical pore, and hexagonal pore geometries. We use the forward feature selection method to determine the most useful features (i.e., descriptors) for developing an artificial neutral network (ANN) model with an emphasis on easily acquired features. Our model shows good predictive ability with a coefficient of determination (i.e., R 2 ) of ~0.99 and a mean squared error of ~2.9 × 10 –5 . Finally, we propose an alteration to our feature set that will allow the ANN model to be applied to nonideal pore geometries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhanced predictability of Eastern North Pacific Tropical cyclone activity using the ENSO Longitude Index

Past studies have indicated that El Nino-Southern Oscillation (ENSO) plays a major role in the interannual variability of Eastern Pacific hurricane activity. The primary mechanism being the eastward displacement of the warm pool during an El Nino, which carries warm water into that basin thereby creating favorable oceanic conditions. Despite this, the question of whether an accurate knowledge of ENSO enhances seasonal predictabiity of Eastern Pacific hurricanes has not been addressed specifically. In this study, we show that unlike traditional indices of ENSO, the ENSO Longitude Index (ELI) is able to predict Eastern Pacific hurricane activity at significant lead times. By capturing changes in the location of deep convection and associated thermocline processes more accurately, ELI explains the most variability in the upper-ocean heat content in the main development region of the Eastern Pacific basin compared to other ENSO indices. These results have substantial implications for operational seasonal forecasts of Eastern Pacific hurricanes.

Balaguru, Karthik↗

Predicting images for the dynamics of stellar clusters ( π-DOC ): a deep learning framework to predict mass, distance, and age of globular clusters

ABSTRACT Dynamical mass estimates of simple systems such as globular clusters (GCs) still suffer from up to a factor of 2 uncertainty. This is primarily due to the oversimplifications of standard dynamical models that often neglect the effects of the long-term evolution of GCs. Here, we introduce a new approach to measure the dynamical properties of GCs, based on the combination of a deep-learning framework and the large amount of data from direct N-body simulations. Our algorithm, π-DOC (Predicting Images for the Dynamics Of stellar Clusters) is composed of two convolutional networks, trained to learn the non-trivial transformation between an observed GC luminosity map and its associated mass distribution, age, and distance. The training set is made of V-band luminosity and mass maps constructed as mock observations from N-body simulations. The tests on π-DOC demonstrate that we can predict the mass distribution with a mean error per pixel of 27 per cent, and the age and distance with an accuracy of 1.5 Gyr and 6 kpc, respectively. In turn, we recover the shape of the mass-to-light profile and its global value with a mean error of 12 per cent, which implies that we efficiently trace mass segregation. A preliminary comparison with observations indicates that our algorithm is able to predict the dynamical properties of GCs within the limits of the training set. These encouraging results demonstrate that our deep-learning framework and its forward modelling approach can offer a rapid and adaptable tool competitive with standard dynamical models.

Chardin, Jonathan↗