Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Traditional Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Deploying and Tracking Software with NCCS Software Provisioning

The National Center for Computational Sciences (NCCS) at Oak Ridge National Laboratory has a long history of deploying ground-breaking leadership-class supercomputers for the U.S. Department of Energy. The latest in this line of supercomputers is Frontier, the first supercomputer to break the exascale barrier (1018 floating-point operations per second) on the TOP500 list. Frontier serves a wide array of scientific domains, from traditional simulation-based workloads to newer AI and Machine Learning workloads. To best serve the NCCS user community, NCCS uses Spack to deploy a comprehensive software stack of scientific software packages, providing straightforward access to these packages through Lmod Environment Modules. Maintaining a large software stack while also including multiple new compiler releases each year is a very time-consuming task. Additionally, it is not straightforward to provide a software stack alongside existing vendor-provided software such as the HPE/Cray Programming Environment (CPE), and existing CPE, Spack, and Lmod integration does not allow for multiple versions of GPU libraries such as AMD’s ROCm to be used. To address these challenges and shortcomings, NCCS has developed the NCCS Software Provisioning tool (NSP)1, a tool for deploying and monitoring software stacks on HPC systems. NSP allows NCCS to quickly and effectively provision software stacks from the ground up using template-driven recipes and configuration files. NSP is successfully deployed on Frontier and several other NCCS clusters, enabling the NCCS software team to quickly deploy software stacks for newly-released compilers, expand current software offerings, better support GPU-based software, and monitor Lmod module usage to identify unused software packages that can be removed from the software stack. In this work, we discuss the shortcomings of the previous CPE, Spack, and Lmod usage at NCCS, provide further details on the implementation and structure of NSP, then discuss the benefits that NSP provides.

Rentschler, Asa [ORNL] (ORCID:0009000597694743)↗

Tensor Text-Mining Methods for Malware Identification and Detection, Malware Dynamics Characterization, and Hosts Ranking

Malware is one of the most persistent and costly cyber threats endangering reputation, confidentiality, integrity, and availability for organizations and national security. Consequently, many of the incident detection and prevention systems, and incident responders have begun to utilize machine learning as a helper in the fight against malware and other cyber threats. However, cyber defenders rely on interpretability and generalizability, yet the popular machine learning methods are black-box and often use traditional supervised solutions that do not generalize to novel malware. Therefore, there is a need to improve the existing solutions. At the same time, the majority of the prior research ignored essential evaluation criteria when reporting the results of their methods, which disables the safe reproducibility of the methods in a production environment. Tensor decomposition, on the other hand, enables interpretable unsupervised analysis of the large-scale data for the discovery of hidden patterns. Our findings, performed on real-world and large-scale experiments, show that tensor factorization-based methods yield performance results that surpasses or competes with existing supervised solutions with the added benefit of interpretability and generalizability. With the ability to analyze complex and large-scale data using tensors, we report results that reflect real-world production environments. We propose to develop new game- changing tools for malware identification and characterization that can trace malware evolution, rank the infected or malicious hosts, and streamline the work of incident response teams, malware analysts, and incident detection and prevention systems.

97 MATHEMATICS AND COMPUTING↗

Understanding Twinning and Deformation in High Entropy Alloys

On the one hand, multi-principal element alloys (MPEAs) have created a paradigm shift in alloy design due to large compositional space, whereas on the other, they have presented enormous computational challenges for theory-based materials design, especially density functional theory (DFT), which is inherently computationally expensive even for traditional dilute alloys. In this project, we developed a machine learning framework, namely PREDICT ( PR edict properties from E xisting D atabase I n C omplex alloys T erritory), that opens a pathway to predict elastic constants in large compositional space with little computational expense. The framework only relies on the DFT database of binary alloys and predicts Voigt–Reuss–Hill Young’s modulus, shear modulus, bulk modulus, elastic constants, and Poisson’s ratio in MPEAs. We show that the key descriptors of elastic constants are the A–B bond length and cohesive energy. The framework can predict elastic constants in hypothetical compositions as long as the constituent elements are present in the database, thereby enabling property exploration in multi-compositional systems. We illustrate predictions in a FCC Ni-Cu-Au-Pd-Pt system.

36 MATERIALS SCIENCE↗

Preliminary Target Selection for the DESI Quasar (QSO) Sample

The DESI survey will measure large-scale structure using quasars as direct tracers of dark matter in the redshift range 0.9 < z < 2.1 and using quasar Lyα forests at z > 2.1. In this work, we present two methods to select candidate quasars for DESI based on imaging in three optical (g, r, z) and two infrared (W1, W2) bands. The first method uses traditional color cuts and the second utilizes a machine-learning algorithm.

79 ASTRONOMY AND ASTROPHYSICS↗

Learning in Artificial Neural Systems

This paper presents an overview and analysis of learning in Artificial Neural Systems (ANS's). It begins with a general introduction to neural networks and connectionist approaches to information processing. The basis for learning in ANS's is then described, and compared with classical Machine learning. While similar in some ways, ANS learning deviates from tradition in its dependence on the modification of individual weights to bring about changes in a knowledge representation distributed across connections in a network. This unique form of learning is analyzed from two aspects: the selection of an appropriate network architecture for representing the problem, and the choice of a suitable learning rule capable of reproducing the desired function within the given network. The various network architectures are classified, and then identified with explicit restrictions on the types of functions they are capable of representing. The learning rules, i.e., algorithms that specify how the network weights are modified, are similarly taxonomized, and where possible, the limitations inherent to specific classes of rules are outlined.

Matheus, Christopher J.↗

Natural Language Understanding and Extraction of Flight Constraints Recorded in Letters of Agreement

This paper presents an automated information extraction and inference technique using natural language processing for extracting flight operational procedures and constraints embedded in heritage air traffic management documents. The extracted flight constraints can be digitized and fit into existing airspace information exchange models such as the Aeronautical Information Exchange Model (AIXM). This approach offers a digitized solution to disseminate airspace operating conditions to diverse air users and stakeholders in the National Airspace System (NAS). Furthermore, the digitized flight procedures can provide operational flexibility for emerging advanced air mobility providers and reduce traffic controller workload while maintaining current safety standards. To demonstrate this process, 1,972 Letters of Agreement (LOAs) have been selected for processing, named entity extraction, constraint identification and extraction. This dataset is derived from a subset of documents related to Air Route Traffic Control Centers (ARTCC) operations. We experimented with various traditional information extraction techniques, state-of-the-art machine learning and deep learning models to perform named entity recognition and pattern recognition on our dataset. We present the results from our experiments and demonstrate 99.0% F-1 score for named entity recognition, and a 96.6% accuracy for our entire workflow up to named entity recognition. We also discuss constraint definitions using generic patterned templates and extensions to this work in applying entity linking to digitally extracting relevant constraints.

Natural Language Processing↗

Natural Language Understanding and Extraction of Flight Constraints Recorded in Letters of Agreement

This paper presents an automated information extraction and inference technique using natural language processing for extracting flight operational procedures and constraints embedded in heritage air traffic management documents. The extracted flight constraints can be digitized and fit into existing airspace information exchange models such as the Aeronautical Information Exchange Model (AIXM). This approach offers a digitized solution to disseminate airspace operating conditions to diverse air users and stakeholders in the National Airspace System (NAS). Furthermore, the digitized flight procedures can provide operational flexibility for emerging advanced air mobility providers and reduce traffic controller workload while maintaining current safety standards. To demonstrate this process, 1,972 Letters of Agreement (LOAs) have been selected for processing, named entity extraction, constraint identification and extraction. This dataset is derived from a subset of documents related to Air Route Traffic Control Centers (ARTCC) operations. We experimented with various traditional information extraction techniques, state-of-the-art machine learning and deep learning models to perform named entity recognition and pattern recognition on our dataset. We present the results from our experiments and demonstrate 99.0% F-1 score for named entity recognition, and a 96.6% accuracy for our entire workflow up to named entity recognition. We also discuss constraint definitions using generic patterned templates and extensions to this work in applying entity linking to digitally extracting relevant constraints.

Natural Language Processing↗

Dense Feature Tracking of Atmospheric Winds with Deep Optical Flow

Atmospheric winds are a key physical phenomenon impacting natural hazards, energy transport, ocean currents, large-scale circulation, and ecosystem fluxes. Observing winds is a complex process and presents a large gap in NASA’s Earth Observation System. Atmospheric motion vectors (AMVs) aim to fill this gap by making numerical estimates of cloud movement between sequences of multi-spectral satellite images, tracking clouds and water vapor. Recent imaging hardware and software advancements have enabled the use of numerical optical flow techniques to produce accurate and dense vector fields outperforming traditional methods. This work presents WindFlow as the first machine learning based system for feature tracking atmospheric motion using optical flow. Due to the lack of large-scale satellite-based observations, we leverage high-resolution numerical simulations from NASA's GEOS-5 Nature Run to perform supervised learning and transfer to satellite images. We demonstrate that our approach using deep learning based optical flow scales to ultra-high-resolution images of size 2881x5760 with less than 1 m/s bias and 2.5 m/s average error. Four network and learning architectures are compared and it is found that recurrent all-pairs field transforms (RAFT) produces the lowest errors on all metrics for wind speed and direction. Results on held out numerical outputs shows RAFT's good performance in each of the spatial, temporal, and physical dimensions. A comparison between WindFlow and an operational AMV product against rawinsonde observations show that RAFT transfers across simulations and thermal infrared satellite observations. This work shows that machine learning based optical flow is an efficient approach to generating robust feature tracking for AMVs consistently over large regions.

Atmospheric winds↗

Topological Regularization via Persistence-Sensitive Optimization

Optimization, a key tool in machine learning and statistics, relies on regularization to reduce overfitting. Traditional regularization methods control a norm of the solution to ensure its smoothness. Recently, topological methods have emerged as a way to provide a more precise and expressive control over the solution, relying on persistent homology to quantify and reduce its roughness. All such existing techniques back-propagate gradients through the persistence diagram, which is a summary of the topological features of a function. Their downside is that they provide information only at the critical points of the function. We propose a method that instead builds on persistence-sensitive simplification and translates the required changes to the persistence diagram into changes on large subsets of the domain, including both critical and regular points. This approach enables a faster and more precise topological regularization, the benefits of which we illustrate with experimental evidence.

Nigmetov, Arnur↗

Machine learning-assisted profiling of a kinked ladder polymer structure using scattering

Ladder polymers consisting of fused rings in the backbone have very limited conformational freedom, which results in very different properties from traditional linear polymers. However, accurately determining their size and chain conformations from solution scattering remains a challenge. Their chain conformations of kinked ladder polymers are largely governed by the structures and relative orientations or configurations of the repeat units, unlike conventional polymer chains whose bending angles between repeat units follow a unimodal Gaussian distribution. Meanwhile, traditional scattering models for polymer chains do not account for these unique structural features. This work introduces a novel approach that integrates machine learning with Monte Carlo simulations to construct a model that can describe the geometry of a type of kinked CANAL ladder polymers. We first develop a Monte Carlo simulation model for sampling the configuration space of CANAL ladder polymers, where each repeat unit is modeled as a biaxial segment. Then, we establish a machine learning-assisted scattering analysis framework based on Gaussian Process Regression. Finally, we conduct small-angle neutron scattering experiments on a CANAL ladder polymer solution to apply our approach. Our method uncovers structural features of such ladder polymers that conventional methods fail to capture.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),↗

Discovery of unconventional and nonintuitive self-assembling peptide materials using experiment-driven machine learning

Prediction of peptide secondary structure is challenging because of complex molecular interactions, sequence-specific behavior, and environmental factors. Traditional design strategies, based on hydrophobicity and structural propensity, can be biased and could indeed prevent discovery of interesting, diverse, and unconventional peptides with desired nanostructure assembly. Using β sheet formation in pentapeptides as a case study, we used an integrated high-throughput experimental workflow and an artificial intelligence–driven active learning framework to improve prediction accuracy of self-assembly. By focusing on sequences where machine learning (ML) predictions deviate from conventional design strategies, we synthesized and tested 268 pentapeptides, successfully finding 96 forming β sheet assemblies, including unconventional sequences (e.g., ILFSM, LMISI, MITIY, MISIW, and WKIYI) not predicted by traditional methods. Our ML models outperformed conventional β sheet propensity tables, revealing useful chemical design rules. A web interface is provided to facilitate community access to these models. This work highlights the value of ML-driven approaches in overcoming the limitations of current peptide design strategies.

Talluri, Y. Nissi [Indian Inst. of Technology (IIT↗

Machine learning based rate optimization under geologic uncertainty

We propose a novel approach for rate optimization during a waterflood under geologic uncertainty in reservoir properties such as permeability and porosity. The traditional approach typically involves several runs of the forward simulator. This may not scale well when the optimization is to be performed at the full field-level and over multiple geologic realizations. A machine-learning (ML) based approach which is quick and scalable for rate optimization over multiple geologic realizations is proposed instead. The training data for the model is generated by running the forward simulator with randomly assigned well rates using multiple geologic realizations. A reduced order representation of the permeability heterogeneity in each of the realizations is derived using a grid connectivity transformation (GCT). This step involves finding basis functions corresponding to the different modal frequencies of the grid connectivity represented by the grid Laplacian. The projection of the heterogeneous property field along these basis functions gives the basis coefficients that form the reduced order representation. Subsequently, for each training datapoint, streamlines are traced and the minimum time of flight (TOF) representing the tracer breakthrough time at each producer is recorded. The basis coefficients and well rates are fed to a machine learning model as input and the minimum TOF at the producers forms the output of the model. This trained model can then be used along with an optimizer for computing the optimal injection rates to maximize the injection sweep efficiency. This corresponds to minimizing the variance in the minimum TOF within each well group. Different architectures of neural network are tested using 5-fold cross validation to decide the best ML model to compute the streamline time of flight. The trained model is used to perform well rate optimization over multiple realizations of geology by using a risk tolerance penalty. The optimal well rates thus obtained are compared with two cases: a) equal well rates assigned to all injectors and producers and b) well rates obtained by optimizing over a single realization without considering the uncertainty in geology. The optimal well rates are seen to offer better oil recovery and sweep efficiency than both cases.

02 PETROLEUM↗

Machine Learning for Multipactor Susceptibility Prediction in Planar RF Gaps

Multipactor discharge is a nonlinear electron avalanche that limits the performance of high-power radio-frequency (RF) and vacuum electronic devices. Predicting multipactor susceptibility traditionally relies on Monte Carlo or particle-in-cell (PIC) simulations, which become computationally expensive for large parametric studies. In this work, we present a supervised machine-learning (ML) framework for prediction of multipactor susceptibility in a two-surface planar geometry. The models are trained using high-fidelity PIC simulation generated susceptibility data and learn the relationship between operational parameters, geometry, and material-dependent secondary electron emission properties. The proposed approach enables rapid reconstruction of susceptibility charts while preserving the physical structure of multipactor growth regions.

43 PARTICLE ACCELERATORS↗

Autonomous Output‐Oriented Aerosol Jet Printing Enabled by Hybrid Machine Learning

Additive manufacturing (AM) is rapidly revolutionizing modern manufacturing with recent progress in advanced printing methods and improved properties of printed materials. However, traditional AM methods are limited by their input‐oriented nature, which demands tedious trial‐and‐error tuning of printing parameters to achieve desired output properties. Here, in this work, an output‐oriented artificial intelligence‐integrated AM (AIAM) method is reported that enables an user to specify desired output properties while the printer autonomously discovers the optimal input printing parameters by integrating hybrid machine learning models and in situ measurements. Based on a predictive mapping between the input printing parameters and the output properties of interests established with <20 experiments designed by active learning, inverse design tasks are performed to intelligently generate the printing parameter settings that lead to desired outcomes using reinforcement learning. This method is demonstrated by autonomous aerosol jet printing (AJP) of conductive polymer films and achieving user‐defined electrical resistances with an ultralow error of 3.7%. The AIAM method, with its output‐oriented nature, holds the potential to significantly improve the autonomy, predictability, efficiency, and accessibility of the AM processes, which will unlock new possibilities in the autonomous and intelligent printing of a broad range of functional materials and devices.

36 MATERIALS SCIENCE↗

A Comparative Study of Machine Learning Algorithms for Industry-Specific Freight Generation Model

According to Bureau of Transportation Statistics, the U.S. transportation system handled 14,329 million ton-miles of freight per day in 2020. Understanding the generation of these freight shipments is crucial for transportation researchers, planners, and policymakers to design and plan for a more efficient and connected freight transportation system. Traditionally, the freight generation modeling has been based on Ordinary Least Square (OLS) regression, although more advanced Machine Learning (ML) algorithms have been evaluated and proven to have excellent performance in various transportation applications in recent years. Furthermore, one modeling approach applied for one industry might not always be applicable for another as their freight generation logics can be quite different. The objective of this study is to apply and evaluate alternative ML algorithms in the estimation of freight generation for each of 45 industry types. Seven alternative ML algorithms, along with the base OLS regression, were evaluated and compared. In addition, the study considered different combinations of variables in both the original and logarithmic form as well as hyperparameters of those ML algorithms in the model selection for each industry type. The results showed statistically significant improvements in the root mean square error reduction by the alternative ML algorithms over the OLS for over 80% of cases. The study suggests utilizing the alternative ML algorithms can reduce the root mean square error by about 30%, depending on industry types.

97 MATHEMATICS AND COMPUTING↗

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML↗

Capturing the Physics of MaNGA Galaxies with Self-supervised Machine Learning

As available data sets grow in size and complexity, advanced visualization tools enabling their exploration and analysis become more important. In modern astronomy, integral field spectroscopic galaxy surveys are a clear example of increasing high dimensionality and complex data sets, which challenges the traditional methods used to extract the physical information they contain. Here, we present the use of a novel self-supervised machine-learning method to visualize the multidimensional information on stellar population and kinematics in the MaNGA survey in a 2D plane. Our framework is insensitive to nonphysical properties such as the size of the integral field unit and is therefore able to order galaxies according to their resolved physical properties. Using the extracted representations, we study how galaxies distribute based on their resolved and global physical properties. We show that even when exclusively using information about the internal structure, galaxies naturally cluster into two well-known categories, rotating main-sequence disks and massive slow rotators, from a purely data-driven perspective, hence confirming distinct assembly channels. Low-mass rotation-dominated quenched galaxies appear as a third cluster only if information about the integrated physical properties is preserved, suggesting a mixture of assembly processes for these galaxies without any particular signature in their internal kinematics that distinguishes them from the two main groups. The framework for data exploration is publicly released with this publication, ready to be used with the MaNGA or other integral field data sets.

79 ASTRONOMY AND ASTROPHYSICS↗