Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Machine learning assisted modeling of mixing timescale for LES/PDF of high-Karlovitz turbulent premixed combustion

Accurate modeling of mixing in the transported probability density function (PDF) method remains a great challenge, especially for turbulent premixed combustion under extreme conditions such as high Karlovitz number Ka. Recently, a power-law based mixing timescale model was developed for the large-eddy simulations (LES)/PDF modeling of high-Ka number turbulent premixed flames. It is found in this work that the power-law mixing timescale model is highly sensitive to the model parameters. It is thus critically needed to develop accurate calibration of these model parameters. The empirical specification of the model parameters developed in Zhang et. al. is found to be inadequate for accurate modeling of the mixing timescale. Here, machine learning is introduced as an attractive alternative in this work for the specification of the model parameters. A high-Ka number DNS jet flame is used as the training and validation of the machine learning models. The choices of the input parameters are discussed and compared for the machine learning models. The effect of differential molecular diffusion on mixing is examined by including the effect of the Lewis number in the training of the machine learning models. The performance of different machine learning algorithms is compared for the specification of the mixing model parameters. Overall, excellent performance of the machine learning models is observed for assisting the mixing modeling. The feasibility, interpretability, applicability, generality, and portability of using machine learning are discussed in general to provide a perspective on applying data-driven machine learning for turbulent combustion modeling studies.

42 ENGINEERING↗

Scaling Inference Using Triton to Accelerate Particle Physics at the LHC and DUNE

DUNE and the LHC experiments consist of unique and cutting-edge particle detectors that create massive, complex, and rich datasets with billions of events. They require sophisticated algorithms to reconstruct and interpret the data. Modern machine learning algorithms provide a powerful toolset to detect and classify particles, from familiar image processing convolutional neural networks to newer graph neural network architectures. A full reconstruction of these particle collisions requires novel approaches to handle the computing challenge of processing so much raw data. In a series of studies, physicists from Fermilab, CERN, and university groups explored how to accelerate their data processing using the Triton Inference Server.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Combining mechanistic and machine learning models for predictive engineering and optimization of tryptophan metabolism

Abstract Through advanced mechanistic modeling and the generation of large high-quality datasets, machine learning is becoming an integral part of understanding and engineering living systems. Here we show that mechanistic and machine learning models can be combined to enable accurate genotype-to-phenotype predictions. We use a genome-scale model to pinpoint engineering targets, efficient library construction of metabolic pathway designs, and high-throughput biosensor-enabled screening for training diverse machine learning algorithms. From a single data-generation cycle, this enables successful forward engineering of complex aromatic amino acid metabolism in yeast, with the best machine learning-guided design recommendations improving tryptophan titer and productivity by up to 74 and 43%, respectively, compared to the best designs used for algorithm training. Thus, this study highlights the power of combining mechanistic and machine learning models to effectively direct metabolic engineering efforts.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient graph representation framework for chemical molecule similarity tasks

Graph data has emerged in numerous scientific domains and machine learning techniques have been widely used for analysis and learning of diverse data for prediction and decision. Machine learning techniques can readily address complex problems by leveraging their structural information. But graphs cannot be directly used for existing machine learning algorithms unless encoded as vectors. The problem of efficient representation of graphs is a substantial challenge in graph machine learning. In this paper, we propose a novel two-stage framework for the representation of chemical molecule graphs based on the strengths of Graph Isomorphism Networks (GINs) and Siamese autoencoders. In the first stage, the GIN model is constructed and trained using the structural information of chemical molecule graphs. Node attributes, edge attributes, and edge indices are used as input data, while graph attributes are used as labels. The GIN model effectively captures the structural characteristics of graphs and can accurately predict graph attributes, i.e., molecular properties. It also generates Graph Embeddings, represented as vectors that encode the structural information of graphs. In the second stage, Graph Embedding vectors are further optimized for downstream similarity tasks while preserving the graph structural information. The Siamese autoencoder is constructed and trained, which reduces the dimensionality of the Graph Embedding vectors, while maximizing the preservation of structural information in the original high-dimensional vectors. The resulting low-dimensional Graph Embeddings can be effectively utilized for tasks such as approximate nearest neighbor search. The experimental results demonstrate the effectiveness of our proposed framework in accurately predicting graph similarity.

Ma, Jiaji↗

System and method for error detection and correction in virtual reality and augmented reality environments

Embodiments of the present disclosure are related to training one or more of machine learning algorithms in a virtual reality environment for error detection and correction and/or for employing one or more trained machine learning models in an augmented reality environment to detect and/or correct user errors associated the performance of one or more tasks.

97 MATHEMATICS AND COMPUTING↗

Active Learning for Metamaterial Optimization on HPC and QC Integrated Systems

Active learning algorithms, integrating machine learning, quantum computing and optics simulation in an iterative loop, offer a promising approach to optimizing metamaterials. However, these algorithms can face difficulties in optimizing highly complex structures due to computational limitations. High-performance computing (HPC) and quantum computing (QC) integrated systems can address these issues by enabling parallel computing. In this study, we develop an active learning algorithm working on HPC-QC integrated systems. We evaluate the performance of optimization processes within active learning (i.e., training a machine learning model, problem-solving with quantum computing, and evaluating optical properties through wave-optics simulation) for highly complex metamaterial cases. Our results showcase that utilizing multiple cores on the integrated system can significantly reduce computational time, thereby enhancing the efficiency of optimization processes. Therefore, we expect that leveraging HPC-QC integrated systems helps effectively tackle large-scale optimization challenges in general.

Kim, Seongmin↗

Identifications and classifications of human locomotion using Rayleigh-enhanced distributed fiber acoustic sensors with deep neural networks

Abstract This paper reports on the use of machine learning to delineate data harnessed by fiber-optic distributed acoustic sensors (DAS) using fiber with enhanced Rayleigh backscattering to recognize vibration events induced by human locomotion. The DAS used in this work is based on homodyne phase-sensitive optical time-domain reflectometry (φ-OTDR). The signal-to-noise ratio (SNR) of the DAS was enhanced using femtosecond laser-induced artificial Rayleigh scattering centers in single-mode fiber cores. Both supervised and unsupervised machine-learning algorithms were explored to identify people and specific events that produce acoustic signals. Using convolutional deep neural networks, the supervised machine learning scheme achieved over 76.25% accuracy in recognizing human identities. Conversely, the unsupervised machine learning scheme achieved over 77.65% accuracy in recognizing events and human identities through acoustic signals. Through integrated efforts on both sensor device innovation and machine learning data analytics, this paper shows that the DAS technique can be an effective security technology to detect and to identify highly similar acoustic events with high spatial resolution and high accuracies.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

IoT Intrusion Detection Taxonomy, Reference Architecture, and Analyses

This paper surveys the deep learning (DL) approaches for intrusion-detection systems (IDSs) in Internet of Things (IoT) and the associated datasets toward identifying gaps, weaknesses, and a neutral reference architecture. A comparative study of IDSs is provided, with a review of anomaly-based IDSs on DL approaches, which include supervised, unsupervised, and hybrid methods. All techniques in these three categories have essentially been used in IoT environments. To date, only a few have been used in the anomaly-based IDS for IoT. For each of these anomaly-based IDSs, the implementation of the four categories of feature(s) extraction, classification, prediction, and regression were evaluated. We studied important performance metrics and benchmark detection rates, including the requisite efficiency of the various methods. Four machine learning algorithms were evaluated for classification purposes: Logistic Regression (LR), Support Vector Machine (SVM), Decision Tree (DT), and an Artificial Neural Network (ANN). Therefore, we compared each via the Receiver Operating Characteristic (ROC) curve. The study model exhibits promising outcomes for all classes of attacks. The scope of our analysis examines attacks targeting the IoT ecosystem using empirically based, simulation-generated datasets (namely the Bot-IoT and the IoTID20 datasets).

97 MATHEMATICS AND COMPUTING↗

A machine learning-based fast frequency response control for a VSC-HVDC system

An HVDC system can realize a very fast frequency response to the disturbed system under a contingency because its active power control is decoupled from the frequency deviation. However, most of existing HVDC frequency control strategies are coupled with system primary frequency control and secondary frequency control. Since the traditional system frequency control is dominated by the thermal generators, the advantage of the fast response of the HVDC system is not made fully used. The development of a frequency response estimation based on a machine learning algorithm provides another approach to improve the frequency response capability of the HVDC system. Different from other frequency deviation tracking strategies, a machine learning based HVDC frequency response control can directly increase the power flow of a HVDC system by estimation of the system generator or load lost. In this paper, a fast frequency response control using a HVDC system for a large power system disturbance based on the multivariate random forest regression (MRFR) algorithm is proposed. The simulation is carried out with an integrated power system model based on the North American interconnections. The simulation results indicate that the proposed MRFR based frequency response control can significantly improve the frequency low point during an event, while stabilizing the frequency in advance.

42 ENGINEERING↗

Simultaneous prediction of structural properties in epitaxially–grown GaN with quantum and conventional multi–output learning algorithms

Hundreds of GaN thin film crystal plasma–assisted molecular beam epitaxy synthesis experiment records spanning two decades were organized into a dataset correlating the growth experiment design parameters with discrete, binary determinations of crystallinity and surface morphology. Conventional data science techniques as well as both quantum and classical multi–output supervised machine learning algorithms were implemented to investigate the relationships between the operating parameter data and the structural figures of merit. Correlation coefficients, decision tree nodes, p–values, and SHAP values all support substrate temperature and gallium effusion cell conditions as being statistically significant for simultaneously influencing GaN crystallinity and surface morphology. Here, a conventional deep neural network learned best from the data, followed by a quantum–classical hybrid gradient boosting algorithm. When combined with calculations of uncertainty intervals based on VennAbers predictors, machine learning predictions of both structural properties show good agreement with results reported in published experimental literature.

36 MATERIALS SCIENCE↗

The Feasibility of Incorporating a 3D Velocity Model Into Earthquake Location Around Salt Lake City, UT Using a Physics Informed Neural Network

Earthquake location algorithms typically require travel time calculation. Doing this calculation in 3D, despite advances in algorithm efficiency and computational power, can still be prohibitively expensive in terms of resources and storage. Implementation of high-resolution 3D models in routine earthquake location would be a significant step forward in most of the world. Machine learning algorithms have potential to act as substitutes for travel time calculation algorithms or stored travel time tables. We investigate EikoNet - a physics informed neural network machine learning model that estimates travel times very quickly and comes with negligible memory-overhead. Specifically, we apply EikoNet to the Wasatch Fault Community Velocity Model (WFCVM), a highly detailed and complex 3D velocity model of the Salt Lake City, UT region. While routine locations in the area and studies of the 2020 Magna, UT earthquake sequence used a 1D velocity model, a 3D model may help better our understanding the structure of the major fault in the region. Our primary goal was to test the speed, memory requirements, and accuracy of EikoNet compared to a reference eikonal solver. We find that while the EikoNet is exceedingly fast and requires little memory overhead, achieving acceptable accuracy in estimated travel times is difficult and requires extensive computational resources.

58 GEOSCIENCES↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Dimensionality Reduction with Variational Encoders Based on Subsystem Purification

Efficient methods for encoding and compression are likely to pave the way toward the problem of efficient trainability on higher-dimensional Hilbert spaces, overcoming issues of barren plateaus. Here, we propose an alternative approach to variational autoencoders to reduce the dimensionality of states represented in higher dimensional Hilbert spaces. To this end, we build a variational algorithm-based autoencoder circuit that takes as input a dataset and optimizes the parameters of a Parameterized Quantum Circuit (PQC) ansatz to produce an output state that can be represented as a tensor product of two subsystems by minimizing $Tr(ρ^2)$. The output of this circuit is passed through a series of controlled swap gates and measurements to output a state with half the number of qubits while retaining the features of the starting state in the same spirit as any dimension-reduction technique used in classical algorithms. The output obtained is used for supervised learning to guarantee the working of the encoding procedure thus developed. We make use of the Bars and Stripes (BAS) dataset for an 8 × 8 grid to create efficient encoding states and report a classification accuracy of 95% on the same. Thus, the demonstrated example provides proof for the working of the method in reducing states represented in large Hilbert spaces while maintaining the features required for any further machine learning algorithm that follows.

97 MATHEMATICS AND COMPUTING↗

Multi-fidelity modeling to predict the rheological properties of a suspension of fibers using neural networks and Gaussian processes

Unveiling the rheological properties of fiber suspensions is of paramount interest to many industrial applications. There are multiple factors, such as fiber aspect ratio and volume fraction, that play a significant role in altering the rheological behavior of suspensions. Three-dimensional (3D) numerical simulations of coupled differential equations of the suspension of fibers are computationally expensive and time-consuming. Machine learning algorithms can be trained on the available data and make predictions for the cases where no numerical data are available. However, some widely used machine learning surrogates, such as neural networks, require a relatively large training dataset to produce accurate predictions. Multi-fidelity models, which combine high-fidelity data from numerical simulations and less expensive lower fidelity data from resources such as simplified constitutive equations, can pave the way for more accurate predictions. Here, we focus on neural networks and the Gaussian processes with two levels of fidelity, i.e., high and low fidelity networks, to predict the steady-state rheological properties, and compare them to the single-fidelity network. High-fidelity data are obtained from direct numerical simulations based on an immersed boundary method to couple the fluid and solid motion. The low-fidelity data are produced by using constitutive equations. Multiple neural networks and the Gaussian process structures are used for the hyperparameter tuning purpose. Results indicate that with the best choice of hyperparameters, both the multi-fidelity Gaussian processes and neural networks are capable of making predictions with a high level of accuracy with neural networks demonstrating marginally better performance.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine learning opportunities for nucleosynthesis studies

Nuclear astrophysics is an interdisciplinary field focused on exploring the impact of nuclear physics on the evolution and explosions of stars and the cosmic creation of the elements. While researchers in astrophysics and in nuclear physics are separately using machine learning approaches to advance studies in their fields, there is currently little use of machine learning in nuclear astrophysics. We briefly describe the most common types of machine learning algorithms, and then detail their numerous possible uses to advance nuclear astrophysics, with a focus on simulation-based nucleosynthesis studies. We show that machine learning offers novel, complementary, creative approaches to address many important nucleosynthesis puzzles, with the potential to initiate a new frontier in nuclear astrophysics research.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SplitML

SplitML (Signal Processing Library for Interference rejecTion by Machine Learning) is a code repository for a set of tools for interference rejection in complex time-domain signals. The goal of the tools is to provide machine learning modeling capabilities for rejecting interference. Recent related machine learning algorithms for signal processing have focused mostly on speech enhancement or multi-speaker speech separation; the tools in SplitML will extend these innovations to generic time-domain signals of interest. SplitML includes tools for generating synthetic noisy signals, code for customized machine learning models applicable to interference rejection, and tools for evaluating such algorithms against standard signal processing techniques. These components are written in Python, a high-level programming language that takes advantage of the Python ecosystem of high-quality open-source packages for machine learning and signal processing.

Klein, Natalie↗

Recursive Use of the Short-Time Fast Fourier Transform for Signature Analysis in Continuous Processes

Although a nuclear reactor is a hostile environment for sensing and electrical communications, the reactor core is amenable to acoustic communication. An acoustic measurement infrastructure (AMI) has been installed in the Advanced Test Reactor (ATR) to record acoustic signals that can capture its different operating regimes. AMI uses coolant pumps as continuous signal sources, coolant and structural components as transmission lines, and accelerometers to capture system motion. A recursive signal processing technique based on the short-time fast Fourier transform (STFFT) for continuous processes provides unique signatures for diagnostic and prognostic analyses from the system motion data. Here this article presents a recursive STFFT methodology that processes acoustic signals from continuous industrial processes. The article first discusses the initial STFFT use with simulated data to elucidate the basic principles necessary to understand and interpret the STFFT results from actual pump vibration data. Each repetitive use of the STFFT on pump vibration data using the results from the prior STFFT processing will generate additional complimentary time-frequency-based signatures. These signatures are generated by the coolant pumps operating under different process conditions. After each use of the STFFT, the resulting signatures provide exemplary examples of the diversity and intuitive nature of recursively using the STFFT. This article focuses on recursively using the STFFT to provide numerous complimentary and diverse signatures that will ultimately be inputs for machine learning algorithms that provide predictive data analytics. The intuitive nature of the information and signatures from recursive STFFT processing will also bring intuitive interpretation capabilities to machine learning and predictive data analytic techniques.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗