Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

An Assessment of Machine Learning Applied to Ultrasonic Nondestructive Evaluation

In the United States, the nuclear industry performs inservice inspection (ISI) through nondestructive examination (NDE) methods in accordance with guidelines specified in the American Society of Mechanical Engineers (ASME) Boiler and Pressure Vessel Code (BPVC), Section XI, Rules for Inservice Inspection of Nuclear Power Plant Components. Ultrasonic nondestructive testing and evaluation (NDT&E) is one of the more commonly used techniques for inspecting Class 1 structural components in nuclear power systems. As the number of qualified NDE inspectors declines, the nuclear industry is looking to take advantage of advances in automation to enhance inspection capabilities. Advances in computational power, cloud-based computing, and machine learning algorithms make automated data analysis possible. Machine learning (ML) has shown huge potential in automated data analyses for ultrasonic NDE in the context of weld inspections.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Predictive Modeling of NOx Emissions from Lean Direct Injection of Hydrogen and Hydrogen/Natural Gas Blends Using Flame Imaging and Machine Learning

This research paper explores the use of machine learning to relate images of flame structure and luminosity to measured NOx emissions. Images of reactions produced by 16 aero-engine derived injectors for a ground-based turbine operated on a range of fuel compositions, air pressure drops, preheat temperatures and adiabatic flame temperatures were captured and postprocessed. The experimental investigations were conducted under atmospheric conditions, capturing CO, NO and NOx emissions data and OH* chemiluminescence images from 27 test conditions. The injector geometry and test conditions were based on a statistically designed test plan. These results were first analyzed using the traditional analysis approach of analysis of variance (ANOVA). The statistically based test plan yielded 432 data points, leading to a correlation for NOx emissions as a function of injector geometry, test conditions and imaging responses, with 70.2% accuracy. As an alternative approach to predicting emissions using imaging diagnostics as well as injector geometry and test conditions, a random forest machine learning algorithm was also applied to the data and was able to achieve an accuracy of 82.6%. This study offers insights into the factors influencing emissions in ground-based turbines while emphasizing the potential of machine learning algorithms in constructing predictive models for complex systems.

08 HYDROGEN↗

GPU-Accelerated Machine Learning Inference as a Service for Computing in Neutrino Experiments

Machine learning algorithms are becoming increasingly prevalent and performant in the reconstruction of events in accelerator-based neutrino experiments. These sophisticated algorithms can be computationally expensive. At the same time, the data volumes of such experiments are rapidly increasing. The demand to process billions of neutrino events with many machine learning algorithm inferences creates a computing challenge. We explore a computing model in which heterogeneous computing with GPU coprocessors is made available as a web service. The coprocessors can be efficiently and elastically deployed to provide the right amount of computing for a given processing task. With our approach, Services for Optimized Network Inference on Coprocessors (SONIC), we integrate GPU acceleration specifically for the ProtoDUNE-SP reconstruction chain without disrupting the native computing workflow. With our integrated framework, we accelerate the most time-consuming task, track and particle shower hit identification, by a factor of 17. This results in a factor of 2.7 reduction in the total processing time when compared with CPU-only production. For this particular task, only 1 GPU is required for every 68 CPU threads, providing a cost-effective solution.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Proactive Intrusion Detection and Mitigation System

SAND2023-05661O The proactive intrusion detection and mitigation system (PIDMS) provides grid-edge situational awareness for cybersecurity defense by capturing real-time distributed energy resource (DER) network traffic and performance data with a novel approach that improves the detection and prevention of cyber-physical attacks. The PIDMS addresses the grid-edge security gap with real-time analysis of both network traffic and photovoltaic performance data to deliver a novel, cyber-physical intrusion detection system (IDS) approach that increases the accuracy and effectiveness of detection and mitigation. This hybrid IDS analysis enables dual monitoring that increases the workload of the adversary; both cyber and physical data would have to be simultaneously spoofed to evade detection. Furthermore, monitoring and analyzing cyber data are insufficient in some cases. For example, in an insider threat aimed at disrupting inverter grid-support functions where proper credentials and authentication are achieved, only the altered PV performance would indicate abnormal behavior. All in all, the PIDMS provides novel capabilities for: • Distributed, real-time cyber-physical detection and mitigation analysis • Cybersecurity defense for grid-edge systems • Analysis framework that can provide situational awareness across the transmission, distribution, and DER systems The PIDMS sensor is designed to collect cyber-physical data, process the data using machine-learning algorithms, detect abnormal events, and deploy mitigations. With these goals, the main functional PIDMS objectives are: • Capability to collect cyber-physical data • Onboard storage of cyber-physical data • Peer-to-peer communication • Computationally efficient machine-learning algorithms • Online cyber-physical data analysis • Alerting/visualization capabilities • Mitigation deployment capability with bump-in-the-wire (BITW) implementation Each of these functional objectives enable PIDMS to perform effective cyber-physical intrusion detection and mitigation. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Jones, Christian↗

Optimizing a magnitude-limited spectroscopic training sample for photometric classification of supernovae

ABSTRACT In preparation for photometric classification of transients from the Legacy Survey of Space and Time (LSST) we run tests with different training data sets. Using estimates of the depth to which the 4-m Multi-Object Spectroscopic Telescope (4MOST) Time Domain Extragalactic Survey (TiDES) can classify transients, we simulate a magnitude-limited sample reaching rAB ≈ 22.5 mag. We run our simulations with the software snmachine, a photometric classification pipeline using machine learning. The machine-learning algorithms struggle to classify supernovae when the training sample is magnitude limited, in contrast to representative training samples. Classification performance noticeably improves when we combine the magnitude-limited training sample with a simulated realistic sample of faint high-redshift supernovae observed from larger spectroscopic facilities; the algorithms’ range of average area under receiver operator characteristic curve (AUC) scores over 10 runs increases from 0.547–0.628 to 0.946–0.969 and purity of the classified sample reaches 95 per cent in all runs for two of the four algorithms. By creating new, artificial light curves using the augmentation software avocado, we achieve a purity in our classified sample of 95 per cent in all 10 runs performed for all machine-learning algorithms considered. We also reach a highest average AUC score of 0.986 with the artificial neural network algorithm. Having ‘true’ faint supernovae to complement our magnitude-limited sample is a crucial requirement in optimization of a 4MOST spectroscopic sample. However, our results are a proof of concept that augmentation is also necessary to achieve the best classification results.

79 ASTRONOMY AND ASTROPHYSICS↗

hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices

Accessible machine learning algorithms, software, and diagnostic tools for energy-efficient devices and systems are extremely valuable across a broad range of application domains. In scientific domains, real-time near-sensor processing can drastically improve experimental design and accelerate scientific discoveries. To support domain scientists, we have developed hls4ml, an open-source software-hardware codesign workflow to interpret and translate machine learning algorithms for implementation with both FPGA and ASIC technologies. We expand on previous hls4ml work by extending capabilities and techniques towards low-power implementations and increased usability: new Python APIs, quantization-aware pruning, end-to-end FPGA workflows, long pipeline kernels for low power, and new device backends include an ASIC workflow. Taken together, these and continued efforts in hls4ml will arm a new generation of domain scientists with accessible, efficient, and powerful tools for machine-learning-accelerated discovery.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine Learning Pattern Recognition Algorithm With Applications to Coherent Laser Combination

Herein we analyze a new kind of machine learning algorithm designed to feedback stabilize coherently combined lasers. This algorithm learns differential, rather than absolute, values of action in phase space, in order to facilitate learning on initially unstable systems. Experiments have shown that this approach can control small-scale spatial beam combination with high stability. In this paper we analyze the algorithm's performance and limitations in depth, showing that it can continuously learn during operation in order to track changes. Using simulation, we extend the application to temporal combination, and show that it scales to more complex instances by combining 81 beams.

97 MATHEMATICS AND COMPUTING↗

Integration of scanning probe microscope with high-performance computing: Fixed-policy and reward-driven workflows implementation

The rapid development of computation power and machine learning algorithms has paved the way for automating scientific discovery with a scanning probe microscope (SPM). The key elements toward operationalization of the automated SPM are the interface to enable SPM control from Python codes, availability of high computing power, and development of workflows for scientific discovery. Here, we build a Python interface library that enables controlling an SPM from either a local computer or a remote high-performance computer, which satisfies the high computation power need of machine learning algorithms in autonomous workflows. We further introduce a general platform to abstract the operations of SPM in scientific discovery into fixed-policy or reward-driven workflows. Furthermore, our work provides a full infrastructure to build automated SPM workflows for both routine operations and autonomous scientific discovery with machine learning.

47 OTHER INSTRUMENTATION↗

Machine learning based algorithms for uncertainty quantification in numerical weather prediction models

Complex numerical weather prediction models incorporate a variety of physical processes, each described by multiple alternative physical schemes with specific parameters. The selection of the physical schemes and the choice of the corresponding physical parameters during model configuration can significantly impact the accuracy of model forecasts. There is no combination of physical schemes that works best for all times, at all locations, and under all conditions. It is therefore of considerable interest to understand the interplay between the choice of physics and the accuracy of the resulting forecasts under different conditions. This paper demonstrates the use of machine learning techniques to study the uncertainty in numerical weather prediction models due to the interaction of multiple physical processes. The first problem addressed herein is the estimation of systematic model errors in output quantities of interest at future times, and the use of this information to improve the model forecasts. The second problem considered is the identification of those specific physical processes that contribute most to the forecast uncertainty in the quantity of interest under specified meteorological conditions. In order to address these questions we employ two machine learning approaches, random forests and artificial neural networks. The discrepancies between model results and observations at past times are used to learn the relationships between the choice of physical processes and the resulting forecast errors. Numerical experiments are carried out with the Weather Research and Forecasting (WRF) model. The output quantity of interest is the model precipitation, a variable that is both extremely important and very challenging to forecast. The physical processes under consideration include various micro-physics schemes, cumulus parameterizations, short wave, and long wave radiation schemes. The experiments demonstrate the strong potential of machine learning approaches to aid the study of model errors.

97 MATHEMATICS AND COMPUTING↗

Automating Bug Report Classification with Few Shot Learning

Orthogonal defect classification (ODC) is a method used to categorize software defects, providing valuable insights into the development process. This study focuses on automating the classification of software bug reports into different ODC defect types using few shot learning, a machine learning approach that requires minimal labeled data. Previous research has manually classified bug reports or used traditional machine learning algorithms like linear support vector machine, achieving limited success. Our approach uses few shot learning to improve classification accuracy and efficiency. The results show a harmonic mean of recall and precision (i.e., the F1 score) of around 0.6 which is a performance improvement over previous methods. The results highlight the potential benefit of few shot learning techniques and their application in enhancing the safety and reliability of nuclear digital instrumentation and control (DI&C) systems. Future work will explore incorporating advanced techniques to supplement the model's training data and achieve better results.

42 - ENGINEERING↗

Process Image Analysis using Big Data, Machine Learning, and Computer Vision

The development of algorithms for machine learning and data analysis for the 3013 MIS corrosion surveillance program is a collaborative effort by SRNL, USC and GT. For corrosion detection, LCM image data is extracted from large binary files, with software written to convert the data to physical attributes (i.e. height, color and grayscale values; all as functions of a location in a plane projection). The user interface for the software permits selective downloading of binary data and interrogation of attributes. User input thresholds are used to flag attributes of interest. Machine learning algorithms, developed for this application, are used to determine whether the features are the result of corrosion. To address the fundamental mechanisms of corrosion, machine learning algorithms are being developed to derive interatomic potential force-fields from ab-initio DFT calculations. The goal is to apply molecular modeling on a large enough scale to guide the design of resistant materials.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

ECNet is an evolutionary context-integrated deep learning framework for protein engineering

Abstract Machine learning has been increasingly used for protein engineering. However, because the general sequence contexts they capture are not specific to the protein being engineered, the accuracy of existing machine learning algorithms is rather limited. Here, we report ECNet (evolutionary context-integrated neural network), a deep-learning algorithm that exploits evolutionary contexts to predict functional fitness for protein engineering. This algorithm integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. As such, it enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-order mutants. We show that ECNet predicts the sequence-function relationship more accurately as compared to existing machine learning algorithms by using ~50 deep mutational scanning and random mutagenesis datasets. Moreover, we used ECNet to guide the engineering of TEM-1 β-lactamase and identified variants with improved ampicillin resistance with high success rates.

59 BASIC BIOLOGICAL SCIENCES↗

Code for the manuscript "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Mode

We disclose a python/pytorch implementation of the physics-informed machine learning algorithm described in "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Modeling", LA-UR-24-30678. Direct numerical simulation (DNS) of ubiquitous turbulence phenomena is computationally infeasible for realistic flows. As a result, reduced modeling for turbulent flows aim to reduce the number of resolved scales while retaining accurate representations of the small-scale physics. The dynamics of the velocity gradient tensor (VGT) is a key ingredient in reduced or subgrid turbulence models. The evolution equation for the VGT involves nonlocal terms, requiring closure modeling. This implementation of the novel methodology of Lagrangian Attention Tensor Networks (LATN), utilizes a structured representation of the history of the VGT to inform a physics-informed machine learning algorithm. This addition of structured memory terms is shown to outperform previous models when trained and evaluated on DNS data.

Livescu, Daniel [LANL]↗

Interpreting Write Performance of Supercomputer I/O Systems with Regression Models

This work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10× on some samples for both of the target systems.

Xie, Bing↗

A Wrapper to Use a Machine-Learning-Based Algorithm for Earthquake Monitoring

Seismology is one of the main sciences used to monitor volcanic activity worldwide. Fast, efficient, and accurate seismicity detectors are crucial to assess the activity level of a volcano in near–real time and to issue timely warnings. Traditional real–time seismic processing software uses phase onset pickers followed by a phase association algorithm to declare an event and estimate its location. The pickers typically do not identify whether the detected phase is a P or S arrival, which can have a negative impact on hypocentral location quality and complicates phase association. We implemented the deep–neural–network–based method PhaseNet to identify in real time P and S seismic waves on data from one– and three–component seismometers. We tuned the Earthworm binder_ew associator module to use the phase identification from PhaseNet to detect and locate the events, which we archive in a SeisComP3 database. We assessed the performance of the algorithm by comparing the results with existing catalogs built to monitor seismic and volcanic activity in Mayotte and the Lesser Antilles region. Our algorithm, which we refer to as PhaseWorm, showed promising results in both contexts and clearly outperformed the previous automatic method implemented in Mayotte. As a result, this innovative real–time processing system is now operational for seismicity monitoring in Mayotte and Martinique.

58 GEOSCIENCES↗

Anomaly detection in the Zwicky Transient Facility DR3

We present results from applying the SNAD anomaly detection pipeline to the third public data release of the Zwicky Transient Facility (ZTF DR3). The pipeline is composed of three stages: feature extraction, search of outliers with machine learning algorithms, and anomaly identification with followup by human experts. Our analysis concentrates in three ZTF fields, comprising more than 2.25 million objects. A set of four automatic learning algorithms was used to identify 277 outliers, which were subsequently scrutinized by an expert. From these, 188 (68 per cent) were found to be bogus light curves – including effects from the image subtraction pipeline as well as overlapping between a star and a known asteroid, 66 (24 per cent) were previously reported sources whereas 23 (8 per cent) correspond to non-catalogued objects, with the two latter cases of potential scientific interest (e.g. one spectroscopically confirmed RS Canum Venaticorum star, four supernovae candidates, one red dwarf flare). Moreover, using results from the expert analysis, we were able to identify a simple bi-dimensional relation that can be used to aid filtering potentially bogus light curves in future studies. We provide a complete list of objects with potential scientific application so they can be further scrutinised by the community. These results confirm the importance of combining automatic machine learning algorithms with domain knowledge in the construction of recommendation systems for astronomy. Our code is publicly available.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning for Improving Surface-Layer-Flux Estimates

Abstract Flows in the atmospheric boundary layer are turbulent, characterized by a large Reynolds number, the existence of a roughness sublayer and the absence of a well-defined viscous layer. Exchanges with the surface are therefore dominated by turbulent fluxes. In numerical models for atmospheric flows, turbulent fluxes must be specified at the surface; however, surface fluxes are not known a priori and therefore must be parametrized. Atmospheric flow models, including global circulation, limited area models, and large-eddy simulation, employ Monin–Obukhov similarity theory (MOST) to parametrize surface fluxes. The MOST approach is a semi-empirical formulation that accounts for atmospheric stability effects through universal stability functions. The stability functions are determined based on limited observations using simple regression as a function of the non-dimensional stability parameter representing a ratio of distance from the surface and the Obukhov length scale (Obukhov in Trudy Inst Theor Geofiz AN SSSR 1:95–115, 1946), $$z/L$$ z / L . However, simple regression cannot capture the relationship between governing parameters and surface-layer structure under the wide range of conditions to which MOST is commonly applied. We therefore develop, train, and test two machine-learning models, an artificial neural network (ANN) and random forest (RF), to estimate surface fluxes of momentum, sensible heat, and moisture based on surface and near-surface observations. To train and test these machine-learning algorithms, we use several years of observations from the Cabauw mast in the Netherlands and from the National Oceanic and Atmospheric Administration’s Field Research Division tower in Idaho. The RF and ANN models outperform MOST. Even when we train the RF and ANN on one set of data and apply them to the second set, they provide more accurate estimates of all of the fluxes compared to MOST. Estimates of sensible heat and moisture fluxes are significantly improved, and model interpretability techniques highlight the logical physical relationships we expect in surface-layer processes.

Meteorology & Atmospheric Sciences↗