Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

GPU-Accelerated Machine Learning Inference as a Service for Computing in Neutrino Experiments

Machine learning algorithms are becoming increasingly prevalent and performant in the reconstruction of events in accelerator-based neutrino experiments. These sophisticated algorithms can be computationally expensive. At the same time, the data volumes of such experiments are rapidly increasing. The demand to process billions of neutrino events with many machine learning algorithm inferences creates a computing challenge. We explore a computing model in which heterogeneous computing with GPU coprocessors is made available as a web service. The coprocessors can be efficiently and elastically deployed to provide the right amount of computing for a given processing task. With our approach, Services for Optimized Network Inference on Coprocessors (SONIC), we integrate GPU acceleration specifically for the ProtoDUNE-SP reconstruction chain without disrupting the native computing workflow. With our integrated framework, we accelerate the most time-consuming task, track and particle shower hit identification, by a factor of 17. This results in a factor of 2.7 reduction in the total processing time when compared with CPU-only production. For this particular task, only 1 GPU is required for every 68 CPU threads, providing a cost-effective solution.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Proactive Intrusion Detection and Mitigation System

SAND2023-05661O The proactive intrusion detection and mitigation system (PIDMS) provides grid-edge situational awareness for cybersecurity defense by capturing real-time distributed energy resource (DER) network traffic and performance data with a novel approach that improves the detection and prevention of cyber-physical attacks. The PIDMS addresses the grid-edge security gap with real-time analysis of both network traffic and photovoltaic performance data to deliver a novel, cyber-physical intrusion detection system (IDS) approach that increases the accuracy and effectiveness of detection and mitigation. This hybrid IDS analysis enables dual monitoring that increases the workload of the adversary; both cyber and physical data would have to be simultaneously spoofed to evade detection. Furthermore, monitoring and analyzing cyber data are insufficient in some cases. For example, in an insider threat aimed at disrupting inverter grid-support functions where proper credentials and authentication are achieved, only the altered PV performance would indicate abnormal behavior. All in all, the PIDMS provides novel capabilities for: • Distributed, real-time cyber-physical detection and mitigation analysis • Cybersecurity defense for grid-edge systems • Analysis framework that can provide situational awareness across the transmission, distribution, and DER systems The PIDMS sensor is designed to collect cyber-physical data, process the data using machine-learning algorithms, detect abnormal events, and deploy mitigations. With these goals, the main functional PIDMS objectives are: • Capability to collect cyber-physical data • Onboard storage of cyber-physical data • Peer-to-peer communication • Computationally efficient machine-learning algorithms • Online cyber-physical data analysis • Alerting/visualization capabilities • Mitigation deployment capability with bump-in-the-wire (BITW) implementation Each of these functional objectives enable PIDMS to perform effective cyber-physical intrusion detection and mitigation. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Jones, Christian↗

Machine Learning to Predict Joint Performance in Epoxy Composites Based on Process Parameters

Polymer matrix composites are gaining popularity in the aerospace industry due to their high specific strength, fatigue properties, and processability. However, based on current FAA certification guidelines, manufacturers utilizing current state-of-the-art composites made with adhesive bonds commonly install redundant fasteners to guarantee the strength of these adhesively bonded composite parts. The number of fasteners in a single-aisle commercial transport aircraft is typically on the order of 105, which reduces manufacturing rate, increases cost tremendously, and reduces the advantage of the specific strength composites provide. Due to this, the Adhesive Free Bonding of Composites (AERoBOND) project at NASA Langley Research Center has developed a novel assembly process to manufacture complex composite parts without the use of adhesives and fasteners. However, optimization of the process is currently challenging due to the complex and interdependent process parameters. To assist with the optimization, four machine learning algorithms utilizing gradient boosting decision trees were created to provide predictions for the mechanical and characterization properties of the composite parts. Approximately 200 random states from each algorithm were tested, and the models from each state were isolated and analyzed based on their accuracy, a validation process, and their feature importance. This analysis concluded that the models created from the machine learning algorithms could accelerate a parametric study for the AERoBOND process by rapidly optimizing process parameters to achieve desired performance characteristics.

Brennen Michael Middleton↗

Machine Learning to Predict Joint Performance in Epoxy Composites Based on Process Parameters

Polymer matrix composites are gaining popularity in the aerospace industry due to their high specific strength, fatigue properties, and processability. However, based on current FAA certification guidelines, manufacturers utilizing current state-of-the art composites made with adhesive bonds commonly install redundant fasteners to guarantee the strength of these adhesively bonded composite parts.1,2 The number of fasteners in a single-aisle commercial transport aircraft is typically on the order of 105, which reduces manufacturing rate, increases cost tremendously, and reduces the advantage of the specific strength composites provide. Due to this, the Adhesive Free Bonding of Composites (AERoBOND) project at NASA Langley Research Center has developed a novel assembly process to manufacture complex composite parts without the use of adhesives and fasteners.1 However, optimization of the process is currently challenging due to the complex and interdependent process parameters. To assist with the optimization, four machine learning algorithms utilizing gradient boosting decision trees were created to provide predictions for the mechanical and characterization properties of the composite parts. Approximately 200 random states from each algorithm were tested, and the models from each state were isolated and analyzed based on their accuracy, a validation process, and their feature importance. This analysis concluded that the models created from the machine learning algorithms could accelerate a parametric study for the AERoBOND process by rapidly optimizing process parameters to achieve desired performance characteristics.

Brennen M Middleton↗

A Global Land Cover Training Dataset From 1984 to 2020

State-of-the-art cloud computing platforms such as Google Earth Engine (GEE) enable regional-to-global land cover and land cover change mapping with machine learning algorithms. However, collection of high-quality training data, which is necessary for accurate land cover mapping, remains costly and labor-intensive. To address this need, we created a global database of nearly 2 million training units spanning the period from 1984 to 2020 for seven primary and nine secondary land cover classes. Our training data collection approach leveraged GEE and machine learning algorithms to ensure data quality and biogeographic representation. We sampled the spectral-temporal feature space from Landsat imagery to efficiently allocate training data across global ecoregions and incorporated publicly available and collaborator-provided datasets to our database. To reflect the underlying regional class distribution and post-disturbance landscapes, we strategically augmented the database. We used a machine learning-based cross-validation procedure to remove potentially mis-labeled training units. Our training database is relevant for a wide array of studies such as land cover change, agriculture, forestry, hydrology, urban development, among many others.

Radost Stanimirova↗

Optimizing a magnitude-limited spectroscopic training sample for photometric classification of supernovae

ABSTRACT In preparation for photometric classification of transients from the Legacy Survey of Space and Time (LSST) we run tests with different training data sets. Using estimates of the depth to which the 4-m Multi-Object Spectroscopic Telescope (4MOST) Time Domain Extragalactic Survey (TiDES) can classify transients, we simulate a magnitude-limited sample reaching rAB ≈ 22.5 mag. We run our simulations with the software snmachine, a photometric classification pipeline using machine learning. The machine-learning algorithms struggle to classify supernovae when the training sample is magnitude limited, in contrast to representative training samples. Classification performance noticeably improves when we combine the magnitude-limited training sample with a simulated realistic sample of faint high-redshift supernovae observed from larger spectroscopic facilities; the algorithms’ range of average area under receiver operator characteristic curve (AUC) scores over 10 runs increases from 0.547–0.628 to 0.946–0.969 and purity of the classified sample reaches 95 per cent in all runs for two of the four algorithms. By creating new, artificial light curves using the augmentation software avocado, we achieve a purity in our classified sample of 95 per cent in all 10 runs performed for all machine-learning algorithms considered. We also reach a highest average AUC score of 0.986 with the artificial neural network algorithm. Having ‘true’ faint supernovae to complement our magnitude-limited sample is a crucial requirement in optimization of a 4MOST spectroscopic sample. However, our results are a proof of concept that augmentation is also necessary to achieve the best classification results.

79 ASTRONOMY AND ASTROPHYSICS↗

hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices

Accessible machine learning algorithms, software, and diagnostic tools for energy-efficient devices and systems are extremely valuable across a broad range of application domains. In scientific domains, real-time near-sensor processing can drastically improve experimental design and accelerate scientific discoveries. To support domain scientists, we have developed hls4ml, an open-source software-hardware codesign workflow to interpret and translate machine learning algorithms for implementation with both FPGA and ASIC technologies. We expand on previous hls4ml work by extending capabilities and techniques towards low-power implementations and increased usability: new Python APIs, quantization-aware pruning, end-to-end FPGA workflows, long pipeline kernels for low power, and new device backends include an ASIC workflow. Taken together, these and continued efforts in hls4ml will arm a new generation of domain scientists with accessible, efficient, and powerful tools for machine-learning-accelerated discovery.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine Learning Pattern Recognition Algorithm With Applications to Coherent Laser Combination

Herein we analyze a new kind of machine learning algorithm designed to feedback stabilize coherently combined lasers. This algorithm learns differential, rather than absolute, values of action in phase space, in order to facilitate learning on initially unstable systems. Experiments have shown that this approach can control small-scale spatial beam combination with high stability. In this paper we analyze the algorithm's performance and limitations in depth, showing that it can continuously learn during operation in order to track changes. Using simulation, we extend the application to temporal combination, and show that it scales to more complex instances by combining 81 beams.

97 MATHEMATICS AND COMPUTING↗

Integration of scanning probe microscope with high-performance computing: Fixed-policy and reward-driven workflows implementation

The rapid development of computation power and machine learning algorithms has paved the way for automating scientific discovery with a scanning probe microscope (SPM). The key elements toward operationalization of the automated SPM are the interface to enable SPM control from Python codes, availability of high computing power, and development of workflows for scientific discovery. Here, we build a Python interface library that enables controlling an SPM from either a local computer or a remote high-performance computer, which satisfies the high computation power need of machine learning algorithms in autonomous workflows. We further introduce a general platform to abstract the operations of SPM in scientific discovery into fixed-policy or reward-driven workflows. Furthermore, our work provides a full infrastructure to build automated SPM workflows for both routine operations and autonomous scientific discovery with machine learning.

47 OTHER INSTRUMENTATION↗

Automating Bug Report Classification with Few Shot Learning

Orthogonal defect classification (ODC) is a method used to categorize software defects, providing valuable insights into the development process. This study focuses on automating the classification of software bug reports into different ODC defect types using few shot learning, a machine learning approach that requires minimal labeled data. Previous research has manually classified bug reports or used traditional machine learning algorithms like linear support vector machine, achieving limited success. Our approach uses few shot learning to improve classification accuracy and efficiency. The results show a harmonic mean of recall and precision (i.e., the F1 score) of around 0.6 which is a performance improvement over previous methods. The results highlight the potential benefit of few shot learning techniques and their application in enhancing the safety and reliability of nuclear digital instrumentation and control (DI&C) systems. Future work will explore incorporating advanced techniques to supplement the model's training data and achieve better results.

42 - ENGINEERING↗

ECNet is an evolutionary context-integrated deep learning framework for protein engineering

Abstract Machine learning has been increasingly used for protein engineering. However, because the general sequence contexts they capture are not specific to the protein being engineered, the accuracy of existing machine learning algorithms is rather limited. Here, we report ECNet (evolutionary context-integrated neural network), a deep-learning algorithm that exploits evolutionary contexts to predict functional fitness for protein engineering. This algorithm integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. As such, it enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-order mutants. We show that ECNet predicts the sequence-function relationship more accurately as compared to existing machine learning algorithms by using ~50 deep mutational scanning and random mutagenesis datasets. Moreover, we used ECNet to guide the engineering of TEM-1 β-lactamase and identified variants with improved ampicillin resistance with high success rates.

59 BASIC BIOLOGICAL SCIENCES↗

Code for the manuscript "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Mode

We disclose a python/pytorch implementation of the physics-informed machine learning algorithm described in "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Modeling", LA-UR-24-30678. Direct numerical simulation (DNS) of ubiquitous turbulence phenomena is computationally infeasible for realistic flows. As a result, reduced modeling for turbulent flows aim to reduce the number of resolved scales while retaining accurate representations of the small-scale physics. The dynamics of the velocity gradient tensor (VGT) is a key ingredient in reduced or subgrid turbulence models. The evolution equation for the VGT involves nonlocal terms, requiring closure modeling. This implementation of the novel methodology of Lagrangian Attention Tensor Networks (LATN), utilizes a structured representation of the history of the VGT to inform a physics-informed machine learning algorithm. This addition of structured memory terms is shown to outperform previous models when trained and evaluated on DNS data.

Livescu, Daniel [LANL]↗

Evaluation of Technology Concepts for Traffic Data Management and Relevant Audio for Datalink in Commercial Airline Flight Decks

Datalink is currently operational for departure clearances and in oceanic environments and is currently being tested in high altitude domestic enroute airspace. Interaction with even simple datalink clearances may create more workload for flight crews than the voice system they replace if not carefully designed. Datalink may also introduce additional complexity for flight crews with hundreds of uplink messages now defined for use. Finally, flight crews may lose airspace awareness and operationally relevant information that they normally pickup from Air Traffic Control (ATC) voice communications with other aircraft (i.e., “party-line” transmissions). Once again, automation may be poised to increase workload on the flight deck for incremental benefit. Datalink implementation to support future air traffic management concepts needs to be carefully considered, understanding human communication norms and especially, the change from voice- to text-based communications modality and its effect on pilot workload and situation awareness. Increasingly autonomous systems, where autonomy is designed to support human-autonomy teaming, may be suited to solve these issues. NASA is conducting research and development of increasingly autonomous systems, utilizing machine-learning algorithms seamlessly integrated with humans whereby task performance of the combined system is significantly greater than the individual components. Increasingly autonomous systems offer the potential for significantly improved levels of performance and safety that are superior to either human or automation alone. Two increasingly autonomous systems concepts - a traffic data manager and a conversational co-pilot - were developed to intelligently address the datalink issues in a complex, future state environment with significant levels of traffic. The system was tested for suitability of datalink usage for terminal airspace. The traffic data manager allowed for automated declutter of the Automatic Dependent Surveillance-Broadcast (ADS-B) display. The system determined relevant traffic for display based on machine learning algorithms trained by experienced human pilot behaviors. The conversational co-pilot provided relevant audio air traffic control messages based on context and proximity to ownship. Both systems made use of the connected aircraft concepts to provide intelligent context to determine relevancy above and beyond proximity to ownship. A human-in-the-loop test was conducted in NASA Langley Research Center’s Integration Flight Deck B-737-800 simulator to evaluate the traffic data manager and the conversational co-pilot. Twelve airline crews flew various normal and non-normal procedures and their actions and performance were recorded in response to the procedural events. This paper details the flight crew performance and evaluation during the events.

Etherington, Timothy↗

A Wrapper to Use a Machine-Learning-Based Algorithm for Earthquake Monitoring

Seismology is one of the main sciences used to monitor volcanic activity worldwide. Fast, efficient, and accurate seismicity detectors are crucial to assess the activity level of a volcano in near–real time and to issue timely warnings. Traditional real–time seismic processing software uses phase onset pickers followed by a phase association algorithm to declare an event and estimate its location. The pickers typically do not identify whether the detected phase is a P or S arrival, which can have a negative impact on hypocentral location quality and complicates phase association. We implemented the deep–neural–network–based method PhaseNet to identify in real time P and S seismic waves on data from one– and three–component seismometers. We tuned the Earthworm binder_ew associator module to use the phase identification from PhaseNet to detect and locate the events, which we archive in a SeisComP3 database. We assessed the performance of the algorithm by comparing the results with existing catalogs built to monitor seismic and volcanic activity in Mayotte and the Lesser Antilles region. Our algorithm, which we refer to as PhaseWorm, showed promising results in both contexts and clearly outperformed the previous automatic method implemented in Mayotte. As a result, this innovative real–time processing system is now operational for seismicity monitoring in Mayotte and Martinique.

58 GEOSCIENCES↗

Anomaly detection in the Zwicky Transient Facility DR3

We present results from applying the SNAD anomaly detection pipeline to the third public data release of the Zwicky Transient Facility (ZTF DR3). The pipeline is composed of three stages: feature extraction, search of outliers with machine learning algorithms, and anomaly identification with followup by human experts. Our analysis concentrates in three ZTF fields, comprising more than 2.25 million objects. A set of four automatic learning algorithms was used to identify 277 outliers, which were subsequently scrutinized by an expert. From these, 188 (68 per cent) were found to be bogus light curves – including effects from the image subtraction pipeline as well as overlapping between a star and a known asteroid, 66 (24 per cent) were previously reported sources whereas 23 (8 per cent) correspond to non-catalogued objects, with the two latter cases of potential scientific interest (e.g. one spectroscopically confirmed RS Canum Venaticorum star, four supernovae candidates, one red dwarf flare). Moreover, using results from the expert analysis, we were able to identify a simple bi-dimensional relation that can be used to aid filtering potentially bogus light curves in future studies. We provide a complete list of objects with potential scientific application so they can be further scrutinised by the community. These results confirm the importance of combining automatic machine learning algorithms with domain knowledge in the construction of recommendation systems for astronomy. Our code is publicly available.

79 ASTRONOMY AND ASTROPHYSICS↗

Advancing Open Science in Atmospheric Research: Integrating Data Usability and Machine Learning

In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences." "In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences.

Jennifer Wei↗

Machine Learning for Improving Surface-Layer-Flux Estimates

Abstract Flows in the atmospheric boundary layer are turbulent, characterized by a large Reynolds number, the existence of a roughness sublayer and the absence of a well-defined viscous layer. Exchanges with the surface are therefore dominated by turbulent fluxes. In numerical models for atmospheric flows, turbulent fluxes must be specified at the surface; however, surface fluxes are not known a priori and therefore must be parametrized. Atmospheric flow models, including global circulation, limited area models, and large-eddy simulation, employ Monin–Obukhov similarity theory (MOST) to parametrize surface fluxes. The MOST approach is a semi-empirical formulation that accounts for atmospheric stability effects through universal stability functions. The stability functions are determined based on limited observations using simple regression as a function of the non-dimensional stability parameter representing a ratio of distance from the surface and the Obukhov length scale (Obukhov in Trudy Inst Theor Geofiz AN SSSR 1:95–115, 1946), $$z/L$$ z / L . However, simple regression cannot capture the relationship between governing parameters and surface-layer structure under the wide range of conditions to which MOST is commonly applied. We therefore develop, train, and test two machine-learning models, an artificial neural network (ANN) and random forest (RF), to estimate surface fluxes of momentum, sensible heat, and moisture based on surface and near-surface observations. To train and test these machine-learning algorithms, we use several years of observations from the Cabauw mast in the Netherlands and from the National Oceanic and Atmospheric Administration’s Field Research Division tower in Idaho. The RF and ANN models outperform MOST. Even when we train the RF and ANN on one set of data and apply them to the second set, they provide more accurate estimates of all of the fluxes compared to MOST. Estimates of sensible heat and moisture fluxes are significantly improved, and model interpretability techniques highlight the logical physical relationships we expect in surface-layer processes.

Meteorology & Atmospheric Sciences↗

Machine-Learning-based Algorithms for Automated Image Segmentation Techniques of Transmission X-ray Microscopy (TXM)

Four state-of-the-art Deep Learning-based Convolutional Neural Networks (DCNN) were applied to automate the semantic segmentation of a 3D Transmission x-ray Microscopy (TXM) nanotomography image data. The standard U-Net architecture as baseline along with UNet++, PSPNet, and DeepLab v3+ networks were trained to segment the microstructural features of an AA7075 micropillar. A workflow was established to evaluate and compare the DCNN prediction dataset with the manually segmented features using the Intersection of Union (IoU) scores, time of training, confusion matrix, and visual assessment. Comparing all model segmentation accuracy metrics, it was found that using pre-trained models as a backbone along with appropriate training encoder-decoder architecture of the Unet++ can robustly handle large volumes of x-ray radiographic images in a reasonable amount of time. This opens a new window for handling accurate and efficient image segmentation of in situ time-dependent 4D x-ray microscopy experimental datasets.

36 MATERIALS SCIENCE↗