Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “active machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Studying the hadron structure with PANDA and CLAS using machine learning techniques

The hadron spectroscopy and structure are currently very active fields of research to study the non-perturbative regime of quantum chronodynamics. The first one studies the complex structure of excited hadrons by looking at their decay products, while the latter uses lepton scattering on nucleons. Both methods require reconstruction algorithms with great efficiency and good particle identification and background rejection rates. This work aims to provide these by either improving the existing methods or developing new ones. The first part of this document presents a feasibility study of a predicted hybrid charmonium state for the PANDA experiment. Lattice QCD calculations predict the ground state hybrid charmonium to be a spin exotic with quantum numbers of JP C = 1?+ at a mass of around 4.3 GeV with a width to be around 20 MeV. A machine learning based data analysis scheme is proposed to further improve the signal efficiency and the background reduction, alongside with improvements of the analysis software (PandaRoot), that are vital for this study. These improvements include a reworked clustering algorithm for the electromagnetic calorimeter (EMC) and an optimized monte carlo matching for neutral particles. The second part of this document is about studying the proton structure. A multidimensional study of the structure function ratio Fsin(?)LU /FUU has been performed for K±, based on the measurement of beam-spin asymmetries. It uses the high statistics data recorded with the CLAS12 spectrometer at Jefferson Laboratory. Fsin(?)LU is a twist-3 quantity that provides information about the quark gluon correlations in the proton. This document will present for the first time a simultaneous analysis of two kaon channels over a large kinematic range of z, xB , PT and Q2 with virtualities Q2 ranging from 1 GeV2 up to 8 GeV2 using machine learning techniques for improved particle identification.

Kripko, Aron↗

Machine Learning the COSMO Model for Predicting Thermodynamics of Electrolyte Mixtures

Bottom-up design of electrolyte mixtures for battery systems requires predicting macro thermodynamic properties from molecular constituents. For instance, molten salt electrolyte batteries require conditions far above room temperature to operate. Therefore, discovering mixtures with increasingly lower eutectic melting points is desirable. A model that can approximate chemical activity is a valuable tool to search through the vast compositional design space. Machine learning can predict properties of materials such as vibrational free energies, electronic energy gaps, and thermal conductivities. Moreover, they can learn physical models such as interatomic potentials. The COSMO-SAC model uses theory and empirical parameterization to predict liquid-vapor and liquid-solid properties using first-principles calculations. However, obtaining activity coefficients required for parameterizing the COSMO-SAC model is costly and limited to a select chemical space. In this work, we explored if machine learning methods could improve the COSMO-SAC model and bridge density functional theory calculations to liquid phase thermodynamic properties. Our data-driven approach uses existing databases for sigma-profiles of organic solvents and reconciles their methodological differences via ensemble averaging. First, an optimal machine learning model is constructed for each dataset. Our machine learning algorithms use the sigma-profile as an input feature to predict binary mixtures' activity coefficients using multi-output regression. Each dataset uses different choices of functionals, methods, and basis sets. Therefore, our ensemble model attempts to predict corrected activity coefficients given the combination of all the model outputs. The activity coefficients used for training are generated using the COSMO-SAC model. This approach enables the extraction of meaningful information from the existing datasets to improve the COSMO-SAC model for obtaining thermodynamic properties of electrolyte mixtures. With the liquid phase activities, we can identify electrolyte mixtures that meet desired phase equilibria conditions.

Thermodynamics↗

Clostridioides difficile Toxin A Remodels Membranes and Mediates DNA Entry Into Cells to Activate Toll-Like Receptor 9 Signaling

Background & Aims: Clostridioides difficile toxin A (TcdA) activates the innate immune response. TcdA co-purifies with DNA. Toll-like receptor 9 (TLR9) recognizes bacterial DNA to initiate inflammation. We investigated whether DNA bound to TcdA activates an inflammatory response in murine models of C difficile infection via activation of TLR9. Methods: We performed studies with human colonocytes and monocytes and macrophages from wild-type and TLR9 knockout mice incubated with TcdA or its antagonist (ODN TTAGGG) or transduced with vectors encoding TLR9 or small-interfering RNAs. Cytokine production was measured with enzyme-linked immunosorbent assay. We studied a transduction domain of TcdA (TcdA 57-80 ), which was predicted by machine learning to have cell-penetrating activity and confirmed by synchrotron small-angle X-ray scattering. Intestines of CD1 mice, C57BL6J mice, and mice that express a form of TLR9 that is not activated by CpG DNA were injected with TcdA, TLR9 antagonist, or both. Enterotoxicity was estimated based on loop weight to length ratios. A TLR9 antagonist was tested in mice infected with C difficile. We incubated human colon explants with an antagonist of TLR9 and measured TcdA-induced production of cytokines. Results: The TcdA 57-80 protein transduction domain had membrane remodeling activity that allowed TcdA to enter endosomes. TcdA-bound DNA entered human colonocytes. TLR9 was required for production of cytokines by cultured cells and in human colon explants incubated with TcdA. TLR9 was required in TcdA-induced mice intestinal secretions and in the survival of mice infected by C difficile. Even in a protease-rich environment, in which only fragments of TcdA exist, the TcdA 57-80 domain organized DNA into a geometrically ordered structure that activated TLR9. Conclusions: TcdA from C difficile can bind and organize bacterial DNA to activate TLR9. TcdA and TcdA fragments remodel membranes, which allows them to access endosomes and present bacterial DNA to and activate TLR9. Rather than inactivating the ability of DNA to bind TLR9, TcdA appears to chaperone and organize DNA into an inflammatory, spatially periodic structure.

59 BASIC BIOLOGICAL SCIENCES↗

Machine-Learning-Based Rotating Detonation Engine Diagnostics: Evaluation for Application in Experimental Facilities

Real-time monitoring of combustion behavior is a crucial step toward actively controlled rotating detonation engine (RDE) operation in laboratory and industrial environments. Various machine learning methods have been developed to advance diagnostic efficiencies from conventional postprocessing efforts to real-time methods. Here this work evaluates and compares conventional techniques alongside convolutional neural network (CNN) architectures trained in previous studies, including image classification, object detection, and time series classification, according to metrics affecting diagnostic feasibility, external applicability, and performance. Real-time, capable diagnostics are deployed and evaluated using an altered experimental setup. Image-based CNNs are applied to externally provided images to approximate dataset restrictions. Image classification using high-speed chemiluminescence images and time series classification using high-speed flame ionization and pressure measurements achieve classification speeds enabling real-time diagnostic capabilities, averaging laboratory-deployed diagnostic feedback rates of 4–5 Hz. Object detection achieves the most refined resolution of 20 μs in postprocessing. Image and time series classification require the additional correlation of sensor data, extending their time-step resolutions to 80 ms. Comparisons show that no single diagnostic approach outperforms its competitors across all metrics. This finding justifies the need for a machine learning portfolio containing a host of networks to address specific needs throughout the RDE research community.

33 ADVANCED PROPULSION SYSTEMS↗

Developing a Deep Learning-Computer Vision Framework to Monitor Avian Interactions with Solar Energy Facility Infrastructure (Final Technical Report)

The project addressed an inability to monitor avian interactions with photovoltaic (PV) solar energy facilities necessary for understanding PV solar impacts on birds. In the project, machine-vision technology that continuously monitors avian activities at PV solar facilities was developed. The technology includes four machine-learning (ML) models, each of which accomplishes a specific task in detecting birds and classifying their activities in live or recorded videos—detecting and tracking moving objects, differentiating birds from other objects, detecting bird collisions with solar panels, and classifying non-collision bird activities around PV facilities. Major project outcomes include adoption by two of DOE SETO’s SolWEB projects, providing novel observational data on birds to promote co-location of PV solar development and habitat conservation, known as ecovoltaics.

14 SOLAR ENERGY↗

Deep active learning for classifying cancer pathology reports

Abstract Background Automated text classification has many important applications in the clinical setting; however, obtaining labelled data for training machine learning and deep learning models is often difficult and expensive. Active learning techniques may mitigate this challenge by reducing the amount of labelled data required to effectively train a model. In this study, we analyze the effectiveness of 11 active learning algorithms on classifying subsite and histology from cancer pathology reports using a Convolutional Neural Network as the text classification model. Results We compare the performance of each active learning strategy using two differently sized datasets and two different classification tasks. Our results show that on all tasks and dataset sizes, all active learning strategies except diversity-sampling strategies outperformed random sampling, i.e., no active learning. On our large dataset (15K initial labelled samples, adding 15K additional labelled samples each iteration of active learning), there was no clear winner between the different active learning strategies. On our small dataset (1K initial labelled samples, adding 1K additional labelled samples each iteration of active learning), marginal and ratio uncertainty sampling performed better than all other active learning techniques. We found that compared to random sampling, active learning strongly helps performance on rare classes by focusing on underrepresented classes. Conclusions Active learning can save annotation cost by helping human annotators efficiently and intelligently select which samples to label. Our results show that a dataset constructed using effective active learning techniques requires less than half the amount of labelled data to achieve the same performance as a dataset constructed using random sampling.

59 BASIC BIOLOGICAL SCIENCES↗

Machine-learning Solution for Automatic Spacesuit Motion Recognition and Measurement from Conventional Video

Extravehicular Activity (EVA) spacesuits exhibit unique movement patterns due to their design characteristics. Mobility assessments using traditional motion capture systems are cost prohibitive and not feasible for some training conditions (e.g., simulated lunar outdoor terrain). This paper aims to present the ongoing development of machine learning solutions to quantify suit motions from conventional videos without special sensors or hardware. Preliminary work into this field was promising but given the fast growth in deep/machine learning technologies, external expertise was sought from open-source communities. Partnerships were formed with the NASA JSC Center of Excellence for Collaborative Innovation (CoCEI) and an execution crowdsourcing platform partner to solicit machine learning framework developments from external contenders. NASA provided contenders with images and video clips of spacesuits with simultaneously measured motion capture data during EVA simulation tasks. The contenders used this data to train and develop generalized algorithms to predict motions. At the end of the crowdsourcing event, the top five solutions were selected from 250 submissions. Each submission was tested and scored using video clips not previously disclosed to the contenders. The weighted scoring metrics measured how well the algorithm detected the suit shape, the 2D suit joint detection accuracy, and 3D joint detection accuracy. The winning solution was able to achieve roughly 85% prediction accuracy. Overall, the algorithms could efficiently detect various types of spacesuits and motions across different EVA environments such as the NASA Active Response Gravity Offload System (ARGOS). After continued improvements and validation, the fully developed system will enable EVA stakeholders to quantify suit kinematic patterns, which can help optimize suit, hardware, and task designs.

Linh Vu↗

Machine Learning-Accelerated First-Principles Molecular Dynamics Reveals C–C Coupling Mechanisms toward Ethylene on Cu(100)

Here, the Cu(100) termination has been identified as the most effective facet for converting CO and CO 2 into ethylene. To enhance both the activity and selectivity of ethylene production, we perform machine-learning-accelerated, first-principles molecular dynamics simulations at 298 K in an explicit solvent at pH 7 to elucidate the C–C coupling mechanism─the critical reaction step in forming C 2+ products. Among the six potential C–C coupling pathways, the most feasible are CO* dimerization and CO – CHO* and CHO* – CHO* couplings. Using the computational hydrogen electrode method, we demonstrate that all three pathways are equally accessible at −0.6 V vs RHE. At a potential below −1.0 V vs RHE, the thermodynamic barriers for the CO – CHO* and CHO* – CHO* pathways become negligible. Our computational findings explain the experimental observations, particularly the absence of C 2+ products above −0.4 V vs RHE and the peaks in ethylene production near −0.6 and −1.0 V vs RHE. Since CHO* acts as a key intermediate common to both C–C coupling and CH 4 formation, we propose that suppressing CHO* hydrogenation would inhibit CH 4 pathways, thereby maximizing ethylene selectivity.

CO2 reduction↗

Machine-learning Solution for Automatic Spacesuit Motion Recognition and Measurement from Conventional Video

Extravehicular Activity (EVA) spacesuits exhibit unique movement patterns due to their design characteristics. Mobility assessments using traditional motion capture systems are cost prohibitive and not feasible for some training conditions (e.g., simulated lunar outdoor terrain). This paper aims to present the ongoing development of machine learning solutions to quantify suit motions from conventional videos without special sensors or hardware. Given the fast growth in deep/machine learning technologies, external expertise was sought from open-source communities. This was expected to accelerate development and provide more cost-effective, time-saving solutions. This work was selected for a NASA Crowdsourcing project through an agency-wide solicitation. Partnerships were formed with the NASA JSC Center of Excellence for Collaborative Innovation and an execution crowdsourcing platform partner to solicit framework developments from external contenders. NASA provided contenders with video clips of spacesuits and simultaneously measured motion capture data during EVA simulation tasks. The contenders used this data to train and develop generalized algorithms to predict motions. At the end of the crowdsourcing event, five solutions were selected from 250 submissions. Each submission was tested and scored using video clips not previously disclosed to the contenders. The scoring metrics measured how well the algorithm detected the suit shape, the 2D suit joint detection accuracy, and 3D joint detection accuracy. The winning solution was able to achieve roughly 85% prediction accuracy (weighted combination of scoring metrics). Overall, the algorithms could efficiently detect various types of spacesuits and motions across different EVA simulation environments such as the Neutral Buoyancy Lab (NBL). However, 3D joint identification is less reliable when parts of the suit were obstructed in the image. After continued improvements and validation, the fully developed system will enable EVA stakeholders to quantify suit kinematic patterns, which can help optimize suit, hardware, and task designs.

Linh Vu↗

Graphene-supported single atom catalysts for high performance lithium-oxygen batteries

The optimal choice of d-block metals in single atom catalysts (SACs) is crucial for designing efficient electrocatalysts for activating the Oxygen reduction reaction (ORR)/ Oxygen evolution reaction (OER) in lithiumoxygen batteries (LOBs). Herein, we used the Quantum Mechanics methods to understand the origin of reactivity for a series of 16 d-block metals supported on nitrogen-doped graphene as SACs for ORR and OER in LOBs. Based on the Gibbs free energy calculations, we found that among the 16 SACs investigated, Zn-SAC exhibits the highest electrochemical activity with the lowest overpotential of 0.17 V. Here we then used machine learning (ML) to develop an intrinsic descriptor, phi, that correlates the catalytic activity with electronic and chemical properties of the catalytic centers at the M-N 4 active site on graphene surface. We established a linear relationship between phi and the catalytic activity that provides guidance for designing efficient SACs for electrocatalysis in LOBs. To validate these predictions, we report electrochemical measurements showing that Zn-SAC exhibits an ultra-stable cyclability with reduced overpotentials over Mo-SAC and nitrogen-doped graphene (NG), confirming our theoretical prediction. This fundamental work provides a deep understanding on the rational design of efficient SACs for OER/ ORR in LOBs.

25 ENERGY STORAGE↗

Machine Learning Based Resilience Testing of an Address Randomization Cyber Defense

Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]↗

Co-design Center for Exascale Machine Learning Technologies (ExaLearn)

We report rapid growth in data, computational methods, and computing power is driving a remarkable revolution in what variously is termed machine learning (ML), statistical learning, computational learning, and artificial intelligence. In addition to highly visible successes in machine-based natural language translation, playing the game Go, and self-driving cars, these new technologies also have profound implications for computational and experimental science and engineering, as well as for the exascale computing systems that the Department of Energy (DOE) is developing to support those disciplines. Not only do these learning technologies open up exciting opportunities for scientific discovery on exascale systems, they also appear poised to have important implications for the design and use of exascale computers themselves, including high-performance computing (HPC) for ML and ML for HPC. The overarching goal of the ExaLearn co-design project is to provide exascale ML software for use by Exascale Computing Project (ECP) applications, other ECP co-design centers, and DOE experimental facilities and leadership class computing facilities.

97 MATHEMATICS AND COMPUTING↗

Back-to-back high category atmospheric river landfalls occur more often on the west coast of the United States

Abstract The catastrophic December 2022-January 2023 nine atmospheric rivers in California underscore the urgent need to better understand such high-risk weather extremes. Here we applied a machine learning clustering tool to understand the activity of atmospheric river clusters. Reanalysis results show that clusters with high density, that is the time fraction under atmospheric river conditions within a cluster, exhibit more frequent high-category atmospheric rivers, alongside an increased likelihood for extreme precipitation and severe land surface responses. The key circulation patterns of atmospheric river clusters are primarily attributed to subseasonal variability. Furthermore, the occurrence and density of atmospheric river clusters are modulated by the daily variability of the geopotential height field. Climate model projections suggest that atmospheric river clusters with higher density and higher categories will be more frequent as warming level increases. Our findings emphasize the important role of atmospheric river clusters in the development of climate adaptation and resilience strategies.

54 ENVIRONMENTAL SCIENCES↗

Philympics 2021: Prophage Predictions Perplex Programs

Most bacterial genomes contain integrated bacteriophages—prophages—in various states of decay. Many are active and able to excise from the genome and replicate, while others are cryptic prophages, remnants of their former selves. Over the last two decades, many computational tools have been developed to identify the prophage components of bacterial genomes, and it is a particularly active area for the application of machine learning approaches. However, progress is hindered and comparisons thwarted because there are no manually curated bacterial genomes that can be used to test new prophage prediction algorithms. Here, we present a library of gold-standard bacterial genome annotations that include manually curated prophage annotations, and a computational framework to compare the predictions from different algorithms. We use this suite to compare all extant stand-alone prophage prediction algorithms to identify their strengths and weaknesses. We provide a FAIR dataset for prophage identification, and demonstrate the accuracy, precision, recall, and f 1 score from the analysis of seven different algorithms for the prediction of prophages. We discuss caveats and concerns in this analysis and how those concerns may be mitigated.

Roach, Michael J.↗

QuCNN : A Quantum Convolutional Neural Network with Entanglement Based Backpropagation

Quantum Machine Learning continues to be a highly active area of interest within Quantum Computing. Many of these approaches have adapted classical approaches to the quantum settings, such as QuantumFlow, etc. We push forward this trend, and demonstrate an adaption of the Classical Convolutional Neural Networks to quantum systems - namely QuCNN. QuCNN is a parameterised multi-quantum-state based neural network layer computing similarities between each quantum filter state and each quantum data state. With QuCNN, back propagation can be achieved through a single-ancilla qubit quantum routine. QuCNN is validated by applying a convolutional layer with a data state and a filter state over a small subset of MNIST images, comparing the backpropagated gradients, and training a filter state against an ideal target state.

high performance computing, quantum computing↗

Anomalous And Normal High Performance Computing Datacenter Activities

This code provides annotation and an interface to 10+ hours of video activities in a high performance datacenter with 20+ different types of anomalous activities. The purpose is to enable machine learning for video surveillance systems in high performance computing centers. This is the first code of this type addressing the space of high performance computing datacenters.

Anderson, Matthew↗