Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Quantum machine learning for chemistry and physics

Machine learning (ML) has emerged as a formidable force for identifying hidden but pertinent patterns within a given data set with the objective of subsequent generation of automated predictive behavior. In recent years, it is safe to conclude that ML and its close cousin, deep learning (DL), have ushered in unprecedented developments in all areas of physical sciences, especially chemistry. Not only classical variants of ML, even those trainable on near-term quantum hardwares have been developed with promising outcomes. Such algorithms have revolutionized materials design and performance of photovoltaics, electronic structure calculations of ground and excited states of correlated matter, computation of force-fields and potential energy surfaces informing chemical reaction dynamics, reactivity inspired rational strategies of drug designing and even classification of phases of matter with accurate identification of emergent criticality. In this review we shall explicate a subset of such topics and delineate the contributions made by both classical and quantum computing enhanced machine learning algorithms over the past few years. We shall not only present a brief overview of the well-known techniques but also highlight their learning strategies using statistical physical insight. The objective of the review is not only to foster exposition of the aforesaid techniques but also to empower and promote cross-pollination among future research in all areas of chemistry which can benefit from ML and in turn can potentially accelerate the growth of such algorithms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Art of Automation: Translating Electron Microscopy Workflows Into Automated Processes

Acquiring data using a scanning transmission electron microscope (STEM) is a complex, multi-step process. The intricacy of the process depends on the type of sample, composition of the material, desired results of the experiment, resolution requirement and other experimental factors. Each experiment presents unique complications, such as sample drift and contamination, that the microscopist must consider when acquiring data. All these challenges are handled fluidly and expertly by experienced microscopists, but to reach new levels of innovation in material development, including greater reproducibility, throughput, and precision, the automation of these workflows is essential. The initial phase of this work involved translating intuition-based workflows into discrete, programmable steps. Some common key stages in STEM workflows are the initial tuning, scanning the sample for areas of interest, and then acquiring the data. Each stage can be broken further into specific parameter adjustments, such as aberration correction and dwell time optimization, depending on the experiment. When deconstructing various experiments each step was assessed for automation feasibility based on the amount of real time operator decisions. There are steps that lend themselves to automation more readily than others, such as course focusing and sample screening, but there is potential for full automation of all stages with time. As an initial step, an automated montage routine was developed, allowing for the efficient acquisition of large portions of the sample without requiring continuous intervention from the operator. The automation of this small process of the procedure demonstrates the value of this capability. A major challenge in automation arises from discrepancies between commanded, reported and actual stage movements. Using systematic tests, stage movement was quantified. This error can be corrected algorithmically for more accurate workflows in the future. Expanding automation capabilities would result in larger, more efficient data acquisition which allows for more robust statistical analysis. Additionally, this work lays the groundwork for a closed loop system where machine learning algorithms would intake automatically acquired data and make real time decisions. By progressively automating this instrument, this work establishes the foundation for fully automated experimentation in transmission electron microscopy.

97 MATHEMATICS AND COMPUTING↗

A Review on Machine Learning for Neutrino Experiments

Neutrino experiments study the least understood of the Standard Model particles by observing their direct interactions with matter or searching for ultra-rare signals. The study of neutrinos typically requires overcoming large backgrounds, elusive signals, and small statistics. The introduction of state-of-the-art machine learning tools to solve analysis tasks has made major impacts to these challenges in neutrino experiments across the board. Machine learning algorithms have become an integral tool of neutrino physics, and their development is of great importance to the capabilities of next generation experiments. An understanding of the roadblocks, both human and computational, and the challenges that still exist in the application of these techniques is critical to their proper and beneficial utilization for physics applications. This review presents the current status of machine learning applications for neutrino physics in terms of the challenges and opportunities that are at the intersection between these two fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Geothermal Operational Optimization with Machine Learning

The Geothermal Operational Optimization with Machine Learning (GOOML) project has developed a generic and extensible component-based system modeling framework to study complex geothermal fields using a data-driven approach. Through building a digital twin of a geothermal steam field with the GOOML modeling framework, operators can analyze historical and forecasted power production, explore possible steam field configurations, and optimize real world operations, all in a cost-effective digital environment. The GOOML modeling software is based on a historical data-assimilation framework that uses first-principal thermodynamics to model steam field components using historical data, and a forecast framework that uses machine-learning-driven models of steam field components to predict future operations. This modeling framework creates countless new opportunities for digital exploration of steam field design and operations. To date, digital twins have been developed for several steam fields in New Zealand and the United States. These digital twins have been validated by comparing hindcast predictions against historical production data. Field design and operations have been explored using genetic optimization and reinforcement learning. Initial results show compelling and often surprising opportunities for improved design and operation of fields with 2 to 5 percent improvements in annual energy production. GOOML is driving a step-change in geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and a first-of-its-kind intelligent geothermal systems model.

40 EE - Geothermal Technologies Office (EE-4G)↗

A Multi-Analysis Approach for Estimating Regional Health Impacts from the 2017 Northern California Wildfires

In the evening of October 8 and early hours of October 9, 2017, high winds in Northern California downed trees and power lines, igniting some of the most devastating wildfires the state had seen, and within hours unhealthy air quality impacted millions of people. We simulated these air quality conditions using fire detection information from the MODIS, VIIRS, and GOES-16 ABI satellite instruments, and applying a set of three WRF–CMAQ simulations, one data fusion, and three machine learning methods. We investigated using the 5-min available GOES-16 fire detection data to simulate timing of fire activity to allocate emissions hourly for the WRF-CMAQ air quality modeling system. Interestingly, this approach did not necessarily improve results compared to the baseline case, which used a default time profile. However, this approach was key to simulating the initial 12-hr explosive fire activity and smoke impacts. The WRF-CMAQ simulations compared well with observational data for the October 8-15 time period and tended to overestimate concentrations October 16-20. To improve these results, we applied three machine learning algorithms. We also had a unique opportunity to evaluate results with temporary monitors deployed specifically for wildfires, and performance was markedly different. For example, at the permanent monitoring locations, the WRF-CMAQ simulations had a Pearson correlation of 0.65, and the data fusion approach improved this (Pearson correlation = 0.95), while at the temporary monitor locations across the WRF-CMAQ, data fusion, and machine learning datasets, the best Pearson correlation was 0.5. The data fusion and machine learning results were biased low and WRF-CMAQ results were biased high. Finally, we applied the optimized PM2.5 exposure estimate in a short-term exposure-response function. Total estimated mortality attributable to PM2.5 exposure during the smoke episode was 83 (95% confidence interval: 0, 196) with 47% of these deaths attributable to wildland fire smoke.

O'Neill, Susan↗

Neural network accelerator for quantum control

Efficient quantum control is necessary for practical quantum computing implementations with current technologies. Conventional algorithms for determining optimal control parameters are computationally expensive, largely excluding them from use outside of the simulation. Existing hardware solutions structured as lookup tables are imprecise and costly. By designing a machine learning model to approximate the results of traditional tools, a more efficient method can be produced. Such a model can then be synthesized into a hardware accelerator for use in quantum systems. In this study, we demonstrate a machine learning algorithm for predicting optimal pulse parameters. This algorithm is lightweight enough to fit on a low-resource FPGA and perform inference with a latency of 175 ns and pipeline interval of 5 ns with > 0.99 gate fidelity. In the long term, such an accelerator could be used near quantum computing hardware where traditional computers cannot operate, enabling quantum control at a reasonable cost at low latencies without incurring large data bandwidths outside of the cryogenic environment.

43 PARTICLE ACCELERATORS↗

Science Autonomy for Ocean Worlds Astrobiology: A Perspective

Astrobiology missions to ocean worlds in our solar system must overcome both scientific and technological challenges due to extreme temperature and radiation conditions, long communication times, and limited bandwidth. While such tools could not replace ground-based analysis by science and engineering teams, machine learning algorithms could enhance the science return of these missions through development of autonomous science capabilities. Examples of science autonomy include onboard data analysis and subsequent instrument optimization, data prioritization (for transmission), and real-time decision-making based on data analysis. Similar advances could be made to develop streamlined data processing software for rapid ground-based analyses. Here we discuss several ways machine learning and autonomy could be used for astrobiology missions, including landing site selection, prioritization and targeting of samples, classification of “features” (e.g., proposed biosignatures) and novelties (uncharacterized, “new” features, which may be of most interest to agnostic astrobiological investigations), and data transmission.

ocean worlds↗

Devices and methods for increasing the speed or power efficiency of a computer when performing machine learning using spiking neural networks

A method for increasing a speed and efficiency of a computer when performing machine learning using spiking neural networks. The method includes computer-implemented operations; that is, operations that are solely executed on a computer. The method includes receiving, in a spiking neural network, a plurality of input values upon which a machine learning algorithm is based. The method also includes correlating, for each input value, a corresponding response speed of a corresponding neuron to a corresponding equivalence relationship between the input value to a corresponding latency of the corresponding neuron. Neurons that trigger faster than other neurons represent close relationships between input values and neuron latencies. Latencies of the neurons represent data points used in performing the machine learning. A plurality of equivalence relationships are formed as a result of correlating. The method includes performing the machine learning using the plurality of equivalence relationships.

Vineyard, Craig Michael↗

Conotoxin Prediction: New Features to Increase Prediction Accuracy

Conotoxins are toxic, disulfide-bond-rich peptides from cone snail venom that target a wide range of receptors and ion channels with multiple pathophysiological effects. Conotoxins have extraordinary potential for medical therapeutics that include cancer, microbial infections, epilepsy, autoimmune diseases, neurological conditions, and cardiovascular disorders. Despite the potential for these compounds in novel therapeutic treatment development, the process of identifying and characterizing the toxicities of conotoxins is difficult, costly, and time-consuming. This challenge requires a series of diverse, complex, and labor-intensive biological, toxicological, and analytical techniques for effective characterization. While recent attempts, using machine learning based solely on primary amino acid sequences to predict biological toxins (e.g., conotoxins and animal venoms), have improved toxin identification, these methods are limited due to peptide conformational flexibility and the high frequency of cysteines present in toxin sequences. This results in an enumerable set of disulfide-bridged foldamers with different conformations of the same primary amino acid sequence that affect function and toxicity levels. Consequently, a given peptide may be toxic when its cysteine residues form a particular disulfide-bond pattern, while alternative bonding patterns (isoforms) or its reduced form (free cysteines with no disulfide bridges) may have little or no toxicological effects. Similarly, the same disulfide-bond pattern may be possible for other peptide sequences and result in different conformations that all exhibit varying toxicities to the same receptor or to different receptors. We present here new features, when combined with primary sequence features to train machine learning algorithms to predict conotoxins, that significantly increase prediction accuracy.

collisional cross section↗

Automated identification of dominant physical processes

The identification of processes that locally and approximately dominate dynamical system behavior has enabled significant advances in understanding and modeling nonlinear differential dynamical systems. Conventional methods of dominant process identification involve piecemeal and ad hoc (non-rigorous, informal) scaling analyses to identify dominant balances of governing equation terms and to delineate the spatiotemporal boundaries (boundaries in space and/or time) of each dominant balance. For the first time, we present an objective global measure of the fit of dominant balances to observations, which is desirable for automation, and was previously undefined. Furthermore, we propose a formal definition of the dominant balance identification problem in the form of an optimization problem. Here, we show that the optimization can be performed by various machine learning algorithms, enabling the automatic identification of dominant balances. Our method is algorithm agnostic and it eliminates reliance upon expert knowledge to identify dominant balances which are not known beforehand.

42 ENGINEERING↗

Considerations for using Privacy Preserving Machine Learning Techniques for Safeguards

In international nuclear safeguards, the International Atomic Energy Agency (IAEA) is tasked with inspecting and verifying nuclear facilities and their activities. Data analytics and machine learning to support inspections require large amounts of data that nuclear facility operators may consider proprietary or sensitive, so the IAEA may not have full access. Allowing computation over private data without compromising its security therefore has value for safeguards inspections and analysis. Privacy-preserving machine learning (PPML) consists of security-focused techniques that allow data analytics and machine learning algorithms to run on sensitive data without revealing it. This includes ideas like homomorphic encryption (HE), secure multiparty computation (SMPC), and secure enclaves. HE allows algorithms and mathematical operations to be conducted directly on the encrypted data instead of first decrypting it. With SMPC, multiple entities collaboratively compute over distributed data such that no party is able to directly view any others’ original data. Secure enclaves allow computation to take place in a separate and heavily blocked-off section of a CPU. Techniques like these allow for several potential use cases in which the security of data is essential. With SMPC, machine learning models can be trained over the input data from multiple entities, resulting in a model that all users can benefit from without leaking the input data from any particular entity. With SMPC or a zero-knowledge proof (ZKP), an algorithm returning some single answer or truth value can be run on someone else’s data without ever needing to see that data, potentially allowing for verification or proof of some underlying question. HE can allow for outsourcing computation on data to a hostile or untrusted environment. Although most of the research in this field resides within the health and financial domains, tools from PPML may have similar applications in nuclear safeguards. Allowing the IAEA to compute over proprietary information, such as process models and raw sensor data using PPML techniques, provides the baseline for running complex analytics without needing direct unencrypted access to the underlying data, maintaining its privacy. Important limitations to consider for these techniques include the efficiency and level of security required. The security of HE and SMPC come at the cost of speed—the significant amount of overhead means that algorithms implemented in these protocols and encryption schemes are slower than when run on plaintext. Additionally, several important parameters determine what techniques or protocols are used based on the security requirements. SMPC protocols may need to be selected for resistance against a party that attempts to deviate from the protocol to distort the result or gain access to additional information, and a protocol secure against these attacks may further increase the overhead of the algorithm.

97 MATHEMATICS AND COMPUTING↗

Exploratory analysis and performance prediction of big data transfer in High-performance Networks

Big data transfer in large-scale scientific and business applications is increasingly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) via advance bandwidth reservation. Provisioning agents need to carefully schedule data transfer requests, compute network paths, and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized, could be simply wasted due to the exclusive access during the approved time window, and cause extra overhead and complexity for resource management. This calls for accurate performance prediction to reserve bandwidths that match actual needs and avoid over-provisioning. We employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements collected in the past several years from data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated testbeds. We first analyze the performance patterns in response to a comprehensive list of parameters in end-host systems, network connections, and data transfer applications, which motivate the use of machine learning and also help us identify the effects of latent factors. We then propose threshold- and clustering-based methods to eliminate negative effects of latent factors in data preprocessing and build a robust performance predictor based on customized domain-oriented loss functions. The performance of the proposed methods is verified by extensive experiments using SVR and RFR as well as theoretical analysis of the general performance bound.

97 MATHEMATICS AND COMPUTING↗

ESD: Ethernet Signal Differentiator [Poster]

Can Machine Learning Algorithms be trained to interpret and decode passively observed Automative Ethernet full-duplex signals without access to the original signals transmitted by either endpoint?

97 - MATHEMATICS AND COMPUTING↗

Robust Molecular Predictive Methods for Novel Polymer Discovery and Applications

Polymeric materials are ubiquitous in modern society and they play an instrumental role in almost all industries, undoubtedly including the energy and environment sectors. Increased demand of energy and awareness to sustainability both necessitates the development of novel polymers with enhanced properties. Unfortunately, their structural and behavioral complexity render such discovery challenging and impeded. To address this problem, scientists are developing various computational modeling techniques and leveraging their power to depict the relationship between structural characteristics of polymers and their properties (such as rheological behaviors), and use such prediction to guide the design and syntheses of novel polymeric materials with enhanced performances. Unfortunately, predicting the relationships between polymer structure and composition with rheological properties via atomistic modeling is still a major challenge because of the extended time and length scales involved. Studying dynamic shear viscosity and linear viscoelasticity using molecular models requires capabilities that have been elusive, including representation of large molecular weight chains with an effective internal scale capable of describing entanglement, shear-rates that are in the s-1 scale with accurate quantitative stresses, and chemically-realistic combinations of both homogeneous and heterogeneous systems. Motivated by these unmet challenges, the overall technical objective of this DOE-STTR Phase II project is to develop robust molecular predictive methods for advanced polymer discovery and applications and especially for designing and demonstrating the “smart” polymer-based waterflooding enhanced oil recovery (EOR) process. In particular, we apply state-of-the-art molecular modeling methods developed by our academic partner, Materials Stimulation Center (MSC) at California Institute of Technology (Caltech), to facilitate and accelerate the experimental discovery processes. During the Phase I of this project, we had focused on development and demonstration of the molecular modeling methods to describe rheological properties of non-Newtonian polymer fluids, and to improve our fundamental understandings of shear-thickening mechanism and kinetics. In Phase II, we further apply the theoretical models to guide our experimental programs to improve our design of smart rheology modifier (SRM) polymers and their optimization for EOR. Specifically, we have three objectives in the Phase II study: (1) to further improve out computational modeling methods, coupling with the advanced machine learning algorithms; (2) to develop cost-effective and efficient SRM-flooding process suitable for EOR applications under typical reservoir conditions; and (3) to further explore the application of our molecular predictive models for innovative material discovery in other industrial applications. The recent development of our multiscale predictive framework allows the successful prediction of rheological properties from the chemical structure for polymers of experimentally relevant molecular weights, and provides an in-silico machine learning engine for screening novel compositions and structures with optimized non-Newtonian response, required for both shear-thinning and shear-thickening applications. Our framework provides: (1) procedures and tools for systematic coarsening from atomistic models and reverse mapping of coarse-grain models to atomistic, (2) unique ab initio methods to characterize the atomistic origin of colloidal and interfacial interactions and phenomena, (3) systematic structure and composition builders based on practical descriptors that drive rheological changes in polymer melts and diluted polymer mixtures, (4) a rheological properties engine capable of predicting viscosity in the zero-shear limit and under realistic dynamic conditions (for shear-rates commensurate with experiments) for large heterogeneous systems, (5) coarse-grain force fields with improved non-bond descriptions based on accurate quantum mechanics, (6) an in-silico screening machine learning engine that feeds from the systematic model builders to cover the descriptors search space, computes the rheological properties from converged trajectories spanning sub-milliseconds and ranks them for each structure/composition using an automated viscosity-vs-shear rate fitness function that can be tuned for shear-thickening, shear-thinning and other rheological responses.

02 PETROLEUM↗

Overcoming the field-of-view to diameter trade-off in microendoscopy via computational optrode-array microscopy

High-resolution microscopy of deep tissue with large field-of-view (FOV) is critical for elucidating organization of cellular structures in plant biology. Microscopy with an implanted probe offers an effective solution. However, there exists a fundamental trade-off between the FOV and probe diameter arising from aberrations inherent in conventional imaging optics (typically, FOV < 30% of diameter). Here, we demonstrate the use of microfabricated non-imaging probes (optrodes) that when combined with a trained machine-learning algorithm is able to achieve FOV of 1x to 5x the probe diameter. Further increase in FOV is achieved by using multiple optrodes in parallel. With a 1 × 2 optrode array, we demonstrate imaging of fluorescent beads (including 30 FPS video), stained plant stem sections and stained living stems. Our demonstration lays the foundation for fast, high-resolution microscopy with large FOV in deep tissue via microfabricated non-imaging probes and advanced machine learning.

59 BASIC BIOLOGICAL SCIENCES↗

A Census of Young Stellar Objects in Two Line-of-Sight Star-Forming Regions Toward IRAS 22147+5948 in the Outer Galaxy

Context. Star formation in the outer Galaxy, namely, outside of the Solar circle, has not been extensively studied in part due to the low CO brightness of the molecular clouds linked with the negative metallicity gradient. Recent infrared surveys provide an overview of dust emission in large sections of the Galaxy, but they suffer from cloud confusion and poor spatial resolution at far-infrared wavelengths. Aims. We aim to develop a methodology to identify and classify young stellar objects (YSOs) in star-forming regions in the outer Galaxy and use it to resolve a long-standing disparity in terms of the distance and evolutionary status of IRAS 22147+5948. Methods. We used a support vector machine learning algorithm to complement standard color–color and color–magnitude diagrams in our search for YSOs in the IRAS 22147 region, based on publicly available data from the Spitzer Mapping of the Outer Galaxy survey. The agglomerative hierarchical clustering algorithm was used to identify clusters. Then the physical properties of individual YSOs were calculated. The distances were determined using CO 1–0 from the Five College Radio Astronomy Observatory survey. Results. We identified 13 Class I and 13 Class II YSO candidates using the color–color diagrams, along with an additional 2 and 21 sources, respectively, using the applied machine learning techniques. The spectral energy distributions of 23 sources were modeled with a star and a passive disk, corresponding to Class II objects. The models of three sources include envelopes that are typical for Class I objects. The objects were grouped into two clusters located at a distance of 2:2 kpc and 5 clusters at 5:6 kpc. The spatial extent of CO, radio continuum, and dust emission confirms the origin of YSOs in two distinct star-forming regions along a similar line of sight. Conclusions. The outer Galaxy may serve as a unique laboratory for exploring star formation across environments, on the condition that complementary methods and ancillary data are used to properly account for cloud confusion and distance uncertainties.

Agata Karska↗

Application of machine learning in the determination of impact parameter in the 132 Sn+ 124 Sn system

Here, 132 Sn + 124 Sn collisions at a beam energy of 270 MeV/nucleon were performed at the Radioactive Isotope Beam Factory (RIBF) in RIKEN to investigate the nuclear equation of state. Reconstructing the impact parameter is one of the important tasks in the experiment as it relates to many observable. In this work, we employ three commonly used algorithms in machine learning, the artificial neural network (ANN), the convolutional neural network (CNN), and the light gradient boosting machine (LightGBM), to determine the impact parameter by analyzing either the charged particle spectra or several features simulated with events from the ultrarelativistic quantum molecular dynamics (UrQMD) model. To closely imitate experimental data and investigate the generalizability of the trained machine learning algorithms, incompressibility of nuclear equation of state and the in-medium nucleon-nucleon cross sections are varied in the UrQMD model to generate the training data. The mean absolute error Δb between the true and the predicted impact parameter is smaller than 0.45 fm if training and testing sets are sampled from the UrQMD model with the same parameter set. However, if training and testing sets are sampled with different parameter sets, Δb would increase to 0.8 fm. The generalizability of the trained machine learning algorithms suggests that these machine learning algorithms can be used reliably to reconstruct the impact parameter in experiment.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗