Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “abstract machine model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Equation‐Free Surrogate Modeling of Geophysical Flows at the Intersection of Machine Learning and Data Assimilation

Abstract There is a growing interest in developing data‐driven reduced‐order models for atmospheric and oceanic flows that are trained on data obtained either from high‐resolution simulations or satellite observations. The data‐driven models are non‐intrusive in nature and offer significant computational savings compared to large‐scale numerical models. These low‐dimensional models can be utilized to reduce the computational burden of generating forecasts and estimating model uncertainty without losing the key information needed for data assimilation (DA) to produce accurate state estimates. This paper aims at exploring an equation‐free surrogate modeling approach at the intersection of machine learning and DA in Earth system modeling. With this objective, we introduce an end‐to‐end non‐intrusive reduced‐order modeling (NIROM) framework equipped with contributions in modal decomposition, time series prediction, optimal sensor placement, and sequential DA. Specifically, we use proper orthogonal decomposition (POD) to identify the dominant structures of the flow, and a long short‐term memory network to model the dynamics of the POD modes. The NIROM is integrated within the deterministic ensemble Kalman filter (DEnKF) to incorporate sparse and noisy observations at optimal sensor locations obtained through QR pivoting. The feasibility and the benefit of the proposed framework are demonstrated for the NOAA Optimum Interpolation Sea Surface Temperature (SST) V2 data set. Our results indicate that the NIROM is stable for long‐term forecasting and can model dynamics of SST with a reasonable level of accuracy. Furthermore, the prediction accuracy of the NIROM gets improved by almost one order of magnitude by the DEnKF algorithm.

Pawar, Suraj↗

Synthetic data generation for machine learning model training for energy theft scenarios using cosimulation

Abstract Technical and non‐technical losses in distribution circuits result in significant economic costs to power utilities. One type of non‐technical loss is energy theft by various means including illegal tapping of feeders, bypassing the meter, and billing fraud. These losses are usually hard to detect, and can remain undetected for long periods of time. Machine learning models have been proven effective in detecting these conditions, but rely on the availability of large, good‐quality training data sets. The problem is exacerbated by the imbalanced nature of data related to these conditions—energy theft, though costly, is very rare. The available data sets generally have very few samples of theft with most of the data pertaining to normal operation. Such data sets are generally not suitable to train machine learning models. In this paper, an overview of energy theft detection techniques, the challenges with their data needs, and the limitations of current techniques to bridge such data limitations is presented. A co‐simulation framework is proposed to generate reliable training data for machine learning algorithms for theft detection. An example scenario is presented and a machine learning model is built to detect certain kinds of energy theft.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Spatially Local Surrogate Modeling of Subgrid-Scale Effects in Idealized Atmospheric Flows: A Deep Learned Approach Using High-Resolution Simulation Data

Abstract We introduce a machine learned surrogate model from high-resolution simulation data to capture the subgrid-scale effects in dry, stratified atmospheric flows. We use deep neural networks (NNs) to model the spatially local state differences between a coarse-resolution simulation and a high-resolution simulation. The setup enables the capture of both dissipative and antidissipative effects in the state differences. The NN model is able to accurately capture the state differences in offline tests outside the training regime. In online tests intended for production use, the NN-coupled coarse simulation has higher accuracy over a significant period of time compared to the coarse-resolution simulation without any correction. We provide evidence of the capability of the NN model to accurately capture high-gradient regions in the flow field. With the accumulation of the errors, the NN-coupled simulation becomes computationally unstable after approximately 90 coarse simulation time steps. Insights gained from these surrogate models further pave the way for formulating stable, complex, physics-based spatially local NN models which are driven by traditional subgrid-scale turbulence closure models. Significance Statement Flows in the atmosphere are highly chaotic and turbulent, comprising flow structures of broad scales. For effective computational modeling of atmospheric flows, the effects of the small- and large-scale structures need to be captured by the simulations. Capturing the small-scale structures requires fine-resolution simulations. Even with the current state-of-the-art supercomputers, it can be prohibitively expensive to simulate these flows when computed for the entire earth over climate time scales. Thus, it is necessary to focus on the larger-scale structures using a coarse-resolution simulation while capturing the effects of the smaller-scale structures using some parameterization (approximation) scheme and incorporating it into the coarse-resolution simulation. We use machine learning to model the effects of the small-scale structures (subgrid-scale effects) in atmospheric flows. Data from a fine-resolution simulation is used to compute the missing subgrid-scale effects in coarse-resolution simulations. We then use machine learning models to approximate these differences between the coarse- and fine-resolution simulations. We see improved accuracy for the coarse-resolution simulations when corrected using these machine learned models.

54 ENVIRONMENTAL SCIENCES↗

Machine learning models for segmentation and classification of cyanobacterial cells

Abstract Timelapse microscopy has recently been employed to study the metabolism and physiology of cyanobacteria at the single-cell level. However, the identification of individual cells in brightfield images remains a significant challenge. Traditional intensity-based segmentation algorithms perform poorly when identifying individual cells in dense colonies due to a lack of contrast between neighboring cells. Here, we describe a newly developed software package called Cypose which uses machine learning (ML) models to solve two specific tasks: segmentation of individual cyanobacterial cells, and classification of cellular phenotypes. The segmentation models are based on the Cellpose framework, while classification is performed using a convolutional neural network named Cyclass. To our knowledge, these are the first developed ML-based models for cyanobacteria segmentation and classification. When compared to other methods, our segmentation models showed improved performance and were able to segment cells with varied morphological phenotypes, as well as differentiate between live and lysed cells. We also found that our models were robust to imaging artifacts, such as dust and cell debris. Additionally, the classification model was able to identify different cellular phenotypes using only images as input. Together, these models improve cell segmentation accuracy and enable high-throughput analysis of dense cyanobacterial colonies and filamentous cyanobacteria.

Huffine, Clair A.↗

Multivariate Machine Learning Models of Nanoscale Porosity from Ultrafast NMR Relaxometry

Abstract Nanoporous materials are of great interest in many applications, such as catalysis, separation, and energy storage. The performance of these materials is closely related to their pore sizes, which are inefficient to determine through the conventional measurement of gas adsorption isotherms. Nuclear magnetic resonance (NMR) relaxometry has emerged as a technique highly sensitive to porosity in such materials. Nonetheless, streamlined methods to estimate pore size from NMR relaxometry remain elusive. Previous attempts have been hindered by inverting a time domain signal to relaxation rate distribution, and dealing with resulting parameters that vary in number, location, and magnitude. Here we invoke well‐established machine learning techniques to directly correlate time domain signals to BET surface areas for a set of metal‐organic frameworks (MOFs) imbibed with solvent at varied concentrations. We employ this series of MOFs to establish a correlation between NMR signal and surface area via partial least squares (PLS), following screening with principal component analysis, and apply the PLS model to predict surface area of various nanoporous materials. This approach offers a high‐throughput, non‐destructive way to assess porosity in c.a. one minute. We anticipate this work will contribute to the development of new materials with optimized pore sizes for various applications.

Fricke, Sophia N.↗

Multivariate Machine Learning Models of Nanoscale Porosity from Ultrafast NMR Relaxometry

Abstract Nanoporous materials are of great interest in many applications, such as catalysis, separation, and energy storage. The performance of these materials is closely related to their pore sizes, which are inefficient to determine through the conventional measurement of gas adsorption isotherms. Nuclear magnetic resonance (NMR) relaxometry has emerged as a technique highly sensitive to porosity in such materials. Nonetheless, streamlined methods to estimate pore size from NMR relaxometry remain elusive. Previous attempts have been hindered by inverting a time domain signal to relaxation rate distribution, and dealing with resulting parameters that vary in number, location, and magnitude. Here we invoke well‐established machine learning techniques to directly correlate time domain signals to BET surface areas for a set of metal‐organic frameworks (MOFs) imbibed with solvent at varied concentrations. We employ this series of MOFs to establish a correlation between NMR signal and surface area via partial least squares (PLS), following screening with principal component analysis, and apply the PLS model to predict surface area of various nanoporous materials. This approach offers a high‐throughput, non‐destructive way to assess porosity in c.a. one minute. We anticipate this work will contribute to the development of new materials with optimized pore sizes for various applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AttentionFire_v1.0: interpretable machine learning fire model for burned-area predictions over tropics

Abstract. African and South American (ASA) wildfires account for more than 70 % of global burned areas and have strong connection to local climate for sub-seasonal to seasonal wildfire dynamics. However, representation of the wildfire–climate relationship remains challenging due to spatiotemporally heterogenous responses of wildfires to climate variability and human influences. Here, we developed an interpretable machine learning (ML) fire model (AttentionFire_v1.0) to resolve the complex controls of climate and human activities on burned areas and to better predict burned areas over ASA regions. Our ML fire model substantially improved predictability of burned areas for both spatial and temporal dynamics compared with five commonly used machine learning models. More importantly, the model revealed strong time-lagged control from climate wetness on the burned areas. The model also predicted that, under a high-emission future climate scenario, the recently observed declines in burned area will reverse in South America in the near future due to climate changes. Our study provides a reliable and interpretable fire model and highlights the importance of lagged wildfire–climate relationships in historical and future predictions.

58 GEOSCIENCES↗

Utilizing ensemble learning for performance and power modeling and improvement of parallel cancer deep learning CANDLE benchmarks

Abstract Machine learning (ML) continues to grow in importance across nearly all domains in modeling to learn from data. Often a tradeoff exists between a model's ability to minimize bias and variance. In this article, we utilize ensemble learning to combine linear, nonlinear, and tree‐/rule‐based ML methods to cope with the bias‐variance tradeoff and result in more accurate models. We use the datasets collected for two parallel cancer deep learning CANDLE benchmarks, NT3 and P1B2, to build performance and power models based on hardware performance counters using single‐object and multiple‐objects ensemble learning to identify the most important counters for improvement on the Cray XC40 Theta at Argonne National Laboratory. Based on the insights from these models, we improve the performance and energy of P1B2 and NT3 by optimizing the deep learning environments TensorFlow, Keras, Horovod, and Python under the huge page size of 8 MB. Experimental results show that ensemble learning not only produces more accurate models but also provides more robust performance counter ranking. We achieve up to 61.15% performance improvement and up to 62.58% energy saving for P1B2 and up to 55.81% performance improvement and up to 52.60% energy saving for NT3 on up to 24,576 cores.

Wu, Xingfu↗

Automated algorithms to build active galactic nucleus classifiers

ABSTRACT We present a machine learning model to classify active galactic nuclei (AGNs) and galaxies (AGN-galaxy classifier) and a model to identify type 1 (optically unabsorbed) and type 2 (optically absorbed) AGN (type 1/2 classifier). We test tree-based algorithms, using training samples built from the X-ray Multi-Mirror Mission–Newton (XMM–Newton) catalogue and the Sloan Digital Sky Survey (SDSS), with labels derived from the SDSS survey. The performance was tested making use of simulations and of cross-validation techniques. With a set of features including spectroscopic redshifts and X-ray parameters connected to source properties (e.g. fluxes and extension), as well as features related to X-ray instrumental conditions, the precision and recall for AGN identification are 94 and 93 per cent, while the type 1/2 classifier has a precision of 74 per cent and a recall of 80 per cent for type 2 AGNs. The performance obtained with photometric redshifts is very similar to that achieved with spectroscopic redshifts in both test cases, while there is a decrease in performance when excluding redshifts. Our machine learning model trained on X-ray features can accurately identify AGN in extragalactic surveys. The type 1/2 classifier has a valuable performance for type 2 AGNs, but its ability to generalize without redshifts is hampered by the limited census of absorbed AGN at high redshift.

Falocco, S. (ORCID:0000000299841103)↗

Improving BC Mixing State and CCN Activity Representation With Machine Learning in the Community Atmosphere Model Version 6 (CAM6)

Abstract Representing mixing state of black carbon (BC) is challenging for global climate models (GCMs). The Community Atmosphere Model version 6 (CAM6) with the four‐mode version of the Modal Aerosol Module (MAM4) represents aerosols as fully internal mixtures with uniform composition within each aerosol mode, resulting in high degree of internal mixing of BC with non‐BC species and large mass ratio of coating to BC ( R BC , the mass ratio of non‐BC species to BC in BC‐containing particles). To improve BC mixing state representation, we coupled a machine learning (ML) model of BC mixing state index trained on particle‐resolved simulations to the CAM6 with MAM4 (MAM4‐ML). In MAM4‐ML, we use R BC to partition accumulation mode particles into two new modes, BC‐free particles and BC‐containing particles. We adjust R BC to make the modeled BC mixing state index ( χ mode ) match the one predicted by the ML model ( χ ML ). On a global average, the mass fraction of BC‐containing particles in accumulation mode decreases from 100% (MAM4‐default) to 48% (MAM4‐ML). The globally averaged χ mode decreases from 78% (MAM4‐default) to 63% (MAM4‐ML, 19% reduction) and agrees well with χ ML (66%). The R BC decreases by 52% for accumulation mode and better agrees with observations. The hygroscopicity drops by 9% for BC‐containing particles in accumulation mode, leading to a 20% reduction in the BC activation fraction. The surface BC concentration increases most (6.9%) in the Arctic, and the BC burden increases by 4%, globally. Our study highlights the application of the ML model for improving key aerosol processes in GCMs.

Meteorology & Atmospheric Sciences↗

Automated Verification of Specifications with Typestates and Access Permissions

We propose an approach to formally verify Plural specifications based on access permissions and typestates, by model-checking automatically generated abstract state-machines. Our exhaustive approach captures all the possible behaviors of abstract concurrent programs implementing the specification. We describe the formal methodology employed by our technique and provide an example as proof of concept for the state-machine construction rules. The implementation of a fully automated algorithm to generate and verify models, currently underway, provides model checking support for the Plural tool, which currently supports only program verification via data flow analysis (DFA).

Siminiceanu, Radu I.↗