Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Predicting protein functions from redundancies in large-scale protein interaction networks

Interpreting data from large-scale protein interaction experiments has been a challenging task because of the widespread presence of random false positives. Here, we present a network-based statistical algorithm that overcomes this difficulty and allows us to derive functions of unannotated proteins from large-scale interaction data. Our algorithm uses the insight that if two proteins share significantly larger number of common interaction partners than random, they have close functional associations. Analysis of publicly available data from Saccharomyces cerevisiae reveals >2,800 reliable functional associations, 29% of which involve at least one unannotated protein. By further analyzing these associations, we derive tentative functions for 81 unannotated proteins with high certainty. Our method is not overly sensitive to the false positives present in the data. Even after adding 50% randomly generated interactions to the measured data set, we are able to recover almost all (approximately 89%) of the original associations.

Proteins/chemistry/metabolism↗

Characterization and thermometry of dissipatively stabilized steady states

In this work we study the properties of dissipatively stabilized steady states of noisy quantum algorithms, exploring the extent to which they can be well approximated as thermal distributions, and proposing methods to extract the effective temperature T. We study an algorithm called the relaxational quantum eigensolver (RQE), which is one of a family of algorithms that attempt to find ground states and balance error in noisy quantum devices. In RQE, we weakly couple a second register of auxiliary ‘shadow’ qubits to the primary system in Trotterized evolution, thus engineering an approximate zero-temperature bath by periodically resetting the auxiliary qubits during the algorithm’s runtime. Balancing the infinite temperature bath of random gate error, RQE returns states with an average energy equal to a constant fraction of the ground state. We probe the steady states of this algorithm for a range of base error rates, using several methods for estimating both T and deviations from thermal behavior. In particular, we both confirm that the steady states of these systems are often well-approximated by thermal distributions, and show that the same resources used for cooling can be adopted for thermometry, yielding a fairly reliable measure of the temperature. These methods could be readily implemented in near-term quantum hardware, and for stabilizing and probing Hamiltonians where simulating approximate thermal states is hard for classical computers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Informing nuclear physics via machine learning methods with differential and integral experiments

Information from differential nuclear-physics experiments and theory is often too uncertain to accurately define nuclear-physics observables such as cross sections or energy spectra. Integral experimental data, representing the applications of these observables, are often more precise but depend simultaneously on too many of them to unambiguously identify issues in the observable with human expert analysis alone. Here, we explore how we can leverage physics knowledge gained from differential experimental data, nuclear theory, integral experiments, and neutron-transport calculations to better understand nuclear-physics observables in the context of the application area represented by integral experiments. We support this task with machine-learning methods to discern trends in a large amount of convoluted data. Differential and integral information was used in an analysis augmented by the random forest and the Shapley additive explanations metric. We chose as an application area one that is represented by criticality measurements and pulsed-sphere neutron-leakage spectra. We show one representative example ( 241 Pu fission observables) where the combination of differential and integral information allowed to resolve issues in data representing these observables. As a starting point, the machine learning (ML) algorithms highlighted several observables as leading potentially to bias in simulating integral experiments. Differential information, paired with sensitivity to integral quantities, allowed us then to pinpoint one specific observable ( 241 Pu fission cross section) as the main driver of bias. The comparison to integral experiments, on the other hand, allowed us to indicate a likely reliable experiment among several discrepant ones for this observables. In other cases (e.g., 239 Pu observables), we were not able to resolve the confounding introduced by integral experiments but instead highlighted the need for targeted new experiments and theory developments to better constrain the nuclear-physics space for the application area represented by integral experiments. We were able to combine information from differential experimental data, nuclear-physics theory, integral experiments, and neutron-transport simulations of the latter experiments with the help of the random forest algorithm and expert judgment. This combination of knowledge allows to improve our description of nuclear-physics observables as applied to a particular application area.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Artificial Diversity and Defense Security (ADDSec)

Artificial Diversity and Defense Security (ADDSec) machine learning algorithms are used to classify and cluster threats so that an appropriate response can be initiated as a mitigation strategy. The package includes an ensemble of machine learning algorithms such as Support Vector Machines, naïve bayes, logistic regression, and random forest that evolve with the data to recognize anomalous behavior at the host and network levels. Inputs into the machine learning algorithms include end host system calls, system utilization, packet captures, and syslog messages. The machine learning algorithms can be retrained based on user defined intervals or on the number of packets received. ADDSEC's threat responses include Internet Protocol (IP) Address randomization, application port number randomization, and application library randomization. The IP randomization implementation is built on top of a Software Defined Networking (SDN) framework. The SDN controller installs flows on each of the SDN switches with randomized source and destination IP addresses. The application port numbers are randomized using iptables. The application library randomization is created with a LLVM compiler. All randomization schemes are transparent to the endpoints on the network. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-3379 O

Cox, RebeccaE.↗

A distributed scheduling algorithm for heterogeneous real-time systems

Much of the previous work on load balancing and scheduling in distributed environments was concerned with homogeneous systems and homogeneous loads. Several of the results indicated that random policies are as effective as other more complex load allocation policies. The effects of heterogeneity on scheduling algorithms for hard real time systems is examined. A distributed scheduler specifically to handle heterogeneities in both nodes and node traffic is proposed. The performance of the algorithm is measured in terms of the percentage of jobs discarded. While a random task allocation is very sensitive to heterogeneities, the algorithm is shown to be robust to such non-uniformities in system components and load.

Zeineldine, Osman↗

Long-term missing value imputation for time series data using deep neural networks

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods ImputeTS and mtsdi. The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.

97 MATHEMATICS AND COMPUTING↗

Mitigating Cascading Outages in Severe Weather Using Simulation-Based Optimization

Severe weather events can trigger cascading power outages and lead to significant losses. In this work, we investigate cascading outage mitigation under severe weather conditions. Given day-ahead weather forecasts and component failure models, we aim to identify a set of power lines that can be hardened to minimize the expected impact of potential cascading outages. Since the expected load shedding cannot be expressed as an explicit function of line hardening decisions and system states, we developed a cascading outage simulator to estimate the expected value of load shedding under various initial weather-related disruption scenarios generated using a weather forecast. To avoid massive enumeration of all possible combinations of line hardening decisions and reduce the simulation efforts, we employed an efficient simulation-based optimization approach that quickly identifies the (near) optimal line hardening decisions in the presence of both large simulation noises due to the highly variable initial disturbances and system states, and significant randomness in the subsequent cascades. Furthermore, the algorithm is also able to utilize parallel computing to dramatically reduce computation time to support decision making in preparation for severe weather conditions. We performed a case study on the Northeast Power Coordinating Council (NPCC) 140-bus system model to demonstrate that our approach can significantly improve power grid resilience to adverse weather events.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An Approach to Bayesian Optimization for Design Feasibility Check on Discontinuous Black-Box Functions

The paper presents a novel approach to applying Bayesian Optimization (BO) in predicting an unknown constraint boundary, also representing the discontinuity of an unknown function, for a feasibility check on the design space, thereby representing a classification tool to discern between a feasible and infeasible region. Bayesian optimization is a low-cost black-box global optimization tool in the Sequential Design Methods where one learns and updates knowledge from prior evaluated designs, and proceeds to the selection of new designs for future evaluation. However, BO is best suited to problems with the assumption of a continuous objective function and does not guarantee true convergence when having a discontinuous design space. This is because of the insufficient knowledge of the BO about the nature of the discontinuity of the unknown true function. In this paper, we have proposed to predict the location of the discontinuity using a BO algorithm on an artificially projected continuous design space from the original discontinuous design space. The proposed approach has been implemented in a thin tube design with the risk of creep-fatigue failure under constant loading of temperature and pressure. The stated risk depends on the location of the designs in terms of safe and unsafe regions, where the discontinuities lie at the transition between those regions; therefore, the discontinuity has also been treated as an unknown creep-fatigue failure constraint. The proposed BO algorithm has been trained to maximize sampling toward the unknown transition region, to act as a high accuracy classifier between safe and unsafe designs with minimal training cost. The converged solution has been validated for different design parameters with classification error rate and function evaluations at an average of <1% and ~150, respectively. Finally, the performance of our proposed approach in terms of training cost and classification accuracy of thin tube design is shown to be better than the existing machine learning (ML) algorithms such as Support Vector Machine (SVM), Random Forest (RF), and Boosting.

Engineering↗

Heat Equation Neural Simulation

This code solves a steady state heat equation problem on a wire with random walks implemented using a spiking neural algorithm. SAND2020-12194 M

Reeder, Leah↗

rBahadur: efficient simulation of structured high-dimensional genotype data with applications to assortative mating

Existing methods for generating synthetic genotype data are ill-suited for replicating the effects of assortative mating (AM). We propose rb_dplr, a novel and computationally efficient algorithm for generating high-dimensional binary random variates that effectively recapitulates AM-induced genetic architectures using the Bahadur order-2 approximation of the multivariate Bernoulli distribution. The rBahadur R library is available through the Comprehensive R Archive Network at https://CRAN.R-project.org/package=rBahadur.

59 BASIC BIOLOGICAL SCIENCES↗

A Novel Machine Learning Approach to Disentangle Multitemperature Regions in Galaxy Clusters

The hot intracluster medium (ICM) surrounding the heart of galaxy clusters is a complex medium that comprises various emitting components. Although previous studies of nearby galaxy clusters, such as the Perseus, the Coma, or the Virgo cluster, have demonstrated the need for multiple thermal components when spectroscopically fitting the ICM’s X-ray emission, no systematic methodology for calculating the number of underlying components currently exists. In turn, underestimating or overestimating the number of components can cause systematic errors in the emission parameter estimations. In this paper, we present a novel approach to determining the number of components using an amalgam of machine learning techniques. Synthetic spectra containing a various number of underlying thermal components were created using well-established tools available from the Chandra X-ray Observatory. The dimensions of the training set was initially reduced using principal component analysis and then categorized based on the number of underlying components using a random forest classifier. Our trained and tested algorithm was subsequently applied to Chandra X-ray observations of the Perseus cluster. Our results demonstrate that machine learning techniques can efficiently and reliably estimate the number of underlying thermal components in the spectra of galaxy clusters, regardless of the thermal model (MEKAL versus APEC). We also confirm that the core of the Perseus cluster contains a mix of differing underlying thermal components. We emphasize that although this methodology was trained and applied on Chandra X-ray observations, it is readily portable to other current (e.g., XMM-Newton, eROSITA) and upcoming (e.g., Athena, Lynx, XRISM) X-ray telescopes. The code is publicly available at https://github.com/XtraAstronomy/Pumpkin.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluation of the procedure 1A component of the 1980 US/Canada wheat and barley exploratory experiment

Several techniques which use clusters generated by a new clustering algorithm, CLASSY, are proposed as alternatives to random sampling to obtain greater precision in crop proportion estimation: (1) Proportional Allocation/relative count estimator (PA/RCE) uses proportional allocation of dots to clusters on the basis of cluster size and a relative count cluster level estimate; (2) Proportional Allocation/Bayes Estimator (PA/BE) uses proportional allocation of dots to clusters and a Bayesian cluster-level estimate; and (3) Bayes Sequential Allocation/Bayesian Estimator (BSA/BE) uses sequential allocation of dots to clusters and a Bayesian cluster level estimate. Clustering in an effective method in making proportion estimates. It is estimated that, to obtain the same precision with random sampling as obtained by the proportional sampling of 50 dots with an unbiased estimator, samples of 85 or 166 would need to be taken if dot sets with AI labels (integrated procedure) or ground truth labels, respectively were input. Dot reallocation provides dot sets that are unbiased. It is recommended that these proportion estimation techniques are maintained, particularly the PA/BE because it provides the greatest precision.

Chapman, G. M.↗

Geosat altimeter observations of the surface circulation of the Southern Ocean

Using Geosat altimeter data for 26 months from November 1986 to December 1988 and a newly developed technique for the analysis of height data, the variability of the sea level and the surface geostrophic currents in the Southern Ocean is investigated. The processed Geosat data are used to examine the relationship between the mesoscale variability and the values of mean circulation, determined from historical hydrographic data. It is shown that the geographical patterns of both the mean flow and the mesoscale variability are correlated. An efficient objective-analysis algorithm for generating smoothed fields from observations randomly distributed in time and two space dimensions is developed and applied to 26 months of Geosat data. The smoothed fields are then used to investigate the large-scale low-frequency variability of the sea level and the surface geostrophic velocity in the Southern Ocean, in order to identify the mode of the observed variations.

Chelton, Dudley B.↗

Moisture Forecast Bias Correction in GEOS DAS

Data assimilation methods rely on numerous assumptions about the errors involved in measuring and forecasting atmospheric fields. One of the more disturbing of these is that short-term model forecasts are assumed to be unbiased. In case of atmospheric moisture, for example, observational evidence shows that the systematic component of errors in forecasts and analyses is often of the same order of magnitude as the random component. we have implemented a sequential algorithm for estimating forecast moisture bias from rawinsonde data in the Goddard Earth Observing System Data Assimilation System (GEOS DAS). The algorithm is designed to remove the systematic component of analysis errors and can be easily incorporated in an existing statistical data assimilation system. We will present results of initial experiments that show a significant reduction of bias in the GEOS DAS moisture analyses.

Dee, D.↗

Understanding the Scalability of Bayesian Network Inference Using Clique Tree Growth Curves

One of the main approaches to performing computation in Bayesian networks (BNs) is clique tree clustering and propagation. The clique tree approach consists of propagation in a clique tree compiled from a Bayesian network, and while it was introduced in the 1980s, there is still a lack of understanding of how clique tree computation time depends on variations in BN size and structure. In this article, we improve this understanding by developing an approach to characterizing clique tree growth as a function of parameters that can be computed in polynomial time from BNs, specifically: (i) the ratio of the number of a BN s non-root nodes to the number of root nodes, and (ii) the expected number of moral edges in their moral graphs. Analytically, we partition the set of cliques in a clique tree into different sets, and introduce a growth curve for the total size of each set. For the special case of bipartite BNs, there are two sets and two growth curves, a mixed clique growth curve and a root clique growth curve. In experiments, where random bipartite BNs generated using the BPART algorithm are studied, we systematically increase the out-degree of the root nodes in bipartite Bayesian networks, by increasing the number of leaf nodes. Surprisingly, root clique growth is well-approximated by Gompertz growth curves, an S-shaped family of curves that has previously been used to describe growth processes in biology, medicine, and neuroscience. We believe that this research improves the understanding of the scaling behavior of clique tree clustering for a certain class of Bayesian networks; presents an aid for trade-off studies of clique tree clustering using growth curves; and ultimately provides a foundation for benchmarking and developing improved BN inference and machine learning algorithms.

Mengshoel, Ole J.↗

Optimization of Microphone Locations for Acoustic Liner Impedance Eduction

Two impedance eduction methods are explored for use with data acquired in the NASA Langley Grazing Flow Impedance Tube. The first is an indirect method based on the convected Helmholtz equation, and the second is a direct method based on the Kumaresan and Tufts algorithm. Synthesized no-flow data, with random jitter to represent measurement error, are used to evaluate a number of possible microphone locations. Statistical approaches are used to evaluate the suitability of each set of microphone locations. Given the computational resources required, small sample statistics are employed for the indirect method. Since the direct method is much less computationally intensive, a Monte Carlo approach is employed to gather its statistics. A comparison of results achieved with full and reduced sets of microphone locations is used to determine which sets of microphone locations are acceptable. For the indirect method, each array that includes microphones in all three regions (upstream and downstream hard wall sections, and liner test section) provides acceptable results, even when as few as eight microphones are employed. The best arrays employ microphones well away from the leading and trailing edges of the liner. The direct method is constrained to use microphones opposite the liner. Although a number of arrays are acceptable, the optimum set employs 14 microphones positioned well away from the leading and trailing edges of the liner. The selected sets of microphone locations are also evaluated with data measured for ceramic tubular and perforate-over-honeycomb liners at three flow conditions (Mach 0.0, 0.3, and 0.5). They compare favorably with results attained using all 53 microphone locations. Although different optimum microphone locations are selected for the two impedance eduction methods, there is significant overlap. Thus, the union of these two microphone arrays is preferred, as it supports usage of both methods. This array contains 3 microphones in the upstream hard wall section, 14 microphones opposite the liner, and 3 microphones in the downstream hard wall section.

Jones, M. G.↗

Amazonia Disasters: Assessing Methods for Gold Mining-related Deforestation Detection in Amazonia Using NASA Earth Observations

Artisanal and small-scale gold mining (ASGM) is responsible for a large fraction of deforestation and disturbance in Amazonia. These activities cause severe impacts on the rainforest ecosystem and socioeconomic state of the region. NASA DEVELOP partnered with the Asociación para la Conservación de la Cuenca Amazónica (ACCA), NASA SERVIR Science Coordination Office, and the Spatial Informatics Group to enhance ASGM-related deforestation detection methods. ACCA currently uses the Omnibus Q-test Change Point Detection Algorithm to identify changes in Synthetic Aperture Radar (SAR) monthly-aggregated temporal data from the Sentinel-1 satellite. The team determined the algorithm's accuracy by comparing a stratified random sample of change points against data from January 2019 to June 2020 identified using PlanetScope and Landsat 8 Operational Land Imager (OLI) Earth observations through Collect Earth Online. Our results indicated a users' accuracy of 55% for temporal change detection and producer's and user's accuracies of 99% and 97%, respectively, for detecting when change did not occur. Of the labeled change points, only 19% were due to mining activity. This research can help our partners have a more accurate understanding of where illegal gold mining may be taking place and inform decisions to remediate this activity.

DEVELOP Project Summary↗