Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Big Data For Operation and Maintenance Cost Reduction

The purpose of this research is to develop a first-of-a-kind framework for integrating Big Data capability into the daily activities of our current fleet of nuclear power plants. Big Data is traditionally defined as data sets with high volume, velocity, and heterogeneity, and the existing Big Data analytics capabilities are now widely popular in fields such as finance, weather, e-commerce, healthcare and sports. In the nuclear industry, while the volume and velocity of data may present computational challenges for existing analytics capabilities, data heterogeneity are seen to present the major challenge. This research project mainly focuses on incorporating the wide range of data heterogeneities in nuclear power plants into an integrated Big Data Analytics capability. The primary end-product of this project is a Big Data framework that is capable of dealing with the large volume and heterogeneity of the data found in nuclear power plants to extract timely and valuable information on equipment performance. The framework can generate system insights that are actionable relations between measurable impacts and the corresponding maintenance action plans and enable optimization of plant operation and maintenance based on the extracted information. The developed framework is capable of handling heterogeneous data including both image data and time-series sensor data. Specifically, this developed framework includes the following components. The first component is an overarching maintenance ontology which includes system insights required by maintenance optimization. The maintenance ontology interacts with other components in the developed framework. The second component handles Piping & Instrumentation Diagram (P&ID) data. It can be used to extract system components and their relations automatically from the P&IDs. This extracted information is stored in the first component, i.e., maintenance ontology, and is also used as input to the third component, i.e., a tool for generating the fault tree for the corresponding system. The generated fault tree in turn is stored in the ontology for assessing risk that is used as a criterion in maintenance policy optimization. The fourth component is a tool for inferring the parameters in the Markov degradation model for a nuclear system. It uses basic information from the ontology. The fifth component is a tool for assessing the degradation level using sensor measurement data, for example, pressure, flowrate. This tool can be used for determining corrective maintenance actions. The results obtained from components four and five are returned to the ontology. The sixth component of the framework is a tool for optimizing the maintenance policy for a nuclear system of interest. It takes certain basic information from the ontology, e.g., costs of maintenance actions and system failures, as input, and returns the optimal maintenance policy to the ontology. This tool can be used for determining predictive maintenance actions. A set of experiments have also been conducted to verify the algorithms developed in this project for nuclear system degradation monitoring. The experiments are based on four solenoid valves, similar to the ones used in nuclear power plants. The analyses based on the experimental data using two algorithms, i.e., the Randomized Window Decomposition (RWD) algorithm and the particle filtering algorithm, and the results are introduced in the report. The Big Data framework developed in this project can be used as a support tool in daily activities of plant operation and maintenance and will reduce current costs while maintaining or improving safety levels. Overall, the project will not only benefit existing reactors, however it will open new frontiers to realize the long overdue value of Big Data Analytics in the nuclear sphere.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Real-time capable modeling of ICRF heating on NSTX and WEST via machine learning approaches

Abstract A real-time capable core Ion Cyclotron Range of Frequencies (ICRF) heating model on NSTX and WEST is developed. The model is based on two nonlinear regression algorithms, the random forest ensemble of decision trees and the multilayer perceptron neural network. The algorithms are trained on TORIC ICRF spectrum solver simulations of the expected flat-top operation scenarios in NSTX and WEST assuming Maxwellian plasmas. The surrogate models are shown to successfully capture the multi-species core ICRF power absorption predicted by the original model for the high harmonic fast wave and the ion cyclotron minority heating schemes while reducing the computational time by six orders of magnitude. Although these models can be expanded, the achieved regression scoring, computational efficiency and increased model robustness suggest these strategies can be implemented into integrated modeling frameworks for real-time control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Robust wind farm layout optimization

Wake interactions in wind farms cause losses in annual energy production (AEP) on the order of 10%. Wind farm designers optimize the layout of the farm to mitigate wake losses, especially in the dominant site-specific wind directions. As wind turbines and wind farms grow in scale, optimization becomes more complex. Offshore wind farms regularly comprise more than 100 wind turbines and are characterized by complex boundaries due to shipping lanes, neighboring wind farms, and other constraints. Layout optimization methods are broadly split between gradient-based and gradient-free approaches. Gradient-based approaches can converge quickly and perform well for smaller, academic problems but are often sensitive to initial conditions and tuning parameters and require expert knowledge to use. On the other hand, gradient-free approaches can be more robust to problem complexities. We present a robust layout optimization approach based on a random search algorithm. The algorithm is intended for those who are not optimization experts and has few tuning parameters that need specification to achieve satisfactory results. Unlike off-the-shelf methods, which use generally available, non-domain-specific optimization routines that accept as inputs an optimization function and constraint definitions, this approach takes advantage of the relative computational costs of the different evaluations by evaluating cheaper computations first (boundary and minimum distance constraints) and running expensive AEP evaluations only if all other checks pass. Moreover, an outer genetic algorithm allows multiple solutions to evolve in parallel, enabling rapid solution development on high-performance computers. We discuss the relative ease of selecting necessary tuning parameters and demonstrate the efficacy of the genetic random search on a complex layout problem consisting of placing 70 turbines in a nonconvex and unconnected boundary region.

17 WIND ENERGY↗

iRF v2.0

A predictive, stable, and interpretable machine learning tool: the iterative random forest algorithm (iRF). iRF discovers high-order interactions among variables with the same order of computational cost as random forests (RF). We have demonstrated the utility of iRF in several applications in the biological and environmental sciences. It is a general purpose machine learning framework for building "explainable" predictive engines.

Brown, JamesB.↗

Leveraging Randomized Compiling for the QITE Algorithm

The success of the current generation of Noisy Intermediate-Scale Quantum (NISQ) hardware shows that quantum hardware may be able to tackle complex problems even without error correction. One outstanding issue is that of coherent errors arising from the increased complexity of these devices. These errors can accumulate through a circuit, making their impact on algorithms hard to predict and mitigate. Iterative algorithms like Quantum Imaginary Time Evolution are susceptible to these errors. This article presents the combination of both noise tailoring using Randomized Compiling and error mitigation with a purification. We also show that Cycle Benchmarking gives an estimate of the reliability of the purification. We apply this method to the Quantum Imaginary Time Evolution of a Transverse Field Ising Model and report an energy estimation and a ground state infidelity both below 1\%. Our methodology is general and can be used for other algorithms and platforms. We show how combining noise tailoring and error mitigation will push forward the performance of NISQ devices.

Ville, Jean-Loup↗

An extended numerical manifold method for unsaturated soil-water interaction analysis at micro-scale

To investigate unsaturated soil-water interaction at micro-scale, this work extends the numerical manifold method (NMM) by incorporating a soil-water coupling model considering specific capillary water distribution and capillary force calculation. The soil skeleton is constructed by a soil skeleton generation algorithm with random polygons. To more realistically capture the interaction between soil grains and capillary water, a capillary mechanics-based geometric algorithm is proposed to iteratively calculate the capillary water distribution. The capillary forces corresponding to the capillary water distribution are calculated based on the Young-Laplace equation. The proposed capillary water solving framework is first verified by reproducing the soil-water characteristic curve and the capillary water distribution of an ideal contact-disk model against analytical solutions. To further validate the ability of the capillary water solving framework to predict hydraulic behavior of the real soil, a laboratory test on the Toyoura sand is reproduced numerically. Then an ideal direct shear test is performed to further validate the two-way soil-water coupling procedure, in which a comparison between the numerical and analytical results regarding the shear strength and matric suction is presented. Finally, microscopic hydraulic and compression tests are conducted on two soil specimens with the same porosity and mean grain diameter but different uniformity coefficients. The results elucidate that the extended method is a potential tool to explore unsaturated soil behaviors at micro-scale.

58 GEOSCIENCES↗

Alert Classification for the ALeRCE Broker System: The Light Curve Classifier

We present the first version of the Automatic Learning for the Rapid Classification of Events (ALeRCE) broker light curve classifier. ALeRCE is currently processing the Zwicky Transient Facility (ZTF) alert stream, in preparation for the Vera C. Rubin Observatory. The ALeRCE light curve classifier uses variability features computed from the ZTF alert stream and colors obtained from AllWISE and ZTF photometry. We apply a balanced random forest algorithm with a two-level scheme where the top level classifies each source as periodic, stochastic, or transient, and the bottom level further resolves each of these hierarchical classes among 15 total classes. This classifier corresponds to the first attempt to classify multiple classes of stochastic variables (including core- and host-dominated active galactic nuclei, blazars, young stellar objects, and cataclysmic variables) in addition to different classes of periodic and transient sources, using real data. We created a labeled set using various public catalogs (such as the Catalina Surveys and Gaia DR2 variable stars catalogs, and the Million Quasars catalog), and we classify all objects with ≥6 g-band or ≥6 r-band detections in ZTF (868,371 sources as of 2020 June 9), providing updated classifications for sources with new alerts every day. For the top level we obtain macro-averaged precision and recall scores of 0.96 and 0.99, respectively, and for the bottom level we obtain macro-averaged precision and recall scores of 0.57 and 0.76, respectively. Updated classifications from the light curve classifier can be found at the ALeRCE Explorer website (http://alerce.online).

47 OTHER INSTRUMENTATION↗

Inexact Newton-CG algorithms with complexity guarantees

Abstract We consider variants of a recently developed Newton-CG algorithm for nonconvex problems (Royer, C. W. & Wright, S. J. (2018) Complexity analysis of second-order line-search algorithms for smooth nonconvex optimization. SIAM J. Optim., 28, 1448–1477) in which inexact estimates of the gradient and the Hessian information are used for various steps. Under certain conditions on the inexactness measures, we derive iteration complexity bounds for achieving $\epsilon $-approximate second-order optimality that match best-known lower bounds. Our inexactness condition on the gradient is adaptive, allowing for crude accuracy in regions with large gradients. We describe two variants of our approach, one in which the step size along the computed search direction is chosen adaptively, and another in which the step size is pre-defined. To obtain second-order optimality, our algorithms will make use of a negative curvature direction on some steps. These directions can be obtained, with high probability, using the randomized Lanczos algorithm. In this sense, all of our results hold with high probability over the run of the algorithm. We evaluate the performance of our proposed algorithms empirically on several machine learning models. Our approach is a first attempt to introduce inexact Hessian and/or gradient information into the Newton-CG algorithm of Royer & Wright (2018, Complexity analysis of second-order line-search algorithms for smooth nonconvex optimization. SIAM J. Optim., 28, 1448–1477).

Mathematics↗

Gradient Coding With Iterative Block Leverage Score Sampling

Gradient coding is a method for mitigating straggling servers in a centralized computing network that uses erasure-coding techniques to distributively carry out first-order optimization methods. Randomized numerical linear algebra uses randomization to develop improved algorithms for large-scale linear algebra computations. In this study, we propose a method for distributed optimization that combines gradient coding and randomized numerical linear algebra. The proposed method uses a randomized ℓ 2 -subspace embedding and a gradient coding technique to distribute blocks of data to the computational nodes of a centralized network, and at each iteration the central server only requires a small number of computations to obtain the steepest descent update. The novelty of our approach is that the data is replicated according to importance scores, called block leverage scores, in contrast to most gradient coding approaches that uniformly replicate the data blocks. Furthermore, we do not require a decoding step at each iteration, avoiding a bottleneck in previous gradient coding schemes. We show that our approach results in a valid ℓ 2 -subspace embedding, and that our resulting approximation converges to the optimal solution.

97 MATHEMATICS AND COMPUTING↗

Automatic Multiple Experiment Simulation and Fitting (Ames-Fit)

AMES-Fit is a program used to automatically fit the multi-field solid-state NMR spectra of half-integer quadrupolar nuclei. Due to the high dimensional space, gradient algorithms have failed to address the fitting of such data, which is at present done manually. AMES-Fit diverges from these approaches by using an adaptive step size random search algorithm to fit the NMR spectra to consistently find the global best fit parameters.

Perras, Frederic↗

Stochastic minibatch approach to the ptychographic iterative engine

The ptychographic iterative engine (PIE) is a widely used algorithm that enables phase retrieval at nanometer-scale resolution over a wide range of imaging experiment configurations. By analyzing diffraction intensities from multiple scanning locations where a probing wavefield interacts with a sample, the algorithm solves a difficult optimization problem with constraints derived from the experimental geometry as well as sample properties. The effectiveness at which this optimization problem is solved is highly dependent on the ordering in which we use the measured diffraction intensities in the algorithm, and random ordering is widely used due to the limited ability to escape from stagnation in poor-quality local solutions. In this study, we introduce an extension to the PIE algorithm that uses ideas popularized in recent machine learning training methods, in this case minibatch stochastic gradient descent. Our results demonstrate that these new techniques significantly improve the convergence properties of the PIE numerical optimization problem.

47 OTHER INSTRUMENTATION↗

COWALKER:EFFECTIVE TRANSPORT PROPERTIES OF COMPOSITE MATERIALS

SF-23-026 This software computes effective transport properties of composite materials involving fibers and nanoparticles using a random-walk algorithm that efficiently scales to an arbitrary number of processes and cores. Effective transport properties (thermal, electrical) are key to bridge the microstructure of complex materials with its macroscopic behavior. Traditional approaches either use effective medium approximations (closed mathematical expressions that are approximation for certain conditions) or continuum simulation models such as finite element or finite volume, which require the generation of a mesh for each configuration explored. cowalker leverages the equivalence between laplacian or heat equation-based models and random walks to compute the asymptotic transport properties from an ensemble of first sojourn times of a random walker moving through the composite material. This allows us to directly define a composite material as a collection of particles and use algorithms developed for molecular dynamics to quickly compute the intersection of the walker with the different interfaces in the material. cowalker is developed in C++, and it relies on the GNU Scientific Library for random generation. cowalker is currently delivered as source code, so the GSL library is not included in cowalker's distribution. A more userfriendly version, cowalker.jl is currently in development and will be released as part of cowalker.

YANGUAS-GIL, ANGEL↗

Developing a SARS-CoV-2 main protease binding prediction random forest model for drug repurposing for COVID-19 treatment

The coronavirus disease 2019 (COVID-19) global pandemic resulted in millions of people becoming infected with the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) virus and close to seven million deaths worldwide. It is essential to further explore and design effective COVID-19 treatment drugs that target the main protease of SARS-CoV-2, a major target for COVID-19 drugs. In this study, machine learning was applied for predicting the SARS-CoV-2 main protease binding of Food and Drug Administration (FDA)-approved drugs to assist in the identification of potential repurposing candidates for COVID-19 treatment. Ligands bound to the SARS-CoV-2 main protease in the Protein Data Bank and compounds experimentally tested in SARS-CoV-2 main protease binding assays in the literature were curated. These chemicals were divided into training (516 chemicals) and testing (360 chemicals) data sets. To identify SARS-CoV-2 main protease binders as potential candidates for repurposing to treat COVID-19, 1188 FDA-approved drugs from the Liver Toxicity Knowledge Base were obtained. A random forest algorithm was used for constructing predictive models based on molecular descriptors calculated using Mold2 software. Model performance was evaluated using 100 iterations of fivefold cross-validations which resulted in 78.8% balanced accuracy. The random forest model that was constructed from the whole training dataset was used to predict SARS-CoV-2 main protease binding on the testing set and the FDA-approved drugs. Model applicability domain and prediction confidence on drugs predicted as the main protease binders discovered 10 FDA-approved drugs as potential candidates for repurposing to treat COVID-19. Our results demonstrate that machine learning is an efficient method for drug repurposing and, thus, may accelerate drug development targeting SARS-CoV-2.

Research & Experimental Medicine↗

SPLENDAQ: A Detector-Agnostic Data Acquisition System for Small-Scale Physics Experiments

Many scientific applications from rare-event searches to condensed matter system characterization to high-rate nuclear experiments require time-domain triggering on a raw stream of data, where the triggering is generally threshold-based or randomly acquired. When carrying out detector R &D, there is a need for a general data acquisition (DAQ) system to quickly and efficiently process such data. In the SPLENDOR collaboration, we are developing the Python-based SPLENDAQ package for this exact purpose—it offers two main features for offline analysis of continuous data: a threshold triggering algorithm based on the time-domain optimal filter formalism and an algorithm for randomly choosing nonoverlapping segments for noise measurements. Further, combined with the commercially available Moku platform, developed by Liquid Instruments, we have a full pipeline of event building off raw data with minimal setup. Here, we review the underlying principles of this detector-agnostic DAQ package and give concrete examples of its utility in various applications.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Using machine learning to identify extragalactic globular cluster candidates from ground-based photometric surveys of M87

Globular clusters (GCs) have been at the heart of many longstanding questions in many sub-fields of astronomy and, as such, systematic identification of GCs in external galaxies has immense impacts. In this study, we take advantage of M87’s well-studied GC system to implement supervised machine learning (ML) classification algorithms – specifically random forest and neural networks – to identify GCs from foreground stars and background galaxies, using ground-based photometry from the Canada–France–Hawaii Telescope (CFHT). We compare these two ML classification methods to studies of ‘human-selected’ GCs and find that the best-performing random forest model can reselect 61.2 per cent ± 8.0 per cent of GCs selected from HST data (ACSVCS) and the best-performing neural network model reselects 95.0 per cent ± 3.4 per cent. When compared to human-classified GCs and contaminants selected from CFHT data – independent of our training data – the best-performing random forest model can correctly classify 91.0 per cent ± 1.2 per cent and the best-performing neural network model can correctly classify 57.3 per cent ± 1.1 per cent. ML methods in astronomy have been receiving much interest as Vera C. Rubin Observatory prepares for first light. The observables in this study are selected to be directly comparable to early Rubin Observatory data and the prospects for running ML algorithms on the upcoming data set yields promising results.

79 ASTRONOMY AND ASTROPHYSICS↗

Diverging climate response of corn yield and carbon use efficiency across the U.S.

Abstract In this paper, we developed an open-source package to analyze the overall trend and responses of both carbon use efficiency (CUE) and corn yield to climate factors for the contiguous United States. Our algorithm enables automatic retrieval of remote sensing data through the Google Earth Engine (GEE) and U.S. Department of Agriculture (USDA) agricultural production data at the county level through application programming interface (API). Firstly, we integrated satellite products of net primary productivity and gross primary productivity based on the Moderate Resolution Imaging Spectroradiometer (MODIS) sensor, and climatic variables from the European Centre for Medium-Range Weather Forecasts. Secondly, we calculated CUE and commonly used climate metrics. Thirdly, we investigated the spatial heterogeneity of these variables. We applied a random forest algorithm to identify the key climate drivers of CUE and crop yield, and estimated the responses of CUE and yield to climate variability using the spatial moving window regression across the U.S. Our results show that growing degree days (GDD) has the highest predictive power for both CUE and yield, while extreme degree days (EDD) is the least important explanatory variable. Moreover, we observed that in most areas of the U.S., yield increases or stays the same with higher GDD and precipitation. However, CUE decreases with higher GDD in the north and shows more mixed and fragmented interactions in the south. Notably, there are some exceptions where yield is negatively correlated with precipitation in the Missouri and Mississippi River Valleys. As global warming continues, we anticipate a decrease in CUE throughout the vast northern part of the country, despite the possibility of yield remaining stable or increasing.

54 ENVIRONMENTAL SCIENCES↗

Automatic fitting of multiple-field solid-state NMR spectra

The NMR lineshapes produced by half-integer quadrupolar nuclei are sensitive to 11 distinct fit parameters per inequivalent site. To date, automatic fitting routines have failed to replace manual parameter insertion and evaluation due to the importance of local minima and the need for fitting multiple-field magic-angle spinning (MAS) and static spectra simultaneously. Herein we introduce a new tool, AMES-Fit (Automatic Multiple Experiment Simulation and Fitting), to automatically find the global best-fit simulation parameters for a series of multiple-field NMR lineshapes. AMES-Fit uses an adaptive step size random search algorithm to dynamically probe parameter space and requires minimal human input. Importantly, the best fits are obtained in a few minutes of computation time that would otherwise have required several person-hours of work. The program is freely available and open-source.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning Analysis of Direct Dynamics Trajectory Outcomes for Thermal Deazetization of 2,3-Diazabicyclo[2.2.1]hept-2-ene

Experimentally, the thermal gas-phase deazetization of 2,3-diazabicyclo[2.2.1]hept-2-ene (1) results in the loss of N2 and the formation of bicyclo products 3 (exo) and 4 (endo) in a nonstatistical ratio, with preference for the exo product. Here, we report unrestricted M06-2X quasiclassical trajectories initialized from the concerted N2 ejection transition state that were able to replicate the experimental preference to form 3. We found that the 3:4 ratio results from the relative amounts of very fast (ballistic) exotype trajectories versus trajectories that lead to the 1,3-diradical intermediate 2. These quasiclassical trajectories provided a set of transition-state vibrational, velocity, momenta, and geometric features for the machine learning analysis. Additionally, a selection of popular supervised classification algorithms (e.g., random forest) provided poor prediction of trajectory outcomes based on only transition-state vibrational quanta and energy features. However, these machine learning models provided more accurate predictions using atomic velocities and atomic positions, attaining ~70% accuracy using initial conditions and between 85 and 95% accuracy at later reaction time steps. This increased accuracy allowed the feature importance analysis to reveal that, at the later-time analysis, the methylene bridge out-of-plane bending is correlated with trajectory outcomes for the formation of either the exo product or toward the diradical intermediate. Possible reasons for the struggle of machine learning algorithms to classify trajectories based on transition-state features is the heavily overlapping feature values, the finite but very large possible vibrational mode combinations, and the possibility of chaos as trajectories propagate. We examined this chaos by comparing a set of nearly identical trajectories that differed by only a very small scaling of the kinetic energies resulting from the transition-state reaction coordinate.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗