Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Further adoption of conservation tillage can increase maize yields in the western US Corn Belt

Conservation tillage can reduce soil erosion, increase soil health, and decrease labor and fuel input costs. Despite these benefits, potential yield impacts remain an important concern for farmers considering adoption. Previous research suggests that conservation tillage is likely to have the largest yield benefits in more arid conditions, but a lack of field-level analyses across climatic, management and soil conditions limits confidence in such predictions. Satellite imagery provides the opportunity to monitor agricultural lands at sub-field resolution across large spatial scales and wide environmental gradients. Here we investigate the maize yield impacts of conservation tillage in the semi-arid western US Corn Belt, using sub-field resolution datasets on tillage practices and crop yields derived from satellite data spanning four states (Nebraska, Kansas, South Dakota, and North Dakota) between 2008 and 2020. On these datasets, we estimate heterogenous yield outcomes for several thousand maize fields across gradients in climate, soil quality and irrigation status by using a causal forests analysis, an adaptation of the random forests machine-learning algorithm for causal inference on observational data. We find that long-term adoption of conservation tillage increased rainfed maize yields by an average of 9.9% in the region. Impacts on irrigated yields were small and not statistically significant. These results, along with an analysis of variables related to greater than average yield benefits, indicate that improved water infiltration and retention are the primary reasons for conservation tillage benefits. Despite yield benefits, many fields estimated to see increased yields under long term low till have not adopted the practice. Therefore, we identify specific counties likely to benefit most from increased levels of adoption. Our results strengthen the understanding of the impacts of conservation agriculture on crop yields and help define environments and counties most likely to benefit from conservation tillage.

54 ENVIRONMENTAL SCIENCES↗

EcoPLOT: dynamic analysis of biogeochemical data

Motivation: We have created EcoPLOT (parameterized linkage of omics-driven technologies), a web-app for the dynamic, interactive analysis of biogeochemical datasets that combines state-of-the-art analysis tools to statistically and graphically explore environmental, geochemical and microbiome datasets. Using the iterative random forest, a machine learning algorithm, EcoPLOT allows for the de novo discovery of drivers which exhibit significant impact on plant, microbial or soil dynamics. Availability and implementation: EcoPLOT is built entirely within the R language. It can be accessed through any system where R is installed, including Windows, Mac and most Linux systems. EcoPLOT is free to use and can be accessed at https://github.com/cdsanchez18/EcoPLOT.

59 BASIC BIOLOGICAL SCIENCES↗

Methods for the robust computation of the long-period seismic spectrum of broad-band arrays

SUMMARY We describe array methods to search for low signal-to-noise ratio (SNR) signals in long-period seismic data using Fourier analysis. This is motivated by published results that find evidence of solar free oscillations in the Earth's seismic hum. Previous work used data from only one station. In this paper, we describe methods for computing spectra from array data. Arrays reduce noise level through averaging and provide redundancy that we use to distinguish coherent signal from a random background. We describe two algorithms for calculating a robust spectrum from seismic arrays, an algorithm that automatically removes impulsive transient signals from data, a jackknife method for estimating the variance of the spectrum, and a method for assessing the significance of an entire spectral band. We show examples of their application to data recorded by the Homestake Mine 3-D array in Lead, SD and the Piñon Flats PY array. These are two of the quietest small aperture arrays ever deployed in North America. The underground Homestake data has exceptionally low noise, and the borehole sensors of the PY array also have very low noise, making these arrays well suited to finding very weak signals. We find that our methods remove transient signals effectively from the data so that even low-SNR signals in the seismic background can be found and tested. Additionally, we find that the jackknife variance estimate is comparable to the noise floor, and we present initial evidence for solar g-modes in our data through the T2 test, a multivariate generalization of Student's t-test.

Caton, Ross C.↗

Stochastic quantum Krylov protocol with double-factorized Hamiltonians

Here we propose a class of randomized quantum Krylov diagonalization (rQKD) algorithms capable of solving the eigenstate estimation problem with modest quantum resource requirements. Compared to previous real-time evolution quantum Krylov subspace methods, our approach expresses the time evolution operator e –i$\widehat{H}$$\tau$ as a linear combination of unitaries and subsequently uses a stochastic sampling procedure to reduce circuit depth requirements. While our methodology applies to any Hamiltonian with fast-forwardable subcomponents, we focus on its application to the explicitly double-factorized electronic-structure Hamiltonian. To demonstrate the potential of the proposed rQKD algorithm on near-term quantum devices, we provide numerical benchmarks for a variety of molecular systems with circuit-based state-vector simulators including the effects of sampling noise, achieving ground-state energy errors of less than 1 kcal mol -1 with circuit depths orders of magnitude shallower than those required for low-rank deterministic Trotter-Suzuki decompositions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Composite Qdrift-product formulas for quantum and classical simulations in real and imaginary time

Recent study has shown that it can be advantageous to implement a composite channel that partitions the Hamiltonian H for a given simulation problem into subsets A and B such that H = A + B , where the terms in A are simulated with a Trotter-Suzuki channel and the B terms are randomly sampled via the Qdrift algorithm. Here we extend Qdrift and composite product formulas to imaginary time, formulating candidate classical algorithms for quantum Monte Carlo calculations. We upper bound the induced Schatten- 1 → 1 norm on both imaginary-time Qdrift and composite channels. Another recent result demonstrated that simulations of lattice Hamiltonians containing geometrically local interactions can be improved using a Lieb-Robinson argument to decompose H into subsets that contain only terms supported on that subset of the lattice. Here, we provide a quantum algorithm by unifying this result with the composite approach into “local composite channels” and we upper bound the diamond distance. We provide exact numerical simulations of algorithmic cost by counting the number of gates of the form e − i H j t and e − H j β to meet a certain error tolerance ε . In doing so, we optimize the partitioning into sets A and B using gradient boosted tree models from machine learning. These numerical studies are important given that product formulas have been historically known to outperform analytic upper bounds. We show constant factor advantages for a variety of interesting Hamiltonians, the maximum of which is a ≈ 20 -fold speedup that occurs in the simulation of Jellium. Published by the American Physical Society 2024

Pocrnic, Matthew (ORCID:0000000203089376)↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Detecting Arsenic Contamination Using Satellite Imagery and Machine Learning

Arsenic, a potent carcinogen and neurotoxin, affects over 200 million people globally. Current detection methods are laborious, expensive, and unscalable, being difficult to implement in developing regions and during crises such as COVID-19. This study attempts to determine if a relationship exists between soil’s hyperspectral data and arsenic concentration using NASA’s Hyperion satellite. It is the first arsenic study to use satellite-based hyperspectral data and apply a classification approach. Four regression machine learning models are tested to determine this correlation in soil with bare land cover. Raw data are converted to reflectance, problematic atmospheric influences are removed, characteristic wavelengths are selected, and four noise reduction algorithms are tested. The combination of data augmentation, Genetic Algorithm, Second Derivative Transformation, and Random Forest regression (R 2 =0.840 and normalized root mean squared error (re-scaled to [0,1]) = 0.122) shows strong correlation, performing better than past models despite using noisier satellite data (versus lab-processed samples). Three binary classification machine learning models are then applied to identify high-risk shrub-covered regions in ten U.S. states, achieving strong accuracy (=0.693) and F1-score (=0.728). Overall, these results suggest that such a methodology is practical and can provide a sustainable alternative to arsenic contamination detection.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Sparse Approximate Multifrontal Factorization with Butterfly Compression for High-Frequency Wave Equations

In this work, we present a fast and approximate multifrontal solver for large-scale sparse linear systems arising from finite-difference, finite-volume or finite-element discretization of high-frequency wave equations. The proposed solver leverages the butterfly algorithm and its hierarchical matrix extension for compressing and factorizing large frontal matrices via graph-distance guided entry evaluation or randomized matrix-vector multiplication-based schemes. Complexity analysis and numerical experiments demonstrate $\mathcal{O}(N\log^2 N)$ computation and $\mathcal{O}(N)$ memory complexity when applied to an $N\times N$ sparse system arising from 3D high-frequency Helmholtz and Maxwell problems.

97 MATHEMATICS AND COMPUTING↗

Reconstructions from randomly generated longitudinal electron bunch profiles with Gaussian envelopes using the Gerchberg–Saxton algorithm

Knowledge of longitudinal electron bunch profiles is vital to optimize the performance of plasma wakefield accelerators and x-ray free electron laser linacs. Because of their importance to these novel applications, noninvasive frequency domain techniques are often employed to reconstruct longitudinal bunch profiles from coherent synchrotron, transition, or undulator radiation measurements. In this paper, we detail several common reconstruction techniques involving the Kramers–Kronig phase relationship and Gerchberg–Saxton algorithm. Additionally, through statistical analysis, we draw general conclusions about the accuracy of these reconstruction techniques and the most suitable candidate for reconstructing well-isolated longitudinal bunch profiles from spectroscopic data.

47 OTHER INSTRUMENTATION↗

Learning From User Behavior: A Survey-Assist Algorithm for Longitudinal Mobility Data Collection

GPS-based travel surveys are widely used in mobility studies to gather crucial qualitative data, like purpose, transportation mode and replaced mode. However, survey response still poses a burden to users, especially in long-term mobility studies, leading to response fatigue. We explore a survey-assist strategy to ease this burden by a novel, user-level modeling approach that leverages past responses from each user to predict responses for new trips, without relying on external data sources like GIS data. We investigate three main algorithms for predicting responses: (i) clustering trips and extrapolating responses for similar trips, (ii) using random forest classification, and (iii) clustering that uses a hybrid algorithm to determine spatial structure, which is then fed as input to a classic random forest classifier. The clustering approach can flexibly predict responses for even complex qualitative survey questions; it achieved F-scores of 65%. The random forest pipeline uses architecture that restricts it to predicting three predetermined survey questions: trip purpose, mode, and replaced mode. However, it achieved F-scores of 78%. While the survey-assist approach has been implemented by several proprietary systems, to our knowledge, this is the first exploration in the academic literature. It follows that this is also the first rigorous evaluation of multiple algorithms that can implement the approach. The evaluation uses a large scale, publicly available, longitudinal dataset consisting of ~ 92k trips from 235 users over a period of roughly one and a half years. With this approach, travel surveys can be pre-filled with the predicted responses for each trip, thus streamlining the survey process for users. Combined with an active learning system that requests user input on low-confidence predictions, models can be updated and improved over time to better support the long-term collection of longitudinal qualitative data.

clustering↗

Predicting Biomass Yields of Advanced Switchgrass Cultivars for Bioenergy and Ecosystem Services Using Machine Learning

The production of advanced perennial bioenergy crops within marginal areas of the agricultural landscape is gaining interest due to its potential to sustainably produce feedstocks for biofuels and bioproducts while also improving the sustainability and resilience of commodity crop production. However, predicting the biomass yields of this production system is challenging because marginal areas are often relatively small and spread around agricultural fields and are typically associated with various abiotic conditions that limit crop production. Machine learning (ML) offers a viable solution as a biomass yield prediction tool because it is suited to predicting relationships with complex functional associations. The objectives of this study were to (1) evaluate the accuracy of commonly applied ML algorithms in agricultural applications for predicting the biomass yields of advanced switchgrass cultivars for bioenergy and ecosystem services and (2) determine the most important biomass yield predictors. Datasets on biomass yield, weather, land marginality, soil properties, and agronomic management were generated from three field study sites in two U.S. Midwest states (Illinois and Iowa) over three growing seasons. The ML algorithms evaluated in the study included random forests (RFs), gradient boosting machines (GBMs), artificial neural networks (ANNs), K-neighbors regressor (KNR), AdaBoost regressor (ABR), and partial least squares regression (PLSR). Coefficient of determination (R 2 ) and mean absolute error (MAE) were used to evaluate the predictive accuracy of the tested algorithms. Results showed that the ensemble methods, RF (R 2 = 0.86, MAE = 0.62 Mg/ha), GBM (R 2 = 0.88, MAE = 0.57 Mg/ha), and GBM (R 2 = 0.78, MAE = 0.66 Mg/ha), were the most accurate in predicting biomass yields of the Independence, Liberty, and Shawnee switchgrass cultivars, respectively. This is in agreement with similar studies that apply ML to multi-feature problems where traditional statistical methods are less applicable and datasets used were considered to be relatively small for ANNs. Consistent with previous studies on switchgrass, the most important predictors of biomass yield included average annual temperature, average growing season temperature, sum of the growing season precipitation, field slope, and elevation. This study helps pave the way for applying ML as a management tool for alternative bioenergy landscapes where understanding agronomic and environmental performance of a multifunctional cropping system seasonally and interannually at the sub-field scale is critical.

09 BIOMASS FUELS↗

Fast GPU-Based Generation of Large Graph Networks From Degree Distributions

Synthetically generated, large graph networks serve as useful proxies to real-world networks for many graph-based applications. The ability to generate such networks helps overcome several limitations of real-world networks regarding their number, availability, and access. Here, we present the design, implementation, and performance study of a novel network generator that can produce very large graph networks conforming to any desired degree distribution. The generator is designed and implemented for efficient execution on modern graphics processing units (GPUs). Given an array of desired vertex degrees and number of vertices for each desired degree, our algorithm generates the edges of a random graph that satisfies the input degree distribution. Multiple runtime variants are implemented and tested: 1) a uniform static work assignment using a fixed thread launch scheme, 2) a load-balanced static work assignment also with fixed thread launch but with cost-aware task-to-thread mapping, and 3) a dynamic scheme with multiple GPU kernels asynchronously launched from the CPU. The generation is tested on a range of popular networks such as Twitter and Facebook, representing different scales and skews in degree distributions. Results show that, using our algorithm on a single modern GPU (NVIDIA Volta V100), it is possible to generate large-scale graph networks at rates exceeding 50 billion edges per second for a 69 billion-edge network. GPU profiling confirms high utilization and low branching divergence of our implementation from small to large network sizes. For networks with scattered distributions, we provide a coarsening method that further increases the GPU-based generation speed by up to a factor of 4 on tested input networks with over 45 billion edges.

97 MATHEMATICS AND COMPUTING↗

Active Learning Accelerated Discovery of Stable Iridium Oxide Polymorphs for the Oxygen Evolution Reaction

The discovery of high-performing and stable materials for sustainable energy applications is a pressing goal in catalysis and materials science. Understanding the relationship between a material’s structure and functionality is an important step in the process, such that viable polymorphs for a given chemical composition need to be identified. Machine-learning-based surrogate models have the potential to accelerate the search for polymorphs that target specific applications. Herein, we report a readily generalizable active-learning (AL) accelerated algorithm for identification of electrochemically stable iridium oxide polymorphs of IrO 2 and IrO 3 . The search is coupled to a subsequent analysis of the electrochemical stability of the discovered structures for the acidic oxygen evolution reaction (OER). Structural candidates are generated by identifying all 956 structurally unique AB2 and AB3 prototypes in existing materials databases (more than 38000). Next, using an active learning approach, we find 196 IrO 2 polymorphs within the thermodynamic amorphous synthesizability limit and reaffirm the global stability of the rutile structure. We find 75 synthesizable IrO 3 polymorphs and report a previously unknown FeF 3 -type structure as the most stable, termed α-IrO 3 . To test the algorithms performance, we compare to a random search of the candidate space and report at least a 2-fold increase in the rate of discovery. Additionally, the AL approach can acquire the most stable polymorphs of IrO 2 and IrO 3 with fewer than 30 density functional theory optimizations. Analysis of the structural properties of the discovered polymorphs reveals that octahedral local coordination environments are preferred for nearly all low-energy structures. Subsequent Pourbaix Ir–H 2 O analysis shows that α-IrO3 is the globally stable solid phase under acidic OER conditions and supersedes the stability of rutile IrO 2 . Calculation of theoretical OER surface activities reveal ideal weaker binding of the OER intermediates on α-IrO 3 than on any other considered iridium oxide. We emphasize that the proposed AL algorithm can be easily generalized to search for any binary metal oxide structure with a defined stoichiometry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

How Should Machine Learning Be Successfully Used for Wind Speed Vertical Extrapolation?

An accurate characterization of the wind resource available at hub-height is required for an efficient and bankable wind farm project. However, direct measurement of wind speed at the constantly increasing height of the hub of commercial wind turbines is oftentimes challenging and expensive, so that it is common practice to vertically extrapolate the wind resource from lower and more easily accessible levels. Conventional techniques for wind speed vertical extrapolation include the use of a power law and a logarithmic profile. While simple, the limits in accuracy of these methods have been shown in various studies. Recently, machine learning has been proposed as a new method to vertically extrapolate winds. All the published studies on the topic assess the performance of machine learning techniques in vertically extrapolating the wind resource at the same location where the algorithm has been trained. However, in real-world applications, the wind resource is measured at the instrument location, but it then needs to be extrapolated at hub height at the location of the wind turbines within the find farm. To be able to fully recommend the use of machine learning techniques over the simple power law and logarithmic law, the spatial variability of the performance improvements of the machine learning approaches needs to be assessed. Here, we propose a round-robin validation of a machine learning-based method for wind speed extrapolation. We use 20 months of observations at four locations spanning a 100 km wide region at the Southern Great Plains (SGP) atmospheric observatory, in north-central Oklahoma. At each location, we train a random forest to predict 30-min average wind speed at 143 m AGL. We use as input features lidar wind speed at 65 m AGL, time of day, sonic anemometer wind speed at 4 m AGL, turbulent kinetic energy, and Obukhov length. First, we perform a same-site comparison of the performance of the proposed random forest against the conventional techniques for wind speed extrapolation (namely power law and logarithmic profile, with widely accepted stability corrections). We find that the random forest outperforms the power law in vertically extrapolating wind speed in all the considered stability regimes, with a 33% reduction in MAE for stable conditions, and a 31% reduction in unstable conditions. Similar results are found when comparing predictions of extrapolated winds from the logarithmic profile and the random forest with the observed values. Next, we propose a round-robin validation, to use the random forest trained at each site to extrapolate wind speed at the remaining three sites. We find that the performance of the random forest approach degrades when the algorithm is tested at a site different than the training one. However, even under those circumstances, the machine learning-based approach still outperforms the conventional techniques for wind speed extrapolation, with, on average, a reduction in mean absolute error between 15 and 20% over the conventional methods, with the largest benefits obtained under stable conditions.

Monte Carlo↗

Survey of Time Shift Detection Algorithms for Measured PV Data

In this research, three variations of time shift detection algorithms were tested for their ability to detect time shift issues (including daylight savings time and random time shifts) in measured PV data sets. Two algorithms from the Python PVAnalytics package were assessed, and one algorithm from the Solar-Data-Tools package was assessed. Each algorithm's ability to accurately detect and measure time shifts was assessed.

automated preprocessing↗

Quantum mixed state compiling

The task of learning a quantum circuit to prepare a given mixed state is a fundamental quantum subroutine. We present a variational quantum algorithm (VQA) to learn mixed states which is suitable for near-term hardware. Our algorithm represents a generalization of previous VQAs that aimed at learning preparation circuits for pure states. We consider two different ansätze for compiling the target state; the first is based on learning a purification of the state and the second on representing it as a convex combination of pure states. In both cases, the resources required to store and manipulate the compiled state grow with the rank of the approximation. Thus, by learning a lower rank approximation of the target state, our algorithm provides a means of compressing a state for more efficient processing. As a byproduct of our algorithm, one effectively learns the principal components of the target state, and hence our algorithm further provides a new method for principal component analysis. We investigate the efficacy of our algorithm through extensive numerical implementations, showing that typical random states and thermal states of many body systems may be learnt this way. Additionally, we demonstrate on quantum hardware how our algorithm can be used to study hardware noise-induced states.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Illumination correction of dyed fabric based on extreme learning machine with improved ant lion optimizer

Abstract In order to eliminate the influence of scene illumination on the evaluation of the color difference of dyed fabrics, this paper proposes a dyed fabric illumination correction algorithm based on the extreme learning machine (ELM) with grey wolf optimizer (GWO)‐optimized ant lion optimizer (ALO). Firstly, the Grey Edge framework is used to extract the features of the dyed fabric image as the input vector. Then, to improve the optimization ability of the ALO algorithm, the GWO algorithm is used to provide a set of optimized initial populations to the ALO algorithm, and then the improved ALO algorithm is used to optimize the parameters of the ELM. Finally, the proposed GWO‐ALO‐ELM algorithm is used to correct the illumination of the dyed fabric, and restore the graphics to the effect display under standard illumination through the diagonal reduction model. Compared with the experimental results of GWO‐ELM, ALO‐ELM, backpropagation (BP), ELM, random vector function link (RVFL), and other algorithms, it can be seen that the GWO‐ALO‐ELM algorithm proposed in this paper has good predictive value and quasi‐bias effect, and good stability.

Zhou, Zhiyu↗

Hybrid quantum-classical algorithms for approximate graph coloring

We show how to apply the recursive quantum approximate optimization algorithm (RQAOA) to MAX- k -CUT, the problem of finding an approximate k -vertex coloring of a graph. We compare this proposal to the best known classical and hybrid classical-quantum algorithms. First, we show that the standard (non-recursive) QAOA fails to solve this optimization problem for most regular bipartite graphs at any constant level p : the approximation ratio achieved by QAOA is hardly better than assigning colors to vertices at random. Second, we construct an efficient classical simulation algorithm which simulates level- 1 QAOA and level- 1 RQAOA for arbitrary graphs. In particular, these hybrid algorithms give rise to efficient classical algorithms, and no benefit arising from the use of quantum mechanics is to be expected. Nevertheless, they provide a suitable testbed for assessing the potential benefit of hybrid algorithm: We use the simulation algorithm to perform large-scale simulation of level- 1 QAOA and RQAOA with up to 300 qutrits applied to ensembles of randomly generated 3 -colorable constant-degree graphs. We find that level- 1 RQAOA is surprisingly competitive: for the ensembles considered, its approximation ratios are often higher than those achieved by the best known generic classical algorithm based on rounding an SDP relaxation. This suggests the intriguing possibility that higher-level RQAOA may be a potentially useful algorithm for NISQ devices.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗