Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Data trustworthiness signatures for nuclear reactor dynamics simulation

With the increased reliance on digitization in industrial control systems, the need for effective monitoring techniques has risen dramatically. Specifically, there is now a growing concern about the so-called false data injection (FDI) attacks. These attacks aim to alter the raw sensors’ data to cause malicious outcomes. Any serious FDI algorithm is based on an intimate knowledge of the system and its associated physics models, which renders conventional outlier/anomaly detection techniques almost obsolete in the face of such attacks. Thus, a critical need has emerged to develop a new class of defense methods that are capable of detecting FDI attacks under the assumption that the attacker has a strong familiarity with the system and its physics modeling. This class of defense methods are denoted by model-based defenses which are premised on the assumption that the attacker, while having a good understanding of the system, does not have full privileged access to all proprietary data and historical records of operation. However, (s)he is assumed to be capable of learning system behavior using self-learning techniques during an initial lie-in-wait period. To defend against this scenario, we propose a new model-based randomized window algorithm that searches time-series data for signatures that can serve as classifiers between normal and FDI scenarios. The classifiers are based on the correlations between the dominant degrees of freedom (DOFs) and the less-dominant DOFs (expected to be very sensitive to the system details that are unknown to the attacker). For demonstration, RELAP5 models are employed to calculate representative nuclear reactor behavior during a number of transient scenarios. Finally, falsified data are injected into the RELAP5-simulated behavior, and the proposed signature-identification algorithm is employed to detect the injected data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Artificial intelligence driven laser parameter search: Inverse design of photonic surfaces using greedy surrogate-based optimization

Photonic surfaces designed with specific optical characteristics are becoming increasingly crucial for novel energy harvesting and storage systems. The design of these surfaces can be achieved by texturing materials using lasers. The optimal adjustment of laser fabrication parameters to achieve target surface optical properties is an open challenge. Thus, we develop a surrogate-based optimization approach. Our framework employs the Random Forest algorithm to model the forward relationship between the laser fabrication parameters and the resulting optical characteristics. During the optimization process, we use a greedy, prediction-based exploration strategy that iteratively selects batches of laser parameters to be used in experimentation by minimizing the predicted discrepancy between the surrogate model’s outputs and the user-defined target optical characteristics. This strategy allows for efficient identification of optimal fabrication parameters without the need to model the error landscape directly. We demonstrate the efficiency and effectiveness of our approach on two synthetic benchmarks and two specific experimental applications of photonic surface inverse design targets. By calculating the average performance of our algorithm compared to other state of the art optimization methods, we show that our algorithm performs, on average, twice as well across all benchmarks. Additionally, a warm starting inverse design technique for changed target optical characteristics enhances the performance of the introduced approach.

97 MATHEMATICS AND COMPUTING↗

Pass-efficient methods for compression of high-dimensional turbulent flow data

The future of high-performance computing, specifically on future Exascale computers, will presumably see memory capacity and bandwidth fail to keep pace with data generated, for instance, from massively parallel partial differential equation (PDE) systems. Current strategies proposed to address this bottleneck entail the omission of large fractions of data, as well as the incorporation of in situ compression algorithms to avoid overuse of memory. To ensure that post-processing operations are successful, this must be done in a way that a sufficiently accurate representation of the solution is stored. Moreover, in situations where the input/output system becomes a bottleneck in analysis, visualization, etc., or the execution of the PDE solver is expensive, the number of passes made over the data must be minimized. In the interest of addressing this problem, this work focuses on the utility of pass-efficient, parallelizable, low-rank, matrix decomposition methods in compressing high-dimensional simulation data from turbulent flows. Additionally, a particular emphasis is placed on using coarse representation of the data – compatible with the PDE discretization grid – to accelerate the construction of the low-rank factorization. This includes the presentation of a novel single-pass matrix decomposition algorithm for computing the so-called interpolative decomposition. The methods are described extensively and numerical experiments on two turbulent channel flow data are performed. In the first (unladen) channel flow case, compression factors exceeding 400 are achieved while maintaining accuracy with respect to first- and second-order flow statistics. In the particle-laden case, compression factors of 100 are achieved and the compressed data is used to recover particle velocities. These results show that these compression methods can enable efficient computation of various quantities of interest in both the carrier and disperse phases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine learning enables identification of an alternative yeast galactose utilization pathway

How genomic differences contribute to phenotypic differences is a major question in biology. The recently characterized genomes, isolation environments, and qualitative patterns of growth on 122 sources and conditions of 1,154 strains from 1,049 fungal species (nearly all known) in the yeast subphylum Saccharomycotina provide a powerful, yet complex, dataset for addressing this question. We used a random forest algorithm trained on these genomic, metabolic, and environmental data to predict growth on several carbon sources with high accuracy. Known structural genes involved in assimilation of these sources and presence/absence patterns of growth in other sources were important features contributing to prediction accuracy. By further examining growth on galactose, we found that it can be predicted with high accuracy from either genomic (92.2%) or growth data (82.6%) but not from isolation environment data (65.6%). Prediction accuracy was even higher (93.3%) when we combined genomic and growth data. After the GALactose utilization genes, the most important feature for predicting growth on galactose was growth on galactitol, raising the hypothesis that several species in two orders, Serinales and Pichiales (containing the emerging pathogen Candida auris and the genus Ogataea, respectively), have an alternative galactose utilization pathway because they lack the GAL genes. Growth and biochemical assays confirmed that several of these species utilize galactose through an alternative oxidoreductive D-galactose pathway, rather than the canonical GAL pathway. Machine learning approaches are powerful for investigating the evolution of the yeast genotype–phenotype map, and their application will uncover novel biology, even in well-studied traits.

59 BASIC BIOLOGICAL SCIENCES↗

Randomized Preconditioned Solvers for Strong Constraint 4D-Var Data Assimilation

The Strong Constraint 4D Variational (SC-4DVAR) data assimilation method is widely used in climate and weather applications. SC-4DVAR involves solving a minimization problem to compute the maximum a posteriori estimate, which we tackle using the Gauss-Newton method. The computation of the descent direction is expensive since it involves the solution of a large-scale and potentially ill-conditioned linear system, solved using the preconditioned conjugate gradient (PCG) method. Here, to address this cost, we efficiently construct scalable preconditioners using three different randomization techniques, which all rely on a certain low-rank structure involving the Gauss-Newton Hessian. The proposed techniques come with theoretical guarantees on the condition number, and at the same time, are amenable to parallelization. We also develop an adaptive approach to estimate the sketch size and choose between the reuse or recomputation of the preconditioner. We demonstrate the performance and effectiveness of our methodology on two representative model problems—the Burgers and barotropic vorticity equation—showing a drastic reduction in both the number of PCG iterations and the number of Gauss-Newton Hessian products after including the preconditioner construction cost.

Gauss-Newton↗

A learning-augmented approach for AC optimal power flow

Because of the high nonlinearity of AC optimal power flow (OPF), numerous efforts have been made in recent decades to find efficient methods. Machine learning (ML) has proven to significantly reduce the computational costs in many real-world problems. Thus, this paper develops a learning-augmented method for solving AC OPF, which integrates both power network equations and ML to yield near-optimal solutions. More specifically, ML models are developed to first predict bus voltage magnitudes and angles. Then, physics-based network equations are employed to calculate the power injection at different buses. Three ML algorithms, i.e., random forest, multi-target decision tree, and extreme learning machine, are explored and compared. To evaluate the efficiency of the proposed learning-augmented AC OPF solver, the MATPOWER Interior Point Solver is adopted as a baseline. Case studies on both 500-bus and 4918-bus test networks show that the proposed learning-augmented method has reduced the computational time by 15–100 times depending on the network size with a minimal loss in optimality.

42 ENGINEERING↗

Resilient information and inference networks under mixed-trust sensing

With ubiquitous digitization, sensing, and computational intelligence deployed in increasingly more and broader domains, including critical infrastructure, potentially misleading and destabilizing effects of multimodal anomalies and adversarial behavior are growing in importance. Here, we develop randomized and reinforcement learning-based strategies for strategically recruiting and utilizing deployed (and, thus, vulnerable and potentially faulty and/or compromised) nodes from information and inference networks, while defending against adversaries that attempt to misguide assessments of inferred variables. Recognizing that, besides communication and other costs, sampling from any observable node can either provide true data or dangerously expose our inference to misinformation (without being easily distinguishable what actually happens), the proposed strategies proceed by progressively recruiting nodes and cautiously scaling their information contribution based on assumed, or, in our reinforcement learning approach, intelligently weighed trustworthiness, with the learning approach also considering network-wide, threat-inclusive risk/value tradeoffs. While avoiding the hardware, communication, analytical and computational burden of explicit redundancy, the proposed defensive schemes enable on-the-fly assessments of underlying processes, and system-wide situational awareness with demonstrable resilience against adversarial activities.

97 - MATHEMATICS AND COMPUTING↗

A data-driven global soil heterotrophic respiration dataset and the drivers of its inter-annual variability

Soil heterotrophic respiration (SHR), one of the primary carbon fluxes from terrestrial ecosystems to the atmosphere, is important for carbon-climate feedbacks because of its sensitivity to available litter and soil carbon, climatic conditions, and nutrient availability. However, until recently limited SHR data were available, and most published global SHR estimates have either a short time span, coarse spatial resolution, or reply on overly-simple model formulations. To better understand and quantify the global distribution of SHR and its sensitivity to climate variability, we produced a new global SHR dataset using Random Forest algorithms, up-scaling 455 point data from the Global Soil Respiration Database (SRDB 4.0) with gridded fields of climatic, edaphic and productivity as explanatory variables. We estimated a global total SHR of 46.8 Pg C yr-1 over 1985-2013 (95% confidence interval: 38.6-56.3 Pg C yr-1), with a significant increasing trend of 0.03 Pg C yr-2 during this period. We found that the choice of soil moisture datasets contributes more to the difference among these data-driven SHR members rather than that of productivity, temperature and precipitation data sources. We also analyzed the influence of climatic variables on the inter-annual variability (IAV) of our SHR product. Water availability was the dominant driver of IAV at global scales, although the inferred sensitivity depends on the choice of the soil moisture gridded dataset. At the ecosystem scale, temperature strongly controls the IAV of SHR in tropical forests, while water availability dominates in extra-tropical forest and semi-arid regions. Our machine-learning gridded SHR dataset and outputs from process-based land surface models (TRENDYv6) show agreement for a strong association between water variability and SHR IAV at the global scale, but the two approaches lead to different temporal trend globally and different controlling variables for IAV at the ecosystem scale. Our study provides evidence for the pervasive and important role of water availability in driving SHR, indicating both a direct effect limiting decomposition rates and an indirect effect through the amount of fresh organic matter made available to SHR from productivity. In consideration of potential limitations and uncertainties remaining in our data-driven SHR datasets, we call for a more scientifically designed observation network for SHR, more observation data compilation, and increased use of deep learning methods making maximum use of observation data in hand. This will benefit process-based models, and improve our understanding of SHR response to future anomalous environmental conditions.

Yao, Yitong↗

Curiosity driven exploration to optimize structure–property learning in microscopy

Rapidly determining structure–property correlations in materials is an important challenge in better understanding fundamental mechanisms and greatly assists in materials design. In microscopy, imaging data provides a direct measurement of the local structure, while spectroscopic measurements provide relevant functional property information. Deep kernel active learning approaches have been utilized to rapidly map local structure to functional properties in microscopy experiments, but are computationally expensive for multi-dimensional and correlated output spaces. Here, we present an alternative lightweight curiosity algorithm which actively samples regions with unexplored structure–property relations, utilizing a deep-learning based surrogate model for error prediction. We show that the algorithm outperforms random sampling for predicting properties from structures, and provides a convenient tool for efficient mapping of structure–property relationships in materials science.

36 MATERIALS SCIENCE↗

Jacobian sparsity detection using Bloom filters

Determining Jacobian sparsity structure is an important step in the efficient computation of sparse Jacobians. We introduce a new method for determining Jacobian sparsity patterns by combining bit vector probing with Bloom filters. In conclusion, we further refine Bloom filter probing by combining it with hierarchical probing to yield a highly effective strategy for Jacobian sparsity pattern determination.

Bloom filter↗

Universality class for loopless invasion percolation models and a percolation avalanche burst model for hydraulic fracturing

Invasion percolation is a model that was originally proposed to describe growing networks of fractures. In this work, we describe a loopless algorithm on random lattices, coupled with an avalanche-based model for bursts. The model reproduces the characteristic b-value seismicity and spatial distribution of bursts consistent with earthquakes resulting from hydraulic fracturing (“fracking”). We test models for both site invasion percolation and bond invasion percolation. These have differences on the scale of site and bond lengths l. But since the networks are characterized by their large-scale behavior, l << L, we find small differences between scaling exponents. Though data may not differentiate between models, our results suggest that both models belong to different universality classes.

58 GEOSCIENCES↗

A Machine Learning-Based Method to Estimate Transformer Primary-Side Voltages with Limited Customer-Side AMI Measurements

Distribution control applications such as volt/var optimization, network reconfiguration, and distribution automation require accurate knowledge of the distribution system state. The lack of sufficient sensors on the primary side of distribution networks often limits the accuracy of the control decisions by these applications. The deployment of advanced metering infrastructure (AMI) provides utilities an opportunity to translate the AMI data on the secondary onto the primary so that it can be used as pseudo-measurements to augment the limited existing measurements on the primary. This paper develops a machine learning based approach for estimating service transformer primary-side voltages by using limited secondary-side AMI measurement. The machine learning model is developed by using random forest algorithm. The estimated primary-side voltages can be used by utilities as pseudo-measurements for distribution control applications. The detailed secondary model topology, which is an essential input data for many existing algorithms, is not required for the proposed method. The performance of the proposed method is validated by using AMI measurements from the field and an actual distribution feeder model of San Diego Gas & Electric Company.

advanced metering infrastructure↗

A Genome-Based Model to Predict the Virulence of Pseudomonas aeruginosa Isolates

ABSTRACT: Variation in the genome of Pseudomonas aeruginosa , an important pathogen, can have dramatic impacts on the bacterium’s ability to cause disease. We therefore asked whether it was possible to predict the virulence of P. aeruginosa isolates based on their genomic content. We applied a machine learning approach to a genetically and phenotypically diverse collection of 115 clinical P. aeruginosa isolates using genomic information and corresponding virulence phenotypes in a mouse model of bacteremia. We defined the accessory genome of these isolates through the presence or absence of accessory genomic elements (AGEs), sequences present in some strains but not others. Machine learning models trained using AGEs were predictive of virulence, with a mean nested cross-validation accuracy of 75% using the random forest algorithm. However, individual AGEs did not have a large influence on the algorithm’s performance, suggesting instead that virulence predictions are derived from a diffuse genomic signature. These results were validated with an independent test set of 25 P. aeruginosa isolates whose virulence was predicted with 72% accuracy. Machine learning models trained using core genome single-nucleotide variants and whole-genome k-mers also predicted virulence. Our findings are a proof of concept for the use of bacterial genomes to predict pathogenicity in P. aeruginosa and highlight the potential of this approach for predicting patient outcomes. IMPORTANCE Pseudomonas aeruginosa is a clinically important Gram-negative opportunistic pathogen. P. aeruginosa shows a large degree of genomic heterogeneity both through variation in sequences found throughout the species (core genome) and through the presence or absence of sequences in different isolates (accessory genome). P. aeruginosa isolates also differ markedly in their ability to cause disease. In this study, we used machine learning to predict the virulence level of P. aeruginosa isolates in a mouse bacteremia model based on genomic content. We show that both the accessory and core genomes are predictive of virulence. This study provides a machine learning framework to investigate relationships between bacterial genomes and complex phenotypes such as virulence.

59 BASIC BIOLOGICAL SCIENCES↗

Dependence of Convective Cloud Microphysical Properties on Environmental Conditions during the TRACER and ESCAPE Field Campaigns: A Synergistic Approach of Observations, Machine Learning and Parcel Models

The sensitivity of convective clouds to aerosols and their interactions with environment, combined with limited observational constraints in parameterizations, introduces significant uncertainties in atmospheric models. Here, this study investigates the dependence of convective cloud microphysical properties on environmental conditions using a synergistic approach that combines unique observations from the TRACER and ESCAPE field campaigns, machine learning techniques, and parcel model simulations with a super-droplet microphysics scheme. A random forest algorithm identifies in-situ vertical velocity (w), temperature (T), and surface fine-mode aerosol mass concentration as the three most important environmental conditions influencing cloud properties including liquid water content (LWC), number concentration for particles with D max < 50 μm (N c ,<50), 50 μm ≤ D max ≤ 3000 μm (N c,50–3000 ), and droplet effective diameter (D e ). Results show that LWC, N c,<50 , and N c,50–3000 significantly increase with w in updrafts. Across w bins, as T decreases, LWC, D e , and N c,50–3000 increase, while N c,<50 decreases, which are closely linked to the distance above cloud bases. Warmer cloud bases yield higher LWC, greater N c,50–3000 , and smaller N c,<50 , while polluted environments produce greater N c,<50 . Parcel model simulations successfully replicate these observed dependencies. The simulation results indicate that warmer cloud bases enhance condensation generating larger droplets, and differences in droplet sizes are then amplified through collision-coalescence, resulting in a greater N c,50–3000 . Polluted conditions result in a greater N c,<50 primarily due to enhanced cloud condensation nuclei activation despite increased collision-coalescence rates compared to pristine conditions. This study provides observed quantitative patterns characterizing cloud microphysical properties as a function of key environmental parameters, offering valuable constraints for improving physics parameterizations and numerical models.

54 ENVIRONMENTAL SCIENCES↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill↗