Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Curiosity driven exploration to optimize structure–property learning in microscopy

Rapidly determining structure–property correlations in materials is an important challenge in better understanding fundamental mechanisms and greatly assists in materials design. In microscopy, imaging data provides a direct measurement of the local structure, while spectroscopic measurements provide relevant functional property information. Deep kernel active learning approaches have been utilized to rapidly map local structure to functional properties in microscopy experiments, but are computationally expensive for multi-dimensional and correlated output spaces. Here, we present an alternative lightweight curiosity algorithm which actively samples regions with unexplored structure–property relations, utilizing a deep-learning based surrogate model for error prediction. We show that the algorithm outperforms random sampling for predicting properties from structures, and provides a convenient tool for efficient mapping of structure–property relationships in materials science.

36 MATERIALS SCIENCE↗

Jacobian sparsity detection using Bloom filters

Determining Jacobian sparsity structure is an important step in the efficient computation of sparse Jacobians. We introduce a new method for determining Jacobian sparsity patterns by combining bit vector probing with Bloom filters. In conclusion, we further refine Bloom filter probing by combining it with hierarchical probing to yield a highly effective strategy for Jacobian sparsity pattern determination.

Bloom filter↗

Universality class for loopless invasion percolation models and a percolation avalanche burst model for hydraulic fracturing

Invasion percolation is a model that was originally proposed to describe growing networks of fractures. In this work, we describe a loopless algorithm on random lattices, coupled with an avalanche-based model for bursts. The model reproduces the characteristic b-value seismicity and spatial distribution of bursts consistent with earthquakes resulting from hydraulic fracturing (“fracking”). We test models for both site invasion percolation and bond invasion percolation. These have differences on the scale of site and bond lengths l. But since the networks are characterized by their large-scale behavior, l << L, we find small differences between scaling exponents. Though data may not differentiate between models, our results suggest that both models belong to different universality classes.

58 GEOSCIENCES↗

A Machine Learning-Based Method to Estimate Transformer Primary-Side Voltages with Limited Customer-Side AMI Measurements

Distribution control applications such as volt/var optimization, network reconfiguration, and distribution automation require accurate knowledge of the distribution system state. The lack of sufficient sensors on the primary side of distribution networks often limits the accuracy of the control decisions by these applications. The deployment of advanced metering infrastructure (AMI) provides utilities an opportunity to translate the AMI data on the secondary onto the primary so that it can be used as pseudo-measurements to augment the limited existing measurements on the primary. This paper develops a machine learning based approach for estimating service transformer primary-side voltages by using limited secondary-side AMI measurement. The machine learning model is developed by using random forest algorithm. The estimated primary-side voltages can be used by utilities as pseudo-measurements for distribution control applications. The detailed secondary model topology, which is an essential input data for many existing algorithms, is not required for the proposed method. The performance of the proposed method is validated by using AMI measurements from the field and an actual distribution feeder model of San Diego Gas & Electric Company.

advanced metering infrastructure↗

A Genome-Based Model to Predict the Virulence of Pseudomonas aeruginosa Isolates

ABSTRACT: Variation in the genome of Pseudomonas aeruginosa , an important pathogen, can have dramatic impacts on the bacterium’s ability to cause disease. We therefore asked whether it was possible to predict the virulence of P. aeruginosa isolates based on their genomic content. We applied a machine learning approach to a genetically and phenotypically diverse collection of 115 clinical P. aeruginosa isolates using genomic information and corresponding virulence phenotypes in a mouse model of bacteremia. We defined the accessory genome of these isolates through the presence or absence of accessory genomic elements (AGEs), sequences present in some strains but not others. Machine learning models trained using AGEs were predictive of virulence, with a mean nested cross-validation accuracy of 75% using the random forest algorithm. However, individual AGEs did not have a large influence on the algorithm’s performance, suggesting instead that virulence predictions are derived from a diffuse genomic signature. These results were validated with an independent test set of 25 P. aeruginosa isolates whose virulence was predicted with 72% accuracy. Machine learning models trained using core genome single-nucleotide variants and whole-genome k-mers also predicted virulence. Our findings are a proof of concept for the use of bacterial genomes to predict pathogenicity in P. aeruginosa and highlight the potential of this approach for predicting patient outcomes. IMPORTANCE Pseudomonas aeruginosa is a clinically important Gram-negative opportunistic pathogen. P. aeruginosa shows a large degree of genomic heterogeneity both through variation in sequences found throughout the species (core genome) and through the presence or absence of sequences in different isolates (accessory genome). P. aeruginosa isolates also differ markedly in their ability to cause disease. In this study, we used machine learning to predict the virulence level of P. aeruginosa isolates in a mouse bacteremia model based on genomic content. We show that both the accessory and core genomes are predictive of virulence. This study provides a machine learning framework to investigate relationships between bacterial genomes and complex phenotypes such as virulence.

59 BASIC BIOLOGICAL SCIENCES↗

Dependence of Convective Cloud Microphysical Properties on Environmental Conditions during the TRACER and ESCAPE Field Campaigns: A Synergistic Approach of Observations, Machine Learning and Parcel Models

The sensitivity of convective clouds to aerosols and their interactions with environment, combined with limited observational constraints in parameterizations, introduces significant uncertainties in atmospheric models. Here, this study investigates the dependence of convective cloud microphysical properties on environmental conditions using a synergistic approach that combines unique observations from the TRACER and ESCAPE field campaigns, machine learning techniques, and parcel model simulations with a super-droplet microphysics scheme. A random forest algorithm identifies in-situ vertical velocity (w), temperature (T), and surface fine-mode aerosol mass concentration as the three most important environmental conditions influencing cloud properties including liquid water content (LWC), number concentration for particles with D max < 50 μm (N c ,<50), 50 μm ≤ D max ≤ 3000 μm (N c,50–3000 ), and droplet effective diameter (D e ). Results show that LWC, N c,<50 , and N c,50–3000 significantly increase with w in updrafts. Across w bins, as T decreases, LWC, D e , and N c,50–3000 increase, while N c,<50 decreases, which are closely linked to the distance above cloud bases. Warmer cloud bases yield higher LWC, greater N c,50–3000 , and smaller N c,<50 , while polluted environments produce greater N c,<50 . Parcel model simulations successfully replicate these observed dependencies. The simulation results indicate that warmer cloud bases enhance condensation generating larger droplets, and differences in droplet sizes are then amplified through collision-coalescence, resulting in a greater N c,50–3000 . Polluted conditions result in a greater N c,<50 primarily due to enhanced cloud condensation nuclei activation despite increased collision-coalescence rates compared to pristine conditions. This study provides observed quantitative patterns characterizing cloud microphysical properties as a function of key environmental parameters, offering valuable constraints for improving physics parameterizations and numerical models.

54 ENVIRONMENTAL SCIENCES↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill↗

Advancing Multiscale Simulation of Plasma-Surface Interfaces

We report the development of an atomistic-informed, surface-state-dependent predictive model for particle exchange in a carbon-tungsten plasma-surface interface. The predictive model uses machine learning (ML) techniques to learn the energy and angular distributions for particle exchange and rate functions for surface state evolution from molecular dynamics simulations of cumulative bombardment of tungsten by energetic carbon ions. Each predictive component is sensitive to the energy and trajectory of incident plasma species and the surface state. The surface state is represented by a set of surface state descriptors, which were derived from the atomistic surface state for each independent carbon bombardment event. These descriptors are representative of the composition and degree of amorphization of the outermost angstrom of surface material and were chosen to optimize predictive performance for particle exchange at the interface. The distributions for particle exchange (reflection/sputtering) are demonstrated to vary with each surface state descriptor, motivating the development of surface-state-dependent particle exchange models for plasma simulations. The performance of various ML methods was compared, including polynomial quantile regression, artificial neural networks, k-nearest neighbors, and random forest algorithms, with polynomial regression performing the best for interpolation and extrapolation of learned relationships. In addition to the particle exchange model, a neutral network was developed and used to identify data sufficiency throughout surface descriptor space, which will enable real-time feedback during future data production to ensure data is produced where it is most needed, and we provide commentary on improvements to the data production workflow for future endeavors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The direct and indirect drivers shaping RNA viral communities in grassland soil

Recent studies have revealed diverse RNA viral communities in soils. Yet, how environmental factors influence soil RNA viruses remains largely unknown. Here, we recovered RNA viral communities from 24 metatranscriptomes sequenced from grassland soils managed under a range of environmental conditions including 1) water content: 100% and 25% water holding capacity, 2) plant presence: planted with tall wheatgrass (Thinopyrum ponticum) and bare soil, 3) cultivar type: Alkar and Jose, and 4) soil depth: 0-5 cm and 15-25 cm. The recovered RNA viral communities were novel with nearly one-third of the RNA viral contigs uniquely detected in the studied grassland. The classified RNA viral contigs are mostly known as eukaryotic RNA viruses (74.7%) belonging to Phyla Duplornaviricota, Kitrinoviricota, Lenarviricota, and Pisuviricota. Eukaryotic RNA viruses of Family Mitoviridae as well as their natural hosts, Fungi, are one of the most dominant taxa. Consistent with the results of nonmetric multidimensional scaling analysis, the four environmental conditions (water content, plant presence, cultivar, and soil depth) significantly influence the assemblages of soil RNA viral communities as suggested by the correlation analysis and the random forest algorithm. The modularity analysis of the factor network and the structural equation modeling further support the hierarchical associations among the four environmental factors and the community factors representing the co-existing eukaryotic, prokaryotic, and RNA viral communities. The soil water content, plant presence, and type of cultivar demonstrate a significant positive impact on eukaryotic RNA viral richness directly as well as indirectly on eukaryotic RNA viral abundance via influencing the co-existing eukaryotic members in this soil. Our data also provide statistical support for the negative influence of soil depth on soil eukaryotic richness and abundances resulting in its indirect impact on soil eukaryotic RNA viral communities. This study provides field-relevant information on how environmental and community factors collectively shape soil RNA communities and contribute to ecological understanding of RNA viral survival under various environmental conditions and virus-host interactions in soil.

Wu, Ruonan↗

Microbial communities in the liver and brain are informative for postmortem submersion interval estimation in the late phase of decomposition: A study in mouse cadavers recovered from freshwater

Introduction Bodies recovered from water, especially in the late phase of decomposition, pose difficulties to the investigating authorities. Various methods have been proposed for postmortem submersion interval (PMSI) estimation and drowning identification, but some limitations remain. Many recent studies have proved the value of microbiota succession in viscera for postmortem interval estimation. Nevertheless, the visceral microbiota succession and its application for PMSI estimation and drowning identification require further investigation. Methods In the current study, mouse drowning and CO 2 asphyxia models were developed, and cadavers were immersed in freshwater for 0 to 14 days. Microbial communities in the liver and brain were characterized via 16S rDNA high-throughput sequencing. Results Only livers and brains collected from 5 to 14 days postmortem were qualified for sequencing. There was significant variation between microbiota from liver and brain. Differences in microbiota between the cadavers of mice that had drowned and those only subjected to postmortem submersion decreased over the PMSI. Significant successions in microbial communities were observed among the different subgroups within the late phase of the PMSI in livers and brains. Eighteen taxa in the liver which were mainly related to Clostridium_sensu_stricto and Aeromonas , and 26 taxa in the brain which were mainly belonged to Clostridium_sensu_stricto , Acetobacteroides , and Limnochorda , were selected as potential biomarkers for PMSI estimation based on a random forest algorithm. The PMSI estimation models established yielded accurate prediction results with mean absolute errors ± the standard error of 1.282 ± 0.189 d for the liver and 0.989 ± 0.237 d for the brain. Conclusions The present study provides novel information on visceral postmortem microbiota succession in corpses submerged in freshwater which sheds new light on PMSI estimation based on the liver and brain in forensic practice.

Wang, Linlin↗

Deep Learning for In-Situ Layer Quality Monitoring during Laser-Based Directed Energy Deposition (LB-DED) Additive Manufacturing Process

Defects are a leading issue for the rejection of parts manufactured through the Directed Energy Deposition (DED) Additive Manufacturing (AM) process. In an attempt to illuminate and advance in situ quality monitoring and control of workpieces, we present an innovative data-driven method that synchronously collects sensing data and AM process parameters with a low sampling rate during the DED process. The proposed data-driven technique determines the important influences that individual printing parameters and sensing features have on prediction at the inter-layer qualification to perform feature selection. Three Machine Learning (ML) algorithms including Random Forest (RF), Support Vector Machine (SVM), and Convolutional Neural Network (CNN) are used. During post-production, a threshold is applied to detect low-density occurrences such as porosity sizes and quantities from CT scans that render individual layers acceptable or unacceptable. This information is fed to the ML models for training. Training/testing are completed offline on samples deemed “high-quality” and “low-quality”, utilizing only features recorded from the build process. CNN results show that the classification of acceptable/unacceptable layers can reach between 90% accuracy while training/testing on a “high-quality” sample and dip to 65% accuracy when trained/tested on “low-quality”/“high-quality” (respectively), indicating over-fitting but showing CNN as a promising inter-layer classifier.

36 MATERIALS SCIENCE↗

Machine Learning-Based Classification of Lignocellulosic Biomass from Pyrolysis-Molecular Beam Mass Spectrometry Data

High-throughput analysis of biomass is necessary to ensure consistent and uniform feedstocks for agricultural and bioenergy applications and is needed to inform genomics and systems biology models. Pyrolysis followed by mass spectrometry such as molecular beam mass spectrometry (py-MBMS) analyses are becoming increasingly popular for the rapid analysis of biomass cell wall composition and typically require the use of different data analysis tools depending on the need and application. Here, the authors report the py-MBMS analysis of several types of lignocellulosic biomass to gain an understanding of spectral patterns and variation with associated biomass composition and use machine learning approaches to classify, differentiate, and predict biomass types on the basis of py-MBMS spectra. Py-MBMS spectra were also corrected for instrumental variance using generalized linear modeling (GLM) based on the use of select ions relative abundances as spike-in controls. Machine learning classification algorithms e.g., random forest, k-nearest neighbor, decision tree, Gaussian Naïve Bayes, gradient boosting, and multilayer perceptron classifiers were used. The k-nearest neighbors (k-NN) classifier generally performed the best for classifications using raw spectral data, and the decision tree classifier performed the worst. After normalization of spectra to account for instrumental variance, all the classifiers had comparable and generally acceptable performance for predicting the biomass types, although the k-NN and decision tree classifiers were not as accurate for prediction of specific sample types. Gaussian Naïve Bayes (GNB) and extreme gradient boosting (XGB) classifiers performed better than the k-NN and the decision tree classifiers for the prediction of biomass mixtures. The data analysis workflow reported here could be applied and extended for comparison of biomass samples of varying types, species, phenotypes, and/or genotypes or subjected to different treatments, environments, etc. to further elucidate the sources of spectral variance, patterns, and to infer compositional information based on spectral analysis, particularly for analysis of data without a priori knowledge of the feedstock composition or identity.

59 BASIC BIOLOGICAL SCIENCES↗

Image-Based Methods to Score Fungal Pathogen Symptom Progression and Severity in Excised Arabidopsis Leaves

Image-based symptom scoring of plant diseases is a powerful tool for associating disease resistance with plant genotypes. Advancements in technology have enabled new imaging and image processing strategies for statistical analysis of time-course experiments. There are several tools available for analyzing symptoms on leaves and fruits of crop plants, but only a few are available for the model plant Arabidopsis thaliana (Arabidopsis). Arabidopsis and the model fungus Botrytis cinerea (Botrytis) comprise a potent model pathosystem for the identification of signaling pathways conferring immunity against this broad host-range necrotrophic fungus. Here, we present two strategies to assess severity and symptom progression of Botrytis infection over time in Arabidopsis leaves. Thus, a pixel classification strategy using color hue values from red-green-blue (RGB) images and a random forest algorithm was used to establish necrotic, chlorotic, and healthy leaf areas. Secondly, using chlorophyll fluorescence (ChlFl) imaging, the maximum quantum yield of photosystem II (Fv/Fm) was determined to define diseased areas and their proportion per total leaf area. Both RGB and ChlFl imaging strategies were employed to track disease progression over time. This has provided a robust and sensitive method for detecting sensitive or resistant genetic backgrounds. A full methodological workflow, from plant culture to data analysis, is described.

59 BASIC BIOLOGICAL SCIENCES↗

Target Selection and Validation of DESI Quasars

The Dark Energy Spectroscopic Instrument (DESI) survey will measure large-scale structures using quasars as direct tracers of dark matter in the redshift range 0.9 < z < 2.1 and using Lyα forests in quasar spectra at z > 2.1. We present several methods to select candidate quasars for DESI, using input photometric imaging in three optical bands (g, r, z) from the DESI Legacy Imaging Surveys and two infrared bands (W1, W2) from the Wide-field Infrared Survey Explorer. These methods were extensively tested during the Survey Validation of DESI. In this paper, we report on the results obtained with the different methods and present the selection we optimized for the DESI main survey. The final quasar target selection is based on a random forest algorithm and selects quasars in the magnitude range of 16.5 < r < 23. Visual selection of ultra-deep observations indicates that the main selection consists of 71% quasars, 16% galaxies, 6% stars, and 7% inconclusive spectra. Using the spectra based on this selection, we build an automated quasar catalog that achieves a fraction of true QSOs higher than 99% for a nominal effective exposure time of ~1000 s. With a 310 deg -2 target density, the main selection allows DESI to select more than 200 deg -2 quasars (including 60 deg -2 quasars with z > 2.1), exceeding the project requirements by 20%. The redshift distribution of the selected quasars is in excellent agreement with quasar luminosity function predictions.

79 ASTRONOMY AND ASTROPHYSICS↗

The analysis of the pilot's cognitive and decision processes

Articles are presented on pilot performance in zero-visibility precision approach, failure detection by pilots during automatic landing, experiments in pilot decision-making during simulated low visibility approaches, a multinomial maximum likelihood program, and a random search algorithm for laboratory computers. Other topics discussed include detection of system failures in multi-axis tasks and changes in pilot workload during an instrument landing.

Curry, R. E.↗