Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Evolutionary Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

MOOSE ProbML: Parallelized probabilistic machine learning and uncertainty quantification for computational energy applications

Here, this paper presents the development and demonstration of massively parallel probabilistic machine learning (ML) and uncertainty quantification (UQ) capabilities within the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source computational platform for parallel finite element and finite volume analyses. In addressing the computational expense and uncertainties inherent in complex multiphysics simulations, this paper integrates Gaussian process (GP) variants, active learning, Bayesian inverse UQ, adaptive forward UQ, Bayesian optimization, evolutionary optimization, and Markov chain Monte Carlo (MCMC) within MOOSE. It also elaborates on the interaction among key MOOSE systems — Sampler, MultiApp, Reporter, and Surrogate — in enabling these capabilities. The modularity offered by these systems enables development of a multitude of probabilistic ML and UQ algorithms in MOOSE. Example code demonstrations include parallel active learning and parallel Bayesian inference via active learning. The impact of these developments is illustrated through five applications relevant to computational energy applications: UQ of nuclear fuel fission product release, using parallel active learning Bayesian inference; very rare events analysis in nuclear microreactors using active learning; advanced manufacturing process modeling using multi-output GPs (MOGPs) and dimensionality reduction; fluid flow using deep GPs (DGPs); and tritium transport model parameter optimization for fusion energy, using batch Bayesian optimization. These capabilities are part of the MOOSE framework.

97 - MATHEMATICS AND COMPUTING↗

Discovery of peculiar radio morphologies with ASKAP using unsupervised machine learning

Abstract We present a set of peculiar radio sources detected using an unsupervised machine learning method. We use data from the Australian Square Kilometre Array Pathfinder (ASKAP) telescope to train a self-organizing map (SOM). The radio maps from three ASKAP surveys, Evolutionary Map of Universe pilot survey (EMU-PS), Deep Investigation of Neutral Gas Origins pilot survey (DINGO), and Survey With ASKAP of GAMA-09 + X-ray (SWAG-X), are used to search for the rarest or unknown radio morphologies. We use an extension of the SOM algorithm that implements rotation and flipping invariance on astronomical sources. The SOM is trained using the images of all ‘complex’ radio sources in the EMU-PS which we define as all sources catalogued as ‘multi-component’. The trained SOM is then used to estimate a similarity score for complex sources in all surveys. We select 0.5% of the sources that are most complex according to the similarity metric and visually examine them to find the rarest radio morphologies. Among these, we find two new odd radio circle (ORC) candidates and five other peculiar morphologies. We discuss multiwavelength properties and the optical/infrared counterparts of selected peculiar sources. In addition, we present examples of conventional radio morphologies including: diffuse emission from galaxy clusters, and resolved, bent-tailed, and FR-I and FR-II type radio galaxies. We discuss the overdense environment that may be the reason behind the circular shape of ORC candidates.

Astronomy & Astrophysics↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Analysis of TRGBs (CATs) from Unsupervised, Multi-halo-field Measurements: Contrast is Key

The tip of the red giant branch (TRGB) is an apparent discontinuity of the luminosity function (LF) due to the end of the red giant evolutionary phase and is used to measure distances in the local universe. In practice, tip localization via edge detection response (EDR) relies on several methods applied on a case-by-case basis. It is hard to evaluate how individual choices affect a distance estimation using only a single host field while also avoiding confirmation bias. To devise a standardized approach, we compare unsupervised, algorithmic analyses of the TRGB in multiple halo fields per galaxy. We first optimize methods for the lowest field-to-field dispersion, including spatial filtering, smoothing, and weighting of LF, color band selection, and tip selection based on the number of likely RGB stars and the ratio of stars below versus above the tip (R). We find R, which we call the tip contrast, to be the most important indicator of the quality of EDR measurements; higher R selection can decrease field-to-field dispersion. Further, since R is found to correlate with the age or metallicity of the stellar population based on theoretical modeling, it might result in a displacement of the detected tip magnitude. We find a tip-contrast relation with a slope of -0.023 ± 0.0046 mag/ratio, an ~5σ result that can be used to correct these variations in the detections. When using TRGB to establish a distance ladder, consistent TRGB standardization using tip-contrast relation across rungs is vital to make robust cosmological measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

Recovered supernova Ia rate from simulated LSST images

Aims.TheVera C. RubinObservatory’s Legacy Survey of Space and Time (LSST) will revolutionize time-domain astronomy by detecting millions of different transients. In particular, it is expected to increase the number of known type Ia supernovae (SN Ia) by a factor of 100 compared to existing samples up to redshift ∼1.2. Such a high number of events will dramatically reduce statistical uncertainties in the analysis of the properties and rates of these objects. However, the impact of all other sources of uncertainty on the measurement of the SN Ia rate must still be evaluated. The comprehension and reduction of such uncertainties will be fundamental both for cosmology and stellar evolution studies, as measuring the SN Ia rate can put constraints on the evolutionary scenarios of different SN Ia progenitors. Methods.We used simulated data from the Dark Energy Science Collaboration (DESC) Data Challenge 2 (DC2) and LSST Data Preview 0 to measure the SN Ia rate on a 15 deg 2 region of the “wide-fast-deep” area. We selected a sample of SN candidates detected in difference images, associated them to the host galaxy with a specially developed algorithm, and retrieved their photometric redshifts. We then tested different light-curve classification methods, with and without redshift priors (albeit ignoring contamination from other transients, as DC2 contains only SN Ia). We discuss how the distribution in redshift measured for the SN candidates changes according to the selected host galaxy and redshift estimate. Results.We measured the SN Ia rate, analyzing the impact of uncertainties due to photometric redshift, host-galaxy association and classification on the distribution in redshift of the starting sample. We find that we are missing 17% of the SN Ia, on average, with respect to the simulated sample. As 10% of the mismatch is due to the uncertainty on the photometric redshift alone (which also affects classification when used as a prior), we conclude that this parameter is the major source of uncertainty. We discuss possible reduction of the errors in the measurement of the SN Ia rate, including synergies with other surveys, which may help us to use the rate to discriminate different progenitor models.

Astronomy & Astrophysics↗

GenomeFace v1.0

GenomeFace is meta-genome binning software. Metagenomic binning, the process of grouping DNA sequences into taxonomic units, is critical for understanding the functions, interactions, and evolutionary dynamics of microbial communities. We propose a deep learning approach to binning using two neural networks, one based on composition and another on environmental abundance, dynamically weighting the contribution of each based on characteristics of the input data. Trained on over 43,000 prokaryotic genomes, our network for composition-based binning is inspired by metric learning techniques used for facial recognition. Using a task-specific, multi-GPU accelerated algorithm to cluster the embeddings produced by our network, our binner leverages marker genes observed to be universally present in nearly all taxa to grade and select optimal clusters of sequences from a hierarchy of candidates. We evaluate our approach on four simulated datasets with known ground truth. Our linear time integration of marker genes recovers more near complete genomes than state of the art but computationally infeasible solutions using them, while being over an order of magnitude faster. Finally, we demonstrate the scalability and acuity of our approach by testing it on three of the largest metagenome assemblies ever performed. Compared to other binners, we produced 47%-183% more near complete genomes. From these datasets, we find over the genomes of over 3000 new candidate species which have never been previously cataloged, representing a potential 4% expansion of the known bacterial tree of life.

Lettich, Richard [Lawrence Berkeley National Labor↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗

Detecting macroevolutionary genotype–phenotype associations using error-corrected rates of protein convergence

On macroevolutionary timescales, extensive mutations and phylogenetic uncertainty mask the signals of genotype–phenotype associations underlying convergent evolution. To overcome this problem, we extended the widely used framework of non-synonymous to synonymous substitution rate ratios and developed the novel metric ω C , which measures the error-corrected convergence rate of protein evolution. While ω C distinguishes natural selection from genetic noise and phylogenetic errors in simulation and real examples, its accuracy allows an exploratory genome-wide search of adaptive molecular convergence without phenotypic hypothesis or candidate genes. Using gene expression data, we explored over 20 million branch combinations in vertebrate genes and identified the joint convergence of expression patterns and protein sequences with amino acid substitutions in functionally important sites, providing hypotheses on undiscovered phenotypes. We further extended our method with a heuristic algorithm to detect highly repetitive convergence among computationally non-trivial higher-order phylogenetic combinations. Our approach allows bidirectional searches for genotype–phenotype associations, even in lineages that diverged for hundreds of millions of years.

59 BASIC BIOLOGICAL SCIENCES↗

Deep representation learning improves prediction of LacI-mediated transcriptional repression

Significance The understanding of protein function increases with new experimental and evolutionary datasets. A major challenge is to apply machine learning to these datasets to capture essential features of protein function. Here, we analyze the experimentally determined repression function for tens of thousands of mutants of the LacI protein. This study provides a continuous, noncategorical repression value across a majority of all single mutations and for thousands of higher-order mutations. To develop a top-performing model for the prediction of repression by LacI, we compare several leading variant effect prediction algorithms. A deep representation learning paradigm, first trained across millions of proteins from all known protein families and then fine-tuned using LacI experimental data, offers the highest predictive performance of repression function.

42 ENGINEERING↗

Finding Your Niche: An Evolutionary Approach to HPC Topologies

Traditional interconnection network design approaches focus on building general network topologies by optimizing the bisection bandwidth or minimizing the network’s diameter to reduce the maximum distance between any two nodes, thus amortizing the overall execution time of the HPC workloads. While such network topologies may accommodate a wide variety of applications in general, this may result in sub-optimal performance for many frequently-executed or dynamic workloads. In this paper, instead of focusing on designing an all-encompassing, general-purpose network topology, we develop a methodology to design customized network interconnects, evolved by “finding” the optimal topologies for a particular target workload given by its communication and contention profiles. To this end, we implement a Genetic Algorithm (GA)-based approach for network topology design tailored to improve the overall execution time of a particular workload of interest. We conducted extensive experiments with well-known motifs in physics-based workloads (Sweep3D and FFT), as well as with a representative graph application (MiniVite), using the well-known Structural Simulation Toolkit (SST) Macroscale Element Library (SST/macro) simulator for network interconnect evaluation. We demonstrate that our genetic algorithm-based approach is robust enough to find the underlying optimal topology of a particular workload.

network interconnects, graph search, meta-heuristi↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

Structural characterization and computational analysis of PDZ domains in Monosiga brevicollis

Abstract Identification of the molecular networks that facilitated the evolution of multicellular animals from their unicellular ancestors is a fundamental problem in evolutionary cellular biology. Choanoflagellates are recognized as the closest extant nonmetazoan ancestors to animals. These unicellular eukaryotes can adopt a multicellular‐like “rosette” state. Therefore, they are compelling models for the study of early multicellularity. Comparative studies revealed that a number of putative human orthologs are present in choanoflagellate genomes, suggesting that a subset of these genes were necessary for the emergence of multicellularity. However, previous work is largely based on sequence alignments alone, which does not confirm structural nor functional similarity. Here, we focus on the PDZ domain, a peptide‐binding domain which plays critical roles in myriad cellular signaling networks and which underwent a gene family expansion in metazoan lineages. Using a customized sequence similarity search algorithm, we identified 178 PDZ domains in the Monosiga brevicollis proteome. This includes 11 previously unidentified sequences, which we analyzed using Rosetta and homology modeling. To assess conservation of protein structure, we solved high‐resolution crystal structures of representative M. brevicollis PDZ domains that are homologous to human Dlg1 PDZ2, Dlg1 PDZ3, GIPC, and SHANK1 PDZ domains. To assess functional conservation, we calculated binding affinities for mbGIPC, mbSHANK1, mbSNX27, and mbDLG‐3 PDZ domains from M. brevicollis . Overall, we find that peptide selectivity is generally conserved between these two disparate organisms, with one possible exception, mbDLG‐3. Overall, our results provide novel insight into signaling pathways in a choanoflagellate model of primitive multicellularity.

Gao, Melody↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Parallel derivative-free optimization for simulation-based design of behind-the-meter energy systems

In this work, the integrated design and dispatch of behind-the-meter or distributed resources (e.g. stationary battery storage and solar PV generation) is considered. A simulation-based framework is employed, generating high-fidelity results with closed-loop predictive control at a fine resolution, at the expense of high computational cost (several minutes to a few hours per design point). To address this challenge, parallel derivative-free design methods are considered. Four methods are compared, including state-of-the-art surrogate-based methods (Radial-Basis Functions and Gaussian processes) and sampling strategies, an evolutionary-based method, and a simple sequential grid refinement method. As a case study, two types of design problem with increasing complexity are considered, namely, the design of behind-the-meter resources (three design variables) and the inclusion of grid capacity (four design variables). The second yields a constrained design problem for which violations can only be determined after solving the computationally expensive simulation. For the three-dimensional case, all methods present a good performance, achieving a solution within 1% of the optimum after the first iteration, with the sequential grid refinement exhibiting the fastest convergence and achieving the best final objective value. This indicates that the parallel evaluation of multiple sampling points may be more important than the choice of method for small decision spaces. For the four-dimensional constrained case, the Genetic Algorithm presents the best tradeoff between performance and computational effort, while the rough objective function terrain generated by constraint violation penalties reduces the performance of surrogate-based methods. Contour plots with flat regions indicate flexibility in the optimal design and highlight the importance of characterizing the solution space.

24 POWER TRANSMISSION AND DISTRIBUTION↗

SCOPe: improvements to the structural classification of proteins – extended database to facilitate variant interpretation and machine learning

Abstract The Structural Classification of Proteins—extended (SCOPe, https://scop.berkeley.edu) knowledgebase aims to provide an accurate, detailed, and comprehensive description of the structural and evolutionary relationships amongst the majority of proteins of known structure, along with resources for analyzing the protein structures and their sequences. Structures from the PDB are divided into domains and classified using a combination of manual curation and highly precise automated methods. In the current release of SCOPe, 2.08, we have developed search and display tools for analysis of genetic variants we mapped to structures classified in SCOPe. In order to improve the utility of SCOPe to automated methods such as deep learning classifiers that rely on multiple alignment of sequences of homologous proteins, we have introduced new machine-parseable annotations that indicate aberrant structures as well as domains that are distinguished by a smaller repeat unit. We also classified structures from 74 of the largest Pfam families not previously classified in SCOPe, and we improved our algorithm to remove N- and C-terminal cloning, expression and purification sequences from SCOPe domains. SCOPe 2.08-stable classifies 106 976 PDB entries (about 60% of PDB entries).

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced resonant ultrasound spectroscopy for measurement of the elastic properties of multi-material systems

Understanding the elastic properties of materials is critical for their safe incorporation and predictable performance. Current methods of bulk elastic characterization often have notable limitations for in situ structural applications, with usage restricted to simple geometries and material distributions. To address these existing issues, this study sought to expand the capabilities of resonant ultrasound spectroscopy (RUS), an established nondestructive evaluation method, to include the characterization of isotropic multi-material samples. In this work, finite-element-based RUS analysis consisted of numerical simulations and experimental testing of composite samples comprised of material pairs with varying elasticity and density contrasts. Utilizing genetic algorithm inversion and mode matching, our results demonstrate that elastic properties of multi-material samples can be reliably identified within several percent of known or nominal values using a minimum number of identified resonance modes, given sample mass is held consistent. The accurate recovery of material properties for composite samples of varying material similarity and geometry expands the pool of viable samples for RUS and advances the method towards in situ inspection and evaluation.

36 MATERIALS SCIENCE↗

Measuring photometric redshifts for high-redshift radio source surveys

With the advent of deep, all-sky radio surveys, the need for ancillary data to make the most of the new, high-quality radio data from surveys like the Evolutionary Map of the Universe (EMU), GaLactic and Extragalactic All-sky Murchison Widefield Array survey eXtended, Very Large Array Sky Survey, and LOFAR Two-metre Sky Survey is growing rapidly. Radio surveys produce significant numbers of Active Galactic Nuclei (AGNs) and have a significantly higher average redshift when compared with optical and infrared all-sky surveys. Thus, traditional methods of estimating redshift are challenged, with spectroscopic surveys not reaching the redshift depth of radio surveys, and AGNs making it difficult for template fitting methods to accurately model the source. Machine Learning (ML) methods have been used, but efforts have typically been directed towards optically selected samples, or samples at significantly lower redshift than expected from upcoming radio surveys. This work compiles and homogenises a radio-selected dataset from both the northern hemisphere (making use of Sloan Digital Sky Survey optical photometry) and southern hemisphere (making use of Dark Energy Survey optical photometry). We then test commonly used ML algorithms such as k-Nearest Neighbours (kNN), Random Forest, ANNz, and GPz on this monolithic radio-selected sample. We show that kNN has the lowest percentage of catastrophic outliers, providing the best match for the majority of science cases in the EMU survey. We note that the wider redshift range of the combined dataset used allows for estimation of sources up to z = 3 before random scatter begins to dominate. When binning the data into redshift bins and treating the problem as a classification problem, we are able to correctly identify ≈ 76% of the highest redshift sources—sources at redshift z > 2.51 —as being in either the highest bin (z > 2.51) or second highest (z = 2.25).

79 ASTRONOMY AND ASTROPHYSICS↗