Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evolutionary computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Deep representation learning improves prediction of LacI-mediated transcriptional repression

Significance The understanding of protein function increases with new experimental and evolutionary datasets. A major challenge is to apply machine learning to these datasets to capture essential features of protein function. Here, we analyze the experimentally determined repression function for tens of thousands of mutants of the LacI protein. This study provides a continuous, noncategorical repression value across a majority of all single mutations and for thousands of higher-order mutations. To develop a top-performing model for the prediction of repression by LacI, we compare several leading variant effect prediction algorithms. A deep representation learning paradigm, first trained across millions of proteins from all known protein families and then fine-tuned using LacI experimental data, offers the highest predictive performance of repression function.

42 ENGINEERING↗

The influence of Hickson-like compact group environment on galaxy luminosities

Compact groups of galaxies are devised as extreme environments where interactions may drive galaxy evolution. In this work, we analysed whether the luminosities of galaxies inhabiting compact groups differ from those of galaxies in loose galaxy groups. We computed the luminosity functions of galaxy populations inhabiting a new sample of 1412 Hickson-like compact groups of galaxies identified in the Sloan Digital Sky Survey Data Release 16. We observed a characteristic absolute magnitude for galaxies in compact groups brighter than that observed in the field or loose galaxy systems. We also observed a deficiency of faint galaxies in compact groups in comparison with loose systems. Our analysis showed that the brightening is mainly due to galaxies inhabiting the more massive compact groups. In contrast to what is observed in loose systems, where only the luminosities of Red (and Early) galaxies show a dependency with group mass, luminosities of Red and Blue (also Early and Late) galaxies in compact groups are affected similarly as a function of group virial mass. When using Hubble types, we observed that elliptical galaxies in compact groups are the brightest galaxy population, and groups dominated by an elliptical galaxy also display the brightest luminosities in comparison with those dominated by spiral galaxies. Moreover, we show that the general luminosity trends can be reproduced using a mock catalogue obtained from a semi-analytical model of galaxy formation. These results suggest that the inner extreme environment in compact groups prompts a different evolutionary history for their galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

Design of diverse, functional mitochondrial targeting sequences across eukaryotic organisms using variational autoencoder

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

59 BASIC BIOLOGICAL SCIENCES↗

Naturally ornate RNA-only complexes revealed by cryo-EM

The structures of natural RNAs remain poorly characterized and may hold numerous surprises. Here we report three-dimensional structures of three large ornate bacterial RNAs using cryo-electron microscopy (cryo-EM). GOLLD (Giant, Ornate, Lake- and Lactobacillales-Derived), ROOL (Rumen-Originating, Ornate, Large) and OLE (Ornate Large Extremophilic) RNAs form homo-oligomeric complexes whose stoichiometries are retained at lower concentrations than measured in cells. OLE RNA forms a dimeric complex with long co-axial pipes spanning two monomers. Both GOLLD and ROOL form distinct RNA-only multimeric nanocages with diameters larger than the ribosome, each empty except for a disordered loop. Extensive intramolecular and intermolecular A-minor interactions, kissing loops, an unusual A–A helix and other interactions stabilize the three complexes. Sequence covariation analysis of these large RNAs reveals evolutionary conservation of intermolecular interactions, supporting the biological importance of large, ornate RNA quaternary structures that can assemble without any involvement of proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Toward Consistent High-Fidelity Quantum Learning on Unstable Devices via Efficient In-Situ Calibration

In the near-term noisy intermediate-scale quantum (NISQ) era, high noise will significantly reduce the fidelity of quantum computing. What's worse, recent works reveal that the noise on quantum devices is not stable, that is, the noise is dynamically changing over time. This leads to an imminent challenging problem: At run-time, is there a way to efficiently achieve a consistent high-fidelity quantum system on unstable devices? To study this problem, we take quantum learning (a.k.a., variational quantum algorithm) as a vehicle, which has a wide range of applications, such as combinatorial optimization and machine learning. A straightforward approach is to optimize a variational quantum circuit (VQC) with a parameter-shift approach on the target quantum device before using it; however, the optimization has an extremely high time cost, which is not practical at run-time. To address the pressing issue, in this paper, we proposed a novel quantum pulse-based noise adaptation framework, namely QuPAD. In the proposed framework, first, we identify that the CNOT gate is the fidelity bottleneck of the conventional VQC, and we employ a more robust parameterized multi-qubit gate (i.e., Rzx gate) to replace CNOT gate. Second, by benchmarking Rzx gate with different parameters, we build a fitting function for each coupling qubit pair, such that the deviation between the theoretic output of Rzx gate and its on-device output under a given pulse amplitude and duration can be efficiently predicted. On top of this, an evolutionary algorithm is devised to identify the pulse amplitude and duration of each Rzx gate (i.e., calibration) and find the quantum circuits with high fidelity. Experiments show that the runtime on quantum devices of QuPAD with 8–10 qubits is less than 15 minutes, which is up to 270 x faster than the parameter-shift approach. In addition, compared to the vanilla VQC as a baseline, QuPAD can achieve 59.33% accuracy gain on a classification task, and average 66.34% closer to ground state energy for molecular simulation.

Hu, Zhirui↗

The Swan: Data-driven Inference of Stellar Surface Gravities for Cool Stars from Photometric Light Curves

Stellar light curves are well known to encode physical stellar properties. Precise, automated, and computationally inexpensive methods to derive physical parameters from light curves are needed to cope with the large influx of these data from space-based missions such as Kepler and TESS. Here we present a new methodology that we call “The Swan,” a fast, generalizable, and effective approach for deriving stellar surface gravity (logg) for main-sequence, subgiant, and red giant stars from Kepler light curves using local linear regression on the full frequency content of Kepler long-cadence power spectra. With this inexpensive data-driven approach, we recover logg to a precision of ~0.02 dex for 13,822 stars with seismic logg values between 0.2 and 4.4 dex and ~0.11 dex for 4646 stars with Gaia-derived logg values between 2.3 and 4.6 dex. We further develop a signal-to-noise metric and find that granulation is difficult to detect in many cool main-sequence stars (T {sub eff} ≲ 5500 K), in particular K dwarfs. By combining our logg measurements with Gaia radii, we derive empirical masses for 4646 subgiant and main-sequence stars with a median precision of ~7%. Finally, we demonstrate that our method can be used to recover logg to a similar mean absolute deviation precision for a TESS baseline of 27 days. Our methodology can be readily applied to photometric time series observations to infer stellar surface gravities to high precision across evolutionary states.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Stellar migration and chemical enrichment in the milky way disc: a hybrid model

ABSTRACT We develop a hybrid model of galactic chemical evolution that combines a multiring computation of chemical enrichment with a prescription for stellar migration and the vertical distribution of stellar populations informed by a cosmological hydrodynamic disc galaxy simulation. Our fiducial model adopts empirically motivated forms of the star formation law and star formation history, with a gradient in outflow mass loading tuned to reproduce the observed metallicity gradient. With this approach, the model reproduces many of the striking qualitative features of the Milky Way disc’s abundance structure: (i) the dependence of the [O/Fe]–[Fe/H] distribution on radius Rgal and mid-plane distance |z|; (ii) the changing shapes of the [O/H] and [Fe/H] distributions with Rgal and |z|; (iii) a broad distribution of [O/Fe] at sub-solar metallicity and changes in the [O/Fe] distribution with Rgal, |z|, and [Fe/H]; (iv) a tight correlation between [O/Fe] and stellar age for [O/Fe] > 0.1; (v) a population of young and intermediate-age α-enhanced stars caused by migration-induced variability in the Type Ia supernova rate; (vi) non-monotonic age–[O/H] and age–[Fe/H] relations, with large scatter and a median age of ∼4 Gyr near solar metallicity. Observationally motivated models with an enhanced star formation rate ∼2 Gyr ago improve agreement with the observed age–[Fe/H] and age–[O/H] relations, but worsen agreement with the observed age–[O/Fe] relation. None of our models predict an [O/Fe] distribution with the distinct bimodality seen in the observations, suggesting that more dramatic evolutionary pathways are required. All code and tables used for our models are publicly available through the Versatile Integrator for Chemical Evolution (VICE; https://pypi.org/project/vice).

79 ASTRONOMY AND ASTROPHYSICS↗

Geometry-Driven Detection, Tracking and Visual Analysis of Viscous and Gravitational Fingers

Viscous and gravitational flow instabilities cause a displacement front to break up into finger-like fluids. The detection and evolutionary analysis of these fingering instabilities are critical in multiple scientific disciplines such as fluid mechanics and hydrogeology. However, previous detection methods of the viscous and gravitational fingers are based on density thresholding, which provides limited geometric information of the fingers. The geometric structures of fingers and their evolution are important yet little studied in the literature. In this work, we explore the geometric detection and evolution of the fingers in detail to elucidate the dynamics of the instability. Additionally, we propose a ridge voxel detection method to guide the extraction of finger cores from three-dimensional (3D) scalar fields. After skeletonizing finger cores into skeletons, we design a spanning tree based approach to capture how fingers branch spatially from the finger skeletons. Finally, we devise a novel geometric-glyph augmented tracking graph to study how the fingers and their branches grow, merge, and split over time. Feedback from earth scientists demonstrates the usefulness of our approach to performing spatio-temporal geometric analyses of fingers.

97 MATHEMATICS AND COMPUTING↗

Origin and evolution of HIV-1 subtype A6

Background: HIV outbreaks in the Former Soviet Union (FSU) countries were characterized by repeated transmission of the HIV variant AFSU, which is now classified as a distinct subtype A sub-subtype called A6. The current study used phylogenetic/phylodynamic and signature mutation analyses to determine likely evolutionary relationship between subtype A6 and other subtype A sub-subtypes. Methods: For this study, an initial Maximum Likelihood phylogenetic analysis was performed using a total of 553 full-length, publicly available, reverse transcriptase sequences, from A1, A2, A3, A4, A5, and A6 sub-subtypes of subtype A. For phylogenetic clustering and signature mutation analysis, a total of 5961 and 3959 pol and env sequences, respectively, were used. Results: Phylogenetic and signature mutation analysis showed that HIV-1 sub-subtype A6 likely originated from sub-subtype A1 of African origin. A6 and A1 pol and env genes shared several signature mutations that indicate genetic similarity between the two subtypes. For A6, tMRCA dated to 1975, 15 years later than that of A1. Conclusion: The current study provides insights into the evolution and diversification of A6 in the backdrop of FSU countries and indicates that A6 in FSU countries evolved from A1 of African origin and is getting bridged outside the FSU region.

60 APPLIED LIFE SCIENCES↗

Early prediction of antigenic transitions for influenza A/H3N2

Influenza A/H3N2 is a rapidly evolving virus which experiences major antigenic transitions every two to eight years. Anticipating the timing and outcome of transitions is critical to developing effective seasonal influenza vaccines. Using a published phylodynamic model of influenza transmission, we identified indicators of future evolutionary success for an emerging antigenic cluster and quantified fundamental trade-offs in our ability to make such predictions. The eventual fate of a new cluster depends on its initial epidemiological growth rate––which is a function of mutational load and population susceptibility to the cluster––along with the variance in growth rate across co-circulating viruses. Logistic regression can predict whether a cluster at 5% relative frequency will eventually succeed with ~80% sensitivity, providing up to eight months advance warning. As a cluster expands, the predictions improve while the lead-time for vaccine development and other interventions decreases. However, attempts to make comparable predictions from 12 years of empirical influenza surveillance data, which are far sparser and more coarse-grained, achieve only 56% sensitivity. By expanding influenza surveillance to obtain more granular estimates of the frequencies of and population-wide susceptibility to emerging viruses, we can better anticipate major antigenic transitions. This provides added incentives for accelerating the vaccine production cycle to reduce the lead time required for strain selection.

59 BASIC BIOLOGICAL SCIENCES↗

DEVELOPMENT AND APPLICATION OF RISK ANALYSIS TOOLKIT FOR PLANT RESOURCE OPTIMIZATION

This paper presents the development of methods and tools that are being designed to optimize plant operations (e.g., maintenance/replacement schedules and optimal maintenance postures for plant components) in a manner that is more cost effective than current approaches and makes better use of available component health and cost data. These methods include both data- and model-based optimization methods. Model-based optimization methods directly include reliability and cost models to determine an optimal plant operational strategy. We consider gradient-based and evolutionary (based on genetic algorithms) optimization methods. The second class of methods target more specific use cases (e.g., project schedule optimization) and are not based on reliability models directly, but they require specific component reliability and cost data. This class of methods is based on variants of the knapsack problem with an aim to determine an optimal project schedule that maximizes the overall NPV. This paper also presents multi-objective methods designed to identify an optimal maintenance posture based on a Pareto frontier analysis. Rather than dictating the “right” tradeoff (i.e., identify the absolute best posture), we show how it is possible to perform a trade space exploration approach (i.e., identify value and costs of several postures and let the analysis account for desired value and cost metrics). This is performed by identifying maintenance postures that maximize value (e.g., system availability) and minimize operational costs, i.e., the Pareto frontier in a value-cost trade space. For all these methods we present detailed applicative examples that show their validity from a decision-making perspective.

97 - MATHEMATICS AND COMPUTING↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗

A deep dilated convolutional residual network for predicting interchain contacts of protein homodimers

Abstract Motivation Deep learning has revolutionized protein tertiary structure prediction recently. The cutting-edge deep learning methods such as AlphaFold can predict high-accuracy tertiary structures for most individual protein chains. However, the accuracy of predicting quaternary structures of protein complexes consisting of multiple chains is still relatively low due to lack of advanced deep learning methods in the field. Because interchain residue–residue contacts can be used as distance restraints to guide quaternary structure modeling, here we develop a deep dilated convolutional residual network method (DRCon) to predict interchain residue–residue contacts in homodimers from residue–residue co-evolutionary signals derived from multiple sequence alignments of monomers, intrachain residue–residue contacts of monomers extracted from true/predicted tertiary structures or predicted by deep learning, and other sequence and structural features. Results Tested on three homodimer test datasets (Homo_std dataset, DeepHomo dataset and CASP-CAPRI dataset), the precision of DRCon for top L/5 interchain contact predictions (L: length of monomer in a homodimer) is 43.46%, 47.10% and 33.50% respectively at 6 Å contact threshold, which is substantially better than DeepHomo and DNCON2_inter and similar to Glinter. Moreover, our experiments demonstrate that using predicted tertiary structure or intrachain contacts of monomers in the unbound state as input, DRCon still performs well, even though its accuracy is lower than using true tertiary structures in the bound state are used as input. Finally, our case study shows that good interchain contact predictions can be used to build high-accuracy quaternary structure models of homodimers. Availability and implementation The source code of DRCon is available at https://github.com/jianlin-cheng/DRCon. The datasets are available at https://zenodo.org/record/5998532#.YgF70vXMKsB. Supplementary information Supplementary data are available at Bioinformatics online.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reference-free structural variant detection in microbiomes via long-read co-assembly graphs

Motivation: The study of bacterial genome dynamics is vital for understanding the mechanisms underlying microbial adaptation, growth, and their impact on host phenotype. Structural variants (SVs), genomic alterations of 50 base pairs or more, play a pivotal role in driving evolutionary processes and maintaining genomic heterogeneity within bacterial populations. While SV detection in isolate genomes is relatively straightforward, metagenomes present broader challenges due to the absence of clear reference genomes and the presence of mixed strains. In response, our proposed method rhea, forgoes reference genomes and metagenome-assembled genomes (MAGs) by encompassing all metagenomic samples in a series (time or other metric) into a single co-assembly graph. The log fold change in graph coverage between successive samples is then calculated to call SVs that are thriving or declining. Results: We show rhea to outperform existing methods for SV and horizontal gene transfer (HGT) detection in two simulated mock metagenomes, particularly as the simulated reads diverge from reference genomes and an increase in strain diversity is incorporated. We additionally demonstrate use cases for rhea on series metagenomic data of environmental and fermented food microbiomes to detect specific sequence alterations between successive time and temperature samples, suggesting host advantage. Our approach leverages previous work in assembly graph structural and coverage patterns to provide versatility in studying SVs across diverse and poorly characterized microbial communities for more comprehensive insights into microbial gene flux.

59 BASIC BIOLOGICAL SCIENCES↗

StructuredFuzzer: Fuzzing Structured Text-Based Control Logic Applications

Rigorous testing methods are essential for ensuring the security and reliability of industrial controller software. Fuzzing, a technique that automatically discovers software bugs, has also proven effective in finding software vulnerabilities. Unsurprisingly, fuzzing has been applied to a wide range of platforms, including programmable logic controllers (PLCs). However, current approaches, such as coverage-guided evolutionary fuzzing implemented in the popular fuzzer American Fuzzy Lop Plus Plus (AFL++), are often inadequate for finding logical errors and bugs in PLC control logic applications. They primarily target generic programming languages like C/C++, Java, and Python, and do not consider the unique characteristics and behaviors of PLCs, which are often programmed using specialized programming languages like Structured Text (ST). Furthermore, these fuzzers are ill suited to deal with complex input structures encapsulated in ST, as they are not specifically designed to generate appropriate input sequences. This renders the application of traditional fuzzing techniques less efficient on these platforms. To address this issue, this paper presents a fuzzing framework designed explicitly for PLC software to discover logic bugs in applications written in ST specified by the IEC 61131-3 standard. The proposed framework incorporates a custom-tailored PLC runtime and a fuzzer designed for the purpose. We demonstrate its effectiveness by fuzzing a collection of ST programs that were crafted for evaluation purposes. We compare the performance against a popular fuzzer, namely, AFL++. The proposed fuzzing framework demonstrated its capabilities in our experiments, successfully detecting logic bugs in the tested PLC control logic applications written in ST. On average, it was at least 83 times faster than AFL++, and in certain cases, for example, it was more than 23,000 times faster.

47 OTHER INSTRUMENTATION↗

Inference of Chromosome-Length Haplotypes Using Genomic Data of Three or a Few More Single Gametes

Compared with genomic data of individual markers, haplotype data provide higher resolution for DNA variants, advancing our knowledge in genetics and evolution. Although many computational and experimental phasing methods have been developed for analyzing diploid genomes, it remains challenging to reconstruct chromosome-scale haplotypes at low cost, which constrains the utility of this valuable genetic resource. Gamete cells, the natural packaging of haploid complements, are ideal materials for phasing entire chromosomes because the majority of the haplotypic allele combinations has been preserved. Therefore, compared with the current diploid-based phasing methods, using haploid genomic data of single gametes may substantially reduce the complexity in inferring the donor’s chromosomal haplotypes. In this study, we developed the first easy-to-use R package, Hapi, for inferring chromosome-length haplotypes of individual diploid genomes with only a few gametes. Hapi outperformed other phasing methods when analyzing both simulated and real single gamete cell sequencing data sets. The results also suggested that chromosome-scale haplotypes may be inferred by using as few as three gametes, which has pushed the boundary to its possible limit. The single gamete cell sequencing technology allied with the cost-effective Hapi method will make large-scale haplotype-based genetic studies feasible and affordable, promoting the use of haplotype data in a wide range of research.

59 BASIC BIOLOGICAL SCIENCES↗

A User’s Guide to the PLTEMP/ANL Code

PLTEMP/ANL V4.4 is a program that obtains a steady-state flow and temperature solution, in geometrical extent, for a nuclear reactor core, or a single fuel assembly, including the effect of manufacturing and modelling uncertainties by the means of hot channel factors. It is based on an evolutionary sequence of codes originally used for calculating plate temperatures, hence “PLTEMP”, developed at Argonne National Laboratory over the past 30 years. Fueled and non-fueled regions are modeled. Each fuel assembly consists of one or more plates or tubes separated by coolant channels. The fuel plates may have one to five layers of different materials, each with heat generation. The width of a fuel plate may be divided into multiple longitudinal stripes, each with its own axial power shape. Depending upon the heat transfer method selected, the temperature solution is effectively 2-dimensional or 3-dimensional. The geometry may be either slab or radial, corresponding to fuel assemblies made of a series of flat (or slightly curved) plates, or of nested tubes. A variety of thermal-hydraulic correlations are available to determine safety margins such as onset of nucleate boiling ratio (ONBR), departure from nucleate boiling ratio (DNBR), and onset of flow instability ratio (OFIR). Coolant properties for either light or heavy water are obtained from FORTRAN functions rather than from tables. The code is intended for thermal-hydraulic analysis of research reactor performance in the sub-cooled boiling regime. Both turbulent and laminar flow regimes can be modeled. Options to calculate both forced flow and natural circulation are available. A general search capability for select design and safety parameters is available to greatly reduce the reactor analyst’s time.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗