Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evolutionary computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗

The landscape of regulatory element evolution in a C4 perennial grass

Gene regulatory evolution is a well-known source of phenotypic diversity and adaptive evolution. Although cis-regulatory elements (CREs) play a vital role in gene expression evolution, the molecular evolution of CREs remains mostly unknown due to the difficulty in identifying and characterizing these functional elements. Comparative genomic analyses of noncoding DNA can be leveraged to identify conserved noncoding sequences (CNS), many of which may harbor functional CREs conserved by purifying selection. However, purely computational inference of CREs from putative CNS can be erroneous due to the complex genomic architecture in plants. One promising experimental approach to identify CREs is by profiling accessible chromatin regions (ACRs) that are often associated with the location of CREs. In this study, we use comparative genomics along with the profiling of ACRs to study the molecular evolution of putative functional noncoding regulatory regions in Panicoid grasses. We identified sets of CNS that varied in relationship to the degree of evolutionary divergence among the studied taxa, including identifying core-Panicoid-CNS. We augmented this analysis by profiling ACRs in Panicum hallii ecotypes using ATAC-seq. ACRs had low SNP density at the summit, harbored a high frequency of core-Panicoid-CNS, and were enriched with expression QTL. These data help to annotate the P. hallii genome for putative functional elements and suggest that a large proportion of these ACRs are evolving under purifying selection. Turnover in CNS and ACR between ecotypes of P. hallii identifies a small set of putatively divergent CREs that may underlie differences in gene regulation between genotypes from inland and coastal habitats. In summary, we profiled ACRs in Panicoid grasses and integrated this data with our putative CNS prediction framework, which provides unique insight into patterns of polymorphism and divergence in CREs in C4 perennial grasses.

59 BASIC BIOLOGICAL SCIENCES↗

MAL33 drives natural variation in maltose metabolism in Saccharomyces eubayanus

Maltose is one of the most abundant sugars in brewer’s wort, and its efficient utilization is critical for successful fermentation. However, maltose consumption varies naturally among Saccharomyces eubayanus strains isolated from different host trees, such as Quercus and Nothofagus. To identify the genetic determinants underlying these phenotypic differences, we performed bulk segregant analysis (BSA) and quantitative trait loci (QTL) mapping using an F 2 offspring derived from QC18 (Quercus-associated) and CL467.1 (Nothofagus-associated) strains. QTL mapping identified two significant genomic regions on subtelomeric loci of chromosomes V-R and XVI-L, each containing complete MAL loci composed of MAL32 (encoding maltase), MAL31 (transporter), and MAL33 (transcriptional activator) genes. Comparative polymorphism analyses identified mutations in MAL32 and MAL33 of QC18, including frameshift mutations resulting in premature stop codons. Functional validation demonstrated that the heterologous expression of MAL33 ChrV from CL467.1 fully restored maltose utilization in QC18, indicating the functional presence of MAL33 cis-regulatory sequences and MAL32 and MAL31 genes in QC18. While structural protein predictions identified truncation and impaired functionality in the maltose-responsive activation domain of Mal33p from QC18, overexpression of QC18’s own MAL33 ChrV allele also improved maltose metabolism, suggesting dosage-dependent transcriptional limitations rather than complete functional loss. These results indicate that allelic variations in the maltose-responsive activation domain of Mal33p result in differences in maltose consumption between strains. Here, we hypothesized that reduced maltose metabolism in QC18 is an adaptive response to the distinct sugar composition in Quercus robur bark, contrasting with the starch-rich environment of Nothofagus pumilio. These findings highlight subtelomeric MAL gene diversity as a reservoir of genetic variation, representing a key evolutionary mechanism that influences maltose adaptation among natural Saccharomyces isolates.

evolutionary plasticity↗

Top-down design of protein architectures with reinforcement learning

As a result of evolutionary selection, the subunits of naturally occurring protein assemblies often fit together with substantial shape complementarity to generate architectures optimal for function in a manner not achievable by current design approaches. We describe a “top-down” reinforcement learning–based design approach that solves this problem using Monte Carlo tree search to sample protein conformers in the context of an overall architecture and specified functional constraints. Cryo–electron microscopy structures of the designed disk-shaped nanopores and ultracompact icosahedra are very close to the computational models. The icosohedra enable very-high-density display of immunogens and signaling molecules, which potentiates vaccine response and angiogenesis induction. Our approach enables the top-down design of complex protein nanomaterials with desired system properties and demonstrates the power of reinforcement learning in protein design.

Science & Technology - Other Topics↗

Chemical Complexity of Phosphorous-bearing Species in Various Regions of the Interstellar Medium

Phosphorus-related species are not known to be as omnipresent in space as hydrogen, carbon, nitrogen, oxygen, and sulfur-bearing species. Astronomers spotted very few P-bearing molecules in the interstellar medium and circumstellar envelopes. Limited discovery of the P-bearing species imposes severe constraints in modeling the P-chemistry. In this paper, we carry out extensive chemical models to follow the fate of P-bearing species in diffuse clouds, photon-dominated or photodissociation regions (PDRs), and hot cores/corinos. We notice a curious correlation between the abundances of PO and PN and atomic nitrogen. Since N atoms are more abundant in diffuse clouds and PDRs than in the hot core/corino region, PO/PN reflects <1 in diffuse clouds, ≪1 in PDRs, and >1 in the late warm-up evolutionary stage of the hot core/corino regions. During the end of the post-warm-up stage, we obtain PO/PN > 1 for hot core and <1 for its low-mass analog. We employ a radiative transfer model to investigate the transitions of some of the P-bearing species in diffuse cloud and hot core regions and estimate the line profiles. Our study estimates the required integration time to observe these transitions with ground-based and space-based telescopes. We also carry out quantum chemical computation of the infrared features of PH{sub 3}, along with various impurities. We notice that SO{sub 2} overlaps with the PH{sub 3} bending-scissoring modes around ∼1000–1100 cm{sup −1}. We also find that the presence of CO{sub 2} can strongly influence the intensity of the stretching modes around ∼2400 cm{sup −1} of PH{sub 3}.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Global Archaeal Diversity Revealed Through Massive Data Integration: Uncovering Just Tip of Iceberg

The domain of Archaea has gathered significant interest for its ecological and biotechnological potential and its role in helping us to understand the evolutionary history of Eukaryotes. In comparison to the bacterial domain, the number of adequately described members in Archaea is relatively low, with less than 1000 species described. It is not clear whether this is solely due to the cultivation difficulty of its members or, indeed, the domain is characterized by evolutionary constraints that keep the number of species relatively low. Based on molecular evidence that bypasses the difficulties of formal cultivation and characterization, several novel clades have been proposed, enabling insights into their metabolism and physiology. Given the extent of global sampling and sequencing efforts, it is now possible and meaningful to question the magnitude of global archaeal diversity based on molecular evidence. To do so, we extracted all sequences classified as Archaea from 500 thousand amplicon samples available in public repositories. After processing through our highly conservative pipeline, we named this comprehensive resource the ‘Global Archaea Diversity’ (GAD), which encompassed nearly 3 million molecular species clusters at 97% similarity, and organized it into over 500 thousand genera and nearly 100 thousand families. Saline environments have contributed the most to the novel taxa of this previously unseen diversity. The majority of those 16S rRNA gene sequence fragments were verified by matches in metagenomic datasets from IMG/M. These findings reveal a vast and previously overlooked diversity within the Archaea, offering insights into their ecological roles and evolutionary importance while establishing a foundation for the future study and characterization of this intriguing domain of life.

59 BASIC BIOLOGICAL SCIENCES↗

Divergent selection and climate adaptation fuel genomic differentiation between sister species of Sphagnum (peat moss)

Abstract Background and Aims New plant species can evolve through the reinforcement of reproductive isolation via local adaptation along habitat gradients. Peat mosses (Sphagnaceae) are an emerging model system for the study of evolutionary genomics and have well-documented niche differentiation among species. Recent molecular studies have demonstrated that the globally distributed species Sphagnum magellanicum is a complex of morphologically cryptic lineages that are phylogenetically and ecologically distinct. Here, we describe the architecture of genomic differentiation between two sister species in this complex known from eastern North America: the northern S. diabolicum and the largely southern S. magniae. Methods We sampled plant populations from across a latitudinal gradient in eastern North America and performed whole genome and restriction-site associated DNA sequencing. These sequencing data were then analyzed computationally. Key Results Using sliding-window population genetic analyses we find that differentiation is concentrated within ‘islands’ of the genome spanning up to 400 kb that are characterized by elevated genetic divergence, suppressed recombination, reduced nucleotide diversity and increased rates of non-synonymous substitution. Sequence variants that are significantly associated with genetic structure and bioclimatic variables occur within genes that have functional enrichment for biological processes including abiotic stress response, photoperiodism and hormone-mediated signalling. Demographic modelling demonstrates that these two species diverged no more than 225 000 generations ago with secondary contact occurring where their ranges overlap. Conclusions We suggest that this heterogeneity of genomic differentiation is a result of linked selection and reflects the role of local adaptation to contrasting climatic zones in driving speciation. This research provides insight into the process of speciation in a group of ecologically important plants and strengthens our predictive understanding of how plant populations will respond as Earth’s climate rapidly changes.

58 GEOSCIENCES↗

GENESPACE R Package (GENESPACE) v1.0

In short, the GENESPACE pipeline conducts analysis of orthology networks, constrained within syntenic regions. Since analyses are limited to local tests conducted within syntenic blocks, GENESPACE is agnostic to ploidy, duplicated regions, inversions or other whole-genome chromosomal complexities that are common across many evolutionary lineages. This advantage allows for evolutionary tests in polyploids (e.g. switchgrass, manuscript in review), species with ancient, but retained whole-genome duplications (e.g. pecan, manuscript in prep), high levels of tandem array proliferation (e.g. eukalypts, manuscript in review) and many other factors that can confound comparative genomic analyses. The major advances of GENESPACE are three-fold: First, this is the first R package to integrate visualization and analysis of large-scale comparative genomics. R, which offers a high-level environment for graphical and statistical exploration of data, is often speed- and memory-limited and not used for computationally intensive tasks such as comparative genomics. The highly efficient C++ scripts used in GENESPACE (via data.table) permit a much faster and computationally lightweight implementation of comparative genomics than is currently available. Second, the pipeline itself is novel. To the best of our knowledge, no other program accomplishes synteny-constrained and ploidy-agnostic comparative genomics. Since nearly all plants and many animals have a history of whole-genome duplications, this is a major and necessary advance to the field. Third, GENESPACE offers high-level and intuitive multi-genome graphical outputs. The dotplots and 'riparian' plots produced herein, which are produced entirely through original R code, are publication-ready and easily customizable.

Schmutz, Jeremy↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Evolutionary reinforcement learning of dynamical large deviations

In this work, we show how to bound and calculate the likelihood of dynamical large deviations using evolutionary reinforcement learning. An agent, a stochastic model, propagates a continuous-time Monte Carlo trajectory and receives a reward conditioned upon the values of certain path-extensive quantities. Evolution produces progressively fitter agents, potentially allowing the calculation of a piece of a large-deviation rate function for a particular model and path-extensive quantity. For models with small state spaces, the evolutionary process acts directly on rates, and for models with large state spaces, the process acts on the weights of a neural network that parameterizes the model's rates. This approach shows how path-extensive physics problems can be considered within a framework widely used in machine learning.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evolutionary multi-objective optimization and Pareto-frontal uncertainty quantification of interatomic forcefields for thermal conductivity simulations

Predictive Molecular Dynamics simulations of thermal transport require forcefields that can simultaneously reproduce several structural, thermodynamic and vibrational properties of materials like lattice constants, phonon density of states, and specific heat. This requires a multi-objective optimization approach for forcefield parameterization. Existing methodologies for forcefield parameterization use ad-hoc and empirical weighting schemes to convert this into a single-objective optimization problem. Here, we provide and describe software to perform multi-objective optimization of Stillinger–Weber forcefields (SWFF) for two-dimensional layered materials using the recently developed 3rd generation non-dominated sorting genetic algorithm (NSGA-III). NSGA-III converges to the set of optimal forcefields lying on the Pareto front in the multi-dimensional objective space. This set of forcefields is used for uncertainty quantification of computed thermal conductivity due to variability in the forcefield parameters. We demonstrate this new optimization scheme by constructing a SWFF for a representative two-dimensional material, 2H-MoSe 2 and quantifying the uncertainty in their computed thermal conductivity.

97 MATHEMATICS AND COMPUTING↗

The histone code of the fungal genus Aspergillus uncovered by evolutionary and proteomic analyses

Chemical modifications of DNA and histone proteins impact the organization of chromatin within the nucleus. Changes in these modifications, catalysed by different chromatin-modifying enzymes, influence chromatin organization, which in turn is thought to impact the spatial and temporal regulation of gene expression. While combinations of different histone modifications, the histone code, have been studied in several model species, we know very little about histone modifications in the fungal genus Aspergillus, whose members are generally well studied due to their importance as models in cell and molecular biology as well as their medical and biotechnological relevance. Here, we used phylogenetic analyses in 94 Aspergilli as well as other fungi to uncover the occurrence and evolutionary trajectories of enzymes and protein complexes with roles in chromatin modifications or regulation. We found that these enzymes and complexes are highly conserved in Aspergilli, pointing towards a complex repertoire of chromatin modifications. Nevertheless, we also observed few recent gene duplications or losses, highlighting Aspergillus species to further study the roles of specific chromatin modifications. SET7 (KMT6) and other components of PRC2 (Polycomb Repressive Complex 2), which is responsible for methylation on histone H3 at lysine 27 in many eukaryotes including fungi, are absent in Aspergilli as well as in closely related Penicillium species, suggesting that these lost the capacity for this histone modification. We corroborated our computational predictions by performing untargeted MS analysis of histone post-translational modifications in Aspergillus nidulans. This systematic analysis will pave the way for future research into the complexity of the histone code and its functional implications on genome architecture and gene regulation in fungi.

59 BASIC BIOLOGICAL SCIENCES↗

Nearby stellar substructures in the Galactic halo from DESI Milky Way Survey Year 1 Data Release

We report five nearby ($d_{\mathrm{helio}} < 5$ kpc) stellar substructures in the Galactic halo from a subset of 138 661 stars in the Dark Energy Spectroscopic Instrument (DESI) Milky Way Survey Year 1 Data Release. With an unsupervised clustering algorithm, HDBSCAN*, these substructures are independently identified in Integrals of Motion ($E_{\rm tot}$, $L_{\rm z}$, $\log {J_r}$, $\log {J_z}$) space and Galactocentric cylindrical velocity space ($V_{R}$, $V_{\phi }$, $V_{z}$). We associate all identified clusters with known nearby substructures (Helmi streams, M18-Cand10/MMH-1, Sequoia, Antaeus, and ED-2) previously reported in various studies. With metallicities precisely measured by DESI, we confirm that the Helmi streams, M18-Cand10, and ED-2 are chemically distinct from local halo stars. We have characterized the chemodynamic properties of each dynamic group, including their metallicity dispersions, to associate them with their progenitor types (globular cluster or dwarf galaxy). Our approach for searching substructures with HDBSCAN* reliably detects real substructures in the Galactic halo, suggesting that applying the same method can lead to the discovery of new substructures in future DESI data. With more stars from future DESI data releases and improved astrometry from the upcoming Gaia Data Release 4, we will have a more detailed blueprint of the Galactic halo, offering a significant improvement in our understanding of the formation and evolutionary history of the Milky Way Galaxy.

dynamics↗

Multidimensional low-Mach number time-implicit hydrodynamic simulations of convective helium shell burning in a massive star

A realistic parametrization of convection and convective boundary mixing in conventional stellar evolution codes is still the subject of ongoing research. Furthermore, to improve the current situation, multidimensional hydrodynamic simulations are used to study convection in stellar interiors. Such simulations are numerically challenging, especially for flows at low Mach numbers which are typical for convection during early evolutionary stages. We explore the benefits of using a low-Mach hydrodynamic flux solver and demonstrate its usability for simulations in the astrophysical context. Simulations of convection for a realistic stellar profile are analyzed regarding the properties of convective boundary mixing. The time-implicit Seven-League Hydro (SLH) code was used to perform multidimensional simulations of convective helium shell burning based on a 25 M ⊙ star model. The results obtained with the low-Mach AUSM + -up solver were compared to results when using its non low-Mach variant AUSM B + -up. We applied well-balancing of the gravitational source term to maintain the initial hydrostatic background stratification. The computational grids have resolutions ranging from 180 × 90 2 to 810 × 540 2 cells and the nuclear energy release was boosted by factors of 3 × 10 3 , 1 × 10 4 , and 3 × 10 4 to study the dependence of the results on these parameters.

79 ASTRONOMY AND ASTROPHYSICS↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Quartet-based inference is statistically consistent under the unified duplication-loss-coalescence model

Abstract Motivation The classic multispecies coalescent (MSC) model provides the means for theoretical justification of incomplete lineage sorting-aware species tree inference methods. This has motivated an extensive body of work on phylogenetic methods that are statistically consistent under MSC. One such particularly popular method is ASTRAL, a quartet-based species tree inference method. Novel studies suggest that ASTRAL also performs well when given multi-locus gene trees in simulation studies. Further, Legried et al. recently demonstrated that ASTRAL is statistically consistent under the gene duplication and loss model (GDL). GDL is prevalent in evolutionary histories and is the first core process in the powerful duplication-loss-coalescence evolutionary model (DLCoal) by Rasmussen and Kellis. Results In this work, we prove that ASTRAL is statistically consistent under the general DLCoal model. Therefore, our result supports the empirical evidence from the simulation-based studies. More broadly, we prove that the quartet-based inference approach is statistically consistent under DLCoal. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Neuro-Spark: A Submicrosecond Spiking Neural Networks Architecture for In-Sensor Filtering

Neuro-Spark, which is a new neuromorphic architecture with a field-programmable gate array (FPGA) implementation for ultrafast spiking neural network (SNN) inference at the edge, facilitates smart-pixel in-sensor filtering for high-energy physics experiments at the Large Hadron Collider (LHC). Utilizing the evolutionary optimization for neuromorphic systems (EONS) training method, we generate compact SNN models with 91% signal efficiency, akin to convolutional neural networks but with half the parameters. However, deploying near the detector poses a challenge because the SNN must handle a sustained input data rate exceeding 1013 GB/s. To overcome this, we propose a novel hardware architecture that uses high-level synthesis to construct a tuned architecture for the EONS-trained SNN. In addition to the analysis and validation with an AMD Xilinx Artix-A7 FPGA, our solution consumes only ç24% of FPGA LUT and flipflops. We also introduce an innovative quantization method that reduces FPGA resource utilization by ç15% without compromising accuracy. Our FPGA implementation achieves computing latency of ç10 ns for smart-pixel application inference on an edge FPGA.

Miniskar, Narasinga Rao↗

Standardized phylogenetic and molecular evolutionary analysis applied to species across the microbial tree of life

There is growing interest in reconstructing phylogenies from the copious amounts of genome sequencing projects that target related viral, bacterial or eukaryotic organisms. To facilitate the construction of standardized and robust phylogenies for disparate types of projects, we have developed a complete bioinformatic workflow, with a web-based component to perform phylogenetic and molecular evolutionary (PhaME) analysis from sequencing reads, draft assemblies or completed genomes of closely related organisms. Furthermore, the ability to incorporate raw data, including some metagenomic samples containing a target organism (e.g. from clinical samples with suspected infectious agents), shows promise for the rapid phylogenetic characterization of organisms within complex samples without the need for prior assembly.

59 BASIC BIOLOGICAL SCIENCES↗