Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evolutionary computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Robust Carbon Dioxide Plume Imaging Using Joint Tomographic Inversion of Seismic Onset Time and Distributed Pressure and Temperature Measurements (Final Report)

We develop and demonstrate rapid and cost-effective methodologies for spatiotemporal tracking of CO2 plumes during geologic sequestration using joint inversion of seismic data and distributed pressure and temperature measurements. Key elements of our methodology are: (a) a computationally efficient approach to pressure and temperature propagation, (b) analysis of time lapse seismic data using a novel ‘seismic onset time’ approach to detect fluid front propagation, and (c) data assimilation and uncertainty assessment via joint inversion of pressure, temperature and time lapse seismic data, and (d) validating the numerical tomographic inversion using a CO2 injection demonstration projects, specifically data collected from the from the Petra Nova Parish Holdings CCUS project in the West Ranch Field, Texas and the Chester-16 reef CO2 injection site in Northern Michigan which is part of the DOE Midwestern Carbon Sequestration Project. The research team is led by Texas A&M University and includes Battelle as a subcontractor with support from Shell, Anadarko, Chevron and JX Nippon. A carbon dioxide (CO2) water-alternating-gas (WAG) pilot was conducted to gain insights into tertiary oil recovery potential via CO2 flood in the West Ranch Field as part of the Petra Nova project, the world’s largest post-combustion CO2 capture and utilization initiative. With a fluvial formation geology and large contrasts in permeability, this is a challenging and novel application of CO2 enhanced oil recovery (EOR). We build a predictive dynamic model of the subsurface that incorporates the multiphase and compositional data acquired during the pilot operation. The calibrated model is used for the carbon dioxide plume imaging. The study began with an initialization of the pilot sector model extracted from a calibrated full-field model. The pilot model calibration follows a two-step hierarchical workflow. First, we performed a large-scale update of the permeability distribution by integrating available bottomhole pressure and multiphase production data. In the second step, local permeability field is fine-tuned using a streamline-based method to match CO2 breakthrough times at the producers. The predictive capability of the calibrated model was verified through two blind validation tests: (1) the model showed good agreement with saturation logs acquired at two observation wells; and (2) the model reproduced the CO2 recovery as a fraction of the injected CO2. The use of seismic onset times has shown great promise for integrating near-continuous seismic surveys for updating geologic models. In this study, we analyze the impact of seismic survey frequency on the onset time approach aiming to extend the application of onset time to infrequent seismic surveys. In addition, we quantitatively examine the nonlinearity of the onset time method and compare it to the commonly used amplitude inversion method. We carry out a sensitivity analysis of seismic survey frequency based on the complete seismic survey data (over 175 surveys) of steam injection in a heavy oil reservoir (Peace River Unit) in Canada. Our results show that an adequate onset time map can be obtained from the infrequent seismic surveys by interpolation between seismic surveys as long as there is no change in the dominant underlying physics between the successive surveys. The study also shows that nonlinearity of the onset time method can be -smaller than that of the amplitude inversion method by several orders of magnitude. Application to the Brugge benchmark case shows that the onset time method obtains comparable permeability update as the traditional seismic amplitude inversion method with faster computation and improved convergence characteristics. We extend the streamline-based data integration approach to incorporate distributed temperature sensor (DTS) data using the concept of thermal tracer travel time. Then, a hierarchical workflow composed of evolutionary and streamline methods is employed to jointly history match the DTS and pressure data. Finally, CO2 saturation and streamline maps are used to visualize the CO2 plume movement during the sequestration process. The hierarchical workflow is applied to a carbon sequestration project in a carbonate reef reservoir within the Northern Niagaran Pinnacle Reef Trend in Michigan, USA. The monitoring data set consists of distributed temperature sensing (DTS) data acquired at the injection well and a monitoring well, flowing bottom-hole pressure data at the injection well, and time-lapse pressure measurements at several locations along the monitoring well. The history matching results indicate that the CO2 movement is mostly restricted to the intended zones of injection which is consistent with an independent warm-back analysis of the temperature data. In addition to employing simulation models and inverse methods for CO2 plume imaging, we also initialized a data-driven technology for detecting inter-well connectivity based on production and pressure data. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO2 EOR projects utilizing the water-alternating-gas (WAG) process. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. Texas A&M University, the lead organization in the project, was primarily responsible for the development of tomographic approaches for CO2 plume mapping in conjunction with distributed pressure, temperature and seismic onset time data. Battelle, as a subcontractor, was primarily responsible for the development of analytical and empirical methods for analyzing transient injection rate and pressure data from point/line sources such as injection and monitoring wells. An additional area of emphasis for Battelle was the use of machine learning for such tasks as inferring reservoir connectivity information from injection-production data, and identifying variable importance for machine learning-based proxy models developed from full-physics simulations. The two organizations also collaborated on the application of the tomographic inversion methodology for a field data set.

02 PETROLEUM↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Deep Reinforcement Learning Based Control of Wind Turbines for Fast Frequency Response

In order to fulfill vital auxiliary grid services, such as load regulation, spin and non-spin reserve provision, and frequency support during emergencies, there is often a requirement for certain wind farms to operate in de-loaded modes. Leveraging the swift response capabilities of wind farms, this study demonstrates that reserving power in de-loaded modes can significantly enhance power grid stability and reliability during system contingencies. Controlling wind farms optimally for frequency support is intricate due to the nonlinearity of models and controllers and the complexity of wind farm interactions with power systems. Here, to address this challenge, this paper introduces a novel approach that integrates wind turbines into reinforcement learning-based solutions for frequency response. This innovative methodology utilizes the state-of-the-art reinforcement learning algorithm known as the surrogate-gradient-based evolutionary strategy. The proposed learning-based algorithm provides continuous control of wind farm output to rapidly stabilize system frequency and prevent unnecessary trips of under-frequency load shedding relays. To facilitate efficient training, parallel computing techniques are employed. The proposed methodology is evaluated on a modified IEEE-39 bus system, and simulation results reveal its efficacy in reliably supporting power system frequency and preventing the need for unnecessary load shedding.

Gao, Wei [Argonne National Laboratory (ANL), Argon↗

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis↗

Betelgeuse as a Merger of a Massive Star with a Companion

We investigate the merger between a 16M ⊙ star, on its way to becoming a red supergiant (RSG), and a 4M ⊙ main-sequence companion. Our study employs three-dimensional hydrodynamic simulations using the state-of-the-art adaptive mesh refinement code Octo-Tiger. The initially corotating binary undergoes interaction and mass transfer, resulting in the accumulation of mass around the companion and its subsequent loss through the second Lagrangian point (L2). The companion eventually plunges into the envelope of the primary, leading to its spin-up and subsequent merger with the helium core. We examine the internal structural properties of the post-merger star, as well as the merger environment and the outflow driven by the merger. Our findings reveal the ejection of approximately ∼0.6 M ⊙ of material in an asymmetric and somewhat bipolar outflow. We import the post-merger stellar structure into the MESA stellar evolution code to model its long-term nuclear evolution. In certain cases, the post-merger star exhibits persistent rapid equatorial surface rotation as it evolves in the H–R diagram toward the observed location of Betelgeuse. These cases demonstrate surface rotation velocities of a similar magnitude to those observed in Betelgeuse, along with a chemical composition resembling that of Betelgeuse. In other cases, efficient rotationally induced mixing leads to slower surface rotation. This pioneering study aims to model stellar mergers across critical timescales, encompassing dynamical, thermal, and nuclear evolutionary stages.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimizing Grain Boundary Structures with LAMMPS Using Evolutionary Algorithms

Grain boundary structure optimization is an important part of materials modeling. Current methods for grain boundary structure optimization involve inefficient, time-consuming processes that do not fully explore the interface parameter space. Evolutionary algorithms have recently been demonstrated to be effective at determining both stable and metastable grain boundary interface structures. In this work, we demonstrate the use of GBOpt, a grain boundary structure optimization software designed to use the Large-scale Atomic/Molecular Massively Parallel Simulation (LAMMPS) software to efficiently determine grain boundary structures. We demonstrate that a only a few manipulations, namely atom insertion, atom removal, and relative grain displacement, are sufficient to explore much of the grain boundary structure parameter space. The efficacy of this approach is demonstrated on an FCC Ni system, and a BCC Fe system. The computational cost is compared against the gamma-surface sampling approach to demonstrate performance improvement.

Evolutionary algorithms↗

P finder: genomic and metagenomic annotation of RNase P RNA gene (rnpB)

Abstract Background The rnpB gene encodes for an essential catalytic RNA (RNase P). Like other essential RNAs, RNase P’s sequence is highly variable. However, unlike other essential RNAs (i.e. tRNA, 16 S, 6 S,...) its structure is also variable with at least 5 distinct structure types observed in prokaryotes. This structural variability makes it labor intensive and challenging to create and maintain covariance models for the detection of RNase P RNA in genomic and metagenomic sequences. The lack of a facile and rapid annotation algorithm has led to the rnpB gene being the most grossly under annotated essential gene in completed prokaryotic genomes with only a 24% annotation rate. Here we describe the coupling of the largest RNase P RNA database with the local alignment scoring algorithm to create the most sensitive and rapid prokaryote rnpB gene identification and annotation algorithm to date. Results Of the 2772 completed microbial genomes downloaded from GenBank only 665 genomes had an annotated rnpB gene. We applied P Finder to these genomes and were able to identify 2733 or nearly 99% of the 2772 microbial genomes examined. From these results four new rnpB genes that encode the minimal T-type P RNase P RNAs were identified computationally for the first time. In addition, only the second C-type RNase P RNA was identified in Sphaerobacter thermophilus . Of special note, no RNase P RNAs were detected in several obligate endosymbionts of sap sucking insects suggesting a novel evolutionary adaptation. Conclusions The coupling of the largest RNase P RNA database and associated structure class identification with the P Finder algorithm is both sensitive and rapid, yielding high quality results to aid researchers annotating either genomic or metagenomic data. It is the only algorithm to date that can identify challenging RNAse P classes such as C-type and the minimal T-type RNase P RNAs. P Finder is written in C# and has a user-friendly GUI that can run on multiple 64-bit windows platforms (Windows Vista/7/8/10). P Finder is free available for download at https://github.com/JChristopherEllis/P-Finder as well as a small sample RNase P RNA file for testing.

59 BASIC BIOLOGICAL SCIENCES↗

HUNTRESS: a fast heuristic for reconstructing phylogenetic trees of tumor evolution (HUNTRESS) v0.1

We introduce HUNTRESS (Histogrammed UNion Tree REconStruction heuriStic), a computational method for tumor phylogeny reconstruction from noisy genotype matrices derived from single-cell sequencing data, whose running time is linear with the number of cells and quadratic with the number of mutations. Provided that the input genotype matrix includes no false positives, each cellular subpopulation is at least a user defined fraction of the total number of cells, and the number of cells are much bigger than the number of mutations considered, HUNTRESS computes the ground truth tumor phylogeny with high probability. On simulated data HUNTRESS is faster than available alternatives with comparable or better accuracy. Additionally, the phylogenies reconstructed by HUNTRESS on two single-cell sequencing data sets agree with the best known evolutionary scenarios for the associated tumors.

Buluc, Aydin↗

A metabolic modeling platform for the computation of microbial ecosystems in time and space (COMETS)

Genome-scale stoichiometric modeling of metabolism has become a standard systems biology tool for modeling cellular physiology and growth. Extensions of this approach are emerging as a valuable avenue for predicting, understanding and designing microbial communities. Computation of microbial ecosystems in time and space (COMETS) extends dynamic flux balance analysis to generate simulations of multiple microbial species in molecularly complex and spatially structured environments. Here we describe how to best use and apply the most recent version of COMETS, which incorporates a more accurate biophysical model of microbial biomass expansion upon growth, evolutionary dynamics and extracellular enzyme activity modules. In addition to a command-line option, COMETS includes user-friendly Python and MATLAB interfaces compatible with the well-established COBRA models and methods, as well as comprehensive documentation and tutorials. Overall, this protocol provides a detailed guideline for installing, testing and applying COMETS to different scenarios, generating simulations that take from a few minutes to several days to run, with broad applicability to microbial communities across biomes and scales.

59 BASIC BIOLOGICAL SCIENCES↗

Benchmarking the Performance of Neuromorphic and Spiking Neural Network Simulators

Software simulators play a critical role in the development of new algorithms and system architectures in any field of engineering. Neuromorphic computing, which has shown potential in building brain-inspired energy-efficient hardware, suffers a slow-down in the development cycle due to a lack of flexible and easy-to-use simulators of either neuromorphic hardware itself or of spiking neural networks (SNNs), the type of neural network computation executed on most neuromorphic systems. While there are several openly available neuromorphic or SNN simulation packages developed by a variety of research groups, they have mostly targeted computational neuroscience simulations, and only a few have targeted small-scale machine learning tasks with SNNs. Evaluations or comparisons of these simulators have often targeted computational neuroscience-style workloads. In this work, we seek to evaluate the performance of several publicly available SNN simulators with respect to non-computational neuroscience workloads, in terms of speed, flexibility, and scalability. We evaluate the performance of the NEST, Brian2, Brian2GeNN, BindsNET and Nengo packages under a common front-end neuromorphic framework. Our evaluation tasks include a variety of different network architectures and workload types to mimic the computation common in different algorithms, including feed-forward network inference, genetic algorithms, and reservoir computing. We also study the scalability of each of these simulators when running on different computing hardware, from single core CPU workstations to multi-node supercomputers. Our results show that the BindsNET simulator has the best speed and scalability for most of the SNN workloads (sparse, dense, and layered SNN architectures) on a single core CPU. However, when comparing the simulators leveraging the GPU capabilities, Brian2GeNN outperforms the others for these workloads in terms of scalability. NEST performs the best for small sparse networks and is also the most flexible simulator in terms of reconfiguration capability NEST shows a speedup of at least 2x compared to the other packages when running evolutionary algorithms for SNNs. The multi-node and multi-thread capabilities of NEST show at least 2x speedup compared to the rest of the simulators (single core CPU or GPU based simulators) for large and sparse networks. We conclude our work by providing a set of recommendations on the suitability of employing these simulators for different tasks and scales of operations. We also present the characteristics for a future generic ideal SNN simulator for different neuromorphic computing workloads.

97 MATHEMATICS AND COMPUTING↗

An evolutionary algorithm for designing microbial communities via environmental modification

Despite a growing understanding of how environmental composition affects microbial communities, it remains difficult to apply this knowledge to the rational design of synthetic multispecies consortia. This is because natural microbial communities can harbour thousands of different organisms and environmental substrates, making up a vast combinatorial space that precludes exhaustive experimental testing and computational prediction. Here, we present a method based on the combination of machine learning and metabolic modelling that selects optimal environmental compositions to produce target community phenotypes. In this framework, dynamic flux balance analysis is used to model the growth of a community in candidate environments. A genetic algorithm is then used to evaluate the behaviour of the community relative to a target phenotype, and subsequently adjust the environment to allow the organisms to approach this target. We apply this iterative process to thousands of in silico communities of varying sizes, showing how it can rapidly identify environments that yield desired taxonomic compositions and patterns of metabolic exchange. Moreover, this combination of approaches produces testable predictions for the assembly of experimental microbial communities with specific properties and can facilitate rational environmental design processes for complex microbiomes.

59 BASIC BIOLOGICAL SCIENCES↗

Keck/NIRC2 L’-band Imaging of Jovian-mass Accreting Protoplanets around PDS 70

We present L’-band imaging of the PDS 70 planetary system with Keck/NIRC2 using the new infrared pyramid wave front sensor. We detected both PDS 70 b and c in our images, as well as the front rim of the circumstellar disk. After subtracting off a model of the disk, we measured the astrometry and photometry of both planets. Placing priors based on the dynamics of the system, we estimated PDS 70 b to have a semimajor axis of 20{sub −4}{sup +3} au and PDS 70 c to have a semimajor axis of 34{sub −6}{sup +12} au (95% credible interval). We fit the spectral energy distribution (SED) of both planets. For PDS 70 b, we were able to place better constraints on the red half of its SED than previous studies and inferred the radius of the photosphere to be 2–3 R {sub Jup}. The SED of PDS 70 c is less well constrained, with a range of total luminosities spanning an order of magnitude. With our inferred radii and luminosities, we used evolutionary models of accreting protoplanets to derive a mass of PDS 70 b between 2 and 4 M {sub Jup} and a mean mass accretion rate between 3 × 10{sup −7} and 8 × 10{sup −7} M {sub Jup}/yr. For PDS 70 c, we computed a mass between 1 and 3 M {sub Jup} and mean mass accretion rate between 1 × 10{sup −7} and 5 × 10{sup −7} M {sub Jup}/yr. The mass accretion rates imply dust accretion timescales short enough to hide strong molecular absorption features in both planets’ SEDs.

79 ASTRONOMY AND ASTROPHYSICS↗

Three-dimensional structure-guided evolution of a ribosome with tethered subunits

We report RNA-based macromolecular machines, such as the ribosome, have functional parts reliant on structural interactions spanning sequence-distant regions. These features limit evolutionary exploration of mutant libraries and confound three-dimensional structure-guided design. To address these challenges, we describe Evolink (evolution and linkage), a method that enables high-throughput evolution of sequence-distant regions in large macromolecular machines, and library design guided by computational RNA modeling to enable exploration of structurally stable designs. Using Evolink, we evolved a tethered ribosome with a 58% increased activity in orthogonal protein translation and a 97% improvement in doubling times in SQ171 cells compared to a previously developed tethered ribosome, and reveal new permissible sequences in a pair of ribosomal helices with previously explored biological function. The Evolink approach may enable enhanced engineering of macromolecular machines for new and improved functions for synthetic biology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Tad and toxin-coregulated pilus structures reveal unexpected diversity in bacterial type IV pili

<Type IV pili (T4P) are ubiquitous in both bacteria and archaea. They are polymers of the major pilin protein, which has an extended and protruding N-terminal helix, α1, and a globular C-terminal domain. Cryo-EM structures have revealed key differences between the bacterial and archaeal T4P in their C-terminal domain structure and in the packing and continuity of α1. This segment forms a continuous α-helix in archaeal T4P but is partially melted in all published bacterial T4P structures due to a conserved helix breaking proline at position 22. The tad (tight adhesion) T4P are found in both bacteria and archaea and are thought to have been acquired by bacteria through horizontal transfer from archaea. Tad pilins are unique among the T4 pilins, being only 40 to 60 residues in length and entirely lacking a C-terminal domain. They also lack the Pro22 found in all high-resolution bacterial T4P structures. We show using cryo-EM that the bacterial tad pilus from Caulobacter crescentus is composed of continuous helical subunits that, like the archaeal pilins, lack the melted portion seen in other bacterial T4P and share the packing arrangement of the archaeal T4P. We further show that a bacterial T4P, the Vibrio cholerae toxin coregulated pilus, which lacks Pro22 but is not in the tad family, has a continuous N-terminal α-helix, yet its α1 s are arranged similar to those in other bacterial T4P. Our results highlight the role of Pro22 in helix melting and support an evolutionary relationship between tad and archaeal T4P.

59 BASIC BIOLOGICAL SCIENCES↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning enabled discovery of superhard and ultrahard carbon polymorphs

The demand for multifunctional materials has motivated the move from near-equilibrium materials to metastable i.e. out-of-equilibrium phases that can meet several desired target properties. The search for such metastable phases with exotic properties is non-trivial and often serendipitous. Inverse design approaches based on evolutionary search have been powerful tools, but such traditional searches have focused on identifying primarily stable and metastable materials with the lowest enthalpy. The inverse design of materials, with a focus on a desired property such as, for example, hardness is a challenging task because of the expensive computational cost involved in sampling multiple structures. The recent advances in machine learning have brought new powerful AI techniques to the forefront which can potentially revolutionize the inverse design and discovery of materials, especially metastable phases capable of meeting multifunctionality. Here, in this work, we develop and apply an automated reinforcement learning workflow for inverse design that integrates first principles physics and atomistic simulations with machine learning (ML), and high-performance computing to allow rapid exploration of the superhard and ultrahard metastable phases of Carbon. We demonstrate an automatic machine learning based inverse design workflow to map new undiscovered metastable states ranging from near equilibrium to those far-from-equilibrium that satisfy multiple property objectives, specifically bulk moduli, shear moduli and hardness. We create a comprehensive library of carbon stable and metastable phases with varying hardness and subsequently shortlist 10 top performing candidate carbon structures, including two newly reported phases, based on their hardness and characterize their temperature dependent mechanical properties. A neural network model is built using featurization of allotropes of carbon to predict the quasi-harmonic Gibbs free energies. The Gibbs free energies of the top performing phases are analyzed to get an estimate of the experimental synthesizability of these superhard and ultrahard carbon phases. In general, we show using machine learning based inverse design approaches how hitherto inaccessible metastable states can be identified and potentially synthesized to meet the demand for multifunctional materials.

Balasubramanian, Karthik [Univ. of Illinois, Chica↗

Structural basis of the stereoselective formation of the spirooxindole ring in the biosynthesis of citrinadins

Prenylated indole alkaloids featuring spirooxindole rings possess a 3 R or 3 S carbon stereocenter, which determines the bioactivities of these compounds. Despite the stereoselective advantages of spirooxindole biosynthesis compared with those of organic synthesis, the biocatalytic mechanism for controlling the 3 R or 3 S -spirooxindole formation has been elusive. Here, we report an oxygenase/semipinacolase CtdE that specifies the 3 S -spirooxindole construction in the biosynthesis of 21 R -citrinadin A. High-resolution X-ray crystal structures of CtdE with the substrate and cofactor, together with site-directed mutagenesis and computational studies, illustrate the catalytic mechanisms for the possible β-face epoxidation followed by a regioselective collapse of the epoxide intermediate, which triggers semipinacol rearrangement to form the 3 S -spirooxindole. Comparing CtdE with PhqK, which catalyzes the formation of the 3 R -spirooxindole, we reveal an evolutionary branch of CtdE in specific 3 S spirocyclization. Our study provides deeper insights into the stereoselective catalytic machinery, which is important for the biocatalysis design to synthesize spirooxindole pharmaceuticals.

59 BASIC BIOLOGICAL SCIENCES↗

Vertex finding in neutrino-nucleus interaction: a model architecture comparison

We compare different neural network architectures for machine learning algorithms designed to identify the neutrino interaction vertex position in the MINERvA detector. The architectures developed and optimized by hand are compared with the architectures developed in an automated way using the package “Multi-node Evolutionary Neural Networks for Deep Learning” (MENNDL), developed at Oak Ridge National Laboratory. While the domain-expert hand-tuned network was the best performer, the differences were negligible and the auto-generated networks performed as well. There is always a trade-off between human, and computer resources for network optimization and this work suggests that automated optimization, assuming resources are available, provides a compelling way to save significant expert time.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗