Engineering Papers⌕ Search

Engineering topics

Ramanathan, Arvind

Publications and source records attributed to Ramanathan, Arvind.

A Deep Learning-Driven Sampling Technique to Explore the Phase Space of an RNA Stem-Loop

The folding and unfolding of RNA stem-loops are critical biological processes; however, their computational studies are often hampered by the ruggedness of their folding landscape, necessitating long simulation times at the atomistic scale. Here, we adapted DeepDriveMD (DDMD), an advanced deep learning-driven sampling technique originally developed for protein folding, to address the challenges of RNA stem-loop folding. Although tempering- and order parameter-based techniques are commonly used for similar rare-event problems, the computational costs or the need for a priori knowledge about the system often present a challenge in their effective use. DDMD overcomes these challenges by adaptively learning from an ensemble of running MD simulations using generic contact maps as the raw input. DeepDriveMD enables on-the-fly learning of a low-dimensional latent representation and guides the simulation toward the undersampled regions while optimizing the resources to explore the relevant parts of the phase space. We showed that DDMD estimates the free energy landscape of the RNA stem-loop reasonably well at room temperature. Our simulation framework runs at a constant temperature without external biasing potential, hence preserving the information on transition rates, with a computational cost much lower than that of the simulations performed with external biasing potentials. Here, we also introduced a reweighting strategy for obtaining unbiased free energy surfaces and presented a qualitative analysis of the latent space. This analysis showed that the latent space captures the relevant slow degrees of freedom for the RNA folding problem of interest. Finally, throughout the manuscript, we outlined how different parameters are selected and optimized to adapt DDMD for this system. We believe this compendium of decision-making processes will help new users adapt this technique for the rare-event sampling problems of their interest.

Gupta, Ayush↗

Rapid Design and Engineering of Smart and Secure Microbiological Systems (Final Report)

The design and application of successfully engineered biosystems requires an understanding of how engineered microbes will interact with other organisms – either as one-on-one competitors or in the context of microbial consortia. Engineering microorganisms from first principles for non-laboratory, environmental applications is inherently challenging because: (1) engineered systems tend to quickly revert back to their wild-type behaviors; and (2) these systems typically pay a price in reduced fitness, making them uncompetitive against invasive contaminating species (i.e., metabolic burden). For this project, we used a synthetic biology-based strategy to investigate the organization, control, stabilization, and destabilization of natural and engineered microbes. This approach enabled development of (1) single-strain systems capable of detecting and responding to target organisms in the environment; (2) a pipeline for refining and engineering biological constructs in new, non-model host organisms; and (3) improved systems for rapidly designing, engineering, and assaying new biological modules. This coupled approach to safeguard system design is predictable and portable across bacterial species and is focused on microbes that are part of the beneficial plant microbiome. A long-term goal beyond the proposed research is to enable the rational engineering of microbial communities based on first principles of biological design that mimic the smart performance of microorganisms observed in natural systems.

59 BASIC BIOLOGICAL SCIENCES↗

Epitopes recognition of SARS-CoV-2 nucleocapsid RNA binding domain by human monoclonal antibodies

Coronavirus nucleocapsid protein (NP) of SARS-CoV-2 plays a central role in many functions important for virus proliferation including packaging and protecting genomic RNA. The protein shares sequence, structure, and architecture with nucleocapsid proteins from betacoronaviruses. The N-terminal domain (NP RBD ) binds RNA and the C-terminal domain is responsible for dimerization. After infection, NP is highly expressed and triggers robust host immune response. The anti-NP antibodies are not protective and not neutralizing but can effectively detect viral proliferation soon after infection. Two structures of SARS-CoV-2 NP RBD were determined providing a continuous model from residue 48 to 173, including RNA binding region and key epitopes. Five structures of NP RBD complexes with human mAbs were isolated using an antigen-bait sorting. Complexes revealed a distinct complement-determining regions and unique sets of epitope recognition. This may assist in the early detection of pathogens and designing peptide-based vaccines. Mutations that significantly increase viral load were mapped on developed, full length NP model, likely impacting interactions with host proteins and viral RNA.

59 BASIC BIOLOGICAL SCIENCES↗

Towards a modular architecture for science factories

Advances in robotic automation, high-performance computing, and artificial intelligence encourage us to propose large, general-purpose science factories with the scale needed to tackle large discovery problems and to support thousands of scientists.

97 MATHEMATICS AND COMPUTING↗

GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics

We seek to transform how new and emergent variants of pandemic-causing viruses, specifically SARS-CoV-2, are identified and classified. By adapting large language models (LLMs) for genomic data, we build genome-scale language models (GenSLMs) which can learn the evolutionary landscape of SARS-CoV-2 genomes. By pre-training on over 110 million prokaryotic gene sequences and fine-tuning a SARS-CoV-2-specific model on 1.5 million genomes, we show that GenSLMs can accurately and rapidly identify variants of concern. Thus, to our knowledge, GenSLMs represents one of the first whole-genome scale foundation models which can generalize to other prediction tasks. We demonstrate scaling of GenSLMs on GPU-based supercomputers and AI-hardware accelerators utilizing 1.63 Zettaflops in training runs with a sustained performance of 121 PFLOPS in mixed precision and peak of 850 PFLOPS. We present initial scientific insights from examining GenSLMs in tracking evolutionary dynamics of SARS-CoV-2, paving the path to realizing this on large biological data.

Zvyagin, Maxim↗