Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “biologists”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Deep Learning Coordinate-Free Quantum Chemistry

Computing quantum chemical properties of small molecules and polymers can provide insights valuable to physicists, chemists and biologists when designing new materials, catalysts, biological probes and drugs. Deep learning can compute quantum chemical properties accurately in a fraction of the time required by commonly used methods such as density functional theory (DFT). However, many of these deep learning architectures require energy minimized molecular geometries as input, which is also computationally expensive, and decreasing the reproducibility and throughput of these methods. In this study, we demonstrate that accurate quantum chemical computations can be performed without optimized geometries by operating in the coordinate-free domain using deep learning on graph encodings. Furthermore, we also find that the choice of graph-encoding architecture substantially affects the performance of these methods. The Wave architecture outperforms graph convolution architectures, particularly on complex molecules. Furthermore, the structures of these graph encoding architectures provide an opportunity to probe an important, outstanding question in quantum mechanics: What types of quantum chemical properties can be represented by local-variable models? We find that Wave, a local-variable model, is more accurately calculates quantum chemical properties. Graph convolutional architectures require global variables, and are not as effective as as Wave. We anticipate that coordinate-free, deep-learning models of quantum chemistry will become valuable tools in chemistry and biology, enabling researchers to rapidly screen chemical databases or identify new molecules using automated, de-novo design algorithms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Strain Construction for Biosynthetic Pathway Screening in Yeast

Automation accelerates the Design-Build-Test-Learn (DBTL) cycle for synthetic biology; however, most strain construction pipelines lack robotic integration. Here, in this study, we present the workflow design and source code for a modular, integrated protocol that automates the Build step in Saccharomyces cerevisiae. We programmed the Hamilton Microlab VANTAGE to integrate off-deck hardware via its central robotic arm, enabling automated steps that increased throughput to 2,000 transformations per week. We developed a user interface with the Hamilton VENUS software to support on-demand parameter customization. As a proof of concept, we screened a gene library in an engineered yeast strain producing verazine, a key intermediate in the biosynthesis of steroidal alkaloids. Our pipeline rapidly identified pathway bottlenecks and genes that enhanced verazine production by 2.0- to 5-fold. This technical note provides resources for synthetic biologists designing yeast workflows for biofoundries to screen libraries for pathway discovery/optimization, combinatorial biosynthesis, and protein engineering.

automation↗

Impact of structural biology and the protein data bank on us fda new drug approvals of low molecular weight antineoplastic agents 2019–2023

Abstract Open access to three-dimensional atomic-level biostructure information from the Protein Data Bank (PDB) facilitated discovery/development of 100% of the 34 new low molecular weight, protein-targeted, antineoplastic agents approved by the US FDA 2019–2023. Analyses of PDB holdings, the scientific literature, and related documents for each drug-target combination revealed that the impact of structural biologists and public-domain 3D biostructure data was broad and substantial, ranging from understanding target biology (100% of all drug targets), to identifying a given target as likely druggable (100% of all targets), to structure-guided drug discovery (>80% of all new small-molecule drugs, made up of 50% confirmed and >30% probable cases). In addition to aggregate impact assessments, illustrative case studies are presented for six first-in-class small-molecule anti-cancer drugs, including a selective inhibitor of nuclear export targeting Exportin 1 (selinexor, Xpovio), an ATP-competitive CSF-1R receptor tyrosine kinase inhibitor (pexidartinib,Turalia), a non-ATP-competitive inhibitor of the BCR-Abl fusion protein targeting the myristoyl binding pocket within the kinase catalytic domain of Abl (asciminib, Scemblix), a covalently-acting G12C KRAS inhibitor (sotorasib, Lumakras or Lumykras), an EZH2 methyltransferase inhibitor (tazemostat, Tazverik), and an agent targeting the basic-Helix-Loop-Helix transcription factor HIF-2α (belzutifan, Welireg).

60 APPLIED LIFE SCIENCES↗

Leveraging public AI tools to explore systems biology resources in mathematical modeling

Predictive mathematical modeling is an essential part of systems biology and is interconnected with information management. Systems biology information is often stored in specialized formats to facilitate data storage and analysis. These formats are not designed for easy human readability and thus require specialized software to visualize and interpret results. Therefore, comprehending modeling and underlying networks and pathways is contingent on mastering systems biology tools, which is particularly challenging for users with no or little background in data science or system biology. To address this challenge, we investigated the usage of public Artificial Intelligence (AI) tools in exploring systems biology resources in mathematical modeling. We tested public AI’s understanding of mathematics in models, related systems biology data, and the complexity of model structures. Our approach can enhance the accessibility of systems biology for non-system biologists and help them understand systems biology without a deep learning curve.

59 BASIC BIOLOGICAL SCIENCES↗

Mapping protein dynamics at high spatial resolution with temperature-jump X-ray crystallography

Understanding and controlling protein motion at atomic resolution is a hallmark challenge for structural biologists and protein engineers because conformational dynamics are essential for complex functions such as enzyme catalysis and allosteric regulation. Time-resolved crystallography offers a window into protein motions, yet without a universal perturbation to initiate conformational changes the method has been limited in scope. Here we couple a solvent-based temperature jump with time-resolved crystallography to visualize structural motions in lysozyme, a dynamic enzyme. We observed widespread atomic vibrations on the nanosecond timescale, which evolve on the submillisecond timescale into localized structural fluctuations that are coupled to the active site. An orthogonal perturbation to the enzyme, inhibitor binding, altered these dynamics by blocking key motions that allow energy to dissipate from vibrations into functional movements linked to the catalytic cycle. Because temperature jump is a universal method for perturbing molecular motion, the method demonstrated here is broadly applicable for studying protein dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

RNA-Puzzles Round V: blind predictions of 23 RNA structures

RNA-Puzzles is a collective endeavor dedicated to the advancement and improvement of RNA three-dimensional structure prediction. With agreement from structural biologists, RNA structures are predicted by modeling groups before publication of the experimental structures. We report a large-scale set of predictions by 18 groups for 23 RNA-Puzzles: 4 RNA elements, 2 Aptamers, 4 Viral elements, 5 Ribozymes and 8 Riboswitches. We describe automatic assessment protocols for comparisons between prediction and experiment. Our analyses reveal some critical steps to be overcome to achieve good accuracy in modeling RNA structures: identification of helix-forming pairs and of non-Watson–Crick modules, correct coaxial stacking between helices and avoidance of entanglements. Three of the top four modeling groups in this round also ranked among the top four in the CASP15 contest.

59 BASIC BIOLOGICAL SCIENCES↗

Disturbance of hibernating bats due to researchers entering caves to conduct hibernacula surveys

Estimating population changes of bats is important for their conservation. Population estimates of hibernating bats are often calculated by researchers entering hibernacula to count bats; however, the disturbance caused by these surveys can cause bats to arouse unnaturally, fly, and lose body mass. We conducted 17 hibernacula surveys in 9 caves from 2013 to 2018 and used acoustic detectors to document cave-exiting bats the night following our surveys. We predicted that cave-exiting flights (i.e., bats flying out and then back into caves) of Townsend’s big-eared bats (Corynorhinus townsendii) and western small-footed myotis (Myotis ciliolabrum) would be higher the night following hibernacula surveys than on nights following no surveys. Those two species, however, did not fly out of caves more than predicted the night following 82% of surveys. Nonetheless, the activity of bats flying out of caves following surveys was related to a disturbance factor (i.e., number of researchers × total time in a cave). We produced a parsimonious model for predicting the probability of Townsend’s big-eared bats flying out of caves as a function of disturbance factor and ambient temperature. That model can be used to help biologists plan for the number of researchers, and the length of time those individuals are in a cave to minimize disturbing bats.

59 BASIC BIOLOGICAL SCIENCES↗

BioCARS: Synchrotron facility for probing structural dynamics of biological macromolecules

A major goal in biomedical science is to move beyond static images of proteins and other biological macromolecules to the internal dynamics underlying their function. This level of study is necessary to understand how these molecules work and to engineer new functions and modulators of function. Stemming from a visionary commitment to this problem by Keith Moffat decades ago, a community of structural biologists has now enabled a set of x-ray scattering technologies for observing intramolecular dynamics in biological macromolecules at atomic resolution and over the broad range of timescales over which motions are functionally relevant. Many of these techniques are provided by BioCARS, a cutting-edge synchrotron radiation facility built under Moffat leadership and located at the Advanced Photon Source at Argonne National Laboratory. BioCARS enables experimental studies of molecular dynamics with time resolutions spanning from 100 ps to seconds and provides both time-resolved x-ray crystallography and small- and wide-angle x-ray scattering. Structural changes can be initiated by several methods—UV/Vis pumping with tunable picosecond and nanosecond laser pulses, substrate diffusion, and global perturbations, such as electric field and temperature jumps. Studies of dynamics typically involve subtle perturbations to molecular structures, requiring specialized computational techniques for data processing and interpretation. In this review, we present the challenges in experimental macromolecular dynamics and describe the current state of experimental capabilities at this facility. As Moffat imagined years ago, BioCARS is now positioned to catalyze the scientific community to make fundamental advances in understanding proteins and other complex biological macromolecules.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid and automated design of two-component protein nanomaterials using ProteinMPNN

The design of protein–protein interfaces using physics-based design methods such as Rosetta requires substantial computational resources and manual refinement by expert structural biologists. Deep learning methods promise to simplify protein–protein interface design and enable its application to a wide variety of problems by researchers from various scientific disciplines. Here, we test the ability of a deep learning method for protein sequence design, ProteinMPNN, to design two-component tetrahedral protein nanomaterials and benchmark its performance against Rosetta. ProteinMPNN had a similar success rate to Rosetta, yielding 13 new experimentally confirmed assemblies, but required orders of magnitude less computation and no manual refinement. The interfaces designed by ProteinMPNN were substantially more polar than those designed by Rosetta, which facilitated in vitro assembly of the designed nanomaterials from independently purified components. Crystal structures of several of the assemblies confirmed the accuracy of the design method at high resolution. Our results showcase the potential of deep learning–based methods to unlock the widespread application of designed protein–protein interfaces and self-assembling protein nanomaterials in biotechnology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Convergent evolution of NFP -facilitated root nodule symbiosis

The origin and phylogenetic distribution of symbiotic associations between nodulating angiosperms and nitrogen-fixing bacteria have long intrigued biologists. Recent comparative evolutionary analyses have yielded alternative hypotheses: a multistep pathway of independent gains and losses of root nodule symbiosis vs. a single gain followed by numerous losses. A detailed reconstruction of the history of genes involved in signaling between nitrogen-fixing bacteria and potential hosts, particularly lipo-chitooligosaccharide (LCO) signaling, is needed to distinguish between these hypotheses. LCO recognition by plants involves the Nod Factor Perception ( NFP ) gene family; in the legume model Medicago truncatula (Fabales), MtNFP is essential for establishing rhizobial symbiosis. Here, we document convergent evolution of NFP , indicating multiple origins of LCO-driven symbiosis. In contrast to previous models that explain the recruitment of NFP via a single duplication in the ancestor of the nitrogen-fixing clade, our phylogenomic and synteny results suggest this duplication does not span the entire clade. Tandem duplication in a common ancestor of Cucurbitales and Rosales resulted in the NFP1 and NFP2 groups. In contrast, the phylogenetically closest paralog of MtNFP is MtLYR1 , located on a different chromosome within a large syntenic block. All available data indicate that a large-scale duplication resulted in MtNFP and MtLYR1 , likely corresponding to a whole-genome duplication in an ancestor of subfamily Papilionoideae of Fabaceae. We show that MtNFP and the NFP2 -like group are not orthologous, indicating multiple independent gains of NFP -based LCO signaling. This molecular convergence provides a possible mechanism for multiple gains of root nodule symbiosis across the nitrogen-fixing clade.

LCO signaling↗

A roadmap for the functional annotation of protein families: a community perspective

Over the last 25 years, biology has entered the genomic era and is becoming a science of ‘big data’. Most interpretations of genomic analyses rely on accurate functional annotations of the proteins encoded by more than 500 000 genomes sequenced to date. By different estimates, only half the predicted sequenced proteins carry an accurate functional annotation, and this percentage varies drastically between different organismal lineages. Such a large gap in knowledge hampers all aspects of biological enterprise and, thereby, is standing in the way of genomic biology reaching its full potential. A brainstorming meeting to address this issue funded by the National Science Foundation was held during 3–4 February 2022. Bringing together data scientists, biocurators, computational biologists and experimentalists within the same venue allowed for a comprehensive assessment of the current state of functional annotations of protein families. Further, major issues that were obstructing the field were identified and discussed, which ultimately allowed for the proposal of solutions on how to move forward.

59 BASIC BIOLOGICAL SCIENCES↗

Testing for the Genomic Footprint of Conflict Between Life Stages in an Angiosperm and Moss Species

Abstract The maintenance of genetic variation by balancing selection is of considerable interest to evolutionary biologists. An important but understudied potential driver of balancing selection is antagonistic pleiotropy between diploid and haploid stages of the plant life cycle. Despite sharing a common genome, sporophytes (2n) and gametophytes (n) may undergo differential or even opposing selection. Theoretical work suggests antagonistic pleiotropy between life stages can generate balancing selection and maintain genetic variation. Despite the potential for far-reaching consequences of gametophytic selection, empirical tests of its pleiotropic effects (neutral, synergistic, or antagonistic) on sporophytes are generally lacking. Here, we examined the population genomic signals of selection across life stages in the angiosperm Rumex hastatulus and the moss Ceratodon purpureus. We compared gene expression between life stages and sexes, combined with neutral diversity statistics and the analysis of the distribution of fitness effects. In contrast to what would be predicted under balancing selection due to antagonistic pleiotropy, we found that unbiased genes between life stages were under stronger purifying selection, likely explained by a predominance of synergistic pleiotropy between life stages and strong purifying selection on broadly expressed genes. In addition, we found that 30% of candidate genes under balancing selection in R. hastatulus were located within inversion polymorphisms. Our findings provide novel insights into the genome-wide characteristics and consequences of plant gametophytic selection.

Evolutionary Biology↗

Machine learning-enabled computer vision for plant phenotyping: a primer on AI/ML and a case study on stomatal patterning

Abstract Artificial intelligence and machine learning (AI/ML) can be used to automatically analyze large image datasets. One valuable application of this approach is estimation of plant trait data contained within images. Here we review 39 papers that describe the development and/or application of such models for estimation of stomatal traits from epidermal micrographs. In doing so, we hope to provide plant biologists with a foundational understanding of AI/ML and summarize the current capabilities and limitations of published tools. While most models show human-level performance for stomatal density (SD) quantification at superhuman speed, they are often likely to be limited in how broadly they can be applied across phenotypic diversity associated with genetic, environmental, or developmental variation. Other models can make predictions across greater phenotypic diversity and/or additional stomatal/epidermal traits, but require significantly greater time investment to generate ground-truth data. We discuss the challenges and opportunities presented by AI/ML-enabled computer vision analysis, and make recommendations for future work to advance accelerated stomatal phenotyping.

Plant Sciences↗

MolViewSpec: a Mol* extension for describing and sharing molecular visualizations

Data visualization is a pivotal component of a structural biologist’s arsenal. The Mol* Viewer makes molecular visualizations available to broader audiences via most web browsers. While Mol* provides a wide range of functionality, it has a steep learning curve and is only available via a JavaScript interface. To enhance the accessibility and usability of web-based molecular visualization, we introduce MolViewSpec (molstar.org/mol-view-spec), a standardized approach for defining molecular visualizations that decouples the definition of complex molecular scenes from their rendering. Scene definition can include references to commonly used structural, volumetric, and annotation data formats together with a description of how the data should be visualized and paired with optional annotations specifying colors, labels, measurements, and custom 3D geometries. Developed as an open standard, this solution paves the way for broader interoperability and support across different programming languages and molecular viewers, enabling more streamlined, standardized, and reproducible visual molecular analyses. MolViewSpec is freely available as a Mol* extension and a standalone Python package.

Midlik, Adam [European Bioinformatics Institute (U↗

A rich and bountiful harvest: Key discoveries in plant cell biology

Abstract The field of plant cell biology has a rich history of discovery, going back to Robert Hooke’s discovery of cells themselves. The development of microscopes and preparation techniques has allowed for the visualization of subcellular structures, and the use of protein biochemistry, genetics, and molecular biology has enabled the identification of proteins and mechanisms that regulate key cellular processes. In this review, seven senior plant cell biologists reflect on the development of this research field in the past decades, including the foundational contributions that their teams have made to our rich, current insights into cell biology. Topics covered include signaling and cell morphogenesis, membrane trafficking, cytokinesis, cytoskeletal regulation, and cell wall biology. In addition, these scientists illustrate the pathways to discovery in this exciting research field.

59 BASIC BIOLOGICAL SCIENCES↗

DIRT/3D: 3D root phenotyping for field-grown maize ( Zea mays )

The development of crops with deeper roots holds substantial promise to mitigate the consequences of climate change. Deeper roots are an essential factor to improve water uptake as a way to enhance crop resilience to drought, to increase nitrogen capture, to reduce fertilizer inputs, and to increase carbon sequestration from the atmosphere to improve soil organic fertility. A major bottleneck to achieving these improvements is high-throughput phenotyping to quantify root phenotypes of field-grown roots. We address this bottleneck with Digital Imaging of Root Traits (DIRT)/3D, an image-based 3D root phenotyping platform, which measures 18 architecture traits from mature field-grown maize (Zea mays) root crowns (RCs) excavated with the Shovelomics technique. DIRT/3D reliably computed all 18 traits, including distance between whorls and the number, angles, and diameters of nodal roots, on a test panel of 12 contrasting maize genotypes. The computed results were validated through comparison with manual measurements. Overall, we observed a coefficient of determination of r 2 > 0.84 and a high broad-sense heritability of H$_{mean}^{2}$ < 0.6 for all but one trait. The average values of the 18 traits and a developed descriptor to characterize complete root architecture distinguished all genotypes. DIRT/3D is a step toward automated quantification of highly occluded maize RCs. Therefore, DIRT/3D supports breeders and root biologists in improving carbon sequestration and food security in the face of the adverse effects of climate change.

59 BASIC BIOLOGICAL SCIENCES↗

DNA Sequence-Based Identification of Fusarium : A Work in Progress

Accurate species-level identification of an etiological agent is crucial for disease diagnosis and management because knowing the agent’s identity connects it with what is known about its host range, geographic distribution, and toxin production potential. This is particularly true in publishing peer-reviewed disease reports, where imprecise and/or incorrect identifications weaken the public knowledge base. This can be a daunting task for phytopathologists and other applied biologists that need to identify Fusarium in particular, because published and ongoing multilocus molecular systematic studies have highlighted several confounding issues. Paramount among these are: (i) this agriculturally and clinically important genus is currently estimated to comprise more than 400 phylogenetically distinct species (i.e., phylospecies), with more than 80% of these discovered within the past 25 years; (ii) approximately one-third of the phylospecies have not been formally described; (iii) morphology alone is inadequate to distinguish most of these species from one another; and (iv) the current rapid discovery of novel fusaria from pathogen surveys and accompanying impact on the taxonomic landscape is expected to continue well into the foreseeable future. To address the critical need for accurate pathogen identification, our research groups are focused on populating two web-accessible databases (FUSARIUM-ID v.3.0 and the nonredundant National Center for Biotechnology Information nucleotide collection that includes GenBank) with portions of three phylogenetically informative genes (i.e., TEF1, RPB1, and RPB2) that resolve at or near the species level in every Fusarium species. The objectives of this Special Report, and its companion in this issue ( Torres-Cruz et al. 2022 ), are to provide a progress report on our efforts to populate these databases and to outline a set of best practices for DNA sequence-based identification of fusaria.

Plant Sciences↗

RCSB Protein Data Bank: supporting research and education worldwide through explorations of experimentally determined and computationally predicted atomic level 3D biostructures

The Protein Data Bank (PDB) was established as the first open-access digital data resource in biology and medicine in 1971 with seven X-ray crystal structures of proteins. Today, the PDB houses >210 000 experimentally determined, atomic level, 3D structures of proteins and nucleic acids as well as their complexes with one another and small molecules ( e.g. approved drugs, enzyme cofactors). These data provide insights into fundamental biology, biomedicine, bioenergy and biotechnology. They proved particularly important for understanding the SARS-CoV-2 global pandemic. The US-funded Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) and other members of the Worldwide Protein Data Bank (wwPDB) partnership jointly manage the PDB archive and support >60 000 `data depositors' (structural biologists) around the world. wwPDB ensures the quality and integrity of the data in the ever-expanding PDB archive and supports global open access without limitations on data usage. The RCSB PDB research-focused web portal at https://www.rcsb.org/ (RCSB.org) supports millions of users worldwide, representing a broad range of expertise and interests. In addition to retrieving 3D structure data, PDB `data consumers' access comparative data and external annotations, such as information about disease-causing point mutations and genetic variations. RCSB.org also provides access to >1 000 000 computed structure models (CSMs) generated using artificial intelligence/machine-learning methods. To avoid doubt, the provenance and reliability of experimentally determined PDB structures and CSMs are identified. Related training materials are available to support users in their RCSB.org explorations.

59 BASIC BIOLOGICAL SCIENCES↗