SEARCH · Engineering Papers
Results for “structural modeling”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Semi-Analytical Hierarchical Bayesian Inference of Nonlinear Model Structure in Stochastic Dynamics: Applied to Compartmental Models of Infectious Diseases
A Bayesian computational framework for parsimonious inference in stochastic nonlinear dynamical systems is presented. This framework enables the concurrent estimation of system states, time-varying parameters, time-invariant parameters, and the optimal sparsity structure of the model parameters. Because differential equation-based models are often simplified mechanistic or phenomenological representations, robust inference from noisy measurement data requires explicit treatment of model error and uncertainty. Model error and time-varying parameters can be represented as random processes, enabling inference while making minimal assumptions about the underlying sources of discrepancy and variability. Adopting stochastic differential equation representations affords the model significant flexibility, but can also render it susceptible to overfitting during statistical inversion, where the inferred model may track noise rather than the underlying signal. To alleviate the effects of overfitting and to enable the discovery of the optimal sparse representation of the time-invariant parameters, a Bayesian sparse learning algorithm is embedded within the framework. This sparse learning framework adopts an approximate hierarchical Bayesian setting defined by a series of semi-analytical expressions. The model structure inference framework is validated using a stochastic compartmental model for tracking and forecasting active cases of an infectious disease. Compartmental models describe population-level infectious disease dynamics through interactions among population fractions grouped by disease state. Mathematically, such models consist of a system of coupled ordinary differential equations. This example adopts an expressive compartmental model that includes multiple possible interactions between disease states, motivated by early uncertainty surrounding COVID-19 reinfection dynamics and their implications for long-term epidemic forecasting. The sparse learning exercise permits the inference of a priori unknown epidemiological dynamics from simulated public health data, discovering the nested compartmental model that optimizes the trade-off between average data-fit and model complexity. It is shown that inducing sparsity among the model parameters eliminates redundant interactions between compartments, equivalently revealing the optimal coupling structure between differential equations.
De novo atomic protein structure modeling for cryoEM density maps using 3D transformer and HMM
Accurately building 3D atomic structures from cryo-EM density maps is a crucial step in cryo-EM-based protein structure determination. Converting density maps into 3D atomic structures for proteins lacking accurate homologous or predicted structures as templates remains a significant challenge. Here, we introduce Cryo2Struct, a fully automated de novo cryo-EM structure modeling method. Cryo2Struct utilizes a 3D transformer to identify atoms and amino acid types in cryo-EM density maps, followed by an innovative Hidden Markov Model (HMM) to connect predicted atoms and build protein backbone structures. Cryo2Struct produces substantially more accurate and complete protein structural models than the widely used ab initio method Phenix. Additionally, its performance in building atomic structural models is robust against changes in the resolution of density maps and the size of protein structures.
Feedback density and causal complexity of simulation model structure
Measures of simulation model complexity generally focus on outputs; we propose measuring the complexity of a model’s causal structure to gain insight into its fundamental character. This article introduces tools for measuring causal complexity. First, we introduce a method for developing a model’s causal structure diagram, which characterises the causal interactions present in the code. Causal structure diagrams facilitate comparison of simulation models, including those from different paradigms. Next, we develop metrics for evaluating a model’s causal complexity using its causal structure diagram. We discuss cyclomatic complexity as a measure of the intricacy of causal structure and introduce two new metrics that incorporate the concept of feedback, a fundamental component of causal structure. The first new metric introduced here is feedback density, a measure of the cycle-based interconnectedness of causal structure. The second metric combines cyclomatic complexity and feedback density into a comprehensive causal complexity measure. Finally, we demonstrate these complexity metrics on simulation models from multiple paradigms and discuss potential uses and interpretations. These tools enable direct comparison of models across paradigms and provide a mechanism for measuring and discussing complexity based on a model’s fundamental assumptions and design.
OpenMDlr: parallel, open-source tools for general protein structure modeling and refinement from pairwise distances
Easy-to-use, open-source, general-purpose programs for modeling a protein structure from inter-atomic distances are needed for modeling from experimental data and refinement of predicted protein structures. OpenMDlr is an open-source Python package for modeling protein structures from pairwise distances between any atoms, and optionally, dihedral angles. Finally, we provide a user-friendly input format for harnessing modern biomolecular force fields in an easy-to-install package that can efficiently make use of multiple compute cores.
Structural Models of the Rhodopseudomonas palustris Proteome
This dataset contains the structural models for the primary transcripts of the Rhodopseudomonas palustris proteome. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. palustris proteome to those available in the AlphaFold Protein Structure Database.
Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome
This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.
Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome
This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.
Updated resources for exploring experimentally-determined PDB structures and Computed Structure Models at the RCSB Protein Data Bank
The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.
RCSB Protein Data Bank (RCSB.org): delivery of experimentally-determined PDB structures alongside one million computed structure models of proteins from artificial intelligence/machine learning
Abstract The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB), founding member of the Worldwide Protein Data Bank (wwPDB), is the US data center for the open-access PDB archive. As wwPDB-designated Archive Keeper, RCSB PDB is also responsible for PDB data security. Annually, RCSB PDB serves >10 000 depositors of three-dimensional (3D) biostructures working on all permanently inhabited continents. RCSB PDB delivers data from its research-focused RCSB.org web portal to many millions of PDB data consumers based in virtually every United Nations-recognized country, territory, etc. This Database Issue contribution describes upgrades to the research-focused RCSB.org web portal that created a one-stop-shop for open access to ∼200 000 experimentally-determined PDB structures of biological macromolecules alongside >1 000 000 incorporated Computed Structure Models (CSMs) predicted using artificial intelligence/machine learning methods. RCSB.org is a ‘living data resource.’ Every PDB structure and CSM is integrated weekly with related functional annotations from external biodata resources, providing up-to-date information for the entire corpus of 3D biostructure data freely available from RCSB.org with no usage limitations. Within RCSB.org, PDB structures and the CSMs are clearly identified as to their provenance and reliability. Both are fully searchable, and can be analyzed and visualized using the full complement of RCSB.org web portal capabilities.
Structural models and functional annotations for the Sphagnum divinum proteome
This dataset contains the structural models for the primary transcripts of the Sphagnum divinum proteome. Additionally, for a subset of these proteins, sequence and structural alignment results are provided. This dataset represents the most thorough structural study of a Sphagnum species, also known as peat mosses, by providing three-dimensional atomic resolution structures of the majority of the encoded proteins as well as structural alignment results used in the application of annotating the proteome. References (DOI) AlphaFold v2 Monomer: https://doi.org/10.1038/s41586-021-03819-2. References (DOI) US-align2: https://doi.org/10.1038/s41592-022-01585-1
RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models
Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.
RootSlice —A novel functional-structural model for root anatomical phenotypes
Root anatomy is an important determinant of root metabolic costs, soil exploration, and soil resource capture. Root anatomy varies substantially within and among plant species. RootSlice is a multicellular functional-structural model of root anatomy developed to facilitate the analysis and understanding of root anatomical phenotypes. RootSlice can capture phenotypically accurate root anatomy in three dimensions of different root classes and developmental zones, of both monocotyledonous and dicotyledonous species. Several case studies are presented illustrating the capabilities of the model. For maize nodal roots, the model illustrated the role of vacuole expansion in cell elongation; and confirmed the individual and synergistic role of increasing root cortical aerenchyma and reducing the number of cortical cell files in reducing root metabolic costs. Integration of RootSlice for different root zones as the temporal properties of the nodal roots in the whole-plant and soil model OpenSimRoot/maize enabled the multiscale evaluation of root anatomical phenotypes, highlighting the role of aerenchyma formation in enhancing the utility of cortical cell files for improving plant performance over varying soil nitrogen supply. Such integrative in silico approaches present avenues for exploring the fitness landscape of root anatomical phenotypes.
Molecular structure models of amorphous bismuth and cerium carboxylate catalyst precursors
As our societal need for materials and energy has grown, so has our need for catalyst processes in hydrogen production. A major function in these applications, for both homogenous and heterogeneous catalysis processes, is the synthesis of an active metal catalyst. It must first be soluble to control the physical properties of the metal being used. Recent work in metal precursors has begun to turn toward these metal carboxylate types of material. Here, structural models are proposed for bismuth 2-ethylhexanoate and 2,2-dimethyloctanoate and cerium 2-ethylhexanoate. The bismuth compounds have been characterized at different ratios of bismuth to carboxylate as solutions of the free acids. Their structures are most consistent with a Bi 4 (RCO 2 ) 12 motif where the Bi ions are arranged in a flattened tetrahedron with Bi – Bi distances of about 4.3 Å. There is evidence for Bi – O – Bi linkages at low free acid concentrations. The cerium compound is most consistent with a linear tetracerium molecule where the Ce – Ce distances repeat at about 4.3 Å out to 16.4 Å. The models were generated by analogy with known crystal structures and compared to high-energy x-ray scattering data. To further evaluate the models, DFT calculations were made, and the equilibrium geometries were compared. The vibrational spectra calculated from those geometries are presented and compared to the experimental results. Magnetization vs. temperature data was collected on the cerium compound, and its behavior was consistent with the proposed model. A geometrical approach to determining the dimensionality and relative positions of the metal ions in these structures is presented.
Integrated structural model of the palladin–actin complex using XL ‐ MS , docking, NMR , and SAXS
Abstract Palladin is an actin‐binding protein that accelerates actin polymerization and is linked to the metastasis of several types of cancer. Previously, three lysine residues in an immunoglobulin‐like domain of palladin have been identified as essential for actin binding. However, it is still unknown where palladin binds to F‐actin. Evidence that palladin binds to the sides of actin filaments to facilitate branching is supported by our previous study showing that palladin was able to compensate for Arp2/3 in the formation of Listeria actin comet tails. Here, we used chemical crosslinking to covalently link palladin and F‐actin residues based on spatial proximity. Samples were then enzymatically digested, separated by liquid chromatography, and analyzed by tandem mass spectrometry. Peptides containing the crosslinks and specific residues involved were then identified for input to the HADDOCK docking server to model the most likely binding conformation. Small‐angle x‐ray scattering was used to provide further insight into palladin flexibility and the binding interface, and NMR spectra identified potential interactions between palladin's Ig domains. Our final structural model of the F‐actin:palladin complex revealed how palladin interacts with and stabilizes F‐actin at the interface between two actin monomers. Three actin residues that were identified in this study also appear commonly in the actin‐binding interface with other proteins such as myotilin, myosin, and tropomodulin. An accurate structural representation of the complex between palladin and actin extends our understanding of palladin's role in promoting cancer metastasis through the regulation of actin dynamics.
Multilevel atomic structural model for interstratified opal materials
The structure of opal has long fascinated scientists. It occurs in a number of structural states, ranging from amorphous to exhibiting features of stacking disorder. Opal-CT, where C and T signify cristobalite- and tridymite-like interstratification, represents an important link in the length scales between amorphous and crystalline states. However, details about local atomic (dis)order and arrangements extending to long-range stacking faults in opal polymorphs remain incompletely understood. Here, a multilevel modeling approach is reported that considers stacking states in correlation with the abundance of C and T segments as a high-level structural parameter (i.e. not each atom). Optimization accounting for inter-tetrahedral bond lengths and angles and the regularity of the silicate tetrahedra is included as lower levels of structural parameters. Together, a set of parameters with both coarse-grained and atomistic features for different levels of structural details is refined. Structural disorder at the ~10–100 Å distance scale is evaluated using experimental pair distribution function and diffraction datasets, comparing peak intensities, widths and asymmetry. Here this work presents a complete multilevel structural description of natural opal-CT and explains many of the unusual features observed in X-ray powder diffraction patterns. This modeling approach can be adopted generally for analyzing layered materials and their assembly into 3D structures.
ModelCIF: An Extension of PDBx/mmCIF Data Representation for Computed Structure Models
Not Available