Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Models, Biological”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Roles for epigenetics in wood formation and stress response intrees–from basic biology to forest management

Annual model and crop species have been the subject of most epigenetic studies for plants. In contrast to annuals, forest trees persist on natural landscapes and experience environmental variation within and across seasons, years, and decades or even centuries. Most forest trees species are undomesticated and typically grown on variable landscapes with no irrigation or application of agricultural chemicals. Forest trees must thus rely on their inherent ability to alter growth and physiology to mitigate the effects of changing abiotic and biotic stressors. Like other plants, trees have mechanisms encoded in their genomic DNA sequence that can respond directly to stress events such as drought or heat. Hypothetically, it would be highly advantageous to join these mechanisms with a dynamic “memory” of past exposure to stress. It is now well established that annual model and crop plants can establish epigenetic-based memory of stress events that support more rapid and robust response to stress in the future. Here, evidence is discussed for epigenetic regulation and “memory” in two fundamental biological processes in trees, wood formation and abiotic stress response. Wood formation is an ideal trait for epigenetic research in trees, as wood formation is highly responsive to environmental conditions and includes multiple rapid developmental changes as cells adopt distinct fates within complex tissues. This is followed by a discussion of research needs that would provide the foundation for new epigenetic applications for forestry.

Groover, Andrew↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

pyDiSCaMB : enabling the use of multipolar scattering factors in Phenix

Multipolar scattering models, such as the transferable aspherical atom model, account for atomic chemical interactions and provide a more accurate representation of experimental data. However, the simpler independent atom model (IAM), which assumes non-interacting atoms, is the only model available in the most widely used macromolecular refinement programs. This is primarily because IAM offers a hard-to-beat combination of computational efficiency and modelling power at typical macromolecular resolutions. By contrast, more accurate multipolar modelling has historically been limited due to its computational cost and the absence of an interface between software capable of calculating structure factors and gradients based on multipolar models and software designed for macromolecular refinement. This work introduces pyDiSCaMB , a Python software package designed to integrate between the computational crystallography toolbox ( cctbx ) and the quantum crystallography library DiSCaMB ( Densities in Structural Chemistry and Molecular Biology ), thus enabling multipolar scattering models in Phenix 's toolkit. The implementation, features and capabilities of pyDiSCaMB are presented, the runtimes for the calculation of structure factor and target gradients with respect to atomic parameters are explored, and Fourier images of electrostatic potential, electron density and deformation maps are computed as illustrative examples. The pyDiSCaMB library will make multipolar modelling widely available to the structural biology community, potentially transforming refinement and model-building for both crystallography and cryogenic electron microscopy (cryoEM).

MATTS data bank↗

An efficient hybrid downscaling framework to estimate high-resolution river hydrodynamics

Flow depth and velocity are the most important hydrodynamic variables that govern various river functions, including water resources, navigation, sediment transport, and biogeochemical cycling. Existing high-resolution flow depth simulations rely on either computationally expensive river hydrodynamic models (RHMs) or data-driven models with formidable training costs, whereas data-driven modeling of flow velocity has rarely been explored. Here, using the hybrid Low-fidelity, Spatial analysis, and Gaussian process learning (LSG) model, we developed a downscaling approach to construct high-resolution flow depth and velocity from a two-dimensional (2-D) RHM simulation at coarse resolution. The LSG models were trained and tested in an urban watershed in Houston using two different hurricane-driven flood events. The high-resolution (as fine as 30 m resolution) and low-resolution (mostly 1000 m resolution) meshes include 664 724 and 14 536 grid cells, respectively. The results showed that through downscaling, the simulation errors were reduced to less than one-fourth and one-third of the errors of the low-resolution 2-D RHM for flow depth and velocity, respectively. Our analysis further revealed that the dominant uncertainty sources of the downscaled hydrodynamics are different, with flow velocity dominated by the dimensionality reduction error, which we reduced by using a regionalized training procedure. The downscaling approach achieves an 84-fold acceleration in computational time compared to the high-resolution 2-D RHM, making high-fidelity ensemble flood modeling feasible. More importantly, the developed method provides an opportunity to couple large-scale hydrodynamical processes with local physical, chemical, and biological processes in river models.

Tan, Zeli [Pacific Northwest National Laboratory (↗

A Corrected Score Function Framework for Modelling Circadian Gene Expression

Many biological processes display oscillatory behaviour based on an approximately 24 h internal timing system specific to each individual. One process of particular interest is gene expression, for which several circadian transcriptomic studies have identified associations between gene expression during a 24 h period and an individual's health. A challenge with analysing data from these studies is that each individual's internal timing system is offset relative to the 24 h day-night cycle, where day–night cycle time is recorded for each collected sample. Laboratory procedures can accurately determine each individual's offset and determine the internal time of sample collection. However, these laboratory procedures are labour-intensive and expensive. Here, in this paper, we propose a corrected score function framework to obtain a regression model of gene expression given internal time when the offset of each individual is too burdensome to determine. A feature of this framework is that it does not require the probability distribution generating offsets to be symmetric with a mean of zero. Simulation studies validate the use of this corrected score function framework for cosinor regression, which is prevalent in circadian transcriptomic studies. Illustrations with data from three circadian transcriptomic studies further demonstrate that the proposed framework consistently mitigates bias relative to using a score function that does not account for this offset.

59 BASIC BIOLOGICAL SCIENCES↗

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

Transcriptional response of Methanosarcina acetivorans to repression of the energy-conserving methanophenazine: CoM-CoB heterodisulfide reductase enzyme HdrED

ABSTRACT Methane-producing archaea are key organisms in the anaerobic carbon cycle. These organisms, also called methanogens, grow by converting substrate to methane gas in a process called methanogenesis. Previous research showed that the reduction of the terminal electron acceptor is the rate-limiting step in methanogenesis by Methanosarcina acetivorans . In order to gain insight into how the cells sense and respond to the availability of the terminal electron acceptor, we designed an experiment to deplete cells of the essential terminal oxidase enzyme, HdrED. We found that the depletion of HdrED in vivo results in a higher abundance of transcripts for methyltransferases ( mtaC2, mtaB3, mtaC3 ), coenzyme B biosynthesis, C1 metabolism, and pyrimidine compounds. In most cases, these changes were distinct from transcript abundance changes observed during the transition from exponential growth to stationary phase cultures. These data implicate the methylotrophic methanogenesis regulator MsrC (MA4383) in CoM-S-S-CoB heterodisulfide sensing and indicate cells have a specific mechanism to sense intracellular ratio of CoM-S-S-CoB, coenzyme M, and coenzyme B thiols and further suggest transcripts encoding translation and methanogenesis functions are controlled by feed-forward regulation depending on substrate availability. IMPORTANCE Methanosarcina is an emerging model archaeon and synthetic biology platform for the production of renewable energy and sustainable chemicals to reduce dependence on petroleum. Research into metabolic networks and gene regulation in this organism and other methanogens will inform genome-scale metabolic modeling and microbial function prediction in uncultured or non-model anaerobes and archaea. This study suggests methanogens use unknown mechanisms to efficiently couple methanogenesis to gene regulation via CoM-S-S-CoB and ATP availability.

Buan, Nicole R. (ORCID:000000027560973X)↗

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics↗

Physicochemical and biological characterization of a bispecific antibody in a CrossMab/KIH format that targets EGFR and VEGF-A

Introduction Bispecific antibodies (BsAbs) are a class of antibody therapeutics engineered in various molecular formats to bind two distinct antigens and potentially mediate multiple biological effects. These molecular formats are tailored to mediate specific mechanisms of action and possess unique physicochemical and biological properties that are necessary to assure product quality. In ovarian cancer (OC), both EGFR- and VEGF-A-mediated signaling pathways are often upregulated and cooperate to promote tumor growth and angiogenesis. Thus, inhibiting of EGFR- and VEGF-A pathways with a BsAb may provide synergistic anti-tumor activity. Methods Using publicly available sequences and applying immunoglobulin domain crossover (CrossMab) and knobs-into-holes (KIH) technologies, we generated a BsAb to simultaneously bind EGFR and VEGF-A (designated as anti-EGFR/VEGF-A BsAb). This BsAb served as a model for physiochemical and biological characterization of quality attributes that would be critical for the BsAb’s mechanisms of action. Our goal was to gain fundamental insights into BsAbs designed to target a receptor with one arm and a soluble ligand with the other, to support bioassay development and inform quality control strategies. Results Our data demonstrated that the CrossMab/KIH platform successfully produced a correctly assembled BsAb during cell culture. Characterization confirmed that the anti-EGFR/VEGF-A BsAb bound both EGFR and VEGF-A with comparable activity and affinity to the respective parental monoclonal antibodies. Functionally, the BsAb disrupted both EGF/EGFR and VEGF-A/VEGFR2 signaling pathways in OC and human umbilical vein endothelial cell (HUVEC) models. Furthermore, the BsAb effectively blocked angiogenic signaling driven by VEGF-A secreted from OC cells in a paracrine manner. Discussion Based on the combinatorial mechanism of action and our characterization findings, we concluded that two or more bioassays may be needed to accurately assess the activity of both arms of this type of BsAb.

Immunology↗

Advocating Feedback Control for Human-Earth System Applications

This paper proposes a feedback control perspective for Human-Earth Systems (HESs) which essentially are complex systems that capture the interactions between humans and nature. Recent attention in HES research has been directed towards devising strategies for climate change mitigation and adaptation, aimed at achieving environmental and societal objectives. However, existing approaches heavily rely on HES models, which inherently suffer from inaccuracies due to the complexity of the system. Moreover, overly detailed models often prove impractical for optimization tasks. We propose a framework inheriting from feedback control strategies the robustness against model errors, because inaccuracies are mitigated using measurements retrieved from the field. The framework comprises two nested control loops. The outer loop computes the optimal inputs to the HES, which are then implemented by actuators controlled in the inner loop. Potential fields of applications are also identified and a numerical example is provided.

biological system modeling↗

Comparative genomics of Aspergillus nidulans and section Nidulantes

Aspergillus nidulans is an important model organism for eukaryotic biology and the reference for the section Nidulantes in comparative studies. In this study, we de novo sequenced the genomes of 25 species of this section. Whole-genome phylogeny of 34 Aspergillus species and Penicillium chrysogenum clarifies the position of clades inside section Nidulantes. Comparative genomics reveals a high genetic diversity between species with 684 up to 2433 unique protein families. Furthermore, we categorized 2118 secondary metabolite gene clusters (SMGC) into 603 families across Aspergilli, with at least 40 % of the families shared between Nidulantes species. Genetic dereplication of SMGC and subsequent synteny analysis provides evidence for horizontal gene transfer of a SMGC. Proteins that have been investigated in A. nidulans as well as its SMGC families are generally present in the section Nidulantes, supporting its role as model organism. The set of genes encoding plant biomass-related CAZymes is highly conserved in section Nidulantes, while there is remarkable diversity of organization of MAT-loci both within and between the different clades. This study provides a deeper understanding of the genomic conservation and diversity of this section and supports the position of A. nidulans as a reference species for cell biology.

Theobald, Sebastian [Technical University of Denma↗

Wormholes without averaging

After averaging over fermion couplings, SYK has a collective field description that sometimes has “wormhole” solutions. We study the fate of these wormholes when the couplings are fixed. Working mainly in a simple model, we find that the wormhole saddles persist, but that new saddles also appear elsewhere in the integration space — “half-wormholes.” The wormhole contributions depend only weakly on the specific choice of couplings, while the half-wormhole contributions are strongly sensitive. The half-wormholes are crucial for factorization of decoupled systems with fixed couplings, but they vanish after averaging, leaving the non-factorizing wormhole behind.

1/N Expansion↗

Native Top-Down Mass Spectrometry Characterization of Model Integral Membrane Protein Bacteriorhodopsin

Bacteriorhodopsin (bR) from Halobacterium salinarum has been a model system for structural biology and is a structural template for the characterization of membrane G-protein couple receptors (GPCRs) in particular. Here, in this study, wild-type bacteriorhodopsin and two single-residue mutants were characterized by native top-down mass spectrometry (nTD-MS) with Orbitrap-based high-energy collision dissociation (HCD) and electron capture dissociation (ECD). After in-source dissociation ejected the membrane protein from detergent micelles, high-resolution native MS measurement allowed for identification of multiple proteoforms as well as lipid-bound forms. Further top-down MS measurements by HCD produced a large number of product ions for in-depth sequencing and unambiguous localization of post-translational modifications. For the first time, native TD-MS with ECD was used to characterize an integral membrane protein. ECD yielded fragments originating from all helices and loop regions, even accessing a sequence stretch that HCD could not. Combining HCD and ECD fragmentation patterns significantly enhanced the sequence coverage of bR. We propose bR to be a model analyte for testing nTD-MS performance for membrane proteins.

crystal cleavage↗

Structure analysis of the telomere resolvase from the Lyme disease spirochete Borrelia garinii reveals functional divergence of its C-terminal domain

Borrelia spirochetes are the causative agents of Lyme disease and relapsing fever, two of the most common tick-borne illnesses. A characteristic feature of these spirochetes is their highly segmented genomes which consists of a linear chromosome and a mixture of up to approximately 24 linear and circular extrachromosomal plasmids. The complexity of this genomic arrangement requires multiple strategies for efficient replication and partitioning during cell division, including the generation of hairpin ends found on linear replicons mediated by the essential enzyme ResT, a telomere resolvase. Using an integrative structural biology approach employing advanced modelling, circular dichroism, X-ray crystallography and small-angle X-ray scattering, we have generated high resolution structural data on ResT from B. garinii. Our data provides the first high-resolution structures of ResT from Borrelia spirochetes and revealed active site positioning in the catalytic domain. We also demonstrate that the C-terminal domain of ResT is required for both transesterification steps of telomere resolution, and is a requirement for DNA binding, distinguishing ResT from other telomere resolvases from phage and bacteria. These results advance our understanding of the molecular function of this essential enzyme involved in genome maintenance in Borrelia pathogens.

59 BASIC BIOLOGICAL SCIENCES↗

Quantifying the basic reproduction number and underestimated fraction of Mpox cases worldwide at the onset of the outbreak

In 2022, there was a global resurgence of mpox, with different clinical-epidemiological features compared with previous outbreaks. Sexual contact was hypothesized as the primary transmission route, and the community of men having sex with men (MSM) was disproportionately affected. Because of the stigma associated with sexually transmitted infections, the real burden of mpox could be masked. We quantified the basic reproduction number (R 0 ) and the underestimated fraction of mpox cases in 16 countries, from the onset of the outbreak until early September 2022, using Bayesian inference and a compartmentalized, risk-structured (high-/low-risk populations) and two-route (sexual/non-sexual transmission) mathematical model. Machine learning (ML) was harnessed to identify underestimation determinants. Estimated R 0 ranged between 1.37 (Canada) and 3.68 (Germany). The underestimation rates for the high- and low-risk populations varied between 25–93% and 65–85%, respectively. The estimated total number of mpox cases, relative to the reported cases, is highest in Colombia (3.60) and lowest in Canada (1.08). In the ML analysis, two clusters of countries could be identified, differing in terms of attitudes towards the 2SLGBTQIAP+ community and the importance of religion. Given the substantial mpox underestimation, surveillance should be enhanced, and country-specific campaigns against the stigmatization of MSM should be organized, leveraging community-based interventions.

60 APPLIED LIFE SCIENCES↗

Design of diverse, functional mitochondrial targeting sequences across eukaryotic organisms using variational autoencoder

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

59 BASIC BIOLOGICAL SCIENCES↗