Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “KB”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Multiplex Editing of the Nucleoredoxin1 Tandem Array in Poplar: From Small Indels to Translocations and Complex Inversions

The CRISPR-Cas9 system has been deployed for precision mutagenesis in an ever-growing number of species, including agricultural crops and forest trees. Its application to closely linked genes with extremely high sequence similarities has been less explored. In this study, we used CRISPR-Cas9 to mutagenize a tandem array of seven Nucleoredoxin1 (NRX1) genes spanning ~100 kb in Populus tremula × Populus alba. We demonstrated efficient multiplex editing with one single guide RNA in 42 transgenic lines. The mutation profiles ranged from small insertions and deletions and local deletions in individual genes to large genomic dropouts and rearrangements spanning tandem genes. We also detected complex rearrangements including translocations and inversions resulting from multiple cleavage and repair events. Target capture sequencing was instrumental for unbiased assessments of repair outcomes to reconstruct unusual mutant alleles. The work highlights the power of CRISPR-Cas9 for multiplex editing of tandemly duplicated genes to generate diverse mutants with structural and copy number variations to aid future functional characterization.

59 BASIC BIOLOGICAL SCIENCES↗

Divergent selection and climate adaptation fuel genomic differentiation between sister species of Sphagnum (peat moss)

Abstract Background and Aims New plant species can evolve through the reinforcement of reproductive isolation via local adaptation along habitat gradients. Peat mosses (Sphagnaceae) are an emerging model system for the study of evolutionary genomics and have well-documented niche differentiation among species. Recent molecular studies have demonstrated that the globally distributed species Sphagnum magellanicum is a complex of morphologically cryptic lineages that are phylogenetically and ecologically distinct. Here, we describe the architecture of genomic differentiation between two sister species in this complex known from eastern North America: the northern S. diabolicum and the largely southern S. magniae. Methods We sampled plant populations from across a latitudinal gradient in eastern North America and performed whole genome and restriction-site associated DNA sequencing. These sequencing data were then analyzed computationally. Key Results Using sliding-window population genetic analyses we find that differentiation is concentrated within ‘islands’ of the genome spanning up to 400 kb that are characterized by elevated genetic divergence, suppressed recombination, reduced nucleotide diversity and increased rates of non-synonymous substitution. Sequence variants that are significantly associated with genetic structure and bioclimatic variables occur within genes that have functional enrichment for biological processes including abiotic stress response, photoperiodism and hormone-mediated signalling. Demographic modelling demonstrates that these two species diverged no more than 225 000 generations ago with secondary contact occurring where their ranges overlap. Conclusions We suggest that this heterogeneity of genomic differentiation is a result of linked selection and reflects the role of local adaptation to contrasting climatic zones in driving speciation. This research provides insight into the process of speciation in a group of ecologically important plants and strengthens our predictive understanding of how plant populations will respond as Earth’s climate rapidly changes.

58 GEOSCIENCES↗

efam: an e xpanded, metaproteome-supported HMM profile database of viral protein fam ilies

Viruses infect, reprogram and kill microbes, leading to profound ecosystem consequences, from elemental cycling in oceans and soils to microbiome-modulated diseases in plants and animals. Although metagenomic datasets are increasingly available, identifying viruses in them is challenging due to poor representation and annotation of viral sequences in databases. Here, we establish efam, an expanded collection of Hidden Markov Model (HMM) profiles that represent viral protein families conservatively identified from the Global Ocean Virome 2.0 dataset. This resulted in 240 311 HMM profiles, each with at least 2 protein sequences, making efam >7-fold larger than the next largest, pan-ecosystem viral HMM profile database. Adjusting the criteria for viral contig confidence from ‘conservative’ to ‘eXtremely Conservative’ resulted in 37 841 HMM profiles in our efam-XC database. To assess the value of this resource, we integrated efam-XC into VirSorter viral discovery software to discover viruses from less-studied, ecologically distinct oxygen minimum zone (OMZ) marine habitats. This expanded database led to an increase in viruses recovered from every tested OMZ virome by ~24% on average (up to ~42%) and especially improved the recovery of often-missed shorter contigs (<5 kb). Additionally, to help elucidate lesser-known viral protein functions, we annotated the profiles using multiple databases from the DRAM pipeline and virion-associated metaproteomic data, which doubled the number of annotations obtainable by standard, single-database annotation approaches. Together, these marine resources (efam and efam-XC) are provided as searchable, compressed HMM databases that will be updated bi-annually to help maximize viral sequence discovery and study from any ecosystem.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CUT&RUN identifies centromeric DNA regions of Rhodotorula toruloides IFO0880

ABSTRACT Rhodotorula toruloides has been increasingly explored as a host for bioproduction of lipids, fatty acid derivatives and terpenoids. Various genetic tools have been developed, but neither a centromere nor an autonomously replicating sequence (ARS), both necessary elements for stable episomal plasmid maintenance, has yet been reported. In this study, cleavage under targets and release using nuclease (CUT&RUN), a method used for genome-wide mapping of DNA–protein interactions, was used to identify R. toruloides IFO0880 genomic regions associated with the centromeric histone H3 protein Cse4, a marker of centromeric DNA. Fifteen putative centromeres ranging from 8 to 19 kb in length were identified and analyzed, and four were tested for, but did not show, ARS activity. These centromeric sequences contained below average GC content, corresponded to transcriptional cold spots, were primarily nonrepetitive and shared some vestigial transposon-related sequences but otherwise did not show significant sequence conservation. Future efforts to identify an ARS in this yeast can utilize these centromeric DNA sequences to improve the stability of episomal plasmids derived from putative ARS elements.

59 BASIC BIOLOGICAL SCIENCES↗

MINE: maximally informative next experiment—toward a new GWAS experimental design and methodology

Abstract The computational methodology of Genome Wide Association Studies (GWAS) currently has several limitations: (i) the number of observations (rows) on a quantitative trait tends to be smaller than the number of single nucleotide polymorphisms (SNPs) (columns) in the design matrix; (ii) each SNP is usually modeled separately, failing to acknowledge interaction between each other (ie epistasis); (iii) there is implicit linkage disequilibrium (LD) between neighboring SNPs due to their linkage. To overcome these issues, we developed a tool that uses ensemble methods to fit mixed linear models to GWAS data, and these ensemble methods include the development of a new experimental design approach in GWAS, which uses the resultant models and data to select the next informative experiment over time. This new adaptive and staged approach for GWAS experimental design was developed and tested in a 3 yr adaptive model-guided discovery experiment against a fixed classical design. In Sorghum bicolor a total of 79, 86, and 78 accessions were tested in years 1, 2, and 3, respectively out of 343 accessions available in the Bioenergy Association Panel (BAP) each identified for 232,303 SNPs, 1 every 2–3 kb in the genomes. We demonstrated the feasibility of MINE enacted with 8 people in the field per year over 3 yr vs in 1 large classical design enacted with 20 people in 1 yr. The MINE results for chromosomal regions identified controlling dry weight were confirmed against results from previous sorghum GWAS experiments and 1 large classical design for the BAP panel.

Genetics & Heredity↗

Viroid-like “obelisk” agents are widespread in the ocean and exceed the abundance of RNA viruses in the prokaryotic fraction

Abstract “Obelisks” are recently discovered ribonucleic acid (RNA) viroid-like elements present in diverse environments with no phylogenetic similarity to any known biological agent. obelisks were first identified in the human gut and in a commensal bacterium acting as a replicative host. They have a circular ∼1 kb RNA genome, rod-like secondary structures, and the encoding of a protein superfamily called “Oblins”. We performed a large-scale search of obelisks in the ocean using the Pebblescout program and the transcriptomic Sequence Archive Read databases, revealing the biogeography and abundance of these viroid-like RNA elements. We detected 55 obelisk genomes resulting in 35 marine clusters at the species level. These obelisks were detected in the prokaryotic fraction and to a lesser extent in the eukaryotic fraction, and distributed across all the oceans from surface to mesopelagic including the Arctic, and even in the coldest seawater of Earth beneath the Antarctic Ross Ice Shelf. The obelisk hallmark protein Oblin-1 confirmed by 3D models was found in various marine samples. Some of the detected marine obelisks harbor hammerhead self-cleaving ribozymes in both polarities. In the prokaryotic, but not the eukaryotic, fraction of the Tara Ocean dataset, relative abundance of obelisks calculated by transcriptomic fragment recruitment indicated that they are abundant in marine samples, reaching or even exceeding the relative abundance of the previously discovered uncultured RNA viruses. In conclusion, obelisks are abundant and widespread viroid-like elements that should be included in ocean biogeochemical models.

Environmental Sciences & Ecology↗

Linezolid-resistant (Tn 6246 :: fexB - poxtA ) Enterococcus faecium strains colonizing humans and bovines on different continents: similarity without epidemiological link

Abstract Objectives poxtA is the most recently described gene conferring acquired resistance to linezolid, a relevant antibiotic for treating enterococcal infections. We retrospectively screened for poxtA in diverse enterococci and aimed to characterize its genetic/genomic contexts. Methods poxtA was screened by PCR in 812 enterococci from 458 samples (hospitals/healthy humans/wastewater/animals/retail food) obtained in Portugal/Angola/Tunisia (1996–2019). Antimicrobial susceptibility testing was performed for 13 antibiotics (EUCAST/CLSI). poxtA stability (∼500 generations), transfer (filter mating), clonality (SmaI-PFGE) and location (S1-PFGE/hybridization) were tested. WGS (Illumina-HiSeq) was performed for clonal representatives. Results poxtA was detected in Enterococcus faecium from six samples (1.3%): a healthy human (rectal swab) in Porto, Portugal (ST32/2001); four farm cows (milk) in Mateur, Tunisia (ST1058/2015); and a hospitalized patient (faeces) in Matosinhos, Portugal (ST1058/2015). All expressed resistance to linezolid (MIC = 8 mg/L), chloramphenicol, tetracycline and erythromycin, with variable resistance to ciprofloxacin and streptomycin. ST1058-poxtA-carrying isolates from Tunisia and Portugal differed by two SNPs and had similar plasmid content. poxtA, located in an IS1216-flanked Tn6246-like element, co-hybridized with fexB on one or more plasmids per isolate (one to three plasmids of 30–100 kb), was stable after several generations and transferred only from ST1058. ST1058 strains carried resistance/virulence genes (Efmqnr/acm) possibly induced under selective quinolone treatment. Conclusions poxtA has been circulating in Portugal since at least 2001, corresponding to the oldest description worldwide to date. We also extend the reservoir of poxtA to bovines. The similar linezolid-resistant poxtA-carrying strains colonizing humans and livestock on different continents, and without a noticeable relationship, suggests a recent transmission event or convergent evolution of E. faecium populations in different hosts and geographic regions.

Freitas, Ana R.↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗

Museomics Dissects the Genetic Basis for Adaptive Seasonal Coloration in the Least Weasel

Abstract Dissecting the link between genetic variation and adaptive phenotypes provides outstanding opportunities to understand fundamental evolutionary processes. Here, we use a museomics approach to investigate the genetic basis and evolution of winter coat coloration morphs in least weasels (Mustela nivalis), a repeated adaptation for camouflage in mammals with seasonal pelage color moults across regions with varying winter snow. Whole-genome sequence data were obtained from biological collections and mapped onto a newly assembled reference genome for the species. Sampling represented two replicate transition zones between nivalis and vulgaris coloration morphs in Europe, which typically develop white or brown winter coats, respectively. Population analyses showed that the morph distribution across transition zones is not a by-product of historical structure. Association scans linked a 200-kb genomic region to coloration morph, which was validated by genotyping museum specimens from intermorph experimental crosses. Genotyping the wild populations narrowed down the association to pigmentation gene MC1R and pinpointed a candidate amino acid change cosegregating with coloration morph. This polymorphism replaces an ancestral leucine residue by lysine at the start of the first extracellular loop of the protein in the vulgaris morph. A selective sweep signature overlapped the association region in vulgaris, suggesting that past adaptation favored winter-brown morphs and can anchor future adaptive responses to decreasing winter snow. Using biological collections as valuable resources to study natural adaptations, our study showed a new evolutionary route generating winter color variation in mammals and that seasonal camouflage can be modulated by changes at single key genes.

Miranda, Inês↗

Deeplasmid: deep learning accurately separates plasmids from bacterial chromosomes

Plasmids are mobile genetic elements that play a key role in microbial ecology and evolution by mediating horizontal transfer of important genes, such as antimicrobial resistance genes. Many microbial genomes have been sequenced by short read sequencers and have resulted in a mix of contigs that derive from plasmids or chromosomes. New tools that accurately identify plasmids are needed to elucidate new plasmid-borne genes of high biological importance. We have developed Deeplasmid, a deep learning tool for distinguishing plasmids from bacterial chromosomes based on the DNA sequence and its encoded biological data. It requires as input only assembled sequences generated by any sequencing platform and assembly algorithm and its runtime scales linearly with the number of assembled sequences. Deeplasmid achieves an AUC–ROC of over 89%, and it was more accurate than five other plasmid classification methods. Finally, as a proof of concept, we used Deeplasmid to predict new plasmids in the fish pathogen Yersinia ruckeri ATCC 29473 that has no annotated plasmids. Deeplasmid predicted with high reliability that a long assembled contig is part of a plasmid. Using long read sequencing we indeed validated the existence of a 102 kb long plasmid, demonstrating Deeplasmid's ability to detect novel plasmids.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative mitogenomics of kingdom Fungi – evolutionary insights and metagenomic applications

Mitochondria are essential components of eukaryotic cells, responsible for ATP production through oxidative phosphorylation. Despite their biological importance, unique challenges have hindered the adoption of automated mitochondrial genome (mitogenome) annotation methods, obstructing mitochondrial comparative genomics in a broad evolutionary context. Using Fungi as a study system and a Joint Genome Institute (JGI) annotated high-quality reference set, we observed broad patterns of mitochondrial evolution across the kingdom. We found that the median fungal mitogenome size is 58 kb and identified exceptionally large examples over 1 Mb in Pezizomycetes. All 14 expected oxidative phosphorylation protein-coding genes, plus rps3, were generally conserved. We found evidence of major evolutionary transitions within the Ascomycota, including the transfer of mitochondrially encoded atp8 and atp9 to the nuclear genomes across the Pezizomycotina and shifts in mitogenome tRNA patterns across the kingdom. We found substantial concordance between mitochondrial and nuclear evolution, enabling us to document 3131 total fungal mitogenomes from JGI-derived metagenomic datasets. We also identified 6467 total undeclared mitogenomes embedded in Genbank fungal nuclear assemblies. We provide interactive tools for mitogenome analysis through the JGI MycoCosm platform. Collectively, this work generated nearly 10 000 new fungal mitogenome annotations, providing a foundation and resources for future exploration of comparative fungal mitogenomics.

Ahrendt, Steven R. [USDOE Joint Genome Institute (↗

Topologically Associating Domain Boundaries are Commonly Required for Normal Genome Function

Topologically associating domain (TAD) boundaries are thought to partition the genome into distinct regulatory territories. Anecdotal evidence suggests that their disruption may interfere with normal gene expression and cause disease phenotype, but the overall extent to which this occurs remains unknown. Here we show that TAD boundary deletions commonly disrupt normal genome function in vivo . We used CRISPR genome editing in mice to individually delete eight TAD boundaries (11-80kb in size) from the genome in mice. All deletions examined resulted in at least one detectable molecular or organismal phenotype, which included altered chromatin interactions or gene expression, reduced viability, and anatomical phenotypes. For 5 of 8 (62%) loci examined, boundary deletions were associated with increased embryonic lethality or other developmental phenotypes. For example, a TAD boundary deletion near Smad3/Smad6 caused complete embryonic lethality, while a deletion near Tbx5/Lhx5 resulted in a severe lung malformation. Our findings demonstrate the importance of TAD boundary sequences for in vivo genome function and suggest that noncoding deletions affecting TAD boundaries should be carefully considered for potential pathogenicity in clinical genetics screening.

Rajderkar, Sudha↗

Enhancing EV Motor Design Through Knowledge-Based AI and Hierarchical Fuzzy Logic Model

This work presents a novel approach to optimizing electric vehicle motor design through the integration of Knowledge-Based Artificial Intelligence (KB-AI) and Hierarchical Fuzzy Logic. Traditional motor design processes are time-intensive, relying heavily on iterative simulations and domain-specific expertise. These processes are further complicated by the nonlinear relationships between key design parameters. The proposed framework addresses these challenges by systematically encoding expert knowledge from scientific literature into a fuzzy logic system, allowing for the efficient handling of complex design variables. The hierarchical fuzzy logic model reduces computational complexity by decomposing the nonlinear relationships into manageable rule sets while maintaining design accuracy. The proposed methodology was applied to the design of a 100 kW motor, yielding optimal values for key parameters. This resulted in a compact motor design with a volume of 2.2 liters, showcasing the framework’s ability to deliver high-performance, application-specific motor configurations.

Kumar, Praveen [ORNL] (ORCID:0000000291877857)↗

Silencing of Dicer‐like protein 2a restores the resistance phenotype in the rice mutant, sxi4 ( suppressor of Xa21‐mediated immunity 4 )

SUMMARY The rice immune receptor XA21 confers resistance to Xanthomonas oryzae pv. oryzae ( Xoo ), and upon recognition of the RaxX21‐sY peptide produced by Xoo , XA21 activates the plant immune response. Here we screened 21 000 mutant plants expressing XA21 to identify components involved in this response, and reported here the identification of a rice mutant, sxi4, which is susceptible to Xoo. The sxi4 mutant carries a 32‐kb translocation from chromosome 3 onto chromosome 7 and displays an elevated level of DCL2a transcript, encoding a Dicer‐like protein . Silencing of DCL2a in the sxi4 genetic background restores resistance to Xoo . RaxX21‐sY peptide‐treated leaves of sxi4 retain the hallmarks of XA21‐mediated immune response. However, WRKY45‐1 , a known negative regulator of rice resistance to Xoo , is induced in the sxi4 mutant in response to RaxX21‐sY peptide treatment. A CRISPR knockout of a short interfering RNA (TE‐siRNA815) in the intron of WRKY45‐1 restores the resistance phenotype in sxi4 . These results suggest a model where DCL2a accumulation negatively regulates XA21‐mediated immunity by altering the processing of TE‐siRNA815.

Liu, Furong↗

Sorghum Association Panel whole‐genome sequencing establishes cornerstone resource for dissecting genomic diversity

SUMMARY Association mapping panels represent foundational resources for understanding the genetic basis of phenotypic diversity and serve to advance plant breeding by exploring genetic variation across diverse accessions. We report the whole‐genome sequencing (WGS) of 400 sorghum ( Sorghum bicolor (L.) Moench) accessions from the Sorghum Association Panel (SAP) at an average coverage of 38× (25–72×), enabling the development of a high‐density genomic marker set of 43 983 694 variants including single‐nucleotide polymorphisms (approximately 38 million), insertions/deletions (indels) (approximately 5 million), and copy number variants (CNVs) (approximately 170 000). We observe slightly more deletions among indels and a much higher prevalence of deletions among CNVs compared to insertions. This new marker set enabled the identification of several novel putative genomic associations for plant height and tannin content, which were not identified when using previous lower‐density marker sets. WGS identified and scored variants in 5‐kb bins where available genotyping‐by‐sequencing (GBS) data captured no variants, with half of all bins in the genome falling into this category. The predictive ability of genomic best unbiased linear predictor (GBLUP) models was increased by an average of 30% by using WGS markers rather than GBS markers. We identified 18 selection peaks across subpopulations that formed due to evolutionary divergence during domestication, and we found six F st peaks resulting from comparisons between converted lines and breeding lines within the SAP that were distinct from the peaks associated with historic selection. This population has served and continues to serve as a significant public resource for sorghum research and demonstrates the value of improving upon existing genomic resources.

59 BASIC BIOLOGICAL SCIENCES↗

Multiplexed CRISPR-Cas9 mutagenesis of rice PSBS1 noncoding sequences for transgene-free overexpression

Understanding CRISPR-Cas9’s capacity to produce native overexpression (OX) alleles would accelerate agronomic gains achievable by gene editing. To generate OX alleles with increased RNA and protein abundance, we leveraged multiplexed CRISPR-Cas9 mutagenesis of noncoding sequences upstream of the rice PSBS1 gene. We isolated 120 gene-edited alleles with varying non-photochemical quenching (NPQ) capacity in vivo—from knockout to overexpression—using a high-throughput screening pipeline. Overexpression increased OsPsbS1 protein abundance two- to threefold, matching fold changes obtained by transgenesis. Increased PsbS protein abundance enhanced NPQ capacity and water-use efficiency. Across our resolved genetic variation, we identify the role of 5'UTR indels and inversions in driving knockout/knockdown and overexpression phenotypes, respectively. Complex structural variants, such as the 252-kb duplication/inversion generated here, evidence the potential of CRISPR-Cas9 to facilitate significant genomic changes with negligible off-target transcriptomic perturbations. Our results may inform future gene-editing strategies for hypermorphic alleles and have advanced the pursuit of gene-edited, non-transgenic rice plants with accelerated relaxation of photoprotection.

60 APPLIED LIFE SCIENCES↗

A New Pneumococcal Capsule Type, 10D, is the 100th Serotype and Has a Large cps Fragment from an Oral Streptococcus

Streptococcus pneumoniae (pneumococcus) is a major human pathogen producing structurally diverse capsular polysaccharides. Widespread use of highly successful pneumococcal conjugate vaccines (PCVs) targeting pneumococcal capsules has greatly reduced infections by the vaccine types but increased infections by nonvaccine serotypes. Herein, we report a new and the 100th capsule type, named serotype 10D, by determining its unique chemical structure and biosynthetic roles of all capsule synthesis locus (cps) genes. The name 10D reflects its serologic cross-reaction with serotype 10A and appearance of cross-opsonic antibodies in response to immunization with 10A polysaccharide in a 23-valent pneumococcal vaccine. Genetic analysis showed that 10D cps has three large regions syntenic to and highly homologous with cps loci from serotype 6C, serotype 39, and an oral streptococcus strain (S. mitis SK145). The 10D cps region syntenic to SK145 is about 6 kb and has a short gene fragment of wciNα at the 5' end. The presence of this nonfunctional wciNα fragment provides compelling evidence for a recent interspecies genetic transfer from oral streptococcus to pneumococcus. Since oral streptococci have a large repertoire of cps loci, widespread PCV usage could facilitate the appearance of novel serotypes through interspecies recombination.

59 BASIC BIOLOGICAL SCIENCES↗

Complete Genome Sequence of Sulfurospirillum Strain ACS TCE , a Tetrachloroethene-Respiring Anaerobe Isolated from Contaminated Soil

Here, we report the complete genome sequence of the tetrachloroethene-to-trichloroethene dechlorinator Sulfurospirillum sp. strain ACS TCE . The genome consists of a 38.05-kb circular plasmid and a 2.69-Mb circular chromosome, which encodes 3 identical reductive dehalogenases with 91.47% amino acid identity to the PceA of Sulfurospirillum multivorans strain DSM 12446.

59 BASIC BIOLOGICAL SCIENCES↗