Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genome sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat↗

Data for Engineering and Evolution of Yarrowia lipolytica for Producing Lipids from Lignocellulosic Hydrolysates

Yarrowia lipolytica , an oleaginous yeast, shows promise for industrial fermentation due to its robust acetyl-CoA flux and well-developed genetic engineering tools. However, its lack of an active xylose metabolism restricts the conversion of cellulosic sugars to valuable products. To address this, metabolic engineering, and adaptive laboratory evolution (ALE) were applied to the Y. lipolytica PO1f strain, resulting in an efficient xylose-assimilating strain (XEV). Whole-genome sequencing (WGS) of the XEV followed by reverse engineering revealed that the amplification of the heterologous oxidoreductase pathway and a mutation in the GTPase-activating protein gene (YALI0B12100g) might be the primary reasons for improved xylose assimilation in the XEV strain. When a sorghum hydrolysate was used, the XEV strain showed superior xylose consumption and lipid production compared to its parental strain (X123). This study advances our understanding of xylose metabolism in Y. lipolytica and proposes effective metabolic engineering strategies for optimizing lignocellulosic hydrolysates.

Hydrolysate↗

Data for Promoter Deletion in the Soybean Compact Mutant Leads to Overexpression of a Gene with Homology to the C20-Gibberellin 2-Oxidase Family

Height is a critical component of plant architecture, significantly affecting crop yield. The genetic basis of this trait in soybean remains unclear. In this study, we report the characterization of the Compact mutant of soybean, which has short internodes. The candidate gene was mapped to chromosome 17, and the interval containing the causative mutation was further delineated using biparental mapping. Whole-genome sequencing of the mutant revealed an 8.7 kb deletion in the promoter of the Glyma.17g145200 gene, which encodes a member of the class III gibberellin (GA) 2-oxidases. The mutation has a dominant effect, likely via increased expression of the GA 2-oxidase transcript observed in green tissue, as a result of the deletion in the promoter of Glyma.17g145200. We further demonstrate that levels of GA precursors are altered in the Compact mutant, supporting a role in GA metabolism, and that the mutant phenotype can be rescued with exogenous GA3. We also determined that overexpression of Glyma.17g145200 in Arabidopsis results in dwarfed plants. Thus, gain of promoter activity in the Compact mutant leads to a short internode phenotype in soybean through altered metabolism of gibberellin precursors. These results provide an example of how structural variation can control an important crop trait and a role for Glyma.17g145200 in soybean architecture, with potential implications for increasing crop yield.

Biomass Analytics↗

Expanded genome and proteome reallocation in a novel, robust Bacillus coagulans strain capable of utilizing pentose and hexose sugars

Bacillus coagulans, a Gram-positive thermophilic bacterium, is recognized for its probiotic properties and recent development as a microbial cell factory. Despite its importance for biotechnological applications, the current understanding of B. coagulans’ robustness is limited, especially for undomesticated strains. To fill this knowledge gap, we characterized the metabolic capability and performed functional genomics and systems analysis of a novel, robust strain, B. coagulans B-768. Genome sequencing revealed that B-768 has the largest B. coagulans genome known to date (3.94 Mbp), about 0.63 Mbp larger than the average genome of sequenced B. coagulans strains, with expanded carbohydrate metabolism and mobilome. Functional genomics identified a well-equipped genetic portfolio for utilizing a wide range of C5 (xylose, arabinose), C6 (glucose, mannose, galactose), and C12 (cellobiose) sugars present in biomass hydrolysates, which was validated experimentally. For growth on individual xylose and glucose, the dominant sugars in biomass hydrolysates, B-768 exhibited distinct phenotypes and proteome profiles. Faster growth and glucose uptake rates resulted in lactate overflow metabolism, which makes B. coagulans a lactate overproducer; however, slower growth and xylose uptake diminished overflow metabolism due to the high energy demand for sugar assimilation. Carbohydrate Transport and Metabolism (COG-G), Translation (COG-J), and Energy Conversion and Production (COG-C) made up 60%–65% of the measured proteomes but were allocated differently when growing on xylose and glucose. The trade-off in proteome reallocation, with high investment in COG-C over COG-G, explains the xylose growth phenotype with significant upregulation of xylose metabolism, pyruvate metabolism, and tricarboxylic acid (TCA) cycle. Strain B-768 tolerates and effectively utilizes inhibitory biomass hydrolysates containing mixed sugars and exhibits hierarchical sugar utilization with glucose as the preferential substrate.

carbohydrate metabolism↗

Poplar

SAND2025-00683O Poplar is a software tool that generates a phylogenetic tree from input gene and genome sequences. It integrates established tools to identify genes within genomes, group sequences, construct gene trees, and infer a species tree. Poplar processes nucleotide sequences, identifies similar sequences using Nucleotide BLAST, groups them with DBSCAN, aligns sequences with MAFFT, constructs gene trees with RAxML-NG, and infers a species tree using ASTRAL-Pro3. This pipeline provides a structured approach to phylogenetic analysis, facilitating the study of evolutionary relationships among species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Krishnakumar, Raga↗

Unveiling the Arsenal of Apple Bitter Rot Fungi: Comparative Genomics Identifies Candidate Effectors, CAZymes, and Biosynthetic Gene Clusters in Colletotrichum Species

The bitter rot of apple is caused by Colletotrichum spp. and is a serious pre-harvest disease that can manifest in postharvest losses on harvested fruit. In this study, we obtained genome sequences from four different species, C. chrysophilum, C. noveboracense, C. nupharicola, and C. fioriniae, that infect apple and cause diseases on other fruits, vegetables, and flowers. Our genomic data were obtained from isolates/species that have not yet been sequenced and represent geographic-specific regions. Genome sequencing allowed for the construction of phylogenetic trees, which corroborated the overall concordance observed in prior MLST studies. Bioinformatic pipelines were used to discover CAZyme, effector, and secondary metabolic (SM) gene clusters in all nine Colletotrichum isolates. We found redundancy and a high level of similarity across species regarding CAZyme classes and predicted cytoplastic and apoplastic effectors. SM gene clusters displayed the most diversity in type and the most common cluster was one that encodes genes involved in the production of alternapyrone. Our study provides a solid platform to identify targets for functional studies that underpin pathogenicity, virulence, and/or quiescence that can be targeted for the development of new control strategies. With these new genomics resources, exploration via omics-based technologies using these isolates will help ascertain the biological underpinnings of their widespread success and observed geographic dominance in specific areas throughout the country.

59 BASIC BIOLOGICAL SCIENCES↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Laminarin stimulates single cell rates of sulfate reduction whereas oxygen inhibits transcriptomic activity in coastal marine sediment

Abstract The chemical cycles carried out by bacteria and archaea living in coastal sediments are vital aspects of benthic ecology. These ecosystems are subject to physical disruption, which may allow for increased respiration and complex carbon consumption—impacting chemical cycling in this environment often thought to be a terminal place of deposition. We use the redox-enzyme sensitive probe RedoxSensor Green to measure rates of electron transfer physiology in individual sulfate reducer cells residing in anoxic sediment, subjected to transient exposure of oxygen and laminarin. We use index fluorescence activated cell sorting and single cell genomics sequencing to link those measurements to genomes of respiring cells. We measure per-cell sulfate reduction rates in marine sediments (0.01–4.7 fmol SO42− cell−1 h−1) and determine that cells within the Chloroflexota phylum are the most active in respiration. Chloroflexota respiration activity is also stimulated with the addition of laminarin, even in marine sediments already rich in organic matter. Evaluating metatranscriptomic data alongside this respiration-based technique, Chloroflexota genomes encode laminarinases indicating a likely ability to degrade laminarin. We also provide evidence that abundant Patescibacteria cells do not use electron transport pathways for energy, and instead likely carry out fermentation of polysaccharides. There is a decoupling of respiration-related activity rates from transcription, as respiration rates increase while transcription decreases with oxygen exposure. Overall, we reveal an active community of respiring Chloroflexota that cycles sulfate at potential rates of 23–40 nmol h−1 per cm3 sediment in incubation settings, and non-respiratory Patescibacteria that can cycle complex polysaccharides.

Lindsay, Melody R.↗

Elucidation of odd-chain dicarboxylate metabolism in Acinetobacter baylyi and application to polyethylene upcycling

Polyethylene (PE) is a versatile polymer, but its end-of-life management is challenging due to its recalcitrant structure. We present a promising approach combining chemical degradation and bio-upcycling to convert postconsumer PE waste into a value-added bioproduct. Specifically, PE was degraded into acetic acid and C 4 –C 7 dicarboxylic acids by nitric acid. We then elucidated the catabolic pathways for glutarate (C 5 ) and pimelate (C 7 ) in the nonmodel bacterium Acinetobacter baylyi ADP1 through RNA sequencing, phenotyping, and enzymatic assays. Whole-genome sequencing of evolved isolates also identified a crucial IclR family transcriptional regulator, DcaS, which acts as a repressor of dicarboxylate metabolism. The reverse-engineered strain exhibited enhanced substrate utilization compared to the wild-type strain. Using rational metabolic engineering, the PE deconstruction products were bioconverted into the valuable chemical lycopene, highlighting the potential of this microbial chassis to produce value-added bioproducts from postconsumer PE waste, thus promoting a circular economy for plastics.

metabolic engineering↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Engineering and evolution of Yarrowia lipolytica for producing lipids from lignocellulosic hydrolysates

Yarrowia lipolytica, an oleaginous yeast, shows promise for industrial fermentation due to its robust acetyl-CoA flux and well-developed genetic engineering tools. However, its lack of an active xylose metabolism restricts the conversion of cellulosic sugars to valuable products. To address this, metabolic engineering, and adaptive laboratory evolution (ALE) were applied to the Y. lipolytica PO1f strain, resulting in an efficient xylose-assimilating strain (XEV). Whole-genome sequencing (WGS) of the XEV followed by reverse engineering revealed that the amplification of the heterologous oxidoreductase pathway and a mutation in the GTPase-activating protein gene (YALI0B12100g) might be the primary reasons for improved xylose assimilation in the XEV strain. When a sorghum hydrolysate was used, the XEV strain showed superior xylose consumption and lipid production compared to its parental strain (X123). This study advances our understanding of xylose metabolism in Y. lipolytica and proposes effective metabolic engineering strategies for optimizing lignocellulosic hydrolysates.

60 APPLIED LIFE SCIENCES↗

Ocelot: An Interactive, Efficient Distributed Compression-As-a-Service Platform With Optimized Data Compression Techniques

Large volumes of data generated by scientific simulations, genome sequencing, and other applications need to be moved among clusters for data collection/analysis. Data compression techniques have effectively reduced data storage and transfer costs. However, users' requirements on interactively controlling both data quality and compression ratios are non-trivial to fulfill. Here, we propose a novel Compression-as-a-Service (CaaS) platform called Ocelot with four important contributions: (1) It offers real-time visualization, interactive compression, and transfer of scientific datasets. (2) It incorporates new strategies for compressing diverse types of datasets more effectively than traditional methods. (3) It provides an effective method for estimating the compression ratio and execution time of compression tasks. (4) Experiments on multiple real-world datasets on geographically distributed computers show that Ocelot can significantly improve data transfer efficiency with a performance gain of more than 10x in computing clusters with relatively slow networks.

compression as a service (CaaS)↗

Fine-scale contemporary recombination variation and its fitness consequences in adaptively diverging stickleback fish

Despite deep evolutionary conservation, recombination rates vary greatly across the genome and among individuals, sexes and populations. Yet the impact of this variation on adaptively diverging populations is not well understood. Here we characterized fine-scale recombination landscapes in an adaptively divergent pair of marine and freshwater populations of threespine stickleback from River Tyne, Scotland. Through whole-genome sequencing of large nuclear families, we identified the genomic locations of almost 50,000 crossovers and built recombination maps for marine, freshwater and hybrid individuals at a resolution of 3.8 kb. We used these maps to quantify the factors driving variation in recombination rates. We found strong heterochiasmy between sexes but also differences in recombination rates among ecotypes. Hybrids showed evidence of significant recombination suppression in overall map length and in individual loci. Recombination rates were lower not only within individual marine–freshwater-adaptive loci, but also between loci on the same chromosome, suggesting selection on linked gene ‘cassettes’. Through temporal sampling along a natural hybrid zone, we found that recombinants showed traits associated with reduced fitness. Our results support predictions that divergence in cis-acting recombination modifiers, whose functions are disrupted in hybrids, may play an important role in maintaining differences among adaptively diverging populations.

59 BASIC BIOLOGICAL SCIENCES↗

Aeromonas in South Asia: genomic insights into an environmental pathogen and reservoir of antimicrobial resistance

Aeromonads are an ecologically versatile group of bacteria that cause infections in aquatic animals and are recognised as emerging human pathogens. Despite this, our understanding of Aeromonas diversity, especially the relationship between clinical and environmental strains, remains limited. Here, we present a genomic analysis of the Aeromonas genus, comprising 1853 genomes, and a detailed comparison of clinical and environmental strains from South Asia, including 996 newly sequenced genomes from Bangladesh and India. Phylogenetic analyses revealed that Aeromonas is a highly diverse genus, with no distinct clade separating clinical and environmental isolates. We identified 28 Aeromonas species and 905 novel sequence types, comprising 72.5% of the genomes. Notably, we show a high incidence of antimicrobial resistance (AMR) genes across all isolates, including against front and last-line antibiotics. Finally, we highlight frequent misidentification of Aeromonas as Vibrio cholerae, which is relevant to cholera-endemic regions where both genera co-exist and are associated with diarrhoeal disease. Our study underscores Aeromonas as an important environmental AMR reservoir and emerging multi-species pathogen capable of spilling over into human populations.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient mutagenesis and genotyping of maize inbreds using biolistics, multiplex CRISPR/Cas9 editing, and Indel-Selective PCR

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Efficient Mutagenesis and Genotyping of Maize Inbreds Using Biolistics, Multiplex CRISPR/Cas9 Editing, and Indel-Selective PCR"

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

gene editing↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗