Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sequence alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Structural models and functional annotations for the Sphagnum divinum proteome

This dataset contains the structural models for the primary transcripts of the Sphagnum divinum proteome. Additionally, for a subset of these proteins, sequence and structural alignment results are provided. This dataset represents the most thorough structural study of a Sphagnum species, also known as peat mosses, by providing three-dimensional atomic resolution structures of the majority of the encoded proteins as well as structural alignment results used in the application of annotating the proteome. References (DOI) AlphaFold v2 Monomer: https://doi.org/10.1038/s41586-021-03819-2. References (DOI) US-align2: https://doi.org/10.1038/s41592-022-01585-1

59 BASIC BIOLOGICAL SCIENCES↗

Metatransciptomic Analysis Data for Interactive effects of depth and differential irrigation on soil microbiome composition and functioning

RNA was collected from soil at different depths and after three different levels of irrigation T1 100% of normal field irrigation, T4: 18.75% or normal irrigation and T5: unirrigated controls. Total RNA was isolated using the Zymo Quick-RNA fecal/soil microbe miniprep (catalog no. R2040), incorporating the DNase I treatment using Zymo’s DNase I kit (catalog no. E1010). To increase the yield of RNA, we modified the manufacturer’s instructions by first doubling the amount of soil per extraction (from 0.25 g to 0.5 g) and by performing extractions in triplicate before pooling separate extractions together. Certain soil samples (largely those from deeper soil layers) had low yield (< 100 ng per extraction) so additional rounds of extraction were performed to obtain sufficient RNA. RNA concentration was assessed using a Qubit RNA HS assay kit (Thermo Fisher) and RNA quality was determined using an Agilent 2100 BioAnalyzer (Agilent; Santa Clara, CA). The resultant RNA samples were then sequenced by GENEWIZ using Illumina technology (GENEWIZ; South Plainfield, NJ). Sequences were then aligned to a soil metagenome previously obtained from the same site using the Burrows-Wheeler aligner (BWA). SAM files were then converted to raw counts using HTSeq.

Soil microbiome, metatranscriptomics↗

RCSB Protein Data Bank: visualizing groups of experimentally determined PDB structures alongside computed structure models of proteins

Recent advances in Artificial Intelligence and Machine Learning (e.g., AlphaFold, RosettaFold, and ESMFold) enable prediction of three-dimensional (3D) protein structures from amino acid sequences alone at accuracies comparable to lower-resolution experimental methods. These tools have been employed to predict structures across entire proteomes and the results of large-scale metagenomic sequence studies, yielding an exponential increase in available biomolecular 3D structural information. Given the enormous volume of this newly computed biostructure data, there is an urgent need for robust tools to manage, search, cluster, and visualize large collections of structures. Equally important is the capability to efficiently summarize and visualize metadata, biological/biochemical annotations, and structural features, particularly when working with vast numbers of protein structures of both experimental origin from the Protein Data Bank (PDB) and computationally-predicted models. Moreover, researchers require advanced visualization techniques that support interactive exploration of multiple sequences and structural alignments. This paper introduces a suite of tools provided on the RCSB PDB research-focused web portal RCSB. org, tailor-made for efficient management, search, organization, and visualization of this burgeoning corpus of 3D macromolecular structure data.

3D visualization↗

Comprehensive Genetic Characterization of Four Novel HIV-1 Circulating Recombinant Forms (CRF129_56G, CRF130_A1B, CRF131_A1B, and CRF138_cpx): Insights from Molecular Epidemiology in Cyprus

Molecular investigations of the HIV-1 pol region (2253–5250 in the HXB2 genome) were conducted on sequences obtained from 331 individuals infected with HIV-1 in Cyprus between 2017 and 2021. This study unveiled four distinct HIV-1 putative transmission clusters, encompassing 19 previously unidentified HIV-1 recombinants. These recombinants, each comprising eight, three, four, and four sequences, respectively, did not align with previously established Circulating Recombinant Forms (CRFs). To characterize these novel HIV-1 recombinants, near-full-length genome sequences were successfully obtained for 16 of the 19 recombinants (790–8795 in the HXB2 genome) using an in-house-developed RT-PCR assay. Phylogenetic analyses, employing MEGAX and Cluster-Picker, along with confirmatory neighbor-joining tree analyses of subregions, were conducted to identify distinct clusters and determine subtypes. The uniqueness of the HIV-1 recombinants was evident in their exclusive clustering within generated maximum likelihood trees. Recombination analyses highlighted the distinct chimeric nature of these recombinants, with consistent mosaic patterns observed across all sequences within each of the four putative transmission clusters. Conclusive genetic characterization identified four novel HIV-1 CRFs: CRF129_56G, CRF130_A1B, CRF131_A1B, and CRF138_cpx. CRF129_56G exhibited two recombination breakpoints and three fragments of subtypes CRF56_cpx and G. Both CRF130_A1B and CRF131_A1B featured seven recombination breakpoints and eight fragments of subtypes A1 and B. CRF138_cpx displayed five recombination breakpoints and six fragments of subtypes CRF22_01A1 and F2, along with an unclassified fragment. Additional BLAST analyses identified a Unique Recombinant Form (URF) of CRF138_cpx with three additional recombination sites, involving subtype F2, a fragment of unknown subtype origin, and CRF138_cpx. Post-identification, all putative transmission clusters remained active, with CRF130_A1B, CRF131_A1B, and CRF138_cpx clusters exhibiting further growth. Furthermore, international connections were identified through BLAST analyses, linking one sequence from the USA to the CRF130_A1B strain, and three sequences from Belgium and Cameroon to the CRF138_cpx strain. This study contributes valuable insights into the dynamic landscape of HIV-1 diversity and transmission patterns, emphasizing the need for ongoing molecular surveillance and global collaboration in tracking emerging viral variants.

60 APPLIED LIFE SCIENCES↗

When Machine Learning Meets 2D Materials: A Review

The availability of an ever-expanding portfolio of 2D materials with rich internal degrees of freedom (spin, excitonic, valley, sublattice, and layer pseudospin) together with the unique ability to tailor heterostructures made layer by layer in a precisely chosen stacking sequence and relative crystallographic alignments, offers an unprecedented platform for realizing materials by design. However, the breadth of multi-dimensional parameter space and massive data sets involved is emblematic of complex, resource-intensive experimentation, which not only challenges the current state of the art but also renders exhaustive sampling untenable. To this end, machine learning, a very powerful data-driven approach and subset of artificial intelligence, is a potential game-changer, enabling a cheaper – yet more efficient – alternative to traditional computational strategies. It is also a new paradigm for autonomous experimentation for accelerated discovery and machine-assisted design of functional 2D materials and heterostructures. Here, the study reviews the recent progress and challenges of such endeavors, and highlight various emerging opportunities in this frontier research area.

2D materials↗

A genome-scale phylogeny of the kingdom Fungi

Phylogenomic studies using genome-scale amounts of data have greatly improved understanding of the tree of life. Despite the diversity, ecological significance, and biomedical and industrial importance of fungi, evolutionary relationships among several major lineages remain poorly resolved, especially those near the base of the fungal phylogeny. To examine poorly resolved relationships and assess progress toward a genome-scale phylogeny of the fungal kingdom, we compiled a phylogenomic data matrix of 290 genes from the genomes of 1,644 species that includes representatives from most major fungal lineages. We also compiled 11 data matrices by subsampling genes or taxa from the full data matrix based on filtering criteria previously shown to improve phylogenomic inference. Analyses of these 12 data matrices using concatenation- and coalescent-based approaches yielded a robust phylogeny of the fungal kingdom, in which ~85% of internal branches were congruent across data matrices and approaches used. We found support for several historically poorly resolved relationships as well as evidence for polytomies likely stemming from episodes of ancient diversification. By examining the relative evolutionary divergence of taxonomic groups of equivalent rank, we found that fungal taxonomy is broadly aligned with both genome sequence divergence and divergence time but also identified lineages where current taxonomic circumscription does not reflect their levels of evolutionary divergence. Finally, our results provide a robust phylogenomic framework to explore the tempo and mode of fungal evolution and offer directions for future fungal phylogenetic and taxonomic studies.

59 BASIC BIOLOGICAL SCIENCES↗

A Digital Three Level Space Vector Modulator for High Frequency Vector Sequence Generation

This letter proposes a digital high-speed three-level space vector pulse width modulator (3L-SVPWM). A conventional 3L-SVPWM is typically computation-based, involving a sequential execution of sub-tasks on a digital signal processor (DSP) based controller. The resulting high computation time of 5.4 μs limits the implementation of additional control blocks for switching frequencies greater than 100 kHz. This is overcome by transforming sub-tasks into digital blocks with 1-0 decisions and simpler arithmetic operations. The sub-task blocks are executed concurrently on a programmable logic device (PLD). Hence, a fast 3L-SVPWM execution in 140 ns is achieved. The proposed digital 3L-SVPWM enables high switching frequency operation of wide bandgap (WBG) device-based 3 L inverters to generate high fundamental frequency waveforms. A finite state machine is an integral part of the proposed implementation with the ability to generate any vector sequence, maximizing the usage of redundant vector states in 3L-SVPWM. Here, the proposed digital 3L-SVPWM operation is demonstrated with a GaN-based 3 L active neutral point clamped (3L-ANPC) inverter. Experimental results are presented at 250 kHz switching frequency to generate vector sequences for center-aligned SVPWM (CA-SVPWM) and common mode voltage reduced SVPWM (CMVR-SVPWM). The results also showcase a high fundamental frequency generation capability of 10 kHz.

active neutral point clamped inverter↗

Genesearch: a Gene Homology Search Service (Genesearch) v1.0

This is a web service that allows users to query a set of genome databases to find matches for specific genes. Genesearch provides endpoints for interactive and batch homology searches that identify alignments between a query sequence and the sequences in a genomics database.

Johnson, Jeffrey↗

Single-nuclei transcriptome analysis of IgM+ cells isolated from channel catfish (Ictalurus punctatus) spleen

Catfish production is the primary aquaculture sector in the United States, and the key cultured species is channel catfish (Ictalurus punctatus). The major causes of production losses are pathogenic diseases, and the spleen, an important site of adaptive immunity, is implicated in these diseases. To examine the channel catfish immune system, single-nuclei transcriptomes of sorted and captured IgM + cells were produced from adult channel catfish. Three channel catfish (~1 kg) were euthanized, the spleen dissected, and the tissue dissociated. The lymphocytes were isolated using a Ficoll gradient and IgM + cells were then sorted with flow cytometry. The IgM + cells were lysed and single-nuclei libraries generated using a Chromium Next GEM Single Cell 3’ GEM Kit and the Chromium X Instrument (10x Genomics) and sequenced with the Illumina NovaSeq X Plus sequencer. The reads were aligned to theI. punctatusreference assembly (Coco_2.0) using Cell Ranger, and normalization, cluster analysis, and differential gene expression analysis were carried out with Seurat. Across the three samples, approximately 753.5 million reads were generated for 18,686 cells. After filtering, 10,637 cells remained for the cluster analysis. The cluster analysis identified 16 clusters which were classified as B cells (10,276), natural killer-like (NK-like) cells (178), T cells or natural killer cells (45), hematopoietic stem and progenitor cells (HSPC)/megakaryocytes (MK) (66), myeloid/epithelial cells (40), and plasma cells (32). The B cell clusters were further defined as different populations of mature B cells, cycling B cells, and plasma cells. The plasma cells highly expressedighmand we demonstrated that the secreted form of the transcript was largely being expressed by these cells. This atlas provides insight into the gene expression of IgM + immune cells in channel catfish. The atlas is publicly available and could be used garner more important information regarding the gene expression of splenic immune cells.

Immunology↗

Discovery of photosynthesis genes through whole-genome sequencing of acetate-requiring mutants of Chlamydomonas reinhardtii

Large-scale mutant libraries have been indispensable for genetic studies, and the development of next-generation genome sequencing technologies has greatly advanced efforts to analyze mutants. In this work, we sequenced the genomes of 660 Chlamydomonas reinhardtii acetate-requiring mutants, part of a larger photosynthesis mutant collection previously generated by insertional mutagenesis with a linearized plasmid. We identified 554 insertion events from 509 mutants by mapping the plasmid insertion sites through paired-end sequences, in which one end aligned to the plasmid and the other to a chromosomal location. Nearly all (96%) of the events were associated with deletions, duplications, or more complex rearrangements of genomic DNA at the sites of plasmid insertion, and together with deletions that were unassociated with a plasmid insertion, 1470 genes were identified to be affected. Functional annotations of these genes were enriched in those related to photosynthesis, signaling, and tetrapyrrole synthesis as would be expected from a library enriched for photosynthesis mutants. Systematic manual analysis of the disrupted genes for each mutant generated a list of 253 higher-confidence candidate photosynthesis genes, and we experimentally validated two genes that are essential for photoautotrophic growth, CrLPA3 and CrPSBP4 . The inventory of candidate genes includes 53 genes from a phylogenomically defined set of conserved genes in green algae and plants. Altogether, 70 candidate genes encode proteins with previously characterized functions in photosynthesis in Chlamydomonas , land plants, and/or cyanobacteria; 14 genes encode proteins previously shown to have functions unrelated to photosynthesis. Among the remaining 169 uncharacterized genes, 38 genes encode proteins without any functional annotation, signifying that our results connect a function related to photosynthesis to these previously unknown proteins. This mutant library, with genome sequences that reveal the molecular extent of the chromosomal lesions and resulting higher-confidence candidate genes, will aid in advancing gene discovery and protein functional analysis in photosynthesis.

59 BASIC BIOLOGICAL SCIENCES↗

Haploidy and aneuploidy in switchgrass mediated by misexpression of CENH3

Cross bred species such as switchgrass may benefit from advantageous breeding strategies requiring inbred lines. Doubled haploid production methods offer several ways that these lines can be produced that often involve uniparental genome elimination as the rate limiting step. We have used a centromere-mediated genome elimination strategy in which modified CENH3 is expressed to induce the process. Transgenic tetraploid switchgrass lines coexpressed Cas9, a poly-cistronic tRNA-gRNA tandem array containing eight guide RNAs that target two CENH3 genes, and different chimeric versions of CENH3 with alterations to the N-terminal tail region. Genotyping of CENH3 genes in transgenics identified edits including frameshift mutations and deletions in one or both copies of the two CENH3 genes. Flow cytometry of T 1 seedlings identified two T 0 lines that produced five haploid individuals representing an induction rate of 0.5% and 1.4%. Eight different T0 lines produced aneuploids at rates ranging from 2.1 to 14.6%. A sample of aneuploid lines were sequenced at low coverage and aligned to the reference genome, revealing missing chromosomes and chromosome arms.

59 BASIC BIOLOGICAL SCIENCES↗

Data-driven prediction of geometry- and toolpath sequence-dependent intra-layer process conditions variations in laser powder bed fusion

Geometrical features and toolpath sequence are two important factors that cause process condition variations, such as variations in the meltpool temperature or meltpool size, that might lead to undesired material properties in the laser powder bed fusion (LPBF) process. Due to the high dynamics and complex physics of the LPBF process, it is difficult to predict variations in process conditions with simulations alone. Advances in measurement technology and computational technologies open up new possibilities for smart manufacturing. In this paper, a data-driven method to predict intra-layer variations in the processing conditions that source from the toolpath sequence and part geometry is presented. The approach is demonstrated using two-color on-axis pyrometer measurements. Three demonstration cases are presented in which it is demonstrated (1) how the trained predictive model can be used as a filter to ease the interpretation of process variations and discover patterns related to toolpath and part geometry, and (2) how to generate predictions that can be used for feedforward control, i.e., for adjusting laser power or scanning speed along the toolpath using a meltpool temperature prediction model generated based on on-axis measurements. Results show that the developed prediction model is able to meaningfully predict process variations resulted from toolpath sequence and geometry. Predictions are aligned with the results from the related work of others and for the case of 180° laser path turnarounds in our high-speed X-ray imaging experiments. In conclusion, the potential issues related to the current maturity status of the process and measuring equipment that could in practice affect the performance of the proposed solutions are also discussed.

42 ENGINEERING↗

Single-nuclei transcriptome analysis of channel catfish spleen provides insight into the immunome of an aquaculture-relevant species

The catfish industry is the largest sector of U.S. aquaculture production. Given its role in food production, the catfish immune response to industry-relevant pathogens has been extensively studied and has provided crucial information on innate and adaptive immune function during disease progression. To further examine the channel catfish immune system, we performed single-cell RNA sequencing on nuclei isolated from whole spleens, a major lymphoid organ in teleost fish. Libraries were prepared using the 10X Genomics Chromium X with the Next GEM Single Cell 3’ reagents and sequenced on an Illumina sequencer. Each demultiplexed sample was aligned to the Coco_2.0 channel catfish reference assembly, filtered, and counted to generate feature-barcode matrices. From whole spleen samples, outputs were analyzed both individually and as an integrated dataset. The three splenic transcriptome libraries generated an average of 278,717,872 reads from a mean 8,157 cells. The integrated data included 19,613 cells, counts for 20,121 genes, with a median 665 genes/cell. Cluster analysis of all cells identified 17 clusters which were classified as erythroid, hematopoietic stem cells, B cells, T cells, myeloid cells, and endothelial cells. Subcluster analysis was carried out on the immune cell populations. Here, distinct subclusters such as immature B cells, mature B cells, plasma cells, γδ T cells, dendritic cells, and macrophages were further identified. Differential gene expression analyses allowed for the identification of the most highly expressed genes for each cluster and subcluster. This dataset is a rich cellular gene expression resource for investigation of the channel catfish and teleost splenic immunome.

Science & Technology - Other Topics↗

Primordial Black Hole Triggered Type Ia Supernovae. I. Impact on Explosion Dynamics and Light Curves

Primordial black holes (PBHs) in the asteroid-mass window are compelling dark matter candidates, made plausible by the existence of black holes and by the variety of mechanisms of their production in the early Universe. If a PBH falls into a white dwarf (WD), the strong tidal forces can generate enough heat to trigger a thermonuclear runaway explosion, depending on the WD’s mass and the PBH’s orbital parameters. In this work, we investigate the WD explosion triggered by the passage of a PBH. We perform 2D simulations of the WD undergoing thermonuclear explosion in this scenario, with the predicted ignition site as a parameter assuming the deflagration–detonation transition model. We study the explosion dynamics, and predict the associated light curves and nucleosynthesis. We find that the model sequence predicts light curves which align with the Phillips relation (B max versus ΔM 15 ). Our models hint at a unifying approach in triggering Type Ia supernovae without involving two distinctive evolutionary tracks.

Dark matter↗

A Systematic Framework for Tuning Open-Source Multifunctional IBR Models To Emulate OEM Black-Box Fault Dynamics

This paper presents a systematic framework to tune a generic IBR EMT model to match with an OEM provided balckbox inverter model based on the fault current responses. The key learnings and findings are summarized as follows: The tunable key parameters include inner control loops and current limiters to align the fault current magnitude, sequence content, and phase trajectories with the OEM models across diverse fault type and locations. The tuned model's fidelity is validated through comparative analysis with an OEM blackbox model, assessing both the fault current response and the responses of multiple relay elements. The results demonstrate the tuned generic model can trigger relay decision logic that is identical or near identical to that of the OEM model, thus generating very good match model for fault studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Expanding the Scope of Genomic Security: Targeted Genome Editing within Microbiomes through Designer Bacteriophage Vectors

The ability to engineer the genome of a bacterial strain, not as an isolate, but while present among other microbes in a microbiome, would open new technological possibilities in the areas of medicine, energy and biomanufacturing. Our approach is to develop sets of phages (bacterial viruses) active on the target strain and themselves engineered to act not as killers but as vectors for gene delivery. This approach is rooted in our bioinformatic tools that map prophages accurately within bacterial genomes. We present new bioinformatic results in cross-contig search, design of phage genome assemblies, satellites that embed within prophages, alignment of large numbers of biological sequences, and improvement of reference databases for prophage discovery. We targeted a Pseudomonas putida strain within a lignin-degrading microbiome, but were unable to obtain active phages, and turned toward a defined microbiome of the mouse gut.

59 BASIC BIOLOGICAL SCIENCES↗

SSGUI v1.0

SSGUI is a web-based application that integrates the integrative genomics browser (IGV) with upstream alignment pipelines enabling rapid analysis of a batch of next-generation sequencing (NGS) samples. The input to SSGUI is an NGS file system directory that is organized by experiment and reference sequence. The output is an online dashboard containing various sequencing statistics for each sample and an integrated IGV plugin enabling rapid analysis of aligned NGS reads.

Kulawik, Mark↗

Distributed Berkeley Efficient Long-Read to Long-Read Aligner and Overlapper (DiBELLA) v1.0.0

We present a parallel algorithm and scalable implementation for genome analysis, specifically the problem of finding overlaps and alignments for data from "third generation" long read sequencers. While long sequences of DNA offer enormous advantages for biological analysis and insight, current long read sequencing instruments have high error rates and therefore require different approaches to analysis than their short read counterparts. Our work focuses on an efficient distributed-memory parallelization of an accurate single-node algorithm for overlapping and aligning long reads. We achieve scalability of this irregular algorithm by addressing the competing issues of increasing parallelism, minimizing communication, constraining the memory footprint, and ensuring good load balance. The resulting application, DiBELLA, is the first distributed memory overlapper and aligner specifically designed for long reads and parallel scalability.

Ellis, Marquita↗