Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Base Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Performance Evaluation of a Novel Sequence-Based Directional Detection Strategy for Protection of Active Distribution Networks

Directional elements are relied on to achieve selectivity in fault detection in power systems. Although such elements have been deployed successfully for many years, there is an increased need for novel methods to deal with the unique challenges of directional protection in modern distribution networks. This article analyzes the impact of inverter-based resources (IBRs) on existing directional protection methods in distribution systems. It identifies parts of such elements that pose a risk of misoperation when IBRs are used in distribution networks. The authors have developed a new directional detection method for unbalanced faults in such networks using superimposed symmetrical sequence quantities. The phase angle of the superimposed negative sequence admittance is used to determine fault direction. The paper also presents a real-time co-simulation platform between a simulated distribution system and physical protection relay, using OPAL-RT. An SEL-411L relay is used to program the detection algorithm. This hardware-in-the-loop (HIL) setup is used to verify the performance of the method and the results are compared with existing directional methods

24 POWER TRANSMISSION AND DISTRIBUTION

Sequence-based generative AI design of versatile tryptophan synthases

Enzymes are powerful and sustainable catalysts, but their widespread application is limited by the difficulty of identifying functional starting points for optimization, creating a major bottleneck in early- stage biocatalyst discovery. Designing libraries of such starting enzymes remains particularly challenging. Here, we use the GenSLM protein language model to generate novel β-subunit of tryptophan synthase (TrpB) enzymes that express in Escherichia coli and are both stable and catalytically active. Many generated TrpBs also display significant substrate promiscuity, outperforming their natural counterparts on non-native substrates. Some even surpass laboratory-evolved TrpBs. Comparison of the most-active and most-promiscuous generated TrpB to its closest natural homolog confirms that the enhanced versatility is absent from the natural enzyme, highlighting the creative potential of generative models. These results demonstrate that the generated TrpBs not only preserve natural structure and function but also acquire non-natural properties, establishing generative models as powerful tools for biocatalyst discovery and engineering.

biocatalysis

Sequence-Based Anomaly Detection in Critical Infrastructure Networks

United States critical infrastructure faces new cyber threats from adversarial nation-state actors in the form of malware-free attacks. Traditional cybersecurity techniques use rules-based methods to identify indicators of compromise on networks, often missing these sophisticated attacks. Our approach leverages multiple state of the art machine learning models in a pipeline to identify abnormal network events through sequential analysis. We combine both device and packet-level information into individual events to characterize anomalous network actions. The model is trained and tested on real network traffic from the Idaho National Lab High Performance Computing (HPC) with greater than 98% precision. It is capable of flagging malicious tactics used by adversaries in malware-free attacks, severe changes to the network, and abnormal user activity by network devices.

99 - GENERAL AND MISCELLANEOUS

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan

Comparison of Sequence Component-Based Fault Detection and Relay Coordination Algorithms in Inverter-Based Networks

Protection of inverter-based microgrids using sequence component-based relaying schemes is a promising solution. These methods offer several advantages, including lower computational requirements, compatibility with commercial relay systems, and cost-effectiveness compared to communication-based approaches. This article investigate the performance of various sequence component based schemes with the objective of identifying the algorithms that provide the best fault detection and relay coordination, solely relying on local voltages and current at relay terminals. Positive, negative and zero sequence impedance, admittance and power detection algorithms were tested on modified IEEE 13 bus test network for various shunt faults (LG, LL, LLG, LLL). Hardware-in-the-loop validation was achieved using the Typhoon real-time simulator, interfacing with a SEL 751 relay. This research demonstrates that while several algorithms are capable of detecting faults with sufficient accuracy, only a few are effective in achieving proper coordination. Validation results indicate that the negative sequence power approach provides the best performance in both fault detection and coordination.

Patel, Deepika [ORNL] (ORCID:0000000341099994)

A 1-year study on SARS-CoV-2 variant shifts in wastewater using dPCR: comparison with clinical and GISAID data

Wastewater testing can be used to monitor SARS-CoV-2 infections in communities. Data from PCR-based wastewater testing are usually available to public health authorities within 5–7 days after excreta and other body fluids enter the sewer. While PCR-based methods can accurately detect and quantify SARS-CoV-2, sequencing-based methods are usually required to distinguish between variants, delaying the results and adding cost to the process. We developed and assessed a novel, customizable digital PCR (dPCR)-based genotyping method for SARS-CoV-2 variant detection in wastewater, which is more cost-effective, faster, and more accessible than sequencing. This approach was applied to more than 1,400 wastewater samples

Wilton, Rose

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES

Viromics approaches for the study of viral diversity and ecology in microbiomes

Viruses are found across all ecosystems and infect every type of organism on Earth. Traditional culture-based methods have proven insufficient to explore this viral diversity at scale, driving the development of viromics, the sequence-based analysis of uncultivated viruses. Viromics approaches have been particularly useful for studying viruses of microorganisms, which can act as crucial regulators of microbiomes across ecosystems. They have already revealed the broad geographic distribution of viral communities and are progressively uncovering the expansive genetic and functional diversity of the global virome. Moving forward, large-scale viral ecogenomics studies combined with new experimental and computational approaches to identify virus activity and host interactions will enable a more complete characterization of global viral diversity and its effects.

Ecology

wavess 1.2: presenting an HLA-aware within-host virus sequence simulation framework

Motivation Understanding how virus sequences are shaped by selection can inform vaccine design and transmission inference. Modeling within-host evolution to interrogate these questions requires a detailed mechanistic framework that accurately captures sequence diversification. The CD8 + cytotoxic T-lymphocyte (CTL) response plays an important role in immune-mediated selection and can leave strong signatures in virus sequences; however, existing sequence-based within-host virus modeling frameworks do not explicitly include a human leukocyte antigen (HLA)-aware CTL response. Results We extended our previously published within-host sequence evolution simulator, wavess, to include an explicit CTL response, and share a method for identifying HLA-specific CTL epitopes given a founder virus sequence. We also updated the model to permit a variable recombination rate, which allows for modeling non-adjacent genes, segmented genomes, and recombination hotspots. These extensions to wavess allow for more accurate simulation of viruses and virus genes, particularly in regions of the genome where the immune response is dominated by CTLs (rather than antibodies). It also provides the foundation for investigations of how these newly-added biological mechanisms influence within-host evolution. Availability and implementation The core of wavess is written in Python 3, with helper functions written in R. It is available at https://github.com/MolEvolEpid/wavess.

60 APPLIED LIFE SCIENCES

Geologic Characterization of the South Georgia Rift Basin for Source Proximal CO2 Storage

The project Geologic Characterization of the South Georgia Rift Basin for Source Proximal CO2 Storage is one of 9 site characterization projects that were implemented as part of ARRA (American Recovery and Reinvestment Act). Data from this project was used to improve resolution of data in NATCARB in the area of study. Data related to this study has already been incorporated in NATCARB Atlas. The South Carolina Research Foundation and partners evaluated the feasibility of CCS in the Jurassic/ Triassic (J / TR) saline formations of the buried Mesozoic South Georgia Rift (SGR) Basin that extends from South Carolina into Georgia. The J / TR sequence, based on preliminary assessment of limited geologic and geophysical data, appears to have both the appropriate areal extent and multiple horizons to permanently and safely store CO2 The presence of several igneous rock layers within the sequence may potentially provide adequate seals to prevent upward CO2 migration into the Coastal Plain aquifer systems. Approximately 81 kilometers of 2-D seismic reflection data were collected by Bay Geophysical, Inc. to explore a portion of the SGR located in southern Georgia. The 81 kilometers were divided into two lines approximately 40.5 kilometers each, with Line 1 intersecting Georgia well GGS 3457. Line 2 intersects Line 1 at the southern portion of Line 1 to maximize the extent of coverage away from GGS-3457 (a deep well drilled in the 1980s for oil and gas exploration). This well had a set of usable logs, including gamma and neutron logs that provided promising results related to CO2 storage. Results showed sandstone with porosity values greater than 10 percent and a thickness of 120 meters. The design of the seismic shot was to extrapolate information away from the well and to better define the extent of the SGR and the necessary reservoir and caprock for a successful CO2 injection. A numerical simulation model of CO2 Injection and migration was developed based on the geology log for the GGS-3457 well. The simulation model was used to investigate the feasibility of injecting 30 million metric tons of CO2 into SGR J / TA sediments and integrity of the diabase layers as seals to prevent CO2 migration.

2-D seismic

Biomass yields, reproductive fertility, compositional analysis, and genetic diversity of newly developed triploid giant miscanthus hybrids

Abstract Miscanthus × giganteus (giant miscanthus), first found as a naturally occurring hybrid, has shown promise as a bioenergy/biomass crop throughout much of the temperate world. This allotriploid (2 n = 3 x = 57) hybrid resulted from a cross between tetraploid Miscanthus sacchariflorus (2 n = 4 x = 76) and diploid Miscanthus sinensis (2 n = 2 x = 38) and is particularly desirable due to its low fertility that minimizes reseeding and potential invasiveness. However, there is limited genetic diversity in commonly grown cultivars of triploid M. × giganteus and breeding and development efforts to improve and domesticate this crop have been minimal. Here, we report on newly developed M. × giganteus hybrids compared with the industry standard M. × giganteus '1993‐1780'. Dry biomass yields of new hybrids ranged from 19.5 to 32.4 Mg/ha/year for the fourth growing season, compared with 21.0 Mg/ha/year for M. × giganteus '1993‐1780'. Plant reproductive fertility remained low for all accessions with overall fertility [(seed set × seed germination)/100] ranging from 0.3% to 4.5% for new hybrids compared to 0.4% for M. × giganteus '1993‐1780'. Culm density and height varied among accessions and were positively correlated with increased biomass. Based on compositional analyses, theoretical ethanol yields ranged from 9, 740 to 16,278 L/ha/year for new hybrids compared to 10,406 L/ha/year for M. × giganteus '1993‐1780'. Relative feed value indices were low overall and ranged between 66.0 and 72.8 for new hybrids compared to M. × giganteus '1993‐1780' with 71.3. The genetic diversity of new hybrids, compared with existing cultivars, was characterized using whole genome sequences. Based on pair‐wise distances, cluster analysis clearly showed increased diversity of new hybrids compared with earlier selections. These results document new triploid hybrids of M. × giganteus with enhanced biomass and theoretical ethanol yields in combination with broader genetic diversity and lowreproductive fertility.

Touchell, Darren H.

Prediction of Specificity of α-Conotoxins to Subtypes of Human Nicotinic Acetylcholine Receptors with Semi-supervised Machine Learning

Conotoxins are a family of highly toxic neurotoxins composed of cysteine-rich peptides produced by marine cone snails. The most lethal cone snail species to humans is Conus geographus, with fatality rates of up to ∼65% from a single sting, which is caused mostly by the activity of α-conotoxins against human nicotinic acetylcholine receptors (nAChRs). While sequence-based machine learning (ML) classifiers have been trained to identify targets of conotoxins binding voltage-gated ion channels, no ML model has been built to predict the subtype-specific nAChR targets of α-conotoxins. Here, we trained an ML model in a semi-supervised manner to predict the specificity of α-conotoxin binding toward different human nAChR subtypes to overcome the challenge of limited data in subtype-specific nAChR targets of α-conotoxins and the issue that one α-conotoxin can bind multiple nAChR subtypes with high selectivity. We considered additional features of sequences of α-conotoxins in training our ML model, including the secondary structure propensities and electrostatic properties, which resulted in better prediction capability for the ML model. Notably, we identify that most α-conotoxins bind to α3β2, α1γδ, and α7 subtypes of human nAChRs. Our findings from this study provide a framework for predicting targets of various kinds of toxins.

59 BASIC BIOLOGICAL SCIENCES

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State

Extracellular DNA Alters Detection of Subtle Bacterial Responses to Soil Rewetting

Microbial communities are often characterized using DNA-based sequencing, but these approaches also capture extracellular DNA (exDNA) released from dead cells, potentially altering inference about microbial responses to environmental change. This may be especially important during pulse disturbances, such as soil drying–rewetting, which can increase microbial mortality and transient necromass pools. We assessed whether exDNA altered inference about bacterial responses to drying–rewetting (an 80 mm simulated rainfall event following a 28-day drought) in conventionally tilled corn and perennial switchgrass soils. We quantified bacterial abundance (16 S rRNA gene copies), alpha diversity, and community composition in paired soil samples with exDNA included (+ exDNA) and in samples treated with propidium monoazide (PMAxx) to reduce amplification of exDNA (− exDNA). At our level of replication (n = 4), PMAxx treatment did not significantly alter overall temporal response patterns (i.e., no significant main effect of DNA treatment or DNA × time interaction). However, PMAxx treatment increased sensitivity to detect some pairwise temporal changes in bacterial abundance and community composition in corn soils following rewetting. exDNA pools were proportionally highest immediately after rewetting in corn soils, suggesting transient extracellular DNA may contribute to masking during disturbance recovery. In contrast, PMAxx treatment had comparatively small effects in switchgrass soils, which exhibited weaker temporal responses overall. Inclusion of exDNA also changed which taxa appeared most responsive to rewetting. Together, our results suggest that exDNA does not uniformly bias soil microbial inference, but may reduce detectability of subtle disturbance-driven shifts in certain soils. Future studies should advance knowledge of microbial turnover and necromass dynamics, particularly using multiple complementary methods, to help predict when exDNA is most likely to influence ecological inference.

drying-rewetting

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES

Machine learning approaches for influenza A virus risk assessment identifies predictive correlates using ferret model in vivo data

In vivo assessments of influenza A virus (IAV) pathogenicity and transmissibility in ferrets represent a crucial component of many pandemic risk assessment rubrics, but few systematic efforts to identify which data from in vivo experimentation are most useful for predicting pathogenesis and transmission outcomes have been conducted. To this aim, we aggregated viral and molecular data from 125 contemporary IAV (H1, H2, H3, H5, H7, and H9 subtypes) evaluated in ferrets under a consistent protocol. Three overarching predictive classification outcomes (lethality, morbidity, transmissibility) were constructed using machine learning (ML) techniques, employing datasets emphasizing virological and clinical parameters from inoculated ferrets, limited to viral sequence-based information, or combining both data types. Among 11 different ML algorithms tested and assessed, gradient boosting machines and random forest algorithms yielded the highest performance, with models for lethality and transmission consistently better performing than models predicting morbidity. Comparisons of feature selection among models was performed, and highest performing models were validated with results from external risk assessment studies. Our findings show that ML algorithms can be used to summarize complex in vivo experimental work into succinct summaries that inform and enhance risk assessment criteria for pandemic preparedness that take in vivo data into account.

59 BASIC BIOLOGICAL SCIENCES

The SRG/eROSITA All-Sky Survey: Optical identification and properties of galaxy clusters and groups in the western galactic hemisphere

The first SRG/eROSITA All-Sky Survey (eRASS1) provides the largest intracluster medium-selected galaxy cluster and group catalog covering the western Galactic hemisphere. Compared to samples selected purely on X-ray extent, the sample purity can be enhanced by identifying cluster candidates using optical and near-infrared data from the DESI Legacy Imaging Surveys. Using the red-sequence-based cluster findereROMaPPer, we measured individual photometric properties (redshiftz λ , richnessλ, optical center, and BCG position) for 12000 eRASS1 clusters over a sky area of 13 116 deg 2 , augmented by 247 cases identified by matching the candidates with known clusters from the literature. The median redshift of the identified eRASS1 sample isz= 0.31, with 10% of the clusters atz> 0.72. The photometric redshifts have an accuracy ofδz/(1 +z) ≲ 0.005 for 0.05 specand velocity dispersionσ) were measured a posteriori for a subsample of 3210 and 1499 eRASS1 clusters, respectively, using an extensive compilation of spectroscopic redshifts of galaxies from the literature. We infer that the primary eRASS1 sample has a purity of 86% and optical completeness >95% forz> 0.05. For these and further quality assessments of the eRASS1 identified catalog, we applied our identification method to a collection of galaxy cluster catalogs in the literature, as well as blindly on the full Legacy Surveys covering 24069 deg 2 . Using a combination of these cluster samples, we investigated the velocity dispersion-richness relation, finding that it scales with richness as log(λ norm ) = 2.401 × log(σ) − 5.074 with an intrinsic scatter ofδ in = 0.10 ± 0.01 dex. The primary product of our work is the identified eRASS1 cluster catalog with high purity and a well-defined X-ray selection process, opening the path for precise cosmological analyses presented in companion papers.

Astronomy & Astrophysics