Engineering PapersSearch

SEARCH · Engineering Papers

Results for “sequence development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

RNA language models predict mutations that improve RNA function

Structured RNA lies at the heart of many central biological processes, from gene expression to catalysis. RNA structure prediction is not yet possible due to a lack of high-quality reference data associated with organismal phenotypes that could inform RNA function. We present GARNET (Gtdb Acquired RNa with Environmental Temperatures), a new database for RNA structural and functional analysis anchored to the Genome Taxonomy Database (GTDB). GARNET links RNA sequences to experimental and predicted optimal growth temperatures of GTDB reference organisms. Using GARNET, we develop sequence- and structure-aware RNA generative models, with overlapping triplet tokenization providing optimal encoding for a GPT-like model. Leveraging hyperthermophilic RNAs in GARNET and these RNA generative models, we identify mutations in ribosomal RNA that confer increased thermostability to the Escherichia coli ribosome. The GTDB-derived data and deep learning models presented here provide a foundation for understanding the connections between RNA sequence, structure, and function.

59 BASIC BIOLOGICAL SCIENCES

Development of high throughput and in vitro assays for analyzing RNA modifications

Modifications on RNAs play major roles in their stability, translation, and enzymatic activity. Despite its importance, the current techniques are insufficient to study the structure and function of RNA modifications. Indeed, the National Academies of Science, Engineering and Medicine indicate that developing new tools and further study the function of RNA modifications is strategically a high priority for advancing science in the coming years (https://www.nationalacademies.org/our-work/toward-sequencing-and-mapping-of-rna-modifications). RNA modifications occur in all domains of life controlling processes such as RNA turnover, translation regulation, cellular defenses and bioproduction. Our preliminary data indicated that the insulin mRNA might get ADP-ribosylated by the ADP-ribosyltransferase PARP12. RNA ADP-ribosylation has been described in Escherichia coli. Combined to the fact that ADP-ribosyltransferase (PARP) genes are conserved throughout evolution we hypothesize that this modification might play essential roles in cells. Therefore, we proposed to develop sequencing techniques and in vitro enzymatic assays to identify and validate ADP-ribosylation motifs and sites. Here we report the development of RNA-seq and qPCR assays to identify ADP-ribosylated RNAs, in addition to a nicotinamide adenosine dinucleotide (NAD – ADP-ribosylation donor) consumption assay and an enzyme-linked immunosorbent assay (ELISA) to measure ADP-ribosyltransferase activity. Testing these assays with the insulin mRNA confirmed that this transcript is ADP-ribosylated. These assays will not only enable studying the function of ADP-ribosylation but can be easily adapted for studying other RNA modifications. This will open opportunities to study RNA modifications in different model systems from bacteria to viruses to plants, bringing insights into their cellular functions and the possibility of targeting them for biotechnological applications.

59 BASIC BIOLOGICAL SCIENCES

Developing Near Optimal Control Sequences for Chiller Plants with Water-side Economizers: A Case Study in a Warm and Marine Climate

Various advanced control sequences for chiller plants with water-side economizers (WSE) have been proposed in literature, but the optimization of those controls is limited. It is possible to maximize energy savings by developing near-optimal control sequences, which are dependent on several factors such as the load profile. To address these gaps, we first identify an advanced control sequence and three key control parameters for chiller plants with WSE. Next, optimizations are performed to minimize energy consumption for seven combinations of control parameters. A chiller plant with WSE system in a warm and marine climate is studied and two load profiles are considered. The system and controls are modeled using the Modelica Buildings library. The results show optimizing the selected control parameters can reduce energy consumption by up to 11% depending on the load profile. Specifically, optimizing the cooling tower efficiency threshold in the condenser water reset control can significantly reduce energy savings for the variable load profile by efficiently shifting the load from the cooling tower to the chiller. This paper provides practical guidance for developing near-optimal control sequences for chiller plant with WSE systems considering impacts such as the load profile.

chiller plant

Two deeply conserved non-coding sequences control PLETHORA1/2 expression and coordinate embryo and root development

Conserved non-coding sequences (CNSs) are integral elements of transcriptional regulation. Transcriptional tuning of PLETHORA (PLT) genes that encode master regulators of plant development is vital for embryogenesis and meristematic function. However, how the expression of PLT genes is modulated through CNSs remains unclear. Through motif-based mining of upstream sequences in 120 angiosperm genomes, we identified 21 conserved and lineage-specific CNSs, two of which are unusually long, similar, and colinear within eudicots. Using Arabidopsis thaliana, we demonstrate that these two deeply conserved elements, which we named BOX1 and BOX2, control PLT1 and PLT2 expression. CRISPR mutants within these elements specifically reduced PLT expression levels, and reporter lines revealed that deletion of either or both BOXes altered and/or abrogated the PLT2 expression pattern in the root tip, affecting the ability to rescue the plt1 plt2 double mutant. We further show that the influence of these elements on expression patterns is already exerted during embryogenesis and functional in the context of the early embryo. Finally, we reveal the existence of a BOX-mediated autoregulatory feedback loop that, in large part, explains CNS influence on expression patterns. We thus uncover a transcriptional mechanism by which genes encoding master regulators of embryo and root meristem development are regulated.

PLETHORA

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277

Hamiltonian simulation in Zeno subspaces

Here, we investigate the quantum Zeno effect as a framework for designing and analyzing quantum algorithms for Hamiltonian simulation. We show that frequent projective measurements of an ancilla qubit register can be used to simulate quantum dynamics on a target qubit register with a circuit complexity similar to randomized approaches. The classical sampling overhead in the latter approaches is traded for ancilla qubit overhead in Zeno-based approaches. A second-order Zeno sequence is developed to improve scaling and implementations through unitary kicks are discussed. We derive rigorous error bounds that allow for identifying the associated circuit complexities for the first- and second-order Zeno sequences. We show that the circuits over the combined register can be identified as a subroutine commonly used in post-Trotter Hamiltonian simulation methods. We build on this observation to reveal connections between different Hamiltonian simulation algorithms.

Hamiltonian simulation

Full ribosomal operon sequencing of anaerobic gut fungi (phylum Neocallimastigomycota ): insights on its markers and phylogenetic resolution

The phylogenetic affiliations of anaerobic gut fungi (Neocallimastigomycota) are typically evaluated using single-gene markers. However, this approach often fails to resolve relationships between closely related lineages. To address this issue and identify alternative markers, we created a curated database comprising the complete ribosomal operon sequences of 156 isolates, representing 20 of the 22 recognized genera and two new genus-level clades. Using long-read sequencing, we obtained ~9 kbp operon sequences and developed a robust analysis pipeline. Incorporating both coding genes and non-coding regions (excluding IGS1) improved phylogenetic resolution. This phylogenetic approach successfully resolved the Cyllamyces and Caecomyces clades (hard-to-distinguish genetically), as well as seven analysed Piromyces species. We also scanned the operon for markers that are suitable for short-read sequencing platforms, with the aim of enhancing biodiversity and phylogenetic studies. Notably, the ETS1 genetic region also enabled the distinction between these lineages, indicating its phylogenetic value within the ribosomal operon. The resulting database is a valuable resource for expanding and strengthening phylogenetic frameworks.

High-throughput sequencing

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram

Overview of IMPACT Data Acquisition System and Data Reduction Process

This report documents the development of the data acquisition system (DAS) and data reduction methodologies for the Irradiated Material Property Accelerated Characterization Test (IMPACT) experiment at the Advanced Test Reactor (ATR). The IMPACT experiment is designed to enable in-pile measurement of thermal conductivity in metallic nuclear fuels, specifically U-10Zr, using an instrumented thermal conductivity probe. The DAS supports both passive temperature monitoring and active thermal interrogation of the probe through controlled AC and DC excitation. Significant modifications to laboratory-scale systems were required to accommodate the higher resistance paths associated with the in-pile application. Custom electronics and relay-controlled measurement sequencing were developed to enable the measurement and sufficient power delivery to the sensing region. A reduced-order, axisymmetric thermal model based on the thermal quadrupoles method is presented to support data interpretation. This model enables efficient evaluation of transient heat transfer behavior and facilitates solution of the inverse problem required to extract thermal properties from measured signals. Multiple boundary condition formulations are discussed to address varying experimental time scales and geometries. Additionally, machine learning techniques are introduced to support data reduction and improve confidence in inverse solutions. Convolutional neural networks are applied to identify the presence of gas gaps and other evolving geometric features that significantly impact thermal response during irradiation. These efforts contribute to the broader integration of digital twin frameworks and real-time modeling capabilities within the Advanced Fuels Campaign.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN

Development of near-optimal advanced control sequences for chiller plants with water-side economizers in U.S. Climates (ASHRAE RP-1661)

Various advanced control sequences for chiller plants with water-side economizers (WSE) have been proposed in literature, but the evaluation and optimization of those controls is limited. It is possible to maximize energy savings by selecting different sequences and related parameters based on the plant configuration, load, and climate. This paper addresses this gap by developing near-optimal advanced control sequences for chiller plants with WSEs. First, advanced control sequences for chiller plants with WSEs are categorized into condenser water, chilled water, and hybrid controls and representative sequences from each category are identified. Next, 504 different scenarios are optimized. These scenarios represent all possible combinations of two plant configurations, a constant or variable load profile, three advanced control sequences, and seven optimization parameter combinations in six climate zones. The results show the recommended near-optimal sequences can reduce energy consumption by up to 15% relative to the baseline depending on the configuration, load profile, and climate. Specifically, the CW-CHW sequence is recommended for the majority of systems because it is often the most energy efficient and/or reduces the runtime of chillers. The methodology in this paper provides practical guidance for achieving energy savings through near-optimal control of chiller plants with WSEs.

42 ENGINEERING

CDL2PLC translator v0.1.0

The CDL-PLC translator aims at translating control sequences for building energy systems from the CDL CXF format to the PLCopen XML format. The CDL CXF developed at LBL within the OpenBuildingControl project, and now being standardized via ASHRAE Standard 231P, enables expressing control sequences developed in the simulation environment Modelica in a JSON format. The PLCopen XML is an existing exchange format standardized in IEC 61131-10 for Programmable Logic Controllers (PLCs) following the IEC 61131 standard as one target system of CDL among others. The translation from the CDL CXF to the PLCopen XML contributes to a seamless workflow from the model-based development of control sequences in simulation environments, which is not building practice today, and their digital implementation on building controllers, which replaces graphical and textual documents used for this purpose today. The translator is at a prototypical stage and enables, as a proof of concept, the translation of very simple control sequences composed of 4 selected function blocks out of 137 function blocks defined in CDL. The translation includes the connection of inputs and outputs of function blocks and the expression of a control function in CDL to the equivalent code in IEC 61131-3.

Walther, Karl

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis

In Vitro Selection of Antibodies Targeting Yersinia pestis Membrane Lipids Using Nanodisc-Based Antigen Presentation

Proteins are the most common targets for antibody discovery and vaccine development, but their sequence variability can limit the breadth of resulting antigens. Lipids represent an alternative class of antigens due to their structural conservation and roles in host–pathogen interactions. Here, we describe the development and optimization of an in vitro antibody selection workflow using lipid-containing nanodiscs as antigen presentation platforms to enable phage and yeast display selections under conditions adapted for these non-protein targets. Lipopolysaccharide (LPS) nanodiscs were first used as a model system to evaluate selection strategies, including competitive and subtractive approaches to reduce non-specific binders, yielding peptide and single-chain variable fragment (scFv) binders that were affinity matured to improve binding signals. The same approach was subsequently used to select scFv antibodies that recognize lipid nanodiscs prepared from Yersinia pestis membrane lipid extracts. These antibodies show binding to lipid nanodiscs derived from Y. pestis, with evidence of selectivity relative to control nanodiscs. Overall, this work establishes a workflow for antibody selection against lipid-containing nanodisc antigens and highlights practical considerations associated with these targets. The approach may be useful for generating affinity reagents to membrane-associated lipids, although further characterization is required to define antigen specificity and functional activity.

59 BASIC BIOLOGICAL SCIENCES

Telomere-to-telomere assemblies of chromosome 10 reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by Helitrons

The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases ( KCS ), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomere-to-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.

09 BIOMASS FUELS

Predicting river turbidity in Pine Island Bayou using machine learning techniques coupled with variational mode decomposition

Elevated turbidity levels pose significant public health risks by facilitating the transport of harmful pollutants, including metals, organic compounds, and pathogenic microorganisms into the surface water. These conditions create serious challenges for public recreational water use and drinking water treatment, leading to economic losses and health risks. This study utilizes water monitoring data in Pine Island Bayou, Texas, and develops a Sequence-to-Sequence (S2S) model to predict turbidity using Attention-based Gated Recurrent Units with Encoder-Decoder (AT-GRU-ED) and Long Short-Term Memory (LSTM), coupled with Variational Mode Decomposition (VMD). Compared to the model without VMD, the model demonstrates satisfactory 72-hour turbidity prediction performance, achieving MAEs of 2.60 and 3.29 NTU (reductions of 53% and 58%), RMSEs of 21.08 and 31.49 NTU (reductions of 82% and 80%), and R² values of 0.96 and 0.84 on the validation and test sets, respectively. Feature importance analysis reveals that water temperature is the dominant factor influencing seasonal turbidity patterns, while real-time hourly rainfall significantly contributes to short-term variability. Turbidity typically peaks within 48 hours after rainfall events due to lagged effects from surface runoff and upstream flow. Findings suggest suspending recreational water use and water supply pumping for three days after heavy rainfall can benefit public health and improve water treatment processes. Discharges above 100 m3/s are found to accelerate sediment dilution and transport, reducing turbidity levels more quickly after the peak. In conclusion, the proposed model demonstrates reliable 72-hour turbidity prediction, supporting decision-making for water treatment plant operations and providing early warning for public recreational water use.

Deep learning

Small-Signal Stability of Grid-Forming Converters Under Fault Conditions

Threshold virtual impedance (TVI)-based current limiting for grid-forming converters (GFMs) has gained great interest due to its ability to maintain voltage source behaviour during faults. However, sequence component extraction (SCE) and negative-sequence control (NSC) are often overlooked in small-signal stability assessments during faults. This paper develops small-signal sequence impedance models for GFMs under four well-known SCE methods based on TVI current limiting control during symmetrical fault conditions. Using the developed impedance models, the impacts of SCE and NSC, and the voltage and current control loop bandwidths, on system stability during faults are investigated. Additionally, since negative-sequence TVI (TVI-) is typically added along with its positive-sequence counterpart, which is often inductive, inductive and resistive TVI- are examined. The findings suggest that a higher voltage or current control loop bandwidth has a negative impact on system stability, while SCE and NSC largely reduce the stable range for voltage and current control loop bandwidth during faults, and that the severity of such impacts is determined by the particular SCE method. Furthermore, it is observed that inductive TVI- significantly degrades system stability, while resistive TVI- can enhance stability when suitable SCE methods are appropriately selected and designed. Matlab/Simulink electromagnetic transient simulations validate these analytical results.

24 POWER TRANSMISSION AND DISTRIBUTION

Data for Transposon Signatures of Allopolyploid Genome Evolution

Hybridization brings together chromosome sets from two or more distinct progenitor species. Genome duplication associated with hybridization, or allopolyploidy, allows these chromosome sets to persist as distinct subgenomes during subsequent meioses. Here, we present a general method for identifying the subgenomes of a polyploid based on shared ancestry as revealed by the genomic distribution of repetitive elements that were active in the progenitors. This subgenome-enriched transposable element signal is intrinsic to the polyploid, allowing broader applicability than other approaches that depend on the availability of sequenced diploid relatives. We develop the statistical basis of the method, demonstrate its applicability in the well-studied cases of tobacco, cotton, and Brassica napus, and apply it to several cases: allotetraploid cyprinids, allohexaploid false flax, and allooctoploid strawberry. These analyses provide insight into the origins of these polyploids, revise the subgenome identities of strawberry, and provide perspective on subgenome dominance in higher polyploids.

Genomics

Higher-order Zeno sequences

The quantum Zeno effect typically refers to freezing the dynamics of a quantum system through frequent observations. In general, quantum Zeno dynamics is obtained with an error of order 𝒪⁢(1/𝑁), where 𝑁 is the number of projective measurements performed within a fixed evolution time. In this work, we develop higher-order Zeno sequences that achieve faster convergence to Zeno dynamics, yielding an improved error scaling of 𝒪⁢(1/𝑁 2⁢𝑘 ), where 𝑘 describes the order of the Zeno sequence. This is achieved by relating higher-order Zeno sequences to higher-order Trotter formulas that achieve similar convergence behavior. We leverage this relation to develop higher-order Zeno sequences for different manifestations of the quantum Zeno effect, including frequent projective measurements and unitary kicks. We go on to discuss achieving quantum Zeno dynamics through periodic control fields of high frequency. We explicitly develop control fields that yield a second-order type improvement in the Zeno error scaling and present shorter Zeno sequences. Finally, we discuss the connection to randomized and Uhrig dynamical decoupling to develop more efficient implementations in the weak-coupling regime.

Quantum Zeno dynamics