Engineering PapersSearch

SEARCH · Engineering Papers

Results for “sequence development”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Develop High-Throughput Workflows for Whole-Genome Sequencing and Insertion Site Screening (CRADA Final Report)

The engineering of microbes for biomanufacturing (e.g. of fuels, chemicals, materials) applications has advanced to a stage where researchers screen genetic libraries with millions of variations each for those with enhanced productivity. This screening, however, can be slow and expensive, as screening individual variants in a high-throughput yet cost-effective manner is challenging. In this project, we aimed to reduce by 3-fold costs associated with the sequencing aspects of the screening process (to determine which genetic variant is responsible for an observed change in productivity), while being able to process over 1,000 samples per batch.

60 APPLIED LIFE SCIENCES

Develop High-Throughput Workflows for Whole-Genome Sequencing and Insertion Site Screening

The engineering of microbes for biomanufacturing (e.g. of fuels, chemicals, materials) applications has advanced to a stage where researchers screen genetic libraries with millions of variations each for those with enhanced productivity. This screening, however, can be slow and expensive, as screening individual variants in a high-throughput yet cost-effective manner is challenging. In this project, we aimed to reduce by 3-fold costs associated with the sequencing aspects of the screening process (to determine which genetic variant is responsible for an observed change in productivity), while being able to process over 1,000 samples per batch.

60 APPLIED LIFE SCIENCES

Experimental Validation of Thermal Hydraulic Behavior in Sodium Fast Reactors (SFR) with the Thermal Hydraulic Experimental Test Article (THETA)

Thermal stratification and transition to natural circulation pose two of the largest sources of uncertainty in systems-level modeling of liquid metal-cooled fast reactors. As these phenomena typically develop during transient event sequences, licensing-basis events analyzed using systemslevel models may have considerable uncertainties associated with thermal-hydraulic parameters of the system to account for these phenomena. As a result, the validation basis for these phenomena for systems-level codes is insufficient to fully support the wide range of liquid metal fast reactors being developed in the US. Currently, the most viable path for licensing a design is to take significant conservatisms and maintain sufficiently large safety margins to account for this uncertainty.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Towards verifiable cancer digital twins: tissue level modeling protocol for precision medicine

Cancer exhibits substantial heterogeneity, manifesting as distinct morphological and molecular variations across tumors, which frequently undermines the efficacy of conventional oncological treatments. Developments in multiomics and sequencing technologies have paved the way for unraveling this heterogeneity. Nevertheless, the complexity of the data gathered from these methods cannot be fully interpreted through multimodal data analysis alone. Mathematical modeling plays a crucial role in delineating the underlying mechanisms to explain sources of heterogeneity using patient-specific data. Intra-tumoral diversity necessitates the development of precision oncology therapies utilizing multiphysics, multiscale mathematical models for cancer. This review discusses recent advancements in computational methodologies for precision oncology, highlighting the potential of cancer digital twins to enhance patient-specific decision-making in clinical settings. We review computational efforts in building patient-informed cellular and tissue-level models for cancer and propose a computational framework that utilizes agent-based modeling as an effective conduit to integrate cancer systems models that encode signaling at the cellular scale with digital twin models that predict tissue-level response in a tumor microenvironment customized to patient information. Furthermore, we discuss machine learning approaches to building surrogates for these complex mathematical models. These surrogates can potentially be used to conduct sensitivity analysis, verification, validation, and uncertainty quantification, which is especially important for tumor studies due to their dynamic nature.

60 APPLIED LIFE SCIENCES

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES

Spatial description of dislocation nucleation in the shock response of single-crystal aluminum

Nonequilibrium molecular dynamics simulations of shock loaded single-crystal Al in the $\langle$100$\rangle$, $\langle$110$\rangle$, $\langle$111$\rangle$, and $\langle$123$\rangle$ orientations are conducted to study elastic and plastic shockwave formation and details associated with dislocation activity. A computer vision-based approach is implemented to capture the presence of dislocations and describe their spatial characteristics in the zone of nucleation behind the propagating shockwave. The methodology developed relies on the sequences of images extracted during shock loading that show dislocation activity within a cross section of the sample. Results reveal that the spacing between activated slip systems is orientation dependent and exhibits a modest reduction for the $\langle$100$\rangle$ and $\langle$111$\rangle$ orientations as shock pressure increases. Comparisons are made to existing theoretical models. Such relationships between shock pressure and dislocation activity, extracted from molecular dynamics simulations, can be used to inform higher length scale simulations or modeling of dislocation-based plasticity during shock.

36 MATERIALS SCIENCE

Automated Generation of Weather and Climate Analysis Products

Wind roses are an important part to site operations as they depict the wind speed and direction percentage over a time period. In this project, I automated the production of wind roses for SRS meteorological towers. I developed a script to sequence through the dataset by period and create a wind rose for each of those period sets. Graphics were generated for every 4-hour period of the day, for each of the 4 heights of the instruments, for every half of each month. A 10-year climatological period consisted of wind speed and directions measurements at 15-minute intervals. We used the years 2014-2024 as the wind instruments on the tower were upgraded to sonic anemometers in early 2014. I then compared these wind roses to those from a previous study done in 2003 to analyze the differences and similarities in the wind patterns. The wind roses are also used to understand environmental transport conditions at SRS. When comparing the different heights of the anemometers to the 2003 report, there is a similarity in the fact that the direction is relatively the same across the levels, with the wind speeds increasing as you get higher in the air. When comparing the season of our graphs against the 2003 report, we noticed that the winter and summer months appear to have very similar wind directions. However, in the spring we noticed that there were more southerly winds compared to the 2003 report which had more westerly winds. Finally, in the fall months, we noticed that there was a higher percentage of northeasterly winds, while the 2003 report showed more southeasterly winds. We can use these summaries to estimate the directional dependence of dose that the surrounding areas receive throughout the year according to time of day.

47 OTHER INSTRUMENTATION

NW-BRaVE T3 Hydroplane Project Close: Project Close-out for T3 Hydroplane Analysis

Thrust 3 of the Northwest Biopreparedness Research in a Virtual Environment was an expansive project including method development, sample collection and sequence analysis. The sampling occurred over a multi-year period to generate metagenomic datasets that inform cyanophage-picocyanobacterial interactions in the Salish Sea across a moderate timeframe and geographical range. Part of the thrust’s aim was to validate how much experimental structural and multiomics work in a model organism (Prochlorococcus Marinus, str. MED4) from thrusts 1 and 2 would carry over into a broader range of related organisms in the natural world, to address a fundamental question in scientific preparation for epidemics: whether and how much experimental information from known and experimentally tractable species can translate to actionable biological information in unknown species. In other words, thrust 3 aimed to find out whether the model organism experiments matter in terms of how organisms interact. This report updates work described in Johnson and Pollock 2025 (1).

54 ENVIRONMENTAL SCIENCES

Hierarchical semi-Markov models with duration-aware dynamics for activity sequences

Residential electricity demand at granular scales is driven by what people do and for how long. Accurately forecasting this demand for applications like microgrid management and demand response therefore requires generative models for activities that can produce realistic daily activity sequences, capturing both the timing and duration of human behavior. This paper develops a generative model of human activity sequences using nationally representative time-use diaries at a 10-min resolution. We use this model to quantify which demographic factors are most critical for improving predictive performance. We propose a hierarchical semi-Markov framework that addresses two key modeling challenges. First, a time-inhomogeneous Markov router learns the patterns of “which activity comes next.” Second, a semi-Markov hazard component explicitly models activity durations, capturing “how long” activities realistically last. To ensure statistical stability when data are sparse, the model pools information across related demographic groups and time blocks. The entire framework is trained and evaluated using survey design weights to ensure our findings are representative of the U.S. population. On a held-out test set, we demonstrate that explicitly modeling durations with the hazard component provides a substantial and statistically significant improvement over purely Markovian models. Furthermore, our analysis reveals a clear hierarchy of demographic factors: Sex, Day-Type, and Household Size provide the largest predictive gains, while Region and Season, though important for energy calculations, contribute little to predicting the activity sequence itself. The result is an interpretable and robust generator of synthetic activity traces, providing a high-fidelity foundation for downstream energy systems modeling.

24 POWER TRANSMISSION AND DISTRIBUTION

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES

Predicting synthetic mRNA stability using massively parallel kinetic measurements, biophysical modeling, and machine learning

Abstract mRNA degradation is a central process that affects all gene expression levels, though it remains challenging to predict the stability of a mRNA from its sequence, due to the many coupled interactions that control degradation rate. Here, we carried out massively parallel kinetic decay measurements on over 50,000 bacterial mRNAs, using a learn-by-design approach to develop and validate a predictive sequence-to-function model of mRNA stability. mRNAs were designed to systematically vary translation rates, secondary structures, sequence compositions, G-quadruplexes, i-motifs, and RppH activity, resulting in mRNA half-lives from about 20 seconds to 20 minutes. We combined biophysical models and machine learning to develop steady-state and kinetic decay models of mRNA stability with high accuracy and generalizability, utilizing transcription rate models to identify mRNA isoforms and translation rate models to calculate ribosome protection. Overall, the developed model quantifies the key interactions that collectively control mRNA stability in bacterial operons and predicts how changing mRNA sequence alters mRNA stability, which is important when studying and engineering bacterial genetic systems.

Cetnar, Daniel P.

Kölliker's Organ Functions as a Developmental Hub in Mouse Cochlea Regulating Spiral Limbus and Tectorial Membrane Development

Kölliker's organ is a transient developmental structure in the mouse cochlea that undergoes significant remodeling postnatally. Utilizing an epithelial-specific conditional deletion mouse model of Prdm16 (marker and regulator of Kölliker's organ), we show that Prdm16 is required for interdental cell development, and thereby the development of the limbal domain of the tectorial membrane and its medial anchorage to the spiral limbus. Additionally, we show that Kölliker's organ is involved in normal tectorial membrane collagen fibril development and maturation. Interestingly, mesenchymal cells of the spiral limbus underneath Prdm16 -deficient Kölliker's organ failed to produce interstitial matrix proteins, resulting in a hypoplastic and truncated spiral limbus, indicating a non-cell autonomous role of Prdm16 in regulating spiral mesenchymal matrix development. Single-cell RNA sequencing identified differentially expressed genes in Prdm16 -deficient Kölliker's organ suggesting a role for connective tissue growth factor (CTGF) downstream Prdm16 in epithelial-mesenchymal signaling involved in spiral limbus matrix deposition. Prdm16 -deficient mice showed a hearing deficit, as indicated by elevated auditory brainstem response thresholds at most frequencies, consistent with the cochlear structural defects. Both sexes were studied. This work establishes Prdm16 as a deafness gene in mice through its role in regulating Kölliker's organ development. Such understanding recognizes Kölliker's organ as a developmental hub regulating multiple surrounding cochlear structures.

Zhang, Hongji

Comparison of the spatial statistics of random and defined-sequence photoresist films

The resolution-line edge roughness-sensitivity tradeoff has motivated the exploration of potential improvements using defined sequence polymers and polymer-bound photoacid generators and quenchers. We characterize the internal structures of positive tone photoresist polymer films formed from defined sequence polymers and compare them with random copolymers of the same composition. We model their imaging to connect initially to developable film structures. We use a polymer packing algorithm to simulate films of diverse compositions and locations of photoacid generators and quenchers, using the composition of an ESCAP photoresist. We use a simple extreme ultraviolet exposure-deprotection algorithm to model developable image formation within them. In all cases, the spatial distribution of chemical moieties in the film for defined sequence polymers is nearly indistinguishable from random copolymers. We evaluate several exposure-deprotection scenarios and find that a defined sequence copolymer has a distinctive developable image under certain circumstances. The use of defined sequence polymers within a photoresist layer does not automatically result in improved imaging; however, they do have some characteristics different from random polymers of the same composition. Further study of these characteristics may provide a route to improved control over the nanoscale imaging process.

36 MATERIALS SCIENCE

Combinatorial transcription factor binding encodes cis -regulatory wiring of mouse forebrain GABAergic neurogenesis

Transcription factors (TFs) bind combinatorially to cis-regulatory elements, orchestrating transcriptional programs. Although studies of chromatin state and chromosomal interactions have demonstrated dynamic neurodevelopmental cis-regulatory landscapes, parallel understanding of TF interactions lags. To elucidate combinatorial TF binding driving mouse basal ganglia development, we integrated chromatin immunoprecipitation sequencing (ChIP-seq) for twelve TFs, H3K4me3-associated enhancer-promoter interactions, chromatin and gene expression data, and functional enhancer assays. We identified sets of putative regulatory elements with shared TF binding (TF-pRE modules) that orchestrate distinct processes of GABAergic neurogenesis and suppress other cell fates. The majority of pREs were bound by one or two TFs; however, a small proportion were extensively bound. These sequences had exceptional evolutionary conservation and motif density, complex chromosomal interactions, and activity as in vivo enhancers. Our results provide insights into the combinatorial TF-pRE interactions that activate and repress expression programs during telencephalon neurogenesis and demonstrate the value of TF binding toward modeling developmental transcriptional wiring.

59 BASIC BIOLOGICAL SCIENCES

Viromics approaches for the study of viral diversity and ecology in microbiomes

Viruses are found across all ecosystems and infect every type of organism on Earth. Traditional culture-based methods have proven insufficient to explore this viral diversity at scale, driving the development of viromics, the sequence-based analysis of uncultivated viruses. Viromics approaches have been particularly useful for studying viruses of microorganisms, which can act as crucial regulators of microbiomes across ecosystems. They have already revealed the broad geographic distribution of viral communities and are progressively uncovering the expansive genetic and functional diversity of the global virome. Moving forward, large-scale viral ecogenomics studies combined with new experimental and computational approaches to identify virus activity and host interactions will enable a more complete characterization of global viral diversity and its effects.

Ecology

A modular and extensible CHARMM-compatible model for all-atom simulation of polypeptoids

Peptoids (N-substituted glycines) are a class of sequence-defined synthetic peptidomimetic polymers with applications including drug delivery, catalysis, and biomimicry. Classical molecular simulations have been used to predict and understand the conformational dynamics of single chains and their self-assembly into morphologies including sheets, tubes, spheres, and fibrils. The CGenFF-NTOID model based on the CHARMM General Force Field has demonstrated success in accurate all-atom molecular modeling of peptoid structure and thermodynamics. Extension of this force field to new peptoid side chains has historically required reparameterization of side chain bonded interactions against ab initio data. This fitting protocol improves the accuracy of the force field but is also burdensome and precludes modular extensibility of the model to arbitrary peptoid sequences. In this work, we develop and demonstrate a Modular Side Chain CGenFF-NTOID (MoSiC-CGenFF-NTOID) as an extension of CGenFF-NTOID employing a modular decomposition of the peptoid backbone and side chain parameterizations, wherein arbitrary side chains within the large family of substituted methyl groups (i.e., –CH 3 , –CH 2 R, –CHRR', and –CRR'R") are directly ported from CGenFF. We validate this approach against ab initio calculations and experimental data to develop a MoSiC-CGenFF-NTOID model for all 20 natural amino acid side chains along with 13 commonly used synthetic side chains and present an extensible paradigm to efficiently determine whether a novel side chain can be directly incorporated into the model or whether refitting of the CGenFF parameters is warranted. We make the model freely available to the community along with a tool to perform automated initial structure generation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Fermilab's controls development with virtual accelerator

Control Systems development is often the last thing considered when designing and building new equipment, e.g. a new detector or superconducting RF LINAC; however when the new equipment is installed, it is the first thing desired to be operational for testing. Due to frequent delays in building new equipment and project deadlines, control system development and testing is often curtailed. A way to alleviate this problem is to simulate the control system, though this will be challenging for complex systems.The Fermilab PIP-II (proton improvement plan - II) project is being constructed at Fermilab to deliver $800\,MeV$ protons of $>1\,MW$ beam power to replace the present LINAC for the remainder of the existing accelerator complex. The new LINAC consists of a warm front end (WFE), 23 superconducting RF cryomodules (of 5 types), and a beam transfer line (BTL) to the existing complex.The accelerator physics group has a parallel project to create a digital twin (DT) of the PIP-II accelerator. We have coupled the EPICS controls to this DT and are developing both the DT and EPICS software in parallel. This will allow us to develop the EPICS software framework, the HMIs, sequences, high level physics applications, and other services for use in a fully functional control system.This presentation will detail the work that we have performed to date and show demonstrations of controlling and monitoring the status of the accelerator, as well as future plans for this work.

Hanlet, Pierrick [Fermilab]

Nucleic Acid-Based Detection Protease Activity

Proteases include clinically relevant markers for clotting disorders, certain cancers as well as toxins. Assays for protease activity often use designed peptides mimicking natural substrates and detection with colorometric and fluorescence-based detection that is difficult to multiplex without expensive and resource demanding instruments. This work demonstrates detection of proteolytic activity using PCR and sequencing-readable reporter molecules. The assay development focused on binding the constructed peptide-oligonucleotide chimera to immobilized streptavidin. Thrombin, an essential component of the clotting cascade, was used as a model system for testing peptide substrate recognition and release of a designed oligonucleotide for detection. Detection of protease activity was demonstrated in a concentration-dependent manner using MALDI-MS, RT-PCR and DNA sequencing.

Wunschel, David S [Pacific Northwest National Labo