Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Direct Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Decoding substrate specificity determining factors in glycosyltransferase-B enzymes – insights from machine learning models

Substrate specificity is an essential characteristic of any enzyme's function and an understanding of the factors that determine this specificity is crucial for enzyme engineering. Unlike the structure of an enzyme which is directly impacted by its sequence, substrate specificity as an enzyme attribute involves a rather indirect relationship with sequence as it also depends on structural aspects that dictate substrate accessibility and active site dynamics. In this study, we explore the performance of classifier-based machine learning models trained on curated sequence and structural data for a class of glycosyltransferases (GTs), namely GT-Bs, to understand their substrate specificity determining factors. GTs enable the transfer of sugar moieties to other biomolecules such as oligosaccharides or proteins and are found in all kingdoms of life. In plants, GTs participate in the biosynthesis of plant cell wall biopolymers (e.g.: hemicelluloses and pectins) and are an integral part of the enzymatic machinery that enables the storage of carbon and energy as plant biomass. To elucidate the substrate specificity of uncharacterized GT-Bs, we constructed multi-label machine learning models (Support Vector Classifier, K-Nearest Neighbors, Gaussian Naïve-Bayes, Random Forest) that incorporate both sequence and structural features. These models achieve good predictive accuracies on test datasets. However, despite our use of structural information, we highlight that there is further scope for improvement in training these models to draw interpretable relationships between sequence, structure and substrate specificity determining motifs in GT-Bs.

97 MATHEMATICS AND COMPUTING↗

Detecting Spin-Bath Polarization with Quantum Quench Phase Shifts of Single Spins in Diamond

Single-qubit sensing protocols can be used to measure qubit-bath coupling parameters. However, for sufficiently large coupling, the sensing protocol itself perturbs the bath, which is predicted to result in a characteristic response in the sensing measurements. Here, we observe this bath perturbation, also known as a quantum quench, by preparing the nuclear spin bath of a nitrogen-vacancy (NV) center in polarized initial states and performing phase-resolved spin-echo measurements on the NV electron spin. These measurements reveal a time-dependent phase determined by the initial state of the bath. We derive the relationship between the sensor phase and the Gaussian spin-bath polarization and apply it to reconstruct both the axial and transverse polarization components. Using this insight, we optimize the transfer efficiency of our dynamic nuclear polarization sequence. This technique for directly measuring bath polarization may assist in preparing high-fidelity quantum memory states, improving nanoscale NMR methods, and investigating non-Gaussian quantum baths.

36 MATERIALS SCIENCE↗

Roles for epigenetics in wood formation and stress response intrees–from basic biology to forest management

Annual model and crop species have been the subject of most epigenetic studies for plants. In contrast to annuals, forest trees persist on natural landscapes and experience environmental variation within and across seasons, years, and decades or even centuries. Most forest trees species are undomesticated and typically grown on variable landscapes with no irrigation or application of agricultural chemicals. Forest trees must thus rely on their inherent ability to alter growth and physiology to mitigate the effects of changing abiotic and biotic stressors. Like other plants, trees have mechanisms encoded in their genomic DNA sequence that can respond directly to stress events such as drought or heat. Hypothetically, it would be highly advantageous to join these mechanisms with a dynamic “memory” of past exposure to stress. It is now well established that annual model and crop plants can establish epigenetic-based memory of stress events that support more rapid and robust response to stress in the future. Here, evidence is discussed for epigenetic regulation and “memory” in two fundamental biological processes in trees, wood formation and abiotic stress response. Wood formation is an ideal trait for epigenetic research in trees, as wood formation is highly responsive to environmental conditions and includes multiple rapid developmental changes as cells adopt distinct fates within complex tissues. This is followed by a discussion of research needs that would provide the foundation for new epigenetic applications for forestry.

Groover, Andrew↗

Genome Extraction from Shotgun Metagenome Sequence Data

Uncultivated Bacteria and Archaea comprise the vast majority of species on Earth, but obtaining their genomes directly from the environment, using shotgun sequencing, has only recently become possible. To realize the hope of capturing Earth’s microbial genetic complement, technologies that accelerate recovery of high-quality genomes are necessary. We present a series of analysis steps and data products for the extraction of high quality metagenome-assembled genomes (MAGs) from microbiomes using the U.S. Department of Energy Systems Biology Knowledgebase (KBase) platform (http://www.kbase.us/). In KBase, the process is end-to-end, allowing a user to go from the initial sequencing reads all the way through to MAG genomes, which can then be analyzed with other KBase capabilities such as phylogenetic placement, functional assignment, metabolic modeling, pangenome functional profiling, RNA-Seq, and others. While portions of such capabilities are individually available from other resources, the combination of the intuitive usability, data interoperability, and integration of tools in a freely available compute resource makes KBase a uniquely powerful platform for obtaining MAGs from microbiomes. While this workflow offers tools for each of the key steps in the genome extraction process, it also provides a scaffold that can be easily extended, with additional MAG recovery and analysis tools, via the KBase SDK (Software Development Kit).

Chivian, Dylan↗

A Laplace-Domain Circuit Model for Fault and Stability Analysis Considering Unbalanced Topology

For systems subject to unbalanced faults, analytical model building for stability assessment is a challenging task. This letter presents a straightforward modeling approach. A generalized dynamic circuit representation is achieved by use of the Laplacian transform variable s . Here, we translate the voltage and current relationship at the fault location into the relationship of three subsystems. The final circuit model is an interconnected sequence network with impedances in the Laplace domain. This circuit can be directly converted from a steady-state sequence network. This modeling procedure is illustrated by an example case of an induction motor served by a grid through a series compensated line. Electromagnetic transient simulation results demonstrate that sub-synchronous oscillations can be mitigated when a single-line to ground fault is applied at the motor terminal. Stability analysis results based on the dynamic circuit corroborate the simulation results. What's more, the derived circuit effortlessly reveals why unbalance can enhance stability.

42 ENGINEERING↗

High-quality Acinetobacter genomes recovered from combat wounds via metagenomic sequencing resemble cultured isolate genomes

The ability to accurately characterize wound pathogens is critical to informing clinical decisions for wound infections with complex treatment requirements. Acinetobacter baumannii is an impactful nosocomial pathogen in combat wounds and civilian hospital-acquired infections. An informed understanding of the phylogenetics and epidemiology of A. baumannii infections in military and civilian environments could guide approaches that improve antibiotic treatment regimens for both military and civilian patients. Whole-genome data for bacterial strains can be difficult to obtain due to challenges in culturing isolates from preserved military specimens. Metagenomic sequencing and assembly create opportunities for genomic analysis of pathogens directly from clinical specimens. The ability to perform comparative analyses between metagenome-derived genomes and culture-derived genomes would support a range of comparative bacterial genomic studies. Wound tissue biopsy and effluent samples from combat injuries were subjected to metagenomic sequencing and assembly. In total, 42 microbial metagenome-assembled genomes (MAGs) were obtained directly from metagenomic sequence data, 36 of which were designated “high” quality. Thirty of these genomes corresponded to Acinetobacter, with 29 mapping specifically to A. baumannii. Other observed genera included Bordetella, Citrobacter, Escherichia, and Pseudomonas. Single-copy and multi-copy orthologs were identified across Acinetobacter MAGs and publicly available isolate genomes derived from military and civilian sources. Both MAG and military isolate genomes were annotated with antimicrobial resistance data, and MAG genomes were statistically comparable to genomes obtained from isolates. Our results highlight the potential of de novo metagenome assembly for enabling high-resolution characterization directly from clinical specimens, thereby improving diagnostic precision, guiding antimicrobial stewardship, and enhancing understanding of pathogen evolution across diverse healthcare and battlefield environments.

Acinetobacter baumannii↗

16S and ITS Amplicon Sequencing Fastq files and metadata from PARCHED Panama Tropical Forest soils, 2019-2020,

Model projections predict tropical forests will experience longer periods of drought and more intense precipitation cycles under a changing climate. Such transitions have implications for structure-function relationships within microbial communities. We examine how chronic drying might reshape prokaryotic and fungal communities across four lowland forests in Panama with a wide variation in mean annual precipitation and soil fertility. Four sites were established across a 1000 mm span in mean annual precipitation (2335 to 3300 mm). We expected microbial communities at sites with lower MAP to be less sensitive to chronic drying than sites with higher MAP; while fungal communities to be more resistant to disturbance than prokaryotes. At each location, partial throughfall exclusion structures were established over 10 x 10 m plots to reduce direct precipitation input. Raw demultiplexed sequences (bacteria, archaea, fungal) from soil samples taken from PARCHED Panama Tropical Forest throughfall exclusion experiments. Files that contain 16S are sequences from prokaryotes, ITS indicates sequences from fungi. Compressed fastq files are contained in Field_PARCHED_ITS_fastq_2020.zip, Field_PARCHED_ITS_fastq_2019.zip, Field_PARCHED_16S_fastq_2020.zip, Field_PARCHED_16S_fastq_2019.zip.There are four sites across the isthmus of Panama, each with 4 control and 4 exclusion plots. Throughfall exclusion shelters were built to intercept 50% of throughfall precipitation that hits the soil. Samples were taken on May 2019 and Jan 2020 from 0-10 cm and 10-20 cm depth approximately 9 and 18 months after shelter installation. Metadata and sample IDs for fastq files for 2019 sampling are within Field_PARCHED_Metadata_2019_ITS.csv and Field_PARCHED_Metadata_2019_16S.csv. The metadata for both 16S and ITS fastq files for January 2020 is included in Field_PARCHED_Metadata_2020.csv. No data processing or QA/QC was done on the raw data. Data processing example provided in R notebook file.

54 ENVIRONMENTAL SCIENCES↗

A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries

In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.

Qi, Pearl [Mork Family Department of Chemical Engi↗

SPARC: Structural properties associated with residue constraints

SPARC facilitates the generation of plausible hypotheses regarding underlying biochemical mechanisms by structurally characterizing protein sequence constraints. Such constraints appear as residues co-conserved in functionally related subgroups, as subtle pairwise correlations (i.e., direct couplings), and as correlations among these sequence features or with structural features. SPARC performs three types of analyses. First, based on pairwise sequence correlations, it estimates the biological relevance of alternative conformations and of homomeric contacts, as illustrated here for death domains. Second, it estimates the statistical significance of the correspondence between directly coupled residue pairs and interactions at heterodimeric interfaces. Third, given molecular dynamics simulated structures, it characterizes interactions among constrained residues or between such residues and ligands that: (a) are stably maintained during the simulation; (b) undergo correlated formation and/or disruption of interactions with other constrained residues; or (c) switch between alternative interactions. We illustrate this for two homohexameric complexes: the bacterial enhancer binding protein (bEBP) NtrC1, which activates transcription by remodeling RNA polymerase (RNAP) containing σ 54 , and for DnaB helicase, which opens DNA at the bacterial replication fork. Based on the NtrC1 analysis, we hypothesize possible mechanisms for inhibiting ATP hydrolysis until ADP is released from an adjacent subunit and for coupling ATP hydrolysis to restructuring of σ 54 binding loops. Based on the DnaB analysis, we hypothesize that DnaB ‘grabs’ ssDNA by flipping every fourth base and inserting it into cavities between subunits and that flipping of a DnaB-specific glutamine residue triggers ATP hydrolysis.

97 MATHEMATICS AND COMPUTING↗

Hierarchical Chiral Self-Assembly of Nanocylinders Composed of Sequence-Defined Mesogenic Dimers

Chiral ensembles can arise through supramolecular curvature that resolves geometric frustrations in the packing of bent, achiral molecular or colloidal building blocks. Here, we leverage orthogonal protection−deprotection click chemistry to create sequence-defined mesogenic heterodimers exhibiting emergent chirality. We compare the hierarchical self-assembly of the synthesized asymmetric, achiral heterodimers, which differ only in the position of a methyl substituent. Both dimers form chiral spherulites composed of nanocylinders. However, the detailed arrangement of nanocylinders depends on the position of the methyl substituent and the crystallization conditions. Despite the chemical similarity, in one dimer, two crystalline forms are optically active. They form conglomerates of dextrorotatory and levorotatory spherulites. The other dimer forms more highly anisotropic spherulites that mask circular birefringence arising from the misorientation of nanocylinders, while mapping of nanocylinder directors reveals a sense at the spherulite surface. We propose that differences in nanocylinder arrangements may arise from changes in nanocylinder curvature and dimensions dictated by the methyl substituent position, inducing chirality. These results demonstrate multiscale hierarchical assembly relevant to dense systems of tubular structures and highlight the role of sequence and molecular design in directing the bottom-up hierarchical self-assembly and chirality of mesogenic systems.

Alkyls↗

Sequential Stress Identifies Processing Defects in Bifacial Photovoltaic Modules That Limit Durability

Here, we use sequential stress to investigate hurdles to bifacial photovoltaic (PV) module durability from lamination defects. We test mini-modules with glass/glass (G/G) and glass/transparent-backsheet (G/TB) constructions using either ethylene vinyl acetate or polyolefin elastomer (POE) based encapsulants under a modified IEC 63209-2 sequential stress. This sequence includes multiple iterations of damp heat (DH200), full spectrum light exposure (A3), thermal cycling (TC50), and humidity/freeze (HF10). We compare indoor stress with outdoor exposure. Results show similar relative trends in degradation after a year outdoors compared to our first stress cycle. Subsequent stress cycles impart more severe damage than outdoor exposure for the short outdoor duration used here. Edge-pinch lamination defects in G/G mini-modules limit durability causing delamination and cell cracks. Conversely, we observe greater degradation in G/TB mini-modules compared to G/G in the later stages of the stress sequence when the backsheets are directly exposed to UV-containing light. Our results highlight: 1) the utility of sequential stress testing to uncover degradation modes in bifacial PV, 2) implications of using mini-modules for testing PV quality, and 3) the importance of lamination defects that must be avoided to ensure durability as the industry adopts G/G or G/TB packaging.

14 SOLAR ENERGY↗

A multifunctional sesquiterpene synthase integrates with cytochrome P450s to reinforce the terpenoid defense network in maize

Terpenoids, the largest and most structurally diverse class of plant natural products, play essential roles in maize defense and ecological interactions. In this study, we identified and functionally characterized a sesquiterpenoid-based defense pathway in maize centered on α-santalenoic acid, a pathogen-inducible sesquiterpenoid antibiotic. Using a combination of metabolite-based genome-wide association studies (mGWAS), linkage mapping, and heterologous expression assays, we identified ZmTPS9 as a multiproduct terpene synthase that primarily produces α-santalene and β-bisabolene. Sequence analysis and site-directed mutagenesis revealed that threonine at position 413 is critical for enzyme activity, with its deletion resulting in a complete loss of enzyme activity. The sesquiterpene hydrocarbons produced by ZmTPS9 are further oxidized by three cytochrome P450 monooxygenases, ZmCYP71Z16, ZmCYP71Z18, and ZmCYP71Z19, to yield antimicrobial metabolites including α-santalenoic acid, zealexin D1 (ZD1), and zealexin D2 (ZD2). Together, these findings demonstrate a convergent biosynthetic strategy in maize, where multiproduct terpene synthases and promiscuous P450s collaboratively generate a flexible and robust terpenoid defense network.

a-santalenoic acid↗

EvoProtGrad (Directed Evolution for Proteins with Gradients) [SWR-23-48]

A Python package for directed evolution on a protein sequence with gradient-based discrete Markov chain monte carlo (MCMC). Users are able to compose custom models that map sequence to function with pretrained models, including protein language models (PLMs), to guide and constrain search. Our package natively integrates with the HuggingFace platform and supports PLMs from transformers. Our MCMC sampler identifies promising amino acids to mutate via model gradients taken with respect to the input (i.e., sensitivity analysis). We allow users to compose their own custom target function for MCMC by leveraging the Product of Experts MCMC paradigm. Each model is an "expert" that contributes its own knowledge about the protein's fitness landscape to the overall target function. The sampler is designed to be more efficient and effective than brute force and random search while maintaining most of the generality and flexibility. Additional information can be found in the related publication: https://iopscience.iop.org/article/10.1088/2632-2153/accacd

Emami, Patrick↗

Quantification of Sub-Pixel Dynamics in High-Speed Neutron Imaging

The high penetration depth of neutrons through many metals and other common materials makes neutron imaging an attractive method for non-destructively probing the internal structure and dynamics of objects or systems that may not be accessible by conventional means, such as X-ray or optical imaging. While neutron imaging has been demonstrated to achieve a spatial resolution below 10 μm and temporal resolution below 10 μs, the relatively low flux of neutron sources and the limitations of existing neutron detectors have, until now, dictated that these cannot be achieved simultaneously, which substantially restricts the applicability of neutron imaging to many fields of research that could otherwise benefit from its unique capabilities. In this work, we present an attenuation modeling approach to the quantification of sub-pixel dynamics in cyclic ensemble neutron image sequences of an automotive gasoline direct injector at a 5 μs time scale with a spatial noise floor in the order of 5 μm.

47 OTHER INSTRUMENTATION↗

Image feature extraction and galaxy classification: a novel and efficient approach with automated machine learning

ABSTRACT In this work, we explore the possibility of applying machine learning methods designed for 1D problems to the task of galaxy image classification. The algorithms used for image classification typically rely on multiple costly steps, such as the point spread function deconvolution and the training and application of complex Convolutional Neural Networks of thousands or even millions of parameters. In our approach, we extract features from the galaxy images by analysing the elliptical isophotes in their light distribution and collect the information in a sequence. The sequences obtained with this method present definite features allowing a direct distinction between galaxy types. Then, we train and classify the sequences with machine learning algorithms, designed through the platform Modulos AutoML. As a demonstration of this method, we use the second public release of the Dark Energy Survey (DES DR2). We show that we are able to successfully distinguish between early-type and late-type galaxies, for images with signal-to-noise ratio greater than 300. This yields an accuracy of $86{{\ \rm per\ cent}}$ for the early-type galaxies and $93{{\ \rm per\ cent}}$ for the late-type galaxies, which is on par with most contemporary automated image classification approaches. The data dimensionality reduction of our novel method implies a significant lowering in computational cost of classification. In the perspective of future data sets obtained with e.g. Euclid and the Vera Rubin Observatory, this work represents a path towards using a well-tested and widely used platform from industry in efficiently tackling galaxy classification problems at the peta-byte scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Protein sequence design with a learned potential

The task of protein sequence design is central to nearly all rational protein engineering problems, and enormous effort has gone into the development of energy functions to guide design. Here, we investigate the capability of a deep neural network model to automate design of sequences onto protein backbones, having learned directly from crystal structure data and without any human-specified priors. The model generalizes to native topologies not seen during training, producing experimentally stable designs. We evaluate the generalizability of our method to a de novo TIM-barrel scaffold. The model produces novel sequences, and high-resolution crystal structures of two designs show excellent agreement with in silico models. Our findings demonstrate the tractability of an entirely learned method for protein sequence design.

59 BASIC BIOLOGICAL SCIENCES↗

Deterministic-Monte Carlo Hybrid Methods for Eigenvalue Sensitivity Coefficient Calculations

The TSUNAMI suite within the SCALE code package includes several methods for generating sensitivity data, including multigroup (MG) and continuous-energy (CE) capabilities. For generating sensitivities with CE data, three methods are available in SCALE 6.3.0: (1) the iterated fission probability (IFP) method with the KENO Monte Carlo transport solver, (2) IFP with the Shift Monte Carlo transport solver, and (3) the Contributon-Linked eigenvalue sensitivity/Uncertainty estimation via Tracklength importance Characterization (CLUTCH) with the KENO Monte Carlo transport solver. Currently, it is difficult to generate accurate sensitivities with large reflectors when using the CLUTCH method, specifically with fissionable and hydrogenous materials. To address this issue, the work presented herein examines a methodology to calculate the adjoint flux externally with the 3D deterministic SN transport code DENOVO in SCALE; the result is then read directly into the CLUTCH-TSUNAMI sequence. This hybridization method replaces the Monte Carlo F*(r) calculation in CLUTCH while still utilizing the forward calculation. The critical benchmark HEU-MET-FAST-028-001 is used to generate sensitivities based on the inability of CLUTCH to generate accurate sensitivities. Results from the hybrid method appear to generate sensitivity values that are in excellent agreement with direct perturbations. Although further testing is needed, the method provides promising results for the development and utility of a hybrid method for use in TSUNAMI.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

UAE6 - Wind Tunnel Tests Data - UAE6 - Sequence 5 - Raw Data

Sequence 5: Sweep Wind Speed (F,P) This test sequence used an upwind, rigid turbine with a 0° cone angle. The wind speed was ramped from 5 m/s to 25 m/s by the wind tunnel operator. This was repeated with a decreasing ramp. The yaw angle was maintained at 0°. The blade tip pitch was 3° or 6°. The rotor rotated at 72 RPM. Blade pressure and probe measurements were collected for both pitch angles. The five-hole probes were removed and the plugs were installed for another 3° pitch case. Plastic tape 0.03 mm thick was used to smooth the interface between the plugs and the blade. The teeter dampers were replaced with rigid links, and these two channels were flagged as not applicable by setting the measured values in the data file to –99999.99 Nm. The teeter link load cell was pre-tensioned to 40,000 N. During post-processing, the probe channels were set to read -99999.99. The 6- minute campaigns were named using the sequence designation 5, followed by DN or UP, which indicates the wind speed ramp direction. The next four digits are 0000, and the sequence digit is at the end.

17 WIND ENERGY↗