Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Sequence Function Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES

Data for "Design of Diverse, Functional Mitochondrial Targeting Sequences Across Eukaryotic Organisms Using Variational Autoencoder"

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

AI/ML

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab

Implementation of a blade element UH-60 helicopter simulation on a parallel computer architecture in real-time

A high-performance platform for development of real-time helicopter flight simulations based on a simulation development and analysis platform combining a parallel simulation development and analysis environment with a scalable multiprocessor computer system is described. Simulation functional decomposition is covered, including the sequencing and data dependency of simulation modules and simulation functional mapping to multiple processors. The multiprocessor-based implementation of a blade-element simulation of the UH-60 helicopter is presented, and a prototype developed for a TC2000 computer is generalized in order to arrive at a portable multiprocessor software architecture. It is pointed out that the proposed approach coupled with a pilot's station creates a setting in which simulation engineers, computer scientists, and pilots can work together in the design and evaluation of advanced real-time helicopter simulations.

Moxon, Bruce C.

Monitoring technology

A process for infrared spectroscopic monitoring of insitu compositional changes in a polymeric material comprises the steps of providing an elongated infrared radiation transmitting fiber that has a transmission portion and a sensor portion, embedding the sensor portion in the polymeric material to be monitored, subjecting the polymeric material to a processing sequence, applying a beam of infrared radiation to the fiber for transmission through the transmitting portion to the sensor portion for modification as a function of properties of the polymeric material, monitoring the modified infrared radiation spectra as the polymeric material is being subjected to the processing sequence to obtain kinetic data on changes in the polymeric material during the processing sequence, and adjusting the processing sequence as a function of the kinetic data provided by the modified infrared radiation spectra information.

Stevenson, William A.

Monitoring technology

A process for infrared spectroscopic monitoring of insitu compositional changes in a polymeric material comprises the steps of providing an elongated infrared radiation transmitting fiber that has a transmission portion and a sensor portion, embedding the sensor portion in the polymeric material to be monitored, subjecting the polymeric material to a processing sequence, applying a beam of infrared radiation to the fiber for transmission through the transmitting portion to the sensor portion for modification as a function of properties of the polymeric material, monitoring the modified infrared radiation spectra as the polymeric material is being subjected to the processing sequence to obtain kinetic data on changes in the polymeric material during the processing sequence, and adjusting the processing sequence as a function of the kinetic data provided by the modified infrared radiation spectra information.

Stevenson, William A.

Biomolecule Sequencer: Next-Generation DNA Sequencing Technology for In-Flight Environmental Monitoring, Research, and Beyond

On the International Space Station (ISS), technologies capable of rapid microbial identification and disease diagnostics are not currently available. NASA still relies upon sample return for comprehensive, molecular-based sample characterization. Next-generation DNA sequencing is a powerful approach for identifying microorganisms in air, water, and surfaces onboard spacecraft. The Biomolecule Sequencer payload, manifested to SpaceX-9 and scheduled on the Increment 4748 research plan (June 2016), will assess the functionality of a commercially-available next-generation DNA sequencer in the microgravity environment of ISS. The MinION device from Oxford Nanopore Technologies (Oxford, UK) measures picoamp changes in electrical current dependent on nucleotide sequences of the DNA strand migrating through nanopores in the system. The hardware is exceptionally small (9.5 x 3.2 x 1.6 cm), lightweight (120 grams), and powered only by a USB connection. For the ISS technology demonstration, the Biomolecule Sequencer will be powered by a Microsoft Surface Pro3. Ground-prepared samples containing lambda bacteriophage, Escherichia coli, and mouse genomic DNA, will be launched and stored frozen on the ISS until experiment initiation. Immediately prior to sequencing, a crew member will collect and thaw frozen DNA samples, connect the sequencer to the Surface Pro3, inject thawed samples into a MinION flow cell, and initiate sequencing. At the completion of the sequencing run, data will be downlinked for ground analysis. Identical, synchronous ground controls will be used for data comparisons to determine sequencer functionality, run-time sequence, current dynamics, and overall accuracy. We will present our latest results from the ISS flight experiment the first time DNA has ever been sequenced in space and discuss the many potential applications of the Biomolecule Sequencer for environmental monitoring, medical diagnostics, higher fidelity and more adaptable Space Biology Human Research Program investigations, and even life detection experiments for astrobiology missions.

Sequencer

A lignin-specific peroxidase in tobacco whose antisense suppression leads to vascular tissue modification

A tobacco peroxidase isoenzyme (TP60) was down-regulated in tobacco using an antisense strategy, this affording transformants with lignin reductions of up to 40-50% of wild type (control) plants. Significantly, both guaiacyl and syringyl levels decreased in essentially a linear manner with the reductions in lignin amounts, as determined by both thioacidolysis and nitrobenzene oxidative analyses. These data provisionally suggest that a feedback mechanism is operative in lignifying cells, which prevents build-up of monolignols should oxidative capacity for their subsequent metabolism be reduced. Prior to this study, the only known rate-limiting processes in the monolignol/lignin pathways involved that of Phe supply and the relative activities of cinnamate-4-hydroxylase/p-coumarate-3-hydroxylase, respectively. These transformants thus provide an additional experimental means in which to further dissect and delineate the factors involved in monolignol targeting to precise regions in the cell wall, and of subsequent lignin assembly. Interestingly, the lignin down-regulated tobacco phenotypes displayed no readily observable differences in overall growth and development profiles, although the vascular apparatus was modified.

NASA Program Fundamental Space Biology

Origin of the Eumetazoa: testing ecological predictions of molecular clocks against the Proterozoic fossil record

Molecular clocks have the potential to shed light on the timing of early metazoan divergences, but differing algorithms and calibration points yield conspicuously discordant results. We argue here that competing molecular clock hypotheses should be testable in the fossil record, on the principle that fundamentally new grades of animal organization will have ecosystem-wide impacts. Using a set of seven nuclear-encoded protein sequences, we demonstrate the paraphyly of Porifera and calculate sponge/eumetazoan and cnidarian/bilaterian divergence times by using both distance [minimum evolution (ME)] and maximum likelihood (ML) molecular clocks; ME brackets the appearance of Eumetazoa between 634 and 604 Ma, whereas ML suggests it was between 867 and 748 Ma. Significantly, the ME, but not the ML, estimate is coincident with a major regime change in the Proterozoic acritarch record, including: (i) disappearance of low-diversity, evolutionarily static, pre-Ediacaran acanthomorphs; (ii) radiation of the high-diversity, short-lived Doushantuo-Pertatataka microbiota; and (iii) an order-of-magnitude increase in evolutionary turnover rate. We interpret this turnover as a consequence of the novel ecological challenges accompanying the evolution of the eumetazoan nervous system and gut. Thus, the more readily preserved microfossil record provides positive evidence for the absence of pre-Ediacaran eumetazoans and strongly supports the veracity, and therefore more general application, of the ME molecular clock.

NASA Discipline Evolutionary Biology

Neuromorphic learning of continuous-valued mappings from noise-corrupted data. Application to real-time adaptive control

The ability of feed-forward neural network architectures to learn continuous valued mappings in the presence of noise was demonstrated in relation to parameter identification and real-time adaptive control applications. An error function was introduced to help optimize parameter values such as number of training iterations, observation time, sampling rate, and scaling of the control signal. The learning performance depended essentially on the degree of embodiment of the control law in the training data set and on the degree of uniformity of the probability distribution function of the data that are presented to the net during sequence. When a control law was corrupted by noise, the fluctuations of the training data biased the probability distribution function of the training data sequence. Only if the noise contamination is minimized and the degree of embodiment of the control law is maximized, can a neural net develop a good representation of the mapping and be used as a neurocontroller. A multilayer net was trained with back-error-propagation to control a cart-pole system for linear and nonlinear control laws in the presence of data processing noise and measurement noise. The neurocontroller exhibited noise-filtering properties and was found to operate more smoothly than the teacher in the presence of measurement noise.

Troudet, Terry

DIALOG: An executive computer program for linking independent programs

A very large scale computer programming procedure called the DIALOG Executive System has been developed for the Univac 1100 series computers. The executive computer program, DIALOG, controls the sequence of execution and data management function for a library of independent computer programs. Communication of common information is accomplished by DIALOG through a dynamically constructed and maintained data base of common information. The unique feature of the DIALOG Executive System is the manner in which computer programs are linked. Each program maintains its individual identity and as such is unaware of its contribution to the large scale program. This feature makes any computer program a candidate for use with the DIALOG Executive System. The installation and use of the DIALOG Executive System are described at Johnson Space Center.

Glatt, C. R.

DIALOG: An executive computer program for linking independent programs

A very large scale computer programming procedure called the DIALOG executive system was developed for the CDC 6000 series computers. The executive computer program, DIALOG, controls the sequence of execution and data management function for a library of independent computer programs. Communication of common information is accomplished by DIALOG through a dynamically constructed and maintained data base of common information. Each computer program maintains its individual identity and is unaware of its contribution to the large scale program. This feature makes any computer program a candidate for use with the DIALOG executive system. The installation and uses of the DIALOG executive system are described.

Glatt, C. R.

ALTKAL: An optimum linear filter for GEOS-3 altimeter data

ALTKAL is a computer program designed to smooth sea surface height data obtained from the GEOS 3 altimeter, and to produce minimum variance estimates of sea surface height and sea surface slopes, along with their standard derivations. The program operates by processing the data through a Kalman filter in both the forward and backward directions, and optimally combining the results. The sea surface height signal is considered to have a geoid signal, modeled by a third order Gauss-Markov process, corrupted by additive white noise. The governing parameters for the signal and noise processes are the signal correlation length and the signal-to-noise ratio. Mathematical derivations of the filtering and smoothing algorithms are presented. The smoother characteristics are illustrated by giving the frequency response, the data weighting sequence and the transfer function of a realistic steady-state smoother example. Based on nominal estimates for geoidal undulation amplitude and correlation length, standard deviations for the estimated sea surface height and slope are 12 cm and 3 arc seconds, respectively.

Fang, B. T.

Functional language and data flow architectures

This is a tutorial article about language and architecture approaches for highly concurrent computer systems based on the functional style of programming. The discussion concentrates on the basic aspects of functional languages, and sequencing models such as data-flow, demand-driven and reduction which are essential at the machine organization level. Several examples of highly concurrent machines are described.

Ercegovac, M. D.

Discovering methylated DNA motifs in bacterial nanopore sequencing data with MIJAMP

Abstract Bacterial DNA methylation is involved in diverse cellular functions, including modulation of gene expression, DNA repair, and restriction–modification systems for defense against viruses and other foreign DNA. Restriction systems hinder efforts to engineer organisms to produce fuels and chemicals from waste and renewable feedstocks by degrading DNA during transformation. Methylome analysis allows identification of motifs within a bacterial chromosome that may be targeted by native restriction enzymes. Further expression of the corresponding methyltransferases in Escherichia coli allows plasmid DNA to be protected from restriction in the target organism, thereby drastically enhancing transformation efficiency. Nanopore sequencing can detect methylated bases, but software is needed to transform modified base coordinates into methylated motifs. Here, we develop MIJAMP (MIJAMP Is Just A MethylBED Parser), a software package that was developed to discover methylated motifs from the output of ONT’s Modkit or other data in the methylBED format. MIJAMP employs a human-driven refinement strategy that empirically validates all motifs against genome-wide methylation data, thus eliminating incorrect motifs. MIJAMP also reports methylation data on specific, user-defined motifs. Using MIJAMP, we determined the methylated motifs both in a control strain (wild-type E. coli) and in Synecococcus sp. strain PCC7002, laying the foundation for improved transformation in this organism. MIJAMP is available at https://code.ornl.gov/alexander-public/mijamp/. One Sentence Summary: Here we describe software written to discover DNA methylation motifs from nanopore sequencing data.

59 BASIC BIOLOGICAL SCIENCES

GRC MILab Software: Quick Start Guide

This document provides detailed installation and operating instructions for the GRC MILab Excel Add-In software developed at the NASA Glenn Research Center. The software described has been implemented to facilitate the process of importing into Microsoft Excel and analyzing materials test data from a wide range of materials tests. All resulting data is then ready for automated upload to the relevant table of the GRC Materials Intelligence (MI) database. This new software represents an update to the original MILab software developed by Granta Design Ltd.—a company specializing in materials software, data, and databases—for members of the Materials Data Management Consortium (MDMC), a collaboration between Granta, ASM International, NASA Glenn, and several other materials-oriented corporations and government agencies in the aerospace and defense industries. The updated software consists of the addition of two test type modules, the Generic and Generic Cyclic modules, with both representing a generalization of the original software. The Generic module supports the import and analysis of multiaxial data from any sequence of tensile, compression, relaxation, and/or creep test stages; and the Generic Cyclic module expands the functionality to include repeated sequences during cyclic testing. During processing, all imported data and analysis results are formatted by the software so as to be ready for immediate automated upload to the MI database, ensuring minimal overhead on the part of the user and access to persistent and reliable data for all relevant personnel.

Quick Start Guide