Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Direct Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Whole genome sequencing of Mycobacterium bovis directly from clinical tissue samples without culture

Advancement in next generation sequencing offers the possibility of routine use of whole genome sequencing (WGS) for Mycobacterium bovis (M. bovis) genomes in clinical reference laboratories. To date, the M. bovis genome could only be sequenced if the mycobacteria were cultured from tissue. This requirement for culture has been due to the overwhelmingly large amount of host DNA present when DNA is prepared directly from a granuloma. To overcome this formidable hurdle, we evaluated the usefulness of an RNA-based targeted enrichment method to sequence M. bovis DNA directly from tissue samples without culture. Initial spiking experiments for method development were established by spiking DNA extracted from tissue samples with serially diluted M. bovis BCG DNA at the following concentration range: 0.1 ng/μl to 0.1 pg/μl (10 –1 to 10 –4 ). Library preparation, hybridization and enrichment was performed using SureSelect custom capture library RNA baits and the SureSelect XT HS2 target enrichment system for Illumina paired-end sequencing. The method validation was then assessed using direct WGS of M. bovis DNA extracted from tissue samples from naturally (n = 6) and experimentally (n = 6) infected animals with variable Ct values. Direct WGS of spiked DNA samples achieved 99.1% mean genome coverage (mean depth of coverage: 108×) and 98.8% mean genome coverage (mean depth of coverage: 26.4×) for tissue samples spiked with BCG DNA at 10 –1 (mean Ct value: 20.3) and 10 –2 (mean Ct value: 23.4), respectively. The M. bovis genome from the experimentally and naturally infected tissue samples was successfully sequenced with a mean genome coverage of 99.56% and depth of genome coverage ranging from 9.2× to 72.1×. The spoligoyping and M. bovis group assignment derived from sequencing DNA directly from the infected tissue samples matched that of the cultured isolates from the same sample. Our results show that direct sequencing of M. bovis DNA from tissue samples has the potential to provide accurate sequencing of M. bovis genomes significantly faster than WGS from cultures in research and diagnostic settings.

59 BASIC BIOLOGICAL SCIENCES↗

Human RNome Project draft human RNome sequence of GM12878, B-cell line, obtained by mass-spectrometry sequencing, long-read sequencing and short-read sequencing.

Here we report the first draft of the human RNome sequence, a reference map of RNA chemical modifications in a human B-cell line. RNA carries a diverse repertoire of chemical modifications that regulate gene expression, cellular function, and responses to physiological and pathological cues. Yet, unlike the genome, no reference map of RNA modifications is available for any human cell. To generate this resource, the Human RNome Project Consortium analyzed a shared RNA preparation from the well-characterized GM12878 B-cell line using short-read sequencing, long-read direct RNA sequencing, and mass spectrometry, generating more than 7.1 billion sequencing reads spanning approximately 1.2 trillion nucleotides. The resulting maps of the human RNome reveal that RNA modifications are organized according to function, transcript architecture, and cellular identity. Modifications concentrate at functional centers of ribosomal and transfer RNAs, follow the canonical topology of N6-methyladenosine in coding transcripts, and form coordinated hotspots in immune regulatory genes. This first reference human RNome provides a foundation for understanding how RNA chemistry shapes cellular identity, human disease, and the development of RNA-based therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

Nonenzymatic template-directed synthesis on oligodeoxycytidylate sequences in hairpin oligonucleotides

We have developed a novel method for studying template-directed synthesis in hairpin oligonucleotides. An unpaired segment at the 5'-terminus of the hairpin acts as an intramolecular template for the extension of the paired 3'-terminus. Products are analyzed by denaturing gel electrophoresis of [32P]-labeled hairpins. Using this system, we have studied the synthesis of oligoguanylates on an oligodeoxycytidylate template. We find that guanosine 5'-phosphoro(2-methyl)imidazolide adds efficiently to a terminal riboguanylate residue at temperatures in the range 0-37 degrees C but not at 50 degrees C. At 0 degree C, the half-time for addition of the first G residue is about 3 h, and the reaction rate is independent of pH in the range 6.5-8.0. The first addition reaction results in the formation of a predominantly 3'-5'-internucleotide bond. When the 3'-terminal riboguanylate residue is placed by a deoxyguanylate residue, the half-time for the first addition increases from about 3 to about 30 h.

NASA Discipline Exobiology↗

Engine With Regression and Neural Network Approximators Designed

At the NASA Glenn Research Center, the NASA engine performance program (NEPP, ref. 1) and the design optimization testbed COMETBOARDS (ref. 2) with regression and neural network analysis-approximators have been coupled to obtain a preliminary engine design methodology. The solution to a high-bypass-ratio subsonic waverotor-topped turbofan engine, which is shown in the preceding figure, was obtained by the simulation depicted in the following figure. This engine is made of 16 components mounted on two shafts with 21 flow stations. The engine is designed for a flight envelope with 47 operating points. The design optimization utilized both neural network and regression approximations, along with the cascade strategy (ref. 3). The cascade used three algorithms in sequence: the method of feasible directions, the sequence of unconstrained minimizations technique, and sequential quadratic programming. The normalized optimum thrusts obtained by the three methods are shown in the following figure: the cascade algorithm with regression approximation is represented by a triangle, a circle is shown for the neural network solution, and a solid line indicates original NEPP results. The solutions obtained from both approximate methods lie within one standard deviation of the benchmark solution for each operating point. The simulation improved the maximum thrust by 5 percent. The performance of the linear regression and neural network methods as alternate engine analyzers was found to be satisfactory for the analysis and operation optimization of air-breathing propulsion engines (ref. 4).

Patnaik, Surya N.↗

Identification of defective illegitimate recombinational repair of oxidatively-induced DNA double-strand breaks in ataxia-telangiectasia cells

Ataxia-telangiectasia (A-T) is an autosomal-recessive lethal human disease. Homozygotes suffer from a number of neurological disorders, as well as very high cancer incidence. Heterozygotes may also have a higher than normal risk of cancer, particularly for the breast. The gene responsible for the disease (ATM) has been cloned, but its role in mechanisms of the disease remain unknown. Cellular A-T phenotypes, such as radiosensitivity and genomic instability, suggest that a deficiency in the repair of DNA double-strand breaks (DSBs) may be the primary defect; however, overall levels of DSB rejoining appear normal. We used the shuttle vector, pZ189, containing an oxidatively-induced DSB, to compare the integrity of DSB rejoining in one normal and two A-T fibroblast cells lines. Mutation frequencies were two-fold higher in A-T cells, and the mutational spectrum was different. The majority of the mutations found in all three cell lines were deletions (44-63%). The DNA sequence analysis indicated that 17 of the 17 plasmids with deletion mutations in normal cells occurred between short direct-repeat sequences (removing one of the repeats plus the intervening sequences), implicating illegitimate recombination in DSB rejoining. The combined data from both A-T cell lines showed that 21 of 24 deletions did not involve direct-repeats sequences, implicating a defect in the illegitimate recombination pathway. These findings suggest that the A-T gene product may either directly participate in illegitimate recombination or modulate the pathway. Regardless, this defect is likely to be important to a mechanistic understanding of this lethal disease.

Non-NASA Center↗

On blue straggler information by direct collisions of main sequence stars

We report the results of new smoothed particle hydrodynamics calculations of parabolic collisions between main-sequence (MS) stars. The stars are assumed to be close the MS turnoff point in a globular cluster and are therefore modeled as n = 3, Gamma = 5/3 polytropes. We find that the high degree of central mass concentration in these stars has a profound effect on the hydrodynamics. In particular, very little hydrodynamic mixing occurs between the dense, helium-rich inner cores and the outer envelopes. As a result, and in contrast to what has been assumed in previous studies, blue stragglers formed by direct stellar collisions are not necessarily expected to have anomalously high helium abundances in their envelopes or to have their cores replenished with fresh hydrogen fuel.

Lombardi, James, C. jr.↗

Decoding co-/post-transcriptional complexities of plant transcriptomes and epitranscriptome using next-generation sequencing technologies

Next-generation sequencing (NGS) technologies - Illumina RNA-seq, Pacific Biosciences isoform sequencing (PacBio Iso-seq), and Oxford Nanopore direct RNA sequencing (DRS) - have revealed the complexity of plant transcriptomes and their regulation at the co-/post-transcriptional level. Global analysis of mature mRNAs, transcripts from nuclear run-on assays, and nascent chromatin-bound mRNAs using short as well as full-length and single-molecule DRS reads have uncovered potential roles of different forms of RNA polymerase II during the transcription process, and the extent of co-transcriptional pre-mRNA splicing and polyadenylation. These tools have also allowed mapping of transcriptome-wide start sites in cap-containing RNAs, poly(A) site choice, poly(A) tail length, and RNA base modifications. The emerging theme from recent studies is that reprogramming of gene expression in response to developmental cues and stresses at the co-/post-transcriptional level likely plays a crucial role in eliciting appropriate responses for optimal growth and plant survival under adverse conditions. Although the mechanisms by which developmental cues and different stresses regulate co-/post-transcriptional splicing are largely unknown, a few recent studies indicate that the external cues target spliceosomal and splicing regulatory proteins to modulate alternative splicing. In this review, we provide an overview of recent discoveries on the dynamics and complexities of plant transcriptomes, mechanistic insights into splicing regulation, and discuss critical gaps in co-/post-transcriptional research that need to be addressed using diverse genomic and biochemical approaches.

Biochemistry & Molecular Biology↗

Novel End-to-End Molecular Biology Approach for Direct Nanopore 1D cDNA Sequencing of Reverse Transcribed mRNAs Purified from Cell Cultures by the NASA ISS WetLab2 SPM

Continued space bioscience research onboard the International Space Station (ISS) and future long-duration flight missions to the Moon or Mars will require the ability to conduct on-orbit molecular analysis of biological samples independently from Earth. In the last year two new molecular analytic technologies have been installed and the technologies demonstrated onboard the ISS: The Sample Prep Module (SPM) WetLab-2 (WL2) qRT-PCR toolbox and the Oxford Nanopore MinIon Biomolecule Sequencer. Here we describe protocol development and integration into existing ISS technology for end-to-end on-orbit biological sample processing and molecular analysis with real time results generated utilizing only field offline analytic software. For this experiment we isolated primary cells from bone marrow flushes of wild type B6129SF2 mice (Jackson Labs) long bones. The cell isolate was then processed using the SPM to produce total 147nanograms of RNA. The total RNA was purified to only messenger RNA (mRNA) and transferred to Smartcycler Thermocycle ISS kit consumable tube using Eppendorf gel loading pipette tips for further processing. Complementary first strand cDNA was synthesized using OLIGO dT priming followed by addition of SuperScript II Reverse Transcriptase and thermal cycling as per manufacturers instruction. All thermal cycling was conducted using the ISS WetLab-2 Cephid Smarcycler real time thermal cycler. Our protocol takes advantage of mRNAs native poly(A) tail, synthesized in vivo to protect the mRNA from degradation by endonucleases, to eliminate end-prep for adapter ligation. The adapted library is purified using MyOne C1 Streptavidin beads before elution in buffer. The pre-sequencing library is diluted in the loading buffer and injected into the MinIon sample port, drawn into the nanopore window by capillary action, and sequenced using the MinKnown software with local basecalling. The sequencing read produced 34.5 million events and local basecalling produced 117,301 successful reads. NCBI Blast of the data for the mouse genome resulted in 2,462 successful nucleotide collection matches (gene sequences) exceeding 70 homology. These results demonstrate the viability of this novel flight ready end-to-end sample analytic methodology and provide a real time homolog for flight experimentation utilizing supply kits and technologies that have already been demonstrated on ISS.

MinIon↗

Implied alignment: a synapomorphy-based multiple-sequence alignment method and its use in cladogram search

A method to align sequence data based on parsimonious synapomorphy schemes generated by direct optimization (DO; earlier termed optimization alignment) is proposed. DO directly diagnoses sequence data on cladograms without an intervening multiple-alignment step, thereby creating topology-specific, dynamic homology statements. Hence, no multiple-alignment is required to generate cladograms. Unlike general and globally optimal multiple-alignment procedures, the method described here, implied alignment (IA), takes these dynamic homologies and traces them back through a single cladogram, linking the unaligned sequence positions in the terminal taxa via DO transformation series. These "lines of correspondence" link ancestor-descendent states and, when displayed as linearly arrayed columns without hypothetical ancestors, are largely indistinguishable from standard multiple alignment. Since this method is based on synapomorphy, the treatment of certain classes of insertion-deletion (indel) events may be different from that of other alignment procedures. As with all alignment methods, results are dependent on parameter assumptions such as indel cost and transversion:transition ratios. Such an IA could be used as a basis for phylogenetic search, but this would be questionable since the homologies derived from the implied alignment depend on its natal cladogram and any variance, between DO and IA + Search, due to heuristic approach. The utility of this procedure in heuristic cladogram searches using DO and the improvement of heuristic cladogram cost calculations are discussed. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

Non-NASA Center↗

Biosensors for DNA sequence detection

DNA biosensors are being developed as alternatives to conventional DNA microarrays. These devices couple signal transduction directly to sequence recognition. Some of the most sensitive and functional technologies use fibre optics or electrochemical sensors in combination with DNA hybridization. In a shift from sequence recognition by hybridization, two emerging single-molecule techniques read sequence composition using zero-mode waveguides or electrical impedance in nanoscale pores.

Review↗

Rapid Diagnostics of Onboard Sequences

Keeping track of sequences onboard a spacecraft is challenging. When reviewing Event Verification Records (EVRs) of sequence executions on the Mars Exploration Rover (MER), operators often found themselves wondering which version of a named sequence the EVR corresponded to. The lack of this information drastically impacts the operators diagnostic capabilities as well as their situational awareness with respect to the commands the spacecraft has executed, since the EVRs do not provide argument values or explanatory comments. Having this information immediately available can be instrumental in diagnosing critical events and can significantly enhance the overall safety of the spacecraft. This software provides auditing capability that can eliminate that uncertainty while diagnosing critical conditions. Furthermore, the Restful interface provides a simple way for sequencing tools to automatically retrieve binary compiled sequence SCMFs (Space Command Message Files) on demand. It also enables developers to change the underlying database, while maintaining the same interface to the existing applications. The logging capabilities are also beneficial to operators when they are trying to recall how they solved a similar problem many days ago: this software enables automatic recovery of SCMF and RML (Robot Markup Language) sequence files directly from the command EVRs, eliminating the need for people to find and validate the corresponding sequences. To address the lack of auditing capability for sequences onboard a spacecraft during earlier missions, extensive logging support was added on the Mars Science Laboratory (MSL) sequencing server. This server is responsible for generating all MSL binary SCMFs from RML input sequences. The sequencing server logs every SCMF it generates into a MySQL database, as well as the high-level RML file and dictionary name inputs used to create the SCMF. The SCMF is then indexed by a hash value that is automatically included in all command EVRs by the onboard flight software. Second, both the binary SCMF result and the RML input file can be retrieved simply by specifying the hash to a Restful web interface. This interface enables command line tools as well as large sophisticated programs to download the SCMF and RMLs on-demand from the database, enabling a vast array of tools to be built on top of it. One such command line tool can retrieve and display RML files, or annotate a list of EVRs by interleaving them with the original sequence commands. This software has been integrated with the MSL sequencing pipeline where it will serve sequences useful in diagnostics, debugging, and situational awareness throughout the mission.

Starbird, Thomas W.↗

Protein remote homology detection and structural alignment using deep learning

Exploiting sequence–structure–function relationships in biotechnology requires improved methods for aligning proteins that have low sequence similarity to previously annotated proteins. We develop two deep learning methods to address this gap, TM-Vec and DeepBLAST. TM-Vec allows searching for structure–structure similarities in large sequence databases. It is trained to accurately predict TM-scores as a metric of structural similarity directly from sequence pairs without the need for intermediate computation or solution of structures. Once structurally similar proteins have been identified, DeepBLAST can structurally align proteins using only sequence information by identifying structurally homologous regions between proteins. It outperforms traditional sequence alignment methods and performs similarly to structure-based alignment methods. We show the merits of TM-Vec and DeepBLAST on a variety of datasets, including better identification of remotely homologous proteins compared with state-of-the-art sequence alignment and structure prediction methods.

59 BASIC BIOLOGICAL SCIENCES↗

Io's sodium directional features - Evidence for a magnetospheric-wind-driven gas escape mechanism

Elongated features in Io's sodium cloud, directed away from Jupiter and inclined both to the north and to the south of the satellite's orbital plane, have been observed. The north/south directions of the features are correlated with Io's magnetic longitude, suggesting a formation mechanism involving the oscillating plasma torus. It is shown by means of a model analysis that the features can result from a source of high-velocity (about 20 km/s) sodium combined with the oscillating neutral sodium sink provided by the plasma. The phase relationship between the features' directions and Io's magnetic longitude can be understood if escaping sodium is initially directed at near right angles to Io's orbital motion. The directionality of the features requires that the sodium flux from equatorial regions be higher than that from the poles. The initial directions and speeds of sodium atoms escaping Io to form the directional features can be understood in terms of a magnetospheric-wind-driven escape mechanism. The one sequence of directional feature observations that has been analyzed in detail implies a high-speed sodium source rate of about 10 to the 26th atoms/s.

Pilcher, C. B.↗

Deletions at short direct repeats and base substitutions are characteristic mutations for bleomycin-induced double- and single-strand breaks, respectively, in a human shuttle vector system

Using the radiomimetic drug, bleomycin, we have determined the mutagenic potential of DNA strand breaks in the shuttle vector pZ189 in human fibroblasts. The bleomycin treatment conditions used produce strand breaks with 3'-phosphoglycolate termini as > 95% of the detectable dose-dependent lesions. Breaks with this end group represent 50% of the strand break damage produced by ionizing radiation. We report that such strand breaks are mutagenic lesions. The type of mutation produced is largely determined by the type of strand break on the plasmid (i.e. single versus double). Mutagenesis studies with purified DNA forms showed that nicked plasmids (i.e. those containing single-strand breaks) predominantly produce base substitutions, the majority of which are multiples, which presumably originate from error-prone polymerase activity at strand break sites. In contrast, repair of linear plasmids (i.e. those containing double-strand breaks) mainly results in deletions at short direct repeat sequences, indicating the involvement of illegitimate recombination. The data characterize the nature of mutations produced by single- and double-strand breaks in human cells, and suggests that deletions at direct repeats may be a 'signature' mutation for the processing of DNA double-strand breaks.

NASA Discipline Radiation Health↗

Data for Comparison of Genotyping Assays for Detection of Targeted CRISPR/Cas Mutagenesis in Highly Polyploid Sugarcane

Sugarcane ( Saccharum spp.) is an important biofuel feedstock and a leading source of global table sugar. Saccharum hybrid cultivars are highly polyploid (2n = 100–130), containing large numbers of functionally redundant hom(e)ologs in their genomes. Genome editing with sequence-specific nucleases holds tremendous promise for sugarcane breeding. However, identification of plants with the desired level of co-editing within a pool of primary transformants can be difficult. While DNA sequencing provides direct evidence of targeted mutagenesis, it is cost-prohibitive as a primary screening method in sugarcane and most other methods of identifying mutant lines have not been optimized for use in highly polyploid species. In this study, non-sequencing methods of mutant screening, including capillary electrophoresis (CE), Cas9 RNP assay, and high-resolution melt analysis (HRMA), were compared to assess their potential for CRISPR/Cas9-mediated mutant screening in sugarcane. These assays were used to analyze sugarcane lines containing mutations at one or more of six sgRNA target sites. All three methods distinguished edited lines from wild type, with co-mutation frequencies ranging from 2% to 100%. Cas9 RNP assays were able to identify mutant sugarcane lines with as low as 3.2% co-mutation frequency, and samples could be scored based on undigested band intensity. CE was highlighted as the most comprehensive assay, delivering precise information on both mutagenesis frequency and indel size to a 1 bp resolution across all six targets. This represents an economical and comprehensive alternative to sequencing-based genotyping methods which could be applied in other polyploid species.

Genomics↗