Engineering PapersSearch

Engineering topics

Liang, Shoudan

Publications and source records attributed to Liang, Shoudan.

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis/genetics

Predicting protein functions from redundancies in large-scale protein interaction networks

Interpreting data from large-scale protein interaction experiments has been a challenging task because of the widespread presence of random false positives. Here, we present a network-based statistical algorithm that overcomes this difficulty and allows us to derive functions of unannotated proteins from large-scale interaction data. Our algorithm uses the insight that if two proteins share significantly larger number of common interaction partners than random, they have close functional associations. Analysis of publicly available data from Saccharomyces cerevisiae reveals >2,800 reliable functional associations, 29% of which involve at least one unannotated protein. By further analyzing these associations, we derive tentative functions for 81 unannotated proteins with high certainty. Our method is not overly sensitive to the false positives present in the data. Even after adding 50% randomly generated interactions to the measured data set, we are able to recover almost all (approximately 89%) of the original associations.

Proteins/chemistry/metabolism

cWINNOWER Algorithm for Finding Fuzzy DNA Motifs

The cWINNOWER algorithm detects fuzzy motifs in DNA sequences rich in protein-binding signals. A signal is defined as any short nucleotide pattern having up to d mutations differing from a motif of length l. The algorithm finds such motifs if multiple mutated copies of the motif (i.e., the signals) are present in the DNA sequence in sufficient abundance. The cWINNOWER algorithm substantially improves the sensitivity of the winnower method of Pevzner and Sze by imposing a consensus constraint, enabling it to detect much weaker signals. We studied the minimum number of detectable motifs qc as a function of sequence length N for random sequences. We found that qc increases linearly with N for a fast version of the algorithm based on counting three-member sub-cliques. Imposing consensus constraints reduces qc, by a factor of three in this case, which makes the algorithm dramatically more sensitive. Our most sensitive algorithm, which counts four-member sub-cliques, needs a minimum of only 13 signals to detect motifs in a sequence of length N = 12000 for (l,d) = (15,4).

Liang, Shoudan

Website on Protein Interaction and Protein Structure Related Work

In today's world, three seemingly diverse fields - computer information technology, nanotechnology and biotechnology are joining forces to enlarge our scientific knowledge and solve complex technological problems. Our group is dedicated to conduct theoretical research exploring the challenges in this area. The major areas of research include: 1) Yeast Protein Interactions; 2) Protein Structures; and 3) Current Transport through Small Molecules.

Samanta, Manoj

Simple Math is Enough: Two Examples of Inferring Functional Associations from Genomic Data

Non-random features in the genomic data are usually biologically meaningful. The key is to choose the feature well. Having a p-value based score prioritizes the findings. If two proteins share a unusually large number of common interaction partners, they tend to be involved in the same biological process. We used this finding to predict the functions of 81 un-annotated proteins in yeast.

Liang, Shoudan

Systematic Underutilization of Glutamine In Thermophile Proteins

Rapid racemization above 100 C of L-amino acids to Domino acids, as well as deamidation, is probably a hazard for high temperature life. For example, the half-life of some asparaginyl peptides can be as short as 10 minutes at 100 C. High temperature organisms could protect themselves by reducing usage of amino acids that are easily racemized/deamidazed, by having a rapid rate of protein turnover which requires energy, or by adapting special cis-peptide conformations. We have searched eight completely sequenced thermophile genomes, and compare them to mesophile genomes, in order to identify underutilized amino acids. To our surprise, asparagine, the most unstable amino acid to deamidation, is used at about the same level in thermophile proteins in comparison to mesophiles whereas it is the second most unstable amino acid, glutamine, that is underutilized in all of eight thermophile species. Glutamines are present at 2% level in a typical thermophile protein, instead of 4% in mesophile. We argue that it is easier to protect asparagines from deamidation by cis-peptide conformations. We discuss statistical as well as structural evidence in support of our conclusions.

Liang, Shoudan

Research in Computational Astrobiology

We present results from several projects in the new field of computational astrobiology, which is devoted to advancing our understanding of the origin, evolution and distribution of life in the Universe using theoretical and computational tools. We have developed a procedure for calculating long-range effects in molecular dynamics using a plane wave expansion of the electrostatic potential. This method is expected to be highly efficient for simulating biological systems on massively parallel supercomputers. We have perform genomics analysis on a family of actin binding proteins. We have performed quantum mechanical calculations on carbon nanotubes and nucleic acids, which simulations will allow us to investigate possible sources of organic material on the early earth. Finally, we have developed a model of protobiological chemistry using neural networks.

Chaban, Galina

Modeling the Normal and Neoplastic Cell Cycle with 'Realistic Boolean Genetic Networks': Their Application for Understanding Carcinogenesis and Assessing Therapeutic Strategies

In this paper we show how Boolean genetic networks could be used to address complex problems in cancer biology. First, we describe a general strategy to generate Boolean genetic networks that incorporate all relevant biochemical and physiological parameters and cover all of their regulatory interactions in a deterministic manner. Second, we introduce 'realistic Boolean genetic networks' that produce time series measurements very similar to those detected in actual biological systems. Third, we outline a series of essential questions related to cancer biology and cancer therapy that could be addressed by the use of 'realistic Boolean genetic network' modeling.

Szallasi, Zoltan

Theoretical Modeling and Computer Simulations for the Origins and Evolution of Reproducing Molecular Systems and Complex Systems with Many Interactive Parts

Our research effort has produced nine publications in peer-reviewed journals listed at the end of this report. The work reported here are in the following areas: (1) genetic network modeling; (2) autocatalytic model of pre-biotic evolution; (3) theoretical and computational studies of strongly correlated electron systems; (4) reducing thermal oscillations in atomic force microscope; (5) transcription termination mechanism in prokaryotic cells; and (6) the low glutamine usage in thennophiles obtained by studying completely sequenced genomes. We discuss the main accomplishments of these publications.

Liang, Shoudan

Genetic Network Inference: From Co-Expression Clustering to Reverse Engineering

Advances in molecular biological, analytical, and computational technologies are enabling us to systematically investigate the complex molecular processes underlying biological systems. In particular, using high-throughput gene expression assays, we are able to measure the output of the gene regulatory network. We aim here to review datamining and modeling approaches for conceptualizing and unraveling the functional relationships implicit in these datasets. Clustering of co-expression profiles allows us to infer shared regulatory inputs and functional pathways. We discuss various aspects of clustering, ranging from distance measures to clustering algorithms and multiple-duster memberships. More advanced analysis aims to infer causal connections between genes directly, i.e., who is regulating whom and how. We discuss several approaches to the problem of reverse engineering of genetic networks, from discrete Boolean networks, to continuous linear and non-linear models. We conclude that the combination of predictive modeling with systematic experimental verification will be required to gain a deeper insight into living organisms, therapeutic targeting, and bioengineering.

Dhaeseleer, Patrik

Phase Diagram of the Two-Chain Hubbard Model

We have calculated the charge gap and spin gap for the two-chain Hubbard model as a function of the on-site Coulomb interaction and the interchain hopping amplitude. We used the density matrix renormalization group method and developed a method to calculate separately the gaps numerically for the symmetric and antisymmetric modes with respect to the exchange of the chain indices. We have found very different behaviors for the weak and strong interaction cases. Our calculated phase diagram is compared to the one obtained by Balents and Fisher using the weak coupling renormalization group technique.

Park, Youngho

Spatial Autocatalytic Dynamics: An Approach to Modeling Prebiotic Evolution

This paper addresses the origin of robust and evolvable metabolic functions, and the conditions under which it took place. We propose that spatial considerations, traditionally ignored, are essential to answering these important questions in prebiotic evolution. Our probabilistic cellular automaton model, based on work on autocatalytic metabolisms by Eigen, Kauffman, and others, has biologically interesting dynamical behavior that is missed if spatial extension is ignored.

Stassinopoulos, Dimitris

Charge and Spin Dynamics of the Hubbard Chains

We calculate the local correlation functions of charge and spin for the one-chain and two-chain Hubbard model using density matrix renormalization group method and the recursion technique. Keeping only finite number of states we get good accuracy for the low energy excitations. We study the charge and spin gaps, bandwidths and weights of the spectra for various values of the on-site Coulomb interaction U and the electron filling. In the low energy part, the local correlation functions are different for the charge and spin. The bandwidths are proportional to t for the charge and J for the spin respectively.

Park, Youngho

Thermal Noise Reduction of Mechanical Oscillators by Actively Controlled External Dissipative Forces

We show that the thermal fluctuations of very soft mechanical oscillators, such as the cantilever in an atomic force microscope (AFM), can be reduced without changing the stiffness of the spring or having to lower the environment temperature. We derive a theoretical relationship between the thermal fluctuations of an oscillator and an actively external-dissipative force. This relationship is verified by experiments with an AFM cantilever where the external active force is coupled through a magnetic field. With simple instrumentation, we have reduced the thermal noise amplitude of the cantilever by a factor of 3.4, achieving an apparent temperature of 25 K with the environment at 295K. This active noise reduction approach can significantly improve the accuracy of static position or static force measurements in a number of practical applications.

Liang, Shoudan

Reveal, A General Reverse Engineering Algorithm for Inference of Genetic Network Architectures

Given the immanent gene expression mapping covering whole genomes during development, health and disease, we seek computational methods to maximize functional inference from such large data sets. Is it possible, in principle, to completely infer a complex regulatory network architecture from input/output patterns of its variables? We investigated this possibility using binary models of genetic networks. Trajectories, or state transition tables of Boolean nets, resemble time series of gene expression. By systematically analyzing the mutual information between input states and output states, one is able to infer the sets of input elements controlling each element or gene in the network. This process is unequivocal and exact for complete state transition tables. We implemented this REVerse Engineering ALgorithm (REVEAL) in a C program, and found the problem to be tractable within the conditions tested so far. For n = 50 (elements) and k = 3 (inputs per element), the analysis of incomplete state transition tables (100 state transition pairs out of a possible 10(exp 15)) reliably produced the original rule and wiring sets. While this study is limited to synchronous Boolean networks, the algorithm is generalizable to include multi-state models, essentially allowing direct application to realistic biological data sets. The ability to adequately solve the inverse problem may enable in-depth analysis of complex dynamic systems in biology and other fields.

Liang, Shoudan