Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Complex Network Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Distributing Data to Hand-Held Devices in a Wireless Network

ADROIT is a developmental computer program for real-time distribution of complex data streams for display on Web-enabled, portable terminals held by members of an operational team of a spacecraft-command-and-control center who may be located away from the center. Examples of such terminals include personal data assistants, laptop computers, and cellular telephones. ADROIT would make it unnecessary to equip each terminal with platform- specific software for access to the data streams or with software that implements the information-sharing protocol used to deliver telemetry data to clients in the center. ADROIT is a combination of middleware plus software specific to the center. (Middleware enables one application program to communicate with another by performing such functions as conversion, translation, consolidation, and/or integration.) ADROIT translates a data stream (voice, video, or alphanumerical data) from the center into Extensible Markup Language, effectuates a subscription process to determine who gets what data when, and presents the data to each user in real time. Thus, ADROIT is expected to enable distribution of operations and to reduce the cost of operations by reducing the number of persons required to be in the center.

Hodges, Mark↗

Geodata Modeling and Query in Geographic Information Systems

Geographic information systems (GIS) deal with collecting, modeling, man- aging, analyzing, and integrating spatial (locational) and non-spatial (attribute) data required for geographic applications. Examples of spatial data are digital maps, administrative boundaries, road networks, and those of non-spatial data are census counts, land elevations and soil characteristics. GIS shares common areas with a number of other disciplines such as computer- aided design, computer cartography, database management, and remote sensing. None of these disciplines however, can by themselves fully meet the requirements of a GIS application. Examples of such requirements include: the ability to use locational data to produce high quality plots, perform complex operations such as network analysis, enable spatial searching and overlay operations, support spatial analysis and modeling, and provide data management functions such as efficient storage, retrieval, and modification of large datasets; independence, integrity, and security of data; and concurrent access to multiple users. It is on the data management issues that we devote our discussions in this monograph. Traditionally, database management technology have been developed for business applications. Such applications require, among other things, capturing the data requirements of high-level business functions and developing machine- level implementations; supporting multiple views of data and yet providing integration that would minimize redundancy and maintain data integrity and security; providing a high-level language for data definition and manipulation; allowing concurrent access to multiple users; and processing user transactions in an efficient manner. The demands on database management systems have been for speed, reliability, efficiency, cost effectiveness, and user-friendliness. Significant progress have been made in all of these areas over the last two decades to the point that many generalized database platforms are now available for developing data intensive applications that run in real-time. While continuous improvement is still being made at a very fast-paced and competitive rate, new application areas such as computer aided design, image processing, VLSI design, and GIS have been identified by many as the next generation of database applications. These new application areas pose serious challenges to the currently available database technology. At the core of these challenges is the nature of data that is manipulated. In traditional database applications, the database objects do not have any spatial dimension, and as such, can be thought of as point data in a multi-dimensional space. For example, each instance of an entity EMPLOYEE will have a unique value corresponding to every attribute such as employee id, employee name, employee address and so on. Thus, every Employee instance can be thought of as a point in a multi-dimensional space where each dimension is represented by an attribute. Furthermore, all operations on such data are one-dimensional. Thus, users may retrieve all entities satisfying one or more constraints. Examples of such constraints include employees with addresses in a certain area code, or salaries within a certain range. Even though constraints can be specified on multiple attributes (dimensions), the search for such data is essentially orthogonal across these dimensions.

Adam, Nabil↗

SAIL-Net CloudPuck CCN Data

SAIL-Net is a DOE funded project in the East River Watershed near Crested Butte, Colorado with the goal of advancing our understanding of aerosol-cloud interactions in complex, mountainous regions. Through the deployment of a network of six low cost microphysics nodes in Fall 2021 in the same domain at the SAIL campaign, SAIL-Net provides data on aerosol size distributions, cloud condensation nuclei (CCN), and ice nucleations particles (INP). This network enables the investigation of small-scale variations in complex terrain. This specific dataset provides the cleaned data recorded from the CloudPuck, an in house instrument made by Handix Scientific that counts CCN concentrations. The CloudPuck was deployed at the sites for the summer/fall of 2022 before winter conditions were too harsh to maintain the instrument. For more information on the the CloudPuck or how the raw data are processed, see the read me.

54 ENVIRONMENTAL SCIENCES↗

RWRtoolkit: multi-omic network analysis using random walks on multiplex networks in any species

Abstract We introduce RWRtoolkit, a multiplex generation, exploration, and statistical package built for R and command-line users. RWRtoolkit enables the efficient exploration of large and highly complex biological networks generated from custom experimental data and/or from publicly available datasets, and is species agnostic. A range of functions can be used to find topological distances between biological entities, determine relationships within sets of interest, search for topological context around sets of interest, and statistically evaluate the strength of relationships within and between sets. The command-line interface is designed for parallelization on high-performance cluster systems, which enables high-throughput analysis such as permutation testing. Several tools in the package have also been made available for use in reproducible workflows via the KBase web application.

Kainer, David (ORCID:0000000172714676)↗

Robotic tape library system level testing at NSA: Present and planned

In the present of declining Defense budgets, increased pressure has been placed on the DOD to utilize Commercial Off the Shelf (COTS) solutions to incrementally solve a wide variety of our computer processing requirements. With the rapid growth in processing power, significant expansion of high performance networking, and the increased complexity of applications data sets, the requirement for high performance, large capacity, reliable and secure, and most of all affordable robotic tape storage libraries has greatly increased. Additionally, the migration to a heterogeneous, distributed computing environment has further complicated the problem. With today's open system compute servers approaching yesterday's supercomputer capabilities, the need for affordable, reliable secure Mass Storage Systems (MSS) has taken on an ever increasing importance to our processing center's ability to satisfy operational mission requirements. To that end, NSA has established an in-house capability to acquire, test, and evaluate COTS products. Its goal is to qualify a set of COTS MSS libraries, thereby achieving a modicum of standardization for robotic tape libraries which can satisfy our low, medium, and high performance file and volume serving requirements. In addition, NSA has established relations with other Government Agencies to complete this in-house effort and to maximize our research, testing, and evaluation work. While the preponderance of the effort is focused at the high end of the storage ladder, considerable effort will be extended this year and next at the server class or mid range storage systems.

Shields, Michael F.↗

Earth Science System of the Future: Observing, Processing, and Delivering Data Products Directly to Users

Advancement of our predictive capabilities will require new scientific knowledge, improvement of our modeling capabilities, and new observation strategies to generate the complex data sets needed by coupled modeling networks. New observation strategies must support remote sensing from a variety of vantage points and will include "sensorwebs" of small satellites in low Earth orbit, large aperture sensors in Geostationary orbits, and sentinel satellites at L1 and L2 to provide day/night views of the entire globe. Onboard data processing and high speed computing and communications will enable near real-time tailoring and delivery of information products (i.e., predictions) directly to users.

Crisp, David↗

Computationally efficient Bayesian estimation of graphical networks for omics data

Graphical networks are useful, widely-used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts, such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple hundred biomolecules due to prohibitive computational time, but omics data often contains tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized datasets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. On a SARS-CoV2 proteomics data with 7000 proteins, the competing algorithm takes three times as long to complete the first iteration as BPlane takes to converge after over 100 iterations.

EM algorithm↗

FlbB forms a distinctive ring essential for periplasmic flagellar assembly and motility in Borrelia burgdorferi

Spirochetes are a widespread group of bacteria with a distinct morphology. Some spirochetes are important human pathogens that utilize periplasmic flagella to achieve motility and host infection. The motors that drive the rotation of periplasmic flagella have a unique spirochete-specific feature, termed the collar, crucial for the flat-wave morphology and motility of the Lyme disease spirochete Borrelia burgdorferi. Here, we deploy cryo-electron tomography and subtomogram averaging to determine high-resolution in-situ structures of the B. burgdorferi flagellar motor. Comparative analysis and molecular modeling of in-situ flagellar motor structures from B. burgdorferi mutants lacking each of the known collar proteins (FlcA, FlcB, FlcC, FlbB, and Bb0236/FlcD) uncover a complex protein network at the base of the collar. Importantly, our data suggest that FlbB forms a novel periplasmic ring around the rotor but also acts as a scaffold supporting collar assembly and subsequent recruitment of stator complexes. The complex protein network based on the FlbB ring effectively bridges the rotor and 16 torque-generating stator complexes in each flagellar motor, thus contributing to the specialized motility and lifestyle of spirochetes in complex environments.

59 BASIC BIOLOGICAL SCIENCES↗

Prediction of Aerodynamic Coefficient using Genetic Algorithm Optimized Neural Network for Sparse Data

Wind tunnels use scale models to characterize aerodynamic coefficients, Wind tunnel testing can be slow and costly due to high personnel overhead and intensive power utilization. Although manual curve fitting can be done, it is highly efficient to use a neural network to define the complex relationship between variables. Numerical simulation of complex vehicles on the wide range of conditions required for flight simulation requires static and dynamic data. Static data at low Mach numbers and angles of attack may be obtained with simpler Euler codes. Static data of stalled vehicles where zones of flow separation are usually present at higher angles of attack require Navier-Stokes simulations which are costly due to the large processing time required to attain convergence. Preliminary dynamic data may be obtained with simpler methods based on correlations and vortex methods; however, accurate prediction of the dynamic coefficients requires complex and costly numerical simulations. A reliable and fast method of predicting complex aerodynamic coefficients for flight simulation I'S presented using a neural network. The training data for the neural network are derived from numerical simulations and wind-tunnel experiments. The aerodynamic coefficients are modeled as functions of the flow characteristics and the control surfaces of the vehicle. The basic coefficients of lift, drag and pitching moment are expressed as functions of angles of attack and Mach number. The modeled and training aerodynamic coefficients show good agreement. This method shows excellent potential for rapid development of aerodynamic models for flight simulation. Genetic Algorithms (GA) are used to optimize a previously built Artificial Neural Network (ANN) that reliably predicts aerodynamic coefficients. Results indicate that the GA provided an efficient method of optimizing the ANN model to predict aerodynamic coefficients. The reliability of the ANN using the GA includes prediction of aerodynamic coefficients to an accuracy of 110% . In our problem, we would like to get an optimized neural network architecture and minimum data set. This has been accomplished within 500 training cycles of a neural network. After removing training pairs (outliers), the GA has produced much better results. The neural network constructed is a feed forward neural network with a back propagation learning mechanism. The main goal has been to free the network design process from constraints of human biases, and to discover better forms of neural network architectures. The automation of the network architecture search by genetic algorithms seems to have been the best way to achieve this goal.

Rajkumar, T.↗

SAIL-Net Raw and Post Corrected POPS Data Fall 2021 - Summer 2023

SAIL-Net is a DOE funded project in the East River Watershed near Crested Butte, Colorado with the goal of advancing our understanding of aerosol-cloud interactions in complex, mountainous regions. Through the deployment of a network of six low cost microphysics nodes in Fall 2021 in the same domain at the SAIL campaign, SAIL-Net provides data on aerosol size distributions, cloud condensation nuclei (CCN), and ice nucleation particles (INP). This network enables the investigation of small-scale variations in complex terrain. Two datasets are provided - one containing raw data and the other containing post-corrected data. The raw dataset provides the raw data recorded from the POPS which were deployed at each of the six sites. These data are organized by site and broken down into daily data files. The six site names used here are: “gothic”, “irwin”, “cbtop”, “cbmid”, “pumphouse”, and “snodgrass”. These data are not cleaned or post-corrected, but some flags have been added. The data are reported at 1 second time resolution. The post-corrected dataset provides the post-corrected and cleaned data recorded from the POPS which were deployed at each of the six sites. This data are also organized by site (same as those found in the raw data) and broken down into daily data files. Unlike the raw POPS data, these data have already been cleaned to remove what we believe are bad values. These data should be ready to use with no cleaning. For a full description of the cleaning and post-correction process, see the readme.

54 ENVIRONMENTAL SCIENCES↗

Experimental Observations of the Topology of Convolutional Neural Network Activations

Topological data analysis (TDA) is a branch of computational mathematics, bridging algebraic topology and data science, that provides compact, noise-robust representations of complex structures. Deep neural networks (DNNs) learn millions of parameters associated with a series of transformations defined by the model architecture resulting in high-dimensional, difficult to interpret internal representations of input data. As DNNs become more ubiquitous across multiple sectors of our society, there is increasing recognition that mathematical methods are needed to aid analysts, researchers, and practitioners in understanding and interpreting how these models' internal representations relate to the final classification. In this paper we apply cutting edge techniques from TDA with the goal of gaining insight towards interpretability of convolutional neural networks used for image classification. We use two common TDA approaches to explore several methods for modeling hidden layer activations as high-dimensional point clouds, and provide experimental evidence that these point clouds capture valuable structural information about the model's process. First, we demonstrate that a distance metric based on persistent homology can be used to quantify meaningful differences between layers and discuss these distances in the broader context of existing representational similarity metrics for neural network interpretability. Second, we show that a mapper graph can provide semantic insight as to how these models organize hierarchical class knowledge at each layer. These observations demonstrate that TDA is a useful tool to help deep learning practitioners unlock the hidden structures of their models.

topological data analysis, deep learning↗

Single-cell and spatial omics in plants: from cellular atlases to regulatory mechanisms

Single-cell RNA sequencing (scRNA-seq) has transformed transcriptomic studies by enabling gene expression profiling at the resolution of individual cells within and across a broad range of tissue types, revealing cellular heterogeneity that is obscured in bulk tissue transcriptomes. Over the past decade, improvements in microfluidics and library preparation have drastically increased throughput, allowing tens of thousands of cells to be assayed in a single experiment. Although initially developed in animal systems, scRNA-seq has rapidly emerged as a powerful and widely adopted approach in plant biology. Beyond transcriptomics, the integration of single-cell data with chromatin accessibility, proteomics, metabolomics, and spatial omics is enabling a system-level understanding of plant gene regulation and cellular organization. Network-based analytical frameworks further support the reconstruction of gene regulatory networks and the interpretation of complex single-cell data. In this review, we summarize the current technological landscape of plant single-cell studies, discuss key experimental and analytical challenges, and review emerging strategies for validating single-cell discoveries. We also discuss future directions in applying single-cell technologies to woody perennials plants and bioenergy-relevant crops, emphasizing their potential to accelerate the discovery of cell type-specific regulatory mechanisms underlying growth, stress resilience, and biomass production.

Li, Miaomiao [ORNL] (ORCID:0000000321326168)↗

Delay-Throughput Performance of the Deep-Space Ka-band Link

In this paper, performance of a first-in, first-out (FIFO), selective retransmission scheme for the deep-space Ka-band link is presented and compared to the performance of a comparable X-band link. In this analysis, 16 months of water vapor radiometer (WVR) and advanced water vapor radiometer (AWVR) data from the three Deep Space Network (DSN) Communication Complexes (DSCC) were used to emulate weather effects on X-band and Ka-band links from Mars. Mars Reconnaissance Orbiter (MRO) X-band and Ka-band telecommunications parameters were used for spacecraft telecommunications capabilities. One pass per week per complex was selected from MRO's Deep Space Network (DSN) schedule from April 1, 2006 to August 31, 2007 for a total of 207 passes (69 passes per complex) for this analysis. For each pass both X-band and Ka-band links were designed using at most two data rates so that the expected pass capacity would be maximized subject to a minimum availability requirement (MAR). In conjunction with the WVR/AWVR data, elevation profiles of the selected passes and models for the performance of the antennas in the DSN were used to emulate the performance of both links. It was assumed that the retransmission of the data takes place not on the same pass as the original transmission but during subsequent passes. The data collected before a pass was assumed to be a fraction of the expected capacity of the pass as calculated through the link design process. Infinite spacecraft storage was assumed to obtain an upper bound on the spacecraft storage requirement. The independent parameters of this analysis were MAR and the ratio of data collected before a pass to the expected pass capacity. Since the selected passes did not occur at regular intervals, the delay in this analysis was measured in terms of number of passes. The throughput was measured in terms of number of bits received successfully on the ground. The results indicate that reasonable delay performance could be achieved with very high throughput for relatively low MAR values for data collection to expected pass capacity ratio of around 97% for Ka-band. The results indicate that, except for very low average delay requirements, the Ka-band link provides more than twice the throughput of the X-band link for the same amount of power consumed by the spacecraft. In addition, the results indicate that the required storage onboard the spacecraft is not prohibitive and good performance could be achieved by using a buffer size less than three times the maximum amount of data collected before a pass.

Shambayati, Shervin↗

Delay-Throughput Performance the Deep-Space Ka-Band Link

In this paper, performance of a first-in, first-out (FIFO), selective retransmission scheme for the deep-space Ka-band link is presented and compared to the performance of a comparable X-band link. In this analysis, 16 months of water vapor radiometer (WVR) and advanced water vapor radiometer (AWVR) data from the three Deep Space Network (DSN) Communication Complexes (DSCC) were used to emulate weather effects on X-band and Ka-band links from Mars. Mars Reconnaissance Orbiter (MRO) X-band and Ka-band telecommunications parameters were used for spacecraft telecommunications capabilities. One pass per week per complex was selected from MRO's Deep Space Network (DSN) schedule from April 1, 2006 to August 31, 2007 for a total of 207 passes (69 passes per complex) for this analysis. For each pass both X-band and Ka-band links were designed using at most two data rates so that the expected pass capacity would be maximized subject to a minimum availability requirement (MAR). In conjunction with the WVR/AWVR data, elevation profiles of the selected passes and models for the performance of the antennas in the DSN were used to emulate the performance of both links. It was assumed that the retransmission of the data takes place not on the same pass as the original transmission but during subsequent passes. The data collected before a pass was assumed to be a fraction of the expected capacity of the pass as calculated through the link design process. Infinite spacecraft storage was assumed to obtain an upper bound on the spacecraft storage requirement. The independent parameters of this analysis were MAR and the ratio of data collected before a pass to the expected pass capacity. Since the selected passes did not occur at regular intervals, the delay in this analysis was measured in terms of number of passes. The throughput was measured in terms of number of bits received successfully on the ground. The results indicate that reasonable delay performance could be achieved with very high throughput for relatively low MAR values for data collection to expected pass capacity ratio of around 97% for Ka-band. The results indicate that, except for very low average delay requirements, the Ka-band link provides more than twice the throughput of the X-band link for the same amount of power consumed by the spacecraft. In addition, the results indicate that the required storage onboard the spacecraft is not prohibitive and good performance could be achieved by using a buffer size less than three times the maximum amount of data collected before a pass.

Shambayati, Shervin↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

X-Band Radar and Surface-Based Observations of Cold-Season Precipitation in Western Colorado’s Complex Terrain

Abstract Hydrologic processes associated with intermountain cold-season precipitation in the Upper Colorado River basin have important impacts on avalanche forecasting and water resource management. However, traditional weather radar networks struggle with observations in this complex terrain. Data collected during the Study of Precipitation, the Lower Atmosphere, and the Surface for Hydrometeorology (SPLASH) and its sister campaign, Surface Atmosphere Integrated Field Laboratory (SAIL) in the East River watershed of western Colorado, are used to examine a multistorm period from 23 December 2021 to 1 January 2022 that contributed 35% of the total winter precipitation in this watershed. Dual-polarization X-band radar and disdrometer measurements show ∼30-mm differences in precipitation amount at two sites in proximity over four distinct storm events within the period. Wind patterns, synoptic forcings, microphysical characteristics of precipitation, and surface meteorology are analyzed to explain the observed spatial variability of cold-season precipitation in complex mountainous terrain. Analysis shows that differences over time within this event are mainly accounted for by synoptic forcings, such as frontal passages; differences between sites are accounted for by the impact of variations in local wind patterns on precipitation microphysics. Patterns of surface precipitation intensity are compared and found to be correlated with X-band radar signatures; a relationship between a strong dendritic growth stage and intense low-density surface precipitation is reinforced by this study. This relationship demonstrates the importance of particle growth mechanisms on surface snowfall patterns in high-altitude complex terrain, underscoring the importance of realistic microphysical parameterizations. Significance Statement The amount and density of snowpack from western Colorado winter storms have significant impacts on water resources in the Upper Colorado River basin. Snowpack characteristics are affected by small-scale differences in how snow forms in the atmosphere. These differences are hard to study in the complex terrain of the Rockies, but data from the SPLASH and SAIL field campaigns allows us to investigate how snow crystal formation and mountain-driven wind patterns affect snow near the surface. Our study finds that snow crystal growth varies over small space and time scales and is likely controlled by the terrain beneath a given location and resultant local wind patterns. These results imply that predicting snowpack in the Rockies requires properly representing local wind patterns and crystal growth processes in models.

Heflin, Stella↗

Network Models of Active Degradation Mechanisms and Pathways for Service Life Prediction of Indoor and Outdoor PV Modules

ct: PV service lifetime prediction (SLP) enables accurate calculation of levelized cost of energy (LCOE), which is crucial to rationalizing PV investment and installation. However, SLP is challeging since PV reliability in the field is affected by many combined factors, including various environmental stresses and module quality. In order to map out the active degradation mechanisms and pathways that best resemble real world conditions, we introduce the framework of a study protocol and use network models fitted to data, to enable analysis and SLP of complex PV systems with multiple active degradation mechanisms. The study protocol is the experimental design, including module variants and different exposure conditions, selection of evaluation methods, time-series data acquisition and training of network models to these data. We present SLP of minimodules in the lab and PV systems in the field. For lab SLP, minimodules with 8 variants based on manufacturer, architecture, and encapsulation were prepared and aged in modified damp heat with or without full spectrum light exposure. Stepwise I-V and Suns-Voc data acquisition tracks changes in electrical properties including Rs,IV, Isc,IV, Vmp,PIV providing insights into power loss of minimodules. Network structural equation modeling (netSEM) was utilized to construct degradation pathway models that identify active degradation mechanisms and predict power loss over time. For field SLP, datastreams of Pmp values and I-V curve datastreams of two types of modules installed in three distinctly different Köppen-Geiger climate zones for 9 years were acquired. With power loss modes corresponding to uniform current loss (ΔPIsc), recombination (ΔPVoc), series resistance (ΔPRs), and current mismatch (ΔPImis) determined, the performance loss rates (PLR) were determined using PVplr. We show how to establish a study protocol framework to ensure appropriate parametric variations and valid data collection from the variants of your complex systems. Then the data-driven netSEM model fitting provides a comprehensive mapping of multiple active degradation mechanisms, and accurate service life prediction.

network model, degradation, photovoltaic, solar↗