Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Complex Network Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Multifidelity deep operator networks for data-driven and physics-informed problems

Operator learning for complex nonlinear systems is increasingly common in modeling multi-physics and multi-scale systems. However, training such high-dimensional operators requires a large amount of expensive, high-fidelity data, either from experiments or simulations. In this work, we present a composite Deep Operator Network (DeepONet) for learning using two datasets with different levels of fidelity to accurately learn complex operators when sufficient high-fidelity data is not available. Additionally, we demonstrate that the presence of low-fidelity data can improve the predictions of physics-informed learning with DeepONets. We demonstrate the new multi-fidelity training in diverse examples, including modeling of the ice-sheet dynamics of the Humboldt glacier, Greenland, using two different fidelity models and also using the same physical model at two different resolutions.

97 MATHEMATICS AND COMPUTING↗

Data-Driven Kinetic Reaction Networks for Separation Chemistry

Understanding complex, multistep chemical reactions at the molecular level is a major challenge whose solution would greatly benefit the design and optimization of numerous chemical processes. The separation of rare-earth (4f) and actinide (5f) elements is an example where improving our chemical understanding is important for designing and optimizing new chemistries, even with a limited number of observations. Here, in this work, we leverage data-driven artificial intelligence and machine-learning approaches to develop kinetic reaction networks that describe the liquid–liquid extraction mechanism of uranium using N,N-di-2-ethylhexyl-isobutyramide (DEHiBA). Specifically, we compare and contrast the properties of two classes of models: (1) purely data-driven models that are regularized using chemistry-agnostic, L1 regression and (2) chemistry-informed models that are regularized using relative reaction energies provided by quantum mechanical calculations. We observe that purely data-driven models are unbiased, simple, and accurate in their predictions of experimental measurements when provided with sufficient data but are difficult to fully constrain and interpret. In contrast, chemistry-informed models exhibit significantly improved chemical interpretability and consistency, providing a detailed description of the separation process while achieving high accuracy through ensemble averaging. Overall, the dominant species predicted to be extracted into the organic phase is UO 2 (NO 3 ) 2 (DEHiBA) 2 , agreeing with experimental slope analysis, thermodynamic modeling, EXAFS, and crystal structures. This work demonstrates that leveraging the fundamental structure of the problem can lead to efficient learning schemes that provide both accurate predictions and chemical insights at a low computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Neural network-based classification and regression of magnetohydrodynamic modes in tokamaks

We present a machine learning-based magnetohydrodynamic (MHD) classifier and regressor that utilizes real or complex-valued 3D magnetic sensor array data to determine neoclassical tearing mode (NTM) onset times in tokamaks with millisecond accuracy. The input dataset consists of poloidal profiles of complex Fourier amplitudes with an n = 1 toroidal mode number from 144 human-labeled ITER Baseline Scenario discharges in the DIII-D tokamak, spanning both tearing-dominated and sawtooth-dominated regimes. Since m, n = 2,1 NTMs frequently emerge alongside sawteeth at the same frequency in this scenario, the focus is on isolating the m = 1 and m = 2 components of the n = 1 MHD mode near the tearing onset. To improve model regularization and prediction stability, singular value decomposition was applied to balance the sawtooth and tearing datasets. The enriched datasets facilitated training neural networks that learn the key distinguishing features of sawtooth and tearing modes in the poloidal profiles of their magnetic amplitude and phase. When the modes occur independently, the networks achieve perfect classification due to the modes’ distinct characteristics and low measurement noise. In the more experimentally relevant case where both modes coexist, the networks maintain exceptional performance across key metrics. Tests on synthetic data with known ground truth demonstrate the superior accuracy of the neural network trained on complex-valued input compared to models using real amplitude, phase, or pseudo-complex data, achieving both a mean time delay and standard deviation below 1 ms. Notably, standard linear regression methods fitting the dominant singular modes to the data closely match the neural network’s performance. Applying these methods across a broad range of H-mode scenarios will enable future studies to systematically identify dominant NTM triggers as scenario-specific variables, paving the way for more effective tearing mode avoidance strategies in future fusion reactor designs.

machine learning↗

Multisource Data Fusion Outage Location in Distribution Systems via Probabilistic Graphical Models

Efficient outage location is critical to enhancing the resilience of power distribution systems. However, accurate outage location requires combining massive evidence received from diverse data sources, including smart meter (SM) last gasp signals, customer trouble calls, social media messages, weather data, vegetation information, and physical parameters of the network. This is a computationally complex task due to the high dimensionality of data in distribution grids. In this paper, we propose a multi-source data fusion approach to locate outage events in partially observable distribution systems using Bayesian networks (BNs). A novel aspect of the proposed approach is that it takes multi-source evidence and the complex structure of distribution systems into account using a probabilistic graphical method. Our method can radically reduce the computational complexity of outage location inference in high-dimensional spaces. The graphical structure of the proposed BN is established based on the network’s topology and the causal relationship between random variables, such as the states of branches/customers and evidence. Utilizing this graphical model, accurate outage locations are obtained by leveraging a Gibbs sampling (GS) method, to infer the probabilities of de-energization for all branches. Compared with commonly-used exact inference methods that have exponential complexity in the size of the BN, GS quantifies the target conditional probability distributions in a timely manner. As a result, a case study of several real-world distribution systems is presented to validate the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Generating synthetic signaling networks for in silico modeling studies

Predictive models of signaling pathways have proven to be difficult to develop. Reasons include the uncertainty in the number of species, the complexity in species’ interactions, and the sparseness and uncertainty in experimental data. Traditional approaches to developing mechanistic models rely on collecting experimental data and fitting a single model to that data. This approach works for simple systems but has proven unreliable for complex systems such as biological signaling networks. For example, uncertainty and sparseness of the data often result in overfitted models that have little predictive value beyond recapitulating the experimental data itself. Thus, there is a need to develop new approaches to create predictive mechanistic models of complex systems. However, to determine the effectiveness of any new algorithm, a baseline model is needed to test its performance. To meet this need, we developed a method for generating artificial synthetic networks that are reasonably realistic and thus can be treated as ground truth models. These synthetic models can then be used to generate synthetic data for developing and testing algorithms designed to recover the underlying network topology and associated parameters. Here, we describe a simple approach for generating synthetic signaling networks that can be used for this purpose.

42 ENGINEERING↗

Characterizing the acceleration time of laser-driven ion acceleration with data-informed neural networks

Peak ion energy is an important figure-of-merit in short-pulse, laser-driven ion acceleration and is dependent on an associated acceleration time. Standard metrics for these quantities depend on analytical results such as the self-similar fluid model or empirical models based on relatively small experimental and simulation datasets. In this work we attempt to use a data-informed neural network (NN) as a surrogate model for a large ensemble of PIC simulations to investigate an effective acceleration time. We explore the application of a stacked convolutional and recurrent NN architecture for improved regression by incorporating the time dependencies of the data into the training process. Of particular note is how pretraining a network on lower fidelity data, e.g. 1D analytical results, greatly improves the network's ability to learn more complex, higher fidelity data. Finally, the dependency of the acceleration time on various laser and plasma parameters is explored.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Electric Utility Industry Standards Landscape

The electric utility industry relies on robust communication protocols to manage complex electrical grid data. The inherent networked nature of electrical grids, coupled with the radial structure of the “last mile” portion delivering power to end-use customers, presents difficulties in describing electrical models using simple data constructs. The paper provides an overview of communication protocols that address electric utility data, including grid data, and classifies these protocols through identification key characteristics that make them suitable for different electric grid data domains. Because of their variety, grid edge devices and their associated dedicated protocols are assessed by groups: those primarily designed for energy production and storage, those related to flexible loads, and those related to electric vehicles. The intent of this report is to provide guidance for stakeholders to navigate the challenges posed by the numerous overlapping protocols available to address electric grid data. Although further industry review, refinement, and validation of the categorization of these protocols is recommended, the authors propose this categorization as a start to improve electric grid awareness and understanding. In addition, this report includes three recommended industry actions regarding protocols to improve communications on the electric grid.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440

ABSTRACT There is growing interest in engineering Pseudomonas putida KT2440 as a microbial chassis for the conversion of renewable and waste-based feedstocks, and metabolic engineering of P. putida relies on the understanding of the functional relationships between genes. In this work, independent component analysis (ICA) was applied to a compendium of existing fitness data from randomly barcoded transposon insertion sequencing (RB-TnSeq) of P. putida KT2440 grown in 179 unique experimental conditions. ICA identified 84 independent groups of genes, which we call fModules (“functional modules”), where gene members displayed shared functional influence in a specific cellular process. This machine learning-based approach both successfully recapitulated previously characterized functional relationships and established hitherto unknown associations between genes. Selected gene members from fModules for hydroxycinnamate metabolism and stress resistance, acetyl coenzyme A assimilation, and nitrogen metabolism were validated with engineered mutants of P. putida . Additionally, functional gene clusters from ICA of RB-TnSeq data sets were compared with regulatory gene clusters from prior ICA of RNAseq data sets to draw connections between gene regulation and function. Because ICA profiles the functional role of several distinct gene networks simultaneously, it can reduce the time required to annotate gene function relative to manual curation of RB-TnSeq data sets. IMPORTANCE This study demonstrates a rapid, automated approach for elucidating functional modules within complex genetic networks. While Pseudomonas putida randomly barcoded transposon insertion sequencing data were used as a proof of concept, this approach is applicable to any organism with existing functional genomics data sets and may serve as a useful tool for many valuable applications, such as guiding metabolic engineering efforts in other microbes or understanding functional relationships between virulence-associated genes in pathogenic microbes. Furthermore, this work demonstrates that comparison of data obtained from independent component analysis of transcriptomics and gene fitness datasets can elucidate regulatory-functional relationships between genes, which may have utility in a variety of applications, such as metabolic modeling, strain engineering, or identification of antimicrobial drug targets.

09 BIOMASS FUELS↗

Reducing uncertainty of high-latitude ecosystem models through identification of key parameters

Abstract Climate change is having significant impacts on Earth’s ecosystems and carbon budgets, and in the Arctic may drive a shift from an historic carbon sink to a source. Large uncertainties in terrestrial biosphere models (TBMs) used to forecast Arctic changes demonstrate the challenges of determining the timing and extent of this possible switch. This spread in model predictions can limit the ability of TBMs to guide management and policy decisions. One of the most influential sources of model uncertainty is model parameterization. Parameter uncertainty results in part from a mismatch between available data in databases and model needs. We identify that mismatch for three TBMs, DVM-DOS-TEM, SIPNET and ED2, and four databases with information on Arctic and boreal above- and belowground traits that may be applied to model parametrization. However, focusing solely on such data gaps can introduce biases towards simple models and ignores structural model uncertainty, another main source for model uncertainty. Therefore, we develop a causal loop diagram (CLD) of the Arctic and boreal ecosystem that includes unquantified, and thus unmodeled, processes. We map model parameters to processes in the CLD and assess parameter vulnerability via the internal network structure. One important substructure, feed forward loops (FFLs), describe processes that are linked both directly and indirectly. When the model parameters are data-informed, these indirect processes might be implicitly included in the model, but if not, they have the potential to introduce significant model uncertainty. We find that the parameters describing the impact of local temperature on microbial activity are associated with a particularly high number of FFLs but are not constrained well by existing data. By employing ecological models of varying complexity, databases, and network methods, we identify the key parameters responsible for limited model accuracy. They should be prioritized for future data sampling to reduce model uncertainty.

54 ENVIRONMENTAL SCIENCES↗

SAIL-Net CloudPuck CCN Data

SAIL-Net is a DOE funded project in the East River Watershed near Crested Butte, Colorado with the goal of advancing our understanding of aerosol-cloud interactions in complex, mountainous regions. Through the deployment of a network of six low cost microphysics nodes in Fall 2021 in the same domain at the SAIL campaign, SAIL-Net provides data on aerosol size distributions, cloud condensation nuclei (CCN), and ice nucleations particles (INP). This network enables the investigation of small-scale variations in complex terrain. This specific dataset provides the cleaned data recorded from the CloudPuck, an in house instrument made by Handix Scientific that counts CCN concentrations. The CloudPuck was deployed at the sites for the summer/fall of 2022 before winter conditions were too harsh to maintain the instrument. For more information on the the CloudPuck or how the raw data are processed, see the read me.

54 ENVIRONMENTAL SCIENCES↗

RWRtoolkit: multi-omic network analysis using random walks on multiplex networks in any species

Abstract We introduce RWRtoolkit, a multiplex generation, exploration, and statistical package built for R and command-line users. RWRtoolkit enables the efficient exploration of large and highly complex biological networks generated from custom experimental data and/or from publicly available datasets, and is species agnostic. A range of functions can be used to find topological distances between biological entities, determine relationships within sets of interest, search for topological context around sets of interest, and statistically evaluate the strength of relationships within and between sets. The command-line interface is designed for parallelization on high-performance cluster systems, which enables high-throughput analysis such as permutation testing. Several tools in the package have also been made available for use in reproducible workflows via the KBase web application.

Kainer, David (ORCID:0000000172714676)↗

Computationally efficient Bayesian estimation of graphical networks for omics data

Graphical networks are useful, widely-used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts, such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple hundred biomolecules due to prohibitive computational time, but omics data often contains tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized datasets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. On a SARS-CoV2 proteomics data with 7000 proteins, the competing algorithm takes three times as long to complete the first iteration as BPlane takes to converge after over 100 iterations.

EM algorithm↗

FlbB forms a distinctive ring essential for periplasmic flagellar assembly and motility in Borrelia burgdorferi

Spirochetes are a widespread group of bacteria with a distinct morphology. Some spirochetes are important human pathogens that utilize periplasmic flagella to achieve motility and host infection. The motors that drive the rotation of periplasmic flagella have a unique spirochete-specific feature, termed the collar, crucial for the flat-wave morphology and motility of the Lyme disease spirochete Borrelia burgdorferi. Here, we deploy cryo-electron tomography and subtomogram averaging to determine high-resolution in-situ structures of the B. burgdorferi flagellar motor. Comparative analysis and molecular modeling of in-situ flagellar motor structures from B. burgdorferi mutants lacking each of the known collar proteins (FlcA, FlcB, FlcC, FlbB, and Bb0236/FlcD) uncover a complex protein network at the base of the collar. Importantly, our data suggest that FlbB forms a novel periplasmic ring around the rotor but also acts as a scaffold supporting collar assembly and subsequent recruitment of stator complexes. The complex protein network based on the FlbB ring effectively bridges the rotor and 16 torque-generating stator complexes in each flagellar motor, thus contributing to the specialized motility and lifestyle of spirochetes in complex environments.

59 BASIC BIOLOGICAL SCIENCES↗

SAIL-Net Raw and Post Corrected POPS Data Fall 2021 - Summer 2023

SAIL-Net is a DOE funded project in the East River Watershed near Crested Butte, Colorado with the goal of advancing our understanding of aerosol-cloud interactions in complex, mountainous regions. Through the deployment of a network of six low cost microphysics nodes in Fall 2021 in the same domain at the SAIL campaign, SAIL-Net provides data on aerosol size distributions, cloud condensation nuclei (CCN), and ice nucleation particles (INP). This network enables the investigation of small-scale variations in complex terrain. Two datasets are provided - one containing raw data and the other containing post-corrected data. The raw dataset provides the raw data recorded from the POPS which were deployed at each of the six sites. These data are organized by site and broken down into daily data files. The six site names used here are: “gothic”, “irwin”, “cbtop”, “cbmid”, “pumphouse”, and “snodgrass”. These data are not cleaned or post-corrected, but some flags have been added. The data are reported at 1 second time resolution. The post-corrected dataset provides the post-corrected and cleaned data recorded from the POPS which were deployed at each of the six sites. This data are also organized by site (same as those found in the raw data) and broken down into daily data files. Unlike the raw POPS data, these data have already been cleaned to remove what we believe are bad values. These data should be ready to use with no cleaning. For a full description of the cleaning and post-correction process, see the readme.

54 ENVIRONMENTAL SCIENCES↗

Experimental Observations of the Topology of Convolutional Neural Network Activations

Topological data analysis (TDA) is a branch of computational mathematics, bridging algebraic topology and data science, that provides compact, noise-robust representations of complex structures. Deep neural networks (DNNs) learn millions of parameters associated with a series of transformations defined by the model architecture resulting in high-dimensional, difficult to interpret internal representations of input data. As DNNs become more ubiquitous across multiple sectors of our society, there is increasing recognition that mathematical methods are needed to aid analysts, researchers, and practitioners in understanding and interpreting how these models' internal representations relate to the final classification. In this paper we apply cutting edge techniques from TDA with the goal of gaining insight towards interpretability of convolutional neural networks used for image classification. We use two common TDA approaches to explore several methods for modeling hidden layer activations as high-dimensional point clouds, and provide experimental evidence that these point clouds capture valuable structural information about the model's process. First, we demonstrate that a distance metric based on persistent homology can be used to quantify meaningful differences between layers and discuss these distances in the broader context of existing representational similarity metrics for neural network interpretability. Second, we show that a mapper graph can provide semantic insight as to how these models organize hierarchical class knowledge at each layer. These observations demonstrate that TDA is a useful tool to help deep learning practitioners unlock the hidden structures of their models.

topological data analysis, deep learning↗

Single-cell and spatial omics in plants: from cellular atlases to regulatory mechanisms

Single-cell RNA sequencing (scRNA-seq) has transformed transcriptomic studies by enabling gene expression profiling at the resolution of individual cells within and across a broad range of tissue types, revealing cellular heterogeneity that is obscured in bulk tissue transcriptomes. Over the past decade, improvements in microfluidics and library preparation have drastically increased throughput, allowing tens of thousands of cells to be assayed in a single experiment. Although initially developed in animal systems, scRNA-seq has rapidly emerged as a powerful and widely adopted approach in plant biology. Beyond transcriptomics, the integration of single-cell data with chromatin accessibility, proteomics, metabolomics, and spatial omics is enabling a system-level understanding of plant gene regulation and cellular organization. Network-based analytical frameworks further support the reconstruction of gene regulatory networks and the interpretation of complex single-cell data. In this review, we summarize the current technological landscape of plant single-cell studies, discuss key experimental and analytical challenges, and review emerging strategies for validating single-cell discoveries. We also discuss future directions in applying single-cell technologies to woody perennials plants and bioenergy-relevant crops, emphasizing their potential to accelerate the discovery of cell type-specific regulatory mechanisms underlying growth, stress resilience, and biomass production.

Li, Miaomiao [ORNL] (ORCID:0000000321326168)↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

X-Band Radar and Surface-Based Observations of Cold-Season Precipitation in Western Colorado’s Complex Terrain

Abstract Hydrologic processes associated with intermountain cold-season precipitation in the Upper Colorado River basin have important impacts on avalanche forecasting and water resource management. However, traditional weather radar networks struggle with observations in this complex terrain. Data collected during the Study of Precipitation, the Lower Atmosphere, and the Surface for Hydrometeorology (SPLASH) and its sister campaign, Surface Atmosphere Integrated Field Laboratory (SAIL) in the East River watershed of western Colorado, are used to examine a multistorm period from 23 December 2021 to 1 January 2022 that contributed 35% of the total winter precipitation in this watershed. Dual-polarization X-band radar and disdrometer measurements show ∼30-mm differences in precipitation amount at two sites in proximity over four distinct storm events within the period. Wind patterns, synoptic forcings, microphysical characteristics of precipitation, and surface meteorology are analyzed to explain the observed spatial variability of cold-season precipitation in complex mountainous terrain. Analysis shows that differences over time within this event are mainly accounted for by synoptic forcings, such as frontal passages; differences between sites are accounted for by the impact of variations in local wind patterns on precipitation microphysics. Patterns of surface precipitation intensity are compared and found to be correlated with X-band radar signatures; a relationship between a strong dendritic growth stage and intense low-density surface precipitation is reinforced by this study. This relationship demonstrates the importance of particle growth mechanisms on surface snowfall patterns in high-altitude complex terrain, underscoring the importance of realistic microphysical parameterizations. Significance Statement The amount and density of snowpack from western Colorado winter storms have significant impacts on water resources in the Upper Colorado River basin. Snowpack characteristics are affected by small-scale differences in how snow forms in the atmosphere. These differences are hard to study in the complex terrain of the Rockies, but data from the SPLASH and SAIL field campaigns allows us to investigate how snow crystal formation and mountain-driven wind patterns affect snow near the surface. Our study finds that snow crystal growth varies over small space and time scales and is likely controlled by the terrain beneath a given location and resultant local wind patterns. These results imply that predicting snowpack in the Rockies requires properly representing local wind patterns and crystal growth processes in models.

Heflin, Stella↗