Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Complex Network Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

X-Band Radar and Surface-Based Observations of Cold-Season Precipitation in Western Colorado’s Complex Terrain

Abstract Hydrologic processes associated with intermountain cold-season precipitation in the Upper Colorado River basin have important impacts on avalanche forecasting and water resource management. However, traditional weather radar networks struggle with observations in this complex terrain. Data collected during the Study of Precipitation, the Lower Atmosphere, and the Surface for Hydrometeorology (SPLASH) and its sister campaign, Surface Atmosphere Integrated Field Laboratory (SAIL) in the East River watershed of western Colorado, are used to examine a multistorm period from 23 December 2021 to 1 January 2022 that contributed 35% of the total winter precipitation in this watershed. Dual-polarization X-band radar and disdrometer measurements show ∼30-mm differences in precipitation amount at two sites in proximity over four distinct storm events within the period. Wind patterns, synoptic forcings, microphysical characteristics of precipitation, and surface meteorology are analyzed to explain the observed spatial variability of cold-season precipitation in complex mountainous terrain. Analysis shows that differences over time within this event are mainly accounted for by synoptic forcings, such as frontal passages; differences between sites are accounted for by the impact of variations in local wind patterns on precipitation microphysics. Patterns of surface precipitation intensity are compared and found to be correlated with X-band radar signatures; a relationship between a strong dendritic growth stage and intense low-density surface precipitation is reinforced by this study. This relationship demonstrates the importance of particle growth mechanisms on surface snowfall patterns in high-altitude complex terrain, underscoring the importance of realistic microphysical parameterizations. Significance Statement The amount and density of snowpack from western Colorado winter storms have significant impacts on water resources in the Upper Colorado River basin. Snowpack characteristics are affected by small-scale differences in how snow forms in the atmosphere. These differences are hard to study in the complex terrain of the Rockies, but data from the SPLASH and SAIL field campaigns allows us to investigate how snow crystal formation and mountain-driven wind patterns affect snow near the surface. Our study finds that snow crystal growth varies over small space and time scales and is likely controlled by the terrain beneath a given location and resultant local wind patterns. These results imply that predicting snowpack in the Rockies requires properly representing local wind patterns and crystal growth processes in models.

Heflin, Stella↗

Network Models of Active Degradation Mechanisms and Pathways for Service Life Prediction of Indoor and Outdoor PV Modules

ct: PV service lifetime prediction (SLP) enables accurate calculation of levelized cost of energy (LCOE), which is crucial to rationalizing PV investment and installation. However, SLP is challeging since PV reliability in the field is affected by many combined factors, including various environmental stresses and module quality. In order to map out the active degradation mechanisms and pathways that best resemble real world conditions, we introduce the framework of a study protocol and use network models fitted to data, to enable analysis and SLP of complex PV systems with multiple active degradation mechanisms. The study protocol is the experimental design, including module variants and different exposure conditions, selection of evaluation methods, time-series data acquisition and training of network models to these data. We present SLP of minimodules in the lab and PV systems in the field. For lab SLP, minimodules with 8 variants based on manufacturer, architecture, and encapsulation were prepared and aged in modified damp heat with or without full spectrum light exposure. Stepwise I-V and Suns-Voc data acquisition tracks changes in electrical properties including Rs,IV, Isc,IV, Vmp,PIV providing insights into power loss of minimodules. Network structural equation modeling (netSEM) was utilized to construct degradation pathway models that identify active degradation mechanisms and predict power loss over time. For field SLP, datastreams of Pmp values and I-V curve datastreams of two types of modules installed in three distinctly different Köppen-Geiger climate zones for 9 years were acquired. With power loss modes corresponding to uniform current loss (ΔPIsc), recombination (ΔPVoc), series resistance (ΔPRs), and current mismatch (ΔPImis) determined, the performance loss rates (PLR) were determined using PVplr. We show how to establish a study protocol framework to ensure appropriate parametric variations and valid data collection from the variants of your complex systems. Then the data-driven netSEM model fitting provides a comprehensive mapping of multiple active degradation mechanisms, and accurate service life prediction.

network model, degradation, photovoltaic, solar↗

Ice-nucleating particles (INPs) concentrations from SAIL-Net

This data set contains ice-nucleating particle (INP) concentration spectra collected during the SAIL-Net sampling period, which complemented the Surface Atmosphere Integrated Field Laboratory (SAIL) campaign in the East River watershed near Crested Butte, Colorado. SAIL-Net was a distributed aerosol measurement network designed to investigate aerosol variability across complex mountainous terrain. The data set includes samples from multiple SAIL-Net sites, including AOS, Gothic, Snodgrass, Pumphouse, Irwin, and Top. INP concentrations are reported as a function of freezing temperature, together with confidence limits, sampling times, site location, elevation, sampled air volume, and treatment information. These data provide an analysis-ready record of the INPs across the SAIL-Net network.

activation temperature↗

Inferring microbial co-occurrence networks from amplicon data: a systematic evaluation

Microbes commonly organize into communities consisting of hundreds of species involved in complex interactions with each other. 16S ribosomal RNA (16S rRNA) amplicon profiling provides snapshots that reveal the phylogenies and abundance profiles of these microbial communities. These snapshots, when collected from multiple samples, can reveal the co-occurrence of microbes, providing a glimpse into the network of associations in these communities. However, the inference of networks from 16S data involves numerous steps, each requiring specific tools and parameter choices. Moreover, the extent to which these steps affect the final network is still unclear. In this study, we perform a meticulous analysis of each step of a pipeline that can convert 16S sequencing data into a network of microbial associations. Through this process, we map how different choices of algorithms and parameters affect the co-occurrence network and identify the steps that contribute substantially to the variance. We further determine the tools and parameters that generate robust co-occurrence networks and develop consensus network algorithms based on benchmarks with mock and synthetic data sets. The Microbial Co-occurrence Network Explorer, or MiCoNE (available at https://github.com/segrelab/MiCoNE) follows these default tools and parameters and can help explore the outcome of these combinations of choices on the inferred networks. We envisage that this pipeline could be used for integrating multiple data sets and generating comparative analyses and consensus networks that can guide our understanding of microbial community assembly in different biomes.

16S rRNA↗

HERMES: A cyber-energy SAAS platform to provide real-time 2d and 3d interactive visualizations [SWR-23-15]

HERMES is a novel visualization SAAS platform for achieving real-time visualization of large-scale environments involving cyber-energy devices. The platform receives data from emulated environments which include physical hardware and network devices that represent a complex IT / OT system. The platform is capable of streaming, filtering, storing, and visualizing all data within the environment and scaling vertically and horizontally. This capability enables high fidelity visual analysis of events to be performed in real time as well as the collection and storage of historical data for forensic analysis.

Van Natta, Joshua↗

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Rapid Inference of Logic Gate Neural Networks for Anomaly Detection in High Energy Physics

The increasing data rates and complexity of detectors at the Large Hadron Collider (LHC) necessitate fast and efficient machine learning models, particularly for rapid selection of what data to store, known as triggering. Building on recent work in differentiable logic gates, we present a public implementation of a Convolutional Differentiable Logic Gate Neural Network (CLGN). We apply this to detecting anomalies at the Level-1 Trigger at CMS using public data from the CICADA project. We demonstrate that the CLGN achieves physics performance on par with or superior to conventional quantized neural networks. We also synthesize an LGN for a Field-Programmable Gate Array (FPGA) and show highly promising FPGA characteristics, notably zero Digital Signal Processor (DSP) resource usage. This work highlights the potential of logic gate networks for high-speed, on-detector inference in High Energy Physics and beyond.

FOS: Physical sciences↗

Deep structural clustering for single-cell RNA-seq data jointly through autoencoder and graph neural network

Abstract Single-cell RNA sequencing (scRNA-seq) permits researchers to study the complex mechanisms of cell heterogeneity and diversity. Unsupervised clustering is of central importance for the analysis of the scRNA-seq data, as it can be used to identify putative cell types. However, due to noise impacts, high dimensionality and pervasive dropout events, clustering analysis of scRNA-seq data remains a computational challenge. Here, we propose a new deep structural clustering method for scRNA-seq data, named scDSC, which integrate the structural information into deep clustering of single cells. The proposed scDSC consists of a Zero-Inflated Negative Binomial (ZINB) model-based autoencoder, a graph neural network (GNN) module and a mutual-supervised module. To learn the data representation from the sparse and zero-inflated scRNA-seq data, we add a ZINB model to the basic autoencoder. The GNN module is introduced to capture the structural information among cells. By joining the ZINB-based autoencoder with the GNN module, the model transfers the data representation learned by autoencoder to the corresponding GNN layer. Furthermore, we adopt a mutual supervised strategy to unify these two different deep neural architectures and to guide the clustering task. Extensive experimental results on six real scRNA-seq datasets demonstrate that scDSC outperforms state-of-the-art methods in terms of clustering accuracy and scalability. Our method scDSC is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/DHUDBlab/scDSC.

Gan, Yanglan↗

Ultrafast Bragg coherent diffraction imaging of epitaxial thin films using deep complex-valued neural networks

Abstract Domain wall structures form spontaneously due to epitaxial misfit during thin film growth. Imaging the dynamics of domains and domain walls at ultrafast timescales can provide fundamental clues to features that impact electrical transport in electronic devices. Recently, deep learning based methods showed promising phase retrieval (PR) performance, allowing intensity-only measurements to be transformed into snapshot real space images. While the Fourier imaging model involves complex-valued quantities, most existing deep learning based methods solve the PR problem with real-valued based models, where the connection between amplitude and phase is ignored. To this end, we involve complex numbers operation in the neural network to preserve the amplitude and phase connection. Therefore, we employ the complex-valued neural network for solving the PR problem and evaluate it on Bragg coherent diffraction data streams collected from an epitaxial La 2-x Sr x CuO 4 (LSCO) thin film using an X-ray Free Electron Laser (XFEL). Our proposed complex-valued neural network based approach outperforms the traditional real-valued neural network methods in both supervised and unsupervised learning manner. Phase domains are also observed from the LSCO thin film at an ultrafast timescale using the complex-valued neural network.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

SymProp: Scaling Sparse Symmetric Tucker Decomposition via Symmetry Propagation

Sparse symmetric tensors are an important class of tensors, and their decompositions serve as powerful tools for revealing low-rank structures. This paper introduces SymProp, a novel approach for scaling sparse symmetric Tucker decomposition by propagating symmetry through intermediate computations. SymProp optimizes two key computational kernels: Sparse Symmetric Tensor Times Same Matrix chain (S3 TTMc) for Higher-Order Orthogonal Iteration (HOOI) and Sparse Symmetric Tensor Times Same Matrix chain Times Core (S3 TTMcTC) for Higher-Order QR Iteration (HOQRI). Our method employs a metaprogramming-based index iteration approach to efficiently handle the upper triangular parts of intermediate dense symmetric tensors. SymProp achieves up to 50.9× speedup over SPLATT and up to 360.8× over Compressed Sparse Symmetric (CSS) format on the S3 TTMc operation. Moreover, our S3 TTMc and S3 TTMcTC implementations support tensor orders four levels higher than state-of-the-art methods. Our HOQRI demonstrates superior scalability and up to a 33.6× speedup over optimized HOOI. By enabling more scalable Tucker decompositions for higher orders, decomposition ranks, and dimension sizes, SymProp opens new possibilities for analyzing complex hypergraph structures in fields such as network science, data mining, and machine learning.

Li, Zecheng [North Carolina State University]↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

RuralAI in Tomato Farming: Integrated Sensor System, Distributed Computing, and Hierarchical Federated Learning for Crop Health Monitoring

Precision horticulture is evolving due to scalable sensor deployment and machine learning (ML) integration. These advancements boost the operational efficiency of individual farms, balancing the benefits of analytics with autonomy requirements. However, given concerns that affect wide geographic regions (e.g., climate change), there is a need to apply models that span farms. Federated learning (FL) has emerged as a potential solution. FL enables decentralized ML across different farms without sharing private data. Traditional FL assumes simple two-tier network topologies and, thus, falls short of operating on more complex networks found in real-world agricultural scenarios. Networks vary across crops and farms and encompass various sensor data modes, extending across jurisdictions. New hierarchical FL (HFL) approaches are needed for more efficient and context-sensitive model sharing, accommodating regulations across multiple jurisdictions. Here, we present the RuralAI architecture deployment for tomato crop monitoring, featuring sensor field units for soil, crop, and weather data collection. HFL with personalization is used to offer localized and adaptive insights. Model management, aggregation, and transfers are facilitated via a flexible approach, enabling seamless communication between local devices, edge nodes, and the cloud.

60 APPLIED LIFE SCIENCES↗

Phonon-informed Neural Thermal Scattering (NeTS) Optimization for Crystalline Graphite and Beryllium Metal

Fast neutrons born from fission lose energy through scattering interactions in the process of slowing-down. As neutrons thermalize to the order of $k$ $b$ $T$ (where $k$ $b$ is the Boltzmann constant, and $T$ is the temperature of the medium), their de Broglie wavelength and energy approaches the order of inter-atomic spacing and quantized lattice vibrations, i.e., phonons. At thermal energies, the thermal scattering law (TSL), i.e., $S$($α, β$), captures crystal binding contributions to the total reaction rate, or cross section. This dimensionless material property describes the energy ($β$) and momentum ($α$) exchanges available in a medium. Currently, $S$($α, β$) is evaluated in the Full Law Analysis Scattering System Hub (FLASSH) code for discrete inputs and stored as ENDF/B File 7 for 0-phonon elastic (MT 2) and n-phonon inelastic (MT 4) processes. Further processing recasts $S$($α, β$) into cumulative distribution functions for sampling post-collision scattering kinematics. In practice, interpolation schemes are employed to access data between tabulated values. An improvement to this juncture of the nuclear data pipeline is supplying cross sections on-the-fly (OTF), as has been developed for the un-resolved resonance region to minimize non-physical interpolation errors. This capability may improve simulation accuracy for accident and transient analyses, where rapidly varying changes in temperature and pressure are difficult to predict beforehand. To do so, deep artificial neural networks (ANNs) can be employed which collapse non-linear, complex data into a lightweight dictionary of neural weights and biases. This has been successfully demonstrated for the hydrogen in light water $S$($α, β$) dataset in the form of a Neural Thermal Scattering (NeTS) module. In this work, the NeTS framework is extended to consider the impact of material-dependent dynamical features on optimal neural pre-processing and architecture design decisions, such as number of neurons per hidden layer, residual skip connections and neural depth. New NeTS modules for crystalline graphite and beryllium metal illuminate a novel correlation between dynamical nonlinearity and optimal neural parametrization when deploying $S$($α, β$) on-the-fly.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Impact of Domain Knowledge on the Property Prediction of Specialized Machine Learning Models

Developing transferable machine learning models is trending in data-driven materials research. However, how to apply such models to a specific research domain remains unclear. Here, in this work, we choose high-entropy materials as a platform with a specialized data set containing 145,323 DFT-relaxed materials. This data set is used to explore the role of domain-specific knowledge in training effective models. Our tests with three representative graph neural network architectures indicate the model complexity has much smaller influence on performance than the data itself. Specifically, the consideration of low-energy atomic ordering, structures with diverse elemental coverage, and high-order interactions significantly influences the model performance. We also find that domain knowledge-driven sampling can greatly enhance unsupervised learning techniques. This research highlights that developing specialized data sets is more beneficial than further complicating deep learning architectures. Additionally, physics-inspired sampling algorithms are crucially needed for better machine learning models for a specific materials research domain.

36 MATERIALS SCIENCE↗

Neural network-based control of an ultrafast laser

With the recent advances in machine learning (ML) and data science (DS), the control, modeling, and analysis of these complex systems continues to improve. In this work, we report on the optimization of the intensity of a femtosecond laser using feedforward neural networks (FFNN) that model the input–output relationships of the data. The input parameters of the system were optimized to achieve the required performance of the femtosecond laser. We propose a neural network-based control system to model the relationship between the spectral amplitude and phase of the input laser pulse at the amplifier input and the shape of the output pulse. Low-jitter laser parameter inputs and the resulting laser pulse duration were modeled, and the resulting correlation between the input and output data was used to optimize the laser pulse. Here, we demonstrate improved processing and laser control performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

CTGAN-TVAE

SAND2026-18914O CTGAN-TVAE (Conditional Tabular Generative Adversarial Networks-Tabular Variational Autoencoders) generates extensive sets of variable generation data through a hybrid framework. It enhances latent space representation by combining TVAE's robust feature-embedding with CTGAN's ability to condition categorical variables such as time. CTGAN-TVAE employs a fully connected neural network within a conditional generative adversarial network framework to manage continuous and categorical data effectively, capturing complex feature interactions without needing sequential modeling. This was developed as part of NNSA-MSIPP: Minority Serving Institution Partnership Program, Grant Number DE-NA0004016. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Newlun, Cody [Sandia National Lab. (SNL-CA), Liver↗