Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Complex Network Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

SymProp: Scaling Sparse Symmetric Tucker Decomposition via Symmetry Propagation

Sparse symmetric tensors are an important class of tensors, and their decompositions serve as powerful tools for revealing low-rank structures. This paper introduces SymProp, a novel approach for scaling sparse symmetric Tucker decomposition by propagating symmetry through intermediate computations. SymProp optimizes two key computational kernels: Sparse Symmetric Tensor Times Same Matrix chain (S3 TTMc) for Higher-Order Orthogonal Iteration (HOOI) and Sparse Symmetric Tensor Times Same Matrix chain Times Core (S3 TTMcTC) for Higher-Order QR Iteration (HOQRI). Our method employs a metaprogramming-based index iteration approach to efficiently handle the upper triangular parts of intermediate dense symmetric tensors. SymProp achieves up to 50.9× speedup over SPLATT and up to 360.8× over Compressed Sparse Symmetric (CSS) format on the S3 TTMc operation. Moreover, our S3 TTMc and S3 TTMcTC implementations support tensor orders four levels higher than state-of-the-art methods. Our HOQRI demonstrates superior scalability and up to a 33.6× speedup over optimized HOOI. By enabling more scalable Tucker decompositions for higher orders, decomposition ranks, and dimension sizes, SymProp opens new possibilities for analyzing complex hypergraph structures in fields such as network science, data mining, and machine learning.

Li, Zecheng [North Carolina State University]↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

RuralAI in Tomato Farming: Integrated Sensor System, Distributed Computing, and Hierarchical Federated Learning for Crop Health Monitoring

Precision horticulture is evolving due to scalable sensor deployment and machine learning (ML) integration. These advancements boost the operational efficiency of individual farms, balancing the benefits of analytics with autonomy requirements. However, given concerns that affect wide geographic regions (e.g., climate change), there is a need to apply models that span farms. Federated learning (FL) has emerged as a potential solution. FL enables decentralized ML across different farms without sharing private data. Traditional FL assumes simple two-tier network topologies and, thus, falls short of operating on more complex networks found in real-world agricultural scenarios. Networks vary across crops and farms and encompass various sensor data modes, extending across jurisdictions. New hierarchical FL (HFL) approaches are needed for more efficient and context-sensitive model sharing, accommodating regulations across multiple jurisdictions. Here, we present the RuralAI architecture deployment for tomato crop monitoring, featuring sensor field units for soil, crop, and weather data collection. HFL with personalization is used to offer localized and adaptive insights. Model management, aggregation, and transfers are facilitated via a flexible approach, enabling seamless communication between local devices, edge nodes, and the cloud.

60 APPLIED LIFE SCIENCES↗

Phonon-informed Neural Thermal Scattering (NeTS) Optimization for Crystalline Graphite and Beryllium Metal

Fast neutrons born from fission lose energy through scattering interactions in the process of slowing-down. As neutrons thermalize to the order of $k$ $b$ $T$ (where $k$ $b$ is the Boltzmann constant, and $T$ is the temperature of the medium), their de Broglie wavelength and energy approaches the order of inter-atomic spacing and quantized lattice vibrations, i.e., phonons. At thermal energies, the thermal scattering law (TSL), i.e., $S$($α, β$), captures crystal binding contributions to the total reaction rate, or cross section. This dimensionless material property describes the energy ($β$) and momentum ($α$) exchanges available in a medium. Currently, $S$($α, β$) is evaluated in the Full Law Analysis Scattering System Hub (FLASSH) code for discrete inputs and stored as ENDF/B File 7 for 0-phonon elastic (MT 2) and n-phonon inelastic (MT 4) processes. Further processing recasts $S$($α, β$) into cumulative distribution functions for sampling post-collision scattering kinematics. In practice, interpolation schemes are employed to access data between tabulated values. An improvement to this juncture of the nuclear data pipeline is supplying cross sections on-the-fly (OTF), as has been developed for the un-resolved resonance region to minimize non-physical interpolation errors. This capability may improve simulation accuracy for accident and transient analyses, where rapidly varying changes in temperature and pressure are difficult to predict beforehand. To do so, deep artificial neural networks (ANNs) can be employed which collapse non-linear, complex data into a lightweight dictionary of neural weights and biases. This has been successfully demonstrated for the hydrogen in light water $S$($α, β$) dataset in the form of a Neural Thermal Scattering (NeTS) module. In this work, the NeTS framework is extended to consider the impact of material-dependent dynamical features on optimal neural pre-processing and architecture design decisions, such as number of neurons per hidden layer, residual skip connections and neural depth. New NeTS modules for crystalline graphite and beryllium metal illuminate a novel correlation between dynamical nonlinearity and optimal neural parametrization when deploying $S$($α, β$) on-the-fly.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Impact of Domain Knowledge on the Property Prediction of Specialized Machine Learning Models

Developing transferable machine learning models is trending in data-driven materials research. However, how to apply such models to a specific research domain remains unclear. Here, in this work, we choose high-entropy materials as a platform with a specialized data set containing 145,323 DFT-relaxed materials. This data set is used to explore the role of domain-specific knowledge in training effective models. Our tests with three representative graph neural network architectures indicate the model complexity has much smaller influence on performance than the data itself. Specifically, the consideration of low-energy atomic ordering, structures with diverse elemental coverage, and high-order interactions significantly influences the model performance. We also find that domain knowledge-driven sampling can greatly enhance unsupervised learning techniques. This research highlights that developing specialized data sets is more beneficial than further complicating deep learning architectures. Additionally, physics-inspired sampling algorithms are crucially needed for better machine learning models for a specific materials research domain.

36 MATERIALS SCIENCE↗

Neural network-based control of an ultrafast laser

With the recent advances in machine learning (ML) and data science (DS), the control, modeling, and analysis of these complex systems continues to improve. In this work, we report on the optimization of the intensity of a femtosecond laser using feedforward neural networks (FFNN) that model the input–output relationships of the data. The input parameters of the system were optimized to achieve the required performance of the femtosecond laser. We propose a neural network-based control system to model the relationship between the spectral amplitude and phase of the input laser pulse at the amplifier input and the shape of the output pulse. Low-jitter laser parameter inputs and the resulting laser pulse duration were modeled, and the resulting correlation between the input and output data was used to optimize the laser pulse. Here, we demonstrate improved processing and laser control performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

CTGAN-TVAE

SAND2026-18914O CTGAN-TVAE (Conditional Tabular Generative Adversarial Networks-Tabular Variational Autoencoders) generates extensive sets of variable generation data through a hybrid framework. It enhances latent space representation by combining TVAE's robust feature-embedding with CTGAN's ability to condition categorical variables such as time. CTGAN-TVAE employs a fully connected neural network within a conditional generative adversarial network framework to manage continuous and categorical data effectively, capturing complex feature interactions without needing sequential modeling. This was developed as part of NNSA-MSIPP: Minority Serving Institution Partnership Program, Grant Number DE-NA0004016. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Newlun, Cody [Sandia National Lab. (SNL-CA), Liver↗

Distributed Real-time Plume Monitoring for Deep Sea Mineral Extraction​

In the emerging industry of deep-sea mining for minerals and deposits (e.g. polymetallic nodules for nickel, cobalt, copper, and manganese), more data is required to understand the effects of sediment plume generation and predict the distribution of disturbed sediment. There are two main sources of plume generation, the first being at the active mining site where the “collector” directly removes the top layer of the sea floor. The other is the “midwater plume” consisting of unwanted sediment that was collected during extraction that is pumped back into the aphotic zone. The vast majority of plume generation is caused by the collector, causing detrimental and long-lasting impacts on seafloor ecosystems due to the lack of wave activity or strong currents at the sea floor. Therefore, it is crucial to invest in the infrastructure to support the study and constant monitoring over a large area of the sea floor where plume generation is present. Due to the limited number of usable channels and power requirements, current subsea wireless communications technologies are not well suited to instrumenting the large areas of the sea floor needed to monitor plume migration. The scope of this effort is to transition experimental demonstrations of high-bandwidth, full-duplex scalable underwater laser communications to the seafloor in an open ocean environment. Specifically tackling challenges associated with the dynamic nature of the subsea world, including but not limited to, deployment logistics, sustainability, and range. The goal is to enable the internet of underwater things for deep sea industries by broadening the capabilities of subsea communications. By using high-precision laser transmitters, many of the challenges current subsea optical systems face can be circumvented, such as power consumption, interference, and bandwidth limitations. This approach lends itself to wireless interlinking multi-node networks, in series or parallel, facilitating the implementation of a wide array of sensor types. This interlinking allows all the data gathered from the network to be processed through a single hardline uplink to the surface, lowering the complexity required for near real-time data processing. Additionally, the laser control systems produce metadata that can be used to help characterize the water column between the nodes. Combining data from various sensors such as turbidity, temperature, current velocity with metadata such as beam attenuation and deflection can produce a high-resolution model of sea floor conditions around an active mining zone. The resulting near real-time model can be used to optimize location and flow rate of the mining operation to minimize and quantify the environmental impact.

Mons, Ishan↗

Data-driven modeling of coarse mesh turbulence for reactor transient analysis using convolutional recurrent neural networks

Advanced nuclear reactors often exhibit complex thermal-fluid phenomena during transients. To accurately capture such phenomena, a coarse-mesh three-dimensional (3-D) modeling capability is desired for modern nuclear-system code. In the coarse-mesh 3-D modeling of advanced-reactor transients that involve flow and heat transfer, accurately predicting the turbulent viscosity is a challenging task that requires an accurate and computationally efficient model to capture the unresolved fine-scale turbulence. In this work, we propose a data-driven coarse-mesh turbulence model based on local flow features for the transient analysis of thermal mixing and stratification in a sodium-cooled fast reactor. The model has a coarse mesh setup to ensure computational efficiency, while it is trained by fine-mesh computational fluid dynamics (CFD) data to ensure accuracy. A novel neural network architecture, combining a densely connected convolutional network and a long-short-term-memory network, is developed that can efficiently learn from the spatial temporal CFD transient simulation results. The neural network model was trained and optimized on a loss-of flow transient and demonstrated high accuracy in predicting the turbulent viscosity field during the whole transient. The trained model's generalization capability was also investigated on two other transients with different inlet conditions. The study demonstrates the potential of applying the proposed data-driven approach to support the coarse-mesh multi-dimensional modeling of advanced reactors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]↗

Vitamin interdependencies predicted by metagenomics-informed network analyses and validated in microbial community microcosms

Abstract Metagenomic or metabarcoding data are often used to predict microbial interactions in complex communities, but these predictions are rarely explored experimentally. Here, we use an organism abundance correlation network to investigate factors that control community organization in mine tailings-derived laboratory microbial consortia grown under dozens of conditions. The network is overlaid with metagenomic information about functional capacities to generate testable hypotheses. We develop a metric to predict the importance of each node within its local network environments relative to correlated vitamin auxotrophs, and predict that a Variovorax species is a hub as an important source of thiamine. Quantification of thiamine during the growth of Variovorax in minimal media show high levels of thiamine production, up to 100 mg/L. A few of the correlated thiamine auxotrophs are predicted to produce pantothenate, which we show is required for growth of Variovorax , supporting that a subset of vitamin-dependent interactions are mutualistic. A Cryptococcus yeast produces the B-vitamin pantothenate, and co-culturing with Variovorax leads to a 90-130-fold fitness increase for both organisms. Our study demonstrates the predictive power of metagenome-informed, microbial consortia-based network analyses for identifying microbial interactions that underpin the structure and functioning of microbial communities.

59 BASIC BIOLOGICAL SCIENCES↗

A Nuclear Security Enterprise Study of High-Reliability Systems, Collaboration, and Data

It may seem simple and trivial, but defining the difference between data and information is contested and has implications that may affect the security of United States interests and even cost lives. For security, data are raw facts or figures without context, while information is the compilation or articulation of data that forms context. Security depends on clarity in the differences between data and information and controlling them. Control is necessary to ensure that data and information are not inadvertently released to foreign governments, the public, or those without Need-to-Know. A primary concern in the practice of security is the control of data to avoid the inadvertent conversion to sensitive information. The complexity of this concern is further augmented when institutions are part of tightly coupled networks that informally share data and information. Additionally, those that share data as a function of legislative action—and/or formally integrate data and information system infrastructures—may be a higher security risk. This paper will present a case study that utilizes elements of literature from Knowledge Management and networks to tell a story of an issue in security—specifically, controlling the conversion of data to information.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Object detection with deep learning for rare event search in the GADGET II TPC

In the pursuit of identifying rare two-particle events within the GADGET II Time Projection Chamber (TPC), this paper presents a comprehensive approach for leveraging Convolutional Neural Networks (CNNs) and various data processing methods. To address the inherent complexities of 3D TPC track reconstructions, the data is expressed in 2D projections and 1D quantities. This approach capitalizes on the diverse data modalities of the TPC, allowing for the efficient representation of the distinct features of the 3D events, with no loss in topology uniqueness. Additionally, it leverages the computational efficiency of 2D CNNs and benefits from the extensive availability of pre-trained models. Given the scarcity of real training data for the rare events of interest, simulated events are used to train the models to detect real events. To account for potential distribution shifts when predominantly depending on simulations, significant perturbations are embedded within the simulations. This produces a broad parameter space that works to account for potential physics parameter and detector response variations and uncertainties. These parameter-varied simulations are used to train sensitive 2D CNN object detectors. When combined with 1D histogram peak detection algorithms, this multi-modal detection framework is highly adept at identifying rare, two-particle events in data taken during experiment 21072 at the Facility for Rare Isotope Beams (FRIB), demonstrating a 100% recall for events of interest. Here, we present the methods and outcomes of our investigation and discuss the potential future applications of these techniques.

Convolutional neural network↗

Mauka Energy FEVER Tool Dataset

Mauka Energy’s dataset, developed under the Forestry Electric Vehicle Energy Routing (FEVER) project and funded by the U.S. Department of Energy’s Small Business Innovation Research program, is a high-resolution geospatial resource designed to support energy modeling for electric log trucks in complex forestry environments. The dataset integrates detailed spatial and road network data to enable accurate simulation of vehicle performance across varied terrain. At its core, the dataset incorporates lidar-derived elevation models, road alignments, and surface classifications from Oregon State University’s McDonald-Dunn Research Forest. These data capture fine-scale variations in slope, curvature, and surface conditions across forest road systems, allowing for vehicle-level analysis of energy consumption and recovery. The dataset also includes data collected on the surrounding public and private road networks in Benton County, Oregon, used in real-world haul routes. These connecting segments provide critical context for modeling transitions between forest operations and regional transportation infrastructure, incorporating attributes such as grade profiles, elevation change, and speed constraints. This combined dataset underpins the development of Mauka Energy’s rolldown tool, which quantifies energy use and regenerative braking potential on downhill and variable-grade segments. By leveraging high-resolution terrain and road data, the FEVER project enables more accurate assessment of electric vehicle feasibility and performance in forestry applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Inference, Prediction, & Entropy-Rate Estimation of Continuous-Time, Discrete-Event Processes

Inferring models, predicting the future, and estimating the entropy rate of discrete-time, discrete-event processes is well-worn ground. However, a much broader class of discrete-event processes operates in continuous-time. Here, we provide new methods for inferring, predicting, and estimating them. The methods rely on an extension of Bayesian structural inference that takes advantage of neural network’s universal approximation power. Based on experiments with complex synthetic data, the methods are competitive with the state-of-the-art for prediction and entropy-rate estimation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Predicting Geologic Behavior in Carbon Storage Projects Using Graph Neural Network

This study was invited to presented at NVIDIA's GTC conference to highlight the potential of Graph Neural Network as a novel and promising methodology for predicting pressure and saturation evolution in carbon storage projects. Carbon capture and storage (CCS) technology plays a pivotal role in mitigating greenhouse gas emissions, facilitating the transition to a low-carbon future. Effective management of subsurface reservoirs is essential to ensure the safe and efficient storage of captured carbon dioxide (CO₂). Accurate predictions of pressure and saturation over time are critical for evaluating the long-term performance and integrity of CCS projects. In recent years, Graph Neural Network (GNN) has emerged as a powerful framework for analyzing complex data in graph-structured domains. This abstract explores the application of GNN to forecast pressure and saturation evolution in carbon storage projects. Traditional numerical simulations of subsurface reservoirs have proven successful in providing pressure and saturation forecasts. However, these simulations involve massive amounts of computational effort and require extensive domain expertise for proper model calibration and validation. Graph Neural Operator offers an alternative approach that harnesses the inherent graph structure of reservoirs, where nodes represent reservoir grid cells and edges represent the geological connectivity between them.

Shih, Chung Yan↗

Domain Adaptive Graph Neural Networks for Constraining Cosmological Parameters Across Multiple Data Sets

State of the art astronomical simulations have provided datasets which enabled the training of novel deep learning techniques for constraining cosmological parameters. However, differences in subgrid physics implementation and numerical approximations among simulation suites lead to differences in simulated datasets, which pose a hard challenge when trying to generalize across diverse data domains and ultimately when applying models to observational data. Recent work reveals deep learning algorithms are able to extract more information from complex cosmological simulations than summary statistics like power spectra. We introduce Domain Adaptive Graph Neural Networks (DA-GNNs), trained on CAMELS data, inspired by CosmoGraphNet (Villanueva-Domingo et al 2023). By utilizing GNNs, we can capitalize on their capacity to capture both astrophysical and topological features of galaxy distributions. Mixing these capabilities with domain adaptation techniques such as Maximum Mean Discrepancy (MMD), which enable extraction of domain-invariant features, our framework demonstrates enhanced accuracy and robustness. We present experimental results, including the alignment of distributions across domains through data visualization. These findings suggest that DA-GNNs are an efficient way of extracting domain independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Prediction of plant complex traits via integration of multi-omics data

The formation of complex traits is the consequence of genotype and activities at multiple molecular levels. However, connecting genotypes and these activities to complex traits remains challenging. Here, we investigate whether integrating genomic, transcriptomic, and methylomic data can improve prediction for six Arabidopsis traits. We find that transcriptome- and methylome-based models have performances comparable to those of genome-based models. However, models built for flowering time using different omics data identify different benchmark genes. Nine additional genes identified as important for flowering time from our models are experimentally validated as regulating flowering. Gene contributions to flowering time prediction are accession-dependent and distinct genes contribute to trait prediction in different genotypes. Models integrating multi-omics data perform best and reveal known and additional gene interactions, extending knowledge about existing regulatory networks underlying flowering time determination. These results demonstrate the feasibility of revealing molecular mechanisms underlying complex traits through multi-omics data integration.

59 BASIC BIOLOGICAL SCIENCES↗