Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Complex Network Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Discovery of Probabilistic Dirichlet-to-Neumann Maps on Graphs

Dirichlet-to-Neumann maps enable the coupling of multiphysics simulations across computational subdomains by ensuring continuity of state variables and fluxes at artificial interfaces. We present a novel method for learning Dirichlet-to-Neumann maps on graphs using Gaussian processes, specifically for problems where the data obey a conservation law arising from an underlying partial differential equation. Our approach combines discrete exterior calculus and nonlinear optimal recovery to infer relationships between vertex and edge values. This framework yields data-driven predictions with uncertainty quantification across the entire graph, even when observations are limited to a subset of vertices and edges. By minimizing the reproducing kernel Hilbert space norm while penalizing kernel complexity through maximum likelihood estimation, our method ensures that the resulting surrogate strictly enforces conservation laws without overfitting. We demonstrate our method on two representative applications: subsurface flow in fracture networks and arterial blood flow. Finally, the results demonstrate that the method maintains high accuracy and well-calibrated uncertainty estimates even under severe data scarcity, highlighting its potential for scientific applications where limited data and reliable uncertainty quantification are critical.

Dirichlet-to-Neumann map↗

Optimizing Cryo-Focused Pyrolysis GC/MS for Tracing Soil Organic Matter Across Diverse Ecosystems

The cycling of organic matter in terrestrial soils and sediments is central to a range of biogeochemical processes that regulate nutrient cycling, crop productivity, trace gas emissions, and contaminant transport. Pyrolysis-gas chromatography/mass spectrometry (py-GC/MS) is a powerful tool for characterizing bulk soil organic matter (SOM) at the molecular level. In this study, we used a cryo-focused py-GC/MS system to analyze soil samples from seven diverse ecosystems: vernal pool, prairie pothole, temperate forest, tropical forest, tundra, wildfire-affected boreal forest, and grassland. We addressed a key bottleneck in molecular-level SOM characterization by developing an automated data analysis pipeline to optimize py-GC/MS and complementary evolved gas analysis/mass spectrometry (EGA/MS) methods, incorporating advanced tools for peak deconvolution, developing a custom compound class library, and implementing fragmentation spectrum-based molecular networking for the first time. This improved workflow was applied to soil samples from all seven ecosystems, including multiple depths and density fractions. Our findings demonstrate that ecosystem type plays a dominant role in shaping compositional differences in SOM. We also identified trends in the source of SOM compounds (e.g., microbial vs plantderived) across soil depth and density fractions, which are critical for understanding persistence and turnover of SOM. Our molecular networking analysis indicated that although many compounds are widespread across ecosystems, others are restricted to specific environments, such as wetlands. This underscores the utility of molecular-level data in elucidating the complexity of SOM composition and the environmental drivers that shape it. Such molecular-level insights can deepen our knowledge of biogeochemical SOM cycles.

54 ENVIRONMENTAL SCIENCES↗

Multi-head physics-informed neural networks for learning functional priors and uncertainty quantification

In numerous applications, the integration of prior knowledge and historical information is essential, particularly for tasks requiring the solution of ordinary or partial differential equations (ODEs/PDEs) in data-sparse or noisy environments. For instance, achieving accurate solutions to time-dependent PDEs with limited initial condition measurements necessitates an effective strategy for embedding prior knowledge. Hard-parameter sharing architectures in neural networks (NNs) have demonstrated success in both traditional and scientific machine learning domains, facilitating the learning of informative representations. Here, in this study, we introduce a novel, yet efficient, method to enhance physics-informed neural networks (PINNs) by incorporating a multi-head structure that enables the learning of functional priors from both empirical data and governing physical laws. This prior information can then be used to address data sparsity and high-level noise in solving ODE/PDE problems with uncertainty quantification (UQ). The approach, termed Multi-Head PINN (MH-PINN), consists of a shared body NN and multiple head NNs, each corresponding to an individual PINN instance. Our framework for functional prior learning is carried out in two stages: (1) training the MH-PINNs to develop a shared body NN alongside multiple head NNs, and (2) employing these trained head NNs to estimate a prior distribution through a normalizing flow-based density estimator. The learned functional prior can then be applied as a regularization mechanism in deterministic contexts or as an informative prior within a Bayesian inference framework, aiding in the resolution of subsequent ODE/PDE tasks. We evaluate the efficacy of MH-PINNs across five benchmark problems, including a high-dimensional parametric PDE, all characterized by data sparsity or substantial noise levels. Our findings reveal that MH-PINNs deliver accurate solutions and robust UQ, demonstrating adaptability across a range of complex and challenging scenarios.

Bayesian inference↗

Machine-Learning-Based High-Resolution Earthquake Catalog Reveals How Complex Fault Structures Were Activated during the 2016–2017 Central Italy Sequence

The 2016–2017 central Italy seismic sequence occurred on an 80 km long normal-fault system. The sequence initiated with the Mw 6.0 Amatrice event on 24 August 2016, followed by the Mw 5.9 Visso event on 26 October and the Mw 6.5 Norcia event on 30 October. We analyze continuous data from a dense network of 139 seismic stations to build a high-precision catalog of ~900,000 earthquakes spanning a 1 yr period, based on arrival times derived using a deep-neural-network-based picker. Our catalog contains an order of magnitude more events than the catalog routinely produced by the local earthquake monitoring agency. Aftershock activity reveals the geometry of complex fault structures activated during the earthquake sequence and provides additional insights into the potential factors controlling the development of the largest events. Activated fault structures in the northern and southern regions appear complementary to faults activated during the 1997 Colfiorito and 2009 L’Aquila sequences, suggesting that earthquake triggering primarily occurs on critically stressed faults. Delineated major fault zones are relatively thick compared to estimated earthquake location uncertainties, and a large number of kilometer-long faults and diffuse seismicity were activated during the sequence. These properties might be related to fault age, roughness, and the complexity of inherited structures. The rich details resolvable in this catalog will facilitate continued investigation of this energetic and well-recorded earthquake sequence.

58 GEOSCIENCES↗

Inverse Analysis with Variational Autoencoders: A Comparison of Shallow and Deep Networks

Inverse problems are applied to determine unknown properties by matching observational data with a physical model that often takes many parameters as input. To overcome the underconstrained nature of inverse problems and achieve good performance, an approach is presented involving regularization with a technique known as a variational autoencoder (VAE), which is trained to map a high-dimensional parameter space with a complex structure to a low-dimensional latent space with a simple structure. We apply this approach to unconditioned realizations of the parameters (heterogeneous hydraulic fields) for a hydrogeological inverse problem. Two types of hydraulic conductivity fields are used to evaluate the characterization for the different levels of heterogeneity complexity of the physical inputs. This approach keeps the computational cost of generating the training data low. The reason is unconditioned realizations neither rely on the observational data used to perform the inverse analysis nor require any groundwater flow model forward runs. In addition, this approach applies regularization on a low-dimensional latent space from the VAE and increases optimization efficiency through automatic differentiation. Furthermore, two different neural network (NN) structures are tested for their utility in using the VAE for inverse analysis. The performance of a deep, convolutional neural network strongly depends on the dimensionality of the latent space, which requires tuning. In contrast, a shallow, dense neural network provides consistently accurate characterization without tuning. Furthermore, our approach evaluates the advantages of the shallow, dense neural network over the deep, convolutional one and enables future application to a wide range of inverse problems.

97 MATHEMATICS AND COMPUTING↗

Neural-network learning of SPOD latent dynamics

Here, we aim to reconstruct the latent space dynamics of high dimensional, quasi-stationary systems using model order reduction via the spectral proper orthogonal decomposition (SPOD). The proposed method is based on three fundamental steps: in the first, once that the mean flow field has been subtracted from the realizations (also referred to as snapshots), we compress the data from a high-dimensional representation to a lower dimensional one by constructing the SPOD latent space; in the second, we build the time-dependent coefficients by projecting the snapshots containing the fluctuations onto the SPOD basis and we learn their evolution in time with the aid of recurrent neural networks; in the third, we reconstruct the high-dimensional data from the learnt lower -dimensional representation. The proposed method is demonstrated on two different test cases, namely, a compressible jet flow, and a geophysical problem known as the Madden-Julian Oscillation. An extensive comparison between SPOD and the equivalent POD-based counterpart is provided and differences between the two approaches are highlighted. The numerical results suggest that the proposed model is able to provide low rank predictions of complex statistically stationary data and to provide insights into the evolution of phenomena characterized by specific range of frequencies. The comparison between POD and SPOD surrogate strategies highlights the need for further work on the characterization of the interplay of error between data reduction techniques and neural network forecasts.

97 MATHEMATICS AND COMPUTING↗

Design, execution, and interpretation of plant RNA-seq analyses

Genomics has transformed our understanding of the genetic architecture of traits and the genetic variation present in plants. Here, we present a review of how RNA-seq can be performed to tackle research challenges addressed by plant sciences. We discuss the importance of experimental design in RNA-seq, including considerations for sampling and replication, to avoid pitfalls and wasted resources. Approaches for processing RNA-seq data include quality control and counting features, and we describe common approaches and variations. Though differential gene expression analysis is the most common analysis of RNA-seq data, we review multiple methods for assessing gene expression, including detecting allele-specific gene expression and building co-expression networks. With the production of more RNA-seq data, strategies for integrating these data into genetic mapping pipelines is of increased interest. Finally, special considerations for RNA-seq analysis and interpretation in plants are needed, due to the high genome complexity common across plants. By incorporating informed decisions throughout an RNA-seq experiment, we can increase the knowledge gained.

59 BASIC BIOLOGICAL SCIENCES↗

Improving High-Energy Particle Detectors with Machine Learning

Microseconds after the Big Bang, the universe existed in a state called the quark-gluon plasma (QGP). To experimentally study its properties, the QGP is recreated in high-energy nuclear collisions at the LHC, and the particles produced from the QGP are reconstructed from their energy deposition in the ATLAS calorimeter. This requires both classifying the particles and calibrating their deposited energy. The objective of this project is to improve the reconstruction by using machine learning techniques, where the energy depositions of clusters of cells, formed by ATLAS topo-clustering methods, are treated as three-dimensional images when inputted to neural networks. This approach significantly improves the calibration of deposited energies when cross-validating while training, and models trained on idealized data predict the calibrated energies of particles in more complex data sets well. Additionally, implementation of a data generator using uproot allows the program to load input data into memory as needed while training or predicting, significantly reducing the amount of memory used. The data generator also allows for use of multiprocessing to speed up training and evaluating. This work illustrates that using machine learning methods for both classification and calibration has the potential to significantly improve particle reconstruction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DECAL MDN Resolution Calibration

The simulated pixelated calorimeter uses 0.1 mm × 0.1 mm silicon pixels (0.0001 cm²) silicon pixels as active layers — far finer than current concepts like HGCAL (~0.5 cm²) — improving energy resolution through much finer segmentation. While the standard resolution formalism fits three numbers to a handful of discrete test-beam energies, this work is a first try at a more data-efficient alternative: learning the full response shape continuously in energy with a mixture density network. This continuous, differentiable surrogate is a natural building block for fast simulation of extremely complex calorimeters.

Wang, Peter [U. Chicago (main)]↗

Wabash CarbonSAFE. Final Report

This document summarizes work detailed in separate Wabash CarbonSAFE reports; the report describes the data collection efforts of the project and consolidates the geologic characterization, well testing, and storage complex modeling results for the Mt. Simon Sandstone and Potosi Dolomite, two distinct reservoirs characterized at the Wabash CarbonSAFE project site. Also presented are summaries of work to characterize the CO 2 source and infrastructure network, as well as summaries of reports analyzing stakeholder engagement options, policy, regulatory and legal considerations, and risk assessment associated with the Wabash CarbonSAFE project site. The report then presents recommendations for the next steps for site characterization, identifies data gaps for future activities, and provides an overall assessment of site potential.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Inpainting radar missing data regions with deep learning

Abstract. Missing and low-quality data regions are a frequent problem for weather radars. They stem from a variety of sources: beam blockage, instrument failure, near-ground blind zones, and many others. Filling in missing data regions is often useful for estimating local atmospheric properties and the application of high-level data processing schemes without the need for preprocessing and error-handling steps – feature detection and tracking, for instance. Interpolation schemes are typically used for this task, though they tend to produce unrealistically spatially smoothed results that are not representative of the atmospheric turbulence and variability that are usually resolved by weather radars. Recently, generative adversarial networks (GANs) have achieved impressive results in the area of photo inpainting. Here, they are demonstrated as a tool for infilling radar missing data regions. These neural networks are capable of extending large-scale cloud and precipitation features that border missing data regions into the regions while hallucinating plausible small-scale variability. In other words, they can inpaint missing data with accurate large-scale features and plausible local small-scale features. This method is demonstrated on a scanning C-band and vertically pointing Ka-band radar that were deployed as part of the Cloud Aerosol and Complex Terrain Interactions (CACTI) field campaign. Three missing data scenarios are explored: infilling low-level blind zones and short outage periods for the Ka-band radar and infilling beam blockage areas for the C-band radar. Two deep-learning-based approaches are tested, a convolutional neural network (CNN) and a GAN that optimize pixel-level error or combined pixel-level error and adversarial loss respectively. Both deep-learning approaches significantly outperform traditional inpainting schemes under several pixel-level and perceptual quality metrics.

54 ENVIRONMENTAL SCIENCES↗

A general-purpose method for Pareto optimal placement of flow rate and concentration sensors in networked systems – With application to wastewater treatment plants

The advent of affordable computing, low-cost sensor hardware, and high-speed and reliable communications have spurred installation of ubiquitous sensors in complex engineered systems. However, ensuring reliable data quality remains a challenge. Exploitation of redundancy among sensor signals can help improving the precision of measured variables, detecting the presence of gross errors, and identifying faulty sensors. The cost of sensor ownership, maintenance efforts in particular, can still be cost-prohibitive however. Maximizing the ability to assess and control data quality while minimizing the cost of ownership thus requires a careful sensor placement. To solve this challenge, in this work we develop a generally applicable method to solve the multi-objective sensor placement problem in systems governed by linear and bilinear balance equations. Importantly, the method computes all Pareto-optimal sensor layouts with conventional computational resources and requires no information about the expected sensor quality.

42 ENGINEERING↗

Performance Prediction of Big Data Transfer Through Experimental Analysis and Machine Learning

Big data transfer in next-generation scientific applications is now commonly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) through advance bandwidth reservation. To use HPN resources efficiently, provisioning agents need to carefully schedule data transfer requests and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized by the requesting user, could be simply wasted or cause extra overhead and complexity in management due to exclusive access. This calls for the capability of performance prediction to reserve bandwidth resources that match actual needs. Towards this goal, we employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements, which are collected over a span of several years from a large number of data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated HPN testbeds. We first identify a comprehensive list of attributes involved in a typical big data transfer process, including end host system configurations, network connection properties, and control parameters of data transfer methods. We then conduct an in-depth exploratory analysis of their impacts on application-level throughput, which provides insights into big data transfer performance and motivates the use of machine learning. We also investigate the applicability of machine learning algorithms and derive their general performance bounds for performance prediction of big data transfer in HPNs. Experimental results show that, with appropriate data preprocessing, the proposed machine learning-based approach achieves 95% or higher prediction accuracy in up to 90% of the cases with very noisy real-life performance measurements.

Yun, Daqing↗

Constructing Neural Network Based Models for Simulating Dynamical Systems

Dynamical systems see widespread use in natural sciences like physics, biology, and chemistry, as well as engineering disciplines such as circuit analysis, computational fluid dynamics, and control. For simple systems, the differential equations governing the dynamics can be derived by applying fundamental physical laws. However, for more complex systems, this approach becomes exceedingly difficult. Data-driven modeling is an alternative paradigm that seeks to learn an approximation of the dynamics of a system using observations of the true system. In recent years, there has been an increased interest in applying data-driven modeling techniques to solve a wide range of problems in physics and engineering. Here this article provides a survey of the different ways to construct models of dynamical systems using neural networks. In addition to the basic overview, we review the related literature and outline the most significant challenges from numerical simulations that this modeling paradigm must overcome. Based on the reviewed literature and identified challenges, we provide a discussion on promising research areas.

97 MATHEMATICS AND COMPUTING↗

MetaboDirect: an analytical pipeline for the processing of FT-ICR MS-based metabolomic data

Background: Microbiomes are now recognized as the main drivers of ecosystem function ranging from the oceans and soils to humans and bioreactors. However, a grand challenge in microbiome science is to characterize and quantify the chemical currencies of organic matter (i.e., metabolites) that microbes respond to and alter. Critical to this has been the development of Fourier transform ion cyclotron resonance mass spectrometry (FT-ICR MS), which has drastically increased molecular characterization of complex organic matter samples, but challenges users with hundreds of millions of data points where readily available, user-friendly, and customizable software tools are lacking. Results: Here, we build on years of analytical experience with diverse sample types to develop MetaboDirect, an open-source, command-line-based pipeline for the analysis (e.g., chemodiversity analysis, multivariate statistics), visualization (e.g., Van Krevelen diagrams, elemental and molecular class composition plots), and presentation of direct injection high-resolution FT-ICR MS data sets after molecular formula assignment has been performed. When compared to other available FT-ICR MS software, MetaboDirect is superior in that it requires a single line of code to launch a fully automated framework for the generation and visualization of a wide range of plots, with minimal coding experience required. Among the tools evaluated, MetaboDirect is also uniquely able to automatically generate biochemical transformation networks (ab initio) based on mass differences (mass difference network-based approach) that provide an experimental assessment of metabolite connections within a given sample or a complex metabolic system, thereby providing important information about the nature of the samples and the set of microbial reactions or pathways that gave rise to them. Finally, for more experienced users, MetaboDirect allows users to customize plots, outputs, and analyses. Conclusion: Application of MetaboDirect to FT-ICR MS-based metabolomic data sets from a marine phage-bacterial infection experiment and a Sphagnum leachate microbiome incubation experiment showcase the exploration capabilities of the pipeline that will enable the research community to evaluate and interpret their data in greater depth and in less time. It will further advance our knowledge of how microbial communities influence and are influenced by the chemical makeup of the surrounding system. The source code and User’s guide of MetaboDirect are freely available through (https://github.com/Coayala/MetaboDirect) and (https://metabodirect.readthedocs.io/en/latest/), respectively.

54 ENVIRONMENTAL SCIENCES↗

Locating Seismic Events with Local-Distance Data

As the seismic monitoring community advances toward detecting, identifying, and locating ever-smaller natural and anthropogenic events, the need is constantly increasing for higher resolution, higher fidelity data, models, and methods for accurately characterizing events. Local-distance seismic data provide robust constraints on event locations, but also introduce complexity due to the significant geologic heterogeneity of the Earth’s crust and upper mantle, and the relative sparsity of data that often occurs with small events recorded on regional seismic networks. Identifying the critical characteristics for improving local-scale event locations and the factors that impact location accuracy and reliability is an ongoing challenge for the seismic community. Using Utah as a test case, we examine three data sets of varying duration, finesse, and magnitude to investigate the effects of local earth structure and modeling parameters on local-distance event location precision and accuracy. We observe that the most critical elements controlling relocation precision are azimuthal coverage and local-scale velocity structure, with tradeoffs based on event depth, type, location, and range.

42 ENGINEERING↗

Inferring adversarial behaviour in cyber‐physical power systems using a Bayesian attack graph approach

Abstract Highly connected smart power systems are subject to increasing vulnerabilities and adversarial threats. Defenders need to proactively identify and defend new high‐risk access paths of cyber intruders that target grid resilience. However, cyber‐physical risk analysis and defense in power systems often requires making assumptions on adversary behaviour, and these assumptions can be wrong. Thus, this work examines the problem of inferring adversary behaviour in power systems to improve risk‐based defense and detection. To achieve this, a Bayesian approach for inference of the Cyber‐Adversarial Power System (Bayes‐CAPS) is proposed that uses Bayesian networks (BNs) to define and solve the inference problem of adversarial movement in the grid infrastructure towards targets of physical impact. Specifically, BNs are used to compute conditional probabilities to queries, such as the probability of observing an event given a set of alerts. Bayes‐CAPS builds initial Bayesian attack graphs for realistic power system cyber‐physical models. These models are adaptable using collected data from the system under study. Then, Bayes‐CAPS computes the posterior probabilities of the occurrence of a security breach event in power systems. Experiments are conducted that evaluate algorithms based on time complexity, accuracy and impact of evidence for different scales and densities of network. The performance is evaluated and compared for five realistic cyber‐physical power system models of increasing size and complexities ranging from 8 to 300 substations based on computation and accuracy impacts.

Sahu, Abhijeet↗

Data for Reshaping the 2-Pyrone Synthase Active Site for Chemoselective Biosynthesis of Polyketides

Engineering enzymes with novel reactivity and applying them in metabolic pathways to produce valuable products are quite challenging due to the intrinsic complexity of metabolic networks and the need for high in vivo catalytic efficiency. Triacetic acid lactone (TAL), naturally generated by 2-pyrone synthase (2PS), is a platform molecule that can be produced via microbial fermentation and further converted into value-added products. However, these conversions require extra synthetic steps under harsh conditions. We herein report a biocatalytic system for direct generation of TAL derivatives under mild conditions with controlled chemoselectivity by rationally engineering the 2PS active site and then rewiring the biocatalytic pathway in the metabolic network of E. coli to produce high-value products, such as kavalactone precursors, with yields up to 17 mg/L culture. Computer modeling indicates sterics and hydrogen-bond interactions play key roles in tuning the selectivity, efficiency, and yield.

Conversion↗