Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Complex Network Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

New trends in photonic switching and optical networking architectures for data centers and computing systems [Invited]

The rapid increases in data traffic coupled with user preferences are driving the data center and computing system service providers to offer energy-efficient, intelligent, flexible, cost-effective, high-capacity, and low-latency data services without added complexity to the users. Disaggregated heterogeneous reconfigurable computing systems realized by photonic switching and interconnects can enhance throughput and energy efficiency for artificial intelligence/machine learning (AI/ML) workloads, especially when aided by the AI/ML-enhanced control plane. Photonic switching and new optical networking architectures are expected to solve many of these challenging problems. This paper discusses new trends in photonic switching and optical network architectures for future data centers and computing systems summarized as follows: (1) flat reconfigurable disaggregated computing enabled by high-radix photonic switching and interconnects in data centers; (2) chiplet-based computing architectures empowered by embedded photonics toward heterogeneous reconfigurable computing; (3) nanosecond-scale photonic switching in data centers and computing systems; (4) AI/ML in self-driving, application-aware, and situation-aware data centers; (5) the emergence of flexible networking for cloud computing, edge computing, and split computing, as well as flexible networking for 5G/6G RF-optical networks; and (6) the deployment of embedded co-designed silicon photonics being considered for future data centers.

Yoo, S. J. Ben (ORCID:0000000274201871)↗

When less is more: How increasing the complexity of machine learning strategies for geothermal energy assessments may not lead toward better estimates

Previous moderate- and high-temperature geothermal resource assessments of the western United States utilized data-driven methods and expert decisions to estimate resource favorability. Although expert decisions can add confidence to the modeling process by ensuring reasonable models are employed, expert decisions also introduce human and, thereby, model bias. This bias can present a source of error that reduces the predictive performance of the models and confidence in the resulting resource estimates. Our study aims to develop robust data-driven methods with the goals of reducing bias and improving predictive ability. We present and compare nine favorability maps for geothermal resources in the western United States using data from the U.S. Geological Survey's 2008 geothermal resource assessment. Two favorability maps are created using the expert decision-dependent methods from the 2008 assessment (i.e., weight-of-evidence and logistic regression). With the same data, we then create six different favorability maps using logistic regression (without underlying expert decisions), XGBoost, and support-vector machines paired with two training strategies. The training strategies are customized to address the inherent challenges of applying machine learning to the geothermal training data, which have no negative examples and severe class imbalance. We also create another favorability map using an artificial neural network. We demonstrate that modern machine learning approaches can improve upon systems built with expert decisions. We also find that XGBoost, a non-linear algorithm, produces greater agreement with the 2008 results than linear logistic regression without expert decisions, because the expert decisions in the 2008 assessment rendered the otherwise linear approaches non-linear despite the fact that the 2008 assessment used only linear methods. The F1 scores for all approaches appear low (F1 score < 0.10), do not improve with increasing model complexity, and, therefore, indicate the fundamental limitations of the input features (i.e., training data). Until improved feature data are incorporated into the assessment process, simple non-linear algorithms (e.g., XGBoost) perform equally well or better than more complex methods (e.g., artificial neural networks) and remain easier to interpret.

15 GEOTHERMAL ENERGY↗

Topological structure of complex predictions

Abstract Current complex prediction models are the result of fitting deep neural networks, graph convolutional networks or transducers to a set of training data. A key challenge with these models is that they are highly parameterized, which makes describing and interpreting the prediction strategies difficult. We use topological data analysis to transform these complex prediction models into a simplified topological view of the prediction landscape. The result is a map of the predictions that enables inspection of the model results with more specificity than dimensionality-reduction methods such as tSNE and UMAP. The methods scale up to large datasets across different domains. We present a case study of a transformer-based model previously designed to predict expression levels of a piece of DNA in thousands of genomic tracks. When the model is used to study mutations in the BRCA1 gene, our topological analysis shows that it is sensitive to the location of a mutation and the exon structure of BRCA1 in ways that cannot be found with tools based on dimensionality reduction. Moreover, the topological framework offers multiple ways to inspect results, including an error estimate that is more accurate than model uncertainty. Further studies show how these ideas produce useful results in graph-based learning and image classification.

Computer Science↗

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

Experimental Testing of Data Fusion in A Distributed Ground-Based Sensing Network for Advanced Air Mobility

Advanced Air Mobility (AAM) is an active area of development which foresees the integration of autonomous uncrewed aircraft into the civil airspace for air transportation of people and cargo. Safe integration requires significant technological developments and extensive testing phases of sensing and surveillance strategies in dense airspace. Compared to well-assessed manned aviation systems scenarios, surveillance strategies in the AAM and small Uncrewed Aircraft Vehicles(UAVs) context need to detect smaller platforms flying at lower altitude against cluttered backgrounds in dense airspace. Fusion of data provided by a network of distributed sensing nodes is a powerful tool to enable detection and tracking in such complex conditions. This paper contributes to this research direction by proposing a surveillance strategy for the AAM environment based on sensor fusion of data acquired by distributed ground-based radars. Specifically, experimental data collected with two independent radars, observing the flight of two small UAVs, are used. Data fusion at tracking level is based on a leader-helper strategy where the leader radar uses the helper’s measurements to increase the lifespan of its generated tracks. This solution shows promising results with a 10%increase in track coverage with respect to the standalone leader radar tracking solution. The paper also proposes an interference removal processing method which is applied on the data collected by one of the two radars.

Federica Vitiello↗

Demonstration of Data Processing and Fusion from Distributed Radars for AAM Surveillance

Advanced Air Mobility (AAM) is an active area of development which foresees the integration of autonomous uncrewed aircraft into the civil airspace for air transportation of people and cargo. Safe integration requires significant technological developments and extensive testing phases of sensing and surveillance strategies in dense airspace. Compared to well-assessed manned aviation systems scenarios, surveillance strategies in the AAM and small Uncrewed Aircraft Vehicles (UAVs) context need to detect smaller platforms flying at lower altitude against cluttered backgrounds in dense airspace. Fusion of data provided by a network of distributed sensing nodes is a powerful tool to enable detection and tracking in such complex conditions. This paper contributes to this research direction by proposing a surveillance strategy for the AAM environment based on sensor fusion of data acquired by distributed ground-based radars. Specifically, experimental data collected with three independent radars, observing the flight of two small UAVs, are used. Data fusion at tracking level is based on a leader-helper strategy where the leader radar uses the helper’s measurements to increase the lifespan of its generated tracks. This solution shows promising results with a 10% increase in track coverage with respect to the standalone leader radar tracking solution. The paper also proposes an interference removal processing method which is applied on the data collected by two of the radars.

Federica Vitiello↗

Data-Efficient Dimensionality Reduction and Surrogate Modeling of High-Dimensional Stress Fields

Tensor datatypes representing field variables like stress, displacement, velocity, etc., have increasingly become a common occurrence in data-driven modeling and analysis of simulations. Numerous methods [such as convolutional neural networks (CNNs)] exist to address the meta-modeling of field data from simulations. As the complexity of the simulation increases, so does the cost of acquisition, leading to limited data scenarios. Modeling of tensor datatypes under limited data scenarios remains a hindrance for engineering applications. Here, in this article, we introduce a direct image-to-image modeling framework of convolutional autoencoders enhanced by information bottleneck loss function to tackle the tensor data types with limited data. The information bottleneck method penalizes the nuisance information in the latent space while maximizing relevant information making it robust for limited data scenarios. The entire neural network framework is further combined with robust hyperparameter optimization. We perform numerical studies to compare the predictive performance of the proposed method with a dimensionality reduction-based surrogate modeling framework on a representative linear elastic ellipsoidal void problem with uniaxial loading. The data structure focuses on the low-data regime (fewer than 100 data points) and includes the parameterized geometry of the ellipsoidal void as the input and the predicted stress field as the output. The results of the numerical studies show that the information bottleneck approach yields improved overall accuracy and more precise prediction of the extremes of the stress field. Additionally, an in-depth analysis is carried out to elucidate the information compression behavior of the proposed framework.

artificial intelligence↗

Applying a Space-Based Security Recovery Scheme for Critical Homeland Security Cyberinfrastructure Utilizing the NASA Tracking and Data Relay (TDRS) Based Space Network

Protection of the national infrastructure is a high priority for cybersecurity of the homeland. Critical infrastructure such as the national power grid, commercial financial networks, and communications networks have been successfully invaded and re-invaded from foreign and domestic attackers. The ability to re-establish authentication and confidentiality of the network participants via secure channels that have not been compromised would be an important countermeasure to compromise of our critical network infrastructure. This paper describes a concept of operations by which the NASA Tracking and Data Relay (TDRS) constellation of spacecraft in conjunction with the White Sands Complex (WSC) Ground Station host a security recovery system for re-establishing secure network communications in the event of a national or regional cyberattack. Users would perform security and network restoral functions via a Broadcast Satellite Service (BSS) from the TDRS constellation. The BSS enrollment only requires that each network location have a receive antenna and satellite receiver. This would be no more complex than setting up a DIRECTTV-like receiver at each network location with separate network connectivity. A GEO BSS would allow a mass re-enrollment of network nodes (up to nationwide) simultaneously depending upon downlink characteristics. This paper details the spectrum requirements, link budget, notional assets and communications requirements for the scheme. It describes the architecture of such a system and the manner in which it leverages off of the existing secure infrastructure which is already in place and managed by the NASAGSFC Space Network Project.

Cybersecurity↗

Recent Troposheric Ozone Observations from Satellite and In-Situ Measurements: An Overview

The past 4-5 years have seen an unprecedented increase in tropospheric ozone (and related) observations thanks to focused campaigns, new satellite data, and to innovative approaches to climatology (commercial aircraft sampling, networks). The result has been new insights into the complex chemistry and dynamics affecting tropospheric ozone in the non-urban environment: global pollution, subtropical stratospheric folds, disturbances associated with the 1997 El-Nino. A synthesis based on selected examples of field and satellite data will be presented.

Thompson, A.↗

New Insights into Tropospheric Ozone from Satellites and Soundings

The past 4-5 years have seen an unprecedented increase in tropospheric ozone (and related) observations thanks to focused campaigns, new satellite data, and to innovative approaches to climatology (commercial aircraft sampling, networks). The result has been new insights into the complex chemistry and dynamics affecting tropospheric ozone in the non-urban environment: global pollution, subtropical stratospheric folds, disturbances associated with the 1997 El-Nino. A synthesis based on selected examples of field and satellite data, much of it from Goddard's tropospheric work, will be presented.

Thompson, A.↗

Detectability of Varied Hybridization Scenarios Using Genome-Scale Hybrid Detection Methods

Hybridization events complicate the accurate reconstruction of phylogenies, as they lead to patterns of genetic heritability that are unexpected under traditional, bifurcating models of species trees. This phenomenon has led to the development of methods to infer these varied hybridization events, both methods that reconstruct networks directly, as well as summary methods that predict individual hybridization events from a subset of taxa. However, a lack of empirical comparisons between methods – especially those pertaining to large networks with varied hybridization scenarios – hinders their practical use. Here, we provide a comprehensive review of popular summary methods: TICR, MSCquartets, HyDe, Patterson’s D-Statistic (ABBA-BABA), D3, and Dp. TICR and MSCquartets are based on quartet concordance factors gathered from gene tree topologies and HyDe, Patterson’s D-Statistic, D3, and Dp use site pattern frequencies to identify hybridization events between sets of three taxa. We then use simulated data to address questions of method accuracy and ideal use scenarios by testing methods against complex networks which depict gene flow events that differ in depth (timing), quantity (single vs. multiple, overlapping hybridizations), and rate of gene flow (γ). We find that deeper or multiple hybridization events may introduce noise and weaken the signal of hybridization, leading to higher relative false negative rates across all methods. Despite some forms of hybridization eluding quartet-based detection methods, MSCquartets displays high precision in most scenarios. While HyDe results in high false negative rates when tested on hybridizations involving extinct or unsampled ghost lineages, HyDe is the only method able to identify the direction of hybridization, distinguishing the source parental lineages from recipient hybrid lineages. Lastly, we test the methods on a dataset of ultraconserved elements from the bee subfamily Nomiinae, finding possible hybridization events between clades which correspond to regions of poor support in the species tree estimated in a previous study.

Bjorner, Marianne B.↗

Cognitive network organization and cockpit automation

Attention is given to a technique for the derivation of pilot cognitive networks from empirical data, which has been successfully used to guide the redesign of the Control Display Unit that serves as the primary interface of the complex flight management system being developed by NASA's Advanced Concepts Flight Simulator program. The 'pathfinder' algorithm of Schvaneveldt et al. (1985) is used to obtain the conceptual organization of four pilots by generating a family of link-weighted networks from a set of psychological distance data derived through similarity ratings. The degree of conceptual agreement between pilots is assessed, and the means of translating a cognitive network into a menu structure are noted.

Roske-Hofstrand, R. J.↗

Deep Learning Analysis of Polaritonic Wave Images

Deep learning (DL) is an emerging analysis tool across the sciences and engineering. Encouraged by the successes of DL in revealing quantitative trends in massive imaging data, we applied this approach to nanoscale deeply subdiffractional images of propagating polaritonic waves in complex materials. Utilizing the convolutional neural network (CNN), we developed a practical protocol for the rapid regression of images that quantifies the wavelength and the quality factor of polaritonic waves. Using simulated near-field images as training data, the CNN can be made to simultaneously extract polaritonic characteristics and material parameters in a time scale that is at least 3 orders of magnitude faster than common fitting/processing procedures. The CNN-based analysis was validated by examining the experimental near-field images of charge-transfer plasmon polaritons at graphene/α-RuCl3 interfaces. Our work provides a general framework for extracting quantitative information from images generated with a variety of scanning probe methods.

97 MATHEMATICS AND COMPUTING↗

A Modular and Transferable Reinforcement Learning Framework for the Fleet Rebalancing Problem

Mobility on demand (MoD) systems show great promise in realizing flexible and efficient urban transportation. However, significant technical challenges arise from operational decision making associated with MoD vehicle dispatch and fleet rebalancing. For this reason, operators tend to employ simplified algorithms that have been demonstrated to work well in a particular setting. To help bridge the gap between novel and existing methods, we propose a modular framework for fleet rebalancing based on model-free reinforcement learning (RL) that can leverage an existing dispatch method to minimize system cost. In particular, by treating dispatch as part of the environment dynamics, a centralized agent can learn to intermittently direct the dispatcher to reposition free vehicles and mitigate against fleet imbalance. We formulate RL state and action spaces as distributions over a grid partitioning of the operating area, making the framework scalable and avoiding the complexities associated with multiagent RL. Numerical experiments, using real-world trip and network data, demonstrate that RL reduces waiting time by 28% to 38% for the same-day evaluation, 17% to 44% for cross-day evaluation, and 22% to 25% for cross-season evaluation compared with no rebalancing scenarios. This approach has several distinct advantages over baseline methods including: improved system cost; high degree of adaptability to the selected dispatch method; and the ability to perform scale-invariant transfer learning between problem instances with similar vehicle and request distributions.

33 ADVANCED PROPULSION SYSTEMS↗

Modeling and Calibration of Supplier Selection Problem in Freight Agent-Based Simulations

Freight transportation modeling often struggles with data limitations, especially in accurately representing complex supplier selection processes and their impact on network flows. This research addresses this critical gap by developing a large-scale, calibrated agent-based model for supplier selection, complemented by a probabilistic heuristic for international shipments. Our approach integrates trade relationships between industry sectors, transportation costs, and a supplier-rating model adapted from existing literature. The model’s core objective is to minimize the discrepancy between modeled and observed commodity flows while ensuring a close match to regional shipping distance distributions. Implemented and tested across four major U.S. metropolitan areas—Atlanta, Chicago, Dallas–Fort Worth, and Los Angeles—the model demonstrates high fidelity in replicating observed freight patterns. Key findings reveal consistent alignment with national shipping distance trends and highlight significant spatial variations in commodity trade assignments and demand across the study regions. This behaviorally informed and transport-sensitive framework is designed to approximate real-world decision making, providing a robust tool for policymakers and planners to evaluate targeted interventions, assess infrastructure investments, and enhance supply chain resilience in the face of disruptions.

Ismael, Abdelrahman (ORCID:0000000303712110)↗

Portability and the National Energy Software Center

The software portability problem is examined from the viewpoint of experience gained in the operation of a software exchange and information center. First, the factors contributing to the program interchange to date are identified, then major problem areas remaining are noted. The import of the development of programming language and documentation standards is noted, and the program packaging procedures and dissemination practices employed by the Center to facilitate successful software transport are described. Organization, or installation, dependencies of the computing environment, often hidden from the program author, and data interchange complexities are seen as today's primary issues with dedicated processors and network communications offering an alternative solution.

Butler, M. K.↗

Quantifying rupture characteristics of microearthquakes in the Parkfield Area using a high-resolution borehole network

It is well known that large earthquakes often exhibit significant rupture complexity such as well separated subevents. With improved recording and data processing techniques, small earthquakes have been found to exhibit rupture complexity as well. Studying these small earthquakes offers the opportunity to better understand the possible causes of rupture complexities. Specifically, if they are random or are related to fault properties. We examine microearthquakes (M < 3) in the Parkfield, California, area that are recorded by a high-resolution borehole network. We quantify earthquake complexity by the deviation of source time functions and source spectra from simple circular (omega-square) source models. We establish thresholds to declare complexity, and find that it can be detected in earthquakes larger than magnitude 2, with the best resolution above M2.5. Comparison between the two approaches reveals good agreement (>90 per cent), implying both methods are characterizing the same source complexity. For the two methods, 60–80 per cent (M 2.6–3) of the resolved events are complex depending on the method. The complex events we observe tend to cluster in areas of previously identified structural complexity; a larger fraction of the earthquakes exhibit complexity in the days following the M w 6 2004 Parkfield earthquake. Ignoring the complexity of these small events can introduce artefacts or add uncertainty to stress drop measurements. Focusing only on simple events however could lead to systematic bias, scaling artefacts and the lack of measurements of stress in structurally complex regions.

58 GEOSCIENCES↗

Regularization via f -Divergence: An Application to Multi-Oxide Spectroscopic Analysis

In this paper, we explore the application of convolutional neural networks (CNNs) for predicting the chemical composition of complex geologic samples in a simulated Martian atmospheric environment. Specifically, we aim to characterize oxide weight percentages (wt.%) of rock samples analyzed by remote Laser-Induced Breakdown Spectroscopy (LIBS), framing the problem as a multi-target regression task . Neural networks trained on LIBS spectra are prone to overfitting due to high spectral complexity, limited labeled data, and measurement noise. While regularization is critical for improving generalization, common methods (e.g., ℓ 2 regularization) impose constraints not directly tied to data distribution properties. We propose a novel regularization method based on a specific ƒ-divergence induced by a graph-based estimator, designed to constrain the distributional discrepancy between predictions and targets. This regularizer serves a dual purpose: (a) mitigating overfitting by enforcing a constraint on the distributional difference between predictions and noisy targets, and (b) acting as an auxiliary loss that penalizes large divergences. To enable backpropagation, we develop a differentiable approximation of this particular ƒ-divergence, making the method feasible for neural networks. Experiments on ChemCam and SuperCam LIBS calibration spectra show that mathematical equation-divergence regularization outperforms or matches standard regularization methods (ℓ 1 , ℓ 2 , dropout) and the classical baseline, partial least squares (PLS). Combining ƒ-divergence regularization with standard regularization yields further performance gains, indicating that distributional regularization is useful in this context giving a promising direction for robust model training in planetary science applications. Source code is publicly available at Klein and Li (2025), https://doi.org/10.11578/dc.20250530.7.

58 GEOSCIENCES↗