Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network partition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Transparent bridge for monitoring crypto-partitioned wide-area network

This disclosure is directed to monitoring a crypto-partitioned, or cipher-text, wide-area network (WAN). A first computing device may be situated in a plain-text portion of a first enclave behind a first inline network encryptor (INE). A second device may be positioned in a plain-text portion of a second enclave behind a second INE. The two enclaves may be separated by a cipher-text WAN, over which the two enclaved may communicate. The first computing device may receive a data packet from the second computing device. The first computing device may then determine contents of a header of the data packet. The first computing device may, based at least in part on the contents of the header of the data packet, determine a status of the cipher-text WAN.

97 MATHEMATICS AND COMPUTING↗

Probabilistic resource allocation system with self-adaptive capability

A probabilistic resource allocation system is disclosed containing a low capacity computational module (Short Term Memory or STM) and a self-organizing associative network (Long Term Memory or LTM) where nodes represent elementary resources, terminal end nodes represent goals, and weighted links represent the order of resource association in different allocation episodes. Goals and their priorities are indicated by the user, and allocation decisions are made in the STM, while candidate associations of resources are supplied by the LTM based on the association strength (reliability). Weights are automatically assigned to the network links based on the frequency and relative success of exercising those links in the previous allocation decisions. Accumulation of allocation history in the form of an associative network in the LTM reduces computational demands on subsequent allocations. For this purpose, the network automatically partitions itself into strongly associated high reliability packets, allowing fast approximate computation and display of allocation solutions satisfying the overall reliability and other user-imposed constraints. System performance improves in time due to modification of network parameters and partitioning criteria based on the performance feedback.

Yufik, Yan M.↗

Probabilistic resource allocation system with self-adaptive capability

A probabilistic resource allocation system is disclosed containing a low capacity computational module (Short Term Memory or STM) and a self-organizing associative network (Long Term Memory or LTM) where nodes represent elementary resources, terminal end nodes represent goals, and directed links represent the order of resource association in different allocation episodes. Goals and their priorities are indicated by the user, and allocation decisions are made in the STM, while candidate associations of resources are supplied by the LTM based on the association strength (reliability). Reliability values are automatically assigned to the network links based on the frequency and relative success of exercising those links in the previous allocation decisions. Accumulation of allocation history in the form of an associative network in the LTM reduces computational demands on subsequent allocations. For this purpose, the network automatically partitions itself into strongly associated high reliability packets, allowing fast approximate computation and display of allocation solutions satisfying the overall reliability and other user-imposed constraints. System performance improves in time due to modification of network parameters and partitioning criteria based on the performance feedback.

Yufik, Yan M.↗

Experiments to Determine Whether Recursive Partitioning (CART) or an Artificial Neural Network Overcomes Theoretical Limitations of Cox Proportional Hazards Regression

New computationally intensive tools for medical survival analyses include recursive partitioning (also called CART) and artificial neural networks. A challenge that remains is to better understand the behavior of these techniques in effort to know when they will be effective tools. Theoretically they may overcome limitations of the traditional multivariable survival technique, the Cox proportional hazards regression model. Experiments were designed to test whether the new tools would, in practice, overcome these limitations. Two datasets in which theory suggests CART and the neural network should outperform the Cox model were selected. The first was a published leukemia dataset manipulated to have a strong interaction that CART should detect. The second was a published cirrhosis dataset with pronounced nonlinear effects that a neural network should fit. Repeated sampling of 50 training and testing subsets was applied to each technique. The concordance index C was calculated as a measure of predictive accuracy by each technique on the testing dataset. In the interaction dataset, CART outperformed Cox (P less than 0.05) with a C improvement of 0.1 (95% Cl, 0.08 to 0.12). In the nonlinear dataset, the neural network outperformed the Cox model (P less than 0.05), but by a very slight amount (0.015). As predicted by theory, CART and the neural network were able to overcome limitations of the Cox model. Experiments like these are important to increase our understanding of when one of these new techniques will outperform the standard Cox model. Further research is necessary to predict which technique will do best a priori and to assess the magnitude of superiority.

Kattan, Michael W.↗

High-performance parallel analysis of coupled problems for aircraft propulsion

This research program deals with the application of high-performance computing methods for the analysis of complete jet engines. We have entitled this program by applying the two dimensional parallel aeroelastic codes to the interior gas flow problem of a bypass jet engine. The fluid mesh generation, domain decomposition, and solution capabilities were successfully tested. We then focused attention on methodology for the partitioned analysis of the interaction of the gas flow with a flexible structure and with the fluid mesh motion that results from these structural displacements. This is treated by a new arbitrary Lagrangian-Eulerian (ALE) technique that models the fluid mesh motion as that of a fictitious mass-spring network. New partitioned analysis procedures to treat this coupled three-component problem are developed. These procedures involved delayed corrections and subcycling. Preliminary results on the stability, accuracy, and MPP computational efficiency are reported.

Felippa, C. A.↗

Graph Partitioning and Sparse Matrix Ordering using Reinforcement Learning and Graph Neural Networks

We present a novel method for graph partitioning, based on reinforcement learning and graph convolutional neural networks. Our approach is to recursively partition coarser representations of a given graph. The neural network is implemented using SAGE graph convolution layers, and trained using an advantage actor critic (A2C) agent. We present two variants, one for finding an edge separator that minimizes the normalized cut or quotient cut, and one that finds a small vertex separator. The vertex separators are then used to construct a nested dissection ordering to permute a sparse matrix so that its triangular factorization will incur less fill-in. The partitioning quality is compared with partitions obtained using METIS and SCOTCH, and the nested dissection ordering is evaluated in the sparse solver SuperLU. Our results show that the proposed method achieves similar partitioning quality as METIS and SCOTCH. Furthermore, the method generalizes across different classes of graphs, and works well on a variety of graphs from the SuiteSparse sparse matrix collection.

97 MATHEMATICS AND COMPUTING↗

Distributionally Robust Decentralized Volt-Var Control With Network Reconfiguration

Here, this paper presents a decentralized volt-var optimization (VVO) and network reconfiguration strategy to address the challenges arising from the growing integration of distributed energy resources, particularly photovoltaic (PV) generation units, in active distribution networks. To reconcile control measures with different time resolutions and empower local control centers to handle intermittency locally, the proposed approach leverages a two-stage distributionally robust optimization; decisions on slow-responding control measures and set points that link neighboring subnetworks are made in advance while considering all plausible distributions of uncertain PV outputs. We present a decomposition algorithm with an acceleration scheme for solving the proposed model. Numerical experiments on the IEEE 123 bus distribution system are given to demonstrate its outstanding out-of-sample performance and computational efficiency, which suggests that the proposed method can effectively localize uncertainty via risk-informed proactive timely decisions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

PhILMs: Collaboratory on Mathematics and Physics-Informed Learning Machines for Multiscale and Multiphysics Problems

The landscape of computational science and engineering is continually evolving, with the challenge of high-dimensional regression problems standing as a significant hurdle in numerous scientific endeavors. Addressing this challenge, our research, funded by this award, has led to the development of an innovative computational framework known as Probabilistic Partition of Unity Networks (PPOU-Nets). This initiative represents a collaborative effort to harness the potential of mathematics and physics-informed machine learning in tackling multiscale and multiphysics problems prevalent in high-dimensional spaces. Through this work, we have proposed a novel methodology that seamlessly integrates adaptive dimensionality reduction and a mixture of experts model, thereby facilitating a more efficient and accurate approximation of complex functions. This research effort has not only advanced the state of computational science but also opened new avenues for exploration in quantum computing and beyond. This report outlines the motivation, methodology, key findings, and implications of our work, underscoring our contributions to the broader scientific community and the potential pathways for future research.

97 MATHEMATICS AND COMPUTING↗

PCAfold 2.0—Novel tools and algorithms for low-dimensional manifold assessment and optimization

We describe an update to our open-source Python package, PCAfold, designed to help researchers generate, analyze and improve low-dimensional data manifolds. In the current version, PCAfold 2.0, we introduce novel tools and algorithms for assessing and optimizing low-dimensional manifolds. This includes a method that generates a “map” of local feature sizes that can help pinpoint researchers to problematic regions on a manifold. We introduce a novel cost function that characterizes the quality of a manifold topology with a single number. We develop two algorithms for feature selection based on principal component analysis (PCA) that use the cost function as an objective function to minimize. We introduce a quantity of interest (QoI)-aware dimensionality reduction strategy where data projections are computed using an artificial neural network and are directly optimized towards representing various projection-independent and projection-dependent QoIs. We also introduce an implementation of partition of unity networks (POUnets) for efficient reconstruction of QoIs from low-dimensional manifolds based on combining neural network classification with localized polynomial regression. Our software can be broadly applicable in all domains of science and engineering that aim to reduce data dimensionality, as well as in the fundamental research on representation learning.

97 MATHEMATICS AND COMPUTING↗

Partitioning of ocean and land uptake of CO2 as inferred by delta C-13 measurements from the NOAA Climate Monitoring and Diagnostics Laboratory Global Air Sampling Network

Using delta C-13 measurements in atmospheric CO2 from a cooperative global air sampling network, we determined the partitioning of the net uptake of CO2 between ocean and land as a function of latitude and time. The majority of delta C-13 measurements were made at the Institute of Arctic and Alpine Research (INSTAAR) of the University of Colorado. We perform an inverse deconvolution of both CO2 and delta C-13 observations, using a two-dimensional model of atmospheric transport. Also, the discrimination against C-13 by plant photosynthesis, as a function of latitude and time, is calculated from global runs of the simple biosphere (SiB) model. Uncertainty due to the longitudinal structure of the data, which is not represented by the model, is studied through a bootstrap analysis by adding and omitting measurement sites. The resulting error estimates for our inferred sources and sinks are of the order of 1 GTC (1 GTC = 10(exp 15) gC). Such error bars do not reflect potential systematic errors arising from our estimates of the isotopic disequilibria between the atmosphere and the oceans and biosphere, which are estimated in a separate sensitivity analysis. With respect to global totals for 1992 we found that 3.2 GTC of carbon dissolved into the ocean and that 1.5 GTC were sequestered by land ecosystems. Northern hemisphere ocean gyres north of 15 deg N absorbed 2.7 GTC. The equatorial oceans between 10 deg S and 10 deg N were a net source to the atmosphere of 0.9 GTC. We obtained a sink of 1.6 GTC in southern ocean gyres south of 20 deg S, although the deconvolution is poorly constrained by sparse data coverage at high southern latitudes. The seasonal uptake of CO2 in the northern gyres appears to be correlated with a bloom of phytoplankton in surface waters. On land, northern temperate and boreal ecosystems between 35 deg N and 65 deg N were found to be a major sink of CO2 in 1992, as large as 3.5 GTC. Northern tropical ecosystems (equator-30 deg N) appear to be a net source to the source to the atmosphere of 2 GTC which could reflect biomass burning. A small sink, 0.3 GTC, was inferred for southern tropical ecosystems (30 deg S-equator).

Ciais, Philippe↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Wireless Avionics Packet to Support Fault Tolerance for Flight Applications

In this protocol and packet format, data traffic is monitored by all network interfaces to determine the health of transmitter and subsystems. When failures are detected, the network inter face applies its recover y policies to provide continued service despite the presence of faults. The protocol, packet format, and inter face are independent of the data link technology used. The current demonstration system supports both commercial off-the-shelf wireless connections and wired Ethernet connections. Other technologies such as 1553 or serial data links can be used for the network backbone. The Wireless Avionics packet is divided into three parts: a header, a data payload, and a checksum. The header has the following components: magic number, version, quality of service, time to live, sending transceiver, function code, payload length, source Application Data Interface (ADI) address, destination ADI address, sending node address, target node address, and a sequence number. The magic number is used to identify WAV packets, and allows the packet format to be updated in the future. The quality of service field allows routing decisions to be made based on this value and can be used to route critical management data over a dedicated channel. The time to live value is used to discard misrouted packets while the source transceiver is updated at each hop. This information is used to monitor the health of each transceiver in the network. To identify the packet type, the function code is used. Besides having a regular data packet, the system supports diagnostic packets for fault detection and isolation. The payload length specifies the number of data bytes in the payload, and this supports variable-length packets in the network. The source ADI is the address of the originating interface. This can be used by the destination application to identify the originating source of the packet where the address consists of a subnet, subsystem class within the subnet, a subsystem unit, and the local ADI number. The destination ADI is used to route the packet to its ultimate destination. At each hop, the sending interface uses the destination address to determine the next node for the data. The sending node is the node address of the interface that is broadcasting the packet. This field is used to determine the health of the subsystem that is sending the packet. In the case of a packet that traverses several intermediate nodes, it may be the node address of the intermediate node. The target node is the node address of the next hop for the packet. It may be an intermediate node, or the final destination for the packet. The sequence number is used to identify duplicate packets. Because each interface has multiple transceivers, the same packet will appear at both receivers. The sequence number allows the interface to correlate the reception and forward a single, unique packet for additional processing. The subnet field allows data traffic to be partitioned into segregated local networks to support large networks while keeping each subnet at a manageable size. This also keeps the routing table small enough so routing can be done by a simple table lookup in an FPGA device. The subsystem class identifies members of a set of redundant subsystems, and, in a hot standby configuration, all members of the subsystem class will receive the data packets. Only the active subsystem will generate data traffic. Specific units in a class of redundant units can be identified and, if the hot standby configuration is not used, packets will be directed to a specific subsystem unit.

Block, Gary L.↗

Global Aerodynamic Modeling Using Automated Local Model Networks in Real Time

A novel method is presented for automated real-time global aerodynamic modeling using local model networks, known as Smoothed Partitioning with Localized Trees in Real Time (SPLITR), as part of NASA’s Learn-to-Fly technology development initiative. The global nonlinear aerodynamics are partitioned into several local regions known as cells, with the dimension, location, and timing of each partition automatically selected based on a residual characterization procedure, under the constraints of real-time operation. Regression trees represent the successive partitioning of the global flight envelope and describe the evolution of the cell structure. Recursive equation-error least-squares parameter estimation in the time domain is used to estimate a model that represents the local aerodynamics in each region, so that it can be updated independently with non-contiguous data in the range of each cell over time. A weighted superposition of these piecewise local models across the flight envelope forms a global nonlinear model that also accurately captures the local aerodynamics. The SPLITR approach is demonstrated using both simulation and flight data, and the results are analyzed in terms of model predictive capabilities as well as interpretability. The results show that SPLITR can be used to automatically partition complex nonlinear aerodynamic behavior, produce an accurate model, and provide valuable physical insight into the local and global aerodynamics.

Weinstein, Rose↗

Machine Learning Applications to Metal-Silicate Equilibria and their Insights into Core Formation

An extensive number of studies have experimentally investigated how elements distribute between metal and silicate phases, to better constrain core-mantle chemical equilibrium. Here, we present a new database compiling all (to our knowledge) experimental data on liquid metal-silicate partitioning from 118 peer-reviewed publications. We applied various machine learning techniques to gain further insights into these partitioning equilibria and their dependencies. We performed a network analysis to investigate the relationship between experiments and partition coefficients, which enables visualizing gaps in the experimental dataset and biases related to varying experimental conditions and analytical setup. In addition, semi-empirical thermodynamic models are commonly used to extrapolate these chemical reactions to the wide range of pressure, temperature and compositional conditions of planetary differentiation. These models are based on linear regressions that assume continuous relationship between partition coefficients and experimental variables. Here, we considered random forest regressions, which are algorithms based on ensembles of decision trees and does not consider continuous effects of each variable. The application of this regression significantly improves the prediction of metal-silicate partitioning for several elements including Ni, Si and Cr. We will show how this new approach improves our understanding of elemental exchange between metal and silicate and their implications for the Earth’s core formation.

siderophile element↗

Algorithms for parallel flow solvers on message passing architectures

The purpose of this project has been to identify and test suitable technologies for implementation of fluid flow solvers -- possibly coupled with structures and heat equation solvers -- on MIMD parallel computers. In the course of this investigation much attention has been paid to efficient domain decomposition strategies for ADI-type algorithms. Multi-partitioning derives its efficiency from the assignment of several blocks of grid points to each processor in the parallel computer. A coarse-grain parallelism is obtained, and a near-perfect load balance results. In uni-partitioning every processor receives responsibility for exactly one block of grid points instead of several. This necessitates fine-grain pipelined program execution in order to obtain a reasonable load balance. Although fine-grain parallelism is less desirable on many systems, especially high-latency networks of workstations, uni-partition methods are still in wide use in production codes for flow problems. Consequently, it remains important to achieve good efficiency with this technique that has essentially been superseded by multi-partitioning for parallel ADI-type algorithms. Another reason for the concentration on improving the performance of pipeline methods is their applicability in other types of flow solver kernels with stronger implied data dependence. Analytical expressions can be derived for the size of the dynamic load imbalance incurred in traditional pipelines. From these it can be determined what is the optimal first-processor retardation that leads to the shortest total completion time for the pipeline process. Theoretical predictions of pipeline performance with and without optimization match experimental observations on the iPSC/860 very well. Analysis of pipeline performance also highlights the effect of uncareful grid partitioning in flow solvers that employ pipeline algorithms. If grid blocks at boundaries are not at least as large in the wall-normal direction as those immediately adjacent to them, then the first processor in the pipeline will receive a computational load that is less than that of subsequent processors, magnifying the pipeline slowdown effect. Extra compensation is needed for grid boundary effects, even if all grid blocks are equally sized.

Vanderwijngaart, Rob F.↗

Method for Constructing Composite Response Surfaces by Combining Neural Networks with Polynominal Interpolation or Estimation Techniques

A method and system for data modeling that incorporates the advantages of both traditional response surface methodology (RSM) and neural networks is disclosed. The invention partitions the parameters into a first set of s simple parameters, where observable data are expressible as low order polynomials, and c complex parameters that reflect more complicated variation of the observed data. Variation of the data with the simple parameters is modeled using polynomials; and variation of the data with the complex parameters at each vertex is analyzed using a neural network. Variations with the simple parameters and with the complex parameters are expressed using a first sequence of shape functions and a second sequence of neural network functions. The first and second sequences are multiplicatively combined to form a composite response surface, dependent upon the parameter values, that can be used to identify an accurate mode

Rai, Man Mohan↗