Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network partition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Use of networked workstations for parallel nonlinear structural dynamic simulations of rotating bladed-disk assemblies

The principal objective of this research is to investigate, develop and demonstrate coarse-grained, parallel-processing strategies for nonlinear dynamic simulations for rotating bladed-disk assemblies. The parallel -processing strategies addressed include numerical algorithms for parallel nonlinear solutions and techniques to effect load balancing among processors. The parallel environment employed is a distributed-memory, coarse-grained one consisting of networked workstations. A parallel explicit time integration method has been implemented for transient nonlinear solutions of rotationg bladed-disk assemblies. Automatic domain partitioning techniques have been investigated for load balancing among processors. Advanced computing environments, data structures and interactive computer graphics all contribute to an integrated parallel finite element analysis system to facilitate more efficient and powerful dynamic simulations.

Hsieh, Shang-Hsien↗

Software defined grid energy storage

Today, consumer battery installations are isolated, physical devices. Virtual power plants (VPPs) allow consumer devices to aggregate for grid services, but they are are vertically integrated, vendor controlled systems (e.g., Tesla’s VPP). Consumer batteries are therefore unable to participate in energy markets or other grid services outside what their vendor provides. We describe a software system that provides software control of multiple, networked battery energy storage systems in the electric grid. The system introduces two new ideas that enable flexible and dependable management of energy storage. The first is a virtual battery, which can either partition a battery or aggregate multiple batteries. The second is a reservation-based API which allows asynchronous control of batteries to meet contractual guarantees in a safe and dependable manner. Virtual batteries and a reservation-based API address the unique challenges of achieving high and efficient utilization of energy storage systems, including heterogeneity of battery systems such as varying C-rates, participation in energy markets, utility bill management systems, community resource sharing, and reliability. Using a testbed comprised of sonnen Inc. storage units installed in several homes and a lab, we demonstrate that virtualized batteries can seamlessly replace physical batteries, flexibly manage energy storage resources, isolate multiple clients using a shared battery, and create new energy storage applications.

25 ENERGY STORAGE↗

Phenolic Polymer Interactions with Water and Ethylene Glycol Solvents

Interactions between pre-cured phenolic polymer chains and a solvent have a significant impact on the structure and properties of the final post-cured phenolic resin. Developing an understanding of the nature of these interactions is important and will aid in the selection of the proper solvent that will lead to the desired final product. Here, we investigate the role of the phenolic chain structure and the solvent type on the overall solvation performance of the system through ab initio techniques and molecular dynamics simulations. Two types of solvents are considered: ethylene glycol (EGL) and H2O. Three phenolic chain structures are considered, including two novolac-type chains with either an ortho-ortho (OON) or an ortho-para (OPN) backbone network and a resole-type (RES) chain with an ortho-ortho network. Each system is characterized through a structural analysis of the solvation shell and the hydrogen-bonding environment as well as through a quantification of the solvation free energy along with partitioned interaction energies between specific molecular species. The combination of simulations and the analyses indicate that EGL provides a higher solvation free energy than H2O due to more energetically favorable hydrophilic interactions as well as favorable hydrophobic interactions between CH element groups. In addition, the phenolic chain structure significantly affects the solvation performance, with OON having limited intermolecular hydrogen-bond formations, while OPN and RES interact more favorably with the solvent molecules. The results suggest that a resole-type phenolic chain with an ortho-para network should have the best solvation performance in EGL, H2O, and other similar solvents.

Haskins, Justin B.↗

Traffic study of a computer system.

A study which may guide the operations of existing computer installations, as well as the design of future networks, is described. Performance data and evaluations are considered with reference to interarrival time, users' habits, waiting time for execution, time spent in a partition, figures of merit, and states of the system. The analysis of the variables proceeds from examination of typical data with appropriate statistical tests to conclusions about the possible state of nature.

Cramer, R. L.↗

A differentiable approach to the maximum independent set problem using dataless neural networks

The success of machine learning solutions for reasoning about discrete structures has brought attention to its adoption within combinatorial optimization algorithms. Such approaches generally rely on supervised learning by leveraging datasets of the combinatorial structures of interest drawn from some distribution of problem instances. Reinforcement learning has also been employed to find such structures. Here, in this paper, we propose a different approach in that no data is required for training the neural networks that produce the solution. In this sense, what we present is not a machine learning solution, but rather one that is dependent on neural networks and where backpropagation is applied to a loss function defined by the structure of the neural network architecture as opposed to a training dataset. In particular, we reduce the popular combinatorial optimization problem of finding a maximum independent set to a neural network and employ a dataless training scheme to refine the parameters of the network such that those parameters yield the structure of interest. Additionally, we propose a universal graph reduction procedure to handle large-scale graphs. The reduction exploits community detection for graph partitioning and is applicable to any graph type and/or density. Experimental results on both real and synthetic graphs demonstrate that our proposed method performs on par or outperforms state-of-the-art learning-based methods in terms of the size of the found set without requiring any training data.

97 MATHEMATICS AND COMPUTING↗

A Modular and Transferable Reinforcement Learning Framework for the Fleet Rebalancing Problem

Mobility on demand (MoD) systems show great promise in realizing flexible and efficient urban transportation. However, significant technical challenges arise from operational decision making associated with MoD vehicle dispatch and fleet rebalancing. For this reason, operators tend to employ simplified algorithms that have been demonstrated to work well in a particular setting. To help bridge the gap between novel and existing methods, we propose a modular framework for fleet rebalancing based on model-free reinforcement learning (RL) that can leverage an existing dispatch method to minimize system cost. In particular, by treating dispatch as part of the environment dynamics, a centralized agent can learn to intermittently direct the dispatcher to reposition free vehicles and mitigate against fleet imbalance. We formulate RL state and action spaces as distributions over a grid partitioning of the operating area, making the framework scalable and avoiding the complexities associated with multiagent RL. Numerical experiments, using real-world trip and network data, demonstrate that RL reduces waiting time by 28% to 38% for the same-day evaluation, 17% to 44% for cross-day evaluation, and 22% to 25% for cross-season evaluation compared with no rebalancing scenarios. This approach has several distinct advantages over baseline methods including: improved system cost; high degree of adaptability to the selected dispatch method; and the ability to perform scale-invariant transfer learning between problem instances with similar vehicle and request distributions.

33 ADVANCED PROPULSION SYSTEMS↗

Fracture Network Prediction Using Physics-based Machine Learning Algorithms

In recent years, systematic CO2 injection into geological reservoirs across the U.S. has gained traction as a strategy to mitigate greenhouse gas emissions. This approach necessitates precise monitoring to ensure secure containment, minimize risks, and optimize storage management. Our study leverages machine learning (ML) techniques to advance the understanding of CO2 injection processes, focusing on the Illinois Basin. Over a three-year injection period, we analyzed microseismic data, identifying 19 temporal intervals with significant bottom-hole pressure changes. By partitioning microseismic events into these intervals and estimating b-values, we revealed over 100 clusters of events related to fracture initiation or reactivation. Advanced spatial analysis highlighted horizontally-oriented fractures along the NNW-SSE axis. This quantification of fracture networks informs dynamic injection scheduling, work-over strategies, and risk assessments, enhancing carbon capture, utilization, and storage (CCUS) operations. Additionally, our methodology offers valuable insights for oil and gas operations and geothermal development, supporting fracture-based monitoring and risk mitigation.

Kumar, Abhash↗

Cation–Ligand Interactions Dictate Salt Partitioning and Diffusivity in Ligand-Functionalized Polymer Membranes

Membranes are an attractive alternative to current thermal separations due to their scalability and energy efficiency in desalinating water. Unfortunately, many of the conventional membrane materials available today are unable to differentiate between ionic solutes, especially alkali cations, compromising their use in ion–ion separations. Inspired by the ion-specific interactions exhibited by biological ion channels, recent research efforts have focused on synthesizing and characterizing new polymeric materials that incorporate ligands into polymer networks to bias solubility and/or diffusivity of one cationic species over another. Despite these efforts, little is known about the influence of incorporating ligands into polymer membranes on solubility and diffusivity of the complexing species. In this study, we first build a qualitative model of salt partitioning, diffusivity, and permeability in generic cation-complexing ligand-functionalized polymer membranes. Next, to validate our model and hypotheses, we perform atomistic molecular dynamics simulations of a 12-crown-4-functionalized membrane in the presence of alkali halide salts at low concentration. Generally, cation complexation enhances cation solubility but decreases diffusivity. Interestingly, the reduction in diffusivity is predicted to be larger than the enhancement in solubility for materials which operate by the mechanisms proposed in our physical picture, ultimately resulting in a reduction in the permeability of the selectively complexing ion.

36 MATERIALS SCIENCE↗

Deep Space Networking Experiments on the EPOXI Spacecraft

NASA's Space Communications & Navigation Program within the Space Operations Directorate is operating a program to develop and deploy Disruption Tolerant Networking [DTN] technology for a wide variety of mission types by the end of 2011. DTN is an enabling element of the Interplanetary Internet where terrestrial networking protocols are generally unsuitable because they rely on timely and continuous end-to-end delivery of data and acknowledgments. In fall of 2008 and 2009 and 2011 the Jet Propulsion Laboratory installed and tested essential elements of DTN technology on the Deep Impact spacecraft. These experiments, called Deep Impact Network Experiment (DINET 1) were performed in close cooperation with the EPOXI project which has responsibility for the spacecraft. The DINET 1 software was installed on the backup software partition on the backup flight computer for DINET 1. For DINET 1, the spacecraft was at a distance of about 15 million miles (24 million kilometers) from Earth. During DINET 1 300 images were transmitted from the JPL nodes to the spacecraft. Then, they were automatically forwarded from the spacecraft back to the JPL nodes, exercising DTN's bundle origination, transmission, acquisition, dynamic route computation, congestion control, prioritization, custody transfer, and automatic retransmission procedures, both on the spacecraft and on the ground, over a period of 27 days. The first DINET 1 experiment successfully validated many of the essential elements of the DTN protocols. DINET 2 demonstrated: 1) additional DTN functionality, 2) automated certain tasks which were manually implemented in DINET 1 and 3) installed the ION SW on nodes outside of JPL. DINET 3 plans to: 1) upgrade the LTP convergence-layer adapter to conform to the international LTP CL specification, 2) add convergence-layer "stewardship" procedures and 3) add the BSP security elements [PIB & PCB]. This paper describes the planning and execution of the flight experiment and the validation results.

automated data communication↗

Daily evapotranspiration changes during heatwaves at 32 NEON sites, 2019-2021

This dataset provides partitioned evapotranspiration (ET, the combined loss of water from soil and plant surfaces) anomalies during heatwave events—soil evaporation (E) and transpiration (T)—for 268 heatwave events across 32 National Ecological Observatory Network (NEON) flux sites in the contiguous United States from 2019–2021. Using an ensemble of four high-frequency turbulence methods (Flux-variance Similarity, Conditional Eddy Covariance [CEC], CEC with Water-Use Efficiency, and Conditional Eddy Accumulation; see Zahn and Bou-Zeid 2024), half-hourly transpiration-to-evapotranspiration (T/ET) ratios were derived from 20 hertz (Hz, cycles per second) eddy covariance measurements of carbon dioxide (CO₂) and water vapor (H₂O) concentrations. The dataset spans six vegetation types including evergreen and deciduous forests, grasslands, cultivated crops, shrublands, and emergent herbaceous wetlands. Data Package Contents: The dataset includes a single CSV (comma-separated values) file containing daily anomalies (deviations from baseline conditions) for transpiration (Delta_T), evaporation (Delta_E), total evapotranspiration (Delta_ET), and T/ET ratio (Delta_T_ET) during each day of identified heatwave events. The file also includes site codes, dates, heatwave event identifiers, and day-of-heatwave indicators. The CSV file can be opened with spreadsheet software (Microsoft Excel, Google Sheets) or programming environments (Python, R, MATLAB). This resource enables researchers to investigate ecosystem-specific responses to thermal extremes, validate land surface model partitioning of ET fluxes, and examine feedbacks between water cycling and surface energy balance during heatwaves. The dataset is particularly valuable for studies linking vegetation hydraulic strategies to climate resilience, as it captures the divergent responses of shallow-rooted versus deep-rooted ecosystems. Potential applications include improving drought early warning systems, informing irrigation management strategies, and advancing our mechanistic understanding of land-atmosphere interactions under extreme heat conditions.

Day of Heatwave↗

A parallel algorithm for multi-level logic synthesis using the transduction method

The Transduction Method has been shown to be a powerful tool in the optimization of multilevel networks. Many tools such as the SYLON synthesis system (X90), (CM89), (LM90) have been developed based on this method. A parallel implementation is presented of SYLON-XTRANS (XM89) on an eight processor Encore Multimax shared memory multiprocessor. It minimizes multilevel networks consisting of simple gates through parallel pruning, gate substitution, gate merging, generalized gate substitution, and gate input reduction. This implementation, called Parallel TRANSduction (PTRANS), also uses partitioning to break large circuits up and performs inter- and intra-partition dynamic load balancing. With this, good speedups and high processor efficiencies are achievable without sacrificing the resulting circuit quality.

Lim, Chieng-Fai↗

Latency Hiding in Dynamic Partitioning and Load Balancing of Grid Computing Applications

The Information Power Grid (IPG) concept developed by NASA is aimed to provide a metacomputing platform for large-scale distributed computations, by hiding the intricacies of highly heterogeneous environment and yet maintaining adequate security. In this paper, we propose a latency-tolerant partitioning scheme that dynamically balances processor workloads on the.IPG, and minimizes data movement and runtime communication. By simulating an unsteady adaptive mesh application on a wide area network, we study the performance of our load balancer under the Globus environment. The number of IPG nodes, the number of processors per node, and the interconnected speeds are parameterized to derive conditions under which the IPG would be suitable for parallel distributed processing of such applications. Experimental results demonstrate that effective solution are achieved when the IPG nodes are connected by a high-speed asynchronous interconnection network.

Das, Sajal K.↗

Continuous Lidar Monitoring of Polar Stratospheric Clouds at the South Pole

Polar stratospheric clouds (PSC) play a primary role in the formation of annual ozone holes over Antarctica during the austral sunrise. Meridional temperature gradients in the lower stratosphere and upper troposphere, caused by strong radiative cooling, induce a broad dynamic vortex centered near the South Pole that decouples and insulates the winter polar airmass. PSC nucleate and grow as vortex temperatures gradually fall below equilibrium saturation and frost points for ambient sulfate, nitrate, and water vapor concentrations (generally below 197 K). Cloud surfaces promote heterogeneous reactions that convert stable chlorine and bromine-based molecules into photochemically active ones. As spring nears, and the sun reappears and rises, photolysis decomposes these partitioned compounds into individual halogen atoms that react with and catalytically destroy thousands of ozone molecules before they are stochastically neutralized. Despite a generic understanding of the ozone hole paradigm, many key components of the system, such as cloud occurrence, phase, and composition; particle growth mechanisms; and denitrification of the lower stratosphere have yet to be fully resolved. Satellite-based observations have dramatically improved the ability to detect PSC and quantify seasonal polar chemical partitioning. However, coverage directly over the Antarctic plateau is limited by polar-orbiting tracks that rarely exceed 80 degrees S. In December 1999, a NASA Micropulse Lidar Network instrument (MPLNET) was first deployed to the NOAA Earth Systems Research Laboratory (ESRL) Atmospheric Research Observatory at the Amundsen-Scott South Pole Station for continuous cloud and aerosol profiling. MPLNET instruments are eye-safe, capable of full-time autonomous operation, and suitably rugged and compact to withstand long-term remote deployment. With only brief interruptions during the winters of 2001 and 2002, a nearly continuous data archive exists to the present.

OZONE DESTRUCTION↗

An Adaptive Flow Solver for Air-Borne Vehicles Undergoing Time-Dependent Motions/Deformations

This report describes a concurrent Euler flow solver for flows around complex 3-D bodies. The solver is based on a cell-centered finite volume methodology on 3-D unstructured tetrahedral grids. In this algorithm, spatial discretization for the inviscid convective term is accomplished using an upwind scheme. A localized reconstruction is done for flow variables which is second order accurate. Evolution in time is accomplished using an explicit three-stage Runge-Kutta method which has second order temporal accuracy. This is adapted for concurrent execution using another proven methodology based on concurrent graph abstraction. This solver operates on heterogeneous network architectures. These architectures may include a broad variety of UNIX workstations and PCs running Windows NT, symmetric multiprocessors and distributed-memory multi-computers. The unstructured grid is generated using commercial grid generation tools. The grid is automatically partitioned using a concurrent algorithm based on heat diffusion. This results in memory requirements that are inversely proportional to the number of processors. The solver uses automatic granularity control and resource management techniques both to balance load and communication requirements, and deal with differing memory constraints. These ideas are again based on heat diffusion. Results are subsequently combined for visualization and analysis using commercial CFD tools. Flow simulation results are demonstrated for a constant section wing at subsonic, transonic, and a supersonic case. These results are compared with experimental data and numerical results of other researchers. Performance results are under way for a variety of network topologies.

Singh, Jatinder↗

Graph neural networks for mechanical property prediction of 2D fiber composites

This work investigates the ability of graph neural networks (GNNs) to homogenize 2D fiber composite microstructures. We use different inhomogeneity and anisotropy indices to motivate and show that the Volume Elements (VEs) used in ML methods should ideally be far from their Representative Volume Element (RVE) size limit and, consequently, are notably anisotropic. Hence, training only the isotropic limit properties may not be acceptable. Another aspect is the need to normalize elastic stiffness values for ML, especially when high elastic contrast ratios are encountered between composite phases or in the material set. We introduce a normalization technique based on the mean-field method (MFM) to handle such high contrast ratios and train for the entire stiffness tensor. We show that the proposed GNN approaches exhibit high accuracy and efficiency compared to traditional methods and convolutional neural networks, utilizing unstructured graphs constructed from microstructure topology. Our model successfully predicts the stiffness tensor, peak strength under bulk damage, and brittle fracture initiation strength across diverse microstructure configurations while maintaining high accuracy even for extreme material contrasts and volume fractions. We also present a method to improve prediction accuracy for small dataset sizes using Voronoi partitioning.

Brittle strength↗

Hybrid electric buses fuel consumption prediction based on real-world driving data

Estimating fuel consumption by hybrid diesel buses is challenging due to its diversified operations and driving cycles. Here, long-term transit bus monitoring data were utilized to empirically compare fuel consumption of diesel and hybrid buses under various driving conditions. Artificial neural network (ANN) based high-fidelity microscopic (1 Hz) and mesoscopic (5–60 min) fuel consumption models were developed for hybrid buses. The microscopic model contained 1 Hz driving, grade, and environment variables. The mesoscopic model aggregated 1 Hz data into 5 to 60-minute traffic pattern factors and predicted average fuel consumption over its duration. The prediction results show mean absolute percentage errors of 1–2% for microscopic models and 5–8% for mesoscopic models. The data were partitioned by different driving speeds, vehicle engine demand, and road grade to investigate their impacts on prediction performance.

33 ADVANCED PROPULSION SYSTEMS↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗