Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network partition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Implementing Directive-Based Deferred Execution for Effective Network Aggregation

Remote direct memory access technology provides an efficient mechanism for one-sided communication that can be leveraged to implement a distributed shared memory programming model. However, when applications generate large numbers of small, irregular messages, network congestion often arises. Existing solutions address this small message problem by facilitating message aggregation but typically require disruptive code transformations that detract from the algorithmic intent of applications, or can be limited by dependent operations on aggregated data between synchronisation points. A solution is to use a directive-assisted approach that enables compilers to transform code dependent on aggregated communication for deferred execution. This paper presents an algorithm that a compiler can use to implement and optimise deferred execution for code dependent on aggregated data, based on an "aggregation context" extension for the OpenSHMEM partitioned global address space library. This new capability addresses a key challenge of message aggregation, allowing its full potential to reduce network congestion and enhance programmability to be realised.

Welch, Aaron [ORNL]↗

Configuration Space Integration for Adsorbate Partition Functions: The Effect of Anharmonicity on the Thermophysical Properties of CO–Pt(111) and CH 3 OH–Cu(111)

A method for computing anharmonic thermophysical properties for adsorbates on metal surfaces has been extended to include libration, or frustrated rotation. Classical phase space integration is used with Monte Carlo sampling of the configuration space to obtain the partition function of CO on Pt(111) and CH 3 OH on Cu(111). A minima-preserving neural network potential energy surrogate is used within the integration routines. Direct state counting using discrete variable representation is used to benchmark the results. We find that the phase space integration approach is in excellent agreement with the direct state counting results. Comparison with standard models such as the harmonic oscillator indicates that anharmonicity contributes significantly to the thermodynamic properties of CH 3 OH on Cu(111). We find that there is also a considerable difference between the harmonic oscillator and phase space integration for CO on Pt(111), although the discrepancy can largely be attributed to the presence of multiple binding sites within the unit cell. We demonstrate that a multisite harmonic oscillator model might be sufficient for CO-Pt(111). A more thorough description of the potential energy surface, which can be achieved with phase space integration, is necessary for weakly bound adsorbates such as CH 3 OH. In conclusion, the thermophysical properties were used to calculate free energies of adsorption on the respective metals, and subsequently the equilibrium constants and Langmuir isotherms in relevant temperature ranges. The results show that the choice of model to obtain partition functions greatly affects the resulting surface coverages in kinetic models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Instrumentation, Monitoring and Visualization of PVM Programs Using AIMS

We present views and analysis of the execution of several PVM codes for Computational Fluid Dynamics on a network of Sparcstations, including (a) NAS Parallel benchmarks CG and MG (White, Alund and Sunderam 1993); (b) a multi-partitioning algorithm for NAS Parallel Benchmark SP (Wijngaart 1993); and (c) an overset grid flowsolver (Smith 1993). These views and analysis were obtained using our Automated Instrumentation and Monitoring System (AIMS) version 3.0, a toolkit for debugging the performance of PVM programs. We will describe the architecture, operation and application of AIMS. The AIMS toolkit contains (a) Xinstrument, which can automatically instrument various computational and communication constructs in message-passing parallel programs; (b) Monitor, a library of run-time trace-collection routines; (c) VK (Visual Kernel), an execution-animation tool with source-code clickback; and (d) Tally, a tool for statistical analysis of execution profiles. Currently, Xinstrument can handle C and Fortran77 programs using PVM 3.2.x; Monitor has been implemented and tested on Sun 4 systems running SunOS 4.1.2; and VK uses X11R5 and Motif 1.2. Data and views obtained using AIMS clearly illustrate several characteristic features of executing parallel programs on networked workstations: (a) the impact of long message latencies; (b) the impact of multiprogramming overheads and associated load imbalance; (c) cache and virtual-memory effects; and (4significant skews between workstation clocks. Interestingly, AIMS can compensate for constant skew (zero drift) by calibrating the skew between a parent and its spawned children. In addition, AIMS' skew-compensation algorithm can adjust timestamps in a way that eliminates physically impossible communications (e.g., messages going backwards in time). Our current efforts are directed toward creating new views to explain the observed performance of PVM programs. Some of the features planned for the near future include: (a) ConfigView, showing the physical topology of the virtual machine, inferred using specially formatted IP (Internet Protocol) packets; and (b) LoadView, synchronous animation of PVM-program execution and resource-utilization patterns.

Mehra, Pankaj↗

Distributed PV Hosting Capacity Evaluation Considering Equitable PV Accommodation

Distributed photovoltaic (DPV) hosting capacity is widely used by distribution system operators to evaluate how much DPV the distribution networks can accommodate without violating operational constraints. If the proliferation of DPV deployment is appropriately managed, it offers an excellent opportunity to mitigate the energy inequity issue in electric grids. Hosting capacity improvement plans that neglect energy equity can exacerbate unequal DPV access. To consider energy equity in distribution network planning to enhance DPV accommodation, hosting capacity evaluations should be improved to consider equity constraints endogenously. This paper proposes a DPV hosting capacity evaluation method that considers equitable DPV access among sub-regions of a distribution feeder partitioned by socioeconomic indicators, such as income level. The results demonstrate that considering equity in DPV hosting capacity evaluations can provide a valuable reference in the equitable energy transition. This work provides guidance for future distribution network upgrade planning for equitable DPV deployment.

distributed PV↗

Scalable Approaches to Selecting Key Entities in Large Networked Infrastructure Systems

This work aims at bringing advances in discrete optimization algorithms to solving practical engineering problems at scale. Often times, in many engineering design problems, there is a need to select a small set of influential or representative elements from a large ground set of entities in an optimal fashion. Submodular optimization provides for a formal way to solve such problems. Common examples with infrastructure systems involve sensor placement and identification of key entities with certain objectives. However, scaling these approaches to large infrastructure systems can be challenging because of the high computational complexity of the overall framework that include the optimization algorithms as well as high-complexity compute-oracles that provide the necessary objective function values. In this work, we explore a well-studied and widely-applicable paradigm, namely leader-selection in a multi-agent networked setting in the context of scalable methodologies. We demonstrate novel frameworks that utilize variations of accelerated submodular optimization algorithms along with linear-algebraic methods that can help accelerate the oracle computations. We further explore this combination in conjunction with graph partitioning paradigms to take advantage of the accelerated algorithms in a distributed setting. Finally we demonstrate the key findings on a practical problem in an operational setting. For this, we leverage an example road network with approximately 18k nodes and 27k edges in a traffic control application, where we seek a limited number of k=200 key intersections. This problem can be solved in a serial setting in just under 5 hours providing more than 2 orders of magnitude speed-up over methods that do not consider acceleration techniques.

Visweswara Sathanur, Arun↗

Assembly of polyelectrolyte star block copolymers at the oil–water interface

To understand and resolve adsorption, reconfiguration, and equilibrium conformations of charged star copolymers, we carried out an integrated experimental and coarse-grained molecular dynamics simulation study of the assembly process at the oil–water interface. This is important to guide development of novel surfactants or amphiphiles for chemical transformations and separations. The star block copolymer consisted of arms that are comprised of hydrophilic–hydrophobic block copolymers that are covalently tethered via the hydrophobic blocks to one point. The hydrophobic core represents polystyrene (PS) chains, while the hydrophilic corona represents quaternized poly(2-vinylpyridine) (P2VP) chains. The P2VP is modeled to become protonated when in contact with an acidic aqueous phase, thereby massively increasing the hydrophilicity of this block, and changing the nature of the star at the oil–water interface. This results in a configurational change whereby the chains comprising the hydrophilic corona are significantly stretched into the aqueous phase, while the hydrophobic core remains solubilized in the oil phase. In the simulations, we followed the kinetics of the anchoring and assembly of the star block copolymer at the interface, monitoring the lateral assembly, and the subsequent reconfiguration of the star via changes in the interfacial tension that varies as the degree-of-protonation increases. At low fractions of protonation, the arm cannot fully partition into the aqueous side of the interface and instead interacts with other arms in the oil phase forming a network near the interface. These insights were used to interpret the non-monotonic dependence of pH with the asymptotic interfacial tension from pendant drop tensiometry experiments and spectral signatures of aromatic stretches seen in vibrational sum frequency generation (SFG) spectroscopy. We describe the relationship of interfacial tension to the star assembly via the Frumkin isotherm, which phenomenologically describes anti-cooperativity in adsorbing stars to the interface due to crowding. Although our model explicitly considers long-range electrostatics, the contribution of electrostatics to interfacial tension is small and brought about by strong counterion condensation at the interface. Finally, these results provide key insights into resolving the adsorption, reconfiguration, and equilibrium conformations of charged star block copolymers as surfactants.

36 MATERIALS SCIENCE↗

Modeling and measurement of fault-tolerant multiprocessors

The workload effects on computer performance are addressed first for a highly reliable unibus multiprocessor used in real-time control. As an approach to studing these effects, a modified Stochastic Petri Net (SPN) is used to describe the synchronous operation of the multiprocessor system. From this model the vital components affecting performance can be determined. However, because of the complexity in solving the modified SPN, a simpler model, i.e., a closed priority queuing network, is constructed that represents the same critical aspects. The use of this model for a specific application requires the partitioning of the workload into job classes. It is shown that the steady state solution of the queuing model directly produces useful results. The use of this model in evaluating an existing system, the Fault Tolerant Multiprocessor (FTMP) at the NASA AIRLAB, is outlined with some experimental results. Also addressed is the technique of measuring fault latency, an important microscopic system parameter. Most related works have assumed no or a negligible fault latency and then performed approximate analyses. To eliminate this deficiency, a new methodology for indirectly measuring fault latency is presented.

Shin, K. G.↗

JPL control/structure interaction test bed real-time control computer architecture

The Control/Structure Interaction Program is a technology development program for spacecraft that exhibit interactions between the control system and structural dynamics. The program objectives include development and verification of new design concepts - such as active structure - and new tools - such as combined structure and control optimization algorithm - and their verification in ground and possibly flight test. A focus mission spacecraft was designed based upon a space interferometer and is the basis for design of the ground test article. The ground test bed objectives include verification of the spacecraft design concepts, the active structure elements and certain design tools such as the new combined structures and controls optimization tool. In anticipation of CSI technology flight experiments, the test bed control electronics must emulate the computation capacity and control architectures of space qualifiable systems as well as the command and control networks that will be used to connect investigators with the flight experiment hardware. The Test Bed facility electronics were functionally partitioned into three units: a laboratory data acquisition system for structural parameter identification and performance verification; an experiment supervisory computer to oversee the experiment, monitor the environmental parameters and perform data logging; and a multilevel real-time control computing system. The design of the Test Bed electronics is presented along with hardware and software component descriptions. The system should break new ground in experimental control electronics and is of interest to anyone working in the verification of control concepts for large structures.

Briggs, Hugh C.↗

Distributed PV Hosting Capacity Evaluation Considering Equitable PV Accommodation: Preprint

In distribution systems, distributed photovoltaic (DPV) hosting capacity is widely used by system operators to evaluate how much DPV the distribution networks can accommodate without violating operational constraints. With the proliferation of DPV deployment, if managed properly, it offers a great opportunity to mitigate the persistent energy inequity issue in electric grid. Otherwise, energy inequity can be exacerbated with unequal DPV access when hosting capacity improvement plans ignore the equity factor. To consider energy equity in distribution network planning for enhancing DPV accommodation, the hosting capacity evaluation should be improved to consider equity constraints endogenously. In this paper, a DPV hosting capacity evaluation method is proposed considering equitable DPV access among sub-regions of a distribution feeder partitioned by social-economic indicators such as the income levels. The results demonstrate that considering equity in DPV hosting capacity evaluation can provide valuable reference in the equitable energy transition. This work provides a valuable guidance for future distribution network upgrade planning for equitable DPV deployment.

distributed PV↗

Using Hyperoptimized Tensor Networks and First-Principles Electronic Structure to Simulate the Experimental Properties of the Giant {Mn 84 } Torus

The single-molecule magnet {Mn 84 } is a challenge to theory because of its high nuclearity. Here, we directly compute two experimentally accessible observables, the field-dependent magnetization up to 75 T and the temperature-dependent heat capacity, using parameter-free theory. In particular, we use first-principles calculations to derive short- and long-range exchange interactions and compute the exact partition function of the resulting classical Potts and Ising spin models for all 84 Mn S = 2 spins to obtain observables. The latter computation is made possible by using hyperoptimized tensor network contractions, a technique developed to simulate quantum supremacy circuits. We also synthesize the magnet and measure its heat capacity and magnetization, observing qualitative agreement between theory and experiment and identifying an unusual bump in the heat capacity and a plateau in the magnetization. Our work also identifies some limitations of current theoretical modeling in large magnets, such as sensitivity to small, long-range exchange couplings.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Transforming AdaPT to Ada

This paper describes how the main features of the proposed Ada language extensions intended to support distribution, and offered as possible solutions for Ada9X can be implemented by transformation into standard Ada83. We start by summarizing the features proposed in a paper (Gargaro et al, 1990) which constitutes the definition of the extensions. For convenience we have called the language in its modified form AdaPT which might be interpreted as Ada with partitions. These features were carefully chosen to provide support for the construction of executable modules for execution in nodes of a network of loosely coupled computers, but flexibly configurable for different network architectures and for recovery following failure, or adapting to mode changes. The intention in their design was to provide extensions which would not impact adversely on the normal use of Ada, and would fit well in style and feel with the existing standard. We begin by summarizing the features introduced in AdaPT.

Goldsack, Stephen J.↗

On the Generalizability of Time-of-Flight Convolutional Neural Networks for Noninvasive Acoustic Measurements

Bulk wave acoustic time-of-flight (ToF) measurements in pipes and closed containers can be hindered by guided waves with similar arrival times propagating in the container wall, especially when a low excitation frequency is used to mitigate sound attenuation from the material. Convolutional neural networks (CNNs) have emerged as a new paradigm for obtaining accurate ToF in non-destructive evaluation (NDE) and have been demonstrated for such complicated conditions. However, the generalizability of ToF-CNNs has not been investigated. In this work, we analyze the generalizability of the ToF-CNN for broader applications, given limited training data. We first investigate the CNN performance with respect to training dataset size and different training data and test data parameters (container dimensions and material properties). Furthermore, we perform a series of tests to understand the distribution of data parameters that need to be incorporated in training for enhanced model generalizability. This is investigated by training the model on a set of small- and large-container datasets regardless of the test data. We observe that the quantity of data partitioned for training must be of a good representation of the entire sets and sufficient to span through the input space. The result of the network also shows that the learning model with the training data on small containers delivers a sufficiently stable result on different feature interactions compared to the learning model with the training data on large containers. To check the robustness of the model, we tested the trained model to predict the ToF of different sound speed mediums, which shows excellent accuracy. Furthermore, to mimic real experimental scenarios, data are augmented by adding noise. We envision that the proposed approach will extend the applications of CNNs for ToF prediction in a broader range.

47 OTHER INSTRUMENTATION↗

Using an Isotope Enabled Mass Balance to Evaluate Existing Land Surface Models

Abstract Land surface models (LSMs) play a crucial role in elucidating water and carbon cycles by simulating processes such as plant transpiration and evaporation from bare soil, yet calibration often relies on comparing LSM outputs of landscape total evapotranspiration ( ET ) and discharge with measured bulk fluxes. Discrepancies in partitioning into component fluxes predicted by various LSMs have been noted, prompting the need for improved evaluation methods. Stable water isotopes serve as effective tracers of component hydrologic fluxes, but data and model integration challenges have hindered their widespread application. Leveraging National Ecological Observation Network measurements of water isotope ratios at 16 US sites over 3 years combined with LSM‐modeled fluxes, we employed an isotope‐enabled mass balance framework to simulate ET isotope values ( δET ) within three operational LSMs (Mosaic, Noah, and VIC) to evaluate their partitioning. Models simulating δET values consistent with observations were deemed more reflective of water cycling in these ecosystems. Mosaic exhibited the best overall performance (Kling‐Gupta Efficiency of 0.28). For both Mosaic and Noah there were robust correlations between bare soil evaporation fraction and error (negative) as well as transpiration fraction and error (positive). We found the point at which errors are smallest ( x ‐intercept of the multi‐site regression) is at a higher transpiration fraction than is currently specified in the models. Which means that transpiration fraction is underestimated on average. Stable isotope tracers offer an additional tool for model evaluation and identifying areas for improvement, potentially enhancing LSM simulations and our understanding of land‐surface hydrologic processes.

58 GEOSCIENCES↗

Overexpression of plasma membrane SUT1 in poplar alters lateral sucrose partitioning in stem and promotes leaf necrosis

Abstract In Populus and many other tree species, photoassimilate sucrose diffuses down a concentration gradient via symplastically connected mesophyll cells to minor vein phloem for long‐distance transport. There is no evidence for apoplastic phloem‐loading in Populus . However, plasma membrane sucrose transporters (SUT1 and SUT3) orthologous to those associated with apoplastic phloem loading are expressed in vascular tissues of poplar. While SUT3 functions in sucrose import into developing xylem, the role of SUT1 remains unclear. Here, we overexpressed PtaSUT1 in Populus tremula x P. alba to examine the effects on sucrose partitioning in transgenic plants. Overall leaf sucrose levels were similar between wild type and transgenic lines. Stem sucrose levels were not changed in bark but were significantly reduced in the adjacent xylem, suggesting hindered intercellular sucrose trafficking from the phloem to the developing xylem. Fully expanded leaves of transgenic plants deteriorated prematurely with declining photosynthesis prior to severe necrotic spotting. Necrotic spotting advanced most rapidly in the distal portion of mature leaves and was accompanied by sharp hexose increases and sharp sucrose decreases there. Leaf transcriptome profiling and network inference revealed the down‐regulation of copper proteins and elevated expression of copper microRNAs prior to noticeable leaf injury. Our results suggest ectopic expression of PtaSUT1 altered sucrose partitioning in stems with systemic effects on leaf health and copper homeostasis mediated in part by sucrose‐sensitive copper miRNAs.

59 BASIC BIOLOGICAL SCIENCES↗

Coarse-grained fixed-point tensor networks and holographic reflected entropy in 3D gravity

We use the framework of fixed-point BCFT tensor networks to present a microscopic CFT derivation of the correspondence between reflected entropy (RE) and entanglement wedge cross section (EW) in AdS 3 /CFT 2 , for both bipartite and multipartite settings. These fixed-point tensor networks, obtained by triangulating Euclidean CFT path integrals, allow us to explicitly construct the canonical purification via cutting-and-gluing CFT path integrals. Employing modular flow in the large-c limit, we demonstrate that these intrinsic CFT manipulations reproduce bulk geometric prescriptions, without assuming the AdS/CFT dictionary. The emergence of bulk geometry is traced to coarse-graining over heavy states in the large-c limit. Universal coarse-grained BCFT data for compact 2D CFTs, through the relation to Liouville theory with ZZ boundary conditions, yields hyperbolic geometry on the Cauchy slice. The corresponding averaged replica partition functions reproduce all candidate EWs, arising from different averaging patterns, with the dominant one providing the correct RE and EW. In this way, many heuristic tensor-network intuitions in toy models are made precise and established directly from intrinsic CFT data.

AdS-CFT correspondence↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Use of networked workstations for parallel nonlinear structural dynamic simulations of rotating bladed-disk assemblies

The principal objective of this research is to investigate, develop and demonstrate coarse-grained, parallel-processing strategies for nonlinear dynamic simulations for rotating bladed-disk assemblies. The parallel -processing strategies addressed include numerical algorithms for parallel nonlinear solutions and techniques to effect load balancing among processors. The parallel environment employed is a distributed-memory, coarse-grained one consisting of networked workstations. A parallel explicit time integration method has been implemented for transient nonlinear solutions of rotationg bladed-disk assemblies. Automatic domain partitioning techniques have been investigated for load balancing among processors. Advanced computing environments, data structures and interactive computer graphics all contribute to an integrated parallel finite element analysis system to facilitate more efficient and powerful dynamic simulations.

Hsieh, Shang-Hsien↗