Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Performance and usability enhancements for continuous subgraph matching queries on graph-structured data

A query graph, which includes vertices and edges, represents a query on graph-structured data. The query graph is decomposed into query subgraphs. A network analysis tool performs continuous subgraph matching queries to facilitate analysis of computer network traffic, social media events, or other streams of data represented as a dynamic data graph (graph-structured data). This can help identify emerging trends in the data. Some features of the network analysis tool enhance performance by effectively utilizing distributed computing resources (including processing cores and memory at different nodes of a cluster) to speed up the process of updating the dynamic data graph and detecting matches of query subgraphs. Features of a query graph building tool enhance usability by providing intuitive ways to specify query graphs and their subgraphs. Features of a results visualization tool enhance usability by providing an intuitive way to present the results of continuous subgraph matching queries.

Choudhury, Sutanay↗

Harness the Power of AI and CI/CD to Fuel Scientific Discovery

The "Harness the Power of AI and CI/CD to Fuel Scientific Discovery" project aims to enhance and automate critical scientific computing systems used in large-scale experiments like CMS at LHC and DUNE at Fermilab. By leveraging GlideinWMS and HEPCloud, this initiative focuses on developing containerized CI/CD pipelines, integrating AI for code quality improvement, and automating security verifications. Participants will gain hands-on experience with distributed computing systems and implement secure communications, contributing to real-world scientific progress and the open-source community.

Nurcellari, Tea↗

Trigger-based Incremental Data Processing with Unified Sync and Async Model

In recent years, more and more applications in the cloud have needs to process large-scale on-line datasets, which evolve over time as new entries are added and existing entries are modified. Several programming frameworks, such as Percolator and Oolong, are proposed for such incremental data processing and can achieve efficient processing with an event-driven abstraction. However, these frameworks are inherently asynchronous, leaving the heavy burden of managing synchronization to applications' developers, which further significantly restricts their usabilities. In this study, we propose a trigger-based incremental computing framework in the cloud, called Domino, with both synchronous and asynchronous mechanisms to coordinate parallel triggers. With this new framework, both synchronous and asynchronous applications can be seamlessly developed. Use cases and extensive evaluation results confirm that it can deliver sufficient performance, and also is easy to use for incremental applications in large-scale distributed computing.

97 MATHEMATICS AND COMPUTING↗

Distributionally Robust Bilevel Optimization Model for Distribution Network With Demand Response Under Uncertain Renewables Using Wasserstein Metrics

Here, we consider a distribution network integrating demand response (DR) participants in the presence of uncertain renewable suppliers and outdoor temperatures. A bilevel optimization model is proposed to capture the intricate dynamics between price-incentivized DR participants and distribution system operations, including energy procurement and active/reactive power flows. The model is formulated as a distributional robust bilevel optimization using Wasserstein metrics. We show favorable data-driven properties including out-of-sample guarantee and asymptotic consistency. Furthermore, we present a tractable mixed-integer linear programming reformulation and characterize the worst-case distribution. Computational experiments are conducted on a modified 33-bus system. Our findings underscore the efficacy of the pricing strategies derived from the proposed bilevel optimization model. These strategies not only effectively manage DR participants' behavior but also bring equity considerations among households with various characteristics to light. The results contribute to a deeper understanding of the interplay between distribution system operators and DR participants.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Simulation Study of Quantum Clock Synchronization Using Teleportation

An important requirement in implementing distributed computing and sensing application is the synchronization of clocks at various locations. The Internet relies on the Network Time Protocol (NTP), which synchronizes clocks with accuracy in the order of milliseconds. More recently an ensemble of atomic clocks is used for navigation based on GPS. These clocks are highly accurate and provide time with very low uncertainty. Even so, many physics experiments such as distributed LIGO-based systems may require more accurate clock synchronization that is achievable using quantum entanglement. This requires the deployment of a network of quantum clocks synchronized by exploiting entangled atomic clock qubits. In this paper, we carry out a simulation study of synchronizing a network of quantum clocks interconnected by a fiber plant that supports the ESnet; the latter is used to support the classical communication needed for teleportation. We consider an existing protocol for synchro-nizing the atomic clock qubits that relies on the GHZ states. To assess the performance of the protocol we developed a discrete-event simulation of the network using IBM Qiskit framework for underlying quantum gate operations and measurements. The simulation results shed light on the resources required in terms of the number entangled qubits and the time needed to achieve the synchronization of different number of nodes in ESnet.

Kiran, Mariam↗

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

A method for characterization of multiple dynamic constitutive parameters of FRCs

We propose a method to measure multiple dynamic material constitutive parameters of unidirectional fiber reinforced composites (FRCs) in a single experiment. Dynamic short-beam shear (DSBS) experiments were performed on a modified Kolsky compression bar, with integration of high-speed imaging and digital image correlation (DIC). The unidirectional FRCs investigated were S-2 glass fiber reinforced matrix of TGDDM-Jeffamine® D230 with monoamine functionalized partially reacted substructures (mPRS) and commercially available SC-15. Analytical solutions of normal and shear strains of a composite beam were derived based on Timoshenko beam theory, assuming material to be transversely isotropic and have different moduli in tension and compression in each principle material orientation. Tensile and compressive moduli were inversely computed through monitoring normal strain slope when specimen was constantly loaded at a speed of ~7.3 m/s within a small deflection. Non-linear shear stress-strain behavior of the composite was described via Ramberg–Osgood equation. Finite element (FE) analysis was conducted in ABAQUS, simultaneously defining via user subroutine UMAT the transverse isotropy of material, bi-modulus constitutive model, and non-linear shear stress-strain relation. The method proposed in this work was validated by comparing strain distributions computed by FE model and DIC measurements. Comparing with traditional dynamic tensile, compressive, and shear experiments on FRCs, this method significantly simplifies the specimen preparation and design of complicated gripping fixtures for multiple experiments. Here, systematic errors resulting from variations of specimen geometry and dimension, loading direction, and instrumentation are reduced, thereby providing compatible data for numerical studies on impact behavior of composites.

36 MATERIALS SCIENCE↗

Potential of the Julia Programming Language for High Energy Physics Computing

Research in high energy physics (HEP) requires huge amounts of computing and storage, putting strong constraints on the code speed and resource usage. To meet these requirements, a compiled high-performance language is typically used; while for physicists, who focus on the application when developing the code, better research productivity pleads for a high-level programming language. A popular approach consists of combining Python, used for the high-level interface, and C++, used for the computing intensive part of the code. A more convenient and efficient approach would be to use a language that provides both high-level programming and high-performance. The Julia programming language, developed at MIT especially to allow the use of a single language in research activities, has followed this path. In this paper the applicability of using the Julia language for HEP research is explored, covering the different aspects that are important for HEP code development: runtime performance, handling of large projects, interface with legacy code, distributed computing, training, and ease of programming. The study shows that the HEP community would benefit from a large scale adoption of this programming language. The HEP-specific foundation libraries that would need to be consolidated are identified.

97 MATHEMATICS AND COMPUTING↗

Automated Image Segmentation and Processing Pipeline Applied to X–Ray Computed Tomography Studies of Pitting Corrosion in Aluminum Wires

Understanding pitting corrosion is critical, yet its kinetics and morphology remain challenging to study from X-ray computed tomography (XCT) due to manual segmentation barriers. To address this, an automated pipeline leveraging deep learning for efficient large-scale XCT analysis is developed, revealing new corrosion insights. The pipeline enables pit segmentation, 3D reconstruction, statistical characterization, and a topological transformation for visualization. Here, the pipeline is applied to 87 648 XCT images capturing commercial purity aluminum (1100 Al) wire exposed to sodium chloride (NaCl) salt particles over a period of 122 h. The pipeline achieves complete feature extraction and statistical quantification across the entire XCT dataset, leveraging distributed computing environment for high efficiency. Global growth kinetics such as high-level stepwise sigmoidal volume loss patterns and granular individual pit developments are both captured for 36 detected pits. By combining automation, computer vision, and extensive XCT datasets, this research accelerates precise corrosion assessment to enable materials science discoveries at scale.

36 MATERIALS SCIENCE↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Multinode Multi-GPU Two-Electron Integrals: Code Generation Using the Regent Language

The computation of two-electron repulsion integrals (ERIs) is often the most expensive step of integral-direct self-consistent field methods. Formally it scales as O(N 4 ), where N is the number of Gaussian basis functions used to represent the molecular wave function. In practice, this scaling can be reduced to O(N 2 ) or less by neglecting small integrals with screening methods. The contributions of the ERIs to the Fock matrix are of Coulomb (J) and exchange (K) type and require separate algorithms to compute matrix elements efficiently. We previously implemented highly efficient GPU-accelerated J-matrix and K-matrix algorithms in the electronic structure code TeraChem. Although these implementations supported the use of multiple GPUs on a node, they did not support the use of multiple nodes. This presents a key bottleneck to cutting-edge ab initio simulations of large systems, e.g., excited state dynamics of photoactive proteins. We present our implementation of multinode multi-GPU J- and K-matrix algorithms in TeraChem using the Regent programming language. Regent directly supports distributed computation in a task-based model and can generate code for a variety of architectures, including NVIDIA GPUs. We demonstrate multinode scaling up to 45 GPUs (3 nodes) and benchmark against hand-coded TeraChem integral code. Finally, we also outline our metaprogrammed Regent implementation, which enables flexible code generation for integrals of different angular momenta.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Framework for International Collaboration on ITER Using Large-Scale Data Transfer to Enable Near-Real-Time Analysis

The global nature of the ITER project along with its projected ~ petabyte per day data generation presents a unique challenge, but also an opportunity for the fusion community to rethink, optimize, and enhance our scientific discovery process. Recognizing this, collaborative research with computational scientists was undertaken over the past several years to create a framework for large-scale data movement across wide-area networks (WANs), to enable global near-real time analysis of fusion data. This would broaden the available computational resources for analysis/simulation, and increase the number of researchers actively participating in experiments. An official demonstration of this framework for fast, large data transfer and real-time analysis was carried out between the KSTAR tokamak in Daejeon, Korea and PPPL in Princeton, USA. Streaming large data transfer, with near real-time movie creation and analysis of the KSTAR Electron Cyclotron Emission imaging (ECEI) data, was performed using the I/O framework ADIOS, and comparisons made at PPPL with simulation results from the XGC1 code. These demonstrations were made possible utilizing an optimized network configuration at PPPL, which achieved over 8.8 Gbps (88% utilization) in throughput tests from NFRI to PPPL. This demonstration showed the feasibility for large-scale data analysis of KSTAR data, and provides a nascent framework to enable use of globally distributed computational and personnel resources in pursuit of scientific knowledge from the ITER experiment.

43 PARTICLE ACCELERATORS↗

Quantifying uncertainties due to irreducible three-body forces in deuteron-nucleus reactions

Deuteron-induced nuclear reactions are an essential tool for probing the structure of nuclei as well as astrophysical information such as (n, γ) cross sections. The deuteron-nucleus system is typically described within a Faddeev three-body model consisting of a neutron (n), a proton (p), and the target nucleus (A) interacting through pairwise phenomenological potentials. While Faddeev techniques enable the exact description of the three-body dynamics, their predictive power is limited in part by the omission of irreducible neutron-proton-nucleus three-body force (n–p–A 3BF). Here, our goal is to quantify systematic uncertainties stemming from the reduction of deuteron-nucleus (d + A) dynamics to a picture of three pointlike nuclear clusters interacting via pairwise nucleon-nucleus forces, using as testing grounds d + α scattering and the 6 Li ground state. We particularly focus on quantifying uncertainties arising from the full antisymmetrization of the (A + 2)-body system with the target nucleus fixed in its ground state. We adopt the ab initio no-core shell model coupled with the resonating group method (NCSM/RGM) to compute microscopic n–α and p–α interactions, and use them in a three-body description of the d + α system by means of momentum-space Faddeev-type equations. Simultaneously, we also carry out ab initio calculations of d + α scattering and 6 Li ground state by means of six-body NCSM/RGM calculations to serve as a benchmark for the three-body model predictions given by the Faddeev calculations. By comparing the Faddeev and NCSM/RGM results, we show that the irreducible n–p–α 3BF has a non-negligible effect on bound state and scattering observables alike. Specifically, the Faddeev approach yields a 6 Li ground state that is approximately 600 keV shallower than the one obtained with the NCSM/RGM. Additionally, the Faddeev calculations for d + α scattering yield a 3 + resonance that is located approximately 400 keV higher in energy compared to the NCSM/RGM result. The shape of the d + α angular distributions computed using the two approaches also differ, owing to the discrepancy in the predictions of the 3 + resonance energy. The Faddeev three-body model predictions for d + α scattering and 6 Li using microscopic n–α and p–α potentials differ from those computed microscopically with the NCSM/RGM. These discrepancies are due to the n–p–α 3BF, which arises from two-nucleon exchange terms in the microscopic d–α interaction and are not accounted for in the three-body model Faddeev calculations. This study lays the foundation for future parametrizations of the 3BF due to Pauli exclusion principle effects in improved three-body calculations of deuteron-induced reactions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

HIPPO – A Software Platform for Electricity Market Research and Development

The goal of this project is to provide Regional transmission organizations (RTOs) and independent system operators (ISOs) a market design and prototyping software, High-Performance Power-Grid Optimization (HIPPO), that they can evaluate electricity market design options, calculate market planning strategies and operational performance. With the high standards and strict reliability requirements for operating power systems, impacts of new technologies need to be fully investigated prior to any consideration for adoption. A market design and prototyping software tool which can be used to prototype electricity market design options, to calculate market planning strategies and operational performance with high precision, and to investigate the impacts for integrating future power grid technologies will be valuable to RTOs/ISOs who operate power systems, to vendors like GE and ABB who provide the market solvers, and to market participants and researchers who are actively doing market research. HIPPO is a such tool that can be used to improve the current market operations and provide capabilities for rigorous forward-looking design and prototyping of next-generation energy markets. HIPPO has a high-resolution model for the day-ahead SCUC, which was validated with MISO and GE-Grid Solutions. HIPPO is built with parallel and distributed computing capabilities and can be executed in both multi-thread and high-performance computing (HPC) settings. This capability provides fast solution speed necessary to handle the larger and more complex SCUC problems of real-world cases and the potentially growing size and complexity of future scenarios. In addition, HIPPO has a concurrent optimizer (CO) which manages multiple algorithm executions simultaneously and leverages the advantages from different algorithms. This structure provides flexibility to better benchmark competing approaches. Highly accurate market model, fast solution technologies and flexible model and algorithm control are the features which will make HIPPO an extensible platform for developing and testing multiple approaches to meet a wide range of future market needs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Resilient Information Architecture Platform for Smart Grid (RIAPS)

A number of emerging trends will substantially alter the operation and control of the electric grid over the next several decades. These trends include ensuring resiliency under severe weather events, increasing integration of renewable electricity generation, supporting changing electricity demand patterns, and the improving cost effectiveness of distributed energy resources. To address these challenges, the future “Smart Grid” management will need to transition from centralized to coordinated distributed control paradigm. Reliable operation of the Smart Grid depends on distributed intelligence realized through software applications that run on distributed computing devices attached to the power system to collect data and collaboratively manage resources. However, much of the existing software for Smart Grid-enabled devices is either proprietary or developed with custom solutions, which limits interoperability among the heterogeneous devices and hinders the ability to manage system-level reliability, security, and resiliency requirements. Additionally, this approach makes Smart Grid applications hard to maintain, evolve, verify, and replace; resulting in high development and deployment costs. Further development of the Smart Grid requires a reusable software base-layer to move from hard-coded functionality to a plug-and-play architecture capable of managing system-level objectives and constraints in addition to providing consistent common services across heterogeneous devices and applications. Vanderbilt University, in collaboration with North Carolina State University and Washington State University has developed a foundation ‘software platform’ for developing and deploying robust, reliable, effective and secure software applications for the Smart Grid. The Resilient Information Architecture Platform for the Smart Grid (RIAPS) provides core services for building effective and powerful smart grid applications. It offers unique services for real-time data dissemination, fault tolerance, and coordination across apps distributed over the network.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Proximity Portability and in Transit , M-to-N Data Partitioning and Movement in SENSEI [Book Chapter]

In high-performance parallel in situ processing, the term in transit processing refers to those configurations where data must move from a producer to a consumer that runs on separate resources. In the context of parallel and distributed computing on an HPC platform one of the central challenges is to determine a mapping of data from producer ranks to consumer ranks. This problem is complicated by the heterogeneity that arises in producer-consumer pairs, such as when producer and consumer codes have different levels of concurrency, different scaling characteristics, or different data models. The resulting mapping and movement of data from M producer to N consumer ranks can have a significant impact on aggregate application performance, particularly when the data consumer requires only a subset of the overall data for its task. This chapter focuses on the design considerations that underlie SENSEI’s implementation to this challenging problem. These design considerations extend the core SENSEI architecture and include ideas like the need to accommodate flexibility in the choice of different partitioning methods, the ability for a data consumer to request and receive only the subset of data needed for its particular operation, and the ability to leverage any of several different data transport tools. The idea of proximity portability, being able to use different data transport methods as part of an in transit workflow, is illustrated through the use of three different transport layers where switching from one transport tool to another is accomplished with only a configuration file change. Here, the chapter also includes a performance analysis summary showing the performance gains that are possible in terms of multiple metrics, such as memory footprint, time to solution, and amount of data moved, when using optimized partitioners in an in transit setting, gains that are made possible by the implementation shaped by specific design considerations.

Bethel, E. Wes↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗