Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Coupling hydrodynamics to a rigid-body motion solver for fluid-structure interaction [Slides]

Quinoa is a massively parallel computational fluid dynamics with multi-material and programmed burn capabilities developed from Programmatic and LDRD funds. Overset is a well-established method of using an “inset” mesh that communicates with a background mesh. Solutions transfer freely from one to another and operate as boundary conditions on the opposite mesh. Either mesh can move at specified velocity! This project couples the existing Mesh-to-mesh transfer with Quinoa and demonstrates use of resulting Overset solver in computational simulation of blast effects on a re-entry body for survivability assessments.

42 ENGINEERING↗

High-dimensional discrete Fourier transform gates with a quantum frequency processor

The discrete Fourier transform (DFT) is of fundamental interest in photonic quantum information, yet the ability to scale it to high dimensions depends heavily on the physical encoding, with practical recipes lacking in emerging platforms such as frequency bins. In this article, we show that d -point frequency-bin DFTs can be realized with a fixed three-component quantum frequency processor (QFP), simply by adding to the electro-optic modulation signals one radio-frequency harmonic per each incremental increase in d . We verify gate fidelity F W > 0.9997 and success probability P W > 0.965 up to d = 10 in numerical simulations, and experimentally implement the solution for d = 3, utilizing measurements with parallel DFTs to quantify entanglement and perform tomography of multiple two-photon frequency-bin states. Our results furnish new opportunities for high-dimensional frequency-bin protocols in quantum communications and networking.

97 MATHEMATICS AND COMPUTING↗

A structural description of the evolution of stakeholders and risk communication in the Department of Energy's defense nuclear facilities: Historical perspective, major stakeholders, and external events

Regulators and policymakers are routinely challenged with explaining complex concepts concerning risk. Part of the challenge is helping external and internal stakeholders to understand the context behind risk-related information and decisions. Here, this paper will describe the historical evolution of the safety and regulatory framework for an important category in the nuclear industry—defense nuclear facilities owned and operated by the US Department of Energy. In parallel with describing this evolution, three major events which occurred external to the complex of defense nuclear facilities will be summarized, and their impact on the maturation of the Department's safety and regulatory framework will be discussed. Finally, integrated with these two threads of discussion will be a chronicle of the changing set of involved organizations and the expanding set of external stakeholders involved in risk decisions—and therefore, the risk communications ecosystem surrounding defense nuclear facilities. It will be noted that this system was once describable as a classic “iron triangle,” but now has progressed to a complex network of federal and state organizations, numerous congressional committees, and expanding sets of external stakeholders. It is hoped that a comprehensive discussion of the context of risk assessment in the defense nuclear facilities complex—addressing historical insights, organizational evolution, and the maturing structure of regulation—will provide enhanced opportunities for building trust and understanding in this complex environment.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Domain decomposition in the GPU-accelerated Shift Monte Carlo code

The GPU solver within the Shift continuous-energy Monte Carlo neutron transport code has been extended to provide domain decomposition in addition to domain replication to enable the solution of problems with memory requirements exceeding the capacity of a single GPU. The strategy follows the Multiple Set, Overlapping Domain (MSOD) approach that is used in Shift’s CPU solver and integrates into the event-based algorithm used for Shift’s GPU solver. Furthermore, the ability to assign processors to spatial domains non-uniformly has been maintained. In this work, two different approaches for communicating particle data between domains are considered, and multiple criteria for load balancing problems have been investigated. Numerical results are presented for both fresh and depleted small modular nuclear reactor (SMR) cores. A parallel efficiency of approximately 80% was achieved with up to 16 spatial domains measured relative to full domain replication. A scaling study on the Summit supercomputer demonstrates a weak scaling parallel efficiency of over 90% on over 24000 GPUs.

97 MATHEMATICS AND COMPUTING↗

Synthesis of (oxo)chlorin dimers chelated with thallium(III)

Two target dimers have been prepared for fundamental studies of hole/electron transfer. Metalation with thallium(III) enables clocking of the rate of hole/electron transfer between the two macrocycles. Each dimer contains a diphenylethyne linker joining two identical hydroporphyrins (chlorin or oxochlorin). The linker is substituted at the 4,4′-positions whereas each (oxo)chlorin is joined at the 10-position. Each (oxo)chlorin is equipped with a gem-dimethyl group at the 18-position to stabilize the hydroporphyrin chromophore toward adventitious dehydrogenation and a 3,5-di-tert-butyl group at the 5-position to achieve increased solubilization in organic media. The dimers parallel a prior set of diphenylethyne-linked (oxo)chlorin constructs containing zinc-free base, zinc-zinc, and copper-copper metalation states that have been examined in studies of electronic communication. The building block (oxo)chlorins for preparing the thallium-containing dimers have been prepared in quantities of 32–404 mg, a scale up to 14-fold larger than previously. Thallation of the free base (oxo)chlorin dimers was achieved with excess TlCl 3 ⋅4H 2 O in CH 2 Cl 2 /CH 3 OH (3–4:1) upon overnight reaction at room temperature. The long-wavelength (Q[Formula: see text] absorption band of (oxo)chlorins lies between that of the zinc(II) and free base counterparts. Absorption spectral comparisons are provided of the thallium(III) and free base (oxo)chlorin monomers and dimers.

Chemistry↗

Device for controlling additive manufacturing machinery

A computing device for controlling the operation of an additive manufacturing machine comprises a memory element and a processing element. The memory element is configured to store a three-dimensional model of a part to be manufactured, wherein the three-dimensional model defines a plurality of cross sections of the part. The processing element is in communication with the memory element. The processing element is configured to receive the three-dimensional model, determine a path across a surface of each cross section, wherein the path includes a plurality of parallel lines, calculate a power for a radiation beam to scan each of the lines, such that the power varies from line to line non-linearly according to a length of the line, and calculate a scan speed for the radiation beam for each of the lines, such that the scan speed varies line to line non-linearly according to the power of the radiation beam.

Barr, Christian↗

Enhancing scalability of a matrix-free eigensolver for studying many-body localization

We propose several techniques to enhance the parallel scalability of a matrix-free eigensolver designed for studying many-body localization (MBL) of quantum spin chain models with nearest-neighbor interactions and on-site disorder. This type of problem is computationally challenging because the dimension of the associated Hamiltonian matrix grows exponentially with respect to the number of spins L, and we need to average over different realizations of the random disorder to obtain relevant statistical behavior. For each disorder realization, we need to compute eigenvalues from different regions of the spectrum and their corresponding eigenvectors. In previous work, the interior eigenstates for a single eigenvalue problem are computed via the shift-and-invert Lanczos algorithm. Due to the extremely high memory footprint of the LU factorizations, this technique is not well suited for large L’s. For example, we need thousands of compute nodes on modern high performance computing infrastructures to go beyond L = 24. The matrix-free approach does not suffer from this memory bottleneck, however, its scalability is limited by a computation and communication load imbalance. To reduce this imbalance and to significantly enhance the scalability of the matrix-free eigensolver, we reorder the matrix and leverage the consistent space runtime, CSPACER. We also show its efficiency in managing irregular communication patterns at scale compared to optimized MPI non-blocking two-sided and one-sided RMA implementation variants. This effort enables us to study MBL for spin chains with a larger number of spins. The efficiency and effectiveness of the proposed algorithm is demonstrated by computing eigenstates on a massively parallel many-core high performance computer.

METIS↗

PoliMOR

PoliMOR is a scalable, automated, and customizable policy engine framework for multi-tiered parallel file systems. It is composed of single-purpose agents that handle tasks such as gathering file metadata, making policy decisions, and then executing actions based on those policies. These agents are designed to communicate using distributed message queues, allowing the number of individual agents to be scaled up as needed. PoliMOR automates the data management tasks by precluding the need for admin intervention. The agents in PoliMOR can be customized to integrate any utilities/tools that perform tasks like metadata scanning and data placement management.

Brumgard, Christopher↗

GSoFa: Scalable Sparse Symbolic LU Factorization on GPUs

Decomposing a matrix $\mathbf {A}$ into a lower matrix $\mathbf {L}$ and an upper matrix $\mathbf {U}$, which is also known as LU decomposition, is an essential operation in numerical linear algebra. For a sparse matrix, LU decomposition often introduces more nonzero entries in the $\mathbf {L}$ and $\mathbf {U}$ factors than in the original matrix. A symbolic factorization step is needed to identify the nonzero structures of $\mathbf {L}$ and $\mathbf {U}$ matrices. Attracted by the enormous potentials of the Graphics Processing Units (GPUs), an array of efforts have surged to deploy various LU factorization steps except for the symbolic factorization, to the best of our knowledge, on GPUs. This article introduces gSoFa, the first GPU-based symbolic factorization design with the following three optimizations to enable scalable LU symbolic factorization for nonsymmetric pattern sparse matrices on GPUs. First, here we introduce a novel fine-grained parallel symbolic factorization algorithm that is well suited for the Single Instruction Multiple Thread (SIMT) architecture of GPUs. Second, we tailor supernode detection into a SIMT friendly process and strive to balance the workload, minimize the communication and saturate the GPU computing resources during supernode detection. Third, we introduce a three-pronged optimization to reduce the excessive space consumption problem faced by multi-source concurrent symbolic factorization. Taken together, gSoFa achieves up to 31× speedup from 1 to 44 Summit nodes (6 to 264 GPUs) and outperforms the state-of-the-art CPU project, on average, by 5×. Notably, gSoFa also achieves up to 47 percent of the peak memory throughput of a V100 GPU in the Summit Supercomputer.

97 MATHEMATICS AND COMPUTING↗

OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model

With the sharp increasing volume of user data, Deep Learning Recommendation Model (DLRM) becomes an indispensable infrastructure in large technology companies. However, large-scale DLRM on the multi-GPU platform is still inefficient due to unbalanced workload partitioning and intensive inter-GPU communication. To this end, we propose OPER, an OPtimality guided Embedding table placement for large-scale Recommendation model training and inference. OPER explores the potential of mitigating remote memory access latency in DLRM through fine-grained embedding table placement. Specifically, OPER proposes a theoretical modeling that builds up the relationship between EMT placement and the embedding communication latency in both training and inference. OPER proves the NP hardness of finding the optimal embedding table placement and proposes a heuristic algorithm that yields near optimal placement. OPER implements a SHMEM-based embedding table training system and a unified embedding index mapping to support fine-grained embedding table sharding and placement. Comprehensive experiments reveal that OPER achieves on average 3.4× and 5.1× speedup on training and inference respectively over state-of-the-art DLRM frameworks.

Wang, Zheng↗

PythonFOAM: In-situ data analyses with OpenFOAM and Python

Here, we outline the development of a general-purpose Python-based data analysis tool for OpenFOAM. Our implementation relies on the construction of OpenFOAM applications that have bindings to data analysis libraries in Python. Double precision data in OpenFOAM is cast to a NumPy array using the NumPy C-API and Python modules may then be used for arbitrary data analysis and manipulation on flow-field information. We highlight how the proposed wrapper may be used for an in-situ online singular value decomposition (SVD) implemented in Python and accessed from the OpenFOAM solver PimpleFOAM. Here, 'in-situ' refers to a programming paradigm that allows for a concurrent computation of the data analysis on the same computational resources utilized for the partial differential equation solver. In addition, to demonstrate parallel deployments, we deploy a distributed SVD, which collects snapshot data across the ranks of a distributed simulation to compute the global left singular vectors. Crucially, both OpenFOAM and Python share the same message passing interface (MPI) communicator for this deployment which allows Python objects and functions to exchange NumPy arrays across ranks. Subsequently, we provide scaling assessments of this distributed SVD on multiple nodes of Intel Broadwell and KNL architectures for canonical test cases such as the large eddy simulations of a backward facing step and a channel flow at friction Reynolds number of 395. Finally, we demonstrate the deployment of a deep neural network for compressing the flow-field information using an autoencoder to demonstrate an ability to use state-of-the-art machine learning tools in the Python ecosystem.

97 MATHEMATICS AND COMPUTING↗

Computable and Operationally Meaningful Multipartite Entanglement Measures

Multipartite entanglement is an essential resource for quantum communication, quantum computing, quantum sensing, and quantum networks. The utility of a quantum state |ψ$\rangle$ for these applications is often directly related to the degree or type of entanglement present in |ψ$\rangle$. Therefore, efficiently quantifying and characterizing multipartite entanglement is of paramount importance. Here, in this work, we introduce a family of multipartite entanglement measures, called concentratable entanglements. Several well-known entanglement measures are recovered as special cases of our family of measures, and hence we provide a general framework for quantifying multipartite entanglement. We prove that the entire family does not increase, on average, under local operations and classical communications. We also provide an operational meaning for these measures in terms of probabilistic concentration of entanglement into Bell pairs. Finally, we show that these quantities can be efficiently estimated on a quantum computer by implementing a parallelized SWAP test, opening up a research direction for measuring multipartite entanglement on quantum devices.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Control And Optimization Modular Modeling Application For Nuclear Deployment

The purpose of the COMMAND code is to provide a flexible, scalable tool for use in developing, integrating, and testing the technologies necessary for achieving autonomous operations of advanced nuclear reactors. The code enables users to efficiently implement custom simulations and experiments by combining key methods from different software modules. These modules are focused on: modeling and simulation tools, such as nuclear simulation tools used for high-fidelity modeling (e.g., Reactor Excursion and Leak Analysis Program [RELAP5-3D] and Monte Carlo N-Particle [MCNP]); machine learning and optimization tools (e.g., anomaly detection and data-driven modeling techniques); advanced control in its digital, high-performance, and supervisory control forms (e.g., proportional integral derivative (PID) control and model predictive control (MPC); and integration with hardware through industrial communication protocols. To ensure flexibility and scalability, COMMAND was designed to be both modular—the software “pieces” all inherit from generic building blocks and can be combined and connected to create complicated simulations—and high performing—designed for parallel processing, enabling simulations and experiments to take advantage of multi-core computers, servers, and nodes. The code is written in the Python programming language due to the language's popularity, active community, and open-source and cross-platform nature. Maintaining consistency with other simulation tools used within the nuclear energy community, users implement simulations and experiments through text input files, which define components, parameters, connections, etc., through lines of text. Given that COMMAND is written in Python, these input files are native Python scripts, and so use the standard Python structure and formatting. This also enables users to take advantage of Python's extensive package library to develop custom capabilities for their specific use cases.

Faber, Jacob [Idaho National Laboratory (INL), Ida↗

Exaflops Biomedical Knowledge Graph Analytics

We are motivated by newly proposed methods for mining large-scale corpora of scholarly publications (e.g., full biomedical literature), which consists of tens of millions of papers spanning decades of research. In this setting, analysts seek to discover relationships among concepts. They construct graph representations from annotated text databases and then formulate the relationship-mining problem as an all-pairs shortest paths (APSP) and validate connective paths against curated biomedical knowledge graphs (e.g., Spoke). In this context, we present Coast (Exascale Communication-Optimized All-Pairs Shortest Path) and demonstrate 1.004 EF/s on 9,200 Frontier nodes (73,600 GCDs). We develop hyperbolic performance models (HYPERMOD), which guide optimizations and parametric tuning. The proposed Coast algorithm achieved the memory constant parallel efficiency of 99% in the single-precision tropical semiring. Looking forward, Coast will enable the integration of scholarly corpora like PubMed into the Spoke biomedical knowledge graph.

Kannan, Ramakrishnan {ramki}↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Parameter Optimization Toolbox for NS-3 network optimization, NS-3 Parameter Optimization Framework [SWR-18-60]

This simulation-based parameter optimization framework is proposed to tune parameters of different types of communication networks using ns-3 to achieve the optimal network performance. It consists of three main components: an ns-3 packet reporting module; a sampler running simulations with all possible parameter sets for the input parameter variables by using a parallel executor at each generation; and a hybrid optimization algorithm for tuning configurable parameters of hybrid designs and application parameter variables. The proposed hybrid metaheuristic optimization algorithm combines an evolutionary algorithm with a gradient descent function to quickly achieve an approximate globally optimum solution. This software is designed to be used in a multi-core processing Linux environment and run over a long duration of time. The execution time varies depending mainly upon the nature of the ns-3 configuration being simulated. This software includes a custom ns-3 QoS measurement application which must be included with the ns-3 source code during installation of the software.

Hasandka, Adarsh↗

torch-einshard v1.0

torch-einshard is a Python library for describing local and distributed PyTorch tensor computations with compact, einsum-like notation. Its expressions name logical axes, specify how they are sharded across a PyTorch DeviceMesh, and represent partial reductions. The library automatically performs contractions, permutations, reshaping, splitting, gathering, reduction, reduce-scatter, and repartitioning while preserving autograd. Additional features include sharding-aware FFTs, tensor rolls, halo exchange, sliding windows, 1D–3D convolutions, uneven-shard handling, parameter initialization and gradient management, and cost-based execution planning. It is designed for scientific machine learning and large-model workloads, including tensor-, sequence-, and spatial-parallel MLPs, attention, convolutions, and spectral operations. Compared with manually combining torch.einsum and distributed collectives, torch-einshard expresses both the mathematical operation and data placement in one readable formula. This reduces boilerplate and synchronization errors, keeps forward and backward communication consistent, and allows the library to select optimized collective strategies without changing model code.

Morozov, Dmitriy [Lawrence Berkeley National Labor↗

Distributed fiber sensing systems for 3D combustion temperature field monitoring in coal-fired boilers using optically generated acoustic waves (Final Report)

In this project, we have developed and tested three kinds of fiber optic sensing systems for real time monitoring of temperature variations within an industrial scale boiler furnace. The fiber optic sensing systems target spatial and temporal distributions of high temperature profiles in a boiler furnace in fossil power plants. The reconstructed temperature profile will provide critical input for the control mechanisms to optimize the combustion process. This temperature profile will address the essential problem for fossil power plants in achieving higher efficiency and fewer pollutant emissions. Acoustic pyrometer systems have been used to reconstruct temperature field of power plant boilers based on measuring TOF (times-of-flight) of sound waves along some straight paths in a 2D cross-section of the boiler. In this project, optically generated acoustic signals from a fiber optic sensing system have replaced the acoustic signals generated from an electrical transducer. A 3D reconstruction algorithm replaced the previous 2D model. In this project, three kinds of fiber optic sensing systems have been developed and tested. They are fiber optic sensing system I, fiber optic sensing system II (Distributed Sensing System I) and fiber optic sensing system III (Distributed Sensing System II). For fiber optic sensing system I, the fiber optic ultrasound generator acts as a signal generator. A microphone, hydrophone or other electronic devices serve as a signal receiver. In this system, there are one generator and one receiver. Distance test, water temperature test, air temperature test, air temperature reconstruction, and GE ISBF pilot test were performed by Fiber optic sensing system I. The fiber optic sensing system I successfully detected temperature in all these tests. We got 2D temperature reconstruction results by using the fiber optic sensing system I and it matched the reference data. The fiber optic sensing system I successfully survived in GE ISBF boiler environment (480 °F). For fiber optic sensing system II (Distributed Sensing System I), it is an all optical ultrasound system. The fiber optic ultrasound generator acts as a signal generator. Fiber Bragg Grating (FBG) and Fabry-Perot (FP) sensor act as a signal receiver. In this system, there is one generator and one receiver. Aluminum plate temperature test, furnace high temperature test, and GE ISBF pilot test were performed by the fiber optic sensing system II. Fiber optic sensing system II successfully detected the temperature in all these tests. The fiber optic sensing system II successfully survived at up to 700 °C furnace environment and 320 °C GE ISBF boiler environment. For fiber optic sensing system III (Distributed Sensing System II), it is also an all optical ultrasound system. The fiber optic ultrasound generator acts as a signal generator. Multiple FP fiber sensors act as signal receivers. In this system, there is one generator and three receivers. Three GE ISBF pilot tests were performed by fiber optic sensing system III. The fiber optic sensing system III survived in the cold flow tests in GE’s ISBF pilot test facility. However, we didn’t get high temperature data by using this system since the nanosecond laser issues. During the period of the project, test trials, data simulation and algorithm optimization was performed successfully. For real time temperature field construction, the sampling rate must be fast enough to capture the field variations. The technology of Code-division multiple access (CDMA) is well studied which could allow parallel multiplexing, even if signals overlap in time or frequencies. Moreover, it has been known that extending the length of signal significantly improves SNR. For acoustic signals, these multiplexing techniques have also been widely used, mainly for sonar and acoustic communications. The CDMA modulation technique has been proposed and studied to guarantee high network throughput, low channel access delay and low energy consumption. We have studied the temperature field reconstruction using Gaussian Radial Basis Functions (GRBF)-based approximation approach. Reconstruction of 3D temperature field using Neural Networks with measured TOF and known propagation paths is feasible. 2D and 3D temperature field reconstruction simulation results are achieved. The milestone status is shown in Table 1. We finished milestone 1-8 and milestone 10. For milestone 9, we did three pilot tests by using the fiber optic sensing system III (Distributed Sensing System II) at GE Power. However, due to the failure of the ns laser, we did not get the temperature results. We conducted some additional tasks that were not originally proposed: 1) We fabricated a fiber optic sensing system I and did a pilot test based on this system. 2) In the proposal, we proposed two pilot tests at GE Power. In reality, we finished at least seven pilot tests at GE Power. GE Power has made a lot of efforts for supporting the pilot tests. 3) We got a simulation results based on CDMA. In summary, most of the tasks have been accomplished. The outcome of this project removed a few barriers that hinder the achievement of the final product of the distributed sensing systems. With the successful accomplishment of this project, a prototype of the fiber optic sensing system can be fabricated to attract more interests from companies and other funding agencies.

47 OTHER INSTRUMENTATION↗