Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

GSoFa: Scalable Sparse Symbolic LU Factorization on GPUs

Decomposing a matrix $\mathbf {A}$ into a lower matrix $\mathbf {L}$ and an upper matrix $\mathbf {U}$, which is also known as LU decomposition, is an essential operation in numerical linear algebra. For a sparse matrix, LU decomposition often introduces more nonzero entries in the $\mathbf {L}$ and $\mathbf {U}$ factors than in the original matrix. A symbolic factorization step is needed to identify the nonzero structures of $\mathbf {L}$ and $\mathbf {U}$ matrices. Attracted by the enormous potentials of the Graphics Processing Units (GPUs), an array of efforts have surged to deploy various LU factorization steps except for the symbolic factorization, to the best of our knowledge, on GPUs. This article introduces gSoFa, the first GPU-based symbolic factorization design with the following three optimizations to enable scalable LU symbolic factorization for nonsymmetric pattern sparse matrices on GPUs. First, here we introduce a novel fine-grained parallel symbolic factorization algorithm that is well suited for the Single Instruction Multiple Thread (SIMT) architecture of GPUs. Second, we tailor supernode detection into a SIMT friendly process and strive to balance the workload, minimize the communication and saturate the GPU computing resources during supernode detection. Third, we introduce a three-pronged optimization to reduce the excessive space consumption problem faced by multi-source concurrent symbolic factorization. Taken together, gSoFa achieves up to 31× speedup from 1 to 44 Summit nodes (6 to 264 GPUs) and outperforms the state-of-the-art CPU project, on average, by 5×. Notably, gSoFa also achieves up to 47 percent of the peak memory throughput of a V100 GPU in the Summit Supercomputer.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

Dispersion-enhanced sequential batch sampling for adaptive contour estimation

In computer simulation and optimal design, sequential batch sampling offers an appealing way to iteratively stipulate optimal sampling points based upon existing selections and efficiently construct surrogate modeling. Nonetheless, the issue of near duplicates poses tremendous quandary for sequential learning. It refers to the situation that selected critical points cluster together in each sampling batch, which are individually but not collectively informative towards the optimal design. Near duplicates severely diminish the computational efficiency as they barely contribute extra information towards update of the surrogate. To address this issue, we impose a dispersion criterion on concurrent selection of sampling points, which essentially forces a sparse distribution of critical points in each batch, and demonstrate the effectiveness of this approach in adaptive contour estimation. Specifically, we adopt Gaussian process surrogate to emulate the simulator, acquire variance reduction of the critical region from new sampling points as a dispersion criterion, and combine it with the modified expected improvement (EI) function for critical batch selection. The critical region here is the proximity of the contour of interest. This proposed approach is vindicated in numerical examples of a two-dimensional four-branch function, a four-dimensional function with a disjoint contour of interest and a time-delay dynamic system.

97 MATHEMATICS AND COMPUTING↗

SUNDIALS time integrators for exascale applications with many independent systems of ordinary differential equations

Many complex systems can be accurately modeled as a set of coupled time-dependent partial differential equations (PDEs). However, solving such equations can be prohibitively expensive, easily taxing the world’s largest supercomputers. One pragmatic strategy for attacking such problems is to split the PDEs into components that can more easily be solved in isolation. This operator splitting approach is used ubiquitously across scientific domains, and in many cases leads to a set of ordinary differential equations (ODEs) that need to be solved as part of a larger “outer-loop” time-stepping approach. The SUNDIALS library provides a plethora of robust time integration algorithms for solving ODEs, and the U.S. Department of Energy Exascale Computing Project (ECP) has supported its extension to applications on exascale-capable computing hardware. In this paper, we highlight some SUNDIALS capabilities and its deployment in combustion and cosmology application codes (Pele and Nyx, respectively) where operator splitting gives rise to numerous, small ODE systems that must be solved concurrently.

97 MATHEMATICS AND COMPUTING↗

Stable parallel training of Wasserstein conditional generative adversarial neural networks

In this work, we propose a stable, parallel approach to train Wasserstein conditional generative adversarial neural networks (W-CGANs) under the constraint of a fixed computational budget. Differently from previous distributed GANs training techniques, our approach avoids inter-process communications, reduces the risk of mode collapse and enhances scalability by using multiple generators, each one of them concurrently trained on a single data label. The use of the Wasserstein metric also reduces the risk of cycling by stabilizing the training of each generator. We illustrate the approach on the CIFAR10, CIFAR100, and ImageNet1k datasets, three standard benchmark image datasets, maintaining the original resolution of the images for each dataset. Performance is assessed in terms of scalability and final accuracy within a limited fixed computational time and computational resources. To measure accuracy, we use the inception score, the Fréchet inception distance, and image quality. An improvement in inception score and Fréchet inception distance is shown in comparison to previous results obtained by performing the parallel approach on deep convolutional conditional generative adversarial neural networks as well as an improvement of image quality of the new images created by the GANs approach. Weak scaling is attained on both datasets using up to 2000 NVIDIA V100 GPUs on the OLCF supercomputer Summit.

97 MATHEMATICS AND COMPUTING↗

Bayesian Conavigation: Dynamic Designing of the Material Digital Twins via Active Learning

Scientific advancement is universally based on the dynamic interplay between theoretical insights, modeling, and experimental discoveries. However, this feedback loop is often slow, including delayed community interactions and the gradual integration of experimental data into theoretical frameworks. This challenge is particularly exacerbated in domains dealing with high-dimensional object spaces, such as molecules and complex microstructures. Hence, the integration of theory within automated and autonomous experimental setups, or theory in the loop-automated experiment, is emerging as a crucial objective for accelerating scientific research. The critical aspect is to use not only theory but also on-the-fly theory updates during the experiment. Furthermore, we introduce a method for integrating theory into the loop through Bayesian conavigation of theoretical model space and experimentation. Our approach leverages the concurrent development of surrogate models for both simulation and experimental domains at the rates determined by latencies and costs of experiments and computation, alongside the adjustment of control parameters within theoretical models to minimize epistemic uncertainty over the experimental object spaces. This methodology facilitates the creation of digital twins of material structures, encompassing both the surrogate model of behavior that includes the correlative part and the theoretical model itself. While being demonstrated here within the context of functional responses in ferroelectric materials, our approach holds promise for broader applications, such as the exploration of optical properties in nanoclusters, microstructure-dependent properties in complex materials, and properties of molecular systems.

Microscopy↗

HPC Resource Allocation Under Energy Constraints

We discuss the new problem faced by High-Performance Computing (HPC) facilities in allocating resources to users of their facilities: while facilities once allocated a single finite resource—node-hours—now facilities must also concurrently allocate a second scarce resource: electrical energy, which is bounded within each facility's annual operations budget. Current application optimization practices encourage conservation of the first resource, but can be potentially unaffordably wasteful of the second. We describe a framework for reasoning about such allocations that can be utilized by facilities to articulate policy, while encouraging scientific application developers to write code mindfully of both constraints. We outline the requirements on facilities, on developers, and on hardware vendors and integrators that are necessary to enable the implementation of this framework.

97 MATHEMATICS AND COMPUTING↗

Multi-reward reinforcement learning based development of inter-atomic potential models for silica

Abstract Silica is an abundant and technologically attractive material. Due to the structural complexities of silica polymorphs coupled with subtle differences in Si–O bonding characteristics, the development of accurate models to predict the structure, energetics and properties of silica polymorphs remain challenging. Current models for silica range from computationally efficient Buckingham formalisms (BKS, CHIK, Soules) to reactive (ReaxFF) and more recent machine-learned potentials that are flexible but computationally costly. Here, we introduce an improved formalism and parameterization of BKS model via a multireward reinforcement learning (RL) using an experimental training dataset. Our model concurrently captures the structure, energetics, density, equation of state, and elastic constants of quartz (equilibrium) as well as 20 other metastable silica polymorphs. We also assess its ability in capturing amorphous properties and highlight the limitations of the BKS-type functional forms in simultaneously capturing crystal and amorphous properties. We demonstrate ways to improve model flexibility and introduce a flexible formalism, machine-learned ML-BKS, that outperforms existing empirical models and is on-par with the recently developed 50 to 100 times more expensive Gaussian approximation potential (GAP) in capturing the experimental structure and properties of silica polymorphs and amorphous silica.

36 MATERIALS SCIENCE↗

Evolution of microstructures in radiation fields using a coupled binary-collision Monte Carlo phase field approach

The simulation of radiation effects in materials broadly falls into two categories. At short time and length scales lies the modeling of primary radiation damage, such as point defect creation, energy deposition, and ballistic mixing. This is followed by the modeling at longer time scales of thermally activated microstructure evolution and defect reactions, such as recombination, clustering, and coarsening. The binary collision Monte Carlo method is an established, numerically efficient method for the computation of primary radiation damage. Conversely, the phase field method is a state of the art method for microstructure evolution on longer time and length scales. Here we present a concurrent coupling of these two methods, overcoming the difference between the discrete object Monte Carlo paradigm for primary radiation damage and the continuum field variable approach for microstructure evolution. The coupling is bidirectional, in which the microstructure evolution in the MOOSE finite element frame- work provides the spatial scattering data set for the charged particle trans- port and receives point defect, mass transport, and heat source terms from the simulated collision cascades that contribute to the field variable evolution. The concurrent coupling scheme is implemented in the code Magpie and demonstrated by investigating patterning for a model irradiated immiscible binary alloy. The results from the coupled binary collision Monte Carlo/phase field simulations reproduce the results of analytical models for phase separation, phase mixing, and patterning, supporting the approach and indicating its utility for modeling real materials systems.

36 MATERIALS SCIENCE↗

Structural and electronic changes in L⁢i 2 ⁢Ru⁢O 3 induced by lithium intercalation

Despite extensive research on oxide battery cathodes that transcend classical cationic redox activity, the detailed interplay between structural transformations and electronic redox processes remains insufficiently understood. We report a detailed study of the sequential structural and electronic changes in Li 2 RuO 3 upon lithium intercalation, characterized by powder x-ray and neutron diffraction alongside Ru and O K-edge x-ray absorption spectroscopy (XAS), and guided by operando synchrotron x-ray diffraction. During delithiation, Li 2 RuO 3 evolves from a well-defined monoclinic state to a complex trigonal phase via multiple intermediate structures, marked by significant changes in Ru-O bond distances that closely track the transition from a classical cationic redox to an unconventional process centered at oxygen states. Armed with high-quality atomic structural descriptions, computational models of the O K-edge XAS closely reproduce the experimentally observed spectral shifts. Lastly, we relate observations of electrochemical hysteresis with concurrent changes in the pathways of structural and electronic transitions. In conclusion, our results not only clarify the mechanisms underpinning voltage hysteresis in a model for lattice oxygen redox but also underscore the importance of structural fidelity in modeling redox behavior when this type of complex reactivity is present.

Li, Haifeng [Univ. of Illinois, Chicago, IL (Unite↗

Optimal Power Management for Large-Scale Battery Energy Storage Systems via Bayesian Inference

Large-scale battery energy storage systems (BESS) have found ever-increasing use across industry and society to accelerate clean energy transition and improve energy supply reliability and resilience. However, their optimal power management poses significant challenges: the underlying high-dimensional nonlinear nonconvex optimization lacks computational tractability in real-world implementation, and the uncertainty of the exogenous power demand makes exact optimization difficult. This paper presents a new solution framework to address these bottlenecks. The solution pivots on introducing power-sharing ratios to specify each cell’s power quota from the output power demand. To find the optimal power-sharing ratios, we formulate a nonlinear model predictive control (NMPC) problem to achieve power-loss-minimizing BESS operation while complying with safety, cell balancing, and power supply-demand constraints. We then propose a parameterized control policy for the power-sharing ratios, which utilizes only three parameters, to reduce the computational demand in solving the NMPC problem. This policy parameterization allows us to translate the NMPC problem into a Bayesian inference problem for the sake of 1) computational tractability, and 2) overcoming the nonconvexity of the optimization problem. We leverage the ensemble Kalman inversion technique to solve the parameter estimation problem. Concurrently, a low-level control loop is developed to seamlessly integrate our proposed approach with the BESS to ensure practical implementation. This low-level controller receives the optimal power-sharing ratios, generates output power references for the cells, and maintains a balance between power supply and demand despite uncertainty in output power. We conduct extensive simulations and experiments on a 20-cell prototype to validate the proposed approach.

Battery energy storage systems (BESSs)↗

QuComm: Optimizing Collective Communication for Distributed Quantum Computing

Distributed quantum computing (DQC) is a scalable way to build a large-scale quantum computing system while the error-prone nonlocal communication between DQC nodes may heavily degrade the fidelity of the distributed quantum program and thus demands specific compiler optimizations. Previous compilers on DQC communication optimization either assumes unlimited communication resource or a few communication qubits due to the hardware limitation. The former compilers may not be efficient when interfacing with communication-resource-constrained DQC hardware while the latter compilers lose the opportunities of optimizing collective communication and routing concurrent communication as they unnecessarily couple limited communication qubits with the implementation of expensive inter-node operations. In this paper, we invent the communication buffer, a communication facility consisting of idle qubits in each compute node, to decouple the execution of inter-node quantum operations from communication qubits: communication qubits are devoted to generating inter-node entanglement while internode operations are conducted in the communication buffer. The communication buffer provides an intermediate layer for inter-node communication and paves the way for collective communication optimization. We then propose QuComm, a buffer-based compiler framework that first performs smart buffer allocation according to communication characteristics of the distributed quantum program and then optimizes and collectively routes inter-node quantum operations. Experimental results on a hierarchical DQC system show that the proposed QuComm can reduce the most expensive inter-node communication request and the latency of various distributed quantum programs by 50.4% and 47.6% on average, respectively.

Wu, Anbang↗

Towards reverse mode automatic differentiation of Kokkos-based codes

Derivative computation is a key component of optimization, sensitivity analysis, uncertainty quantification, and the solving of nonlinear problems. Automatic differentiation (AD) is a powerful technique for evaluating such derivatives, and in recent years, has been integrated into programming environments such as Jax, PyTorch, and TensorFlow to support derivative computations needed for training of machine learning models, facilitating wide-spread use of these technologies. The C++ language has become the de facto standard for scientific computing due to numerous factors, yet language complexity has made the wide-spread adoption of AD technologies for C++ difficult, hampering the incorporation of powerful differentiable programming approaches into C++ scientific simulations. This is exacerbated by the increasing emergence of architectures, such as GPUs, with limited memory capabilities and requiring massive thread-level concurrency. C++ AD tools must effectively use these environments to bring novel scientific simulations to next-generation DOE experimental and observational facilities. In this project, we investigated source transformation-based automatic differentiation using LLVM compiler infrastructure to automatically generate portable and efficient gradient computations of Kokkos-based code. We have demonstrated that our proposed strategy is feasible by investigating the usage of a prototype LLVM-based source transformation tool to generate gradients of simple functions made of sequences of simple Kokkos parallel regions. Speedups of up to 500x compared to Sacado were observed on NVIDIA V100 GPU.

97 MATHEMATICS AND COMPUTING↗

Component level modeling of materials degradation for insights into operational flexibility of Existing Coal Power Plants

Increasingly, coal-fired power plants are required to balance power grids by compensating for the variable electricity supply from renewable energy sources. Fossil-fueled power plants, originally designed to be base loaded, will increasingly need to operate on a load following or cyclic basis. This demanding requirement for operational flexibility needs insights into accelerated material degradation arising due to the harsh operating conditions (e.g., fatigue, early oxide exfoliation due to stresses) along with current damage mechanisms (fireside corrosion, creep and erosion) observed in service. Our research objective is to develop component level modeling toolkit for materials-based degradation for two key mechanisms that can accelerate with cyclic operations. In more detail, this includes the fireside corrosion/steam oxidation/erosion/creep/fatigue of superheaters/reheaters and steam pipework and also the water droplet erosion/ fatigue of last stage steam turbine blades degradation mechanisms, that demand routine and sometimes unplanned maintenance and repair. The innovation is in developing a computational fluid dynamics/finite element (CFD/FE) modeling toolkit for the component level models of the boilers and low-pressure steam turbines in coal power plants that can tackle multidisciplinary failure mechanisms occurring concurrently for extreme environment materials. Lifetime assessment in such environments also needs to account for the unit-specific analyses, operational history and fuel feedstock; this can only be obtained by destructive analysis of components. This, in turn, enables validation of the model toolkits utilizing service feedback data, improving the probability of time/temperature dependent life prediction.

20 FOSSIL-FUELED POWER PLANTS↗

Probabilistic Discrete‐Time Models for Spreading Processes in Complex Networks: A Review

Abstract Research into network dynamics of spreading processes typically employs both discrete and continuous time methodologies. Although each approach offers distinct insights, integrating them can be challenging, particularly when maintaining coherence across different time scales. This review focuses on the Microscopic Markov Chain Approach (MMCA), a probabilistic f ramework originally designed for epidemic modeling. MMCA uses discrete dynamics to compute the probabilities of individuals transitioning between epidemiological states. By treating each time step—usually a day—as a discrete event, the approach captures multiple concurrent changes within this time frame. The approach allows to estimate the likelihood of individuals or populations being in specific states, which correspond to distinct epidemiological compartments. This review synthesizes key findings from the application of this approach, providing a comprehensive overview of its utility in understanding epidemic spread.

Granell, Clara↗

Understanding the Interplay between Hardware Errors and User Job Characteristics on the Titan Supercomputer

Designing dependable supercomputers begins with an understanding of errors in real-world, large-scale systems. The Titan supercomputer at Oak Ridge National Laboratory provides a unique opportunity to investigate errors when an actual system is actively used by multiple concurrent users and workloads from diverse domains at varying scales. This study presents a thorough analysis of 6, 908, 497 hardware errors from 18, 688 compute nodes of Titan for 312, 215 user jobs over a 3-year time period. Through careful joining of two system logs – the Machine Check Architecture (MCA) log and the job scheduler log – we show the correlated pattern of hardware errors for each job and user, in addition to individual descriptive statistics of errors, jobs, and users. Since the majority of hardware errors are memory errors, this study also shows the importance of error correcting in memory systems.

Lim, Seung-Hwan↗