Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Improving I/O Performance for Exascale Applications through Online Data Layout Reorganization

The applications being developed within the U.S. Exascale Computing Project (ECP) to run on imminent Exascale computers will generate scientific results with unprecedented fidelity and record turn-around time. Many of these codes are based on particle-mesh methods and use advanced algorithms, especially dynamic load-balancing and mesh-refinement, to achieve high performance on Exascale machines. Yet, as such algorithms improve parallel application efficiency, they raise new challenges for I/O logic due to their irregular and dynamic data distributions. Thus, while the enormous data rates of Exascale simulations already challenge existing file system write strategies, the need for efficient read and processing of generated data introduces additional constraints on the data layout strategies that can be used when writing data to secondary storage. We review these I/O challenges and introduce two online data layout reorganization approaches for achieving good tradeoffs between read and write performance. We demonstrate the benefits of using these two approaches for the ECP particle-in-cell simulation WarpX, which serves as a motif for a large class of important Exascale applications. Here, we show that by understanding application I/O patterns and carefully designing data layouts we can increase read performance by more than 80 percent.

97 MATHEMATICS AND COMPUTING↗

Automatic Differentiation of C++ Codes on Emerging Manycore Architectures with Sacado

Automatic differentiation (AD) is a well-known technique for evaluating analytic derivatives of calculations implemented on a computer, with numerous software tools available for incorporating AD technology into complex applications. However, a growing challenge for AD is the efficient differentiation of parallel computations implemented on emerging manycore computing architectures such as multicore CPUs, GPUs, and accelerators as these devices become more pervasive. In this work, we explore forward mode, operator overloading-based differentiation of C++ codes on these architectures using the widely available Sacado AD software package. In particular, we leverage Kokkos, a C++ tool providing APIs for implementing parallel computations that is portable to a wide variety of emerging architectures. Here we describe the challenges that arise when differentiating code for these architectures using Kokkos, and two approaches for overcoming them that ensure optimal memory access patterns as well as expose additional dimensions of fine-grained parallelism in the derivative calculation. We describe the results of several computational experiments that demonstrate the performance of the approach on a few contemporary CPU and GPU architectures. We then conclude with applications of these techniques to the simulation of discretized systems of partial differential equations.

97 MATHEMATICS AND COMPUTING↗

HYPPO v1.0.0

HYPPO is a hyperparameter optimization (HPO) software that uses surrogate modeling and uncertainty quantification to provide reliable neural network architectures. Similar state-of-the-art HPO technologies do not use such features and mostly perform random searches to explore the parameter space, which can easily become computationally expensive. Using HYPPO, the user is able to reduce by an order of magnitude the number of evaluations necessary to identify the most optimal region in the hyperparameter space. Finally, using asynchronous nested parallelism, we are able to decrease by two orders of magnitude the amount of time required to complete the HPO process.

Dumont, Vincent↗

Optimizing multigrid reduction-in-time and Parareal coarse-grid operators for linear advection

Parallel-in-time methods, such as multigrid reduction-in-time (MGRIT) and Parareal, provide an attractive option for increasing concurrency when simulating time-dependent partial differential equations (PDEs) in modern high-performance computing environments. While these techniques have been very successful for parabolic equations, it has often been observed that their performance suffers dramatically when applied to advection-dominated problems or purely hyperbolic PDEs using standard rediscretization approaches on coarse grids. In this paper, we apply MGRIT or Parareal to the constant-coefficient linear advection equation, appealing to existing convergence theory to provide insight into the typically nonscalable or even divergent behavior of these solvers for this problem. To overcome these failings, we replace rediscretization on coarse grids with improved coarse-grid operators that are computed by applying optimization techniques to approximately minimize error estimates from the convergence theory. Therefore, one of our main findings is that, in order to obtain fast convergence as for parabolic problems, coarse-grid operators should take into account the behavior of the hyperbolic problem by tracking the characteristic curves. Our approach is tested for schemes of various orders using explicit or implicit Runge–Kutta methods combined with upwind-finite-difference spatial discretizations. In all cases, we obtain scalable convergence in just a handful of iterations, with parallel tests also showing significant speed-ups over sequential time-stepping.

97 MATHEMATICS AND COMPUTING↗

Parallel Multigrid in Time and Space for Extreme-Scale Computational Science: Chaotic and Hyperbolic Problems

The coming massive parallelism of exascale computing presents a pressing challenge for the many DOE simulations of time-dependent partial differential equations (PDEs), which typically use traditional sequential time stepping methods. Since this traditional approach is inherently serial, it presents a sequential bottleneck when moving to exascale computing, because future performance gains will come through greater concurrency, not faster clock speeds. Thus, the goal of this work is to research parallelism in time, i.e., methods that compute multiple time values simultaneously, not sequentially. The focus will be on hyperbolic and chaotic problems of interest to DOE, with the goal of enabling scalable simulations of time-dependent hyperbolic and chaotic problems on future architectures. The chosen methodology for solving these problems parallel-in-time is multigrid, because multigrid (when it works) is a powerful, optimal, and scalable solver for discretized PDEs. Multigrid is already commonly used in many DOE simulations for scalably and optimally solving space-only PDE problems. The areas of hyperbolic and chaotic problems are chosen because of their relevance to problems of programmatic interest to DOE. However, these problems are also well-known to be difficult for parallelin-time methods, with the most common method, parareal, diverging in many cases. The current stateof-the-art for parallel-in-time at LLNL is the multigrid reduction in time (MGRIT) XBraid package, which also struggles for such problems, while still showing some improvement over parareal. In summary, new methods are needed for an efficient parallel-in-time scheme for hyperbolic and chaotic problems, and this work shall research promising new multigrid methods in this area. In particular, we take inspiration from the Least Squares Shadowing (LSS by Wang) approach for solving chaotic problems. Here, an optimization approach is able to find “well-conditioned” shadow trajectories/solutions to the original “ill-conditioned” chaotic problem. Thus, the new multigrid methods researched here also arise in an optimization context.

97 MATHEMATICS AND COMPUTING↗

Real-time optimization of multi-cell industrial evaporative cooling towers using machine learning and particle swarm optimization

Existing electrical generating stations must operate with greater flexibility due to increasing renewable energy penetration on the electrical grid, and many coal-fired power stations have transitioned away from baseload operation to load-following operation to aid in grid stability. In cases where multiple independently controlled cooling tower cells are used in parallel for the cooling purposes of such stations, there is an opportunity to increase plant efficiency through data-driven optimization across their full load ranges. This work presents a novel application of real-time optimization using machine learning and particle swarm optimization on a multi-cell induced-draft cooling tower servicing a coal-fired power station under variable load. This is the first work to demonstrate simultaneous optimization of a multi-cell cooling tower, in addition to using machine learning for closed-loop control on a cooling tower. A novel control configuration is presented that ensures original control logic is not adversely affected and that the overall plant process is not disrupted using only existing hardware and operational data. To verify this methodology, the 12 independent cooling tower cells are simulated in parallel using historic operating data to demonstrate the effectiveness of real-time optimization compared to current practice. An artificial neural network is trained to predict overall cooling tower power consumption using only operational data and ambient conditions with an R2 value of greater than 0.96. The real-time optimization using particle swarm yields 6.7% annual energy usage savings compared to current practices, although the extent of the real-time savings varies greatly with both plant load and environmental conditions. This is particularly significant for a variable load situation because frequent ramping typically results in reduced overall efficiency. Furthermore, this proposed AI-based solution presents an opportunity to improve the overall heat rate of a load-following coal-fired power plant without the need to perform extensive first-principles modeling or add additional hardware to the cooling tower, resulting in more resources conserved and less overall emissions per unit of electricity generated.

42 ENGINEERING↗

Optimizing Grain Boundary Structures with LAMMPS Using Evolutionary Algorithms

Grain boundary structure optimization is an important part of materials modeling. Current methods for grain boundary structure optimization involve inefficient, time-consuming processes that do not fully explore the interface parameter space. Evolutionary algorithms have recently been demonstrated to be effective at determining both stable and metastable grain boundary interface structures. In this work, we demonstrate the use of GBOpt, a grain boundary structure optimization software designed to use the Large-scale Atomic/Molecular Massively Parallel Simulation (LAMMPS) software to efficiently determine grain boundary structures. We demonstrate that a only a few manipulations, namely atom insertion, atom removal, and relative grain displacement, are sufficient to explore much of the grain boundary structure parameter space. The efficacy of this approach is demonstrated on an FCC Ni system, and a BCC Fe system. The computational cost is compared against the gamma-surface sampling approach to demonstrate performance improvement.

Evolutionary algorithms↗

Non-Intrusive Parallel-in-Time Solvers for Partial Differential Equations (Final Report)

Many time-dependent problems and simulations are often modeled using Partial Differential Equations. Traditional modeling approaches that use sequential time-stepping are reaching a bottleneck in optimizing efficiency. The Center of Applied Science and Computing at Lawrence Livermore National Laboratory extensively works on parallelizing these algorithms to leverage the increasing computational power from the growing number of processors in computer hardware. In particular, they aim to design non-intrusive algorithms that can generalize to a variety of problems and sizes without requiring additional information from or modifications on the original problems. Multigrid Reduction in Time (MGRIT) is a parallel-in-time algorithm that is designed to be non-intrusive. This project focuses on increasing the efficiency of MGRIT by approximating the coarse-grid operator using machine learning approaches as a means to find the most non-intrusive, or general, solution.

97 MATHEMATICS AND COMPUTING↗

Characterization of a Pixelated Cadmium Telluride Detector System Using a Polychromatic X-Ray Source and Gold Nanoparticle-Loaded Phantoms for Benchtop X-Ray Fluorescence Imaging

In this paper, the imaging dose and scan time have been considered as the two major constraints for routine benchtop x-ray fluorescence computed tomography (XFCT) imaging. One way to address this issue is to acquire x-ray fluorescence (XRF) signals in parallel through a 2D array of single-crystal detectors or a pixelated detector along with the cone-beam x-ray source. To identify a detector system suitable for this purpose, a commercially available, fully spectroscopic cadmium telluride (CdTe) pixelated detector, HEXITEC (High-Energy X-ray Imaging Technology), was tested under the experimental conditions optimized for benchtop XFCT imaging of gold nanoparticles (GNPs). Specifically, two different parallel-hole stainless steel collimators were fabricated and coupled with the detector for seamless integration into our existing benchtop cone-beam XFCT system. After the detector deployment, this benchtop XFCT system was used to detect XRF photons from GNP-loaded phantoms. A pixel-merging algorithm was introduced to enhance the sensitivity of XRF photon detection thereby minimizing the scan time. The effect of pixel-level charge sharing correction algorithms was investigated within the context of benchtop XFCT imaging. The detector energy resolution, in terms of the full width at half maximum (FWHM) values at different gold K-shell XRF energies, was also determined. Of the two charge sharing correction algorithms examined, the charge sharing addition gave better sensitivity than the charge sharing discrimination (csd). On the other hand, under the current experimental conditions, the energy resolution of the HEXITEC detector was the best with the csd and estimated to be 1.56 keV FWHM at 66-69 keV photon energy. Overall, despite some degradation of the detector energy resolution (compared with typical single crystal CdTe detectors), the HEXITEC detector enabled parallel data acquisition under the experimental conditions typical of benchtop XFCT imaging and operated well within our benchtop XFCT setup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Implementation and optimization of the PTOLEMY transverse drift electromagnetic filter

The PTOLEMY transverse drift filter is a new concept to enable precision analysis of the energy spectrum of electrons near the tritium β-decay endpoint. Here, we detail the implementation and optimization methods for successful operation of the filter for electrons with a known pitch angle. We present the first demonstrator that produces the required magnetic field properties with an iron return-flux magnet. Two methods for the setting of filter electrode voltages are detailed. The challenges of low-energy electron transport in cases of low field are discussed, such as the growth of the cyclotron radius with decreasing magnetic field, which puts a ceiling on filter performance relative to fixed filter dimensions. Additionally, low pitch angle trajectories are dominated by motion parallel to the magnetic field lines and introduce non-adiabatic conditions and curvature drift. To minimize these effects and maximize electron acceptance into the filter, we present a three-potential-well design to simultaneously drain the parallel and transverse kinetic energies throughout the length of the filter. These optimizations are shown, in simulation, to achieve low-energy electron transport from a 1 T iron core (or 3 T superconducting) starting field with initial kinetic energy of 18.6 keV drained to < 10 eV (< 1 eV) in about 80 cm. This result for low field operation paves the way for the first demonstrator of the PTOLEMY spectrometer for measurement of electrons near the tritium endpoint to be constructed at the Gran Sasso National Laboratory (LNGS) in Italy.

47 OTHER INSTRUMENTATION↗

Multigrid Reduction in Time for Chaotic and Hyperbolic Problems (Final Report)

The coming massive parallelism of exascale computing presents a pressing challenge for the many DOE simulations of time-dependent partial differential equations (PDEs), which typically use traditional sequential time stepping methods. Since this traditional approach is inherently serial, it presents a sequential bottleneck when moving to exascale computing, because future performance gains will come through greater concurrency, not faster clock speeds. Thus, the goal of this work is to research parallelism in time, i.e., methods that compute multiple time values simultaneously, not sequentially. The focus will be on hyperbolic and chaotic problems of interest to DOE, with the goal of enabling scalable simulations of time-dependent hyperbolic and chaotic problems on future architectures. The chosen methodology for solving these problems parallel-in-time is multigrid, because multigrid (when it works) is a powerful, optimal, and scalable solver for discretized PDEs. Multigrid is already commonly used in many DOE simulations for scalably and optimally solving space-only PDE problems. The areas of hyperbolic and chaotic problems are chosen because of their relevance to problems of programmatic interest to DOE. However, these problems are also well-known to be difficult for parallelin-time methods, with the most common method, parareal, diverging in many cases. The current state-of-the-art for parallel-in-time at LLNL is the multigrid reduction in time (MGRIT) XBraid package, which also struggles for such problems, while still showing some improvement over parareal. In summary, new methods are needed for an efficient parallel-in-time scheme for hyperbolic and chaotic problems, and this work shall research promising new multigrid methods in this area. In particular, this work shall continue researching the directions from the current collaboration with Dr. Falgout, which are laid out in the work Toward Parallel in Time for Chaotic Dynamical Systems and showed the first known results of a parallel-in-time speedup for a chaotic problem. This work outlines two key improvements to XBraid for chaotic problems, the so-called “theta” and “delta-correction” methods. Here, these two improvements will be implemented in a high-performance but general way in XBraid and explored for more complicated problems. We will additionally research, as time allows, improvements to these techniques, as well as multigrid relaxation techniques based on Least Squares Shadowing (LSS by Wang) and a nonintrusive block tridiagonal solver based on MGRIT, called TriMGRIT.

97 MATHEMATICS AND COMPUTING↗

A New Configuration of Paralleled Modular ANPC Multilevel Converter Controlled by an Improved Modulation Method for 1 MHz, 1 MW EV Charger

In this work, a new configuration of the modular multilevel converter (MLC) based on the parallel connection of three-level active-neutral-point-clamped (3L-ANPC) cells as well as its improved modulation method is proposed for 1 MHz, 1 MW electric vehicle (EV) megacharger. In the proposed paralleled modular ANPC-MLC, only six high-frequency silicon carbide (SiC) power switches operating at 333 kHz are required to generate 1 MHz switching frequency spectrum. Moreover, the operating voltage of all power devices is halved, the magnitude of the first switching frequency harmonic cluster is decreased by the factor of five, and the load current is equally distributed between the 3L-ANPC legs by employing the proposed improved modulation method. Hence, the modularity, efficiency, and power density of the proposed converter are notably increased, whereas the value of passive components and the overall switching loss are remarkably decreased. In addition, an optimized design of the one 3L-ANPC cell of the proposed paralleled modular ANPC-MLC for 1 MHz, 1MW EV megacharger using Ansys SIwave, Icepak, and Q3D finite element method platforms is presented and analyzed in detail. The provided experimental results of the down-scaled setup verify the feasibility and viability of the proposed configuration as well as its improved switching pattern.

42 ENGINEERING↗

Parallel Algorithms for Computing the Tensor-Train Decomposition

The tensor-train (TT) decomposition expresses a tensor in a data-sparse format used in molecular simulations, high-order correlation functions, and optimization. In this paper, we propose four parallelizable algorithms that compute the TT format from various tensor inputs: (1) Parallel-TTSVD for traditional format, (2) PSTT and its variants for streaming data, (3) Tucker2TT for Tucker format, and (4) TT-fADI for solutions of Sylvester tensor equations. We provide theoretical guarantees of accuracy, parallelization methods, scaling analysis, and numerical results. For example, for a d-dimension tensor in $\mathbb{R}$ $n\times∙∙∙$$\times$$n$ a two-sided sketching algorithm PSTT2 is shown to have a memory complexity of $O(n^{[d/2]})$, improving upon $O(n^{d—1})$ from previous algorithms.

97 MATHEMATICS AND COMPUTING↗

Data-Driven Compositional Optimization in Misspecified Regimes

With a manifold growth in the scale and intricacy of systems, the challenges of parametric misspecification become pronounced. These concerns are further exacerbated in compositional settings, which emerge in problems complicated by modeling risk and robustness. In “Data-Driven Compositional Optimization in Misspecified Regimes,” the authors consider the resolution of compositional stochastic optimization problems, plagued by parametric misspecification. In considering settings where such misspecification may be resolved via a parallel learning process, the authors develop schemes that can contend with diverse forms of risk, dynamics, and nonconvexity. They provide asymptotic and rate guarantees for unaccelerated and accelerated schemes for convex, strongly convex, and nonconvex problems in a two-level regime with extensions to the multilevel setting. Surprisingly, the nonasymptotic rate guarantees show no degradation from the rate statements obtained in a correctly specified regime and the schemes achieve optimal (or near-optimal) sample complexities for general T-level strongly convex and nonconvex compositional problems.

Business & Economics↗

A Case Study of LLVM-Based Analysis for Optimizing SIMD Code Generation

This paper presents a methodology for using LLVM-based tools to tune the DCA++ (dynamical cluster approximation) application that targets the new ARM A64FX processor. The goal is to describe the changes required for the new architecture and generate efficient single instruction/multiple data (SIMD) instructions that target the new Scalable Vector Extension instruction set. During manual tuning, the authors used the LLVM tools to improve code parallelization by using OpenMP SIMD, refactored the code and applied transformation that enabled SIMD optimizations, and ensured that the correct libraries were used to achieve optimal performance. By applying these code changes, code speed was increased by 1.98× and 78 GFlops were achieved on the A64FX processor. The authors aim to automatize parts of the efforts in the OpenMP Advisor tool, which is built on top of existing and newly introduced LLVM tooling.

Huber, Joseph↗

The Persistent Challenge of Data Locality in the Post-Exascale Era

The era of exascale computing, exemplified by systems like Frontier achieving exaflop-level performance, marks a milestone. However, the quest for sheer compute power leads to strong imbalance in system design. Hence, scaling advancements in memory, network bandwidth, and storage are also necessary and pose challenges, with a crucial need to address data locality issues. This article underscores the fundamental importance of data locality as a key abstraction for optimizing application performance. Despite notable software solutions, the growing complexity of parallelism and memory hierarchy demands performance-portable data locality solutions across diverse computing platforms. Additionally, the article revisits data locality aspects, covering hardware considerations, application perspectives, software stack abstractions, and tool support. It concludes with insights into data locality challenges and opportunities, emphasizing the ongoing significance of collaborative research for progress in this critical issue.

Unat, Didem [Koc University, Istanbul (Turkey)] (O↗

Tough Errors are no Match (TEAM): Optimizing the Quantum Compiler for Noise Resilience

This project builds toward a comprehensive error-mitigating toolkit that makes quantum programming more robust and adaptive to the noisy, resource-limited nature of today’s quantum hardware. To that end, it integrates established error-mitigation methods — such as zero-noise extrapolation and dynamical decoupling — directly into compiler infrastructures. These techniques will be packaged as modules that can automatically adjust and combine based on performance analysis, enabling compilers to explore large design spaces and produce optimized, low-noise quantum programs with minimal manual intervention. In parallel, this project also explores new approaches to analog quantum programming or quantum simulation, and has developed the programming language SimuQ which treats quantum Hamiltonian evolution as the central object.

97 MATHEMATICS AND COMPUTING↗

BM3DORNL

BM3DORNL is a high-performance, open-source library for removing streak and ring artifacts from computed-tomography (CT) data, developed for neutron imaging at Oak Ridge National Laboratory's Spallation Neutron Source (VENUS beamline) and applicable to X-ray CT as well. Ring artifacts — concentric rings in reconstructed slices caused by detector pixel-to-pixel response non-uniformities — appear as vertical streaks in the sinogram and degrade both image quality and quantitative analysis. BM3DORNL operates in the sinogram domain using an adaptation of the BM3D (block-matching and 3D collaborative filtering) algorithm (Dabov et al., 2007). It provides a dedicated streak-removal mode, a true multi-scale BM3D variant (after Mäkinen et al., 2021) that suppresses wide streaks single-scale methods miss, and an alternative Fourier–SVD method (~2.6× faster) combining FFT-based energy detection with rank-1 SVD. The computationally intensive core is implemented in Rust with parallel (Rayon) block matching, integral-image pre-screening, and optimized transforms, and is exposed through a simple Python API (with an optional GUI) so it integrates directly into existing tomography reconstruction pipelines. It processes both 2D sinograms and 3D sinogram stacks, is pip-installable for Linux and macOS, and is documented at https://bm3dornl.readthedocs.io.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗