Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Characterization of a Pixelated Cadmium Telluride Detector System Using a Polychromatic X-Ray Source and Gold Nanoparticle-Loaded Phantoms for Benchtop X-Ray Fluorescence Imaging

In this paper, the imaging dose and scan time have been considered as the two major constraints for routine benchtop x-ray fluorescence computed tomography (XFCT) imaging. One way to address this issue is to acquire x-ray fluorescence (XRF) signals in parallel through a 2D array of single-crystal detectors or a pixelated detector along with the cone-beam x-ray source. To identify a detector system suitable for this purpose, a commercially available, fully spectroscopic cadmium telluride (CdTe) pixelated detector, HEXITEC (High-Energy X-ray Imaging Technology), was tested under the experimental conditions optimized for benchtop XFCT imaging of gold nanoparticles (GNPs). Specifically, two different parallel-hole stainless steel collimators were fabricated and coupled with the detector for seamless integration into our existing benchtop cone-beam XFCT system. After the detector deployment, this benchtop XFCT system was used to detect XRF photons from GNP-loaded phantoms. A pixel-merging algorithm was introduced to enhance the sensitivity of XRF photon detection thereby minimizing the scan time. The effect of pixel-level charge sharing correction algorithms was investigated within the context of benchtop XFCT imaging. The detector energy resolution, in terms of the full width at half maximum (FWHM) values at different gold K-shell XRF energies, was also determined. Of the two charge sharing correction algorithms examined, the charge sharing addition gave better sensitivity than the charge sharing discrimination (csd). On the other hand, under the current experimental conditions, the energy resolution of the HEXITEC detector was the best with the csd and estimated to be 1.56 keV FWHM at 66-69 keV photon energy. Overall, despite some degradation of the detector energy resolution (compared with typical single crystal CdTe detectors), the HEXITEC detector enabled parallel data acquisition under the experimental conditions typical of benchtop XFCT imaging and operated well within our benchtop XFCT setup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Implementation and optimization of the PTOLEMY transverse drift electromagnetic filter

The PTOLEMY transverse drift filter is a new concept to enable precision analysis of the energy spectrum of electrons near the tritium β-decay endpoint. Here, we detail the implementation and optimization methods for successful operation of the filter for electrons with a known pitch angle. We present the first demonstrator that produces the required magnetic field properties with an iron return-flux magnet. Two methods for the setting of filter electrode voltages are detailed. The challenges of low-energy electron transport in cases of low field are discussed, such as the growth of the cyclotron radius with decreasing magnetic field, which puts a ceiling on filter performance relative to fixed filter dimensions. Additionally, low pitch angle trajectories are dominated by motion parallel to the magnetic field lines and introduce non-adiabatic conditions and curvature drift. To minimize these effects and maximize electron acceptance into the filter, we present a three-potential-well design to simultaneously drain the parallel and transverse kinetic energies throughout the length of the filter. These optimizations are shown, in simulation, to achieve low-energy electron transport from a 1 T iron core (or 3 T superconducting) starting field with initial kinetic energy of 18.6 keV drained to < 10 eV (< 1 eV) in about 80 cm. This result for low field operation paves the way for the first demonstrator of the PTOLEMY spectrometer for measurement of electrons near the tritium endpoint to be constructed at the Gran Sasso National Laboratory (LNGS) in Italy.

47 OTHER INSTRUMENTATION↗

Multigrid Reduction in Time for Chaotic and Hyperbolic Problems (Final Report)

The coming massive parallelism of exascale computing presents a pressing challenge for the many DOE simulations of time-dependent partial differential equations (PDEs), which typically use traditional sequential time stepping methods. Since this traditional approach is inherently serial, it presents a sequential bottleneck when moving to exascale computing, because future performance gains will come through greater concurrency, not faster clock speeds. Thus, the goal of this work is to research parallelism in time, i.e., methods that compute multiple time values simultaneously, not sequentially. The focus will be on hyperbolic and chaotic problems of interest to DOE, with the goal of enabling scalable simulations of time-dependent hyperbolic and chaotic problems on future architectures. The chosen methodology for solving these problems parallel-in-time is multigrid, because multigrid (when it works) is a powerful, optimal, and scalable solver for discretized PDEs. Multigrid is already commonly used in many DOE simulations for scalably and optimally solving space-only PDE problems. The areas of hyperbolic and chaotic problems are chosen because of their relevance to problems of programmatic interest to DOE. However, these problems are also well-known to be difficult for parallelin-time methods, with the most common method, parareal, diverging in many cases. The current state-of-the-art for parallel-in-time at LLNL is the multigrid reduction in time (MGRIT) XBraid package, which also struggles for such problems, while still showing some improvement over parareal. In summary, new methods are needed for an efficient parallel-in-time scheme for hyperbolic and chaotic problems, and this work shall research promising new multigrid methods in this area. In particular, this work shall continue researching the directions from the current collaboration with Dr. Falgout, which are laid out in the work Toward Parallel in Time for Chaotic Dynamical Systems and showed the first known results of a parallel-in-time speedup for a chaotic problem. This work outlines two key improvements to XBraid for chaotic problems, the so-called “theta” and “delta-correction” methods. Here, these two improvements will be implemented in a high-performance but general way in XBraid and explored for more complicated problems. We will additionally research, as time allows, improvements to these techniques, as well as multigrid relaxation techniques based on Least Squares Shadowing (LSS by Wang) and a nonintrusive block tridiagonal solver based on MGRIT, called TriMGRIT.

97 MATHEMATICS AND COMPUTING↗

A New Configuration of Paralleled Modular ANPC Multilevel Converter Controlled by an Improved Modulation Method for 1 MHz, 1 MW EV Charger

In this work, a new configuration of the modular multilevel converter (MLC) based on the parallel connection of three-level active-neutral-point-clamped (3L-ANPC) cells as well as its improved modulation method is proposed for 1 MHz, 1 MW electric vehicle (EV) megacharger. In the proposed paralleled modular ANPC-MLC, only six high-frequency silicon carbide (SiC) power switches operating at 333 kHz are required to generate 1 MHz switching frequency spectrum. Moreover, the operating voltage of all power devices is halved, the magnitude of the first switching frequency harmonic cluster is decreased by the factor of five, and the load current is equally distributed between the 3L-ANPC legs by employing the proposed improved modulation method. Hence, the modularity, efficiency, and power density of the proposed converter are notably increased, whereas the value of passive components and the overall switching loss are remarkably decreased. In addition, an optimized design of the one 3L-ANPC cell of the proposed paralleled modular ANPC-MLC for 1 MHz, 1MW EV megacharger using Ansys SIwave, Icepak, and Q3D finite element method platforms is presented and analyzed in detail. The provided experimental results of the down-scaled setup verify the feasibility and viability of the proposed configuration as well as its improved switching pattern.

42 ENGINEERING↗

On the Development of an Efficient Parallel Hybrid Solver with Application to Acoustically Treated Aero-Engine Nacelles

A finite element solution to the convected Helmholtz equation in a nonuniform flow is used to model the noise field within 3-D acoustically treated aero-engine nacelles. Options to select linear or cubic Hermite polynomial basis functions and isoparametric elements are included. However, the key feature of the method is a domain decomposition procedure that is based upon the inter-mixing of an iterative and a direct solve strategy for solving the discrete finite element equations. This procedure is optimized to take full advantage of sparsity and exploit the increased memory and parallel processing capability of modern computer architectures. Example computations are presented for the Langley Flow Impedance Test facility and a rectangular mapping of a full scale, generic aero-engine nacelle. The accuracy and parallel performance of this new solver are tested on both model problems using a supercomputer that contains hundreds of central processing units. Results show that the method gives extremely accurate attenuation predictions, achieves super-linear speedup over hundreds of CPUs, and solves upward of 25 million complex equations in a quarter of an hour.

Watson, Willie R.↗

Parallel Algorithms for Computing the Tensor-Train Decomposition

The tensor-train (TT) decomposition expresses a tensor in a data-sparse format used in molecular simulations, high-order correlation functions, and optimization. In this paper, we propose four parallelizable algorithms that compute the TT format from various tensor inputs: (1) Parallel-TTSVD for traditional format, (2) PSTT and its variants for streaming data, (3) Tucker2TT for Tucker format, and (4) TT-fADI for solutions of Sylvester tensor equations. We provide theoretical guarantees of accuracy, parallelization methods, scaling analysis, and numerical results. For example, for a d-dimension tensor in $\mathbb{R}$ $n\times∙∙∙$$\times$$n$ a two-sided sketching algorithm PSTT2 is shown to have a memory complexity of $O(n^{[d/2]})$, improving upon $O(n^{d—1})$ from previous algorithms.

97 MATHEMATICS AND COMPUTING↗

Data-Driven Compositional Optimization in Misspecified Regimes

With a manifold growth in the scale and intricacy of systems, the challenges of parametric misspecification become pronounced. These concerns are further exacerbated in compositional settings, which emerge in problems complicated by modeling risk and robustness. In “Data-Driven Compositional Optimization in Misspecified Regimes,” the authors consider the resolution of compositional stochastic optimization problems, plagued by parametric misspecification. In considering settings where such misspecification may be resolved via a parallel learning process, the authors develop schemes that can contend with diverse forms of risk, dynamics, and nonconvexity. They provide asymptotic and rate guarantees for unaccelerated and accelerated schemes for convex, strongly convex, and nonconvex problems in a two-level regime with extensions to the multilevel setting. Surprisingly, the nonasymptotic rate guarantees show no degradation from the rate statements obtained in a correctly specified regime and the schemes achieve optimal (or near-optimal) sample complexities for general T-level strongly convex and nonconvex compositional problems.

Business & Economics↗

A Case Study of LLVM-Based Analysis for Optimizing SIMD Code Generation

This paper presents a methodology for using LLVM-based tools to tune the DCA++ (dynamical cluster approximation) application that targets the new ARM A64FX processor. The goal is to describe the changes required for the new architecture and generate efficient single instruction/multiple data (SIMD) instructions that target the new Scalable Vector Extension instruction set. During manual tuning, the authors used the LLVM tools to improve code parallelization by using OpenMP SIMD, refactored the code and applied transformation that enabled SIMD optimizations, and ensured that the correct libraries were used to achieve optimal performance. By applying these code changes, code speed was increased by 1.98× and 78 GFlops were achieved on the A64FX processor. The authors aim to automatize parts of the efforts in the OpenMP Advisor tool, which is built on top of existing and newly introduced LLVM tooling.

Huber, Joseph↗

Parallel algorithms for mapping pipelined and parallel computations

Many computational problems in image processing, signal processing, and scientific computing are naturally structured for either pipelined or parallel computation. When mapping such problems onto a parallel architecture it is often necessary to aggregate an obvious problem decomposition. Even in this context the general mapping problem is known to be computationally intractable, but recent advances have been made in identifying classes of problems and architectures for which optimal solutions can be found in polynomial time. Among these, the mapping of pipelined or parallel computations onto linear array, shared memory, and host-satellite systems figures prominently. This paper extends that work first by showing how to improve existing serial mapping algorithms. These improvements have significantly lower time and space complexities: in one case a published O(nm sup 3) time algorithm for mapping m modules onto n processors is reduced to an O(nm log m) time complexity, and its space requirements reduced from O(nm sup 2) to O(m). Run time complexity is further reduced with parallel mapping algorithms based on these improvements, which run on the architecture for which they create the mappings.

Nicol, David M.↗

Multiphase complete exchange on Paragon, SP2 and CS-2

The overhead of interprocessor communication is a major factor in limiting the performance of parallel computer systems. The complete exchange is the severest communication pattern in that it requires each processor to send a distinct message to every other processor. This pattern is at the heart of many important parallel applications. On hypercubes, multiphase complete exchange has been developed and shown to provide optimal performance over varying message sizes. Most commercial multicomputer systems do not have a hypercube interconnect. However, they use special purpose hardware and dedicated communication processors to achieve very high performance communication and can be made to emulate the hypercube quite well. Multiphase complete exchange has been implemented on three contemporary parallel architectures: the Intel Paragon, IBM SP2 and Meiko CS-2. The essential features of these machines are described and their basic interprocessor communication overheads are discussed. The performance of multiphase complete exchange is evaluated on each machine. It is shown that the theoretical ideas developed for hypercubes are also applicable in practice to these machines and that multiphase complete exchange can lead to major savings in execution time over traditional solutions.

Bokhari, Shahid H.↗

The Persistent Challenge of Data Locality in the Post-Exascale Era

The era of exascale computing, exemplified by systems like Frontier achieving exaflop-level performance, marks a milestone. However, the quest for sheer compute power leads to strong imbalance in system design. Hence, scaling advancements in memory, network bandwidth, and storage are also necessary and pose challenges, with a crucial need to address data locality issues. This article underscores the fundamental importance of data locality as a key abstraction for optimizing application performance. Despite notable software solutions, the growing complexity of parallelism and memory hierarchy demands performance-portable data locality solutions across diverse computing platforms. Additionally, the article revisits data locality aspects, covering hardware considerations, application perspectives, software stack abstractions, and tool support. It concludes with insights into data locality challenges and opportunities, emphasizing the ongoing significance of collaborative research for progress in this critical issue.

Unat, Didem [Koc University, Istanbul (Turkey)] (O↗

Tough Errors are no Match (TEAM): Optimizing the Quantum Compiler for Noise Resilience

This project builds toward a comprehensive error-mitigating toolkit that makes quantum programming more robust and adaptive to the noisy, resource-limited nature of today’s quantum hardware. To that end, it integrates established error-mitigation methods — such as zero-noise extrapolation and dynamical decoupling — directly into compiler infrastructures. These techniques will be packaged as modules that can automatically adjust and combine based on performance analysis, enabling compilers to explore large design spaces and produce optimized, low-noise quantum programs with minimal manual intervention. In parallel, this project also explores new approaches to analog quantum programming or quantum simulation, and has developed the programming language SimuQ which treats quantum Hamiltonian evolution as the central object.

97 MATHEMATICS AND COMPUTING↗

BM3DORNL

BM3DORNL is a high-performance, open-source library for removing streak and ring artifacts from computed-tomography (CT) data, developed for neutron imaging at Oak Ridge National Laboratory's Spallation Neutron Source (VENUS beamline) and applicable to X-ray CT as well. Ring artifacts — concentric rings in reconstructed slices caused by detector pixel-to-pixel response non-uniformities — appear as vertical streaks in the sinogram and degrade both image quality and quantitative analysis. BM3DORNL operates in the sinogram domain using an adaptation of the BM3D (block-matching and 3D collaborative filtering) algorithm (Dabov et al., 2007). It provides a dedicated streak-removal mode, a true multi-scale BM3D variant (after Mäkinen et al., 2021) that suppresses wide streaks single-scale methods miss, and an alternative Fourier–SVD method (~2.6× faster) combining FFT-based energy detection with rank-1 SVD. The computationally intensive core is implemented in Rust with parallel (Rayon) block matching, integral-image pre-screening, and optimized transforms, and is exposed through a simple Python API (with an optional GUI) so it integrates directly into existing tomography reconstruction pipelines. It processes both 2D sinograms and 3D sinogram stacks, is pip-installable for Linux and macOS, and is documented at https://bm3dornl.readthedocs.io.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Method of constructing dished ion thruster grids to provide hole array spacing compensation

The center-to-center spacings of a photoresist pattern for an array of holes applied to a thin metal sheet are increased by uniformly stretching the thin metal sheet in all directions along the plane of the sheet. The uniform stretching is provided by securely clamping the periphery of the sheet and applying an annular force against the face of the sheet, within the periphery of the sheet and around the photoresist pattern. The technique is used in the construction of ion thruster grid units where the outer or downstream grid is subjected to uniform stretching prior to convex molding. The technique provides alignment of the holes of grid pairs so as to direct the ion beamlets in a direction parallel to the axis of the grid unit and thereby provide optimization of the available thrust.

Banks, B. A.↗

Regenerative fuel cell energy storage system for a low earth orbit space station

A study was conducted to define characteristics of a Regenerative Fuel Cell System (RFCS) for low earth orbit Space Station missions. The RFCS's were defined and characterized based on both an alkaline electrolyte fuel cell integrated with an alkaline electrolyte water electrolyzer and an alkaline electrolyte fuel cell integrated with an acid solid polymer electrolyte (SPE) water electrolyzer. The study defined the operating characteristics of the systems including system weight, volume, and efficiency. A maintenance philosophy was defined and the implications of system reliability requirements and modularization were determined. Finally, an Engineering Model System was defined and a program to develop and demonstrate the EMS and pacing technology items that should be developed in parallel with the EMS were identified. The specific weight of an optimized RFCS operating at 140 F was defined as a function of system efficiency for a range of module sizes. An EMS operating at a nominal temperature of 180 F and capable of delivery of 10 kW at an overall efficiency of 55.4 percent is described. A program to develop the EMS is described including a technology development effort for pacing technology items.

Martin, R. E.↗

Single-agent parallel window search

Parallel window search is applied to single-agent problems by having different processes simultaneously perform iterations of Iterative-Deepening-A(asterisk) (IDA-asterisk) on the same problem but with different cost thresholds. This approach is limited by the time to perform the goal iteration. To overcome this disadvantage, the authors consider node ordering. They discuss how global node ordering by minimum h among nodes with equal f = g + h values can reduce the time complexity of serial IDA-asterisk by reducing the time to perform the iterations prior to the goal iteration. Finally, the two ideas of parallel window search and node ordering are combined to eliminate the weaknesses of each approach while retaining the strengths. The resulting approach, called simply parallel window search, can be used to find a near-optimal solution quickly, improve the solution until it is optimal, and then finally guarantee optimality, depending on the amount of time available.

Powley, Curt↗

Selection of Lockheed Martin's Preferred TSTO Configurations for the Space Launch Initiative

Lockheed Martin is developing concepts for safe, affordable Two Stage to Orbit (TSTO) reusable launch vehicles as part of NASA s Space Launch Initiaiive. This paper discusses the options considered for the design of the TSTO, the impact of each of these options on the vehicle configuration, the criteria used for selection of preferred configurations, and the results of the selection process. More than twenty configurations were developed in detail in order to compare optioiis such as propellant choice, serial vs. parallel burn sequence, use of propellant crossfeed between stages, bimese or optimized stage designs, and high or low staging velocities. Each configuration was analyzed not only for performance and sizing, but also for cost and reliability. The study concluded that kerosene was the superior fuel for first stages, and that bimese vehicles were not attractive.

Hopkins, Joshua B.↗

Development of the Science Data System for the International Space Station Cold Atom Lab

Cold Atom Laboratory (CAL) is a facility that will enable scientists to study ultra-cold quantum gases in a microgravity environment on the International Space Station (ISS) beginning in 2016. The primary science data for each experiment consists of two images taken in quick succession. The first image is of the trapped cold atoms and the second image is of the background. The two images are subtracted to obtain optical density. These raw Level 0 atom and background images are processed into the Level 1 optical density data product, and then into the Level 2 data products: atom number, Magneto-Optical Trap (MOT) lifetime, magnetic chip-trap atom lifetime, and condensate fraction. These products can also be used as diagnostics of the instrument health. With experiments being conducted for 8 hours every day, the amount of data being generated poses many technical challenges, such as downlinking and managing the required data volume. A parallel processing design is described, implemented, and benchmarked. In addition to optimizing the data pipeline, accuracy and speed in producing the Level 1 and 2 data products is key. Algorithms for feature recognition are explored, facilitating image cropping and accurate atom number calculations.

bose einstein condensate↗