Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “FFT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Implementation of a Mesh refinement algorithm into the quasi-static PIC code QuickPIC

Plasma-based acceleration (PBA) has emerged as a promising candidate for the accelerator technology used to build a future linear collider and/or an advanced light source. In PBA, a trailing or witness particle beam is accelerated in the plasma wave wakefield (WF) created by a laser or particle beam driver. The WF is often nonlinear and involves the crossing of plasma particle trajectories in real space and thus particle-in-cell methods are used. The distance over which the drive beam evolves is several orders of magnitude larger than the wake wavelength. This large disparity in length scales is amenable to the quasi-static approach. Three-dimensional (3D), quasi-static (QS), particle-in-cell (PIC) codes, e.g., QuickPIC, have been shown to provide high fidelity simulation capability with 2-4 orders of magnitude speedup over 3D fully explicit PIC codes. In PBA, the witness beam needs to be matched to the focusing forces of the WF to reduce the emittance growth. In some linear collider designs, the matched spot size of the witness beam can be 2 to 3 orders of magnitude smaller than the spot size (and wavelength) of the wakefield. Such an additional disparity in length scales is ideal for mesh refinement where the WF within the witness beam is described on a finer mesh than the rest of the WF. A mesh refinement scheme is described that has been implemented into the 3D QS PIC code, QuickPIC. Very fine (high) resolution is used in a small spatial region that includes the witness beam and progressively coarser resolutions in the rest of the simulation domain. A fast multigrid Poisson solver has been implemented for the field solve on the refined meshes and a Fast Fourier Transform (FFT) based Poisson solver is used for the coarse mesh. The code has been parallelized with both MPI and OpenMP, and the parallel scalability has also been improved by using pipelining. A preliminary adaptive mesh refinement technique is described to optimize the computational time for simulations with an evolving witness beam size. Several test problems are used to verify that the mesh refinement algorithm provides accurate results. Additionally, the results are benchmarked against highly resolved simulations exhibiting near-azimuthal symmetry, performed using QPAD—a novel hybrid QS PIC code that uses a PIC description in the coordinates (r, ct – z) and a gridless description in the azimuthal angle, Φ.

Linear collider↗

Study of the interplay between lower-order and higher-order energetic strain-gradient effects in polycrystal plasticity

in this report strain-gradient (SG) plasticity refers to a class of non-local theories in which gradients of plastic slip determine the storage of geometrically necessary dislocations, introducing a length-scale dependence in the mechanical behavior of crystalline materials, which is otherwise lacking in local theories. In this work, we incorporate lower-order (LO) and higher-order energetic (HOE) strain-gradient effects into a crystal plasticity fast Fourier transform (FFT)-based formulation to investigate the interplay of the length scale that each strain-gradient term introduces at the microscale, and the mechanical properties that result at the macroscale. For an applicable range of length scales, we consider two systems: a 1-D two-phase face centered cubic (FCC) laminate and a 3-D FCC polycrystal, and two uniaxial deformation modes: monotonic tension and cyclic tension–compression. We show that increases in the individual LO and HOE length scales increase the hardening rate and strength of the material, respectively. When combined, the strong LO hardening is less pronounced than the effect alone due to the lowering of the gradients due to the HOE microstress. We demonstrate that the LO and HOE hardening manifest as “isotropic” (yield surface expansion) and “kinematic” (yield surface shift) effects, respectively, consistent with their theoretical origins. We show that in cyclic loading, the Bauschinger effect emerges in both local and non-local calculations and link its origins and severity to the behavior in the strain field, slip-system rates, and the HOE microforce.

36 MATERIALS SCIENCE↗

Fast GPU 3D diffeomorphic image registration

3D image registration is one of the most fundamental and computationally expensive operations in medical image analysis. Here, we present a mixed-precision, Gauss–Newton–Krylov solver for diffeomorphic registration of two images. Our work extends the publicly available CLAIRE library to GPU architectures. Despite the importance of image registration, only a few implementations of large deformation diffeomorphic registration packages support GPUs. Our contributions are new algorithms to significantly reduce the run time of the two main computational kernels in CLAIRE: calculation of derivatives and scattered-data interpolation. Additionally, we deploy (i) highly-optimized, mixed-precision GPU-kernels for the evaluation of scattered-data interpolation, (ii) replace Fast-Fourier-Transform (FFT)-based first-order derivatives with optimized 8th-order finite differences, and (iii) compare with state-of-the-art CPU and GPU implementations. As a highlight, we demonstrate that we can register clinical images in less than 6 s on a single NVIDIA Tesla V100. This amounts to over 20 speed-up over the current version of CLAIRE and over 30 speed-up over existing GPU implementations.

97 MATHEMATICS AND COMPUTING↗

A discrete integral transform for rapid spectral synthesis

Accurate synthetic spectra that rely on large Line-By-Line (LBL)-databases are used in a wide range of applications such as high temperature combustion, atmospheric re-entry, planetary surveillance and laboratory plasmas. Conventionally synthetic spectra are calculated by computing a lineshape for every spectral line in the database and adding those together, which may take multiple hours for large databases. In this paper we propose a new approach for spectral synthesis based on an integral transform: the synthetic spectrum is calculated as the integral over the product of a Voigt profile and a newly proposed three-dimensional “lineshape distribution function”, which is a function of spectral position and Gaussian- & Lorentzian width coordinates. A fast discrete version of this transform based on the Fast Fourier Transform (FFT) is proposed, which improves performance compared to the conventional approach by several orders of magnitude while maintaining accuracy. Strategies that minimize the discretization error are discussed. A Python implementation of the method is compared against state-of-the-art spectral code RADIS, and is since adopted as RADIS's default synthesis method. The synthesis of a benchmark CO2 spectrum consisting of 1.8 M spectral lines and 200k spectral points took only 3.1 s using the proposed method (1011 lines × spectral points/s), a factor ~300 improvement over the state-of-the-art, with the relative improvement generally increasing for higher number of lines and/or number of spectral points. Finally, an experimental GPU-implementation of the method was also benchmarked, which demonstrated another 2~3 orders performance increase, achieving up to 5 ∙ 10 14 lines × spectral points/s.

42 ENGINEERING↗

A systematic study and framework of fringe projection profilometry with improved measurement performance for in-situ LPBF process monitoring

Fringe Projection Profilometry (FPP) is a cost-effective and non-invasive technology that has been shown to measure finer features. Here, in this work, we developed an in-situ FPP method to measure the dynamic topography of powder bed and printed layer during Laser Powder Bed Fusion (LPBF) additive manufacturing (AM) process. A systematic study towards developing a comprehensive framework of LPBF-specific FPP is demonstrated to enhance and evaluate the performance of applying FPP for in-situ LPBF monitoring, including 1) a modified sensor model with localized correction; 2) improved phase unwrapping with FFT filtering 3) quantitative uncertainty analysis; and 4) experimental validation with ex-situ characterization. The developed LPBF-specific FPP system and methods are implemented on a commercial LPBF-AM machine, achieving better accuracy, more robustness, and increased field of view while maintaining sufficient measurement range and decent resolution, in contrast to literature methods. The established FPP framework will facilitate the development of closed-loop control strategies for advancing LPBF based AM.

42 ENGINEERING↗

The effects of free surfaces on deformation twinning in HCP metals

Deformation twinning is a predominant mode of plastic deformation for hexagonal close packed metals, like Mg and Ti. The heterogenous microstructure and the local stresses associated with twinning play a key role in their mechanical response and fracture. Surface analyses, like electron microscopy, are frequently employed to spatially map microstructure and micromechanical fields in order to study twinning behavior. However, these measurements are inherently influenced by the vicinity of the free surface. In this work, an elasto-visco-plastic fast-Fourier-transform (EVP-FFT) polycrystal modeling approach is employed to investigate the effects of free surfaces on twin development before and after loading. We compare calculated micromechanical fields on free surfaces with those calculated inside the bulk and, in some cases, experimental surface measurements. The results indicate that the creation of free surfaces can promote twin propagation and growth and can influence twin morphology by causing a twin lamella to become larger, more blunted and irregular. The structure along the twin boundaries are also affected, due to the higher driving stresses that extend prismatic-basal and basal-prismatic facets. Furthermore, free surfaces invoke different slip activities in the twin and the surrounding parent crystal by enhancing basal, prismatic and pyramidal slip in some localized regions, while reducing slip in others. We demonstrate that the simulated free-surface effects lead to better qualitative and quantitative agreement with experimental measurements from scanning electron microscopy and digital image correlation.

36 MATERIALS SCIENCE↗

Twinning pathways enabled by precipitates in $\mathrm{AZ91}$

While precipitates have been shown to substantially strengthen magnesium alloys by blocking the glide of dislocations inside the grain, the interactions between these precipitates and deformation twins, which commonly occur in these alloys, are much less understood. Here, in this work, an elasto-viscoplastic fast-Fourier-transform (EVP-FFT) model is used to study the interactions between plate-shaped basal precipitates and propagating $10\bar{1}2$ tensile twins in AZ91 Mg alloy. The results suggest that while precipitates may impede the propagation and thickening of twins, they can also cause stress localizations that can promote the formation of multiple new twins of the same or different crystallographic variants. We show that the location of the twin-precipitate interaction site, whether precipitate-central or precipitate-edge impingement, and the thickness of the precipitate can influence the propensity for twins to expand around the precipitate or nucleate a new twin on the other side of it. Depending on the twin-precipitate impingement site, we propose multiple twinning pathways that can help explain how twins can proliferate in the magnesium alloys in the presence of precipitates.

36 MATERIALS SCIENCE↗

Topological behavior and Zeeman splitting in trigonal PtBi2-x single crystals

Abstract Transition-metal dipnictide PtBi 2 exhibits rich structural and physical properties with topological semimetallic behavior and extremely large magnetoresistance (XMR) at low temperatures. We have investigated the electrical and magnetic properties of trigonal-phase PtBi 2 -x single crystals with x ~ 0.4. Profound de Haas–van Alphen (dHvA) and Shubnikov-de Haas (SdH) oscillations are observed. Through fast Fourier transformation (FFT) analyses, four oscillation frequencies are extracted, which result from α, β, γ, and δ bands. By constructing the Landau fan diagram for each band, the Berry phase is extracted demonstrating the non-trivial nature of the α, β, and δ bands. Despite Bi deficiency, we observe the Zeeman splitting in dHvA and SdH oscillations under moderate magnetic field and the moderate Landé g factor (4.97–6.48) for the α band. Quantitative analysis of the non-monotonic field dependence including the sign change of the Hall resistivity suggests that electrons and holes in our system are not perfectly compensated thus not responsible for the XMR effect.

36 MATERIALS SCIENCE↗

Single-shot electrical detection of short-wavelength magnon pulse transmission in a magnonic thin-film waveguide

Abstract The advance of magnon spintronics requires understanding of time-domain magnon pulse transmission in order to develop high-speed information processing protocols. In this work, we demonstrate single-shot electrical detection of narrow-band magnon pulse transmission in a yttrium iron garnet thin-film delay line. The high signal-to-background ratio of magnon transmission band allows us to directly probe the magnon transmission electrically using a fast oscilloscope and to study its spectral evolution using Fast Fourier Transform (FFT) of the time-domain transmitted signal. At elevated input power, we show a magnon transmission reduction and a spectral distortion, which can be understood by the nonlinear magnon excitation in the transmission band defined by the antenna geometry. In addition, we also find that the higher- (lower-) frequency magnon spectral component exhibits a lower (higher) magnon group velocity, showing a dispersion agreeing with the Damon-Eshbach dependence. Our results provide important guidance of magnon pulse engineering for their applications in spin wave computing and coherent magnon information processing.

Song, Moojune↗

ConKer: An algorithm for evaluating correlations of arbitrary order

Context. High order correlations in the cosmic matter density have become increasingly valuable in cosmological analyses. However, computing these correlation functions is computationally expensive. Aims. We aim to circumvent these challenges by developing a new algorithm called ConKer for estimating correlation functions. Methods. This algorithm performs convolutions of matter distributions with spherical kernels using FFT. Since matter distributions and kernels are defined on a grid, it results in some loss of accuracy in the distance and angle definitions. We study the algorithm setting at which these limitations become critical and suggest ways to minimize them. Results. ConKer is applied to the CMASS sample of the SDSS DR12 galaxy survey and corresponding mock catalogs, and is used to compute the correlation functions up to correlation order n = 5. We compare the n = 2 and n = 3 cases to traditional algorithms to verify the accuracy of the new algorithm. We perform a timing study of the algorithm and find that three of the four distinct processes within the algorithm are nearly independent of the catalog size N , while one subdominant component scales as O ( N ). The dominant portion of the calculation has complexity of O ( N c 4/3 log N c ), where N c is the of cells in a three-dimensional grid corresponding to the matter density. Conclusions. We find ConKer to be a fast and accurate method of probing high order correlations in the cosmic matter density, then discuss its application to upcoming surveys of large-scale structure.

79 ASTRONOMY AND ASTROPHYSICS↗

Welch Method and Bootstrapping Applied to Subcritical Gamma Noise

We measured the prompt neutron decay constant 𝛼 of the CROCUS zero-power reactor at the Swiss Federal Institute of Technology Lausanne using cross-power spectral density (CPSD) analysis of gamma-gamma correlations from two trans-stilbene organic scintillators positioned near the reactor core. We measured critical and subcritical states, with water levels ranging from 960 mm (critical) to 800 mm (𝜌=−1.4 $ subcritical). Our analysis used the Welch method, dividing signal segments for fast Fourier transform (FFT) frequency analysis and applying bootstrapping uncertainty quantification that uses Welch-defined segments. Results demonstrated a clear increase in the measured 𝛼 as reactor reactivity decreased, distinguishing critical from subcritical conditions. At the 960-mm critical level, 𝛼 was estimated at 155.9 ± 0.7 s −1 , and for the 800-mm subcritical level, 𝛼 increased significantly to 367.3 ± 6.9 s –1 . A linear regression of subcritical states yielded a critical estimate of 154.0 ± 3.1 s –1 , aligning with the static 𝛼 estimate at critical. The bootstrapping method produced normally distributed 𝛼 estimates, confirming data consistency. The gamma CPSD 𝛼 estimates clearly distinguish reactor states and improve monitoring of zero-power reactors. The future deployment of modular and microreactors as potential candidates for noise analysis is demonstrated in CROCUS, particularly zero-power mock-ups of new designs. The improvement of noise analysis in the subcritical domain from this work will support experimental data for reactor deployment and procedure.

CROCUS↗

Mock data sets for the Eboss and DESI Lyman-α forest surveys

We present a publicly-available code to generate sets of mock Lyman-α (Lyα) forest data that have realistic large-scale correlations including those due to the Baryonic Acoustic Oscillations (BAO). The primary purpose of these mocks is to test the analysis procedures of the Extended Baryon Oscillation Survey (eBOSS) and the Dark Energy Spectroscopy Instrument (DESI) surveys. The transmitted flux fraction, F(λ), of background quasars due to Lyα absorption in the intergalactic medium (IGM) is simulated using the Fluctuating Gunn-Petterson Approximation (FGPA) applied to Gaussian random fields produced through the use of fast Fourier transforms (FFT). The output includes the IGM-Lyα transmitted flux fraction along quasar lines of sight and a catalog of high-column-density systems appropriately placed at high-density regions of the IGM. This output serves as input to additional code that superimposes the IGM tranmission on realistic quasar spectra, adds absorption by high-column-density systems and metals, and simulates instrumental transmission and noise. Redshift space distortions (RSD) of the flux correlations are implemented by including the large-scale velocity-gradient field in the FGPA resulting in a correlation function of F(λ) that can be accurately predicted. One hundred realizations have been produced over the 14,000 deg 2 DESI survey footprint with 100 quasars per deg 2 . The analysis of these realizations shows that the correlations of F(λ) follows the prediction within the accuracy of eBOSS survey. Here, the most time-consuming part of the mock production occurs before application of the FGPA, and the existing pre-FGPA forests can be used to easily produce new mock sets with modified redshift-dependent bias parameters or observational conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

DESI DR1 Lyα 1D power spectrum: the optimal estimator measurement

The one-dimensional power spectrum P 1D of Lyα forest offers rich insights into cosmological and astrophysical parameters, including constraints on the sum of neutrino masses, warm dark matter models, and the thermal state of the intergalactic medium. We present the measurement of P 1D using the optimal quadratic maximum likelihood estimator applied to over 300,000 Lyα quasars from Data Release 1 (DR1) of the Dark Energy Spectroscopic Instrument (DESI) survey. This sample represents the largest to date for P 1D measurements and is larger than the Extended Baryon Oscillation Spectroscopic Survey (eBOSS) by a factor of 1.7. We conduct a meticulous investigation of instrumental and analysis systematics and quantify their impact on P 1D . This includes the development of a cross-exposure estimator that eliminates the need to model the pipeline noise and has strong potential for future P 1D measurements. We also present new insights into metal contamination through the 1D correlation function. Using a fitting function we measure the evolution of the Lyα forest bias with high precision: b F (z) = (-0.218 ± 0.002) × ((1 + z)/4) 2.96±0.06 . In a companion validation paper, we substantially extend our previous suite of CCD image simulations to quantify the pipeline's exquisite performance accurately. In another companion paper, we present DR1 P 1D measurements using the Fast Fourier Transform (FFT) approach to power spectrum estimation. These two measurements produce a forest bias parameter that differs by 2.2 sigma. However, our model is simplistic, so this disagreement will be investigated in future work.

Lyman alpha forest↗

Waveform resampling with LMN method

In this article, resampling is a common technique applied in digital signal processing. Based on the Fast Fourier Transformation (FFT), we apply an optimization called here the LMN method to achieve fast and robust re-sampling. In addition to performance comparisons with some other popular methods, we illustrate the effectiveness of this LMN method in a particle physics experiment: re-sampling of waveforms from Liquid Argon Time Projection Chambers.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Finding Your Niche: An Evolutionary Approach to HPC Topologies

Traditional interconnection network design approaches focus on building general network topologies by optimizing the bisection bandwidth or minimizing the network’s diameter to reduce the maximum distance between any two nodes, thus amortizing the overall execution time of the HPC workloads. While such network topologies may accommodate a wide variety of applications in general, this may result in sub-optimal performance for many frequently-executed or dynamic workloads. In this paper, instead of focusing on designing an all-encompassing, general-purpose network topology, we develop a methodology to design customized network interconnects, evolved by “finding” the optimal topologies for a particular target workload given by its communication and contention profiles. To this end, we implement a Genetic Algorithm (GA)-based approach for network topology design tailored to improve the overall execution time of a particular workload of interest. We conducted extensive experiments with well-known motifs in physics-based workloads (Sweep3D and FFT), as well as with a representative graph application (MiniVite), using the well-known Structural Simulation Toolkit (SST) Macroscale Element Library (SST/macro) simulator for network interconnect evaluation. We demonstrate that our genetic algorithm-based approach is robust enough to find the underlying optimal topology of a particular workload.

network interconnects, graph search, meta-heuristi↗

An Evaluation of the Effect of Network Cost Optimization for Leadership Class Supercomputers

Dragonfly-based networks are an extensively deployed network topology in large-scale high-performance computing due to their cost-effectiveness and efficiency. The US will soon have three Exascale supercomputers for leadership class workloads deployed using dragonfly networks. Compared to indirect networks of similar scale, the dragonfly network has considerably reduced cable lengths, cable counts, and switch counts, resulting in significant network cost savings for a given system size, however, these cost reductions result in reduced global minimal paths and more challenging routing. Additionally, large scale dragonfly networks often require a taper at the global link level, resulting in less bisection bandwidth than is achievable in other traditional non-blocking topologies of equivalent scale. While dragonfly networks have been extensively studied, they have yet to be fully evaluated in an extreme scale (i.e., exascale) system that targets capability workloads. In this paper, we present the results of the first large scale evaluation of a dragonfly network on an exascale system (Frontier) and compare its behavior to a similar scale fat-tree network on a previous generation TOP500 system (Summit). This evaluation aims to determine the effect of network cost optimizations by measuring a tapered topology’s impact on capability workloads. Our evaluation is based on a collection of synthetic microbenchmarks, mini-apps, and full scale applications. It compares the scaling efficiencies of each benchmark between the dragonfly-based Frontier and the fat-tree-based Summit systems. Our results show that a dragonfly network is $\sim \mathbf{3 0 \%}$ more cost efficient than a fat-tree topology, which amortizes to $\sim 3 \%$ of an exascale system cost. Furthermore, while tapered dragonfly networks impose significant tradeoffs, the impacts are not as broad as initially thought and are mostly seen in applications with global communication patterns, particularly all-to-all (e.g., FFT-based algorithms), but also local communication patterns (e.g., nearest-neighbor algorithms) that are sensitive to network performance variability.

Khan, Awais↗

GPU-Accelerated Analytic Simulation of Sparse Ionization Signal Formation in Pixelated Projection Detector

This paper presents a GPU-accelerated simulation package, TRED, for next-generation neutrino detectors with pixelated charge readout, leveraging community-driven software ecosystems to ensure adaptability and extensibility. We introduce two generic contributions: (i) an effective-charge representation based on Gaussian quadrature rules, in which the linear- interpolation factors for the field response inside each voxel are absorbed into the effective charge, and (ii) a sparse, block- binned tensor representation that enables efficient FFT-based computation of induced signals on readout electrodes for sparsely activated detector volumes. The former captures structure inside a voxel without dense sampling, while the latter achieves low memory usage and scalable runtime, as demonstrated in bench- mark studies. The underlying data representation is applicable to large-scale detectors and to other computational problems involving sparse activity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗