Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

An empirical method for geometric calibration of a photon counting detector-based cone beam CT system

BACKGROUND: Geometric calibration is essential in developing a reliable computed tomography (CT) system. It involves estimating the geometry under which the angular projections are acquired. Geometric calibration of cone beam CTs employing small area detectors, such as currently available photon counting detectors (PCDs), is challenging when using traditional-based methods due to detectors’ limited areas. OBJECTIVE: This study presented an empirical method for the geometric calibration of small area PCD-based cone beam CT systems. METHODS: Unlike the traditional methods, we developed an iterative optimization procedure to determine geometric parameters using the reconstructed images of small metal ball bearings (BBs) embedded in a custom-built phantom. An objective function incorporating the sphericities and symmetries of the embedded BBs was defined to assess performance of the reconstruction algorithm with the given initial estimated set of geometric parameters. The optimal parameter values were those which minimized the objective function. The TIGRE toolbox was employed for fast tomographic reconstruction. To evaluate the proposed method, computer simulations were carried out using various numbers of spheres placed in various locations. Furthermore, efficacy of the method was experimentally assessed using a custom-made benchtop PCD-based cone beam CT. RESULTS: Computer simulations validated the accuracy and reproducibility of the proposed method. The precise estimation of the geometric parameters of the benchtop revealed high-quality imaging in CT reconstruction of a breast phantom. Within the phantom, the cylindrical holes, fibers, and speck groups were imaged in high fidelity. The CNR analysis further revealed the quantitative improvements of the reconstruction performed with the estimated parameters using the proposed method. CONCLUSION: Apart from the computational cost, we concluded that the method was easy to implement and robust.

Instruments & Instrumentation↗

Fast Local Spatial Verification for Feature-Agnostic Large-Scale Image Retrieval

Images from social media can reflect diverse viewpoints, heated arguments, and expressions of creativity, adding new complexity to retrieval tasks. Researchers working on Content-Based Image Retrieval (CBIR) have traditionally tuned their algorithms to match filtered results with user search intent. However, we are now bombarded with composite images of unknown origin, authenticity, and even meaning. With such uncertainty, users may not have an initial idea of what the search query results should look like. For instance, hidden people, spliced objects, and subtly altered scenes can be difficult for a user to detect initially in a meme image, but may contribute significantly to its composition. It is pertinent to design systems that retrieve images with these nuanced relationships in addition to providing more traditional results, such as duplicates and near-duplicates — and to do so with enough efficiency at large scale. In this work, we propose a new approach for spatial verification that aims at modeling object-level regions using image keypoints retrieved from an image index, which is then used to accurately weight small contributing objects within the results, without the need for costly object detection steps. We call this method the Objects in Scene to Objects in Scene (OS2OS) score, and it is optimized for fast matrix operations, which can run quickly on either CPUs or GPUs. It performs comparably to state-of-the-art methods on classic CBIR problems (Oxford 5K, Paris 6K, and Google-Landmarks), and outperforms them in emerging retrieval tasks such as image composite matching in the NIST MFC2018 dataset and meme-style imagery from Reddit.

42 ENGINEERING↗

Elastic distributed training with fast convergence and efficient resource utilization

Distributed learning is now routinely conducted on cloud as well as dedicated clusters. Training with elastic resources brings new challenges and design choices. Prior studies focus on runtime performance and assume a static algorithmic behavior. In this work, by analyzing the impact of of resource scaling on convergence, we introduce schedules for synchronous stochastic gradient descent that proactively adapt the number of learners to reduce training time and improve convergence. Our approach no longer assumes a constant number of processors throughout training. In our experiment, distributed stochastic gradient descent with dynamic schedules and reduction momentum achieves better convergence and significant speedups over prior static ones. Numerous distributed training jobs running on cloud may benefit from our approach.

Cong, Guojing↗

Real‐time XFEL data analysis at SLAC and NERSC: A trial run of nascent exascale experimental data analysis

X‐ray scattering experiments using free electron lasers (XFELs) are a powerful tool to determine the molecular structure and function of unknown samples (such as COVID‐19 viral proteins). XFEL experiments are a challenge to computing in two ways: (i) due to the high cost of running XFELs, a fast turnaround time from data acquisition to data analysis is essential to make informed decisions on experimental protocols; (ii) data‐collection rates are growing exponentially, requiring new scalable algorithms. Here we report our experiences analyzing data from two experiments at the Linac Coherent Light Source (LCLS) during September 2020. Raw data were analyzed on NERSC's Cori XC40 system, using the Superfacility paradigm: our workflow automatically moves raw data between LCLS and NERSC, where it is analyzed using the software package CCTBX. We achieved real time data analysis with a turnaround time from data acquisition to full molecular reconstruction in as little as 10 min—sufficient time for the experiment's operators to make informed decisions. By hosting the data analysis on Cori, and by automating LCLS‐NERSC interoperability, we achieved a data analysis rate which matches the data acquisition rate. Completing data analysis within 10 min is a first for XFEL experiments and an important milestone if we are to keep up with data‐collection trends.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

GX: a GPU-native gyrokinetic turbulence code for tokamak and stellarator design

GX is a code designed to solve the nonlinear gyrokinetic system for low-frequency turbulence in magnetized plasmas, particularly tokamaks and stellarators. In GX, our primary motivation and target is a fast gyrokinetic solver that can be used for fusion reactor design and optimization along with wide-ranging physics exploration. Here, this has led to several code and algorithm design decisions, specifically chosen to prioritize time to solution. First, we have used a discretization algorithm that is pseudospectral in the entire phase space, including a Laguerre–Hermite pseudospectral formulation of velocity space, which allows for smooth interpolation between coarse gyrofluid-like resolutions and finer conventional gyrokinetic resolutions and efficient evaluation of a model collision operator. Additionally, we have built GX to natively target graphics processors (GPUs), which are among the fastest computational platforms available today. Finally, we have taken advantage of the reactor-relevant limit of small $\rho _*$ by using the radially local flux-tube approach. In this paper we present details about the gyrokinetic system and the numerical algorithms used in GX to solve the system. We then present several numerical benchmarks against established gyrokinetic codes in both tokamak and stellarator magnetic geometries to verify that GX correctly simulates gyrokinetic turbulence in the small $\rho _*$. Moreover, we show that the convergence properties of the Laguerre–Hermite spectral velocity formulation are quite favourable for nonlinear problems of interest. Coupled with GPU acceleration, which we also investigate with scaling studies, this enables GX to be able to produce useful turbulence simulations in minutes on one (or a few) GPUs and higher fidelity results in a few hours using several GPUs. GX is open-source software that is ready for fusion reactor design studies.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Dispersion and the speed-limited particle-in-cell algorithm

This paper discusses temporally continuous and discrete forms of the speed-limited particle-in-cell (SLPIC) method first treated by Werner et al. [Phys. Plasmas 25, 123512 (2018)]. The dispersion relation for a 1D1V electrostatic plasma whose fast particles are speed-limited is derived and analyzed. By examining the normal modes of this dispersion relation, we show that the imposed speed-limiting substantially reduces the frequency of fast electron plasma oscillations while preserving the correct physics of lower-frequency plasma dynamics (e.g. ion acoustic wave dispersion and damping). We then demonstrate how the timestep constraints of conventional electrostatic particle-in-cell methods are relaxed by the speed-limiting approach, thus enabling larger timesteps and faster simulations. Here, these results indicate that the SLPIC method is a fast, accurate, and powerful technique for modeling plasmas wherein electron kinetic behavior is nontrivial (such that a fluid/Boltzmann representation for electrons is inadequate) but evolution is on ion timescales.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Fast methods for multisite charge transfer processes. I. Constrained, state averaged CASSCF(1,n) and CASSCF(2n − 1,n) simulations

We design a dynamically weighted state-averaged constrained complete active space self-consistent field (DW-SA-cCASSCF) algorithm to treat electrons or holes moving between n molecular fragments (where n can be larger than 2). Within such a so-called eDSCn/hDSCn approach, we consider configurations that are mutually single excitations of each other, and we apply a generalized set of constraints to tailor the method for studying charge transfer problems. The constrained optimization problem is efficiently solved using a DIIS-SQP algorithm, thus maintaining computational efficiency. We demonstrate the method for a finite Su–Schrieffer–Heeger chain, successfully reproducing the expected exponential decay of diabatic couplings with distance. When combined with a gradient, the current extension immediately enables efficient nonadiabatic dynamics simulations of complex multi-state charge transfer processes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Robust Communication-Free Protection Scheme for Islanded Microgrids with Relay Logic and Hardware-in-the-Loop Validation

This paper presents the Imbalance Square Factor (ISF) detection algorithm, an effective, computationally lightweight method for detecting faults in inverter-based microgrids. ISF provides a high magnitude at the time of the fault, which allows for fast detection and coordination between primary and backup relays. The imbalance squared factor, calculated using local voltage and currents, is used for fault detection and coordination among three relays. ISF is validated in hardware-in-the-loop (HIL) implementation inside commercial-grade relay logic (SEL-751). The HIL validation shows that ISF can coordinate primary, secondary, and tertiary relays in the 13-bus system in islanded operation.

Ferrari Maglia, Max [ORNL]↗

Real-Time XFEL Data Analysis at SLAC and NERSC: a Trial Run of Nascent Exascale Experimental Data Analysis

X-ray scattering experiments using Free Electron Lasers (XFELs) are a powerful tool to determine the molecular structure and function of unknown samples (such as COVID-19 viral proteins). XFEL experiments are a challenge to computing in two ways: i) due to the high cost of running XFELs, a fast turnaround time from data acquisition to data analysis is essential to make informed decisions on experimental protocols; ii) data collection rates are growing exponentially, requiring new scalable algorithms. Here we report our experiences analyzing data from two experiments at the Linac Coherent Light Source (LCLS) during September 2020. Raw data were analyzed on NERSC's Cori XC40 system, using the Superfacility paradigm: our workflow automatically moves raw data between LCLS and NERSC, where it is analyzed using the software package CCTBX. We achieved real time data analysis with a turnaround time from data acquisition to full molecular reconstruction in as little as 10 min -- sufficient time for the experiment's operators to make informed decisions. By hosting the data analysis on Cori, and by automating LCLS-NERSC interoperability, we achieved a data analysis rate which matches the data acquisition rate. Furthermore, completing data analysis with 10 mins is a first for XFEL experiments and an important milestone if we are to keep up with data collection trends.

Blaschke, Johannes P.↗

Space-Split Algorithm for Sensitivity Analysis of Discrete Chaotic Systems With Multidimensional Unstable Manifolds

Accurate approximations of the change of a system's output and its statistics with respect to the input are highly desired in computational dynamics. Ruelle's linear response theory provides breakthrough mathematical machinery for computing the linear response of chaotic dynamical systems. In this paper, we propose an algorithm for sensitivity analysis of discrete chaos with an arbitrary number of positive Lyapunov exponents. We combine the concept of perturbation space-splitting, which regularizes Ruelle's original expression, together with measure-based parameterization of the expanding subspace. We use these tools to rigorously derive trajectory-following recursive relations that converge exponentially fast, and construct a memory-efficient Monte Carlo scheme for derivatives of the output statistics. Thanks to the regularization and lack of simplifying assumptions on the system's behavior, our method is immune to the common problems of other popular methods such as the exploding tangent solutions and unphysical shadowing directions. Here, we provide a ready-to-use algorithm, analyze its complexity, and demonstrate several numerical examples of sensitivity computation using physically-inspired low-dimensional systems.

97 MATHEMATICS AND COMPUTING↗

Distilling particle knowledge for fast reconstruction at high-energy physics experiments

Knowledge distillation is a form of model compression that allows artificial neural networks of different sizes to learn from one another. Its main application is the compactification of large deep neural networks to free up computational resources, in particular on edge devices. In this article, we consider proton-proton collisions at the High-Luminosity Large Hadron Collider (HL-LHC) and demonstrate a successful knowledge transfer from an event-level graph neural network (GNN) to a particle-level small deep neural network (DNN). Our algorithm, DistillNet, is a DNN that is trained to learn about the provenance of particles, as provided by the soft labels that are the GNN outputs, to predict whether or not a particle originates from the primary interaction vertex. The results indicate that for this problem, which is one of the main challenges at the HL-LHC, there is minimal loss during the transfer of knowledge to the small student network, while improving significantly the computational resource needs compared to the teacher. This is demonstrated for the distilled student network on a CPU, as well as for a quantized and pruned student network deployed on a field programmable gate array. Our study proves that knowledge transfer between networks of different complexity can be used for fast artificial intelligence (AI) in high-energy physics that improves the expressiveness of observables over non-AI-based reconstruction algorithms. Such an approach can become essential at the HL-LHC experiments, e.g. to comply with the resource budget of their trigger stages.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Distributed approximate minimal Steiner trees with millions of seed vertices on billion-edge graphs

In this report, we present a parallel 2-approximation Steiner minimal tree algorithm and its MPI-based distributed implementation. In place of expensive distance computations between all pairs of seed vertices, the solution we employ exploits a cheaper Voronoi cell computation. Our design leverages asynchronous processing and message prioritization to accelerate convergence of distance computations, and harnesses vertex and edge centric processing to offer fast time-to-solution. We demonstrate scalability and performance using real-world graphs with up to 128 billion edges and 512 compute nodes, and show the ability to find Steiner trees with up to one million seed vertices. Using 12 data instances, we present comparison with the state-of-the-art exact solver, SCIP-Jack, and two sequential 2-approximate algorithms. We empirically show that, on average, the total distance of the Steiner tree identified by our solution is 1.1290 times greater than the Steiner minimal tree – well within the theoretical approximation bound of 2.

97 MATHEMATICS AND COMPUTING↗

Updates to the Regional Seismic Travel Time (RSTT) Model: 1. Tomography

Abstract A function of global monitoring of nuclear explosions is the development of Earth models for predicting seismic travel times for more accurate calculation of event locations. Most monitoring agencies rely on fast, distance-dependent one-dimensional (1D) Earth models to calculate seismic event locations quickly and in near real-time. RSTT (Regional Seismic Travel Time) is a seismic velocity model and computer software package that captures the major effects of three-dimensional crust and upper mantle structure on regional seismic travel times, while still allowing for fast prediction speed (milliseconds). We describe updates to the RSTT model using a refined data set of regional phases (i.e., Pn, Pg, Sn, Lg) using the Bayesloc relative relocation algorithm. The tomographic inversion shown here acts to refine the previous RSTT public model ( rstt201404um ) and displays significant features related to areas of global tectonic complexity as well as further reduction in arrival residual values. Validation of the updated RSTT model demonstrates significant reduction in median epicenter mislocation (15.3 km) using all regional phases compared to the iasp91 1D model (22.1 km) as well as to the current station correction approach used at the Comprehensive Nuclear-Test-Ban Treaty Organization International Data Centre (18.9 km).

58 GEOSCIENCES↗

Sparse Approximate Multifrontal Factorization with Composite Compression Methods

This article presents a fast and approximate multifrontal solver for large sparse linear systems. In a recent work by Liu et al., we showed the efficiency of a multifrontal solver leveraging the butterfly algorithm and its hierarchical matrix extension, HODBF (hierarchical off-diagonal butterfly) compression to compress large frontal matrices. The resulting multifrontal solver can attain quasi-linear computation and memory complexity when applied to sparse linear systems arising from spatial discretization of high-frequency wave equations. To further reduce the overall number of operations and especially the factorization memory usage to scale to larger problem sizes, in this article we develop a composite multifrontal solver that employs the HODBF format for large-sized fronts, a reduced-memory version of the nonhierarchical block low-rank format for medium-sized fronts, and a lossy compression format for small-sized fronts. This allows us to solve sparse linear systems of dimension up to 2.7 × larger than before and leads to a memory consumption that is reduced by 70% while ensuring the same execution time. The code is made publicly available in GitHub.

97 MATHEMATICS AND COMPUTING↗

Extension of the high-resolution thermal-hydraulics code ESCOT to hexagonal core geometries for multi-physics calculations

The extension of the capabilities of the pin-level nuclear reactor core thermal-hydraulics (T/H) code ESCOT to analyze hexagonal fueled cores and its performance are presented. ESCOT is an accurate yet fast core thermal-hydraulics solution aiming at high-fidelity and high-resolution multi-physics core analysis in the framework of massively parallel computing platforms. Its algorithm solution is based on the four-equation drift-flux model for two-phase calculations, these are numerically solved by applying the Finite Volume Method (FVM) and the Semi-Implicit Method for Pressure-Linked Equation (SIMPLE)-like algorithm in a staggered grid system. Constitutive models such as turbulent mixing, pressure drop, and vapor generation are employed to simulate key phenomena in subchannel-scale analysis. ESCOT is parallelized by a double (radial and axial) domain decomposition that enables its highly parallelized execution. The coupling of the code with the neutronics whole core solver for hexagonal geometries nTRACER is described. The newly implemented ESCOT features are validated by comparing single assembly and full core steady state nTRACER-ESCOT solutions with nTRACER standalone internal one-dimensional T/H solver results. The validation problems are based on the VVER 440 and VVER 1000 cores. ESCOT results show differences within an acceptable range with respect to the simple 1D nTRACER built-in solver. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Application of Monte Carlo Algorithms to Cardiac Imaging Reconstruction

Monte Carlo algorithms have a growing impact on nuclear medicine reconstruction processes. One ofthe main limitations of myocardial perfusion imaging (MPI) is the effective mitigation of the scattering component,which is particularly challenging in Single Photon Emission Computed Tomography (SPECT). In SPECT,no timing information can be retrieved to locate the primary source photons. Monte Carlo methods allow anevent-by-event simulation of the scattering kinematics, which can be incorporated into a model of the imagingsystem response. This approach was adopted in the late Nineties by several authors, and recently took advantageof the increased computational power made available by high-performance CPUs and GPUs. These recent developmentsenable a fast image reconstruction with improved image quality, compared to deterministic approaches.Deterministic approaches are based on energy-windowing of the detector response, and on the cumulative estimateand subtraction of the scattering component. In this paper, we review the main strategies and algorithms tocorrect the scattering effect in SPECT and focus on Monte Carlo developments, which nowadays allow the threedimensionalreconstruction of SPECT cardiac images in a few seconds.

Pharmacology & Pharmacy↗

Fast widefield imaging of neuronal structure and function with optical sectioning in vivo

Optical microscopy, owing to its noninvasiveness and subcellular resolution, enables in vivo visualization of neuronal structure and function in the physiological context. Optical-sectioning structured illumination microscopy (OS-SIM) is a widefield fluorescence imaging technique that uses structured illumination patterns to encode in-focus structures and optically sections 3D samples. However, its application to in vivo imaging has been limited. In this study, we optimized OS-SIM for in vivo neural imaging. We modified OS-SIM reconstruction algorithms to improve signal-to-noise ratio and correct motion-induced artifacts in live samples. Incorporating an adaptive optics (AO) module to OS-SIM, we found that correcting sample-induced optical aberrations was essential for achieving accurate structural and functional characterizations in vivo. With AO OS-SIM, we demonstrated fast, high-resolution in vivo imaging with optical sectioning for structural imaging of mouse cortical neurons and zebrafish larval motor neurons, and functional imaging of quantal synaptic transmission at Drosophila larval neuromuscular junctions.

42 ENGINEERING↗

Investigation of a digitizer for the plastic scintillation detectors of time-of-flight mass measurements

A CAEN DT5742 digitizer has been investigated to process the fast signals from the photomultiplier tubes of the time-of-flight detectors for fast ion beams. A test setup consisting of two plastic scintillation detectors and a pulsed laser source provided signals which were recorded by the digitizer and systematically analyzed with different algorithms to derive the amplitude, rise time and arrival time of the detection signal. To obtain the best amplitude and time resolutions, various optimization techniques including peak fitting and signal smoothing, interpolation improvement, and time-walk correction by amplitude and rise time have been performed and compared in detail. Here we compared the time resolutions obtained by three digital algorithms of leading-edge, zero-crossing constant-fraction (ZC-CFD) and direct constant-fraction discriminations. Finally, we found that the best time-of-flight resolution between two PMTs can be achieved as 12 ps by the method using sample minimum for peak location and using the time-walk-resistant ZC-CFD with 4-point interpolation for timing.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗