Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Proton discrimination in CLYC for fast neutron spectroscopy

The Cs 2 LiYCl 6 :Ce (CLYC) elpasolite scintillator is known for its response to fast and thermal neutrons along with good γ-ray energy resolution. While the 35 Cl(n,p) reaction has been identified as a potential means for CLYC-based fast neutron spectroscopy in the absence of time-of-flight (TOF), previous efforts to functionalize CLYC as a fast neutron spectrometer have been thwarted by the inability to isolate proton interactions from 6 Li(n,α) and 35 Cl(n,α) signals. This work introduces a new approach to particle discrimination in CLYC for fission spectrum neutrons using a multi-gate charge integration algorithm that provides excellent separation between protons and heavier charged particles. Neutron TOF data were collected using a 252 Cf source, an array of EJ-309 organic liquid scintillators, and a 6 Li-enriched CLYC scintillator outfitted with fast electronics. Modal waveforms were constructed corresponding to the different reaction channels, revealing significant differences in the pulse characteristics of protons and heavier charged particles at ultrafast, fast, and intermediate time scales. These findings informed the design of a pulse shape discrimination algorithm, which was validated using the TOF data. This study also proposes an iterative subtraction method to mitigate contributions from confounding reaction channels in proton and heavier charged particle pulse height spectra, opening the door for CLYC-based fast neutron and γ-ray spectroscopy while preserving sensitivity to thermal neutron capture signals.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Low-latency Jet Tagging for HL-LHC Using Transformer Architectures

Transformers are the state-of-the-art model architectures and widely used in application areas of machine learning. However the performance of such architectures is less well explored in the ultra-low latency domains where deployment on FPGAs or ASICs is required. Such domains include the trigger and data acquisition systems of the LHC experiments. We present a transformer-based algorithm for jet tagging built with the HGQ2 framework, which is able to produce a model with heterogeneous bitwidths for fast inference on FPGAs, as required in the trigger systems at the LHC experiments. The bitwidths are acquired during training by minimizing the total bit operations as an additional parameter. By allowing a bitwidth of zero, the model is pruned in-situ during training. Using this quantization-aware approach, our algorithm achieves state-of-the-art performance while also retaining permutation invariance which is a key property for particle physics applications. Due to the strength of transformers in representation learning, our work also serves as a stepping stone for the development of a larger foundation model for trigger applications.

Laatu, Lauri [Imperial Coll., London]↗

Application of automated iterative target detection for standoff hyperspectral imaging

The utility of hyperspectral imaging (HSI) has been well established for a wide array of applications but has generated a need for automated screening of high volumes of large HSI cubes. We report two important automated algorithms for more efficient standoff processing: atmospheric correction and target detection. The atmospheric correction method is based on a fast asymmetric least squares approach that is applied on a pixel-by-pixel basis. Here, the correction can be applied to entire images without manually identifying regions of interest and utilizes only in-scene information, no ancillary modeling of the atmosphere is required. An iterative target detection approach is also introduced which demonstrates faster speeds relative to moving window approaches. The target detection algorithm classifies each pixel as true target detections, near target detections, clutter, and no-calls. The algorithms were tested on forty images of twenty-two solid mineral targets placed at a 14-meter standoff distance allowing general observations on expected detection performance for a variety of minerals. In addition to identifying anomalous pixels, the inclusion of “no-calls” reduced the number of false detections significantly.

47 OTHER INSTRUMENTATION↗

BigNeuron: a resource to benchmark and predict performance of algorithms for automated tracing of neurons in light microscopy datasets

BigNeuron is an open community bench-testing platform with the goal of setting open standards for accurate and fast automatic neuron tracing. We gathered a diverse set of image volumes across several species that is representative of the data obtained in many neuroscience laboratories interested in neuron tracing. Here, we report generated gold standard manual annotations for a subset of the available imaging datasets and quantified tracing quality for 35 automatic tracing algorithms. The goal of generating such a hand-curated diverse dataset is to advance the development of tracing algorithms and enable generalizable benchmarking. Together with image quality features, we pooled the data in an interactive web application that enables users and developers to perform principal component analysis, t-distributed stochastic neighbor embedding, correlation and clustering, visualization of imaging and tracing data, and benchmarking of automatic tracing algorithms in user-defined data subsets. The image quality metrics explain most of the variance in the data, followed by neuromorphological features related to neuron size. Furthermore, we observed that diverse algorithms can provide complementary information to obtain accurate results and developed a method to iteratively combine methods and generate consensus reconstructions. The consensus trees obtained provide estimates of the neuron structure ground truth that typically outperform single algorithms in noisy datasets. However, specific algorithms may outperform the consensus tree strategy in specific imaging conditions. Finally, to aid users in predicting the most accurate automatic tracing results without manual annotations for comparison, we used support vector machine regression to predict reconstruction quality given an image volume and a set of automatic tracings.

97 MATHEMATICS AND COMPUTING↗

Deep Reinforcement Learning Enabled Physical-Model-Free Two-Timescale Voltage Control Method for Active Distribution Systems

Active distribution networks are being challenged by frequent and rapid voltage violations due to renewable energy integration. Conventional model-based voltage control methods rely on accurate parameters of the distribution networks, which are difficult to achieve in practice. This paper proposes a novel physical-model-free two-timescale voltage control framework for active distribution systems. To achieve fast control of PV inverters, the whole network is first partitioned into several subnetworks using voltage-reactive power sensitivity. Then, the scheduling of PV inverters in the multiple sub-networks is formulated as Markov games and solved by a multi-agent soft actor-critic (MASAC) algorithm, where each subnetwork is modeled as an intelligent agent. All agents are trained in a centralized manner to learn a coordinated strategy while being executed based on only local information for fast response. For the slower time-scale control, OLTCs and switched capacitors are coordinated by a single agent-based SAC algorithm using the global information with considering control behaviors of the inverters. Particularly, the two-level agents are trained concurrently with information exchange according to the reward signal calculated from the data-driven surrogate model. Comparative tests with different benchmark methods on IEEE 33-and 123-bus systems and 342-node low voltage distribution system demonstrate that the proposed method can effectively mitigate the fast voltage violations and achieve systematical coordination of different voltage regulation assets without the knowledge of accurate system model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toucan: A performance portable, scalable implementation of the DECA algorithm

In the field of additive manufacturing (AM), cellular automata (CA) is extensively used to simulate microstructural evolution during solidification. However, while traditional CA approaches are relatively fast, they still require a substantial number of time steps, are limited to moderate volumes, and are relatively difficult to improve through parallelism due to the highly localized nature of the solidification front. Here, to address these issues of time to solution and load balancing, we introduce Toucan, a parallel, performance-portable, and scalable code written in C++ with the Kokkos library that leverages the discrete event inspired cellular automata (DECA) algorithm to perform parallel-in-time (PinT) grain growth simulations. Toucan effectively mitigates load balancing issues by distributing the computational workload more evenly across processors, enhancing scalability and efficiency. We conduct both strong and weak scaling studies on up to 64 GPUs on the Frontier supercomputer, demonstrating that Toucan significantly outperforms the current state-of-the-art, time-stepped CA code, ExaCA, on both single and multi-GPU simulations. Even in AM-specific weak scaling scenarios, Toucan maintains near-ideal scaling, in contrast to the linear increase observed with ExaCA due to the moving laser raster pattern. This study highlights Toucan’s potential to transform microstructural simulations in AM by radically improving both efficiency and scalability over existing methods.

36 MATERIALS SCIENCE↗

QuanSimBench

QuanSimBench is a software program that performs a gate-by-gate simulation of a quantum computer with a full state vector formulation. It simulates a simplified version of Shor's algorithm with increasing number of qubits until resources are exhausted. The score (states/s) is how fast the Approximate Quantum Fourier Transform can be computed (the modular exponentiation is not timed). The goal is to quantify the ability of a computer to simulate ideal quantum circuits.

Pakin, Scott↗

OpenCGRA: An Open-Source Unified Framework for Modeling,Testing, and Evaluating CGRAs

Coarse-grained reconfigurable arrays (CGRAs),loosely defined as arrays of functional units (e.g, adder, sub-tractor, multiplier, divider, or larger multi-operation units, butsmaller than a general-purpose core) interconnected through aNetwork-on-Chip, provide higher flexibility than domain-specificASIC accelerators while offering increased hardware efficiencywith respect to fine-grained reconfigurable devices, such as FieldProgrammable Gate Arrays (FPGAs). The fast evolving fieldsof machine learning and edge computing, which are seeing acontinuous flow of novel algorithms and larger models, makeCGRAs ideal target architectures to allow domain specializationwithout loosing too much generality. They also generally offerquicker and more effective reconfigurability than FPGAs, po-tentially allowing adaptation during actual algorithm execution,and implement a dataflow programming paradigm that adaptswell to these emerging workloads. Designing and generating aCGRA, however, still requires to define the type and number ofthe specific functional units, implement their interconnect andthe network topology, and perform its simulation and validation,given a variety of workloads of interest.In this paper, we propose OpenCGRA, a Python-based unifiedframework that integrates generation, modeling, testing and eval-uation for CGRAs. OpenCGRA is the first open-source integratedframework able to support the full top-to-bottom design flow forspecializing and implementing CGRAs: modeling at different ab-straction levels (functional level, cycle level, register-transfer level),generation, simulation, testing at different granularities (unit test-ing, integration testing, property-based testing), and characteriza-tion (area, power, and timing). OpenCGRAs will be made availableon GitHub.

CGRA, synthesis↗

Progress in modelling fast-ion D-alpha spectra and neutral particle analyzer fluxes using FIDASIM

FIDASIM is a code that models signals produced by charge-exchange reactions between neutrals and ions (both fast and thermal) in magnetically confined plasmas. With the ion distribution function as input, the code predicts the efflux to a neutral particle analyzer diagnostic and the photon radiance of Balmer-alpha light to a fast-ion D α diagnostic, in addition to many other related quantities. A new, parallelized version of the Monte Carlo code FIDASIM has been developed in Fortran90 that is substantially faster than the original interactive data language version. Modified algorithms include more accurate treatments of the time dependent collisional-radiative equations that describe neutral energy levels, of the cloud of ‘halo’ neutrals that surround the injected neutral beam, and of finite Larmor radius effects. Enhanced physics capabilities include modelling ‘passive’ signals from cold edge neutrals, the ability to treat general three-dimensional magnetic confinement configurations, and calculations of diagnostic-specific weight functions that enable tomographic reconstructions of the fast-ion distribution function. Neutral beam attenuation, beam emission, and fast-ion birth profiles are also modelled. Finally, the new algorithms have been successfully validated against experimental data and new features have been tested through benchmarks between two independently developed versions of the code.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Front-End Design for SiPM-Based Monolithic Neutron Double Scatter Imagers

Neutron double scatter imaging exploits the kinematics of neutron elastic scattering to enable emission imaging of neutron sources. Due to the relatively low coincidence detection efficiency of fast neutrons in organic scintillator arrays, imaging efficiency for double scatter cameras can also be low. One method to realize significant gains in neutron coincidence detection efficiency is to develop neutron double scatter detectors which employ monolithic blocks of organic scintillator, instrumented with photosensor arrays on multiple faces to enable 3D position and multi-interaction time pickoff. Silicon photomultipliers (SiPMs) have several advantageous characteristics for this approach, including high photon detection efficiency (PDE), good single photon time resolution (SPTR), high gain that translates to single photon counting capabilities, and ability to be tiled into large arrays with high packing fraction and photosensitive area fill factor. However, they also have a tradeoff in high uncorrelated and correlated noise rates (dark counts from thermionic emissions and optical photon crosstalk generated during avalanche) which may complicate event positioning algorithms. We have evaluated the noise characteristics and SPTR of Hamamatsu S13360-6075 SiPMs with low noise, fast electronic readout for integration into a monolithic neutron scatter camera prototype. The sensors and electronic readout were implemented in a small-scale prototype detector in order to estimate expected noise performance for a monolithic neutron scatter camera and perform proof-of-concept measurements for scintillation photon counting and three-dimensional event positioning.

47 OTHER INSTRUMENTATION↗

Avoiding excess computation in asynchronous evolutionary algorithms

Abstract Asynchronous evolutionary algorithms are becoming increasingly popular as a means of making full use of many processors while solving computationally expensive search and optimization problems. These algorithms excel at keeping large clusters fully utilized, but may sometimes inefficiently sample an excess of fast‐evaluating solutions at the expense of higher‐quality, slow‐evaluating ones. We have previously introduced a steady‐state parent selection strategy, SWEET (“Selection whilE EvaluaTing”), that sometimes selects individuals that are still being evaluated and allows them to reproduce early. We perform a takeover‐time analysis that confirms that this strategy gives slow‐evaluating individuals that have higher fitnesses an increased ability to multiply in the population. We also find that SWEET appears effective at improving optimization performance on problems in which solution quality is positively correlated with evaluation time. We evaluate our approach on six simulated real‐valued optimization problems and three real‐world applications: an autonomous vehicle controller problem that involves tuning a spiking neural network and two adversarial EA problems. We further evaluate SWEET versus a basic asynchronous process in a simulated setting. We present evidence that SWEET outperforms basic asynchronous processes in a use‐case in which performance is positively correlated with evaluation time, and performs comparably (and often better) than basic asynchronous processes in several use‐cases where performance is negatively correlated with evaluation time. That said, in the cases where performance and evaluation time are negatively correlated the variance of outcomes for SWEET is notably high.

97 MATHEMATICS AND COMPUTING↗

Automated Algorithms for Screening Electronic Parts for Aging using Power Spectra Analysis (PSA) Data

Understanding the age of semiconductor parts being built into devices and systems is of interest for manufacturing quality control. Power spectrum analysis (PSA) is a fast, non-destructive, sensitive method for examining semiconductor parts. This talk will cover the use of multivariate analysis on both PSA data and conventional current-voltage data generated prior to PSA analysis to create algorithms that can be automated to screen semiconductor parts for aging.

Multari, Rosalie A↗

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Online Optimization for Networked Distributed Energy Resources With Time-Coupling Constraints

This paper proposes a Lyapunov optimization-based online distributed (LOOD) algorithmic framework for active distribution networks (ADNs) with numerous photovoltaic inverters and inverter air conditionings (IACs). In the proposed scheme, ADNs can track an active power setpoint reference at the substation in response to transmission-level requests while concurrently minimizing the social utility loss and ensuring the security of voltages. Conventional distributed optimization methods are rarely feasible to track the optimal solutions in fast variable environments using a fine-grained sampling interval where the underlying optimization problem evolves with the iterations of the algorithms. In contrast, based on the framework of online convex optimization (OCO), the developed approach uses a distributed algebraic update to compute the next round decisions relying on the current feedback of measurements. Notably, the time-coupling constraints of IACs are decoupled for online implementation with Lyapunov optimization technique. An incentive scheme is tailored to coordinate the customer-owned assets in lieu of the direct control from network operators. Optimality and convergency are characterized analytically. Finally, we corroborate the proposed method on a modified version of 33-node test feeder. Benchmark tests show that the proposed method is computationally and economically efficient, and outperforming existing algorithms.

active distribution networks↗

encore : an O ( N g2) estimator for galaxy N -point correlation functions

ABSTRACT We present a new algorithm for efficiently computing the N-point correlation functions (NPCFs) of a 3D density field for arbitrary N. This can be applied both to a discrete spectroscopic galaxy survey and a continuous field. By expanding the statistics in a separable basis of isotropic functions built from spherical harmonics, the NPCFs can be estimated by counting pairs of particles in space, leading to an algorithm with complexity $\mathcal {O}(N_\mathrm{g}^2)$ for Ng particles, or $\mathcal {O}(N_\mathrm{FFT}\log N_\mathrm{FFT})$ when using a Fast Fourier Transform with NFFT grid-points. In practice, the rate-limiting step for N > 3 will often be the summation of the histogrammed spherical harmonic coefficients, particularly if the number of radial and angular bins is large. In this case, the algorithm scales linearly with Ng. The approach is implemented in the encore code, which can compute the 3PCF, 4PCF, 5PCF, and 6PCF of a BOSS-like galaxy survey in ${\sim}100$ CPU-hours, including the corrections necessary for non-uniform survey geometries. We discuss the implementation in depth, along with its GPU acceleration, and provide practical demonstration on realistic galaxy catalogues. Our approach can be straightforwardly applied to current and future data sets to unlock the potential of constraining cosmology from the higher point functions.

79 ASTRONOMY AND ASTROPHYSICS↗

ORCA: Outlier detection and Robust Clustering for Attributed graphs

Here, a framework is proposed to simultaneously cluster objects and detect anomalies in attributed graph data. Our objective function along with the carefully constructed constraints promotes interpretability of both the clustering and anomaly detection components, as well as scalability of our method. In addition, we developed an algorithm called Outlier detection and Robust Clustering for Attributed graphs (ORCA) within this framework. ORCA is fast and convergent under mild conditions, produces high quality clustering results, and discovers anomalies that can be mapped back naturally to the features of the input data. The efficacy and efficiency of ORCA is demonstrated on real world datasets against multiple state-of-the-art techniques.

97 MATHEMATICS AND COMPUTING↗

Robust constrained tension control for high-precision roll-to-roll processes

Tension control is critical for maintaining good product quality in most roll-to-roll (R2R) production systems. Previous work has primarily focused on improving the disturbance rejection performance of tension controllers. Here, a robust linear parameter-varying model predictive control (LPV-MPC) scheme is designed to enhance the tension tracking performance of a pilot R2R system for deposition of materials used in flexible thin film applications. The performance of a tension controller may degrade due to disturbances associated with model uncertainties and the slowly-changing dynamics in R2R systems. We introduce a method that separately treats these two sources of disturbance. The controller utilizes an incremental model to eliminate the errors caused by the mismatch between the nominal model and the actual system. A tube-based MPC formulation combined with scheduled parameters adequately updates models and corrects for the time-varying dynamics. Constraints on the rated motor torque are incorporated in the MPC to maintain the controller reliability and avoid machine failures. We illustrate the operation of our control algorithm through simulation of an actual R2R system. The controller outperforms the benchmarks in terms of fast transient response and offset-free tension tracking. Furthermore, it also demonstrates immunity from variations due to parametric uncertainties.

42 ENGINEERING↗

Single-shot in-line x-ray phase-contrast imaging of void-shockwave interactions in fusion energy materials

Recent breakthroughs in nuclear fusion, specifically the report of reactions exceeding scientific breakeven at the National Ignition Facility (NIF), highlight the potential of inertial fusion energy (IFE) as a sustainable and virtually limitless energy source. However, further progress in IFE requires characterization of defects in ablator materials and how they affect fuel capsule compression. Voids within the ablator can degrade energy yield, but their impact on the density distribution has primarily been studied through simulations, with limited high-resolution experimental validation. To address this, we used the x-ray free-electron laser (XFEL) at the matter in extreme conditions (MECs) instrument at the Linac coherent light source (LCLS) to capture 2D x-ray phase-contrast (XPC) images of a void-bearing sample with a composition similar to inertial confinement fusion (ICF) ablators. By driving a compressive shockwave through the sample using MEC's long-pulse laser system, we analyzed how voids influence shockwave propagation and density distribution during compression. To quantify this impact, we extracted phase information using two phase retrieval algorithms. First, we applied the contrast transfer function (CTF) method, paired with Tikhonov regularization and a fast optimization approach to generate an initial phase estimate. We then refined the result using a projected gradient descent (PGD) method that works directly with the sample's refractive index. Comparing these results with radiation adaptive grid Eulerian (xRAGE) radiation hydrodynamic simulations enables identification of model validation needs or improvements. By calculating phase maps in situ, it becomes possible to reconstruct areal density maps, improving understanding of laser-capsule interactions and advancing IFE research.

Hodge, D. S. [Colorado State Univ., Fort Collins, ↗