Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Developing extreme fast charge battery protocols – A review spanning materials to systems

Extreme fast charging (XFC) has become a focal research point in the lithium-battery community over the last several years. As adoption of electric vehicles increases, fast charging has become a key driver in enhancing consumer recharge experience. Recently, the research community has made significant improvements in developing charge protocols to support XFC. New charge protocol designs derived using a combination of advanced, physically derived models, and electrochemical and secondary characterization methods, increase charge acceptance and decrease aging. By coordinating these methods and modifying protocols to account for different material constraints, including lithium plating and cathode particle degradation, novel charge protocols have increased the energy accepted during charging by over 25% in 10 min and increased the charge acceptance prior to a constant-voltage step by approximately 3x. Here, we review several charge-protocol advances, aging factors which are enhanced by XFC and advances which will enable adoption of XFC capable vehicles. These advances include implementing machine learning and other detection algorithms to reduce and classify lithium plating, which is known to significantly degrade cell performance and reduce cell life. The review concludes by discussing full-system fast charge requirements, including electric vehicle service equipment needs for implementing XFC protocols.

25 ENERGY STORAGE↗

Online Adaptive Algorithm for Constraint Energy Minimizing Generalized Multiscale Discontinuous Galerkin Method

Here in this research, we propose an online basis enrichment strategy within the framework of a recently developed constraint energy minimizing generalized multiscale discontinuous Galerkin method. Combining the technique of oversampling, one makes use of the information of the current residuals to adaptively construct basis functions in the online stage to reduce the error of multiscale approximation. A complete analysis of the method is presented, which shows the proposed online enrichment leads to a fast convergence from multiscale approximation to the fine-scale solution. The error reduction can be made sufficiently large by suitably selecting oversampling regions and the number of oversampling layers. Further, the convergence rate of the enrichment algorithm depends on a factor of exponential decay regarding the number of oversampling layers and a user-defined parameter. Numerical results are provided to demonstrate the effectiveness and efficiency of the proposed online adaptive algorithm.

97 MATHEMATICS AND COMPUTING↗

Machine learning without a processor: Emergent learning in a nonlinear analog network

Standard deep learning algorithms require differentiating large nonlinear networks, a process that is slow and power-hungry. Electronic contrastive local learning networks (CLLNs) offer potentially fast, efficient, and fault-tolerant hardware for analog machine learning, but existing implementations are linear, severely limiting their capabilities. These systems differ significantly from artificial neural networks as well as the brain, so the feasibility and utility of incorporating nonlinear elements have not been explored. Here, we introduce a nonlinear CLLN—an analog electronic network made of self-adjusting nonlinear resistive elements based on transistors. We demonstrate that the system learns tasks unachievable in linear systems, including XOR (exclusive or) and nonlinear regression, without a computer. We find our decentralized system reduces modes of training error in order (mean, slope, curvature), similar to spectral bias in artificial neural networks. The circuitry is robust to damage, retrainable in seconds, and performs learned tasks in microseconds while dissipating only picojoules of energy across each transistor. This suggests enormous potential for fast, low-power computing in edge systems like sensors, robotic controllers, and medical devices, as well as manufacturability at scale for performing and studying emergent learning.

Science & Technology - Other Topics↗

Enabling Extreme Fast Charging with Energy Storage Final Report

The objective of this project was to develop and demonstrate an extreme fast charging (XFC) station with three 350 kW charging ports that operate at a combined power exceeding 1 MW and mitigate the grid impact by employing smart charging algorithms, using an energy storage system, and connecting directly to a medium-voltage (15 kV-class three-phase) distribution feeder. This final report describes the team's achievements in all aspects of the technology development.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Recovered supernova Ia rate from simulated LSST images

Aims.TheVera C. RubinObservatory’s Legacy Survey of Space and Time (LSST) will revolutionize time-domain astronomy by detecting millions of different transients. In particular, it is expected to increase the number of known type Ia supernovae (SN Ia) by a factor of 100 compared to existing samples up to redshift ∼1.2. Such a high number of events will dramatically reduce statistical uncertainties in the analysis of the properties and rates of these objects. However, the impact of all other sources of uncertainty on the measurement of the SN Ia rate must still be evaluated. The comprehension and reduction of such uncertainties will be fundamental both for cosmology and stellar evolution studies, as measuring the SN Ia rate can put constraints on the evolutionary scenarios of different SN Ia progenitors. Methods.We used simulated data from the Dark Energy Science Collaboration (DESC) Data Challenge 2 (DC2) and LSST Data Preview 0 to measure the SN Ia rate on a 15 deg 2 region of the “wide-fast-deep” area. We selected a sample of SN candidates detected in difference images, associated them to the host galaxy with a specially developed algorithm, and retrieved their photometric redshifts. We then tested different light-curve classification methods, with and without redshift priors (albeit ignoring contamination from other transients, as DC2 contains only SN Ia). We discuss how the distribution in redshift measured for the SN candidates changes according to the selected host galaxy and redshift estimate. Results.We measured the SN Ia rate, analyzing the impact of uncertainties due to photometric redshift, host-galaxy association and classification on the distribution in redshift of the starting sample. We find that we are missing 17% of the SN Ia, on average, with respect to the simulated sample. As 10% of the mismatch is due to the uncertainty on the photometric redshift alone (which also affects classification when used as a prior), we conclude that this parameter is the major source of uncertainty. We discuss possible reduction of the errors in the measurement of the SN Ia rate, including synergies with other surveys, which may help us to use the rate to discriminate different progenitor models.

Astronomy & Astrophysics↗

Sub-system quantum dynamics using coupled cluster downfolding techniques

In this paper, we discuss extending the sub-system embedding sub-algebra coupled cluster (SES-CC) formalism and the double unitary coupled cluster (DUCC) ansatz to the time domain. As we demonstrated in earlier studies, it is possible, using these formalisms, to calculate the energy of the entire system as an eigenvalue of downfolded/effective Hamiltonian in the active space, that is identifiable with the sub-system of the composite system. In these studies, we demonstrated that downfolded Hamiltonians integrate out Fermionic degrees of freedom that do not correspond to the physics encapsulated by the active space. We extend these results to the time-dependent Schrödinger equation, showing that a similar construct is possible to partition a system into a sub-system that varies slowly in time and a remaining subsystem that corresponds to fast oscillations. This time dependent formalism allows coupled cluster quantum dynamics to be extended to larger systems and for the formulation of novel quantum algorithms based on the quantum Lanczos approach, which have recently been considered in the literature.

coupled cluster, Electron correlation, quantum dyn↗

Quantitative phase retrieval and characterization of magnetic nanostructures via Lorentz (scanning) transmission electron microscopy

Magnetic materials phase reconstruction using Lorentz transmission electron microscopy (LTEM) measurements have traditionally been achieved using longstanding methods such as off-axis holography (OAH) fast-Fourier transform technique and the transport-of-intensity equation (TIE). The increase in access to processing power alongside the development of advanced algorithms have allowed for phase retrieval of nanoscale magnetic materials with greater efficacy and resolution. Specifically, reverse-mode automatic differentiation (RMAD) and the extended electron ptychography iterative engine (ePIE) are two recent developments of phase retrieval that can be applied to analyzing micro-to-nano- scale magnetic materials. This work evaluates phase retrieval using TIE, RMAD, and ePIE in simulations of Permalloy (Ni 80 Fe 20 ) nanoscale islands, or nanomagnets. Extending beyond simulations, we demonstrate total phase retrieval and image reconstructions of a NiFe nanowire using OAH and RMAD in LTEM and ePIE in Lorentz-mode-4D scanning transmission electron microscopy experiments and determine the saturation magnetization through corroborations with micromagnetic modeling. Finally, we demonstrate the efficacy of these methods in retrieving the total phase and highlight its use in characterizing and analyzing the proximity effect of the magnetic nanostructures.

Lorentz transmission electron microscopy↗

A Survey of Singular Value Decomposition Methods for Distributed Tall/Skinny Data

The Singular Value Decomposition (SVD) is one of the most important matrix factorizations, enjoying a wide variety of applications across numerous application domains. In statistics and data analysis, the common applications of SVD inclue Principal Components Analysis (PCA) and regression. Usually these applications arise on data that has far more rows than columns, so-called "tall/skinny" matrices. In the big data analytics context, this may take the form of hundreds of millions to billions of rows with only a few hundred columns. There is a need, therefore, for fast, accurate, and scalable tall/skinny SVD implementations which can fully utilize modern computing resources. To that end, we present a survey of three different algorithms for computing the SVD for these kinds of tall/skinny data layouts using MPI for communication. We contextualize these with common big data analytics techniques. Finally, we present both CPU and GPU timing results from the Summit supercomputer, and discuss possible alternative approaches.

Schmidt, Drew↗

Combining Spike Time Dependent Plasticity (STDP) and Backpropagation (BP) for Robust and Data Efficient Spiking Neural Networks (SNN)

National security applications require artificial neural networks (ANNs) that consume less power, are fast and dynamic online learners, are fault tolerant, and can learn from unlabeled and imbalanced data. We explore whether two fundamentally different, traditional learning algorithms from artificial intelligence and the biological brain can be merged. We tackle this problem from two directions. First, we start from a theoretical point of view and show that the spike time dependent plasticity (STDP) learning curve observed in biological networks can be derived using the mathematical framework of backpropagation through time. Second, we show that transmission delays, as observed in biological networks, improve the ability of spiking networks to perform classification when trained using a backpropagation of error (BP) method. These results provide evidence that STDP could be compatible with a BP learning rule. Combining these learning algorithms will likely lead to networks more capable of meeting our national security missions.

97 MATHEMATICS AND COMPUTING↗

Fast Extraction and Characterization of Fundamental Frequency Events from a Large PMU Dataset using Big Data Analytics

A novel method for fast extraction of fundamental frequency events (FFE) based on measurements of frequency and rate of change of frequency by Phasor Measurement Units (PMU) is introduced. The method is designed to work with exceptionally large historical PMU datasets. Statistical analysis was used to extract the features and train Random Forest and Catboost classifiers. The method is capable of fast extraction of FFE from a historical dataset containing measurements from hundreds of PMUs captured over multiple years. The reported accuracy of the best algorithm for classification expressed as Area Under the receiver operating Characteristic curve reaches 0.98, which was obtained in out-of-sample evaluations on 109 system-wide events over 2 years observed at 43 PMUs. Then Minimum Volume Enclosing Ellipsoid Algorithm was used to further analyze the events. 93.72% events were correctly characterized, where average duration of the event as seen by the PMU was 9.93 sec.

Baembitov, Rashid↗

Fast model-based scenario optimization in NSTX-U enabled by analytic gradient computation

Model-based optimization offers a systematic approach to advanced scenario planning. In this case, the feedforward-control inputs (actuator trajectories) that are needed to attain and sustain a desired scenario are obtained by solving a nonlinear constrained optimization problem. This class of problems generally minimize a cost function that measures the difference between desired and actual plasma states. Several numerical optimization algorithms, such as sequential quadratic programming, require repeated calculation of the cost function gradients with respect to the input trajectories. Calculating these gradients numerically can be computationally intensive, increasing the time needed to solve the feedforward-control optimization problem. Here, this work introduces a method to analytically calculate these cost function gradients from the current profile evolution model. This can significantly reduce the computational time and allow for fast feedforward-control optimization, which would eventually enable optimal scenario planning between discharges. The performance of the feedforward optimizer with analytical gradients is compared to a traditional optimization algorithm based on numerical gradients for different NSTX-U scenarios. The plasma dynamics in the optimization algorithm are simulated using the Control Oriented Transport SIMulator (COTSIM). Results of the work show that analytical gradients consistently reduce the computation time while achieving trajectories that are comparable to those obtained by traditional optimization algorithms based on numerical gradients.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modifying the Asynchronous Jacobi Method for Data Corruption Resilience

Moving scientific computation from high-performance computing (HPC) and cloud computing (CC) environments to devices on the edge, i.e., physically near instruments of interest, has received tremendous interest in recent years. Such edge computing environments can operate on data in situ, offering enticing benefits over data aggregation to HPC and CC facilities that include avoiding costs of transmission, increased data privacy, and real-time data analysis. Because of the inherent unreliability of edge computing environments, new fault-tolerant approaches must be developed before the benefits of edge computing can be realized. Motivated by algorithm-based fault tolerance, a variant of the asynchronous Jacobi (ASJ) method is developed that achieves resilience to data corruption by rejecting solution approximations from neighbor devices according to a bound derived from convergence theory. Numerical results on a two-dimensional Poisson problem show that the new rejection criterion, along with a novel approximation to the shortest path length on which the criterion depends, restores convergence for the ASJ variant in the presence of certain types data corruption. Numerical results are obtained for when the singular values in the analytic bound are approximated. Additional linear systems are also explored, one with a more dense sparsity pattern and one that includes advection. All results indicate that successful resilience to data corruption depends on whether the bound tightens fast enough to reject corrupted data before the iteration evolution deviates significantly from that predicted by the convergence theory defining the bound. This observation generalizes to future work on algorithm-based fault tolerance for other asynchronous algorithms, including upcoming approaches that leverage Krylov subspaces.

97 MATHEMATICS AND COMPUTING↗

A scalable multidimensional fully implicit solver for Hall magnetohydrodynamics

We propose an optimally performant fully implicit algorithm for the Hall magnetohydrodynamics (HMHD) equations based on multigrid-preconditioned Jacobian-free Newton-Krylov methods. HMHD is a challenging system to solve numerically because it supports stiff fast dispersive waves. The preconditioner is formulated using an operator-split approximate block factorization (Schur complement), informed by physics insight. We use a vector-potential formulation (instead of a magnetic field one) to allow a clean segregation of the problematic $\nabla$ x $\nabla$ x operator in the electron Ohm's law subsystem. This segregation allows the formulation of an effective damped block-Jacobi smoother for multigrid. We demonstrate by analysis that our proposed block-Jacobi iteration is convergent and has the smoothing property. The resulting HMHD solver is verified linearly with wave propagation examples, and nonlinearly with the GEM challenge reconnection problem by comparison against another HMHD code. We demonstrate the excellent algorithmic and parallel performance of the algorithm up to 16384 MPI tasks in two dimensions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Fast GPU 3D diffeomorphic image registration

3D image registration is one of the most fundamental and computationally expensive operations in medical image analysis. Here, we present a mixed-precision, Gauss–Newton–Krylov solver for diffeomorphic registration of two images. Our work extends the publicly available CLAIRE library to GPU architectures. Despite the importance of image registration, only a few implementations of large deformation diffeomorphic registration packages support GPUs. Our contributions are new algorithms to significantly reduce the run time of the two main computational kernels in CLAIRE: calculation of derivatives and scattered-data interpolation. Additionally, we deploy (i) highly-optimized, mixed-precision GPU-kernels for the evaluation of scattered-data interpolation, (ii) replace Fast-Fourier-Transform (FFT)-based first-order derivatives with optimized 8th-order finite differences, and (iii) compare with state-of-the-art CPU and GPU implementations. As a highlight, we demonstrate that we can register clinical images in less than 6 s on a single NVIDIA Tesla V100. This amounts to over 20 speed-up over the current version of CLAIRE and over 30 speed-up over existing GPU implementations.

97 MATHEMATICS AND COMPUTING↗

Rapid CFD Using Machine Learning Algorithms (CRADA Final Report)

This is a collaborative effort between Lawrence Livermore National Security, LLC as manager and operator of Lawrence Livermore National Laboratory (“LLNL”) and Guardian Glass, LLC ("Guardian Glass") to develop a fast-running emulator of the reactive Computational Fluid Dynamics (“CFD”) simulations needed to understand the complex reactions and flows in the glass melting, fining, and forming subprocesses. This CRADA project is sponsored under the High-Performance Computing for Manufacturing (“HPC4Mfg”) Program of the Department of Energy’s Advanced Manufacturing Office (“AMO”) within the Energy Efficiency and Renewable Energy (“EERE”) Office.

36 MATERIALS SCIENCE↗

Data reduction through optimized scalar quantization for more compact neural networks

Raw data generation for several existing and planned large physics experiments now exceeds TB/s rates, generating untenable data sets in very little time. Those data often demonstrate high dimensionality while containing limited information. Meanwhile, Machine Learning algorithms are now becoming an essential part of data processing and data analysis. Those algorithms can be used offline for post processing and post data analysis, or they can be used online for real time processing providing ultra low latency experiment monitoring. Both use cases would benefit from data throughput reduction while preserving relevant information: one by reducing the offline storage requirements by several orders of magnitude and the other by allowing ultra fast online inferencing with low complexity Machine Learning models. Moreover, reducing the data source throughput also reduces material cost, power and data management requirements. In this work we demonstrate optimized nonuniform scalar quantization for data source reduction. This data reduction allows lower dimensional representations while preserving the relevant information of the data, thus enabling high accuracy Tiny Machine Learning classifier models for online fast inferences. We demonstrate this approach with an initial proof of concept targeting the CookieBox, an array of electron spectrometers used for angular streaking, that was developed for LCLS-II as an online beam diagnostic tool. We used the Lloyd-Max algorithm with the CookieBox dataset to design an optimized nonuniform scalar quantizer. Optimized quantization lets us reduce input data volume by 69% with no significant impact on inference accuracy. When we tolerate a 2% loss on inference accuracy, we achieved 81% of input data reduction. Finally, the change from a 7-bit to a 3-bit input data quantization reduces our neural network size by 38%.

97 MATHEMATICS AND COMPUTING↗

Scaling pair count to next galaxy surveys

ABSTRACT Counting pairs of galaxies or stars according to their distance is at the core of real-space correlation analyses performed in astrophysics and cosmology. Upcoming galaxy surveys (LSST, Euclid) will measure properties of billions of galaxies challenging our ability to perform such counting in a minute-scale time relevant for the usage of simulations. The problem is only limited by efficient access to the data, hence belongs to the big data category. We use the popular Apache Spark framework to address it and design an efficient high-throughput algorithm to deal with hundreds of millions to billions of input data. To optimize it, we revisit the question of non-hierarchical sphere pixelization based on cube symmetries and develop a new one dubbed the ‘Similar Radius Sphere Pixelization’ (SARSPix) with very close to square pixels. It provides the most adapted indexing over the sphere for all distance-related computations. Using LSST-like fast simulations, we compute autocorrelation functions on tomographic bins containing between a hundred million to one billion data points. In each case, we achieve the construction of a standard pair-distance histogram in about 2 min, using a simple algorithm that is shown to scale, over a moderate number of nodes (16–64). This illustrates the potential of this new techniques in the field of astronomy where data access is becoming the main bottleneck. They can be easily adapted to other use-cases as nearest-neighbours search, catalogue cross-match or cluster finding. The software is publicly available from https://github.com/astrolabsoftware/SparkCorr.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluation of smart charging for electric vehicle-to-building integration: A case study

Higher electric vehicle (EV) adoption will stress the importance of demand flexibility to achieve more economic, efficient, and reliable grid operation. Charging technologies will be paramount in shifting temporally to better fit the variable generation of wind and solar. As such, analysis is warranted on the benefits of EV charge scheduling with respect to installation cost, operation cost, difficulty of implementation, and grid flexibility. We tackle this by analyzing the cost savings of implementing an EV charge scheduling infrastructure to reduce demand charges and installation costs. In this paper, we analyze a case study for operation of 16 level 2 chargers and 1 fast charger for two different building types. We then evaluate various test phases for controlling building and charging loads using an adaptive charging network (ACN) algorithm to characterize the ACN’s potential to reduce overall project cost.

30 DIRECT ENERGY CONVERSION↗