Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel Matrix Multiplication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Multiscale and Multifidelity Modeling of a 3D Woven Composite Thermal Protection System

Complex three-dimensional (3D) woven composites have been considered by multiple NASA projects in recent years as a means of offering improved mechanical and thermal performance over traditional laminated composite systems. Parallel efforts have focused on developing simulation capabilities for these systems, which have traditionally and heavily relied on experimental testing to evaluate composite performance. One system is the Heatshield for Extreme Entry Environment Technology (HEEET), which is being considered for the thermal protection system on reentry spacecraft. Optical microscopy was used to characterize the blended carbon and phenolic fiber tows. A section of HEEET insulation layer was imaged with high-resolution micro-computed tomography (microCT) and segmented to separate individual tows, porous matrix, and voids. These data were used to develop multiscale thermomechanical computational models within the NASA Multiscale Analysis Tool (NASMAT). Two NASMAT modeling approaches were considered to capture the details of the 3D woven architecture: a coarse model appropriate for inclusion in multiscale structural analyses and a high-fidelity model created by downsampling the microCT data. Both elastic and thermal properties were computed and compared. The feasibility and challenges associated with modeling complex, hybrid 3D woven composites were also addressed.

NASMAT↗

GSoFa: Scalable Sparse Symbolic LU Factorization on GPUs

Decomposing a matrix $\mathbf {A}$ into a lower matrix $\mathbf {L}$ and an upper matrix $\mathbf {U}$, which is also known as LU decomposition, is an essential operation in numerical linear algebra. For a sparse matrix, LU decomposition often introduces more nonzero entries in the $\mathbf {L}$ and $\mathbf {U}$ factors than in the original matrix. A symbolic factorization step is needed to identify the nonzero structures of $\mathbf {L}$ and $\mathbf {U}$ matrices. Attracted by the enormous potentials of the Graphics Processing Units (GPUs), an array of efforts have surged to deploy various LU factorization steps except for the symbolic factorization, to the best of our knowledge, on GPUs. This article introduces gSoFa, the first GPU-based symbolic factorization design with the following three optimizations to enable scalable LU symbolic factorization for nonsymmetric pattern sparse matrices on GPUs. First, here we introduce a novel fine-grained parallel symbolic factorization algorithm that is well suited for the Single Instruction Multiple Thread (SIMT) architecture of GPUs. Second, we tailor supernode detection into a SIMT friendly process and strive to balance the workload, minimize the communication and saturate the GPU computing resources during supernode detection. Third, we introduce a three-pronged optimization to reduce the excessive space consumption problem faced by multi-source concurrent symbolic factorization. Taken together, gSoFa achieves up to 31× speedup from 1 to 44 Summit nodes (6 to 264 GPUs) and outperforms the state-of-the-art CPU project, on average, by 5×. Notably, gSoFa also achieves up to 47 percent of the peak memory throughput of a V100 GPU in the Summit Supercomputer.

97 MATHEMATICS AND COMPUTING↗

TUMME: Tsinghua University Minnesota Master Equation program

We report that TUMME is a program for assembling and solving master equations for gas-phase chemical kinetics based on chemically significant eigenmodes. TUMME has interfaces to the Gaussian, Polyrate, and/or MSTor output files that allow the master equation code to obtain the microcanonical flux coefficients needed for the coefficient matrix of the master equation. The flux coefficients for reactions with barriers can be calculated by multi-structural variational transition state theory with small-curvature tunneling (MS-VTST/SCT) or by simpler approximations to this such as conventional transition state theory without tunneling (also called RRKM theory). The flux coefficients for barrierless reactions are provided by a hard-sphere model. TUMME is written in double precision with Python 3; quadruple and octuple precision are also available for some subtasks in C++. The Python code can run in serial or parallel (MP or MPI), and the C++ code can run on a single processor or on multiple processors with OpenMP.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Accelerating Binarized Neural Networks via Bit-Tensor-Cores in Turing GPUs

Despite foreseeing tremendous speedups over conventional deep neural networks, the performance advantage of binarized neural networks (BNNs) has merely been showcased on general-purpose processors such as CPUs and GPUs. In fact, due to being unable to leverage bit-level-parallelism with a word-based architecture, GPUs have been criticized for extremely low utilization (1%) when executing BNNs. Consequently, the latest tensorcores in NVIDIA Turing GPUs start to experimentally support bit computation. In this work, we look into this brand new bit computation capability and characterize its unique features. We show that the stride of memory access can significantly affect performance delivery and a data-format co-design is highly desired to support the tensorcores for achieving superior performance than existing software solutions without tensorcores. We realize the tensorcore-accelerated BNN design, particularly the major functions for fully-connect and convolution layers — bit matrix multiplication and bit convolution. Evaluations on two NVIDIA Turing GPUs show that, with ResNet-18, our BTC-BNN design can process ImageNet at a rate of 5.6K images per second, 77% faster than state-of-the-art. Our BNN approach is released on https://github.com/pnnl/TCBNN.

Li, Ang↗

Exploring the Connection Between Sampling Problems in Bayesian Inference and Statistical Mechanics

The Bayesian and statistical mechanical communities often share the same objective in their work - estimating and integrating probability distribution functions (pdfs) describing stochastic systems, models or processes. Frequently, these pdfs are complex functions of random variables exhibiting multiple, well separated local minima. Conventional strategies for sampling such pdfs are inefficient, sometimes leading to an apparent non-ergodic behavior. Several recently developed techniques for handling this problem have been successfully applied in statistical mechanics. In the multicanonical and Wang-Landau Monte Carlo (MC) methods, the correct pdfs are recovered from uniform sampling of the parameter space by iteratively establishing proper weighting factors connecting these distributions. Trivial generalizations allow for sampling from any chosen pdf. The closely related transition matrix method relies on estimating transition probabilities between different states. All these methods proved to generate estimates of pdfs with high statistical accuracy. In another MC technique, parallel tempering, several random walks, each corresponding to a different value of a parameter (e.g. "temperature"), are generated and occasionally exchanged using the Metropolis criterion. This method can be considered as a statistically correct version of simulated annealing. An alternative approach is to represent the set of independent variables as a Hamiltonian system. Considerab!e progress has been made in understanding how to ensure that the system obeys the equipartition theorem or, equivalently, that coupling between the variables is correctly described. Then a host of techniques developed for dynamical systems can be used. Among them, probably the most powerful is the Adaptive Biasing Force method, in which thermodynamic integration and biased sampling are combined to yield very efficient estimates of pdfs. The third class of methods deals with transitions between states described by rate constants. These problems are isomorphic with chemical kinetics problems. Recently, several efficient techniques for this purpose have been developed based on the approach originally proposed by Gillespie. Although the utility of the techniques mentioned above for Bayesian problems has not been determined, further research along these lines is warranted

Pohorille, Andrew↗

Programmable remapper with single flow architecture

The invention relates to image processing systems and methods and in particular to a machine which accepts a real time video image in the form of a matrix of picture elements (pixels) and remaps such image according to a selectable one of a plurality of mapping functions to create an output matrix of pixels. Such mapping functions, or transformations, may be any one of a number of different transformations depending on the objective of the user of the system. The system remaps input images from one coordinate system to another using a set of look-up tables for the data necessary for the transform. The transforms, which are operator selectable, are precomputed and loaded into massive look-up tables. Input pixels, via the look-up tables of any particular transform selected, are mapped into output pixels with the radiance information of the input pixels being appropriately weighted. An earlier embodiment of the system included two parallel processors: a collective processor which mapped multiple input pixels into a single output pixel and an interpolative processor. The interpolative processor performed an interpolation among pixels in the input image where a given input pixel may affect the value of many output pixels. Several advantages are provided over previous embodiments in that the two distinct processors are replaced by a single processor capable of performing both types of operations (collective and interpolative) with no more complexity. Previously, there has existed no image processor or 'remapper' that can operate with sufficient speed and flexibility to permit investigating different transformation patterns in real time.

Fisher, Timothy E.↗

Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation

This Final Scientific and Technical Report summarizes work performed under the Phase IIA SBIR project “Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation” (DE-SC0020906) from August 2023 through August 2025. The objective of the Phase IIA effort was to mature and harden capabilities developed during Phase II, with the goal of enabling practical interoperability between Coreform’s isogeometric analysis (IGA) technologies and the Multiphysics Object-Oriented Simulation Environment (MOOSE), while improving robustness, performance, and scalability for complex, nuclear-relevant geometries. Over the course of Phase IIA, the project established and validated an extraction-based interoperability pathway between Coreform tools and MOOSE. A combined mesh and matrix format was defined collaboratively with MOOSE developers and integrated into the solver, enabling standard MOOSE workflows to operate on data exported from Coreform’s IGA and Flex Representation Method (FRM) pipelines. Early demonstrations validated architectural compatibility using linear solid mechanics problems, while later efforts focused on benchmark testing and external use. By the end of the project period, engineers at BWXT were able to independently set up and execute a simulation using the Coreform–MOOSE workflow and provide direct feedback that informed further refinement. In parallel, substantial effort was devoted to improving the robustness of trimmed U-spline construction for complex CAD geometries. A growing test suite of nuclear-relevant models was compiled through collaboration with multiple stakeholders and used to drive extensive bug fixing and reliability improvements. These efforts resulted in improved robustness and performance, including the addition of fallback capabilities that enhance reliability when the underlying commercial CAD kernel fails. Performance-oriented work progressed later in the project, with the development and demonstration of methods to decompose complex geometries into structured subregions and updated data representations to support more efficient solver processing. Additionally, extensive enhancements to threadsafe parallel data structures and trimming operations established a foundation for scalable processing of large assemblies. Collaboration with Sandia National Laboratories on the SGM geometric modeling kernel advanced to a functioning interface test case, positioning the workflow for future kernel integration. Overall, the Phase IIA effort successfully transitioned the project from architectural proof-of-concept to externally exercised, solver-integrated capability, while clarifying remaining technical challenges related to standardization, performance optimization, and kernel integration.

42 ENGINEERING↗

Implementation and Assessment of Advanced Analog Vector-Matrix Processor

This paper discusses the design and implementation of an analog optical vecto-rmatrix coprocessor with a throughput of 128 Mops for a personal computer. Vector matrix calculations are inherently parallel, providing a promising domain for the use of optical calculators. However, to date, digital optical systems have proven too cumbersome to replace electronics, and analog processors have not demonstrated sufficient accuracy in large scale systems. The goal of the work described in this paper is to demonstrate a viable optical coprocessor for linear operations. The analog optical processor presented has been integrated with a personal computer to provide full functionality and is the first demonstration of an optical linear algebra processor with a throughput greater than 100 Mops. The optical vector matrix processor consists of a laser diode source, an acoustooptical modulator array to input the vector information, a liquid crystal spatial light modulator to input the matrix information, an avalanche photodiode array to read out the result vector of the vector matrix multiplication, as well as transport optics and the electronics necessary to drive the optical modulators and interface to the computer. The intent of this research is to provide a low cost, highly energy efficient coprocessor for linear operations. Measurements of the analog accuracy of the processor performing 128 Mops are presented along with an assessment of the implications for future systems. A range of noise sources, including cross-talk, source amplitude fluctuations, shot noise at the detector, and non-linearities of the optoelectronic components are measured and compared to determine the most significant source of error. The possibilities for reducing these sources of error are discussed. Also, the total error is compared with that expected from a statistical analysis of the individual components and their relation to the vector-matrix operation. The sufficiency of the measured accuracy of the processor is compared with that required for a range of typical problems. Calculations resolving alloy concentrations from spectral plume data of rocket engines are implemented on the optical processor, demonstrating its sufficiency for this problem. We also show how this technology can be easily extended to a 100 x 100 10 MHz (200 Cops) processor.

Gary, Charles K.↗

Propulsion Electrification Architecture Selection Process and Cost of Carbon Abatement Analysis for Heavy-Duty Off-Road Material Handler

The heavy-duty off-road industry continues to expand efforts to reduce fuel consumption and CO 2 e (carbon dioxide equivalent) emissions. Many manufacturers are pursuing electrification to decrease fuel consumption and emissions. Future policies will likely require electrification for CO2e savings, as seen in light-duty on-road vehicles. Electrified architectures vary widely in the heavy-duty off-road space, with parallel hybrids in some applications and series hybrids in others. The diverse applications for different types of equipment mean different electrified configurations are required. Companies must also determine the value in pursuing electrified architectures; this work analyzes a range of electrified architectures, from micro hybrids to parallel hybrids to series hybrids to a BEV, looking at the total cost, total CO 2 e, and cost per CO 2 e (cost of carbon abatement, or cost of carbon reduction) using data for the year 2021. This study is focused on a heavy-duty off-road material handler, the Pettibone Cary-Lift 204i. This machine’s specialty application, including events like unloading large oil pipes from a railcar, requires a unique electrified architecture that suits its specific needs. However, the results from this study may be extrapolated to similar machinery to inform fuel savings options across the heavy-duty off-road industry. In this study, a unique electrified architecture is determined for the Cary-Lift. This architecture is informed by multiple rounds of a Pugh matrix decision analysis to select a shortened list of desirable electrified architectures. The shortened list is modeled and simulated to determine CO 2 e, cost, and cost per CO 2 e. A final architecture is determined as a plug-in series hybrid that reduces fuel consumption by 65%, targeting the large fuel and CO 2 e savings that are likely to be required for the future of the heavy-duty off-road industry.

33 ADVANCED PROPULSION SYSTEMS↗

Design of a LQR Controller of Reduced Inputs for Multiple Spacecraft Formation Flying

Regarding multiple spacecraft formation flying, the observation is made that control thrust need only be applied coplanar to the local horizon to achieve complete controllability of a two-satellite formation. Without the need for zenith-nadir (radial) thrust, simplifications and reduction of the weight of the propulsion system may be accomplished. This work focuses on the validation of this radial-excluding control system on its own merits, and in comparison to a related system which does provide thrust parallel to the orbital radius. Simulations are performed using commercial ODE solvers to propagate the Keplerian dynamics of a controlled satellite relative to an uncontrolled, leader satellite. The conclusion is drawn that, despite the exclusion of the radial thrust axis, the remaining control thrust available still provides enough control to design a gain matrix of adequate performance using linear-quadratic regulator (LQR) techniques.

Starin, Scott R.↗

Porting fragmentation methods to GPUs using an OpenMP API: Offloading the resolution-of-the-identity second-order Møller–Plesset perturbation method

Here, using an OpenMP Application Programming Interface, the resolution-of-the-identity second-order Møller–Plesset perturbation (RI-MP2) method has been off-loaded onto graphical processing units (GPUs), both as a standalone method in the GAMESS electronic structure program and as an electron correlation energy component in the effective fragment molecular orbital (EFMO) framework. First, a new scheme has been proposed to maximize data digestion on GPUs that subsequently linearizes data transfer from central processing units (CPUs) to GPUs. Second, the GAMESS Fortran code has been interfaced with GPU numerical libraries (e.g., NVIDIA cuBLAS and cuSOLVER) for efficient matrix operations (e.g., matrix multiplication, matrix decomposition, and matrix inversion). The standalone GPU RI-MP2 code shows an increasing speedup of up to 7.5× using one NVIDIA V100 GPU with one IBM 42-core P9 CPU for calculations on fullerenes of increasing size from 40 to 260 carbon atoms using the 6-31G(d)/cc-pVDZ-RI basis sets. A single Summit node with six V100s can compute the RI-MP2 correlation energy of a cluster of 175 water molecules using the correlation consistent basis sets cc-pVDZ/cc-pVDZ-RI containing 4375 atomic orbitals and 14 700 auxiliary basis functions in ~0.85 h. In the EFMO framework, the GPU RI-MP2 component shows near linear scaling for a large number of V100s when computing the energy of an 1800-atom mesoporous silica nanoparticle in a bath of 4000 water molecules. The parallel efficiencies of the GPU RI-MP2 component with 2304 and 4608 V100s are 98.0% and 96.1%, respectively.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MATRIX-VBS (v1.0): Implementing an Evolving Organic Aerosol Volatility in an Aerosol Microphysics Model

The gas-particle partitioning and chemical aging of semi-volatile organic aerosol are presented in a newly developed box model scheme, where its effect on the growth, composition, and mixing state of particles is examined. The volatility-basis set (VBS) framework is implemented into the aerosol microphysical scheme MATRIX (Multiconfiguration Aerosol TRacker of mIXing state), which resolves mass and number aerosol concentrations and in multiple mixing-state classes. The new scheme, MATRIX-VBS, has the potential to significantly advance the representation of organic aerosols in Earth system models by improving upon the conventional representation as non-volatile particulate organic matter, often also with an assumed fixed size distribution. We present results from idealized cases representing Beijing, Mexico City, a Finnish forest, and a southeastern US forest, and investigate the evolution of mass concentrations and volatility distributions for organic species across the gas and particle phases, as well as assessing their mixing state among aerosol populations. Emitted semi-volatile primary organic aerosols evaporate almost completely in the intermediate-volatility range, while they remain in the particle phase in the low-volatility range. Their volatility distribution at any point in time depends on the applied emission factors, oxidation by OH radicals, and temperature. We also compare against parallel simulations with the original scheme, which represented only the particulate and non-volatile component of the organic aerosol, examining how differently the condensed-phase organic matter is distributed across the mixing states in the model. The results demonstrate the importance of representing organic aerosol as a semi-volatile aerosol, and explicitly calculating the partitioning of organic species between the gas and particulate phases.

volatility-basis set↗

Generalizing mkFit and its Application to HL-LHC

mkFit is an implementation of the Kalman filter-based track reconstruction algorithm that exploits both thread- and data-level parallelism. In the past few years the project transitioned from the R&D phase to deployment in the Run-3 offline workflow of the CMS experiment. The CMS tracking performs a series of iterations, targeting reconstruction of tracks of increasing difficulty after removing hits associated to tracks found in previous iterations. mkFit has been adopted for several of the tracking iterations, which contribute to the majority of reconstructed tracks. When tested in the standard conditions for production jobs, speedups in track pattern recognition are on average of the order of 3.5x for the iterations where it is used (3-7x depending on the iteration). Multiple factors contribute to the observed speedups, including vectorization and a lightweight geometry description, as well as improved memory management and single precision. Efficient vectorization is achieved with both the icc and the gcc (default in CMSSW) compilers and relies on a dedicated library for small matrix operations, Matriplex, which has recently been released in a public repository. While the mkFit geometry description already featured levels of abstraction from the actual Phase-1 CMS tracker, several components of the implementations were still tied to that specific geometry. We have further generalized the geometry description and the configuration of the run-time parameters, in order to enable support for the Phase-2 upgraded tracker geometry for the HL-LHC and potentially other detector configurations. The implementation strategy and high-level code changes required for the HL-LHC geometry are presented. Speedups in track building from mkFit imply that track fitting becomes a comparably time consuming step of the tracking chain.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Ordering Unstructured Meshes for Sparse Matrix Computations on Leading Parallel Systems

The ability of computers to solve hitherto intractable problems and simulate complex processes using mathematical models makes them an indispensable part of modern science and engineering. Computer simulations of large-scale realistic applications usually require solving a set of non-linear partial differential equations (PDES) over a finite region. For example, one thrust area in the DOE Grand Challenge projects is to design future accelerators such as the SpaHation Neutron Source (SNS). Our colleagues at SLAC need to model complex RFQ cavities with large aspect ratios. Unstructured grids are currently used to resolve the small features in a large computational domain; dynamic mesh adaptation will be added in the future for additional efficiency. The PDEs for electromagnetics are discretized by the FEM method, which leads to a generalized eigenvalue problem Kx = AMx, where K and M are the stiffness and mass matrices, and are very sparse. In a typical cavity model, the number of degrees of freedom is about one million. For such large eigenproblems, direct solution techniques quickly reach the memory limits. Instead, the most widely-used methods are Krylov subspace methods, such as Lanczos or Jacobi-Davidson. In all the Krylov-based algorithms, sparse matrix-vector multiplication (SPMV) must be performed repeatedly. Therefore, the efficiency of SPMV usually determines the eigensolver speed. SPMV is also one of the most heavily used kernels in large-scale numerical simulations.

Oliker, Leonid↗

Versatile soil gas concentration and isotope monitoring: optimization and integration of novel soil gas probes with online trace gas detection

Abstract. Gas concentrations and isotopic signatures can unveil microbial metabolisms and their responses to environmental changes in soil. Currently, few methods measure in situ soil trace gases such as the products of nitrogen and carbon cycling or volatile organic compounds (VOCs) that constrain microbial biochemical processes like nitrification, methanogenesis, respiration, and microbial communication. Versatile trace gas sampling systems that integrate soil probes with sensitive trace gas analyzers could fill this gap with in situ soil gas measurements that resolve spatial (centimeters) and temporal (minutes) patterns. We developed a system that integrates new porous and hydrophobic sintered polytetrafluoroethylene (sPTFE) diffusive soil gas probes that non-disruptively collect soil gas samples with a transfer system to direct gas from multiple probes to one or more central gas analyzer(s) such as laser and mass spectrometers. Here, we demonstrate the feasibility and versatility of this automated multiprobe system for soil gas measurements of isotopic ratios of nitrous oxide (δ18O, δ15N, and the 15N site preference of N2O), methane, carbon dioxide (δ13C), and VOCs. First, we used an inert silica matrix to challenge probe measurements under controlled gas conditions. By changing and controlling system flow parameters, including the probe flow rate, we optimized recovery of representative soil gas samples while reducing sampling artifacts on subsurface concentrations. Second, we used this system to provide a real-time window into the impact of environmental manipulation of irrigation and soil redox conditions on in situ N2O and VOC concentrations. Moreover, to reveal the dynamics in the stable isotope ratios of N2O (i.e., 14N14N16O, 14N15N16O, 15N14N16O, and 14N14N18O), we developed a new high-precision laser spectrometer with a reduced sample volume demand. Our integrated system – a tunable infrared laser direct absorption spectrometry (TILDAS) in parallel with Vocus proton transfer reaction mass spectrometry (PTR-MS), in line with sPTFE soil gas probes – successfully quantified isotopic signatures for N2O, CO2, and VOCs in real time as responses to changes in the dry–wetting cycle and redox conditions. Broadening the collection of trace gases that can be monitored in the subsurface is critical for monitoring biogeochemical cycles, ecosystem health, and management practices at scales relevant to the soil system.

54 ENVIRONMENTAL SCIENCES↗

Exploring the Use of Novel Spatial Accelerators in Scientific Applications

Driven by the need to find alternative accelerators which can viably replace GPUs in next-generation Supercomputing systems, this paper proposes a methodology to enable agile application/hardware co-design. The application-first methodology provides the ability to come up with design of accelerators while working with real-world workloads, available accelerators, and system software. The iterative design process targets a set of kernels in a workload for performance estimates that can prune the design space for later phases of detailed architectural evaluations. To this effect, in this paper, a novel data-parallel device model is introduced that simulates the latency of performance-sensitive operations in an accelerator including data transfers and kernel computation using multi-core CPUs. The use of off-the-shelf simulators, such as pre-RTL simulator Aladdin or multiple tools available for exploring the design of deep neural network accelerators (e.g., Timeloop) is demonstrated for evaluation of various accelerator designs using applications with realistic inputs. Examples of multiple device configurations that are instantiable in a system are explored to evaluate the performance benefit of deploying novel accelerators. The proposed device is integrated with a programming model and system software to potentially explore the impacts of high-level programming languages/compilers and low-level effects such as task scheduling on multiple accelerators. We analyze our methodology for a set of applications that represent high-performance computing (HPC) and graph analytics. The applications include a computational chemistry kernel realized using tensor contractions, triangle counting, GraphSAGE and Breadth-first Search. These applications include kernels such as dense matrix-dense matrix multiplication, sparse matrix-spare matrix multiplication, and sparse matrix-dense vector multiplication. Our results indicate potential performance benefits and insights for system design by including accelerators that realize these kernels along-side general purpose accelerators.

AI, codesign, Accelerated Computing, Modeling and ↗

Sensitivity and reliability of key electrochemical markers for detecting lithium plating during extreme fast charging

Lithium plating is one of the key challenges for enabling extreme fast charging (XFC, ≤10 to 15 min charging at ≥6C) in graphite-based lithium-ion batteries. Significant R&D effort has been focused on how to mitigate Li plating. Parallel effort is also being devoted to developing methods to detect Li plating when and if it happens during fast charging. In that regard, electrochemical (EC) signature-based detection techniques are less resource intensive, more convenient, and more practical from an end-user application perspective. However, a comprehensive understanding of key plating related EC signatures for extreme fast charging is presently unavailable. In particular, there exist distinct issues of unreliability with key plating-related EC signatures—e.g., incremental capacity (dQ.dV -1 ), differential OCV (dOCV.dt -1 ), end of lithiation (EOL) rest voltage—at XFC conditions, and the underlying reasons have not been explored and identified methodically. Using a comprehensive test matrix and XFC conditions with Li/graphite half cells, this article highlights the unreliability issues associated with the EC Li plating diagnostics and explains the underlying root cause. This study finds distinct sensitivity and unreliability issues with plating related dQ.dV -1 , dOCV.dt -1 , and EOL rest voltage signatures with charging rates. Furthermore, the complex interaction between graphite and plated Li that happens through multiple competing mechanisms —Li stripping and chemical intercalation— at different charging rates is at the core of the sensitivity and unreliability issue.

25 ENERGY STORAGE↗

Surrogate models for plasma displacement and current in 3D perturbed magnetohydrodynamic equilibria in tokamaks

Abstract A numerical database of over one thousand perturbed three-dimensional (3D) equilibria has been generated, constructed based on the MARS-F (Liu et al 2000 Phys. Plasmas 7 3681) computed plasma response to the externally applied 3D field sources in multiple tokamak devices. Perturbed 3D equilibria with the n = 1–4 ( n is the toroidal mode number) toroidal periodicity are computed. Surrogate models are created for the computed perturbed 3D equilibrium utilizing model order reduction (MOR) techniques. In particular, retaining the first few eigenstates from the singular value decomposition (SVD) of the data is found to produce reasonably accurate MOR-representations for the key perturbed quantities, such as the perturbed parallel plasma current density and the plasma radial displacement. SVD also helps to reveal the core versus edge plasma response to the applied 3D field. For the database covering the conventional aspect ratio devices, about 95% of data can be represented by the truncated SVD-series with inclusion of only the first five eigenstates, achieving a relative error (RE) below 20%. The MOR-data is further utilized to train neural networks (NNs) to enable fast reconstruction of perturbed 3D equilibria, based on the two-dimensional equilibrium input and the 3D source field. The best NN-training is achieved for the MOR-data obtained with a global SVD approach, where the full set of samples used for NN training and testing are stretched and form a large matrix which is then subject to SVD. The fully connected multi-layer perceptron, with one or two hidden layers, can be trained to predict the MOR-data with less than 10% RE. As a key insight, a better strategy is to train separate NNs for the plasma response fields with different toroidal mode numbers. It is also better to apply MOR and to subsequently train NNs separately for conventional and low aspect ratio devices, due to enhanced toroidal coupling of Fourier spectra in the plasma response in the latter case.

3D equilibrium↗