Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed numerical linear algebra”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

34 records · Page 2

Numerical simulation of the nonlinear response of composite plates under combined thermal and acoustic loading

A time-domain study of the random response of a laminated plate subjected to combined acoustic and thermal loads is carried out. The features of this problem also include given uniform static inplane forces. The formulation takes into consideration a possible initial imperfection in the flatness of the plate. High decibel sound pressure levels along with high thermal gradients across thickness drive the plate response into nonlinear regimes. This calls for the analysis to use von Karman large deflection strain-displacement relationships. A finite element model that combines the von Karman strains with the first-order shear deformation plate theory is developed. The development of the analytical model can accommodate an anisotropic composite laminate built up of uniformly thick layers of orthotropic, linearly elastic laminae. The global system of finite element equations is then reduced to a modal system of equations. Numerical simulation using a single-step algorithm in the time-domain is then carried out to solve for the modal coordinates. Nonlinear algebraic equations within each time-step are solved by the Newton-Raphson method. The random gaussian filtered white noise load is generated using Monte Carlo simulation. The acoustic pressure distribution over the plate is capable of accounting for a grazing incidence wavefront. Numerical results are presented to study a variety of cases.

Mei, Chuh↗

Research in Computational Aeroscience Applications Implemented on Advanced Parallel Computing Systems

Improving the numerical linear algebra routines for use in new Navier-Stokes codes, specifically Tim Barth's unstructured grid code, with spin-offs to TRANAIR is reported. A fast distance calculation routine for Navier-Stokes codes using the new one-equation turbulence models is written. The primary focus of this work was devoted to improving matrix-iterative methods. New algorithms have been developed which activate the full potential of classical Cray-class computers as well as distributed-memory parallel computers.

Wigton, Larry↗

The effect of adhesive layer on crack propagation in laminates

The effect of the adhesive layer on crack propagation in composite materials is investigated. The composite medium consists of parallel load carrying laminates and buffer strips arranged periodically and bonded with thin adhesive layers. The strips, assumed to be isotropic and linearly elastic, contain symmetric cracks of arbitrary lengths located normal to the interfaces. Two problems are considered: (1) thin adhesive layers are approximated by uncoupled tension and shear springs distributed along the interfaces of the strips for which only the case of internal cracks can be treated rigorously; (2) broken laminates and the true singular behavior in the presence of the adhesive layer are studied. The adhesive is then treated as an isotropic, linearly elastic continuum. General expressions for field quantities are obtained in terms of infinite Fourier integrals. These expressions give a system of singular integral equations in terms of the crack surface displacement derivatives. By using appropriate quadrature formulas, the integral equations reduce to a system of linear algebraic equations which are solved numerically.

Gecit, M. R.↗

Failure Probability Constrained AC Optimal Power Flow

Despite cascading failures being the central cause of blackouts in power transmission systems, existing operational and planning decisions are made largely by ignoring their underlying cascade potential. This paper posits a reliability-aware AC Optimal Power Flow formulation that seeks to design a dispatch point which has a low operator-specified likelihood of triggering a cascade starting from any single component outage. By exploiting a recently developed analytical model of the probability of component failure, our Failure Probability-constrained ACOPF (FP-ACOPF) utilizes the system's expected first failure time as a smoothly tunable and interpretable signature of cascade risk. Here, we use techniques from bilevel optimization and numerical linear algebra to efficiently formulate and solve the FP-ACOPF using off-the-shelf solvers. Extensive simulations on the IEEE 118-bus case show that, when compared to the unconstrained and N-1 security-constrained ACOPF, our probability-constrained dispatch points can significantly lower the probabilities of long severe cascades and of large demand losses, while incurring only minor increases in total generation costs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An efficient algorithm for estimating noise covariances in distributed systems

An efficient computational algorithm for estimating the noise covariance matrices of large linear discrete stochatic-dynamic systems is presented. Such systems arise typically by discretizing distributed-parameter systems, and their size renders computational efficiency a major consideration. The proposed adaptive filtering algorithm is based on the ideas of Belanger, and is algebraically equivalent to his algorithm. The earlier algorithm, however, has computational complexity proportional to p to the 6th, where p is the number of observations of the system state, while the new algorithm has complexity proportional to only p-cubed. Further, the formulation of noise covariance estimation as a secondary filter, analogous to state estimation as a primary filter, suggests several generalizations of the earlier algorithm. The performance of the proposed algorithm is demonstrated for a distributed system arising in numerical weather prediction.

Dee, D. P.↗

MAPPRAISER: A massively parallel map-making framework for multi-kilo pixel CMB experiments

Forthcoming cosmic microwave background (CMB) polarized anisotropy experiments have the potential to revolutionize our understanding of the Universe and fundamental physics. The sought-after, tale-telling signatures will be however distributed over voluminous data sets which these experiments will collect. These data sets will need to be efficiently processed and unwanted contributions due to astrophysical, environmental, and instrumental effects characterized and efficiently mitigated in order to uncover the signatures. This poses a significant challenge to data analysis methods, techniques, and software tools which will not only have to be able to cope with huge volumes of data but to do so with unprecedented precision driven by the demanding science goals posed for the new experiments. A keystone of efficient CMB data analysis is solvers of very large linear systems of equations. Such systems appear in very diverse contexts throughout CMB data analysis pipelines, however they typically display similar algebraic structures and can therefore be solved using similar numerical techniques. Linear systems arising in the so-called map-making problem are one of the most prominent and common ones. In this work we present a massively parallel, flexible and extensible framework, comprised of a numerical library, MIDAPACK, and a high level code, MAPPRAISER, which provide tools for solving efficiently such systems. Here, the framework implements iterative solvers based on conjugate gradient techniques: enlarged and preconditioned using different preconditioners. We demonstrate the framework on simulated examples reflecting basic characteristics of the forthcoming data sets issued by ground-based and satellite-borne instruments, executing it on as many as 16,384 compute cores. The software is developed as an open source project freely available to the community at: https://github.com/B3Dcmb/midapack.

79 ASTRONOMY AND ASTROPHYSICS↗

A Model Predictive Control to Improve Grid Resilience

The following article details a model predictive control (MPC) to improve grid resilience when faced with variable generation resources. This topic is of significant interest to utility power systems where distributed intermittent energy sources will increase significantly and be relied on for electric grid ancillary services. Previous work on MPCs has focused on narrowly targeted control applications such as improving electric vehicle (EV) charging infrastructure or reducing the cost of integrating Energy Storage Systems (ESSs) into the grid. In contrast, this article develops a comprehensive treatment of the construction of an MPC tailored to electric grids and then applies it integration of intermittent energy resources. To accomplish this, the following article includes a description of a reduced order model (ROM) of an electric power grid based on a circuit model, an optimization formulation that describes the MPC, a collocation method for solving linear time-dependent differential algebraic equations (DAEs) that result from the ROM, and an overall strategy for iteratively refining the behavior of the MPC. Next, the algorithm is validated using two separate numerical experiments. First, the algorithm is compared to an existing MPC code and the results are verified by a numerically precise simulation. It is shown that this algorithm produces a control comparable to existing algorithms and the behavior of the control carefully respects the bounds specified. Second, the MPC is applied to a small nine bus system that contains a mix of turbine-spinning-machine-based and intermittent generation in order to demonstrate the algorithm’s utility for resource planning and control of intermittent resources. This study demonstrates how the MPC can be tuned to change the behavior of the control, which can then assist with the integration of intermittent resources into the grid. The emphasis throughout the paper is to provide systematic treatment of the topic and produce a novel nonlinear control compatible design framework applicable to electric grids and the control of variable resources. This differs from the more targeted application-based focus in most presentations.

microgrid↗

Finite-analytic numerical method for unsteady two-dimensional Navier-Stokes equations

A finite analytic (FA) numerical solution is developed for unsteady two-dimensional Navier-Stokes equations. The FA method utilizes the analytic solution in a small local element to formulate the algebraic representation of partial differential equations. The combination of linear and exponential functions that satisfy the governing equation is adopted as the boundary function, thereby improving the accuracy of the finite analytic solution. Two flows, one a starting cavity flow and the other a vortex shedding flow behind a rectangular block, are solved by the FA method. The starting square cavity flow is solved for Reynolds number of 400, 1000, and 2000 to show the accuracy and stability of the FA solution. The FA solution for flow over a rectangular block (H x H/4) predicts the Strouhal number for Reynolds numbers of 100 and 500 to be 0.156 and 0.125. Details of the flow patterns are given. In addition to streamlines and vorticity distribution, rest-streamlines are given to illustrate the vortex motion downstream of the block.

Chen, C.-J.↗

Mathematical solutions in internal dose assessment: A comparison of Python-based differential equation solvers in biokinetic modeling

Abstract In biokinetic modeling systems employed for radiation protection, biological retention and excretion have been modeled as a series of discretized compartments representing the organs and tissues of the human body. Fractional retention and excretion in these organ and tissue systems have been mathematically governed by a series of coupled first-order ordinary differential equations (ODEs). The coupled ODE systems comprising the biokinetic models are usually stiff due to the severe difference between rapid and slow transfers between compartments. In this study, the capabilities of solving a complex coupled system of ODEs for biokinetic modeling were evaluated by comparing different Python programming language solvers and solving methods with the motivation of establishing a framework that enables multi-level analysis. The stability of the solvers was analyzed to select the best performers for solving the biokinetic problems. A Python-based linear algebraic method was also explored to examine how the numerical methods deviated from an analytical or semi-analytical method. Results demonstrated that customized implicit methods resulted in an enhanced stable solution for the inhaled 60 Co (Type M) and 131 I (Type F) exposure scenarios for the inhalation pathway of the International Commission on Radiological Protection (ICRP) Publication 130 Human Respiratory Tract Model (HRTM). The customized implementation of the Python-based implicit solvers resulted in approximately consistent solutions with the Python-based matrix exponential method ( expm ). The differences generally observed between the implicit solvers and expm are attributable to numerical precision and the order of numerical approximation of the numerical solvers. This study provides the first analysis of a list of Python ODE solvers and methods by comparing their usage for solving biokinetic models using the ICRP Publication 130 HRTM and provides a framework for the selection of the most appropriate ODE solvers and methods in Python language to implement for modeling the distribution of internal radioactivity.

61 RADIATION PROTECTION AND DOSIMETRY↗

GPU acceleration of all-electron electronic structure theory using localized numeric atom-centered basis functions

We present an implementation of all-electron density-functional theory for massively parallel GPU-based platforms, using localized atom-centered basis functions and real-space integration grids. Special attention is paid to domain decomposition of the problem on non-uniform grids, which enables compute- and memory-parallel execution across thousands of nodes for real-space operations, e.g. the update of the electron density, the integration of the real-space Hamiltonian matrix, and calculation of Pulay forces. To assess the performance of our GPU implementation, we performed benchmarks on three different architectures using a 103-material test set. We find that operations which rely on dense serial linear algebra show dramatic speedups from GPU acceleration: in particular, SCF iterations including force and stress calculations exhibit speedups ranging from 4.5 to 6.6. For the architectures and problem types investigated here, this translates to an expected overall speedup between 3–4 for the entire calculation (including non-GPU accelerated parts), for problems featuring several tens to hundreds of atoms. Additional calculations for a 375-atom Bi2Se3 bilayer show that the present GPU strategy scales for large-scale distributed-parallel simulations.

42 ENGINEERING↗

A Lightning Channel Retrieval Algorithm for the North Alabama Lightning Mapping Array (LMA)

A new multi-station VHF time-of-arrival (TOA) antenna network is, at the time of this writing, coming on-line in Northern Alabama. The network, called the Lightning Mapping Array (LMA), employs GPS timing and detects VHF radiation from discrete segments (effectively point emitters) that comprise the channel of lightning strokes within cloud and ground flashes. The network will support on-going ground validation activities of the low Earth orbiting Lightning Imaging Sensor (LIS) satellite developed at NASA Marshall Space Flight Center (MSFC) in Huntsville, Alabama. It will also provide for many interesting and detailed studies of the distribution and evolution of thunderstorms and lightning in the Tennessee Valley, and will offer many interesting comparisons with other meteorological/geophysical wets associated with lightning and thunderstorms. In order to take full advantage of these benefits, it is essential that the LMA channel mapping accuracy (in both space and time) be fully characterized and optimized. In this study, a new revised channel mapping retrieval algorithm is introduced. The algorithm is an extension of earlier work provided in Koshak and Solakiewicz (1996) in the analysis of the NASA Kennedy Space Center (KSC) Lightning Detection and Ranging (LDAR) system. As in the 1996 study, direct algebraic solutions are obtained by inverting a simple linear system of equations, thereby making computer searches through a multi-dimensional parameter domain of a Chi-Squared function unnecessary. However, the new algorithm is developed completely in spherical Earth-centered coordinates (longitude, latitude, altitude), rather than in the (x, y, z) cartesian coordinates employed in the 1996 study. Hence, no mathematical transformations from (x, y, z) into spherical coordinates are required (such transformations involve more numerical error propagation, more computer program coding, and slightly more CPU computing time). The new algorithm also has a more realistic definition of source altitude that accounts for Earth oblateness (this can become important for sources that are hundreds of kilometers away from the network). In addition, the new algorithm is being applied to analyze computer simulated LMA datasets in order to obtain detailed location/time retrieval error maps for sources in and around the LMA network. These maps will provide a more comprehensive analysis of retrieval errors for LMA than the 1996 study did of LDAR retrieval errors. Finally, we note that the new algorithm can be applied to LDAR, and essentially any other multi-station TWA network that depends on direct line-of-site antenna excitation.

Koshak, William↗

Compressed basis GMRES on high-performance graphics processing units

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

97 MATHEMATICS AND COMPUTING↗

HHL algorithm with mapping function and enhanced sampling for model predictive control in microgrids

Here, this paper presents a refined quantum Harrow Hassidim Lloyd (HHL) algorithm for microgrid control. The first novelty of the developed method is that a mapping shift function enables the original HHL algorithm to handle general linear equations with non-singular and indefinite matrix. Second, a method of Matrix Extension for Amplifying Sampling Probabilities of Intended Solution (ME-ASPI) is proposed to design the reformulated linear algebraic equations, allowing for improved sampling efficiency of the quantum tomography in the refined HHL algorithm. Then, we applied the method to solve the model predictive control (MPC) problem in nonlinear dynamical microgrids. Specifically, with the ME-ASPI method, the refined HHL algorithm can effectively obtain the intended partial optimal control inputs for MPC. The optimization of quadratic programming problem in each time step of MPC is transformed into a linear system problem, which is addressed by the proposed quantum solver through using only partial information, with the time complexity improved from $\mathscr{O}(\mathscr{N}^{2.37286})$ classically to $\mathscr{O}(\mathscr{N}^{2} log \mathscr{N}$ x $p$ log $p)$ in quantum. Numerical examples have validated the effectiveness of the refined HHL algorithm with the proposed mapping function and the ME-ASPI method. By leveraging quantum properties, the proposed method provides a hybrid quantum–classical framework for microgrid control. This generic method can also potentially tackle many other challenges in analyzing and controlling general complex engineered systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A parallel, distributed memory implementation of the adaptive sampling configuration interaction method

The many-body simulation of quantum systems is an active field of research that involves several different methods targeting various computing platforms. Many methods commonly employed, particularly coupled cluster methods, have been adapted to leverage the latest advances in modern high-performance computing. Selected configuration interaction (sCI) methods have seen extensive usage and development in recent years. However, the development of sCI methods targeting massively parallel resources has been explored only in a few research works. Here, we present a parallel, distributed memory implementation of the adaptive sampling configuration interaction approach (ASCI) for sCI. In particular, we will address the key concerns pertaining to the parallelization of the determinant search and selection, Hamiltonian formation, and the variational eigenvalue calculation for the ASCI method. Load balancing in the search step is achieved through the application of memory-efficient determinant constraints originally developed for the ASCI-PT2 method. The presented benchmarks demonstrate near optimal speedup for ASCI calculations of Cr 2 (24e, 30o) with 10 6 , 10 7 , and 3 × 10 8 variational determinants on up to 16 384 CPUs. Importantly, to the best of the authors’ knowledge, this is the largest variational ASCI calculation to date.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GPU-acceleration of the ELPA2 distributed eigensolver for dense symmetric and hermitian eigenproblems

The solution of eigenproblems is often a key computational bottleneck that limits the tractable system size of numerical algorithms, among them electronic structure theory in chemistry and in condensed matter physics. Large eigenproblems can easily exceed the capacity of a single compute node, thus must be solved on distributed-memory parallel computers. We here present GPU-oriented optimizations of the ELPA two-stage tridiagonalization eigensolver (ELPA2). On top of cuBLAS-based GPU offloading, we add a CUDA kernel to speed up the back-transformation of eigenvectors, which can be the computationally most expensive part of the two-stage tridiagonalization algorithm. Furthermore, we benchmark the performance of this GPU-accelerated eigensolver on two hybrid CPU–GPU architectures, namely a compute cluster based on Intel Xeon Gold CPUs and NVIDIA Volta GPUs, and the Summit supercomputer based on IBM POWER9 CPUs and NVIDIA Volta GPUs. Consistent with previous benchmarks on CPU-only architectures, the GPU-accelerated two-stage solver exhibits a parallel performance superior to the one-stage counterpart. Finally, we demonstrate the performance of the GPU-accelerated eigensolver developed in this work for routine semi-local KS-DFT calculations comprising thousands of atoms.

97 MATHEMATICS AND COMPUTING↗

Machine learning Frenkel Hamiltonian parameters to accelerate simulations of exciton dynamics

In this manuscript, we develop multiple machine learning (ML) models to accelerate a scheme for parameterizing site-based models of exciton dynamics from all-atom configurations of condensed phase sexithiophene systems. This scheme encodes the details of a system’s specific molecular morphology in the correlated distributions of model parameters through the analysis of many single-molecule excited-state electronic-structure calculations. These calculations yield excitation energies for each molecule in the system and the network of pair-wise intermolecular electronic couplings. In this work, we demonstrate that the excitation energies can be accurately predicted using a kernel ridge regression (KRR) model with Coulomb matrix featurization. We present two ML models for predicting intermolecular couplings. The first one utilizes a deep neural network and bi-molecular featurization to predict the coupling directly, which we find to perform poorly. The second one utilizes a KRR model to predict unimolecular transition densities, which can subsequently be analyzed to compute the coupling. We find that the latter approach performs excellently, indicating that an effective, generalizable strategy for predicting simple bimolecular properties is through the indirect application of ML to predict higher-order unimolecular properties. Such an approach necessitates a much smaller feature space and can incorporate the insight of well-established molecular physics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗