Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse grids”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Porting the Nonlinear Optimization Library HiOp to Accelerator-Based Hardware Architectures

While interior point method has been the centerpiece of nonlinear programming tools used in science and engineering, its reliance on linear solvers that can tackle sparse symmetric indefinite and highly ill-conditioned problems made it difficult to implement it effectively on hardware accelerators. HiOp optimization package attempts to provide an implementation of the interior point method suitable for hardware accelerators by compressing the original sparse problem to produce an underlying linear problem that is dense and of manageable size. Implementations of dense linear solvers are more mature and utilize hardware accelerators better than their sparse counterparts. There is a number of important domain problems, such as optimal power flow analysis for power grids, where the sparse problem can be effectively compressed and deploying dense linear solver within the interior point method can improve performance. Here we describe a portable implementation of HiOp optimization engine, which uses a linear solver from Magma library and runs entirely on hardware accelerators. To compress the problem, HiOp uses customized mixed dense-sparse linear algebra. All HiOp kernels are implemented using Umpire and RAJA portability libraries. We describe details of the implementation and discuss trade-offs between performance, portability and development cost.

97 MATHEMATICS AND COMPUTING↗

Spatial and temporal variation in forest transpiration across a forested boreal peatland complex

Transpiration is a globally important component of evapotranspiration. Careful upscaling of transpiration from point measurements is thus crucial for quantifying water and energy fluxes. In spatially heterogeneous landscapes common across the boreal biome, upscaled transpiration estimates are difficult to determine due to variation in local environmental conditions (e.g., basal area, soil moisture, permafrost). Here, we sought to determine stand-level attributes that influence transpiration scalars for a forested boreal peatland complex consisting of sparsely treed wetlands and densely treed permafrost plateaus as land cover types. The objectives were to quantify spatial and temporal variability in stand-level transpiration, and to identify sources of uncertainty when scaling point measurements to the stand-level. Using heat ratio method sap flow sensors, we determined sap velocity for black spruce and tamarack for 2-week periods during peak growing season in 2013, 2017 and 2018. We found greater basal area, drier soils, and the presence of permafrost increased daily sap velocity in individual trees, suggesting that local environmental conditions are important in dictating sap velocity. When sap velocity was scaled to stand-level transpiration using gridded 20 × 20 m resolution data across the ~10 ha Scotty Creek ForestGEO plot, we observed significant differences in daily plot transpiration among years (0.17–0.30 mm), and across land cover types. Daily transpiration was lowest in grid-cells with sparsely treed wetlands compared to grid-cells with well-drained and densely treed permafrost plateaus, where daily transpiration reached 0.80 mm, or 30% of the daily evapotranspiration. When transpiration scalars (i.e., sap velocity) were not specific to the different land cover types (i.e., permafrost plateaus and wetlands), scaled stand-level transpiration was overestimated by 42%. To quantify the relative contribution of tree transpiration to ecosystem evapotranspiration, we recommend that sampling designs stratify across local environmental conditions to accurately represent variation associated with land cover types, especially with different hydrological functioning as encountered in rapidly thawing boreal peatland complexes.

54 ENVIRONMENTAL SCIENCES↗

Algorithms and Application of Sparse Matrix Assembly and Equation Solvers for Aeroacoustics

An algorithm for symmetric sparse equation solutions on an unstructured grid is described. Efficient, sequential sparse algorithms for degree-of-freedom reordering, supernodes, symbolic/numerical factorization, and forward backward solution phases are reviewed. Three sparse algorithms for the generation and assembly of symmetric systems of matrix equations are presented. The accuracy and numerical performance of the sequential version of the sparse algorithms are evaluated over the frequency range of interest in a three-dimensional aeroacoustics application. Results show that the solver solutions are accurate using a discretization of 12 points per wavelength. Results also show that the first assembly algorithm is impractical for high-frequency noise calculations. The second and third assembly algorithms have nearly equal performance at low values of source frequencies, but at higher values of source frequencies the third algorithm saves CPU time and RAM. The CPU time and the RAM required by the second and third assembly algorithms are two orders of magnitude smaller than that required by the sparse equation solver. A sequential version of these sparse algorithms can, therefore, be conveniently incorporated into a substructuring for domain decomposition formulation to achieve parallel computation, where different substructures are handles by different parallel processors.

Watson, W. R.↗

Glassoformer: A Query-Sparse Transformer for Post-Fault Power Grid Voltage Prediction

Here, we propose GLassoformer, a novel and efficient transformer architecture leveraging group Lasso regularization to reduce the number of queries of the standard self-attention mechanism. Due to the sparsified queries, GLassoformer is more computationally efficient than the standard transformers. On the power grid post-fault voltage prediction task, GLasso-former shows remarkably better prediction than many existing benchmark algorithms in terms of accuracy and stability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Physics-Informed Gaussian Process Regression for States Estimation and Forecasting in Power Grids

Real-time state estimation and forecasting are critical for the efficient operation of power grids. In this paper, a physics-informed Gaussian process regression (PhI-GPR) method is presented and used for forecasting and estimating the phase angle, angular speed, and wind mechanical power of a three-generator power grid system using sparse measurements. In standard data-driven Gaussian process regression (GPR), parameterized models for the prior statistics are fit by maximizing the marginal likelihood of observed data. In the PhI-GPR method, we propose to compute the prior statistics offline by solving stochastic differential equations (SDEs) governing the power grid dynamics. The short-term forecast of a power grid system dominated by wind generation is complicated by the stochastic nature of the wind and the resulting uncertainty in wind mechanical power. Here, we assume that the power grid dynamics are governed by swing equations, with the wind mechanical power fluctuating randomly in time. We solve these equations for the mean and covariances of the power grid states using the Monte Carlo simulation method. We demonstrate that the proposed PhI-GPR method can accurately forecast and estimate observed and unobserved states. For the considered problem, PhI-GPR has computational advantages over the ensemble Kalman filter (EnKF) method: In PhI-GPR, ensembles are computed offline and independently of the data acquisition process, whereas for EnFK, ensembles are computed online with data acquisition, rendering real-time forecast more challenging. We also demonstrate that the PhI-GPR forecast is more accurate than the EnKF forecast when the random mechanical wind power is non-Markovian. In contrast, the two methods produce similar forecasts for the Markovian mechanical wind power. For observed states, we show that PhI-GPR provides a forecast comparable to the standard data-driven GPR; both forecasts are significantly more accurate than the autoregressive integrated moving average (ARIMA) forecast. We also show that the ARIMA forecast is more sensitive to observation frequency and measurement errors than the PhI-GPR forecast.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Two-dimensional radiative transfer in cloudy atmospheres - The spherical harmonic spatial grid method

A new two-dimensional monochromatic method that computes the transfer of solar or thermal radiation through atmospheres with arbitrary optical properties is described. The model discretizes the radiative transfer equation by expanding the angular part of the radiance field in a spherical harmonic series and representing the spatial part with a discrete grid. The resulting sparse coupled system of equations is solved iteratively with the conjugate gradient method. A Monte Carlo model is used for extensive verification of outgoing flux and radiance values from both smooth and highly variable (multifractal) media. The spherical harmonic expansion naturally allows for different levels of approximation, but tests show that the 2D equivalent of the two-stream approximation is poor at approximating variations in the outgoing flux. The model developed here is shown to be highly efficient so that media with tens of thousands of grid points can be computed in minutes. The large improvement in efficiency will permit quick, accurate radiative transfer calculations of realistic cloud fields and improve our understanding of the effect of inhomogeneity on radiative transfer in cloudy atmospheres.

Evans, K. F.↗

Greedy permanent magnet optimization

Abstract A number of scientific fields rely on placing permanent magnets in order to produce a desired magnetic field. We have shown in recent work that the placement process can be formulated as sparse regression. However, binary, grid-aligned solutions are desired for realistic engineering designs. We now show that the binary permanent magnet problem can be formulated as a quadratic program with quadratic equality constraints, the binary, grid-aligned problem is equivalent to the quadratic knapsack problem with multiple knapsack constraints, and the single-orientation-only problem is equivalent to the unconstrained quadratic binary problem. We then provide a set of simple greedy algorithms for solving variants of permanent magnet optimization, and demonstrate their capabilities by designing magnets for stellarator plasmas. The algorithms can a-priori produce sparse, grid-aligned, binary solutions. Despite its simple design and greedy nature, we provide an algorithm that compares with or even outperforms the state-of-the-art algorithms while being substantially faster, more flexible, and easier to use.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Physics-informed graphical neural network for power system state estimation

State estimation is highly critical for accurately observing the dynamic behavior of the power grids and minimizing risks from cyber threats. However, existing state estimation methods encounter challenges in accurately capturing power system dynamics, primarily because of limitations in encoding the grid topology and sparse measurements. Here, this paper proposes a physics-informed graphical learning state estimation method to address these limitations by leveraging both domain physical knowledge and a graph neural network (GNN). We employ a GNN architecture that can handle the graph-structured data of power systems more effectively than traditional data-driven methods. The physics-based knowledge is constructed from the branch current formulation, making the approach adaptable to both transmission and distribution systems. The validation results of three IEEE test systems show that the proposed method can achieve lower mean square error more than 20% than the conventional methods.

43 PARTICLE ACCELERATORS↗

Lightning Detection Efficiency Analysis Process: Modeling Based on Empirical Data

A ground based lightning detection system employs a grid of sensors, which record and evaluate the electromagnetic signal produced by a lightning strike. Several detectors gather information on that signal s strength, time of arrival, and behavior over time. By coordinating the information from several detectors, an event solution can be generated. That solution includes the signal s point of origin, strength and polarity. Determination of the location of the lightning strike uses algorithms based on long used techniques of triangulation. Determination of the event s original signal strength relies on the behavior of the generated magnetic field over distance and time. In general the signal from the event undergoes geometric dispersion and environmental attenuation as it progresses. Our knowledge of that radial behavior together with the strength of the signal received by detecting sites permits an extrapolation and evaluation of the original strength of the lightning strike. It also limits the detection efficiency (DE) of the network. For expansive grids and with a sparse density of detectors, the DE varies widely over the area served. This limits the utility of the network in gathering information on regional lightning strike density and applying it to meteorological studies. A network of this type is a grid of four detectors in the Rondonian region of Brazil. The service area extends over a million square kilometers. Much of that area is covered by rain forests. Thus knowledge of lightning strike characteristics over the expanse is of particular value. I have been developing a process that determines the DE over the region [3]. In turn, this provides a way to produce lightning strike density maps, corrected for DE, over the entire region of interest. This report offers a survey of that development to date and a record of present activity.

Rompala, John T.↗

Exploration of Domain Aware Machine Learning for Grid Analytics: Transfer-Learnt Energy Models to Assist Buildings Control with Sparse Field Data

Buildings are a primary consumer of energy in the United States and are also increasingly being perceived as providers of grid services such as load shifting, shedding and modulation. High fidelity models of building energy consumption are needed to set appropriate baselines for measurement and verification (M&V) of controllers designed for energy efficient operation of buildings and to enable buildings to provide grid services via. participation in demand response programs. State-of-the-art building energy modeling techniques either rely on Physics based models, or extensive instrumentation of the building envelope to gather “big” data to train machine learning based models such as deep neural networks. While Physics based models are often limited by their accuracy, it is not always feasible to gather a significant amount of field data required to train machine learning based models with sufficient accuracy. In this paper, we explore the use of transfer learning-based strategies to address unsatisfactory accuracy of models for estimating building energy consumption when available field data for training is sparse or of unacceptable quality. In particular, we transfer knowledge in the form of data and parameters, from Physics based simulation frameworks to the field to improve the model accuracy, thus resulting in a Physics-informed Machine Learning framework. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that the proposed transfer learning based models provide comparative (and in some cases better) accuracy than state-of-the-art machine learning and deep learning solutions, with just one month of field data.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Unified Communication Optimization Strategies for Sparse Triangular Solver on CPU and GPU Clusters

This paper presents a unified communication optimization framework for sparse triangular solve (SpTRSV) algorithms on CPU and GPU clusters. The framework builds upon a 3D communication-avoiding (CA) layout of Px × Py × Pz processes that divides a sparse matrix into Pz submatrices, each handled by a Px × Py 2D grid with block-cyclic distribution. We propose three communication optimization strategies: First, a new 3D SpTRSV algorithm is developed, which trades the inter-grid communication and synchronization with replicated computation. This design requires only one inter-grid synchronization, and the inter-grid communication is efficiently implemented with sparse allreduce operations. Second, broadcast and reduction communication trees are used to reduce message latency of the intra-grid 2D communication on CPU clusters. Finally, we leverage GPU-initiated one-sided communication to implement the communication trees on GPU clusters. With these nested inter- and intra-grid communication optimization strategies, the proposed 3D SpTRSV algorithm can attain up to 3.45x speedups compared to the baseline 3D SpTRSV algorithm using up to 2048 Cori Haswell CPU cores. In addition, the proposed GPU 3D SpTRSV algorithm can achieve up to 6.5x speedups compared to the proposed CPU 3D SpTRSV algorithm with Pz up to 64. Finally it is remarkable that the proposed GPU 3D SpTRSV can scale to 256 GPUs using the Perlmutter system while the existing 2D SpTRSV algorithm can only scale up to 4 GPUs.

Sao, Piyush↗

Line Faults Classification Using Machine Learning on Three Phase Voltages Extracted from Large Dataset of PMU Measurements

An end-to-end supervised learning method is developed to classify transmission line faults in a twoyear field-recorded dataset that includes synchronized measurements of three-phase voltages recorded by 38 Phasor Measurement Units (PMU) sparsely located in in the US Western Grid interconnection. Statistical analysis is performed to extract features from this large dataset to train Support Vector Machine (SVM), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) classifiers initially. The training further leverages a simulated dataset from a synthetic grid with 12 PMUs to increase the number of faults of types infrequently seen in the field-recorded dataset. Training the classification models with the combined dataset resulted in a classification accuracy of 97.7%. This is a significant improvement over 89.7% to 92.5% accuracy obtained by relying on the field-recorded dataset alone.

47 OTHER INSTRUMENTATION↗

Interactive solution-adaptive grid generation

TURBO-AD is an interactive solution-adaptive grid generation program under development. The program combines an interactive algebraic grid generation technique and a solution-adaptive grid generation technique into a single interactive solution-adaptive grid generation package. The control point form uses a sparse collection of control points to algebraically generate a field grid. This technique provides local grid control capability and is well suited to interactive work due to its speed and efficiency. A mapping from the physical domain to a parametric domain was used to improve difficulties that had been encountered near outwardly concave boundaries in the control point technique. Therefore, all grid modifications are performed on a unit square in the parametric domain, and the new adapted grid in the parametric domain is then mapped back to the physical domain. The grid adaptation is achieved by first adapting the control points to a numerical solution in the parametric domain using control sources obtained from flow properties. Then a new modified grid is generated from the adapted control net. This solution-adaptive grid generation process is efficient because the number of control points is much less than the number of grid points and the generation of a new grid from the adapted control net is an efficient algebraic process. TURBO-AD provides the user with both local and global grid controls.

Choo, Yung K.↗

Incomplete nested dissection for solving n by n grid problems

Nested dissection orderings are known to be very effective for solving sparse positive definite linear systems which arise from n by n grid problems. In this paper we consider incomplete nested dissection, an ordering which corresponds to the premature termination of nested dissection. Analyses of the arithmetic and storage requirements for incomplete nested dissection are given and the ordering is shown to be competitive with nested dissection with regard to arithmetic operations and superior to that ordering in storage requirements.

George, A.↗

Monthly Mean In Situ Surface Flux Observations Paired with Satellite-Derived and Reanalysis-Based Flux Data for the Great Lakes Region, 2001–2020

Surface radiative and turbulent heat fluxes over the Great Lakes strongly influence regional hydrological and meteorological processes, and their accurate representation is critical for numerical weather prediction and coupled atmosphere–lake modeling. However, direct flux observations are spatially sparse across the region, so gridded reanalysis and satellite-derived products are often used for climatological analyses and model evaluation despite differences in their flux representations. This dataset provides processed, quality-controlled, monthly mean surface flux observations from the Great Lakes Evaporation Network (GLEN), AmeriFlux, and the National Data Buoy Center, paired with spatiotemporally matched flux estimates from two reanalysis products, the fifth generation European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis dataset (ERA5) and the Modern Era Reanalysis for Research and Applications, version 2 (MERRA-2), and two satellite-derived products, the Clouds and Earth's Radiant Energy Systems Energy Balanced and Filled (CERES-EBAF) and the Cloud, Albedo and Surface Radiation dataset from AVHRR data - Edition 3 (CLARA-A3). The dataset includes sixteen observational stations with variable temporal coverage within 2001–2020. For each station, a CSV file contains monthly time series of available flux variables, including surface downwelling shortwave radiation (SW), surface downwelling longwave radiation (LW), sensible heat (SH) flux, and latent heat flux (LH), alongside matched gridded product values where available. Columns in the CSV file correspond to different variables sourced from each dataset, with column titles structured as "{dataset}_{variable}". Columns with relevant metadata are also provided in each CSV file, including station latitude and longitude, monthly timestamps, and the name of the sourced observational data. These files are structured for direct use in common analysis tools, including Microsoft Excel, Python pandas, and Python matplotlib. This dataset supports climatological analysis of the Great Lakes regional surface energy budget, evaluation of satellite-derived and reanalysis-based flux products, and development or validation of flux representations in numerical weather prediction and coupled atmosphere–lake models.

Great Lakes↗

Progress on a generalized coordinates tensor product finite element 3DPNS algorithm for subsonic

A generalized coordinates form of the penalty finite element algorithm for the 3-dimensional parabolic Navier-Stokes equations for turbulent subsonic flows was derived. This algorithm formulation requires only three distinct hypermatrices and is applicable using any boundary fitted coordinate transformation procedure. The tensor matrix product approximation to the Jacobian of the Newton linear algebra matrix statement was also derived. Tne Newton algorithm was restructured to replace large sparse matrix solution procedures with grid sweeping using alpha-block tridiagonal matrices, where alpha equals the number of dependent variables. Numerical experiments were conducted and the resultant data gives guidance on potentially preferred tensor product constructions for the penalty finite element 3DPNS algorithm.

Baker, A. J.↗

A variant of nested dissection for solving n by n grid problems

Nested dissection orderings are known to be very effective for solving the sparse positive definite linear systems which arise from n by n grid problems. In this paper nested dissection is shown to be the final step of incomplete nested dissection, an ordering which corresponds to the premature termination of dissection. Analyses of the arithmetic and storage requirements for incomplete nested dissection are given, and the ordering is shown to be competitive with nested dissection under certain conditions.

George, A.↗

Impact of Reordering on the LU Factorization Performance of Bordered Block-Diagonal Sparse Matrix

Power engineers rely on computer-based simulation tools to assess grid performance and ensure security. At the core of these tools are solvers for sparse linear equations. When transformed into a bordered block-diagonal (BBD) structure, part of the sparse linear equation solving can be parallelized. This work focuses on using the Schur-complement-based method for LU factorization on BBD matrices, specifically, Jacobian matrices from large-scale systems. Our findings show that the natural ordering method outperforms the default ordering method in computational performance for each block of the BBD matrix. This observation is validated using synthetic 25k-bus and 70k-bus cases, showing a speedup of up to 38% when using natural ordering without permutation. Additionally, the impact of the number of partitions is studied, and the result shows that computational performance improves with more, smaller partitions in the BBD matrices.

BBD matrix↗