Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Learning model combining convolutional deep neural network with a self-attention mechanism for AC optimal power flow

Alternating current optimal power flow (OPF) analysis is critical for efficient and reliable operation of power systems. For large systems or repetitive computations, the traditional methods such as the direct and gradient methods, or non-traditional methods, such as the genetic algorithm and simulating annealing, are time-consuming and unsuitable for real-time computing. The work in this paper proposes a novel framework to obtain the optimal solution of power flow in real-time using a combination of convolutional neural networks and a self-attention mechanism. All parameters of the power networks are rearranged in an image-like shape of a multi-channel image where each channel is a two-dimensional matrix. The proposed approach is adaptive with every input size of power systems as well as frequent variations of network topologies without intervention to the framework core. The encompassment of all power system contexts in which all parameters of internal elements, generation costs, and topology information are included, contributes to the higher accuracy of inference compared to other current machine-learning-based OPF-solving methods. Besides, the proposed framework established on ubiquitous platforms is effortlessly integrated into current infrastructures of power systems, and the great efficiency along with the computation speed may serve as a critical point for practical implications, such as enabling faster decision-making during real-time operations, predicting system contingencies, and remedial actions based on an offline pre-trained model. Furthermore, this supervised learning process is applied to the dataset of four case studies of meshed power systems: the IEEE 5-bus system (IEEE-5), the IEEE 30-bus system (IEEE-30), the IEEE 39-bus system (IEEE-39), and the IEEE 57-bus system (IEEE-57) to prove the efficacy of the proposed method.

42 ENGINEERING↗

Subtropical Jet in Reanalysis Data from STJ_PV

Subtropical jet position from a new method for locating the subtropical jet, called the tropopause gradient method. It is based on the peak gradient in potential temperature along the dynamic tropopause. This data has the identified subtropical jet latitude, level, and intensity across four different reanalysis products (CFSR-2, ERA-Interim, JRA-55, and MERRA-2), at both daily and monthly output frequency.

58 GEOSCIENCES↗

Optimal surface-tension isotropy in the Rothman-Keller color-gradient lattice Boltzmann method for multiphase flow

The Rothman-Keller color-gradient (CG) lattice Boltzmann method is a popular method to simulate two-phase flow because of its ability to deal with fluids with large viscosity contrasts and a wide range of interfacial tensions. Here, two fluids are labeled red and blue, and the gradient in the color difference is used to compute the effect of interfacial tension. It is well known that finite-difference errors in the color-gradient calculation lead to anisotropy of interfacial tension and errors such as spurious currents. Here, we investigate the accuracy of the CG calculation for interfaces between fluids with several radii of curvature and find that the standard CG calculations lead to significant inaccuracy. Specifically, we observe significant anisotropy of the color gradient of order 7% for high curvature of an interface such as when a pinchout occurs. We derive a second order accurate color gradient and find that the diagonal nearest neighbors can be weighted differently than in the usual color-gradient calculation such that anisotropy is minimized to a fraction of a percent. The optimal weights that minimize anisotropy for the smallest radius of curvature interface are found to be w = (0.298, 0.284, 0.275) for diagonal nearest neighbors for the cases of the interface smoothing parameter β = (0.5, 0.7, 0.99), somewhat higher than the w = 0.25 value derived by Leclaire et al. [Leclaire, Reggio, and Trepanier, Computers and Fluids 48, 98 (2011)] based on obtaining isotropic errors to second order. We find that use of these optimal w values yields over a factor of 10 decrease in anisotropy and over a factor of 30 decrease in mean anisotropy relative to using the standard w = 1 value. And we find a factor of about 2 decrease in the anisotropic error and up to factor 15 decrease in mean anisotropic error relative to the choice of w = 0.25 for small radius of curvature interfaces. The improved CG calculations will allow the method to be more reliably applied to studies of phenomenology and pore scale processes such as viscous and capillary fingering, and droplet formation where surface-tension isotropy of narrow fingers and small droplets plays a crucial role in correctly capturing phenomenology. We present an example illustrating how different phenomena can be captured using the improved color-gradient method. Namely, we present simulations of a wetting fluid invading a fluid filled pipe where the viscosity ratio of fluids is unity in which droplets form at the transition to fingering using the improved CG calculations that are not captured using the standard CG calculations. We present an explanation of why this is so which relates to anisotropy of the surface tension, which inhibits the pinchouts needed to form droplets.

58 GEOSCIENCES↗

Estimating SHmax azimuth with P sources and vertical geophones: Use P-P reflection amplitudes or use SV-P reflection times?

We compared two methods for extracting the azimuth of maximum horizontal stress (SHmax) from 3D land-based seismic data generated by a P source and recorded with vertical geophones. In the first method, we used the direct-SV mode that is produced by all land-based P sources. P sources generate SV illumination that radiates in all azimuth directions from a source station and creates SV-P reflections that are recorded by vertical geophones. Unless stratigraphy has steep dip, SV-P raypaths recorded by vertical geophones are the reverse of P-SV raypaths recorded by horizontal geophones. Thus, SV-P data provide the same S-wave sensitivity to stress fields as popular P-SV data do. In the second method, we retrieved P-P reflections and then performed an amplitude-variation-with-azimuth (AVA) analysis of the amplitude-gradient behavior of P-P reflection wavelets. We did this analysis in narrow azimuth corridors to determine the gradient of reflection-wavelet amplitudes as a function of azimuth. This P-P AVA amplitude-gradient method has been of great interest in the reflection seismology community since it was introduced in the late 1990s. Each of these methods, AVA analysis of the gradient of P-P reflection amplitudes and azimuth-dependent arrival times of SV-P reflections, can be used to determine the azimuth of SHmax stress. We compare the results of the two methods with ground truth measurements of SHmax azimuth at a CO 2 sequestration site in the Michigan Basin. SHmax azimuths were determined from P-P and SV-P data at three major boundaries at depths of approximately 3500 ft (1067 m), 5500 ft (1676 m), and 7500 ft (2286 m). Two estimates of SHmax azimuth (one using SV-P data and one using P-P data) were made at each stacking bin inside a 24 mi 2 (62 km 2 ) image space. The result was approximately 98,000 estimates of SHmax azimuth across each of these three boundaries for each of these two prediction strategies. Histogram displays of PP AVA gradient estimates had peaks at correct azimuths of SHmax at all three depths, but the spread of the distributions widened with depth and split into two peaks at the deepest boundary. In contrast, each histogram of SHmax azimuth predicted by azimuth-dependent SV-P traveltimes had a single, definitive peak that was positioned at the correct SHmax azimuth at all three boundary depths.

Geochemistry & Geophysics↗

Integral boundary conditions in phase field models

Modeling the chemical, electric and thermal transport as well as phase transitions and the accompanying mesoscale microstructure evolution within a material in an electronic device setting involves the solution of partial differential equations often with integral boundary conditions. Employing the familiar Poisson equation describing the electric potential evolution in a material exhibiting insulator to metal transitions, we exploit a special property of such an integral boundary condition, and we properly formulate the variational problem and establish its well-posedness. Next, we compare our method with the commonly-used Lagrange multiplier method that can also handle such boundary conditions. Numerical experiments demonstrate that our new method achieves optimal convergence rate in contrast to the conventional Lagrange multiplier method. Furthermore, the linear system derived from our method is symmetric positive definite, and can be efficiently solved by Conjugate Gradient method with algebraic multigrid preconditioning.

97 MATHEMATICS AND COMPUTING↗

Towards Enhancing Coding Productivity for GPU Programming Using Static Graphs

The main contribution of this work is to increase the coding productivity of GPU programming by using the concept of Static Graphs. GPU capabilities have been increasing significantly in terms of performance and memory capacity. However, there are still some problems in terms of scalability and limitations to the amount of work that a GPU can perform at a time. To minimize the overhead associated with the launch of GPU kernels, as well as to maximize the use of GPU capacity, we have combined the new CUDA Graph API with the CUDA programming model (including CUDA math libraries) and the OpenACC programming model. We use as test cases two different, well-known and widely used problems in HPC and AI: the Conjugate Gradient method and the Particle Swarm Optimization. In the first test case (Conjugate Gradient) we focus on the integration of Static Graphs with CUDA. In this case, we are able to significantly outperform the NVIDIA reference code, reaching an acceleration of up to 11x thanks to a better implementation, which can benefit from the new CUDA Graph capabilities. In the second test case (Particle Swarm Optimization), we complement the OpenACC functionality with the use of CUDA Graph, achieving again accelerations of up to one order of magnitude, with average speedups ranging from 2x to 4x, and performance very close to a reference and optimized CUDA code. Our main target is to achieve a higher coding productivity model for GPU programming by using Static Graphs, which provides, in a very transparent way, a better exploitation of the GPU capacity. The combination of using Static Graphs with two of the current most important GPU programming models (CUDA and OpenACC) is able to reduce considerably the execution time w.r.t. the use of CUDA and OpenACC only, achieving accelerations of up to more than one order of magnitude. Finally, we propose an interface to incorporate the concept of Static Graphs into the OpenACC Specifications.

58 GEOSCIENCES↗

Random Phase Approximation Correlation Energy Using Real-Space Density Functional Perturbation Theory

We present a real-space method for computing the random phase approximation (RPA) correlation energy within Kohn–Sham density functional theory, leveraging the low-rank nature of the frequency-dependent density response operator. In particular, we employ a cubic-scaling formalism based on density functional perturbation theory that circumvents the calculation of the response function matrix, instead relying on the ability to compute its product with a vector through the solution of the associated Sternheimer linear systems. We develop a large-scale parallel implementation of this formalism using the subspace iteration method in conjunction with the spectral quadrature method while employing the Kronecker product-based method for the application of the Coulomb operator and the conjugate orthogonal conjugate gradient method for the solution of the linear systems. We demonstrate convergence with respect to key parameters and verify the method’s accuracy by comparing with plane-wave results. We show that the framework achieves good strong scaling to many thousands of processors, reducing the time to solution for a lithium hydride system with 128 electrons to around 150 s on 4608 processors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Drift‐Kinetic Method for Obtaining Gradients in Plasma Properties From Single‐Point Distribution Function Data

Abstract In this paper, we derive a new drift‐kinetic method for estimating gradients in the plasma properties through a velocity space distribution at a single point. The gradients are intrinsically related to agyrotropic features of the distribution function. This method predicts the gradients in the magnetized distribution function and can predict gradients of arbitrary moments of the gyrotropic background distribution function. The method allows for estimates on density and pressure gradients on the scale of a Larmor radius, proving to resolve smaller scales than any method currently available to spacecraft. The model is verified with a set of fully kinetic VPIC particle‐in‐cell simulations.

Wetherton, B. A.↗

CRCNS22 Learning Rules in the Hippocampus and their Mapping to Neuromorphic Systems (Final Technical Report)

Large scale biologically-realistic computational models are key to investigating the interplay between structure and function in nervous systems, thus paving the way to new clinical methods and neuro-inspired computing solutions. This project focuses on the hippocampus, in particular the CA3-CA1 regions, due to their role in associative learning and memory, pattern separation and completion, and spatial navigation. Investigations into the neuronal organization and learning rule(s) of this circuit can shed light into how declarative memories are formed, stored, recalled and forgotten and inform computational, experimental and clinical neuroscience work. Our project aims at developing a novel data-driven methodology supported by a broad heterogeneous base of neuroscience experimental knowledge and inspired from advances in computer science and engineering. Specifically, this work will benchmark existing and new learning rules within a full-scale spiking neural network simulation of the CA3-CA1 region. The model will be based on an open-source repository, called the Hippocampome, which contains neuronal morphologies, firing patterns, synapse probabilities, and most other required parameters for all known neuron types in the rodent hippocampal formation. The model will be first trained in a supervised fashion for associative memory tasks using backpropagation through time traditionally used in computer science, enhanced with a new technique called the surrogate gradient method. This optimization method will be used to obtain a global loss minimization, but it is not biologically inspired as it assumes the use of data not locally available to the synapses. However, we propose its use as a benchmarking tool, to compare the training performance of local biologically plausible and hardware-mappable learning rules at scale. New rules or combinations will be proposed and tested as needed, based on the obtained results. Progress in this area will also drive the development of novel hardware-mappable algorithms for continual lifelong learning and categorization of new events from few presented examples. This project goes beyond the existing state-of-the-art by looking at large scale realistic neuronal circuits as networks trainable via global optimization methods such as surrogate gradient descent. The objective function of the brain that supports learning is largely unknown, but it is likely that it operates through local learning rules. Studying network trajectories around local minima as proposed in this work represents a useful strategy for understanding whether a network is training by using a specific (set of) learning rule(s). Starting from a completely untrained network is a challenging test since it is difficult to determine how the learning rule affects the trajectory of the network. This interdisciplinary project will help understand what rule governs learning in these regions or if multiple learning rules are involved. The work will develop a robust methodology to measure if the network is converging to the target solution, oscillating around it, or diverging away.

59 BASIC BIOLOGICAL SCIENCES↗

Estimating turbulent energy flux vertical profiles from uncrewed aircraft system measurements: exemplary results for the MOSAiC campaign

This study analyzes turbulent energy fluxes in the Arctic atmospheric boundary layer (ABL) using measurements with a small uncrewed aircraft system (sUAS). Turbulent fluxes constitute a major part of the atmospheric energy budget and influence the surface heat balance by distributing energy vertically in the atmosphere. However, only few in situ measurements of the vertical profile of turbulent fluxes in the Arctic ABL exist. The study presents a method to derive turbulent heat fluxes from DataHawk2 sUAS turbulence measurements, based on the flux gradient method with a parameterization of the turbulent exchange coefficient. This parameterization is derived from high-resolution horizontal wind speed measurements in combination with formulations for the turbulent Prandtl number and anisotropy depending on stability. Measurements were taken during the MOSAiC (Multidisciplinary drifting Observatory for the Study of Arctic Climate) expedition in the Arctic sea ice during the melt season of 2020. For three example cases from this campaign, vertical profiles of turbulence parameters and turbulent heat fluxes are presented and compared to balloon-borne, radar, and near-surface measurements. The combination of all measurements draws a consistent picture of ABL conditions and demonstrates the unique potential of the presented method for studying turbulent exchange processes in the vertical ABL profile with sUAS measurements.

54 ENVIRONMENTAL SCIENCES↗

SymbolNet: neural symbolic regression with adaptive dynamic pruning for compression

Abstract Compact symbolic expressions have been shown to be more efficient than neural network (NN) models in terms of resource consumption and inference speed when implemented on custom hardware such as field-programmable gate arrays (FPGAs), while maintaining comparable accuracy (Tsoi et al 2024 EPJ Web Conf. 295 09036). These capabilities are highly valuable in environments with stringent computational resource constraints, such as high-energy physics experiments at the CERN Large Hadron Collider. However, finding compact expressions for high-dimensional datasets remains challenging due to the inherent limitations of genetic programming (GP), the search algorithm of most symbolic regression (SR) methods. Contrary to GP, the NN approach to SR offers scalability to high-dimensional inputs and leverages gradient methods for faster equation searching. Common ways of constraining expression complexity often involve multistage pruning with fine-tuning, which can result in significant performance loss. In this work, we propose S y m b o l N e t , a NN approach to SR specifically designed as a model compression technique, aimed at enabling low-latency inference for high-dimensional inputs on custom hardware such as FPGAs. This framework allows dynamic pruning of model weights, input features, and mathematical operators in a single training process, where both training loss and expression complexity are optimized simultaneously. We introduce a sparsity regularization term for each pruning type, which can adaptively adjust its strength, leading to convergence at a target sparsity ratio. Unlike most existing SR methods that struggle with datasets containing more than O ( 10 ) inputs, we demonstrate the effectiveness of our model on the LHC jet tagging task (16 inputs), MNIST (784 inputs), and SVHN (3072 inputs).

Tsoi, Ho Fung (ORCID:0000000225502184)↗

Model-Free Primal-Dual Methods for Network Optimization with Application to Real-Time Optimal Power Flow

This paper examines the problem of real-time optimization of networked systems and develops online algorithms that steer the system towards the optimal trajectory without explicit knowledge of the system model. The problem is modeled as a dynamic optimization problem with time-varying performance objectives and engineering constraints. The design of the algorithms leverages the online zero-order primal-dual projected-gradient method. In particular, the primal step that involves the gradient of the objective function (and hence requires a networked systems model) is replaced by its zero-order approximation with two function evaluations using a deterministic perturbation signal. The evaluations are performed using the measurements of the system output, hence giving rise to a feedback interconnection, with the optimization algorithm serving as a feedback controller. The paper provides some insights on the stability and tracking properties of this interconnection. Finally, the paper applies this methodology to a real-time optimal power flow problem in power systems, and shows its efficacy on the IEEE 37-node distribution test feeder for reference power tracking and voltage regulation.

61 RADIATION PROTECTION AND DOSIMETRY↗

Plant physical defenses contribute to a latitudinal gradient in resistance to insect herbivory within a widespread perennial grass

Premise: Herbivore pressure can vary across the range of a species, resulting in different defensive strategies. If herbivory is greater at lower latitudes, plants may be better defended there, potentially driving a latitudinal gradient in defense. However, relationships that manifest across the entire range of a species may be confounded by differences within genetic subpopulations, which may obscure the drivers of these latitudinal gradients. Methods: We grew plants of the widespread perennial grass Panicum virgatum in a common garden that included genotypes from three genetic subpopulations spanning an 18.5° latitudinal gradient. We then assessed defensive strategies of these plants by measuring two physical resistance traits—leaf mass per area (LMA) and leaf ash, a proxy for silica—and multiple measures of herbivory by caterpillars of the generalist herbivore fall armyworm (Spodoptera frugiperda). Results: Across all genetic subpopulations, low-latitude plants experienced less herbivory than high-latitude plants. Within genetic subpopulations, however, this relationship was inconsistent—the most widely distributed and phenotypically variable subpopulation (Atlantic) exhibited more consistent latitudinal trends than either of the other two subpopulations. The two physical resistance traits, LMA and leaf ash, were both highly heritable and positively associated with resistance to different measures of herbivory across all subpopulations, indicating their importance in defense against herbivores. Again, however, these relationships were inconsistent within subpopulations. Conclusions: Defensive gradients that occur across the entire species range may not arise within localized subpopulations. Thus, identifying the drivers of latitudinal gradients in herbivory defense may depend on adequately sampling the diversity within a species.

09 BIOMASS FUELS↗

Inference of main ion particle transport coefficients with experimentally constrained neutral ionization during edge localized mode recovery on DIII-D

Abstract The plasma and neutral density dynamics after an edge localized mode are investigated and utilized to infer the plasma transport coefficients for the density pedestal. The Lyman-Alpha Measurement Apparatus (LLAMA) diagnostic provides sub-millisecond profile measurements of the ionization and neutral density and shows significant poloidal asymmetries in both. Exploiting the absolute calibration of the LLAMA diagnostic allows quantitative comparison to the electron and main ion density profiles determined by charge-exchange recombination, Thomson scattering and interferometry. Separation of diffusion and convection contributions to the density pedestal transport are investigated through flux gradient methods and time-dependent forward modeling with Bayesian inference by adaptation of the Aurora transport code and IMPRAD framework to main ion particle transport. Both methods suggest time-dependent transport coefficients and are consistent with an inward particle pinch on the order of 1 m s −1 and diffusion coefficient of 0.05 m 2 s −1 in the steep density gradient region of the pedestal. While it is possible to recreate the experimentally observed phenomena with no pinch in the pedestal, low diffusion in the core and high outward convection in the near scrape-off layer are required without an inward pedestal pinch.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Extending XACC for Quantum Optimal Control

Quantum computing vendors are beginning to open up application programming interfaces for direct pulse-level quantum control. With this, programmers can begin to describe quantum kernels of execution via sequences of arbitrary pulse shapes. This opens new avenues of research and development with regards to smart quantum compilation routines that enable direct translation of higher-level digital assembly representations to these native pulse instructions. In this work, we present an extension to the XACC system-level quantum-classical software framework that directly enables this compilation lowering phase via user-specified quantum optimal control techniques. This extension enables the translation of digital quantum circuit representations to equivalent pulse sequences that are optimal with respect to the backend system dynamics. Our work is modular and extensible, enabling third party optimal control techniques and strategies in both C++ and Python. We demonstrate this extension with familiar gradient-based methods like gradient ascent pulse engineering (GRAPE), gradient optimization of analytic controls (GOAT), and Krotov's method. Our work serves as a foundational component of future quantum-classical compiler designs that lower high-level programmatic representations to low-level machine instructions.

Nguyen, Thien↗

Improved Guarantees for Optimal Nash Equilibrium Seeking and Bilevel Variational Inequalities

We consider a class of hierarchical variational inequality (VI) problems that subsumes VI-constrained optimization and several other problem classes, including the optimal solution selection problem and the optimal Nash equilibrium (NE) seeking problem. Our main contribution is threefold. (i) We consider bilevel VIs with monotone and Lipschitz continuous mappings and devise a single-timescale iteratively regularized extragradient method, named IR-EG 𝚖,𝚖 . We improve the existing iteration complexity results for addressing both bilevel VI and VI-constrained convex optimization problems. (ii) Under the strong monotonicity of the outer-level mapping, we develop a method named IR-EG 𝚜,𝚖 and derive faster guarantees than those in (i). We also study the iteration complexity of this method under a constant regularization parameter. These results appear to be new for both bilevel VIs and VI-constrained optimization. (iii) To our knowledge, complexity guarantees for computing the optimal NE in nonconvex settings do not exist. Motivated by this lacuna, we consider VI-constrained nonconvex optimization problems and devise an inexactly projected gradient method, named IPR-EG, where the projection onto the unknown set of equilibria is performed using IR-EG 𝚜,𝚖 with a prescribed termination criterion and an adaptive regularization parameter. We obtain new complexity guarantees in terms of a residual map and an infeasibility metric for computing a stationary point. Here, we validate the theoretical findings using preliminary numerical experiments for computing the best and the worst NEs.

bilevel optimization↗

Stochasticity and robustness in spiking neural networks

Despite drawing inspiration from biological systems which are inherently noisy and variable, artificial neural networks have been shown to require precise weights to carry out the task which they are trained to accomplish. This creates a challenge when adapting these artificial networks to specialized execution platforms which may encode weights in a manner which restricts their accuracy and/or precision.Reflecting back on the non-idealities which are observed in biological systems, we investigated the effect these properties have on the robustness of spiking neural networks under perturbations to weights. First, we examined techniques extant in conventional neural networks which resemble noisy processes, and postulated they may produce similar beneficial effects in spiking neural networks. Second, we evolved a set of spiking neural networks utilizing biological non-idealities to solve a pole-balancing task, and estimated their robustness. We showed it is higher in networks using noisy neurons, and demonstrated that one of these networks can perform well under the variance expected when a hafnium-oxide based resistive memory is used to encode synaptic weights. Lastly, we trained a series of networks using a surrogate gradient method on the MNIST classification task. We confirmed that these networks demonstrate similar trends in robustness to the evolved networks. We discuss these results and argue that they display empirical evidence supporting the role of noise as a regularizer which can increase network robustness.

97 MATHEMATICS AND COMPUTING↗

A machine learning framework for accurate and robust analysis of radiation detector pulses

The microscopic properties of atomic nuclei are used to study various scientific questions. They are essential for understanding the fundamental forces of nature and the chemical evolution of the universe. Detecting decay radiation from radioactive nuclei makes it possible to probe these fundamental nuclear properties. Detector waveform traces may contain additional information about the radiation. Generally, advanced signal processing techniques are needed to extract this additional information, often involving fitting the waveform with model response functions using non-linear least-squares optimization with second-order gradient methods. While this is a powerful technique, it is also computationally expensive, leading to slow processing time, which scales with the volume of data. To address this problem, we have developed a machine learning (ML) approach that infers the characteristics of traces from a model detector response function. In particular, we are interested in classifying whether a single recorded trace consists of one or two pulse constituents and estimating the pulse parameters. Furthermore, our proposed ML method can precisely extract the pulses’ parameters, such as energy and timing information, and accurately classify the pulse multiplicity of a trace. Unlike non-learning-based approaches, our ML approach uses neural networks that are significantly faster at inference, as they do not require any optimization during this stage.

Curve fitting↗