Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Tucker”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Improved Particle Swarm Algorithm Using Rubik’s Cube Topology for Bilevel Building Energy Transaction

Following the rapid growth of distributed energy resources (e.g., renewables, battery), localized peer-to-peer energy transactions are receiving more attention for multiple benefits, such as reducing power loss and stabilizing the main power grid. To promote distributed renewables locally, the local trading price is usually set to be within the external energy purchasing and selling price range. Consequently, building prosumers are motivated to trade energy through a local transaction center. This local energy transaction is modeled in bilevel optimization game. A selfish upper level agent is assumed with the privilege to set the internal energy transaction price with an objective of maximizing its arbitrage profit. Meanwhile, the building prosumers at the lower level will response to this transaction price and make decisions on electricity transaction amount. Therefore, this non-cooperative leader-follower trading game is seeking for equilibrium solutions on the energy transaction amount and prices. Additionally, a uniform local transaction price structure (purchase price equals selling price) is considered here. Aiming at reducing the computational burden from classical Karush–Kuhn–Tucker (KKT) transformation and protecting the private information of each stakeholder (e.g., building), swarm intelligence-based solution approach is employed for upper level agent to generate trading price and coordinate the transactive operations. On one hand, to decrease the chance of premature convergence in global-best topology, Rubik’s Cube topology is proposed in this study based on further improvement of a two-dimensional square lattice model (i.e., one local-best topology-Von Neumann topology). Rotating operation of the cube is introduced to dynamically changing the neighborhood and enhancing information flow at the later searching state. Several groups of experiments are designed to evaluate the performance of proposed Rubik’s Cube topology-based particle swarm algorithm. The results have validated the effectiveness of proposed topology and operators comparing with global-best version PSO and Von Neumann topology-based PSO and its scalability on larger scale applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

BI-LEVEL OPTIMIZATION FOR ELECTRICITY TRANSACTION IN SMART COMMUNITY WITH MODULAR PUMP HYDRO STORAGE

Grid integration of the increasing distributed energy resources could be challenging in terms of new infrastructure investment, power grid stability, etc. To resolve more renewables locally and reduce the need for extensive electricity transmission, a community energy transaction market is assumed with market operator as the leader whose responsibility is to generate local energy prices and clear the energy transaction payment among the prosumers (followers). The leader and multi-followers have competitive objectives of revenue maximization and operational cost minimization. This non-cooperative leader-follower (Stackelberg) game is formulated using a bi-level optimization framework, where a novel modular pump hydro storage technology (GLIDES system) is set as an upper level market operator, and the lower level prosumers are nearby commercial buildings. The best responses of the lower level model could be derived by necessary optimality conditions, and thus the bi-level model could be transformed into single level optimization model via replacing the lower level model by its Karush-Kuhn-Tucker (KKT) necessary conditions. Several experiments have been designed to compare the local energy transaction behavior and profit distribution with the different demand response levels and different local price structures. The experimental results indicate that the lower level prosumers could benefit the most when local buying and selling prices are equal, while maximum revenue potential for the upper level agent could be reached with non-equal trading prices.

Chen, Yang↗

Running Primal-Dual Gradient Method for Time-Varying Nonconvex Problems

This paper focuses on a time-varying constrained nonconvex optimization problem, and considers the synthesis and analysis of online regularized primal-dual gradient methods to track a Karush-Kuhn-Tucker (KKT) trajectory. The proposed regularized primal-dual gradient method is implemented in a running fashion, in the sense that the underlying optimization problem changes during the execution of the algorithms. In order to study its performance, we first derive its continuous-time limit as a system of differential inclusions. We then study sufficient conditions for tracking a KKT trajectory, and also derive asymptotic bounds for the tracking error (as a function of the time-variability of a KKT trajectory). Further, we provide a set of sufficient conditions for the KKT trajectories not to bifurcate or merge, and also investigate the optimal choice of the parameters of the algorithm. Illustrative numerical results for a time-varying nonconvex problem are provided.

differential inclusion↗

Towards Compact Neural Networks via End-to-End Training: A Bayesian Tensor Approach with Automatic Rank Determination

Post-training model compression can reduce the inference costs of deep neural networks, but uncompressed training still consumes enormous hardware resources and energy. To enable low-energy training on edge devices, it is highly desirable to directly train a compact neural network from scratch with a low memory cost. Low-rank tensor decomposition is an effective approach to reduce the memory and computing costs of large neural networks. However, directly training low-rank tensorized neural networks is a very challenging task because it is hard to determine a proper tensor rank a priori, and the tensor rank controls both model complexity and accuracy. Here, this paper presents a novel end-to-end framework for low-rank tensorized training. We first develop a Bayesian model that supports various low-rank tensor formats (e.g., CANDECOMP/PARAFAC, Tucker, tensor-train, and tensor-train matrix) and reduces neural network parameters with automatic rank determination during training. Then we develop a customized Bayesian solver to train large-scale tensorized neural networks. Our training methods shows orders-of-magnitude parameter reduction and little accuracy loss (or even better accuracy) in the experiments. On a very large deep learning recommendation system with over 4.2 ×10 9 model parameters, our method can reduce the parameter number to 1.6 ×10 5 automatically in the training process (i.e., by 2.6 ×10 4 times) while achieving almost the same accuracy. Code is available at https://github.com/colehawkins/bayesian-tensor-rank-determination.

compact neural networks↗

Parallel Algorithms for Computing the Tensor-Train Decomposition

The tensor-train (TT) decomposition expresses a tensor in a data-sparse format used in molecular simulations, high-order correlation functions, and optimization. In this paper, we propose four parallelizable algorithms that compute the TT format from various tensor inputs: (1) Parallel-TTSVD for traditional format, (2) PSTT and its variants for streaming data, (3) Tucker2TT for Tucker format, and (4) TT-fADI for solutions of Sylvester tensor equations. We provide theoretical guarantees of accuracy, parallelization methods, scaling analysis, and numerical results. For example, for a d-dimension tensor in $\mathbb{R}$ $n\times∙∙∙$$\times$$n$ a two-sided sketching algorithm PSTT2 is shown to have a memory complexity of $O(n^{[d/2]})$, improving upon $O(n^{d—1})$ from previous algorithms.

97 MATHEMATICS AND COMPUTING↗

Communication Lower Bounds and Optimal Algorithms for Multiple Tensor-Times-Matrix Computation

Multiple tensor-times-matrix (Multi-TTM) is a key computation in algorithms for computing and operating with the Tucker tensor decomposition, which is frequently used in multidimensional data analysis. Here, we establish communication lower bounds that determine how much data movement is required (under mild conditions) to perform the Multi-TTM computation in parallel. The crux of the proof relies on analytically solving a constrained, nonlinear optimization problem. We also present a parallel algorithm to perform this computation that organizes the processors into a logical grid with twice as many modes as the input tensor. We show that, with correct choices of grid dimensions, the communication cost of the algorithm attains the lower bounds and is therefore communication optimal. Finally, we show that our algorithm can significantly reduce communication compared to the straightforward approach of expressing the computation as a sequence of tensor-times-matrix operations when the input and output tensors vary greatly in size.

HBL-inequalities↗

A Linear-Complexity Tensor Butterfly Algorithm for Compressing High-Dimensional Oscillatory Integral Operators

This paper presents a multilevel tensor compression algorithm called tensor butterfly algorithm for efficiently representing large-scale and high-dimensional oscillatory integral operators, including Green's functions for wave equations and integral transforms such as Radon transforms and Fourier transforms. The proposed algorithm leverages a tensor extension of the so-called complementary low-rank property of existing matrix butterfly algorithms. The algorithm partitions the discretized integral operator tensor into subtensors of multiple levels and factorizes each subtensor at the middle level as a Tucker-type interpolative decomposition, whose factor matrices are formed in a multilevel fashion. For a d-dimensional (d > 1) integral operator discretized into a 2d-mode tensor with n2d entries, the overall CPU time and memory requirement scale as O(nd), in stark contrast to the O(nd log n) complexity of existing matrix algorithms such as matrix butterfly algorithms and fast Fourier transforms (FFTs), where n is the number of points per direction. When comparing with other tensor algorithms such as quantized tensor train (QTT), the proposed algorithm also shows superior CPU and memory performance for tensor contraction. Remarkably, the tensor butterfly algorithm can efficiently model high-frequency Green's function interactions between two unit cubes, each spanning 512 wavelengths per direction, which represents problems of scale over 512× larger than that existing butterfly algorithms can handle, with the same amount of computation resources. On the other hand, for a problem representing 64 wavelengths per direction, which is the largest size existing algebraic matrix algorithms can handle, our tensor butterfly algorithm exhibits 200x speedups and 30× memory reduction compared with existing ones. Moreover, the tensor butterfly algorithm also permits O(nd)-complexity FFTs and Radon transforms up to d = 6 dimensions.

Kielstra, P Michael↗

HyKKT

HyKKT (pronounced as "hiked") is a package for solving systems of linear equations of Karush-Kuhn-Tucker (KKT) form, which typically arise in optimization problems, such as optimal power flow analysis. HyKKT uses Cholesky instead of LDL^T factorization and solves the general KKT system to a desired numerical precision via block reduction and conjugate gradient on the Schur complement. Such implementation is more suitable for implementation on graphic processing units (GPUs).

Regev, Shaked↗

pnnl/HiParTI

A Hierarchical Parallel Tensor Infrastructure (HiParTI), is to support fast essential sparse tensor operations and tensor decompositions on multicore CPU and GPU architectures. It consists of sparse tensor decompositions, CANDECOMP/PARAFAC (CP) and Tucker decompositions, fundamental tensor operations, and tensor transformations

Li, Jiajia↗

Planetary Boundary-Layer Height (PBLHT) Value-Added Product: Remote-Sensing Retrievals

The planetary boundary layer (PBL) is fundamental to numerous atmospheric processes, including aerosol mixing and transport, cloud evolution, and precipitation formation. A critical parameter in these studies is the PBL height (PBLHT). This vertical depth is essential for characterizing PBL structures in numerical simulations and serves as a primary metric for estimating flux exchanges between the Earth’s surface and the atmosphere. Radiosonde (SONDE) observations provide high-vertical-resolution measurements of temperature and moisture profiles and are widely used to estimate PBLHT (Liu and Liang 2010, Seidel et al. 2010). The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility’s PBLHT value-added product (VAP) for radiosonde measurements, known as PBLHTSONDE, applies three commonly used methods—the Heffter (1980) method, the Liu and Liang (2010) method, and the bulk Richardson number approach (Seibert et al. 2000)—to derive PBLHT. The PBLHTSONDE VAP operates routinely at ARM observatories and mobile facilities, with data available from the ARM Data Center shortly after sounding observations are collected (Sivaraman et al. 2013). However, radiosonde observations are limited by their low temporal resolution. Most stations launch soundings only twice daily, which constrains the ability to investigate and characterize the temporal evolution of the PBL using radiosonde data alone. The use of continuous remote-sensing observations provides high temporal resolution of PBLHT estimates. These observations include aerosol lidars (Dang et al. 2019, Su et al. 2020), Doppler lidar (DL; Tucker et al. 2009, Krishnamurthy et al. 2021), and water vapor and/or temperature lidars and radiometers (Turner et al. 2014). These observations provide valuable data on the PBL’s thermodynamic properties (e.g., water vapor and/or temperature lidars and radiometers), dynamic properties (e.g., DL), and distribution of tracer substances (e.g., aerosol lidars), all of which can be used to estimate PBLHT. ARM developed PBLHT estimates from the micropulse lidar (MPL; PBLHTMPL), Doppler lidar (PBLHTDL), and combined Raman lidar (RL)/atmospheric emitted radiance interferometer (AERI) thermodynamic profiles (PBLHTTHERMO). Each estimate captures different physical characteristics of the boundary layer—aerosol tracers, vertical velocity turbulence, and thermodynamic structure—and exhibits distinct strengths and limitations depending on the PBL regime and time of day. In addition, the ARM ceilometer (CEIL) provides three potential PBLHT candidates derived from the vendor's built-in algorithm. Building on these individual retrievals, ARM developed the PBLHTBEML VAP, which combines the four remote-sensing-based estimates with ancillary meteorological variables using the machine learning approach of Zhang et al. (2025) to produce a best-estimate PBLHT at 10-minute resolution.

54 ENVIRONMENTAL SCIENCES↗

On the estimation of boundary layer heights: a machine learning approach

Abstract. The planetary boundary layer height (zi) is a key parameter used in atmospheric models for estimating the exchange of heat, momentum, and moisture between the surface and the free troposphere. Near-surface atmospheric and subsurface properties (such as soil temperature, relative humidity, etc.) are known to have an impact on zi. Nevertheless, precise relationships between these surface properties and zi are less well known and not easily discernible from the multi-year dataset. Machine learning approaches, such as random forest (RF), which use a multi-regression framework, help to decipher some of the physical processes linking surface-based characteristics to zi. In this study, a 4-year dataset from 2016 to 2019 at the Southern Great Plains site is used to develop and test a machine learning framework for estimating zi. Parameters derived from Doppler lidars are used in combination with over 20 different surface meteorological measurements as inputs to a RF model. The model is trained using radiosonde-derived zi values spanning the period from 2016 through 2018 and then evaluated using data from 2019. Results from 2019 showed significantly better agreement with the radiosonde compared to estimates derived from a thresholding technique using Doppler lidars only. Noteworthy improvements in daytime zi estimates were observed using the RF model, with a 50 % improvement in mean absolute error and an R2 of greater than 85 % compared to the Tucker method zi. We also explore the effect of zi uncertainty on convective velocity scaling and present preliminary comparisons between the RF model and zi estimates derived from atmospheric models.

54 ENVIRONMENTAL SCIENCES↗

Rubik’s Cube Topology Based Particle Swarm Algorithm for Bilevel Building Energy Transaction

Following the rapid growth of distributed energy resources (e.g. renewables, battery), localized peer-to-peer energy transactions are receiving more attention for multiple benefits, such as, reducing power loss, stabilizing the main power grid, etc. To promote distributed renewables locally, the local trading price is usually set to be within the external energy purchasing and selling price range. Consequently, building prosumers are motivated to trade energy through a local transaction center. This local energy transaction is modeled in bilevel optimization game. A selfish upper level agent is assumed with the privilege to set the internal energy transaction price with an objective of maximizing its arbitrage profit. Meanwhile, the building prosumers at the lower level will response to this transaction price and make decisions on electricity transaction amount. Therefore, this non-cooperative leader-follower trading game is seeking for equilibrium solutions on the energy transaction amount and prices. In addition, a uniform local transaction price structure (purchase price equals selling price) is considered here. Aiming at reducing the computational burden from classical Karush-Kuhn-Tucker (KKT) transformation and protecting the private information of each stakeholder (e.g., building), swarm intelligence based solution approach is employed for upper level agent to generate trading price and coordinate the transactive operations. On one hand, to decrease the chance of premature convergence in global-best topology, Rubiks Cube topology is proposed in this study based on further improvement of a two-dimensional square lattice model (i.e., one local-best topology-Von Neumann topology). Rotating operation of the cube is introduced to dynamically changing the neighborhood and enhancing information flow at the later searching state. Several groups of experiments are designed to evaluate the performance of proposed Rubiks Cube topology based particle swarm algorithm. The results have validated the effectiveness of proposed topology and operators comparing with global-best version PSO and Von Neumann topology based PSO and its scalability on larger scale applications.

Feng, Xiaochun↗

Modeling the Strategic Behavior of an Active Distribution Network in the ISO Markets

With increasing integration of distributed energy resources (DERs), active distribution networks (ADNs) can actively participate in the electricity markets by dispatching their DERs, which can change the existing electricity market paradigm. It is essential to investigate the strategic behaviors of ADNs and their DER dispatch when they participate in the wholesale market as price-makers. This paper proposes a bi-level optimization model to study the strategic behavior of an ADN in both energy and reserve markets. The optimal scheduling of DERs in the ADN is modeled as the upper level problem and the joint energy and reserve market-clearing of the ISO is modeled as the lower-level problem. The two-level optimization models exchange bidding information and energy/reserve prices with each other. The proposed bi-Ievel optimization problem is converted to a mathematical programming with equilibrium constraints (MPEC) by using Karush-Kuhn Tucker (KKT) conditions and strong duality theory. Further, the MPEC problem is reformulated as a computationally-solvable mixed integer second order cone programming (MISOCP) model. The simulation results on an illustrative case demonstrate the impact of the strategic bidding of the ADN on the day-ahead energy and reserve market prices.

Xue, Yaosuo↗

Reverse-mode differentiation in arbitrary tensor network format: with application to supervised learning.

This paper describes an efficient reverse-mode differentiation algorithm for contraction operations of tensor networks that may have arbitrary and unconventional network topologies. The approach leverages the tensor contraction tree of Evenbly and Pfeifer (2014), which provides an instruction set for the contraction sequence of a network. We show that this tree can be efficiently leveraged for differentiation of a full tensor network contraction using a recursive scheme that exploits (1) the bilinear property of contraction and (2) the property that trees have single path from root to leaves. While differentiation of tensor-tensor contraction is already possible in most automatic differentiation packages, we show that exploiting these two additional properties in the specific context of contraction sequences can improve efficiency. Following a description of the algorithm and computational complexity analysis, we investigate its utility for gradient-based supervised learning for low-rank function recovery and for fitting real-world unstructured datasets. We demonstrate improved performance over alternating least-squares optimization approaches and the capability to handle heterogeneous and arbitrary tensor network formats. When compared to alternating minimization algorithms, we find that the gradient-based approach requires a smaller oversampling ratio (number of samples compared to number model parameters) for recovery. This increased efficiency extends to fitting unstructured data of varying dimensionality and when employing a variety of tensor network formats. Here, we show improved learning using the hierarchical Tucker method over the tensor-train in high-dimensional settings on a number of benchmark problems.

97 MATHEMATICS AND COMPUTING↗

In-Flight Gain Monitoring of SPIDER’s Transition-Edge Sensor Arrays

We report experiments deploying large arrays of transition-edge sensors (TESs) often require a robust method to monitor gain variations with minimal loss of observing time. We propose a sensitive and non-intrusive method for monitoring variations in TES responsivity using small square waves applied to the TES bias. We construct an estimator for a TES’s small-signal power response from its electrical response that is exact in the limit of strong electrothermal feedback. We discuss the application and validation of this method using flight data from SPIDER, a balloon-borne telescope that observes the polarization of the cosmic microwave background with more than 2000 TESs. This method may prove useful for future balloon- and space-based instruments, where observing time and ground control bandwidth are limited.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Voltage cycling as a dynamic operation mode for high temperature electrolysis solid oxide cells

Solid Oxide Electrolysis Cells (SOECs) have emerged as a promising technology for the efficient production of H2 via high-temperature electrolysis. However, power input from dynamic energy sources remains a significant challenge for their long-term stability. It is important to analyze the tolerance of cells under dynamic operation conditions. This study focuses on evaluating the impact of voltage cycling on the performance and durability of electrode-supported SOECs. We explore the operational limits and degradation mechanisms of SOECs subjected to various voltage conditions and find that the cells have high tolerance for dynamic voltage. Voltage cycling between 1.3 V and 1.5 V for 9000 cycles does not damage the cell. Conversely, cycling to higher voltages (≥1.7 V) results in accelerated degradation. Advanced characterization is used to screen for various degradation modes post operation. Within the oxygen electrode, XRD and STEM EDS find compositional and phase evolution in all voltage cycled samples including increased decomposition of the air electrode resulting in cation migration. Microstructural analysis of the fuel electrode from nano-CT data shows minimal change throughout the sample set and no evidence of Ni migration, indicating the fuel electrode is stable and not impacted by cycling to higher voltages within the timeframe studied.

Zhu, Zhikuan↗