Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “block term decomposition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Large Scale Tensor Factorization via Parallel Sketches

Tensor factorization methods have recently gained increased popularity. A key feature that renders tensors attractive is the ability to directly model multi-relational data. In this work, we propose ParaSketch, a parallel tensor factorization algorithm that enables massive parallelism, to deal with large tensors. The idea is to compress the large tensor into multiple small tensors, decompose each small tensor in parallel, and combine the results to reconstruct the desired latent factors. Prior art in this direction entails potentially very high complexity in the (Gaussian) compression and final combining stages. Adopting sketching matrices for compression, the proposed method enjoys a dramatic reduction in compression complexity, and features a much lighter combining step. Moreover, theoretical analysis shows that the compressed tensors inherit latent identifiability under mild conditions, hence establishing correctness of the overall approach. Numerical experiments corroborate the theory and demonstrate the effectiveness of the proposed algorithm.

block term decomposition↗

Four-point correlators of light-ray operators in CCFT

We compute the four-point correlator of two gluon light-ray operators and two gluon primaries from the four-gluon celestial amplitude in (2, 2) signature spacetime. The correlator is non-distributional and allows us to verify that light-ray operators appear in the OPE of two gluon primaries. We also carry out a conformal block decomposition of the terms involving the exchange of gluon operators.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Mitigating electrochemical degradation in CsPbBr{sub 3} gamma detectors by organic and inorganic encapsulation.

CsPbBr3 perovskite semiconductors have emerged as a leading candidate for nextgeneration radiation detectors because of their exceptional charge transport properties, defect tolerance, and record-breaking sensitivity and energy resolution. Their long-term stability, however, is hindered by electrode-driven electrochemical decomposition, which is accelerated by moisture- and oxygen-assisted ion migration during operation. Here, we investigated organic and inorganic encapsulation strategies as both environmental barriers and means to suppress interfacial degradation pathways. Atomic layer deposition (ALD) of Al2O3 provided a conformal passivation layer that blocked environmental ingress, suppressed ionic diffusion, reduced leakage current, enhanced energy resolution and expanded the operational electric-field window beyond 5 kV∙cm1 . By contrast, organic encapsulants such as paraffin wax and polystyrene slowed moisture diffusion but did not suppress interfacial reactions, with wax extending stability to over 90 days. These results show that ALD-Al2O3 suppresses dominant interfacial degradation pathways, enabling stable, high-field operation and advancing the practical deployment of CsPbBr3 γ-ray detectors.

Unal, Mustafa↗

Multi-color incomplete Cholesky conjugate gradient methods for vector computers

In this research, we are concerned with the solution on vector computers of linear systems of equations, Ax = b, where A is a larger, sparse symmetric positive definite matrix. We solve the system using an iterative method, the incomplete Cholesky conjugate gradient method (ICCG). We apply a multi-color strategy to obtain p-color matrices for which a block-oriented ICCG method is implemented on the CYBER 205. (A p-colored matrix is a matrix which can be partitioned into a pXp block matrix where the diagonal blocks are diagonal matrices). This algorithm, which is based on a no-fill strategy, achieves O(N/p) length vector operations in both the decomposition of A and in the forward and back solves necessary at each iteration of the method. We discuss the natural ordering of the unknowns as an ordering that minimizes the number of diagonals in the matrix and define multi-color orderings in terms of disjoint sets of the unknowns. We give necessary and sufficient conditions to determine which multi-color orderings of the unknowns correpond to p-color matrices. A performance model is given which is used both to predict execution time for ICCG methods and also to compare an ICCG method to conjugate gradient without preconditioning or another ICCG method. Results are given from runs on the CYBER 205 at NASA's Langley Research Center for four model problems.

Poole, E. L.↗

A finite element algorithm for sound propagation in axisymmetric ducts containing compressible mean flow

An accurate mathematical model for sound propagation in axisymmetric aircraft engine ducts with compressible mean flow is reported. The model is based on the usual perturbation of the basic fluid mechanics equations for small motions. Mean flow parameters are derived in the absence of fluctuating quantities and are then substituted into the equations for the acoustic quantities which were linearized by eliminating higher order terms. Mean swirl is assumed to be zero from the restriction of axisymmetry. A linear rectangular serendipity element is formulated from these equations using a Galerkin procedure and assembled in a special purpose computer program in which the matrix map for a rectangular mesh was specifically coded. Representations of the fluctuating quantities, mean quantities and coordinate transformations are isoparametric. The global matrix is solved by foreward and back substitution following an L-U decomposition with pivoting restricted internally to the blocks. Results from the model were compared with results from several alternative analyses and yielded satisfactory agreement.

Abrahamson, A. L.↗

Massively Parallel Dantzig-Wolfe Decomposition Applied to Traffic Flow Scheduling

Optimal scheduling of air traffic over the entire National Airspace System is a computationally difficult task. To speed computation, Dantzig-Wolfe decomposition is applied to a known linear integer programming approach for assigning delays to flights. The optimization model is proven to have the block-angular structure necessary for Dantzig-Wolfe decomposition. The subproblems for this decomposition are solved in parallel via independent computation threads. Experimental evidence suggests that as the number of subproblems/threads increases (and their respective sizes decrease), the solution quality, convergence, and runtime improve. A demonstration of this is provided by using one flight per subproblem, which is the finest possible decomposition. This results in thousands of subproblems and associated computation threads. This massively parallel approach is compared to one with few threads and to standard (non-decomposed) approaches in terms of solution quality and runtime. Since this method generally provides a non-integral (relaxed) solution to the original optimization problem, two heuristics are developed to generate an integral solution. Dantzig-Wolfe followed by these heuristics can provide a near-optimal (sometimes optimal) solution to the original problem hundreds of times faster than standard (non-decomposed) approaches. In addition, when massive decomposition is employed, the solution is shown to be more likely integral, which obviates the need for an integerization step. These results indicate that nationwide, real-time, high fidelity, optimal traffic flow scheduling is achievable for (at least) 3 hour planning horizons.

Rios, Joseph Lucio↗

IDAES-PSE 2.0 Release

The Institute for the Design of Advanced Energy Systems (IDAES) Integrated Platform is a versatile computational environment offering extensive process systems engineering (PSE) capabilities for optimizing the design and operation of complex, interacting technologies and systems. IDAES enables users to efficiently search vast, complex design spaces to discover the lowest cost, most environmentally sustainable solutions while supporting the full process modeling lifecycle, from conceptual design to dynamic optimization and control. The extensible, open platform empowers users to create models of novel processes and rapidly develop custom analyses, workflows, and end-user applications. IDAES-PSE 2.0.0 Release Highlights Removal of deprecated features from IDAES v1 Update to Pyomo v6.5 – this required a number of updates to support the new NL solver writer and to address some changes in Pyomo Creation of new testing suite for backward compatibility, model robustness and verification More general implementation of the Helmholtz EoS. This brings some new features like standard property diagrams, choice of mass or mole basis, and new state variable options Standardizing names in Heat Exchanger models (breaking change from v2.0.0a2): Control Volumes named hot_side and cold_side Ports names hot_side_inlet, hot_side_outlet, cold_side_inlet and cold_side_outlet Config Blocks names hot_side_config and cold_side_config Config arguments for user provided names for each side: hot_side_name and cold_side_name. Updating Keras surrogate tool to use v1.1 of OMLT New prototype API for model initialization (idaes.core.initialization) The new API uses "Model Initializer" objects instead of class methods, allowing for the definition of multiple initialization routines for a single model A number of common, model agnostic initialization routines have also been defined, including initialization from data, block-decomposition and a general hierarchical approach equivalent to the existing method for common unit models New metadata for thermophysical properties – valid_range This can be used to record the range of values over which a property value can be trusted, such as the range of experimental data used to regress parameters A number of new utility functions have been added to check for properties with values outside the valid range and to set bounds based on this metadata Updated construction of balance expressions in Control Volumes to remove unneeded terms In the past, unneeded terms were added as a constant 0 term, however they will now be dropped entirely from the expression This was necessary due to more strict unit checking in the new Pyomo solver writer which no longer ignores 0 terms Updates to metadata for thermophysical properties to better define known properties and units of measurement This results in more strict enforcement of standard naming for thermophysical and reaction properties Users can still define custom properties, but these must be done explicitly using the define_custom_properties() method instead of being implicitly created by add_property() Updated convergence tester utility tool to support definition of benchmark files (JSON format) and comparison of performance to benchmarks Set default iteration limit for IPOPT in IDAES config to 200 iterations Update scaling of example models to work with new Pyomo NL solver writer Improve testing of extensions and examples infrastructure to avoid need for downloading files Updated distillation column to centralize common functionality and remove a number of Pyomo warnings

IDAES↗

On the Virasoro six-point identity block and chaos

We study six-point correlation functions in two dimensional conformal field theory, where the six operators are grouped in pairs with equal conformal dimension. Assuming large central charge $c$ and a sparse spectrum, the leading contribution to this correlation function is the six-point Virasoro identity block - corresponding to each distinct pair of operators fusing into the identity and its descendants. We call this the star channel. One particular term in the star channel identity block is the stress tensor $SL(2,\mathbb{R})$ (global) block, for which we derive an explicit expression. In the holographic context, this object corresponds to a direct measure of nonlinear effects in pure gravity. We calculate additional terms in the star channel identity block that contribute at the same order at large $c$ as the global block using the novel theory of reparametrizations, which extends the shadow operator formalism in a natural way. We investigate these blocks' relevance to quantum chaos in the form of six-point scrambling in an out-of time ordered correlator. Interestingly, the global block does not contribute to the scrambling mode of this correlator, implying that, to leading order, six-point scrambling is insensitive to the three-point graviton coupling in the bulk dual. Finally, we compare our findings with a different OPE channel, called the comb channel, and find the same result for the chaos exponent in this decomposition.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Vector quantization for efficient coding of upper subbands

This paper examines the application of vector quantization (VQ) to exploit both intra-band and inter-band redundancy in subband coding. The focus here is on the exploitation of inter-band dependency. It is shown that VQ is particularly suitable and effective for coding the upper subbands. Three subband decomposition-based VQ coding schemes are proposed here to exploit the inter-band dependency by making full use of the extra flexibility of VQ approach over scalar quantization. A quadtree-based variable rate VQ (VRVQ) scheme which takes full advantage of the intra-band and inter-band redundancy is first proposed. Then, a more easily implementable alternative based on an efficient block-based edge estimation technique is employed to overcome the implementational barriers of the first scheme. Finally, a predictive VQ scheme formulated in the context of finite state VQ is proposed to further exploit the dependency among different subbands. A VRVQ scheme proposed elsewhere is extended to provide an efficient bit allocation procedure. Simulation results show that these three hybrid techniques have advantages, in terms of peak signal-to-noise ratio (PSNR) and complexity, over other existing subband-VQ approaches.

Zeng, W. J.↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗

Efficient Space–Time Reduced Order Model for Linear Dynamical Systems in Python Using Less than 120 Lines of Code

A classical reduced order model (ROM) for dynamical problems typically involves only the spatial reduction of a given problem. Recently, a novel space–time ROM for linear dynamical problems has been developed [Choi et al., Space–tume reduced order model for large-scale linear dynamical systems with application to Boltzmann transport problems, Journal of Computational Physics, 2020], which further reduces the problem size by introducing a temporal reduction in addition to a spatial reduction without much loss in accuracy. The authors show an order of a thousand speed-up with a relative error of less than 10−5 for a large-scale Boltzmann transport problem. In this work, we present for the first time the derivation of the space–time least-squares Petrov–Galerkin (LSPG) projection for linear dynamical systems and its corresponding block structures. Utilizing these block structures, we demonstrate the ease of construction of the space–time ROM method with two model problems: 2D diffusion and 2D convection diffusion, with and without a linear source term. For each problem, we demonstrate the entire process of generating the full order model (FOM) data, constructing the space–time ROM, and predicting the reduced-order solutions, all in less than 120 lines of Python code. We compare our LSPG method with the traditional Galerkin method and show that the space–time ROMs can achieve O(10−3) to O(10−4) relative errors for these problems. Depending on parameter–separability, online speed-ups may or may not be achieved. For the FOMs with parameter–separability, the space–time ROMs can achieve O(10) online speed-ups. Finally, we present an error analysis for the space–time LSPG projection and derive an error bound, which shows an improvement compared to traditional spatial Galerkin ROM methods.

97 MATHEMATICS AND COMPUTING↗

Block-Structured Operator Inference for Coupled Multiphysics Model Reduction

This work presents a block-structured formulation of Operator Inference as a way to learn structured reduced-order models for multiphysics systems. The approach specifies the governing equation structure for each physics component and the structure of the coupling terms. Once the multiphysics structure is specified, the reduced-order model is learned from snapshot data following the nonintrusive Operator Inference methodology. In addition to preserving physical system structure, which in turn permits preservation of system properties such as stability and second-order structure, the block-structured approach has the advantages of reducing the overall dimensionality of the learning problem and admitting tailored regularization for each physics component. The numerical advantages of the block-structured formulation over a monolithic Operator Inference formulation are demonstrated for aeroelastic analysis, which couples aerodynamic and structural models. For the benchmark test case of the AGARD 445.6 wing, block-structured Operator Inference provides an average 20% online prediction speedup over monolithic Operator Inference across subsonic and supersonic flow conditions in both the stable and fluttering parameter regimes while preserving the accuracy achieved with monolithic Operator Inference.

42 ENGINEERING↗