Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Tensor representation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

85 records · Page 5

A representation for the turbulent mass flux contribution to Reynolds-stress and two-equation closures for compressible turbulence

The turbulent mass flux, or equivalently the fluctuating Favre velocity mean, appears in the first and second moment equations of compressible kappa-epsilon and Reynolds stress closures. Mathematically it is the difference between the unweighted and density-weighted averages of the velocity field and is therefore a measure of the effects of compressibility through variations in density. It appears to be fundamental to an inhomogeneous compressible turbulence, in which it characterizes the effects of the mean density gradients, in the same way the anisotropy tensor characterizes the effects of the mean velocity gradients. An evolution equation for the turbulent mass flux is derived. A truncation of this equation produces an algebraic expression for the mass flux. The mass flux is found to be proportional to the mean density gradients with a tensor eddy-viscosity that depends on both the mean deformation and the Reynolds stresses. The model is tested in a wall bounded DNS at Mach 4.5 with notable results.

Ristorcelli, J. R.↗

Bridging the Gap Between LLMs and LNS with Dynamic Data Format and Architecture Codesign

Deep neural networks (DNNs) have achieved tremendous success in the past few years. However, their training and inference demand exceptional computational and memory resources. Quantization has been shown as an effective approach to mitigate the cost, with the mainstream data types reduced from FP32 to FP16/BF16 and recently FP8 in the latest NVIDIA H100 GPUs. With increasingly aggressive quantization, however, the conventional floating-point formats suffer from limited precision in representing numbers around zero. Recently, NVIDIA demonstrated the potential of using a Logarithmic Number System (LNS) for the next generation of tensor cores. While LNS mitigates the hurdles in representing small numbers, in this work we observed a mismatch between LNS and the emerging Large Language Models (LLM), where LLM exhibits significant outliers when directly adopting the LNS format. In this paper, we present a data-format/architecture codesign to bright this gap. On the format side, we propose a dynamic LNS format to flexibly represent outliers at a higher precision, by exploiting asymmetry in the LNS representation and identifying outliers through a per-vector basis. On the architecture side, for demonstration, we realize the dynamic LNS format in a systolic array, which can handle the irregularity of the outliers at runtime. We implement our approach on an Alveo U280 FPGA as a prototype. Experimental results show that our design can effectively handle the outliers and resolve the mismatch between LNS and LLM, contributing to an accuracy improvement of 15.4% and 16% over the floating-point and the original LNS baselines, using four state-of-the-art LLM models. Our observation and design lay a solid foundation for the large-scale adoption of the LNS format in the next-generation deep learning hardware.

Haghi, Pouya↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

Classical and quantum computing of shear viscosity for ( 2 + 1 ) D SU(2) gauge theory

We perform a nonperturbative calculation of the shear viscosity for ( 2 + 1 )-dimensional SU(2) gauge theory by using the lattice Hamiltonian formulation. The retarded Green’s function of the stress-energy tensor is calculated from real time evolution via exact diagonalization of the lattice Hamiltonian with a local Hilbert space truncation, and the shear viscosity is obtained via the Kubo formula. When taking the continuum limit, we account for the renormalization group flow of the coupling but no additional operator renormalization. We find the ratio of the shear viscosity and the entropy density η s is consistent with a well-known holographic result 1 4 π at several temperatures on a 4 × 4 honeycomb lattice with the local electric representation truncated at j max = 1 2 . We also find the ratio of the spectral function and frequency ρ x y ( ω ) ω exhibits a peak structure when the frequency is small. Both the exact diagonalization method and simple matrix product state classical simulation method beyond j max = 1 2 on bigger lattices require exponentially growing resources. So we develop a quantum computing method to calculate the retarded Green’s function and analyze various systematics of the calculation including j max truncation and finite size effects, Trotter errors and the thermal state preparation efficiency. Our thermal state preparation method still requires resources that grow exponentially with the lattice size, but with a very small prefactor at high temperature. We test our quantum circuit on both the Quantinuum emulator and the IBM simulator for a small lattice and obtain results consistent with the classical computing ones. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Theory of electromagnetic transmission structures. I - Relativistic foundation and network formalisms

A new theorem on a class of four-dimensional skew-symmetric tensors is demonstrated. Coupled with the relativistic covariant form of Maxwell's equations, this theorem consolidates the classifications of guided waves by combining the three types - TE, TM, TEM - under a uniform condition applied to the generating four-potential which is Lorentz invariant. Each type corresponds to a potential of which a pair of the four components vanishes in a particular frame. Through appropriate normalization conditions, the resulting time-domain equations for the field amplitudes are readily reduced to modified telegraphist equations, which in turn lead to distributed network representations for each of the three types. The ambiguity of distributed network formalisms in general is elucidated and the concept of network parameter densities such as traditionally employed in TEM transmission line theory is questioned.

Gabriel, G. J.↗

Spin-orbit correlations in the nucleon in the large- N c limit

We study the twist-3 spin-orbit correlations of quarks described by the nucleon matrix elements of the parity-odd rank-2 tensor QCD operator (the parity-odd partner of the QCD energy-momentum tensor). Our treatment is based on the effective dynamics emerging from the spontaneous breaking of chiral symmetry and the mean-field picture of the nucleon in the large- N c limit. The twist-3 QCD operators are converted to effective operators, in which the QCD interactions are replaced by spin-flavor-dependent chiral interactions of the quarks with the pion field. We compute the nucleon matrix elements of the twist-3 effective operators and discuss the role of the chiral interactions in the spin-orbit correlations. We derive the first-quantized representation in the mean-field picture and develop a quantum-mechanical interpretation. The chiral interactions give rise to new spin-orbit couplings and qualitatively change the correlations compared to the quark model picture. We also derive the twist-3 matrix elements in the topological soliton picture where the quarks are integrated out (skyrmion). The methods used here can be extended to other QCD operators describing higher-twist nucleon structure and generalized parton distributions. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Visual Understanding of Circular Dichroism Spectroscopy

Mapping chemical and structural properties to electronic and magnetic responses is critical to many applications such as quantum information science, where the precise storage and transmission of unique information is paramount. Specifically, constructing molecules and materials that provide strong polarized responses at tunable frequencies and with large anisotropies is key to optical processing of quantum information. Chiral molecules provide chiroptical response to circularly polarized light, making them attractive for quantum information science and other applications related to sensing, polarized photodetectors, and spintronics. Predicting a molecular design, a priori, with large anisotropies to circularly polarized light is challenging due to the complex interplay between electric and magnetic components of the optical response. In this work, we explore a visual representation of the electronic chiroptical response by decomposing the rotary strength into its constituent components. Here, we make use of the intuitive electronic oscillator framework to develop classical intuition regarding the rotary strength and its constituents. We explore three model chemical systems that exhibit local and global chirality. Our analysis reveals that local chirality necessarily exhibits competition between the local chiral center and chirality induced in other fragments of the molecule, resulting in both unexpected nonmonotonic trends and sign flips in chemically adjacent geometries. Furthermore, we can visually distinguish between local and global chirality via examination of the transition chiral tensor. Interestingly, we make strong connections to ferromagnetic and antiferromagnetic spin systems in that chiroptically inactive transitions exhibit antiferromagnetic-like alternating orbital patterns while active transitions show domain formation in an ferromagnetic-like alignment that produces a net chiroptical response.

36 MATERIALS SCIENCE↗

Spatial Signatures of Electron Correlation in Least-Squares Tensor Hypercontraction

Least Squares Tensor Hypercontraction (LS-THC) has received some attention in recent years as an approach to reduce the significant computational costs of wavefunc- tion based methods in quantum chemistry. However, previous work has demonstrated that the LS-THC factorization performs disproportionately worse in the description of wavefunction components (e.g. cluster amplitudes T 2 ) than Hamiltonian compo- nents (e.g. electron repulsion integrals (pq|rs)). This work develops novel theoretical methods to study the source of these errors in the context of the real-space T 2 kernel, and reports, for the first time, the existence of a “correlation feature” in the errors of the LS-THC representation of the “exchange-like” correlation energy EX and T 2 that is remarkably consistent across ten molecular species, three correlated wavefunctions, and four basis sets. This correlation feature portends the existence of a “pair-point kernel” missing in the usual LS-THC representation of the wavefunction, which critically depends upon pairs of grid points situated close to atoms and with inter-pair distances between one and two Bohr radii. These findings point the way for future LS-THC developments to address these shortcomings.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Remote, Noncontact Strain Sensing by Laser Diffraction Developed

A system was developed at the NASA Glenn Research Center for continually monitoring, in real time, the in-plane strain tensor in opaque solids during high-temperature, long-term mechanical testing. The simple, noncontacting, strain-sensing methodology should also be suitable for measurement in hostile environments. This procedure has obvious advantages over traditional, mechanical, contacting techniques, and it is easier to interpret than moir and speckle interferometric approaches. A two-dimensional metallic grid of micrometer dimensions is applied to a metallographically prepared gauge section on the surface of a tensile test specimen by a standard photolithographic process. The grid on the fixtured specimen is interrogated by an He-Ne laser, and the resulting diffraction pattern is projected backwards onto a translucent screen. A charge-coupled device (CCD) camera is used to image the first-order diffraction peaks from the translucent screen. A schematic representation of the system is shown in the figure.

Freedman, Marc R.↗

IRMA

IRMA (In)elastic Representation of Materials As S(α,β) evaluations IRMA turns one phonon model into three outputs that usually require three separate tool chains: an evaluated nuclear-data file, predicted neutron-scattering spectra, and scattering kernels for Monte Carlo transport. The three outputs draw on a single, consistent description of the material, so the evaluation, the spectroscopy that can validate it, and the transport that uses it always agree about the physics. Nuclear data. IRMA writes ENDF-6 File 7 thermal scattering evaluations on automatically constructed (α, β) grids. This part reimplements and generalizes NJOY's LEAPR: the classic kernels reproduce freshly generated NJOY2016 tapes digit for digit and published reference tapes to about 1e-4, and the generalized paths add the exact coherent one-phonon term, anisotropic Debye-Waller tensors, coherent elastic for arbitrary crystals, and a per-species partition for polyatomic materials. The tapes feed NJOY, AMPX, FUDGE, and every transport code downstream of them. Neutron spectroscopy. The irma.spectra forward model projects the same physics onto an instrument's kinematics and resolution: INS spectra for VISION and generic indirect geometries, and 2-D S(Q,E) powder maps for direct-geometry spectrometers, from a phonopy model or straight from a phonon DOS. It can be used to predict a proposed measurement before beam time; in analysis, it supplies the calculated single-scattering counterpart of a measured spectrum, from the same material description the evaluation was built from. Monte Carlo transport. The irma.ncrystal exporter writes per-temperature scattering kernels for the companion NCrystal plugin, so McStas, OpenMC, and other NCrystal-aware codes sample the same physics. The exported kernels carry the per-site anisotropic Debye-Waller tensors, keeping directional coherent-elastic physics that NCrystal's standard scalar treatment does not represent. With the same physics inside a transport code, an entire beamline becomes a virtual experiment: IRMA's end-to-end validation ran a custom McStas implementation of the ARCS spectrometer, assembled from the existing McVine and McStas models, against measured data. From a bare crystal structure. The irma mlip front end builds the phonon model itself: a structure file and a choice of potential are enough. Nine pretrained machine-learned interatomic potentials are supported, on a laptop CPU, with no first-principles calculation; an approximate phonon model for a new material costs minutes, not a DFT campaign, and the build emits prefilled inputs for all three outputs. The result is a good starting point rather than a finished evaluation: survey-quality physics with every parameter exposed for review. A converged atomistic calculation enters the same way, as a phonopy model, when higher fidelity is needed.

Ramic, Kemal [Oak Ridge National Laboratory (ORNL)↗

Neural chaos: A spectral stochastic neural operator

Building surrogate models for operators with uncertainty quantification capabilities is essential for many engineering applications where randomness–such as variability in material properties, boundary conditions, and initial conditions–is unavoidable. Polynomial Chaos Expansion (PCE) is widely recognized as a go-to method for constructing stochastic surrogates in both intrusive and non-intrusive ways, and it has recently been used in the context of operator learning. However, its application becomes challenging for complex or high-dimensional processes, as achieving accuracy requires higher-order polynomials, which can increase computational demand and/or the risk of overfitting. Furthermore, PCE requires specialized treatments to manage random variables that are not independent, and these treatments may be problem-dependent or may fail with increasing complexity. Here, in this work, we adopt the same formalism as the spectral expansion used in PCE; however, we replace the classical polynomial basis functions with neural network (NN) basis functions to leverage their expressivity. To achieve this, we propose an algorithm that identifies NN-parameterized basis functions in a purely data-driven manner, without any prior assumptions about the joint distribution of the random variables involved, whether independent or dependent, or about their marginal distributions. The proposed algorithm identifies each NN-parameterized basis function sequentially, ensuring they are orthogonal with respect to the data distribution. The basis functions are constructed directly on the joint stochastic variables without requiring a tensor product structure or assuming independence of the random variables. This approach may offer greater flexibility for complex stochastic models, while simplifying implementation compared to the tensor product structures typically used in PCE to handle random vectors. This is particularly advantageous given the current state of open-source packages, where building and training neural networks can be done with just a few lines of code and extensive community support. We demonstrate the effectiveness of the proposed scheme through several numerical examples of varying complexity and provide comparisons with classical PCE.

Polynomial chaos expansion↗

Renormalization group methods for the Reynolds stress transport equations

The Yakhot-Orszag renormalization group is used to analyze the pressure gradient-velocity correlation and return to isotropy terms in the Reynolds stress transport equations. The perturbation series for the relevant correlations, evaluated to lowest order in the epsilon-expansion of the Yakhot-Orszag theory, are infinite series in tensor product powers of the mean velocity gradient and its transpose. Formal lowest order Pade approximations to the sums of these series produce a rapid pressure strain model of the form proposed by Launder, Reece, and Rodi, and a return to isotropy model of the form proposed by Rotta. In both cases, the model constants are computed theoretically. The predicted Reynolds stress ratios in simple shear flows are evaluated and compared with experimental data. The possibility is discussed of deriving higher order nonlinear models by approximating the sums more accurately. The Yakhot-Orszag renormalization group provides a systematic procedure for deriving turbulence models. Typical applications have included theoretical derivation of the universal constants of isotropic turbulence theory, such as the Kolmogorov constant, and derivation of two equation models, again with theoretically computed constants and low Reynolds number forms of the equations. Recent work has applied this formalism to Reynolds stress modeling, previously in the form of a nonlinear eddy viscosity representation of the Reynolds stresses, which can be used to model the simplest normal stress effects. The present work attempts to apply the Yakhot-Orszag formalism to Reynolds stress transport modeling.

Rubinstein, R.↗

Computation of Effective Mechanical Properties and Mechanical Erosion Modeling of TPS Materials

The goal of this presentation is to provide a general overview of the multi-scale modeling formulation to determine if there is additional surface recession in Thermal Protection Systems (TPS) materials as a result of mechanical erosion due to high shear conditions during atmospheric entry. This modeling process is performed at different scales by leveraging two computational frameworks developed at NASA: the Porous Microstructure Analysis (PuMA) software, and the Porous material Analysis Toolbox based on OpenFOAM (PATO). PuMA specializes in computing effective macro-scale material properties by performing material response simulations on 3D digital micro-scale representations of porous micro-structures. The modeling of TPS materials at the micro-scale is essential to understand how they behave as part of a heat shield assembly. The first part of the presentation will detail the implementation of PuMA’s cell-centered finite volume elasticity solver, which allows the computation of macro-scale effective mechanical properties of heterogeneous and anisotropic materials such as fibrous and woven TPS composites. These homogenized mechanical properties are used by PATO’s mechanical erosion model to predict the TPS material’s recession at a larger scale. This work will also provide some examples of multi-scale analysis from the fiber level up to the unit cell. The second part of the presentation will focus on the macro-scale approach to determine if erosion at the heat shield’s surface occurs due to mechanical and thermal loads experienced during atmospheric entry. To accomplish this, a solid mechanics module was integrated within PATO enabling it to model the potential mechanical erosion in three steps: first, after obtaining the effective mechanical properties with PuMA, the implemented stress analysis solver computes the stress and the displacement fields for the TPS material using the wall shear stress tensor, computed using a CFD solver, as boundary conditions; then, regions on the surface where the stress meets the failure criteria are identified; finally, the failed material is removed and the mesh is redistributed accordingly. The outcome is a model capable of predicting the total recession in the material due to surface chemistry and mechanical erosion.

Mechanical Properties↗