Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

CrossLink: General Overview [Slides]

Problem: Traditional mesh generation approaches are labor intensive and have limited robustness when applied to parametric design exploration and optimization of complex geometries. While automatic mesh generation approaches exist, they tend to generate tetrahedral or mixed-hybrid meshes which are generally unsuitable for physics applications with strong shock waves, thin boundary layers, and strong gradients. In addition, simulations sizes in the billions of cells are becoming more common with traditional mesh generation methods quickly reaching scalability limits. Solution: CrossLink offers a topology-based mesh generation approach with unstructured block-filling methods and a scalable mesh generation engine. In addition, CrossLink incorporates a python-based API for seamless workflow integration and robust repeatability of the geometry handling and mesh generation process. This makes it ideal for parametric design study and optimization of complex geometries. Finally, future versions of CrossLink will offer a parametric mesh capability that optimizes a high-order mesh and enables reconstruction of the final mesh in memory by the physics solver.

97 MATHEMATICS AND COMPUTING↗

MLP: A Parallel Programming Alternative to MPI for New Shared Memory Parallel Systems

Recent developments at the NASA AMES Research Center's NAS Division have demonstrated that the new generation of NUMA based Symmetric Multi-Processing systems (SMPs), such as the Silicon Graphics Origin 2000, can successfully execute legacy vector oriented CFD production codes at sustained rates far exceeding processing rates possible on dedicated 16 CPU Cray C90 systems. This high level of performance is achieved via shared memory based Multi-Level Parallelism (MLP). This programming approach, developed at NAS and outlined below, is distinct from the message passing paradigm of MPI. It offers parallelism at both the fine and coarse grained level, with communication latencies that are approximately 50-100 times lower than typical MPI implementations on the same platform. Such latency reductions offer the promise of performance scaling to very large CPU counts. The method draws on, but is also distinct from, the newly defined OpenMP specification, which uses compiler directives to support a limited subset of multi-level parallel operations. The NAS MLP method is general, and applicable to a large class of NASA CFD codes.

Taft, James R.↗

Evaluating the Potential and Challenges of an Uncertainty Quantification Method for Long Short–Term Memory Models for Soil Moisture Predictions

Recently, recurrent deep networks have shown promise to harness newly available satellite–sensed data for long–term soil moisture projections. However, to be useful in forecasting, deep networks must also provide uncertainty estimates. Here we evaluated Monte Carlo dropout with an input–dependent data noise term (MCD+N), an efficient uncertainty estimation framework originally developed in computer vision, for hydrologic time series predictions. MCD+N simultaneously estimates a heteroscedastic input–dependent data noise term (a trained error model attributable to observational noise) and a network weight uncertainty term (attributable to insufficiently constrained model parameters). Although MCD+N has appealing features, many heuristic approximations were employed during its derivation, and rigorous evaluations and evidence of its asserted capability to detect dissimilarity were lacking. To address this, we provided an in–depth evaluation of the scheme's potential and limitations. We showed that for reproducing soil moisture dynamics recorded by the Soil Moisture Active Passive (SMAP) mission, MCD+N indeed gave a good estimate of predictive error, provided that we tuned a hyperparameter and used a representative training data set. The input–dependent term responded strongly to observational noise, while the model term clearly acted as a detector for physiographic dissimilarity from the training data, behaving as intended. However, when the training and test data were characteristically different, the input–dependent term could be misled, undermining its reliability. Additionally, due to the data–driven nature of the model, data noise also influences network weight uncertainty, and therefore the two uncertainty terms are correlated. Altogether, this approach has promise, but care is needed to interpret the results.

54 ENVIRONMENTAL SCIENCES↗

Efficient Implementation of an Optimal Interpolator for Large Spatial Data Sets

Scattered data interpolation is a problem of interest in numerous areas such as electronic imaging, smooth surface modeling, and computational geometry. Our motivation arises from applications in geology and mining, which often involve large scattered data sets and a demand for high accuracy. The method of choice is ordinary kriging. This is because it is a best unbiased estimator. Unfortunately, this interpolant is computationally very expensive to compute exactly. For n scattered data points, computing the value of a single interpolant involves solving a dense linear system of size roughly n x n. This is infeasible for large n. In practice, kriging is solved approximately by local approaches that are based on considering only a relatively small'number of points that lie close to the query point. There are many problems with this local approach, however. The first is that determining the proper neighborhood size is tricky, and is usually solved by ad hoc methods such as selecting a fixed number of nearest neighbors or all the points lying within a fixed radius. Such fixed neighborhood sizes may not work well for all query points, depending on local density of the point distribution. Local methods also suffer from the problem that the resulting interpolant is not continuous. Meyer showed that while kriging produces smooth continues surfaces, it has zero order continuity along its borders. Thus, at interface boundaries where the neighborhood changes, the interpolant behaves discontinuously. Therefore, it is important to consider and solve the global system for each interpolant. However, solving such large dense systems for each query point is impractical. Recently a more principled approach to approximating kriging has been proposed based on a technique called covariance tapering. The problems arise from the fact that the covariance functions that are used in kriging have global support. Our implementations combine, utilize, and enhance a number of different approaches that have been introduced in literature for solving large linear systems for interpolation of scattered data points. For very large systems, exact methods such as Gaussian elimination are impractical since they require 0(n(exp 3)) time and 0(n(exp 2)) storage. As Billings et al. suggested, we use an iterative approach. In particular, we use the SYMMLQ method, for solving the large but sparse ordinary kriging systems that result from tapering. The main technical issue that need to be overcome in our algorithmic solution is that the points' covariance matrix for kriging should be symmetric positive definite. The goal of tapering is to obtain a sparse approximate representation of the covariance matrix while maintaining its positive definiteness. Furrer et al. used tapering to obtain a sparse linear system of the form Ax = b, where A is the tapered symmetric positive definite covariance matrix. Thus, Cholesky factorization could be used to solve their linear systems. They implemented an efficient sparse Cholesky decomposition method. They also showed if these tapers are used for a limited class of covariance models, the solution of the system converges to the solution of the original system. Matrix A in the ordinary kriging system, while symmetric, is not positive definite. Thus, their approach is not applicable to the ordinary kriging system. Therefore, we use tapering only to obtain a sparse linear system. Then, we use SYMMLQ to solve the ordinary kriging system. We show that solving large kriging systems becomes practical via tapering and iterative methods, and results in lower estimation errors compared to traditional local approaches, and significant memory savings compared to the original global system. We also developed a more efficient variant of the sparse SYMMLQ method for large ordinary kriging systems. This approach adaptively finds the correct local neighborhood for each query point in the interpolation process.

Memarsadeghi, Nargess↗

Towards Enhancing Coding Productivity for GPU Programming Using Static Graphs

The main contribution of this work is to increase the coding productivity of GPU programming by using the concept of Static Graphs. GPU capabilities have been increasing significantly in terms of performance and memory capacity. However, there are still some problems in terms of scalability and limitations to the amount of work that a GPU can perform at a time. To minimize the overhead associated with the launch of GPU kernels, as well as to maximize the use of GPU capacity, we have combined the new CUDA Graph API with the CUDA programming model (including CUDA math libraries) and the OpenACC programming model. We use as test cases two different, well-known and widely used problems in HPC and AI: the Conjugate Gradient method and the Particle Swarm Optimization. In the first test case (Conjugate Gradient) we focus on the integration of Static Graphs with CUDA. In this case, we are able to significantly outperform the NVIDIA reference code, reaching an acceleration of up to 11x thanks to a better implementation, which can benefit from the new CUDA Graph capabilities. In the second test case (Particle Swarm Optimization), we complement the OpenACC functionality with the use of CUDA Graph, achieving again accelerations of up to one order of magnitude, with average speedups ranging from 2x to 4x, and performance very close to a reference and optimized CUDA code. Our main target is to achieve a higher coding productivity model for GPU programming by using Static Graphs, which provides, in a very transparent way, a better exploitation of the GPU capacity. The combination of using Static Graphs with two of the current most important GPU programming models (CUDA and OpenACC) is able to reduce considerably the execution time w.r.t. the use of CUDA and OpenACC only, achieving accelerations of up to more than one order of magnitude. Finally, we propose an interface to incorporate the concept of Static Graphs into the OpenACC Specifications.

58 GEOSCIENCES↗

Perceptual Repetition Blindness Effects

The phenomenon of repetition blindness (RB) may reveal a new limitation on human perceptual processing. Recently, however, researchers have attributed RB to post-perceptual processes such as memory retrieval and/or reporting biases. The standard rapid serial visual presentation (RSVP) paradigm used in most RB studies is, indeed, open to such objections. Here we investigate RB using a "single-frame" paradigm introduced by Johnston and Hale (1984) in which memory demands are minimal. Subjects made only a single judgement about whether one masked target word was the same or different than a post-target probe. Confidence ratings permitted use of signal detection methods to assess sensitivity and bias effects. In the critical condition for RB a precue of the post-target word was provided prior to the target stimulus (identity precue), so that the required judgement amounted to whether the target did or did not repeat the precue word. In control treatments, the precue was either an unrelated word or a dummy.

Hochhaus, Larry↗

Non-adiabatic approximations in time-dependent density functional theory: progress and prospects

Time-dependent density functional theory continues to draw a large number of users in a wide range of fields exploring myriad applications involving electronic spectra and dynamics. Although in principle exact, the predictivity of the calculations is limited by the available approximations for the exchange-correlation functional. In particular, it is known that the exact exchange-correlation functional has memory-dependence, but in practise adiabatic approximations are used which ignore this. Here we review the development of non-adiabatic functional approximations, their impact on calculations, and challenges in developing practical and accurate memory-dependent functionals for general purposes.

36 MATERIALS SCIENCE↗

Tracking 3-D body motion for docking and robot control

An advanced method of tracking three-dimensional motion of bodies has been developed. This system has the potential to dynamically characterize machine and other structural motion, even in the presence of structural flexibility, thus facilitating closed loop structural motion control. The system's operation is based on the concept that the intersection of three planes defines a point. Three rotating planes of laser light, fixed and moving photovoltaic diode targets, and a pipe-lined architecture of analog and digital electronics are used to locate multiple targets whose number is only limited by available computer memory. Data collection rates are a function of the laser scan rotation speed and are currently selectable up to 480 Hz. The tested performance on a preliminary prototype designed for 0.1 in accuracy (for tracking human motion) at a 480 Hz data rate includes a worst case resolution of 0.8 mm (0.03 inches), a repeatability of plus or minus 0.635 mm (plus or minus 0.025 inches), and an absolute accuracy of plus or minus 2.0 mm (plus or minus 0.08 inches) within an eight cubic meter volume with all results applicable at the 95 percent level of confidence along each coordinate region. The full six degrees of freedom of a body can be computed by attaching three or more target detectors to the body of interest.

Donath, M.↗

Wind Tunnel Tests on Autorotation and the "Flat Spin."

This report deals with the autorotational characteristics of certain differing wing systems as determined from wind tunnel tests made at the Langley Memorial Aeronautical Laboratory. The investigation was confined to autorotation about a fixed axis in the plane of symmetry and parallel to the wind direction. Analysis of the tests leads to the following conclusions: autorotation below 30 degree angle of attack is governed chiefly by wing profile, and above that angle by wing arrangement. The strip method of autorotation analysis gives uncertain results between maximum C subscript L and 35 degrees. The polar curve of a wing system, and to a lower degree of accuracy the polar of a complete airplane model are sufficient for direct determination of the limits of rotary instability, subject to strip method limitations. The results of the investigation indicate that in free flight a monoplane is incapable of flat spinning, whereas an unstaggered biplane has inherent flat-spinning tendencies. The difficulty of maintaining equilibrium in stalled flight is due primarily to rotary instability, a rapid change from stability to instability occurring as the angle of maximum lift is exceeded. (author)

Knight, Montgomery↗

Numerical Prediction Methods (Reynolds-Averaged Navier-Stokes Simulations of Transonic Separated Flows)

During the past five years, numerous pioneering archival publications have appeared that have presented computer solutions of the mass-weighted, time-averaged Navier-Stokes equations for transonic problems pertinent to the aircraft industry. These solutions have been pathfinders of developments that could evolve into a major new technological capability, namely the computational Navier-Stokes technology, for the aircraft industry. So far these simulations have demonstrated that computational techniques, and computer capabilities have advanced to the point where it is possible to solve forms of the Navier-Stokes equations for transonic research problems. At present there are two major shortcomings of the technology: limited computer speed and memory, and difficulties in turbulence modelling and in computation of complex three-dimensional geometries. These limitations and difficulties are the pacing items of the continuing developments, although the one item that will most likely turn out to be the most crucial to the progress of this technology is turbulence modelling. The objective of this presentation is to discuss the state of the art of this technology and suggest possible future areas of research. We now discuss some of the flow conditions for which the Navier-Stokes equations appear to be required. On an airfoil there are four different types of interaction of a shock wave with a boundary layer: (1) shock-boundary-layer interaction with no separation, (2) shock-induced turbulent separation with immediate reattachment (we refer to this as a shock-induced separation bubble), (3) shock-induced turbulent separation without reattachment, and (4) shock-induced separation bubble with trailing edge separation.

Mehta, Unmeel↗

Ensuring the relocatability of programs in the operational system DOS YeS

Specific modifications in the Disk Operational System Unified Series to insure the relocatability of programs stored permanently in the core image library is described. A self-relocating method for loading programs into the working memory with re-editing all the programs recorded in the core image library is presented. The modified linkage editor can be included in a relocation dictionary containing data about each address constant at the assembly stage at the request of the programmer. The relocation dictionary increases the dimension of the RL-phase in comparison with the dimension of this same phase when edited by the standard method, making possible the creation of multiphase program complexes. Generation and use of the modified system using Assembly language is described. An example of the use of the system is given, and limitations of the use of the relocatable programs in the modified system are outlined.

Novoseltsev, S. K.↗

On the Information Content of Program Traces

Program traces are used for analysis of program performance, memory utilization, and communications as well as for program debugging. The trace contains records of execution events generated by monitoring units inserted into the program. The trace size limits the resolution of execution events and restricts the user's ability to analyze the program execution. We present a study of the information content of program traces and develop a coding scheme which reduces the trace size to the limit given by the trace entropy. We apply the coding to the traces of AIMS instrumented programs executed on the IBM SPA and the SCSI Power Challenge and compare it with other coding methods. Our technique shows size of the trace can be reduced by more than a factor of 5.

Frumkin, Michael↗

Implicit fast sweeping method for hyperbolic systems of conservation laws

Implicit time-accurate methods are often used to integrate stiff problems where explicit schemes impose severe time step restrictions. This paper presents an efficient numerical framework based on the Fast Sweeping Method (FSM) for solving linear and nonlinear hyperbolic systems of conservation laws. The solution at each discrete location is computed by sweeping the numerical domain in several predetermined directions that follow the causality of the characteristic families. The use of a fractional step strategy eliminates the need for a solution selection criterion while one-sided stencils limit the number of sweeps to at most 2 d for d space dimensions. This work focuses on the first-order implicit upwind method since it constitutes the building block for high-order conservative schemes. For problems where the degree of stiffness evolves over time, implicit-explicit hybridization can be accomplished with the same algorithm by simply switching the stencil at each time level. As opposed to traditional implicit solvers, the sweeping method does not require a local time linearization of the fluxes thereby preserving the nonlinear stability properties of the original implicit scheme. It also avoids the large computational and memory requirements associated with solving large block-diagonal systems of equations. Here, a series of one- and two-dimensional test cases are presented for the inviscid Burgers' equation and the reactive Euler equations. The results indicate that the implicit FSM can allow a major reduction in the number of time steps even in the presence of discontinuous solution profiles.

74 ATOMIC AND MOLECULAR PHYSICS↗

Application of Fast Multipole Methods to the NASA Fast Scattering Code

The NASA Fast Scattering Code (FSC) is a versatile noise prediction program designed to conduct aeroacoustic noise reduction studies. The equivalent source method is used to solve an exterior Helmholtz boundary value problem with an impedance type boundary condition. The solution process in FSC v2.0 requires direct manipulation of a large, dense system of linear equations, limiting the applicability of the code to small scales and/or moderate excitation frequencies. Recent advances in the use of Fast Multipole Methods (FMM) for solving scattering problems, coupled with sparse linear algebra techniques, suggest that a substantial reduction in computer resource utilization over conventional solution approaches can be obtained. Implementation of the single level FMM (SLFMM) and a variant of the Conjugate Gradient Method (CGM) into the FSC is discussed in this paper. The culmination of this effort, FSC v3.0, was used to generate solutions for three configurations of interest. Benchmarking against previously obtained simulations indicate that a twenty-fold reduction in computational memory and up to a four-fold reduction in computer time have been achieved on a single processor.

Dunn, Mark H.↗

Efficient Gradient-Based Shape Optimization Methodology Using Inviscid/Viscous CFD

The formerly developed preconditioned-biconjugate-gradient (PBCG) solvers for the analysis and the sensitivity equations had resulted in very large error reductions per iteration; quadratic convergence was achieved whenever the solution entered the domain of attraction to the root. Its memory requirement was also lower as compared to a direct inversion solver. However, this memory requirement was high enough to preclude the realistic, high grid-density design of a practical 3D geometry. This limitation served as the impetus to the first-year activity (March 9, 1995 to March 8, 1996). Therefore, the major activity for this period was the development of the low-memory methodology for the discrete-sensitivity-based shape optimization. This was accomplished by solving all the resulting sets of equations using an alternating-direction-implicit (ADI) approach. The results indicated that shape optimization problems which required large numbers of grid points could be resolved with a gradient-based approach. Therefore, to better utilize the computational resources, it was recommended that a number of coarse grid cases, using the PBCG method, should initially be conducted to better define the optimization problem and the design space, and obtain an improved initial shape. Subsequently, a fine grid shape optimization, which necessitates using the ADI method, should be conducted to accurately obtain the final optimized shape. The other activity during this period was the interaction with the members of the Aerodynamic and Aeroacoustic Methods Branch of Langley Research Center during one stage of their investigation to develop an adjoint-variable sensitivity method using the viscous flow equations. This method had algorithmic similarities to the variational sensitivity methods and the control-theory approach. However, unlike the prior studies, it was considered for the three-dimensional, viscous flow equations. The major accomplishment in the second period of this project (March 9, 1996 to March 8, 1997) was the extension of the shape optimization methodology for the Thin-Layer Navier-Stokes equations. Both the Euler-based and the TLNS-based analyses compared with the analyses obtained using the CFL3D code. The sensitivities, again from both levels of the flow equations, also compared very well with the finite-differenced sensitivities. A fairly large set of shape optimization cases were conducted to study a number of issues previously not well understood. The testbed for these cases was the shaping of an arrow wing in Mach 2.4 flow. All the final shapes, obtained either from a coarse-grid-based or a fine-grid-based optimization, using either a Euler-based or a TLNS-based analysis, were all re-analyzed using a fine-grid, TLNS solution for their function evaluations. This allowed for a more fair comparison of their relative merits. From the aerodynamic performance standpoint, the fine-grid TLNS-based optimization produced the best shape, and the fine-grid Euler-based optimization produced the lowest cruise efficiency.

Baysal, Oktay↗

A long short-term memory embedding for hybrid uplifted reduced order models

In this paper, we introduce an uplifted reduced order modeling (UROM) approach through the integration of standard projection based methods with long short-term memory (LSTM) embedding. Our approach has three modeling layers or components. In the first layer, we utilize an intrusive projection approach to model dynamics represented by the largest modes. The second layer consists of an LSTM model to account for residuals beyond this truncation. This closure layer refers to the process of including the residual effect of the discarded modes into the dynamics of the largest scales. However, the feasibility of generating a low rank approximation tails off for higher Kolmogorov n -width systems due to the underlying nonlinear processes. The third uplifting layer, called super-resolution, addresses this limited representation issue by expanding the span into a larger number of modes utilizing the versatility of LSTM. Therefore, our model integrates a physics-based projection model with a memory embedded LSTM closure and an LSTM based super-resolution model. In several applications, we exploit the use of Grassmann manifold to construct UROM for unseen conditions. We performed numerical experiments by using the Burgers and Navier-Stokes equations with quadratic nonlinearity. Finally, our results show robustness of the proposed approach in building reduced order models for parameterized systems and confirm the improved trade-off between accuracy and efficiency.

42 ENGINEERING↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Towards Accurate and Efficient Predictions of Martensitic Transition Temperatures for Shape Memory Alloys from First Principles

Shape memory alloys (SMAs) can remember and recover their original shapes upon heating due to the existence of a reversible martensitic transition (MT) between the high-temperature austenite (A) and low-temperature martensite (M) phases. The martensitic transition temperature (MTT) is a crucial characteristic of an SMA. SMAs have a wide range of potential applications in aerospace, civil engineering, bioengineering, etc., but their operating temperatures are limited by the available SMAs. MTT can be tuned by alloying a binary with other metals, and the multicomponent NiTi-based SMAs have attracted tremendous research efforts recently. It is not efficient to employ the trial-and-error method alone due to the dramatically increased complexity and possibilities in compositions, and thus reliable theory and accurate computations play an indispensable role in creating SMAs with desirable properties.

Zhigang Wu↗