Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Geoid Recovery using Geophysical Inverse Theory Applied to Satellite to Satellite Tracking Data

This report describes a new method for determination of the geopotential. The analysis is aimed at the GRACE mission. This Satellite-to-Satellite Tracking (SST) mission is viewed as a mapping mission The result will be maps of the geoid. The elements of potential theory, celestial mechanics, and Geophysical Inverse Theory are integrated into a computation architecture, and the results of several simulations presented Centimeter accuracy geoids with 50 to 100 km resolution can be recovered with a 30 to 60 day mission.

Gaposchkin, E. M.↗

Advancement of hybrid fluid-kinetic modeling for HEDP and ICF science

We report on the development progress of a hybrid fluid-kinetic code for simulating fluids and plasmas in a wide range of environments, such as laser–matter interactions, inertial confinement fusion, magnetic confinement fusion, and pulsed power. The suite of numerical tools under development utilizes heterogeneous computer architectures and leverages the benefits of particle–based simulation techniques. By working to combine the kinetic particle-in-cell (PIC) model with a particle-based fluid simulation technique, such as smoothed particle hydrodynamics, we are developing a flexible framework capable of accurately modeling complex flows within and between kinetic and fluid regimes. The TriForce code is under development as a C++ framework for parallel, 3D, particle-based, hybrid fluid-kinetic plasma simulations. The fluid half of TriForce will be based upon the meshless smoothed-particle-hydrodynamics (SPH) approach, well-suited for shear, mixing, and turbulence, whereas the kinetic half resembles a traditional particle-in-cell (PIC) code; other particle-based approaches to fluid modeling that do use a mesh are also possible to use and are under investigation. Maxwell’s electromagnetic field equations are solved either via explicit or implicit algorithms or approximated via resistive magnetohydrodynamics (MHD) using an Ohm’s law and resulting induction equation (extended MHD is under development). A primary goal of enabling direct comparisons, from the same code, between results from the variants of MHD and implicit electromagnetic solutions is to improve our fundamental understanding of systems with magnetic fields. The code is under development to recover results from both radiation-MHD and fully kinetic codes in those limits, and is continuing to be developed from other follow-on grants to operate in between where both descriptions may co-exist and interact. For certain applications, it is desired for a simulation to contain fluid ions and electrons as well as kinetic ions and electrons. Typically, it is too computationally intensive to model a full-scale ICF or HEDP experiment fully kinetically since many cycles are expended with very small time steps on modeling the fluid part of a material that is well treated by the fluid approximation. In this case, many traditional PIC particles can be replaced with a single fluid particle representing the thermal part of the distribution function, and there are fewer needed kinetic particles, which describe the non-thermal part and can be sub-cycled relative to the fluid particle advance. Furthermore, a pure fluid code may, depending on the problem, simply lack many physically important details that are beyond the scope of its reduced approximations and assumptions. In this report, we summarize the objectives achieved in the development of the collisional and kinetic half of the code, and the physics problems to which the code has been applied in the areas of advanced and innovative fusion concepts, pulsed power, and magneto-inertial fusion.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Scale-Resolving Simulations of Low-Pressure Turbine Cascades with Wall Roughness Using A Spectral-Element Method

The accurate prediction of wall-roughness effects in turbomachinery is becoming critical as turbine designers address airfoil surface quality and degradation concerns arising from the shift to advanced ceramic matrix composite (CMC) or additively-manufactured airfoils operating in higher temperature environments. In this paper, a recently developed computational capability for accurate and efficient scale-resolving simulations of turbomachinery is extended to analyze the boundary- layer separation and transition characteristics in a rough-wall low-pressure turbine (LPT) cascade. The computational capability is based on an entropy-stable discontinuous-Galerkin spectral-element approach that extends to arbitrarily high orders of spatial and temporal accuracy, and is implemented in an efficient manner for a modern high performance computer architecture. Results from the scale-resolving simulations of both smooth and rough airfoil cascades are presented and compared to previous experiments and numerical simulations. The results show that the suction surface boundary layer undergoes laminar separation, transition, and turbulent reattachment for the smooth airfoil cascade, while in the presence of roughness the separation and transition behavior of the suction surface boundary layer is substantially modified. The differences between the smooth and rough airfoil cascades are then highlighted by a detailed analysis of their respective turbulent flow fields.

Spectral-Element↗

Fast Parallel Computation Of Manipulator Inverse Dynamics

Method for fast parallel computation of inverse dynamics problem, essential for real-time dynamic control and simulation of robot manipulators, undergoing development. Enables exploitation of high degree of parallelism and, achievement of significant computational efficiency, while minimizing various communication and synchronization overheads as well as complexity of required computer architecture. Universal real-time robotic controller and simulator (URRCS) consists of internal host processor and several SIMD processors with ring topology. Architecture modular and expandable: more SIMD processors added to match size of problem. Operate asynchronously and in MIMD fashion.

Fijany, Amir↗

Fast Simulation of the NICER Instrument

The NICER mission uses a complicated physical system to collect information from objects that are, by x-ray timing science standards, rather faint. To get the most out of the data we will need a rigorous understanding of all instrumental effects. We are in the process of constructing a very fast, high fidelity simulator that will help us to assess instrument performance, support simulation-based data reduction, and improve our estimates of measurement error. We will combine and extend existing optics, detector, and electronics simulations. We will employ the Compute Unied Device Architecture (CUDA2) to parallelize these calculations. The price of suitable CUDA-compatible multi-gigaflop cores is about $0.20/core, so this approach will be very cost-effective.

gEDA↗

KinCat v.1.0

SAND2024-02099O The software is designed to allow researchers to perform kinetic Monte Carlo (KMC) simulations of catalytic reactions on a 2D lattice. The code is written efficiently to run on a variety of shared memory computing architectures (e.g. GPU, multi-core) and to natively express the full complexity of lateral interactions on reaction rates. The software allows researchers to perform KMC simulations of catalytic reactions on a 2D lattice. It uses parallel shared-memory computing architectures to reduce run-times and allows for simultaneous simulation of multiple independent runs. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Najm, Habib↗

Project Integration Architecture (PIA) and Computational Analysis Programming Interface (CAPRI) for Accessing Geometry Data from CAD Files

Integration of a supersonic inlet simulation with a computer aided design (CAD) system is demonstrated. The integration is performed using the Project Integration Architecture (PIA). PIA provides a common environment for wrapping many types of applications. Accessing geometry data from CAD files is accomplished by incorporating appropriate function calls from the Computational Analysis Programming Interface (CAPRI). CAPRI is a CAD vendor neutral programming interface that aids in acquiring geometry data directly from CAD files. The benefits of wrapping a supersonic inlet simulation into PIA using CAPRI are; direct access of geometry data, accurate capture of geometry data, automatic conversion of data units, CAD vendor neutral operation, and on-line interactive history capture. This paper describes the PIA and the CAPRI wrapper and details the supersonic inlet simulation demonstration.

Benyo, Theresa L.↗

Scalable multiscale modeling of platelets with 100 million particles

Here, we developed the core components of the AI-aided multiple time stepping algorithm for multiscale modeling of cell dynamics. This algorithm was implemented and analyzed on two supercomputer architectures with an application of simulating the aggregation of 250 platelets, or 102 million particles. To scale on these computers with complex memory and network architectures with GPUs, we devised a biomechanics-informed task mapping scheme to optimize load imbalance, communications, and memory utilization. Our simulations, scaling well up to 192 nodes on a Summit-like supercomputer with a peak speed of 11 petaflops, achieved a rate of 423 μs/day which is 500 times faster than the conventional algorithm using static time step and this has enabled studies of record size blood clots at record spatial–temporal resolutions. Additionally, we discovered the sensitive dependence of the scalability and execution time on the methods of decomposition, CPU–GPU coupling, and task mapping.

97 MATHEMATICS AND COMPUTING↗

Development of the Tensoral Computer Language

The research scientist or engineer wishing to perform large scale simulations or to extract useful information from existing databases is required to have expertise in the details of the particular database, the numerical methods and the computer architecture to be used. This poses a significant practical barrier to the use of simulation data. The goal of this research was to develop a high-level computer language called Tensoral, designed to remove this barrier. The Tensoral language provides a framework in which efficient generic data manipulations can be easily coded and implemented. First of all, Tensoral is general. The fundamental objects in Tensoral represent tensor fields and the operators that act on them. The numerical implementation of these tensors and operators is completely and flexibly programmable. New mathematical constructs and operators can be easily added to the Tensoral system. Tensoral is compatible with existing languages. Tensoral tensor operations co-exist in a natural way with a host language, which may be any sufficiently powerful computer language such as Fortran, C, or Vectoral. Tensoral is very-high-level. Tensor operations in Tensoral typically act on entire databases (i.e., arrays) at one time and may, therefore, correspond to many lines of code in a conventional language. Tensoral is efficient. Tensoral is a compiled language. Database manipulations are simplified optimized and scheduled by the compiler eventually resulting in efficient machine code to implement them.

Ferziger, Joel↗

Utilizing a Common Model Architecture for the Simulation of On-Orbit Human Spaceflight Operations

Advances in computing and simulation technology during the 1980 s promoted the development of multiple simulations for on-orbit human spaceflight operations at the National Aeronautics and Space Administration (NASA) Johnson Space Center (JSC). By the late 1980 s, it became increasingly clear that redundancy risked leading to inefficiencies in dissemination and incorporation of models across simulations, institutional organizations, and spaceflight programs. Software development tools, such as the Trick Simulation Environment, began to emerge providing a generalized framework for the development, integration, and operation of simulations. Throughout the 1990s, engineering, operations, and training simulations began to migrate to Trick in order to take advantage of a common simulation environment. Due to forecasted reductions in funding to the Shuttle and International Space Station (ISS) programs in the early 2000s, the common framework that arose from this simulation migration was formalized as the Common Model Architecture (CMA). The CMA was established in order to further reduce redundancy, while promoting model sharing and enhancing model integrity. A minimal set of standards for data flows, functional interactions, and model organization based on Trick were defined for those features found to be most important in facilitating model interaction. Model-providing organizations within NASA JSC took ownership of models for their specific areas of expertise, and agreements were reached on common model sharing and usage. This paper describes the evolution of commonality enabled by the Trick Simulation Environment at JSC leading to the CMA, the standards defined by the CMA, a recent CMA simulation example, and potential application of the CMA to NASA s new Exploration program.

Leslie J Quiocho↗

Some Problems and Solutions in Transferring Ecosystem Simulation Codes to Supercomputers

Many computer codes for the simulation of ecological systems have been developed in the last twenty-five years. This development took place initially on main-frame computers, then mini-computers, and more recently, on micro-computers and workstations. Recent recognition of ecosystem science as a High Performance Computing and Communications Program Grand Challenge area emphasizes supercomputers (both parallel and distributed systems) as the next set of tools for ecological simulation. Transferring ecosystem simulation codes to such systems is not a matter of simply compiling and executing existing code on the supercomputer since there are significant differences in the system architectures of sequential, scalar computers and parallel and/or vector supercomputers. To more appropriately match the application to the architecture (necessary to achieve reasonable performance), the parallelism (if it exists) of the original application must be exploited. We discuss our work in transferring a general grassland simulation model (developed on a VAX in the FORTRAN computer programming language) to a Cray Y-MP. We show the Cray shared-memory vector-architecture, and discuss our rationale for selecting the Cray. We describe porting the model to the Cray and executing and verifying a baseline version, and we discuss the changes we made to exploit the parallelism in the application and to improve code execution. As a result, the Cray executed the model 30 times faster than the VAX 11/785 and 10 times faster than a Sun 4 workstation. We achieved an additional speed-up of approximately 30 percent over the original Cray run by using the compiler's vectorizing capabilities and the machine's ability to put subroutines and functions "in-line" in the code. With the modifications, the code still runs at only about 5% of the Cray's peak speed because it makes ineffective use of the vector processing capabilities of the Cray. We conclude with a discussion and future plans.

Skiles, J. W.↗

Multi-level Hierarchical Poly Tree computer architectures

Based on the concept of hierarchical substructuring, this paper develops an optimal multi-level Hierarchical Poly Tree (HPT) parallel computer architecture scheme which is applicable to the solution of finite element and difference simulations. Emphasis is given to minimizing computational effort, in-core/out-of-core memory requirements, and the data transfer between processors. In addition, a simplified communications network that reduces the number of I/O channels between processors is presented. HPT configurations that yield optimal superlinearities are also demonstrated. Moreover, to generalize the scope of applicability, special attention is given to developing: (1) multi-level reduction trees which provide an orderly/optimal procedure by which model densification/simplification can be achieved, as well as (2) methodologies enabling processor grading that yields architectures with varying types of multi-level granularity.

Padovan, Joe↗

GPU-friendly surface model for Monte-Carlo detector simulations

The demands for Monte-Carlo simulation are drastically increasing with the Large Hadron Collider’s high-luminosity upgrade, and are expected to exceed the currently available compute resources. At the same time, modern high-performance computing has adopted powerful hardware accelerators, particularly GPUs. The AdePT and Celeritas projects aim to address the demanding computational needs by leveraging these heterogeneous computing architectures. While both have successfully ported realistic detector simulations to GPUs using the VecGeom library, the complexity of geometry modeling emerged as a bottleneck. Thread divergence and high register usage were degrading the GPU performance. Therefore, a new, GPU-friendly surface-based model has been introduced in the VecGeom library that decomposes the divergent code of the 3D primitive solids into simpler and more balanced surface algorithms. In this work, we present the latest developments, focusing on the additions required to efficiently model complex setups like the CMS Phase-2 geometry. This includes memory reduction techniques, and adding accelerating structures for faster traversal.

Diederichs, Severin [CERN]↗

Advances in 3D transient plasma dynamics and control through MHD and hybrid fluid-kinetic simulations with JOREK

Transient phenomena and their control are of high relevance in magnetic confinement fusion plasmas to guarantee a stable and safe plasma operation. Interpretative simulations can maximize the insights gained from experiments on present machines and predictive simulations can help in the preparation of design, mitigation techniques and operational scenarios for future devices. In this article, we provide an overview of recent advances and novel scientific results obtained with the 3D non-linear hybrid fluid-kinetic code JOREK, covering physics of plasma transients from the core to the scrape-off layer (SOL) both for tokamak and stellarator devices. Substantial progress was made in the physics understanding, model validation with experiments and experiment interpretation, thus, giving confidence for predictions to devices like DTT, ITER and DEMO. The topics addressed comprise a wide range: the edge physics of new operation scenarios and edge localized mode suppression; major disruptions with a focus on runaway electrons and vertical displacement events as well as disruption mitigation by shattered pellet injection; the physics mechanisms and operational limits of the flux pumping regime for sawtooth control; MHD limits of stellarators and work towards incorporating advanced edge/SOL/exhaust dynamics; continuing improvements of the code for more efficient hybrid simulations on conventional and accelerated high performance computing architectures.

disruptions↗

Simulation of complex three-dimensional flows

The concept of splitting is used extensively to simulate complex three dimensional flows on modern computer architectures. Used in all aspects, from initial grid generation to the determination of the final converged solution, splitting is used to enhance code vectorization, to permit solution driven grid adaption and grid enrichment, to permit the use of concurrent processing, and to enhance data flow through hierarchal memory systems. Three examples are used to illustrate these concepts to complex three dimensional flow fields: (1) interactive flow over a bump; (2) supersonic flow past a blunt based conical afterbody at incidence to a free stream and containing a centered propulsive jet; and (3) supersonic flow past a sharp leading edge delta wing at incidence to the free stream.

Diewert, G. S.↗

NASA Workshop on Computational Structural Mechanics 1987, part 2

Advanced methods and testbed/simulator development topics are discussed. Computational Structural Mechanics (CSM) testbed architecture, engine structures simulation, applications to laminate structures, and a generic element processor are among the topics covered.

Sykes, Nancy P.↗

Investigating Local Buckling and Plastic Collapse in Multifunctional Sandwich Composite Cores using a Unit Cell Method

The primary goal of developing analytical models of composite sandwich cores has been to obtain simple expressions for elastic and failure response to improve the rapid analysis of composite structures. An earlier investigation by the authors was made to verify the accuracy of widely-used analytical models that calculate equivalent elastic moduli of a uniform homogeneous single layer of material in place of the complex honeycomb core properties in sandwich composites. This transformation of elastic properties permits a coarsening of the finite element model (FEM) of the core to a fraction of the computational cost in simulating the original complex core architecture. In addition, the core was assumed constructed of multi-layered isotropic walls to support potential additional functionality beyond simple load-carrying capability.

Erik Saether↗

xSDK: Building an ecosystem of highly efficient math libraries for exascale

Current efforts to build increasingly powerful computer architectures are opening up new avenues for more complex and higher fidelity simulations coupled with data analytics and learning, leading to new scientific insights and deeper understanding. At one extreme, exascale computers will be much faster than previous computer generations (performing 10 18 operations per second—that is, 1,000 times faster than petascale). To achieve these performance improvements, computer architectures are becoming increasingly complex, with deep memory hierarchies, very high node and core counts, and heterogeneous features such as graphics processing units (GPUs). Such architectural changes impact the full breadth of computing scales, as heterogeneity pervades even current-generation laptops, workstations, and moderate-sized clusters. While emerging advanced architectures provide unprecedented opportunities, they also present significant challenges for developers of scientific applications, such as multiphysics and multiscale codes, who must adapt their software to handle disruptive changes in architectures and new programming models that have not yet stabilized. Developers must consider increasing concurrency while reducing communication and synchronization, and other complexities such as the potential for using mixed precision to leverage the compute power available in low-precision tensor cores. On one hand, developers must implement new scientific capabilities, which in turn increase code complexity. On the other hand, the codes must be ported to new architectures, requiring the inclusion of new programming models and the restructuring of code to achieve good performance. Addressing these issues is beyond the capability of any single person or team—leading to the need for collaboration among many teams, who encapsulate their expertise in reusable software and work together to create sustainable software ecosystems.

97 MATHEMATICS AND COMPUTING↗