Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

NASA Workshop on Computational Structural Mechanics 1987, part 2

Advanced methods and testbed/simulator development topics are discussed. Computational Structural Mechanics (CSM) testbed architecture, engine structures simulation, applications to laminate structures, and a generic element processor are among the topics covered.

Sykes, Nancy P.↗

Investigating Local Buckling and Plastic Collapse in Multifunctional Sandwich Composite Cores using a Unit Cell Method

The primary goal of developing analytical models of composite sandwich cores has been to obtain simple expressions for elastic and failure response to improve the rapid analysis of composite structures. An earlier investigation by the authors was made to verify the accuracy of widely-used analytical models that calculate equivalent elastic moduli of a uniform homogeneous single layer of material in place of the complex honeycomb core properties in sandwich composites. This transformation of elastic properties permits a coarsening of the finite element model (FEM) of the core to a fraction of the computational cost in simulating the original complex core architecture. In addition, the core was assumed constructed of multi-layered isotropic walls to support potential additional functionality beyond simple load-carrying capability.

Erik Saether↗

xSDK: Building an ecosystem of highly efficient math libraries for exascale

Current efforts to build increasingly powerful computer architectures are opening up new avenues for more complex and higher fidelity simulations coupled with data analytics and learning, leading to new scientific insights and deeper understanding. At one extreme, exascale computers will be much faster than previous computer generations (performing 10 18 operations per second—that is, 1,000 times faster than petascale). To achieve these performance improvements, computer architectures are becoming increasingly complex, with deep memory hierarchies, very high node and core counts, and heterogeneous features such as graphics processing units (GPUs). Such architectural changes impact the full breadth of computing scales, as heterogeneity pervades even current-generation laptops, workstations, and moderate-sized clusters. While emerging advanced architectures provide unprecedented opportunities, they also present significant challenges for developers of scientific applications, such as multiphysics and multiscale codes, who must adapt their software to handle disruptive changes in architectures and new programming models that have not yet stabilized. Developers must consider increasing concurrency while reducing communication and synchronization, and other complexities such as the potential for using mixed precision to leverage the compute power available in low-precision tensor cores. On one hand, developers must implement new scientific capabilities, which in turn increase code complexity. On the other hand, the codes must be ported to new architectures, requiring the inclusion of new programming models and the restructuring of code to achieve good performance. Addressing these issues is beyond the capability of any single person or team—leading to the need for collaboration among many teams, who encapsulate their expertise in reusable software and work together to create sustainable software ecosystems.

97 MATHEMATICS AND COMPUTING↗

Efficient Optimization of Energy Recovery From Geothermal Reservoirs With Recurrent Neural Network Predictive Models

Improving the long-term energy production performance of geothermal reservoirs can be accomplished by optimizing field development and management plans. Reliable prediction models, however, are needed to evaluate and optimize the performance of the underlying reservoirs under various operation and development strategies. In traditional frameworks, physics-based simulation models are used to predict the energy production performance of geothermal reservoirs. However, detailed simulation models are not trivial to construct, require a reliable description of the reservoir conditions and properties, and entail high computational complexity. Data-driven predictive models can offer an efficient alternative for use in optimization workflows. This paper presents an optimization framework for net power generation in geothermal reservoirs using a variant of the recurrent neural network (RNN) as a data-driven predictive model. The RNN architecture is developed and trained to replace the simulation model for computationally efficient prediction of the objective function and its gradients with respect to the well control variables. The net power generation performance of the field is optimized by automatically adjusting the mass flow rate of production and injection wells over 12 years, using a gradient-based local search algorithm. Two field-scale examples are presented to investigate the performance of the developed data-driven prediction and optimization framework. Furthermore, the prediction and optimization results from the RNN model are evaluated through comparison with the results obtained by using a numerical simulation model of a real geothermal reservoir.

15 GEOTHERMAL ENERGY↗

Mapping a battlefield simulation onto message-passing parallel architectures

Perhaps the most critical problem in distributed simulation is that of mapping: without an effective mapping of workload to processors the speedup potential of parallel processing cannot be realized. Mapping a simulation onto a message-passing architecture is especially difficult when the computational workload dynamically changes as a function of time and space; this is exactly the situation faced by battlefield simulations. This paper studies an approach where the simulated battlefield domain is first partitioned into many regions of equal size; typically there are more regions than processors. The regions are then assigned to processors; a processor is responsible for performing all simulation activity associated with the regions. The assignment algorithm is quite simple and attempts to balance load by exploiting locality of workload intensity. The performance of this technique is studied on a simple battlefield simulation implemented on the Flex/32 multiprocessor. Measurements show that the proposed method achieves reasonable processor efficiencies. Furthermore, the method shows promise for use in dynamic remapping of the simulation.

Nicol, David M.↗

Direct Energy Conversion for Nuclear Propulsion at Low Specific Mass

The project will continue the FY13 JSC IR&D (October-2012 to September-2013) effort in Travelling Wave Direct Energy Conversion (TWDEC) in order to demonstrate its potential as the core of a high potential, game-changing, in-space propulsion technology. The TWDEC concept converts particle beam energy into radio frequency (RF) alternating current electrical power, such as can be used to heat the propellant in a plasma thruster. In a more advanced concept (explored in the Phase 1 NIAC project), the TWDEC could also be utilized to condition the particle beam such that it may transfer directed kinetic energy to a target propellant plasma for the purpose of increasing thrust and optimizing the specific impulse. The overall scope of the FY13 first-year effort was to build on both the 2012 Phase 1 NIAC research and the analysis and test results produced by Japanese researchers over the past twenty years to assess the potential for spacecraft propulsion applications. The primary objective of the FY13 effort was to create particle-in-cell computer simulations of a TWDEC. Other objectives included construction of a breadboard TWDEC test article, preliminary test calibration of the simulations, and construction of first order power system models to feed into mission architecture analyses with COPERNICUS tools. Due to funding cuts resulting from the FY13 sequestration, only the computer simulations and assembly of the breadboard test article were completed. The simulations, however, are of unprecedented flexibility and precision and were presented at the 2013 AIAA Joint Propulsion Conference. Also, the assembled test article will provide an ion current density two orders of magnitude above that available in previous Japanese experiments, thus enabling the first direct measurements of power generation from a TWDEC for FY14. The proposed FY14 effort will use the test article for experimental validation of the computer simulations and thus complete to a greater fidelity the mission analysis products originally conceived for FY13.

Scott, John H.↗

ERAS: Enabling the Integration of Real-World Intellectual Properties (IPs) in Architectural Simulators

Sandia National Laboratories is investigating scalable architectural simulation capabilities with a focus on simulating and evaluating highly scalable supercomputers for high performance computing applications. There is a growing demand for RTL model integration to provide the capability to simulate customized node architectures and heterogeneous systems. This report describes the first steps integrating the ESSENTial Signal Simulation Enabled by Netlist Transforms (ESSENT) tool with the Structural Simulation Toolkit (SST). ESSENT can emit C++ models from models written in FIRRTL to automatically generate components. The integration workflow will automatically generate the SST component and necessary interfaces to ’plug’ the ESSENT model into the SST framework.

97 MATHEMATICS AND COMPUTING↗

Berkeley eXtensible Environment (BXE) v3

The Berkeley eXtensible Environment (BXE) provides a cloud environment for hardware designers and computer architects to design, build, and simulate their custom architectures on an on-premises FPGA cluster. Utilizing the Chipyard, MoSAIC, and FireSim frameworks, users are provided an environment where they can assemble SoC designs from an existing library of components or import their own source code. Once their designs are ready, they can utilize the FireSim framework provided by BXE to deploy and simulate their designs on the FPGA. Users aren't limited to a single FPGA; they can deploy multiple instances across multiple FPGAs, acting like a rack of servers, or partition their large design across multiple FPGAs, ganging multiple FPGAs into a single simulated system.

Fatollahi-Fard, Farzin↗

Geoid Recovery Using Geophysical Inverse Theory Applied to Satellite to Satellite Tracking Data

This report describes a new method for determination of the geopotential, or the equivalent geoid. It is based on Satellite-to-Satellite Tracking (SST) of two co-orbiting low earth satellites separated by a few hundred kilometers. The analysis is aimed at the GRACE Mission, though it is generally applicable to any SST data. It is proposed that the SST be viewed as a mapping mission. That is, the result will be maps of the geoid or gravity, as contrasted with determination of spherical harmonics or Fourier coefficients. A method has been developed, based on Geophysical Inverse Theory (GIT), that can provide maps at a prescribed (desired) resolution and the corresponding error map from the SST data. This computation can be done area by area avoiding simultaneous recovery of all the geopotential information. The necessary elements of potential theory, celestial mechanics, and Geophysical Inverse Theory are described, a computation architecture is described, and the results of several simulations presented. Centimeter accuracy geoids with 50 to 100 km resolution can be recovered with a 30 to 60 day mission.

Gaposchkin, E. M.↗

Predicting Unreinforced Fabric Mechanical Behavior with Recurrent Neural Networks

Unreinforced woven fabrics are widely employed in various high-performance applications, including parachute deployment systems, airbags, and ballistic armor. The analysis of such materials is inherently complex due to the multiscale structure of these materials, and the dependence of macroscale behavior on changes that occur at lower scales. Previously, NASA’s Multiscale Analysis Tool (NASMAT) showed its capability in predicting unreinforced fabric behavior at the macroscale by capturing finite rotations that occur at the mesoscale. Though effective, the tool can face high computational cost for large, complex problems, motivating the need for the development of a surrogate model that can capture the same behavior. A recurrent neural network (RNN) was developed and trained on virtual NASMAT data to mimic the physics-based solutions while improving the computational runtime. The architecture of the RNN to best simulate the fabric behavior was carefully crafted based on heuristic knowledge of predicting physics-based temporal data, manual hyperparameter case studies, and Hyperband optimization.. The resultant model was able to predict a variety of stress-strain curves for fabrics with different mesoscale geometries, and was further validated by comparing to experimental data for the K706 style Kevlar plain-weave fabric, demonstrating the ability of the model to effectively capture the geometric changes in the fabric without explicitly calculating them, as is done in NASMAT. Furthermore, the tool showed its ability to improve on the runtime by a factor of 10 for fabric solutions compared to the multiscale tool, which would further enable the simulation of complex loading scenarios on unreinforced fabrics.

Fabric↗

Programmable Heisenberg interactions between Floquet qubits

Abstract The trade-off between robustness and tunability is a central challenge in the pursuit of quantum simulation and fault-tolerant quantum computation. In particular, quantum architectures are often designed to achieve high coherence at the expense of tunability. Many current qubit designs have fixed energy levels and consequently limited types of controllable interactions. Here by adiabatically transforming fixed-frequency superconducting circuits into modifiable Floquet qubits, we demonstrate an XXZ Heisenberg interaction with fully adjustable anisotropy. This interaction model can act as the primitive for an expressive set of quantum operations, but is also the basis for quantum simulations of spin systems. To illustrate the robustness and versatility of our Floquet protocol, we tailor the Heisenberg Hamiltonian and implement two-qubit iSWAP, CZ and SWAP gates with good estimated fidelities. In addition, we implement a Heisenberg interaction between higher energy levels and employ it to construct a three-qubit CCZ gate, also with a competitive fidelity. Our protocol applies to multiple fixed-frequency high-coherence platforms, providing a collection of interactions for high-performance quantum information processing. It also establishes the potential of the Floquet framework as a tool for exploring quantum electrodynamics and optimal control.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Computational Controls Workstation: Algorithms and hardware

The Computational Controls Workstation provides an integrated environment for the modeling, simulation, and analysis of Space Station dynamics and control. Using highly efficient computational algorithms combined with a fast parallel processing architecture, the workstation makes real-time simulation of flexible body models of the Space Station possible. A consistent, user-friendly interface and state-of-the-art post-processing options are combined with powerful analysis tools and model databases to provide users with a complete environment for Space Station dynamics and control analysis. The software tools available include a solid modeler, graphical data entry tool, O(n) algorithm-based multi-flexible body simulation, and 2D/3D post-processors. This paper describes the architecture of the workstation while a companion paper describes performance and user perspectives.

Venugopal, R.↗

Porting hypre to heterogeneous computer architectures: Strategies and experiences

We report that linear systems are occurring in many applications, and solving them can take a large amount of the total simulation time. The high performance library hypre provides a variety of interfaces and linear solvers, including various multigrid methods, that have achieved good scalability on a variety of homogeneous parallel computer architectures. Heterogeneous architectures with nodes that have both CPUs and accelerators provide new challenges, since they require more fine-grained parallelism and reduced data movement between different memories on a single node as well as across nodes. We will discuss our experiences and strategies to port hypre to heterogeneous computers with accelerators, including the design of a new memory model, the use of abstractions, the BoxLoop macros in the structured and semi-structured interfaces, and the restructuring of algebraic multigrid (AMG) into modular components. We present numerical experiments comparing CPU and GPU performance for several test problems.

97 MATHEMATICS AND COMPUTING↗

Linear and Nonlinear Solvers for Simulating Multiphase Flow within Large-Scale Engineered Subsurface Systems

Simulation of multiphase flow in the subsurface is well-known to be computationally challenging. While there have been many studies that have explored approaches to overcoming these challenges, they often utilize relatively simple case studies. In this paper, we focus on the unique numerical challenges posed by modeling large-scale engineered subsurface systems, characterized by discrete features embedded in a heterogeneous natural subsurface setting. The man-made features such as shafts, tunnels, and barriers often cause multiple challenges in modeling the domain for multiphase porous media flow. This flow scenario can have a wide range of applications such as nuclear waste repositories, enhanced recovery of a petroleum reservoir, geothermal engineering, and carbon sequestration. An example of these severe numerical challenges is the case of performance assessment (PA) for Waste Isolation Pilot Plant (WIPP), the only operating deep geological repository in the US, which simulates extreme material properties of bedded salt rock formation and extreme contrast due to open excavation next to the formation. The models have extremes not only of permeability and porosity but also of the constitutive models needed for multiphase flow; additionally, they have process models like salt creep closure reducing porosity over time, fracturing in clay and anhydrite interbeds of the bedded salt, gas generation from the waste materials, and unintentional human borehole intrusions in some scenarios. Numerical simulations require the solution of coupled systems of nonlinear PDEs; in our work, we use the open-source simulator PFLOTRAN which is based on Finite Volume discretization. The solution of the nonlinear equations requires use of the Newton-Raphson iteration at each time step, which entails the solution of the linearized Jacobian system at each iteration. The effects of all the processes (i.e., large number of unknowns, highly nonlinear constitutive relations, large contrasts in material properties in short distances) lead to an ill-conditioned Jacobian matrix that severely challenges traditional linear solver, i.e., stabilized biconjugate gradient with block Jacobi incomplete LU preconditioner (BCGS-ILU) leading to non-convergence for traditional Newton-Raphson nonlinear solver causing unacceptably long computation time for each model. This paper presents linear solvers such as constrained pressure residual (CPR) two-stage preconditioner with alternate-block-factorization (ABF) and quasi- implicit pressure and explicit saturation (QIMPES) decouplers and flexible generalized residual solver (FGMRES). The new general-purpose nonlinear solver, Newton trust-region dogleg Cauchy (NTRDC), is also introduced to resolve extreme nonlinearities in the models. We demonstrate the effectiveness of each method relative to the default BCGS-Newton solver. The two best cases had nearly 50 times speed-up and achieved completion of a simulation in 14 hours that never completed due to non-convergence with the default solver. We also investigate the strong scalability of each method and discuss some of the deficiencies found for Block Jacobi preconditioner using parallel domain decomposition, and node packing effects of modern processor architecture.

Preconditioner, Nonlinear, Porous media, Multiphas↗

Mission Simulation Facility: Simulation Support for Autonomy Development

The Mission Simulation Facility (MSF) supports research in autonomy technology for planetary exploration vehicles. Using HLA (High Level Architecture) across distributed computers, the MSF connects users autonomy algorithms with provided or third-party simulations of robotic vehicles and planetary surface environments, including onboard components and scientific instruments. Simulation fidelity is variable to meet changing needs as autonomy technology advances in Technical Readiness Level (TRL). A virtual robot operating in a virtual environment offers numerous advantages over actual hardware, including availability, simplicity, and risk mitigation. The MSF is in use by researchers at NASA Ames Research Center (ARC) and has demonstrated basic functionality. Continuing work will support the needs of a broader user base.

Pisanich, Greg↗

Cluster Computation of Flight Reynolds Number Flows

The performance of a workstation cluster used for the solution of the Reynolds-averaged Navier-Stokes equations is compared with a conventional vector supercomputer architecture. The application simulation of the steady flowfield about a transonic transport was computed using an implicit diagonal scheme in an overset mesh framework. Static load balancing was used, while coarse grain decomposition was achieved by solution of a grid zone per processor. Price/performance ratios are estimated for several scenarios in which such clusters may be utilized.

Atwood, Christopher A.↗

ARES v1.x - Performance Portable Tool to Simulate Supernovae based on Parthenon Framework

Historically, codes for simulating supernovae (such as Arepo, FLASH or LEAFS) have been at the forefront of scientific high-performance computing to the immense computational resources required for full 3D simulations. However, given the shift towards heterogenous HPC architectures, many current-generation codes are at the risk of losing their competitiveness as they are only designed to run on homogeneous CPU-only systems. There exist several efforts to enable these codes for GPU’s, however, these efforts only consider specific architectures or vendors (e.g., implement only CUDA or HIP), limiting themselves to a small range of exascale computing systems. Frameworks such as Kokkos aim to provide a framework which is agnostic of the targeted architecture, enabling the development of performant and portable code. In the Ares code, we develop a performance portable tool to simulate supernovae based on the Parthenon Framework, which in turn uses Kokkos in the background. Here, the Parthenon Framework provides an interface to the underlying mesh-refinement routines, which form the backbone of our code. In addition, we incorporate the already existing Singularity-EOS toolkit to provide us with various equations of state, primarily the Helmholtz equation of state. We also include the JINA Reaclib as a basis for our nuclear network solver. Finally, we implement a gravity solver to complete the required physics. This setup will provide us with a minimal code base to simulate supernova in a similar style to the tried-and-tested Arepo code, but in a futureproof performance portable framework.

Lim, Hyun↗

Ensuring statistical reproducibility of ocean model simulations in the age of hybrid computing

Novel high performance computing systems that feature hybrid architectures require large scale code refactoring to unravel underlying exploitable parallelism. Such redesign can often be accompanied with machine-precision changes as the order of computation cannot always be maintained. For chaotic systems like climate models, these round-off level differences can grow rapidly. Systematic errors may also manifest initially as machine-precision differences. Isolating genuine round off level differences from such errors remains a challenge. Here, we apply two-sample equality of distribution tests to evaluate statistical reproducibility of the ocean model component of US Department of Energy's Energy Exascale Earth System Model (E3SM). A 2-year control simulation ensemble is compared to a modified ensemble as a test case - after a known non-bit-for-bit change in a model component is introduced - to evaluate the null hypothesis that the two ensembles are statistically indistinguishable. To quantify the false negative rates of these tests, we conduct a formal power analysis using a targeted suite of short simulation ensembles. The ensemble suite contains several perturbed ensembles, each with a progressively different climate than the baseline ensemble - obtained by perturbing the magnitude of a single model tuning parameter, the Gent and McWilliams κ, in a controlled manner. The null hypothesis is evaluated for each of perturbed ensembles using these tests. The power analysis informs on the detection limits of the tests for given ensemble size allowing model developers to evaluate the impact of an introduced non-bit-for-bit change to the model.

Mahajan, Salil↗