Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Development and Integration of a Thermal Management Simulation for a Quadrotor Parallel Hybrid Propulsion System

This paper details the development of a propulsion system simulation for a six-passenger parallel hybrid quadrotor and utilizes the Numerical Propulsion System Simulation (NPSS) along with the NPSS Power System Library as the development environment. This simulation integrates an engine power plant with an electrical generation and distribution system and includes the required thermal management system. The thermal management system is comprised of liquid cooling loops that reject the heat load through air to coolant heat exchangers and utilizes a map-based performance estimation method. This method is developed within NPSS and detailed in this paper. The full system model is designed to predict system weight, range, and performance through a proposed mission profile. Results of the paper show an all-engine system maintains the best range, while a mostly electric system that utilizes an engine as a backup or a conditional power contributor offers range benefit.

Vertical lift and take off vehicle↗

Development and Integration of a Thermal Management Simulation for a Quadrotor Parallel Hybrid Propulsion System

This paper details the development of a propulsion system simulation for a six-passenger parallel hybrid quadrotor and utilizes the Numerical Propulsion System Simulation (NPSS) along with the NPSS Power System Library as the development environment. This simulation integrates an engine power plant with an electrical generation and distribution system and includes the required thermal management system. The thermal management system is comprised of liquid cooling loops that reject the heat load through air to coolant heat exchangers and utilizes a map-based performance estimation method. This method is developed within NPSS and detailed in this paper. The full system model is designed to predict system weight, range, and performance through a proposed mission profile. Results of the paper show an all-engine system maintains the best range, while a mostly electric system that utilizes an engine as a backup or a conditional power contributor offers range benefit.

Vertical lift and take off vehicle↗

Influences of Local Sea-Surface Temperatures and Large-scale Dynamics on Monthly Precipitation Inferred from Two 10-year GCM-Simulations

Two parallel sets of 10-year long: January 1, 1982 to December 31, 1991, simulations were made with the finite volume General Circulation Model (fvGCM) in which the model integrations were forced with prescribed sea-surface temperature fields (SSTs) available as two separate SST-datasets. One dataset contained naturally varying monthly SSTs for the chosen period, and the oth& had the 12-monthly mean SSTs for the same period. Plots of evaporation, precipitation, and atmosphere-column moisture convergence, binned by l C SST intervals show that except for the tropics, the precipitation is more strongly constrained by large-scale dynamics as opposed to local SST. Binning data by SST naturally provided an ensemble average of data contributed from disparate locations with same SST; such averages could be expected to mitigate all location related influences. However, the plots revealed: i) evaporation, vertical velocity, and precipitation are very robust and remarkably similar for each of the two simulations and even for the data from 1987-ENSO-year simulation; ii) while the evaporation increased monotonically with SST up to about 27 C, the precipitation did not; iii) precipitation correlated much better with the column vertical velocity as opposed to SST suggesting that the influence of dynamical circulation including non-local SSTs is stronger than local-SSTs. The precipitation fields were doubly binned with respect to SST and boundary-layer mass and/or moisture convergence. The analysis discerned the rate of change of precipitation with local SST as a sum of partial derivative of precipitation with local SST plus partial derivative of precipitation with boundary layer moisture convergence multiplied by the rate of change of boundary-layer moisture convergence with SST (see Eqn. 3 of Section 4.5). This analysis is mathematically rigorous as well as provides a quantitative measure of the influence of local SST on the local precipitation. The results were recast to examine the dependence of local rainfall on local SSTs; it was discernible only in the tropics. Our methodology can be used for computing relationship between any forcing function and its effect(s) on a chosen field.

Sud, Y. C.↗

OpenSn: A massively parallel, open-source simulation environment for discrete ordinates radiation transport

OpenSn is an open-source, massively parallel deterministic radiation transport code for solving the discrete-ordinates ( S N ) form of the Boltzmann transport equation on unstructured, arbitrary polyhedral meshes. It supports high-fidelity simulations involving steady-state, eigenvalue, and adjoint problems for neutral particles (e.g., neutrons, photons, multi-particles), using the multigroup approximation in energy. OpenSn combines angular discretization via discrete ordinates with a discontinuous Galerkin finite element method (DGFEM) in space, enabling accurate resolution of transport physics on arbitrary polyhedral cells, included locally refined spatial grids. It includes multiple angular quadrature types, including locally refined angular quadratures. Written in modern C++ with a Python API, OpenSn runs efficiently on platforms ranging from laptops to supercomputers. The transport sweep algorithm is implemented using a task-based, directed-acyclic-graph (DAG) approach for each angle and supports asynchronous parallelism across thousands of MPI ranks. Group-set aggregation improves compute intensity, and synthetic acceleration techniques (e.g., diffusion synthetic acceleration, second-moment method) enhance solver convergence. OpenSn has been verified on reactor physics problems and demonstrated excellent weak and strong scaling performance on more than 32,768 processes, making it a versatile and robust platform for large-scale transport simulations in complex geometries.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Use of networked workstations for parallel nonlinear structural dynamic simulations of rotating bladed-disk assemblies

The principal objective of this research is to investigate, develop and demonstrate coarse-grained, parallel-processing strategies for nonlinear dynamic simulations for rotating bladed-disk assemblies. The parallel -processing strategies addressed include numerical algorithms for parallel nonlinear solutions and techniques to effect load balancing among processors. The parallel environment employed is a distributed-memory, coarse-grained one consisting of networked workstations. A parallel explicit time integration method has been implemented for transient nonlinear solutions of rotationg bladed-disk assemblies. Automatic domain partitioning techniques have been investigated for load balancing among processors. Advanced computing environments, data structures and interactive computer graphics all contribute to an integrated parallel finite element analysis system to facilitate more efficient and powerful dynamic simulations.

Hsieh, Shang-Hsien↗

Direct numerical simulation of instabilities in parallel flow with spherical roughness elements

Results from a direct numerical simulation of laminar flow over a flat surface with spherical roughness elements using a spectral-element method are given. The numerical simulation approximates roughness as a cellular pattern of identical spheres protruding from a smooth wall. Periodic boundary conditions on the domain's horizontal faces simulate an infinite array of roughness elements extending in the streamwise and spanwise directions, which implies the parallel-flow assumption, and results in a closed domain. A body force, designed to yield the horizontal Blasius velocity in the absence of roughness, sustains the flow. Instabilities above a critical Reynolds number reveal negligible oscillations in the recirculation regions behind each sphere and in the free stream, high-amplitude oscillations in the layer directly above the spheres, and a mean profile with an inflection point near the sphere's crest. The inflection point yields an unstable layer above the roughness (where U''(y) is less than 0) and a stable region within the roughness (where U''(y) is greater than 0). Evidently, the instability begins when the low-momentum or wake region behind an element, being the region most affected by disturbances (purely numerical in this case), goes unstable and moves. In compressible flow with periodic boundaries, this motion sends disturbances to all regions of the domain. In the unstable layer just above the inflection point, the disturbances grow while being carried downstream with a propagation speed equal to the local mean velocity; they do not grow amid the low energy region near the roughness patch. The most amplified disturbance eventually arrives at the next roughness element downstream, perturbing its wake and inducing a global response at a frequency governed by the streamwise spacing between spheres and the mean velocity of the most amplified layer.

Deanna, R. G.↗

SUPREM-DSMC: A New Scalable, Parallel, Reacting, Multidimensional Direct Simulation Monte Carlo Flow Code

An AFRL/NRL team has recently been selected to develop a scalable, parallel, reacting, multidimensional (SUPREM) Direct Simulation Monte Carlo (DSMC) code for the DoD user community under the High Performance Computing Modernization Office (HPCMO) Common High Performance Computing Software Support Initiative (CHSSI). This paper will introduce the JANNAF Exhaust Plume community to this three-year development effort and present the overall goals, schedule, and current status of this new code.

Campbell, David↗

Parallel Implementation of Nonadditive Gaussian Process Potentials for Monte Carlo Simulations

A strategy is presented to implement Gaussian process potentials in molecular simulations through parallel programming. Attention is focused on the three-body nonadditive energy, though all algorithms extend straightforwardly to the additive energy. The method to distribute pairs and triplets between processes is general to all potentials. Results are presented for a simulation box of argon, including full box and atom displacement calculations, which are relevant to Monte Carlo simulation. Data on speed-up are presented for up to 120 processes across four nodes. A 4-fold speed-up is observed over five processes, extending to 20-fold over 40 processes and 30-fold over 120 processes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Partitioning and packing mathematical simulation models for calculation on parallel computers

The development of multiprocessor simulations from a serial set of ordinary differential equations describing a physical system is described. Degrees of parallelism (i.e., coupling between the equations) and their impact on parallel processing are discussed. The problem of identifying computational parallelism within sets of closely coupled equations that require the exchange of current values of variables is described. A technique is presented for identifying this parallelism and for partitioning the equations for parallel solution on a multiprocessor. An algorithm which packs the equations into a minimum number of processors is also described. The results of the packing algorithm when applied to a turbojet engine model are presented in terms of processor utilization.

Arpasi, D. J.↗

Parallelization of Program to Optimize Simulated Trajectories (POST3D)

This paper describes the parallelization of the Program to Optimize Simulated Trajectories (POST3D). POST3D uses a gradient-based optimization algorithm that reaches an optimum design point by moving from one design point to the next. The gradient calculations required to complete the optimization process, dominate the computational time and have been parallelized using a Single Program Multiple Data (SPMD) on a distributed memory NUMA (non-uniform memory access) architecture. The Origin2000 was used for the tests presented.

Hammond, Dana P.↗

Regional-scale fault-to-structure earthquake simulations with the EQSIM framework: Workflow maturation and computational performance on GPU-accelerated exascale platforms

Continuous advancements in scientific and engineering understanding of earthquake phenomena, combined with the associated development of representative physics-based models, is providing a foundation for high-performance, fault-to-structure earthquake simulations. However, regional-scale applications of high-performance models have been challenged by the computational requirements at the resolutions required for engineering risk assessments. The EarthQuake SIMulation (EQSIM) framework, a software application development under the US Department of Energy (DOE) Exascale Computing Project, is focused on overcoming the existing computational barriers and enabling routine regional-scale simulations at resolutions relevant to a breadth of engineered systems. This multidisciplinary software development—drawing upon expertise in geophysics, engineering, applied math and computer science—is preparing the advanced computational workflow necessary to fully exploit the DOE’s exaflop computer platforms coming online in the 2023 to 2024 timeframe. Achievement of the computational performance required for high-resolution regional models containing upward of hundreds of billions to trillions of model grid points requires numerical efficiency in every phase of a regional simulation. This includes run time start-up and regional model generation, effective distribution of the computational workload across thousands of computer nodes, efficient coupling of regional geophysics and local engineering models, and application-tailored highly efficient transfer, storage, and interrogation of very large volumes of simulation data. This article summarizes the most recent advancements and refinements incorporated in the workflow design for the EQSIM integrated fault-to-structure framework, which are based on extensive numerical testing across multiple graphics processing unit (GPU)-accelerated platforms, and demonstrates the computational performance achieved on the world’s first exaflop computer platform through representative regional-scale earthquake simulations for the San Francisco Bay Area in California, USA.

58 GEOSCIENCES↗

National Campaign (NC)-1 Strategic Conflict Management Simulation (X4) Community Based Rules

Projected demand for transportation services in the urban environment has led to the development of several Concepts of Operation for Urban Air Mobility, or UAM. UAM is a concept for the transportation of people and goods in the metropolitan environment using small, efficient aircraft over short distances as part of an expanding multimodal transportation network. UAM will leverage emerging technologies including electric Vertical Takeoff and Landing (eVTOL) aircraft, increasing levels of automation and a new operational paradigm in dense airspace where a set of agreed-upon rules govern the procedures and interactions defining a cooperative environment in which operators are entrusted with a range of functions typically conducted by Air Traffic Control (ATC). These rules, proposed in the FAA NextGen Office’s UAM Concept of Operations [1], were originally termed Community Based Rules or Community Business Rules (CBRs), and will in the future termed Cooperative Operating Practices (COPs); this document uses the original term, CBR. CBRs are a set of rules, developed by the UAM community and (where necessary) approved by the FAA that govern the interactions between UAM entities and limit the need for ATC services including, but not limited to, separation control by ATC, addressing a fundamental challenge to scaling UAM operations. UAM community development of CBRs is anticipated to accelerate the adoption of new practices while retaining the regulatory authority of the FAA within required domains (e.g., NAS safety, security and equal access). However, there currently exists no agreed industry forum or defined procedures for CBR development. Investigation of best practices for the development of UAM CBRs was identified by NASA and the FAA NextGen Office as a research need. In collaboration with seven industry partners, NASA participated in a series of simulations that investigated elements of the envisioned UAM operations, with a primary focus on Strategic Conflict Management (SCM). The development and conduct of cooperative UAM simulations with seven industry partners provided a unique opportunity to investigate CBR development practices. Development of CBRs for the UAM SCM simulations was conducted in parallel with simulation capability development and was closely related to requirements definition for the simulations. As such, the CBR development effort presented herein had two objectives: explore CBR development practices in collaboration with the industry partners and develop an initial set of UAM CBRs to support simulation requirements definition and development. Consensus was achieved among NASA and the industry partners on 24 CBRs that were developed to support the cooperative simulation operations across five topic areas: General (related to test requirements), Operational Intent, Conformance Monitoring, Demand Capacity Balancing, and Airspace Constraint Management. Additional topic areas and CBRs were discussed but were deemed outside the scope of the simulation; these are included in the appendices. A collaborative, iterative process was employed for developing the CBRs engaging both NASA and Industry; because CBR development is envisioned to be community-driven, opportunities were sought that provided industry partners leadership roles in developing CBRs. The following key observations and recommendations may aid the UAM industry in future CBR development efforts: - The lack of a defined process proved challenging initially. Stakeholder engagement in the early stages of CBR development was intermittent and may have been due to the lack of a clear definition of roles and responsibilities of those involved in the effort. - Industry leadership of CBR topic areas proved successful. Discussions in these topic areas were engaging, with alternate viewpoints freely discussed and detailed CBRs resulting. This points to the importance of identifying the best-suited leadership in technical areas for CBR development. - Discussions within a CBR topic area were typically dominated by only a few participants. Whereas all industry partners contributed to CBR development, within each topic area, technical leadership was evident even when not formally established. This observation may indicate that smaller, focused groups may be more effective in initial CBR development than an open forum or large standards development effort (although both maybe required prior to FAA review and approval for some CBRs). - Identifying suitable forums for initial UAM CBR development and identifying the most effective industry participants and leadership will be crucial for successful CBR development. Although the operational need for UAM CBRs may not be immediate, establishing the forums and leadership to define the processes for CBR development is a prudent early step to UAM realization.

Community Based Rules↗

Parallel-in-time quantum simulation via Page and Wootters quantum time

In the past few decades, researchers have created a veritable zoo of quantum algorithms by drawing inspiration from classical computing, information theory, and even from physical phenomena. Here, we present quantum algorithms for parallel-in-time simulations that are inspired by the Page and Wootters formalism. In this framework, and thus in our algorithms, the classical time variable of quantum mechanics is promoted to the quantum realm by introducing a Hilbert space of “clock” qubits that are then entangled with the “system” qubits. We show that our algorithms can compute temporal properties over 𝑁 different times of many-body systems by only using log⁡(𝑁) clock qubits. As such, we achieve an exponential trade-off between time and spatial complexities. In addition, we rigorously prove that the entanglement created between the system qubits and the clock qubits has operational meaning, as it encodes valuable information about the system’s dynamics. We also provide a circuit depth estimation of all the protocols, showing a running time advantage in computation times over traditional sequential-in-time algorithms. In particular, for the case when the dynamics are determined by the Aubry-Andre model, we present a hybrid method for which our algorithms have a depth that only scales as 𝒪⁡(log⁡(𝑁)⁢𝑛). As a by-product, we can relate the previous schemes to the problem of equilibration of an isolated quantum system, thus indicating that our framework enables a new dimension for studying dynamical properties of many-body systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Program For Parallel Discrete-Event Simulation

User does not have to add any special logic to aid in synchronization. Time Warp Operating System (TWOS) computer program is special-purpose operating system designed to support parallel discrete-event simulation. Complete implementation of Time Warp mechanism. Supports only simulations and other computations designed for virtual time. Time Warp Simulator (TWSIM) subdirectory contains sequential simulation engine interface-compatible with TWOS. TWOS and TWSIM written in, and support simulations in, C programming language.

Beckman, Brian C.↗

A parallel algorithm for switch-level timing simulation on a hypercube multiprocessor

The parallel approach to speeding up simulation is studied, specifically the simulation of digital LSI MOS circuitry on the Intel iPSC/2 hypercube. The simulation algorithm is based on RSIM, an event driven switch-level simulator that incorporates a linear transistor model for simulating digital MOS circuits. Parallel processing techniques based on the concepts of Virtual Time and rollback are utilized so that portions of the circuit may be simulated on separate processors, in parallel for as large an increase in speed as possible. A partitioning algorithm is also developed in order to subdivide the circuit for parallel processing.

Rao, Hariprasad Nannapaneni↗

Parallel Signal Processing and System Simulation using aCe

Recently, networked and cluster computation have become very popular for both signal processing and system simulation. A new language is ideally suited for parallel signal processing applications and system simulation since it allows the programmer to explicitly express the computations that can be performed concurrently. In addition, the new C based parallel language (ace C) for architecture-adaptive programming allows programmers to implement algorithms and system simulation applications on parallel architectures by providing them with the assurance that future parallel architectures will be able to run their applications with a minimum of modification. In this paper, we will focus on some fundamental features of ace C and present a signal processing application (FFT).

Dorband, John E.↗

A direct-execution parallel architecture for the Advanced Continuous Simulation Language (ACSL)

A direct-execution parallel architecture for the Advanced Continuous Simulation Language (ACSL) is presented which overcomes the traditional disadvantages of simulations executed on a digital computer. The incorporation of parallel processing allows the mapping of simulations into a digital computer to be done in the same inherently parallel manner as they are currently mapped onto an analog computer. The direct-execution format maximizes the efficiency of the executed code since the need for a high level language compiler is eliminated. Resolution is greatly increased over that which is available with an analog computer without the sacrifice in execution speed normally expected with digitial computer simulations. Although this report covers all aspects of the new architecture, key emphasis is placed on the processing element configuration and the microprogramming of the ACLS constructs. The execution times for all ACLS constructs are computed using a model of a processing element based on the AMD 29000 CPU and the AMD 29027 FPU. The increase in execution speed provided by parallel processing is exemplified by comparing the derived execution times of two ACSL programs with the execution times for the same programs executed on a similar sequential architecture.

Carroll, Chester C.↗

Massively parallel computing for the simulation of unsteady flows in turbomachinery

This paper deals with evaluating the capabilities of the massively parallel Connection Machine CM2 in predicting unsteady flows in turbomachines. The implementation on the CM2 of an implicit, time-accurate, zonal algorithm for the Navier-Stokes equations in two dimensions is described. Programming issues and modifications made to the original sequential algorithm to improve performance on the CM2 are briefly discussed. Performance is compared to a functionally equivalent code for the Cray YMP.

Madavan, Nateri K.↗