Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62

CP2K: An Electronic Structure and Molecular Dynamics Software Package - Quickstep: Efficient and Accurate Electronic Structure Calculations

CP2K is an open source electronic structure and molecular dynamics software package to perform atomistic simulations of solid-state, liquid, molecular and biological systems. It is especially aimed at massively-parallel and linear-scaling electronic structure methods and state-of-the-art ab-initio molecular dynamics simulations. Excellent performance for electronic structure calculations is achieved using novel algorithms implemented for modern high-performance computing systems. This review revisits the main capabilities of CP2K to perform efficient and accurate electronic structure simulations. The emphasis is put on density functional theory and multiple post-Hartree-Fock methods using the Gaussian and plane wave approach and its augmented all-electron extension. TDK has received funding from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (grant agreement No. 716142). VRR has been supported by the Swiss National Science Foundation in the form of Ambizione grant No. PZ00P2 174227 and RZK by the Natural Sciences and Engineering Research Council of Canada (NSERC) through Discovery Grants (RGPIN-2016-0505). GKS and CJM are supported by the US Department of Energy, Office of Science, Office of Basic Energy Sciences, Division of Chemical Sciences, Geosciences, and Biosciences. UK based work was funded under the embedded CSE programme of the ARCHER UK National Supercomputing Service (http://www.archer.ac.uk), grants eCSE03-011, eCSE06-6, eCSE08-9, eCSE13-17 and the EPSRC (EP/P022235/1) grant “Surface and Interface Toolkit for the Materials Chemistry Community". Computational resources were provided by the Swiss National Supercomputing Centre (CSCS) and Compute Canada. The generous allocation of computing time on the FPGA-based supercomputer “Noctua" at PC2 is kindly acknowledged.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Simulation of Relativistic Shocks and Associated Radiation from Turbulent Magnetic Fields

Using our new 3-D relativistic particle-in-cell (PIC) code, we investigated long-term particle acceleration associated with a relativistic electron-positron jet propagating in an unmagnetized ambient electron-positron plasma. The simulations were performed using a much longer simulation system than our previous simulations in order to investigate the full nonlinear stage of the Weibel instability and its particle acceleration mechanism. Cold jet electrons are thermalized and ambient electrons are accelerated in the resulting shocks. Acceleration of ambient electrons leads to a maximum ambient electron density three times larger than the original value as predicted by hydrodynamic compression. Behind the bow shock, in the jet shock, strong electromagnetic fields are generated. These fields may lead to time dependent afterglow emission. In order to go beyond the standard synchrotron model used in astrophysical objects we have used PIC simulations and calculated radiation based on first principles. We calculated radiation from electrons propagating in a uniform parallel magnetic field to verify the technique. We also used the technique to calculate emission from electrons based on simulations with a small system. We obtain spectra which are consistent with those generated from electrons propagating in turbulent magnetic fields. This turbulent magnetic field is similar to the magnetic field generated at an early nonlinear stage of the Weibel instability. A fully developed shock within a larger system may generate a jitter/synchrotron spectrum.

Nishikawa, K.-I.↗

Multigrid Reduction in Time for Chaotic Dynamical Systems

As CPU clock speeds have stagnated and high performance computers continue to have ever higher core counts, increased parallelism is needed to take advantage of these new architectures. Traditional serial time-marching schemes can be a significant bottleneck, as many types of simulations require large numbers of time-steps which must be computed sequentially. Parallel-in-time schemes, such as the Multigrid Reduction in Time (MGRIT) method, remedy this by parallelizing across time-steps and have shown promising results for parabolic problems. However, chaotic problems have proved more difficult, since chaotic initial value problems (IVPs) are inherently ill-conditioned. MGRIT relies on a hierarchy of successively coarser time-grids to iteratively correct the solution on the finest time-grid, but due to the nature of chaotic systems, small inaccuracies on the coarser levels can be greatly magnified and lead to poor coarse-grid corrections. Here we introduce a modified MGRIT algorithm based on an existing quadratically converging nonlinear extension to the multigrid Full Approximation Scheme (FAS), as well as a novel time-coarsening scheme. Together, these approaches better capture long-term chaotic behavior on coarse-grids and greatly improve convergence of MGRIT for chaotic IVPs. Further, we introduce a novel low-memory variant of the algorithm for solving chaotic PDEs with MGRIT which not only solves the IVP, but also provides estimates for the unstable Lyapunov vectors of the system. Finally, we provide supporting numerical results for the Lorenz system and demonstrate parallel speedup for the chaotic Kuramoto–Sivashinsky PDE over a significantly longer time-domain than in previous works.

97 MATHEMATICS AND COMPUTING↗

Fluid dynamics applications of the Illiac IV computer

The Illiac IV is a parallel-structure computer with computing power an order of magnitude greater than that of conventional computers. It can be used for experimental tasks in fluid dynamics which can be simulated more economically, for simulating flows that cannot be studied by experiment, and for combining computer and experimental simulations. The architecture of Illiac IV is described, and the use of its parallel operation is demonstrated on the example of its solution of the one-dimensional wave equation. For fluid dynamics problems, a special FORTRAN-like vector programming language was devised, called CFD language. Two applications are described in detail: (1) the determination of the flowfield around the space shuttle, and (2) the computation of transonic turbulent separated flow past a thick biconvex airfoil.

Maccormack, R. W.↗

Modeling techniques in a parallelizing compiler for the B-HIVE multiprocessor system

The parallelizing compiler for the B-HIVE loosely-coupled multiprocessor system uses a medium grain model to minimize the communication overhead. A medium grain model is shown to be an optimum way of merging fine grain operations into parallel tasks such that the parallelism obtained at the grain level is retained and communication overhead is decreased. A new communication model is introduced in this paper, allowing additional overlap between computation and communication. Simulation results indicate that the medium grain communication model shows promise for automatic parallelization for a loosely-coupled multiprocessor system.

Kim, Sukil↗

High-resolution turbulent simulations using the Connection Machine-2

The spectral method provides an efficient algorithm for solving the 3D incompressible Navier-Stokes equations in periodic boundaries. Most people, so far, have used vectorized machines, such as the CRAY-2, to implement fast Fourier transformations and time integrations in the spectral calculations. In this paper, new results are presented using the spectral calculations on the Connection Machine-2 with a parallel algorithm. The large memory of the Connection Machine-2 and the parallel algorithm allows, of the first time, to implement a 512-cubed mesh resolution for high Reynolds number flows. The computational speed of the present code is about 30 percent faster than the fastest CRAY-2 simulations with four processors. Parallel machines, such as the Connection Machine-2, will possibly provide new computational power for understanding the intermittency and cascade mechanism in fluid turbulence.

Chen, Shiyi↗

Hydrodynamic irreversibility of non-Brownian suspensions in highly confined duct flow

The irreversible behaviour of a highly confined non-Brownian suspension of spherical particles at low Reynolds number in a Newtonian fluid is studied experimentally and numerically. In the experiment, the suspension is confined in a thin rectangular channel that prevents complete particle overlap in the narrow dimension and is subjected to an oscillatory pressure-driven flow. In the small cross-sectional dimension, particles rapidly separate to the walls, whereas in the large dimension, features reminiscent of shear-induced migration in bulk suspensions are recovered. Furthermore, as a consequence of the channel geometry and the development and application of a single-camera particle tracking method, three-dimensional particle trajectories are obtained that allow us to directly associate relative particle proximity with the observed migration. Companion simulations of a steadily flowing suspension highly confined between parallel plates are conducted using the force coupling method, which also show rapid migration to the walls as well as other salient features observed in the experiment. While we consider relatively low volume fractions compared to most prior work in the area, we nevertheless observe significant and rapid migration, which we attribute to the high degree of confinement.

42 ENGINEERING↗

Investigating the effects of electron bounce-cyclotron resonance on plasma dynamics in capacitive discharges operated in the presence of a weak transverse magnetic field

Recently, Patil et al. [Phys. Rev. Res. 4, 013059 (2022)] have reported the existence of an enhanced operating regime when a low-pressure (5 mTorr) capacitively coupled discharge (CCP) is driven by a very high radio frequency (60 MHz) source in the presence of a weak external magnetic field applied parallel to its electrodes. Their particle-in-cell simulations show that a significantly higher bulk plasma density and ion flux can be achieved at the electrode when the electron cyclotron frequency equals half of the applied radio frequency for a given fixed voltage. In the present work, we take a detailed look at this phenomenon and further delineate the effect of this “electron bounce-cyclotron resonance (EBCR)” on the electron and ion dynamics of the system. We find that the ionization collision rate and stochastic heating are maximum under resonance condition. The electron energy distribution function also indicates that the population of tail-end electrons is highest for the case where EBCR is maximum. Formation of electric field transients in the bulk plasma region is also seen at lower values of applied magnetic field. Finally, we demonstrate that the EBCR-induced effect is a low-pressure phenomenon and weakens as the neutral gas pressure increases. The potential utility of this effect to advance the operational performance of CCP devices for industrial purposes is discussed in this report.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A parallel, distributed memory implementation of the adaptive sampling configuration interaction method

The many-body simulation of quantum systems is an active field of research that involves several different methods targeting various computing platforms. Many methods commonly employed, particularly coupled cluster methods, have been adapted to leverage the latest advances in modern high-performance computing. Selected configuration interaction (sCI) methods have seen extensive usage and development in recent years. However, the development of sCI methods targeting massively parallel resources has been explored only in a few research works. Here, we present a parallel, distributed memory implementation of the adaptive sampling configuration interaction approach (ASCI) for sCI. In particular, we will address the key concerns pertaining to the parallelization of the determinant search and selection, Hamiltonian formation, and the variational eigenvalue calculation for the ASCI method. Load balancing in the search step is achieved through the application of memory-efficient determinant constraints originally developed for the ASCI-PT2 method. The presented benchmarks demonstrate near optimal speedup for ASCI calculations of Cr 2 (24e, 30o) with 10 6 , 10 7 , and 3 × 10 8 variational determinants on up to 16 384 CPUs. Importantly, to the best of the authors’ knowledge, this is the largest variational ASCI calculation to date.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MBMS1.0: An Open-Source Code for Modeling and Simulation of Membrane-Based Dehumidification and Energy Recovery

Membrane-based dehumidification is currently being considered as a promising solution for the building application due to its low cost and very limited energy consumption. Developing a simple and efficient open-source code simulation tool is important for boosting the optimization and evaluation of such device in HVAC community. This paper reports a first-order physics based model which accounts for the fundamental heat and mass transfer of humid-air vapor at feed side to flow stream at permeate side. The current model comprises two membrane mass transfer submodels (i.e. microstructure model and performance map model); and it adopts a segment-by-segment methodology for discretizing heat and mass transfer governing equations. The model is capable of simulating both dehumidifiers and energy recovery ventilators with parallel-flow cross-flow, and counter-flow configurations. The model was validated with the measurements at appropriate device. The practices in dehumidification and energy recovery exchangers are also discussed. The model and open-source codes are expected to become a solid fundament for developing a more comprehensive and accurate membrane-based dehumidification in the future.

Gao, Zhiming↗

All-Atom Simulation of 3D Hot Spot Formation in Shocked TATB Explosive

TATB is an insensitive high explosive (IHE) critical to the stockpile that is challenging to model at the continuum scale. Advanced detonation models in the Cheetah high explosive chemistry code require validation though subscale simulations. High explosive initiation is determined by micron-scale physics of hot spots formed a shock-collapsed pores. Pore sizes between 100 nm and 1 μm are believed to be the most important for determining the shock sensitivity of TATB. This range of pore sizes is difficult to access at the atomic scale through allatom molecular dynamics (MD) simulations, even with Sierra-class computers. Quasi-2D simulations are widely used and allow much larger pore sizes (up to 400 nm) to be studied, but the applicability of 2D simulations to the actual 3D pore response is not understood. Resolving these uncertainties through “full physics” MD modeling is key for generalizing, parameterizing, and validating the kinds of continuum models used to inform design, safety, and performance. This work was a continuation of FY20 efforts pushing simulations to full 3D with the largest-ever all-atom simulations of an explosive. These were the first all-atom full-3D simulations of large hot spots thought to govern explosive detonation and required over a billion atoms. Simulations were performed using LAMMPS, an open SNL science code. MD explosive models present unique challenges, even for established codes such as LAMMPS. Their model forms are more complex than typical models for metals, while simulating high temperature-pressure conditions is demanding and increases computational cost. Scaling problems in GPU-enabled MD algorithms initially limited simulations to <100 million atoms but were resolved through collaboration with SNL. An overall 24x speedup was obtained relative to CPU machines. Specialized analysis of these simulations required a bottom-up refactoring and algorithm parallelization of in-house codes and application of computer vision algorithms to extract meaningful information.

36 MATERIALS SCIENCE↗

Porting Classical Approaches for Quantum Simulations to Quantum Computers

Simulating quantum many-body systems is one of the most promising problems in which we might anticipate that quantum computers should show quantum advantage. Unfortunately, there is still a gap between this promise and actual practice. New quantum algorithms need to be developed and the current quantum algorithms have various difficulties - e.g efficient state preparation - which must be overcome and improved upon. In many cases, classical approaches need to be ported over to quantum devices. In this project we have developed a suite of new quantum algorithms which makes progress in this regard. We developed a new optimization scheme for variational quantum eigensolvers, UBOS, which mitigates problems with local minimas and barren plateaus while improving convergence to the ground state by an order of magnitude. We developed a new way to utilize qubitization to find ground states of nearly frustration-free Hamiltonians faster than all previous methods. We developed a series of state preparation techniques which helps initialize parameterized quantum circuits into reasonable starting points on which quantum algorithms are then applied. In addition to the development of novel algorithms, it is critical to have classical simulation techniques for approximately simulating quantum circuits which can be used to benchmark and understand quantum algorithms. Toward that end, we developed a novel POVM formalism to simulate quantum circuits as well as exemplify the massive parallelization of tensor network methodologies. Finally, we developed physical understanding of entanglement phase transitions such as many-body localization and random tensor networks.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

HEP - A semaphore-synchronized multiprocessor with central control

The paper describes the design concept of the Heterogeneous Element Processor (HEP), a system tailored to the special needs of scientific simulation. In order to achieve high-speed computation required by simulation, HEP features a hierarchy of processes executing in parallel on a number of processors, with synchronization being largely accomplished by hardware. A full-empty-reserve scheme of synchronization is realized by zero-one-valued hardware semaphores. A typical system has, besides the control computer and the scheduler, an algebraic module, a memory module, a first-in first-out (FIFO) module, an integrator module, and an I/O module. The architecture of the scheduler and the algebraic module is examined in detail.

Gilliland, M. C.↗

The P/POD project: Programmable/Pilot Oriented Display

A pilot orientated display system was developed for general aviation aircraft in order to reduce cockpit workloads. Emphasis was placed on the optimization of flight procedural aspects (i.e., interpretation of Loran data). Low cost hardware/software were utilized in the system to reduce developmental costs. Parallel development and testing were conducted on the ground (simulator) and in the air using the same hardware.

Littlefield, J. A.↗

Algorithmic commonalities in the parallel environment

The ultimate aim of this project was to analyze procedures from substantially different application areas to discover what is either common or peculiar in the process of conversion to the Massively Parallel Processor (MPP). Three areas were identified: molecular dynamic simulation, production systems (rule systems), and various graphics and vision algorithms. To date, only selected graphics procedures have been investigated. They are the most readily available, and produce the most visible results. These include simple polygon patch rendering, raycasting against a constructive solid geometric model, and stochastic or fractal based textured surface algorithms. Only the simplest of conversion strategies, mapping a major loop to the array, has been investigated so far. It is not entirely satisfactory.

Mcanulty, Michael A.↗

Origin of large magnetic fluctuations in the magnetosheath of Venus

The origin of large-amplitude hydromagnetic waves in the Venus magnetosheath downstream of the quasi-parallel bow shock is investigated by means of numerical simulations. It is shown that the most likely source of these waves is the bow shock itself, rather than an instability involving the solar wind and oxygen ions of planetary origin. Pickup of O(+) ions by these waves is also examined and shown to be in agreement with previous test particle calculations. The effect of mass loading on the structure of the shock is also discussed.

Winske, D.↗

Electronic neural network for dynamic resource allocation

A VLSI implementable neural network architecture for dynamic assignment is presented. The resource allocation problems involve assigning members of one set (e.g. resources) to those of another (e.g. consumers) such that the global 'cost' of the associations is minimized. The network consists of a matrix of sigmoidal processing elements (neurons), where the rows of the matrix represent resources and columns represent consumers. Unlike previous neural implementations, however, association costs are applied directly to the neurons, reducing connectivity of the network to VLSI-compatible 0 (number of neurons). Each row (and column) has an additional neuron associated with it to independently oversee activations of all the neurons in each row (and each column), providing a programmable 'k-winner-take-all' function. This function simultaneously enforces blocking (excitatory/inhibitory) constraints during convergence to control the number of active elements in each row and column within desired boundary conditions. Simulations show that the network, when implemented in fully parallel VLSI hardware, offers optimal (or near-optimal) solutions within only a fraction of a millisecond, for problems up to 128 resources and 128 consumers, orders of magnitude faster than conventional computing or heuristic search methods.

Thakoor, A. P.↗

Spatial awareness comparisons between large-screen, integrated pictorial displays and conventional EFIS displays during simulated landing approaches

An extensive simulation study was performed to determine and compare the spatial awareness of commercial airline pilots on simulated landing approaches using conventional flight displays with their awareness using advanced pictorial 'pathway in the sky' displays. Sixteen commercial airline pilots repeatedly made simulated complex microwave landing system approaches to closely spaced parallel runways with an extremely short final segment. Scenarios involving conflicting traffic situation assessments and recoveries from flight path offset conditions were used to assess spatial awareness (own ship position relative the the desired flight route, the runway, and other traffic) with the various display formats. The situation assessment tools are presented, as well as the experimental designs and the results. The results demonstrate that the integrated pictorial displays substantially increase spatial awareness over conventional electronic flight information systems display formats.

Parrish, Russell V.↗