Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

Leveraging STARE for Co-aligned Data Locality with netCDF and Python MPI

We have leveraged STARE indexing to package partitioned data chunks from diverse datasets into netCDF files, distributed them on a cluster of 16 lightweight nodes with their placements spatiotemporally co-aligned, and demonstrated a few integrative analyses using netCDF parallel I/O and Python MPI, with single-user performance and scalability comparable to, or even better than, that of a parallel array database management system (ADBMS) such as SciDB. However, records of the node location and STARE index ranges for each data chunk, similar to the chunk maps of SciDB, must be maintained and consulted by the I/O and analysis code for coordinating the analytic operations in parallel, in order to achieve the good performance and scalability.

Kwo-Sen Kuo↗

Stratigraphic Identification with Airborne Electromagnetic Methods at the Hanford Site, Washington

Stratigraphic units can influence the fate and transport of subsurface contaminants within groundwater. Units having coarse-grained sediments act as preferential flow pathways, and therefore can accelerate the transport of contaminants to reach human and ecological receptors. At legacy waste sites, detailed knowledge of subsurface stratigraphy can be used for effective monitoring and remediation planning to help minimize risk to human health and the environment. Airborne electromagnetic (AEM) methods can non-invasively provide information on kilometer-scale or larger subsurface stratigraphic features and fill informational gaps in directly sampled data from sparsely located boreholes. In this paper, we present inversion results of a 412 line-km frequency-domain AEM survey to delineate subsurface stratigraphic features at the Hanford Site, located in southeastern Washington State. The inversion was performed using a massively parallel 3D electromagnetic modeling and inversion code, where the modeling is based on solving frequency-domain Maxwell’s equations using an unstructured-mesh finite-element method and the inversion employs a Gauss-Newton optimization scheme. The results are compared to an underlying geologic framework model (GFM), built by interpolating contact depths of stratigraphic units interpreted from site borehole datasets. In areas with good borehole coverage, the inversion results show a good match with the GFM to a depth of about 60 m. Outside of these areas, the inversion results exhibit inconsistencies from the assumptions made to create the GFM, demonstrating that the AEM survey results can be used to improve the understanding of the geological conceptual model.

47 OTHER INSTRUMENTATION↗

Integrated hydrogeophysical modelling and data assimilation for geoelectrical leak detection

Time-lapse electrical resistivity tomography (ERT) measurements provide indirect observations of hydrological processes in the Earth's shallow subsurface at high spatial and temporal resolution. ERT has been used in the past decades to detect leaks and monitor the evolution of associated contaminant plumes. Specifically, inverted resistivity images allow visualization of the dynamic changes in the structure of the plume. However, existing methods do not allow the direct estimation of leak parameters (e.g. leak rate, location, etc.) and their uncertainties. We propose an ensemble-based data assimilation framework that evaluates proposed hydrological models against observed time-lapse ERT measurements without directly inverting for the resistivities. Each proposed hydrological model is run through the parallel coupled hydro-geophysical simulation code PFLOTRAN-E4D to obtain simulated ERT measurements. The ensemble of model proposals is then updated using an iterative ensemble smoother. In this paper, we demonstrate the proposed framework on synthetic and field ERT data from controlled tracer injection experiments. Our results show that the approach allows joint identification of contaminant source location, initial release time, and solute loading from the cross-borehole time-lapse ERT data, alongside with an assessment of uncertainties in these estimates. We demonstrate a reduction in site-wide uncertainty by comparing the prior and posterior plume mass discharges at a selected image plane. This framework is particularly attractive to sites that have previously undergone extensive geological investigation (e.g., nuclear sites). It is well suited to complement ERT imaging and we discuss practical issues in its application to field problems.

58 GEOSCIENCES↗

Parallel algorithms for hyperdynamics and local hyperdynamics

Hyperdynamics (HD) is a method for accelerating the timescale of standard molecular dynamics (MD). It can be used for simulations of systems with an energy potential landscape that is a collection of basins, separated by barriers, where transitions between basins are infrequent. HD enables the system to escape from a basin more quickly while enabling a statistically accurate renormalization of the simulation time, thus effectively boosting the timescale of the simulation. In [Kim, Perez, Voter, J Chem Phys, 139:144110, 2013)1, a local version of HD was formulated, which exploits the intrinsic locality characteristic typical of most systems to mitigate the poor scaling properties of standard HD as the system size is increased. In this paper, we discuss how both HD and local HD can be formulated to run efficiently in parallel. We have implemented these ideas in the LAMMPS MD code, which means HD can be used with any interatomic potential LAMMPS supports. Together, these parallel methods allow simulations of any size to achieve the time acceleration offered by HD (which can be orders of magnitude), at a cost 3-5x that of standard MD. As examples, we performed two simulations of a million-atom system to model the diffusion and clustering of Pt adatoms on a large patch of Pt(100) surface for 80 and 160 μs.

74 ATOMIC AND MOLECULAR PHYSICS↗

A resistive MHD model and simulation on plasma flow evolution in the presence of resonant magnetic perturbation in a tokamak

Nonaxisymmetric magnetic fields such as the intrinsic error field and the externally applied resonant magnetic perturbation (RMP) in a tokamak are known to influence the plasma momentum transport and flow evolution through plasma response, which itself strongly depends on the plasma flow as well. The nonlinear interaction between plasma response and flow has been previously modeled in the conventional error field theory with the “no-slip” condition, which has been recently extended to allow the “free-slip” condition. In this work, we further target this specific process and numerically simulate the nonlinear plasma response and flow evolution in the presence of a single-helicity RMP in a circular-shaped model tokamak configuration, based on the full resistive MHD model in the initial-value code NIMROD. Time evolution of the parallel (to k) flow or “slip frequency” profile and its asymptotic steady state obtained from the NIMROD simulations are compared with both conventional and extended nonlinear response theories. Here, k is the wave vector of the propagating island. Good agreement with the extended theory with free-slip condition has been achieved for the parallel flow profile evolution in response to RMP in all resistive regimes, whereas the difference from the conventional theory with the no-slip condition tends to diminish as the plasma resistivity approaches zero.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Phoebe: a high-performance framework for solving phonon and electron Boltzmann transport equations

Understanding the electrical and thermal transport properties of materials is critical to the design of electronics, sensors, and energy conversion devices. Computational modeling can accurately predict material properties but, in order to be reliable, requires accurate descriptions of electron and phonon states and their interactions. While first-principles methods are capable of describing the energy spectrum of each carrier, using them to compute transport properties is still a formidable task, both computationally demanding and memory intensive, requiring integration of fine microscopic scattering details for estimation of macroscopic transport properties. To address this challenge, we present Phoebe—a newly developed software package that includes the effects of electron–phonon, phonon–phonon, boundary, and isotope scattering in computations of electrical and thermal transport properties of materials with a variety of available methods and approximations. This open source C++ code combines MPI-OpenMP hybrid parallelization with GPU acceleration and distributed memory structures to manage computational cost, allowing Phoebe to effectively take advantage of contemporary computing infrastructures. We demonstrate that Phoebe accurately and efficiently predicts a wide range of transport properties, opening avenues for accelerated computational analysis of complex crystals.

36 MATERIALS SCIENCE↗

Single classical field description of interacting scalar fields

In this report we test the degree to which interacting Bosonic systems can be approximated by a classical field as total occupation number is increased. This is done with our publicly available code repository, QIBS, a new massively parallel solver for these systems. We use a number of toy models well studied in the literature and track when the classical field description admits quantum corrections, called the quantum breaktime. This allows us to test claims in the literature regarding the rate of convergence of these systems to the classical evolution. We test a number of initial conditions, including coherent states, number eigenstates, and field number states. We find that of these initial conditions, only number eigenstates do not converge to the classical evolution as occupation number is increased. We find that systems most similar to scalar field dark matter exhibit a logarithmic enhancement in the quantum breaktime with total occupation number. Systems with contact interactions or with field number state initial conditions, and linear dispersions, exhibit a power law enhancement. Finally, we find that the breaktime scaling depends on both model interactions and initial conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

Ensuring statistical reproducibility of ocean model simulations in the age of hybrid computing

Novel high performance computing systems that feature hybrid architectures require large scale code refactoring to unravel underlying exploitable parallelism. Such redesign can often be accompanied with machine-precision changes as the order of computation cannot always be maintained. For chaotic systems like climate models, these round-off level differences can grow rapidly. Systematic errors may also manifest initially as machine-precision differences. Isolating genuine round off level differences from such errors remains a challenge. Here, we apply two-sample equality of distribution tests to evaluate statistical reproducibility of the ocean model component of US Department of Energy's Energy Exascale Earth System Model (E3SM). A 2-year control simulation ensemble is compared to a modified ensemble as a test case - after a known non-bit-for-bit change in a model component is introduced - to evaluate the null hypothesis that the two ensembles are statistically indistinguishable. To quantify the false negative rates of these tests, we conduct a formal power analysis using a targeted suite of short simulation ensembles. The ensemble suite contains several perturbed ensembles, each with a progressively different climate than the baseline ensemble - obtained by perturbing the magnitude of a single model tuning parameter, the Gent and McWilliams κ, in a controlled manner. The null hypothesis is evaluated for each of perturbed ensembles using these tests. The power analysis informs on the detection limits of the tests for given ensemble size allowing model developers to evaluate the impact of an introduced non-bit-for-bit change to the model.

Mahajan, Salil↗

PFLOTRAN

PFLOTRAN is a GNU LGPL licensed code that simulates single phase water and multiphase (air, co2, water) flow, energy transport, geochemical reaction and biological reaction in the earth's subsurface environment. The problem may be discretized on structured and unstructured grids. The code is designed to run in parallel on supercomputers and leverages the open source PETSc library for solvers and data structures and open source HDF5 for binary file I/O. See www.pflotran.org and/or doc-dev.pflotran.org for further information. SAND2020-12161 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Hammond, Glenn↗

Data-driven closure modeling for hypersonic turbulent flows

The Reynolds-averaged Navier–Stokes (RANS) equations remain a workhorse technology for simulating compressible fluid flows of practical interest. Due to model-form errors, however, RANS models can yield erroneous predictions that preclude their use on mission-critical problems. This report summarizes work performed from FY22-FY24 focused on improving RANS models for hypersonic flows using data-driven modeling and scientific machine learning. In this work we: 1. Investigate the current capabilities of RANS models in Sandia’s parallel aerodynamics and re-entry code (SPARC) for hypersonic flows with a focus on shock boundary layer interactions (SBLIs), 2. Assess several established corrections that exist in the literature aimed at improving predictions for SBLIs, 3. Develop improved models for the Reynolds stress tensor using tensor-basis neural networks, 4. Develop a neural-network-based variable turbulent Prandtl number model to reduce errors in wall heating in SBLIs. 5. Begin future investigations including employing the LIFE framework to improve wall heating predictions in SBLIs as well as the ensemble Kalman filter. We find that current RANS models in SPARC are deficient for complex SBLI flows. In particular, no current model jointly predicts wall heat flux, wall shear stress, and wall pressure with reasonable accuracy. Existing corrections help, but do not alleviate this issue altogether. The development of improved models for the Reynolds stress tensor via tensor-basis neural networks results in more predictive RANS models across a suite of low-speed and high-speed cases. For hypersonic boundary layers, the inclusion of the wall-normal Reynolds stress via TBNNs has an appreciable impact on the wall-normal momentum balance and wall quantities. However, we find that improvements to the Reynolds stress tensor do not address the over-prediction in wall heat flux in SBLIs. We find that a neural-network-based variable turbulent Prandtl number model systematically and substantially improves wall heating predictions for a range of SBLI cases.

97 MATHEMATICS AND COMPUTING↗

Real time animation of space plasma phenomena

In pursuit of real time animation of computer simulated space plasma phenomena, the code was rewritten for the Massively Parallel Processor (MPP). The program creates a dynamic representation of the global bowshock which is based on actual spacecraft data and designed for three dimensional graphic output. This output consists of time slice sequences which make up the frames of the animation. With the MPP, 16384, 512 or 4 frames can be calculated simultaneously depending upon which characteristic is being computed. The run time was greatly reduced which promotes the rapid sequence of images and makes real time animation a foreseeable goal. The addition of more complex phenomenology in the constructed computer images is now possible and work proceeds to generate these images.

Jordan, K. F.↗

The numerical simulation of a high-speed axial flow compressor

The advancement of high-speed axial-flow multistage compressors is impeded by a lack of detailed flow-field information. Recent development in compressor flow modeling and numerical simulation have the potential to provide needed information in a timely manner. The development of a computer program is described to solve the viscous form of the average-passage equation system for multistage turbomachinery. Programming issues such as in-core versus out-of-core data storage and CPU utilization (parallelization, vectorization, and chaining) are addressed. Code performance is evaluated through the simulation of the first four stages of a five-stage, high-speed, axial-flow compressor. The second part addresses the flow physics which can be obtained from the numerical simulation. In particular, an examination of the endwall flow structure is made, and its impact on blockage distribution assessed.

Mulac, Richard A.↗

Modifications of magnetohydrodynamics as applied to the solar wind

The effect of including the Braginskii viscous stress tensor in magnetohydrodynamics is remarked upon. It is shown that semiquantitative agreement with a recently observed anisotropy in the turbulent solar wind spectrum can be achieved in this way. The modifications of the dynamical equations are simple enough to permit their inclusion in numerical codes. The effects of large 'ion parallel viscosities' also may be significant for plasmas in quite different regimes than the solar wind.

Montgomery, David↗

Nonlinear evolution of a large-amplitude circularly polarized Alfven wave: High beta

The nonlinear dynamics following saturation of the parametric instabilities of a monochromatic field-aligned large-amplitude circularly polarized Alfven wave is investigated via direct numerical simulation in the case of high plasma beta and no wave dispersion. The magnetohydrodynamic (MHD) code permits nonlinear couplings in the parallel direction to the ambient magnetic field and one perpendicular direction. Compressibility is included in the form of a polytropic equation of state. Turbulent cascades develop after saturation of two coupled oblique three-wave parametric instabilities; one of which is an oblique filamentationlike instability reported earlier. Remnants of the parametric processes, as well as of the original Alfven pump wave, persist during late nonlinear times. Nearly incompressible MHD features such as spectral anisotropies appear as well.

Ghosh, S.↗

Nonlinear evolution of a large-amplitude circularly polarized Alfven wave: Low beta

The nature of turbulent cascades arising from the parametric instabilities of a monochromatic field-aligned large-amplitude circularly polarized Alfven wave is investigated via direct numerical simulation for the case of low plasma Beta and no wave dispersion. The magnetohydrodynamic code permits nonlinear couplings in the parallel direction to the ambient magnetic field and one perpendicular direction. Compressibility is included in the form of a polytropic equation of state. Anisotropic turbulent cascades, similar to those found in early incompressible two-dimensional simulations, occur after nonlinear saturation of the parallel propagating decay instability. The turbulent spectrum can be divided into three regimes: the lowest wave numbers are dominated by lower sideband remnants of the parametric process, intermediate wave numbers display nearly incompressible dynamics, and the highest wave numbers are dominated by acoustic turbulence.

Ghosh, S.↗

Model atmospheres for M (sub)dwarf stars. 1: The base model grid

We have calculated a grid of more than 700 model atmospheres valid for a wide range of parameters encompassing the coolest known M dwarfs, M subdwarfs, and brown dwarf candidates: 1500 less than or equal to T(sub eff) less than or equal to 4000 K, 3.5 less than or equal to log g less than or equal to 5.5, and -4.0 less than or equal to (M/H) less than or equal to +0.5. Our equation of state includes 105 molecules and up to 27 ionization stages of 39 elements. In the calculations of the base grid of model atmospheres presented here, we include over 300 molecular bands of four molecules (TiO, VO, CaH, FeH) in the JOLA approximation, the water opacity of Ludwig (1971), collision-induced opacities, b-f and f-f atomic processes, as well as about 2 million spectral lines selected from a list with more than 42 million atomic and 24 million molecular (H2, CH, NH, OH, MgH, SiH, C2, CN, CO, SiO) lines. High-resolution synthetic spectra are obtained using an opacity sampling method. The model atmospheres and spectra are calculated with the generalized stellar atmosphere code PHOENIX, assuming LTE, plane-parallel geometry, energy (radiative plus convective) conservation, and hydrostatic equilibrium. The model spectra give close agreement with observations of M dwarfs across a wide spectral range from the blue to the near-IR, with one notable exception: the fit to the water bands. We discuss several practical applications of our model grid, e.g., broadband colors derived from the synthetic spectra. In light of current efforts to identify genuine brown dwarfs, we also show how low-resolution spectra of cool dwarfs vary with surface gravity, and how the high-regulation line profile of the Li I resonance doublet depends on the Li abundance.

Allard, France↗

Global Load Balancing with Parallel Mesh Adaption on Distributed-Memory Systems

Dynamic mesh adaption on unstructured grids is a powerful tool for efficiently computing unsteady problems to resolve solution features of interest. Unfortunately, this causes load imbalance among processors on a parallel machine. This paper describes the parallel implementation of a tetrahedral mesh adaption scheme and a new global load balancing method. A heuristic remapping algorithm is presented that assigns partitions to processors such that the redistribution cost is minimized. Results indicate that the parallel performance of the mesh adaption code depends on the nature of the adaption region and show a 35.5X speedup on 64 processors of an SP2 when 35% of the mesh is randomly adapted. For large-scale scientific computations, our load balancing strategy gives almost a sixfold reduction in solver execution times over non-balanced loads. Furthermore, our heuristic remapper yields processor assignments that are less than 3% off the optimal solutions but requires only 1% of the computational time.

Biswas, Rupak↗