Massively Parallel Adaptive Computational Fluid and Solid Dynamics for Engineering Applications (Final Report)
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
A fluidic cartridge comprises a fluidic disk having a plurality of alignment openings; a fluidic chip comprising a body, one or more channels formed in the body in fluidic communications with input ports and output ports for transferring one or more fluids between the input ports and the output ports, and a plurality of protrusions formed on the body and received in the alignment openings of the fluidic disk for aligning the fluidic chip to the fluidic disk; an actuator operably engaging with the one or more channels for selectively and individually transferring the one or more fluids through the one or more channels from at least one of the input ports to at least one of the output ports at desired flow rates; and a tube member defining a cylindrical housing for accommodating the fluidic disk, the fluidic chip and the actuator therein.
Explore the source record for details and available documents.
In one aspect, the present disclosure provides a nozzle for a 3D printing system. The nozzle may include a flowpath with a material inlet and a material outlet. The nozzle may further include a valve in fluid communication with the flowpath between the material inlet and the material outlet, where the valve includes a closed state and an open state, where in the closed state the valve obstructs the flowpath between the material inlet and the material outlet, and where in the open state the material inlet is in fluid communication with the material outlet. The nozzle may further include a compensator in fluid communication with the flowpath, where the compensator includes a contracted state associated with the open state of the valve and an expanded state associated with the closed state of the valve.
Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.
Programmable accelerators have become commonplace in modern computing systems. Advances in programming models and the availability of unprecedented amounts of data have created a space for massively parallel accelerators capable of maintaining context for thousands of concurrent threads resident on-chip. These threads are grouped and interleaved on a cycle-by-cycle basis among several massively parallel computing cores. One path for the design of future supercomputers relies on an ability to model the performance of these massively parallel cores at scale. The SST framework has been proven to scale up to run simulations containing tens of thousands of nodes. A previous report described the initial integration of the open-source, execution-driven GPU simulator, GPGPU-Sim, into the SST framework. This report discusses the results of the integration and how to use the new GPU component in SST. It also provides examples of what it can be used to analyze and a correlation study showing how closely the execution matches that of a Nvidia V100 GPU when running kernels and mini-apps.
This report presents a specification for the Portals 4 network programming interface. Portals 4 is intended to allow scalable, high-performance network communication between nodes of a parallel computing system. Portals 4 is well suited to massively parallel processing and embedded systems. Portals 4 represents an adaption of the data movement layer developed for massively parallel processing platforms, such as the 4500-node Intel TeraFLOPS machine. Sandia's Cplant cluster project motivated the development of Version 3.0, which was later extended to Version 3.3 as part of the Cray Red Storm machine and XT line. Version 4 is targeted to the next generation of machines employing advanced network interface architectures that support enhanced offload capabilities.
Charged-particle radiography and shadowgraphy data can be directly inverted to obtain a line-integrated transverse Lorentz force or a line-integrated transverse refractive index gradient if intensity modulations due to scattering and absorption are negligible, and angular deflections are small. We develop a new direct-inversion algorithm based on plasma physics and compare it to a new Monge–Ampère code and an existing power diagram code. The measured or source intensity is represented by electrons subject to drag, and the other intensity by fixed ions. The decrease in kinetic plus electrostatic energy determines convergence. The displacement of the electrons from their initial to their equilibrium positions determines the line-integrated force or refractive index gradient. We have implemented two approaches: PIC (particle in cell) and Lagrangian fluid, in 1-D and 2-D. The PIC code works for arbitrary intensities, can work efficiently in parallel, and can make use of existing codes. The Lagrangian code requires less memory and is faster than the PIC code without massively parallel processing, but fails in 2-D for large intensity modulations. The Monge–Ampère code is by far the fastest in 2-D, without massively parallel processing, but fails for intensities with large voids, high contrast ratios and large deflections across the boundaries, and could not obtain the degree of convergence possible with the PIC code. As a result, the power diagram code was by far the slowest and most memory intensive, and failed for large peaks in the measured intensity.
Systems and methods of building massively parallel computing systems using low power computing complexes in accordance with embodiments of the invention are disclosed. A massively parallel computing system in accordance with one embodiment of the invention includes at least one Solid State Blade configured to communicate via a high performance network fabric. In addition, each Solid State Blade includes a processor configured to communicate with a plurality of low power computing complexes interconnected by a router, and each low power computing complex includes at least one general processing core, an accelerator, an I/O interface, and cache memory and is configured to communicate with non-volatile solid state memory.
The central importance of large-scale eigenvalue problems in scientific computation necessitates the development of massively parallel algorithms for their solution. Recent advances in dense numerical linear algebra have enabled the routine treatment of eigenvalue problems with dimensions on the order of hundreds of thousands on the world’s largest supercomputers. In cases where dense treatments are not feasible, Krylov subspace methods offer an attractive alternative due to the fact that they do not require storage of the problem matrices. However, demonstration of scalability of either of these classes of eigenvalue algorithms on computing architectures capable of expressing massive parallelism is non-trivial due to communication requirements and serial bottlenecks, respectively. In this work, we introduce the SISLICE method: a parallel shift-invert algorithm for the solution of the symmetric self-consistent field (SCF) eigenvalue problem. The SISLICE method drastically reduces the communication requirement of current parallel shift-invert eigenvalue algorithms through various shift selection and migration techniques based on density of states estimation and k-means clustering, respectively. This work demonstrates the robustness and parallel performance of the SISLICE method on a representative set of SCF eigenvalue problems and outlines research directions that will be explored in future work.
This report outlines Sandia National Laboratories modeling studies applied to Stage 1 and Stage 2 of the Full-scale Engineered Barriers Experiment in Crystalline Host Rock (FEBEX) in situ test for the SKB EBS Task Force Task 9. The FEBEX test was a full-scale test conducted over ~18 years at the Grimsel, Switzerland Underground Research Laboratory (URL) managed by NAGRA. It involved emplacing simulated waste packages, in the form of welded cylindrical heaters, inside a tunnel in crystalline granitic rock and surrounded by a bentonite barrier and cement plug. Sensors emplaced within the bentonite monitored the wetting-up, heating, and drying out of the bentonite barrier, and the large resulting data set provides an excellent opportunity for validation of multiphysics Thermal-Hydrological (TH), Thermal-Hydrologic-Chemical (THC), and Thermal-Hydrological-Mechanical (THM) modeling approaches for underground nuclear waste storage and the performance of engineered bentonite barriers. The present status of the EBS Task Force is finalizing Task 9, which follows years of modeling studies of the FEBEX test, by many notable modeling teams (Gens et al., 2009; Sanchez et al. 2010; 2012; Samper et al., 2018). These modeling studies generally use two-dimensional axisymmetric meshes, ignoring threedimensional effects, gravity and asymmetric wetting and dry out of the bentonite engineered barrier. This study investigates these effects with use of the PFLOTRAN THC code with massively parallel computational methods in modeling FEBEX Stage 1 and Stage 2 results. The PFLOTRAN numerical code is an open source, state-of-the-art, massively parallel subsurface flow and reactive transport code operating in a high-performance computing environment (Hammond et al., 2014). Section 2 describes the applied partial differential equations describing mass, momentum and energy balance used in this study, considerations derived by assuming phase equilibrium between gas and liquid phases, constitutive equations for granite, cement plug, and bentonite domains, and specific approaches for use inthe PFLOTRAN code. Section 3 describes the geometry, meshing, and model set-up. Section 4 describes modeling results, Section 5 compares modeling results to field testing data, and Section 6 gives conclusions. The Appendix provides detailed information required by the EBSTask Force for final reporting.
Abstract The growth of computing needs for artificial intelligence and machine learning is critically challenging data communications in today’s data-centre systems. Data movement, dominated by energy costs and limited ‘chip-escape’ bandwidth densities, is perhaps the singular factor determining the scalability of future systems. Using light to send information between compute nodes in such systems can dramatically increase the available bandwidth while simultaneously decreasing energy consumption. Through wavelength-division multiplexing with chip-based microresonator Kerr frequency combs, independent information channels can be encoded onto many distinct colours of light in the same optical fibre for massively parallel data transmission with low energy. Although previous high-bandwidth demonstrations have relied on benchtop equipment for filtering and modulating Kerr comb wavelength channels, data-centre interconnects require a compact on-chip form factor for these operations. Here we demonstrate a massively scalable chip-based silicon photonic data link using a Kerr comb source enabled by a new link architecture and experimentally show aggregate single-fibre data transmission of 512 Gb s −1 across 32 independent wavelength channels. The demonstrated architecture is fundamentally scalable to hundreds of wavelength channels, enabling massively parallel terabit-scale optical interconnects for future green hyperscale data centres.
Here, we present modal-based methods for model calibration in structural dynamics, and address several key challenges in the solution of gradient-based optimization problems with eigenvalues and eigenvectors, including the solution of singular Helmholtz problems encountered in sensitivity calculations, non-differentiable objective functions caused by mode swapping during optimization, and cases with repeated eigenvalues. Unlike previous literature that relied on direct solution of the eigenvector adjoint equations, we present a parallel iterative domain decomposition strategy (Adjoint Computation via Modal Superposition with Truncation Augmentation) for the solution of the singular Helmholtz problems. For problems with repeated eigenvalues we present a novel Mode Separation via Projection algorithm, and in order to address mode swapping between inverse iterations we present a novel Injective mode ordering metric. We present the implementation of these methods in a massively parallel finite element framework with the ability to use measured modal data to extract unknown structural model parameters from large complex problems. A series of increasingly complex numerical examples are presented that demonstrate the implementation and performance of the methods in a massively parallel finite element framework [7], [5], using gradient-based optimization techniques in the Rapid Optimization Library (ROL) [21].
With the growing reliance of modern supercomputers on accelerator-based architecture such a graphics processing units (GPUs), the development and optimization of electronic structure methods to exploit these massively parallel resources has become a recent priority. While significant strides have been made in the development GPU accelerated, distributed memory algorithms for many modern electronic structure methods, the primary focus of GPU development for Gaussian basis atomic orbital methods has been for shared memory systems with only a handful of examples pursing massive parallelism. Here in this work, we present a set of distributed memory algorithms for the evaluation of the Coulomb and exact exchange matrices for hybrid Kohn–Sham DFT with Gaussian basis sets via direct density-fitted (DF-J-Engine) and seminumerical (sn-K) methods, respectively. The absolute performance and strong scalability of the developed methods are demonstrated on systems ranging from a few hundred to over one thousand atoms using up to 128 NVIDIA A100 GPUs on the Perlmutter supercomputer.
We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.
Accurate prediction of wind-plant performance relies, in part, on properly characterizing the turbulent atmospheric boundary layer (ABL) flow in which wind turbines operate. Large-eddy simulation (LES) is a powerful tool for simulating ABLs because it resolves the largest, most energetic scales of three-dimensional turbulent motions. Yet LES predictions are well known to depend on modeling choices such as grid resolution, numerical discretization schemes, and closures for unresolved scales of turbulence. Here, we evaluate how these choices influence predictions of ABL winds using Nalu-Wind, a wind-specific fork of the open-source, generalized, unstructured, massively parallel flow solver NaluCFD/Nalu.
We report Tritium Migration Analysis Program version 8 (TMAP8), the latest version of TMAP, was developed within the framework of the Multiphysics Object-Oriented Simulation Environment (MOOSE). Created at Idaho National Laboratory (INL), MOOSE is an open-source, dimension-agnostic, fully coupled, and fully implicit multiphysics platform featuring massively parallel computation capabilities. Using TMAP8, tritium transport in a divertor monoblock was analyzed to elucidate the effects of pulsed operation (up to fifty 1,600 s plasma discharge and cool-down cycles) on the tritium in-vessel inventory source term and ex-vessel release term (i.e., tritium retention and permeation) for safety analysis. With its built-in Message Passing Interface capability, TMAP8 can, in under 2 h, simulate tritium transport in three different layered materials (i.e., tungsten, copper, and copper-chromium-zirconium alloy) in 2D geometry, using a single device/computer with 10 cores. The MOOSE-based TMAP8 code can leverage other MOOSE tools developed under the Nuclear Energy Advanced Modeling and Simulation program to perform tritium and thermal transport in complex geometries and multiphysics environments. And via its massively parallel computation, MOOSE will enable the fusion pilot plant designers to conduct high-fidelity multiphysics modeling for the design of the divertor and blanket systems as well as for the safety analysis.