Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

The Kokkos Ecosystem [Brief]

In 2016/2017, the field of High-Performance Computing (HPC) entered a new era driven by fundamental physics challenges to produce ever more energy and cost-efficient processors. Since the convergence on the Message-Passing Interface (MPI) standard in the mid-1990s, application developers enjoyed a seemingly static view of the underlying machine — that of a distributed collection of homogeneous nodes executing in collaboration. However, after almost two decades of dominance, the sole use of MPI to derive parallelism acted as a limiter to improved future performance. While MPI is widely expected to continue to function as the basic mechanism for communication between compute nodes for the immediate future, additional parallelism is required on the computing node itself if high performance and efficiency goals are to be realized. When reviewing the architectures of the top HPC systems today, the change in paradigm is clear: the compute nodes of the leading machines in the world are either powered by many-core chips with a few dozen cores each, or use heterogeneous designs, where traditional CPUs marshal work to massively parallel compute accelerators which has as many as 200,000 processing threads in flight simultaneously. Complicating matters further for application developers, each processor vendor has its own preferred way of writing code for their architecture.The Kokkos EcoSystem was released by Sandia in 2017 to address this new era in HPC system design by providing a vendor independent performance portable programming system for scientific, engineering, and mathematical software applications written in the C++ programming language. Using Kokkos, application developers can be more productive because they will not have to create and maintain separate versions of their software for each architecture, nor will they have to be experts in each architecture's peculiar requirements. Instead, they will have a single method of programming for the diverse set of modern HPC architectures. While Kokkos started in 2011 as a programming model only, it soon became clear that complex applications needed more. It is also critical to have a portable mathematical functions and developers need tools to debug their applications, gain insight into the performance characteristics of their codes and tune algorithm performance parameters through automated processes. The Kokkos EcoSystem addresses those needs through its three main components: the Kokkos Core programming model, the Kokkos Kernels math library, and the Kokkos Tools project.

97 MATHEMATICS AND COMPUTING↗

X-composer: enabling cross-environments in-situ workflows between HPC and cloud

As large-scale scientific simulations and big data analyses become more popular, it is increasingly more expensive to store huge amounts of raw simulation results to perform post-analysis. To minimize the expensive data I/O, "in-situ" analysis is a promising approach, where data analysis applications analyze the simulation generated data on the fly without storing it first. However, it is challenging to organize, transform, and transport data at scales between two semantically different ecosystems due to the distinct software and hardware difference. To tackle these challenges, we design and implement the X-Composer framework. X-Composer connects cross-ecosystem applications to form an "in-situ" scientific workflow, and provides a unified approach and recipe for supporting such hybrid in-situ workflows on distributed heterogeneous resources. X-Composer reorganizes simulation data as continuous data streams and feeds them seamlessly into the Cloud-based stream processing services to minimize I/O overheads. For evaluation, we use X-Composer to set up and execute a cross-ecosystem workflow, which consists of a parallel Computational Fluid Dynamics simulation running on HPC, and a distributed Dynamic Mode Decomposition analysis application running on Cloud. Our experimental results show that X-Composer can seamlessly couple HPC and Big Data jobs in their own native environments, achieve good scalability, and provide high-fidelity analytics for ongoing simulations in real-time.

Wang, Dali↗

Power Converter Circuit Design Automation using Parallel Monte Carlo Tree Search

The tidal waves of modern electronic/electrical devices have led to increasing demands for ubiquitous application-specific power converters. A conventional manual design procedure of such power converters is computation- and labor-intensive, which involves selecting and connecting component devices, tuning component-wise parameters and control schemes, and iteratively evaluating and optimizing the design. To automate and speed up this design process, we propose an automatic framework that designs custom power converters from design specifications using Monte Carlo Tree Search. Specifically, the framework embraces the upper-confidence-bound-tree (UCT), a variant of Monte Carlo Tree Search, to automate topology space exploration with circuit design specification-encoded reward signals. Moreover, our UCT-based approach can exploit small offline data via the specially designed default policy and can run in parallel to accelerate topology space exploration. Further, it utilizes a hybrid circuit evaluation strategy to substantially reduce design evaluation costs. Empirically, we demonstrated that our framework could generate energy-efficient circuit topologies for various target voltage conversion ratios. Compared to existing automatic topology optimization strategies, the proposed method is much more computationally efficient --- the sequential version can generate topologies with the same quality while being up to 67% faster. Here, the parallelization schemes can further achieve high speedups compared to the sequential version.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A non-cooperative meta-modeling game for automated third-party calibrating, validating and falsifying constitutive laws with parallelized adversarial attacks

The evaluation of constitutive models, especially for high-risk and high-regret engineering applications, requires efficient and rigorous third-party calibration, validation and falsification. While there are numerous efforts to develop paradigms and standard procedures to validate models, difficulties may arise due to the sequential, manual, and often biased nature of the commonly adopted calibration and validation processes, thus slowing down data collections, hampering the progress towards discovering new physics, increasing expenses and possibly leading to misinterpretations of the credibility and application ranges of proposed models. This work attempts to introduce concepts from game theory and machine learning techniques to overcome many of these existing difficulties. Here, we introduce an automated meta-modeling game where two competing AI agents systematically generate experimental data to calibrate a given constitutive model and to explore its weakness such that the experiment design and model robustness can be improved through competitions. The two agents automatically search for the Nash equilibrium of the meta-modeling game in an adversarial reinforcement learning framework without human intervention. In particular, a protagonist agent seeks to find the more effective ways to generate data for model calibrations, while an adversary agent tries to find the most devastating test scenarios that expose the weaknesses of the constitutive model calibrated by the protagonist. By capturing all possible design options of the laboratory experiments into a single decision tree, we recast the design of experiments as a game of combinatorial moves that can be resolved through deep reinforcement learning by the two competing players. Our adversarial framework emulates idealized scientific collaborations and competitions among researchers to achieve a better understanding of the application range of the learned material laws and prevent misinterpretations caused by conventional AI-based third-party validation. Numerical examples are given to demonstrate the wide applicability of the proposed meta-modeling game with adversarial attacks on both human-crafted constitutive models and machine learning models.

97 MATHEMATICS AND COMPUTING↗

The JOREK non-linear extended MHD code and applications to large-scale instabilities and their control in magnetically confined fusion plasmas

JOREK is a massively parallel fully implicit non-linear extended magneto-hydrodynamic (MHD) code for realistic tokamak X-point plasmas. It has become a widely used versatile simulation code for studying large-scale plasma instabilities and their control and is continuously developed in an international community with strong involvements in the European fusion research programme and ITER organization. This article gives a comprehensive overview of the physics models implemented, numerical methods applied for solving the equations and physics studies performed with the code. A dedicated section highlights some of the verification work done for the code. A hierarchy of different physics models is available including a free boundary and resistive wall extension and hybrid kinetic-fluid models. The code allows for flux-surface aligned iso-parametric finite element grids in single and double X-point plasmas which can be extended to the true physical walls and uses a robust fully implicit time stepping. Particular focus is laid on plasma edge and scrape-off layer (SOL) physics as well as disruption related phenomena. Among the key results obtained with JOREK regarding plasma edge and SOL, are deep insights into the dynamics of edge localized modes (ELMs), ELM cycles, and ELM control by resonant magnetic perturbations, pellet injection, as well as by vertical magnetic kicks. Also ELM free regimes, detachment physics, the generation and transport of impurities during an ELM, and electrostatic turbulence in the pedestal region are investigated. Regarding disruptions, the focus is on the dynamics of the thermal quench (TQ) and current quench triggered by massive gas injection and shattered pellet injection, runaway electron (RE) dynamics as well as the RE interaction with MHD modes, and vertical displacement events. Also the seeding and suppression of tearing modes (TMs), the dynamics of naturally occurring TQs triggered by locked modes, and radiative collapses are being studied.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Magnetism and interlayer bonding in pores of Bernal-stacked hexagonal boron nitride

When single-layer h-BN is subjected to a high-energy electron beam, triangular pores with nitrogen edges are formed. Because of the broken sp 2 bonds, these pores are known to possess magnetic states. Here, we report on the magnetism and electronic structure of triangular pores as a function of their size. Moreover, in the Bernal-stacked h-BN (AB-h-BN), multilayer pores with parallel edges can be created, which is not possible in the commonly fabricated multilayer AA'-h-BN. Given that these pores can be manufactured in a well-controlled fashion using an electron beam, it is important to understand the interactions of pores in neighboring layers. We find that in certain configurations, the edges of the neighboring pores remain open and retain their magnetism, and in others, they form interlayer bonds. We present a comprehensive report on these configurations for small nanopores. We find that at low temperatures, these pores have near degenerate magnetic configurations, and may be utilized in magnetoresistance and spintronics applications. In the process of forming larger multilayer nanopores, interlayer bonds can form, reducing the magnetization. Yet, unbonded parallel multilayer edges remain available at all sizes. Understanding these pores is also helpful in a multitude of applications such as DNA sequencing and quantum emission.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enabling Parallel Execution of System-level Simulations in SAM

This report summarizes the recent code updates related to “element ghosting” in SAM to enable the parallel execution of system-level simulations using multiple processors/cores. Unlike typical MOOSE-based applications, for system-level simulations, SAM mostly deals with a collection of discrete small pieces of meshes, and the connection of physics on these meshes are realized by using “connector” types of components/code structures, such as conjugate heat transfer and flow junctions. The required code implementation is to correctly mark the necessary ghost elements for each type of such components/code structures; thus, the lower-level libraries can correctly perform the necessary data transfer between processors (CPUs) when executed in parallel mode. After the code updates, SAM can now run system-level simulations in the parallel mode. The parallel execution capability was then tested with an ABTR input model with 23k DOFs. Significant speedup was demonstrated when the optimal number of CPUs were used in parallel mode. Future systematic studies on parallelization performance using additional test cases covering different physics/scenarios will be needed to provide additional insights into the scalability of SAM parallelization.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Fermilab's controls development with virtual accelerator

Control Systems development is often the last thing considered when designing and building new equipment, e.g. a new detector or superconducting RF LINAC; however when the new equipment is installed, it is the first thing desired to be operational for testing. Due to frequent delays in building new equipment and project deadlines, control system development and testing is often curtailed. A way to alleviate this problem is to simulate the control system, though this will be challenging for complex systems.The Fermilab PIP-II (proton improvement plan - II) project is being constructed at Fermilab to deliver $800\,MeV$ protons of $>1\,MW$ beam power to replace the present LINAC for the remainder of the existing accelerator complex. The new LINAC consists of a warm front end (WFE), 23 superconducting RF cryomodules (of 5 types), and a beam transfer line (BTL) to the existing complex.The accelerator physics group has a parallel project to create a digital twin (DT) of the PIP-II accelerator. We have coupled the EPICS controls to this DT and are developing both the DT and EPICS software in parallel. This will allow us to develop the EPICS software framework, the HMIs, sequences, high level physics applications, and other services for use in a fully functional control system.This presentation will detail the work that we have performed to date and show demonstrations of controlling and monitoring the status of the accelerator, as well as future plans for this work.

Hanlet, Pierrick [Fermilab]↗

PLANC: Parallel Low-rank Approximation with Nonnegativity Constraints

In this work, we consider the problem of low-rank approximation of massive dense nonnegative tensor data, for example, to discover latent patterns in video and imaging applications. As the size of data sets grows, single workstations are hitting bottlenecks in both computation time and available memory. We propose a distributed-memory parallel computing solution to handle massive data sets, loading the input data across the memories of multiple nodes, and performing efficient and scalable parallel algorithms to compute the low-rank approximation. We present a software package called Parallel Low-rank Approximation with Nonnegativity Constraints, which implements our solution and allows for extension in terms of data (dense or sparse, matrices or tensors of any order), algorithm (e.g., from multiplicative updating techniques to alternating direction method of multipliers), and architecture (we exploit GPUs to accelerate the computation in this work). We describe our parallel distributions and algorithms, which are careful to avoid unnecessary communication and computation, show how to extend the software to include new algorithms and/or constraints, and report efficiency and scalability results for both synthetic and real-world data sets.

97 MATHEMATICS AND COMPUTING↗

Parallelized POD-based suboptimal economic model predictive control of a state-constrained Boussinesq approximation

Motivated by an energy efficient building application, we want to optimize a quadratic cost functional subject to the Boussinesq approximation of the Navier-Stokes equations and to bilateral state and control constraints. Since the computation of such an optimal solution is numerically costly, we design an efficient strategy to compute a sub-optimal (but applicationally acceptable) solution with significantly reduced computational effort. We employ an economic Model Predictive Control (MPC) strategy to obtain a feedback control. The MPC sub-problems are based on a linear-quadratic optimal control problem subjected to mixed control and state constraints and a convection-diffusion equation, reduced with proper orthogonal decomposition. Finally, to solve each sub-problem, we apply a primal-dual active set strategy. The method can be fully parallelized, which enables the solution of large problems with real-world parameters.

97 MATHEMATICS AND COMPUTING↗

Hardware-in-the-Loop Investigation of Emissions Challenges in Hybrid Medium- and Heavy-Duty Powertrains Using a Pre-Production Diesel-Electric Parallel Hybrid System With and Without Stop-Start Operation

Hybrid electric powertrains are a growing market in medium- and heavy-duty applications. There is a lack of available information to understand the challenges in the integration of engine platforms into electrified powertrains, such as cold-start, restart, and load-reduction effects on emissions and emission control devices. Results from the Heavy Heavy-Duty Diesel Truck (HHDDT) cycle using a conventional medium-duty diesel engine were compared with those of a parallel hybrid architecture. Oak Ridge National Laboratory in collaboration with the US Department of Energy and Odyne Systems, LLC developed a powertrain in a hardware-in-the-loop environment, integrating the Odyne Systems, LLC medium-duty parallel hybrid system, which was used for the hybrid portion of this study. Experiments under the HHDDT cycle showed increasing improvements in fuel consumption and engine-out emissions with the integration of stop/start, hybrid, and hybrid with stop/start. However, the effects of load reduction and exhaust temperature on the thermal management strategy have shown an increase in fueling in the second part of the HHDDT cycle. Four configurations of medium-duty electrification were studied and contributed to building a unique data set containing combustion, emissions, and system integration data. Each electrification level was compared with the conventional baseline. The calibration of the conventional engine was not altered for this study. Opportunities to tailor the combustion process were identified with the stop/start strategy.

Lerin, Chloe↗

Ensemble Simulation Techniques and Fast Randomized Algorithms

The major goals of the project were to develop and analyze new ensemble simulation techniques, including trajectory stratification and preconditioned MCMC techniques, as well as develop fast numerical linear algebra techniques closely related to ensemble simulation ideas. The trajectory stratification techniques involve simulating in parallel short trajectory fragments of a Markov process confined to a specific region of space‐time and then patching together the statistics gathered to assemble estimates of very general dynamical properties. We have also developed this approach for rare event simulation and extended the techniques to applications requiring a more general framework (such as electronic structure calculations). The preconditioned MCMC techniques involve simulating multiple Markov chains in parallel and then using information from the ensemble to speed the mixing of each individual chain. The fast randomized linear algebra methods are motivated by the diffusion Monte Carlo technique, but are applicable to finding the dominant eigenvalue of (almost) general matrices. For most non‐negative matrices, the schemes result in an error (compared to the power method) that is constant in the dimension of the problem. For more general matrices, we see a very clear sublinear cost trend in computational tests.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

CaTS: Integration of Geant4 and Opticks

CaTS [6]is an advanced example that is part of Geant4 since version 11.0. It demonstrates the use of Opticks to offload the simulation of optical photons to GPUs. Opticks interfaces with the Geant4 toolkit to collect all the necessary information to generate and trace optical photons, re-implements the optical physics processes to be run on the GPU, and automatically translates the Geant4 geometry into a GPU appropriate format. To trace the photons, Opticks uses NVIDIA OptiX®. In this report, we describe CaTS and the integration of Opticks with Geant4. We demonstrate that the generation and tracing of optical photons represents an ideal application to be offloaded to GPUs, fully utilizing the high degree of available parallelism. In a typical liquid argon TPC simulation, a speedup of several hundred times is observed compared to an equivalent simulation using single threaded Geant4.

Wenzel, Hans↗

Multi-mode analysis of surface losses in a superconducting microwave resonator in high magnetic fields

This paper reports on a surface impedance measurement of a bulk metal niobium–titanium superconducting radio frequency (SRF) cavity in a magnetic field (up to 10 T). Here, a novel method is employed to decompose the surface resistance contributions of the cylindrical cavity end caps and walls using measurements from multiple TM cavity modes. The results confirm that quality factor degradation of a NbTi SRF cavity in a high magnetic field is primarily from surfaces perpendicular to the field (the cavity end caps), while parallel surface resistances (the walls) remain relatively constant. This result is encouraging for applications needing high Q cavities in strong magnetic fields, such as the Axion Dark Matter eXperiment because it opens the possibility of hybrid SRF cavity construction to replace conventional copper cavities.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

libEnsemble: A complete Python toolkit for dynamic ensembles of calculations

Almost all science and engineering applications eventually stop scaling: their runtime no longer decreases as available computational resources increase. Therefore, many applications will struggle to efficiently use emerging extreme-scale high-performance, parallel, and distributed systems. libEnsemble is a complete Python toolkit and workflow system for intelligently driving ensembles of experiments or simulations at massive scales. It enables and encourages multidisciplinary design, decision, and inference studies portably running on laptops, clusters, and supercomputers.

97 MATHEMATICS AND COMPUTING↗

Multi-mode Analysis of Surface Losses in a Superconducting Microwave Resonator in High Magnetic Fields

This paper reports on a surface impedance measurement of a niobium titanium superconducting radio frequency (SRF) cavity in a magnetic field (up to $10\,{\rm T}$). A novel method is employed to decompose the surface resistance contributions of the cylindrical cavity end caps and walls using measurements from multiple $TM$ cavity modes. The results confirm that quality factor degradation of a NbTi SRF cavity in a high magnetic field is primarily from surfaces perpendicular to the field (the cavity end caps), while parallel surface resistances (the walls) remain relatively constant. This result is encouraging for applications needing high Q cavities in strong magnetic fields, such as the Axion Dark Matter eXperiment (ADMX), because it opens the possibility of hybrid SRF cavity construction to replace conventional copper cavities.

43 PARTICLE ACCELERATORS↗

Enabling Floating Offshore VAWT Design by Coupling OWENS and OpenFAST

Vertical-axis wind turbines (VAWTs) have a long history, with a wide variety of turbine archetypes that have been designed and tested since the 1970s. While few utility-scale VAWTs currently exist, the placement of the generator near the turbine base could make VAWTs advantageous over tradition horizontal-axis wind turbines for floating offshore wind applications via reduced platform costs and improved scaling potential. However, there are currently few numerical design and analysis tools available for VAWTs. One existing engineering toolset for aero-hydro-servo-elastic simulation of VAWTs is the Offshore Wind ENergy Simulator (OWENS), but its current modeling capability for floating systems is non-standard and not ideal. This article describes how OWENS has been coupled to several OpenFAST modules to update and improve modeling of floating offshore VAWTs and discusses the verification of these new capabilities and features. The results of the coupled OWENS verification test agree well with a parallel OpenFAST simulation, validating the new modeling and simulation capabilities in OWENS for floating VAWT applications. These developments will enable the design and optimization of floating offshore VAWTs in the future.

17 WIND ENERGY↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗