Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

General-Simulator-Intermediary

This application allows parallel development of simulator screens for the Human System Simulation Laboratory and connection of the backend simulators for various power plants

Lehmer, JacobP↗

Dayflow: CONUS Daily Streamflow Reanalysis, Version 2 (DayflowV2)

The DayflowV2 dataset provides multiple meteorologic forcings driven hourly streamflow information for approximately 2.7 million NHDPlusV2 stream reaches in the conterminous US (CONUS). DaymetV4, Stage-IV, and Analysis of Period of Record for Calibration (AORC) forcings and their corresponding hybrids drive a nationally scalable modeling framework integrating the simulated runoff from the Variable Infiltration Capacity (VIC) model with the Routing Application for Parallel computatIon of Discharge (RAPID) routing model. Streamflow with (Assimilated) and without (Naturalized) streamflow assimilation at US Geological Survey (USGS) streamflow monitoring sites are included in DayflowV2. A comprehensive evaluation of streamflow at 7,526 USGS gauges is performed for both streamflow types. The resulting key evaluation metrics are also included in the Dayflow dataset. The reanalysis data are available for variable periods; 36 years (1980-2015) for DaymetV4 (DayflowV1), 18 years (2002-2019) for Stage-IV and its hybrids, and 40 years (1980-2019) for AORC and its hybrids.

13 HYDRO ENERGY↗

CMIP6-based Multi-model Streamflow Projections over the Conterminous US, Version 1.1

This dataset presents an ensemble of streamflow projections covering the conterminous United States (CONUS), developed to support the SECURE Water Act Section 9505 Assessment for the US Department of Energy (DOE) Water Power Technologies Office (WPTO). Multiple Coupled Models Intercomparison Project phase 6 (CMIP6) Global Climate Models (GCMs) were downscaled using either statistical (DBCCA) or dynamical (RegCM) downscaling methods, based on two meteorological reference datasets (Daymet and Livneh). Subsequently, the downscaled precipitation, temperature, and wind speed data were used to drive two calibrated hydrologic models (VIC and PRMS), with total runoff routed through the Routing Application for Parallel computatIon of Discharge (RAPID) routing model, producing an ensemble of streamflow projections across 2.7 million NHDPlusV2 stream reaches across the CONUS. Each ensemble member covers the 1980-2019 baseline and 2020-2059 near-future periods under the high-end (SSP585) emission scenario. Additionally, using only DBCCA and Daymet, the projections extend to the 2060-2099 far-future period and encompass three additional emission scenarios (SSP370, SSP245, and SSP126). This dataset is designed to support the SECURE Water Act Section 9505 Assessment for the US Department of Energy (DOE) Water Power Technologies Office (WPTO). For further details, refer to Kao et al. (2022), Rastogi et al. (2022), and Ghimire et al. (2023).

13 HYDRO ENERGY↗

THz Plasmonics and Topological Optics of Weyl Semimetals

THz magneto-optical properties of 3D topological Weyl semimetals were investigated with both THz spectroscopy and THz pump-probe measurements. The unique THz effects predicted in these materials may have important applications in THz technology. The electronic band structure was characterized spectroscopically through THz zero-field reflectance and/or cyclotron resonance measurements. The studies included the dynamic chiral pumping and the study of the predicted novel magneto-electric effects arising from the underlying Berry curvature and magneto plasmonic-like effects in the absence of an applied magnetic field. Chiral pumping in the extreme quantum limit was studied on Weyl semimetals to directly probe the chiral N=0 Landau level. Non-linear pump-probe measurements were used to measure the chiral pumping lifetime. Solids with topologically robust electronic states exhibit unusual electronic and optical transport properties that do not exist in other materials. A particularly interesting example is chiral charge pumping, the so-called chiral anomaly, in recently discovered topological Weyl and Dirac semimetals, where simultaneous application of parallel DC electric and magnetic fields creates an imbalance in the number of carriers of opposite topological charge (chirality). In an earlier study we investigated the Weyl metals Na 3 Bi and Cd 3 As 2 . In this grant we followed up with magneto-optical studies of TaAs, another Weyl semimetal. More recently, we have also characterized other Weyl and Dirac systems that have come on line. The Physics community is still looking for the “hydrogen atom” of the Weyl semimetal. CoSi is one promising new system which features Weyl node spacing comparable to the Brillouin zone size. This system may be suitable for study of the predicted chiral plasmons that arise from the Berry curvature in Weyl materials. In other experiments gates can be applied to the samples in order to study the Fermi arc surface states by modulation reflectance spectroscopy at THz frequencies.

36 MATERIALS SCIENCE↗

CMIP6-based Multi-model Streamflow Projections over the Conterminous US

This dataset presents an ensemble of streamflow projections based on the hydroclimate projections dataset supporting the SECURE Water Act Section 9505 Assessment for the US Department of Energy (DOE) Water Power Technologies Office (WPTO). The six-member General Climate Model (GCM) ensemble from the Coupled Models Intercomparison Project phase 6 (CMIP6) downscaled using statistical (i.e., DBCCA) and dynamical (i.e., RegCM) and bias-corrected using two meteorological reference observations (Daymet & Livneh) are driven through two calibrated hydrologic models (VIC & PRMS) to simulate projected future hydrologic responses. This leads to the production of an ensemble of hydroclimate projections, including total runoff (surface runoff and baseflow) for 1980–2019 baseline and 2020–2059 near-term future periods under multiple emission scenarios at 1/24° (~4 km) spatial resolution across the CONUS. The total runoff projections are routed through the Routing Application for Parallel computatIon of Discharge (RAPID) routing model to produce an ensemble of streamflow projections for both periods across 2.7 million NHDPlusV2 stream reaches in the CONUS.

13 HYDRO ENERGY↗

Dayflow: CONUS Daily Streamflow Reanalysis, Version 1 (V1)

Dayflow V1 is a historical streamflow reanalysis dataset reconstructed for a 36-year period (1980-2015). The dataset provides both daily and monthly scale streamflow information at about 2.7 million NHDPlusV2 stream reaches in the conterminous US (CONUS). Dayflow is the result of a nationally scalable modeling framework that integrates the simulated runoff from the Variable Infiltration Capacity (VIC) model with the Routing Application for Parallel computatIon of Discharge (RAPID) routing model. Two types of streamflow products, simulated streamflow with or without assimilation of historic US Geological Survey (USGS) streamflow observations, are provided in Dayflow V1. A comprehensive evaluation at 7,526 USGS National Water Information System (NWIS) gauges is performed for both types of streamflow products. The resulting key evaluation metrics are also included in the Dayflow V1 Dataset.

13 HYDRO ENERGY↗

Exploiting Modern C++ for Portable Parallel Programming in Lattice QCD Applications

The evolution of ISO C++ standards increasingly serves the needs of scientific computing, offering potential benefits for developing portable applications. The recent revisions of C++ programming language, for instance, introduces a suite of algorithms capable of being executed on accelerators. Although this approach may not yield best performance, it can present a viable balance between code productivity and computational efficiency. In this report, we discuss the implementation of the HISQ operator utilizing a range of features from the C++17/20/23 standards and include an assessment of their performance.

Strelchenko, Alexei↗

Improving Performance of M-to-N Processing and Data Redistribution in In Transit Analysis and Visualization

In an in transit setting, a parallel data producer, such as a numerical simulation, runs on one set of ranks M, while a data consumer, such as a parallel visualization application, runs on a different set of ranks N. One of the central challenges in this in transit setting is to determine the mapping of data from the set of M producer ranks to the set of N consumer ranks. This is a challenging problem for several reasons, such as the producer and consumer codes potentially having different scaling characteristics and different data models. The resulting mapping from M to N ranks can have a significant impact on aggregate application performance. In this work, we present an approach for performing this M-to-N mapping in a way that has broad applicability across a diversity of data producer and consumer applications. We evaluate its design and performance with a study that runs at high concurrency on a modern HPC platform. By leveraging design characteristics, which facilitate an “intelligent” mapping from M-to-N, we observe significant performance gains are possible in terms of several different metrics, including time-to-solution and amount of data moved.

Loring, Burlen↗

Improving Performance of M-to-N Processing and Data Redistribution in In Transit Analysis and Visualization

In an in transit setting, a parallel data producer, such as a numerical simulation, runs on one set of ranks M, while a data consumer, such as a parallel visualization application, runs on a different set of ranks N: One of the central challenges in this in transit setting is to determine the mapping of data from the set of M producer ranks to the set of N consumer ranks. This is a challenging problem for several reasons, such as the producer and consumer codes potentially having different scaling characteristics and different data models. The resulting mapping from M to N ranks can have a significant impact on aggregate application performance. In this work, we present an approach for performing this M-to-N mapping in a way that has broad applicability across a diversity of data producer and consumer applications. We evaluate its design and performance with a study that runs at high concurrency on a modern HPC platform. By leveraging design characteristics, which facilitate an ''intelligent'' mapping from M-to-N, we observe significant performance gains are possible in terms of several different metrics, including time-to-solution and amount of data moved.

Loring, Burlen↗

RMACXX: An Efficient High-Level C++ Interface over MPI-3 RMA

Parallel scientific applications can benefit from de- coupling communication and synchronization. One-sided pro- gramming abstractions, which separate communication from syn- chronization, have in fact served as a motivation for partitioned global address space (PGAS) models. However, the use of PGAS models in application codes in a manner that fully exploits the benefit of these programming models requires significant development effort. Meanwhile, a vast majority of scientific codes already use the Message Passing Interface (MPI) and need convenient features to support application-specific one-sided communication scenarios. MPI Remote Memory Access (RMA) can be employed for this purpose. MPI is a low-level API, however, and developing applications with MPI RMA requires programmers to be well versed in its nuances. We present RMACXX, a compact set of C++ bindings to MPI-3 RMA, to ease the use of MPI RMA. Unlike other PGAS models, which may have interoperability issues with MPI, RMACXX is written on top of MPI and uses the same runtime as MPI. The basic functionality of RMACXX adds only a relatively small number of extra instructions (about 20) to the critical communication path. Moreover, RMACXX provides an intuitive API for building a wide variety of scientific applications while enjoying performance matching handwritten MPI-3 RMA codes.

MPI-3 RMA, one-sided communication, PGAS, C++, exp↗

CCM vs. CRM Design Optimization of a Boost-derived Parallel Active Power Decoupler for Microinverter Applications

Single-phase inverter or rectifier systems often make use of an auxiliary active power decoupler (APD) to balance the mismatch between steady DC power and fluctuating AC power. This paper deals with efficiency and size optimization of a parallel boost-type APD circuit for PV microinverter applications. Specifically, design of an eGaNFET-based, 400 W APD circuit, employing planar inductor and operating in either continuous conduction mode (CCM) or critical conduction mode (CRM) is considered. Available design variables including inductance value, inductor core geometry, capacitor voltage, switching frequency, and modulation scheme (CCM vs. CRM) are explored to identify Pareto-optimal configurations, which can achieve low California Energy Commission (CEC) efficiency drop while also reducing the footprint area of the inductor. The theoretical study predicts that the optimal CRM design can achieve 37% reduced inductor size, while operating with similar efficiency drop, compared to the optimal CCM design. Experimental results, obtained using two separate 40 V, 400 W hardware prototypes for CCM and CRM, are presented to verify the analyses.

14 SOLAR ENERGY↗

MOOSE ProbML: Parallelized probabilistic machine learning and uncertainty quantification for computational energy applications

Here, this paper presents the development and demonstration of massively parallel probabilistic machine learning (ML) and uncertainty quantification (UQ) capabilities within the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source computational platform for parallel finite element and finite volume analyses. In addressing the computational expense and uncertainties inherent in complex multiphysics simulations, this paper integrates Gaussian process (GP) variants, active learning, Bayesian inverse UQ, adaptive forward UQ, Bayesian optimization, evolutionary optimization, and Markov chain Monte Carlo (MCMC) within MOOSE. It also elaborates on the interaction among key MOOSE systems — Sampler, MultiApp, Reporter, and Surrogate — in enabling these capabilities. The modularity offered by these systems enables development of a multitude of probabilistic ML and UQ algorithms in MOOSE. Example code demonstrations include parallel active learning and parallel Bayesian inference via active learning. The impact of these developments is illustrated through five applications relevant to computational energy applications: UQ of nuclear fuel fission product release, using parallel active learning Bayesian inference; very rare events analysis in nuclear microreactors using active learning; advanced manufacturing process modeling using multi-output GPs (MOGPs) and dimensionality reduction; fluid flow using deep GPs (DGPs); and tritium transport model parameter optimization for fusion energy, using batch Bayesian optimization. These capabilities are part of the MOOSE framework.

97 - MATHEMATICS AND COMPUTING↗

Controls Status of Fermilab's PIP-II Project

The Fermilab Proton Improvement Project II (PIP-II) is building a new Super Conducting Linear Accelerator (SCL) accelerating protons to 800 MeV for injection into the rest of the FNAL beam complex. Key progress since the last status report given at ICALEPCS includes the adoption of modern DevOps practices with continuous integration and GitOps-based deployments, commissioning of EPICS-based systems at the Cryomodule Test Facility, and integration of a Virtual Accelerator framework for application development ahead of installation. In parallel, web-based applications using Dart and Flutter have matured, providing secure, unified access to both EPICS and legacy ACNET data. Data acquisition and timing systems have also evolved. This paper presents the current state of controls, emphasizing these recent developments and outlining upcoming milestones as PIP-II approaches commissioning of its cryoplant in 2026 and the Warm Front End in 2027.

Crisp, D. B. [Fermilab]↗

PETSc/TAO Users Manual V.3.21

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and optimizers that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started ). The library is organized hierarchically, enabling users to employ the abstraction level most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users.

97 MATHEMATICS AND COMPUTING↗

DAAP (Data Analytics Application Profiling)

The Data Analytics Application Profiling (DAAP) utility is a software program comprised of a library and scripts to provide basic instrumenting of parallel scientific software applications while they are running on an HPC cluster, and transmit this data over a secure connection as individual records to aggregators. When used in conjunction with additional open source software components (Telegraf, FluentD, and RabbitMQ), the data collected by DAAP can be transmitted off-cluster to an analytics system of the user’s choice for further analysis.

Shereda, Charles↗

PIPER: Performance Insight for Programmers and Exascale Runtimes (Final Technical Report)

This project concentrated on the development of novel performance tools and analysis techniques for large scale parallel systems and applications, eventually targeting exascale platforms. Activities conducted included work to novel root cause detection approaches and an auto-tuning framework for parallel systems. The work has led to several software solutions, which are available as open source.

97 MATHEMATICS AND COMPUTING↗

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗