Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Logically Parallel Communication for Fast MPI+Threads Applications

Supercomputing applications are increasingly adopting the MPI+threads programming model over the traditional "MPI everywhere" approach to better handle the disproportionate increase in the number of cores compared with other on-node resources. In practice, however, most applications observe a slower performance with MPI+threads primarily because of poor communication performance. Recent research efforts on MPI libraries address this bottleneck by mapping logically parallel communication, that is, operations that are not subject to MPI's ordering constraints to the underlying network parallelism. Domain scientists, however, typically do not expose such communication independence information because the existing MPI-3.1 standard's semantics can be limiting. Researchers had initially proposed user-visible endpoints to combat this issue, but such a solution requires intrusive changes to the standard (new APIs). The upcoming MPI-4.0 standard, on the other hand, allows applications to relax unneeded semantics and provides them with many opportunities to express logical communication parallelism. Here in this article, we show how MPI+threads applications can achieve high performance with logically parallel communication. Through application case studies, we compare the capabilities of the new MPI-4.0 standard with those of the existing one and user-visible endpoints (upper bound). Logical communication parallelism can boost the overall performance of an application by over 2x.

97 MATHEMATICS AND COMPUTING↗

Lime slurry treatment of soils developing on abandoned coal mine spoil: Linking contaminant transport from the micrometer to pedon-scale

Historical and abandoned coal mine spoil continues to generate acid- and metal(loid)-rich porewaters and represents a geographically large and diffuse non-point source of contamination to local watersheds. A potentially inexpensive approach to treat these materials and soils developing on them is through the application of lime slurries, to neutralize acidity and encourage the (co)precipitation of metal(loids) with Fe(III)-(oxy)hydroxides, and potentially other metal-oxides and/or Ca-bearing phases. Here, the efficacy of this approach was evaluated through parallel field application and laboratory-based flow-through column experiments. The field site is Huff Run sub-watershed 25 located in Tuscarawas County, Ohio and was chosen in part because it was previously classified as one of the most highly AMD-impacted sub-watersheds in the region. Two locations with historical spoil were chosen and suction lysimeters were installed at 25 and 75 cm depth to monitor porewater composition on two sides at the base of each pile. Half of each slope received seven lime slurry treatments from June through October of 2017. A suite of aqueous (ICP-OES, IC, and TOC-L) and solid phase geochemical and mineralogical approaches (quantitative SEM-EDS and synchrotron μ-XRF) were used to determine how composition, texture, morphology, and spatial distribution of mineral coatings differ in pre- and post-lime treated soils, and how that impacts the distribution and transport of trace metal(loid)s. Mine spoil porewater at site 1 was slightly less alkaline (pH ranging from 7.04 to 7.37) than at site 2 (ranging from 7.55 to 7.71), and average electrical conductivity values at site 1 (316–405 μS cm —1 ) were slightly lower than at site 2 (358–464 μS cm —1 ), although differences between the sites were not significant. Porewater pH and electrical conductivity in all lysimeters decreased over the course of the field season but there was no obvious response to lime treatment at either site or any depth. At site 1, both treatment and depth were significant factors affecting Ca, K, Ni, SO 4 2— , and DOC concentrations while only treatment effects were significant for dissolved Al and Cu (p < 0.05). For all soils, there were no trends in metal concentration observed over time although DOC and SO 4 2— decreased over the field season. Pedon-scale changes in metal porewater concentrations in response to treatment were linked to micrometer-scale changes in mineral surface coatings; specifically, higher concentrations of Ca, Fe, Mn, and Zn were observed in the coatings and no changes were observed in Fe redox speciation, whereas total S decreased likely due to oxidation of S in coal fragments. In contrast to the field experiment, the column experiments exhibited a much greater response in effluent composition with respect to lime treatment. The untreated columns had approximately an order of magnitude more H + leached over the course of the experiment (p < 0.001) and resulted in greater Ca, Al, Cu, and DIC leached and less Mn, Zn, and SO 4 2— . Soils treated with the lime slurry in the column experiments exhibited larger and thicker secondary Fe-coatings, including the addition of Fe-sulfates. Despite clear trends in the laboratory-based column experiment where the lime-to-soil ratio was higher, the effects were either muted or undetected in the field pilot project, suggesting that a higher application rate of lime in the field is needed to achieve a similar effect. This work provides evidence that a less alkaline lime slurry could be a practical and inexpensive method of treating coal mine spoil-impacted soils and represents an important step in linking laboratory-based remediation studies to implemented field-based studies.

58 GEOSCIENCES↗

Developing Multiphysics, Integrated, High-Fidelity, Massively Parallel Computational Capabilities for Fusion Applications Using MOOSE

As the need for fusion as a clean, sustainable, and abundant energy source grows internationally, so does the need for multiphysics, computational tools to model, study, and predict the complex interactions between plasma, materials, and engineering processes. These tools have a crucial role to play in solving scientific and engineering challenges and accelerating fusion energy deployment. To address these needs, modeling capabilities should enable massively parallel, multiphysics, fully integrated high-fidelity simulations of fusion systems. Additional attributes, such as being open source and modular while maintaining high software quality assurance standards will maximize impact by ensuring accessibility for all and wide acceptance, rapid expansion and development, as well as reliability, efficiency, and robustness. In this paper, we describe how the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, which has a track record of success in the fission space thanks to the attributes listed above, can be leveraged in the fusion energy field. We highlight key successes of the MOOSE application in the fission space and describe how MOOSE has been and is being applied to fusion applications in the United States---e.g., Tritium Migration Analysis Program, version 8 (TMAP8), MOOSE Fusion Module, Fusion ENergy Integrated multiphys-X (FENIX)---and the United Kingdom---e.g., AURORA, Achlys, Apollo. These efforts aim to establish a suite of tools that can be further extended to accelerate fusion energy deployment.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Bringing OpenCL to Commodity RISC-V CPUs

The importance of open-source hardware has been increasing in recent years with the introduction of the RISC-V Open ISA. This has also accelerated the push for support of the open-source software stack from compiler tools to full-blown operating systems. Parallel computing with today’s Application Programming Interfaces such as OpenCL has proven to be effective at leveraging the parallelism in commodity multi-core processors and programmable parallel accelerators. However, to the best of our knowledge, there is currently no publicly available implementation of OpenCL targeting commodity RISC-V processors that is accessible to the open-source community. Besides opening RISC-V to the existing rich variety of scientific parallel applications, OpenCL also provides access to a unique genre of benchmarks useful in computer architecture research. In this work, we extended an Open-source implementation of OpenCL to target RISC-V CPUs. Our work not only cover commodity multi-core RISC-V processors, but also plethora of low- profile embedded RISC-V CPUs that often do not support atomic instructions or multi-threading.

Tine, Blaise↗

Fixing Amdahl's Law within the Limits of Accelerated Systems: FALLACY

Closeout report for FALLACY project. The performance of Data Model Convergence Initiative (DMC) applications on parallel machines is far below the limit set by Amdahl’s law. Whether the machine is based on many-core, GPUs, FPGAs, or a heterogeneous combination, usually the most significant bottleneck is accessing data from the memory system. Aligning with DMC’s HW/architecture thrust, this project developed a set of memory-centric tools called ‘MemGaze’ that inform the HW/SW stack about an application’s memory behavior, including data access latency and diagnosing poor data layout and data composition. Our approach uses architectural modeling and analysis of workload data accesses.

97 MATHEMATICS AND COMPUTING↗

Layer-Parallel Training of Deep Residual Neural Networks

Residual neural networks (ResNets) are a promising class of deep neural networks that have shown excellent performance for a number of learning tasks, e.g., image classification and recognition. Mathematically, ResNet architectures can be interpreted as forward Euler discretizations of a nonlinear initial value problem whose time-dependent control variables represent the weights of the neural network. Hence, training a ResNet can be cast as an optimal control problem of the associated dynamical system. For similar time-dependent optimal control problems arising in engineering applications, parallel-in-time methods have shown notable improvements in scalability. This paper demonstrates the use of those techniques for efficient and effective training of ResNets. The proposed algorithms replace the classical (sequential) forward and backward propagation through the network layers with a parallel nonlinear multigrid iteration applied to the layer domain. This adds a new dimension of parallelism across layers that is attractive when training very deep networks. From this basic idea, we derive multiple layer-parallel methods. The most efficient version employs a simultaneous optimization approach where updates to the network parameters are based on inexact gradient information in order to speed up the training process. Finally, using numerical examples from supervised classification, we demonstrate that the new approach achieves a training performance similar to that of traditional methods, but enables layer-parallelism and thus provides speedup over layer-serial methods through greater concurrency.

97 MATHEMATICS AND COMPUTING↗

Compact Ferroelectric Programmable Majority Gate for Compute-in-Memory Applications

In this study, a compact and novel ferroelectric (FE) programmable majority gate is proposed and its novel application in Binary Neural Network (BNNs) is investigated. We demonstrate: i) by integrating N metal-ferroelectric-metal (MFM) capacitors on the gate of a transistor (1T-N-MFM structure), a nonvolatile and programmable majority (MAJ) gate that performs MAJ of AND between the gate input and polarization is realized; ii) validation the functionality of our 3-input MAJ of AND gate through comprehensive theoretical and experimental investigations; iii) a compact implementation of 3-input MAJ of XNOR gate that leverages only five of our 3-input MAJ of AND gates connected in parallel; iv) application of MAJ of XNOR gates to replace the XNOR gates and the first layer of the adder tree in the BNNs for up to 21x area saving on top of eliminating the energy-hungry memory accesses due to the compute-in-memory nature.

97 MATHEMATICS AND COMPUTING↗

I/O Bottleneck Detection and Tuning: Connecting the Dots using Interactive Log Analysis

Using parallel file systems efficiently is a tricky problem due to inter-dependencies among multiple layers of I/O software, including high-level I/O libraries (HDF5, netCDF, etc.), MPI-IO, POSIX, and file systems (GPFS, Lustre, etc.). Profiling tools such as Darshan collect traces to help understand the I/O performance behavior. However, there are significant gaps in analyzing the collected traces and then applying tuning options offered by various layers of I/O software. Seeking to connect the dots between I/O bottleneck detection and tuning, we propose DXT Explorer, an interactive log analysis tool. In this paper, we present a case study using our interactive log analysis tool to identify and apply various I/O optimizations. We report an evaluation of performance improvement achieved for four I/O kernels extracted from science applications.

Bez, Jean Luca↗

Nonlinear optical control of chiral charge pumping in a topological Weyl semimetal

Solids with topologically robust electronic states exhibit unusual electronic and optical properties that do not exist in other materials. A particularly interesting example is chiral charge pumping, the so-called chiral anomaly, in recently discovered topological Weyl semimetals, where simultaneous application of parallel DC electric and magnetic fields creates an imbalance in the number of carriers of opposite topological charge (chirality). Here, using time-resolved terahertz measurements on the Weyl semimetal TaAs in a magnetic field, we optically interrogate the chiral anomaly by dynamically pumping the chiral charges and monitoring their subsequent relaxation of the nonequilibrium state. Theory based on Boltzmann transport shows that the observed effects originate from an optical nonlinearity in the chiral charge pumping process. Our measurements reveal that the nonequilibrium chiral excitation relaxation time is much greater than 1 ns. The observation of terahertz-controlled chiral carriers with long coherence times and topological protection suggests the application of Weyl semimetals for quantum optoelectronic technology.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Insights From Dayflow: A Historical Streamflow Reanalysis Dataset for the Conterminous United States

Abstract Reconstructed historical streamflow time series can supplement limited streamflow gauge observations. However, there are common challenges of typical modeling approaches: process‐based hydrologic models can be data/computation‐intensive, and statistics‐based models can be region/stream‐specific. Here we present a nationally scalable modeling framework integrating the simulated runoff from the Variable Infiltration Capacity (VIC) model with the Routing Application for Parallel computatIon of Discharge (RAPID) routing model leveraging high‐performance computing. We demonstrate an efficient method of assimilating streamflow at US Geological Survey (USGS) streamflow monitoring sites using a simple hierarchical approach in the VIC‐RAPID framework. The result is a reconstructed 36‐year (1980–2015) daily and monthly streamflow dataset (Dayflow) at ∼2.7 million NHDPlusV2 stream reaches in the conterminous US (CONUS). We perform a comprehensive evaluation at 7,526 USGS sites and characterize their error statistics. The results demonstrate that 49% of the USGS sites demonstrate Kling–Gupta Efficiency (KGE) > 0.5 and 58% of the sites show percentage bias within ±20% for the daily naturalized streamflow. Streamflow data assimilation across CONUS shows an overall improvement over naturalized streamflow, notably in the western semiarid‐to‐arid regions. Comparison to other national and global streamflow reanalysis datasets such as the National Water Model and Global Reach‐scale A priori Discharge Estimates for SWOT demonstrates improved KGE, reduced bias, and directions for Dayflow improvements. Investigations of error statistics with key hydrologic, hydroclimatic, and geomorphologic basin characteristics reveal region‐specific patterns which may help improve future framework applications. Overall, Dayflow may enable a better understanding of hydrologic conditions in a changing environment, especially in locations currently not represented by streamflow monitoring networks.

54 ENVIRONMENTAL SCIENCES↗

An exploration of online-simulation-driven portfolio scheduling in Workflow Management Systems

Workflow Management Systems used to automate the execution of scientific workflow applications on parallel and distributed computing platforms must make scheduling decisions at runtime. A large number of workflow scheduling algorithms have been proposed in the literature, but often these algorithms are evaluated based on simplifying assumptions that may not hold in practice. Furthermore, published algorithm evaluation and/or comparison results are necessarily only for a subset of all possible scenarios, and thus may not include scenarios relevant to particular use-cases. Consequently, it is difficult for Workflow Management Systems (WMSs) developers to decide which scheduling algorithm should be implemented. To obviate this difficulty, one possible approach is to implement a portfolio of scheduling algorithms and select the most effective algorithm at runtime. One method for performing this selection is to run an online simulation for each algorithm in the portfolio. The algorithm that leads to the best performance, in simulation, is selected for future use. The above simulation-driven portfolio scheduling (SDPS) approach has been proposed in a few parallel and distributed computing contexts. The main objective of this work is to evaluate the feasibility and potential merit of SDPS if implemented in WMSs. Here we perform this evaluation using simulated WMS executions, where the simulations are instantiated from real-world platform and workflow configurations. Our main finding is that SDPS is on par with or outperforms an approach in which a single algorithm is used, where this algorithm is the one that performs best on average across all our experimental scenarios. Furthermore, we find that SDPS remains an attractive proposition even in the presence of high levels of simulation error and for simulators with relatively low levels of sophistication. In many of our experimental scenarios we find that mitigating simulation error at runtime can further improve performance. Finally, we show that simulation overhead can be made sufficiently low for SDPS to be feasible in practice.

97 MATHEMATICS AND COMPUTING↗

Unraveling the 2021 Central Tennessee flood event using a hierarchical multi-model inundation modeling framework

Flood prediction systems need hierarchical atmospheric, hydrologic, and hydraulic models to predict rainfall, runoff, streamflow, and floodplain inundation. The accuracy of such systems depends on the error propagation through the modeling chain, sensitivity to input data, and choice of models. In this study, we used multiple precipitation forcings (hindcast and forecast) to drive hydrologic and hydrodynamic models to analyze the impacts of various drivers on the estimates of flood inundation depth and extent. We implement this framework to unravel the August 2021 extreme flooding event that occurred in Central Tennessee, USA. We used two radar-based quantitative precipitation estimates (STAGE4 and MRMS) as well as quantitative precipitation forecasts (QPF) from the National Weather Service Weather Prediction Center (WPC) to drive a series of models in the hierarchical framework, including the Variable Infiltration Capacity (VIC) land surface model, the Routing Application for Parallel Computation of Discharge (RAPID) river routing model, and the AutoRoute and TRITON inundation models. An evaluation with observed high-water marks demonstrates that the framework can reasonably simulate flood inundation. Despite the complex error propagation mechanism of the modeling chain, we show that inundation estimates are most sensitive to rainfall estimates. Most notably, QPF significantly underestimates flood magnitudes and inundations leading to unanticipated severe flooding for all stakeholders involved in the event. Finally, we discuss the implications of the hydrodynamic modeling framework for real-time flood forecasting.

54 ENVIRONMENTAL SCIENCES↗

The impact of multi-sensor land data assimilation on river discharge estimation

River discharge is one of the most critical renewable water resources. Accurately estimating river discharge with land surface models (LSMs) remains challenging due to the difficulty in estimating land water storages such as snow, soil moisture, and groundwater. While data assimilation (DA) ingesting optical, microwave, and gravity measurements from space can help constrain theses storage states, its impacts on runoff and eventually river discharge are not fully understood. In this study, by taking advantage of recently published land DA results that jointly assimilate eight different combinations of observations from the Moderate Resolution Imaging Spectroradiometer (MODIS), Gravity Recovery and Climate Experiment (GRACE), and Advanced Microwave Scanning Radiometer for EOS (AMSR-E), we quantify to what degree multi-sensor land DA improves the river discharge simulation skills over 40 global river basins, and investigate the complementary strengths of different satellite measurements on river discharge. To be more specific, river discharge is updated by feeding gridded runoff from the eight multi-sensor DA simulations into a vector-based river routing model named the Routing Application for Parallel computatIon of Discharge (RAPID). Our modeling results, including 7-year simulations at 177,458 river reaches globally, are used to study the seasonal to interannual variability of river discharge. It is found that assimilating GRACE has the greatest impact on global runoff patterns, leading to the most pronounced improvements in spatial river discharge in the middle and high latitudes with the R 2 increased by 0.16. The seasonal variation of spatial discharge is most skillful during the boreal summer. However, our evaluation also shows model and DA still struggle to generate reasonable variability and averaged discharge over permafrost regions. Finally, by assessing how different satellites add value to discharge forecasts, this study paves the way for more advanced multi-sensor satellite data assimilation to predict the terrestrial hydrological cycle.

54 ENVIRONMENTAL SCIENCES↗

Qu8its for quantum simulations of lattice quantum chromodynamics

We explore the utility of d = 8 qudits, qu8its, for quantum simulations of the dynamics of 1+1⁢D SU(3) lattice quantum chromodynamics, including a mapping for arbitrary number of flavors and lattice size and a reorganization of the Hamiltonian for efficient time evolution. Recent advances in parallel gate applications, along with the shorter application times of single-qudit operations compared with two-qudit operations, lead to significant projected advantages in quantum simulation fidelities and circuit depths using qu8its rather than qubits. The number of two-qudit entangling gates required for time evolution using qu8its is found to be more than a factor of 5 fewer than for qubits. Here, we anticipate that the developments presented in this work will enable improved quantum simulations to be performed using emerging quantum hardware.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Modeling pre-Exascale AMR Parallel I/O Workloads via Proxy Applications

The present work investigates the modeling of preexascale input/output (I/O) workloads of Adaptive Mesh Refinement (AMR) simulations through a simple proxy application. We collect data from the AMReX Castro framework running on the Summit supercomputer for a wide range of scales and mesh partitions for the hydrodynamic Sedov case as a baseline to provide sufficient coverage to the formulated proxy model. The non-linear analysis data production rates are quantified as a function of a set of input parameters such as output frequency, grid size, number of levels, and the Courant-Friedrichs-Lewy (CFL) condition number for each rank, mesh level and simulation time step. Linear regression is then applied to formulate a simple analytical model which allows to translate AMReX inputs into MACSio proxy I/O application parameters, resulting in a simple “kernel” approximation for data production at each time step. Results show that MACSio can simulate actual AMReX nonlinear “static” I/O workloads to a certain degree of confidence on the Summit supercomputer using the present methodology. The goal is to provide an initial level of understanding of AMR I/O workloads via lightweight proxy applications models to facilitate autotune data management strategies in anticipation of exascale systems.

Godoy, William↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

A memory-driven mapping algorithm for heterogeneous systems

mpibind is a memory-driven algorithm to map parallel hybrid applications to the underlying hardware resources transparently, efficiently, and portably. There are two fundamental aspects of this algorithm. First, unlike existing mappings, its primary design point is the memory system. Compute elements are selected based on the identified memory components and not vice versa. Second, it embodies a global awareness of hybrid programming abstractions as well as heterogeneous devices.

Leon Borja, EdgarA↗