Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Performance Evaluation of Different Parallel Programming Models in SCALE-Shift Sequences for Criticality and Shielding Applications [Abstract]

The SCALE code system has been widely used for nuclear criticality safety, reactor physics, radiation shielding, source term generation, and inventory analyses by researchers, industry, and regulatory bodies. Although limited support for shared- and distributed-memory parallel processing was introduced via C++ threading, OpenMP, and MPI, a hybrid parallel programming model with both distributed- and shared-memory parallelism has not been fully supported in the SCALE code system.

Nuclear Criticality Safety Program (NCSP)↗

An Investigation of using FleCSI for Monte Carlo Radiation Transport

This document details the attempt to use FleCSI to provide MPI parallelization and domain decomposition for a Monte Carlo radiation transport application. FleCSI is a framework designed to support multi-physics application development with a focus on task-based parallelism and domain decomposition [1]. FleCSI also supports performance portability via a wrapper around a back-end performance portability layer. The goal of this work is to use FleCSI for domain decomposition and MPI parallelization inside of a Monte Carlo radiation transport application and document any pain points and shortcomings. The following section introduce the nomenclature used in FleCSI and its potential benefit for physics code developers, detail the code used to explore using FleCSI in a Monte Carlo radiation transport solver, and list the issues and concerns discovered during the work.

36 MATERIALS SCIENCE↗

A space-time parallel algorithm with adaptive mesh refinement for computational fluid dynamics

This work describes a space-time parallel algorithm with space-time adaptive mesh refinement (AMR). AMR with subcycling is added to multigrid reduction-in-time (MGRIT) in order to provide solution efficient adaptive grids with a reduction in work performed on coarser grids. This algorithm is achieved by integrating two software libraries: XBraid (Parallel time integration with multigrid. https://computation.llnl.gov/projects/parallel-timeintegration-multigrid) and Chombo (Chombo software package for AMR applications—design document, 2014). The former is a parallel time integration library using multigrid and the latter is a massively parallel structured AMR library. Employing this adaptive space-time parallel algorithm is Chord (Comput Fluids 123:202–217, 2015), a computational fluid dynamics (CFD) application code for solving compressible fluid dynamics problems. For the same solution accuracy, speedups are demonstrated from the use of space-time parallelization over the time-sequential integration on Couette flow and Stokes’ second problem. On a transient Couette flow case, at least a 1.5× speedup is achieved, and with a time periodic problem, a speedup of up to 13.7× over the time-sequential case is obtained. In both cases, the speedup is achieved by adding processors and exploring additional parallelization in time. The numerical experiments show the algorithm is promising for CFD applications that can take advantage of the time parallelism. Future work will focus on improving the parallel performance and providing more tests with complex fluid dynamics to demonstrate the full potential of the algorithm.

97 MATHEMATICS AND COMPUTING↗

CompLaB v1.0: a scalable pore-scale model for flow, biogeochemistry, microbial metabolism, and biofilm dynamics

Abstract. Microbial activity and chemical reactions in porous media depend on the local conditions at the pore scale and can involve complex feedback with fluid flow and mass transport. We present a modeling framework that quantitatively accounts for the interactions between the bio(geo)chemical and physical processes and that can integrate genome-scale microbial metabolic information into a dynamically changing, spatially explicit representation of environmental conditions. The model couples a lattice Boltzmann implementation of Navier–Stokes (flow) and advection–diffusion-reaction (mass conservation) equations. Reaction formulations can include both kinetic rate expressions and flux balance analysis, thereby integrating reactive transport modeling and systems biology. We also show that the use of surrogate models such as neural network representations of in silico cell models can speed up computations significantly, facilitating applications to complex environmental systems. Parallelization enables simulations that resolve heterogeneity at multiple scales, and a cellular automaton module provides additional capabilities to simulate biofilm dynamics. The code thus constitutes a platform suitable for a range of environmental, engineering and – potentially – medical applications, in particular ones that involve the simulation of microbial dynamics.

58 GEOSCIENCES↗

Scalable Deep Learning-Based Microarchitecture Simulation on GPUs

Cycle-accurate microarchitecture simulators are essential tools for designers to architect, estimate, optimize, and manufacture new processors that meet specific design expectations. However, conventional simulators based on discrete-event methods often require an exceedingly long time-to-solution for the simulation of applications and architectures at full complexity and scale. Given the excitement around wielding the machine learning (ML) hammer to tackle various architecture problems, there have been attempts to employ ML to perform architecture simulations, such as Ithemal and SimNet. However, the direct application of existing ML approaches to architecture simulation may be even slower due to overwhelming memory traffic and stringent sequential computation logic. This work proposes the first graphics processing unit (GPU)-based microarchitecture simulator that fully unleashes the potential of GPUs to accelerate state-of-the-art ML-based simulators. First, considering the application traces are loaded from central processing unit (CPU) to GPU for simulation, we introduce various designs to reduce the data movement cost between CPUs and GPUs. Second, we propose a parallel simulation paradigm that partitions the application trace into sub-traces to simulate them in parallel with rigorous error analysis and effective error correction mechanisms. Combined, this scalable GPU-based simulator outperforms by orders of magnitude the traditional CPU-based simulators and the state-of-the-art ML-based simulators, i.e., SimNet and Ithemal.

97 MATHEMATICS AND COMPUTING↗

A Tradeoff Analysis of Series / Parallel Three-Phase Converter Topologies for Wireless Extreme Chargers

In this paper, extreme fast charging (XFC) technology is studied considering the charge rates of 300 kW for wireless power transfer (WPT) applications. Tradeoff analysis of series and parallel connection of three-phase WPT system are presented by comparisons of voltage and current stresses on power electronics active / passive components. In addition, star (Y) / delta (Δ) connection configurations for three-phase wireless power transfer coupling coils are analyzed with series and LCC resonant compensation circuits. The system series and parallel connection controllability is also reviewed considering voltage and current balance techniques with output control. In a conclusion of evaluation analysis, it is revealed that each component of 300 kW wireless charging network must be designed for high fast charging system and the overall system operation need to be strategically planned for high power charging and infrastructure deployment.

Asa, Erdem↗

A Scalable Interior‐Point Gauss–Newton Method for PDE‐Constrained Optimization With Bound Constraints

Here, we present a scalable approach to solve a class of partial differential equation (PDE)‐constrained optimization problems with bound constraints. This approach utilizes a robust full‐space interior‐point (IP)‐Gauss–Newton optimization method. To cope with the poorly‐conditioned IP‐Gauss–Newton saddle‐point linear systems that need to be solved approximately, once per optimization step, we propose two spectrally related preconditioners. These preconditioners leverage the limited informativeness of data in regularized PDE‐constrained optimization problems. A block Gauss–Seidel preconditioner is proposed for the GMRES‐based solution of the IP‐Gauss–Newton linear systems. It is shown, for a large‐class of PDE‐ and bound‐constrained optimization problems, that the spectrum of the block Gauss–Seidel preconditioned IP‐Gauss–Newton matrix is asymptotically independent of discretization and is not impacted by the ill‐conditioning that notoriously plagues interior‐point methods. We exploit symmetry of the IP‐Gauss–Newton linear systems and propose a regularization and log‐barrier Hessian preconditioner for the preconditioned conjugate gradient (PCG)‐based solution of the equivalent IP‐Gauss–Newton–Schur complement linear systems. The eigenvalues of the block Gauss–Seidel preconditioned IP‐Gauss–Newton matrix, that are not equal to one, are identical to the eigenvalues of the regularization and log‐barrier Hessian preconditioned Schur complement matrix. The scalability of the approach is demonstrated on two example problems. The numerical solution of these optimization problems is shown to require a discretization independent number of IP‐Gauss–Newton linear solves. Furthermore, the linear systems are solved in a discretization and IP ill‐conditioning independent number of preconditioned Krylov subspace iterations. The parallel scalability of the preconditioner, achieved via algebraic multigrid component solvers when applicable, and the aforementioned algorithmic scalability permits a parallel scalable means to compute solutions of a large class of PDE‐ and bound‐constrained problems.

PDE-constrained optimization↗

TChem-atm v1.0

SAND2024-11300O TChem-atm is a software library that was developed to solve complex kinetic models for atmospheric chemistry applications. TChem-atm interface employs a hierarchical parallelism design to exploit the massive parallelism available from modern computing platforms. It also supports gas atmospheric chemistry applications, e.g., the energy exascale earth system model. TChem can be used as a box model or coupled with a climate model to compute the time evolution of gas tracer species. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin↗

A Task Based Approach for Co-Scheduling Ensemble Workloads on Heterogeneous Nodes

Scientific workflows consist of multiple, connected applications, with data and results flowing from one to another in a pipeline. Traditionally, such workflows are executed in sequential order, storing intermediate data in storage disks. Co-scheduling application workflows concurrently on the same compute nodes would greatly reduce the cost of moving data to/from storage and allow real-time analysis of intermediate results. Nevertheless, most parallel programming runtimes do not allow seamless integration of various applications in a scientific workflow, in part due to the complexity of managing data and resources. The situation is even more complicated for heterogeneous systems. In this work we extend the Minos Computing Library (MCL) runtime to accelerate pipe-lined and parallel workloads where multiple applications are running in the same system. MCL’s asynchronous task library and runtime dynamically manages resources to allow co-scheduling of multiple processes sharing heterogeneous resources. In addition, we design a custom ex- tension of the Open Compute Language (OpenCL) to enable multiple processes to share device memory. We enable MCL to coordinate these shared buffers to allow for easy, fast data sharing between applications. Using malleable micro-benchmarks and two application workflows that combine scientific simulation and AI-based analysis, we show that our method outperforms traditional approaches.

Index Terms—Parallel systems, Scheduling and Task ↗

Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid

Researchers in parallel and distributed computing (PDC) often resort to simulation because experiments conducted using a simulator can be for arbitrary experimental scenarios, are less resource-, labor-, and time-consuming than their real-world counterparts, and are perfectly repeatable and observable. Many frameworks have been developed to ease the development of PDC simulators, and these frameworks provide different levels of accuracy, scalability, versatility, extensibility, and usability. Further, the SimGrid framework has been used by many PDC researchers to produce a wide range of simulators for over two decades. Its popularity is due to a large emphasis placed on accuracy, scalability, and versatility, and is in spite of shortcomings in terms of extensibility and usability. Although SimGrid provides sensible simulation models for the common case, it was difficult for users to extend these models to meet domain-specific needs. Furthermore, SimGrid only provided relatively low-level simulation abstractions, making the implementation of a simulator of a complex system a labor-intensive undertaking. In this work we describe developments in the last decade that have contributed to vastly improving extensibility and usability, thus lowering or removing entry barriers for users to develop custom SimGrid simulators.

97 MATHEMATICS AND COMPUTING↗

Final Technical Report

Statement of the problem or situation that is being addressed in your application. The DOE and its national laboratories developed the Home Energy Score™ (HES) to encourage homeowners to improve their energy performance, lower costs and to share energy information through the MLS listing, appraisal, and financing channels. While the HES is an instrumental tool, it is currently underutilized and consists of technical, structural and sector barriers which need to be addressed in order to scale and many energy efficiency contractors are understandably overwhelmed by the added time and effort and lack of incentive to sell and deliver deep retrofit projects while simultaneously meeting the DOE HES program requirements; consequently, contractors may decide to forgo participation. Home Energy Rating System (HERS) Raters have the opportunity to play the critical Assessor role in producing a Home Energy Score (HES); this role has immense potential but currently is unfulfilled. Lastly, while utilities are interested in their customer base achieving greater energy efficiency, especially to help offset growing residential loads in states like California that are accelerating electrification, utilities do not have access to the market actors who are on the front line of influence to homeowners or review and approve their permits: HERS Raters, assessors, contractors and building departments. General statement of how this problem is being addressed: ConSol will integrate the Home Energy Score™ (HES) to its State of California, approved home energy rating services (HERS) platform (CHEERS) to develop a single tool for contractors nationwide to assess, record and install recommended cost, energy, and emissions saving measures to the 140 million single-family homes throughout the U.S. and 14 million homes in California (CHEERS+HES). The CHEERS high fidelity energy code permitting data will be integrated with HES for simple, accurate, easy-to-use home energy estimation and analysis and will directly gain access to the retrofit and renovations markets with the same upgraded platform. This innovative project will assist the utilities in supporting existing homes in their jurisdictions with HES and develop measures to improve energy efficiency and reduce emissions. How is this problem being addressed? What is the overall project approach? In effort to expand the Home Energy Score™ (HES) by increasing the use of aggregable home energy asset data, ConSol proposes to integrate the DOE HES via Application Programming Interface (API) to its State of California approved home energy rating services platform (CHEERS). Once the CHEERS platform and HES are integrated (CHEERS+HES), this enhanced platform will be instantly available and actively deployed via Phase 1 pilot to HERS Raters, assessors and contractors in California to market-test the solution, understand the rate of adoption and identify opportunities for improvement prior to scaling nationally. The CHEERS high fidelity energy code permitting data will be integrated with HES for simple, accurate, easy-to-use home energy estimation and analysis and will directly gain access to the retrofit and renovations markets with the same upgraded platform. This innovative project will assist the building industry and homeowners with an easy-to-use assessment if energy and carbon impacts of existing homes, and assist the utilities in supporting existing homes in their jurisdictions with HES to improve energy efficiency and reduce emissions. What is to be done in Phase I? During Phase I of this proposed project, ConSol will (1) design software architecture that links CHEERS to the Home Energy ScoreTM via API, (2) solicit partnership from one or more California utilities for a regional pilot, (3) test the new software with its HERS Raters and contractor network in the partnership utility jurisdiction, (4) launch a pilot version of the newly developed software with HERS Raters and contractors in the utility territory, and (5) explore California’s GoGreen energy efficiency homeowner lending program in parallel with the pilot. Commercial Applications and Other Benefits. Summarize the future applications or public benefits if the project is carried over into Phase II or Phase III and beyond. The CHEERS+HES commercialized product will be ready for national market scale following a successful Phase 1 performance. The CHEERS+HES adoption is estimated to reach a 5% adoption growth rate versus the 110,000 baseline, starting in Year 1 after Phase I completion, and continuing each year. As a direct benefit to the DOE, CHEERS will set a goal of 100,000 Home Energy Score assessments for existing home alterations within the first 10 years following Phase 1 performance. The technical benefits of this proposed project include the harmonized, automated, and seamless integration of the DOE HES into the widely used and market leading California energy registry, CHEERS. The social benefits include the aggregate energy, cost and GHG savings by allowing the broader public streamlined access to the CHEERS+HES measurement and the energy efficiency recommended measures that may result. Key Words: Home Energy ScoreTM (HES); Application Programming Interface (API); Home Energy Rating Services (HERS); HERS Raters; contractors; assessors; existing homes, energy asset data; cost, energy, and emissions saving measures; energy code (Title 24) compliance; document repository; utilities; pilot; newly developed software; energy efficiency; homeowner. Summary for Members of Congress: The DOE Home Energy Score™ (HES) is a tool to encourage homeowners to improve their energy performance, lower costs and share energy information but is underutilized and consists of barriers which need to be addressed in order to scale. In effort to expand the HES, CHEERS, Inc. will integrate the HES to its State of California, approved home energy rating services (HERS) platform (CHEERS) to develop a single tool for contractors nationwide to assess, record and install recommended cost, energy, and emissions saving measures to the 140 million single-family homes throughout the U.S. and 14 million homes in California.

Application Programming Interface (API)↗

A parallel variable population multi-objective optimizer for accelerator beam dynamics optimization

The simultaneous optimization of multiple objective functions is needed in many particle accelerator applications. In this paper, we present a parallel evolution based multi-objective optimizer that uses a variable population from generation to generation and an external storage to save good solutions. Two heuristic optimization methods, one uses the unified differential evolution and the other uses the real-coded genetic algorithm, are included in the optimizer to generate next generation candidate solutions, and are compared in the test examples. Finally, as an application, we applied this optimizer to the beam dynamics design optimization of a photoinjector and attained the optimal front solutions after 200 generations with the unified differential evolution offspring production scheme.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Elucidating the challenges in extracting ultra-slow flame speeds in a closed vessel - A CH 2 F 2 microgravity case study using optical and pressure-rise data

Refrigerants with a low global warming potential (GWP) possess mild flammability. Hence, a fundamental understanding of their combustion characteristics is required to assess their fire-hazardous potential. The laminar burning velocity is one fundamental property to describe these under safety evaluation aspects. However, measuring laminar burning velocities for slow-burning components, such as low-GWP alternatives, is challenging. This is due to influencing effects such as buoyant flame deformation, stretch, radiation, and confinement, especially at the flammability limits. Here, in the present study, we investigated near-limit flames of the representative refrigerant difluoromethane (CH 2 F 2 ) with nitrogen-enriched oxidizer mixtures to elucidate the potentials and limitations at ultra-slow flame speeds of two widely used flame speed measurement methods in the closed vessel configuration: the optical flame speed measurement in the early quasi-isobaric regime and the flame speed determination from the pressure-rise during the later phase of the isochoric combustion process. Experiments were performed under terrestrial- and microgravity, conducted in a specially designed setup suitable for drop tower applications, using both methods in parallel. Recommendations for the extraction of flame speed data are derived from the non-buoyant reference investigations under microgravity and transferred to ultra-slow flames at terrestrial gravity. Near-limit flames under microgravity were found to be significantly affected by stretch effects for the optical method so that flame speed extrapolations to unstretched conditions are prone to errors. Stretch also affects the pressure method’s lower evaluation limit, and errors can be estimated by combining flame speed data from the pressure method and Markstein lengths from optical experiments.

42 ENGINEERING↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

RadAI

A physics-informed neural network for learning the steady-state, one-dimensional radiative transfer equation in plane-parallel geometries for stellar atmosphere applications.

Ristić, Marko [LANL]↗

DeepGRN: prediction of transcription factor binding site across cell-types using attention-based deep neural networks

Abstract Background Due to the complexity of the biological systems, the prediction of the potential DNA binding sites for transcription factors remains a difficult problem in computational biology. Genomic DNA sequences and experimental results from parallel sequencing provide available information about the affinity and accessibility of genome and are commonly used features in binding sites prediction. The attention mechanism in deep learning has shown its capability to learn long-range dependencies from sequential data, such as sentences and voices. Until now, no study has applied this approach in binding site inference from massively parallel sequencing data. The successful applications of attention mechanism in similar input contexts motivate us to build and test new methods that can accurately determine the binding sites of transcription factors. Results In this study, we propose a novel tool (named DeepGRN) for transcription factors binding site prediction based on the combination of two components: single attention module and pairwise attention module. The performance of our methods is evaluated on the ENCODE-DREAM in vivo Transcription Factor Binding Site Prediction Challenge datasets. The results show that DeepGRN achieves higher unified scores in 6 of 13 targets than any of the top four methods in the DREAM challenge. We also demonstrate that the attention weights learned by the model are correlated with potential informative inputs, such as DNase-Seq coverage and motifs, which provide possible explanations for the predictive improvements in DeepGRN. Conclusions DeepGRN can automatically and effectively predict transcription factor binding sites from DNA sequences and DNase-Seq coverage. Furthermore, the visualization techniques we developed for the attention modules help to interpret how critical patterns from different types of input features are recognized by our model.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗