Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Phlex: Parallel, Hierarchical, and Layered EXecution of data-processing algorithms

Phlex is a computing framework supporting the parallel, hierarchical, and layered execution of data-processing algorithms. It is based on the functional-programming paradigm, thus guaranteeing thread-safety when invoking user-defined pure functions. Phlex allows users to specify arbitrary graph-based hierarchies of data organization, enabling more flexible processing of data as required by the constraints of the program.

Knoepfel, KyleJ. [Fermi National Accelerator Labor↗

PRISMA: PARALLEL REFINEMENT AND INTEGRATION SYSTEM FOR MULTI-AZIMUTHAL ANALYSIS

The Parallel Refinement and Integration System for Multi-azimuthal Analysis (PRISMA, version 1.1.0) is a Python application for processing X-ray diffraction (XRD) image data. PRISMA wraps GSAS-II to perform azimuthally-binned peak refinement, computes per-frame strain and d-spacing from those fits, and provides three PyQt5 graphical interfaces: (1) a Recipe Builder for selecting GSAS-II control (.imctrl) files, optional mask (.immask) files or threshold-ased masking, reference and experiment image sets, peaks, zimuthal range and bin size, and an optional ceria-based auto-calibration; (2) a Batch Processor that uses Dask on local workstations and pure MPI (mpi4py.futures.MPICommExecutor) on HPC to distribute GSAS-II refinement across cores or compute nodes and write results to a 4-dimensional (peaks x frames x azimuths x measurements) Zarr dataset; and (3) a Data Analyzer that renders heatmaps of fit parameters, strain, frame-to-frame deltas, and percent-change-vs-reference, and exports user-defined subsections to CSV or Excel. The peak-refinement algorithm is deterministic. Benchmark on ALCF Crux: a 20,000-image set, single-peak fit in frame mode with 44 azimuthal bins on 128 nodes x 128 workers, 48 seconds total wall time.

Lorenzo Martin, Maria De La Cinta [Argonne Nationa↗

ConnectIt: a framework for static and incremental parallel graph connectivity algorithms

Connected components is a fundamental kernel in graph applications. The fastest existing multicore algorithms for solving graph connectivity are based on some form of edge sampling and/or linking and compressing trees. However, many combinations of these design choices have been left unexplored. In this paper, we design the ConnectIt framework, which provides different sampling strategies as well as various tree linking and compression schemes. ConnectIt enables us to obtain several hundred new variants of connectivity algorithms, most of which extend to computing spanning forest. In addition to static graphs, we also extend ConnectIt to support mixes of insertions and connectivity queries in the concurrent setting. We present an experimental evaluation of ConnectIt on a 72-core machine, which we believe is the most comprehensive evaluation of parallel connectivity algorithms to date. Compared to a collection of state-of-the-art static multicore algorithms, we obtain an average speedup of 12.4x (2.36x average speedup over the fastest existing implementation for each graph). Using ConnectIt, we are able to compute connectivity on the largest publicly-available graph (with over 3.5 billion vertices and 128 billion edges) in under 10 seconds using a 72-core machine, providing a 3.1x speedup over the fastest existing connectivity result for this graph, in any computational setting. For our incremental algorithms, we show that our algorithms can ingest graph updates at up to several billion edges per second. To guide the user in selecting the best variants in ConnectIt for different situations, we provide a detailed analysis of the different strategies. Finally, we show how the techniques in ConnectIt can be used to speed up two important graph applications: approximate minimum spanning forest and SCAN clustering.

Computer Science↗

Nyx: A Massively Parallel AMR Code for Computational Cosmology

Nyx is a highly parallel, adaptive mesh, finite-volume N-body compressible hydrodynamics solver for cosmological simulations. It has been used to simulate different cosmological scenarios with a recent focus on the intergalactic medium and Lyman alpha forest. Together, Nyx, the compressible astrophysical simulation code, Castro, and the low Mach number code MAESTROeX, make up the AMReX-Astrophysics Suite of open-source, adaptive mesh, performance-portable astrophysical simulation codes. Other examples of cosmological simulation research codes include Enzo, Enzo-P/Cello, RAMSES, ART, FLASH, Cholla, as well as Gadget, Gasoline, Arepo, Gizmo, and SWIFT.

79 ASTRONOMY AND ASTROPHYSICS↗

DeepHyper: A Python Package for Massively Parallel Hyperparameter Optimization in Machine Learning

Machine learning models are increasingly applied across scientific disciplines, yet their effectiveness often hinges on heuristic decisions—such as data transformations, training strategies, and model architectures—that are not learned by the models themselves. Automating the selection of these heuristics and analyzing their sensitivity is crucial for building robust and efficient learning workflows. DeepHyper addresses this challenge by democratizing hyperparameter optimization, providing accessible tools to streamline and enhance machine learning workflows from a laptop to the largest supercomputer in the world. Building on top of hyperparameter optimization, it unlocks new capabilities around ensembles of models for improved accuracy and uncertainty quantification. All of these organized around efficient parallel computing.

ensemble↗

ALEGRA Parallel Scaling for Shock in a Heterogeneous Structure

We investigate the strong and weak parallel scaling performance of the ALEGRA multiphysics finite element program when solving a problem involving shock propagation through a heterogeneous material. We determine that ALEGRA scales well over a wide range of problem sizes, cores, and element sizes, and that scaling generally improves as the minimum element size in the mesh increases.

36 MATERIALS SCIENCE↗

Parallelization and Performance Portability in Hydrodynamics Codes

With the eve of Exascale computing, performance and portability are at the forefront of all scientific codes. Adding more cores and more energy to a system is no longer a sustainable way to achieve performance, and extra effort must now be made to improve performance in all areas of code and code development. Using hydrodynamic codes as a basis, this work explores numerous techniques to achieve performance in different ways. Adaptive mesh refinement (AMR) is a necessary technique to improve memory optimization in mesh-based simulations. However it is invasive and conventionally difficult to integrate into existing applications, so we present a new branch of AMR to create a smooth transition to these optimizations, which not only improves performance, but also greatly reduces developer effort. We introduce the concept of this improvement as Phantom-Cell AMR, and assess theoretically the improvements, as well as present an application of its use. Other work included involves and investigation into an efficient data structure that ensures optimal memory layout for cache performance, with a target of making codes performant and portable across all architectures. All of the work targets both performance and portability, not just on CPU hardware, but specifically across GPU architectures. Parallel performance is key to all of the methods presented, but the research makes a great effort to improve the portability of all applications to prepare for current high performance computing systems and those on the horizon.

97 MATHEMATICS AND COMPUTING↗

Xyce Parallel Electronic Simulator Reference Guide (Version 7.2)

This document is a reference guide to the Xyce Parallel Electronic Simulator, and is a companion document to the Xyce Users Guide. The focus of this document is (to the extent possible) exhaustively list device parameters, solver options, parser options, and other usage details of Xyce. This document is not intended to be a tutorial. Users who are new to circuit simulation are better served by the Xyce Users Guide.

42 ENGINEERING↗

Xyce Parallel Electronic Simulator Reference Guide (V.7.1)

This document is a reference guide to the Xyce Parallel Electronic Simulator, and is a companion document to the Xyce Users' Guide. The focus of this document is (to the extent possible) exhaustively list device parameters, solver options, parser options, and other usage details of Xyce. This document is not intended to be a tutorial. Users who are new to circuit simulation are better served by the Xyce Users' Guide.

42 ENGINEERING↗

Massively Parallel Capability in Sierra/SD for Simulation Vibration with Piezoelectrics

Sierra/SD is an engineering structural dynamics code that provides Sandia and other customers a tool to model structural and acoustic physics on large complex physical systems using massively parallel processing. This report provides a detailed overview on Sierra/SD’s most recent physics package: coupled electro-mechanical physics. This capability uses the finite element method to model coupled electro-mechanical physics exhibited by piezoelectric materials. This report provides an applications overview, theory overview, and verification examples demonstrating the electro-mechanical physics modeling capabilities of Sierra/SD.

97 MATHEMATICS AND COMPUTING↗

Xyce™ Parallel Electronic Simulator Reference Guide, Version 7.3

This document is a reference guide to the Xyce Parallel Electronic Simulator, and is a companion document to the Xyce Users' Guide. The focus of this document is (to the extent possible) exhaustively list device parameters, solver options, parser options, and other usage details of Xyce. This document is not intended to be a tutorial. Users who are new to circuit simulation are better served by the Xyce Users' Guide.

42 ENGINEERING↗

Xyce™ Parallel Electronic Simulator Reference Guide (V.7.4)

This document is a reference guide to the Xyce Parallel Electronic Simulator, and is a companion document to the Xyce Users' Guide. The focus of this document is (to the extent possible) exhaustively list device parameters, solver options, parser options, and other usage details of Xyce. This document is not intended to be a tutorial. Users who are new to circuit simulation are better served by the Xyce Users' Guide.

97 MATHEMATICS AND COMPUTING↗

Experience of Migrating a Parallel Graph Coloring Program from CUDA to SYCL

We describe the experience of converting a CUDA implementation of a parallel graph coloring algorithm to SYCL. The goals are for our work to be useful to application and compiler developers by providing a detailed description of migration paths between CUDA and SYCL. We will describe how CUDA functions are mapped to SYCL functions. Evaluating the CUDA and SYCL implementations of the algorithm shows that the performance of SYCL and CUDA kernels are comparable over the test graph set on NVIDIA P100 and V100 GPUs. The SYCL program also allows for performance evaluation with the OpenCL and Level Zero interfaces and power profiling on an Intel GPU computing platform.

97 MATHEMATICS AND COMPUTING↗

Xyce™ Parallel Electronic Simulator Reference Guide, Version 7.5

This document is a reference guide to the Xyce Parallel Electronic Simulator, and is a companion document to the Xyce Users' Guide. The focus of this document is (to the extent possible) exhaustively list device parameters, solver options, parser options, and other usage details of Xyce. This document is not intended to be a tutorial. Users who are new to circuit simulation are better served by the Xyce Users' Guide.

97 MATHEMATICS AND COMPUTING↗

Xyce™ Parallel Electronic Simulator Reference Guide (V.7.6)

This document is a reference guide to the Xyce™ Parallel Electronic Simulator, and is a companion document to the Xyce™ Users' Guide. The focus of this document is (to the extent possible) exhaustively list device parameters, solver options, parser options, and other usage details of Xyce™. This document is not intended to be a tutorial. Users who are new to circuit simulation are better served by the Xyce™ Users' Guide.

97 MATHEMATICS AND COMPUTING↗

Parallel Simulation of Beam Dynamics in Particle Accelerators [Slides]

Particle accelerators are among the most versatile and important tools of scientific discovery. The Nation's accelerators are responsible for a wealth of advances in materials science, chemistry, the biosciences, particle physics, and nuclear physics. They also have important applications to national security, the environment, energy, medicine, and on the quality of people's lives. LANL has a long history of making pioneering contributions to Accelerator Science including key contributions to the field of Computational Accelerator Physics. These include the development of early beam dynamics codes with space charge (such as PARMILA and PARMELA), the development of rf cavity codes and magnet codes (including Poisson and Superfish), and the development and distribution of codes to the accelerator community through the Los Alamos Accelerator Code Group. LANL researchers also helped pioneer the development of massively parallel space-charge codes. In project t22_accelsim we have moved beyond electrostatic models of collective effects (i.e., solving the Poisson equation in the bunch frame) to fully electromagnetic models based on the Lienard-Wiechert formalism. This approach enables the large-scale simulation of radiation production and collective effects in high brightness electron beams. This is highly relevant to LANL given its future goal of developing an X-ray Free Electron Laser (XFEL). It also directly impacts a LANL LDRD project to develop an undulator-based non-invasive beam profile monitor for beams created in laser-plasma accelerator systems.

43 PARTICLE ACCELERATORS↗