Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multithreading”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

49 records · Page 3

nuclear-score-maximization v1.0

This software library presents efficient and multithreaded implementations of matrix low rank approximation via column selection in C++17 code. The algorithms are described in Fornace, Mark, and Michael Lindsey. "Column and row subset selection using nuclear scores: algorithms and theory for Nystro m approximation, CUR decomposition, and graph Laplacian reduction." arXiv preprint arXiv:2407.01698 (2024). The presented methods are by-and-large ver novel, have provable approximation guarantees, multiple use-cases, and exhibit higher quality approximations on a variety of studied examples.

Fornace, Mark↗

Elucidator

SAND2025-01917O Elucidator is a software tool that allows for the storage and interpretation of arbitrary metadata for use in an HPC context. It provides a flexible specification of data members uses facilities to convert buffers of information into the appropriate data types to ensure type safety. The software supports multithreading and database support. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Teves, Joshua↗

Rhythm

Rhythm is a small Python framework that automates Cadence Spectre simulations from the command line. Instead of creating an ADE testbench, you'll write a Rhythm Recipe, a short, readable Python script that contains each step needed to test your circuit from setting up libraries and stimulus waveforms to setting up analyses and reading the results (in Rhythm, these steps are called stages). Need to run more than one simulation? Use a Rhythm Campaign to easily sweep across corners and test conditions with a live dashboard and multithreading support.

Quinn, Adam [Fermi National Accelerator Laboratory↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Cardinal: Seismic and Geoacoustic Array Processing

Data collected via seismic and infrasound array deployments are leveraged in the geosciences to detect and characterize a myriad of natural and anthropogenic sources. These deployments consist of numerous sensors placed in a predetermined configuration to amplify signal strength and improve the efficacy of array processing techniques used to measure signal directionality and waveform coherence. High‐fidelity feature extraction is often predicated on interstation distance as well as the frequency content and wavelength of an incident signal. Numerous array processing softwares analyze data in sequential frequency bands to obtain a more detailed characterization of a signal. However, current algorithms are limited in their ability to determine optimal array configuration for each band. We introduce an open‐source Python code, called Cardinal, to process seismic and infrasound array data in discretized time–frequency space with the option of applying an adaptive array design to determine optimal subarray configuration for each frequency band. To reduce computational time, the array processing step can be run in parallel using multithreading. Furthermore, the software has the capability to aggregate array processing results from different time–frequency pixels to produce separate sets of detections, or families, with added utility via the application of an adaptive semblance threshold, which aids in isolating signals‐of‐interest from coherent background noise. Upon appropriate configuration, Cardinal exhibits the potential to combine distinct seismic and infrasound phases into separate families.

Adaptive Array↗

Subcritical Multiplication Ex-Core Calculations with Shift

This report details the development and testing of Virtual Environment for Reactor Applications (VERA) as applied to full core subcritical multiplication ex-core calculations with Shift. The limiting factor in being able to run these detector calculations with Monte Carlo is the amount of computational memory used by a fully-loaded full core problem. The main development features added to VERA in this effort were enabling source biasing and separable weight windows and integrating multithreading in Shift. With these capabilities, a subcritical multiplication detector response calculation on a full core Watts Bar Nuclear Plant Unit 1 (WBN1) ex-core model was successfully run on a moderately sized computing cluster.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

NEAMS Milestone Report: M2MS–20OR030102 FW–CADIS PWR Ex-Core Analysis with Shift Through VERA

This report presents the work completed for the NEAMS milestone M2MS–20OR030102 titled "FW–CADIS PWR Ex-Core Analysis with Shift through VERA." The work completed for this milestone includes the implementation, integration, and optimization of memory and performance improvement methods in Shift for fully coupled ex-core calculations through Virtual Environment for Reactor Applications (VERA). Fully coupled in this context means the transfer of moderator boron concentration, pin-wise fission source, depleted compositions, temperatures, and moderator densities from MPACT (with COBRA-TF (CTF)) to Shift. The ability to run ex-core calculations with VERA has been enabled and used for several years by Consortium for Advanced Simulation of Light Water Reactors (CASL) partners. However, this implementation was limited and potentially computationally burdensome. This work has enabled the ability to run higher-fidelity ex-core calculations on moderate computing clusters by focusing on multithreading, domain decomposition, and Forward-Weighted CADIS (FW-CADIS) variance reduction. Tests performed on small cores, a small modular reactor (SMR), and CASL progression problems show very promising memory reduction and computational performance. Recommendations for settings when running fully coupled high-fidelity ex-core calculations with VERA are documented. Without these optimization methods, many processors on a compute node would be left unused for the entire ex-core calculation. Therefore, these methods enable the user to better use the resources available and reduce computation time.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Celeritas R&D Report: Accelerating Geant4

Celeritas is a new Monte Carlo (MC) detector simulation code designed for computationally intensive applications on high-performance heterogeneous architectures. In the past two years Celeritas has advanced from prototyping a Graphics Processing Unit (GPU)-based single physics model in infinite medium to implementing a full set of electromagnetic (EM) physics processes in complex geometries. The current release of Celeritas, version 0.4, has incorporated full device-based navigation, an event loop in the presence of magnetic fields, and detector hit scoring. New functionality incorporates a scheduler to offload electromagnetic physics to the GPU within a Geant4-driven simulation, enabling straightforward integration of Celeritas into the high energy physics (HEP) experimental frameworks CMSSW and ATLAS FullSimLight. On the Perlmutter supercomputer, Celeritas performs EM physics between 3× and 18× faster using the machine’s Nvidia GPUs compared to using only CPUs, corresponding to an electrical power efficiency up to a factor of 5. When running a multithreaded Geant4 ATLAS test beam application with full hadronic physics, using Celeritas to accelerate the EM physics results in an overall simulation speedup of 1.7–2.2× on GPU and 1.2× on CPU. In a CMS test application using tt¯ events and the prototype Run 4 configuration, compared to Geant4 CPU, Celeritas with a Nvidia A100 improves overall throughput up to a factor of 2.7× but cannot be efficiently shared with more than 8 cores.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Developing a GUI for the Robotic Test Stand

The introduction of this poster explains the technology behind DUNE’s far and near detectors and how passing neutrinos generate electrons that drift into a wire grid. I then explain how 3 ASICs manage signals received from electron interception. Next, the poster states how COLDATA chips are undergoing quality control by a Robotic Test Stand using a state machine. I further explained how earlier tests were done via a command line script and the necessity to implement a user-friendly Graphical User Interface with new features a command line can’t implement. For the implementation section, tools and methods for implementation are listed such as Python, tkinter, and GitHub as well as how multithreading and queue implementation was necessary for GUI functionality. Then, I elaborated on the GUIs new features. Finally, I explain how the GUI will be distributed across multiple institutions and future changes planned for the GUI. Photos of the RTS, far detector cave, diagram of anode assembly plane, COLDATA chips, set up tab, result tab, and legacy command line interface are shown.

Gutierrez Villanueva, Jaziel [DuPage Coll.]↗

CMSSW Scaling Limits on Many-Core Machines

Today the LHC offline computing relies heavily on CPU resources, despite the interest in compute accelerators, such as GPUs, for the longer term future. The number of cores per CPU socket has continued to increase steadily, reaching the levels of 64 cores (128 threads) with recent AMD EPYC processors, and 128 cores on Ampere Altra Max ARM processors. Over the course of the past decade, the CMS data processing framework, CMSSW, has been transformed from a single-threaded framework into a highly concurrent one. The first multithreaded version was brought into production by the start of the LHC Run 2 in 2015. Since then, the framework's threading efficiency has gradually been improved by adding more levels of concurrency and reducing the amount of serial code paths. The latest addition was support for concurrent Runs. In this work we review the concurrency model of the CMSSW, and measure its scalability with real CMS applications, such as simulation and reconstruction, on mode rn many-core machines. We show metrics such as event processing throughput and application memory usage with and without the contribution of I/O, as I/O has been the major scaling limitation for the CMS applications.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

EXCLUSIVE NEUTRAL PION ELECTROPRODUCTION CROSS SECTION MEASUREMENTSWITHANEUTRALPARTICLE SPECTROMETER

Deep Virtual Compton Scattering (DVCS), the exclusive electron-proton scattering process ep ¿e'p'¿, provides access to generalized parton distributions (GPDs), which correlate information about the longitudinal momentum and transverse spatial structure of quarks inside the nucleon. Experiment E12-13-010 in Hall C at Jefferson Lab was designed to take high-precision measurements of the DVCS cross section over an extended kinematic range using the newly commissioned Neutral Particle Spectrometer (NPS). The NPS features a high-resolution electromagnetic calorimeter and a streaming data acquisition system optimized for operation at high luminosities. This thesis presents the detector and analysis work carried out to support the NPS DVCS program. In particular, it focuses on the hardware design, calibration, and performance of the calorimeter. A development of a waveform reconstruction analysis of the calorimeter signals enabled improved extraction of pulse amplitudes and times. The waveform analysis was also extended to operate in a multithreaded environment, substantially reducing processing time for large datasets. Analysis of exclusive neutral pion electroproduction events in the calorimeter gives a strong validation of the calorimeter’s performance and resolution. Together these developments establish a foundation for future analyses and extraction of the DVCS cross section and its use in constraining the GPDs.

Kerver, Mitchell [Old Dominion Univ., Norfolk, VA ↗

Performance of Julia for High Energy Physics Analyses

We argue that the Julia programming language is a compelling alternative to currently more common implementations in Python and C++ for common data analysis workflows in high energy physics. We compare the speed of implementations of different workflows in Julia with those in Python and C++. Furthermore, our studies show that the Julia implementations are competitive for tasks that are dominated by computational load rather than data access. For work that is dominated by data access, we demonstrate an application with concurrent file reading and parallel data processing.

97 MATHEMATICS AND COMPUTING↗

Analyzing the Performance Trade-Off in Implementing User-Level Threads

User-level threads have been widely adopted as a means of achieving lightweight concurrent execution without the costs of OS-level threads. Nevertheless, the costs of managing user-level threads represent a performance barrier that dictates how fine grained the concurrency exposed by an application can be without incurring significant overheads; this in turn may translate into insufficient parallelism to exploit highly parallel systems. This article is a deep dive into the fundamental costs in implementing user-level threads. We first identify that one of the highest sources of fork-join overheads stems from deviations, events that incur context switching during the execution of a thread and disrupt a run-to-completion execution. We then conduct an in-depth investigation of a wide spectrum of methods with respect to how they handle deviations while covering both parent- and child-first scheduling policies. Our methodology involves a comprehensive instruction- and cache-level analysis of all methods on several modern CPU architectures. Finally, the primary finding of our evaluation is that dynamic promotion methods that assume the absence of deviation and dynamically provide context-switching support offer the best trade-off between performance and capability when the likelihood of deviation is low.

97 MATHEMATICS AND COMPUTING↗