Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “runtime”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Characterizing the Impact of GPU Power Management on an Exascale System

As GPU-accelerated high-performance computing (HPC) systems approach exascale performance, controlling energy consumption without compromising throughput is essential. Architectures such as the AMD MI250X-based Frontier supercomputer provide runtime mechanisms like frequency and power capping, enabling energy tuning without modifying application code. Although both target energy reduction, they operate via distinct hardware control paths and influence workloads differently. We present a comprehensive evaluation of these strategies on a leadership-class system using diverse HPC proxy applications representative of production workloads. Our study analyzes performance–energy trade-offs across multiple capping levels, node counts (1 and 32), and application profiles. Results show that frequency capping generally achieves higher energy efficiency and scalability, with gains of up to 13.2% without performance loss, while power capping is more effective for single-node runs or bursty GPU utilization. We also provide practical guidelines to help system administrators and users balance energy efficiency and performance in large-scale scientific workloads.

Costa, Mariana [Universidade Federal do Rio Grande↗

A Scalable Multi-Modal Framework for High-Fidelity Distributed Human Mobility Simulations

The development of data-driven models for human mobility in urban settings requires access to substantial and diverse real-world data. However, existing historical data often presents challenges such as limited volume, variety, and veracity, as well as missing data and privacy preservation concerns. Also, urban mobility modeling is inherently time-variant, complex, and multi-modal, encompassing everything from individual walking and running to private road travel and large-scale public transportation. These challenges call for innovative solutions to overcome data limitations and compute needs to model mobility behaviors accurately. To address these challenges, we propose a distributed, co-simulation-based architecture DURMOSim that integrates real-world data with scalable, high-fidelity simulations, demonstrating distributed co-simulation feasibility with existing mobility models. DURMOSim underpins a modular integration that would enable using any available mobility simulators for greater extensibility and scalability in performing various urban scenarios. In this paper, we present the design, implementation, and performance evaluation of DURMOSim, highlighting its capability to model population-scale mobility patterns. Our initial results show its ability to dynamically synchronize multiple simulation models at runtime with negligible computational overhead. We believe DURMOSim could be a robust tool for advancing urban mobility research and intelligent transportation systems.

Yoginath, Srikanth [ORNL] (ORCID:0000000184236050)↗

Adapting CLUTCH methodology to multigroup TSUNAMI-3D for eigenvalue sensitivity calculations

The sensitivity of the eigenvalue to uncertainties in nuclear data and its evaluation are important for nuclear criticality safety. TSUNAMI-3D sequences within the SCALE code system offer several options to the user community for calculating eigenvalue sensitivity coefficients with multigroup (MG) and continuous energy (CE) 3D transport capabilities. TSUNAMI-3D sequences implement the adjoint-based perturbation theory with MG KENO code, the Contributon Linked eigenvalue sensitivity/Uncertainty estimation via Track length importance CHaracterization (CLUTCH) method with CE KENO code, and the Iterated Fission Probability (IFP) method with CE KENO and Shift codes. Each method has benefits and limitations depending on the problem that is run. The work presented here aims to adapt the CLUTCH method, which enables the Contributon method's mesh-free, memory-efficient approach for calculating adjoint-weighted tallies for sensitivity calculations, to the MG TSUNAMI-3D sequence. This application would eliminate the explicit adjoint KENO calculation, as well as the memory-consuming mesh flux moment tallies required by the conventional MG TSUNAMI-3D. Smaller memory footprints in the CLUTCH methodology and relatively shorter runtimes in MG KENO transport can make MG TSUNAMI-3D a viable method for some complex problems. Moreover, this adaptation allows MG sensitivity calculations with Shift, ORNL's next-generation high-performance Monte Carlo transport code, which currently does not offer any sensitivity capabilities with MG particle transport simulations. Initial implementation of the new MG TSUNAMI-3D sequence and its preliminary results with a selected critical benchmark experiment in the Verified, Archived Library of Inputs and Data (VALID) are presented in this study.

KENO↗

Bindee

Bindee is a clang tool that outputs a simple pybind11 template given a C++ file for efficient generation of C++-Python bindings. Bindee is intended to be a helper tool for minimizing initial user effort and safeguarding against common runtime errors. Bindee relies on two open-source software to produce bindings. Clang's LibTooling enables bindee to traverse a C++ file's AST to pick out bindable variables and functions, or "bindees." PyBind11 is templated, header-only library for generating C++-Python bindings for variables and functions. Operating purely in C++, bindee does not require learning any new API for accomplishing its task. Additionally, picking the correct pybind11 API for a given bindable element is handled without interaction from the user. Any user input is denoted by '@TEXT@' string substitution. Bindee is capable of generating modular bindings for public class methods, public class variables, enumerations, and free functions and variables. Bindings for templates are also supported. For those familiar with pybind11, bindee does not handle trampolines, C++ extensions through lambdas, or custom type casters.

Kimkno, JasonS.↗

Greggd

greg(g)d - Global runtime for eBPF-enabled gathering (w/ gumption) daemon Recently the linux kernel has added support for low-level kernel monitoring and profiling through a in-kernel virtual machine. The tooling around these new features (the extended Berkley Packet Filter or eBPF for short) is not mature and is difficult to use. Benefits from eBPF are especially hard to realize while trying to do large scale deployments and integrate with existing metric analysis stacks. A tool was needed to enable loading and collecting data from eBPF programs on large scale HPC systems. Given the problems above it was obvious we needed some wrapper program to compile, load, and collect data from eBPF programs running in the kernel. This tool needed to be lightweight without a heavy set of dependencies, relatively stable between different kernel versions, and integrate nicely with existing widely used metric collection tools. We wrote a program that wraps the eBPF tooling and sends data to our metric gathering tool. eBPF programs are either compiled using the host compiler stack, or loaded in the kernel directly from an object file. These programs are then attached to the system calls that we want to profile. Whenever these system calls are run, the eBPF program collects information of interest and writes that to memory. Our wrapper program polls these memory locations, reads and formats the data, then sends the information to a local unix socket. Our other monitoring tools are configured to read from that socket and send it to the rest of our metric monitoring stack for analysis.

Voss, Joseph [Oak Ridge National Lab. (ORNL), Oak ↗

Java Software Extensibility Library

The Java Software Extensibility Library provides tools that help Java software developers produce flexible and extensible Java applications. The major components include a Plugin Library and a Polymorphic Map data structure. The Plugin Library provides developers with a Plugin Registry and a standard pattern for designing, versioning, registering, retrieving, and annotating plugins. It also enables advanced capabilities such as loading plugins at runtime, and creating/loading encrypted plugin packages. The Polymorphic Map data structure facilitates the storage and retrieval of objects of varying types within a single data structure; while maintaining type safety. The library also provides a standard design pattern for producing application-specific "Domain Maps" and "Closed Domain Maps".SAND2020-3816 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Fleschute, Christopher↗

HALOS (Heliostat Aimpoint and Layout Optimization Software) [SWR-21-41]

Heliostat Aimpoint and Layout Optimization Software (HALOS) is an open-source software package that allows users to explore solar field layout optimization, aimpoint strategy optimization, and performance characterization of concentrating solar power tower plants. Users interface with the tool through python, and results are reported in time series tables, plots, runtime logs, and flat-file outputs. Users choose from a list of variables such as tower height, receiver capacity, flux limits, design-point irradiance, etc., and specify information about the system using a small collection of flat files. The software can then optimize the specified variables (e.g., aimpoints for each heliostat) to maximize the thermal energy delivered to the receiver while adhering to flux limits. HALOS is implemented to be flexible with respect to flux characterization methods, but includes a direct connection to NREL's SolarPILOT™ software via its python API so that users can utilize high-fidelity flux simulation methods that have already been developed.

Zolan, Alexander↗

Joint Genome Institute Analysis Workflow Service (JAWS) v2.7

The U.S. Department of Energy Joint Genome Institute (JGI) has developed the JGI Analysis Workflow Service (JAWS) as a distributed framework to run computational workflows across diverse high-performance computing (HPC) and cloud environments. JAWS enhances the reusability, scalability, and robustness of scientific workflows by orchestrating data movement, code execution, and results retrieval across multiple DOE facilities. At its core, JAWS integrates the Cromwell workflow engine to run workflows expressed in the Workflow Description Language (WDL), ensuring portability and interoperability. To provide consistent runtime environments, JAWS employs container technologies such as Shifter, Apptainer, and Docker. Workflow tasks are managed via HTCondor on HPC backends, while Globus ensures secure and efficient data transfer between sites. JAWS is deployed as a multi-site workflow manager across national laboratory computing facilities, with dedicated instances supporting community projects such as the National Microbiome Data Collaborative (NMDC) and KBase. This distributed, service-oriented architecture enables users to "write once, run anywhere," providing scalable, production-quality workflow execution.

Kirton, Edward↗

ZFS Interface For Accelerators

As data volume increases in HPC centers, storing data in reasonable amounts of space and time while maintaining data integrity and recoverability becomes ever more difficult. ZFS is a filesystem that provides these features, among many others, that make it attractive for usage in HPC environments. It has the ability to compress data, compute checksums, and compute redundancy codes within the same runtime, rather than compute each separately, without knowledge of the underlying structure of the filesystem. However, measurements have shown that compression on CPUs can be incredibly inefficient and significantly reduce the performance of ZFS. The ZFS Interface for Accelerators (Z.I.A.) was developed to provide an interface to route data to accelerators while being processed by ZFS, so that the ZFS infrastructure and features are maintained, while allowing for ZFS administrators to provide faster implementations of compression, or other features, to their instance of ZFS.

Lee, Jason↗

JobQueue-PG: A Task Queue for Coordinating Varied Tasks Across Multiple HPC Resources and HPC Jobs

The software allows for queueing and dispatch of tasks of small, varied, or uncertain runtimes across multiple HPC jobs, resources, and other computing systems. The software was designed to allow scientists to enqueue, run, and accumulate results from computational experiments in an efficient, manageable manner. For example, the software can be used to enqueue many small computational experiments and run them using several long-running multi-node HPC jobs that may or may not run simultaneously.

Tripp, Charles↗

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

Autonomous MultiScale Library

AMSLib provides infrastructure to tightly couple multi-scale physics simulation code with ML surrogate model inference. It provides a wholistic runtime execution paradigm to supports uncertainty quantification, surrogate model inference, persistent data storing throughout the execution of a simulation.

Bhatia, Harsh↗

Caffeine v0.1.0

Caffeine is the CoArray Fortran Framework of Efficient Interfaces to Network Environments. Caffeine aims to produce a parallel runtime library that will support Fortran compilers with a programming-model-agnostic application binary interface (ABI) to various lower-level communication libraries. The current version of Caffeine uses the GASNet-EX networking middleware, also developed at Berkeley Lab. On many combinations of applications and platforms, GASNet-EX outperforms the widely used Message Passing Interface (MPI). Through GASNet-EX's support for communicating between graphics processing units (GPU), GASNet-EX has features that specifically target the emerging, leading-edge exascale computing platforms.

Rouson, Damian↗

AGILE GETRS

AGILE GETRS provides an optimized implementation of DETRS that is faster than vendor optimized libraries for small (< 4000) problem sizes. These small problems sizes are of interest to researchers in the Grid Research Integration and Deployment Center (GRID-C) researching real-time power electronics simulations. In addition to our implementation of DETRS, this code provides a test using Google Benchmark which generates matrices of various sizes, fills them with random numbers, runs GETRS, verifies the solution is correct and reports average runtime and other performance metrics.

Hahn, Steven [Oak Ridge National Laboratory (ORNL)↗

Lamellar

asynchronous task based runtime developed using the Rust programming language for HPC

Friese, Ryan↗

pnnl/arena

CFA ARENA is a novel programming model with the support of a runtime targeting asynchronous data-centric execution paradigm in a distributed system. All the machine nodes in ARENA are connected by a ring network to bring the specialized computation to the data rather than the reverse to minimize data movement. The programming interfaces are implemented using C++

Tan, Cheng↗