Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC Trends”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

49 records · Page 3

HPC Analytics of Fused Thermal Plants Data to Optimize Operating Envelope

In this project, ORNL extensively reviewed the ORAP RAM data, and it guided us to develop machine learning models that can predict time to next failures and forecast failure trends, which will be useful for optimizing power plant operation strategies. More specifically, we trained multiple random forest models and evaluated the model accuracy to validate with 10+ years of historical data. In addition, we implemented a web-based graphical user interface system for the models to show how our models can be used in more intuitive ways. This proof of concept allowed exploration of model use with power plant operators in mind. Developed machine learning models will be helpful for managing risks, planning maintenance and operation, ultimately reducing the down time and increasing the service hours. For future work, there are several interesting research topics including but not limited to model enhancement, creating synergy with traditional failure modeling approaches, and data-driven actionable recommendation and suggestions.

20 FOSSIL-FUELED POWER PLANTS↗

Evaluation of a scientific data search infrastructure

The ability to search over large scientific datasets has become crucial to next-generation scientific discoveries as data generated from scientific facilities grow dramatically. In previous work, we developed and deployed ScienceSearch, a search infrastructure for scientific data which uses machine learning to automate metadata creation. Our current deployment is deployed atop a container based platform at a HPC center. In this article, we present an evaluation and discuss our experiences with the ScienceSearch infrastructure. Specifically, we present a performance evaluation of ScienceSearch's infrastructure focusing on scalability trends. The obtained results show that ScienceSearch is able to serve up to 130 queries/min with latency under 3 s. We discuss our infrastructure setup and evaluation results to provide our experiences and a perspective on opportunities and challenges of our search infrastructure.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Unifying Framework to Enable Artificial Intelligence in High-Performance Computing Workflows

Current trends point to a future where large-scale scientific applications are tightly coupled high-performance computing/artificial intelligence (HPC/AI) hybrids. Hence, we urgently need to invest in creating a seamless, scalable framework where HPC and AI/machine learning can efficiently work together and adapt to novel hardware and vendor libraries without starting from scratch every few years. Finally, the current ecosystem and sparsely connected community are not sufficient to tackle these challenges, and we require a breakthrough catalyst for science similar to what PyTorch enabled for AI.

high-performance computing↗

Adaptive job and resource management for the growing quantum cloud

As the popularity of quantum computing continues to grow, efficient quantum machine access over the cloud is critical to both academic and industry researchers across the globe. And as cloud quantum computing demands increase exponentially, the analysis of resource consumption and execution characteristics are key to efficient management of jobs and resources at both the vendor-end as well as the client-end. While the analysis and optimization of job / resource consumption and management are popular in the classical HPC domain, it is severely lacking for more nascent technology like quantum computing.This paper proposes optimized adaptive job scheduling to the quantum cloud taking note of primary characteristics such as queuing times and fidelity trends across machines, as well as other characteristics such as quality of service guarantees and machine calibration constraints. Key components of the proposal include a) a prediction model which predicts fidelity trends across machine based on compiled circuit features such as circuit depth and different forms of errors, as well as b) queuing time prediction for each machine based on execution time estimations. Altogether, this proposal is evaluated on simulated IBM machines across a diverse set of quantum applications and system loading scenarios, and is able to reduce wait times by over 3x and improve fidelity by over 40% on specific usecases, when compared to traditional job schedulers.

97 MATHEMATICS AND COMPUTING↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

EVs@Scale Next-Gen Profiles - Fleet Utilization 2023

As U.S. fleet operators begin transitioning to electric vehicles (EVs), critical questions arise regarding how to manage this shift without disrupting fleet operations or placing undue stress on the electric grid. A major challenge for fleets is maintaining effective operational schedules while accommodating charging requirements, particularly with high-power charging (HPC) infrastructure, which presents grid stability concerns for utilities. Proposed solutions such as charging substations, megawatt charging systems (MCS), and smart charge management systems (SCMS) offer potential pathways forward, but their effectiveness depends on alignment with real-world fleet behavior and operational constraints. This report investigates the charging and utilization behavior of EV and EVSE fleets actively employing HPC technologies by conducting detailed case study analyses based on telematics data. A suite of predefined metrics—covering charging, routing, and other operational behaviors—is developed to evaluate the impact of fleet activities on grid infrastructure and identify opportunities for optimization. Results highlight variations in charging behavior across fleets, such as weekday versus weekend usage, diurnal charging trends, and the role of operational predictability in enabling SCMS effectiveness. While SCMS can help lower costs and improve energy efficiency for fleets with stable schedules, they may be insufficient for fleets with highly variable or long-haul operations, which may require more robust solutions like MCS. Visualization of aggregated hourly energy metrics reveals that while fleet behaviors are diverse, there are common temporal patterns that could inform infrastructure planning and energy management. These insights emphasize the need for fleet-specific charging strategies that minimize grid impact while supporting reliable fleet operations. Additionally, the report underscores the broader economic stakes of electrification, particularly in high-value markets such as freight, where misaligned transitions could stall EV adoption. By examining current EV and EVSE fleet deployments using predetermined standardized metrics, this study offers a foundation for developing technologies and operational frameworks that support scalable, grid-compatible electrification across a variety of fleet types while establishing a baseline understanding of operational behaviors. In doing so, we aim to ensure that future charging solutions reflect actual fleet needs and grid constraints—an essential step toward maintaining operational continuity and achieving a successful transition to electric fleet operations.

Charging↗

Providing Geospatial Intelligence through a Scalable Imagery Pipeline

This chapter describes ORNL’s (Oak Ridge National Laboratory’s) contributions to imagery preprocessing for geospatial intelligence research and development (R&D) in four sections. First, we discuss challenges involved in building an effective imagery preprocessing workflow and the world-class high-performance computing (HPC) resources at ORNL available to process petabytes of imagery data. Second, we highlight how we developed imagery preprocessing tools over three decades while paving the way for our current cutting-edge machine learning and computer vision algorithms that are impacting humanitarian and disaster response efforts. Third, we discuss how PIPE modules work together to turn raw images into analysis-ready datasets. Fourth, we look toward the future and discuss planned advancements to PIPE and computing trends that will affect geospatial intelligence R&D.

Reith, Andrew↗

Brochure for the DOE Office of Science Workshop on Envisioning Frontiers in AI and Computing for Biological Research

In February of 2025 a joint ASCR/BER workshop was held to identify key transformational research directions for understanding biology using artificial intelligence (AI), digital twins and high-performance (HPC) computational methods to facilitate scientific discovery and innovation in support of the Department of Energy mission. AI technologies offer exciting new groundbreaking methods to analyze large volumes of complex biological data, thereby greatly accelerating the ability to understand, predict, and design biological processes for beneficial purposes. In the laboratory, the bridging of AI-enabled automated experimental technologies, HPC and digital twins will provide potent tools for researchers to explore the fundamental nature of biology and harness its inherent metabolic potential for a variety of beneficial purposes. The focus of this workshop was on how high-performance computational methods can impact this objective by exploring digital twins, foundational models, and data-driven approaches with applications to advance automated laboratory experiments, modeling of complex living systems and engineering new functions into plants and microbial systems relevant to DOE mission. Workshop attendees with expertise in plant science, microbiology, mathematics, computer science, and AI assessed the current state of the science, trends, and AI challenges at the interface of plant and microbial systems biology and computational science to identify opportunities for high-impact research. This collaborative effort capitalized on ASCR's advancements in applied mathematics, computer science, and Exascale systems, and BER's expertise in basic genomics-enabled research on DOE relevant plant and microbial systems. The workshop culminated in four key priority research directions to guide future research and development within DOE Office of Science programs.

59 BASIC BIOLOGICAL SCIENCES↗

VTK-m: Visualization for the Exascale Era and Beyond

A recent trend in modern high-performance computing is the increasing use of hybrid architectures, where the vast majority of performance comes from accelerators. Modern accelerators are based on Graphics Processing Units (GPU) that contain many low power cores that in their aggregate provides an extremely high computation rate. Current and future CPU processors are requiring more explicit parallelism as each successive version of the hardware packs in more cores, and technologies like hyperthreading and vector operations require even more parallel processing to leverage each core’s full potential. As an example, the Frontier supercomputer installed at Oak Ridge National Laboratories recently hit a record breaking 1.1 exaflops1 on the LINPACK HPC benchmark [Shoemaker 2022]. The system contains 37632 AMD MI250x GPUs which requires more than half a billion threads to keep the system fully utilized [Khizeran 2022].VTK-m is a toolkit of scientific visualization algorithms for these emerging processor architectures. VTK-m supports the fine-grained concurrency for data analysis and visualization algorithms required to drive extreme scale computing by providing abstract models for data and execution that can be applied to a variety of algorithms across many different processor architectures.

Bolstad, Mark↗

Transportation Hub Infrastructure Expansion: Decision Support Under Uncertainty

The Athena project (www.athena-mobility.org) has worked to investigate the relationship between the Dallas-Fort Worth Airport (DFW) and the greater Dallas area in order to better understand and therefore better inform future decision-making regarding the critical infrastructure that influence mobility between the airport and the city. Through this work, infrastructure related to curbside pickup and drop-off, parking, public transit, and the road network congestion were identified as critical to the operation of the DFW transportation hub. The infrastructure analysis and expansion aspect of the Athena project is focused on the restructuring of the CTA curb as a hierarchical curb and the building or repurposing of parking infrastructure as the interplay between these two areas. Many sources of uncertainty exist that may impact future airport and transportation hub operations, such as passenger volume growth, population demographic changes over time, electric vehicle (EV) adoption rates, and autonomous vehicle (AV) adoption rates. Due to these sources of uncertainty, we have selected for our research a modeling framework that can capture various types of uncertainty and hedge against those uncertainties in the optimization process. We analyze road network and curb congestion, the rise of transportation networking companies, trends in parking usage, existing policies around this infrastructure, airport revenue streams, and other contributing factors to enable infrastructure decision making with less uncertainty. To accomplish this wholistic analysis, we have developed a novel multi-stage, multi-period stochastic optimization model which considers the airport's decisions from 2025-2045 under different possible future macro trajectories and day-to-day variations in operational conditions captured as "annual representation of operations" scenarios with respective probabilities. This model has also been designed to leverage the outputs of various efforts under the Athena project to create a combined decision framework for infrastructure decisions. These various efforts include the route optimization model, the ASPIRES simulation, the mode choice model, and the SUMO traffic simulation. Our computational experiments of this system at scale have resulted in a working version of our infrastructure model which enables the explicit representation and consideration of various sources of uncertainty in the decision process to enable robust, flexible decision-making. This model has been effectively run on NREL's HPC system, Eagle, with large numbers of stochastic scenarios and shows promise as a scalable tool for robust consideration of uncertainties in airport planning. We have tested our model using 30,240 operational circumstances in total, resulting in a problem with more 200 million variables. This model was solved in several different configurations, and a workflow to simulate the performance of the infrastructure model results was developed and deployed. In general, our results indicate that a combination of remote parking, remote curb infrastructure, and dynamic pricing can generate revenue, reduce emissions, accommodate emerging technologies such as AVs and EVs, and manage airport passenger growth over time. We note the success of the proposed strategy depends on the data collection and forecasting abilities of DFW. We have also seen that the AV adoption by TNCs might necessitate larger amounts of remote curb. The results of this work inform strategies for airport infrastructure decision making, as well as demonstrate the value of an adaptable model, but also indicate that there are avenues remaining where further research would be of value.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Using Pilot Jobs and CernVM File System for Simplified Use of Containers and Software Distribution

High Energy Physics (HEP) experiments entail an abundance of computing resources, i.e. sites, to run simulations and analyses by processing data. This requirement is fulfilled by local batch farms, grid sites, private/commercial clouds, and supercomputing centers via High Throughput Computing (HTC). The growing needs of such experiments and resources being prone to trends of heterogeneity make it difficult for physicists to handle these resources directly. Additionally, HEP collaborations heavily rely on data and software releases, typically in the order of tens of gigabytes, while conducting simulations and analyses. Hence, aspects of scalability, reliability, and maintenance become crucial with regards to the distribution of the necessary data and software stack. The GlideinWMS [4] framework helps with the resource management problem by using pilot jobs, aka Glideins, to provision reliable elastic virtual clusters. Glideins are submitted to unreliable heterogeneous resources which are validated and customized by the Glideins to make the worker nodes available for end-user job execution. On the other hand, the CernVM File System (CernVM-FS or CVMFS) [1] helps with data distribution. It is a write-once, read-everywhere filesystem used to deploy scientific software to thousands of nodes on a worldwide distributed computing infrastructure. CVMFS is based on the Hyper Text Transfer Protocol and has been widely used within the particle physics community for (1) distributing experiment software and data such as calibrations, and (2) facilitating containerization by efficiently hosting container images along with providing containerization software, especially Singularity [3] GlideinWMS relies on CVMFS installed locally on the computing resources to satisfy the experiments' software needs. This requires system administrators' effort to install and maintain CVMFS at the sites and limits the use of sites, especially HPC resources, that do not have CVMFS installed. This poster presents a solution, taking advantage of Glideins to provide CVMFS at most sites without the need for a local installation. Doing so expands the pool of resources available for HEP experiments and reduces the effort of system administrators for current resources. Additionally, the proposed solution allows GlideinWMS to also start Singularity [3], a containerization software that can run unprivileged, on sites where neither CVMFS nor Singularity are available, including HPC sites. The benefits provided by this solution are: (1) lower overhead for site administrators in that they have less software to install, (2) an expanded pool of resources that run user jobs with easy access to software and data provided by CVMFS, thus making life easier for the scientists, and (3) improved flexibility to use HPC resources by enabling GlideinWMS pilot jobs to support HPC sites.

Urs, Namratha↗