Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “BeAu”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous Systems

Performance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems.

Fujita, Norihisa↗

IRIS: A Portable Runtime System Exploiting Multiple Heterogeneous Programming Systems

Across embedded, mobile, enterprise, and HPC systems, computer architectures are becoming more heterogeneous and complex. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to specialize for each architecture. As we show, all of these approaches critically depend on their runtime system for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive runtime system is essential to increase performance portability and improve user productivity. In this regard, we have designed and implemented IRIS: a portable runtime system exploiting multiple heterogeneous programming systems. IRIS can discover available resources, manage multiple diverse programming systems (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

Kim, Jungwon↗

IRIS-MEMFLOW: Data Flow-Enabled Portable Memory Orchestration in IRIS Runtime for Diverse Heterogeneity

Task-based programming models and execution paradigms provide a means to decompose a computation by expressing it as a graph in which each node represents a specific computation operating on memory objects and the edges define the dependencies in the execution flow. In this execution model, independent nodes in the graph can be executed concurrently in different computing devices, making it suitable for heterogeneous systems in which computing devices with different architectures coexist. However, careful memory orchestration across heterogeneous devices is needed because copies of the same memory object may reside in multiple devices during execution. Manually ensuring such an orchestration is quite challenging. Not only must an application developer guard against race conditions, but they must also optimize data movement between the host and devices because unnecessary data movement significantly impacts performance. To mitigate these challenges, we enhance the IRIS heterogeneous runtime and introduce IRIS-MEMFLOW–a data flow–enabled portable memory abstraction for seamlessly orchestrating memory in diverse heterogeneous computing environments. By using data-flow analysis, IRIS-MEMFLOW guards against race conditions while multiple heterogeneous devices access memory objects. IRIS-MEMFLOW also optimizes data movement between the host and devices without manual intervention. As a result, IRIS provides improved programming productivity, performance, and portability for multidevice heterogeneous executions in high-performance computing and cloud systems that run diverse architectures from different vendors. The efficacy of IRIS-MEMFLOW is evaluated through experiments that show its capability in terms of programming productivity, multidevice heterogeneity, portability, and low overhead versus the state of the art.

Monil, M. A. H. [ORNL] (ORCID:0000000334194037)↗

IRIS: Exploring Performance Scaling of the Intelligent Runtime System and its Dynamic Scheduling Policies

High-Performance Computing is becoming increasingly heterogeneous, relying on a diverse mix of hardware to achieve good performance. Paradoxically, current drivers and frameworks for these devices typically require separate languages and implementations for each vendor. Furthermore, there are few tools and little support to schedule codes between these devices in a truly heterogeneous manner-partly because of this fragmentation between vendors and the languages each supports. To overcome both limitations, the Intelligent Runtime System (IRIS) was developed. It allows a common task abstraction to automatically be shared among contemporary vendors and is run from a single host-side API. At runtime, IRIS queries the host system and registers which frameworks and drivers are available, these determine which kernels can be used by the scheduler-CPUs via OpenMP, Nvidia GPUs (CUDA), AMD GPUs (HIP), and Intel and Xilinx FPGAs with OpenCL. IRIS enables tasks to be scheduled to any heterogeneous device and resolves to the appropriate kernel binary at runtimeit only uses the devices supported by the system on which it is run. IRIS supports single-task and graph-based expressions of dependencies of tasks. Additionally, IRIS features a range of dynamic scheduling policies, allowing complex chains of tasks and interactions to be executed, relieving the programmer/user from considering the system to assign tasks to devices optimally. This paper presents the peak performance attainable by IRIS over a range of systems-each with different numbers and types of accelerator devices, it highlights the flexibility of IRIS since these devices are truly heterogeneous, relying on different backends (drivers, frameworks, and languages) which historically required unique implementations to utilize them. We then use this peak performance as a baseline to compare increasingly complex chains of tasks (with increasingly complex task dependencies) and evaluate how IRIS copes. Finally, we consider the performance of different IRIS scheduling policies on this range of task graphs.

Johnston, Beau↗

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Nitrogen Source Governs Community Carbon Metabolism in a Model Hypersaline Benthic Phototrophic Biofilm

Increasing anthropogenic inputs of fixed nitrogen are leading to greater eutrophication of aquatic environments, but it is unclear how this impacts the flux and fate of carbon in lacustrine and riverine systems. Here, we present evidence that the form of nitrogen governs the partitioning of carbon among members in a genome-sequenced, model phototrophic biofilm of 20 members. Consumption of NO 3 – as the sole nitrogen source unexpectedly resulted in more rapid transfer of carbon to heterotrophs than when NH 4 + was also provided, suggesting alterations in the form of carbon exchanged. The form of nitrogen dramatically impacted net community nitrogen, but not carbon, uptake rates. Furthermore, this alteration in nitrogen form caused very large but focused alterations to community structure, strongly impacting the abundance of only two species within the biofilm and modestly impacting a third member species. Our data suggest that nitrogen metabolism may coordinate coupled carbon-nitrogen biogeochemical cycling in benthic biofilms and, potentially, in phototroph-heterotroph consortia more broadly. It further indicates that the form of nitrogen inputs may significantly impact the contribution of these communities to carbon partitioning across the terrestrial-aquatic interface.

carbon cycling↗

MetFish: a Metabolomics Pipeline for Studying Microbial Communities in Chemically Extreme Environments

Metabolites have essential roles in microbial communities, including as mediators of nutrient and energy exchange, cell-to-cell communication, and antibiosis. However, detecting and quantifying metabolites and other chemicals in samples having extremes in salt or mineral content using liquid chromatography-mass spectrometry (LC-MS)-based methods remains a significant challenge. Here, we report a facile method based on in situ chemical derivatization followed by extraction for analysis of metabolites and other chemicals in hypersaline samples, enabling for the first time direct LC-MS-based exometabolomics analysis in sample matrices containing up to 2 M total dissolved salts. The method, MetFish, is applicable to molecules containing amine, carboxylic acid, carbonyl, or hydroxyl functional groups, and it can be integrated into either targeted or untargeted analysis pipelines. In targeted analyses, MetFish provided limits of quantification as low as 1 nM, broad linear dynamic ranges (up to 5 to 6 orders of magnitude) with excellent linearity, and low median interday reproducibility (e.g., 2.6%). MetFish was successfully applied in targeted and untargeted exometabolomics analyses of microbial consortia, quantifying amino acid dynamics in the exometabolome during community succession; in situ in a native prairie soil, whose exometabolome was isolated using a hypersaline extraction; and in input and produced fluids from a hydraulically fractured well, identifying dramatic changes in the exometabolome over time in the well.

59 BASIC BIOLOGICAL SCIENCES↗

CHARM-SYCL: New Unified Programming Environment for Multiple Accelerator Types

Addressing performance portability across diverse accelerator architectures has emerged as a major challenge in the development of application and programming systems for high-performance computing environments. Although recent programming systems that focus on performance portability have significantly improved productivity in an effort to meet this challenge, the problem becomes notably more complex when compute nodes are equipped with multiple accelerator types—each with unique performance attributes, optimal data layout, and binary formats. To navigate the intricacies of multi-accelerator programming, we propose CHARM-SYCL as an extension of our CHARM multi-accelerator execution environment [27]. This environment will combine our SYCL-based performance-portability programming front end with a back end for extremely heterogeneous architectures as implemented with the IRIS runtime from Oak Ridge National Laboratory. Our preliminary evaluation indicates potential productivity boost and reasonable performance compared to vendor-specific programming system and runtimes.

Fujita, Norihisa↗

IRIS-MASH: Efficient Multi-device Asynchronous Multi-Stream Heterogeneous Computing

In the rapidly evolving field of high-performance computing (HPC), effectively leveraging heterogeneous devices through asynchronous task programming is paramount. This paper presents a robust asynchronous task programming model tailored for a multi-device, multi-stream execution environment that incorporates a diverse array of heterogeneous computing units, including GPUs from various vendors and other accelerators. Current state-of-the-art task programming models provide methodologies to support asynchronous task executions, but they typically handle homogeneous devices using native programming languages, while support for heterogeneous devices is limited to frameworks like OpenCL. This gap presents significant challenges in abstracting heterogeneous devices to harness their true asynchronous capabilities effectively using their native programming languages. By implementing asynchronous task execution, our model significantly boosts the performance of tiled algorithm task graphs through overlapping data transfers with computation and enabling the simultaneous execution of multiple kernels. We integrate this approach into a heterogeneous Intelligent Runtime System (IRIS) and assess its performance using a suite of tiled algorithm benchmarks from the heterogeneous math kernels library (MatRIS) based on IRIS. Experimental results demonstrate a performance improvement ranging from 1.6 × to 2 × over IRIS without asynchronous support, and a notable 22% performance enhancement compared to established runtime systems such as StarPU and PaRSEC. This approach significantly improves computation efficiency of HPC workflows and provides a solid base for future exploration and development in the area of asynchronous task programming in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Fermilab Phoebus

The Fermilab Phoebus code contains FNAL site-specific code to customize the Control Systems Studio Phoebus application. Phoebus is a tool used to create and present graphical displays based on current real-time and historic values of sensors, gauges, alarms, and many other devices in accelerators and physics experiments here and at other laboratories around the world.

Velarde, Mariana González [Fermi National Accelera↗

Flutter

The following repositories are a set of libraries that are needed for applications that are being developed for users of the Accelerator control system. The programming language used is Dart and for the user interface use the Flutter framework. URL for code repositories: - https://github.com/fermi-ad/flutter-controls-core - https://github.com/fermi-ad/flutter-controls-plotting - https://github.com/fermi-ad/flutter-controls-auth - https://github.com/fermi-ad/flutter-gql-acsys - https://github.com/fermi-ad/flutter-gql-faas - https://github.com/fermi-ad/dart-gql-acsys - https://github.com/fermi-ad/dart-explicit-imports - https://github.com/fermi-ad/dart-gql-faas - https://github.com/fermi-ad/design-system

Neswold, Rich [Fermi National Accelerator Laborato↗

fermi-ad/flutter-template

This code is a collection of repositories for code that will help modernize the AD controls system. This is a collection of backend services for applications, libraries with common code, and templates to help initialize applications. The individual repositories are: 1. https://github.com/fermi-ad/flutter-template 2. https://github.com/fermi-ad/grpc-alarms-db 3. https://github.com/fermi-ad/alarms-actions-synchronizer 4. https://github.com/fermi-ad/aeolisten-async 5. https://github.com/fermi-ad/aeolisten 6. https://github.com/fermi-ad/acorn-alarms-service 7. https://github.com/fermi-ad/ap-python-template 8. https://github.com/fermi-ad/rust-template 9. https://github.com/fermi-ad/acsys-python-container-example 10. https://github.com/fermi-ad/dev-containers 11. https://github.com/fermi-ad/grpc-reflection-catalog 12. https://github.com/fermi-ad/rust-db-lib 13. https://github.com/fermi-ad/ap-python-launcher

Neswold, Rich [Fermi National Accelerator Laborato↗

Managing technical debt across large-scale control systems

This presentation provides a comprehensive overview of technical debt in the context of control systems for large-scale physics facilities. We will explore its various forms, common causes, and potential consequences on system reliability, maintainability, and upgradability. Drawing on experiences from multiple projects, we will present a range of strategies for proactively managing technical debt, including best practices in design, development, testing, and documentation, as well as reactive approaches for identifying and mitigating existing issues. The presentation will also emphasize the importance of a collaborative approach and the use of modern tools, like fault tracking systems, to ensure the long-term health and success of control systems in our field.

Watts, Adam [Fermilab]↗

Artificial Intelligence for improved facilities operation in the FNAL LINAC

The energy consumption in accelerator structures during beam downtimes is a significant fraction of the overall energy budget. Accurate prediction of downtime duration could inform actions to reduce this energy consumption. The LCAPE project started in 2020 and develops artificial intelligence to improve operations in the FNAL control room by reducing the time to identify the cause of a beam outage, improving the reproducibility of labeling it, predicting their duration and forecasting their occurrence. We present our solution for incorporating information from ~2.5k monitored devices in near-real time to distinguish between dozens of different causes of down time. We discuss the performance of different techniques for modeling the state of health of the facility and we compare unsupervised clustering techniques to distinguish between different causes of down time.

43 PARTICLE ACCELERATORS↗

Beam Commissioning and Integrated Test of the PIP-II Injector Test Facility

The PIP-II Injector Test (PIP2IT) facility is a near-complete low energy portion of the Superconducting PIP-II linac driver. PIP2IT comprises the warm front end and the first two PIP-II superconducting cryomodules. PIP2IT is designed to accelerate a 2 mA H⁻ beam to an energy of 20 MeV. The facility serves as a testbed for a number of advanced technologies required to operate PIP-II and provides an opportunity to gain experience with commissioning of the superconducting linac, significantly reducing project technical risks. Some PIP2IT components are contributions from international partners, who also lend their expertise to the accelerator project. The project has been successfully commissioned with the beam in 2021, demonstrating the performance required for the LBNF/DUNE. In this paper, we describe the facility and its critical systems. We discuss our experience with the integrated testing and beam commissioning of PIP2IT, and present commissioning results. This important milestone ushers in a new era at Fermilab of proton beam delivery using superconducting radio-frequency accelerators.

43 PARTICLE ACCELERATORS↗

The L-CAPE Project at FNAL

The controls system at FNAL records data asynchronously from several thousand Linac devices at their respective cadences, ranging from 15Hz down to once per minute. In case of downtimes, current operations are mostly reactive, investigating the cause of an outage and labeling it after the fact. However, as one of the most upstream systems at the FNAL accelerator complex, the Linac’s foreknowledge of an impending downtime as well as its duration could prompt downstream systems to go into standby, potentially leading to energy savings. The goals of the Linac Condition Anomaly Prediction of Emergence (L-CAPE) project that started in late 2020 are (1) to apply data-analytic methods to improve the information that is available to operators in the control room, and (2) to use machine learning to automate the labeling of outage types as they occur and discover patterns in the data that could lead to the prediction of outages. We present an overview of the challenges in dealing with time-series data from 2000+ devices, our approach to developing an ML-based automated outage labeling system, and the status of augmenting operations by identifying the most likely devices predicting an outage.

43 PARTICLE ACCELERATORS↗

Hydrothermal Liquefaction: Path to Sustainable Aviation Fuel

A variety of technologies are currently under development for producing sustainable aviation fuel (SAF), and one of the most promising technologies is the hydrothermal liquefaction (HTL) of low-cost waste feedstocks. Organizations across the globe are conducting research and development on HTL technology at scales varying from the laboratory to demonstrations focused on various aspects of technology development. The identification of more advanced and sustainable solutions to maximize the fuel yield while optimizing fuel properties and achieving cost parity with conventional fuels is the major focus in developing HTL technology. This workshop on the application of HTL to produce SAF, coordinated by the U.S. Department of Energy Bioenergy Technologies Office, Commercial Aviation Alternative Fuels Initiative, and Pacific Northwest National Laboratory (PNNL), was held virtually on November 17–19, 2020. A broad spectrum of experts from industry, academia, national laboratories, and government from across the globe participated in the workshop, contributing their ideas, insights, and perspectives. A series of keynote presentations, plenary presentations, and breakout sessions provided an interdisciplinary framework for sharing information and building collaboration. This document provides an overview of the content discussed in the presentations as well as the breakout session discussion. Diverse stakeholder perspectives were gathered, and this wealth of information collectively provides an update on the current state of the field and identifies the key research opportunities.

09 BIOMASS FUELS↗