Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Compiler frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Accelerated Charged Particle Tracking with Graph Neural Networks on FPGAs

We develop and study FPGA implementations of algorithms for charged particle tracking based on graph neural networks. The two complementary FPGA designs are based on OpenCL, a framework for writing programs that execute across heterogeneous platforms, and hls4ml, a high-level-synthesis-based compiler for neural network to firmware conversion. We evaluate and compare the resource usage, latency, and tracking performance of our implementations based on a benchmark dataset. We find a considerable speedup over CPU-based execution is possible, potentially enabling such algorithms to be used effectively in future computing workflows and the FPGA-based Level-1 trigger at the CERN Large Hadron Collider.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

OpenACC Unified Programming Environment for Multi-hybrid Acceleration with GPU and FPGA

Accelerated computing in HPC such as with GPU, plays a central role in HPC nowadays. However, in some complicated applications with partially different performance behavior is hard to solve with a single type of accelerator where GPU is not the perfect solution in these cases. We are developing a framework and transpiler allowing the users to program the codes with a single notation of OpenACC to be compiled for multi-hybrid accelerators, named MHOAT (Multi-Hybrid OpenACC Translator) for HPC applications. MHOAT parses the original code with directives to identify the target accelerating devices, currently supporting NVIDIA GPU and Intel FPGA, dispatching these specific partial codes to background compilers such as NVIDIA HPC SDK for GPU and OpenARC research compiler for FPGA, then assembles binaries for the final object with FPGA bitstream file. In this paper, we present the concept, design, implementation, and performance evaluation of a practical astrophysics simulation code where we successfully enhanced the performance up to 10 times faster than the GPU-only solution.

Boku, Taisuke↗

NEML2: A High Performance Library for Constitutive Modeling

NEML2, the New Engineering Material model Library, version 2, is an offshoot of NEML, an earlier material modeling code developed at Argonne National Laboratory. NEML2 extends the key philosophy of its predecessor, i.e., material models are flexible, modular, and can be built from smaller blocks. It also provides modern features that do not exist in the framework of its predecessor such as material model vectorization, automatic differentiation, device-portable just-in-time compilation, operator fusion, lazy tensor evaluation, etc. Moreover, NEML2 can seamlessly integrate with the popular machine learning package PyTorch to take advantage of modern and fast-growing machine learning techniques. In this fiscal year, the development of core library features and capabilities are complete. The purpose of this report is not to serve as a verbatim copy of the software API reference (which is available online at https://reverendbedford.github.io/neml2/). Instead, this report documents the motivation, implementation, design choices, and usage of each core capability as well as their applications in solving practical engineering problems. This report is compiled based on the NEML2 major release 2.0.0.

36 MATERIALS SCIENCE↗

AstraAI v1

AstraAI is an open-source, structure-aware AI coding agent designed for large scientific and DOE-HPC codebases such as AMReX-based applications. Unlike general-purpose coding assistants, AstraAI combines retrieval-augmented generation (RAG) with compiler-level Abstract Syntax Tree (AST) analysis to perform precise, scope-constrained code modifications. It identifies exact function spans, enforces locality of edits, and maintains cross-file invariants, enabling deterministic and build-safe transformations in complex C++/GPU environments. AstraAI is intended for developers working on large, evolving HPC frameworks where correctness, reproducibility, and structural integrity are critical. Typical use cases include modifying physics kernels, updating GPU device lambdas, and performing multi-file refactors without breaking compilation or runtime semantics. Compared to conventional LLM-based coding agents - even those with repository access - AstraAI provides structural guarantees rather than free-form text patches. It minimizes unintended diffs, prevents scope drift, preserves formatting and build stability, and reduces structural hallucinations. By integrating compiler tooling directly into the generation loop, AstraAI transforms AI-assisted coding from probabilistic text editing into deterministic, structure-preserving program transformation suitable for mission-critical scientific software.

Natarajan, Mahesh [Lawrence Berkeley National Labo↗

MLIR loop optimizations for High-Level Synthesis: a case study

High-Level Synthesis (HLS) tools simplify the design of hardware accelerators by automatically generating Verilog/VHDL code starting from a general purpose software programming language. They include a wide range of optimization techniques in the process, most of them performed on a low-level intermediate representation (IR) of the code. Introducing optimizations on a higher level of abstraction could significantly contribute to the automated design process results; for example, polyhedral techniques for the manipulation of loops could have a significant impact on the generated accelerators when applied on a specialized IR. We use loop pipelining as a case study to explore the introduction of compiler-based transformations on top of an existing HLS process. We leverage the Multi-Level Intermediate Representation (MLIR) framework and an external scheduler to implement the required transformations, and couple them with existing HLS tools to evaluate the improvements that loop pipelining brings to the performance of generated accelerators. The proposed approach can be integrated with other high-level transformations on the MLIR representation, combining different techniques to obtain pre-optimized inputs for HLS that do not have to rely on a specific backend tool.

Curzel, Serena↗

Caffeine v0.1.0

Caffeine is the CoArray Fortran Framework of Efficient Interfaces to Network Environments. Caffeine aims to produce a parallel runtime library that will support Fortran compilers with a programming-model-agnostic application binary interface (ABI) to various lower-level communication libraries. The current version of Caffeine uses the GASNet-EX networking middleware, also developed at Berkeley Lab. On many combinations of applications and platforms, GASNet-EX outperforms the widely used Message Passing Interface (MPI). Through GASNet-EX's support for communicating between graphics processing units (GPU), GASNet-EX has features that specifically target the emerging, leading-edge exascale computing platforms.

Rouson, Damian↗

Meta-analysis of biogas upgrading to renewable natural gas through biological CO 2 conversion

Biogas upgrading through CO 2 conversion by hydrogenotrophic methanogenesis is receiving an increasing attention worldwide because of the demand for renewable natural gas. Herein, a holistic and statistical study of the operation conditions, driving forces, performances, and potential implementation of biogas upgrading via biological CO 2 conversion was conducted. Based on a systematic review and meta-analysis of 46 existing publications that were selected from 1475 papers, we have compiled a global dataset of CO 2 bioconversion biogas upgrading, encompassing 308 study cases. Subsequently, we employed a rigorous analytical framework incorporating data processing and mixed effects linear regression analysis to examine the dataset. This analysis revealed a significant positive relationship between the H 2 :CO 2 ratio and the methane percentage in the upgraded biogas when using the study as a random effect. Furthermore, we performed meta-analysis on observations taken when the ratio was close to 4:1 and found that ex situ reactors (91.93% [88.11%, 95.75%]) can perform better than in situ reactors (84.74% [80.69%, 88.80%]). No evidence of differential performance was found based on the present dataset between different temperature regimes or operation modes. Furthermore, those findings establish a database that will contribute to a deeper understanding of the biogas upgrading via biological CO 2 conversion.

Biogas upgrading↗

Boosting RDataFrame performance with transparent bulk event processing

RDataFrame is ROOT’s high-level interface for Python and C++ data analysis. Since it first became available, RDataFrame adoption has grown steadily and it is now poised to be a major component of analysis software pipelines for LHC Run 3 and beyond. Thanks to its design inspired by declarative programming principles, RDataFrame enables the development of highperformance, highly parallel analyses without requiring expert knowledge of multi-threading and I/O: user logic is expressed in terms of self-contained, small computation kernels tied together by a high-level API. This design completely decouples analysis logic from its actual execution, and opens several interesting avenues for workflow optimization. In particular, in this work we explore the benefits of moving internal data processing from an event-by-event to a bulkby-bulk loop. This refactoring dramatically reduces the framework’s runtime overheads; in collaboration with the I/O layer it improves data access patterns; it exposes information that optimizing compilers might use to auto-vectorize the invocation of user-defined computations; finally, while existing user-facing interfaces remain unaffected, it becomes possible to additionally offer interfaces that explicitly expose bulks of events, useful e.g. for the injection of GPU kernels into the analysis workflow. In order to inform similar future R&D, design challenges will be presented, as well as an investigation of the relevant timememory trade-off backed by novel performance benchmarks.

Guiraud, Enrico↗

Invited: Bambu: an Open-Source Research Framework for the High-Level Synthesis of Complex Applications

This paper presents the open-source High-Level Synthesis research framework Bambu. The framework provides an open-source starting point to experiment with new ideas across High-Level Synthesis, high-level verification and debugging, FPGA/ASIC design, design flow space exploration, and parallel hardware accelerator design. The tool accepts as input standard C/C++ specifications and compiler intermediate representations (IRs) coming from the well-known Clang/LLVM and GCC com- pilers. The broad spectrum and flexibility of input formats allow the electronic design automation (EDA) research community to explore and integrate new transformations and optimizations. The easily extendable modular framework already includes many op- timizations and HLS benchmarks. The integration with synthesis and verification backends (commercial and open-source) allows researchers to quickly test any new finding and easily obtain performance and resource usage metrics for a given application. Different FPGA devices are supported from several different vendors: AMD/XILINX, Intel/Altera, Lattice Semiconductor, and NanoXplore. Finally, integration with the OpenRoad open-source end-to-end silicon compiler perfectly fits with the recent push towards open-source EDA.

Ferrandi, Fabrizio↗

U.S. Industry Opportunities for Advanced Nuclear Technology Development (Phase III)

This is the third phase of a work scope which has been focused on providing a framework for and preserving key experimental programs and experiences which are critical to the licensing basis of currently operating nuclear reactors or which could be used in the safety basis for the next generation of reactors. The first phase of this effort, documented in Reference 1, focused on compiling a list of key experimental programs through an international survey of reactor safety professionals working in licensing, design, and academia. The second phase in this effort, documented in Reference 2, focused on creating a searchable database framework in which to organize the results from the international survey and to perform detailed research on several key programs to provide the framework for how to categorize the references which could be located. The purpose of the third phase is to perform a high-level research effort on each experiment/experience and to determine if sufficient data, reports, and results have already been captured to consider the program archived for future generators of nuclear professionals.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Risk-based area of review estimation in overpressured reservoirs to support injection well storage facility permit requirements for CO 2 storage projects

This paper by the Energy & Environmental Research Center presents a workflow and modeling approach for delineating a risk-based area of review (AOR) to support a U.S. Environmental Protection Agency (EPA) Class VI permit for a carbon dioxide (CO 2 ) storage project. The approach combines semianalytical solutions for estimating formation fluid leakage through a hypothetical leaky wellbore with the results of numerical reservoir simulations to define the AOR. The modeling utilizes 1) semianalytical solutions from the peer-reviewed literature for formation fluid leakage through abandoned wellbores by Raven (1990) and Avci (1994), 2) a FORTRAN model compiled and described in Cihan et al. (2011, 2012) called ASLMA (Analytical Solution for Leakage in Multilayered Aquifers), and 3) a computational framework for estimating a risk-based AOR first proposed by Oldenburg et al. (2014, 2016). Therefore, the approach builds upon well-established research and underlying hydrogeological principles that have been upheld for nearly three decades. Moreover, the ASLMA model has been broadly applied to an array of storage projects. The work presented herein extends these earlier works using a custom wrapper written in the software environment, R (R Core Team, 2020), which was developed to perform multiple runs of the ASLMA model using given ranges for one or more input parameters. In addition, the current work simulates the pressure buildup within the storage reservoir in response to CO 2 injection using a compositional simulator to better accommodate the temporospatial evolution of pressure buildup within the storage reservoir that is more accurately modeled using a heterogeneous geologic model and a compositional simulator that accounts for the multiphase interactions. The workflow is demonstrated using a case study for a 180,000-metric-ton-per-year storage project located in the PCOR (Plains CO 2 Reduction) Partnership region. For the storage project evaluated here, under the scenario where the leaky wellbore is open to a saline aquifer (thief zone) between the overlying seal (cap rock) and the underground sources of drinking water (USDW), the risk-based AOR essentially collapses to the areal extent of the CO 2 plume in the storage reservoir because the pressure buildup in the storage reservoir beyond the CO 2 plume is insufficient to drive formation fluids up a hypothetical leaky wellbore into the USDW. However, even under the conservative assumption that the leaky wellbore is not open to a thief zone, beyond the areal extent of the CO 2 plume, the incremental leakage is less than 400 m 3 over 20 years, which represents ~0.0001% or less of the total volume of water contained within the USDW rock volume. As discussed in the text, the threshold criterion for defining the risk-based AOR is site-specific and should be informed by the results of the sensitivity analysis and available site characterization data. The approach outlined in this paper is designed to be protective of USDWs and, therefore, comply with the Safe Drinking Water Act requirements and provisions for the U.S. EPA Class VI Underground Injection Control (UIC) Program (Class VI Rule) and North Dakota Administrative Code Chapter 43-05-01.

54 ENVIRONMENTAL SCIENCES↗

Central Asia Seismic Hazard Assessment (CASHA): A Probabilistic Seismic Hazard Assessment for Kazakhstan, Kyrgyzstan and Tajikistan

Probabilistic seismic hazard assessments (PSHA) underpin the calculation of earthquake loads in most building codes around the world. In Central Asia, the building codes are slowly being updated to incorporate some of the contemporary concepts of seismic hazard representation. There is also a regional desire to coordinate hazard assessments and building code modernization. However, some challenges remain. Expertise in the region related to seismic hazard assessments is still largely compartmentalised, requiring a significant amount of training and capacity building in seismic hazard assessment related topics. In addition, there are vast amounts of seismic data (bulletin and waveforms), both from analogue and digital eras, that the region’s countries stored but until recently did not use or share among themselves or with the broader seismological community around the world. Finally, after the collapse of the Soviet Union in the 1990s, many of the countries’ seismic networks suffered a major setback with the lack of attention and budget to update existing equipment and installation of new instruments. In order to address these issues, the United States Department of Energy through Lawrence Livermore National Laboratory (LLNL) initiated a project in 2016 to engage and train local scientists in Central Asia to install new equipment, to enhance the quality of seismic monitoring and reporting, to improve and harmonise the regional earthquake catalogue, and to conduct national probabilistic seismic hazard assessments using the new and improved datasets. To achieve the seismic hazard assessment related goals, a series of workshops were held in Almaty, Kazakhstan; Bishkek, Kyrgyzstan; and Dushanbe, Tajikistan from 2016 until 2020. During the time that the COVID-19 pandemic restricted travel, workshops continued online (22 online workshops were hosted in two years). Finally, in May 2022, an in-person workshop in Istanbul, Turkey brought together all project participants along with civil engineers engaged with building code activities in their respective countries, providing a platform to discuss the implementation of the hazard models into updates of building codes in each country, as well as to discuss model parameters, sensitivity analyses and model results in terms of hazard maps, uniform hazard spectra and hazard deaggregations. The workshops were a combination of lectures and hands-on exercises, and included international participation as well as local scientists and engineers. The workshops served several purposes, including training, coordination of data collection, interactions between local earth scientists and engineers, and brainstorming and knowledge exchange among local and international experts. This report outlines the new earthquake catalogue compilation effort and the PSHA project undertaken in Kyrgyzstan, Tajikistan, and Kazakhstan as part of this initiative. The southern part of this region is tectonically active with moderate to high levels of both shallow crustal seismic activity and occurrence of deeper earthquakes under the Hindu Kush and Pamir mountain ranges. Deeper earthquakes also occur near southwestern Kazakhstan, under the eastern Greater Caucasus and Caspian Sea. Large portions of central and northern Kazakhstan, on the other hand, are in stable continental regions with low levels of seismic activity. This study systematically compiles and improves all available data on local seismicity, active faults, and ground motion attenuation characteristics of the region; and builds a framework to enable a contemporary PSHA to be carried out with the engagement of local scientists. While the project was regional, the seismic hazard assessments are primarily driven by the countries’ own national preferences and understanding of data collection, interpretation, and validation of results.

58 GEOSCIENCES↗

A Pulse Generation Framework with Augmented Program-aware Basis Gates and Criticality Analysis

Near-term intermediate scale quantum (NISQ) de- vices are subject to considerable noise and short coherence time. Consequently, it is critical to minimize circuit execution latency. Traditionally, each basis gate of a transpiled circuit is decoded into a fixed episode of the device control pulses. Recently, people started to investigate merged pulse generation for customized gates through quantum optimal control (QOC). However, existing QOC approaches face the challenges of (i) restricted search space due to prohibitive compilation overhead; (ii) suboptimal end-to-end performance due to aggressive local optimization and falsely introduced dependency among the customized gates; (iii) inadequate adaptivity towards system calibration, which is critical for NISQ devices. In this work, we propose PAQOC, a novel QOC framework that can (i) automatically detect frequently encountered gate patterns in the logical circuit by modeling the problem as a subgraph mining process and reuse these patterns to enable much larger search space exploration (i.e., program aware); (ii) systemically construct customized gate-set based on the impact to the overall program latency (i.e., criticality-aware); and (iii) quickly adapt to system re-calibration thanks to the small-scale pattern-based gate generation (i.e., adaptivity-aware). PAQOC achieves a good tradeoff between circuit performance and compilation time, allowing fully automatic, single stop, ad- hoc customized pulse generation for more efficient execution of user programs on NISQ devices. Evaluations using fifteen applications show that PAQOC can achieve on average 1.95× speedup of the circuit latency and achieve on average 36.7% reduction in compilation overhead. With PAQOC, circuits can run faster with reduced noise, allowing deeper circuits to be tested within the coherence time of present NISQ platforms.

Chen, Yanhao↗

Recommendations for Distributed Energy Resource Access Control

Cybersecurity for internet - connected Distributed Energy Resources (DER) is essential for the safe and reliable operation of the US power system. Many facets of DER cybersecurity are currently being investigated within different standards development organizations, research communities, and industry committees to address this critical need. This report covers DER access control guidance compiled by the Access Controls Subgroup of the SunSpec/Sandia DER Cybersecurity Workgroup. The goal of the group was to create a consensus - based technical framework to minimize the risk of unauthorized access to DER systems. The subgroup set out to define a strict control environment where users are authorized to access DER monitoring and control features through three steps: (a) user is identified using a proof-of-identity, (b) the user is authenticated by a managed database, (c) and the user is authorized for a specific level of access. DER access control also provides accountability and nonrepudiation within the power system control environment that can be used for forensic analysis and attribution in the event of a cyber-attack. This paper covers foundational requirements for a DER access control environment as well as offering a collection of possible policy, model, and mechanism implementation approaches for IEEE 1547-mandated communication protocols.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

SPEL: Software tool for Porting E3SM Land Model with OpenACC in a Function Unit Test Framework

Most high-end computers adopt hybrid architecture, porting a large-scale scientific code onto accelerators is necessary. The paper presents a generic method for porting large-scale scientific code onto accelerators using compiler directives within a modularized function unit test platform. We have implemented the method and designed a software tool (SPEL) to port the E3SM Land Model (ELM) onto the GPUs in the Summit computer. SPEL automatically generates GPU-ready test modules for all ELM functions, such as CanopyFlux, SoilTemperature, and EcosystemDynamics. SPEL breaks the ELM into a collection of standalone unit test programs for easy code verification and further performance improvement. We further optimize several ELM test modules with advanced techniques, including memory reduction, reconstructed parallel loops, and asynchronous GPU kernel launch. We hope our study will inspire new toolkit developments that expedite large-scale scientific code porting with compiler directives.

Schwartz, Peter↗

Enabling Pulse-level Programming, Compilation, and Execution in XACC

Noisy gate-model quantum processing units (QPUs) are currently available from vendors over the cloud, and digital quantum programming approaches exist to run low-depth circuits on physical hardware. These digital representations are ultimately lowered to pulse-level instructions by vendor quantum control systems to affect unitary evolution representative of the submitted digital circuit. Vendors are beginning to open this pulse-level control system to the public via specified interfaces. Robust programming methodologies, software frameworks, and backend simulation technologies for this analog model of quantum computation will prove critical to advancing pulse-level control research and development. Prototypical use cases for this include error mitigation, optimal pulse control, and physics-inspired pulse construction. Here we present an extension to the XACC quantum-classical software framework that enables pulse-level programming for superconducting, gate-model quantum computers, and a novel, general, and extensible pulse-level simulation backend for XACC that scales on classical compute clusters via MPI. Our work enables custom backend Hamiltonian definitions and gate-level compilation to available pulses with a focus on performance and scalability. We end with a demonstration of this capability, and show how to use XACC for pertinent pulse-level programming tasks.

97 MATHEMATICS AND COMPUTING↗

Let Each Quantum Bit Choose Its Basis Gates

Near-term quantum computers are primarily limited by errors in quantum operations (or gates) between two quantum bits (or qubits). A physical machine typically provides a set of basis gates that include primitive 2-qubit (2Q) and 1-qubit (1Q) gates that can be implemented in a given technology. 2Q entangling gates, coupled with some 1Q gates, allow for universal quantum computation. In superconducting technologies, the current state of the art is to implement the same 2Q gate between every pair of qubits (typically an XX-or XY-type gate). This strict hardware uniformity requirement for 2Q gates in a large quantum computer has made scaling up a time and resource-intensive endeavor in the lab. We propose a radical idea – allow the 2Q basis gate(s) to differ between every pair of qubits, selecting the best entangling gates that can be calibrated between given pairs of qubits. This work aims to give quantum scientists the ability to run meaningful algorithms with qubit systems that are not perfectly uniform. Scientists will also be able to use a much broader variety of novel 2Q gates for quantum computing. We develop a theoretical framework for identifying good 2Q basis gates on “nonstandard” Cartan trajectories that deviate from “standard” trajectories like XX. We then introduce practical methods for calibration and compilation with nonstandard 2Q gates, and discuss possible ways to improve the compilation. To demonstrate our methods in a case study, we simulated both standard XY-type trajectories and faster, nonstandard trajectories using an entangling gate architecture with far-detuned transmon qubits. We identify efficient 2Q basis gates on these nonstandard trajectories and use them to compile a number of standard benchmark circuits such as QFT and QAOA. Furthermore, our results demonstrate an 8x improvement over the baseline 2Q gates with respect to speed and coherence-limited gate fidelity.

quantum computing↗

arco (Assembled Resource-Constrained Optimization) [SWR-26-030]

Arco (Assembled Resource-Constrained Optimization) is a memory-smart optimization DSL and solver for LP and MIP problems on constrained hardware. The software is an optimization framework built around a KDL-based domain-specific language and a CLI compiler/solver. You write optimization models in .kdl files, and the arco CLI compiles, validates, inspects, and solves them. Language bindings (Python today, more planned) provide programmatic access to the same engine. Built for harder optimization problems on constrained resources, Arco is intentional about every allocation, careful with stack and heap behavior, and relentless about minimizing memory usage so more systems can run real workloads. Arco is built primarily for internal use within our organization. You are welcome to try it, but we make no guarantees about API stability or robustness at this stage

Sanchez Perez, Pedro Andres [National Laboratory o↗