Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Compiler frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Framework to Assess Advanced Reactor Spent Fuel Management Facility Deployment

Previous planning and prioritization for LWR SNF management investigated the risks and uncertainties of deploying facilities such as consolidated interim storage [1, 2, 3, 4]. As part of that work, activities and milestones were collected into success precedence diagrams that charted a path to achieving facility deployment [1]. In that framework, activities are any research, development, design, or decision required to achieve an intermediate goal; milestones are activity endpoints and mark the completion of deliverables. Milestones can be thought of as achievements required to reach the final goal of facility deployment; activities are the means by which milestones are accomplished. In planning, activities and milestones are compiled into comprehensive flow charts that visualize the steps necessary for deployment. This framework has been used to quantify risks, timelines, and costs of deploying SNF management facilities.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Autotuning PolyBench benchmarks with LLVM Clang/Polly loop optimization pragmas using Bayesian optimization

Here, we develop a ytopt autotuning framework that leverages Bayesian optimization to explore the parameter space search and compare four different supervised learning methods within Bayesian optimization and evaluate their effectiveness. We select six of the most complex PolyBench benchmarks and apply the newly developed LLVM Clang/Polly loop optimization pragmas to the benchmarks to optimize them. We then use the autotuning framework to optimize the pragma parameters to improve their performance. The experimental results show that our autotuning approach outperforms the other compiling methods to provide the smallest execution time for the benchmarks syr2k, 3mm, heat-3d, lu, and covariance with two large datasets in 200 code evaluations for effectively searching the parameter spaces with up to 170,368 different configurations. We find that the Floyd-Warshall benchmark did not benefit from autotuning. To cope with this issue, we provide some compiler option solutions to improve the performance. Then we present loop autotuning without a user's knowledge using a simple mctree autotuning framework to further improve the performance of the Floyd-Warshall benchmark. We also extend the ytopt autotuning framework to tune a deep learning application.

79 ASTRONOMY AND ASTROPHYSICS↗

Fugu v.0.1

SAND2021-15052 O Fugu provides a common software framework for designing and prototyping algorithms for spiking neuromorphic hardware and compiling to multiple hardware platforms. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Aimone, James↗

A Backend-agnostic, Quantum-classical Framework for Simulations of Chemistry in C ++

As quantum computing hardware systems continue to advance, the research and development of performant, scalable, and extensible software architectures, languages, models, and compilers is equally as important to bring this novel coprocessing capability to a diverse group of domain computational scientists. For the field of quantum chemistry, applications and frameworks exist for modeling and simulation tasks that scale on heterogeneous classical architectures, and we envision the need for similar frameworks on heterogeneous quantum-classical platforms. Furthermore, we present the XACC system-level quantum computing framework as a platform for prototyping, developing, and deploying quantum-classical software that specifically targets chemistry applications. We review the fundamental design features in XACC, with special attention to its extensibility and modularity for key quantum programming workflow interfaces and provide an overview of the interfaces most relevant to simulations of chemistry. A series of examples demonstrating some of the state-of-the-art chemistry algorithms currently implemented in XACC are presented, while also illustrating the various APIs that would enable the community to extend, modify, and devise new algorithms and applications in the realm of chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Retargetable Optimizing Compilers for Quantum Accelerators via a Multi-Level Intermediate Representation

In this work, we present a multi-level quantum-classical intermediate representation (IR) that enables an optimizing, retargetable compiler for available quantum languages. Our work builds upon the Multi-level Intermediate Representation (MLIR) framework and leverages its unique progressive lowering capabilities to map quantum languages to the LLVM machine-level IR. We provide both quantum and classical optimizations via the MLIR pattern rewriting sub-system and standard LLVM optimization passes, and demonstrate the programmability, compilation, and execution of our approach via standard benchmarks and test cases. In comparison to other standalone language and compiler efforts available today, our work results in compile times that are 1000x faster than standard Pythonic approaches, and 5-10x faster than comparative standalone quantum language compilers. Our compiler provides quantum resource optimizations via standard programming patterns that result in a 10x reduction in entangling operations, a common source of program noise. We see this work as a vehicle for rapid quantum compiler prototyping.

43 PARTICLE ACCELERATORS↗

MAPredict: Static Analysis Driven Memory Access Prediction Framework for Modern CPUs

Application memory access patterns are crucial in deciding how much traffic is served by the cache and forwarded to the dynamic random-access memory (DRAM). However, predicting such memory traffic is difficult because of the interplay of prefetchers, compilers, parallel execution, and innovations in manufacturer-specific micro-architectures. This research introduced MAPredict, a static analysis-driven framework that addresses these challenges to predict last-level cache (LLC)-DRAM traffic. By exploring and analyzing the behavior of modern Intel processors, MAPredict formulates cache-aware analytical models. MAPredict invokes these models to predict LLC-DRAM traffic by combining the application model, machine model, and user-provided hints to capture dynamic information. MAPredict successfully predicts LLC-DRAM traffic for different regular access patterns and provides the means to combine static and empirical observations for irregular access patterns. Evaluating 130 workloads from six applications on recent Intel micro-architectures, MAPredict yielded an average accuracy of 99% for streaming, 91% for strided, and 92% for stencil patterns. By coupling static and empirical methods, up to 97% average accuracy was obtained for random access patterns on different micro-architectures.

Monil, M. A. H.↗

GCAM Regional Tuning: A framework to tune GCAM parameters

GCAM assumptions typically generate scenarios that are designed to be internally consistent and globally coherent. The gcamdata tool which facilitates the compilation of data sets and user assumptions is not well suited to tailoring to specific country or regional realities, sponsor requirements, or perform harmonization for model intercomparison needs. As described in this report, the GCAM Regional Tuning project develops a computational framework that enables users to adjust GCAM parameters, so model outputs match targeted outcomes at user-defined spatial, temporal, and sectoral resolutions. The framework integrates GCAM, gcamdata, and gcamwrapper with a set of flexible “tuning directives” and an iterative numerical solver. Users can define targets (e.g., technology shares in power generation, BEV uptake, sectoral service demands), select tuners that manipulate relevant GCAM parameters (e.g., share weights, cost adders, elasticities), and export tuned parameters as reusable GCAM XML inputs for future runs. We demonstrate the approach and document usage, diagnostics, and known limitations, and we outline potential future directions.

97 MATHEMATICS AND COMPUTING↗

Bridging Python to Silicon: The SODA Toolchain

Systems performing scientific computing, data analysis, and machine learning tasks have a growing demand for application-specific accelerators that can provide high computational performance while meeting strict size and power requirements. However, the algorithms and applications that need to be accelerated are evolving at a rate that is incompatible with manual design processes based on hardware description languages. Agile hardware design tools based on compiler techniques can help by quickly producing an application-specific integrated circuit (ASIC) accelerator starting from a high-level algorithmic description. Here, we present the software-defined accelerator (SODA) synthesizer, a modular and open-source hardware compiler that provides automated end-to-end synthesis from high-level software frameworks to ASIC implementation, relying on multilevel representations to progressively lower and optimize the input code. Our approach does not require the application developer to write any register-transfer level code, and it is able to reach up to 364 giga floating point operations per second (GFLOPS)/W efficiency (32-bit precision) on typical convolutional neural network operators.

97 MATHEMATICS AND COMPUTING↗

A Full-Stack Exploration of Language-Based Parallelism in Fortran 2023

This poster explores native parallel features in Fortran 2023 through the lens of supporting applications with libraries, compilers, and parallel runtimes. The language revision informally named Fortran 2008 introduced parallelism in the form of Single Program Multiple Data (SPMD) execution with two broad feature sets: (1) loop-level parallelism via do concurrent and (2) a Partitioned Global Address Space (PGAS) comprised of distributed “coarray” data structures. Fortran’s native parallelism has demonstrated high performance [1] and reduced the burden of inserting what sometimes amounts to more directives than code. Several compilers support both feature sets, typically by translating do concurrent into serial do loops annotated by parallel directives and by translating SPMD/PGAS features into direct calls to a communication library. Our research focuses primarily on two questions: (1) can the compiler’s parallel runtime library be developed in the language being compiled (Fortran) and (2) can we define an interface to the runtime that liberates compilers from being hardwired to one runtime and vice versa. We are answering these questions by developing the Parallel Runtime Interface for Fortran (PRIF) [2] and the Co-Array Fortran Framework of Efficient Interfaces to Network Environments (Caffeine) [3]. Caffeine is initially targeting adoption by LLVM Flang, a new open-source Fortran compiler developed by a broad community in industry, academia, and government labs. We are also exploring the use of these features in Inference-Engine, a deep learning library designed to facilitate neural network training and inference for high-performance computing applications written in modern Fortran.

Rasmussen, Katherine↗

The Aerosol Model Benchmarking Repository: A toolkit for model intercomparison

The Aerosol Model Benchmarking Repository and Standards (AMBRS) project was initiated to provide tools and to establish community standards for benchmarking aerosol models. This report describes a set of open-source tools for building, running, and analyzing aerosol box model simulations in a standardized framework. The framework consists of three core components: AMBuilder, a CMake-based build system that compiles supported models consistently; AMBRS, a Python module that defines unified numerical experiments and executes them with aligned inputs; and PyParticle, an aerosol analysis package that standardizes output, computes diagnostics, and visualizes simulation results. Together, these tools enable reproducible intercomparison of aerosol schemes and support process-level evaluation of how model simplifications affect predictions of size distributions, cloud condensation nuclei activity, and other relevant properties relevant for the Earth-Energy system. Beyond its role in benchmarking, AMBRS provides a platform for studying aerosol processes across scales and can be used to generate training data for AI/ML applications in support of a broader hierarchical aerosol modeling strategy.

54 ENVIRONMENTAL SCIENCES↗

A Bayesian approach to evaluation of soil biogeochemical models

Abstract. To make predictions about the carbon cycling consequences of rising global surface temperatures, Earth system scientists rely on mathematical soil biogeochemical models (SBMs). However, it is not clear which models have better predictive accuracy, and a rigorous quantitative approach for comparing and validating the predictions has yet to be established. In this study, we present a Bayesian approach to SBM comparison that can be incorporated into a statistical model selection framework. We compared the fits of linear and nonlinear SBMs to soil respiration data compiled in a recent meta-analysis of soil warming field experiments. Fit quality was quantified using Bayesian goodness-of-fit metrics, including the widely applicable information criterion (WAIC) and leave-one-out cross validation (LOO). We found that the linear model generally outperformed the nonlinear model at fitting the meta-analysis data set. Both WAIC and LOO computed higher overfitting risk and effective numbers of parameters for the nonlinear model compared to the linear model, conditional on the data set. Goodness of fit for both models generally improved when they were initialized with lower and more realistic steady-state soil organic carbon densities. Still, testing whether linear models offer definitively superior predictive performance over nonlinear models on a global scale will require comparisons with additional site-specific data sets of suitable size and dimensionality. Such comparisons can build upon the approach defined in this study to make more rigorous statistical determinations about model accuracy while leveraging emerging data sets, such as those from long-term ecological research experiments.

54 ENVIRONMENTAL SCIENCES↗

Automated Generation of Integrated Digital and Spiking Neuromorphic Machine Learning Accelerators

The growing numbers of application areas for artificial intelligence (AI) methods have led to an explosion of domain-specific accelerators that could support every new machine learning (ML) algorithm advancement, clearly highlighting the need for a capability to quickly and automatically transition from algorithm definition to hardware implementation and explore design space along a variety of SWaP (size, weight and Power). The software defined architectures (SODA) synthesizer implements a compiler-based modular infrastructure for the end-to-end generation of machine learning accelerators from high-level frameworks to hardware description language. At the same time, neuromorphic computing, by mimicking how the brain operates, promises to perform artificial intelligence tasks at efficiencies orders of magnitude higher than the current conventional tensor-processing based accelerators, as demonstrated by a variety of specialized designs leveraging Spiking Neural Networks (SNNs). Nevertheless, the mapping of an artificial neural network (ANN) to solutions supporting SNNs is still a non-trivial and very device-specific task, and completely lack the possibility to design hybrid systems that integrate conventional and spiking neural models. In this paper we discuss the support for such an integrated generation leveraging the SODA Synthesizer framework and its modular structure. In particular, we present a new MLIR dialect (part of the SODA frontend) that allows expressing spiking neural network features (e.g., available resources, spiking sequences, analog signal reading, etc.) and illustrate how it enables mapping to Spiking Neurons and deployment to the related specialized hardware (which, in the digital domain, could be generated through the other existing layers of the SODA Synthesizer). We then discuss the opportunities for even deeper integration afforded by the hardware compilation infrastructure, providing a path towards the generation of complex heterogeneous artificial intelligence systems.

Curzel, Serena↗

Development of a Framework and Methodology for an Advanced Reactor Materials Environmental Effects Design Guide

Advanced non-light-water reactor components may operate at elevated temperature while experiencing cyclic loading, significant neutron irradiation, and exposure to reactor coolant. ASME Boiler and Pressure Vessel Code, Section III, Division 5, provides design rules for elevated-temperature service but does not include specific procedures to account for environmental effects on material properties. This report develops an initial framework and methodology for an Environmental Effects Design Guide (EEDG) focused on neutron irradiation; coolant-environment effects are reserved for future work. The proposed approach treats irradiation as a property-based overlay on the existing Division 5 design process, with two routes: a sparse-data route applying two reduction factors — FCR on creep-rupture strength and FF on fatigue life — for the creep-fatigue evaluations that typically control the design of advanced high-temperature reactor components, and a fuller framework developing the property-to-rule chain across the four Division 5 checks (primary load, strain limits and ratcheting, creep-fatigue, and buckling), together with swelling and weldments as scope items. Both routes are scoped by an in-pile qualification that restricts the use of post-irradiation-examination-derived properties in regimes where an in-pile mechanism could control the design outcome. Illustrative outputs derived on a compiled annealed Type 316 database — FCR ≈ 0.78–0.86 and FF ≈ 0.4 — demonstrate the calculation method within that specific dataset. The framework is an initial, testable design-rule concept; it identifies a practical path for preliminary design evaluations under sparse data and the material data and testing needed to develop the framework further.

Barua, Bipul (ORCID:0000000247184113)↗

Using OpenMP for HEP framework algorithm scheduling

The OpenMP standard is the primary mechanism used at high performance computing facilities to allow intra-process parallelization. In contrast, many HEP specific software packages (such as CMSSW, GaudiHive, and ROOT) make use of Intel’s Threading Building Blocks (TBB) library to accomplish the same goal. In these proceedings we will discuss our work to compare TBB and OpenMP when used for scheduling algorithms to be run by a HEP style data processing framework. This includes both scheduling of different interdependent algorithms to be run concurrently as well as scheduling concurrent work within one algorithm. As part of the discussion we present an overview of the OpenMP threading model. We also explain how we used OpenMP when creating a simplified HEP-like processing framework. Using that simplified framework, and a similar one written using TBB, we will present performance comparisons between TBB and different compiler versions of OpenMP.

97 MATHEMATICS AND COMPUTING↗

Simple, Secure, Internet Delivery of MOOSE-based Applications

Application packaging and distribution are the final steps for delivering software to end-users; both are frequently neglected when creating scientific software. Commercial businesses rely on electronic distribution systems that have rendered disk drives obsolete. Still, national laboratories continue to rely heavily on removable media to distribute and limit access to controlled applications. With increasing concerns of unauthorized copying of sensitive applications, a modern distribution system that utilizes cryptographically secure communication and authentication protocols has been developed. This new distribution system will secure the chain of custody for nuclear software while simultaneously simplifying access to these tools. This report summarizes four primary advancements made toward the secure distribution of Nuclear Energy Advanced Modeling and Simulation (NEAMS)-developed, Multiphysics Object Oriented Simulation Environment (MOOSE)-based applications: application installation, package distribution, automated package building, and distribution of documentation. NEAMS is currently developing more than ten separate applications based on the open-source MOOSE Framework. Distribution of these applications has primarily been accomplished by distributing source code, with end-users compiling the applications themselves. This work created a mechanism where MOOSE applications can be installed in a similar way to any other software. This allows both administrators and end-users simplified access to runnable executables. With this new installation capability, it was then possible to rethink distribution. A new, secure capability for delivering MOOSE-based applications over the internet has been created. This system requires unique cryptographic tokens for authentication, greatly securing the custody chain for software. Once granted access, installation of any NEAMS code can be accomplished with these terminal commands: "conda install ncrc" "ncrc install ncrc-bison." After these two commands (and authenticating) the BISON application will be securely down- loaded from Idaho National Laboratory (INL)’s servers, installed, and ready to use. To enable this new distribution capability to be successful, the open-source Continuous Integration, Verification, Enhancement, and Testing (CIVET) Continuous Integration (CI) capability was augmented to add Continuous Delivery (CD). CD enables the automated building and packaging of MOOSE-based applications as they are modified by development teams, ensuring that our customers can obtain up-to-date versions of the software at any time. The need for instruction on how to use these applications was addressed through modifications to the MOOSE documentation system. The MooseDocs capability, which enables robust documentation of MOOSE-based applications, has been extended to allow both for the installation of documentation and the packaging of documentation with installed applications. Together, these enhancements form the core of a new, secure distribution mechanism for nuclear simulation tools. In concert with the Nuclear Computational Resource Center (NCRC), NEAMS- developed applications will now be straightforward to obtain securely.

97 MATHEMATICS AND COMPUTING↗

Accelerated Charged Particle Tracking with Graph Neural Networks on FPGAs

We develop and study FPGA implementations of algorithms for charged particle tracking based on graph neural networks. The two complementary FPGA designs are based on OpenCL, a framework for writing programs that execute across heterogeneous platforms, and hls4ml, a high-level-synthesis-based compiler for neural network to firmware conversion. We evaluate and compare the resource usage, latency, and tracking performance of our implementations based on a benchmark dataset. We find a considerable speedup over CPU-based execution is possible, potentially enabling such algorithms to be used effectively in future computing workflows and the FPGA-based Level-1 trigger at the CERN Large Hadron Collider.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

OpenACC Unified Programming Environment for Multi-hybrid Acceleration with GPU and FPGA

Accelerated computing in HPC such as with GPU, plays a central role in HPC nowadays. However, in some complicated applications with partially different performance behavior is hard to solve with a single type of accelerator where GPU is not the perfect solution in these cases. We are developing a framework and transpiler allowing the users to program the codes with a single notation of OpenACC to be compiled for multi-hybrid accelerators, named MHOAT (Multi-Hybrid OpenACC Translator) for HPC applications. MHOAT parses the original code with directives to identify the target accelerating devices, currently supporting NVIDIA GPU and Intel FPGA, dispatching these specific partial codes to background compilers such as NVIDIA HPC SDK for GPU and OpenARC research compiler for FPGA, then assembles binaries for the final object with FPGA bitstream file. In this paper, we present the concept, design, implementation, and performance evaluation of a practical astrophysics simulation code where we successfully enhanced the performance up to 10 times faster than the GPU-only solution.

Boku, Taisuke↗

NEML2: A High Performance Library for Constitutive Modeling

NEML2, the New Engineering Material model Library, version 2, is an offshoot of NEML, an earlier material modeling code developed at Argonne National Laboratory. NEML2 extends the key philosophy of its predecessor, i.e., material models are flexible, modular, and can be built from smaller blocks. It also provides modern features that do not exist in the framework of its predecessor such as material model vectorization, automatic differentiation, device-portable just-in-time compilation, operator fusion, lazy tensor evaluation, etc. Moreover, NEML2 can seamlessly integrate with the popular machine learning package PyTorch to take advantage of modern and fast-growing machine learning techniques. In this fiscal year, the development of core library features and capabilities are complete. The purpose of this report is not to serve as a verbatim copy of the software API reference (which is available online at https://reverendbedford.github.io/neml2/). Instead, this report documents the motivation, implementation, design choices, and usage of each core capability as well as their applications in solving practical engineering problems. This report is compiled based on the NEML2 major release 2.0.0.

36 MATERIALS SCIENCE↗