Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Compiler frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The SODA Approach: Leveraging High-Level Synthesis for Hardware/Software Co-design and Hardware Specialization: Invited

Novel "converged" applications combine phases of scientific simulation with data analysis and machine learning. Each computational phase can benefit from specialized accelerators. However, algorithms evolve so quickly that mapping them on existing accelerators is suboptimal or even impossible. This paper presents the SODA (Software Defined Accelerators) framework, a modular, multi-level, open-source, no-human-in-the-loop, hardware synthesizer that enables end-to-end generation of specialized accelerators. SODA is composed of SODA-Opt, a high-level frontend developed in MLIR that interfaces with domain-specific programming frameworks and allows performing system level design, and Bambu, a state-of-the-art high-level synthesis engine that can target different device technologies. The framework implements design space exploration as compiler optimization passes. We show how the modular, yet tight, integration of the high-level optimizer and lower-level HLS tools enables the generation of accelerators optimized for the computational patterns of converged applications. We then discuss some of the research opportunities that such a framework allows, including system-level design, profile driven optimization, and supporting new optimization metrics.

Bohm Agostini, Nicolas↗

Variational Optical Phase Learning on a Continuous-Variable Quantum Compiler

Quantum process learning is a fundamental primitive that draws inspiration from machine learning with the goal of better studying the dynamics of quantum systems. One approach to quantum process learning is quantum compilation, whereby an analog quantum operation is digitized by compiling it into a series of basic gates. While there has been significant focus on quantum compiling for discrete-variable systems, the continuous-variable (CV) framework has received comparatively less attention. We present an experimental implementation of a CV quantum compiler that uses two-mode squeezed light to learn a Gaussian unitary operation. We demonstrate the compiler by learning a parameterized linear phase unitary through the use of target and control phase unitaries to demonstrate a factor of 5.4 increase in the precision of the phase estimation and a 3.6-fold acceleration in the time-to-solution metric when leveraging quantum resources. We further show how our approach can be extended to higher-dimensional compilation tasks. Our results are enabled by the tunable control of our cost landscape via variable squeezing, thus providing a critical framework to simultaneously increase precision and reduce time-to-solution.

97 MATHEMATICS AND COMPUTING↗

LISE$^{++}_{cute}$, the latest generation of the LISE ++ package, to simulate rare isotope production with fragment-separators

The LISE ++ software for fragment separator simulations has undergone a major update. The package, widely used at rare isotope beam facilities, can be used to predict intensities and purities of rare isotope beams and for planning and running of experiments using in-flight separators. It is especially useful for radioactive beam production as its results can be quickly compared to on-line data. The LISE ++ package has been ported to the Qt-framework in order to support modern compilers and computing methods. The benefits include 64-bit operation and LISE ++ availability on three different platforms: Windows, MacOS and Linux. In addition, the porting provides the ability to take advantage of future computational improvements. The updated package is named LISE$^{++}_{cute}$ to indicate a major step forward from the previous Borland-based versions. In addition to porting to the new platform, new main features and modifications have been added, mostly devoted to improving models and implementing other codes involved in rare isotope production at FRIB. Finally, a summary of modifications completed to improve the functionality of the code are discussed in this work, as well as future plans.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Fortran Compiler Test Suite v0.1.0

There is no currently available, open-source, easily adapted test suite for checking a compiler's conformance to the Fortran Standard. The main innovation of this software is in the framework to make the test suite apply to a new compiler. Making it open-source will enable community contributions of the test cases. Single organization construction of a comprehensive test suite is something that would be financially infeasible.

Richardson, Bradley↗

Application of community data to surface complexation modeling framework development: Iron oxide protolysis

This study presents a comprehensive community data-driven surface complexation modeling framework for simulating potentiometric titration of mineral surfaces. Compiled community data for ferrihydrite, goethite, hematite, and magnetite are fit to produce representative protolysis constants that can reproduce potentiometric titration data collected from multiple literature sources. Using this framework, the impact of surface complexation model type and surface site density (SSD) on the fit quality and protolysis constants can be readily evaluated. For example, the non-electrostatic model yielded a poor data fit compared to diffuse double layer model and constant capacitance models due to the absence of known surface charge effects. Regardless of the choice of iron oxide mineral, pK a1 decreased with increasing SSD while the opposite tendency was observed for pK a2 . This newly developed framework demonstrates a method to reconcile community data-wide potentiometric titration data using Findable, Accessible, Interoperable, Reusable data principles to produce mineral protolysis constants that improve robustness of surface complexation models for applications in metal sorption and reactive transport modeling. The framework is readily expandable (as community data increase) and extensible (as the number of minerals increase). The framework provides a path forward for developing self-consistent, comprehensive, and updateable surface complexation databases for surface complexation and reactive transport modeling.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

An MLIR-based Compiler Flow for System-Level Design and Hardware Acceleration

The generation of custom hardware accelerators for applications implemented within high-level productive programming frameworks requires considerable manual effort. To automate this process, we introduce \sodaopt, a compiler tool that extends the MLIR infrastructure. \sodaopt automatically searches, outlines, tiles, and pre-optimizes relevant code regions to generate high-quality accelerators through high-level synthesis. \sodaopt can support any high-level programming framework and domain-specific language that interface with the MLIR infrastructure. By leveraging MLIR, \sodaopt solves compiler optimization problems with specialized abstractions. Backend synthesis tools connect to \sodaopt through progressive intermediate representation lowerings. \sodaopt interfaces to a design space exploration engine to identify the combination of compiler optimization passes and options that provides high-performance generated designs for different backends and targets. We demonstrate the practical applicability of the compilation flow by exploring the automatic generation of accelerators for deep neural networks operators outlined at arbitrary granularity and by combining outlining with tiling on large convolution layers. Experimental results with kernels from the PolyBench benchmark show that \sodaopt high-level optimizations improve execution delays of synthesized accelerators up to 60x. We also show that for the selected kernels, our solution outperforms the current of state-of-the art in more than 70% of the benchmarks and provides better average speedup in 55% of them.

Bohm Agostini, Nicolas↗

Unified Language Frontend for Physic-Informed AI/ML

Artificial intelligence and machine learning (AI/ML) are becoming important tools for scientific modeling and simulation as in several other fields such as image analysis and natural language processing. ML techniques can leverage the computing power available in modern systems and reduce the human effort needed to configure experiments, interpret and visualize results, draw conclusions from huge quantities of raw data, and build surrogates for physics based models. Domain scientists in fields like fluid dynamics, microelectronics and chemistry can automate many of their most difficult and repetitive tasks or improve the design times by use of the faster ML-surrogates. However, modern ML and traditional scientific highperformance computing (HPC) tend to use completely different software ecosystems. While ML frameworks like PyTorch and TensorFlow provide Python APIs, most HPC applications and libraries are written in C++. Direct interoperability between the two languages is possible but is tedious and error-prone. In this work, we show that a compiler-based approach can bridge the gap between ML frameworks and scientific software with less developer effort and better efficiency. We use the MLIR (multi-level intermediate representation) ecosystem to compile a pre-trained convolutional neural network (CNN) in PyTorch to freestanding C++ source code in the Kokkos programming model. Kokkos is a programming model widely used in HPC to write portable, shared-memory parallel code that can natively target a variety of CPU and GPU architectures. Our compiler-generated source code can be directly integrated into any Kokkosbased application with no dependencies on Python or cross-language interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Towards Ultra-high-resolution E3SM Land Modeling on Exascale Computers

Here we present an ultra-high-resolution E3SM land model (uELM) for high-fidelity land simulations targeting new Exascale computers. After considering modeling infrastructure compatibility and ELM software features, we designed a parallel model for the uELM development targeting hybrid architectures of new US Exascale computers. We also described a function unit test framework to expedite the piece-wise code porting (with compiler directives), verification, and global variable management. Furthermore, in this study, we report an early uELM model development using OpenACC within a function unit test framework on a pre-Exascale computer, demonstrate the performance of a uLEM submodel with a 3.0-time speedup, and summarize the code porting experience regarding global variable handling, deepcopy, memory reduction, and parallel loop reconstruction.

97 MATHEMATICS AND COMPUTING↗

Community Based Data of Potentiometric Titration of Iron Oxides: Ferrihydrite (HFO), Goethite, Hematite, Magnetite

This data release includes experimental data of potentiometric titration for iron oxides. The data in the provided .csv files is not our own experimental data but have been compiled from the multiple literature sources. The master database is L-SCIE (LLNL Surface Complexation/Ion Exchange) database, and the provided .csv files are extracted data from L-SCIE. The .csv files were obtained by using the Lawrence Livermore National Laboratory Surface Complexation Database Converter (SCDC) code written in the R programming language (free licensing available at https://ipo.llnl.gov/technologies/software/llnl-surface-complexation-database-converter-scdc).The released data was used for developing a comprehensive community data-driven surface complexation modeling (SCM) framework for simulating potentiometric titration of mineral surfaces. Compiled community data for ferrihydrite, goethite, hematite, and magnetite are fit to produce representative protolysis constants that can reproduce potentiometric titration data collected from multiple literature sources.

54 ENVIRONMENTAL SCIENCES↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

Ciel

Compiler optimizations can alter the numerical results of scientific computing applications. When numerical results differ significantly between compilers, optimization levels, and floating-point hardware, these numerical inconsistencies can impact programming productivity. Ciel is a framework that helps programmers identify locations in the source code that are affected by compiler optimizations in CPU and GPU code. Ciel uses a floating-point precision enhancement strategy, guided by a recursive bisection search algorithm with increasing search granularity, to identify the program expressions that induce numerical inconsistencies due to compiler optimizations.

Miao, Wenjun↗

Data driven investigation to understand the influence of total solids on biological biogas upgrading

In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.

In situ biogas upgrading↗

pLiner

Compiler optimizations can alter significantly the numerical results of scientific computing applications. When numerical results differ significantly between compilers, optimization levels, and floating-point hardware, these numerical inconsistencies can impact programming productivity. pLiner is a framework that helps programmers identify locations in the source code that are highly affected by compiler optimizations. pLiner uses a novel approach to identify such code locations by enhancing the floatingpoint precision of variables and expressions. Using a guided search to locate the most significant code regions, pLiner can report to users such locations at different granularities, file, function, and line of code.

Laguna Peralta, Ignacio↗

Iodine Immobilization by Materials through Sorption and Redox-Driven Processes: A Literature Review

Radioiodine-129 (129I) in the subsurface is mobile and limited information is available on treatment technologies. Scientific literature was reviewed to compile information on materials that could potentially be used to immobilize 129I through sorption and redox-driven processes, with an emphasis on ex-situ processes. Candidate materials to immobilize 129I include iron minerals, sulfur-based materials, silver-based materials, bismuth-based materials, ion exchange resins, activated carbon, modified clays, and tailored materials (metal organic frameworks (MOFS), layered double hydroxides (LDHs) and aerogels). Where available, compiled information includes material performance in terms of (i) capacity for 129I uptake; (ii) long-term performance (i.e., solubility of a precipitated phase); (iii) technology maturity; (iv) cost; (v) available quantity; (vi) environmental impact; (vii) ability to emplace the technology for in situ use at the field-scale; and (viii) ex situ treatment (for media extracted from the subsurface or secondary waste streams). Because it can be difficult to compare materials due to differences in experimental conditions applied in the literature, Part II of this review describes results of laboratory studies for selected materials using a standardized batch loading test.

Moore, Robert C.↗

Resilience–runtime tradeoff relations for quantum algorithms

Abstract A leading approach to algorithm design aims to minimize the number of operations in an algorithm’s compilation. One intuitively expects that reducing the number of operations may decrease the chance of errors. This paradigm is particularly prevalent in quantum computing, where gates are hard to implement and noise rapidly decreases a quantum computer’s potential to outperform classical computers. Here, we find that minimizing the number of operations in a quantum algorithm can be counterproductive, leading to a noise sensitivity that induces errors when running the algorithm in non-ideal conditions. To show this, we develop a framework to characterize the resilience of an algorithm to perturbative noises (including coherent errors, dephasing, and depolarizing noise). Some compilations of an algorithm can be resilient against certain noise sources while being unstable against other noises. We condense these results into a tradeoff relation between an algorithm’s number of operations and its noise resilience. We also show how this framework can be leveraged to identify compilations of an algorithm that are better suited to withstand certain noises.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Framework for Extensible, Asynchronous Task Scheduling (FEATS) in Fortran

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP, explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI), or compiler-specific language extensions such as those provided by CUDA. By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models. Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data-parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. The paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task-scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

Richardson, Brad↗

Extending C++ for Heterogeneous Quantum-Classical Computing

In this report we present qcor - a language extension to C++ and compiler implementation that enables heterogeneous quantum-classical programming, compilation, and execution in a single-source context. Our work provides a first-of-its-kind C++ compiler enabling high-level quantum kernel (function) expression in a quantum-language agnostic manner, as well as a hardware-agnostic, retargetable compiler workflow targeting a number of physical and virtual quantum computing backends. qcor leverages novel Clang plugin interfaces and builds upon the XACC system-level quantum programming framework to provide a state-of-the-art integration mechanism for quantum-classical compilation that leverages the best from the community at-large. qcor translates quantum kernels ultimately to the XACC intermediate representation, and provides user-extensible hooks for quantum compilation routines like circuit optimization, analysis, and placement. This work details the overall architecture and compiler workflow for qcor, and provides a number of illuminating programming examples demonstrating its utility for near-term variational tasks, quantum algorithm expression, and feed-forward error correction schemes.

97 MATHEMATICS AND COMPUTING↗

Opinion: Coordinated development of emission inventories for climate forcers and air pollutants

Emissions into the atmosphere of fine particulate matter, its precursors, and precursors to tropospheric ozone impact not only human health and ecosystems, but also the climate by altering Earth's radiative balance. Accurately quantifying these impacts across local to global scales historically and in future scenarios requires emission inventories that are accurate, transparent, complete, comparable, and consistent. In an effort to better quantify the emissions and impacts of these pollutants, also called short-lived climate forcers (SLCFs), the Intergovernmental Panel on Climate Change (IPCC) is developing a new SLCF emissions methodology report. This report would supplement existing IPCC reporting guidance on greenhouse gas (GHG) emission inventories, which are currently used by inventory compilers to fulfill national reporting requirements under the United Nations Framework Convention on Climate Change (UNFCCC) and new requirements of the Enhanced Transparency Framework (ETF) under the Paris Agreement starting in 2024. We review the relevant issues, including how air pollutant and GHG inventory activities have historically been structured, as well as potential benefits, challenges, and recommendations for coordinating GHG and air pollutant inventory efforts. We argue that, while there are potential benefits to increasing coordination between air pollutant and GHG inventory development efforts, we also caution that there are differences in appropriate methodologies and applications that must jointly be considered.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗