Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Compiler frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Community Data Mining Approach for Surface Complexation Database Development

This paper presents a comprehensive data-to-model workflow, including a findable, accessible, interoperable, reusable (FAIR) community sorption database (newly developed LLNL Surface Complexation/Ion Exchange (L-SCIE) database) along with a data fitting workflow to efficiently optimize surface complexation reaction constants with multiple surface complexation model (SCM) constructs. This workflow serves as a universal framework to mine, compile, and analyze large numbers of published sorption data as well as to estimate reaction constants for parameterizing reactive transport models. Here the framework includes (1) data digitization from published papers, (2) data unification including unit conversions, and (3) data-model integration and reaction constant estimation using geochemical software PHREEQC coupled with the universal parameter estimation code PEST. We demonstrate our approach using an analysis of U(VI) sorption to quartz based on a first L-SCIE implementation, concluding that a multisite SCM construct with carbonate surface species yielded the best fit to community data. Surface complexation reaction constants extracted from this approach captured all available sorption data available in the literature and provided insight into previously published reaction constants and surface complexation model constructs. The L-SCIE sorption database presented herein allows for automating this approach across a wide range of metals and minerals and implementing novel machine learning approaches to reactive transport in the future.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

ChemComp: Compiling and Computing with Chemical Reaction Networks

The exponential growth in computing demands driven by scientific computing, data analytics, and artificial intelligence is pushing conventional CMOS-based high-performance computing systems to their physical and energy efficiency limits. As we approach the era of post-exascale computing, disruptive approaches are necessary to overcome these barriers and achieve substantial gains in energy efficiency. Analog and hybrid digital-analog computing systems have emerged as promising alternatives, offering the potential for orders-of-magnitude improvements in efficiency. Among these, biochemical computing stands out as a novel paradigm capable of leveraging the natural efficiency of chemical reactions, which have shown promise in solving optimization problems by converging to steady states. By scaling up reaction networks or reaction vessel sizes, biochemical systems present an opportunity to meet the high-performance demands of modern computing tasks. Despite their promise, significant theoretical and practical challenges remain, particularly in formulating and mapping computational problems to chemical reaction networks (CRNs) and designing viable biochemical computing devices. This paper addresses these challenges by introducing new ideas to ChemComp, a compilation and emulation framework for chemical computation. This work describes the mechanisms through which solutions to ordinary differential equations (ODEs) that can be represented as CRN systems can be achieved. Furthermore, we explain the design principles of an ODE dialect implemented as a multi-level intermediate representation (MLIR) compiler extension that will be coupled with existing infrastructure. We demonstrate the potential of our framework through a case study emulating a simplified chemical reservoir computing device. This work establishes foundational tools and methodologies necessary to harness the computational power of chemistry, paving the way for the development of energy-efficient, high-performance computing systems tailored to contemporary and future computational needs.

Bohm Agostini, Nicolas↗

Examples of State and Utility Actions on Proactive Planning and Investments

As states across the U.S. confront rising electricity demand, clean energy deployment, grid modernization imperatives, and the integration of large new loads, some regulators and utilities are shifting away from reactive, “just-in-time” investment approaches toward more proactive planning and investment frameworks. This report compiles examples of jurisdictional and utility actions that reshape planning processes, cost recovery mechanisms, and performance oversight to anticipate—rather than simply respond to—future grid needs. Several themes emerge from state actions examined in this report. First, legislatures and commissions are increasingly directing utilities to proactively upgrade their distribution and transmission systems, reflecting a shift toward a forward-looking system that aligns planning with state policy goals. Second, states are establishing long-term, iterative planning frameworks that often feature multi-year horizons, biannual or annual compliance reporting, structured opportunities for stakeholder engagement, and emphasis on collaboration among utilities, regulators, and stakeholders. Third, states are actively investigating innovative cost recovery mechanisms designed to support accelerated electrification and grid modernization, while balancing consumer advocates’ concerns regarding the ratepayer financial risks of premature investments. Fourth, performance metrics and reporting requirements are being developed to ensure transparency and accountability for proactive investments. Fifth, methodological improvements in planning—such as aligning load forecasting assumptions, incorporating sensitivities, and considering load management potential across building, vehicles, storage, and demand response—are recurring areas of stakeholder focus across jurisdictions. Overall, these developments signify a growing recognition among state regulators, utilities, and stakeholders that proactive planning—supported by clear definitions, consistent and transparent methodologies, robust performance metrics, and innovative cost recovery mechanisms—is a tool that can be used to address the scale and urgency of contemporary grid needs.

electricity market↗

Union: A Unified HW-SW Co-Design Ecosystem in MLIR for Evaluating Tensor Operationson Spatial Accelerators

To meet the extreme compute demands for deep learning across commercial and scientific applications, dataflow accelerators are becoming increasingly popular. While these“domain-specific” accelerators are not fully programmable like CPUs and GPUs, they retain varying levels of flexibility with respect to data orchestration, i.e., dataflow and tiling optimizations to enhance efficiency. There are several challenges when designing new algorithms and mapping approaches to execute the algorithms for a target problem on new hardware. Previous works have addressed these challenges individually. To address this challenge as a whole, in this work, we present an HW-SW co-design ecosystem for spatial accelerators called Union within the popular MLIR compiler infrastructure. Our framework allows exploring different algorithms and their mappings on several accelerator cost models. Union also includes a plug-and-play library of accelerator cost models and mappers which can easily be extended. The algorithms and accelerator cost models are connected via a novel mapping abstraction that captures the map space of spatial accelerators which can be systematically pruned based on constraints from the hardware, workload, and mapper. We demonstrate the value of Union for the community with several case studies which examine offloading different tensor operations (CONV/GEMM/Tensor Contraction) on diverse accelerator architectures using different mapping schemes.

Jeong, Geonhwa↗

Software Defined Architectures for Portability and Performance

The Software Defined Architectures for Portability and Performance (SODAPOP) project developed a co-design framework to partition and map converged applications on specialized heterogeneous architectures. We started from key domain applications that combine scientific simulation with data analytics and machine learning as drivers to integrate our framework. The framework includes high-level compilers that interfaces with high-level programming frameworks, domain-specific optimization passes, and hardware-oriented optimizations. The framework leverages hardware generators to enable specialization and facilitate exploration of custom system designs.

97 MATHEMATICS AND COMPUTING↗

Towards Automatic and Agile AI/ML Accelerator Design with End-to-End Synthesis

Domain-specific designs offer greater energy efficiency and performance gain than general-purpose processors. For this reason, modern system-on-chips have a significant portion of their silicon area with custom accelerators. However, designing hardware by hand is laborious and time-consuming, given the large design space and the performance, power, and area constraints that are not realized in the software. Moreover, domain-specific algorithms (e.g., machine learning models) are evolving quickly, challenging the accelerator design further. To address these issues, this paper presents SODA Synthesizer, an automated open-source high-level ML framework to Verilog modular compiler targeting AI/ML Application-Specific Integrated Circuits (ASICs) accelerators. SODA tightly couples the Multi- Level Intermediate Representation (MLIR) compiler infrastructure [24] and open-source HLS approaches. Thus, SODA can support various ML frameworks and algorithms and can perform optimizations that combine specialized architecture templates and conventional HLS to generate the hardware modules. In addition, SODA’s closed-loop design space exploration (DSE) engine allows developers to perform end-to-end design space explorations on different metrics and technology nodes.

Zhang, Jeff↗

Julienne v1.0.0

Julienne is a compiler-portable unit-testing framework for Fortran software projects, including those that use the parallel/accelerator-programming features of Fortran 2023. Julienne achieves portability across compilers through minimalism and isolation. The minimal design ensures that Julienne uses only features supported by the majority of Fortran compilers. The isolation through zero dependencies ensures that no other projects block Julienne from building with a particular compiler. Julienne also contains with additional services that support its unit-testing code. These include functions for manipulating strings, command lines and input/output format strings; and a user-defined collective subroutine for verifying that all processes pass a test in parallel testing. Julienne's name derives from the term for vegetables sliced into thin strings: julienne vegetables. Julienne captures the authors' most frequently used thin slice of the Veggies and Sourcery software repositories while avoiding certain compiler limitations of the those two packages.

Rouson, Damian↗

Importance of Properties of Solids to Friction and Wear Behaviour

The main properties of solids which influence friction and wear are discussed and published rules which relate material properties to friction and wear are considered. In addition, recent experimental results on the tribological behaviour of metals and polymers illustrating the effect of some important interaction characteristics on friction and wear are presented. Finally, a framework for the systematic compilation and documentation of relevant tribological parameters in experimental friction and wear investigations is given.

Czichos, H.↗

AEGIS (Air Emissions Grouped by Industrial Sectors) [SWR-24-49]

AEGIS is a robust framework designed to build and compile emissions inventories for various industrial sectors. It integrates multiple emissions databases provided by the USEPA – including GHGRP, NEI, and TRI – and leverages the STEWI and STEWICOMBO tools to retrieve, merge, and process data. The framework produces facility- and process-level inventories, identifies discrepancies (e.g., NAICS or FRS mismatches), and performs exploratory analysis including emission concentration calculations and visualization.

Atnoorkar, Swaroop↗

pnnl/CARTS

This compiler for ARTS (CARTS) is framework designed to connect a high productive languages with a distributed fine grained runtime that run effectively across clusters. It is built using MLIR and the LLVM infrastructure and it can be used as a platform to test static analysis ideas mapped towards distributed environments with novel technologies

Manzano Franco, Joseph [Pacific Northwest National↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

FOS: Computer and information sciences↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

Schulte, Jan-Frederik [Purdue U.] (ORCID:000000034↗

GAHLS: an optimized graph analytics based high level synthesis framework

The urgent need for low latency, high-compute and low power on-board intelligence in autonomous systems, cyber-physical systems, robotics, edge computing, evolvable computing, and complex data science calls for determining the optimal amount and type of specialized hardware together with reconfigurability capabilities. With these goals in mind, we propose a novel comprehensive graph analytics based high level synthesis (GAHLS) framework that efficiently analyzes complex high level programs through a combined compiler-based approach and graph theoretic optimization and synthesizes them into message passing domain-specific accelerators. This GAHLS framework first constructs a compiler-assisted dependency graph (CaDG) from low level virtual machine (LLVM) intermediate representation (IR) of high level programs and converts it into a hardware friendly description representation. Next, the GAHLS framework performs a memory design space exploration while account for the identified computational properties from the CaDG and optimizing the system performance for higher bandwidth. The GAHLS framework also performs a robust optimization to identify the CaDG subgraphs with similar computational structures and aggregate them into intelligent processing clusters in order to optimize the usage of underlying hardware resources. Finally, the GAHLS framework synthesizes this compressed specialized CaDG into processing elements while optimizing the system performance and area metrics. Evaluations of the GAHLS framework on several real-life applications (e.g., deep learning, brain machine interfaces) demonstrate that it provides 14.27× performance improvements compared to state-of-the-art approaches such as LegUp 6.2.

97 MATHEMATICS AND COMPUTING↗

The SODA Approach: Leveraging High-Level Synthesis for Hardware/Software Co-design and Hardware Specialization: Invited

Novel "converged" applications combine phases of scientific simulation with data analysis and machine learning. Each computational phase can benefit from specialized accelerators. However, algorithms evolve so quickly that mapping them on existing accelerators is suboptimal or even impossible. This paper presents the SODA (Software Defined Accelerators) framework, a modular, multi-level, open-source, no-human-in-the-loop, hardware synthesizer that enables end-to-end generation of specialized accelerators. SODA is composed of SODA-Opt, a high-level frontend developed in MLIR that interfaces with domain-specific programming frameworks and allows performing system level design, and Bambu, a state-of-the-art high-level synthesis engine that can target different device technologies. The framework implements design space exploration as compiler optimization passes. We show how the modular, yet tight, integration of the high-level optimizer and lower-level HLS tools enables the generation of accelerators optimized for the computational patterns of converged applications. We then discuss some of the research opportunities that such a framework allows, including system-level design, profile driven optimization, and supporting new optimization metrics.

Bohm Agostini, Nicolas↗

Variational Optical Phase Learning on a Continuous-Variable Quantum Compiler

Quantum process learning is a fundamental primitive that draws inspiration from machine learning with the goal of better studying the dynamics of quantum systems. One approach to quantum process learning is quantum compilation, whereby an analog quantum operation is digitized by compiling it into a series of basic gates. While there has been significant focus on quantum compiling for discrete-variable systems, the continuous-variable (CV) framework has received comparatively less attention. We present an experimental implementation of a CV quantum compiler that uses two-mode squeezed light to learn a Gaussian unitary operation. We demonstrate the compiler by learning a parameterized linear phase unitary through the use of target and control phase unitaries to demonstrate a factor of 5.4 increase in the precision of the phase estimation and a 3.6-fold acceleration in the time-to-solution metric when leveraging quantum resources. We further show how our approach can be extended to higher-dimensional compilation tasks. Our results are enabled by the tunable control of our cost landscape via variable squeezing, thus providing a critical framework to simultaneously increase precision and reduce time-to-solution.

97 MATHEMATICS AND COMPUTING↗

LISE$^{++}_{cute}$, the latest generation of the LISE ++ package, to simulate rare isotope production with fragment-separators

The LISE ++ software for fragment separator simulations has undergone a major update. The package, widely used at rare isotope beam facilities, can be used to predict intensities and purities of rare isotope beams and for planning and running of experiments using in-flight separators. It is especially useful for radioactive beam production as its results can be quickly compared to on-line data. The LISE ++ package has been ported to the Qt-framework in order to support modern compilers and computing methods. The benefits include 64-bit operation and LISE ++ availability on three different platforms: Windows, MacOS and Linux. In addition, the porting provides the ability to take advantage of future computational improvements. The updated package is named LISE$^{++}_{cute}$ to indicate a major step forward from the previous Borland-based versions. In addition to porting to the new platform, new main features and modifications have been added, mostly devoted to improving models and implementing other codes involved in rare isotope production at FRIB. Finally, a summary of modifications completed to improve the functionality of the code are discussed in this work, as well as future plans.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automatic array alignment in data-parallel programs

FORTRAN 90 and other data-parallel languages express parallelism in the form of operations on data aggregates such as arrays. Misalignment of the operands of an array operation can reduce program performance on a distributed-memory parallel machine by requiring nonlocal data accesses. Determining array alignments that reduce communication is therefore a key issue in compiling such languages. We present a framework for the automatic determination of array alignments in array-based, data-parallel languages. Our language model handles array sectioning, reductions, spreads, transpositions, and masked operations. We decompose alignment functions into three constituents: axis, stride, and offset. For each of these subproblems, we show how to solve the alignment problem for a basic block of code, possibly containing common subexpressions. Alignments are generated for all array objects in the code, both named program variables and intermediate results. We assign computation to processors by virtue of explicit alignment of all temporaries; the resulting work assignment is in general better than that provided by the 'owner-computes' rule. Finally, we present some ideas for dealing with control flow, replication, and dynamic alignments that depend on loop induction variables.

Chatterjee, Siddhartha↗