Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Compiler frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Variational Optical Phase Learning on a Continuous-Variable Quantum Compiler

Quantum process learning is a fundamental primitive that draws inspiration from machine learning with the goal of better studying the dynamics of quantum systems. One approach to quantum process learning is quantum compilation, whereby an analog quantum operation is digitized by compiling it into a series of basic gates. While there has been significant focus on quantum compiling for discrete-variable systems, the continuous-variable (CV) framework has received comparatively less attention. We present an experimental implementation of a CV quantum compiler that uses two-mode squeezed light to learn a Gaussian unitary operation. We demonstrate the compiler by learning a parameterized linear phase unitary through the use of target and control phase unitaries to demonstrate a factor of 5.4 increase in the precision of the phase estimation and a 3.6-fold acceleration in the time-to-solution metric when leveraging quantum resources. We further show how our approach can be extended to higher-dimensional compilation tasks. Our results are enabled by the tunable control of our cost landscape via variable squeezing, thus providing a critical framework to simultaneously increase precision and reduce time-to-solution.

97 MATHEMATICS AND COMPUTING↗

LISE$^{++}_{cute}$, the latest generation of the LISE ++ package, to simulate rare isotope production with fragment-separators

The LISE ++ software for fragment separator simulations has undergone a major update. The package, widely used at rare isotope beam facilities, can be used to predict intensities and purities of rare isotope beams and for planning and running of experiments using in-flight separators. It is especially useful for radioactive beam production as its results can be quickly compared to on-line data. The LISE ++ package has been ported to the Qt-framework in order to support modern compilers and computing methods. The benefits include 64-bit operation and LISE ++ availability on three different platforms: Windows, MacOS and Linux. In addition, the porting provides the ability to take advantage of future computational improvements. The updated package is named LISE$^{++}_{cute}$ to indicate a major step forward from the previous Borland-based versions. In addition to porting to the new platform, new main features and modifications have been added, mostly devoted to improving models and implementing other codes involved in rare isotope production at FRIB. Finally, a summary of modifications completed to improve the functionality of the code are discussed in this work, as well as future plans.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automatic array alignment in data-parallel programs

FORTRAN 90 and other data-parallel languages express parallelism in the form of operations on data aggregates such as arrays. Misalignment of the operands of an array operation can reduce program performance on a distributed-memory parallel machine by requiring nonlocal data accesses. Determining array alignments that reduce communication is therefore a key issue in compiling such languages. We present a framework for the automatic determination of array alignments in array-based, data-parallel languages. Our language model handles array sectioning, reductions, spreads, transpositions, and masked operations. We decompose alignment functions into three constituents: axis, stride, and offset. For each of these subproblems, we show how to solve the alignment problem for a basic block of code, possibly containing common subexpressions. Alignments are generated for all array objects in the code, both named program variables and intermediate results. We assign computation to processors by virtue of explicit alignment of all temporaries; the resulting work assignment is in general better than that provided by the 'owner-computes' rule. Finally, we present some ideas for dealing with control flow, replication, and dynamic alignments that depend on loop induction variables.

Chatterjee, Siddhartha↗

Fortran Compiler Test Suite v0.1.0

There is no currently available, open-source, easily adapted test suite for checking a compiler's conformance to the Fortran Standard. The main innovation of this software is in the framework to make the test suite apply to a new compiler. Making it open-source will enable community contributions of the test cases. Single organization construction of a comprehensive test suite is something that would be financially infeasible.

Richardson, Bradley↗

Application of community data to surface complexation modeling framework development: Iron oxide protolysis

This study presents a comprehensive community data-driven surface complexation modeling framework for simulating potentiometric titration of mineral surfaces. Compiled community data for ferrihydrite, goethite, hematite, and magnetite are fit to produce representative protolysis constants that can reproduce potentiometric titration data collected from multiple literature sources. Using this framework, the impact of surface complexation model type and surface site density (SSD) on the fit quality and protolysis constants can be readily evaluated. For example, the non-electrostatic model yielded a poor data fit compared to diffuse double layer model and constant capacitance models due to the absence of known surface charge effects. Regardless of the choice of iron oxide mineral, pK a1 decreased with increasing SSD while the opposite tendency was observed for pK a2 . This newly developed framework demonstrates a method to reconcile community data-wide potentiometric titration data using Findable, Accessible, Interoperable, Reusable data principles to produce mineral protolysis constants that improve robustness of surface complexation models for applications in metal sorption and reactive transport modeling. The framework is readily expandable (as community data increase) and extensible (as the number of minerals increase). The framework provides a path forward for developing self-consistent, comprehensive, and updateable surface complexation databases for surface complexation and reactive transport modeling.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

An MLIR-based Compiler Flow for System-Level Design and Hardware Acceleration

The generation of custom hardware accelerators for applications implemented within high-level productive programming frameworks requires considerable manual effort. To automate this process, we introduce \sodaopt, a compiler tool that extends the MLIR infrastructure. \sodaopt automatically searches, outlines, tiles, and pre-optimizes relevant code regions to generate high-quality accelerators through high-level synthesis. \sodaopt can support any high-level programming framework and domain-specific language that interface with the MLIR infrastructure. By leveraging MLIR, \sodaopt solves compiler optimization problems with specialized abstractions. Backend synthesis tools connect to \sodaopt through progressive intermediate representation lowerings. \sodaopt interfaces to a design space exploration engine to identify the combination of compiler optimization passes and options that provides high-performance generated designs for different backends and targets. We demonstrate the practical applicability of the compilation flow by exploring the automatic generation of accelerators for deep neural networks operators outlined at arbitrary granularity and by combining outlining with tiling on large convolution layers. Experimental results with kernels from the PolyBench benchmark show that \sodaopt high-level optimizations improve execution delays of synthesized accelerators up to 60x. We also show that for the selected kernels, our solution outperforms the current of state-of-the art in more than 70% of the benchmarks and provides better average speedup in 55% of them.

Bohm Agostini, Nicolas↗

Unified Language Frontend for Physic-Informed AI/ML

Artificial intelligence and machine learning (AI/ML) are becoming important tools for scientific modeling and simulation as in several other fields such as image analysis and natural language processing. ML techniques can leverage the computing power available in modern systems and reduce the human effort needed to configure experiments, interpret and visualize results, draw conclusions from huge quantities of raw data, and build surrogates for physics based models. Domain scientists in fields like fluid dynamics, microelectronics and chemistry can automate many of their most difficult and repetitive tasks or improve the design times by use of the faster ML-surrogates. However, modern ML and traditional scientific highperformance computing (HPC) tend to use completely different software ecosystems. While ML frameworks like PyTorch and TensorFlow provide Python APIs, most HPC applications and libraries are written in C++. Direct interoperability between the two languages is possible but is tedious and error-prone. In this work, we show that a compiler-based approach can bridge the gap between ML frameworks and scientific software with less developer effort and better efficiency. We use the MLIR (multi-level intermediate representation) ecosystem to compile a pre-trained convolutional neural network (CNN) in PyTorch to freestanding C++ source code in the Kokkos programming model. Kokkos is a programming model widely used in HPC to write portable, shared-memory parallel code that can natively target a variety of CPU and GPU architectures. Our compiler-generated source code can be directly integrated into any Kokkosbased application with no dependencies on Python or cross-language interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Towards Ultra-high-resolution E3SM Land Modeling on Exascale Computers

Here we present an ultra-high-resolution E3SM land model (uELM) for high-fidelity land simulations targeting new Exascale computers. After considering modeling infrastructure compatibility and ELM software features, we designed a parallel model for the uELM development targeting hybrid architectures of new US Exascale computers. We also described a function unit test framework to expedite the piece-wise code porting (with compiler directives), verification, and global variable management. Furthermore, in this study, we report an early uELM model development using OpenACC within a function unit test framework on a pre-Exascale computer, demonstrate the performance of a uLEM submodel with a 3.0-time speedup, and summarize the code porting experience regarding global variable handling, deepcopy, memory reduction, and parallel loop reconstruction.

97 MATHEMATICS AND COMPUTING↗

Run-time parallelization and scheduling of loops

Run time methods are studied to automatically parallelize and schedule iterations of a do loop in certain cases, where compile-time information is inadequate. The methods presented involve execution time preprocessing of the loop. At compile-time, these methods set up the framework for performing a loop dependency analysis. At run time, wave fronts of concurrently executable loop iterations are identified. Using this wavefront information, loop iterations are reordered for increased parallelism. Symbolic transformation rules are used to produce: inspector procedures that perform execution time preprocessing and executors or transformed versions of source code loop structures. These transformed loop structures carry out the calculations planned in the inspector procedures. Performance results are presented from experiments conducted on the Encore Multimax. These results illustrate that run time reordering of loop indices can have a significant impact on performance. Furthermore, the overheads associated with this type of reordering are amortized when the loop is executed several times with the same dependency structure.

Saltz, Joel H.↗

Run-time parallelization and scheduling of loops

Run-time methods are studied to automatically parallelize and schedule iterations of a do loop in certain cases where compile-time information is inadequate. The methods presented involve execution time preprocessing of the loop. At compile-time, these methods set up the framework for performing a loop dependency analysis. At run-time, wavefronts of concurrently executable loop iterations are identified. Using this wavefront information, loop iterations are reordered for increased parallelism. Symbolic transformation rules are used to produce: inspector procedures that perform execution time preprocessing, and executors or transformed versions of source code loop structures. These transformed loop structures carry out the calculations planned in the inspector procedures. Performance results are presented from experiments conducted on the Encore Multimax. These results illustrate that run-time reordering of loop indexes can have a significant impact on performance.

Saltz, Joel H.↗

An Earth-Moon System Trajectory Design Reference Catalog

As demonstrated by ongoing concept designs and the recent ARTEMIS mission, there is, currently, significant interest in exploiting three-body dynamics in the design of trajectories for both robotic and human missions within the Earth-Moon system. The concept of an interactive and 'dynamic' catalog of potential solutions in the Earth-Moon system is explored within this paper and analyzed as a framework to guide trajectory design. Characterizing and compiling periodic and quasi-periodic solutions that exist in the circular restricted three-body problem may offer faster and more efficient strategies for orbit design, while also delivering innovative mission design parameters for further examination.

Libration Orbits↗

Community Based Data of Potentiometric Titration of Iron Oxides: Ferrihydrite (HFO), Goethite, Hematite, Magnetite

This data release includes experimental data of potentiometric titration for iron oxides. The data in the provided .csv files is not our own experimental data but have been compiled from the multiple literature sources. The master database is L-SCIE (LLNL Surface Complexation/Ion Exchange) database, and the provided .csv files are extracted data from L-SCIE. The .csv files were obtained by using the Lawrence Livermore National Laboratory Surface Complexation Database Converter (SCDC) code written in the R programming language (free licensing available at https://ipo.llnl.gov/technologies/software/llnl-surface-complexation-database-converter-scdc).The released data was used for developing a comprehensive community data-driven surface complexation modeling (SCM) framework for simulating potentiometric titration of mineral surfaces. Compiled community data for ferrihydrite, goethite, hematite, and magnetite are fit to produce representative protolysis constants that can reproduce potentiometric titration data collected from multiple literature sources.

54 ENVIRONMENTAL SCIENCES↗

Compiler analysis for irregular problems in FORTRAN D

We developed a dataflow framework which provides a basis for rigorously defining strategies to make use of runtime preprocessing methods for distributed memory multiprocessors. In many programs, several loops access the same off-processor memory locations. Our runtime support gives us a mechanism for tracking and reusing copies of off-processor data. A key aspect of our compiler analysis strategy is to determine when it is safe to reuse copies of off-processor data. Another crucial function of the compiler analysis is to identify situations which allow runtime preprocessing overheads to be amortized. This dataflow analysis will make it possible to effectively use the results of interprocedural analysis in our efforts to reduce interprocessor communication and the need for runtime preprocessing.

Vonhanxleden, Reinhard↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

Ciel

Compiler optimizations can alter the numerical results of scientific computing applications. When numerical results differ significantly between compilers, optimization levels, and floating-point hardware, these numerical inconsistencies can impact programming productivity. Ciel is a framework that helps programmers identify locations in the source code that are affected by compiler optimizations in CPU and GPU code. Ciel uses a floating-point precision enhancement strategy, guided by a recursive bisection search algorithm with increasing search granularity, to identify the program expressions that induce numerical inconsistencies due to compiler optimizations.

Miao, Wenjun↗

Data driven investigation to understand the influence of total solids on biological biogas upgrading

In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.

In situ biogas upgrading↗

Resilience–runtime tradeoff relations for quantum algorithms

Abstract A leading approach to algorithm design aims to minimize the number of operations in an algorithm’s compilation. One intuitively expects that reducing the number of operations may decrease the chance of errors. This paradigm is particularly prevalent in quantum computing, where gates are hard to implement and noise rapidly decreases a quantum computer’s potential to outperform classical computers. Here, we find that minimizing the number of operations in a quantum algorithm can be counterproductive, leading to a noise sensitivity that induces errors when running the algorithm in non-ideal conditions. To show this, we develop a framework to characterize the resilience of an algorithm to perturbative noises (including coherent errors, dephasing, and depolarizing noise). Some compilations of an algorithm can be resilient against certain noise sources while being unstable against other noises. We condense these results into a tradeoff relation between an algorithm’s number of operations and its noise resilience. We also show how this framework can be leveraged to identify compilations of an algorithm that are better suited to withstand certain noises.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Framework for Extensible, Asynchronous Task Scheduling (FEATS) in Fortran

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP, explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI), or compiler-specific language extensions such as those provided by CUDA. By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models. Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data-parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. The paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task-scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

Richardson, Brad↗