Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compilation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Community Based Data of Uranium Adsorption onto Quartz

This data upload includes a compilation of experiments for quantifying uranium adsorption onto quartz. The provided .csv file has been compiled from the literature in a findable, accessible, interoperable, reusable (FAIR) data format. This was accomplished using the Lawrence Livermore National Laboratory Surface Complexation Database Converter (SCDC) code written in the R programming language (free licensing available at https://ipo.llnl.gov/technologies/software/llnl-surface-complexation-database-converter-scdc). This constitutes all current data compiled on uranium-quartz interactions in the L-SCIE (LLNL Surface Complexation/Ion Exchange) database (as of 02/28/2022). This data was used to develop surface complexation models that fit the global community dataset (https://doi.org/10.1021/acs.est.1c07109). The FAIR-formatted dataset also enables the implementation of alternative machine-learning approaches that can be explored in the future.

54 ENVIRONMENTAL SCIENCES↗

Linking resource availability to pantropical forest canopy resistance and resilience to cyclone disturbance

Statement of purpose: Tropical cyclones are intensifying and occurring at higher latitudes in recent decades, but the mechanisms underpinning the resistance (ability to withstand disturbance-induced change) and resilience (pace of return to pre-disturbance reference values) of tropical forests to cyclones remains largely unexplored at the pantropical scale. We conducted a meta-analysis to investigate the role of soil resource availability (i.e., total soil phosphorus concentration) in mediating site-level forest canopy resistance and resilience to cyclones pan-tropically. We evaluated cyclone-induced and post-cyclone litterfall mass (g/m2/day), phosphorus (P) and nitrogen (N) fluxes (mg/m2/day), as well as concentrations (mg/g) across 73 case studies in Australia, Guadeloupe, Hawaii, Mexico, Puerto Rico, and Taiwan. The dataset zip file includes three data and two metadata files: - The compiled Litterfall Mass Flux data from tropical forests across the globe prior to and after varying tropical cyclone disturbances are provided in Litterfall_Mass.csv. This data file also includes site location, geographical characteristics, elevation, soil phosphorus concentration, geology, and several variables related to each tropical cyclone disturbance. - The compiled Litterfall Nitrogen and Phosphorus Flux data from tropical forests across the globe prior to and after varying tropical cyclone disturbances are provided in Litterfall_Nutrients.csv. This data file also includes site location, geographical characteristics, elevation, soil phosphorus concentration, geology, and several variables related to each tropical cyclone disturbance. - Tropical cyclone track data compiled from HURDAT2 and IBTrACS databases and used as input in the HURRECON model (https://github.com/hurrecon-model/HurreconR) to generate wind data is provided in hurdat2-1851-2019-052520.txt. - The metadata file (Metadata_Meta-analysis_Litterfall-Mass.pdf) has the complete information on each variable included in the Litterfall_Mass.csv dataset, the data sources, and data processing information. - The metadata file (Metadata_Meta-analysis_Litterfall-Nutrients.pdf) has the complete information on each variable included in the Litterfall_Nutrients.csv dataset, the data sources, and data processing information.

54 ENVIRONMENTAL SCIENCES↗

Using MLIR Framework for Codesign of ML Architectures Algorithms and Simulation Tools

MLIR (Multi-Level Intermediate Representation), is an extensible compiler framework that supports high-level data structures and operation constructs. These higher-level code representations are particularly applicable to the artificial intelligence and machine learning (AI/ML) domain, allowing developers to more easily support upcoming heterogeneous AI/ML accelerators and develop flexible domain specific compilers/frameworks with higher-level intermediate representations (IRs) and advanced compiler optimizations. The result of using MLIR within the LLVM compiler framework is expected to yield significant improvement in the quality of generated machine code, which in turn will result in improved performance and hardware efficiency

97 MATHEMATICS AND COMPUTING↗

Unified Language Frontend for Physic-Informed AI/ML

Artificial intelligence and machine learning (AI/ML) are becoming important tools for scientific modeling and simulation as in several other fields such as image analysis and natural language processing. ML techniques can leverage the computing power available in modern systems and reduce the human effort needed to configure experiments, interpret and visualize results, draw conclusions from huge quantities of raw data, and build surrogates for physics based models. Domain scientists in fields like fluid dynamics, microelectronics and chemistry can automate many of their most difficult and repetitive tasks or improve the design times by use of the faster ML-surrogates. However, modern ML and traditional scientific highperformance computing (HPC) tend to use completely different software ecosystems. While ML frameworks like PyTorch and TensorFlow provide Python APIs, most HPC applications and libraries are written in C++. Direct interoperability between the two languages is possible but is tedious and error-prone. In this work, we show that a compiler-based approach can bridge the gap between ML frameworks and scientific software with less developer effort and better efficiency. We use the MLIR (multi-level intermediate representation) ecosystem to compile a pre-trained convolutional neural network (CNN) in PyTorch to freestanding C++ source code in the Kokkos programming model. Kokkos is a programming model widely used in HPC to write portable, shared-memory parallel code that can natively target a variety of CPU and GPU architectures. Our compiler-generated source code can be directly integrated into any Kokkosbased application with no dependencies on Python or cross-language interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Oak Ridge National Laboratory Modernizing the Kokkos Build System: Using CMake to Encapsulate the Complexity of Build Instructions for Performance Portable Libraries

Kokkos, a C++ library focused on performance portability, requires a build system that can work with a variety of compilers and hardware. Ideally, users need only select the compiler and architecture and should not have to know or specify how programs using Kokkos are built. CMake can be used to create a flexible, robust build system and automatically configures compilers and settings based on the user’s inputs. Nevertheless, Kokkos’ requirements as a performance portability library for the build system exceed CMake’s current capabilities. This report describes the requirements, solutions, and testing of various implementations to create a CMake-based build system suitable for Kokkos. It compares the strengths and shortcomings of the approaches and evaluates the implementations with respect to the requirements. Because no solution was found to meet all of the requirements, the Kokkos team engaged with the CMake development team to discuss and plan a path toward support for performance-portable build systems in CMake in the future.

97 MATHEMATICS AND COMPUTING↗

Thermal Tolerance Metrics for Freshwater Fish, CONUS, Version 1

This dataset is a compilation of 13 thermal response metrics for 834 freshwater fish species across the conterminous United States (CONUS). The data were extracted from six published sources, many of which are compilations of data from other sources. The data were harmonized for comparison, and additional variables were added to summarize the metrics. The dataset is presented as a spreadsheet containing 17 sheets. The first sheet (datasets) describes the data sources. Other sheets describe the source and compilation variables in detail.

13 HYDRO ENERGY↗

Exploring code portability solutions for HEP with a particle tracking test code

Traditionally, high energy physics (HEP) experiments have relied on x86 CPUs for the majority of their significant computing needs. As the field looks ahead to the next generation of experiments such as DUNE and the High-Luminosity LHC, the computing demands are expected to increase dramatically. To cope with this increase, it will be necessary to take advantage of all available computing resources, including GPUs from different vendors. A broad landscape of code portability tools—including compiler pragma-based approaches, abstraction libraries, and other tools—allow the same source code to run efficiently on multiple architectures. In this paper, we use a test code taken from a HEP tracking algorithm to compare the performance and experience of implementing different portability solutions. While in several cases portable implementations perform close to the reference code version, we find that the performance varies significantly depending on the details of the implementation. Achieving optimal performance is not easy, even for relatively simple applications such as the test codes considered in this work. Several factors can affect the performance, such as the choice of the memory layout, the memory pinning strategy, and the compiler used. The compilers and tools are being actively developed, so future developments may be critical for their deployment in HEP experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SALT3: An Improved Type Ia Supernova Model for Measuring Cosmic Distances

Abstract A spectral-energy distribution (SED) model for Type Ia supernovae (SNe Ia) is a critical tool for measuring precise and accurate distances across a large redshift range and constraining cosmological parameters. We present an improved model framework, SALT3, which has several advantages over current models—including the leading SALT2 model (SALT2.4). While SALT3 has a similar philosophy, it differs from SALT2 by having improved estimation of uncertainties, better separation of color and light-curve stretch, and a publicly available training code. We present the application of our training method on a cross-calibrated compilation of 1083 SNe with 1207 spectra. Our compilation is 2.5× larger than the SALT2 training sample and has greatly reduced calibration uncertainties. The resulting trained SALT3.K21 model has an extended wavelength range 2000–11,000 Å (1800 Å redder) and reduced uncertainties compared to SALT2, enabling accurate use of low- z I and iz photometric bands. Including these previously discarded bands, SALT3.K21 reduces the Hubble scatter of the low- z Foundation and CfA3 samples by 15% and 10%, respectively. To check for potential systematic uncertainties, we compare distances of low (0.01 < z < 0.2) and high (0.4 < z < 0.6) redshift SNe in the training compilation, finding an insignificant 3 ± 14 mmag shift between SALT2.4 and SALT3.K21. While the SALT3.K21 model was trained on optical data, our method can be used to build a model for rest-frame NIR samples from the Roman Space Telescope. Our open-source training code, public training data, model, and documentation are available at https://saltshaker.readthedocs.io/en/latest/ , and the model is integrated into the sncosmo and SNANA software packages.

79 ASTRONOMY AND ASTROPHYSICS↗

pLiner

Compiler optimizations can alter significantly the numerical results of scientific computing applications. When numerical results differ significantly between compilers, optimization levels, and floating-point hardware, these numerical inconsistencies can impact programming productivity. pLiner is a framework that helps programmers identify locations in the source code that are highly affected by compiler optimizations. pLiner uses a novel approach to identify such code locations by enhancing the floatingpoint precision of variables and expressions. Using a guided search to locate the most significant code regions, pLiner can report to users such locations at different granularities, file, function, and line of code.

Laguna Peralta, Ignacio↗

Static versioning in the polyhedral model

An approach is presented to enhancing the optimization process in a polyhedral compiler by introducing compile-time versioning, i.e., the production of several versions of optimized code under varying assumptions on its run-time parameters. We illustrate this process by enabling versioning in the polyhedral processor placement pass. We propose an efficient code generation method and validate that versioning can be useful in a polyhedral compiler by performing benchmarking on a small set of deep learning layers defined for dynamically-sized tensors.

Meister, Benoit J.↗

Bridging Python to Silicon: The SODA Toolchain

Systems performing scientific computing, data analysis, and machine learning tasks have a growing demand for application-specific accelerators that can provide high computational performance while meeting strict size and power requirements. However, the algorithms and applications that need to be accelerated are evolving at a rate that is incompatible with manual design processes based on hardware description languages. Agile hardware design tools based on compiler techniques can help by quickly producing an application-specific integrated circuit (ASIC) accelerator starting from a high-level algorithmic description. Here, we present the software-defined accelerator (SODA) synthesizer, a modular and open-source hardware compiler that provides automated end-to-end synthesis from high-level software frameworks to ASIC implementation, relying on multilevel representations to progressively lower and optimize the input code. Our approach does not require the application developer to write any register-transfer level code, and it is able to reach up to 364 giga floating point operations per second (GFLOPS)/W efficiency (32-bit precision) on typical convolutional neural network operators.

97 MATHEMATICS AND COMPUTING↗

Optimized Quantum Program Execution Ordering to Mitigate Errors in Simulations of Quantum Systems

Simulating the time evolution of a physical system at quantum mechanical levels of detail - known as Hamiltonian Simulation (HS) - is an important and interesting problem across physics and chemistry. For this task, algorithms that run on quantum computers are known to be exponentially faster than classical algorithms; in fact, this application motivated Feynman to propose the construction of quantum computers. Nonetheless, there are challenges in reaching this performance potential. Prior work has focused on compiling circuits (quantum programs) for HS with the goal of maximizing either accuracy or gate cancellation. Our work proposes a compilation strategy that simultaneously advances both goals. At a high level, we use classical optimizations such as graph coloring and travelling salesperson to order the execution of quantum programs. Specifically, we group together mutually commuting terms in the Hamiltonian (a matrix characterizing the quantum mechanical system) to improve the accuracy of the simulation. We then rearrange the terms within each group to maximize gate cancellation in the final quantum circuit. Furthermore, these optimizations work together to improve HS performance and result in an average 40% reduction in circuit depth. This work advances the frontier of HS which in turn can advance physical and chemical modeling in both basic and applied sciences.

97 MATHEMATICS AND COMPUTING↗

SQUARE: Strategic Quantum Ancilla Reuse for Modular Quantum Programs via Cost-Effective Uncomputation

Compiling high-level quantum programs to machines that are size constrained (i.e. limited number of quantum bits) and time constrained (i.e. limited number of quantum operations) is challenging. In this paper, we present SQUARE (Strategic QUantum Ancilla REuse), a compilation infrastructure that tackles allocation and reclamation of scratch qubits (called ancilla) in modular quantum programs. At its core, SQUARE strategically performs uncomputation to create opportunities for qubit reuse. Current Noisy Intermediate-Scale Quantum (NISQ) computers and forward-looking Fault-Tolerant (FT) quantum computers have fundamentally different constraints such as data locality, instruction parallelism, and communication overhead. Our heuristic-based ancilla-reuse algorithm balances these considerations and fits computations into resource-constrained NISQ or FT quantum machines, throttling parallelism when necessary. To precisely capture the workload of a program, we propose an improved metric, the "active quantum volume," and use this metric to evaluate the effectiveness of our algorithm. Furthermore, our results show that SQUARE improves the average success rate of NISQ applications by 1.47X. Surprisingly, the additional gates for uncomputation create ancilla with better locality, and result in substantially fewer swap gates and less gate noise overall. SQUARE also achieves an average reduction of 1.5X (and up to 9.6X) in active quantum volume for FT machines.

compiler optimization↗

Experiences in porting mini-applications to OpenACC and OpenMP on heterogeneous systems

This article studies mini-applications—Minisweep, GenASiS , GPP, and FF—that use computational methods commonly encountered in HPC. We have ported these applications to develop OpenACC and OpenMP versions, and evaluated their performance on Titan (Cray XK7 with K20x GPUs), Cori (Cray XC40 with Intel KNL), Summit (IBM AC922 with Volta GPUs), and Cori-GPU (Cray CS-Storm 500NX with Intel Skylake and Volta GPUs). Our goals are for these new ports to be useful to both application and compiler developers, to document and describe the lessons learned and the methodology to create optimized OpenMP and OpenACC versions, and to provide a description of possible migration paths between the two specifications. Cases where specific directives or code patterns result in improved performance for a given architecture are highlighted. Here, we also include discussions of the functionality and maturity of the latest compilers available on the above platforms with respect to OpenACC or OpenMP implementations.

97 MATHEMATICS AND COMPUTING↗

Autotuning PolyBench benchmarks with LLVM Clang/Polly loop optimization pragmas using Bayesian optimization

Here, we develop a ytopt autotuning framework that leverages Bayesian optimization to explore the parameter space search and compare four different supervised learning methods within Bayesian optimization and evaluate their effectiveness. We select six of the most complex PolyBench benchmarks and apply the newly developed LLVM Clang/Polly loop optimization pragmas to the benchmarks to optimize them. We then use the autotuning framework to optimize the pragma parameters to improve their performance. The experimental results show that our autotuning approach outperforms the other compiling methods to provide the smallest execution time for the benchmarks syr2k, 3mm, heat-3d, lu, and covariance with two large datasets in 200 code evaluations for effectively searching the parameter spaces with up to 170,368 different configurations. We find that the Floyd-Warshall benchmark did not benefit from autotuning. To cope with this issue, we provide some compiler option solutions to improve the performance. Then we present loop autotuning without a user's knowledge using a simple mctree autotuning framework to further improve the performance of the Floyd-Warshall benchmark. We also extend the ytopt autotuning framework to tune a deep learning application.

79 ASTRONOMY AND ASTROPHYSICS↗

The Kokkos OpenMPTarget Backend: Implementation and Lessons Learned

As the supercomputing landscape diversifies, solutions such as Kokkos to write vendor agnostic applications and libraries have risen in popularity. Kokkos provides a programming model designed for performance portability, which allows developers to write a single source implementation that can run efficiently on various architectures. At its heart, Kokkos maps parallel algorithms to architecture and vendor specific backends written in lower level programming models such as CUDA and HIP. Another approach to writing vendor agnostic parallel code is using OpenMP’s directives based approach, which lets developers annotate code to express parallelism. It is implemented at the compiler level and is supported by all major high performance computing vendors, as well as the primary Open Source toolchains GNU and LLVM. Since its inception, Kokkos has used OpenMP to parallelize on CPU architectures. In this paper, we explore leveraging OpenMP for a GPU backend and discuss the challenges we encountered when mapping the Kokkos APIs and semantics to OpenMP target constructs. As an exemplar workload we chose a simple conjugate gradient solver for sparse matrices. We find that performance on NVIDIA and AMD GPUs varies widely based on details of the implementation strategy and the chosen compiler. Furthermore, the performance of the OpenMP implementations decreases with increasing complexity of the investigated algorithms.

Gayatri, Rahulkumar↗

Domain-Specific Type-Safe APIs for Hierarchical Scientific Data with Modern C++

General-purpose library application programming interfaces (APIs) for self-describing hierarchical scientific data storage, such as the HDF5 and NetCDF libraries, are traditionally of runtime nature. Runtime errors for entry existence and data types are typically caught later in the development process of higher-level application-specific APIs. In this paper, we propose exploiting modern C++ metaprogramming features to add compile-time type-safety to improve the interaction with a well-defined metadata-rich scientific schema in domain-specific hierarchical datasets. We tackle two aspects of common use: (i) direct data access, (ii) flexible “in-memory” index models for efficient search and data processing. The proposed APIs use C++17’s template type auto deduction features, C++11’s enum class for type-safety and C-style preprocessor macros for generative templated code. We showcase the pros and cons of our initial work on the standard NeXus schema used for annotating and storing experimental neutron scattering data at several facilities around the world on top of HDF5. Extendable compile-time type-safe APIs are a desirable feature that could be indexed by any modern integrated development environment (IDE). Hence, such APIs can help ease the learning curve for domain scientists using a less error-prone software interaction to enhance the findability of their data without resorting to a domain-specific language (DSL).

Godoy, William↗

Nitrogen and phosphorus cycling in an ombrotrophic peatland: a benchmark for assessing change

Aims Slow decomposition and isolation from groundwater mean that ombrotrophic peatlands store a large amount of soil carbon (C) but have low availability of nitrogen (N) and phosphorus (P). To better understand the role these limiting nutrients play in determining the C balance of peatland ecosystems, we compile comprehensive N and P budgets for a forested bog in northern Minnesota, USA. Methods N and P within plants, soils, and water are quantified based on field measurements. The resulting empirical dataset are then compared to modern-day, site-level simulations from the peatland land surface version of the Energy Exascale Earth System Model (ELM-SPRUCE).

Salmon, Verity G.↗