Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compilation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

GradDFT. A software library for machine learning enhanced density functional theory

Density functional theory (DFT) stands as a cornerstone method in computational quantum chemistry and materials science due to its remarkable versatility and scalability. Yet, it suffers from limitations in accuracy, particularly when dealing with strongly correlated systems. To address these shortcomings, recent work has begun to explore how machine learning can expand the capabilities of DFT: an endeavor with many open questions and technical challenges. In this work, we present GradDFT a fully differentiable JAX-based DFT library, enabling quick prototyping and experimentation with machine learning-enhanced exchange–correlation energy functionals. GradDFT employs a pioneering parametrization of exchange–correlation functionals constructed using a weighted sum of energy densities, where the weights are determined using neural networks. Moreover, GradDFT encompasses a comprehensive suite of auxiliary functions, notably featuring a just-in-time compilable and fully differentiable self-consistent iterative procedure. To support training and benchmarking efforts, we additionally compile a curated dataset of experimental dissociation energies of dimers, half of which contain transition metal atoms characterized by strong electronic correlations. The software library is tested against experimental results to study the generalization capabilities of a neural functional across potential energy surfaces and atomic species, as well as the effect of training data noise on the resulting model accuracy.

Chemistry↗

Temporal covariation of island arc Sr isotopes and seawater chemistry over the past 2 billion years

The chemical compositions of island arc basalts (IAB) reflect contributions from the mantle as well as fluids and melts from the subducting slab. Addition of radiogenic seawater Sr to oceanic crust through hydrothermal alteration and subsequent subduction is often invoked to explain elevated 87 Sr/ 86 Sr signatures in modern IAB. However, changes in the 87 Sr/ 86 Sr of island arc magmatic rocks through time has not been investigated, limiting our understanding of the factors influencing the Sr budgets of arcs throughout Earth’s history. To address this, we compiled 87 Sr/ 86 Sr values from island arc magmatic rocks ranging in age from modern to Paleoproterozoic, only including data from island arc localities that best preserve initial magmatic 87 Sr/ 86 Sr. Median initial 87 Sr/ 86 Sr values are consistently elevated compared to depleted mantle 87 Sr/ 86 Sr over this period, indicating persistent enrichment in radiogenic Sr in island arcs. Moreover, the elevation in island arc 87 Sr/ 86 Sr relative to the depleted mantle is variable. A notable rise in island arc 87 Sr/ 86 Sr during the late Neoproterozoic coincides with a steep increase in seawater 87 Sr/ 86 Sr and Sr concentration. To investigate this potential connectivity, we modeled the 87 Sr/ 86 Sr of island arc magmas between 0 and 830 Ma with inputs of depleted mantle 87 Sr/ 86 Sr, seawater 87 Sr/ 86 Sr, and seawater Sr concentration. The model reproduces the overall trajectory of the compiled data. We interpret the observed temporal variation in island arc 87 Sr/ 86 Sr values and its close association with fluctuations in seawater chemistry as evidence that changes in marine geochemistry have strongly influenced the Sr isotopic record of island arc magmas over time.

Science & Technology - Other Topics↗

HPC-driven computational reproducibility in numerical relativity codes: a use case study with IllinoisGRMHD

Abstract Reproducibility of results is a cornerstone of the scientific method. Scientific computing encounters two challenges when aiming for this goal. Firstly, reproducibility should not depend on details of the runtime environment, such as the compiler version or computing environment, so results are verifiable by third-parties. Secondly, different versions of software code executed in the same runtime environment should produceconsistent numerical results for physical quantities. In this manuscript, we test the feasibility of reproducing scientific results obtained using theIllinoisGRMHDcode that is part of an open-source community software for simulation in relativistic astrophysics, theEinstein Toolkit. We verify that numerical results of simulating a single isolated neutron star withIllinoisGRMHDcan be reproduced, and compare them to results reported by the code authors in 2015. We use two different supercomputers: Expanse at SDSC, and Stampede2 at TACC. By compiling the source code archived along with the paper on both Expanse and Stampede2, we find thatIllinoisGRMHDreproduces results published in its announcement paper up to errors comparable to round-off level changes in initial data parameters. We also verify that a current version ofIllinoisGRMHDreproduces these results once we account for bug fixes which have occurred since the original publication.

Astronomy & Astrophysics↗

XACC: a system-level software infrastructure for heterogeneous quantum–classical computing

Quantum programming techniques and software have advanced significantly over the past five years, with a majority focusing on high-level language frameworks targeting remote REST library APIs. As quantum computing architectures advance and become more widely available, lower-level, system software infrastructures will be needed to enable tighter, co-processor programming and access models. In this work, we present XACC, a system-level software infrastructure for quantum–classical computing that promotes a service-oriented architecture to expose interfaces for core quantum programming, compilation, and execution tasks. Additionally, we detail XACC's interfaces, their interactions, and its implementation as a hardware-agnostic framework for both near-term and future quantum–classical architectures. We provide concrete examples demonstrating the utility of this framework with paradigmatic tasks. Our approach lays the foundation for the development of compilers, associated runtimes, and low-level system tools tightly integrating quantum and classical workflows.

97 MATHEMATICS AND COMPUTING↗

Shearing approach to gauge-invariant Trotterization

Universal quantum simulations of gauge field theories are exposed to the risk of gauge symmetry violations when it is not known how to compile the desired operations exactly using the available gate set. In this article, we show how time evolution can be compiled in an Abelian gauge theory—if only approximately—without compromising gauge invariance, by graphically motivating a block-diagonalization procedure. When gauge-invariant interactions are associated with a “spatial network” in the space of discrete quantum numbers, it is seen that cyclically shearing the spatial network converts simultaneous updates to many quantum numbers into conditional updates of a single quantum number; ultimately, this eliminates any need to pass through (and acquire overlap onto) unphysical intermediate configurations. Shearing is explicitly applied to gauge-matter and magnetic interactions of lattice quantum electrodynamics. The features that make shearing successful at preserving Abelian gauge symmetry may also be found in non-Abelian theories, bringing one closer to gauge-invariant simulations of quantum chromodynamics.

Gauge theories↗

Current data are consistent with flat spatial hypersurfaces in the Λ CDM cosmological model but favor more lensing than the model predicts

Here, we study the performance of three pairs of tilted, and a pair of untilited, ΛCDM cosmological models, with three of these four pairs allowing for non-flat spatial hypersurfaces, against cosmic microwave background (CMB) temperature and polarization power spectrum data (P18), measurements of the Planck 2018 lensing potential power spectrum (lensing), and a large compilation of non-CMB data (non-CMB). For the eight models, we measure cosmological parameters and study whether or not pairs of the data sets (as well as subsets of them) are mutually consistent in these models. Half of these models allow the lensing consistency parameter A L , which re-scales the gravitational potential power spectrum, to be an additional free parameter to be determined from data, while the other three have A L = 1 which is the theoretically expected value. The pair of untilted non-flat ΛCDM models are incompatible with P18 data. The tilted spatially-flat models assume the usual primordial spatial inhomogeneity power spectrum that is a power law in wave number. The tilted non-flat models assume either the primordial power spectrum used in the Planck group anal yses [Planck P(q)], that has recently been numerically shown to be a good approximation to what is quantum-mechanically generated from a particular choice of closed inflation model initial conditions, or a recently computed power spectrum [new P(q)] that quantum-mechanically follows from a different set of non-flat inflation model initial conditions. In the tilted non-flat models with A L = 1 we find differences between P18 data and non-CMB data cosmological parameter constraints, which are large enough to rule out the Planck P(q) model at 3σ but not the new P(q) model. No significant differences are found when cosmological parameter constraints obtained with two different data sets are compared within the standard tilted flat ΛCDM model. While both P18 data and non-CMB data separately favor a closed geometry, with spatial curvature density parameter Ω k < 0, when P18+non-CMB data are jointly analyzed the evidence in favor of non-flat hypersurfaces subsides. Differences between P18 data and non-CMB data cosmological constraints subside when A L is allowed to vary. From the most restrictive P18+lensing+non-CMB data combination we get almost model-independent constraints on the cosmological parameters and find that the A L > 1 option is preferred over the Ω k < 0 one, with the A L parameter, for all models, being larger than unity by ~ 2.5σ. According to the deviance information criterion, in the P18+lensing+non-CMB analysis, the varying A L option is on the verge of being strongly favored over the A L = 1 one, which could indicate a problem for the standard tilted flat ΛCDM model. These data are consistent with flat spatial hypersurfaces but more and better data could improve the constraints on Ω k and might alter this conclusion. Error bars on some cosmological parameters are significantly reduced when non-CMB data are used jointly with P18+lensing data. For example, in the tilted flat ΛCDM model for P18+lensing+non-CMB data the Hubble constant H 0 = 68.09 ± 0.38 km s -1 Mpc -1 , which is consistent with that from a median statistics analysis of a large compilation of H 0 measurements as well as with a number of local measurements of the cosmological expansion rate. This H 0 error bar is 31% smaller than that from P18+lensing data alone.

79 ASTRONOMY AND ASTROPHYSICS↗

AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators

Tensor algebra operations represent an important class of algorithms used across many applications, including machine learning, scientific computing, and data analytics. As a result, the efficient generation of custom accelerators for tensor operations has received increased attention. Previous efforts have produced automated tools enabling users to prototype and explore optimized accelerators. However, little effort has been focused on the host-accelerator interaction in these tools. Efficient use of hardware accelerators requires knowledge about the accelerator's capabilities (operations, data formats, and opcode support), the host CPU microarchitecture (e.g., memory hierarchy), the host-accelerator interface, and the application's features (which code regions should be mapped onto an accelerator). Manually rewriting the original applications to facilitate improved custom accelerator mapping is an error-prone and time-consuming endeavor. To cope with this, we propose AXI4MLIR, a new framework to automatically generate and optimize the communication between the host CPU and arbitrary accelerators that implement linear algebra algorithms. AXI4MLIR extends the MLIR compiler framework to automatically generate efficient host-accelerator driver code for accelerators with AXI-based interfaces. Our compiler extensions enable automatic driver code generation while carefully considering the host's memory hierarchy and target accelerator features. To demonstrate the flexibility and utility of AXI4MLIR, we test it with diverse use cases that include different types of accelerators, tiling scenarios, and dataflow schemes. We compare our experimental results to manual implementations of host-accelerator driver code and find that our approach can reduce CPU cache references by 56% and deliver up to a 1.65x speedup.

Bohm Agostini, Nicolas↗

Invited: Bambu: an Open-Source Research Framework for the High-Level Synthesis of Complex Applications

This paper presents the open-source High-Level Synthesis research framework Bambu. The framework provides an open-source starting point to experiment with new ideas across High-Level Synthesis, high-level verification and debugging, FPGA/ASIC design, design flow space exploration, and parallel hardware accelerator design. The tool accepts as input standard C/C++ specifications and compiler intermediate representations (IRs) coming from the well-known Clang/LLVM and GCC com- pilers. The broad spectrum and flexibility of input formats allow the electronic design automation (EDA) research community to explore and integrate new transformations and optimizations. The easily extendable modular framework already includes many op- timizations and HLS benchmarks. The integration with synthesis and verification backends (commercial and open-source) allows researchers to quickly test any new finding and easily obtain performance and resource usage metrics for a given application. Different FPGA devices are supported from several different vendors: AMD/XILINX, Intel/Altera, Lattice Semiconductor, and NanoXplore. Finally, integration with the OpenRoad open-source end-to-end silicon compiler perfectly fits with the recent push towards open-source EDA.

Ferrandi, Fabrizio↗

Automated Generation of Integrated Digital and Spiking Neuromorphic Machine Learning Accelerators

The growing numbers of application areas for artificial intelligence (AI) methods have led to an explosion of domain-specific accelerators that could support every new machine learning (ML) algorithm advancement, clearly highlighting the need for a capability to quickly and automatically transition from algorithm definition to hardware implementation and explore design space along a variety of SWaP (size, weight and Power). The software defined architectures (SODA) synthesizer implements a compiler-based modular infrastructure for the end-to-end generation of machine learning accelerators from high-level frameworks to hardware description language. At the same time, neuromorphic computing, by mimicking how the brain operates, promises to perform artificial intelligence tasks at efficiencies orders of magnitude higher than the current conventional tensor-processing based accelerators, as demonstrated by a variety of specialized designs leveraging Spiking Neural Networks (SNNs). Nevertheless, the mapping of an artificial neural network (ANN) to solutions supporting SNNs is still a non-trivial and very device-specific task, and completely lack the possibility to design hybrid systems that integrate conventional and spiking neural models. In this paper we discuss the support for such an integrated generation leveraging the SODA Synthesizer framework and its modular structure. In particular, we present a new MLIR dialect (part of the SODA frontend) that allows expressing spiking neural network features (e.g., available resources, spiking sequences, analog signal reading, etc.) and illustrate how it enables mapping to Spiking Neurons and deployment to the related specialized hardware (which, in the digital domain, could be generated through the other existing layers of the SODA Synthesizer). We then discuss the opportunities for even deeper integration afforded by the hardware compilation infrastructure, providing a path towards the generation of complex heterogeneous artificial intelligence systems.

Curzel, Serena↗

Computing with a Chemical Reservoir

Contemporary computation is expensive, with large language models and artificial intelligence becoming more common in daily life. However, high-performance computing is reaching the limits in speed and energy expenditure, and domain science requires ever-increasing computational capacity, with simulations and data analysis pipelines ever-growing in complexity. As we progress towards post-exascale computation, with the associated high energy costs, new methods of energy-conscious computation are required. Novel analog and hybrid digital-analog systems can overcome these challenges, and chemical reactions offer a promising avenue. Computers based on chemistry can provide compact desktop devices with immense computational power. These devices are readily scalable by considering greater reaction systems or vessels, meeting the high-performance requirements for scientific workflows. In this article, we present ChemComp, a compilation pipeline for the conversion of ordinary differential equations into implementable chemical reactions. We then demonstrate the solving capabilities of ChemComp by emulating a potential chemical reservoir device. We leverage the multi-layer intermediate representation (MLIR) compiler framework to implement an expressive chemical reaction abstraction and propose a path for chemical reaction networks (CRNs) to represent mathematical problems effectively. Combined, we demonstrate a potential workflow that can harness chemistry’s computing power to create energy-efficient, high-performance computation systems for contemporary computing needs.

artificial intelligence↗

Integer Sum Reduction with OpenMP on an AMD MI100 GPU

Sum reduction is a primitive operation in parallel computing. Device offload support allows a user to use OpenMP directives to take advantage of a highly capable GPU. In this paper, we present the integer sum reduction annotated with the OpenMP directives and evaluate the performance impacts of tunable parameters with the AOMP and GCC compilers on an AMD MI100 GPU. In addition, we explain the implementations of the OpenMP reduction by the compilers. Sweeping over the pruned parameter space, we find that the speedup is approximately 20 with AOMP, and the reduction performance using AOMP is approximately 11% higher than that using GCC. However, the OpenMP offload performance is approximately 30% lower compared to the performance of the reductions written with rocThrust or hipCUB.

Jin, Zheming↗

Implementing Directive-Based Deferred Execution for Effective Network Aggregation

Remote direct memory access technology provides an efficient mechanism for one-sided communication that can be leveraged to implement a distributed shared memory programming model. However, when applications generate large numbers of small, irregular messages, network congestion often arises. Existing solutions address this small message problem by facilitating message aggregation but typically require disruptive code transformations that detract from the algorithmic intent of applications, or can be limited by dependent operations on aggregated data between synchronisation points. A solution is to use a directive-assisted approach that enables compilers to transform code dependent on aggregated communication for deferred execution. This paper presents an algorithm that a compiler can use to implement and optimise deferred execution for code dependent on aggregated data, based on an "aggregation context" extension for the OpenSHMEM partitioned global address space library. This new capability addresses a key challenge of message aggregation, allowing its full potential to reduce network congestion and enhance programmability to be realised.

Welch, Aaron [ORNL]↗

Really Embedding Domain-Specific Languages into C++

The following topics are dealt with: program compilers; optimising compilers; parallel processing; software engineering; learning (artificial intelligence); multiprocessing systems; shared memory systems; optimisation; computational complexity; specification languages.

Finkel, Hal J.↗

Sampling on NISQ Devices: "Who’s the Fairest One of All?"

Modern NISQ devices are subject to a variety of biases and sources of noise that degrade the solution quality of computations carried out on these devices. A natural question that arises in the NISQ era, is how fairly do these devices sample ground state solutions. To this end, we run five fair sampling problems (each with at least three ground state solutions) that are based both on quantum annealing and, on the Grover Mixer, -QAOA algorithm for gate-based NISQ hardware. In particular, we use seven IBM Q devices, the Aspen-9 Rigetti device, the IonQ device, and three D-Wave quantum annealers. For each of the fair sampling problems, we measure the ground state probability, the relative fairness of the frequency of each ground state solution with respect to the other ground state solutions, and the aggregate error as given by each hardware provider. Overall, our results show that NISQ devices do not achieve fair sampling yet. Furthermore, we also observe differences in the software stack with a particular focus on compilation techniques that illustrate what work will still need to be done to achieve a seamless integration of frontend (i.e., quantum circuit description) and backend compilation.

Computer Science↗

SPEL: Software tool for Porting E3SM Land Model with OpenACC in a Function Unit Test Framework

Most high-end computers adopt hybrid architecture, porting a large-scale scientific code onto accelerators is necessary. The paper presents a generic method for porting large-scale scientific code onto accelerators using compiler directives within a modularized function unit test platform. We have implemented the method and designed a software tool (SPEL) to port the E3SM Land Model (ELM) onto the GPUs in the Summit computer. SPEL automatically generates GPU-ready test modules for all ELM functions, such as CanopyFlux, SoilTemperature, and EcosystemDynamics. SPEL breaks the ELM into a collection of standalone unit test programs for easy code verification and further performance improvement. We further optimize several ELM test modules with advanced techniques, including memory reduction, reconstructed parallel loops, and asynchronous GPU kernel launch. We hope our study will inspire new toolkit developments that expedite large-scale scientific code porting with compiler directives.

Schwartz, Peter↗

Statistical upscaling of ecosystem CO 2 fluxes across the terrestrial tundra and boreal domain: Regional patterns and uncertainties

Abstract The regional variability in tundra and boreal carbon dioxide (CO 2 ) fluxes can be high, complicating efforts to quantify sink‐source patterns across the entire region. Statistical models are increasingly used to predict (i.e., upscale) CO 2 fluxes across large spatial domains, but the reliability of different modeling techniques, each with different specifications and assumptions, has not been assessed in detail. Here, we compile eddy covariance and chamber measurements of annual and growing season CO 2 fluxes of gross primary productivity (GPP), ecosystem respiration (ER), and net ecosystem exchange (NEE) during 1990–2015 from 148 terrestrial high‐latitude (i.e., tundra and boreal) sites to analyze the spatial patterns and drivers of CO 2 fluxes and test the accuracy and uncertainty of different statistical models. CO 2 fluxes were upscaled at relatively high spatial resolution (1 km 2 ) across the high‐latitude region using five commonly used statistical models and their ensemble, that is, the median of all five models, using climatic, vegetation, and soil predictors. We found the performance of machine learning and ensemble predictions to outperform traditional regression methods. We also found the predictive performance of NEE‐focused models to be low, relative to models predicting GPP and ER. Our data compilation and ensemble predictions showed that CO 2 sink strength was larger in the boreal biome (observed and predicted average annual NEE −46 and −29 g C m −2 yr −1 , respectively) compared to tundra (average annual NEE +10 and −2 g C m −2 yr −1 ). This pattern was associated with large spatial variability, reflecting local heterogeneity in soil organic carbon stocks, climate, and vegetation productivity. The terrestrial ecosystem CO 2 budget, estimated using the annual NEE ensemble prediction, suggests the high‐latitude region was on average an annual CO 2 sink during 1990–2015, although uncertainty remains high.

Virkkala, Anna‐Maria↗

Predicting nepheline precipitation in waste glasses using ternary submixture model and machine learning

Nepheline precipitation in nuclear waste glasses during vitrification can be detrimental due to its negative effect on chemical durability. Developing models to accurately predict nepheline precipitation from compositions is important to increase waste loading since existing models can be overly conservative. In this study, an expanded dataset containing 955 glasses was compiled from literature data, where 355 glasses are for high-level waste (HLW). Previously developed submixture models were refitted using the new dataset, where a misclassification rate of 7.8% was achieved. Nine machine learning (ML) algorithms (e.g., k-nearest neighbor, Gaussian process regression, artificial neural network, support vector machine, decision tree, etc.) were applied to evaluate their ability of predicting nepheline precipitation from compositions. Model accuracy, precision, recall/sensitivity, and F1 score were systemically compared between different ML algorithms and modeling protocols. Good model prediction with an accuracy ~0.9 (misclassification rate of ~10%) was observed with different algorithms under certain protocol. This study evaluated various ML models to predict nepheline precipitations in waste glasses, highlighting the importance of data preparation, modeling protocol, and their effect on model stability and reproducibility. The results provide insights into applying ML to predict glass properties and suggest areas for future research on modeling nepheline precipitations.

Lu, Xiaonan↗

A Critical Review of Heliostat Design for Concentrating Solar Thermal Technologies

Heliostat designs have undergone a widespread and eclectic development process, with many unique designs demonstrated. However, recent developments do not show the overall cohort converging toward a globally accepted universal design. Here, this study characterizes heliostats by breaking down and evaluating design traits based on emergent patterns from a comprehensive compilation of known heliostat designs spanning several decades. Four main categories for evaluation emerged: heliostat base, heliostat primary axis, heliostat drive, and facet support. Each of these four categories is further defined by four subtypes so that all heliostats fall into a single subtype within each category. The classified heliostats are ranked, yielding several view slices into the heliostat compilation or a breakdown of heliostats by type. An analysis of the breakdown shows several trends: a scale-up of established designs, new approaches at small to medium scales, and a movement toward greater adoption of linear drives. These trends reflect the most meaningful contributing factors to a proposed trajectory for a new era of heliostat designs striving to meet widely considered cost targets of $\$$50/m 2 or $\$$75/m 2 .

14 SOLAR ENERGY↗