Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Compiler frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A Full-Stack Exploration of Language-Based Parallelism in Fortran 2023

This poster explores native parallel features in Fortran 2023 through the lens of supporting applications with libraries, compilers, and parallel runtimes. The language revision informally named Fortran 2008 introduced parallelism in the form of Single Program Multiple Data (SPMD) execution with two broad feature sets: (1) loop-level parallelism via do concurrent and (2) a Partitioned Global Address Space (PGAS) comprised of distributed “coarray” data structures. Fortran’s native parallelism has demonstrated high performance [1] and reduced the burden of inserting what sometimes amounts to more directives than code. Several compilers support both feature sets, typically by translating do concurrent into serial do loops annotated by parallel directives and by translating SPMD/PGAS features into direct calls to a communication library. Our research focuses primarily on two questions: (1) can the compiler’s parallel runtime library be developed in the language being compiled (Fortran) and (2) can we define an interface to the runtime that liberates compilers from being hardwired to one runtime and vice versa. We are answering these questions by developing the Parallel Runtime Interface for Fortran (PRIF) [2] and the Co-Array Fortran Framework of Efficient Interfaces to Network Environments (Caffeine) [3]. Caffeine is initially targeting adoption by LLVM Flang, a new open-source Fortran compiler developed by a broad community in industry, academia, and government labs. We are also exploring the use of these features in Inference-Engine, a deep learning library designed to facilitate neural network training and inference for high-performance computing applications written in modern Fortran.

Rasmussen, Katherine↗

The Aerosol Model Benchmarking Repository: A toolkit for model intercomparison

The Aerosol Model Benchmarking Repository and Standards (AMBRS) project was initiated to provide tools and to establish community standards for benchmarking aerosol models. This report describes a set of open-source tools for building, running, and analyzing aerosol box model simulations in a standardized framework. The framework consists of three core components: AMBuilder, a CMake-based build system that compiles supported models consistently; AMBRS, a Python module that defines unified numerical experiments and executes them with aligned inputs; and PyParticle, an aerosol analysis package that standardizes output, computes diagnostics, and visualizes simulation results. Together, these tools enable reproducible intercomparison of aerosol schemes and support process-level evaluation of how model simplifications affect predictions of size distributions, cloud condensation nuclei activity, and other relevant properties relevant for the Earth-Energy system. Beyond its role in benchmarking, AMBRS provides a platform for studying aerosol processes across scales and can be used to generate training data for AI/ML applications in support of a broader hierarchical aerosol modeling strategy.

54 ENVIRONMENTAL SCIENCES↗

A Bayesian approach to evaluation of soil biogeochemical models

Abstract. To make predictions about the carbon cycling consequences of rising global surface temperatures, Earth system scientists rely on mathematical soil biogeochemical models (SBMs). However, it is not clear which models have better predictive accuracy, and a rigorous quantitative approach for comparing and validating the predictions has yet to be established. In this study, we present a Bayesian approach to SBM comparison that can be incorporated into a statistical model selection framework. We compared the fits of linear and nonlinear SBMs to soil respiration data compiled in a recent meta-analysis of soil warming field experiments. Fit quality was quantified using Bayesian goodness-of-fit metrics, including the widely applicable information criterion (WAIC) and leave-one-out cross validation (LOO). We found that the linear model generally outperformed the nonlinear model at fitting the meta-analysis data set. Both WAIC and LOO computed higher overfitting risk and effective numbers of parameters for the nonlinear model compared to the linear model, conditional on the data set. Goodness of fit for both models generally improved when they were initialized with lower and more realistic steady-state soil organic carbon densities. Still, testing whether linear models offer definitively superior predictive performance over nonlinear models on a global scale will require comparisons with additional site-specific data sets of suitable size and dimensionality. Such comparisons can build upon the approach defined in this study to make more rigorous statistical determinations about model accuracy while leveraging emerging data sets, such as those from long-term ecological research experiments.

54 ENVIRONMENTAL SCIENCES↗

Runtime Verification of Hard Realtime Systems With Copilot: A Tutorial

This presentation is a tutorial on RV using Copilot, a runtime verification framework for real-time embedded systems. Copilot monitors are written in a compositional, stream-based language with support for a variety of Temporal Logics (TL), which results in robust, high-level specifications that are easier to understand than their traditional counterparts. The framework translates monitor specifications into C code with static memory requirements, which can be compiled to run on embedded hardware.

runtime monitoring↗

pyCRTM: A Python Interface for the Community Radiative Transfer Model

The Community Radiative Transfer Model (CRTM) is a powerful and versatile scalar radiative transfer model for satellite data assimilation and remote sensing applications. It is implemented as an object-oriented Fortran library, enabling flexible code development and optimal runtime performance on clusters. The downsides of the Fortran interface are a steep learning curve for students and the reduced productivity of users that is typical for static compiled languages, in contrast to dynamic interpreted languages like Python. pyCRTM is a new software framework that directly interfaces the CRTM Fortran data structures and procedures in Python, leveraging both the simplicity and ease of use of Python syntax as well as the flexibility arising from the vast contemporary Python ecosystem. The goal of pyCRTM is to lower the barrier of entry for university students to learn and use the CRTM and to boost the productivity of researchers seeking to create new methods in radiative transfer and data assimilation, or seeking to apply the CRTM to study atmospheric phenomena without having to go through the pre-existing complexity of the CRTM Fortran interface.

Python↗

Automated Generation of Integrated Digital and Spiking Neuromorphic Machine Learning Accelerators

The growing numbers of application areas for artificial intelligence (AI) methods have led to an explosion of domain-specific accelerators that could support every new machine learning (ML) algorithm advancement, clearly highlighting the need for a capability to quickly and automatically transition from algorithm definition to hardware implementation and explore design space along a variety of SWaP (size, weight and Power). The software defined architectures (SODA) synthesizer implements a compiler-based modular infrastructure for the end-to-end generation of machine learning accelerators from high-level frameworks to hardware description language. At the same time, neuromorphic computing, by mimicking how the brain operates, promises to perform artificial intelligence tasks at efficiencies orders of magnitude higher than the current conventional tensor-processing based accelerators, as demonstrated by a variety of specialized designs leveraging Spiking Neural Networks (SNNs). Nevertheless, the mapping of an artificial neural network (ANN) to solutions supporting SNNs is still a non-trivial and very device-specific task, and completely lack the possibility to design hybrid systems that integrate conventional and spiking neural models. In this paper we discuss the support for such an integrated generation leveraging the SODA Synthesizer framework and its modular structure. In particular, we present a new MLIR dialect (part of the SODA frontend) that allows expressing spiking neural network features (e.g., available resources, spiking sequences, analog signal reading, etc.) and illustrate how it enables mapping to Spiking Neurons and deployment to the related specialized hardware (which, in the digital domain, could be generated through the other existing layers of the SODA Synthesizer). We then discuss the opportunities for even deeper integration afforded by the hardware compilation infrastructure, providing a path towards the generation of complex heterogeneous artificial intelligence systems.

Curzel, Serena↗

Development of a Framework and Methodology for an Advanced Reactor Materials Environmental Effects Design Guide

Advanced non-light-water reactor components may operate at elevated temperature while experiencing cyclic loading, significant neutron irradiation, and exposure to reactor coolant. ASME Boiler and Pressure Vessel Code, Section III, Division 5, provides design rules for elevated-temperature service but does not include specific procedures to account for environmental effects on material properties. This report develops an initial framework and methodology for an Environmental Effects Design Guide (EEDG) focused on neutron irradiation; coolant-environment effects are reserved for future work. The proposed approach treats irradiation as a property-based overlay on the existing Division 5 design process, with two routes: a sparse-data route applying two reduction factors — FCR on creep-rupture strength and FF on fatigue life — for the creep-fatigue evaluations that typically control the design of advanced high-temperature reactor components, and a fuller framework developing the property-to-rule chain across the four Division 5 checks (primary load, strain limits and ratcheting, creep-fatigue, and buckling), together with swelling and weldments as scope items. Both routes are scoped by an in-pile qualification that restricts the use of post-irradiation-examination-derived properties in regimes where an in-pile mechanism could control the design outcome. Illustrative outputs derived on a compiled annealed Type 316 database — FCR ≈ 0.78–0.86 and FF ≈ 0.4 — demonstrate the calculation method within that specific dataset. The framework is an initial, testable design-rule concept; it identifies a practical path for preliminary design evaluations under sparse data and the material data and testing needed to develop the framework further.

Barua, Bipul (ORCID:0000000247184113)↗

Consumables and wastes estimations for the First Lunar Outpost

The First Lunar Outpost mission is a design reference mission for the first human return to the moon. This paper describes a set of consumables and waste material estimations made on the basis of the First Lunar Outpost mission scenario developed by the NASA Exploration Programs Office. The study includes the definition of a functional interface framework and a top-level set of consumables and waste materials to be evaluated, the compilation of mass flow information from mission developers supplemented with information from the literature, and the analysis of the resulting mass flow information to gain insight about the possibility of material flow integration between the moon outpost elements. The results of the study of the details of the piloted mission and the habitat are used to identify areas where integration of consumables and wastes across different mission elements could provide possible launch mass savings.

Theis, Ronald L. A.↗

Theoretical frameworks for testing relativistic gravity. IV - A compendium of metric theories of gravity and their post-Newtonian limits.

Metric theories of gravity are compiled and classified according to the types of gravitational fields they contain, and the modes of interaction among those fields. The gravitation theories considered are classified as (1) general relativity, (2) scalar-tensor theories, (3) conformally flat theories, and (4) stratified theories with conformally flat space slices. The post-Newtonian limit of each theory is constructed and its Parametrized Post-Newtonian (PPN) values are obtained by comparing it with Will's version of the formalism. Results obtained here, when combined with experimental data and with recent work by Nordtvedt and Will and by Ni, show that, of all theories thus far examined by our group, the only currently viable ones are general relativity, the Bergmann-Wagoner scalar-tensor theory and its special cases (Nordtvedt; Brans-Dicke-Jordan), and a recent, new vector-tensor theory by Nordtvedt, Hellings, and Will.

Ni, W.-T.↗

Operations on Graphical Models with Plates

This paper explains how graphical models, for instance Bayesian or Markov networks, can be extended to model problems in data analysis and learning. This provides a unified framework that combines lessons learned from the artificial intelligence, statistical and connectionist communities. This also offers a set of principles for developing a software generator for data analysis, whereby a learning or discovery system can be compiled from specifications. Many of the popular learning algorithms can be compiled in this way from graphical specifications. While in a sense this paper is a multidisciplinary review of learning, the main contribution here is the presentation of the material within the unifying framework of graphical models, and the observation that, as a result, the process of developing learning algorithms can be partly automated.

Buntine, Wray L.↗

Using OpenMP for HEP framework algorithm scheduling

The OpenMP standard is the primary mechanism used at high performance computing facilities to allow intra-process parallelization. In contrast, many HEP specific software packages (such as CMSSW, GaudiHive, and ROOT) make use of Intel’s Threading Building Blocks (TBB) library to accomplish the same goal. In these proceedings we will discuss our work to compare TBB and OpenMP when used for scheduling algorithms to be run by a HEP style data processing framework. This includes both scheduling of different interdependent algorithms to be run concurrently as well as scheduling concurrent work within one algorithm. As part of the discussion we present an overview of the OpenMP threading model. We also explain how we used OpenMP when creating a simplified HEP-like processing framework. Using that simplified framework, and a similar one written using TBB, we will present performance comparisons between TBB and different compiler versions of OpenMP.

97 MATHEMATICS AND COMPUTING↗

Simple, Secure, Internet Delivery of MOOSE-based Applications

Application packaging and distribution are the final steps for delivering software to end-users; both are frequently neglected when creating scientific software. Commercial businesses rely on electronic distribution systems that have rendered disk drives obsolete. Still, national laboratories continue to rely heavily on removable media to distribute and limit access to controlled applications. With increasing concerns of unauthorized copying of sensitive applications, a modern distribution system that utilizes cryptographically secure communication and authentication protocols has been developed. This new distribution system will secure the chain of custody for nuclear software while simultaneously simplifying access to these tools. This report summarizes four primary advancements made toward the secure distribution of Nuclear Energy Advanced Modeling and Simulation (NEAMS)-developed, Multiphysics Object Oriented Simulation Environment (MOOSE)-based applications: application installation, package distribution, automated package building, and distribution of documentation. NEAMS is currently developing more than ten separate applications based on the open-source MOOSE Framework. Distribution of these applications has primarily been accomplished by distributing source code, with end-users compiling the applications themselves. This work created a mechanism where MOOSE applications can be installed in a similar way to any other software. This allows both administrators and end-users simplified access to runnable executables. With this new installation capability, it was then possible to rethink distribution. A new, secure capability for delivering MOOSE-based applications over the internet has been created. This system requires unique cryptographic tokens for authentication, greatly securing the custody chain for software. Once granted access, installation of any NEAMS code can be accomplished with these terminal commands: "conda install ncrc" "ncrc install ncrc-bison." After these two commands (and authenticating) the BISON application will be securely down- loaded from Idaho National Laboratory (INL)’s servers, installed, and ready to use. To enable this new distribution capability to be successful, the open-source Continuous Integration, Verification, Enhancement, and Testing (CIVET) Continuous Integration (CI) capability was augmented to add Continuous Delivery (CD). CD enables the automated building and packaging of MOOSE-based applications as they are modified by development teams, ensuring that our customers can obtain up-to-date versions of the software at any time. The need for instruction on how to use these applications was addressed through modifications to the MOOSE documentation system. The MooseDocs capability, which enables robust documentation of MOOSE-based applications, has been extended to allow both for the installation of documentation and the packaging of documentation with installed applications. Together, these enhancements form the core of a new, secure distribution mechanism for nuclear simulation tools. In concert with the Nuclear Computational Resource Center (NCRC), NEAMS- developed applications will now be straightforward to obtain securely.

97 MATHEMATICS AND COMPUTING↗

Flight-Determined Subsonic Lift and Drag Characteristics of Seven Lifting-Body and Wing-Body Reentry Vehicle Configurations With Truncated Bases

This paper examines flight-measured subsonic lift and drag characteristics of seven lifting-body and wing-body reentry vehicle configurations with truncated bases. The seven vehicles are the full-scale M2-F1, M2-F2, HL-10, X-24A, X-24B, and X-15 vehicles and the Space Shuttle prototype. Lift and drag data of the various vehicles are assembled under aerodynamic performance parameters and presented in several analytical and graphical formats. These formats unify the data and allow a greater understanding than studying the vehicles individually allows. Lift-curve slope data are studied with respect to aspect ratio and related to generic wind-tunnel model data and to theory for low-aspect-ratio planforms. The proper definition of reference area was critical for understanding and comparing the lift data. The drag components studied include minimum drag coefficient, lift-related drag, maximum lift-to-drag ratio, and, where available, base pressure coefficients. The effects of fineness ratio on forebody drag were also considered. The influence of forebody drag on afterbody (base) drag at low lift is shown to be related to Hoerner's compilation for body, airfoil, nacelle, and canopy drag. These analyses are intended to provide a useful analytical framework with which to compare and evaluate new vehicle configurations of the same generic family.

Saltzman, Edwin J.↗

Accelerated Charged Particle Tracking with Graph Neural Networks on FPGAs

We develop and study FPGA implementations of algorithms for charged particle tracking based on graph neural networks. The two complementary FPGA designs are based on OpenCL, a framework for writing programs that execute across heterogeneous platforms, and hls4ml, a high-level-synthesis-based compiler for neural network to firmware conversion. We evaluate and compare the resource usage, latency, and tracking performance of our implementations based on a benchmark dataset. We find a considerable speedup over CPU-based execution is possible, potentially enabling such algorithms to be used effectively in future computing workflows and the FPGA-based Level-1 trigger at the CERN Large Hadron Collider.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Tectonic evaluation of the Nubian shield of Northeastern Sudan using thematic mapper imagery

Bechtel is nearing completion of a one-year program that uses digitally enhanced LANDSAT Thematic Mapper (TM) data to compile the first comprehensive regional tectonic map of the Proterozoic Nubian Shield exposed in the northern Red Sea Hills of northeastern Sudan. The status of significant objectives of this study are given. Pertinent published and unpublished geologic literature and maps of the northern Red Sea Hills to establish the geologic framework of the region were reviewed. Thematic mapper imagery for optimal base-map enhancements was processed. Photo mosaics of enhanced images to serve as base maps for compilation of geologic information were completed. Interpretation of TM imagery to define and delineate structural and lithogologic provinces was completed. Geologic information (petrologic, and radiometric data) was compiled from the literature review onto base-map overlays. Evaluation of the tectonic evolution of the Nubian Shield based on the image interpretation and the compiled tectonic maps is continuing.

Source record↗

OpenACC Unified Programming Environment for Multi-hybrid Acceleration with GPU and FPGA

Accelerated computing in HPC such as with GPU, plays a central role in HPC nowadays. However, in some complicated applications with partially different performance behavior is hard to solve with a single type of accelerator where GPU is not the perfect solution in these cases. We are developing a framework and transpiler allowing the users to program the codes with a single notation of OpenACC to be compiled for multi-hybrid accelerators, named MHOAT (Multi-Hybrid OpenACC Translator) for HPC applications. MHOAT parses the original code with directives to identify the target accelerating devices, currently supporting NVIDIA GPU and Intel FPGA, dispatching these specific partial codes to background compilers such as NVIDIA HPC SDK for GPU and OpenARC research compiler for FPGA, then assembles binaries for the final object with FPGA bitstream file. In this paper, we present the concept, design, implementation, and performance evaluation of a practical astrophysics simulation code where we successfully enhanced the performance up to 10 times faster than the GPU-only solution.

Boku, Taisuke↗

NEML2: A High Performance Library for Constitutive Modeling

NEML2, the New Engineering Material model Library, version 2, is an offshoot of NEML, an earlier material modeling code developed at Argonne National Laboratory. NEML2 extends the key philosophy of its predecessor, i.e., material models are flexible, modular, and can be built from smaller blocks. It also provides modern features that do not exist in the framework of its predecessor such as material model vectorization, automatic differentiation, device-portable just-in-time compilation, operator fusion, lazy tensor evaluation, etc. Moreover, NEML2 can seamlessly integrate with the popular machine learning package PyTorch to take advantage of modern and fast-growing machine learning techniques. In this fiscal year, the development of core library features and capabilities are complete. The purpose of this report is not to serve as a verbatim copy of the software API reference (which is available online at https://reverendbedford.github.io/neml2/). Instead, this report documents the motivation, implementation, design choices, and usage of each core capability as well as their applications in solving practical engineering problems. This report is compiled based on the NEML2 major release 2.0.0.

36 MATERIALS SCIENCE↗

AstraAI v1

AstraAI is an open-source, structure-aware AI coding agent designed for large scientific and DOE-HPC codebases such as AMReX-based applications. Unlike general-purpose coding assistants, AstraAI combines retrieval-augmented generation (RAG) with compiler-level Abstract Syntax Tree (AST) analysis to perform precise, scope-constrained code modifications. It identifies exact function spans, enforces locality of edits, and maintains cross-file invariants, enabling deterministic and build-safe transformations in complex C++/GPU environments. AstraAI is intended for developers working on large, evolving HPC frameworks where correctness, reproducibility, and structural integrity are critical. Typical use cases include modifying physics kernels, updating GPU device lambdas, and performing multi-file refactors without breaking compilation or runtime semantics. Compared to conventional LLM-based coding agents - even those with repository access - AstraAI provides structural guarantees rather than free-form text patches. It minimizes unintended diffs, prevents scope drift, preserves formatting and build stability, and reduces structural hallucinations. By integrating compiler tooling directly into the generation loop, AstraAI transforms AI-assisted coding from probabilistic text editing into deterministic, structure-preserving program transformation suitable for mission-critical scientific software.

Natarajan, Mahesh [Lawrence Berkeley National Labo↗