Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “source code expressions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Do Programmers Prefer Predictable Expressions in Code?

Source code is a form of human communication, albeit one where the information shared between the programmers reading and writing the code is constrained by the requirement that the code executes correctly. Programming languages are more syntactically constrained than natural languages, but they are also very expressive, allowing a great many different ways to express even very simple computations. Still, code written by developers is highly predictable, and many programming tools have taken advantage of this phenomenon, relying on language model surprisal as a guiding mechanism. Additionally, while surprisal has been validated as a measure of cognitive load in natural language, its relation to human cognitive processes in code is still poorly understood. In this paper, we explore the relationship between surprisal and programmer preference at a small granularity—do programmers prefer more predictable expressions in code? Using meaning-preserving transformations, we produce equivalent alternatives to developer-written code expressions and run a corpus study on Java and Python projects. In general, language models rate the code expressions developers choose to write as more predictable than these transformed alternatives. Then, we perform two human subject studies asking participants to choose between two equivalent snippets of Java code with different surprisal scores (one original and transformed). We find that programmers do prefer more predictable variants, and that stronger language models like the transformer align more often and more consistently with these preferences.

97 MATHEMATICS AND COMPUTING↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

RASPA3

RASPA3, a molecular simulation code for computing adsorption and diffusion in nanoporous materials and thermodynamic and transport properties of fluids. It implements force field based classical Monte Carlo/molecular dynamics in various ensembles. RASPA3 is rewritten from the ground up in C++23 with speed and code readability in mind. Transition-matrix Monte Carlo is added to compute the density of states and free energies. The Monte Carlo code for rigid molecules is based on quaternions, and the atomic positions needed in the energy evaluation are recreated from the center of mass position and quaternion orientation. The expanded ensemble methodology for fractional molecules, with a scaling parameter λ between 0 and 1, now also keeps track of analytic expressions of dU/dλ, allowing independent verification of the chemical potential using thermodynamic integration. The source code is freely available under the MIT license on GitHub.

Dubbeldam, David↗

CodeFlow: A Code Generation System for Flash-X Orchestration Runtime

We propose the CodeFlow toolchain for Flash-X that realizes the “recipe-to-source” code transformation for Flash-X simulations and that is necessary to achive performance portability. We design a high-level language to express operations of simulations in so-called recipes, which are given as input to the toolchain. The tools of the CodeFlow pipeline include code transformation with tree-based source code representation techniques and code orchestration and generation based on control flow graphs. The generated source code utilizes a new runtime, developed for Flash-X, that orchestrates dynamic and asynchronous data movement and task execution. The functionality of CodeFlow is demonstrated using a hydrodynamic problem with a strong shock.

97 MATHEMATICS AND COMPUTING↗

The Kokkos OpenMPTarget Backend: Implementation and Lessons Learned

As the supercomputing landscape diversifies, solutions such as Kokkos to write vendor agnostic applications and libraries have risen in popularity. Kokkos provides a programming model designed for performance portability, which allows developers to write a single source implementation that can run efficiently on various architectures. At its heart, Kokkos maps parallel algorithms to architecture and vendor specific backends written in lower level programming models such as CUDA and HIP. Another approach to writing vendor agnostic parallel code is using OpenMP’s directives based approach, which lets developers annotate code to express parallelism. It is implemented at the compiler level and is supported by all major high performance computing vendors, as well as the primary Open Source toolchains GNU and LLVM. Since its inception, Kokkos has used OpenMP to parallelize on CPU architectures. In this paper, we explore leveraging OpenMP for a GPU backend and discuss the challenges we encountered when mapping the Kokkos APIs and semantics to OpenMP target constructs. As an exemplar workload we chose a simple conjugate gradient solver for sparse matrices. We find that performance on NVIDIA and AMD GPUs varies widely based on details of the implementation strategy and the chosen compiler. Furthermore, the performance of the OpenMP implementations decreases with increasing complexity of the investigated algorithms.

Gayatri, Rahulkumar↗

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING↗

GNET2: an R package for constructing gene regulatory networks from transcriptomic data

Abstract Motivation The Gene Network Estimation Tool (GNET) is designed to build gene regulatory networks (GRNs) from transcriptomic gene expression data with a probabilistic graphical model. The data preprocessing, model construction and visualization modules of the original GNET software were developed on different programming platforms, which were inconvenient for users to deploy and use. Results Here, we present GNET2, an improved implementation of GNET as an integrated R package. GNET2 provides more flexibility for parameter initialization and regulatory module construction based on the core iterative modeling process of the original algorithm. The data exchange interface of GNET2 is handled within an R session automatically. Given the growing demand for regulatory network reconstruction from transcriptomic data, GNET2 offers a convenient option for GRN inference on large datasets. Availability and implementation The source code of GNET2 is available at https://github.com/jianlin-cheng/GNET2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Human Liver Epithelium Response to HCoV-229E Infection Epigenomics (ACS-DP4)

The purpose of this experiment was to evaluate how wild-type Human coronavirus strain 229E (HCoV-229E) infection alters chromatin accessibility in infected cells only. Sample data was obtained for mock and infected (standard and UV-inactivated) immortalized human liver cells (HuH-7) and collected 24 hrs. post infection. Samples were processed using assay for transposase-accessible chromatin using high-throughput sequencing (ATAC-Seq) and generated bar coded library samples were evaluated for RNA sequencing (RNA-Seq) expression analysis. Processed ATAC-Seq datasets are openly accessible from the download button and contain secondary processed RNA-Seq results files and supporting metadata materials. Data download includes a sample naming key, infection titer metadata, normalized counts, and relevant computational source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

Mixed-Linkage Glucan Is the Main Carbohydrate Source and Starch Is an Alternative Source during Brachypodium Grain Germination

Seeds of the model grass Brachypodium distachyon are unusual because they contain very little starch and high levels of mixed-linkage glucan (MLG) accumulated in thick cell walls. It was suggested that MLG might supplement starch as a storage carbohydrate and may be mobilised during germination. In this work, we observed massive degradation of MLG during germination in both endosperm and nucellar epidermis. The enzymes responsible for the MLG degradation were identified in germinated grains and characterized using heterologous expression. By using mutants targeting MLG biosynthesis genes, we showed that the expression level of genes coding for MLG and starch-degrading enzymes was modified in the germinated grains of knocked-out cslf6 mutants depleted in MLG but with higher starch content. Our results suggest a substrate-dependent regulation of the storage sugars during germination. These overall results demonstrated the function of MLG as the main carbohydrate source during germination of Brachypodium grain. More astonishingly, cslf6 Brachypodium mutants are able to adapt their metabolism to the lack of MLG by modifying the energy source for germination and the expression of genes dedicated for its use.

59 BASIC BIOLOGICAL SCIENCES↗

Ciel

Compiler optimizations can alter the numerical results of scientific computing applications. When numerical results differ significantly between compilers, optimization levels, and floating-point hardware, these numerical inconsistencies can impact programming productivity. Ciel is a framework that helps programmers identify locations in the source code that are affected by compiler optimizations in CPU and GPU code. Ciel uses a floating-point precision enhancement strategy, guided by a recursive bisection search algorithm with increasing search granularity, to identify the program expressions that induce numerical inconsistencies due to compiler optimizations.

Miao, Wenjun↗

Evaluation of gamma-ray transmission through rectangular collimator slits for application in nuclear fuel spectrometry

Gamma-ray spectrometry is widely applied in several science fields, and in particular in non-destructive gamma scanning and gamma emission tomography of irradiated nuclear fuel. Usually, a collimator is used in the experimental setup, to selectively interrogate a region of interest in the fuel. For the optimization of instrument design, as well as for planning measurement campaigns, predictive models for the transmitted gamma-ray intensity through the collimator are needed. Commonly, Monte Carlo Radiation Transport tools are used for accurate prediction of gamma-ray transport, however, the long computation time requirements when used in low-efficiency experimental setups present challenges. Here, the full-energy peak intensity transmitted through a rectangular collimator slit was examined. A uniform planar surface source emitting isotropically was considered, and the rate of photons reaching an ideal counter plane on the opposite side of the collimator was evaluated by analytical integration. To find a closed-form primitive function, some idealizations were required, and thereby parametric models were obtained for the optical field of view, dependent on slit dimensions (length, height and width) and source-to-collimator distance. For contributions from outside the optical field of view, where a closed-form expression cannot be found, fast numerical integral methods were instead used. The results were validated using the Monte Carlo code MCNP6 and show an agreement within three percent for the numerical method. For the analytical method, deviations up to tens of percent were obtained, which is deemed to still be sufficient for instrument design and measurement planning, where often the order of magnitude of the count rate is not a priori known. The method is planned for use in iterative optimization routines in the design of Gamma Emission Tomography devices, as well as for the prediction of gamma spectra obtained in the planning of fuel inspections. An application of the proposed method was demonstrated in spectrum prediction for a short cooling-time fuel rod test from the Halden reactor.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

iSPECTRON: a simulation interface for linear and nonlinear spectra with ab-initio quantum chemistry software

We introduce iSPECTRON, an open source (under the Educational Community License version 2.0) program that parses data from common quantum chemistry software (NWChem, OpenMolcas, Gaussian, Cobramm, etc.), produces the input files for the simulation of linear and nonlinear spectroscopy of molecules with the Spectron code, and analyzes the spectra with a broad range of tools. Vibronic spectra are expressed in term of the electronic eigenstates, obtained through quantum chemistry computations, and vibrational/bath effects are incorporated in the framework of the displaced harmonic oscillator model, where all required quantities are computed at the Franck-Condon point. The code capabilities are illustrated by simulating linear absorption, transient absorption and two dimensional electronic spectra of the pyrene molecule. Two levels of electronic structure theory, TDDFT (with NWChem) and RASSCF/RASPT2 (with OpenMolcas), are compared where possible. Acknowledgements: F.S., A.N., D.R.N., N.G., S.M, M.G. acknowledge support from the U.S. Department of Energy, Office of Science, Office of Basic Energy Sciences, Chemical Sciences, Geosciences, and Biosciences Division under Award Nos. DE-SC0019484, KC-030103172684. The Spectron code was developed with support from the National Science Foundation (Grant CHE- 1953045). This research benefited from computational resources provided by EMSL, a DOE Office of Science User Facility sponsored by the Office of Biological and Environmental Research and located at PNNL. PNNL is operated by Battelle Memorial Institute for the United States Department of Energy under DOE Contract No. DE-AC05-76RL1830.

Segatta, Francesco↗

Deciphering enhancer sequence using thermodynamics-based models and convolutional neural networks

Abstract Deciphering the sequence-function relationship encoded in enhancers holds the key to interpreting non-coding variants and understanding mechanisms of transcriptomic variation. Several quantitative models exist for predicting enhancer function and underlying mechanisms; however, there has been no systematic comparison of these models characterizing their relative strengths and shortcomings. Here, we interrogated a rich data set of neuroectodermal enhancers in Drosophila, representing cis- and trans- sources of expression variation, with a suite of biophysical and machine learning models. We performed rigorous comparisons of thermodynamics-based models implementing different mechanisms of activation, repression and cooperativity. Moreover, we developed a convolutional neural network (CNN) model, called CoNSEPT, that learns enhancer ‘grammar’ in an unbiased manner. CoNSEPT is the first general-purpose CNN tool for predicting enhancer function in varying conditions, such as different cell types and experimental conditions, and we show that such complex models can suggest interpretable mechanisms. We found model-based evidence for mechanisms previously established for the studied system, including cooperative activation and short-range repression. The data also favored one hypothesized activation mechanism over another and suggested an intriguing role for a direct, distance-independent repression mechanism. Our modeling shows that while fundamentally different models can yield similar fits to data, they vary in their utility for mechanistic inference. CoNSEPT is freely available at: https://github.com/PayamDiba/CoNSEPT.

59 BASIC BIOLOGICAL SCIENCES↗

PayamDiba/CoNSEPT

Deciphering the sequence-function relationship encoded in enhancers holds the key to interpreting non-coding variants and understanding mechanisms of transcriptomic variation. Several quantitative models exist for predicting enhancer function and underlying mechanisms; however, there has been no systematic comparison of these models characterizing their relative strengths and shortcomings. Here, we interrogated a rich data set of neuroectodermal enhancers in Drosophila, representing cis- and trans- sources of expression variation, with a suite of biophysical and machine learning models. We performed rigorous comparisons of thermodynamics-based models implementing different mechanisms of activation, repression and cooperativity. Moreover, we developed a convolutional neural network (CNN) model, called CoNSEPT, that learns enhancer ‘grammar’ in an unbiased manner. CoNSEPT is the first general-purpose CNN tool for predicting enhancer function in varying conditions, such as different cell types and experimental conditions, and we show that such complex models can suggest interpretable mechanisms. We found model-based evidence for mechanisms previously established for the studied system, including cooperative activation and short-range repression. The data also favored one hypothesized activation mechanism over another and suggested an intriguing role for a direct, distance-independent repression mechanism. Our modeling shows that while fundamentally different models can yield similar fits to data, they vary in their utility for mechanistic inference.

Dibaeinia, Payam↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Transcriptomics (PB-DP3)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Sample data was acquired using a Illumina HiSeq sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis. Transcriptomic differential expression analysis revealed coordinated circadian clock-driven adjustment of the cell cycle and rewiring of energy and carbon metabolism. Processed RNA-Seq datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed RNA-seq results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

Analytic error analysis of cross section interpolation methods in nodal diffusion codes - II: Numerical results

This paper is the second part of a two-part paper that documents the numerical results for the partial derivatives model presented in part I. In this paper, we derive the error bounds for the analytical point-wise error expression and verify our bounds with numerical experiments. The point-wise error expressions make available, and bound, the sources that contribute to the total error of the interpolated cross section in terms of the Lagrange interpolation errors and the model form error. MPACT is used to generate two-group homogenized cross sections for Westinghouse's AP1000 Region 4 lattice to evaluate the accuracy of the bounds. Error bounds calculated over a grid are compared to numerical data for uni-variate and multi-variate interpolation. The point-wise error bounds of a typical case matrix - two branches in each state variable - are displayed for bi-variate interpolation in the state variables: moderator density, fuel temperature, and boron concentration. The error bounds are shown to be highly accurate compared to numerical results, and in accordance with the underlying physics. We then discuss and show how the sources of error contribute to the total error, and consider the improvement of each error source. Finally, we mention future work such as propagating our cross section error bounds through a reactivity calculation. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Realistic Cost to Execute Practical Quantum Circuits using Direct Clifford+T Lattice Surgery Compilation

We report a resource estimation pipeline that explicitly compiles quantum circuits expressed using the Clifford+T gate set into a surface code lattice surgery instruction set. The cadence of magic state requests from the compiled circuit enables the optimization of magic state distillation and storage requirements in a post-hoc analysis. To compile logical circuits into lattice surgery operations, we build upon the open-source Lattice Surgery Compiler. The revised compiler operates in two stages: the first translates logical gates into an abstract, layout-independent instruction set; the second compiles these into local lattice surgery instructions that are allocated to hardware tiles according to a specified resource layout. The second stage retains logical parallelism while avoiding resource contention in the fault-tolerant layer, aiding realism. Additionally, users can specify dedicated tiles at which magic states are replenished, enabling resource costs from the logical computation to be considered independently from magic state distillation and storage. We demonstrate the applicability of our pipeline to large practical quantum circuits by providing resource estimates for the ground state estimation of molecules. Finally, we find that variable magic state consumption rates in real circuits can cause the resource costs of magic state storage to dominate unless production is varied to suit.

97 MATHEMATICS AND COMPUTING↗

Verification of Bison fission product species conservation under TRISO reactor conditions

When assessing the reliability and predictive capabilities of a simulation tool, code verification is used to ensure that the implemented numerical algorithm is a faithful representation of its underlying mathematical model, including partial differential or integral equations, initial and boundary conditions, and auxiliary relationships. During this process, numerical results in a discrete solution are compared to the analytical solution of the mathematical model. Here, in this paper, the code verification process is applied to one-dimensional spatiotemporal problems that exercise partial differential equation governing the conservation of fission product species (or mass diffusion). Numerical experiments were performed in the Bison fuel performance code to evaluate its predictive capability under various TRISO reactor conditions such as base irradiation and safety heating test conditions for either short- or long-lived fission product species, as well as a case concerning evaporation from the outer surface of a particle. The code predictions were compared with the expected exact results obtained from the analytical expressions, and the fact that they demonstrate the correct analytical behavior provides strong evidence of proper numerical algorithm implementation.

07 ISOTOPE AND RADIATION SOURCES↗