Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Computer graphics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

ERF: Energy Research and Forecasting Model

High performance computing (HPC) architectures have undergone rapid development in recent years. As a result, established software suites face an ever increasing challenge to remain performant on and portable across modern systems. Many of the widely adopted atmospheric modeling codes cannot fully (or in some cases, at all) leverage the acceleration provided by General-Purpose Graphics Processing Units, leaving users of those codes constrained to increasingly limited HPC resources. Energy Research and Forecasting (ERF) is a regional atmospheric modeling code that leverages the latest HPC architectures, whether composed of only Central Processing Units (CPUs) or incorporating GPUs. ERF contains many of the standard discretizations and basic features needed to model general atmospheric dynamics. The modular design of ERF provides a flexible platform for exploring different physics parameterizations and numerical strategies. ERF is built on a state-of-the-art, well-supported, software framework (AMReX) that provides a performance portable interface and ensures ERF's long-term sustainability on next generation computing systems. This paper details the numerical methodology of ERF, presents results for a series of verification/validation cases, and documents ERF's performance on current HPC systems. The roughly 5× speed up of ERF (using GPUs) over Weather Research and Forecasting (CPUs only) for a 3D squall line test case highlights the significance of leveraging GPU acceleration.

17 WIND ENERGY

The QUIC Start Guide (V.6.4.9)

QUIC stands for the Quick Urban & Industrial Complex (QUIC) dispersion modeling system. QUIC is a fast response urban dispersion model that runs on a laptop. QUIC is comprised of a 3D wind field model called QUIC-URB, a transport and dispersion model called QUIC-PLUME, and graphical user interface called QUIC-GUI. QUIC also includes QUIC-PRESSURE to solve for pressure fields in and around buildings, a population exposure assessment tool called QUIC-POP, and an indoor infiltration calculator for computing indoor concentrations. Transport and dispersion for different types of airborne contaminants can be computed on building to neighborhood scales in tens of seconds to tens of minutes. QUIC will never give perfect answers, but it will account for the effects of buildings in an approximate way and provide more realism than non-building aware dispersion models.

97 MATHEMATICS AND COMPUTING

Exploring 2D X-ray diffraction phase fraction analysis with convolutional neural networks: Insights from kinematic-diffraction simulations

Abstract Deep-learning models are effective for analyzing the complex information in 2D X-ray diffraction (XRD) patterns. Accurately collecting parameters of the material sample is crucial during model training, significantly impacting model performance. In this study, we employ a kinematic-diffraction simulator to generate simulated 2D XRD patterns for Ti–6Al–4V alloy, allowing precise control of sample parameters. These simulated patterns are used to train convolutional neural networks, predicting $$\upbeta$$ β -phase volume fractions. The training data set consists exclusively of 2D XRD patterns with pure $$\upalpha$$ α - or pure $$\upbeta$$ β -phase, while the testing set incorporates patterns with intermediate phase volume fraction. In particular, we investigate how the architectures of the model influence prediction reliability and computational performance. Experimental results reveal that, with appropriate training, the convolutional neural network accurately detects intermediate phase volume fractions even trained with only pure-phase patterns, achieving a mean square error accuracy of $$9.4 \times 10^{-4}$$ 9.4 × 10 - 4 . Graphical abstract

Yue, Weiqi

GPU acceleration of hybrid functional calculations in the SPARC electronic structure code

We present a Graphics Processing Unit (GPU)-accelerated version of the real-space SPARC electronic structure code for performing hybrid functional calculations in generalized Kohn–Sham density functional theory. In particular, we develop a batch variant of the recently formulated Kronecker product-based linear solver for the simultaneous solution of multiple linear systems. We then develop a modular, math kernel based implementation for hybrid functionals on NVIDIA architectures, where computationally intensive operations are offloaded to the GPUs, while the remaining workload is handled by the central processing units (CPUs). Considering bulk and slab examples, we demonstrate that GPUs enable up to 8× speedup in node-hours and 80× in core-hours compared to CPU-only execution, reducing the time to solution on V100 GPUs to around 300 s for a metallic system with over 6000 electrons, and significantly reducing the computational resources required for a given wall time.

Kohn-Sham density functional theory

gRASPA

GPU Monte Carlo Simulation Code with a taste of RASPA We present enhancements in Monte Carlo simulation speed and functionality within an open-source code, gRASPA, which uses graphical processing units (GPUs) to achieve significant performance improvements compared to serial, CPU implementations of Monte Carlo. The code supports a wide range of Monte Carlo simulations, including canonical ensemble (NVT), grand canonical, NVT Gibbs, Widom test particle insertions, and continuous-fractional component Monte Carlo. Implementation of grand canonical transition matrix Monte Carlo (GC-TMMC) and a novel feature to allow different moves for the different components of metal-organic framework (MOF) structures exemplify the capabilities of gRASPA for precise free energy calculations and enhanced adsorption studies, respectively. The introduction of a High-Throughput Computing (HTC) mode permits many Monte Carlo simulations on a single GPU device for accelerated materials discovery. The code can incorporate machine learning (ML) potentials. The open-source nature of gRASPA promotes reproducibility and openness in science, and users may add features to the code and optimize it for their own purposes. The code is written in CUDA/C++ and SYCL/C++ to support different GPU vendors. The gRASPA code is publicly available at https://github.com/snurr-group/gRASPA.

Li, Zhao [Purdue/Northwestern/Notre Dame Universit

High Performance Computing Peak Shaving for Microreactor Operation

There are multiple nuclear microreactors currently under development that are designed to provide autonomous power for as many as ten or more years without refueling and are designed to power high performance computing (HPC) datacenters. But the load-follow speeds for a nuclear microreactor will be much slower than grid power and slower than the power variance typical of a HPC system. HPC datacenters experience peak power load variance driven by several factors ranging from the operation of cooling systems to remove heat from the servers to supporting a wide range of user application workflows and architectures each with different power signatures. One mechanism to support the limited load-follow of a microreactor is peak shaving where an energy storage mechanism is used to shed peak load and reduce significant power variance. This work explores peak electrical load shaving using uninterruptible power supply (UPS) systems designed for HPC support in the context of peak shaving when operating using a nuclear microreactor with a load-follow limited to 10% of load per minute. Using a self contained HPC datacenter complete with stand-alone cooling system and provisioned with an x86 cluster, an ARM cluster, and a graphics processing unit (GPU) cluster, peak shaving for microreactor operation using the UPS battery backup is explored while running two classes of typical HPC user applications. HPC architecture suitability for microreactor operation under this type of peak shaving is examined.

97 MATHEMATICS AND COMPUTING

Cross-correlation image analysis for real-time single particle tracking

Accurately measuring the translations of objects between images is essential in many fields, including biology, medicine, chemistry, and physics. One important application is tracking one or more particles by measuring their apparent displacements in a series of images. Popular methods, such as the center of mass, often require idealized scenarios to reach the shot noise limit of particle tracking and, therefore, are not generally applicable to multiple image types. More general methods, such as maximum likelihood estimation, reliably approach the shot noise limit, but are too computationally intense for use in real-time applications. These limitations are significant, as real-time, shot-noise-limited particle tracking is of paramount importance for feedback control systems. To fill this gap, we introduce a new cross-correlation-based algorithm that approaches shot-noise-limited displacement detection and a graphics processing unit-based implementation for real-time image analysis of a single particle.

Instruments & Instrumentation

Accelerating Neutrino Event Generation in MARLEY Using CUDA-Based RNG and GPU Parallelization

MARLEY is a simulation tool that helps scientists study how low-energy neutrinos interact with matter. To work properly, MARLEY uses random numbers thousands of times in each simulation. These random numbers are important for modeling things like how neutrinos collide with atoms and what particles they produce. Right now, MARLEY runs on a regular computer processor (CPU) and uses a built-in random number generator called the Mersenne Twister. This setup works, but it can be slow, especially when trying to simulate many events. This research focuses on making MARLEY run faster by moving the random number generation and some of the repetitive calculations from the CPU to a graphics processing unit (GPU), which can handle many tasks at the same time. We use CUDA (a tool for programming NVIDIA GPUs) and cuRAND (a GPU-based random number library) to test faster alternatives to the current random number system. We compare different GPU-based generators, like curand_mtgp32, xorwow, and philox, to see which ones are the quickest and still give reliable results. Early tests show that using the GPU can make MARLEY simulations much faster. This project not only helps improve current simulation performance but also moves closer to a full simulation chain where all stages can run on modern GPU hardware.

Dunkley, Kimieka [Florida A-M]

Differential equations for cosmological correlators

Cosmological fluctuations retain a memory of the physics that generated them in their spatial correlations. The strength of correlations varies smoothly as a function of external kinematics, which is encoded in differential equations satisfied by cosmological correlation functions. In this work, we provide a broader perspective on the origin and structure of these differential equations. As a concrete example, we study conformally coupled scalar fields in a power-law cosmology. The wavefunction coefficients in this model have integral representations, with the integrands being the product of the corresponding flat-space results and “twist factors” that depend on the cosmological evolution. Similar twisted integrals arise for loop amplitudes in dimensional regularization, and their recent study has led to the discovery of rich mathematical structures and powerful new tools for computing multi-loop Feynman integrals in quantum field theory. The integrals of interest in cosmology are also part of a finite-dimensional basis of master integrals, which satisfy a system of first-order differential equations. We develop a formalism to derive these differential equations for arbitrary tree graphs. The results can be represented in graphical form by associating the singularities of the differential equations with a set of graph tubings. Upon differentiation, these tubings grow in a local and predictive fashion. In fact, a few remarkably simple rules allow us to predict — by hand — the equations for all tree graphs. While the rules of this “kinematic flow” are defined purely in terms of data on the boundary of the spacetime, they reflect the physics of bulk time evolution. We also study the analogous structures in tr ϕ 3 theory, and see some glimpses of hidden structure in the sum over planar graphs. This suggests that there is an autonomous combinatorial or geometric construction from which cosmological correlations, and the associated spacetime, emerge.

Cosmological models

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES

Kinematic flow from the flow of cuts

The wavefunction coefficients of conformally coupled scalars in power-law FRW cosmologies satisfy differential equations governed by a set of simple combinatorial rules known as the kinematic flow. In this paper we derive the kinematic flow, expressed using a set of differential forms referred to as the cut basis, from a geometric perspective, relying solely on the cosmological hyperplane arrangement and without invoking bulk physics. Each element of the cut basis corresponds to the positive geometry associated to an independent cut of the physical FRW-form and can be labeled by decorating (minors of) the truncated Feynman graph with an acyclic orientation. We provide a straightforward prescription to associate a logarithmic differential form to each element of the cut basis by considering its corresponding decorated graph. Moreover, we show that the residues of the physical FRW-form are canonical forms of certain graphical zonotopes labeled by the same set of decorated graphs. These zonotopes control the cut combinatorics -- flow of cuts -- of the physical FRW-form and the cut basis (by construction). Using the theory of relative twisted cohomology and intersection theory, we derive a closed form formula for the differential equations of the cut basis. We also introduce combinatorial rules that compute the kinematic differential of any basis element without explicit calculation. The combinatorics of our differential equations is a natural consequence of the flow of cuts and is equivalent (up to rescaling) to the kinematic flow for the recently studied time integral basis. In particular, our differential equations decouple into exponentially many sectors, one for each way of cutting a subset of edges of the graph.

General Relativity and Quantum Cosmology

CRiSPPy: An advanced hydropower scheduling tool for the Colorado River Storage Project

The Western Area Power Administration (WAPA) plays a vital role in delivering reliable and cost-effective hydroelectric power to millions of customers across the western United States. The Colorado River Storage Project (CRSP) carries out WAPA’s mission in Arizona, Utah, Colorado, New Mexico, Nevada, Wyoming and Texas. Achieving this mission requires effective management of the Colorado River system, and depends on the use of advanced analytical tools and modeling methodologies. For many years, CRSP has relied on the Generation and Transmission Maximization Superlite (GTMax SL) model for its mid-term and long-term hydroscheduling needs. However, the evolving energy market, power system operations, environmental rules, and hydrology conditions, coupled with advancements in computational capabilities, have necessitated the development of a more modern and robust solution. This report introduces the Colorado River Storage Project Python-based (CRiSPPy) model, a new, advanced hydropower scheduling tool developed to address CRSP ever-evolving challenges. CRiSPPy represents a significant leap forward in our ability to model and optimize the operation of the Colorado River system. It incorporates state-of-the-art optimization algorithms, enhanced data management capabilities, and an advanced graphical user interface, providing WAPA CRSP personnel with unprecedented insights and decision-making support. This document details the development, capabilities, and implementation of CRiSPPy. It is intended to serve as a comprehensive resource for WAPA staff, stakeholders, and anyone interested in the future of hydropower scheduling in the Colorado River Basin. We are confident that CRiSPPy will enhance WAPA's mission while adapting to the challenges of a dynamic and increasingly complex environment. The version of CRiSPPy described in this report is the version 2.3. New versions of CRiSPPy will be developed as the tool keeps evolving to address CRSP challenges.

13 HYDRO ENERGY

An integrated modeling framework with open architecture for phase field simulation of multi-component alloys

An integrated modeling framework (PanPhaseField) has been developed, which enables a direct and fast coupling between CALPHAD calculations and large-scale phase field simulations for multi-component alloys. Further, it adopts an open architecture allowing for integration of user-defined phase field models in a plug-and-play manner by taking full advantage of the user-friendly graphical interface of Pandat software. The developed modeling platform becomes an enabling tool that can be used to simulate the evolution of spatially varying microstructures of industrial complex alloys for various engineering applications.

36 MATERIALS SCIENCE

Hydropower Capacity Factor Trends & Analytics for the United States

This data repository contains all code, input data, and data generated for Turner et al. (2024)—“Hydropower capacity factors trending down in the United States”. File descriptions: – hydro-cf-trends-inputs.zip: Full set of input data used in this study, organized for direct entry into “/data” directory of hydro-cf-trends data processing pipeline. – hydro-cf-trends.zip: Full data processing pipeline, coded using the R {targets} framework. This is a snapshot release (v1.0) of the code repository stored at https://code.ornl.gov/turnersw/hydro-cf-trends/. – hydro-cf-trends-results.zip: Provides all dam level results required to reproduce results and graphics in Turner et al. (2024). Dams are identified by the “complxID” (root of the hydropower plant ID in the Existing Hydropower Assets Database, inherited from HILARRI). Results include: • dam_CF_trends.csv: Table of long-term trends in annualized capacity factors for 610 dams and modeled annualized capacity factors for 362 modeled dams (naturalized and assimilated flows). • dam_annualized_CF_gen.csv: Annualized time series of the following variables for each of 610 hydropower dams with nameplate > 5MW – Reported nameplate capacity (MW) – Implied maximum annual generation (MWh) – Reported net generation (MWh) – Computed annual capacity factor – Modeled annual capacity factor (362 modeled plants only)

13 HYDRO ENERGY

DEM Modeling and Validation of Pebble Bed Packing Using Chrono::GPU

Accurate prediction of pebble packing structure is important for pebble bed reactors because the spatial distribution of void fraction directly affects coolant flow, pressure drop, heat transfer, and neutronic behavior. However, experimentally validated DEM studies that directly evaluate local void-fraction structure in reactor-relevant pebble beds remain limited. In this work, the pebble bed experiment conducted at Missouri University of Science and Technology is simulated using the graphics processing unit (GPU)-based discrete element method (DEM) code Chrono::GPU. The study focuses on evaluating the ability of Chrono::GPU to reproduce the packing arrangement and void-fraction distribution of a randomly packed spherical pebble bed. The DEM results are first verified against established radial void-fraction correlations, including the Mueller and Vortmeyer-Schuster models, to assess the predicted bulk porosity, near-wall behavior, and oscillatory packing structure. The simulation is then verified against reference DEM data and validated against gamma-ray computed tomography (CT) experimental data at three axial locations. The Chrono::GPU results reproduce the main features of the experimental packing, including the high void fraction near the wall, the first near-wall trough, and the damped oscillatory radial profile caused by wall-induced ordering. Quantitative comparison with DEM data and the CT-based radial profiles shows good agreement, with mean absolute errors on the order of 0.07 and root-mean-square errors below 0.09 for the averaged profiles. These results demonstrate that Chrono::GPU can accurately capture the void-fraction structure of spherical pebble beds and provides a reliable DEM framework for future pebble bed reactor packing, recycling, and thermal-hydraulic studies.

97 - MATHEMATICS AND COMPUTING

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Energetics of the nucleation and glide of disconnection modes in symmetric tilt grain boundaries

Grain boundaries (GBs) evolve by the nucleation and glide of disconnections, which are dislocations with a step character. In this work, motivated by recent success in predicting GB properties such as the shear coupling factor and mobility from the intrinsic properties of disconnections, we develop a systematic method to calculate the energy barriers for the nucleation and glide of individual disconnection modes under arbitrary driving forces and a quasi-2D setting. This method combines tools from bicrystallography to enumerate disconnection modes and the Nudged elastic band (NEB) method to calculate their energetics, yielding minimum energy paths and atomistic mechanisms for the nucleation and glide of each disconnection mode. We apply the method to accurately predict shear coupling factors of $[001]$ symmetric tilt grain boundaries in Cu. Particular attention is paid to the boundaries where the dislocation-based disconnection nucleation model produces incorrect nucleation barriers. We demonstrate that the method can accurately compute energy barriers and predict shear-coupling factors in the low-temperature regime. For certain disconnection modes in which the assumptions underlying our method do not hold, we report upper bounds on the energy barriers for disconnection nucleation and glide. In addition, the NEB trajectories reveal interesting phenomena such as the dissociation of a higher energy mode into lower energy modes, and in some cases, shear coupling being mediated by partial disconnections, wherein the GB structure temporarily changes to a metastable state before reverting back to its original structure. Graphical abstract

36 MATERIALS SCIENCE

Monte Carlo Simulation with CAD Interface for Calculation of 3D Maps of Residual Dose (CRADA)

Objective: To develop an easy-to-use software application to predict and mitigate radiation effects in research environment, space instruments, nuclear plants and medical facilities and help nonproliferation and national security efforts. Tech-X will develop standalone software libraries and command-line tools for ( 1) translating CAD into tessellated surfaces and tetrahedral meshes in GDML (for Geant4 and MARS 15), ROOT (for MARS 15) and HDF5 (for compact representation and for the visualization) formats, (2) healing CAD geometries to make them suitable for Monte Carlo simulations; (3) creating uniform and variable Cartesian and cylindrical meshes for detailed scoring; and ( 4) efficient Monte Carlo navigation in CAD geometries. JLAB will finish automation of simulations of residual dose in CAD geometries and integrate Tech-X software into Geant4 and MARS15. Finally, Tech-X will develop a Graphical User Interface to set up and heal CAD geometries, create input files, run and visualize simulations for residual dose. This application will run on local desktops, local and remote clusters and supercomputers and will be made available through public clouds, such as Amazon Web Services.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND