Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer graphics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

The novel Mechanical Ventilator Milano for the COVID-19 pandemic

This paper presents the Mechanical Ventilator Milano (MVM), a novel intensive therapy mechanical ventilator designed for rapid, large-scale, low-cost production for the COVID-19 pandemic. Free of moving mechanical parts and requiring only a source of compressed oxygen and medical air to operate, the MVM is designed to support the long-term invasive ventilation often required for COVID-19 patients and operates in pressure-regulated ventilation modes, which minimize the risk of furthering lung trauma. The MVM was extensively tested against ISO standards in the laboratory using a breathing simulator, with good agreement between input and measured breathing parameters and performing correctly in response to fault conditions and stability tests. The MVM has obtained Emergency Use Authorization by U.S. Food and Drug Administration (FDA) for use in healthcare settings during the COVID-19 pandemic and Health Canada Medical Device Authorization for Importation or Sale, under Interim Order for Use in Relation to COVID-19. Furthermore, following these certifications, mass production is ongoing and distribution is under way in several countries. The MVM was designed, tested, prepared for certification, and mass produced in the space of a few months by a unique collaboration of respiratory healthcare professionals and experimental physicists, working with industrial partners, and is an excellent ventilator candidate for this pandemic anywhere in the world

60 APPLIED LIFE SCIENCES↗

ROSE LCOM Tools

This project provides a tool to compute LCOM metrics for important Ada abstraction mechanisms (e.g., packages, protected objects). It analyzes data accesses and routine calls to compute an abstract representation of an Ada package (or some other abstraction). Different LCOM metrics (1-5) are computed from the abstract representation. The project also provides support for representing the abstraction graphically (via graphviz). The LCOM tool requires an existing ROSE installation configured with the GNAT 2019 compiler.

Lamar, KennethM↗

PUFFIn

The PUFFIn software package is a graphical user interface that allows for photon and electron beam calculations using the PENELOPE computer code. It prepares the required input files and displays the resulting output data. It can be used by anyone interested in the interaction of photon or electron beams with matter

Schwarz, Randy↗

Tractable minor-free generalization of planar zero-field Ising models

In this work, we present a new family of zero-field Ising models over N binary variables/spins obtained by consecutive 'gluing' of planar and O(1)-sized components and subsets of at most three vertices into a tree. The polynomial time algorithm of the dynamic programming type for solving exact inference (computing partition function) and exact sampling (generating i.i.d. samples) consists of sequential application of an efficient (for planar) or brute-force (for O(1)-sized) inference and sampling to the components as a black box. To illustrate the utility of the new family of tractable graphical models, we first build a polynomial algorithm for inference and sampling of zero-field Ising models over K 33 -minor-free topologies and over K 5 -minor-free topologies—both of which are extensions of the planar zero-field Ising models—which are neither genus- nor treewidth-bounded. Second, we empirically demonstrate an improvement in the approximation quality of the NP-hard problem of inference over the square-grid Ising model in a node-dependent nonzero 'magnetic' field.

97 MATHEMATICS AND COMPUTING↗

A survey of software implementations used by application codes in the Exascale Computing Project

The US Department of Energy Office of Science and the National Nuclear Security Administration initiated the Exascale Computing Project (ECP) in 2016 to prepare mission-relevant applications and scientific software for the delivery of the exascale computers starting in 2023. The ECP currently supports 24 efforts directed at specific applications and six supporting co-design projects. These 24 application projects contain 62 application codes that are implemented in three high-level languages—C, C++, and Fortran—and use 22 combinations of graphical processing unit programming models. The most common implementation language is C++, which is used in 53 different application codes. The most common programming models across ECP applications are CUDA and Kokkos, which are employed in 15 and 14 applications, respectively. This article provides a survey of the programming languages and models used in the ECP applications codebase that will be used to achieve performance on the future exascale hardware platforms.

97 MATHEMATICS AND COMPUTING↗

Fostering Remote Visualization: Experiences in Two Different HPC Sites

Visualization of scientific data is crucial for scientific discovery to gain insight into the results of simulations and experiments. Remote visualization is of crucial importance to access infrastructure, data and computational resources and, to avoid data movement from where data is produced and to where data will be analyzed. Remote visualization enables geographically diverse collaboration and enhances user experience through graphical user interfaces. This paper presents two approaches deployed by two different HPC centers: The SC3 - Supercomputación y Cálculo Científico Center in Colombia and the Oak Ridge Leadership Computing Facility in USA. We overview our remote visualization experiences, adopted technologies, use cases, and challenges encountered. Our contribution is to signal the commonality between approaches in terms of the end goal, showing their fitness for their contexts, while not focusing only on attempting to provide a general picture of remote visualization, given the differences between centers in terms of purposes, needs, resources, and national impact.

Gelvez Cortes, Sergio Augusto↗

Rank-reduced coupled-cluster. III. Tensor hypercontraction of the doubles amplitudes

In this work, we develop a quartic-scaling implementation of coupled-cluster singles and doubles (CCSD) based on low-rank tensor hypercontraction (THC) factorizations of both the electron repulsion integrals (ERIs) and the doubles amplitudes. This extends our rank-reduced (RR) coupled-cluster method to incorporate higher-order tensor factorizations. The THC factorization of the doubles amplitudes accounts for most of the gain in computational efficiency as it is sufficient, in conjunction with a Cholesky decomposition of the ERIs, to reduce the computational complexity of most contributions to the CCSD amplitude equations. Further THC factorization of the ERIs reduces the complexity of certain terms arising from nested commutators between the doubles excitation operator and the two-electron operator. We implement this new algorithm using graphical processing units and demonstrate that it enables CCSD calculations for molecules with 250 atoms and 2500 basis functions using a single computer node. Furthermore, we show that the new method computes correlation energies with comparable accuracy to the underlying RR-CCSD method.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BioRT‐HBV 1.0: A Biogeochemical Reactive Transport Model at the Watershed Scale

Abstract Reactive Transport Models (RTMs) are essential tools for understanding and predicting intertwined ecohydrological and biogeochemical processes on land and in rivers. While traditional RTMs have focused primarily on subsurface processes, recent watershed‐scale RTMs have integrated ecohydrological and biogeochemical interactions between surface and subsurface. These emergent, watershed‐scale RTMs are often spatially explicit and require extensive data, computational power, and computational expertise. There is however a pressing need to create parsimonious models that require minimal data and are accessible to scientists with limited computational background. To that end, we have developed BioRT‐HBV 1.0, a watershed‐scale, hydro‐biogeochemical RTM that builds upon the widely used, bucket‐type HBV model known for its simplicity and minimal data requirements. BioRT‐HBV uses the conceptual structure and hydrology output of HBV to simulate processes including advective solute transport and biogeochemical reactions that depend on reaction thermodynamics and kinetics. These reactions include, for example, chemical weathering, soil respiration, and nutrient transformation. The model uses time series of weather (air temperature, precipitation, and potential evapotranspiration) and initial biogeochemical conditions of subsurface water, soils, and rocks as input, and output times series of reaction rates and solute concentrations in subsurface waters and rivers. This paper presents the model structure and governing equations and demonstrates its utility with examples simulating carbon and nitrogen processes in a headwater catchment. As shown in the examples, BioRT‐HBV can be used to illuminate the dynamics of biogeochemical reactions in the invisible, arduous‐to‐measure subsurface, and their influence on the observed stream or river chemistry and solute export. With its parsimonious structure and easy‐to‐use graphical user interface, BioRT‐HBV can be a useful research tool for users without in‐depth computational training. It can additionally serve as an educational tool that promotes pollination of ideas across disciplines and foster a diverse, equal, and inclusive user community.

Sadayappan, Kayalvizhi↗

A Hybrid Climate Modeling System Using AI-assisted Process Emulators

This white paper addresses Focus Area II. We advocate developing a hybrid modeling system to improve the understanding of decadal- and longer-scale predictability of high impact water cycle components. This hybrid model combines a partial differential equation (PDE)-based dynamic core with AI/ML based emulators to represent many of the computationally expensive processes in Earth’s climate models. The hybrid modeling system has the potential to exploit emerging graphics processing unit (GPU)-accelerated architectures and allows for the generation of large ensemble (~1000’s) simulations to better characterize the model uncertainty and understand predictability.

58 GEOSCIENCES↗

Interactive Exploration of High-Dimensional Phase Diagrams

High-dimensional thermodynamic phase stability databases are becoming increasingly common due to the convergence of three recent trends: (i) the widespread interest in so-called “high-entropy” alloys, (ii) the availability of high-throughput computational assessments of phase stability in broad composition spaces and (iii) the ongoing development of ever-increasingly broad, multicomponent, multiphase CALPHAD databases. Although automated computational tools can readily process such high-dimensional data, scientists are often unable to visualize the relevant phase relations, an ability that is crucial to gaining an intuitive understanding of the stability constraints governing materials design. The present work addresses this need by providing algorithms that enable the interactive exploration of phase equilibria in high-dimensional spaces. These algorithms concentrate the complex nonlinear nonsmooth optimization needed into a preprocessing step that generates a large number of high-dimensional yet elementary graphical primitives. Furthermore, these primitives can then be cross-sectioned to yield 3-dimensional views in a computationally efficient manner that enables an interactive exploration of high-dimensional spaces. All of these operations are highly parallelizable, thus facilitating scaling of this method to large data sets.

36 MATERIALS SCIENCE↗

Collaborative Research: Improved Efficiency and Coupling of the Radiation Code in the ACME Earth System Model. Final Report

This final report details all work performed on the project by both project partners. This project provided support to properly couple RTE+RRTMGP, a high-performance broadband radiation code, within DOE’s Energy Exascale Earth System Model (E3SM). RTE+RRTMGP is a successor to the RRTMG radiation code, which has been widely accepted for its speed and accuracy by the global modeling community, and has been in use in the NCAR CESM for many years and was implemented in the initial version of E3SM. However, the computational cost of RRTMG remains high relative to other components in part due to its complexity and to its inefficient use of modern optimization strategies, issues that were rectified by the development of RTE+RRTMGP. Many of the accomplishment in this project necessitated significant collaboration with the E3SM development team. One focus of the project was to enhance the code’s optimization on the limited number of emerging computing systems on which the model is expected be used, including Many Integrated Core (MIC) architectures and Graphics Processing Unit (GPU) hardware. We also developed additional capabilities for RTE+RRTMGP that E3SM scientists identified as important for the planned applications of the model. The result of our project was optimization of a key physical component (radiative transfer calculations) of E3SM, directly supporting E3SM’s overarching global modeling objectives. More broadly, this project provided overall advancements in the use of radiative transfer calculations in atmospheric modeling and simulation, particularly for climate.

54 ENVIRONMENTAL SCIENCES↗

An efficient numerical model for predicting residual stress and strain in parts manufactured by laser powder bed fusion

Abstract Computational modeling of additively manufactured structures plays an increasingly important role in product design and optimization. For laser powder bed fusion processes, the accurate modeling of stress and distortion requires large amount of computational cost due to very localized heat input and evolving complex geometries. The current study takes advantage of a graphics processing unit accelerated explicit finite element analysis code and approximated heat conduction analysis to predict the macroscopic thermo-mechanical behavior in laser selective melting. Adjacent layers and tracks were lumped to reduce the number of time steps and elements in the finite element model. The effects of track and layer grouping on prediction accuracy and solution efficiency are investigated to provide a guidance for a cost-effective simulation. Thin-wall builds from Inconel alloy 625 (IN625) powders were simulated by applying the developed modeling approach to get the detailed residual stress and distortion at a computational speed 50 times higher than conventional approach. Under repeated heating and cooling cycles, a high tensile stress was produced near surfaces of a build due to a larger shrinkage on surface than that in central area. It is also shown that horizontal stresses concentrate near the root and top layers of the IN625 build. The predicted residual elastic strain distribution was validated by the experimental measurement using x-ray synchrotron diffraction.

36 MATERIALS SCIENCE↗

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C↗

MARBLE: A Multi-GPU Aware Job Scheduler for Deep Learning on HPC Systems

Deep learning (DL) has become a key tool for solving complex scientific problems. However, managing the multi-dimensional large-scale data associated with DL, especially atop extant multiple graphics processing units (GPUs) in modern supercomputers poses significant challenges. Moreover, the latest high-performance computing (HPC) architectures bring different performance trends in training throughput compared to the existing studies. Existing DL optimizations such as larger batch size and GPU locality-aware scheduling have little effect on improving DL training throughput performance due to fast CPU-to-GPU connections. Additionally, DL training on multiple GPUs scales sublinearly. Thus, simply adding more GPUs to a system is ineffective. To this end, we design MARBLE, a first-of-its-kind job scheduler, which considers the non-linear scalability of GPUs at the intra-node level to schedule an appropriate number of GPUs per node for a job. By sharing the GPU resources on a node with multiple DL jobs, MARBLE avoids low GPU utilization in current multi-GPU DL training on HPC systems. Our comprehensive evaluation in the Summit supercomputer shows that MARBLE is able to improve DL training performance by up to 48.3% compared to the popular Platform Load Sharing Facility (LSF) scheduler. Compared to the state-of-the-art of DL scheduler, Optimus, MARBLE reduces the job completion time by up to 47%.

Han, Jingoo↗

Kronecker-structured covariance models for multiway data

Many applications produce multiway data of exceedingly high dimension. Modeling such multi-way data is important in multichannel signal and video processing where sensors produce multi-indexed data, e.g. over spatial, frequency, and temporal dimensions. We will address the challenges of covariance representation of multiway data and review some of the progress in statistical modeling of multiway covariance over the past two decades, focusing on tensor-valued covariance models and their inference. We will illustrate through a space weather application: predicting the evolution of solar active regions over time.

97 MATHEMATICS AND COMPUTING↗

Matrix Product (GEMM) Performance Data from GPUs

Timing data for mixed precision GEMM matrix product operations on several GPU models, including NVIDIA V100 and A100, AMD MI100 and Intel P580. Also data from machine learning model training on this data using Scikit-learn.

97 MATHEMATICS AND COMPUTING↗

HL-LHC Analysis With ROOT: ROOT Project Input to the HL-LHC Computing Review (Stage 2)

ROOT is high energy physics' software for storing and mining data in a statistically sound way, to publish results with scientific graphics. It is evolving since 25 years, now providing the storage format for more than one exabyte of data; virtually all high energy physics experiments use ROOT. With another significant increase in the amount of data to be handled scheduled to arrive in 2027, ROOT is preparing for a massive upgrade of its core ingredients. As part of a review of crucial software for high energy physics, the ROOT team has documented its R&D plans for the coming years.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Breaking the mold: Overcoming the time constraints of molecular dynamics on general-purpose hardware

The evolution of molecular dynamics (MD) simulations has been intimately linked to that of computing hardware. For decades following the creation of MD, simulations have improved with computing power along the three principal dimensions of accuracy, atom count (spatial scale), and duration (temporal scale). Since the mid-2000s, computer platforms have, however, failed to provide strong scaling for MD, as scale-out central processing unit (CPU) and graphics processing unit (GPU) platforms that provide substantial increases to spatial scale do not lead to proportional increases in temporal scale. Important scientific problems therefore remained inaccessible to direct simulation, prompting the development of increasingly sophisticated algorithms that present significant complexity, accuracy, and efficiency challenges. While bespoke MD-only hardware solutions have provided a path to longer timescales for specific physical systems, their impact on the broader community has been mitigated by their limited adaptability to new methods and potentials. In this work, we show that a novel computing architecture, the Cerebras wafer scale engine, completely alters the scaling path by delivering unprecedentedly high simulation rates up to 1.144 M steps/s for 200 000 atoms whose interactions are described by an embedded atom method potential. This enables direct simulations of the evolution of materials using general-purpose programmable hardware over millisecond timescales, dramatically increasing the space of direct MD simulations that can be carried out. In this paper, we provide an overview of advances in MD over the last 60 years and present our recent result in the context of historical MD performance trends.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗