Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance Portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Working with Multiple MPIs: Overcoming ABI Incompatibility

Scientific applications rely on third-party libraries, which may have multiple implementations. While sharing an application programming interface (API), many of these implementations do not have a shared application binary interface (ABI) and require recompiling. Recompiling can be a long and complex process and sometimes not even an option when the application is shipped binary only. ABI incompatibility strikes at the heart of portability, productivity, and performance by (1) impeding application execution across different HPC and Cloud systems; (2) adding developer hours rebuilding an application; and (3) not taking advantage of host-optimized libraries. This tutorial teaches attendees a portable way to address ABI incompatibility in MPI using the Wi4MPI library. Wi4MPI translates the ABI dynamically from the MPI library used to build the application to a different MPI library available at run time. With Wi4MPI, HPC practitioners can break the portability barrier imposed by ABI incompatibility, potentially increase performance, and increase user productivity. The tutorial is broken down in three components: (1) Understanding ABI compatibility in MPI; (2) Translating MPI libraries dynamically; and (3) Applying dynamic translation to key use cases in HPC, including Containers. If you use more than one MPI library or supercomputer, this tutorial is for you.

97 MATHEMATICS AND COMPUTING↗

Effects of Promethazine on Performance During Simulated Shuttle Landings

Promethazine (PMZ) is the antimotion sickness drug of choice in the U.S. Space Shuttle program; however, virtually nothing is known about the bioavailability and performance effects of this drug in the microgravity environment. PMZ has detrimental side effects on human performance on Earth that could affect Shuttle operations. In a recent ground-based study we examined: 1) the effects of promethazine (PMZ) on Shuttle landing performance using the portable inflight landing operations trainer (PILOT), and 2) saliva and urine samples to determine the pharmacokinetics of PMZ. The PILOT performance data is presented here.

Harm, D. L.↗

Improved thermal storage material for portable life support systems

The availability of thermal storage materials that have heat absorption capabilities substantially greater than water-ice in the same temperature range would permit significant improvements in performance of projected portable thermal storage cooling systems. A method for providing increased heat absorption by the combined use of the heat of solution of certain salts and the heat of fusion of water-ice was investigated. This work has indicated that a 30 percent solution of potassium bifluoride (KHF2) in water can absorb approximately 52 percent more heat than an equal weight of water-ice, and approximately 79 percent more heat than an equal volume of water-ice. The thermal storage material can be regenerated easily by freezing, however, a lower temperature must be used, 261 K as compared to 273 K for water-ice. This work was conducted by the United Aircraft Research Laboratories as part of a program at Hamilton Standard Division of United Aircraft Corporation under contract to NASA Ames Research Center.

Kellner, J. D.↗

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor↗

International time transfer and portable clock evaluation using GPS timing receivers: Preliminary results

The overall experiment was designed to test the positioning and navigation capabilities of the GPS timing receivers developed by the Naval Research Laboratory (NRL) for the NASA Goddard Laser Tracking Network (GITN). To perform this experiment, a reliable and redundant time scale was set up onboard the ship, and a back-up on shore. This situation provided the opportunity to perform simultaneously a timing experiment ideally divided into two parts, the main objectives of the experimentation being: (1) To test GPS timing receiver synchronization capabilities on a moving platform, and to perform an intercontinental synchronization via GPS between participating international timing laboratories in Europe and in the United States. (2) To evaluate the performance of cesium portable clocks in the field.

Wardrip, S. C.↗

Rasterization with Data-Parallel Primitives

Parallel rasterization can suffer from race conditions during fragment generation, which is traditionally addressed by using specialized hardware accessible via vendor graphics APIs. Unfortunately, graphics APIs are increasingly problematic on high-performance computers, either because they are not provided or because of concerns about dependencies with in situ visualization. In response, we present a hardware-agnostic rasterization algorithm that handles race conditions using only data-parallel primitives (DPPs), enabling efficient rendering on HPC systems without graphics API dependencies and aligning with recent efforts to deliver visualization software with DPPs. Our evaluation consists of three phases: (1) evaluating portability across different CPU and GPU architectures, (2) evaluating competitiveness with a community standard, and (3) evaluating performance across varying workloads and available parallelism. The supporting experiments run on both AMD and NVIDIA GPUs, considering data sets as large as 460 million triangles and 160 million pixels. While performance generally falls short of graphics API baselines, it achieves interactive frame rates on most workloads. As a result, we conclude our approach is a viable solution for rasterization on high-performance computers since our approach is portably performant across different architectures without the need for specialized vendor support.

Buckley, Makani [University of Oregon] (ORCID:0009↗

Short Exploration Extravehicular Mobility Unit Testing Setup: Evaluation Under Realistic Pressure and Thermal Conditions

The purpose of Short Exploration Extravehicular Mobility Unit (SxEMU) thermal vacuum testing was to verify the functionality of the Design Verification Testing (DVT) prototype xEMU (SxEMU for this test) at vacuum pressures and extreme space and lunar surface temperature conditions. The SxEMU Thermal Vacuum Test was the culmination of the DVT xEMU project. This paper’s focus is on the pre-Extravehicular Activity (EVA) test setup, and general performance of the SxEMU Portable Life Support Subsystem (xPLSS), with focus on the performance of the Primary Oxygen Assembly (POA) and Secondary Oxygen Assembly (SOA), including Secondary Oxygen Regulator (SOR) takeover and the POA and SOA low-setpoint change inhibit. The initial pre-EVA test preparation included recharging the batteries and replenishing consumables, including test-system water, oxygen assemblies (with gaseous nitrogen), and the integrated thermal loops, including the Feedwater Supply Assemblies. xPLSS functionality testing included carbon dioxide (CO2) removal via the Rapid Cycle Amine swingbed system, thermal loop temperature control, and monitoring of suit ventilation loop pressure, temperature, and CO2 percentages. Testing evaluated automatic takeover of suit pressure control by the SOR after the primary oxygen supply is depleted. The Primary Oxygen Regulator and SOR low-setpoint change inhibit function prevents the crewmember from inadvertently setting the primary regulator to a low pressure setpoint during an EVA.

xPLSS↗

First Annual Report on Development of Microwave Resonant Cavity Transducer for Fluid Flow Sensing: Development of Sensor Performance Model of Microwave Cavity Flow Meter for Advanced Reactor High Temperature Fluids

We are investigating a microwave cavity-based transducer for in-core high-temperature fluid flow sensing in molten salt cooled reactors (MSCR) and sodium fast reactors (SFR). This sensor is a hollow metallic cylindrical cavity, which can be fabricated from stainless steel, and as such is expected to be resilient to radiation, high temperature and corrosive environment of MSCR and SFR. The principle of sensing consists of making one wall of the cylindrical cavity flexible enough so that dynamic pressure, which is proportional to fluid velocity, will cause membrane deflection. Membrane deflection causes cavity volume change, which leads to a shift in the resonant frequency. Feasibility of the sensor was initially investigated with analytical derivations and with COMSOL RF Module computer simulations of resonant frequency spectral shift due to uniform load. We also investigated the mechanical integrity of the flowmeter’s membrane through analytical modelling and COMSOL Structural Mechanics Module computer simulations. Both the analytic model and COMSOL model showed that maximum stresses on the plate, which are at the radial boundary of the plate, are three orders of magnitude smaller than the material’s yield strength and ultimate tensile strength. This indicates that the sensor is at a low risk of mechanical failure. Using results from models, we have developed an initial design for a microwave K-band sensor. A cylindrical resonator prototype was fabricated from brass for the initial tests. The external dimensions of the cavity are matched to the flange of a standard WR-42 waveguide. Microwave field is coupled into the resonant cavity through a subwavelength-size aperture. A test article was developed consisting of a piping Tee with a bulkhead WR-42 microwave waveguide installed in a leak-proof assembly. A microwave waveguide circulator was installed in the setup to suppress the effect of reflections at the cavity entrance by increasing the isolation between the input and the output port. Preliminary spectral characterization of cavity spectral response was performed with a portable PXIe chassis microwave VNA with a custom GUI. Preliminary dry tests of the transducer response were conducted with a set of calibrated weights. Transducer frequency shift was shown to be monotonically increasing with increasing pressure. The next steps will involve investigation of the transducer performance for water flow sensing.

42 ENGINEERING↗

A Portable Battery for Objective, Nonobtrusive Measures of Human Performance

A need exists for a standardized battery of human performance tests in order to measure the effects of various treatments. The present paper reports on progress in such a program, funded jointly by NASA and the Navy. Three batteries are available which differ in length (7.5, 15, and 30 minutes), and number of tests in the battery (3, 10, and 15). All tests are implemented on a portable, lap-held, briefcase-sized microprocessor (NEC PC 8201A). Performances measured include information processing, memory, visual perception, reasoning, motor skills, etc. Current programs are underway to determine norms, reliabilities, stabilities, factor structure of tests, comparisons with marker tests, apparatus suitability, etc.

Kennedy, Robert S.↗

Analytical comparisons of handheld LIBS and XRF devices for rapid quantification of gallium in a plutonium surrogate matrix

This work compares a portable laser-induced breakdown spectroscopy (LIBS) analyzer to a portable X-ray fluorescence (XRF) device for quantification of gallium (Ga) in a plutonium surrogate matrix of cerium (Ce) for the first time. Calibration methods are developed with spectra of Ce–Ga samples from both devices. Here, metrics such as limit of detection (LoD) and mean average percent error (MAPE) are examined to evaluate calibration performance. While the portable LIBS device can yield a nearly instantaneous analytical measurement, its accuracy is hampered by self-absorption. By employing a self-absorption correction and increasing gating delay, LIBS calibrations with errors in the low single percents and LoDs of 0.1% Ga were constructed. The XRF device produces calibrations with superlative sensitivity, yielding LoDs for gallium in the low tens of parts-per-million (ppm), two orders of magnitude lower than the corrected LIBS models. However, a clear trade-off of measurement fidelity is established between the instantaneous analysis of the LIBS device and the minutes-long XRF measurement yielding superior detection limits.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

INTEGRATION OF THE H3D-M400 DETECTOR AND SPOT ROBOT FOR AUTOMATED AREA SURVEY MISSIONS

Office of Nuclear Smuggling Detection and Deterrence (NSDD) is charged to identify and develop technologies to detect, disrupt, and investigate smuggling of radiological and nuclear materials •Border protection involves primary inspections using Radiation Portal Monitors (RPM), complimented with secondary and area survey inspections, usually performed manually using portable radiation detectors (PRDs) -RPM rely on fast technologies which provide quick scans of passing cargo/vehicles -Secondary inspections rely on trained, field deployed staff responsible for additional interdiction of screened cargo/vehicle putting them in potentially hazardous environment -Area survey inspections rely on trained, field deployed staff responsible for scanning and identifying presence of radiation while patrolling through a facility or public space •This project explores the use of technology which could reduce risks to field inspectors and add efficiency •The work presented here is a proof of principle of detector-robot integration to perform area survey inspections

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Experimental system for the control of surgically induced infections

The results are presented of the development tests performed on the experimental system for the control of surgically induced infections. Tests were performed on the portable clean room to demonstrate assembly, collapsability, portability and storage. Collapsing, relocating and storing within the surgery room can be accomplished in 12 minutes. The storage envelope dimensions are 1.64 m x 4.24 m x 2.62 m high. The disassembly transfer to another room, and reassembly were demonstrated. The laminar air flow velocity profile within the enclosure was measured. In the undisturbed area of the enclosure the air flow met the Federal Standard 209a requirements of 27.45 meters per minute + or - 6.10 meters per minute. Smoke tests with simulated surgery equipment and personnel in the enclosure did not indicate any detrimental air flow patterns. It is concluded that the system as designed will perform the functions required for its intended use.

Source record↗

Spectrally Tailored Pulsed Thulium Fiber Laser System for Broadband Lidar CO2 Sensing

Thulium doped pulsed fiber lasers are capable of meeting the spectral, temporal, efficiency, size and weight demands of defense and civil applications for pulsed lasers in the eye-safe spectral regime due to inherent mechanical stability, compact "all-fiber" master oscillator power amplifier (MOPA) architectures, high beam quality and efficiency. Thulium fiber's longer operating wavelength allows use of larger fiber cores without compromising beam quality, increasing potential single aperture pulse energies. Applications of these lasers include eye-safe laser ranging, frequency conversion to longer or shorter wavelengths for IR countermeasures and sensing applications with otherwise tough to achieve wavelengths and detection of atmospheric species including CO2 and water vapor. Performance of a portable thulium fiber laser system developed for CO2 sensing via a broadband lidar technique with an etalon based sensor will be discussed. The fielded laser operates with approximately 280 J pulse energy in 90-150ns pulses over a tunable 110nm spectral range and has a uniquely tailored broadband spectral output allowing the sensing of multiple CO2 lines simultaneously, simplifying future potentially space based CO2 sensing instruments by reducing the number and complexity of lasers required to carry out high precision sensing missions. Power scaling and future "all fiber" system configurations for a number of ranging, sensing, countermeasures and other yet to be defined applications by use of flexible spectral and temporal performance master oscillators will be discussed. The compact, low mass, robust, efficient and readily power scalable nature of "all-fiber" thulium lasers makes them ideal candidates for use in future space based sensing applications.

Heaps, William S.↗

Translational research in the MPICH project

The MPICH project is an example of translational research in computer science before that term was well known or even coined. The project began in 1992 as an effort to develop a portable, high-performance implementation of the emerging Message-Passing Interface (MPI) Standard. It has enabled the widespread adoption of MPI as a way to write scalable parallel applications on systems of all sizes including upcoming exascale supercomputers. In this paper, we describe how the translational research process was used in MPICH, how that led to its success, the challenges encountered and lessons learned, and how the process could be applied to other similar projects.

97 MATHEMATICS AND COMPUTING↗

MPICH

MPICH is a high-performance and widely portable implementation of the MPI-4 standard from the Argonne National Laboratory. This release has all MPI 4 functions and features required by the standard with the exception of support for the user-defined data representations for I/O.

ECP↗

PETSc Users Manual (Rev. 3.13)

This manual describes the use of PETSc for the numerical solution of partial differential equations and related problems on high-performance computers. The Portable, Extensible Toolkit for Scientific Computation (PETSc) is a suite of data structures and routines that provide the building blocks for the implementation of large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all message-passing communication. PETSc includes an expanding suite of parallel linear solvers, nonlinear solvers, and time integrators that may be used in application codes written in Fortran, C, C++, and Python. PETSc provides many of the mechanisms needed within parallel application codes, such as parallel matrix and vector assembly routines. The library is organized hierarchically, enabling users to employ the level of abstraction that is most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users. PETSc is a sophisticated set of software tools; as such, for some users it initially has a much steeper learning curve than a simple subroutine library. In particular, for individuals without some computer science background, experience programming in C, C++, python, or Fortran and experience using a debugger such as gdb or dbx, it may require a significant amount of time to take full advantage of the features that enable efficient software use. However, the power of the PETSc design and the algorithms it incorporates may make the efficient implementation of many application codes simpler than \rolling them" yourself.

97 MATHEMATICS AND COMPUTING↗

GASNet-EX Memory Kinds: Support for Device Memory in PGAS Programming Models

There is an emerging need for adaptive, lightweight communication in irregular HPC applications at exascale, where GPU accelerators provide the majority of available compute cycles. To address this need, Lawrence Berkeley National Lab is developing a programming system to support distributed-memory HPC application development using the Partitioned Global Address Space (PGAS) model. This work includes two major components: UPC++ and GASNet-EX. UPC++ is a C++ template library providing Remote Memory Access (RMA) and Remote Procedure Call (RPC) communication interfaces. GASNet-EX is a portable, high-performance communication middleware library, used by the implementations of UPC++ and many other PGAS programming models. We describe recent advances in GASNet-EX to efficiently implement zero-copy Remote Memory Access (RMA) communication to and from memory on accelerator devices such as GPUs. We demonstrate performance improvements via benchmark results from UPC++ (on Summit) and the Legion programming system (on DGX-1), both using GASNet-EX for communication.

Hargrove, Paul H↗