Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “performance portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

PLUM: Parallel Load Balancing for Unstructured Adaptive Meshes

Dynamic mesh adaption on unstructured grids is a powerful tool for computing large-scale problems that require grid modifications to efficiently resolve solution features. Unfortunately, an efficient parallel implementation is difficult to achieve, primarily due to the load imbalance created by the dynamically-changing nonuniform grid. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. First, we present an efficient parallel implementation of a tetrahedral mesh adaption scheme. Extremely promising parallel performance is achieved for various refinement and coarsening strategies on a realistic-sized domain. Next we describe PLUM, a novel method for dynamically balancing the processor workloads in adaptive grid computations. This research includes interfacing the parallel mesh adaption procedure based on actual flow solutions to a data remapping module, and incorporating an efficient parallel mesh repartitioner. A significant runtime improvement is achieved by observing that data movement for a refinement step should be performed after the edge-marking phase but before the actual subdivision. We also present optimal and heuristic remapping cost metrics that can accurately predict the total overhead for data redistribution. Several experiments are performed to verify the effectiveness of PLUM on sequences of dynamically adapted unstructured grids. Portability is demonstrated by presenting results on the two vastly different architectures of the SP2 and the Origin2OOO. Additionally, we evaluate the performance of five state-of-the-art partitioning algorithms that can be used within PLUM. It is shown that for certain classes of unsteady adaption, globally repartitioning the computational mesh produces higher quality results than diffusive repartitioning schemes. We also demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required a fine initial mesh. Results indicate that our parallel load balancing strategy will remain viable on large numbers of processors.

Oliker, Leonid↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Effects of Promethazine on Performance During Simulated Shuttle Landings

Promethazine (PMZ) is the antimotion sickness drug of choice in the U.S. Space Shuttle program; however, virtually nothing is known about the bioavailability and performance effects of this drug in the microgravity environment. PMZ has detrimental side effects on human performance on Earth that could affect Shuttle operations. In a recent ground-based study we examined: 1) the effects of promethazine (PMZ) on Shuttle landing performance using the portable inflight landing operations trainer (PILOT), and 2) saliva and urine samples to determine the pharmacokinetics of PMZ. The PILOT performance data is presented here.

Harm, D. L.↗

Improved thermal storage material for portable life support systems

The availability of thermal storage materials that have heat absorption capabilities substantially greater than water-ice in the same temperature range would permit significant improvements in performance of projected portable thermal storage cooling systems. A method for providing increased heat absorption by the combined use of the heat of solution of certain salts and the heat of fusion of water-ice was investigated. This work has indicated that a 30 percent solution of potassium bifluoride (KHF2) in water can absorb approximately 52 percent more heat than an equal weight of water-ice, and approximately 79 percent more heat than an equal volume of water-ice. The thermal storage material can be regenerated easily by freezing, however, a lower temperature must be used, 261 K as compared to 273 K for water-ice. This work was conducted by the United Aircraft Research Laboratories as part of a program at Hamilton Standard Division of United Aircraft Corporation under contract to NASA Ames Research Center.

Kellner, J. D.↗

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor↗

International time transfer and portable clock evaluation using GPS timing receivers: Preliminary results

The overall experiment was designed to test the positioning and navigation capabilities of the GPS timing receivers developed by the Naval Research Laboratory (NRL) for the NASA Goddard Laser Tracking Network (GITN). To perform this experiment, a reliable and redundant time scale was set up onboard the ship, and a back-up on shore. This situation provided the opportunity to perform simultaneously a timing experiment ideally divided into two parts, the main objectives of the experimentation being: (1) To test GPS timing receiver synchronization capabilities on a moving platform, and to perform an intercontinental synchronization via GPS between participating international timing laboratories in Europe and in the United States. (2) To evaluate the performance of cesium portable clocks in the field.

Wardrip, S. C.↗

Rasterization with Data-Parallel Primitives

Parallel rasterization can suffer from race conditions during fragment generation, which is traditionally addressed by using specialized hardware accessible via vendor graphics APIs. Unfortunately, graphics APIs are increasingly problematic on high-performance computers, either because they are not provided or because of concerns about dependencies with in situ visualization. In response, we present a hardware-agnostic rasterization algorithm that handles race conditions using only data-parallel primitives (DPPs), enabling efficient rendering on HPC systems without graphics API dependencies and aligning with recent efforts to deliver visualization software with DPPs. Our evaluation consists of three phases: (1) evaluating portability across different CPU and GPU architectures, (2) evaluating competitiveness with a community standard, and (3) evaluating performance across varying workloads and available parallelism. The supporting experiments run on both AMD and NVIDIA GPUs, considering data sets as large as 460 million triangles and 160 million pixels. While performance generally falls short of graphics API baselines, it achieves interactive frame rates on most workloads. As a result, we conclude our approach is a viable solution for rasterization on high-performance computers since our approach is portably performant across different architectures without the need for specialized vendor support.

Buckley, Makani [University of Oregon] (ORCID:0009↗

Short Exploration Extravehicular Mobility Unit Testing Setup: Evaluation Under Realistic Pressure and Thermal Conditions

The purpose of Short Exploration Extravehicular Mobility Unit (SxEMU) thermal vacuum testing was to verify the functionality of the Design Verification Testing (DVT) prototype xEMU (SxEMU for this test) at vacuum pressures and extreme space and lunar surface temperature conditions. The SxEMU Thermal Vacuum Test was the culmination of the DVT xEMU project. This paper’s focus is on the pre-Extravehicular Activity (EVA) test setup, and general performance of the SxEMU Portable Life Support Subsystem (xPLSS), with focus on the performance of the Primary Oxygen Assembly (POA) and Secondary Oxygen Assembly (SOA), including Secondary Oxygen Regulator (SOR) takeover and the POA and SOA low-setpoint change inhibit. The initial pre-EVA test preparation included recharging the batteries and replenishing consumables, including test-system water, oxygen assemblies (with gaseous nitrogen), and the integrated thermal loops, including the Feedwater Supply Assemblies. xPLSS functionality testing included carbon dioxide (CO2) removal via the Rapid Cycle Amine swingbed system, thermal loop temperature control, and monitoring of suit ventilation loop pressure, temperature, and CO2 percentages. Testing evaluated automatic takeover of suit pressure control by the SOR after the primary oxygen supply is depleted. The Primary Oxygen Regulator and SOR low-setpoint change inhibit function prevents the crewmember from inadvertently setting the primary regulator to a low pressure setpoint during an EVA.

xPLSS↗

A Portable Battery for Objective, Nonobtrusive Measures of Human Performance

A need exists for a standardized battery of human performance tests in order to measure the effects of various treatments. The present paper reports on progress in such a program, funded jointly by NASA and the Navy. Three batteries are available which differ in length (7.5, 15, and 30 minutes), and number of tests in the battery (3, 10, and 15). All tests are implemented on a portable, lap-held, briefcase-sized microprocessor (NEC PC 8201A). Performances measured include information processing, memory, visual perception, reasoning, motor skills, etc. Current programs are underway to determine norms, reliabilities, stabilities, factor structure of tests, comparisons with marker tests, apparatus suitability, etc.

Kennedy, Robert S.↗

Experimental system for the control of surgically induced infections

The results are presented of the development tests performed on the experimental system for the control of surgically induced infections. Tests were performed on the portable clean room to demonstrate assembly, collapsability, portability and storage. Collapsing, relocating and storing within the surgery room can be accomplished in 12 minutes. The storage envelope dimensions are 1.64 m x 4.24 m x 2.62 m high. The disassembly transfer to another room, and reassembly were demonstrated. The laminar air flow velocity profile within the enclosure was measured. In the undisturbed area of the enclosure the air flow met the Federal Standard 209a requirements of 27.45 meters per minute + or - 6.10 meters per minute. Smoke tests with simulated surgery equipment and personnel in the enclosure did not indicate any detrimental air flow patterns. It is concluded that the system as designed will perform the functions required for its intended use.

Source record↗

Spectrally Tailored Pulsed Thulium Fiber Laser System for Broadband Lidar CO2 Sensing

Thulium doped pulsed fiber lasers are capable of meeting the spectral, temporal, efficiency, size and weight demands of defense and civil applications for pulsed lasers in the eye-safe spectral regime due to inherent mechanical stability, compact "all-fiber" master oscillator power amplifier (MOPA) architectures, high beam quality and efficiency. Thulium fiber's longer operating wavelength allows use of larger fiber cores without compromising beam quality, increasing potential single aperture pulse energies. Applications of these lasers include eye-safe laser ranging, frequency conversion to longer or shorter wavelengths for IR countermeasures and sensing applications with otherwise tough to achieve wavelengths and detection of atmospheric species including CO2 and water vapor. Performance of a portable thulium fiber laser system developed for CO2 sensing via a broadband lidar technique with an etalon based sensor will be discussed. The fielded laser operates with approximately 280 J pulse energy in 90-150ns pulses over a tunable 110nm spectral range and has a uniquely tailored broadband spectral output allowing the sensing of multiple CO2 lines simultaneously, simplifying future potentially space based CO2 sensing instruments by reducing the number and complexity of lasers required to carry out high precision sensing missions. Power scaling and future "all fiber" system configurations for a number of ranging, sensing, countermeasures and other yet to be defined applications by use of flexible spectral and temporal performance master oscillators will be discussed. The compact, low mass, robust, efficient and readily power scalable nature of "all-fiber" thulium lasers makes them ideal candidates for use in future space based sensing applications.

Heaps, William S.↗

Time dissemination: An update

A brief description of some of the improvements in the generation and dissemination of precise time and time interval at the U.S. Naval Observation are described. Details on some of the newer hardware developed for this purpose are presented. Data from tests of two of the more significant items, a small portable clock with performance approaching that of larger units and a GPS Time Transfer Unit capable of time transfer to an accuracy of less than 100 nanoseconds, are also given.

Putkovich, K.↗

Development of a miniaturized gas chromatograph-mass spectrometer with a microbore capillary column and an array detector

A review is presented of the demonstration of a miniaturized focal plane mass spectrograph using a microbore capillary column with an array detector to measure multiple mass spectra from narrow and closely eluted gas chromatograph peaks. It is shown that the system possesses high sensitivity and a high speed of spectral mass measurements. The combination of the capillary column with the miniaturized mass spectrograph is uniquely suited for the development of a field-portable, high-performance system.

Sinha, Mahadeva P.↗

Building the electronic industry's roadmaps

JTEC panelists found a strong consistency among the electronics firms they visited: all the firms had clear visions or roadmaps for their research and development activities and had committed resources to ensure that they achieve targeted results. The overarching vision driving Japan's electronics industry is that of achieving market success through developing appealing, high-quality, low-cost consumer goods - ahead of the competition. Specifics of the vision include improving performance, quality, and portability of consumer electronics products. Such visions help Japanese companies define in detail the roadmaps they will follow to develop new and improved electronic packaging technologies.

Boulton, William R.↗

On Designing Lightweight Threads for Substrate Software

Existing user-level thread packages employ a 'black box' design approach, where the implementation of the threads is hidden from the user. While this approach is often sufficient for application-level programmers, it hides critical design decisions that system-level programmers must be able to change in order to provide efficient service for high-level systems. By applying the principles of Open Implementation Analysis and Design, we construct a new user-level threads package that supports common thread abstractions and a well-defined meta-interface for altering the behavior of these abstractions. As a result, system-level programmers will have the advantages of using high-level thread abstractions without having to sacrifice performance, flexibility or portability.

Haines, Matthew↗

Load Balancing Sequences of Unstructured Adaptive Grids

Mesh adaption is a powerful tool for efficient unstructured grid computations but causes load imbalance on multiprocessor systems. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. This paper makes several important additions to our previous work. First, a new remapping cost model is presented and empirically validated on an SP2. Next, our load balancing strategy is applied to sequences of dynamically adapted unstructured grids. Results indicate that our framework is effective on many processors for both steady and unsteady problems with several levels of adaption. Additionally, we demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required for a fine initial mesh. Finally, we show that the data remapping overhead can be significantly reduced by applying our heuristic processor reassignment algorithm.

Biswas, Rupak↗

Processors, Pipelines, and Protocols for Advanced Modeling Networks

Predictive capabilities arise from our understanding of natural processes and our ability to construct models that accurately reproduce these processes. Although our modeling state-of-the-art is primarily limited by existing computational capabilities, other technical areas will soon present obstacles to the development and deployment of future predictive capabilities. Advancement of our modeling capabilities will require not only faster processors, but new processing algorithms, high-speed data pipelines, and a common software engineering framework that allows networking of diverse models that represent the many components of Earth's climate and weather system. Development and integration of these new capabilities will pose serious challenges to the Information Systems (IS) technology community. Designers of future IS infrastructures must deal with issues that include performance, reliability, interoperability, portability of data and software, and ultimately, the full integration of various ES model systems into a unified ES modeling network.

Coughlan, Joseph↗