Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HIP programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Evaluation of OpenAI Codex for HPC Parallel Programming Models Kernel Generation

We evaluate AI-assisted generative capabilities on fundamental numerical kernels in high-performance computing (HPC), including AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG. We test the generated kernel codes for a variety of language-supported programming models, including (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). We use the GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code as of April 2023 to generate a vast amount of implementations given simple + + prompt variants. To quantify and compare the results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. Results suggest that the OpenAI Codex outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general purpose Python can benefit from adding code keywords, while Julia prompts perform acceptably well for its mature programming models (e.g., Threads and CUDA.jl). We expect for these benchmarks to provide a point of reference for each programming model's community. Overall, understanding the convergence of large language models, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

Godoy, William↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

Powder Metallurgy – Hot Isostatic Pressing of 316H Stainless Steel for Nuclear Components

Overview of the PM-HIP work performed on stainless steel 316H as part of the microreactor program. High temperature mechanical testing was performed, with concerns identified with the creep-fatigue performance. Variations in feedstock and processing parameters did not conclusively identify the primary factor in the reduced creep-fatigue performance.

36 - MATERIALS SCIENCE↗

A Peculiar Class of Debris Disks from Herschel/DUNES: A Steep Fall Off in the Far Infrared

Context. The existence of debris disks around old main sequence stars is usually explained by continuous replenishment of small dust grains through collisions from a reservoir of larger objects. Aims. We present photometric data of debris disks around HIP 103389 (HD199260), HIP 100350 (HN Peg, HD206860), and HIP 114948 (HD 219482), obtained in the context of our Herschel Open TIme Key Program DUNES (DUst around NEarby Stars). Methods. We used Herschel/PACS to detect the thermal emission of the three debris disks with a 30 sigma sensitivity of a few mJy at l00 micron and 160 micron. In addition, we obtained Herschel/PACS photometric data at 70 micron for HIP 103389. These observations are complemented by a large variety of optical to far-infrared photometric data. Two different approaches are applied to reduce the Herschel data to investigate the impact of data reduction on the photometry. We fit analytical models to the available spectral energy distribution (SED) data using the fitting method of simulated therma1 annealing as well as a classical grid search method. Results. The SEDs of the three disks potentially exhibit an unusually steep decrease at wavelengths >= 70 micron. We investigate the significance of the peculiar shape of these SEDs and the impact on models of the disks provided it is real. Using grain compositions that have been applied successfully for modeling of many other debris disks, our modeling reveals that such a steep decrease of the SEDs in the long wavelength regime is inconsistent with a power-law exponent of the grain size distribution -3.5 expected from a standard equilibrium collisional cascade. In contrast, a steep grain size distribution or, alternatively an upper grain size in the range of few tens of micrometers are implied. This suggests that a very distinct range of grain sizes would dominate the thermal. emission of such disks. However, we demonstrate that the understanding of the data of faint sources obtained with Herschel is still incomplete and that the significance of our results depends on the version of the data reduction pipeline used. Conclusions. A new mechanism to produce the dust in the presented debris disks, deviations from the conditions required for a standard equilibrium collisional cascade (grain size exponent of -3.5), and/or significantly different dust properties would be necessary to explain the potentially steep SED shape of the three debris disks presented.

Ertel, S.↗

Recommended Methods for Monitoring Skeletal Health in Astronauts to Distinguish Specific Effects of Prolonged Spaceflight

NASA uses areal bone mineral density (aBMD) by dual-energy X-ray absorptiometry (DXA) to monitor skeletal health in astronauts after typical 180-day spaceflights. The osteoporosis field and NASA, however, recognize the insufficiency of DXA aBMD as a sole surrogate for fracture risk. This is an even greater concern for NASA as it attempts to expand fracture risk assessment in astronauts, given the complicated nature of spaceflight-induced bone changes and the fact that multiple 1-year missions are planned. In the past decade, emerging analyses for additional surrogates have been tested in clinical trials; the potential use of these technologies to monitor the biomechanical integrity of the astronaut skeleton will be presented. OVERVIEW: An advisory panel of osteoporosis policy-makers provided NASA with an evidence-based assessment of astronaut biomedical and research data. The panel concluded that spaceflight and terrestrial bone loss have significant differences and certain factors may predispose astronauts to premature fractures. Based on these concerns, a proposed surveillance program is presented which a) uses Quantitative Computed Tomography (QCT) scans of the hip to monitor the recovery of spaceflight-induced deficits in trabecular BMD by 2 years after return, b) develops Finite Element Models [FEM] of QCT data to evaluate spaceflight effect on calculated hip bone strength and c) generates Trabecular Bone Score [TBS] from serial DXA scans of the lumbar spine to evaluate the effect of age, spaceflight and countermeasures on this novel index of bone microarchitecture. SIGNIFICANCE: DXA aBMD is a widely-applied, evidence-based predictor for fractures but not applicable as a fracture surrogate for premenopausal females and males <50 years. Its inability to detect structural parameters is a limitation for assessing changes in bone integrity with and without countermeasures. Collective use of aBMD, TBS, QCT, and FEM analysis for astronaut surveillance could accommodate NASA's aggressive schedule for risk definition and inform a NASA-developed model which assesses the probability of overloading bones during mechanically-loaded mission tasks and possibly for physical activities after return to Earth.

Vasadi, Lukas J.↗

Advanced single crystal for SSME turbopumps

The objective of this program was to evaluate the influence of high thermal gradient casting, hot isostatic pressing (HIP) and alternate heat treatments on the microstructure and mechanical properties of a single crystal nickel base superalloy. The alloy chosen for the study was PWA 1480, a well characterized, commercial alloy which had previously been chosen as a candidate for the Space Shuttle Main Engine high pressure turbopump turbine blades. Microstructural characterization evaluated the influence of casting thermal gradient on dendrite arm spacing, casting porosity distribution and alloy homogeneity. Hot isostatic pressing was evaluated as a means of eliminating porosity as a preferred fatigue crack initiation site. The alternate heat treatment was chosen to improve hydrogen environment embrittlement resistance and for potential fatigue life improvement. Mechanical property evaluation was aimed primarily at determining improvements in low cycle and high cycle fatigue life due to the advanced processing methods. Statistically significant numbers of tests were conducted to quantitatively demonstrate life differences. High thermal gradient casting improves as-cast homogeneity, which facilitates solution heat treatment of PWA 1480 and provides a decrease in internal pore size, leading to increases in low cycle and high cycle fatigue lives.

Fritzemeier, L. G.↗

Performance Evaluation of Heterogeneous GPU Programming Frameworks for Hemodynamic Simulations

Preparing for the deployment of large scientific and engineering codes on upcoming exascale systems with GPU-dense nodes is made challenging by the unprecedented diversity of device architectures and heterogeneous programming models. In this work, we evaluate the process of porting a massively parallel, fluid dynamics code written in CUDA to SYCL, HIP, and Kokkos with a range of backends, using a combination of automated tools and manual tuning. We use a proxy application along with a custom performance model to inform the results and identify additional optimization strategies. At scale performance of the programming model implementations are evaluated on pre-production GPU node architectures for Frontier and Aurora, as well as on current NVIDIA device-based systems Summit and Polaris. Real-world workloads representing 3D blood flow calculations in complex vasculature are assessed. Our analysis highlights critical trade-offs between code performance, portability, and development time.

Martin, Aristotle↗

FloatGuard: Efficient Whole-Program Detection of Floating-Point Exceptions in AMD GPUs

FloatGuard is a tool that captures floating-point exceptions in AMD HIP kernels. FloatGuard leverages AMD GPU hardware registers to detect floating-point exceptions, overcoming the limitations of AMD's built-in trapping mechanisms through a novel algorithm that combines assembly- and source-level instrumentation with debugger-guided execution.

MIAO, WENJUN [Lawrence Livermore National Laborato↗

Skeletal Adaptations to Different Levels of Eccentric Resistance Following Eight Weeks of Training

Coupled concentric-eccentric resistive exercise maintains bone mineral density (BMD) during bed rest and aging. PURPOSE: We hypothesized that 8 wks of lower body resistive exercise training with higher ratios of eccentric to concentric loading would enhance hip and lumbar BMD. METHODS: Forty untrained male volunteers (34.9+/-7.0 yrs, 80.9+/-9.8 kg, 178.2+/-7.1 cm; mean+/-SD) were matched for leg press (LP) 1-Repetition Maximum (1-RM) strength and randomly assigned to one of 5 training groups. Concentric load (% 1-RM) was constant across groups, but each group trained with different levels of eccentric load (0, 33, 66, 100, or 138% of concentric) for all training sessions. Subjects performed a periodized supine LP and heel raise (HR) training program 3 d wk-1 for 8 wks using a modified Agaton Fitness System (Agaton Fitness AB, Boden, Sweden). Hip and lumbar BMD (g/sq cm) was measured in triplicate pre- and post-training using DXA (Hologic Discovery ). Pre- and post-training means were compared using the appropriate ANOVA and Tukey's post hoc tests. Within group pre- to post-training BMD was compared using paired t-tests with a Bonferroni adjustment. RESULTS: There was a main effect of training on L1, L2, L3, L4, total lumbar, and greater trochanter BMD, but there were no differences between groups. CONCLUSION: Eights wks of lower body resistive exercise increased greater trochanter and lumbar BMD. Inability to detect group differences may have been influenced by a potentially osteogenic vibration associated with device operation in the 0, 33, and 66% groups.

English, Kirk L.↗

Application of superalloy powder metallurgy for aircraft engines

In the last decade, Government/Industry programs have advanced powder metallurgy-near-net-shape technology to permit the use of hot isostatic pressed (HIP) turbine disks in the commercial aircraft fleet. These disks offer a 30% savings of input weight and an 8% savings in cost compared in cast-and-wrought disks. Similar savings were demonstrated for other rotating engine components. A compressor rotor fabricated from hot-die-forged-HIP superalloy billets revealed input weight savings of 54% and cost savings of 35% compared to cast-and-wrought parts. Engine components can be produced from compositions such as Rene 95 and Astroloy by conventional casting and forging, by forging of HIP powder billets, or by direct consolidation of powder by HIP. However, each process produces differences in microstructure or introduces different defects in the parts. As a result, their mechanical properties are not necessarily identical. Acceptance methods should be developed which recognize and account for the differences.

Dreshfield, R. L.↗

Application of superalloy powder metallurgy for aircraft engines

The results of the Materials for Advanced Turbine Engines (MATE) program initiated by NASA are presented. Mechanical properties comparisons are made for superalloy parts produced by as-HIP powder consolidation and by forging of HIP consolidated billets. The effect of various defects on the mechanical properties of powder parts are shown.

Dreshfield, R. L.↗

Evaluating performance and portability of high-level programming models: Julia, Python/Numba, and Kokkos on exascale nodes

We explore the performance and portability of the high-level programming models: the LLVM-based Julia and Python/Numba, and Kokkos on high-performance computing (HPC) nodes: AMD Epyc CPUs and MI250X graphical processing units (GPUs) on Frontier’s test bed Crusher system and Ampere’s Arm-based CPUs and NVIDIA’s A100 GPUs on the Wombat system at the Oak Ridge Leadership Computing Facilities. We compare the default performance of a hand-rolled dense matrix multiplication algorithm on CPUs against vendor-compiled C/OpenMP implementations, and on each GPU against CUDA and HIP. Rather than focusing on the kernel optimization per-se, we select this naive approach to resemble exploratory work in science and as a lower-bound for performance to isolate the effect of each programming model. Julia and Kokkos perform comparably with C/OpenMP on CPUs, while Julia implementations are competitive with CUDA and HIP on GPUs. Performance gaps are identified on NVIDIA A100 GPUs for Julia’s single precision and Kokkos, and for Python/Numba in all scenarios. We also comment on half-precision support, productivity, performance portability metrics, and platform readiness. We expect to contribute to the understanding and direction for high-level, high-productivity languages in HPC as the first-generation exascale systems are deployed.

Godoy, William↗

ISS Squat and Deadlift Kinematics on the Advanced Resistive Exercise Device

Visual assessment of exercise form on the Advanced Resistive Exercise Device (ARED) on orbit is difficult due to the motion of the entire device on its Vibration Isolation System (VIS). The VIS allows for two degrees of device translational motion, and one degree of rotational motion. In order to minimize the forces that the VIS must damp in these planes of motion, the floor of the ARED moves as well during exercise to reduce changes in the center of mass of the system. To help trainers and other exercise personnel better assess squat and deadlift form a tool was developed that removes the VIS motion and creates a stick figure video of the exerciser. Another goal of the study was to determine whether any useful kinematic information could be obtained from just a single camera. Finally, the use of these data may aid in the interpretation of QCT hip structure data in response to ARED exercises performed in-flight. After obtaining informed consent, four International Space Station (ISS) crewmembers participated in this investigation. Exercise was videotaped using a single camera positioned to view the side of the crewmember during exercise on the ARED. One crewmember wore reflective tape on the toe, heel, ankle, knee, hip, and shoulder joints. This technique was not available for the other three crewmembers, so joint locations were assessed and digitized frame-by-frame by lab personnel. A custom Matlab program was used to assign two-dimensional coordinates to the joint locations throughout exercise. A second custom Matlab program was used to scale the data, calculate joint angles, estimate the foot center of pressure (COP), approximate normal and shear loads, and to create the VIS motion-corrected stick figure videos. Kinematics for the squat and deadlift vary considerably for the four crewmembers in this investigation. Some have very shallow knee and hip angles, and others have quite large ranges of motion at these joints. Joint angle analysis showed that crewmembers do not return to a normal upright stance during squat, but remain somewhat bent at the hips. COP excursions were quite large during these exercises covering the entire length of the base of support in most cases. Anterior-posterior shear was very pronounced at the bottom of the squat and deadlift correlating with a COP shift to the toes at this part of the exercise. The stick figure videos showing a feet fixed reference frame have made it visually much easier for exercise personnel and trainers to assess exercise kinematics. Not returning to fully upright, hips extended position during squat exercises could have implications for the amount of load that is transmitted axially along the skeleton. The estimated shear loads observed in these crewmembers, along with a concomitant reduction in normal force, may also affect bone loading. The increased shear is likely due to the surprisingly large deviations in COP. Since the footplate on ARED moves along an arced path, much of the squat and deadlift movement is occurring on a tilted foot surface. This leads to COP movements away from the heel. The combination of observed kinematics and estimated kinetics make squat and deadlift exercises on the ARED distinctly different from their ground-based counterparts. CONCLUSION This investigation showed that some useful exercise information can be obtained at low cost, using a single video camera that is readily available on ISS. Squat and deadlift kinematics on the ISS ARED differ from ground-based ARED exercise. The amount of COP shift during these exercises sometimes approaches the limit of stability leading to modifications in the kinematics. The COP movement and altered kinematics likely reduce the bone loading experienced during these exercises. Further, the stick figure videos may prove to be a useful tool in assisting trainers to identify exercise form and make suggestions for improvements

Newby, N.↗

Spinoff 2010

Topics covered include: Burnishing Techniques Strengthen Hip Implants; Signal Processing Methods Monitor Cranial Pressure; Ultraviolet-Blocking Lenses Protect, Enhance Vision; Hyperspectral Systems Increase Imaging Capabilities; Programs Model the Future of Air Traffic Management; Tail Rotor Airfoils Stabilize Helicopters, Reduce Noise; Personal Aircraft Point to the Future of Transportation; Ducted Fan Designs Lead to Potential New Vehicles; Winglets Save Billions of Dollars in Fuel Costs; Sensor Systems Collect Critical Aerodynamics Data; Coatings Extend Life of Engines and Infrastructure; Radiometers Optimize Local Weather Prediction; Energy-Efficient Systems Eliminate Icing Danger for UAVs; Rocket-Powered Parachutes Rescue Entire Planes; Technologies Advance UAVs for Science, Military; Inflatable Antennas Support Emergency Communication; Smart Sensors Assess Structural Health; Hand-Held Devices Detect Explosives and Chemical Agents; Terahertz Tools Advance Imaging for Security, Industry; LED Systems Target Plant Growth; Aerogels Insulate Against Extreme Temperatures; Image Sensors Enhance Camera Technologies; Lightweight Material Patches Allow for Quick Repairs; Nanomaterials Transform Hairstyling Tools; Do-It-Yourself Additives Recharge Auto Air Conditioning; Systems Analyze Water Quality in Real Time; Compact Radiometers Expand Climate Knowledge; Energy Servers Deliver Clean, Affordable Power; Solutions Remediate Contaminated Groundwater; Bacteria Provide Cleanup of Oil Spills, Wastewater; Reflective Coatings Protect People and Animals; Innovative Techniques Simplify Vibration Analysis; Modeling Tools Predict Flow in Fluid Dynamics; Verification Tools Secure Online Shopping, Banking; Toolsets Maintain Health of Complex Systems; Framework Resources Multiply Computing Power; Tools Automate Spacecraft Testing, Operation; GPS Software Packages Deliver Positioning Solutions; Solid-State Recorders Enhance Scientific Data Collection; Computer Models Simulate Fine Particle Dispersion; Composite Sandwich Technologies Lighten Components; Cameras Reveal Elements in the Short Wave Infrared; Deformable Mirrors Correct Optical Distortions; Stitching Techniques Advance Optics Manufacturing; Compact, Robust Chips Integrate Optical Functions; Fuel Cell Stations Automate Processes, Catalyst Testing; Onboard Systems Record Unique Videos of Space Missions; Space Research Results Purify Semiconductor Materials; and Toolkits Control Motion of Complex Robotics.

Source record↗

Mojo: MLIR-based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem

We explore the performance and portability of the novel Mojo language for scientific computing workloads on GPUs. As the first language based on the LLVM’s Multi-Level Intermediate Representation (MLIR) compiler infrastructure, Mojo aims to close performance and productivity gaps by combining Python’s interoperability and CUDA-like syntax for compile-time portable GPU programming. We target four scientific workloads: a seven-point stencil (memory-bound), BabelStream (memory-bound), miniBUDE (compute-bound), and Hartree–Fock (compute-bound with atomic operations); and compare their performance against vendor baselines on NVIDIA H100 and AMD MI300A GPUs. We show that Mojo’s performance is competitive with CUDA and HIP for memory-bound kernels, whereas gaps exist on AMD GPUs for atomic operations and for fast-math compute-bound kernels on both AMD and NVIDIA GPUs. Although the learning curve and programming requirements are still fairly low-level, Mojo can close significant gaps in the fragmented Python ecosystem in the convergence of scientific computing and AI.

Godoy, William [ORNL] (ORCID:0000000225905178)↗

Potential Multi-Component Structure of the Debris Disk Around HIP 17439 Revealed by Herschel DUNES

Context. The dust observed in debris disks is produced through collisions of larger bodies left over from the planet/planetesimal formation process. Spatially resolving these disks permits to constrain their architecture and thus that of the underlying planetary/planetesimal system. Aims. Our Herschel open time key program DUNES aims at detecting and characterizing debris disks around nearby, sun-like stars. In addition to the statistical analysis of the data, the detailed study of single objects through spatially resolving the disk and detailed modeling of the data is a main goal of the project. Methods. We obtained the first observations spatially resolving the debris disk around the sun-like star HIP 17439 (HD 23484) using the instruments PACS and SPIRE on board the Herschel Space Observatory. Simultaneous multi-wavelength modeling of these data together with ancillary data from the literature is presented. Results. A standard single component disk model fails to reproduce the major axis radial profiles at 70 μm, 100 μm, and 160 μm simultaneously. Moreover, the best-fit parameters derived from such a model suggest a very broad disk extending from few au up to few hundreds of au from the star with a nearly constant surface density which seems physically unlikely. However, the constraints from both the data and our limited theoretical investigation are not strong enough to completely rule out this model. An alternative, more plausible, and better fitting model of the system consists of two rings of dust at approx. 30 au and 90 au, respectively, while the constraints on the parameters of this model are weak due to its complexity and intrinsic degeneracies. Conclusions. The disk is probably composed of at least two components with different spatial locations (but not necessarily detached), while a single, broad disk is possible, but less likely. The two spatially well-separated rings of dust in our best-fit model suggest the presence of at least one high mass planet or several low-mass planets clearing the region between the two rings from planetesimals and dust.

sun-like stars↗

Energy efficient engine. Volume 2. Appendix A: Component development and integration program

The large size and the requirement for precise lightening cavities in a considerable portion of the titanium fan blades necessitated the development of a new manufacturing method. The approach which was selected for development incorporated several technologies including HIP diffusion bonding of titanium sheet laminates containing removable cores and isothermal forging of the blade form. The technology bases established in HIP/DB for composite blades and in isothermal forging for fan blades were applicable for development of the manufacturing process. The process techniques and parameters for producing and inspecting the cored diffusion bonded titanium laminate blade preform were established. The method was demonstrated with the production of twelve hollow simulated blade shapes for evaluation. Evaluations of the critical experiments conducted to establish procedures to produce hollow structures by a laminate/core/diffusion bonding approach are included. In addition the transfer of this technology to produce a hollow fan blade is discussed.

Moracz, D. J.↗