Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Fujitsu”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

Vision and "Hand-Eye" Coordination for the Fujitsu HOAP-2 Humanoid Robot

Software was developed for a Fujitsu HOAP-2 humanoid robot to demonstrate that such "human form" robots can be used to perform useful tasks in an unknown environment; specifically, that they can operate with or in the place of humans to carry out construction operations necessary for space exploration, such as building or maintaining planetary surface habitats for humans. This paper describes the autonomous vision and hand-eye coordination software that enables the HOAP-2 to locate and manipulate objects. Simple vision and pattern detection algorithms including color filtering, noise filtering, image segmentation, orientation filters, and a feedback controlled tracking system enable the robot to locate marked building materials in the workspace. Experimentally determined spatial mappings associate locations in the visual field with the arm joint angles and trajectories necessary to reach these locations, allowing the robot to reach for and manipulate the building materials. These capabilities, when combined with locomotion capabilities, enable the robot to operate autonomously in an unknown workspace.

Moore, Jeffrey D.↗

Satisfiability Test with Synchronous Simulated Annealing on the Fujitsu AP1000 Massively-Parallel Multiprocessor

Solving the hard Satisfiability Problem is time consuming even for modest-sized problem instances. Solving the Random L-SAT Problem is especially difficult due to the ratio of clauses to variables. This report presents a parallel synchronous simulated annealing method for solving the Random L-SAT Problem on a large-scale distributed-memory multiprocessor. In particular, we use a parallel synchronous simulated annealing procedure, called Generalized Speculative Computation, which guarantees the same decision sequence as sequential simulated annealing. To demonstrate the performance of the parallel method, we have selected problem instances varying in size from 100-variables/425-clauses to 5000-variables/21,250-clauses. Experimental results on the AP1000 multiprocessor indicate that our approach can satisfy 99.9 percent of the clauses while giving almost a 70-fold speedup on 500 processors.

Sohn, Andrew↗

Analysis of Vector Particle-In-Cell (VPIC) memory usage optimizations on cutting-edge computer architectures

Vector Particle-In-Cell (VPIC) is one of the fastest plasma simulation codes in the world, with particle numbers ranging from one trillion on the first petascale system, Roadrunner, to ten trillion particles on the more recent Blue Waters supercomputer. As supercomputers continue to grow rapidly in size, so too does the gap between computing capability and memory capability. Current memory systems limit VPIC simulations greatly as the maximum number of particles that can be simulated directly depends on the available memory. In this study, we present a suite of VPIC memory optimizations (i.e., particle weight, half-precision, and fixed-point optimizations) that enable a significant increase in the number of particles in VPIC simulations. Here, we assess the optimizations’ impact on memory and runtime performance for a suite of cutting-edge computer architectures such has the NVIDIA V100 GPU, the IBM Power9, and the Fujitsu A64FX architectures. Our optimizations enable a 31.25% reduction in memory usage and up to 40% increase in the number of particles. This paper extends our work on developing particle storage format optimizations Tan et al.

97 MATHEMATICS AND COMPUTING↗

3-regular three-XORSAT planted solutions benchmark of classical and quantum heuristic optimizers

With current semiconductor technology reaching its physical limits, special-purpose hardware has emerged as an option to tackle specific computing-intensive challenges. Optimization in the form of solving quadratic unconstrained binary optimization problems, or equivalently Ising spin glasses, has been the focus of several new dedicated hardware platforms. These platforms come in many different flavors, from highly-efficient hardware implementations on digital-logic of established algorithms to proposals of analog hardware implementing new algorithms. In this work, we use a mapping of a specific class of linear equations whose solutions can be found efficiently, to a hard constraint satisfaction problem (three-regular three-XORSAT, or an Ising spin glass) with a 'golf-course' shaped energy landscape, to benchmark several of these different approaches. We perform a scaling and prefactor analysis of the performance of Fujitsu's digital annealer unit (DAU), the D-Wave advantage quantum annealer, a virtual MemComputing machine, Toshiba's simulated bifurcation machine (SBM), the SATonGPU algorithm from Bernaschi et al, and our implementation of parallel tempering. We identify the SATonGPU and DAU as currently having the smallest scaling exponent for this benchmark, with SATonGPU having a small scaling advantage and in addition having by far the smallest prefactor thanks to its use of massive parallelism. Furthermore, our work provides an objective assessment and a snapshot of the promise and limitations of dedicated optimization hardware relative to a particular class of optimization problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

On Using Linux Kernel Huge Pages with FLASH, an Astrophysical Simulation Code

We present efforts at improving the performance of FLASH, a multi-scale, multi-physics simulation code principally for astrophysical applications, by using huge pages on Ookami, an HPE Apollo 80 A64FX platform. FLASH is written principally in modern Fortran and makes use of the PARAMESH library to manage a block-structured adaptive mesh. We explored options for enabling the use of huge pages with several compilers, but we were only able to successfully use huge pages when compiling with the Fujitsu compiler. As a result, the use of huge pages substantially reduced the number of translation lookaside buffer misses, but overall performance gains were marginal.

79 ASTRONOMY AND ASTROPHYSICS↗

Performance of an Astrophysical Radiation Hydrodynamics Code under Scalable Vector Extension Optimization

We present results of a performance study of an astrophysical radiation hydrodynamics code, V2D, on the Arm-based A64FX processor developed by Fujitsu. The code solves sparse linear systems, a task for which the A64FX architecture should be well suited. Here, we performed the performance analysis study on Ookami, an Apollo 80 platform utilizing the A64FX processor. We explored several compilers and performance anal-ysis packages and found the code did not perform as expected under scalable vector extension optimization, suggesting that a “deeper dive” into analyzing the code is worthwhile. However, a simple driver program that exercised basic sparse linear algebra routines used by V2D did show significant speedup with the use of the scalable vector extension optimization. We present the initial results from the study which used V2D on a relatively simple test problem that emphasized the repeated solution of sparse linear systems.

79 ASTRONOMY AND ASTROPHYSICS↗

Educating HPC Users in the use of advanced computing technology

We examine a multi-modal approach to educating and training users of an advanced computing technology testbed at the Institute for Advanced Computational Science at Stony Brook University. Ookami provides researchers worldwide with access to 176 Fujitsu A64FX compute nodes, this being the same processor technology powering the Japanese Fugaku supercomputer, the fastest computer in the world since June 2020. However, achieving high-performance on this Arm-based, leadership computing technology requires that users be familiar with details of computer architecture, performance analysis and modeling, and high-performance programming models that are commonly omitted in introductory programming courses. Indeed, regardless of their seniority, many of the testbed users are surprisingly unfamiliar with basic concepts such as vectorization, pipelining, latency/bandwidth, roofline models, computing energy/power, threads, and non-uniform memory access. These same concepts also pervade mainstream x86 technologies, so this is of widespread concern. Due to the national/global nature of our user community that is also very diverse in both discipline and experience, the inability to offer formal classes, and our experience that most people do not tend to read online documentation or training materials in sufficient depth, we have consciously employed multiple approaches that heavily emphasize (online) personal interactions and transfer of skills. Online documentation has been organized around best-practices and FAQs; twice-weekly hackathons and office hours via Zoom enable deep dives by both the team and the user community with multiple broad benefits; a Slack channel provides both real time and archived answers and discussions; and workshops, training and webinars target community needs as they arise. Furthermore, the perspective that these tools are being used in an educational setting rather than just for project communication makes them more effective and contributes to community success.

A64FX↗

Experiences with Porting the FLASH Code to Ookami, an HPE Apollo 80 A64FX Platform

We present initial experiences with running the community simulation code FLASH, developed at the University of Chicago for multi-scale multi-physics applications, on Ookami, a technology testbed featuring the A64FX processor developed by Fujitsu. Our effort focused largely on running FLASH “right out of the box” to see which combinations of compilers and software implementations (e.g. MPI) allowed the code to run with minimal modification. FLASH was one application in a larger effort to deploy Ookami; it served as a test for different versions of newly installed software, and as a cornerstone for the FAQ page of the Ookami website. Here, we report on our results with different compilers and other software, along with our initial scaling results and attempts to utilize the A64FX’s SVE instructions and NUMA architecture. We found that FLASH readily ran with different compilers and MPI implementations, and showed the expected good scaling with no turning. However, more work must be done to fully take advantage of the A64FX’s architectural features and produce a significant speedup for FLASH on Ookami.

79 ASTRONOMY AND ASTROPHYSICS↗

Parthenon—a performance portable block-structured adaptive mesh refinement framework

On the path to exascale the landscape of computer device architectures and corresponding programming models has become much more diverse. While various low-level performance portable programming models are available, support at the application level lacks behind. To address this issue, we present the performance portable block-structured adaptive mesh refinement (AMR) framework Parthenon, derived from the well-tested and widely used Athena++ astrophysical magnetohydrodynamics code, but generalized to serve as the foundation for a variety of downstream multi-physics codes. Parthenon adopts the Kokkos programming model, and provides various levels of abstractions from multidimensional variables, to packages defining and separating components, to launching of parallel compute kernels. Parthenon allocates all data in device memory to reduce data movement, supports the logical packing of variables and mesh blocks to reduce kernel launch overhead, and employs one-sided, asynchronous MPI calls to reduce communication overhead in multi-node simulations. Using a hydrodynamics miniapp, we demonstrate weak and strong scaling on various architectures including AMD and NVIDIA GPUs, Intel and AMD x86 CPUs, IBM Power9 CPUs, as well as Fujitsu A64FX CPUs. At the largest scale on Frontier (the first TOP500 exascale machine), the miniapp reaches a total of 1.7 × 10 13 zone-cycles/s on 9216 nodes (73,728 logical GPUs) at [Formula: see text] weak scaling parallel efficiency (starting from a single node). In combination with being an open, collaborative project, this makes Parthenon an ideal framework to target exascale simulations in which the downstream developers can focus on their specific application rather than on the complexity of handling massively-parallel, device-accelerated AMR.

97 MATHEMATICS AND COMPUTING↗

A cooled 1- to 2-GHz balanced HEMT amplifier

The design details and measurement results for a cooled L-band (1 to 2 GHz) balanced high electron mobility transistor (HEMT) amplifier are presented. The amplifier uses commercially available packaged HEMT devices (Fujitsu FHR02FH). At a physical temperature of 12 K, the amplifier achieves noise temperatures between 3 and 6 K over the 1 to 2 GHz band. The associated gain is approximately 20 dB.

G. G. Ortiz↗

A case for Redundant Arrays of Inexpensive Disks (RAID)

Increasing performance of CPUs and memories will be squandered if not matched by a similar performance increase in I/O. While the capacity of Single Large Expensive Disks (SLED) has grown rapidly, the performance improvement of SLED has been modest. Redundant Arrays of Inexpensive Disks (RAID), based on the magnetic disk technology developed for personal computers, offers an attractive alternative to SLED, promising improvements of an order of magnitude in performance, reliability, power consumption, and scalability. This paper introduces five levels of RAIDs, giving their relative cost/performance, and compares RAID to an IBM 3380 and a Fujitsu Super Eagle.

Patterson, David A.↗

Quality assurance and reliability in the Japanese electronics industry

Quality and reliability are two attributes required for all Japanese products, although the JTEC panel found these attributes to be secondary to customer cost requirements. While our Japanese hosts gave presentations on the challenges of technology, cost, and miniaturization, quality and reliability were infrequently the focus of our discussions. Quality and reliability were assumed to be sufficient to meet customer needs. Fujitsu's slogan, 'quality built-in, with cost and performance as prime consideration,' illustrates this point. Sony's definition of a next-generation product is 'one that is going to be half the size and half the price at the same performance of the existing one'. Quality and reliability are so integral to Japan's electronics industry that they need no new emphasis.

Pecht, Michael↗

NAS Parallel Benchmarks Results 3-95

The NAS Parallel Benchmarks (NPB) were developed in 1991 at NASA Ames Research Center to study the performance of parallel supercomputers. The eight benchmark problems are specified in a "pencil and paper" fashion, i.e., the complete details of the problem are given in a NAS technical document. Except for a few restrictions, benchmark implementors are free to select the language constructs and implementation techniques best suited for a particular system. In this paper, we present new NPB performance results for the following systems: (a) Parallel-Vector Processors: CRAY C90, CRAY T90 and Fujitsu VPP500; (b) Highly Parallel Processors: CRAY T3D, IBM SP2-WN (Wide Nodes), and IBM SP2-TN2 (Thin Nodes 2); and (c) Symmetric Multiprocessors: Convex Exemplar SPPIOOO, CRAY J90, DEC Alpha Server 8400 5/300, and SGI Power Challenge XL (75 MHz). We also present sustained performance per dollar for Class B LU, SP and BT benchmarks. We also mention future NAS plans for the NPB.

Saini, Subhash↗

NAS Parallel Benchmarks Results

The NAS Parallel Benchmarks (NPB) were developed in 1991 at NASA Ames Research Center to study the performance of parallel supercomputers. The eight benchmark problems are specified in a pencil and paper fashion i.e. the complete details of the problem to be solved are given in a technical document, and except for a few restrictions, benchmarkers are free to select the language constructs and implementation techniques best suited for a particular system. In this paper, we present new NPB performance results for the following systems: (a) Parallel-Vector Processors: Cray C90, Cray T'90 and Fujitsu VPP500; (b) Highly Parallel Processors: Cray T3D, IBM SP2 and IBM SP-TN2 (Thin Nodes 2); (c) Symmetric Multiprocessing Processors: Convex Exemplar SPP1000, Cray J90, DEC Alpha Server 8400 5/300, and SGI Power Challenge XL. We also present sustained performance per dollar for Class B LU, SP and BT benchmarks. We also mention NAS future plans of NPB.

Subhash, Saini↗

NAS Parallel Benchmark Results 11-96

The NAS Parallel Benchmarks have been developed at NASA Ames Research Center to study the performance of parallel supercomputers. The eight benchmark problems are specified in a "pencil and paper" fashion. In other words, the complete details of the problem to be solved are given in a technical document, and except for a few restrictions, benchmarkers are free to select the language constructs and implementation techniques best suited for a particular system. These results represent the best results that have been reported to us by the vendors for the specific 3 systems listed. In this report, we present new NPB (Version 1.0) performance results for the following systems: DEC Alpha Server 8400 5/440, Fujitsu VPP Series (VX, VPP300, and VPP700), HP/Convex Exemplar SPP2000, IBM RS/6000 SP P2SC node (120 MHz), NEC SX-4/32, SGI/CRAY T3E, SGI Origin200, and SGI Origin2000. We also report High Performance Fortran (HPF) based NPB results for IBM SP2 Wide Nodes, HP/Convex Exemplar SPP2000, and SGI/CRAY T3D. These results have been submitted by Applied Parallel Research (APR) and Portland Group Inc. (PGI). We also present sustained performance per dollar for Class B LU, SP and BT benchmarks.

Bailey, David H.↗