Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Performance Evaluation and Modeling Techniques for Parallel Processors

In practice, the performance evaluation of supercomputers is still substantially driven by singlepoint estimates of metrics (e.g., MFLOPS) obtained by running characteristic benchmarks or workloads. With the rapid increase in the use of time-shared multiprogramming in these systems, such measurements are clearly inadequate. This is because multiprogramming and system overhead, as well as other degradations in performance due to time varying characteristics of workloads, are not taken into account. In multiprogrammed environments, multiple jobs and users can dramatically increase the amount of system overhead and degrade the performance of the machine. Performance techniques, such as benchmarking, which characterize performance on a dedicated machine ignore this major component of true computer performance. Due to the complexity of analysis, there has been little work done in analyzing, modeling, and predicting the performance of applications in multiprogrammed environments. This is especially true for parallel processors, where the costs and benefits of multi-user workloads are exacerbated. While some may claim that the issue of multiprogramming is not a viable one in the supercomputer market, experience shows otherwise. Even in recent massively parallel machines, multiprogramming is a key component. It has even been claimed that a partial cause of the demise of the CM2 was the fact that it did not efficiently support time-sharing. In the same paper, Gordon Bell postulates that, multicomputers will evolve to multiprocessors in order to support efficient multiprogramming. Therefore, it is clear that parallel processors of the future will be required to offer the user a time-shared environment with reasonable response times for the applications. In this type of environment, the most important performance metric is the completion of response time of a given application. However, there are a few evaluation efforts addressing this issue.

Dimpsey, Robert Tod↗

Ground-based PIV and numerical flow visualization results from the Surface Tension Driven Convection Experiment

The Surface Tension Driven Convection Experiment (STDCE) is a Space Transportation System flight experiment to study both transient and steady thermocapillary fluid flows aboard the United States Microgravity Laboratory-1 (USML-1) Spacelab mission planned for June, 1992. One of the components of data collected during the experiment is a video record of the flow field. This qualitative data is then quantified using an all electric, two dimensional Particle Image Velocimetry (PIV) technique called Particle Displacement Tracking (PDT), which uses a simple space domain particle tracking algorithm. Results using the ground based STDCE hardware, with a radiant flux heating mode, and the PDT system are compared to numerical solutions obtained by solving the axisymmetric Navier Stokes equations with a deformable free surface. The PDT technique is successful in producing a velocity vector field and corresponding stream function from the raw video data which satisfactorily represents the physical flow. A numerical program is used to compute the velocity field and corresponding stream function under identical conditions. Both the PDT system and numerical results were compared to a streak photograph, used as a benchmark, with good correlation.

Pline, Alexander D.↗

Performance Measurement, Visualization and Modeling of Parallel and Distributed Programs

This paper presents a methodology for debugging the performance of message-passing programs on both tightly coupled and loosely coupled distributed-memory machines. The AIMS (Automated Instrumentation and Monitoring System) toolkit, a suite of software tools for measurement and analysis of performance, is introduced and its application illustrated using several benchmark programs drawn from the field of computational fluid dynamics. AIMS includes (i) Xinstrument, a powerful source-code instrumentor, which supports both Fortran77 and C as well as a number of different message-passing libraries including Intel's NX Thinking Machines' CMMD, and PVM; (ii) Monitor, a library of timestamping and trace -collection routines that run on supercomputers (such as Intel's iPSC/860, Delta, and Paragon and Thinking Machines' CM5) as well as on networks of workstations (including Convex Cluster and SparcStations connected by a LAN); (iii) Visualization Kernel, a trace-animation facility that supports source-code clickback, simultaneous visualization of computation and communication patterns, as well as analysis of data movements; (iv) Statistics Kernel, an advanced profiling facility, that associates a variety of performance data with various syntactic components of a parallel program; (v) Index Kernel, a diagnostic tool that helps pinpoint performance bottlenecks through the use of abstract indices; (vi) Modeling Kernel, a facility for automated modeling of message-passing programs that supports both simulation -based and analytical approaches to performance prediction and scalability analysis; (vii) Intrusion Compensator, a utility for recovering true performance from observed performance by removing the overheads of monitoring and their effects on the communication pattern of the program; and (viii) Compatibility Tools, that convert AIMS-generated traces into formats used by other performance-visualization tools, such as ParaGraph, Pablo, and certain AVS/Explorer modules.

Yan, Jerry C.↗

Applications of Flow Control to Wing High-Lift Leading Edge Devices on a Commercial Aircraft

Active flow control was applied to the leading edge region of a representative future short/medium-range twin-engine airplane to improve aerodynamic performance during high-lift operations. The study is aimed at enhanced lift over the practical angle of attack range, including stall, and at reduced drag. These benefits translate to airplane performance improvements, such as longer range or larger payload. Various flow control applications were explored using Computational Fluid Dynamics and the aerodynamic performance enhancements were benchmarked against the baseline configuration. The computational analyses are used to quantify aerodynamic benefits, as well as the input required for actuation. The results were used in a system integration study for identifying potential practical implementations, which are described in a companion paper. Combined with the integration analysis, the objective of this project is to identify the most promising flow control candidates that potentially provide material net airplane level enhancements using onboard fluidic sources. Depending on the implementation of active flow control, the current study indicates that up to 1.5% net improvement in L/D at takeoff and 4% increase in maximum lift during landing are potentially achievable, after accounting for factors of system integration.

CFD↗

Flow Control for Enhanced Aileron Effectiveness on a Commercial Aircraft

Active flow control was applied to the ailerons of a representative future short/medium-range twin-engine airplane to improve aerodynamic performance during high-lift operations. The study is aimed at reduced drag and enhanced lift over the range of practical angles of attack, including stall. These benefits translate to airplane performance improvements, such as longer range or larger payload. Various flow control techniques were explored using Computational Fluid Dynamics and the aerodynamic performance enhancements were benchmarked against the baseline configuration. The computational analyses are used to quantify aerodynamic benefits, as well as the input required for actuation. The results were used in a system integration study for identifying potential practical implementations, which are described in a companion paper. Combined with the integration analysis, the objective of this project is to identify the most promising flow control candidates that potentially provide material net airplane level enhancements using onboard fluidic sources. The current study indicates that up to 5% net improvement in L/D at takeoff is potentially achievable using active flow control on the aileron, after accounting for factors of system integration.

CFD↗

Automated Instrumentation, Monitoring and Visualization of PVM Programs Using AIMS

We present views and analysis of the execution of several PVM codes for Computational Fluid Dynamics on a network of Sparcstations, including (a) NAS Parallel benchmarks CG and MG (White, Alund and Sunderam 1993); (b) a multi-partitioning algorithm for NAS Parallel Benchmark SP (Wijngaart 1993); and (c) an overset grid flowsolver (Smith 1993). These views and analysis were obtained using our Automated Instrumentation and Monitoring System (AIMS) version 3.0, a toolkit for debugging the performance of PVM programs. We will describe the architecture, operation and application of AIMS. The AIMS toolkit contains (a) Xinstrument, which can automatically instrument various computational and communication constructs in message-passing parallel programs; (b) Monitor, a library of run-time trace-collection routines; (c) VK (Visual Kernel), an execution-animation tool with source-code clickback; and (d) Tally, a tool for statistical analysis of execution profiles. Currently, Xinstrument can handle C and Fortran77 programs using PVM 3.2.x; Monitor has been implemented and tested on Sun 4 systems running SunOS 4.1.2; and VK uses X11R5 and Motif 1.2. Data and views obtained using AIMS clearly illustrate several characteristic features of executing parallel programs on networked workstations: (a) the impact of long message latencies; (b) the impact of multiprogramming overheads and associated load imbalance; (c) cache and virtual-memory effects; and (4significant skews between workstation clocks. Interestingly, AIMS can compensate for constant skew (zero drift) by calibrating the skew between a parent and its spawned children. In addition, AIMS' skew-compensation algorithm can adjust timestamps in a way that eliminates physically impossible communications (e.g., messages going backwards in time). Our current efforts are directed toward creating new views to explain the observed performance of PVM programs. Some of the features planned for the near future include: (a) ConfigView, showing the physical topology of the virtual machine, inferred using specially formatted IP (Internet Protocol) packets; and (b) LoadView, synchronous animation of PVM-program execution and resource-utilization patterns.

Mehra, Pankaj↗

Supercomputing '91; Proceedings of the 4th Annual Conference on High Performance Computing, Albuquerque, NM, Nov. 18-22, 1991

Various papers on supercomputing are presented. The general topics addressed include: program analysis/data dependence, memory access, distributed memory code generation, numerical algorithms, supercomputer benchmarks, latency tolerance, parallel programming, applications, processor design, networks, performance tools, mapping and scheduling, characterization affecting performance, parallelism packaging, computing climate change, combinatorial algorithms, hardware and software performance issues, system issues. (No individual items are abstracted in this volume)

Source record↗

Spaceborne autonomous multiprocessor systems

The goal of this task is to provide technology for the specification and integration of advanced processors into the Space Station Freedom data management system environment through computer performance measurement tools, simulators, and an extended testbed facility. The approach focuses on five categories: (1) user requirements--determine the suitability of existing computer technologies and systems for real-time requirements of NASA missions; (2) system performance analysis--characterize the effects of languages, architectures, and commercially available hardware on real-time benchmarks; (3) system architecture--expand NASA's capability to solve problems with integrated numeric and symbolic requirements using advanced multiprocessor architectures; (4) parallel Ada technology--extend Ada software technology to utilize parallel architectures more efficiently; and (5) testbed--extend in-house testbed to support system performance and system analysis studies.

Fernquist, Alan↗

Experiences Using OpenMP Based on Compiler Directed Software DSM on a PC Cluster

In this work we report on our experiences running OpenMP (message passing) programs on a commodity cluster of PCs (personal computers) running a software distributed shared memory (DSM) system. We describe our test environment and report on the performance of a subset of the NAS (NASA Advanced Supercomputing) Parallel Benchmarks that have been automatically parallelized for OpenMP. We compare the performance of the OpenMP implementations with that of their message passing counterparts and discuss performance differences.

Hess, Matthias↗

Parallel 3D Mortar Element Method for Adaptive Nonconforming Meshes

High order methods are frequently used in computational simulation for their high accuracy. An efficient way to avoid unnecessary computation in smooth regions of the solution is to use adaptive meshes which employ fine grids only in areas where they are needed. Nonconforming spectral elements allow the grid to be flexibly adjusted to satisfy the computational accuracy requirements. The method is suitable for computational simulations of unsteady problems with very disparate length scales or unsteady moving features, such as heat transfer, fluid dynamics or flame combustion. In this work, we select the Mark Element Method (MEM) to handle the non-conforming interfaces between elements. A new technique is introduced to efficiently implement MEM in 3-D nonconforming meshes. By introducing an "intermediate mortar", the proposed method decomposes the projection between 3-D elements and mortars into two steps. In each step, projection matrices derived in 2-D are used. The two-step method avoids explicitly forming/deriving large projection matrices for 3-D meshes, and also helps to simplify the implementation. This new technique can be used for both h- and p-type adaptation. This method is applied to an unsteady 3-D moving heat source problem. With our new MEM implementation, mesh adaptation is able to efficiently refine the grid near the heat source and coarsen the grid once the heat source passes. The savings in computational work resulting from the dynamic mesh adaptation is demonstrated by the reduction of the the number of elements used and CPU time spent. MEM and mesh adaptation, respectively, bring irregularity and dynamics to the computer memory access pattern. Hence, they provide a good way to gauge the performance of computer systems when running scientific applications whose memory access patterns are irregular and unpredictable. We select a 3-D moving heat source problem as the Unstructured Adaptive (UA) grid benchmark, a new component of the NAS Parallel Benchmarks (NPB). In this paper, we present some interesting performance results of ow OpenMP parallel implementation on different architectures such as the SGI Origin2000, SGI Altix, and Cray MTA-2.

Feng, Huiyu↗

Predictions of Slat Noise from the 30P30N at High Angles of Attack Using Zonal Hybrid RANS-LES

Aeroacoustic predictions of slat noise from the 30P30N three-element high-lift system at high angles of attack are presented using a zonal hybrid RANS-LES method. The simulations are part of the 5th AIAA Benchmark problems for Airframe Noise Computations (BANC-V) Workshop. An economical approach utilizing structured overset grids with spatially varying span-wise grid resolution and a high-order accurate finite difference method is described. The method is utilized for near-field predictions at three angles of attack: α = 5.5, 9.5, and 14.0 degrees. Far-field noise is obtained by propagating the near-field solution using a permeable surface Ffowcs Williams-Hawkings (FWH) method. Good agreement is obtained with both near-field and far-field Power Spectral Density (PSD) data from an experimental study of the 30P30N in the 2m x 2m Kevlar-wall wind tunnel at the Japan Aerospace Exploration Agency (JAXA). Specifically, the reduction in narrow band peaks and overall broadband noise levels with increasing angle of attack is captured well using the zonal hybrid RANS-LES method.

Housman, Jeffrey A.↗

Testing New Programming Paradigms with NAS Parallel Benchmarks

Over the past decade, high performance computing has evolved rapidly, not only in hardware architectures but also with increasing complexity of real applications. Technologies have been developing to aim at scaling up to thousands of processors on both distributed and shared memory systems. Development of parallel programs on these computers is always a challenging task. Today, writing parallel programs with message passing (e.g. MPI) is the most popular way of achieving scalability and high performance. However, writing message passing programs is difficult and error prone. Recent years new effort has been made in defining new parallel programming paradigms. The best examples are: HPF (based on data parallelism) and OpenMP (based on shared memory parallelism). Both provide simple and clear extensions to sequential programs, thus greatly simplify the tedious tasks encountered in writing message passing programs. HPF is independent of memory hierarchy, however, due to the immaturity of compiler technology its performance is still questionable. Although use of parallel compiler directives is not new, OpenMP offers a portable solution in the shared-memory domain. Another important development involves the tremendous progress in the internet and its associated technology. Although still in its infancy, Java promisses portability in a heterogeneous environment and offers possibility to "compile once and run anywhere." In light of testing these new technologies, we implemented new parallel versions of the NAS Parallel Benchmarks (NPBs) with HPF and OpenMP directives, and extended the work with Java and Java-threads. The purpose of this study is to examine the effectiveness of alternative programming paradigms. NPBs consist of five kernels and three simulated applications that mimic the computation and data movement of large scale computational fluid dynamics (CFD) applications. We started with the serial version included in NPB2.3. Optimization of memory and cache usage was applied to several benchmarks, noticeably BT and SP, resulting in better sequential performance. In order to overcome the lack of an HPF performance model and guide the development of the HPF codes, we employed an empirical performance model for several primitives found in the benchmarks. We encountered a few limitations of HPF, such as lack of supporting the "REDISTRIBUTION" directive and no easy way to handle irregular computation. The parallelization with OpenMP directives was done at the outer-most loop level to achieve the largest granularity. The performance of six HPF and OpenMP benchmarks is compared with their MPI counterparts for the Class-A problem size in the figure in next page. These results were obtained on an SGI Origin2000 (195MHz) with MIPSpro-f77 compiler 7.2.1 for OpenMP and MPI codes and PGI pghpf-2.4.3 compiler with MPI interface for HPF programs.

Jin, H.↗

Identification of Computational and Experimental Reduced-Order Models

The identification of computational and experimental reduced-order models (ROMs) for the analysis of unsteady aerodynamic responses and for efficient aeroelastic analyses is presented. For the identification of a computational aeroelastic ROM, the CFL3Dv6.0 computational fluid dynamics (CFD) code is used. Flutter results for the AGARD 445.6 Wing and for a Rigid Semispan Model (RSM) computed using CFL3Dv6.0 are presented, including discussion of associated computational costs. Modal impulse responses of the unsteady aerodynamic system are computed using the CFL3Dv6.0 code and transformed into state-space form. The unsteady aerodynamic state-space ROM is then combined with a state-space model of the structure to create an aeroelastic simulation using the MATLAB/SIMULINK environment. The MATLAB/SIMULINK ROM is then used to rapidly compute aeroelastic transients, including flutter. The ROM shows excellent agreement with the aeroelastic analyses computed using the CFL3Dv6.0 code directly. For the identification of experimental unsteady pressure ROMs, results are presented for two configurations: the RSM and a Benchmark Supercritical Wing (BSCW). Both models were used to acquire unsteady pressure data due to pitching oscillations on the Oscillating Turntable (OTT) system at the Transonic Dynamics Tunnel (TDT). A deconvolution scheme involving a step input in pitch and the resultant step response in pressure, for several pressure transducers, is used to identify the unsteady pressure impulse responses. The identified impulse responses are then used to predict the pressure responses due to pitching oscillations at several frequencies. Comparisons with the experimental data are then presented.

Silva, Walter A.↗

Industrial Hygiene Issues

This breakout session is a traditional conference instrument used by the NASA industrial hygiene personnel as a method to convene personnel across the Agency with common interests. This particular session focused on two key topics, training systems and automation of industrial hygiene data. During the FY 98 NASA Occupational Health Benchmarking study, the training system under development by the U.S. Environmental Protection Agency (EPA) was deemed to represent a "best business practice." The EPA has invested extensively in the development of computer based training covering a broad range of safety, health and environmental topics. Currently, five compact disks have been developed covering the topics listed: Safety, Health and Environmental Management Training for Field Inspection Activities; EPA Basic Radiation Training Safety Course; The OSHA 600 Collateral Duty Safety and Health Course; and Key program topics in environmental compliance, health and safety. Mr. Chris Johnson presented an overview of the EPA compact disk-based training system and answered questions on its deployment and use across the EPA. This training system has also recently been broadly distributed across other Federal Agencies. The EPA training system is considered "public domain" and, as such, is available to NASA at no cost in its current form. Copies of the five CD set of training programs were distributed to each NASA Center represented in the breakout session. Mr. Brisbin requested that each NASA Center review the training materials and determine whether there is interest in using the materials as it is or requesting that EPA tailor the training modules to suit NASA's training program needs. The Safety, Health and Medical Services organization at Ames Research Center has completed automation of several key program areas. Mr. Patrick Hogan, Safety Program Manager for Ames Research Center, presented a demonstration of the automated systems, which are described by the following: (1) Safety, Health and Environmental Training. This system includes an assessment of training needs for every NASA Center organization, course descriptions, schedules and automated course scheduling, and presentation of training program metrics; (2) Safety and Health Inspection Information. This system documents the findings from each facility inspection, tracks abatement status on those findings and presents metrics on each department for senior management review; (3) Safety Performance Evaluation Profile. The survey system used by NASA to evaluate employee and supervisory perceptions of safety programs is automated in this system; and (4) Documentation Tracking System. Electronic archive and retrieval of all correspondence and technical reports generated by the Safety, Health and Medical Services Office are provided by this system.

Brisbin, Steven G.↗

Feeding Ten Billion People Is Possible Within Four Terrestrial Planetary Boundaries

Global agriculture puts heavy pressure on planetary boundaries, posing the challenge to achieve future food security without compromising Earth system resilience. On the basis of process-detailed, spatially explicit representation of four interlinked planetary boundaries (biosphere integrity, land-system change, freshwater use, nitrogen flows) and agricultural systems in an internally consistent model framework, we here show that almost half of current global food production depends on planetary boundary transgressions. Hotspot regions, mainly in Asia, even face simultaneous transgression of multiple underlying local boundaries. If these boundaries were strictly respected, the present food system could provide a balanced diet (2,355 kcal per capita per day) for 3.4 billion people only. However, as we also demonstrate, transformation towards more sustainable production and consumption patterns could support 10.2 billion people within the planetary boundaries analysed. Key prerequisites are spatially redistributed cropland, improved water–nutrient management, food waste reduction and dietary changes. Adoption of the Sustainable Development Goals by all nations in 2015 is the first ever commitment to a world development path that safeguards the stability of the Earth system as a prerequisite for meeting universal human standards1. The longstanding challenge of achieving food security through sustainable agriculture is particularly acute in this context as world agriculture is a leading cause for the current transgressions of multiple planetary boundaries (PBs) globally and regionally2–5. The PB framework is a comprehensive scientific attempt to synoptically define our planet’s biogeophysical limits to anthropogenic interference. It suggests bounds to nine interacting processes that together delineate a Holocene-like Earth system state. The Holocene is chosen as the reference state as it is the only period known to provide a safe operating space for a world population of several billion people, and according to a precautionary principle, the PBs are set in sufficient distance from processes that may critically undermine Earth system resilience and global sustainability. A challenging question, thus, is whether human development goals such as food security can be met while maintaining multiple PBs along with their subglobal manifestations. Further PB transgressions could jeopardize the chances of providing sufficient food for a world population projected to be wealthier and reach >9 billion by 2050. This conundrum portrays a tradeoff between Earth’s biophysical carrying capacity and humankind’s rising food demand, calling in response for radical rethinking of food production and consumption patterns6–9. Yield gap closures, avoidance of excessive input use, shifts towards less resource-demanding diets, food waste reductions and efficient international trade are crucial options for sustainably increasing the food supply10–15. For example, enhancing water-use efficiency on irrigated and rain-fed farms can triple or quadruple crop yields in low-performing systems, suggesting possible global gains of >20% (ref. 16). Even higher gains appear feasible through globally optimized configurations of the land-use pattern17, and cutting food losses by half could generate food for another billion people18. Thus, collective large-scale implementation of such options could sustain food for a further growing world population19. Yet achieving this within a safe operating space as defined by PBs requires not only a halt to but actually a reversal of existing PB transgressions. Previous studies suggest that such a reconciliation might be possible, but these were based on aggregate representations of PBs (not accounting for the spatial patterns of limits, transgressions and interactions) or considered only one boundary in isolation17,20–23. Here, we systematically quantify to what extent current food production depends on local to global transgressions of the PBs for biosphere integrity, land-system change, freshwater use and nitrogen (N) flows, along with the potential of a range of solutions to avoid these transgressions and still increase food supply (Table 1). To this end, we configured an internally consistent process-based model of the terrestrial biosphere including agriculture (LPJmL) with multiple spatially distributed PBs and their interactions. LPJmL is among the longest-established and best-evaluated biosphere models, showing robust performance regarding simulation of, for example, carbon, water and crop yield dynamics (Supplementary Figs. 1 and 2 and Supplementary Table 1; see ref. 24 for a comprehensive benchmarking and Supplementary Methods for more detail on model evaluations). In principle following established definitions4, we refine the computation of some PBs with respect to their regional patterns and interactions (Methods), providing globally gridded precautionary limits to human interference with the Earth system at a level of great detail. In particular, we account for the evidence that many PBs need to be represented spatially explicitly4 to cover their

Gerten, Dieter↗

CFD Analysis for Assessing the Effect of Wind on the Thermal Control of the Mars Science Laboratory Curiosity Rover

The challenging range of landing sites for which the Mars Science Laboratory Rover was designed, requires a rover thermal management system that is capable of keeping temperatures controlled across a wide variety of environmental conditions. On the Martian surface where temperatures can be as cold as -123 C and as warm as 38 C, the rover relies upon a Mechanically Pumped Fluid Loop (MPFL) Rover Heat Rejection System (RHRS) and external radiators to maintain the temperature of sensitive electronics and science instruments within a -40 C to 50 C range. The RHRS harnesses some of the waste heat generated from the rover power source, known as the Multi Mission Radioisotope Thermoelectric Generator (MMRTG), for use as survival heat for the rover during cold conditions. The MMRTG produces 110 W of electrical power while generating waste heat equivalent to approximately 2000 W. Heat exchanger plates (hot plates) positioned close to the MMRTG pick up this survival heat from it by radiative heat transfer. Winds on Mars can be as fast as 15 m/s for extended periods. They can lead to significant heat loss from the MMRTG and the hot plates due to convective heat pick up from these surfaces. Estimation of this convective heat loss cannot be accurately and adequately achieved by simple textbook based calculations because of the very complicated flow fields around these surfaces, which are a function of wind direction and speed. Accurate calculations necessitated the employment of sophisticated Computational Fluid Dynamics (CFD) computer codes. This paper describes the methodology and results of these CFD calculations. Additionally, these results are compared to simple textbook based calculations that served as benchmarks and sanity checks for them. And finally, the overall RHRS system performance predictions will be shared to show how these results affected the overall rover thermal performance.

wind↗

SpaceCubeX: A Framework for Evaluating Hybrid Multi-Core CPU FPGA DSP Architectures

The SpaceCubeX project is motivated by the need for high performance, modular, and scalable on-board processing to help scientists answer critical 21st century questions about global climate change, air quality, ocean health, and ecosystem dynamics, while adding new capabilities such as low-latency data products for extreme event warnings. These goals translate into on-board processing throughput requirements that are on the order of 100-1,000 more than those of previous Earth Science missions for standard processing, compression, storage, and downlink operations. To study possible future architectures to achieve these performance requirements, the SpaceCubeX project provides an evolvable testbed and framework that enables a focused design space exploration of candidate hybrid CPU/FPGA/DSP processing architectures. The framework includes ArchGen, an architecture generator tool populated with candidate architecture components, performance models, and IP cores, that allows an end user to specify the type, number, and connectivity of a hybrid architecture. The framework requires minimal extensions to integrate new processors, such as the anticipated High Performance Spaceflight Computer (HPSC), reducing time to initiate benchmarking by months. To evaluate the framework, we leverage a wide suite of high performance embedded computing benchmarks and Earth science scenarios to ensure robust architecture characterization. We report on our projects Year 1 efforts and demonstrate the capabilities across four simulation testbed models, a baseline SpaceCube 2.0 system, a dual ARM A9 processor system, a hybrid quad ARM A53 and FPGA system, and a hybrid quad ARM A53 and DSP system.

Hybrid Flight Architectures↗

Comparison of secondary flows predicted by a viscous code and an inviscid code with experimental data for a turning duct

A comparison of the secondary flows computed by the viscous Kreskovsky-Briley-McDonald code and the inviscid Denton code with benchmark experimental data for turning duct is presented. The viscous code is a fully parabolized space-marching Navier-Stokes solver while the inviscid code is a time-marching Euler solver. The experimental data were collected by Taylor, Whitelaw, and Yianneskis with a laser Doppler velocimeter system in a 90 deg turning duct of square cross-section. The agreement between the viscous and inviscid computations was generally very good for the streamwise primary velocity and the radial secondary velocity, except at the walls, where slip conditions were specified for the inviscid code. The agreement between both the computations and the experimental data was not as close, especially at the 60.0 deg and 77.5 deg angular positions within the duct. This disagreement was attributed to incomplete modelling of the vortex development near the suction surface.

Schwab, J. R.↗