Flow-Based Algorithms for Improving Clusters: A Unifying Framework, Software, and Performance
Not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not provided.
Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.
The circuit consists of a digital to analog converter, accumulator, read write memory and UV erasable read only memory. The circuit can convert an analog signal to a digital representation, perform mathematical operations on the digital signal and subsequently convert the digital signal to an analog output. Development software tailored for programming the 2920 is presented.
The development of models to represent human characteristics and behaviors in human factors is broad and general. The term "model" can refer to any metaphor to represent any aspect of the human; it is generally used in research to mean a mathematical tool for the simulation (often in software, which makes the simulation digital) of some aspect of human performance and for the prediction of future outcomes. This section is restricted to the application of human models in physical design, e.g., in human factors engineering. This design effort is typically human interface design, and the digital models used are anthropometric. That is, they are visual models that are the physical shape of humans and that have the capabilities and constraints of humans of a selected population. They are distinct from the avatars used in the entertainment industry (movies, video games, and the like) in precisely that regard: as models, they are created through the application of data on humans, and they are used to predict human response; body stresses workspaces. DHM enable iterative evaluation of a large number of concepts and support rapid analysis, as compared with use of physical mockups. They can be used to evaluate feasibility of escape of a suited astronaut from a damaged vehicle, before launch or after an abort (England, et al., 2012). Throughout most of human spaceflight, little attention has been paid to worksite design for ground workers. As a result of repeated damage to the Space Shuttle which adversely affected flight safety, DHM analyses of ground assembly and maintenance have been developed over the last five years for the design of new flight systems (Stambolian, 2012, Dischinger and Dunn Jackson, 2014). The intent of these analyses is to assure the design supports the work of the ground crew personnel and thereby protect the launch vehicle. They help the analyst address basic human factors engineering questions: can a worker reach the task site from the work platform provided; can she or he see the task site; can she or he control tools, which, if dropped, might damage the system? Figure 7.3.1 provides an example of such analysis for a future NASA launch vehicle. [figure 7.3.1 here] In-space systems for operation by astronauts have long been targets for DHM analysis, given the focus on mission success and concerns for astronaut safety. Figure 7.3.2 illustrates the analysis of the design to support astronaut tasks for an International Space Station glovebox. [Figure 7.3.2 here] Use by
The NASA/JPL goal to reduce payload in future space missions while increasing mission capability demands miniaturization of active and passive sensors, analytical instruments and communication systems among others. Currently, typical system requirements include the detection of particular spectral lines, associated data processing, and communication of the acquired data to other systems. Advances in lithography and deposition methods result in more advanced devices for space application, while the sub-micron resolution currently available opens a vast design space. Though an experimental exploration of this widening design space-searching for optimized performance by repeated fabrication efforts-is unfeasible, it does motivate the development of reliable software design tools. These tools necessitate models based on fundamental physics and mathematics of the device to accurately model effects such as diffraction and scattering in opto-electronic devices, or bandstructure and scattering in heterostructure devices. The software tools must have convenient turn-around times and interfaces that allow effective usage. The first issue is addressed by the application of high-performance computers and the second by the development of graphical user interfaces driven by properly developed data structures. These tools can then be integrated into an optimization environment, and with the available memory capacity and computational speed of high performance parallel platforms, simulation of optimized components can proceed. In this paper, specific applications of the electromagnetic modeling of infrared filtering, as well as heterostructure device design will be presented using genetic algorithm global optimization methods.
Using simulations for scientific discovery requires that the software used in the simulations undergoes a rigorous design and development process similar to that of the lab instruments in the experimental sciences. To devise a good design methodology, it is critical to understand the requirements, constraints, and challenges. Furthermore, this article describes insights from the long-term stewardship of a multiphysics multicomponent software, FLASH, that was designed more than 20 years ago for astrophysics, now serves multiple communities, and has been successful in adapting to the changing world of high-performance computing.
Nonlinear programming, preliminary design problems, performance simulation problems trajectory optimization, flight computer optimization, and linear least squares problems are among the topics covered. The nonlinear programming applications encountered in a large aerospace company are a real challenge to those who provide mathematical software libraries and consultation services. Typical applications include preliminary design studies, data fitting and filtering, jet engine simulations, control system analysis, and trajectory optimization and optimal control. Problem sizes range from single-variable unconstrained minimization to constrained problems with highly nonlinear functions and hundreds of variables. Most of the applications can be posed as nonlinearly constrained minimization problems. Highly complex optimization problems with many variables were formulated in the early days of computing. At the time, many problems had to be reformulated or bypassed entirely, and solution methods often relied on problem-specific strategies. Problems with more than ten variables usually went unsolved.
The Exascale Computing Project (ECP) Software Technology and Co-Design teams addressed the growing complexities in high-performance computing (HPC) by developing scalable software libraries and tools that leverage exascale system capabilities. As we enter the exascale era, the need for reusable, optimized software solutions that can handle the unique challenges posed by these systems becomes increasingly important. The primary challenges the ECP teams faced were to create software libraries and tools that are performant on exascale architectures and portable and usable across diverse hardware platforms. Efforts addressed issues related to concurrent execution, memory management, and the integration of heterogeneous computing resources, such as GPUs from multiple vendors. The ECP’s strategy involved a structured development process encompassing the creation, optimization, and deployment of software in collaboration with industry, academia, and national laboratories. The project was organized into several technical areas: co-design of domain-specific suites with target applications, programming models and runtimes, development tools, mathematical libraries, data and visualization tools, and software ecosystem and delivery mechanisms. ECP has successfully developed a large portfolio of software libraries and tools that demonstrate significant improvements in performance and scalability on exascale systems. These products have been integrated into the Department of Energy’s computing facilities, supporting various scientific applications and ensuring robust performance across different hardware setups. ECP advancements in software development for exascale computing highlight the importance of a collaborative and adaptive approach to handling next-generation HPC systems complexities. The lessons learned emphasize the need for continuous engagement with end-users and vendors, and the importance of maintaining a balance between innovation and practical implementation. Future efforts will focus on ensuring scalability, keeping pace with rapid hardware advancements, and further enhancing the interoperability and usability of the software ecosystem. In conclusion, subsequent articles in this special issue provide in-depth discussions and case studies into specific library and tool efforts.
Optimizing the performance of large-scale parallel codes is critical for efficient utilization of computing resources. Code developers often explore various execution parameters, such as hardware configurations, system software choices, and application parameters, and are interested in detecting and understanding bottlenecks in different executions. They often collect hierarchical performance profiles represented as call graphs, which combine performance metrics with their execution contexts. The crucial task of exploring multiple call graphs together is tedious and challenging because of the many structural differences in the execution contexts and significant variability in the collected performance metrics (e.g., execution runtime). In this paper, we present Ensemble CallFlow to support the exploration of ensembles of call graphs using new types of visualizations, analysis, graph operations, and features. We introduce ensemble-Sankey , a new visual design that combines the strengths of resource-flow (Sankey) and box-plot visualization techniques. Whereas the resource-flow visualization can easily and intuitively describe the graphical nature of the call graph, the box plots overlaid on the nodes of Sankey convey the performance variability within the ensemble. Our interactive visual interface provides linked views to help explore ensembles of call graphs, e.g., by facilitating the analysis of structural differences, and identifying similar or distinct call graphs. Finally, we demonstrate the effectiveness and usefulness of our design through case studies on large-scale parallel codes.
This software provides tools for analyzing and modeling physical systems using mathematical transforms and layered models. It includes (1) functions for performing the Abel transform, which is used to relate measurements of bending angles to properties such as refractive index and radius in a medium. The code can compute bending angles from input profiles and also reconstruct these profiles from observed data; (2) the functions for modeling systems with multiple layers using an onion-peeling approach, allowing users to simulate and analyze the behavior of layered materials or structures. These capabilities are useful for researchers and engineers working in fields such as optics, atmospheric science, and materials analysis, enabling them to interpret and model data from experiments or simulations relates to refraction in spherical symmetric medium.
Geographic data, such as seismic event locations, station locations, etc., are generally given in geographic latitude Φ ’, longitude θ , and depth below sea level, ζ , using the WGS84 ellipsoid as a reference. In software systems that use this type of geographic data, it is necessary to manipulate the data mathematically in order to perform such tasks as finding the angular distance or azimuth from one point to another, to find an array of points along a great circle, to rotate a point about a pole of rotation, to move a point some angular distance in a specified direction, to find the intersections of two great circles or to find the intersections of a great circle and a small circle. In this paper, equations are presented that convert geographic locations first to geocentric coordinates and then to Earth-centered Cartesian coordinates where many mathematical manipulations can be performed conveniently and efficiently.
The MACSYMA program for symbolic and algebraic manipulation enables exact, symbolic mathematical computations to be performed on a computer. This program is rather large, and various approaches to the hardware and software problems are examined.
SCRAM computer program determines, according to one-dimensional mathematical model, cycle performance for airframe-integrated subsonic or supersonic combustion Ramjets. Reliable, efficient, and speedy ramjet designer's software tool usable on all standard computers down to those compatible with IBM PC-AT's. Scramjet cycle analysis code simulates, from nose to tail, hydrogen-fueled, airframe-integrated scramjet in real gas flow with equilibrium thermodynamic properties. Enables rapid scramjet performance estimates. Written in FORTRAN 77.
Here, our exploratory work finds that the SambaNova Reconfigurable Dataflow Architecture (RDA) along with the SambaFlow software stack provides for an attractive system and solution to accelerate AI for science workloads. We have observed the efficacy of using the system with a diverse set of science applications and reasoned their suitability for performance gains over traditional hardware. As the Data-Scale system provides for a very large memory capacity, the system can be used to train models that typically do not fit in a GPU. The architecture also provides for deeper integration with upcoming supercomputers at the Argonne Leadership Computing Facility (ALCF), a US Department of Energy Office of Science user facility, to help advance science insights.
Program verification in scientific computing encompasses the application of formal and mathematical techniques to a scientific computing code for its credibility, accuracy, and validity. Code Verification identifies bugs and performance issues in the software development stage. Solution Verification assesses the applicability of the code and the accuracy of the solution to problems of interest. Both activities utilize application cases and quantify the error against prescribed acceptance criteria. However, simply executing more application cases does not guarantee stronger or more comprehensive credibility. Here, we establish a verification framework that involves Code Verification and Solution Verification, both of which work together such that the overarching goal of “converge to the correct answer for the intended application” can be reasonably inferred. The application of such a verification framework is demonstrated using the pin-resolved neutron transport code MPACT, where standard unit tests and regression tests are covered, and where the Method of Exact Solutions and the Method of Manufactured Solutions are successfully used. Additionally, the applicability of Method of Manufactured Solutions is extended to the OECD/NEA C5G7 benchmark problems of practical material and geometric configurations. Solution Verification activities are demonstrated on a practical hierarchy of application models of increasing complexity ranging from 2D pin cell problems to 3D assembly problems. The convergence behavior and rate of convergence with respect to each individual variable are studied and provided. This framework can be adapted broadly to other fields involving scientific computing codes.
Scientific processes rely on software as an important tool for data acquisition, analysis, and discovery. Over the years, sustainable software development practices have made progress in being considered as an integral component of research. However, management of computation-based scientific studies is often left to individual researchers who design their computational experiments based on personal preferences and the nature of the study. Here, we believe that the quality, efficiency, and reproducibility of computation-based scientific research can be improved by explicitly creating an execution environment that allows researchers to provide a clear record of traceability. This is particularly relevant to complex computational studies in high-performance computing (HPC) environments. In this article, we review the documentation required to maintain a comprehensive record of HPC computational experiments for reproducibility. We also provide an overview of tools and practices that we have developed to perform such studies around Flash-X, a multiphysics scientific software.
Measurements were performed to determine the mass and enrichment of 10 uranium oxide (U3O8) card sources. The measurements and analysis were completed as a verification of the card sources in support of Oak Ridge National Laboratory’s nuclear material control and accountability program. Although these cards are not nationally accredited as a nuclear standard, they are used as a working reference for measurements, such as holdup. The uranium card source measurements were taken with a broad energy germanium detector and the Genie 2000 Gamma Acquisition and Analysis software. A complete characterization of each of the 10 uranium card sources was performed using the 4 characteristic full energy peaks of 235U. Using the In Situ Object Counting System software to determine the mathematical efficiency of the measurement, the mass of 235U in each card was determined. The 235U mass in each card ranged from 10.42 to 12.63 g with a systematic error between 0.76 and 0.95 g and a random error of 0.01 g for each card source. The Multi-Group Analysis for Uranium (MGAU) software and the Fixed-Energy, Response Function Analysis with Multiple Efficiency (FRAM) isotopic analysis software were used to determine the isotopic composition of the uranium cards. The measured enrichment was compared to the declared enrichment for each card, with uncertainties ranging from 2.7% to 4.3% for the MGAU analysis and 2.3% to 4.1% for the FRAM analysis. This is a good example of how a well-benchmarked mathematical calibration method can be useful in characterizing uranium sources.
Hardware architectures become increasingly complex as the compute capabilities grow to exascale. Here, we present the Analytical Memory Model with Pipelines (AMMP) of the Performance Prediction Toolkit (PPT). PPT-AMMP takes high-level source code and hardware architecture parameters as input and predicts runtime of that code on the target hardware platform, which is defined in the input parameters. PPT-AMMP transforms the code to an (architecture-independent) intermediate representation, then (i) analyzes the basic block structure of the code, (ii) processes architecture-independent virtual memory access patterns that it uses to build memory reuse distance distribution models for each basic block, and (iii) runs detailed basic-block level simulations to determine hardware pipeline usage. PPT-AMMP uses machine learning and regression techniques to build the prediction models based on small instances of the input code, then integrates into a higher-order discrete-event simulation model of PPT running on Simian PDES engine. We validate PPT-AMMP on four standard computational physics benchmarks and present a use case of hardware parameter sensitivity analysis to identify bottleneck hardware resources on different code inputs. We further extend PPT-AMMP to predict the performance of a scientific application code, namely, the radiation transport mini-app SNAP. To this end, we analyze multi-variate regression models that accurately predict the reuse profiles and the basic block counts. We validate predicted SNAP runtimes against actual measured times.