Engineering PapersSearch

Engineering topics

Thigpen, William

Publications and source records attributed to Thigpen, William.

The Evolution of NASA’s High-End Computing Capabilities

For over 35 years, the NASA Advanced Supercomputing (NAS) Division at Ames Research Center has housed and managed the U.S. space agency’s largest supercomputing assets. Focused on high-end computing technologies, efficient operations, and user success, the NAS Division has worked with industry to deploy a series of highly successful systems that enable scientific and engineering achievements across NASA. The complementary role of the High-End Computing Capability (HECC) project is evolving to meet NASA’s future challenges in returning to the Moon as a pathway to Mars, while continuing exciting research in aeronautics, space exploration, and Earth science.

Thigpen, William

Electra: A Modular-Based Expansion of NASA's Supercomputing Capability

NASA has increasingly relied on high-performance computing (HPC) re- sources for computational modeling, simulation, and data analysis to meet the science and engineering goals of its missions in space exploration, aeronautics, and Earth and space science. The NASA Advanced Supercomputing (NAS) Division at Ames Research Center in Silicon Valley, Calif., hosts NASA’s premier supercomputing resources, integral to achieving and enhancing the success of the agency’s missions. NAS provides a balanced environment, funded under the High-End Computing Capability (HECC) project, comprised of world-class supercomputers, including its flagship distributed-memory cluster, Pleiades; high-speed networking; and massive data storage facilities, along with multi-disciplinary support teams for user support, code porting and optimization, and large-scale data analysis and scientific visualization. However, as scientists have increased the fidelity of their simulations and engineers are conducting larger parameter-space studies, the requirements for supercomputing resources have been growing by leaps and bounds. With the facility housing the HECC systems reaching its power and cooling capacity, NAS undertook a prototype project to investigate an alternative approach for housing supercomputers. Modular supercomputing, or container-based computing, is an innovative concept for expanding NASA’s HPC capabilities. With modular supercomputing, additional containers—similar to portable storage pods—can be connected together as needed to accommodate the agency’s ever-increasing demand for computing resources. In addition, taking advantage of the local weather permits the use of cooling technologies that would additionally save energy and reduce annual water usage. The first stage of NASA’s Modular Supercomputing Facility (MSF) prototype, which resulted in a 1,000 square-foot module on a concrete pad with room for 16 compute racks, was completed in Fall 2016 and an SGI (now HPE) computer system, named Electra, was deployed there in early 2017. Cooling is performed via an evaporative system built into the module, and preliminary experience shows a Power Usage Effectiveness (PUE) measurement of 1.03. Electra achieved over a petaflop on the LINPACK benchmark, sufficient to rank number 96 on the November 2016 TOP500 list [14]. The system consists of 1,152 InfiniBand-connected Intel Xeon Broadwell-based nodes. Its users access their files on a facility-wide file system shared by all HECC compute assets via Mellanox MetroX InfiniBand extenders, which connect the Electra fabric to Lustre routers in the primary facility over fiber-optic links about 900 feet long. The MSF prototype has exceeded expectations and is serving as a blueprint for future expansions. In the remainder of this chapter, we detail how modular data center technology can be used to expand an existing compute resource. We begin by describing NASA’s requirements for supercomputing and how resources were provided prior to the integration of the Electra module-based system.

Biswas, Rupak

NASA Blazes a Different Path to Energy-Efficient Supercomputing

For years, NASA had a very straightforward process for replacing high-performance computing hardware: over a three-year period, when it became more expensive to operate an older suite of hardware than it did to replace it with new products that could accomplish the same work, we simply replaced the old hardware. For NASA’s High-End Computing Capability (HECC) Project, that process changed when we reached the limits of our facility’s power, cooling, and floor-loading capacity, becoming a strategy of decommissioning the least productive hardware and replacing it with more capable counterparts. The impact was that we provided our users with less supercomputing capability than we would have without the limitations. Additionally, with 25% of our total power consumption going to cool our systems and 50,000 gallons of water per day being evaporated, we wanted a solution that would expand our compute facility while being sensitive to the impact on our environment.

Thigpen, William

Far-Infrared and Millimeter Continuum Studies of K-Giants: Alpha Boo and Alpha Tau

We have imaged two normal, non-coronal, infrared-bright K-giants, alpha Boo and alpha Tau, in the 1.4-millimeter and 2.8-millimeter continuum using BIMA. These stars have been used as important absolute calibrators for several infrared satellites. Our goals are: (1) to probe the structure of their upper photospheres; (2) to establish whether these stars radiate as simple photospheres or possess long-wavelength chromospheres; and (3) to make a connection between millimeter-wave and far-infrared absolute flux calibrations. To accomplish these goals we also present ISO Long Wavelength Spectrometer (LWS) measurements of both these K-giants. The far-infrared and millimeter continuum radiation is produced in the vicinity of the temperature minimum in a Boo and a Tau, offering a direct test of the model photospheres and chromospheres for these two cool giants. We find that current photospheric models predict fluxes in reasonable agreement with those observed for those wavelengths which sample the upper photosphere, namely less than or equal to 170 micrometers in alpha Tau and less than or equal to 125 micrometers in alpha Boo. It is possible that alpha Tau is still radiative as far as 0.9 - 1.4 millimeters. We detect chromospheric radiation from both stars by 2.8 millimeters (by 1.4 millimeters in alpha Boo), and are able to establish useful bounds on the location of the temperature minimum. An attempt to interpret the chromospheric fluxes using the two-component "bifurcation model" proposed by Wiedemann et al. (1994) appears to lead to a significant contradiction.

Cohen, Martin

Accounting and Accountability for Distributed and Grid Systems

While the advent of distributed and grid computing systems will open new opportunities for scientific exploration, the reality of such implementations could prove to be a system administrator's nightmare. A lot of effort is being spent on identifying and resolving the obvious problems of security, scheduling, authentication and authorization. Lurking in the background, though, are the largely unaddressed issues of accountability and usage accounting: (1) mapping resource usage to resource users; (2) defining usage economies or methods for resource exchange; (3) describing implementation standards that minimize and compartmentalize the tasks required for a site to participate in a grid.

Thigpen, William

The State of NASA's Information Power Grid

This viewgraph presentation transfers the concept of the power grid to information sharing in the NASA community. An information grid of this sort would be characterized as comprising tools, middleware, and services for the facilitation of interoperability, distribution of new technologies, human collaboration, and data management. While a grid would increase the ability of information sharing, it would not necessitate it. The onus of utilizing the grid would rest with the users.

Johnston, William E.

NASA's Information Power Grid: Large Scale Distributed Computing and Data Management

Large-scale science and engineering are done through the interaction of people, heterogeneous computing resources, information systems, and instruments, all of which are geographically and organizationally dispersed. The overall motivation for Grids is to facilitate the routine interactions of these resources in order to support large-scale science and engineering. Multi-disciplinary simulations provide a good example of a class of applications that are very likely to require aggregation of widely distributed computing, data, and intellectual resources. Such simulations - e.g. whole system aircraft simulation and whole system living cell simulation - require integrating applications and data that are developed by different teams of researchers frequently in different locations. The research team's are the only ones that have the expertise to maintain and improve the simulation code and/or the body of experimental data that drives the simulations. This results in an inherently distributed computing and data management environment.

Johnston, William E.

Distributed Accounting on the Grid

By the late 1990s, the Internet was adequately equipped to move vast amounts of data between HPC (High Performance Computing) systems, and efforts were initiated to link together the national infrastructure of high performance computational and data storage resources together into a general computational utility 'grid', analogous to the national electrical power grid infrastructure. The purpose of the Computational grid is to provide dependable, consistent, pervasive, and inexpensive access to computational resources for the computing community in the form of a computing utility. This paper presents a fully distributed view of Grid usage accounting and a methodology for allocating Grid computational resources for use on a Grid computing system.

Thigpen, William