Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Pathway Evolution Through a Bottlenecking-Debottlenecking Strategy and Machine Learning-Aided Flux Balancing

The evolution of pathway enzymes enhances the biosynthesis of high-value chemicals, crucial for pharmaceutical, and agrochemical applications. However, unpredictable evolutionary landscapes of pathway genes often hinder successful evolution. Here, the presence of complex epistasis is identifued within the representative naringenin biosynthetic pathway enzymes, hampering straightforward directed evolution. Subsequently, a biofoundry-assisted strategy is developed for pathway bottlenecking and debottlenecking, enabling the parallel evolution of all pathway enzymes along a predictable evolutionary trajectory in six weeks. This study then utilizes a machine learning model, ProEnsemble, to further balance the pathway by optimizing the transcription of individual genes. The broad applicability of this strategy is demonstrated by constructing an Escherichia coli chassis with evolved and balanced pathway genes, resulting in 3.65 g L -1 naringenin. The optimized naringenin chassis also demonstrates enhanced production of other flavonoids. This approach can be readily adapted for any given number of enzymes in the specific metabolic pathway, paving the way for automated chassis construction in contemporary biofoundries.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerating Parallel Applications in Cloud Platforms via Adaptive Time-Slice Control

Cloud platforms can provide flexible and cost-effective environments for parallel applications. However, the resource over-commitment issues, i.e., cloud providers often provide much more executable virtual CPUs than available physical CPUs, still impede the synchronization operations of parallel applications, causing severe performance degradation. Existing methods optimize parallel applications by promoting the priorities of involved VMs. They cannot fully explore the performance of parallel applications, because they ignore the time-slice requirements of different phases of parallel applications. Furthermore, non-parallel applications experience unsatisfied performance because of low scheduling priorities. Given empirical analysis on time-slices of virtual machines (VMs), we find that shortening time-slices can mitigate synchronization overhead which incurs during communication phases, while over-short time-slices cause frequent cache misses in computation phases. Accordingly, we propose an Adaptive Time-slice Control (ATC) mechanism. ATC first detects the phases of parallel applications based on lock latency or cache misses. Then, ATC shortens time-slices during communication phases and prolongs time-slices during computation phases for parallel applications, and sets a uniform time-slice for non-parallel applications. Finally, we evaluate ATC using seven well-known benchmarks with 25+ applications. Experiments show that ATC obtains 1.5-75x performance gain for running parallel applications than state-of-the-art solutions, with nearly unaffected impact on non-parallel applications.

97 MATHEMATICS AND COMPUTING↗

Segmentation of remotely sensed data using parallel region growing

The improved spatial resolution of the new earth resources satellites will increase the need for effective utilization of spatial information in machine processing of remotely sensed data. One promising technique is scene segmentation by region growing. Region growing can use spatial information in two ways: only spatially adjacent regions merge together, and merging criteria can be based on region-wide spatial features. A simple region growing approach is described in which the similarity criterion is based on region mean and variance (a simple spatial feature). An effective way to implement region growing for remote sensing is as an iterative parallel process on a large parallel processor. A straightforward parallel pixel-based implementation of the algorithm is explored and its efficiency is compared with sequential pixel-based, sequential region-based, and parallel region-based implementations. Experimental results from on aircraft scanner data set are presented, as is a discussioon of proposed improvements to the segmentation algorithm.

Tilton, J. C.↗

The Kokkos Ecosystem [Brief]

In 2016/2017, the field of High-Performance Computing (HPC) entered a new era driven by fundamental physics challenges to produce ever more energy and cost-efficient processors. Since the convergence on the Message-Passing Interface (MPI) standard in the mid-1990s, application developers enjoyed a seemingly static view of the underlying machine — that of a distributed collection of homogeneous nodes executing in collaboration. However, after almost two decades of dominance, the sole use of MPI to derive parallelism acted as a limiter to improved future performance. While MPI is widely expected to continue to function as the basic mechanism for communication between compute nodes for the immediate future, additional parallelism is required on the computing node itself if high performance and efficiency goals are to be realized. When reviewing the architectures of the top HPC systems today, the change in paradigm is clear: the compute nodes of the leading machines in the world are either powered by many-core chips with a few dozen cores each, or use heterogeneous designs, where traditional CPUs marshal work to massively parallel compute accelerators which has as many as 200,000 processing threads in flight simultaneously. Complicating matters further for application developers, each processor vendor has its own preferred way of writing code for their architecture.The Kokkos EcoSystem was released by Sandia in 2017 to address this new era in HPC system design by providing a vendor independent performance portable programming system for scientific, engineering, and mathematical software applications written in the C++ programming language. Using Kokkos, application developers can be more productive because they will not have to create and maintain separate versions of their software for each architecture, nor will they have to be experts in each architecture's peculiar requirements. Instead, they will have a single method of programming for the diverse set of modern HPC architectures. While Kokkos started in 2011 as a programming model only, it soon became clear that complex applications needed more. It is also critical to have a portable mathematical functions and developers need tools to debug their applications, gain insight into the performance characteristics of their codes and tune algorithm performance parameters through automated processes. The Kokkos EcoSystem addresses those needs through its three main components: the Kokkos Core programming model, the Kokkos Kernels math library, and the Kokkos Tools project.

97 MATHEMATICS AND COMPUTING↗

The 1987 RIACS annual report

The Research Institute for Advanced Computer Science (RIACS) was established at the NASA Ames Research Center in June of 1983. RIACS is privately operated by the Universities Space Research Association (USRA), a consortium of 64 universities with graduate programs in the aerospace sciences, under several Cooperative Agreements with NASA. RIACS's goal is to provide preeminent leadership in basic and applied computer science research as partners in support of NASA's goals and missions. In pursuit of this goal, RIACS contributes to several of the grand challenges in science and engineering facing NASA: flying an airplane inside a computer; determining the chemical properties of materials under hostile conditions in the atmospheres of earth and the planets; sending intelligent machines on unmanned space missions; creating a one-world network that makes all scientific resources, including those in space, accessible to all the world's scientists; providing intelligent computational support to all stages of the process of scientific investigation from problem formulation to results dissemination; and developing accurate global models for climatic behavior throughout the world. In working with these challenges, we seek novel architectures, and novel ways to use them, that exploit the potential of parallel and distributed computation and make possible new functions that are beyond the current reach of computing machines. The investigation includes pattern computers as well as the more familiar numeric and symbolic computers, and it includes networked systems of resources distributed around the world. We believe that successful computer science research is interdisciplinary: it is driven by (and drives) important problems in other disciplines. We believe that research should be guided by a clear long-term vision with planned milestones. And we believe that our environment must foster and exploit innovation. Our activities and accomplishments for the calendar year 1987 and our plans for 1988 are reported.

Source record↗

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor↗

Radiatively cooled magnetic reconnection experiments driven by pulsed power

We present evidence for strong radiative cooling in a pulsed-power-driven magnetic reconnection experiment. Two aluminum exploding wire arrays, driven by a 20 MA peak current, 300 ns rise time pulse from the Z machine (Sandia National Laboratories), generate strongly driven plasma flows (MA≈7) with anti-parallel magnetic fields, which form a reconnection layer (SL≈120) at the mid-plane. The net cooling rate far exceeds the Alfvénic transit rate (τcool−1/τA−1≫1), leading to strong cooling of the reconnection layer. We determine the advected magnetic field and flow velocity using inductive probes positioned in the inflow to the layer, and inflow ion density and temperature from analysis of visible emission spectroscopy. A sharp decrease in x-ray emission from the reconnection layer, measured using filtered diodes and time-gated x-ray imaging, provides evidence for strong cooling of the reconnection layer after its initial formation. X-ray images also show localized hotspots, regions of strong x-ray emission, with velocities comparable to the expected outflow velocity from the reconnection layer. These hotspots are consistent with plasmoids observed in 3D radiative resistive magnetohydrodynamic simulations of the experiment. X-ray spectroscopy further indicates that the hotspots have a temperature (170 eV) much higher than the bulk layer (≤75 eV) and inflow temperatures (about 2 eV) and that these hotspots generate the majority of the high-energy (>1 keV) emission.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Cloud Services Enable Efficient AI-Guided Simulation Workflows across Heterogeneous Resources

Applications which fuse machine learning and simulation are rarely best served by a single computing resource. Highly parallel simulation codes are best deployed on super- computers, while AI tasks used to decide which simulations to perform may be best suited to specialized accelerators. Here we present a Function-as-a-Service (FaaS) system for executing complex, distributed computational campaigns that achieves performance parity with conventional workflow systems without the complexities of secure network connections between compute providers. One innovation enabling high performance is a subsystem that directly moves task data between sites, separate from the cloud-hosted FaaS system used to distribute task instructions. We also introduce a flexible scheduling system that allows us access factor of 2 trade offs between the amount of resources required to solve a problem at each compute site. We anticipate that this system will upgrade multi-site applications from demonstration projects to routine practice in computational science.

Ward, Logan↗

All-Atom Simulation of 3D Hot Spot Formation in Shocked TATB Explosive

TATB is an insensitive high explosive (IHE) critical to the stockpile that is challenging to model at the continuum scale. Advanced detonation models in the Cheetah high explosive chemistry code require validation though subscale simulations. High explosive initiation is determined by micron-scale physics of hot spots formed a shock-collapsed pores. Pore sizes between 100 nm and 1 μm are believed to be the most important for determining the shock sensitivity of TATB. This range of pore sizes is difficult to access at the atomic scale through allatom molecular dynamics (MD) simulations, even with Sierra-class computers. Quasi-2D simulations are widely used and allow much larger pore sizes (up to 400 nm) to be studied, but the applicability of 2D simulations to the actual 3D pore response is not understood. Resolving these uncertainties through “full physics” MD modeling is key for generalizing, parameterizing, and validating the kinds of continuum models used to inform design, safety, and performance. This work was a continuation of FY20 efforts pushing simulations to full 3D with the largest-ever all-atom simulations of an explosive. These were the first all-atom full-3D simulations of large hot spots thought to govern explosive detonation and required over a billion atoms. Simulations were performed using LAMMPS, an open SNL science code. MD explosive models present unique challenges, even for established codes such as LAMMPS. Their model forms are more complex than typical models for metals, while simulating high temperature-pressure conditions is demanding and increases computational cost. Scaling problems in GPU-enabled MD algorithms initially limited simulations to <100 million atoms but were resolved through collaboration with SNL. An overall 24x speedup was obtained relative to CPU machines. Specialized analysis of these simulations required a bottom-up refactoring and algorithm parallelization of in-house codes and application of computer vision algorithms to extract meaningful information.

36 MATERIALS SCIENCE↗

Interstitial Collimating Holes for Gas-Levitation Microfurnace

Spaces between small rods direct gas flow. Wires for thin rods clamped in square array in precise square gooove. Spaces between wires are long, thin, parallel channels that direct flow of gas. Technique extended to such hard-to-machine refractory metals as tungsten and molybdenum.

Dunn, E. G.↗

Workpiece positioning vise

A pair of jaw assemblies simultaneously driven in opposed reciprocation by a single shaft has oppositely threaded sections to automatically center delicate or brittle workpieces such as lithium fluoride crystal beneath the blade of a crystal cleaving machine. Both jaw assemblies are suspended above the vise bed by a pair of parallel guide shafts attached to the vise bed. Linear rolling bearings, fitted around the guide shafts and firmly held by opposite ends of the jaw assemblies, provide rolling friction between the guide shafts and the jaw assemblies. A belleville washer at one end of the drive shaft and thrust bearings at both drive shaft ends hold the shaft in compression between the vise bed, thereby preventing wobble of the jaw assemblies due to wear between the shaft and vise bed.

Hallberg, F. C.↗

On evaluating parallel computer systems

A workshop was held in an attempt to program real problems on the MIT Static Data Flow Machine. Most of the architecture of the machine was specified but some parts were incomplete. The main purpose for the workshop was to explore principles for the evaluation of computer systems employing new architectures. Principles explored were: (1) evaluation must be an integral, ongoing part of a project to develop a computer of radically new architecture; (2) the evaluation should seek to measure the usability of the system as well as its performance; (3) users from the application domains must be an integral part of the evaluation process; and (4) evaluation results should be fed back into the design process. It is concluded that the general organizational principles are achievable in practice from this workshop.

Adams, George B., III↗

RISC Processors and High Performance Computing

In this tutorial, we will discuss top five current RISC microprocessors: The IBM Power2, which is used in the IBM RS6000/590 workstation and in the IBM SP2 parallel supercomputer, the DEC Alpha, which is in the DEC Alpha workstation and in the Cray T3D; the MIPS R8000, which is used in the SGI Power Challenge; the HP PA-RISC 7100, which is used in the HP 700 series workstations and in the Convex Exemplar; and the Cray proprietary processor, which is used in the new Cray J916. The architecture of these microprocessors will first be presented. The effective performance of these processors will then be compared, both by citing standard benchmarks and also in the context of implementing a real applications. In the process, different programming models such as data parallel (CM Fortran and HPF) and message passing (PVM and MPI) will be introduced and compared. The latest NAS Parallel Benchmark (NPB) absolute performance and performance per dollar figures will be presented. The next generation of the NP13 will also be described. The tutorial will conclude with a discussion of general trends in the field of high performance computing, including likely future developments in hardware and software technology, and the relative roles of vector supercomputers tightly coupled parallel computers, and clusters of workstations. This tutorial will provide a unique cross-machine comparison not available elsewhere.

Saini, Subhash↗

Load Balancing Strategies for Multi-Block Overset Grid Applications

The multi-block overset grid method is a powerful technique for high-fidelity computational fluid dynamics (CFD) simulations about complex aerospace configurations. The solution process uses a grid system that discretizes the problem domain by using separately generated but overlapping structured grids that periodically update and exchange boundary information through interpolation. For efficient high performance computations of large-scale realistic applications using this methodology, the individual grids must be properly partitioned among the parallel processors. Overall performance, therefore, largely depends on the quality of load balancing. In this paper, we present three different load balancing strategies far overset grids and analyze their effects on the parallel efficiency of a Navier-Stokes CFD application running on an SGI Origin2000 machine.

Djomehri, M. Jahed↗

Gilgamesh: A Multithreaded Processor-In-Memory Architecture for Petaflops Computing

Processor-in-Memory (PIM) architectures avoid the von Neumann bottleneck in conventional machines by integrating high-density DRAM and CMOS logic on the same chip. Parallel systems based on this new technology are expected to provide higher scalability, adaptability, robustness, fault tolerance and lower power consumption than current MPPs or commodity clusters. In this paper we describe the design of Gilgamesh, a PIM-based massively parallel architecture, and elements of its execution model. Gilgamesh extends existing PIM capabilities by incorporating advanced mechanisms for virtualizing tasks and data and providing adaptive resource management for load balancing and latency tolerance. The Gilgamesh execution model is based on macroservers, a middleware layer which supports object-based runtime management of data and threads allowing explicit and dynamic control of locality and load balancing. The paper concludes with a discussion of related research activities and an outlook to future work.

management locality load balance↗

TPSAS-NF1676L-32014-DND

The Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP), on-board the Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) is a satellite-borne polarization sensitive lidar. It has been providing the vertical distributions of clouds and aerosols along with their microphysical and optical properties since 2006. One of its important Level 2 products, feature classification, has been determined using the lidar information from 532 nm parallel and perpendicular channels, and 1064 nm channel measurements of layer integrated backscatter. Deep machine learning methods which combine both the channel and texture information to recognize feature patterns is uniquely beneficial when applied to this data. In this study, we will use Convolutional Neural Network (CNN), a deep machine learning method, to classify lidar aerosol subtypes by using the lidar profile observations. This method uses additional information from the vertical texture of the feature instead of using only the layer information. Note that in the integrated layer properties, the texture information has been masked due to averaging. Our results will show how the texture information plays a role in the classification. This preliminary work explores the benefits and potential of deep machine learning methods for lidar retrievals and focuses on the aerosol subtype classification. The broader application extends to the classification of other feature types. Future applications include the developing deep machine learning methods with neural networks to retrieve properties of the features, and studies of indirect effect of cloud-aerosol interaction from lidar measurements.

Shan Zeng Kowalski↗

Visualization of Two-phase Flow Maldistribution in Brazed Plate Heat Exchangers

Brazed plate heat exchangers (BPHEs) are widely used in refrigeration and HVAC applications, but are susceptible to two-phase flow maldistribution especially when operated as evaporators. Existing visualization approaches are either limited to idealized conditions or suffer from poor optical transparency. This paper presents a novel visualization method in which one edge of a BPHE, parallel to the refrigerant inlet or outlet port, is removed by wire electrical discharge machining and replaced with a flat, transparent plate. The planar geometry allows the use of optically and infrared (IR)-transparent materials, enabling both high-speed videography and IR thermography of the two-phase flow at the channel entrances and exits. Preliminary tests with R134a and R1234ze(Z) at saturation temperatures between 5 °C and 15 °C demonstrate that distinct two-phase flow patterns in the inlet header can be clearly identified and differentiated under realistic operating conditions. Potentials of optical flow analysis of high-speed videos are shown to provide objective, quantitative indicators for flow regime characterization and comparison. IR imaging of the outlet port reveals non-uniform temperature distributions at the channel exits, providing independent evidence of maldistribution across the channel stack. Limitations of IR temperature accuracy due to the spectral properties of the sapphire window are discussed, and directions for improvement are identified.

Hausherr, Carsten [Technical University of Berlin ↗

On the wall-normal velocity of the compressible boundary-layer equations

Numerical methods for the compressible boundary-layer equations are facilitated by transformation from the physical (x,y) plane to a computational (xi,eta) plane in which the evolution of the flow is 'slow' in the time-like xi direction. The commonly used Levy-Lees transformation results in a computationally well-behaved problem for a wide class of non-similar boundary-layer flows, but it complicates interpretation of the solution in physical space. Specifically, the transformation is inherently nonlinear, and the physical wall-normal velocity is transformed out of the problem and is not readily recovered. In light of recent research which shows mean-flow non-parallelism to significantly influence the stability of high-speed compressible flows, the contribution of the wall-normal velocity in the analysis of stability should not be routinely neglected. Conventional methods extract the wall-normal velocity in physical space from the continuity equation, using finite-difference techniques and interpolation procedures. The present spectrally-accurate method extracts the wall-normal velocity directly from the transformation itself, without interpolation, leaving the continuity equation free as a check on the quality of the solution. The present method for recovering wall-normal velocity, when used in conjunction with a highly-accurate spectral collocation method for solving the compressible boundary-layer equations, results in a discrete solution which is extraordinarily smooth and accurate, and which satisfies the continuity equation nearly to machine precision. These qualities make the method well suited to the computation of the non-parallel mean flows needed by spatial direct numerical simulations (DNS) and parabolized stability equation (PSE) approaches to the analysis of stability.

Pruett, C. David↗