Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Devastator Parallel Discrete Event Simulation Runtime (Devastator) v1.0

The Devastator runtime is a modern C++ implementation of optimistic parallel discrete event simulation methods. Devastator allows simulation application code to productively specify their component and event functionality with C++14 constructs. It utilizes GASNet-EX for distributed memory communication and includes parallel performance optimizations such as light-weight thread message queues and asynchronous GVT. Furthermore, it supports efficient event broadcasts and pause-rewind-resume functionality to support periodic load balancing and outer loop optimization algorithms.

Chan, Cy↗

SBIR Phase I Final Report, TACO: Distributed and Heterogeneous Sparse Compiler

Tensor algebra is a powerful tool for computing, but writing optimized codes that operate on sparse tensors can be very complex. This project enables a Tensor Algebra Compiler (TACO) that simplifies this task from man-years to man-days and extends TACO to support complex and large distributed systems. This report details the hypotheses, approaches used, and findings in this project.

97 MATHEMATICS AND COMPUTING↗

Cobalt-Free Cathodes for Next Generation Li-Ion Batteries

In this U.S. Department of Energy sponsored project Nexceris, in collaboration with project partners; The Ohio State University and Navitas Advanced Systems have advanced the technical maturity of a non-cobalt containing cathode for next-generation Li-ion batteries. The cathode is based on the lithium manganese nickel-titanium oxide, LiNi 0.5 Mn 1.5 TiO 4 (LNMTO) high voltage spinel. To address limitations with poor cycle and calendar life an microstucturally hierarchical LNMO/LNMTO core-shell cathode powder has been developed that enables the formation of a solid-electrolyte interface that effectively passivates the cathode surface. The microstructural enhancements of the cathode material focus on preferentially enriching the surface with titanium. In parallel, new, optimized binder and electrolyte chemistries have been incorporated to address degradation mechanisms associated with high-voltage systems. Single-layer pouch cell and large-format 2-Ah cell testing have shown that an optimized LNMO/LNMTO core-shell powder significantly improves initial cell capacity and cycle life compared to homogeneous LNMO powder. To support development and the fabrication of 2-Ah cells a novel Hybrid Alternative Wet-Chemical Synthesis (HAWCS) process have been developed. This low-cost, synthesis approach enables the excellent compositional and particle morphology control achieved with co-precipitation without the strict process controls and associated expensive process equipment.

25 ENERGY STORAGE↗

Automatic Generation of Algorithms for High-Speed Reliable Lossy Data Compression (Final Report)

Fast reliable data compression is urgently needed for many leading-edge scientific instruments and for exascale high-performance computing applications because they produce vast amounts of data at extremely high rates. The goal of this project has been to develop a framework named LC that is able to automatically generate high-speed lossless and reliable lossy compression and decompression algorithms that can be customized for different kinds of data. The resulting LC framework is freely available on GitHub. To achieve high-speed operation, LC outputs optimized and parallelized CPU and GPU implementations of the generated algorithms. To ensure the quality of lossily compressed data, LC guarantees the user-provided error bound. To be able to customize the compression algorithm to various use cases, LC can synthesize millions of different algorithms and automatically search for the one that works best for the given data. We have already employed LC to create state-of-the-art lossless and lossy compressors for scientific data as well as leading lossless compressors for images. We hope that LC and the customized, fast, reliable, and CPU/GPU-compatible compression algorithms that it can generate will greatly benefit the many scientific applications that need not only high trustworthiness but also high performance.

97 MATHEMATICS AND COMPUTING↗

Cell-free Scaled Production and Adjuvant Addition to a Recombinant Major Outer Membrane Protein from Chlamydia muridarum for Vaccine Development

Subunit vaccines offer advantages over more traditional inactivated or attenuated whole-cell-derived vaccines in safety, stability, and standard manufacturing. To achieve an effective protein-based subunit vaccine, the protein antigen often needs to adopt a native-like conformation. This is particularly important for pathogensurface antigens that are membrane-bound proteins. Cell-free methods have been successfully used to produce correctly folded functional membrane protein through the co-translation of nanolipoprotein particles (NLPs), commonly known as nanodiscs. This strategy can be used to produce subunit vaccines consisting of membrane proteins in a lipid-bound environment. However, cell-free protein production is often limited to small scale (<1 mL). The amount of protein produced in small-scale production runs is usually sufficient for biochemical and biophysical studies. However, the cell-free process needs to be scaled up, optimized, and carefully tested to obtain enough protein for vaccine studies in animal models. Other processes involved in vaccine production, such as purification, adjuvant addition, and lyophilization, need to be optimized in parallel. This paper reports the development of a scaled-up protocol to express, purify, and formulate a membrane-bound protein subunit vaccine

59 BASIC BIOLOGICAL SCIENCES↗

Geometric registration and rectification of spaceborne SAR imagery

This paper describes the development of automated location and geometric rectification techniques for digitally processed synthetic aperture radar (SAR) imagery. A software package has been developed that is capable of determining the absolute location of an image pixel to within 60 m using only the spacecraft ephemeris data and the characteristics of the SAR data collection and processing system. Based on this location capability algorithms have been developed that geometrically rectify the imagery, register it to a common coordinate system and mosaic multiple frames to form extended digital SAR maps. These algorithms have been optimized using parallel processing techniques to minimize the operating time. Test results are given using Seasat SAR data.

Curlander, J. C.↗

Optimal expression evaluation for data parallel architectures

A data parallel machine represents an array or other composite data structure by allocating one processor per data item. A pointwise operation can be performed between two such arrays in unit time, provided their corresponding elements are allocated in the same processors. If the arrays are not aligned in this fashion, the cost of moving one or both of them is part of the cost of operation. The choice of where to perform the operation then affects this cost. If an expression with several operands is to be evaluated, there may be many choices of where to perform the intermediate operations. An efficient algorithm is given to find the minimum cost way to evaluate an expression, for several different data parallel architectures. The algorithm applies to any architecture in which the metric describing the cost of moving an array has a property called robustness. This encompasses most of the common data parallel communication architectures, including meshes of arbitrary dimension and hypercubes.

Gilbert, J. R.↗

Optimal expression evaluation for data parallel architectures

A data parallel machine represents an array or other composits data structure by allocating one processor per data item. A pointwise operation can be performed between two such arrays in unit time, provided their corresponding elements are allocated in the same processors. If the arrays are not aligned in this fashion, the cost of moving one or both of them is part of the cost of operation. The choice of where to perform the operation then affects this cost. If an expression with several operands is to be evaluated, there may be many choices of where to perform the intermediate operations. An efficient algorithm is given to find the minimum cost way to evaluate an expression, for several different data parallel architectures. The algorithm applies to any architecture in which the metric describing the cost of moving an array has a property called robustness. This encompasses most of the common data parallel communication architectures, including meshes of arbitrary dimension and hypercubes.

Gilbert, John R.↗

Optimal expression evaluation for data parallel architectures

A data parallel machine represents an array or other composite data structure by allocating one processor (at least conceptually) per data item. A pointwise operation can be performed between two such arrays in unit time, provided their corresponding elements are allocated in the same processors. If the arrays are not aligned in this fashion, the cost of moving one or both of them is part of the cost of the operation. The choice of where to perform the operation then affects this cost. If an expression with several operands is to be evaluated, there may be many choices of where to perform the intermediate operations. An efficient algorithm is given to find the minimum-cost way to evaluate an expression, for several different data parallel architectures. This algorithm applies to any architecture in which the metric describing the cost of moving an array is robust. This encompasses most of the common data parallel communication architectures, including meshes of arbitrary dimension and hypercubes. Remarks are made on several variations of the problem, some of which are solved and some of which remain open.

Gilbert, John R.↗

Parallelization of NAS Benchmarks for Shared Memory Multiprocessors

This paper presents our experiences of parallelizing the sequential implementation of NAS benchmarks using compiler directives on SGI Origin2000 distributed shared memory (DSM) system. Porting existing applications to new high performance parallel and distributed computing platforms is a challenging task. Ideally, a user develops a sequential version of the application, leaving the task of porting to new generations of high performance computing systems to parallelization tools and compilers. Due to the simplicity of programming shared-memory multiprocessors, compiler developers have provided various facilities to allow the users to exploit parallelism. Native compilers on SGI Origin2000 support multiprocessing directives to allow users to exploit loop-level parallelism in their programs. Additionally, supporting tools can accomplish this process automatically and present the results of parallelization to the users. We experimented with these compiler directives and supporting tools by parallelizing sequential implementation of NAS benchmarks. Results reported in this paper indicate that with minimal effort, the performance gain is comparable with the hand-parallelized, carefully optimized, message-passing implementations of the same benchmarks.

Waheed, Abdul↗

Large-Scale NASA Science Applications on the Columbia Supercluster

Columbia, NASA's newest 61 teraflops supercomputer that became operational late last year, is a highly integrated Altix cluster of 10,240 processors, and was named to honor the crew of the Space Shuttle lost in early 2003. Constructed in just four months, Columbia increased NASA's computing capability ten-fold, and revitalized the Agency's high-end computing efforts. Significant cutting-edge science and engineering simulations in the areas of space and Earth sciences, as well as aeronautics and space operations, are already occurring on this largest operational Linux supercomputer, demonstrating its capacity and capability to accelerate NASA's space exploration vision. The presentation will describe how an integrated environment consisting not only of next-generation systems, but also modeling and simulation, high-speed networking, parallel performance optimization, and advanced data analysis and visualization, is being used to reduce design cycle time, accelerate scientific discovery, conduct parametric analysis of multiple scenarios, and enhance safety during the life cycle of NASA missions. The talk will conclude by discussing how NAS partnered with various NASA centers, other government agencies, computer industry, and academia, to create a national resource in large-scale modeling and simulation.

Brooks, Walter↗

Simulation of Aerosols and Chemistry with a Unified Global Model

This project is to continue the development of the global simulation capabilities of tropospheric and stratospheric chemistry and aerosols in a unified global model. This is a part of our overall investigation of aerosol-chemistry-climate interaction. In the past year, we have enabled the tropospheric chemistry simulations based on the GEOS-CHEM model, and added stratospheric chemical reactions into the GEOS-CHEM such that a globally unified troposphere-stratosphere chemistry and transport can be simulated consistently without any simplifications. The tropospheric chemical mechanism in the GEOS-CHEM includes 80 species and 150 reactions. 24 tracers are transported, including O3, NOx, total nitrogen (NOy), H2O2, CO, and several types of hydrocarbon. The chemical solver used in the GEOS-CHEM model is a highly accurate sparse-matrix vectorized Gear solver (SMVGEAR). The stratospheric chemical mechanism includes an additional approximately 100 reactions and photolysis processes. Because of the large number of total chemical reactions and photolysis processes and very different photochemical regimes involved in the unified simulation, the model demands significant computer resources that are currently not practical. Therefore, several improvements will be taken, such as massive parallelization, code optimization, or selecting a faster solver. We have also continued aerosol simulation (including sulfate, dust, black carbon, organic carbon, and sea-salt) in the global model to cover most of year 2002. These results have been made available to many groups worldwide and accessible from the website http://code916.gsfc.nasa.gov/People/Chin/aot.html.

Chin, Mian↗

Robotically Assembled Aerospace Structures: Digital Material Assembly using a Gantry-Type Assembler

This paper evaluates the development of automated assembly techniques for discrete lattice structures using a multi-axis gantry type CNC machine. These lattices are made of discrete components called digital materials. We present the development of a specialized end effector that works in conjunction with the CNC machine to assemble these lattices. With this configuration we are able to place voxels at a rate of 1.5 per minute. The scalability of digital material structures due to the incremental modular assembly is one of its key traits and an important metric of interest. We investigate the build times of a 5x5 beam structure on the scale of 1 meter (325 parts), 10 meters (3,250 parts), and 30 meters (9,750 parts). Utilizing the current configuration with a single end effector, performing serial assembly with a globally fixed feed station at the edge of the build volume, the build time increases according to a scaling law of n4, where n is the build scale. Build times can be reduced significantly by integrating feed systems into the gantry itself, resulting in a scaling law of n3. A completely serial assembly process will encounter time limitations as build scale increases. Automated assembly for digital materials can assemble high performance structures from discrete parts, and techniques such as built in feed systems, parallelization, and optimization of the fastening process will yield much higher throughput.

Mechanical Structure↗

Robotically Assembled Aerospace Structures: Digital Material Assembly using a Gantry-Type Assembler

This paper evaluates the development of automated assembly techniques for discrete lattice structures using a multi-axis gantry type CNC machine. These lattices are made of discrete components called "digital materials." We present the development of a specialized end effector that works in conjunction with the CNC machine to assemble these lattices. With this configuration we are able to place voxels at a rate of 1.5 per minute. The scalability of digital material structures due to the incremental modular assembly is one of its key traits and an important metric of interest. We investigate the build times of a 5x5 beam structure on the scale of 1 meter (325 parts), 10 meters (3,250 parts), and 30 meters (9,750 parts). Utilizing the current configuration with a single end effector, performing serial assembly with a globally fixed feed station at the edge of the build volume, the build time increases according to a scaling law of n4, where n is the build scale. Build times can be reduced significantly by integrating feed systems into the gantry itself, resulting in a scaling law of n3. A completely serial assembly process will encounter time limitations as build scale increases. Automated assembly for digital materials can assemble high performance structures from discrete parts, and techniques such as built in feed systems, parallelization, and optimization of the fastening process will yield much higher throughput.

Manufacturing↗

“PowerCell”: The Interface Between Mars Resources and Human Exploration

The barriers to forming human settlements on Mars are high but surmountable within our lifetime. While the Apollo astronauts carried their life support with them, our success in exploring and forming settlements on Mars depends on our ability to use local Martian resources to generate the materials and conditions humans need to survive, so-called in situ resource utilization (ISRU). On Earth, biology provides us with food, shelter, oxygen, and other materials. Off-planet, synthetic biology will enable numerous parallel productions: optimized food production, water treatment, air treatment, environmental monitoring, regolith biomining, waste management, cell based biomaterial production, biocementation, and in situ synthesis based on received DNA sequences. How will the organisms responsible for these synthetic production systems obtain organic carbon and fixed nitrogen in the hostile Martian environment? We envision a synthetic-biology enabled Martian colony and introduce here the critical intermediate component a biological power source needed to transform the in situ resources found on Mars into biological feedstocks to enable growth of production organisms. Here, we present our first PowerCell, a photosynthetic and nitrogen-fixing filamentous cyanobacterium engineered to provide a carbon-rich fuel source for a biological life support system on Mars. We provide a vision of how the PowerCell system will operate in a Martian colony based on ground experiments and preparations for testing in space as a NASA secondary payload aboard the upcoming DLR Eu:CROPIS satellite mission experiments.

Rothschild, Lynn J.↗

A Mixed Integer Efficient Global Optimization Algorithm with Multiple Infill Strategy - Applied to a Wing Topology Optimization Problem

With the advancement in high performance computing and numerical optimization techniques,engineering design optimization problems are becoming more complex, larger scale,higher fidelity, and computationally more demanding, requiring longer run times than ever before. There exists methodologies and techniques that can address some of these challenges but very few can address all, and most are limited in the extent that these concerns can be addressed. With the goal of addressing such challenging engineering problems, we developed anew optimization framework, named AMIEGO, that combines concepts from surrogate-based optimization approaches, gradient-based numerical methods, Partial Least Squares, evolutionary algorithms, and Branch-and-Bound, providing newer capabilities that were not previouslyperceived. However, the original version of this framework, in the process of adaptive samplingto explore and exploit the design space, finds only a single sample point per iteration. The efforthere builds upon this previously developed optimization framework to include multiple infillsampling capability that combines the concept of generalized expected improvement function,unsupervised learning, and multi-objective evolutionary technique. To demonstrate, AMIEGOwith the multiple infill capability (called AMIEGO-MIMOS) solves a series of increasingly difficultengineering design optimization problems. The results reveal the performance of the newapproach is problem dependent. When applied to a ten-bar truss problem, the newly proposedmultiple infill strategy consistently leads to a better design solutions when compared to theexisting CPTV method (implemented with the context of the AMIEGO framework). On theother hand, when applied to a mixed-integer high fidelity wing topology optimization problem- MIMOS, despite showing a steeper convergence at the start, eventually leads to an inferiorsolution as compared to CPTV approach. These results also reveal that a small number ofstarting points, in general, are sufficient to lead to a good overall solution.

Mixed-integer optimization↗

FLASSH 1.0: Thermal scattering law evaluation and cross section generation for reactor physics applications

The Full Law Analysis Scattering System Hub (FLASSH) is a modern, advanced code which evaluates the thermal scattering law (TSL) along with accompanying cross sections. FLASSH features generalized methods which accommodate any material structure. Historical approximations including the incoherent and cubic approximations have been removed. Instead, the latest release of FLASSH features advanced physics options including distinct corrections (1-phonon contributions) and non-cubic formulations. The non-cubic elastic and inelastic contributions are necessary to accurately evaluate 1-phonon contributions. Both non-cubic and 1-phonon calculations require high-density sampling of the various scattering directions. Optimization and parallelization of these routines were therefore necessary to produce results in a reasonable timeframe. With these notable improvements to the generalized TSL, FLASSH 1.0 meets benchmark requirements, demonstrating noticeable agreement with experiment for both TSLs and the resulting integrated cross sections. Additional features including a graphical user interface (GUI), plotting diagnostics, and formatted output options including ACE files allow users to complete a TSL evaluation with minimal input and maximum flexibility. The user GUI creates input files for FLASSH, reducing user error and also providing built-in error checks. Autofill options and suggested input values help make TSL evaluation accessible to novice users. The FLASSH code is compiled to run on both Windows and Linux platforms with automatic parallelization. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗