Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “task parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Modeling techniques in a parallelizing compiler for the B-HIVE multiprocessor system

The parallelizing compiler for the B-HIVE loosely-coupled multiprocessor system uses a medium grain model to minimize the communication overhead. A medium grain model is shown to be an optimum way of merging fine grain operations into parallel tasks such that the parallelism obtained at the grain level is retained and communication overhead is decreased. A new communication model is introduced in this paper, allowing additional overlap between computation and communication. Simulation results indicate that the medium grain communication model shows promise for automatic parallelization for a loosely-coupled multiprocessor system.

Kim, Sukil↗

Parallel adaptive mesh refinement within the PUMAA3D Project

To enable the solution of large-scale applications on distributed memory architectures, we are designing and implementing parallel algorithms for the fundamental tasks of unstructured mesh computation. In this paper, we discuss efficient algorithms developed for two of these tasks: parallel adaptive mesh refinement and mesh partitioning. The algorithms are discussed in the context of two-dimensional finite element solution on triangular meshes, but are suitable for use with a variety of element types and with h- or p-refinement. Results demonstrating the scalability and efficiency of the refinement algorithm and the quality of the mesh partitioning are presented for several test problems on the Intel DELTA.

Freitag, Lori↗

Architectures for reasoning in parallel

The research conducted has dealt with rule-based expert systems. The algorithms that may lead to effective parallelization of them were investigated. Both the forward and backward chained control paradigms were investigated in the course of this work. The best computer architecture for the developed and investigated algorithms has been researched. Two experimental vehicles were developed to facilitate this research. They are Backpac, a parallel backward chained rule-based reasoning system and Datapac, a parallel forward chained rule-based reasoning system. Both systems have been written in Multilisp, a version of Lisp which contains the parallel construct, future. Applying the future function to a function causes the function to become a task parallel to the spawning task. Additionally, Backpac and Datapac have been run on several disparate parallel processors. The machines are an Encore Multimax with 10 processors, the Concert Multiprocessor with 64 processors, and a 32 processor BBN GP1000. Both the Concert and the GP1000 are switch-based machines. The Multimax has all its processors hung off a common bus. All are shared memory machines, but have different schemes for sharing the memory and different locales for the shared memory. The main results of the investigations come from experiments on the 10 processor Encore and the Concert with partitions of 32 or less processors. Additionally, experiments have been run with a stripped down version of EMYCIN.

Hall, Lawrence O.↗

The engine design engine. A clustered computer platform for the aerodynamic inverse design and analysis of a full engine

An application for parallel computation on a combined cluster of powerful workstations and supercomputers was developed. A Parallel Virtual Machine (PVM) is used as message passage language on a macro-tasking parallelization of the Aerodynamic Inverse Design and Analysis for a Full Engine computer code. The heterogeneous nature of the cluster is perfectly handled by the controlling host machine. Communication is established via Ethernet with the TCP/IP protocol over an open network. A reasonable overhead is imposed for internode communication, rendering an efficient utilization of the engaged processors. Perhaps one of the most interesting features of the system is its versatile nature, that permits the usage of the computational resources available that are experiencing less use at a given point in time.

Sanz, J.↗

Formulation of consumables management models. Volume 1: Mission planning

Development of an STS (Space Transportation System) interactive computer program MPP (Mission Planning Processor) working model was conducted. A summary of the computer program development and those supporting tasks conducted is presented. Development of the MPP Computer Program is discussed. This development was supported by several parallel tasks. These tasks either directly supported the program development, or provided information for future application and/or modification to the program in relation to the flight planning and flight operations of the STS and advanced spacecraft. The supporting tasks also included development of a Space Station MPP to demonstrate the applicability of the analytical methods developed under this RTOP to more advanced spacecraft than the STS.

Torian, J. G.↗

Efficient Parallelization of Irregular Applications on GPU Architectures

With the enlarging computation capacity of general Graphics Processing Units (GPUs), leveraging GPUs to accelerate parallel applications has become a critical topic in academia and industry. However, a wide range of irregular applications with the computation-/memory-intensive nature cannot easily achieve high GPU utilization. The challenges mainly involve the following aspects: first, data dependence leads to coarse-grained kernel and inefficient parallelism; second, heavy GPU memory usage may cause frequent memory evictions and extra overhead of I/O; third, specific computation patterns produce memory redundancies; last, workload balance and data reusability conjunctly benefit the overall performance, but there may exist a dynamic trade-off between them. Targeting these challenges, this dissertation proposes multiple optimizations to accelerate two real-world applications: many-body correlation functions to simulate nuclear physics in a large-scale scientific system; the other is the eALS-based matrix factorization recommendation system. To accelerate the calculations of many-body correlation functions, this dissertation presents three frameworks in GPU memory management and multi-GPU scheduling. Firstly, an optimized systematic GPU memory management framework, MemHC, utilizes a series of new memory reduction designs in GPU memory allocation, CPU/GPU communications, and GPU memory oversubscription. Secondly, an enhanced multi-GPU scheduling framework, MICCO, particularly by taking both data dimension (e.g., data reuse and data eviction) and computation dimension into account. MICCO designs a heuristic scheduling algorithm and a machine learning-based regression model to generate the optimal settings of a proposed new concept to manage the trade-off. Thirdly, a locality-aware multi-GPU scheduling framework. This scheduler leverages pipeline batch generation with a looking-ahead strategy by building local dependency graphs for memory transfer reduction and better data reuse, achieving up to 79.92% memory cost reduction and 1.67x speedup. To parallelize the eALS-based recommendation system, this dissertation proposes an efficient CPU/GPU heterogeneous recommendation system, HEALS. HEALS employs newly designed architecture-adaptive data formats to achieve load balance and good data locality on CPU and GPU. To mitigate the data dependence, HEALS presents a CPU/GPU collaboration model for both task parallelism and data parallelism with multiple kernel computation optimizations. In summary, this dissertation efficiently accelerates two typical irregular applications on GPUs by building four frameworks, including CPU/GPU collaboration, GPU memory management, and multi-GPU scheduling.

Wang, Qihan↗

High Throughput Source-less Plasma Deposition of Structured Silicon Anodes for Lithium-Ion Batteries

Amprius developed a manufacturing solution for silicon nanowire anode that relies on an inexpensive, high throughput, and high gas precursor utilization plasma deposition method that uses the anode foils as electrodes for plasma generation. The capacitively couple plasma (CCP) method is used in semiconductor and photovoltaic industry and Amprius modified existing high throughput equipment to use anode foils and to deposit amorphous silicon. The equipment was installed ahead of the program at Amprius site. The rest of the tasks included foil handling and process development. The equipment passed site acceptance tests (SAT) and the process parameter mapping was completed, indicating that the target process window limits produce output materials at the rate and with yield and specifications that meet the manufacturing target criteria. Amprius has hired supporting personnel to optimize processes and run the equipment. A parallel task verified the baseline performance of the silicon anode material, to be used as reference for the new manufacturing method.

25 ENERGY STORAGE↗

Extending HPF for advanced data parallel applications

The stated goal of High Performance Fortran (HPF) was to 'address the problems of writing data parallel programs where the distribution of data affects performance'. After examining the current version of the language we are led to the conclusion that HPF has not fully achieved this goal. While the basic distribution functions offered by the language - regular block, cyclic, and block cyclic distributions - can support regular numerical algorithms, advanced applications such as particle-in-cell codes or unstructured mesh solvers cannot be expressed adequately. We believe that this is a major weakness of HPF, significantly reducing its chances of becoming accepted in the numeric community. The paper discusses the data distribution and alignment issues in detail, points out some flaws in the basic language, and outlines possible future paths of development. Furthermore, we briefly deal with the issue of task parallelism and its integration with the data parallel paradigm of HPF.

Chapman, Barbara↗

DSN Beowulf Cluster-Based VLBI Correlator

The NASA Deep Space Network (DSN) requires a broadband VLBI (very long baseline interferometry) correlator to process data routinely taken as part of the VLBI source Catalogue Maintenance and Enhancement task (CAT M&E) and the Time and Earth Motion Precision Observations task (TEMPO). The data provided by these measurements are a crucial ingredient in the formation of precision deep-space navigation models. In addition, a VLBI correlator is needed to provide support for other VLBI related activities for both internal and external customers. The JPL VLBI Correlator (JVC) was designed, developed, and delivered to the DSN as a successor to the legacy Block II Correlator. The JVC is a full-capability VLBI correlator that uses software processes running on multiple computers to cross-correlate two-antenna broadband noise data. Components of this new system (see Figure 1) consist of Linux PCs integrated into a Beowulf Cluster, an existing Mark5 data storage system, a RAID array, an existing software correlator package (SoftC) originally developed for Delta DOR Navigation processing, and various custom- developed software processes and scripts. Parallel processing on the JVC is achieved by assigning slave nodes of the Beowulf cluster to process separate scans in parallel until all scans have been processed. Due to the single stream sequential playback of the Mark5 data, some ramp-up time is required before all nodes can have access to required scan data. Core functions of each processing step are accomplished using optimized C programs. The coordination and execution of these programs across the cluster is accomplished using Pearl scripts, PostgreSQL commands, and a handful of miscellaneous system utilities. Mark5 data modules are loaded on Mark5 Data systems playback units, one per station. Data processing is started when the operator scans the Mark5 systems and runs a script that reads various configuration files and then creates an experiment-dependent status database used to delegate parallel tasks between nodes and storage areas (see Figure 2). This script forks into three processes: extract, translate, and correlate. Each of these processes iterates on available scan data and updates the status database as the work for each scan is completed. The extract process coordinates and monitors the transfer of data from each of the Mark5s to the Beowulf RAID storage systems. The translate process monitors and executes the data conversion processes on available scan files, and writes the translated files to the slave nodes. The correlate process monitors the execution of SoftC correlation processes on the slave nodes for scans that have completed translation. A comparison of the JVC and the legacy Block II correlator outputs reveals they are well within a formal error, and that the data are comparable with respect to their use in flight navigation. The processing speed of the JVC is improved over the Block II correlator by a factor of 4, largely due to the elimination of the reel-to-reel tape drives used in the Block II correlator.

Rogstad, Stephen P.↗

Evaluation of Steam Cycle Upgrades to Improve the Competitiveness of U.S. Coal Power Plants (Final Scientific / Technical Report)

Increasing the competitiveness of the existing pulverized-coal utility fleet in the United States may be achieved by decreasing heat rate, via increases in steam cycle efficiency through upgraded steam temperatures and use of latest technology available in steam turbine and blading design. The average net plant efficiency of the US coal-fired fleet is approximately 33% (HHV). Plant efficiency increases to approximately 41.4% (HHV) at 1,350°F (732°C) steam temperature. However, achieving these Advanced Ultra-Super Critical (AUSC) steam conditions requires the use of advanced high-temperature materials. While there has been a significant amount of DOE-funded materials R&D, most of the related design work has focused on new (greenfield) units, rather than on opportunities to retrofit this advanced technology to the existing utility fleet. If technology, based upon the advanced materials required for AUSC steam conditions, may be applied to the existing fleet, using an economically viable retrofit, a higher capacity factor can be expected as a result of the increased plant competitiveness. The Electric Power Research Institute (EPRI) was awarded a project by the US Department of Energy to examine the technical and economic feasibility of a series of steam cycle upgrades to the two most prevalent types of U.S. coal power plants: 2,400 psig (16.6 MPa) subcritical and 3,500 psig (24.1 MPa) supercritical pulverized coal units. The nine upgrade options that were originally being considered included increasing the main and reheat steam temperatures from 1,000°F (538°C) to 1,100°, 1,200°, and 1,350°F (593°C, 649°C, and 732°C) while holding the steam pressures constant at their original design values, and cases where just the main steam or reheat steam temperatures were increased. The objective was to minimize the modifications required to the existing power plant while still providing a significant improvement in heat rate. The upgrade options assumed that the boiler enclosure envelope remained unchanged from each base case, and that all applicable OEM design guidelines for normal commercial units were imposed. For the highest temperature supercritical case, an option of using a low-pressure molten salt loop to transfer heat from the furnace to the steam was examined. The first major task of the work scope was designed to examine the technical feasibility of various upgrade options, while the subsequent work determined economic viability of the technically feasible upgrade options. Prior to evaluating the effect of these increased temperatures, a “base case” model of a subcritical and supercritical PC boiler was created, which was used for comparative purposes. Upgrade options were evaluated at full-load, part-load and dynamic transient conditions. Once the technical feasibility of each upgrade option was evaluated, the economic value of the heat rate improvement of each feasible option was determined by detailed modeling of unit dispatch in several regional power markets. The dispatch model was used to estimate the amount of revenue from power sales the upgraded unit would receive in comparison to a non-upgraded version of the same power plant. As a parallel task to the dispatch analysis, the capital cost of implementing the upgrades was estimated. The capital cost estimates were then compared to the increased revenue estimated by the dispatch modeling to determine the economic attractiveness of each upgrade option. Several upgrade options were determined to be technically feasible. The net present value (NPV) of the costs for steam cycle upgrades considered in this study ranged from approximately $\$$111 to $\$$130 million. The economic modeling results show that the unit dispatch changes resulting from steam cycle upgrades are relatively small, due largely to heat rate (and operating cost) changes being relatively small. Additionally, the cost of each upgrade exceeds the net revenue increases associated with the upgrade case. Note that the breakeven values are higher for subcritical retrofits, but the capital costs for the subcritical upgrades are also slightly higher. In typical new pulverized coal plants, fuel accounts for approximately 25% of the cost of electricity (COE), while capital costs represent around 50% of the COE. Therefore, in order to improve the heat rate by 4% one can only afford to increase the capital cost by 2%, at the same cost of electricity. The conclusion of this study is that without a cost for emitting CO 2 , it will be difficult to pay for significant efficiency improvements on plants firing low cost coals.

01 COAL, LIGNITE, AND PEAT↗

A fast, dense Chebyshev solver for electronic structure on GPUs

Matrix diagonalization is almost always involved in computing the density matrix needed in quantum chemistry calculations. In the case of modest matrix sizes (≲4000), performance of traditional dense diagonalization algorithms on modern GPUs is underwhelming compared to the peak performance of these devices. This motivates the exploration of alternative algorithms better suited to these types of architectures. We newly derive, and present in detail, an existing Chebyshev expansion algorithm whose number of required matrix multiplications scales with the square root of the number of terms in the expansion. Focusing on dense matrices of modest size, our implementation on GPUs results in large speed ups when compared to diagonalization. Additionally, we improve upon this existing method by capitalizing on the inherent task parallelism and concurrency in the algorithm. Furthermore, this improvement is implemented on GPUs by using CUDA and HIP streams via the MAGMA library and leads to a significant speed up over the serial-only approach for smaller (≲1000) matrix sizes. Finally, we apply our technique to a model system with a high density of states around the Fermi level, which typically presents significant challenges.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

Performance Enhancement of APW+lo Calculations by Simplest Separation of Concerns

Full-potential linearized augmented plane wave (LAPW) and APW plus local orbital (APW+lo) codes differ widely in both their user interfaces and in capabilities for calculations and analysis beyond their common central task of all-electron solution of the Kohn–Sham equations. However, that common central task opens a possible route to performance enhancement, namely to offload the basic LAPW/APW+lo algorithms to a library optimized purely for that purpose. To explore that opportunity, we have interfaced the Exciting-Plus (“EP”) LAPW/APW+lo DFT code with the highly optimized SIRIUS multi-functional DFT package. This simplest realization of the separation of concerns approach yields substantial performance over the base EP code via additional task parallelism without significant change in the EP source code or user interface. We provide benchmarks of the interfaced code against the original EP using small bulk systems, and demonstrate performance on a spin-crossover molecule and magnetic molecule that are of size and complexity at the margins of the capability of the EP code itself.

Zhang, Long↗

Space applications of artificial intelligence; Proceedings of the Annual Goddard Conference, Greenbelt, MD, May 16, 17, 1989

Theoretical and implementation aspects of AI systems for space applications are discussed in reviews and reports. Sections are devoted to planning and scheduling, fault isolation and diagnosis, data management, modeling and simulation, and development tools and methods. Particular attention is given to a situated reasoning architecture for space repair and replace tasks, parallel plan execution with self-processing networks, the electrical diagnostics expert system for Spacelab life-sciences experiments, diagnostic tolerance for missing sensor data, the integration of perception and reasoning in fast neural modules, a connectionist model for dynamic control, and applications of fuzzy sets to the development of rule-based expert systems.

Rash, James L.↗

ALS turbomachinery technology

Advanced Development Programs are being pursued by Rocketdyne, Aerojet, and Pratt and Whitney to define and validate design approaches toward producing low-cost, reliable liquid-hydrogen and liquid-oxygen turbopumps for a 2580 kN (580 klb) thrust Advanced Launch System. The generic approach, which is evolving after 18 months of trade studies and conceptual and preliminary design efforts, is explained. In addition, the preliminary liquid-hydrogen turbopump designs produced in parallel tasks by Rocketdyne and Aerojet and the liquid-oxygen turbopump design produced by Pratt and Whitney are described, and technology features and issues are discussed.

Csomor, A.↗

Parallel methods for dynamic simulation of multiple manipulator systems

In this paper, efficient dynamic simulation algorithms for a system of m manipulators, cooperating to manipulate a large load, are developed; their performance, using two possible forms of parallelism on a general-purpose parallel computer, is investigated. One form, temporal parallelism, is obtained with the use of parallel numerical integration methods. A speedup of 3.78 on four processors of CRAY Y-MP8 was achieved with a parallel four-point block predictor-corrector method for the simulation of a four manipulator system. These multi-point methods suffer from reduced accuracy, and when comparing these runs with a serial integration method, the speedup can be as low as 1.83 for simulations with the same accuracy. To regain the performance lost due to accuracy problems, a second form of parallelism is employed. Spatial parallelism allows most of the dynamics of each manipulator chain to be computed simultaneously. Used exclusively in the four processor case, this form of parallelism in conjunction with a serial integration method results in a speedup of 3.1 on four processors over the best serial method. In cases where there are either more processors available or fewer chains in the system, the multi-point parallel integration methods are still advantageous despite the reduced accuracy because both forms of parallelism can then combine to generate more parallel tasks and achieve greater effective speedups. This paper also includes results for these cases.

Mcmillan, Scott↗

Transcriptional profiling reveals regulated genes in the hippocampus during memory formation

Transcriptional profiling (TP) offers a powerful approach to identify genes activated during memory formation and, by inference, the molecular pathways involved. Trace eyeblink conditioning is well suited for the study of regional gene expression because it requires the hippocampus, whereas the highly parallel task, delay conditioning, does not. First, we determined when gene expression was most regulated during trace conditioning. Rats were exposed to 200 trials per day of paired and unpaired stimuli each day for 4 days. Changes in gene expression were most apparent 24 h after exposure to 200 trials. Therefore, we profiled gene expression in the hippocampus 24 h after 200 trials of trace eyeblink conditioning, on multiple arrays using additional animals. Of 1,186 genes on the filter array, seven genes met the statistical criteria and were also validated by real-time polymerase chain reaction. These genes were growth hormone (GH), c-kit receptor tyrosine kinase (c-kit), glutamate receptor, metabotropic 5 (mGluR5), nerve growth factor-beta (NGF-beta), Jun oncogene (c-Jun), transmembrane receptor Unc5H1 (UNC5H1), and transmembrane receptor Unc5H2 (UNC5H2). All these genes, except for GH, were downregulated in response to trace conditioning. GH was upregulated; therefore, we also validated the downregulation of the GH inhibitor, somatostatin (SST), even though it just failed to meet criteria on the arrays. By during situ hybridization, GH was expressed throughout the cell layers of the hippocampus in response to trace conditioning. None of the genes regulated in trace eyeblink conditioning were similarly affected by delay conditioning, a task that does not require the hippocampus. These findings demonstrate that transcriptional profiling can exhibit a repertoire of genes sensitive to the formation of hippocampal-dependent associative memories.

Non-NASA Center↗