Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “PARALLEL PROCESSING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Integrating ytopt and libEnsemble to autotune OpenMC

Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.

Autotuning↗

High Throughput Source-less Plasma Deposition of Structured Silicon Anodes for Lithium-Ion Batteries

Amprius developed a manufacturing solution for silicon nanowire anode that relies on an inexpensive, high throughput, and high gas precursor utilization plasma deposition method that uses the anode foils as electrodes for plasma generation. The capacitively couple plasma (CCP) method is used in semiconductor and photovoltaic industry and Amprius modified existing high throughput equipment to use anode foils and to deposit amorphous silicon. The equipment was installed ahead of the program at Amprius site. The rest of the tasks included foil handling and process development. The equipment passed site acceptance tests (SAT) and the process parameter mapping was completed, indicating that the target process window limits produce output materials at the rate and with yield and specifications that meet the manufacturing target criteria. Amprius has hired supporting personnel to optimize processes and run the equipment. A parallel task verified the baseline performance of the silicon anode material, to be used as reference for the new manufacturing method.

25 ENERGY STORAGE↗

Autoencoders on FPGAs for real-time, unsupervised new physics detection at 40 MHz at the Large Hadron Collider

In this paper, we show how to adapt and deploy anomaly detection algorithms based on deep autoencoders, for the unsupervised detection of new physics signatures in the extremely challenging environment of a real-time event selection system at the Large Hadron Collider (LHC). We demonstrate that new physics signatures can be enhanced by three orders of magnitude, while staying within the strict latency and resource constraints of a typical LHC event filtering system. This would allow for collecting datasets potentially enriched with high-purity contributions from new physics processes. Through per-layer, highly parallel implementations of network layers, support for autoencoder-specific losses on FPGAs and latent space based inference, we demonstrate that anomaly detection can be performed in as little as $80\,$ns using less than 3% of the logic resources in the Xilinx Virtex VU9P FPGA. Opening the way to real-life applications of this idea during the next data-taking campaign of the LHC.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

System and method for sub micron additive manufacturing

An apparatus is disclosed for performing an additive manufacturing operation to form a structure by processing a photopolymer resist material. The apparatus may incorporate a laser for generating a laser beam, and a tunable mask for receiving the laser beam which has an optically dispersive element. The mask splits the laser beam into a plurality of emergent beams each having a subplurality of beamlets of varying or identical intensity, with each beamlet emerging from a unique subsection of illuminated regions of the mask. A collimator collimates at least one of the emergent beams to form a collimated beam. One or more focusing elements focuses the collimated beam into a focused beam which is projected as a focused image plane on or within the resist material. The focused beam simultaneously illuminates a layer of the resist material to process an entire layer in a parallel fashion.

Saha, Sourabh Kumar↗

Development of a half-meter scale Traveling-Wave (TW) SRF cavity

Traveling-wave technology can push the accelerator field gradient of niobium SRF cavity to 70MV/m or higher beyond the fundamental limit of 50~60MV/m in Standing-Wave regime. The 1st demonstration of TW resonance excitation in a proof-of-principle 3-cell SRF cavity in 2K liquid helium was successfully carried out at Fermilab in collaboration with Euclid Techlabs. In parallel with that, the RF design process of 0.5~1 meter scale TW cavity was begun at Fermilab for advancing TW technologies necessary more for future accelerator-scale one. Considering the physical dimensions of existing SRF facilities (for fabrication, processing, and cryogenic testing) and the lessons learned from the 3-cell, Fermilab has proposed a preliminary RF design of a half-meter scale TW SRF cavity. It consists of a 7-cell structure and a power feedback waveguide (WG) loop with new RF configurations to control TW resonance. Here we report a preliminary RF design, development plans, and activities.

43 PARTICLE ACCELERATORS↗

Data-flow parallelism for high-energy and nuclear physics frameworks

The processing tasks of an event-processing workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this talk, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. Building on the Meld project as presented at CHEP2023, we demonstrate that all common processing idioms supported by current frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Grain structure and texture selection regimes in metal powder bed fusion

Additive manufacturing (AM) offers opportunities to produce complex part geometries not possible with conventional processing and in some cases even improve part performance. However, adoption has been slowed by difficulties assessing microstructure variability and there is no straightforward approach to relate processing to grain structure characteristics. In this study, datasets from AdditiveFOAM heat transport simulations of laser powder bed fusion (LPBF) are used to drive ExaCA simulations of grain structure. The GPU utilization of ExaCA and an algorithmic update for modeling melt pool overlap region solidification enabled rapid and parallel simulation across a wider range of process conditions than previously explored with cellular automata-based solidification models. A texture selection angle $θ_s$ is defined based on melt pool overlap geometry, and the range of $θ_s$ over which a commonly observed texture transition occurs in characterized AM builds was well-reproduced by ExaCA simulations over a wide range of melt pool shape, hatch spacing, and layer height. ExaCA simulations with 90 degree rotation of the scan direction on every other layer reproduced a number of trends from the AM literature including grain refinement, the dominance of layers with larger melt pools on the final grain structure, and the weakening or strengthening of texture depending on odd and even layer melt pool overlap geometry. EBSD data from a benchmark AM part is used to validate the simulated mechanism of a layer rotation-induced texture strengthening effect. Importantly, these results expand the understanding of the mechanisms for texture selection in alloys with cubic crystal symmetry and offer an approach to easily evaluate processing conditions. With this new understanding, these modeling tools will enable anticipation of previously unexpected variations in grain structure and target specific microstructures and properties.

36 MATERIALS SCIENCE↗

Compressed basis GMRES on high-performance graphics processing units

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

97 MATHEMATICS AND COMPUTING↗

Case Study of Using Kokkos and SYCLs Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs

Six of the top ten supercomputers in the TOP500 list from June 2021 rely on NVIDIA GPUs to achieve their peak compute bandwidth. With the announcement of Aurora, Frontier, and El Capitan, Intel and AMD have also entered the domain of providing GPUs for scientific computing. A consequence of the increased diversity in the GPU landscape is the emergence of portable programming models such as Kokkos, SYCL, OpenCL, and OpenMP, which allow application developers to maintain a single-source code across a diverse range of hardware architectures. While the portable frameworks try to optimize the compute resource usage on a given architecture, it is the programmers responsibility to expose parallelism in an application that can take advantage of thousands of processing elements available on GPUs. In this paper, we introduce a GPU-friendly parallel implementation of Milc-Dslash that exposes multiple hierarchies of parallelism in the algorithm. Milc-Dslash was designed to serve as a benchmark with highly optimized matrix-vector multiplications to measure the resource utilization on the GPU systems. The parallel hierarchies in the Milc-Dslash algorithm are mapped onto a target hardware using Kokkos and SYCL programming models. We present the performance achieved by Kokkos and SYCL implementations of Milc-Dslash on NVIDIA A100 GPU, AMD MI100 GPU, and Intel Gen9 GPU. Additionally, we compare the Kokkos and SYCL performances with those obtained from the versions written in CUDA and HIP programming models on NVIDIA A100 GPU and AMD MI100 GPU, respectively.

Dufek, Amanda S↗

Serial2Parallel

In the era of machine learning, we often need to run the same code/script many times with little or no variations (e.g., performance evaluation, data preprocessing, data generation, hyperparameter tuning, etc.). It is not a problem when you just need to do that a few times, but when the number of repetitions becomes very large, it can be a daunting task. The code “Serial2Parallel” provides an easy way for users to be able to run many numbers of any serial code/scripts in a parallel manner across multiple nodes in an message passing interface (MPI) cluster. The code includes the server program that deals with task pool management and client program that processes task. The server gets the tasks ready and waits for clients' connections. The client code pulls tasks from the server and processes them. The client code will run in parallel.

Sangkeun, MattLee↗

Ensemble Simulation Techniques and Fast Randomized Algorithms

The major goals of the project were to develop and analyze new ensemble simulation techniques, including trajectory stratification and preconditioned MCMC techniques, as well as develop fast numerical linear algebra techniques closely related to ensemble simulation ideas. The trajectory stratification techniques involve simulating in parallel short trajectory fragments of a Markov process confined to a specific region of space‐time and then patching together the statistics gathered to assemble estimates of very general dynamical properties. We have also developed this approach for rare event simulation and extended the techniques to applications requiring a more general framework (such as electronic structure calculations). The preconditioned MCMC techniques involve simulating multiple Markov chains in parallel and then using information from the ensemble to speed the mixing of each individual chain. The fast randomized linear algebra methods are motivated by the diffusion Monte Carlo technique, but are applicable to finding the dominant eigenvalue of (almost) general matrices. For most non‐negative matrices, the schemes result in an error (compared to the power method) that is constant in the dimension of the problem. For more general matrices, we see a very clear sublinear cost trend in computational tests.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Singlet fission for quantum information and quantum computing: the parallel JDE model

Abstract Singlet fission is a photoconversion process that generates a doubly excited, maximally spin entangled pair state. This state has applications to quantum information and computing that are only beginning to be realized. In this article, we construct and analyze a spin-exciton hamiltonian to describe the dynamics of the two-triplet state. We find the selection rules that connect the doubly excited, spin-singlet state to the manifold of quintet states and comment on the mechanism and conditions for the transition into formally independent triplets. For adjacent dimers that are oriented and immobilized in an inert host, singlet fission can be strongly state-selective. We make predictions for electron paramagnetic resonance experiments and analyze experimental data from recent literature. Our results give conditions for which magnetic resonance pulses can drive transitions between optically polarized magnetic sublevels of the two-exciton states, making it possible to realize quantum gates at room temperature in these systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improved bacterial recombineering by parallelized protein discovery

Exploiting bacteriophage-derived homologous recombination processes has offered precise, multiplex editing of microbial genomes and the construction of billions of customized genetic variants in a single day. The techniques that enable this, multiplex automated genome engineering (MAGE) and directed evolution with random genomic mutations (DIvERGE), are however, currently limited to a handful of microorganisms for which single-stranded DNA-annealing proteins (SSAPs) that promote efficient recombineering have been identified. Thus, to enable genome-scale engineering in new hosts, efficient SSAPs must first be found. Here we present a high-throughput method for SSAP discovery that we call “serial enrichment for efficient recombineering” (SEER). By performing SEER in Escherichia coli to screen hundreds of putative SSAPs, we identify highly active variants PapRecT and CspRecT. CspRecT increases the efficiency of single-locus editing to as high as 50% and improves multiplex editing by 5- to 10-fold in E. coli , while PapRecT enables efficient recombineering in Pseudomonas aeruginosa , a concerning human pathogen. CspRecT and PapRecT are also active in other, clinically and biotechnologically relevant enterobacteria. We envision that the deployment of SEER in new species will pave the way toward pooled interrogation of genotype-to-phenotype relationships in previously intractable bacteria.

60 APPLIED LIFE SCIENCES↗

HYPPO: A Surrogate-Based Multi-Level Parallelism Tool for Hyperparameter Optimization

We present a new software, HYPPO, that enables the automatic tuning of hyperparameters of various deep learning (DL) models. Unlike other hyperparameter optimization (HPO) methods, HYPPO uses adaptive surrogate models and directly accounts for uncertainty in model predictions to find accurate and reliable models that make robust predictions. Using asynchronous nested parallelism, we are able to significantly alleviate the computational burden of training complex architectures and quantifying the uncertainty. HYPPO is implemented in Python and can be used with both TensorFlow and PyTorch libraries. We demonstrate various software features on time-series prediction and image classification problems as well as a scientific application in computed tomography image reconstruction. Finally, we show that (1) we can reduce by an order of magnitude the number of evaluations necessary to find the most optimal region in the hyperparameter space and (2) we can reduce by two orders of magnitude the throughput for such HPO process to complete.

adaptation models↗

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Chatter Stability of Machining Operations

Here, this paper reviews the dynamics of machining and chatter stability research since the first stability laws were introduced by Tlusty and Tobias in the 1950s. The paper aims to introduce the fundamentals of dynamic machining and chatter stability, as well as the state of the art and research challenges, to readers who are new to the area. First, the unified dynamic models of mode coupling and regenerative chatter are introduced. The chatter stability laws in both the frequency and time domains are presented. The dynamic models of intermittent cutting, such as milling, are presented and their stability solutions are derived by considering the time-periodic behavior. The complexities contributed by highly intermittent cutting, which leads to additional stability pockets, and the contribution of the tool's flank face to process damping are explained. The stability of parallel machining operations is explained. The design of variable pitch and serrated cutting tools to suppress chatter is presented. The paper concludes with current challenges in chatter stability of machining which remains to be the main obstacle in increasing the productivity and quality of manufactured parts.

42 ENGINEERING↗

Data-flow parallelism for high-energy and nuclear physics computing frameworks

The processing tasks of a scientific workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP computing framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this session, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. After introducing the physics DUNE intends to explore, we will show that all common processing idioms supported by current HENP frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

43 PARTICLE ACCELERATORS↗