Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Flexible and Modular Simultaneous Modeling of Flow and Reactive Transport in Rivers and Hyporheic Zones

Investigations of coupled multiphysics processes in rivers and hyporheic zones have extensively used numerical models. Most existing models use a sequential, one-way coupling between the surface and subsurface domains. Such one-way coupling potentially introduces error. To overcome this, a fully coupled model, hyporheicFoam, was developed using the open-source computational platform OpenFOAM. It captures the coupled flow and multicomponent reactive transport processes within both surface and subsurface domains and across their interface. The coupling between two domains is implemented by mapping conservative flux boundary conditions at the interface through an iterative algorithm. Reactive transport is enabled by specifying a reaction network. To start, we have implemented reaction kinetics following the double Monod-type model with inhibition. The model capability is illustrated through modeling of both conservative and reactive hyporheic flow and transport through dune bedforms. With the novel coupled model, it is now possible to quantify reactions wherein the reactants and products are constantly exchanging between domains and have feedbacks. hyporheicFoam can simulate large, three-dimensional cases owing to the computational flexibility and power offered by the code structure and parallel design of OpenFOAM.

58 GEOSCIENCES↗

Mind the Gap: Building Simulation in the Architectural Design Studio

Building modelling and simulation approaches are increasingly being utilized in architectural design studios to guide and inform the design process and offer evidence-based feedback on proposed building performance. The development of intuitive and simplified simulation interfaces has greatly contributed to achieving this integration. One aspect that is often overlooked is the workflow that governs and regulates integrated design, which can have significant impacts on final design outcomes. Currently there are numerous software packages available for building performance simulations. This makes it challenging to select an appropriate tool that provides accurate results yet allows a designer to make informed architectural decisions with a designer-friendly interface. Furthermore, workflows to incorporate simulations into the design process proved to highly impact student’s project design integration. Yet, it is not clear what type of workflows are successful to achieve this goal, under what conditions, and/or for which building and site typologies. This paper addresses these issues by first reviewing three different workflows for integrating building performance simulation processes and highlighting their strengths and weaknesses. Second, a comparative case study approach was employed to test three of the most common workflows in three different integrated design architectural studios at the senior and vertical studio levels as well as in two courses that run parallel and complementary to the design studios. The workflows, processes, and the resultant student projects were further analyzed based on criteria for better integrated design and architectural excellence as outlined in the American Institute of Architects Committee on the Environment (AIA COTE) Top 10 Student’s Competition. Third, to situate the pedagogical case studies’ results within a larger context, a survey of the AIA COTE Top 10 student competition award recipients over the last five years was conducted. The results are summarized in a pedagogical framework that outlines best strategies of the type of workflows, software, design process used, methods to achieve desired interaction between design process and analytical feedback, and metrics for educators to evaluate the success of this integration and their learning outcomes in the design studio. The goal is to help bridge the gap between the building design and simulation within the design studio’s creative process for more integrated design outcomes.

Elzeyadi, Ihab↗

PETSc/TAO Users Manual (Rev. 3.19)

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for the implementation of large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and opti mization that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started). PETSc provides many of the mechanisms needed within parallel application codes, such as parallel matrix and vector assembly routines. The library is organized hierarchically, enabling users to employ the level of abstraction that is most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users. PETSc is a sophisticated set of software tools; as such, for some users it initially has a much steeper learning curve than packages such as MATLAB or a simple subroutine library. In particular, for individuals without some computer science background, experience programming in C, C++, python, or Fortran and experience using a debugger such as gdb or lldb, it may require a significant amount of time to take full advantage of the features that enable efficient software use. However, the power of the PETSc design and the algorithms it incorporates may make the efficient implementation of many application codes simpler than “rolling them” yourself. For many tasks a package such as MATLAB is often the best tool; PETSc is not intended for the classes of problems for which effective MATLAB code can be written. There are several packages, built on PETSc, that may satisfy your needs without requiring directly using PETSc. We recommend reviewing these packages functionality before starting to code directly with PETSc. PETSc can be used to provide a “MPI parallel linear solver” in an otherwise sequential, or OpenMP parallel code. This approach cannot provide extremely large improvements in the application time by utilizing large numbers of MPI processes but can still improve the performance. Certainly all parts of a previously sequential code need not be parallelized but the matrix generation portion must be parallelized to expect true scalability to large numbers of MPI processes. See PCMPI for details on how to utilize the PETSc MPI linear solver server. Since PETSc is under continued development, small changes in usage and calling sequences of routines will occur. PETSc has been supported for twenty-five years; see mailing list information on our website for information on contacting support.

97 MATHEMATICS AND COMPUTING↗

Columnar-to-equiaxed transition in a laser scan for metal additive manufacturing

In laser powder bed fusion additive manufacturing (LPBFAM), different solidification conditions, e.g., thermal gradient and cooling rate, can be achieved by controlling the process parameters, such as laser power and laser speed. Tailoring the behaviour of the columnar to equiaxed transition (CET) of the printed alloy during fabrication can facilitate the production of highly customized microstructures. In this study, effective analytical solutions for both thermal conduction and solidification are employed to model solidifying melt pools. Microstructure textures and solidification conditions are evaluated for numerous combinations of laser power and laser speed under bead-on-plate conditions. This analytical-based high-throughput tool was demonstrated to select specific process parameters that lead to desired microstructures. Two selected process conditions were examined in detail by a highly parallelized microstructural solidification model to reveal both nucleation and grain growth. Both numerical solutions agree well with experiments that are performed based on bead-on-plate conditions, indicating that these numerical models aid evaluation of the nucleation parameters, providing insights for controlling CET during the LPBFAM processing.

Lang, Yuan↗

Predicting synthetic mRNA stability using massively parallel kinetic measurements, biophysical modeling, and machine learning

Abstract mRNA degradation is a central process that affects all gene expression levels, though it remains challenging to predict the stability of a mRNA from its sequence, due to the many coupled interactions that control degradation rate. Here, we carried out massively parallel kinetic decay measurements on over 50,000 bacterial mRNAs, using a learn-by-design approach to develop and validate a predictive sequence-to-function model of mRNA stability. mRNAs were designed to systematically vary translation rates, secondary structures, sequence compositions, G-quadruplexes, i-motifs, and RppH activity, resulting in mRNA half-lives from about 20 seconds to 20 minutes. We combined biophysical models and machine learning to develop steady-state and kinetic decay models of mRNA stability with high accuracy and generalizability, utilizing transcription rate models to identify mRNA isoforms and translation rate models to calculate ribosome protection. Overall, the developed model quantifies the key interactions that collectively control mRNA stability in bacterial operons and predicts how changing mRNA sequence alters mRNA stability, which is important when studying and engineering bacterial genetic systems.

Cetnar, Daniel P.↗

ETG turbulence in a tokamak pedestal

This paper explores the fundamental characteristics of electron-temperature-gradient (ETG)-driven turbulence in the tokamak pedestal. The extreme gradients in the pedestal produce linear instabilities and nonlinear turbulence that are distinct from the corresponding ETG phenomenology in the core plasma. The linear system exhibits multiple (greater than ten) unstable eigenmodes at each perpendicular wave vector, representing different toroidal and slab branches of the ETG instability. Proper orthogonal decomposition of the nonlinear fluctuations reveals no clear one-to-one correspondence between the linear and nonlinear modes for most wave vectors. Moreover, nonlinear frequencies deviate strongly from those of the linear instabilities, with spectra peaking at positive frequencies, which is opposite the sign of the ETG instability. The picture that emerges is one in which the linear properties are preserved only in a narrow range of k-space. Outside this range, nonlinear processes produce strong deviations from both the linear frequencies and eigenmode structures. This is interpreted in the context of critical balance, which enforces alignment between the parallel scales and fluctuation frequencies. We also investigate the nonlinear saturation processes. We observe a direct energy cascade from the injection scale to smaller scales in both perpendicular directions. However, in the bi-normal direction, there is also nonlocal inverse energy transfer to larger scales. Neither streamers nor zonal flows dominate the saturation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

TeraChem: A graphical processing unit-accelerated electronic structure package for large-scale ab initio molecular dynamics

TeraChem was born in 2008 with the goal of providing fast on-the-fly electronic structure calculations to facilitate ab initio molecular dynamics studies of large biochemical systems such as photoswitchable proteins and multichromophoric antenna complexes. Originally developed for videogaming applications, graphics processing units (GPUs) offered a low-cost parallel computer architecture that became more accessible for general-purpose GPU computing with the release of CUDA in 2007. The evaluation of the electron repulsion integrals (ERIs) is a major bottleneck in electronic structure codes and provides an attractive target for acceleration on GPUs. Thus, highly efficient routines for evaluation of and contractions between the ERIs and density matrices were implemented in TeraChem. Here, electronic structure methods were developed and implemented to leverage these integral contraction routines, resulting in the first quantum chemistry package designed from the ground up for GPUs. This GPU acceleration makes TeraChem capable of performing large-scale ground and excited state calculations in the gas and condensed phase. Today, TeraChem's speed forms the basis for a suite of quantum chemistry applications, including optimization and dynamics of proteins, automated and interactive chemical discovery tools, and large-scale nonadiabatic dynamics simulations.

74 ATOMIC AND MOLECULAR PHYSICS↗

Scalable Graph Analytics and HPC Operational Enhancement: Parallel Computing and ML/DL Innovations

Parallel computing plays a pivotal role in the efficient processing of large-scale graphs. Complex network analysis stands as a capti- vating research frontier, holding promise across diverse scientific domains such as sociology, biology, online media, and recommenda- tion systems. In this era, Machine Learning (ML) and Deep Learning (DL) have emerged as indispensable tools, underpinning remarkable technological achievements. Within this dynamic landscape, my research revolves around advancing parallel algorithms tailored for large-scale graph operations. To achieve this, I harness the power of cutting-edge technologies including OpenMP, MPI, HIP, and CUDA, on the High-Performance Computing (HPC) platforms to unlock optimal performance. I also apply ML/DL techniques to HPC operational data, to streamline the monitoring and maintenance of supercomputers, alleviating the complexities associated with their upkeep and enhancing user support. My research echoes the syn- ergy between parallel computing, large-scale graph analysis, and ML/DL, improving computational efficiency and user experience.

Sattar, Naw Safrin↗

Distributed out-of-memory NMF on CPU/GPU architectures

We propose an efficient distributed out-of-memory implementation of the non-negative matrix factorization (NMF) algorithm for heterogeneous high-performance-computing systems. The proposed implementation is based on prior work on NMFk, which can perform automatic model selection and extract latent variables and patterns from data. In this work, we extend NMFk by adding support for dense and sparse matrix operation on multi-node, multi-GPU systems. The resulting algorithm is optimized for out-of-memory problems where the memory required to factorize a given matrix is greater than the available GPU memory. Memory complexity is reduced by batching/tiling strategies, and sparse and dense matrix operations are significantly accelerated with GPU cores (or tensor cores when available). Input/output latency associated with batch copies between host and device is hidden using CUDA streams to overlap data transfers and compute asynchronously, and latency associated with collective communications (both intra-node and inter-node) is reduced using optimized NVIDIA Collective Communication Library (NCCL) based communicators. Benchmark results show significant improvement, from 32X to 76x speedup, with the new implementation using GPUs over the CPU-based NMFk. Good weak scaling was demonstrated on up to 4096 multi-GPU cluster nodes with approximately 25,000 GPUs when decomposing a dense 340 Terabyte-size matrix and an 11 Exabyte-size sparse matrix of density 10 -6 .

97 MATHEMATICS AND COMPUTING↗

Computing the QRPA level density with the finite amplitude method

Here, we describe a new algorithm to calculate the vibrational nuclear level density of an atomic nucleus. Fictitious perturbation operators that probe the response of the system are generated by drawing their matrix elements from some probability distribution function. We use the Finite Amplitude Method to explicitly compute the response for each such sample. With the help of the Kernel Polynomial Method, we build an estimator of the vibrational level density and provide the upper bound of the relative error in the limit of infinitely many random samples. The new algorithm can give accurate estimates of the vibrational level density. Since it is based on drawing multiple samples of perturbation operators, its computational implementation is naturally parallel and scales like the number of available processing units.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Livermore tomography tools: Accurate, fast, and flexible software for tomographic science

Livermore Tomography Tools (LTT) is a customizable scientific software package that enables a broad range of research and development efforts into computed tomography (CT). Here, it was developed to process x-ray and neutron CT data accurately and rapidly from raw detector counts to reconstructed volumes with the flexibility to handle many special cases. LTT fulfills long-term CT software goals to provide quantitatively accurate results reported in physical units (e.g., mm -1 or cm -1 ) while exploiting all available computational advantages to maximize speed. Written in C/C++ with support for multiple CPUs and GPUs, LTT runs on many computing platforms (Linux/Unix, Windows, and Mac; laptops to supercomputers). As a result, LTT can:process data acquired from various custom-built and commercially available CT scanners, model and simulate x-ray and neutron interactions to encourage algorithm prototyping, and allow for rapid insertion of the latest algorithms.We describe LTT’s software architecture, user interfaces, and its 88 algorithms (as of this writing) for pre-processing, reconstruction, post-processing, and simulation that support many scanner geometries (parallel-, fan-, cone-beam, and custom). Several applications are presented that illustrate LTT’s accuracy, speed, and flexibility relative to other solutions.

36 MATERIALS SCIENCE↗

A High-Performance Discrete-Element Framework for Simulating Flow and Jamming of Moisture Bearing Biomass Feedstocks

We developed and verified a high-performance open-source discrete element method (DEM) solver with simultaneously-supported feedstock-specific interaction models, including bonded-sphere, liquid bridge, cohesion, and non-linear contact models. Our solver uses parallel data structures on hybrid central and graphics processing unit (CPU/GPU) architectures, with favorable strong scaling performance observed for large problem sizes comprised of (100 M particles), and 4X single-node GPU speedup. The particles for corn stover feedstock were conceptualized and calibrated based on experimental measurements and results. Sensitivity analyses demonstrate that the mass flow rate from a wedge hopper is governed primarily by moisture content, friction coefficient, and cohesion energy density. The model is used to reproduce experimentally observed hopper jamming results, highlighting that the experimental no-flow trends can only be achieved by using non-spherical particles, liquid bridge and cohesion models, highlighting the importance of using concurrent feedstock specialized models for the effective representation of biomass material handling problems.

bioenergy↗

Electron energization during merging of self-magnetized, high-beta, laser-produced plasmas

Electron energization during merging of magnetized plasmas is studied using the OMEGA and OMEGA EP laser facilities by colliding two plasma plumes, each containing a Biermann-battery self-generated magnetic field. Two neighbouring plasma plumes are produced by intense laser beams, and the anti-parallel Biermann fields merge and reconnect in the process of the plumes’ expansion and collision. To isolate the merging as an acceleration source, the electron energy spectra obtained from two-plume collision shots are compared with the spectra from single-plume shots. Single-plume shots exhibit an energized electron tail with energies up to ${\sim }250\ \textrm {keV}$ . The electrons in merging experiments are additionally accelerated by ${\sim }50\text {--}100$ keV compared to single-plume shots.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enabling high-throughput enzyme discovery and engineering with a low-cost, robot-assisted pipeline

Abstract As genomic databases expand and artificial intelligence tools advance, there is a growing demand for efficient characterization of large numbers of proteins. To this end, here we describe a generalizable pipeline for high-throughput protein purification using small-scale expression in E. coli and an affordable liquid-handling robot. This low-cost platform enables the purification of 96 proteins in parallel with minimal waste and is scalable for processing hundreds of proteins weekly per user. We demonstrate the performance of this method with the expression and purification of the leading poly(ethylene terephthalate) hydrolases reported in the literature. Replicate experiments demonstrated reproducibility and enzyme purity and yields (up to 400 µg) sufficient for comprehensive analyses of both thermostability and activity, generating a standardized benchmark dataset for comparing these plastic-degrading enzymes. The cost-effectiveness and ease of implementation of this platform render it broadly applicable to diverse protein characterization challenges in the biological sciences.

36 MATERIALS SCIENCE↗

The 4D Camera: An 87 kHz Direct Electron Detector for Scanning/Transmission Electron Microscopy

We describe the development, operation, and application of the 4D Camera—a 576 by 576 pixel active pixel sensor for scanning/transmission electron microscopy which operates at 87,000 Hz. The detector generates data at ~480 Gbit/s which is captured by dedicated receiver computers with a parallelized software infrastructure that has been implemented to process the resulting 10–700 Gigabyte-sized raw datasets. The back illuminated detector provides the ability to detect single electron events at accelerating voltages from 30 to 300 kV. Through electron counting, the resulting sparse data sets are reduced in size by 10--300× compared to the raw data, and open-source sparsity-based processing algorithms offer rapid data analysis. The high frame rate allows for large and complex scanning diffraction experiments to be accomplished with typical scanning transmission electron microscopy scanning parameters.

47 OTHER INSTRUMENTATION↗

Optimization Plugin Library

The Optimization Plugin library ("op") is a lightweight general optimization solver interface. The primary purpose of op is to simplify the process of integrating different optimization solvers (serial or parallel) with scalable parallel physics engines. By design it has several features that help make this a reality. The core abstraction interface was developed to encompass a large class of optimization problems in an optimizer-agnostic way. This enables us to describe the optimization problem once and then use a variety of supported "op" optimizers with ideally no code-changes. The abstraction interface is made up of lightweight wrappers that make it easy to integrate with existing simulation codes. This makes integration less intrusive and should minimize changes to existing physics codes. The "op" interface includes an assortment of utility methods that help specify parallel communication patterns as well as methods to convert from optimization-specific interfaces to the more general "op" interface. Lastly a dynamic library linking interface is provided to allow for use of proprietary optimization engines without explicit reference in the source code, along with standard shared library interfaces for opensource engines.

Jekel, CharlesF↗

Scaled up Process Report – Apparatus and Model

Advanced voloxidation with NO 2 is a proposed process for used nuclear fuel head-end reprocessing scheme that converts UO 2 to higher oxides, and it also converts partitioning volatile fission products into the gas phase, thus facilitating fuel dissolution and actinide recovery. NO 2 voloxidation is being studied on small batches of UO 2 Simfuel, but the real test of process feasibility will be when it is scaled up to work with >100 g of irradiated material. This report discusses the aspects of scale-up that must be considered for NO 2 voloxidation, including development of a stirred reactor to promote agitation of the mixture during processing, online process monitoring, and automation controls. Brief details on parallel efforts are also provided in this report, including (a) demonstration of iodine release from Simfuel made by Spark Plasma Sintering, and (b) development of an order-of-magnitude scale-up to react 100 g of UO 2 Simfuel in a metal reactor.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Model for Pair Production Limit Cycles in Pulsar Magnetospheres

Abstract It was recently proposed that the electric field oscillation as a result of self-consistent e ± pair production may be the source of coherent radio emission from pulsars. Direct particle-in-cell simulations of this process have shown that the screening of the parallel electric field by this pair cascade manifests as a limit cycle, as the parallel electric field is recurrently induced when pairs produced in the cascade escape from the gap region. In this work, we develop a simplified time-dependent kinetic model of e ± pair cascades in pulsar magnetospheres that can reproduce the limit-cycle behavior of pair production and electric field screening. This model includes the effects of a magnetospheric current, the escape of e ± , as well as the dynamic dependence of pair production rate on the plasma density and energy. Using this simple theoretical model, we show that the power spectrum of electric field oscillations averaged over many limit cycles is compatible with the observed pulsar radio spectrum.

Astronomy & Astrophysics↗