Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47

A Parallel Computing Infrastructure for Building Energy Simulation

In order to study grid-interactive efficient buildings, Pacific Northwest National Laboratories (PNNL) needs an infrastructure for urban-scale building energy modeling. Such an infrastructure should be fast, scalable, and easy-to-use. Given a set of data from the Energy Information Administration’s Commercial Building Energy Consumption Survey (CBECS) and tool to translate survey data into simulation inputs, this project aimed to conduct the simulation of the entire dataset in parallel. Before running the simulations, the necessary software was bundled into a container for use on the PNNL supercomputing network. Then, the parallel simulation workflow was designed using GNU Make, a file creation software, and submitted to a supercomputing partition which could run hundreds of simulations simultaneously. The EnergyPlus simulations output hourly electric meter data for each CBECS sample, which represents the electricity consumption of similar commercial buildings across the United States. Analyzing and visualizing the meter data is important to the future of the work, and this project wrote code to make common analysis methods simple, fast, and accessible. Moving forwards, the model will need to be expanded to include data from other sources and its accuracy will need to be improved and eventually validated.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Parallel Time Integration: An Approaching Paradigm Shift for Scientific Computing

This note argues that parallel-in-time methods will be necessary for doing high-fidelity time-dependent simulations in the future. A “proof” is given to support the argument and to provide a framework for debate. The effect of a parallel-in-time paradigm on scientific computing practice is also discussed.

97 MATHEMATICS AND COMPUTING↗

Parallel-in-Time Methods for Method-of-Lines Discretizations of Nonlinear Hyperbolic PDEs and Systems (Final Report)

The work for the subcontract is situated in the area of parallel-in-time integration for hyperbolic partial differential equations (PDEs). Parallel-in-time integration is an active area of research due to its ability to enable faster numerical simulations for applications throughout many areas of science. The work in this subcontract builds on a variety of results that were obtained, as part of the work performed for Subcontract No. B648355, for the Multigrid Reduction-in-Time (MGRIT) method from [1] applied to hyperbolic PDEs. This subcontract extends these results further to more efficient methods and to the case of method-of-lines discretizations for nonlinear hyperbolic PDES and systems of PDEs. The following is a summary of the research performed and results achieved during milestone periods 1, 2 and 3 by the PI (Hans De Sterck) and Postdoctoral Research Associate (Oliver Krzysik), for required tasks 1-4 (as listed in the Statement of Work): Research over the previous year has been split into three main projects: (i) solution of acoustic equation system; (ii) solution of nonlinear scalar hyperbolic PDEs; (iii) solution of nonlinear hyperbolic systems of PDEs.

97 MATHEMATICS AND COMPUTING↗

Design and Analysis of 10” Parallel Plate Relief Device

All Cryomodule (CM) and Cryogenic Distribution System (CDS) relieving into Helium Low Pressure (LP) return header, which is connected to compressor suction so, helium can be preserved during small flow relieving event and recirculated to system. However, during worst case scenario, Helium LP header requires a parallel plate relief device to relieve excess pressure from header. To complete the CDS Warm piping header, a new design for a 10 parallel plate relief device is necessary to relieve outside of the tunnel into atmosphere.

Chicas, Kelly↗

Non-Intrusive Parallel-in-Time Solvers for Partial Differential Equations (Final Report)

Many time-dependent problems and simulations are often modeled using Partial Differential Equations. Traditional modeling approaches that use sequential time-stepping are reaching a bottleneck in optimizing efficiency. The Center of Applied Science and Computing at Lawrence Livermore National Laboratory extensively works on parallelizing these algorithms to leverage the increasing computational power from the growing number of processors in computer hardware. In particular, they aim to design non-intrusive algorithms that can generalize to a variety of problems and sizes without requiring additional information from or modifications on the original problems. Multigrid Reduction in Time (MGRIT) is a parallel-in-time algorithm that is designed to be non-intrusive. This project focuses on increasing the efficiency of MGRIT by approximating the coarse-grid operator using machine learning approaches as a means to find the most non-intrusive, or general, solution.

97 MATHEMATICS AND COMPUTING↗

Portable Parallel Algorithms and Frameworks for Exascale Graph Analytics

Graphs (or networks) are a tool used to model the interactions among various entities. Efficiently processing large graphs has recently attracted significant attention due to the applications of graphs in various domains, such as biology, chemistry, and cyber-security. Analyzing the structure and properties of these graphs is an important component of many scientific computing pipelines. With the explosion in the volume of data, graphs have become very large and can contain hundreds of billions of vertices and trillions of edges. Therefore, it is crucial to develop high-performance methods to enable graph analysis to be done quickly and energy-efficiently. Furthermore, these solutions should be highly parallel in order to take advantage of modern parallel machines. However, designing efficient solutions is not enough. With the wide variety of computing environments available, each with different programmability and performance characteristics, it is necessary to develop solutions that are portable in terms of both performance (i.e., provide theoretical guarantees) and programmability (i.e., provide high level abstractions).

97 MATHEMATICS AND COMPUTING↗

Applying Time-Parallelization to Turbulent Flows

Parallelization of the temporal domain is explored for the solution of turbulent flows. Multigrid reduction-in-time (MGRIT) is used to advance the large-scale fluid dynamics in time sequentially on the coarsest space-time grid but propagate the information in time parallel on all other levels. The goal of this process is to accurately and efficiently resolve the coarse-scale turbulence structure and use that to drive the fine-scales of the turbulent flow. The extra forcing from nonlinear multigrid facilitates the coupling and interaction between fine and coarse scales, through which the multiscale nonlinear physics is properly captured. Adaptive mesh refinement is employed to finely resolve only the regions with strong gradients, which provides further computational efficiency. The underlying computational fluid dynamics solver is a fourth-order finite-volume scheme with the standard 4-stage Runge-Kutta method. An advanced approach is devised and implemented to enable MGRIT to solve highly turbulent flows successfully. Furthermore, the method is applied to solve a Taylor-Green vortex problem and a doubleshear-layer turbulent mixing flow. Results are promising, validating that MGRIT with the filtering approach has the potential to efficiently solve general turbulent flows.

Computational Fluid Dynamics↗

Chromatin Changes in Phytochrome Interacting Factor-Regulated Genes Parallel Their Rapid Transcriptional Response to Light

As sessile organisms, plants must adapt to a changing environment, sensing variations in resource availability and modifying their development in response. Light is one of the most important resources for plants, and its perception by sensory photoreceptors (e.g., phytochromes) and subsequent transduction into long-term transcriptional reprogramming have been well characterized. Chromatin changes have been shown to be involved in photomorphogenesis. However, the initial short-term transcriptional changes produced by light and what factors enable these rapid changes are not well studied. Here, we define rapidly light-responsive, Phytochrome Interacting Factor (PIF) direct-target genes (LRP-DTGs). We found that a majority of these genes also show rapid changes in Histone 3 Lysine-9 acetylation (H3K9ac) in response to the light signal. Detailed time-course analysis of transcript and chromatin changes showed that, for light-repressed genes, H3K9 deacetylation parallels light-triggered transcriptional repression, while for light-induced genes, H3K9 acetylation appeared to somewhat precede light-activated transcript accumulation. However, direct, real-time imaging of transcript elongation in the nucleus revealed that, in fact, transcriptional induction actually parallels H3K9 acetylation. Collectively, the data raise the possibility that light-induced transcriptional and chromatin-remodeling processes are mechanistically intertwined. Histone modifying proteins involved in long term light responses do not seem to have a role in this fast response, indicating that different factors might act at different stages of the light response. This work not only advances our understanding of plant responses to light, but also unveils a system in which rapid chromatin changes in reaction to an external signal can be studied under natural conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Oblique instability of quasi-parallel whistler waves in the presence of cold and warm electron populations

Whistler waves propagating nearly parallel to the ambient magnetic field experience a nonlinear instability due to transverse currents when the background plasma has a population of sufficiently low energy electrons. Intriguingly, this nonlinear process may generate oblique electrostatic waves, including whistlers near the resonance cone with properties resembling oblique chorus waves in the Earth’s magnetosphere. Focusing on the generation of oblique whistlers, earlier analysis of the instability is extended here to the case where low-energy background plasma consists of both a “cold” population with energy of a few eV and a “warm” electron component with energy of the order of 100 eV. This is motivated by spacecraft observations in the Earth’s magnetosphere where oblique chorus waves were shown to interact resonantly with the warm electrons. The main new results are: 1) the instability producing oblique electrostatic waves is sensitive to the shape of the electron distribution at low energies. In the whistler range of frequencies, two distinct peaks in the growth rate are typically present for the model considered: a peak associated with the warm electron population at relatively low wavenumbers and a peak associated with the cold electron population at relatively high wavenumbers; 2) overall, the instability producing oblique whistler waves near the resonance cone persists (with a reduced growth rate) even in the cases where the temperature of the cold population is relatively high, including cases where cold population is absent and only the warm population is included; 3) particle-in-cell simulations show that the instability leads to heating of the background plasma and formation of characteristic plateau and beam features in the parallel electron distribution function in the range of energies resonant with the instability. The plateau/beam features have been previously detected in spacecraft observations of oblique chorus waves. However, they have been attributed to external sources and have been proposed to be the mechanism generating oblique chorus. In the present scenario, the causality link is reversed and the instability generating oblique whistler waves is shown to be a possible mechanism for formation of the plateau and beam features.

79 ASTRONOMY AND ASTROPHYSICS↗

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)↗

Parallel Diffusion Coefficient of Energetic Charged Particles in the Inner Heliosphere from the Turbulent Magnetic Fields Measured by Parker Solar Probe

Diffusion coefficients of energetic charged particles in turbulent magnetic fields are a fundamental aspect of diffusive transport theory but remain incompletely understood. In this work, we use quasi-linear theory to evaluate the spatial variation of the parallel diffusion coefficient κ ∥ from the measured magnetic turbulence power spectra in the inner heliosphere. We consider the magnetic field and plasma velocity measurements from Parker Solar Probe made during Orbits 5–13. The parallel diffusion coefficient is calculated as a function of radial distance from 0.062 to 0.8 au, and the particle energy from 100 keV to 1 GeV. We find that κ ∥ increases exponentially with both heliocentric distance and energy of particles. The fluctuations in κ ∥ are related to the episodes of large-scale magnetic structures in the solar wind. By fitting the results, we also provide an empirical formula of κ ∥ = (5.16 ± 1.22) × 10 18 r 1.17 ± 0.08 E 0.71 ± 0.02 (cm 2 s -1 ) in the inner heliosphere, which can be used as a reference in studying the transport and acceleration of solar energetic particles as well as the modulation of cosmic rays.

79 ASTRONOMY AND ASTROPHYSICS↗

Electron Influence on the Parallel Proton Firehose Instability in 10-moment, Multifluid Simulations

Instabilities driven by pressure anisotropy play a critical role in modulating the energy transfer in space and astrophysical plasmas. For the first time, we simulate the evolution and saturation of the parallel proton firehose instability using a multifluid model without adding artificial viscosity. These simulations are performed using a 10-moment, multifluid model with local and gradient relaxation heat-flux closures in high-β proton–electron plasmas. When these higher-order moments are included and pressure anisotropy is permitted to develop in all species, we find that the electrons have a significant impact on the saturation of the parallel proton firehose instability, modulating the proton pressure anisotropy as the instability saturates. Even for lower β's more relevant to heliospheric plasmas, we observe a pronounced electron energization in simulations using the gradient relaxation closure. Our results indicate that resolving the electron pressure anisotropy is important to correctly describe the behavior of multispecies plasma systems.

79 ASTRONOMY AND ASTROPHYSICS↗

Complementary-MOS binary counter with parallel-set inputs

Metal oxide semiconductor four-stage binary counter contains reset capability as well as four parallel-set inputs gated in by a logic signal. Parallel-set inputs permit setting the counter into any of sixteen possible states.

Keller, K. R.↗

An efficient parallel algorithm for the solution of a tridiagonal linear system of equations

Tridiagonal linear systems of equations are solved on conventional serial machines in a time proportional to N, where N is the number of equations. The conventional algorithms do not lend themselves directly to parallel computations on computers of the ILLIAC IV class, in the sense that they appear to be inherently serial. An efficient parallel algorithm is presented in which computation time grows as log sub 2 N. The algorithm is based on recursive doubling solutions of linear recurrence relations, and can be used to solve recurrence relations of all orders.

Stone, H. S.↗

Analysis and performance of paralleling circuits for modular inverter-converter systems

As part of a modular inverter-converter development program, control techniques were developed to provide load sharing among paralleled inverters or converters. An analysis of the requirements of paralleling circuits and a discussion of the circuits developed and their performance are included in this report. The current sharing was within 5.6 percent of rated-load current for the ac modules and 7.4 percent for the dc modules for an initial output voltage unbalance of 5 volts.

Birchenough, A. G.↗

Monolithic parallel processor, phase 1A

A four-bit parallel processor LSI array was designed and fabricated using COS/MOS integrated-circuit technology. The design features include the provision for interconnecting groups of parallel-processor chips to form an expanded processor of any desired word length. This 800-transistor "computer on a chip' circuit has the logic capability of a medium-size, medium-speed, general-purpose computer suitable for sophisticated scientific data processing. The ability to fabricate this device repetitively was demonstrated.

Source record↗

Automatic recognition of vector and parallel operations in a higher level language

A compiler for recognizing statements of a FORTRAN program which are suited for fast execution on a parallel or pipeline machine such as Illiac-4, Star or ASC is described. The technique employs interval analysis to provide flow information to the vector/parallel recognizer. Where profitable the compiler changes scalar variables to subscripted variables. The output of the compiler is an extension to FORTRAN which shows parallel and vector operations explicitly.

Schneck, P. B.↗