Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

The Influence of Environment on Post-Detonation Chemistry and Debris Formation (Abbreviated Final Report: 20-SI-006)

Predicting, responding to, or interpreting the chemical record preserved in debris derived from nuclear events can be challenging due to chemical fractionation. Chemical fractionation is where different species of the evolving radionuclide inventory segregate and/or are lost from the system over the timescales of debris formation. Both historic data and recent research suggest that the interaction and character of the local environment may exert controls on chemical fractionation by influencing the cooling and evolution of the associated fireball as well as the composition of the vapor term and resultant speciation. Prior to this work, an integrated platform permitting dynamic and concurrent consideration of physical and chemical evolution of early time post-detonation event environments did not exist. Our work merged historic data and experimental approaches to support development of a computational framework able to simulate fundamental processes (e.g., entrainment of local environment, oxidation chemistry, and cooling time scales) that may perturb the radionuclide inventory captured in post-detonation debris. Work with historic debris confirmed that entrained environmental material affect debris composition, structure, and radionuclide incorporation. Complementary work utilizing a readily controllable and tunable benchtop setup (a plasma flow reactor) simulated the late cooling of a nuclear fireball (e.g., T < 6000 K) and bounded the sensitivity of actinide speciation and particle size distribution to variations in oxygen concentration and cooling rates. Concurrent laser ablation and laser heating experiments were used to investigate the chemistry and physics of processes occurring in vaporized and/or rapidly heated actinides and other elements in the presence of oxygen. A more computationally efficient microphysical model was developed for predicting and evolving size distributions of particles forming from mixed vapor terms and simulating particle formation processes under a variety of extreme conditions. Continued study of historic nuclear event film confirmed that shockwave data and physics codes agree to within the uncertainty of the data. Good agreement was achieved for thermal emission from an airburst, however the paucity of low-temperature molecular opacity data for mixtures of air, bomb debris, entrained dirt, and water vapor complicate agreement for more elaborate scenarios. A multiphysics code (ALE3D) was modified to bring the necessary physics and chemistry, including these new data and insights, onto a single platform. Code development included improved initialization of large physical systems, modernization of chemistry capabilities, and modifications to enable inclusion of particle transport.

07 ISOTOPE AND RADIATION SOURCES↗

Classical Benchmarks for Variational Quantum Eigensolver Simulations of the Hubbard Model

Simulating the Hubbard model is of great interest to a wide range of applications within condensed matter physics, however its solution on classical computers remains challenging in dimensions larger than one. The relative simplicity of this model, embodied by the sparseness of the Hamiltonian matrix, allows for its efficient implementation on quantum computers, and for its approximate solution using variational algorithms such as the variational quantum eigensolver. While these algorithms have been shown to reproduce the qualitative features of the Hubbard model, their quantitative accuracy in terms of producing true ground state energies and other properties, and the dependence of this accuracy on the system size and interaction strength, the choice of variational ansatz, and the degree of spatial inhomogeneity in the model, remains unknown. Here we present a rigorous classical benchmarking study, demonstrating the potential impact of these factors on the accuracy of the variational solution of the Hubbard model on quantum hardware, for systems with up to 32 qubits. We find that even when using the most accurate wavefunction ansätze for the Hubbard model, the error in its ground state energy and wavefunction plateaus for larger lattices, while stronger electronic correlations magnify this issue. Concurrently, spatially inhomogeneous parameters and the presence of off-site Coulomb interactions only have a small effect on the accuracy of the computed ground state energies. Our study highlights the capabilities and limitations of current approaches for solving the Hubbard model on quantum hardware, and we discuss potential future avenues of research.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

GSoFa: Scalable Sparse Symbolic LU Factorization on GPUs

Decomposing a matrix $\mathbf {A}$ into a lower matrix $\mathbf {L}$ and an upper matrix $\mathbf {U}$, which is also known as LU decomposition, is an essential operation in numerical linear algebra. For a sparse matrix, LU decomposition often introduces more nonzero entries in the $\mathbf {L}$ and $\mathbf {U}$ factors than in the original matrix. A symbolic factorization step is needed to identify the nonzero structures of $\mathbf {L}$ and $\mathbf {U}$ matrices. Attracted by the enormous potentials of the Graphics Processing Units (GPUs), an array of efforts have surged to deploy various LU factorization steps except for the symbolic factorization, to the best of our knowledge, on GPUs. This article introduces gSoFa, the first GPU-based symbolic factorization design with the following three optimizations to enable scalable LU symbolic factorization for nonsymmetric pattern sparse matrices on GPUs. First, here we introduce a novel fine-grained parallel symbolic factorization algorithm that is well suited for the Single Instruction Multiple Thread (SIMT) architecture of GPUs. Second, we tailor supernode detection into a SIMT friendly process and strive to balance the workload, minimize the communication and saturate the GPU computing resources during supernode detection. Third, we introduce a three-pronged optimization to reduce the excessive space consumption problem faced by multi-source concurrent symbolic factorization. Taken together, gSoFa achieves up to 31× speedup from 1 to 44 Summit nodes (6 to 264 GPUs) and outperforms the state-of-the-art CPU project, on average, by 5×. Notably, gSoFa also achieves up to 47 percent of the peak memory throughput of a V100 GPU in the Summit Supercomputer.

97 MATHEMATICS AND COMPUTING↗

Three dimensional nozzle-exhaust flow field analysis by a reference plane technique.

A numerical method based on reference plane characteristics has been developed for the calculation of highly complex supersonic nozzle-exhaust flow fields. The difference equations have been developed for three coordinate systems. Local reference plane orientations are employed using the three coordinate systems concurrently thus catering to a wide class of flow geometries. Discontinuities such as the underexpansion shock and contact surfaces are computed explicitly for nonuniform vehicle external flows. The nozzles considered may have irregular cross-sections with swept throats and may be stacked in modules using the vehicle undersurface for additional expansion. Results are presented for several nozzle configurations.

Dash, S. M.↗

Remote file inquiry (RFI) system

System interrogates and maintains user-definable data files from remote terminals, using English-like, free-form query language easily learned by persons not proficient in computer programming. System operates in asynchronous mode, allowing any number of inquiries within limitation of available core to be active concurrently.

Source record↗

Digital system for dynamic turbine engine blade displacement measurements

An instrumentation concept for measuring blade tip displacements which employs optical probes and an array of micro-computers is presented. The system represents a hitherto unknown instrumentation capability for the acquisition and direct digitization of deflection data concurrently from all of the blade tips of an operational engine rotor undergoing flutter or forced vibration. System measurements are made using optical transducers which are fixed to the case. Measurements made in this way are the equivalent of those obtained by placing three surface-normal displacement transducers at three positions on each blade of an operational rotor.

Kiraly, L. J.↗

A transputer based finite element solver

The feasibility of performing FEM structural-mechanics analyses on transputer systems is investigated experimentally. Transputers are programmable microprocessors equipped with local memory and point-to-point communication links; they can be joined in a large concurrent system via a programming language which supports distributed processing; this permits parallel processing at relatively low hardware cost. The computational tasks required by FEM programs are reviewed; the hardware (one PC, one master transputer, and 12 slave transputers) employed in the test calculations is described; and results demonstrating the speed and efficiency of the transputer array in assembling a global stiffness matrix and performing Gauss-Jordan matrix inversion are presented in graphs. It is predicted that larger transputer networks could approach the power of supercomputers at minicomputer costs.

Favenesi, J. A.↗

STEP: A Futurevision, Today

STEP (STandard for the Exchange of Product Model Data) is an innovative software tool that allows the exchange of data between different programming systems to occur and helps speed up the designing in various process industries. This exchange occurs easily between those companies that have STEP, and many industries and government agencies are requiring that their vendors utilize STEP in their computer aided design projects, such as in the areas of mechanical, aeronautical, and electrical engineering. STEP allows the process of concurrent engineering to occur and increases the quality of the design product. One example of the STEP program is the Boeing 777, the first paperless airplane.

Source record↗

Computational And Experimental Studies Of Three-Dimensional Flame Spread Over Liquid Fuel Pools

Schiller, Ross, and Sirignano (1996) studied ignition and flame spread above liquid fuels initially below the flashpoint temperature by using a two-dimensional computational fluid dynamics code that solves the coupled equations of both the gas and the liquid phases. Pulsating flame spread was attributed to the establishment of a gas-phase recirculation cell that forms just ahead of the flame leading edge because of the opposing effect of buoyancy-driven flow in the gas phase and the thermocapillary-driven flow in the liquid phase. Schiller and Sirignano (1996) extended the same study to include flame spread with forced opposed flow in the gas phase. A transitional flow velocity was found above which an originally uniform spreading flame pulsates. The same type of gas-phase recirculation cell caused by the combination of forced opposed flow, buoyancy-driven flow, and thermocapillary-driven concurrent flow was responsible for the pulsating flame spread. Ross and Miller (1998) and Miller and Ross (1998) performed experimental work that corroborates the computational findings of Schiller, Ross, and Sirignano (1996) and Schiller and Sirignano (1996). Cai, Liu, and Sirignano (2002) developed a more comprehensive three-dimensional model and computer code for the flame spread problem. Many improvements in modeling and numerical algorithms were incorporated in the three-dimensional model. Pools of finite width and length were studied in air channels of prescribed height and width. Significant three-dimensional effects around and along the pool edge were observed. The same three-dimensional code is used to study the detailed effects of pool depth, pool width, opposed air flow velocity, and different levels of air oxygen concentration (Cai, Liu, and Sirignano, 2003). Significant three-dimensional effects showing an unsteady wavy flame front for cases of wide pool width are found for the first time in computation, after being noted previously by experimental observers (Ross and Miller, 1999). Regions of uniform and pulsating flame spread are mapped for the flow conditions of pool depth, opposed flow velocity, initial pool temperature, and air oxygen concentration under both normal and microgravity conditions. Details can be found in Cai et al. (2002, 2003). Experimental results recently performed at NASA Glenn of flame spread across a wide, shallow pool as a function of liquid temperature are also presented here.

Ross, Howard D.↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

Optical multiple access techniques for on-board routing

The purpose of this research contract was to design and analyze an optical multiple access system, based on Code Division Multiple Access (CDMA) techniques, for on board routing applications on a future communication satellite. The optical multiple access system was to effect the functions of a circuit switch under the control of an autonomous network controller and to serve eight (8) concurrent users at a point to point (port to port) data rate of 180 Mb/s. (At the start of this program, the bit error rate requirement (BER) was undefined, so it was treated as a design variable during the contract effort.) CDMA was selected over other multiple access techniques because it lends itself to bursty, asynchronous, concurrent communication and potentially can be implemented with off the shelf, reliable optical transceivers compatible with long term unattended operations. Temporal, temporal/spatial hybrids and single pulse per row (SPR, sometimes termed 'sonar matrices') matrix types of CDMA designs were considered. The design, analysis, and trade offs required by the statement of work selected a temporal/spatial CDMA scheme which has SPR properties as the preferred solution. This selected design can be implemented for feasibility demonstration with off the shelf components (which are identified in the bill of materials of the contract Final Report). The photonic network architecture of the selected design is based on M(8,4,4) matrix codes. The network requires eight multimode laser transmitters with laser pulses of 0.93 ns operating at 180 Mb/s and 9-13 dBm peak power, and 8 PIN diode receivers with sensitivity of -27 dBm for the 0.93 ns pulses. The wavelength is not critical, but 830 nm technology readily meets the requirements. The passive optical components of the photonic network are all multimode and off the shelf. Bit error rate (BER) computations, based on both electronic noise and intercode crosstalk, predict a raw BER of (10 exp -3) when all eight users are communicating concurrently. If better BER performance is required, then error correction codes (ECC) using near term electronic technology can be used. For example, the M(8,4,4) optical code together with Reed-Solomon (54,38,8) encoding provides a BER of better than (10 exp -11). The optical transceiver must then operate at 256 Mb/s with pulses of 0.65 ns because the 'bits' are now channel symbols.

Mendez, Antonio J.↗

Dispersion-enhanced sequential batch sampling for adaptive contour estimation

In computer simulation and optimal design, sequential batch sampling offers an appealing way to iteratively stipulate optimal sampling points based upon existing selections and efficiently construct surrogate modeling. Nonetheless, the issue of near duplicates poses tremendous quandary for sequential learning. It refers to the situation that selected critical points cluster together in each sampling batch, which are individually but not collectively informative towards the optimal design. Near duplicates severely diminish the computational efficiency as they barely contribute extra information towards update of the surrogate. To address this issue, we impose a dispersion criterion on concurrent selection of sampling points, which essentially forces a sparse distribution of critical points in each batch, and demonstrate the effectiveness of this approach in adaptive contour estimation. Specifically, we adopt Gaussian process surrogate to emulate the simulator, acquire variance reduction of the critical region from new sampling points as a dispersion criterion, and combine it with the modified expected improvement (EI) function for critical batch selection. The critical region here is the proximity of the contour of interest. This proposed approach is vindicated in numerical examples of a two-dimensional four-branch function, a four-dimensional function with a disjoint contour of interest and a time-delay dynamic system.

97 MATHEMATICS AND COMPUTING↗

Solar Radiation Transport in the Cloudy Atmosphere: A 3D Perspective on Observations and Climate Impacts

The interplay of sunlight with clouds is a ubiquitous and often pleasant visual experience, but it conjures up major challenges for weather, climate, environmental science and beyond. Those engaged in the characterization of clouds (and the clear air nearby) by remote sensing methods are even more confronted. The problem comes, on the one hand, from the spatial complexity of real clouds and, on the other hand, from the dominance of multiple scattering in the radiation transport. The former ingredient contrasts sharply with the still popular representation of clouds as homogeneous plane-parallel slabs for the purposes of radiative transfer computations. In typical cloud scenes the opposite asymptotic transport regimes of diffusion and ballistic propagation coexist. We survey the three-dimensional (3D) atmospheric radiative transfer literature over the past 50 years and identify three concurrent and intertwining thrusts: first, how to assess the damage (bias) caused by 3D effects in the operational 1D radiative transfer models? Second, how to mitigate this damage? Finally, can we exploit 3D radiative transfer phenomena to innovate observation methods and technologies? We quickly realize that the smallest scale resolved computationally or observationally may be artificial but is nonetheless a key quantity that separates the 3D radiative transfer solutions into two broad and complementary classes: stochastic and deterministic. Both approaches draw on classic and contemporary statistical, mathematical and computational physics.

Davis, Anthony B.↗

Power converter design optimization

Utilizing the demonstrated capability of nonlinear programming algorithms, a practical design optimization approach for power converters is established to conceive a design to meet all power-circuit performance requirements and concurrently optimize a defined quantity such as weight or losses. In addition, to facilitate a cost-effective design, the computer-aided approach provides a means to readily assess (1) the weight-efficiency tradeoff, (2) impacts of converter requirements and component characteristics on a given design, and (3) optimum power system configurations.

Yu, Y.↗

Reactive system verification case study: Fault-tolerant transputer communication

A reactive program is one which engages in an ongoing interaction with its environment. A system which is controlled by an embedded reactive program is called a reactive system. Examples of reactive systems are aircraft flight management systems, bank automatic teller machine (ATM) networks, airline reservation systems, and computer operating systems. Reactive systems are often naturally modeled (for logical design purposes) as a composition of autonomous processes which progress concurrently and which communicate to share information and/or to coordinate activities. Formal (i.e., mathematical) frameworks for system verification are tools used to increase the users' confidence that a system design satisfies its specification. A framework for reactive system verification includes formal languages for system modeling and for behavior specification and decision procedures and/or proof-systems for verifying that the system model satisfies the system specifications. Using the Ostroff framework for reactive system verification, an approach to achieving fault-tolerant communication between transputers was shown to be effective. The key components of the design, the decoupler processes, may be viewed as discrete-event-controllers introduced to constrain system behavior such that system specifications are satisfied. The Ostroff framework was also effective. The expressiveness of the modeling language permitted construction of a faithful model of the transputer network. The relevant specifications were readily expressed in the specification language. The set of decision procedures provided was adequate to verify the specifications of interest. The need for improved support for system behavior visualization is emphasized.

Crane, D. Francis↗

SUNDIALS time integrators for exascale applications with many independent systems of ordinary differential equations

Many complex systems can be accurately modeled as a set of coupled time-dependent partial differential equations (PDEs). However, solving such equations can be prohibitively expensive, easily taxing the world’s largest supercomputers. One pragmatic strategy for attacking such problems is to split the PDEs into components that can more easily be solved in isolation. This operator splitting approach is used ubiquitously across scientific domains, and in many cases leads to a set of ordinary differential equations (ODEs) that need to be solved as part of a larger “outer-loop” time-stepping approach. The SUNDIALS library provides a plethora of robust time integration algorithms for solving ODEs, and the U.S. Department of Energy Exascale Computing Project (ECP) has supported its extension to applications on exascale-capable computing hardware. In this paper, we highlight some SUNDIALS capabilities and its deployment in combustion and cosmology application codes (Pele and Nyx, respectively) where operator splitting gives rise to numerous, small ODE systems that must be solved concurrently.

97 MATHEMATICS AND COMPUTING↗