Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Availability and redundancy design considerations for Electron-Ion Collider cooling water systems and components

The Electron-Ion Collider (EIC) is composed of several large-scale systems, which all require highly available cooling systems based on the Conceptual Design Report. Since the overall EIC availability goal is 85%, all the subsystems will contribute. Availability is defined by the time that a component or system is functional given a required or scheduled run time. The focus of the project was to determine if redundant or extra cooling water and deionized water pumps were needed in a subsystem to increase the overall EIC availability. The largest subsystem in the ring is the 10 o’clock cooling system, so this was used for redundancy analysis. Relationships between the components were defined as parallel, series, or m-out-of-n for the cooling towers, valves, pumps and heat exchangers to perform calculations. A comparison between an original system versus a redundant one was then done using component availability from industry standards. Cost analysis was also performed to see how much the availability and cost would decrease if the redundancy application was not chosen. The Collider-Accelerator Department (CAD) also has a calculation in place for the Relativistic-Heavy Ion Collider (RHIC), so cooling water failure data was collected from their operational log and analyzed for real-life component availability. A comparison of component availability from the CAD data to the industry standards was made through the 10 o’clock cooling system layout. The results showed that the subsystems need to have over a 98% availability for an 85% goal, and that the CAD data is very close to the industry component standards. The cost differential is miniscule compared to the potential availability increase. Therefore, redundant systems can increase the overall EIC availability by a sizable amount and are worth implementing.

43 PARTICLE ACCELERATORS↗

Development and Analysis of Optimal Multilevel Solvers on Advanced Computers. Final Report

Constrained optimization in the context of time dependent, partial differential equations (PDE) leads to a symmetric, block-tridiagonal system of nonlinear equations that must be solved repeatedly in an iterative solution strategy. The blocks represent spatial discretization, while the connection between the blocks represents a forward and backward integration in time. The focus of this project is to apply a parallel-in-time (PiT) solution technique to the large block-triangular system.

97 MATHEMATICS AND COMPUTING↗

Parallel algorithms for hyperdynamics and local hyperdynamics

Hyperdynamics (HD) is a method for accelerating the timescale of standard molecular dynamics (MD). It can be used for simulations of systems with an energy potential landscape that is a collection of basins, separated by barriers, where transitions between basins are infrequent. HD enables the system to escape from a basin more quickly while enabling a statistically accurate renormalization of the simulation time, thus effectively boosting the timescale of the simulation. In [Kim, Perez, Voter, J Chem Phys, 139:144110, 2013)1, a local version of HD was formulated, which exploits the intrinsic locality characteristic typical of most systems to mitigate the poor scaling properties of standard HD as the system size is increased. In this paper, we discuss how both HD and local HD can be formulated to run efficiently in parallel. We have implemented these ideas in the LAMMPS MD code, which means HD can be used with any interatomic potential LAMMPS supports. Together, these parallel methods allow simulations of any size to achieve the time acceleration offered by HD (which can be orders of magnitude), at a cost 3-5x that of standard MD. As examples, we performed two simulations of a million-atom system to model the diffusion and clustering of Pt adatoms on a large patch of Pt(100) surface for 80 and 160 μs.

74 ATOMIC AND MOLECULAR PHYSICS↗

Label-Free Profiling of up to 200 Single-Cell Proteomes per Day Using a Dual-Column Nanoflow Liquid Chromatography Platform

Single-cell proteomics (SCP) has great potential to advance biomedical research and personalized medicine. The sensitivity of such measurements increases with low-flow separations (<100 nL/min) due to improved ionization efficiency, but the time required for sample loading, column washing, and regeneration in these systems can lead to low measurement throughput and inefficient utilization of the mass spectrometer. Herein, we developed a two-column liquid chromatography (LC) system that dramatically increases the throughput of label-free SCP using two parallel subsystems to multiplex sample loading, online desalting, analysis, and column regeneration. The integration of MS1-based feature matching increased proteome coverage when short LC gradients were used. The high-throughput LC system was reproducible between the columns, with a 4% difference in median peptide abundance and a median CV of 18% across 100 replicate analyses of a single-cell-sized peptide standard. An average of 621, 774, 952, and 1622 protein groups were identified with total analysis times of 7, 10, 15, and 30 min, corresponding to a measurement throughput of 206, 144, 96, and 48 samples per day, respectively. When applied to single HeLa cells, we identified nearly 1000 protein groups per cell using 30 min cycles and 660 protein groups per cell for 15 min cycles. Finally, we explored the possibility of measuring cancer therapeutic targets with a pilot study comparing the K562 and Jurkat leukemia cell lines. This work demonstrates the feasibility of high-throughput label-free single-cell proteomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Physics of the Inverted Harmonic Oscillator: From the lowest Landau level to event horizons

In this work, we present the inverted harmonic oscillator (IHO) Hamiltonian as a paradigm to understand the quantum mechanics of scattering and time-decay in a diverse set of physical systems. As one of the generators of area preserving transformations, the IHO Hamiltonian can be studied as a dilatation generator, squeeze generator, a Lorentz boost generator, or a scattering potential. In establishing these different forms, we demonstrate the physics of the IHO that underlies phenomena as disparate as the Hawking–Unruh effect and scattering in the lowest Landau level (LLL) in quantum Hall systems. We derive the emergence of the IHO Hamiltonian in the LLL in a gauge invariant way and show its exact parallels with the Rindler Hamiltonian that describes quantum mechanics near event horizons. This approach of studying distinct physical systems with symmetries described by isomorphic Lie algebras through the emergent IHO Hamiltonian enables us to reinterpret geometric response in the lowest Landau level in terms of relativistic effects such as Wigner rotation. Further, the analytic scattering matrix of the IHO points to the existence of quasinormal modes (QNMs) in the spectrum, which have quantized time-decay rates. We present a way to access these QNMs through wave packet scattering, thus proposing a novel effect in quantum Hall point contact geometries that parallels those found in black holes.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Slow Shock Formation Upstream of Reconnecting Current Sheets

The formation, development, and impact of slow shocks in the upstream regions of reconnecting current layers are explored. Slow shocks have been documented in the upstream regions of magnetohydrodynamic (MHD) simulations of magnetic reconnection as well as in similar simulations with the kglobal kinetic macroscale simulation model. They are therefore a candidate mechanism for preheating the plasma that is injected into the current layers that facilitate magnetic energy release in solar flares. Of particular interest is their potential role in producing the hot thermal component of electrons in flares. During multi-island reconnection, the formation and merging of flux ropes in the reconnecting current layer drives plasma flows and pressure disturbances in the upstream region. These pressure disturbances steepen into slow shocks that propagate along the reconnecting component of the magnetic field and satisfy the expected Rankine–Hugoniot jump conditions. Plasma heating arises from both compression across the shock and the parallel electric field that develops to maintain charge neutrality in a kinetic system. Shocks are weaker at lower plasma β, where shock steepening is slow. While these upstream slow shocks are intrinsic to the dynamics of multi-island reconnection, their contribution to electron heating remains relatively minor compared with that from Fermi reflection and the parallel electric fields that bound the reconnection outflow.

79 ASTRONOMY AND ASTROPHYSICS↗

Systems and methods for powder bed additive manufacturing anomaly detection

Detection and classification of anomalies for powder bed metal additive manufacturing. Anomalies, such as recoater blade impacts, binder deposition issues, spatter generation, and some porosities, are surface-visible at each layer of the building process. A multi-scaled parallel dynamic segmentation convolutional neural network architecture provides additive manufacturing machine and imaging system agnostic pixel-wise semantic segmentation of layer-wise powder bed image data. Learned knowledge is easily transferrable between different additive manufacturing machines. The anomaly detection can be conducted in real-time and provides accurate and generalizable results.

Scime, Luke R.↗

Coupling Surface Flow with High-performance Subsurface Reactive Flow and Transport Code PFLOTRAN

Water exchange between the surface and subsurface is important for both water resource management and environmental protection. In this paper, we develop coupled surface and subsurface flow simulation capability in a parallel subsurface flow and reactive transport code PFLOTRAN. We sequentially couple the diffusion wave-based surface flow with the subsurface flow governedby the Richards equation in PFLOTRAN. These two flow domains are linked with a boundary condition switching method that ensures continuity of pressure and flux at the surface-subsurface interface. We verify the coupled code against other existing hydrologic models and observation data using a number of numerical experiments. The coupled hydrological model exhibits good performance in strong parallel scaling tests. The new coupled surface and subsurface simulator significantly advance community simulation capability towards improving integrated hydrologic and biogeochemical understanding of complex systems such as watersheds and river corridors. Keywords: Surface flow, Integrated hydrological modeling, Boundary condition switching, Parallel computing

Wu, Runjian↗

Characterizing Output Bottlenecks of a Production Supercomputer: Analysis and Implications

This article studies the I/O write behaviors of the Titan supercomputer and its Lustre parallel file stores under production load. The results can inform the design, deployment, and configuration of file systems along with the design of I/O software in the application, operating system, and adaptive I/O libraries.We propose a statistical benchmarking methodology to measure write performance across I/O configurations, hardware settings, and system conditions. Moreover, we introduce two relative measures to quantify the write-performance behaviors of hardware components under production load. In addition to designing experiments and benchmarking on Titan, we verify the experimental results on one real application and one real application I/O kernel, XGC and HACC IO, respectively. These two are representative and widely used to address the typical I/O behaviors of applications.In summary, we find that Titan’s I/O system is variable across the machine at fine time scales. This variability has two major implications. First, stragglers lessen the benefit of coupled I/O parallelism (striping). Peak median output bandwidths are obtained with parallel writes to many independent files, with no striping or write sharing of files across clients (compute nodes). I/O parallelism is most effective when the application—or its I/O libraries—distributes the I/O load so that each target stores files for multiple clients and each client writes files on multiple targets in a balanced way with minimal contention. Second, our results suggest that the potential benefit of dynamic adaptation is limited. In particular, it is not fruitful to attempt to identify “good locations” in the machine or in the file system: component performance is driven by transient load conditions and past performance is not a useful predictor of future performance. For example, we do not observe diurnal load patterns that are predictable.

97 MATHEMATICS AND COMPUTING↗

Development of Real-Time System Identification to Detect Abnormal Operations in a Gas Turbine Cycle

Here, we present a novel online system identification methodology for monitoring the performance of power systems. This methodology was demonstrated in a gas turbine recuperated power plant designed for a hybrid configuration. A 120-kW Garrett microturbine modified to test dynamic control strategies for hybrid power systems designed at the National Energy Technology Laboratory (NETL) was used to implement and validate this online system identification methodology. The main component of this methodology consists of an empirical transfer function model implemented in parallel to the turbine speed operation and the fuel control valve, which can monitor the process response of the gas turbine system while it is operating. During fully closed-loop operations or automated control, the output of the controller, fuel valve position, and the turbine speed measurements were fed for a given period of time to a recursive algorithm that determined the transfer function parameters during the nominal condition. After the new parameters were calculated, they were fed into the transfer function model for online prediction. The turbine speed measurement was compared against the transfer function prediction, and a control logic was implemented to capture when the system operated at nominal or abnormal conditions. To validate the ability to detect abnormal conditions during dynamic operations, drifting in the performance of the gas turbine system was evaluated. A leak in the turbomachinery working fluid was emulated by bleeding 10% of the airflow from the compressor discharge to the atmosphere, and electrical load steps were performed before and after the leak. This tool could detect the leak 7 s after it had occurred, which accounted for a fuel flow increase of approximately 15.8% to maintain the same load and constant turbine speed operations.

algorithms↗

Generating Massive Scale-free Networks: Novel Parallel Algorithms using the Preferential Attachment Model

Recently, there has been substantial interest in the study of various random networks as mathematical models of complex systems. As real-life complex systems grow larger, the ability to generate progressively large random networks becomes all the more important. This motivates the need for efficient parallel algorithms for generating such networks. Naïve parallelization of sequential algorithms for generating random networks is inefficient due to inherent dependencies among the edges and the possibility of creating duplicate (parallel) edges. In this article, we present message passing interface-based distributed memory parallel algorithms for generating random scale-free networks using the preferential-attachment model. Our algorithms are experimentally verified to scale very well to a large number of processing elements (PEs), providing near-linear speedups. The algorithms have been exercised with regard to scale and speed to generate scale-free networks with one trillion edges in 6 minutes using 1,000 PEs.

97 MATHEMATICS AND COMPUTING↗

Performance Analysis of Speculative Parallel Adaptive Local Timestepping for Conservation Laws

Stable simulation of conservation laws, such as those used to model fluid dynamics and plasma physics applications, requires the satisfaction of the so-called Courant-Friedrichs-Lewy condition. By allowing regions of the mesh to advance with different timesteps that locally satisfy this stability constraint, significant work reduction can be attained when compared to a time integration scheme using a single timestep size. However, parallelizing this algorithm presents considerable difficulty. Since the stability condition depends on the state of the system, dependencies become dynamic and potentially non-local. In this article, we present an adaptive local timestepping algorithm using an optimistic (Timewarp-based) parallel discrete event simulation. We introduce waiting heuristics to limit misspeculation and a semi-static load balancing scheme to eliminate load imbalance as parts of the mesh require finer or coarser timesteps. Last, we outline an interface for separating the physics of the specific conservation law from the temporal integration allowing for productive adoption of our proposed algorithm. We present a misspeculation study for three conservation laws, demonstrating both the productivity of the local timestepping API, for which 74% of the lines of code are reused across different conservation laws, and the robustness of the waiting heuristics—at most 1.5% of element updates are rolled back. Our performance studies demonstrate up to a 2.8× speedup versus a baseline unoptimized local timestepping approach, a 4x improvement in per-node throughput compared to an MPI parallelization of synchronous timestepping, and scalability up to 3,072 cores on NERSC’s Cori Haswell partition.

97 MATHEMATICS AND COMPUTING↗

An agent-based deployment decision-support system for electric vehicle services

METS-R ADDSEVS simulator is a high fidelity, parallel, agent-based evacuation simulator for multi-modal energy-optimal trip scheduling in real-time (METS-R) at transportation hubs. It consists of two modules. The first one is the traffic simulator module; the second one is the high-performance computing (HPC) module. More details can be found at https://umnilab.github.io/METS-R_doc/.

Lei, Zengxiang↗

Instability Issue of Paralleled Dies in an SiC Power Module in Solid-State Circuit Breaker Applications

Paralleled dies in a power module could have instability issues during high current switching transients. Here, the instability is caused by the differential-mode oscillation among paralleled MOSFETs. Conventional analyses of paralleled MOSFETs’ stability are normally limited to a single operating point, which ignores the influences of the switching trajectory and nonlinear device parameters on stability. This article reveals that the switching trajectory can significantly influence parallel stability. The analysis is improved by solving eigenvalues of state-space modeling system matrices of all operating points that the switching trajectory goes through considering nonlinear device parameters. Higher voltage and current stresses result in greater real parts of complex eigenvalues, which explains why the paralleled MOSFETs are more unstable with higher voltage and current stresses. To improve stability in solid-state circuit breaker applications, we propose a method to manipulate the switching trajectory to avoid the unstable region where the conventional hard switching trajectory normally goes through. Experimental results show that the turn- off current capability can be increased from ~five times of rated current with the gate oscillation using the conventional turn- off trajectory to ~ten times of rated current without the gate oscillation using the optimal turn- off trajectory.

42 ENGINEERING↗

Contingency Analysis Based on Partitioned and Parallel Holomorphic Embedding

In the steady-state contingency analysis, the traditional Newton-Raphson method suffers from non-convergence issues when solving post-outage power flow problems, which hinders the integrity and accuracy of security assessment. In this paper, we propose a novel robust contingency analysis approach based on holomorphic embedding (HE). Here, the HE-based simulator provides theoretical convergence guarantee, which is desirable because it avoids the influence of numerical issues and provides a credible security assessment conclusion. In addition, based on the multi-area characteristics of real-world power systems, a partitioned HE (PHE) method is proposed with an interfacebased partitioning of HE formulation. The PHE method does not undermine the numerical robustness of HE and significantly reduces the computation burden in large-scale contingency analysis. The PHE method is further enhanced by parallel or distributed computation to become parallel PHE (P2HE). Tests on a 458-bus system, a synthetic 419-bus system and a large-scale 21447-bus system demonstrate the advantages of the proposed methods in robustness and efficiency.

42 ENGINEERING↗

Carbon and Hydrocarbon Particle Seeding in Air-Breathing Rotating Detonation Engine

Within the power generation community, the rotating detonation engine (RDE) is only growing in popularity with its increased performance, simple mechanism, and operation. Although significant testing is underway to characterize the RDE for integration with conventional gas turbines, this entire system is still at a relatively low technology readiness level. In the midst of RDE research, there is an initiative to understand solid particle seeding effects in the detonation performance. Under investigation at the University of Central Florida is a Department of Energy (DOE) 15.24 cm (6 in.) RDE, with a solid particle seeder in parallel with its H2 and air flow lines. Previous work on this system involved carbon particle detonation; however, the tested particles were taken one step further to include more sustainable, greener hydrocarbon particles. Testing of powdered sugar, peanut flour, and cornstarch, along with previous carbon black tests have shown not only successful detonability, but a noticeable effect on the detonation wave dynamics. Side-by-side with a particle burning model being developed, an operational map can be determined for the hydrocarbon particles particularly, which can be tuned with the local flow conditions to achieve peak operability while replacing fuels with sustainable alternatives that could even be grown.

Engineering↗

A time-parallel method for scalable heat transfer simulations of additive manufacturing

Here, a major challenge in simulating the thermal behavior in additive manufacturing processes is the disparate length and time scales between transport phenomena occurring in the melt pool and the component. A common simulation approach relies on spatial decomposition for parallel computing, but due to the nature of heat transfer in AM, where most of the computational expenditure is localized near the melt pool, the computational speedup from spatial parallelization saturates quickly. Therefore, additional parallelism by means of time-domain decomposition is needed to fully take advantage of high-performance computing (HPC) resources. This work introduces a time-parallel method to improve the computational scalability of additive manufacturing simulations on HPC systems, while maintaining high temporal resolution of heat transfer near the melt pool. The method, inspired by the nonlinear paraexp formalism, performs an iterative superposition of nonlinear solutions to the initial value problem, integrating the heat equation across overlapping time-parallel intervals. For a single layer of the NIST AMB2018–01 L7 benchmark problem, the method achieves a 38.51x speedup in wall-clock time with a maximum error in the global temperature solution of 0.99%. This reduces the total solution time from 196.72 min to 5.11 min on 128 nodes of the ORNL Frontier supercomputer. The tradeoff between accuracy and total wall-clock time is investigated and recommendations for time-parallel deployment for AM problems are made.

Additive manufacturing↗

Domaine-Specific Runtime to Orchestrate Computation on Heterogeneous Platforms

Task-based runtime systems in the past have sought to exploit inherent asynchronicity in the application execution to reduce overall runtime. In the last decade focus shifted to supporting the heterogeneity that is increasingly prevalent in high-performance computing systems. However, existing task-based runtime systems, being general, come with a challenging set of issues such as the complexity of abstractions and overheads. And they still leave much of the burden of exposing the parallelism on the application developers who also have to fit their applications to the runtime systems’ interfaces. In this paper we take a different approach for applications to target heterogeneous systems through domain-specific run-times. Our design aspires to leverage the domain-specific knowledge of a focused class of scientific simulations to pragmatically orchestrate computations in the simulations.

97 MATHEMATICS AND COMPUTING↗