Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Analysis and design of algorithm-based fault-tolerant systems

An important consideration in the design of high performance multiprocessor systems is to ensure the correctness of the results computed in the presence of transient and intermittent failures. Concurrent error detection and correction have been applied to such systems in order to achieve reliability. Algorithm Based Fault Tolerance (ABFT) was suggested as a cost-effective concurrent error detection scheme. The research was motivated by the complexity involved in the analysis and design of ABFT systems. To that end, a matrix-based model was developed and, based on that, algorithms for both the design and analysis of ABFT systems are formulated. These algorithms are less complex than the existing ones. In order to reduce the complexity further, a hierarchical approach is developed for the analysis of large systems.

Nair, V. S. Sukumaran↗

Method development for compensating temperature effects in pressure sensitive paint measurements

Pressure sensitive luminescent paints (PSP) have recently emerged as a viable technique for aerodynamic pressure measurements. The technique uses a surface coating which contains probe molecules that luminesce when excited by light of an appropriate wavelength. The photoluminescence of these materials is known to be quenched by the presence of molecular oxygen. Since oxygen is a fixed mole fraction of the air, the coating's luminescence intensity varies inversely with air pressure. Digital imaging of the luminescence varying across a coated surface produces a pressure distribution map over that surface. One difficulty encountered with this technique is the temperature effect on the luminescence intensity. Present PSP formulations have significant sensitivity to temperature. At the moment, the most practical way of correcting for temperature effects is to calibrate the paint in place at the operating temperatures by using a few well-placed pressure taps. This study is looking at development of temperature indicating coatings that can be applied and measured concurrently with PSP, and use the temperature measurement to compute the correct pressure. Two methods for this dual paint formulation are proposed. One method will use a coating that consists of temperature sensitive phosphors in a polymer matrix. This is similar in construction to PSP, except that the probe molecules used are selected primarily for their temperature sensitivity. Both organic phosphors (e.g., europium thenoyltrifluoroacetonate, bioprobes) and inorganic phosphors (e.g., Mg4(F)GeO6:Mn, La2O2S:Eu, Radelin Type phosphors, Sylvania Type phosphors) will be evaluated for their temperature sensing potential. The next method will involve a novel coating composing of five membered heterocyclic conducting polymers which are known to show temperature dependent luminescence (e.g., poly(3-alkylthiopene), poly(3-alkylselenophene), poly(3-alkylfuran)). Both methods will involve applying a bottom layer of temperature sensitive coating followed by a top coating of PSP. An oxygen-impermeable polymer can be used as the temperature sensitive coating matrix, or it can be layered in between the coatings to prevent oxygen quenching of the bottom coating's luminescence. The probe molecules of the coatings will be excited by a broad band of light, with the different emissions detected and measured at their distinct wavelengths. Developmental research of these coatings is still in progress; however, preliminary results look very promising.

Demandante, Carlo Greg N.↗

Concurrency and discrete event control

Much of discrete event control theory has been developed within the framework of automata and formal languages. An alternative approach inspired by the theories of process-algebra as developed in the computer science literature is presented. The framework, which rests on a new formalism of concurrency, can adequately handle nondeterminism and can be used for analysis of a wide range of discrete event phenomena.

Heymann, Michael↗

Online and Scalable Data Compression Pipeline with Guarantees on Quantities of Interest

Data compression is becoming critical for data-intensive scientific applications. Scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Prior work has shown that a pipeline can be built to guarantee error on the primary data (PD) within user-defined bounds and achieve near-floating point QoI errors. In this paper, we present novel computational approaches for accelerating the pipeline and demonstrate results that enable concurrent execution of compression in parallel with the simulation nodes. This allows compression, including the writing of the required compression data, for the previous time step to be completed while the simulation proceeds with the current time step. Overall, the approach presented in this paper results in a 6–8 times improvement in computational overhead compared to previous work. These results were obtained using data generated by a large-scale fusion code called XGC, which produces hundreds of terabytes of data in a single day.

Banerjee, Tania↗

Partitioning problems in parallel, pipelined and distributed computing

The problem of optimally assigning the modules of a parallel program over the processors of a multiple computer system is addressed. A Sum-Bottleneck path algorithm is developed that permits the efficient solution of many variants of this problem under some constraints on the structure of the partitions. In particular, the following problems are solved optimally for a single-host, multiple satellite system: partitioning multiple chain structured parallel programs, multiple arbitrarily structured serial programs and single tree structured parallel programs. In addition, the problems of partitioning chain structured parallel programs across chain connected systems and across shared memory (or shared bus) systems are also solved under certain constraints. All solutions for parallel programs are equally applicable to pipelined programs. These results extend prior research in this area by explicitly taking concurrency into account and permit the efficient utilization of multiple computer architectures for a wide range of problems of practical interest.

Bokhari, S.↗

Characterizing, Modeling, and Accurately Simulating Power and Energy Consumption of I/O-intensive Scientific Workflows

While distributed computing infrastructures can provide infrastructure-level techniques for managing energy consumption, application-level energy consumption models have also been developed to support energy-efficient scheduling and resource provisioning algorithms. In this work, we analyze the accuracy of a widely-used application-level model that has been developed and used in the context of scientific workflow executions. To this end, we profile two production scientific workflows on a distributed platform instrumented with power meters. We then conduct an analysis of power and energy consumption measurements. This analysis shows that power consumption is not linearly related to CPU utilization and that I/O operations significantly impact power, and thus energy, consumption. We then propose a power consumption model that accounts for I/O operations, including the impact of waiting for these operations to complete, and for concurrent task executions on multi-socket, multi-core compute nodes. We implement our proposed model as part of a simulator that allows us to draw direct comparisons between real-world and modeled power and energy consumption. Here, we find that our model has high accuracy when compared to real-world executions. Furthermore, our model improves accuracy by about two orders of magnitude when compared to the traditional models used in the energy-efficient workflow scheduling literature.

97 MATHEMATICS AND COMPUTING↗

Fault-Tolerant, Real-Time, Multi-Core Computer System

A document discusses a fault-tolerant, self-aware, low-power, multi-core computer for space missions with thousands of simple cores, achieving speed through concurrency. The proposed machine decides how to achieve concurrency in real time, rather than depending on programmers. The driving features of the system are simple hardware that is modular in the extreme, with no shared memory, and software with significant runtime reorganizing capability. The document describes a mechanism for moving ongoing computations and data that is based on a functional model of execution. Because there is no shared memory, the processor connects to its neighbors through a high-speed data link. Messages are sent to a neighbor switch, which in turn forwards that message on to its neighbor until reaching the intended destination. Except for the neighbor connections, processors are isolated and independent of each other. The processors on the periphery also connect chip-to-chip, thus building up a large processor net. There is no particular topology to the larger net, as a function at each processor allows it to forward a message in the correct direction. Some chip-to-chip connections are not necessarily nearest neighbors, providing short cuts for some of the longer physical distances. The peripheral processors also provide the connections to sensors, actuators, radios, science instruments, and other devices with which the computer system interacts.

Gostelow, Kim P.↗

Domain decomposition for aerodynamic and aeroacoustic analyses, and optimization

The overarching theme was the domain decomposition, which intended to improve the numerical solution technique for the partial differential equations at hand; in the present study, those that governed either the fluid flow, or the aeroacoustic wave propagation, or the sensitivity analysis for a gradient-based optimization. The role of the domain decomposition extended beyond the original impetus of discretizing geometrical complex regions or writing modular software for distributed-hardware computers. It induced function-space decompositions and operator decompositions that offered the valuable property of near independence of operator evaluation tasks. The objectives have gravitated about the extensions and implementations of either the previously developed or concurrently being developed methodologies: (1) aerodynamic sensitivity analysis with domain decomposition (SADD); (2) computational aeroacoustics of cavities; and (3) dynamic, multibody computational fluid dynamics using unstructured meshes.

Baysal, Oktay↗

Application of a distributed network in computational fluid dynamic simulations

A general-purpose 3-D, incompressible Navier-Stokes algorithm is implemented on a network of concurrently operating workstations using parallel virtual machine (PVM) and compared with its performance on a CRAY Y-MP and on an Intel iPSC/860. The problem is relatively computationally intensive, and has a communication structure based primarily on nearest-neighbor communication, making it ideally suited to message passing. Such problems are frequently encountered in computational fluid dynamics (CDF), and their solution is increasingly in demand. The communication structure is explicitly coded in the implementation to fully exploit the regularity in message passing in order to produce a near-optimal solution. Results are presented for various grid sizes using up to eight processors.

Deshpande, Manish↗

Fault-tolerant software for aircraft control systems

Concepts for software to implement real time aircraft control systems on a centralized digital computer were discussed. A fault tolerant software structure employing functionally redundant routines with concurrent error detection was proposed for critical control functions involving safety of flight and landing. A degraded recovery block concept was devised to allow collocation of critical and noncritical software modules within the same control structure. The additional computer resources required to implement the proposed software structure for a representative set of aircraft control functions were discussed. It was estimated that approximately 30 percent more memory space is required to implement the total set of control functions. A reliability model for the fault tolerant software was described and parametric estimates of failure rate were made.

Source record↗

Asynchronous Communication Scheme For Hypercube Computer

Scheme devised for asynchronous-message communication system for Mark III hypercube concurrent-processor network. Network consists of up to 1,024 processing elements connected electrically as though were at corners of 10-dimensional cube. Each node contains two Motorola 68020 processors along with Motorola 68881 floating-point processor utilizing up to 4 megabytes of shared dynamic random-access memory. Scheme intended to support applications requiring passage of both polled or solicited and unsolicited messages.

Madan, Herb S.↗

Self-checking self-repairing computer nodes using the mirror processor

Circuitry added to fault-tolerant systems for concurrent error deduction usually reduces performance. Using a technique called micro rollback, it is possible to eliminate most of the performance penalty of concurrent error detection. Error detection is performed in parallel with intermodule communication, and erroneous state changes are later undone. The author reports on the design and implementation of a VLSI RISC microprocessor, called the Mirror Processor (MP), which is capable of micro rollback. In order to achieve concurrent error detection, two MP chips operate in lockstep, comparing external signals and a signature of internal signals every clock cycle. If a mismatch is detected, both processors roll back to the beginning of the cycle when the error occurred. In some cases the erroneous state is corrected by copying a value from the fault-free processor to the faulty processor. The architecture, microarchitecture, and VLSI implementation of the MP, emphasizing its error-detection, error-recovery, and self-diagnosis capabilities, are described.

Tamir, Yuval↗

Anatomical and hydraulic responses to desiccation in emergent conifer seedlings

Premise The young seedling life stage is critical for reforestation after disturbance and for species migration under climate change, yet little is known regarding their basic hydraulic function or vulnerability to drought. Here, we sought to characterize responses to desiccation including hydraulic vulnerability, xylem anatomical traits, and impacts on other stem tissues that contribute to hydraulic functioning. Methods Larix occidentalis , Pseudotsuga menziesii , and Pinus ponderosa (all ≤6 weeks old) were imaged using x‐ray computed microtomography during desiccation to assess seedling biomechanical responses with concurrently measured hydraulic conductivity ( k s ) and water potential ( Ψ ) to assess vulnerability to xylem embolism formation and other tissue damage. Results In non‐stressed samples for all species, pith and cortical cells appeared circular and well hydrated, but they started to empty and deform with decreasing Ψ which resulted in cell tearing and eventual collapse. Despite the severity of this structural damage, the vascular cambium remained well hydrated even under the most severe drought. There were significant differences among species in vulnerability to xylem embolism formation, with 78% xylem embolism in L. occidentalis by Ψ of −2.1 MPa, but only 47.7% and 62.1% in P. ponderosa and P. menziesii at −4.27 and −6.73 MPa, respectively. Conclusions Larix occidentalis seedlings appeared to be more susceptible to secondary xylem embolism compared to the other two species, but all three maintained hydration of the vascular cambium under severe stress, which could facilitate hydraulic recovery by regrowth of xylem when stress is relieved.

Miller, Megan L.↗

From anti-Arrhenius to Arrhenius behavior in a dislocation-obstacle bypass: Atomistic simulations and theoretical investigation

Dislocations are the primary carriers of plasticity in metallic materials. Understanding the basic mechanisms for dislocation movement is paramount to predicting the material mechanical response. Relying on atomistic simulations, we observe a transition from non-Arrhenius to Arrhenius behavior in the rate of an edge dislocation overcoming the long-range elastic interaction with a prismatic loop in tungsten. Close to the critical resolved shear stress, the process shows a non-Arrhenius behavior at low temperatures. However, as the temperature increases, the activation entropy starts to dominate, leading to a traditional Arrhenius-like behavior. Here, we have computed the activation entropy analytically along the minimum energy path following Schoeck’s method, which captures the cross-over between anti-Arrhenius and Arrhenius domains. Also, the Projected Average Force Integrator (PAFI), another simulation method to compute free energies along an initial transition path, exhibits considerable concurrence with Schoeck’s formalism. We conclude that entropic effects need to be considered to understand processes involving dislocations bypassing elastic barriers close to the critical resolved shear stress. More work needs to be performed to fully understand the discrepancies between Schoeck’s and PAFI compared to molecular dynamics.

36 MATERIALS SCIENCE↗

Performance of Heterogeneous Algorithm Scheduling in CMSSW

The CMS experiment started to utilize Graphics Processing Units (GPU) to accelerate the online reconstruction and event selection running on its High Level Trigger (HLT) farm in the 2022 data taking period. The projections of the HLT farm to the High-Luminosity LHC foresee a significant use of compute accelerators in the LHC Run 4 and onwards in order to keep the cost, size, and power budget of the farm under control. This direction of leveraging compute accelerators has synergies with the increasing use of HPC resources in HEP computing, as HPC machines are employing more and more compute accelerators that are predominantly GPUs today. In this work we review the features developed for the CMS data processing framework, CMSSW, to support the effective utilization of both compute accelerators and many-core CPUs within a highly concurrent task-based framework. We measure the impact of various design choices for the scheduling of heterogeneous algorithms on the event processing throughput, using the Run-3 HLT application as a realistic use case.

Bocci, Andrea↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

Some Effects of Bluntness on Boundary-Layer Transition and Heat Transfer at Supersonic Speeds

Large downstream movements of transition observed when the leading edge of a hollow cylinder or a flat plate is slightly blunted are explained in terms of the reduction in Reynolds number at the outer edge of the boundary layer due to the detached shock wave. The magnitude of this reduction is computed for cones and wedges for Mach numbers to 20. Concurrent changes in outer-edge Mach number and temperature occur in the direction that would increase the stability of the laminar boundary layer. The hypothesis is made that transition Reynolds number is substantially unchanged when a sharp leading edge or tip is blunted. This hypothesis leads to the conclusion that the downstream movement of transition is inversely proportional to the ratio of surface Reynolds number with blunted tip or leading edge to surface Reynolds number with sharp tip or leading edge. The conclusion is in good agreement with the hollow-cylinder result at Mach 3.1.

Moeckel, W E↗

Parallelization and visual analysis of multidimensional fields: Application to ozone production, destruction, and transport in three dimensions

Atmospheric modeling is a grand challenge problem for several reasons, including its inordinate computational requirements and its generation of large amounts of data concurrent with its use of very large data sets derived from measurement instruments like satellites. In addition, atmospheric models are typically run several times, on new data sets or to reprocess existing data sets, to investigate or reinvestigate specific chemical or physical processes occurring in the earth's atmosphere, to understand model fidelity with respect to observational data, or simply to experiment with specific model parameters or components.

Schwan, Karsten↗