Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq↗

An asynchronous parallel high-throughput model calibration framework for crystal plasticity finite element constitutive models

Crystal plasticity finite element model (CPFEM) is a powerful numerical simulation in the integrated computational materials engineering toolboxes that relates microstructures to homogenized materials properties and establishes the structure–property linkages in computational materials science. However, to establish the predictive capability, one needs to calibrate the underlying constitutive model, verify the solution and validate the model prediction against experimental data. Bayesian optimization (BO) has stood out as a gradient-free efficient global optimization algorithm that is capable of calibrating constitutive models for CPFEM. Here in this paper, we apply a recently developed asynchronous parallel constrained BO algorithm to calibrate phenomenological constitutive models for stainless steel 304 L, Tantalum, and Cantor high-entropy alloy.

304L stainless steel↗

Scheduling and Performance of Asynchronous Tasks in Fortran 2018 with FEATS

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP (Hermanns in Parallel programming in Fortran 95 using openMP, 2002. School of Aeronautical Engineering, Universidad Politécnica de Madrid, España, 2011), explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI) (in A message-passing interface standard version 4.0, 2021. https://www.mpi-forum.org/docs/mpi-4.0/mpi40-report.pdf), or compiler-specific language extensions such as those provided by CUDA (Ruetsch and Fatica in CUDA Fortran for scientists and engineers: best practices for efficient CUDA Fortran programming, Elsevier, 2013). By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models (Numrich in Parallel programming with co-arrays, CRC Press, 2018, and Curcic in Modern Fortran: building efficient parallel applications, Manning Publications, 2020). Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. Further, the paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

97 MATHEMATICS AND COMPUTING↗

Asynchronous aging and turnover of human circulating and tissue-resident memory T cells across sites

Memory T cells are maintained in tissues as circulating effector-memory (T EM ) and tissue-resident (T RM ) populations for protective immunity, though the role of site and subset in memory persistence remains undefined. Here, in this work, we investigated age-associated dynamics of human T cells in lymphoid organs, mucosal sites, and blood over 10 decades of life using retrospective radiocarbon ( 14 C) birth dating, along with cellular, transcriptome, and epigenetic profiling. Memory T cells across peripheral sites exhibited continuous turnover with mean lifespans of 1–2 years, while the spleen contained longer-lived T cells. Over age, T EM cells expressed senescent markers and a GZMK transcriptional signature, while T RM cells maintained site-specific resident phenotypes without exhibiting features of senescence. Both T EM and T RM cells showed age-associated DNA hypomethylation, though T RM cells exhibited more epigenetically regulated genes. Together, our findings reveal asynchronous aging of human memory T cells by subset and site, as well as persistence of T RM cells without immunosenescence.

T cells↗

Hybrid additive manufacturing of AISI 316L via asynchronous powder and hot-wire laser directed energy deposition

Hybrid Additive Manufacturing (AM) offers a way to leverage the advantages of different AM technologies, enabling the efficient production of sizeable parts without compromising material properties or geometric complexity capabilities. This study presents an asynchronous hybrid Directed Energy Deposition (DED) strategy employing laser powder DED and laser hot-wire DED. AISI 316L parts comprising multiple powder and wire segments were fabricated with optional machining on AISI 316L substrates to investigate how quality is impacted by (i) alternative process sequences (laser powder DED followed by laser hot-wire DED and vice versa), (ii) machined vs. as-printed interfacial conditions, and (iii) material deposition on top vs. alongside previously built segments. Optical microscopy, X-ray computed tomography, and Vickers hardness were used to characterize the morphology and microstructure of the parts, localized porosity and lack of fusion defects, bulk density, and mechanical properties. Interfacial machining was necessary for dimensional control but promoted lack of fusion voids, resulting in a 99.71 ± 0.01% dense part. As-printed interfaces resulted in a denser part (99.82 ± 0.02%) at the expense of dimensional accuracy. The hardness of the parts with as-printed and machined interfaces was 196 ± 0.37 HV and 192 ± 0.40 HV, respectively, compared to 156 ± 1.4 HV for the substrate. Depositing powder alongside or on top of wire sections resulted in interfaces with a hardness of 217 ± 2.2 HV, compared to 185 ± 3.4 HV for the wire-powder interfaces.

36 MATERIALS SCIENCE↗

Hydrodynamic spin-orbit coupling in asynchronous optically driven micro-rotors

Abstract Vortical flows of rotating particles describe interactions ranging from molecular machines to atmospheric dynamics. Yet to date, direct observation of the hydrodynamic coupling between artificial micro-rotors has been restricted by the details of the chosen drive, either through synchronization (using external magnetic fields) or confinement (using optical tweezers). Here we present a new active system that illuminates the interplay of rotation and translation in free rotors. We develop a non-tweezing circularly polarized beam that simultaneously rotates hundreds of silica-coated birefringent colloids. The particles rotate asynchronously in the optical torque field while freely diffusing in the plane. We observe that neighboring particles orbit each other with an angular velocity that depends on their spins. We derive an analytical model in the Stokes limit for pairs of spheres that quantitatively explains the observed dynamics. We then find that the geometrical nature of the low Reynolds fluid flow results in a universal hydrodynamic spin-orbit coupling. Our findings are of significance for the understanding and development of far-from-equilibrium materials.

47 OTHER INSTRUMENTATION↗

Asynchronous domain dynamics and equilibration in layered oxide battery cathode

To improve lithium-ion battery technology, it is essential to probe and comprehend the microscopic dynamic processes that occur in a real-world composite electrode under operating conditions. The primary and secondary particles are the structural building blocks of battery cathode electrodes. Their dynamic inconsistency has profound but not well-understood impacts. In this research, we combine operando coherent multi-crystal diffraction and optical microscopy to examine the chemical dynamics in local domains of layered oxide cathode. Our results not only pinpoint the asynchronicity of the lithium (de)intercalation at the sub-particle level, but also reveal sophisticated diffusion kinetics and reaction patterns, involving various localized processes, e.g., chemical onset, reaction front propagation, domains equilibration, particle deformation and motion. These observations shed new lights onto the activation and degradation mechanisms of state-of-the-art battery cathode materials.

25 ENERGY STORAGE↗

Asynchronous x-ray multiprobe data acquisition for x-ray transient absorption spectroscopy

Laser pump X-ray Transient Absorption (XTA) spectroscopy offers unique insights into photochemical and photophysical phenomena. X-ray Multiprobe data acquisition (XMP DAQ) is a technique that acquires XTA spectra at thousands of pump-probe time delays in a single measurement, producing highly self-consistent XTA spectral dynamics. In this work, we report two new XTA data acquisition techniques that leverage the high performance of XMP DAQ in combination with High Repetition Rate (HRR) laser excitation: HRR-XMP and Asynchronous X-ray Multiprobe (AXMP). HRR-XMP uses a laser repetition rate up to 200 times higher than previous implementations of XMP DAQ and proportionally increases the data collection efficiency at each time delay. This allows HRR-XMP to acquire more high-quality XTA data in less time. AXMP uses a frequency mismatch between the laser and x-ray pulses to acquire XTA data at a flexibly defined set of pump-probe time delays with a spacing down to a few picoseconds. AXMP introduces a novel pump-probe synchronization concept that acquires data in clusters of time delays. Further, the temporally inhomogeneous distribution of acquired data improves the attainable signal statistics at early times, making the AXMP synchronization concept useful for measuring sub-nanosecond dynamics with photon-starved techniques like XTA. In this paper, we demonstrate HRR-XMP and AXMP by measuring the laser-induced spectral dynamics of dilute aqueous solutions of Fe(CN) 6 4₋ and [Fe II (bpy) 3 ] 2+ (bpy: 2,2'-bipyridine), respectively.

47 OTHER INSTRUMENTATION↗

Quantum gates on asynchronous atomic excitations

A method for realising a universal system of quantum gates based on asynchronous excitations of two-level atoms in optical cavities is proposed. The entangling operator of the CSign type is implemented without beam splitters, approximately, using the incommensurability of the Rabi oscillation periods in a cavity with single and double excitations. (quantum technology)

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Lamellar: A Rust-based Asynchronous Tasking and PGAS Runtime for High Performance Computing

Cybersecurity is one of the largest concerns in modern computing, impacting and dictating how governments, private corporations, and individuals interact with and live in an increasingly digital world. The NSA has recently released a memo [ 1] on “Software Memory Safety” where they highlight that both Microsoft and Google have stated around 70% of software vulnerabilities were due to memory safety issues. Although languages such as C and C++ provide freedom and flexibility with memory management, guaran- teeing safety falls mostly on the developer. The NSA recommends using “memory safe” languages whenever possible. In this paper we introduce Lamellar, an asynchronous tasking and PGAS HPC runtime written in Rust, one such "memory safe" language. We describe the entire Lamellar stack, from network interfaces to high- level abstractions such as distributed LamellarArrays and Active Messages. We conclude by showing comparable performance to legacy PGAS runtimes (e.g. OpenSHMEM) on a subset of the BALE kernel suite while maintaining strong memory safety principles.

HPC Software Systems, Rust Programming Language, P↗

Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis

We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. Here, we demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.

97 MATHEMATICS AND COMPUTING↗

Asynchronous Truncated Multigrid-Reduction-in-Time

In this paper, we present the new “asynchronous truncated multigrid-reduction-in-time” (AT-MGRIT) algorithm for introducing time parallelism to the solution of discretized time-dependent problems. The new algorithm is based on the multigrid-reduction-in-time (MGRIT) approach, which, in certain settings, is equivalent to another common multilevel parallel-in-time method, Parareal. In contrast to Parareal and MGRIT that both consider a global temporal grid over the entire time interval on the coarsest level, the AT-MGRIT algorithm uses truncated local time grids on the coarsest level, each grid covering certain temporal subintervals. Further, these local grids can be solved completely in an independent way from each other, which reduces the sequential part of the algorithm and, thus, increases parallelism in the method. Here, we study the effect of using truncated local coarse grids on the convergence of the algorithm, both theoretically and numerically, and show, using challenging nonlinear problems, that the new algorithm consistently outperforms classical Parareal/MGRIT in terms of time to solution.

97 MATHEMATICS AND COMPUTING↗

A Surrogate-Based Asynchronous Decomposition Technique for Realistic Security-Constrained Optimal Power Flow Problems

Here we present a decomposition approach for obtaining good feasible solutions for the security-constrained, alternating-current, optimal power flow (SC-AC-OPF) problem at an industrial scale and under real-world time and computational limits. The approach was designed while preparing and participating in ARPA-E’s Grid Optimization Competition (GOC) Challenge 1. The challenge focused on a near-real-time version of the SC-AC-OPF problem, where a base operating point is optimized, taking into account possible single-element contingencies, after which the system adapts its operating point following the response of automatic frequency droop controllers and voltage regulators. Our solution approach for this problem relies on state-of-the-art nonlinear programming algorithms, and it employs nonconvex relaxations for complementarity constraints, a specialized two-stage decomposition technique with sparse approximations of recourse terms and contingency ranking and prescreening. The paper describes and justifies our approach and outlines the features of its implementation, including functions and derivatives evaluation, warm-starting strategies, and asynchronous parallelism. We discuss the results of the independent benchmark of our approach by ARPA-E’s GOC team in Challenge 1, where it was found to consistently produce high-quality solutions across a wide range of network sizes and difficulty, and conclude by outlining future extensions of the approach.

97 MATHEMATICS AND COMPUTING↗

Asynchronous Ballistic Reversible Computing using Superconducting elements

Computing uses energy. At the bare minimum, erasing information in a computer increases the entropy. Landauer has calculated %7E k B T log(2) Joules is dissipated per bit of energy erased. While the success of Moores law has allowed increasing computing power and efficiency for many years, these improvements are coming to an end. This project asks if there is a way to continue those gains by circumventing Landauer through reversible computing. We explore a new reversible computing paradigm, asynchronous ballistic reversible computing or ABRC. The ballistic nature of data in ABRC matches well with superconductivity which provides a low-loss environment and a quantized bit encoding the fluxon. We discuss both these and our development of a superconducting fabrication process at Sandia. We describe a fully reversible 1-bit memory cell based on fluxon dynamics. Building on this model, we propose several other gates which may also offer reversible operation.

97 MATHEMATICS AND COMPUTING↗

Optimization of Asynchronous Communication Operations through Eager Notifications

UPC++ is a C++ library implementing the Asynchronous Partitioned Global Address Space (APGAS) model. We propose an enhancement to the completion mechanisms of UPC++ used to synchronize communication operations that is designed to reduce overhead for on-node operations. Our enhancement permits eager delivery of completion notification in cases where the data transfer semantics of an operation happen to complete synchronously, for example due to the use of shared-memory bypass. This semantic relaxation allows removing significant overhead from the critical path of the implementation in such cases. We evaluate our results on three different representative systems using a combination of microbenchmarks and five variations of the the HPCChallenge RandomAccess benchmark implemented in UPC++ and run on a single node to accentuate the impact of locality. We find that in RMA versions of the benchmark written in a straightforward manner (without manually optimizing for locality), the new eager notification mode can provide up to a 25% speedup when synchronizing with promises and up to a 13.5x speedup when synchronizing with conjoined futures. We also evaluate our results using a graph matching application written with UPC++ RMA communication, where we measure overall speedups of as much as 11% in single-node runs of the unmodified application code, due to our transparent enhancements.

Kamil, Amir↗

Framework for Extensible, Asynchronous Task Scheduling (FEATS) in Fortran

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP, explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI), or compiler-specific language extensions such as those provided by CUDA. By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models. Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data-parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. The paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task-scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

Richardson, Brad↗