Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

An interface-controlled Mott memristor in α -RuCl 3

Memristor devices have history-dependent charge transport properties that are ideal for neuromorphic computing applications. We reveal a memristor material and mechanism in the layered Mott insulator α-RuCl 3 . The pinched hysteresis loops and S-shaped negative differential resistance in bulk crystals verify memristor behavior and are attributed to a nonlinear coupling between charge injection over a Schottky barrier at the electrical contacts and concurrent Joule heating. Direct simulations of this coupling can reproduce the device characteristics.

Physics↗

Multi-Core Microcontroller Hardware In the Loop System for Electric Machine Control

Hardware in the Loop (HIL) is a simulation technique used to reduce the software development cycle and test control systems in a non-destructive environment. This work describes a cost effective HIL simulator on a dual core microcontroller in which one core acts as a controller and the other emulates the system under control. The emulator runs one step per Pulse Width Modulation (PWM) period in real time. To handle the computational burden and prioritize execution of simulation and control tasks, an interrupt-based software architecture with task prioritization has been developed. As a demonstration, the HIL has been implemented on a Texas Instruments TMS320F28379D dual core microcontroller, which emulates a Permanent Magnet Synchronous Machine (PMSM) with resolver feedback. Hardware peripherals are developed and tested concurrently with the control system, providing higher confidence in the software. By using the peripherals in the HIL development, the controller exercises either the HIL emulation or a pin compatible PMSM testbench. To quantify performance and validate the processor based emulator, the HIL results are compared to the preexisting testbench for accuracy benchmarking at no-load and under load for a range of operating points.

33 ADVANCED PROPULSION SYSTEMS↗

Computational modeling of coupled mechanical damage and electrochemistry in ternary oxide composite electrodes

Performance degradation of ternary layered oxide cathodes largely originates from their loss of structural integrity in cyclic usage. Mechanical damage, such as intergranular fracture of the active particles, is not only a mechanical cleavage process but also interferes with electrochemical kinetics such as infiltration of liquid electrolyte, surface corrosion of the constituent primary particles, and may eventually isolate the primary grains from the electron conducting network. Here, in this work, we develop a computational framework that integrates electrochemistry of a LiNi x Mn y Co 1−x−y O 2 (NMC) composite cathode with mechanical damage of the active particles. To fully examine the intricate chemomechanical behavior of the electrode, we evaluate the effects of the anisotropic material properties, the influence of mechanical potential on Li transport, and the concurrent intergranular fracture and electrolyte penetration along the grain boundaries upon multiple cycles. Electrolyte infiltration benefits capacity retention but aggravates further mechanical damage by corrosion. Structural failure mostly occurs in the first charging due to the anisotropic mechanical strain between the primary grains, while the resulting damage remains stable in the later few cycles. The results are consistent with experimental observations and the integration of electrochemistry and mechanical failure enables a step further understanding of the complex mechanism of battery degradation.

Battery degradation↗

A Pd III Sulfate Dimer Initiates Rapid Methane Monofunctionalization by H Atom Abstraction

An electrogenerated Pd III 2 species in fuming sulfuric acid is competent for rapid and concurrent methane monohydroxylation to methyl bisulfate (CH 3 OSO 3 H) and methane sulfonation to methanesulfonic acid (CH 3 SO 3 H). In situ NMR at 50 °C is used to track methane transformation exclusively to CH 3 OSO 3 H and CH 3 SO 3 H at high conversions. Integrating a set of kinetic and computational studies, the mechanism of methane monofunctionalization by Pd III 2 is examined. Here, experimental rate laws and common kinetic isotope effects for CH 3 OSO 3 H and CH 3 SO 3 H formation suggest that both transformations proceed via a common rate-limiting C-H activation step. Introduction of O 2 or Pd II,III 2 suppresses CH 3 SO 3 H generation, indicating a radical chain sequence. Although the metal-metal bonded Pd III 2 complex is a net two-electron oxidant, our aggregate kinetic data point to a mechanistic model that features rate-limiting H atom abstraction by the Pd III 2 complex to generate a methyl radical intermediate. The CH 3 • intermediate then recombines with Pd II,III 2 to furnish a CH 3 Pd III 2 intermediate that reductively eliminates CH 3 OSO 3 H. Alternatively, the CH 3 intermediate can enter a chain reaction with SO 3 to generate CH 3 SO 3 H. DFT computations support the radical-based C-H activation by Pd III 2 and delineate H atom abstraction pathways with computed reaction barriers and kinetic isotope effects (KIEs) that are consistent with experimental data. These mechanistic investigations challenge the paradigm of electrophilic C-H activation and highlight H atom abstraction as a potent pathway for selective methane C-H oxidative functionalization at high reaction rates.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Rahasak—Scalable blockchain architecture for enterprise applications

Blockchain-based decentralized infrastructure has been adapted in various industries to handle the sensitive data in a privacy-preserving manner without trusting third parties. However, integrating state-of-the-art blockchain platforms with the scalable, enterprise-level applications result in several challenges. Current blockchain platforms do not support high transaction throughput, lack high scalability, and cannot provide real-time transaction processing and back-pressure operation handling in high transaction throughput applications(e.g Big data, IoT). In this paper, we propose a novel permissioned blockchain platform “Rahasak” for highly scalable, enterprise applications. Rahasak blockchain adopts the Apache Kafka-based consensus on top of a “Validate-Execute-Group” blockchain architecture to handle realtime transaction execution on the blockchain. The architecture is equipped with a functional programming and actor-based smart contract platform that enables concurrent execution of transactions in the blockchain. Rahasak supports high transaction throughput, high scalability, concurrent transaction execution, data analytics features. Finally, with Rahasak, we make blockchain more scalable, secure, structured and meaningful for further data analytics.

97 MATHEMATICS AND COMPUTING↗

Enriched immersed finite element and isogeometric analysis: algorithms and data structures

Immersed finite element methods provide a convenient analysis framework for problems involving geometrically complex domains, such as those found in topology optimization and microstructures for engineered materials. However, their implementation remains a major challenge due to, among other things, the need to apply nontrivial stabilization schemes and generate custom quadrature rules. This article introduces the robust and computationally efficient algorithms and data structures comprising an immersed finite element preprocessing framework. The input to the preprocessor consists of a background mesh and one or more geometries defined on its domain. The output is structured into groups of elements with custom quadrature rules formatted such that common finite element assembly routines may be used without or with only minimal modifications. The key to the preprocessing framework is the construction of material topology information, concurrently with the generation of a quadrature rule, which is then used to perform enrichment and generate stabilization rules. While the algorithmic framework applies to a wide range of immersed finite element methods using different types of meshes, integration, and stabilization schemes, the preprocessor is presented within the context of the extended isogeometric analysis. This method utilizes a structured B-spline mesh, a generalized Heaviside enrichment strategy considering the material layout within individual basis functions’ supports, and face-oriented ghost stabilization. Using a set of examples, the effectiveness of the enrichment and stabilization strategies is demonstrated alongside the preprocessor’s robustness in geometric edge cases. Additionally, the performance and parallel scalability of the implementation are evaluated.

Computer implementation↗

Watching the Watchers with Verified Formal-Assurance Tools (Abbreviated Final Report)

The “Watching the Watchers” project studied the problem of establishing assurance cases for tools that are used to assure other things. Specifically, we were interested in understanding the tools and techniques one could apply to software to build an assurance case to evaluate their applicability, difficulty, level of assurance provided, and scalability. To do so we chose a set of use cases of relevance to LLNL and our various DOE and non-DOE partners and developed demonstrators to perform this evaluation. Our key focal point was around additive manufacturing problems and assurance gaps that we identified in the additive manufacturing workflow from start to completion. We also explored other areas related to AI, data analysis, and concurrent programming. Follow-on research is planned to take our prototypes from this project and adapt and mature them to fit LLNL mission applications.

97 MATHEMATICS AND COMPUTING↗

AFRL Additive Manufacturing Modeling Series: Challenge 4, In Situ Mechanical Test of an IN625 Sample with Concurrent High-Energy Diffraction Microscopy Characterization

In this work, we describe 3D characterization of an additively manufactured Inconel 625 nickel-base superalloy specimen conducted during a uniaxial tension test using a suite of nondestructive x-ray techniques. High-energy diffraction microscopy in both near- and far-field modalities are employed in situ to track evolution of the material orientation and stress–strain fields at six points during the mechanical test, and these data streams are registered with micro-computed tomography reconstructions which probe the material density. This data volume was matched to a multi-modal serial sectioning characterization of the specimen taken after loading, described in this article’s companion. Twenty-eight grains which were monitored throughout the experiment were selected to form the basis for AFRL AM Modeling Series Challenge 4, Microscale Structure-to-Properties.

36 MATERIALS SCIENCE↗

A fast, dense Chebyshev solver for electronic structure on GPUs

Matrix diagonalization is almost always involved in computing the density matrix needed in quantum chemistry calculations. In the case of modest matrix sizes (≲4000), performance of traditional dense diagonalization algorithms on modern GPUs is underwhelming compared to the peak performance of these devices. This motivates the exploration of alternative algorithms better suited to these types of architectures. We newly derive, and present in detail, an existing Chebyshev expansion algorithm whose number of required matrix multiplications scales with the square root of the number of terms in the expansion. Focusing on dense matrices of modest size, our implementation on GPUs results in large speed ups when compared to diagonalization. Additionally, we improve upon this existing method by capitalizing on the inherent task parallelism and concurrency in the algorithm. Furthermore, this improvement is implemented on GPUs by using CUDA and HIP streams via the MAGMA library and leads to a significant speed up over the serial-only approach for smaller (≲1000) matrix sizes. Finally, we apply our technique to a model system with a high density of states around the Fermi level, which typically presents significant challenges.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]↗

Distributed Quantum Learning with co-Management in a Multi-tenant Quantum System

The rapid advancement of quantum computing has pushed classical designs into the quantum domain, breaking physical boundaries for computing-intensive and data-hungry applications with the hope that some systems may provide a quantum speedup. For example, variational quantum algorithms have been proposed for quantum neural networks to train deep learning models on qubits, achieving promising results. Existing quantum learning architectures and systems rely on single, monolithic quantum machines with abundant and stable resources, such as qubits. However, fabricating a large, monolithic quantum device is considerably more challenging than producing an array of smaller devices. In this paper, we investigate a distributed quantum system that combines multiple quantum machines into a unified system. We propose DQuLearn, which divides a quantum learning task into multiple subtasks. Each subtask can be executed distributively on individual quantum machines, with the results looping back to classical machines for subsequent training iterations. Additionally, our system supports multiple concurrent clients and dynamically manages their circuits according to the runtime status of quantum workers. Through extensive experiments, we demonstrate that DQuLearn achieves similar accuracies with significant runtime reduction, by up to 68.7% and an increase per-second circuit processing speed, by up to 3.99 times, in a 4-worker multi-tenant setting.

quantum computing↗

Tikiri—Towards a lightweight blockchain for IoT

Internet of Things (IoT) platforms have been deployed in several domains to enhance efficiency of business process and improve productivity. Most IoT platforms comprise of heterogeneous software and hardware components which can potentially introduce security and privacy challenges. Blockchain technology has been proposed as one of the solutions to realize IoT security by leveraging the (a) Immutable ledger, (b) Decentralized architecture and (c) Strong cryptography primitives. However, integrating blockchain platforms with IoT based applications presents several challenges due to lack of (a) acceptable performance on resource-constrained devices, (b) high transaction throughput, (c) keyword-based search and retrieve, (d) transaction back pressure operations, and (e) real-time response. In this paper, we propose a lightweight blockchain platform, “Tikiri”, for resource-constrained IoT devices. Tikiri uses Apache Kafka for the consensus and proposes new blockchain architecture to handle real-time transaction execution on the blockchain. Tikiri is characterized by functional programming and actor-based smart contract platform that realizes concurrent execution of transactions in the blockchain. Tikiri realizes a lightweight and scalable blockchain that can provides performance on the resource-constrained IoT devices.

97 MATHEMATICS AND COMPUTING↗

Photoresponsive Rotaxanes Switch Lipid Bilayer Neuromorphic Behavior with Light

A rotaxane consisting of a macrocycle ring with two azobenzene units mechanically interlocked onto a bolaamphiphilic axle was incorporated into droplet interface bilayers (DIBs). The azobenzene groups on the ring underwent quasi-reversible, photoisomerization-induced cycling between 1-E and 1-Z configurations when irradiated with 370 and 467 nm light, respectively, enabling programmable access to different history-dependent electrical behaviors from the same membrane. In the 1-E configuration, bilayers exhibited type-IIactive memristance that coincided with increasingly elevated ionic conduction, associated with progressively enhanced bilayer permeability during voltage cycling. In the 1-Z configuration, bilayers displayed type-I, passive memcapacitive behavior, reflecting tighter lipid packing and reduced ionic permeability. Photoswitching also yielded a nonvolatile, photoresponsive memcapacitor that could be modulated repetitively with negligible loss, likely via reversible changes in membrane thickness. Concurrent ohmic leakage currents across the membrane were less than 0.3%. These results agree with previous studies with increasing membrane permeability using photoswitchable rotaxanes and provide new insights into the coupling between volatile and nonvolatile memcapacitance during photoisomerization. More broadly, they demonstrate a new strategy for the manipulation of neuromorphic behaviors in soft materials using light, with implications for brain-inspired computation and sensing.

droplet interface bilayers↗

Early Inference of Nuclear Technology-Directed Research Activities of Authors from Scientific Publications

Nuclear research articles can provide information about early nuclear proliferation indicators such as influential research entities and technology capability levels of a country, but detection of nuclear activities typically occurs after they have started. We investigate the extent to which nuclear research articles can be used to infer whether a research entity will acquire or develop a nuclear technology before it happens. Early detection of nuclear proliferation or technology development indicators from data is challenging due to partial observability, sparse and unlabeled information, and confounding signals from multiple concurrent activities. This paper presents the early detection problem as a sequential decision-making, goal inference problem, where the objective is to characterize and predict an individual’s, organization’s, or a country’s intent (unobserved goal-directed behavior) towards developing a nuclear capability from partially observed sequences of their research publications, using inverse reinforcement learning and Bayesian goal inference methods. A computational framework is presented, and its application demonstrated using 29,196 Scopus records for a case study related to a civil nuclear capability. The case study results serve as a proof-of-concept demonstration for inference of technology-directed research activity of authors who publish in the nuclear domain. The inference method, combined with advanced computing, may be used to assess and monitor activities pertaining to early developmental stages of a nuclear technology or capability, which in turn can help to identify and prioritize activities with nuclear proliferation potential for further investigation.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Fiats: Functional inference and training for surrogates

Fiats provides a platform for research on the training and deployment of neural-network surrogate models for computational science. Fiats also supports exploring, advancing, and combining functional, object-oriented, and parallel programming patterns in Fortran 2023. As such, the Fiats name has dual expansions: “Functional Inference And Training for Surrogates” or “Fortran Inference And Training for Science.” Fiats inference and training procedures are pure and therefore satisfy a language constraint imposed on procedure invocations inside Fortran’s parallel loop construct: do concurrent. Furthermore, the Fiats training procedures are built around a do concurrent parallel reduction. Several compilers can automatically parallelize do concurrent on Central Processing Units (CPUs) or Graphics Processing Units (GPUs). Fiats thus aims to achieve performance portability through standard language mechanisms.

Rouson, Damian [Lawrence Berkeley National Laborat↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

SpecSims: A Scalable Speculative Tree-based Simulation Cloning Framework for Finite Memory Machines

Simulation cloning is a technique in which cloned simulations whose state spaces differ partially from their parent simulation due to intervening events are spawned at runtime and concurrently advanced. It is a powerful method to carry out what-if analysis by speculatively exploring and evaluating the impact of various permutations of intervening cascade of events. Due to the exponential growth in the number of possible clones even for a small number of distinct intervening events, the practical efficacy of the approach is often severely limited by the maximum available memory of the computing host. In this paper, we introduce a novel speculative simulation cloning framework that executes a simulation cloning campaign capable of efficiently exploring an exponentially large space of clone simulations created by permutation of intervening events under a finite memory constraint. We provide a theoretical analysis of the runtime characteristics of our proposed approach and highlight its novel advantages such as memory-aware and as-long-as-needed execution. Furthermore, in support of our analytical findings and to demonstrate its practical feasibility, we implement a prototype of the cloning framework on a shared memory system and report its performance characteristics in the context of a heat diffusion simulation, and a power grid simulation subject to cascading disruptions from geomagnetic disturbances.

Simulation framework↗