Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Crack-Free 30% Chromium-Nickel Alloy Welding Products for Nuclear Service

We report prior research in the development of 30% chromium-nickel alloy nuclear welding wires has resulted in the resolution of primary water stress corrosion cracking (PWSCC) and ductility dip cracking (DDC) as well as improvement in solidification cracking (SC) resistance. The resolution of DDC exhibits some Laves phase, which has a negative effect on SC resistance. In this study, the use of an alternate carbide former, tantalum (Ta), in combination with niobium (Nb) was researched. Three heats of recently designed Filler Metal 52MSS-Ta (i.e., HV1648, HV1673A, and VX131WXW) were melted, fabricated, and systematically studied. DDC and SC were evaluated with thermodynamic modeling using the Scheil solidification simulation model, two types of varestraint tests, and strain-to-fracture (STF) testing. The varestraint and STF test results showed an improved SC resistance with reduced Laves phase and concurrent excellent DDC resistance. Optimized compositions with low Laves phase also exhibited high threshold strain values (TSVs) in the STF test. VX131WXW — which contains 2.81 wt-% Ta, 0.6 wt-% Nb, and 6 wt-% iron (Fe) - exhibited a TSV of 24%. Thermo-Calc computed the Laves phase to be 0.24% for VX131WXW compared to 0.06% in HV1673A. This difference in Laves phase resulted in the lower SC resistance of VX131WXW compared to HV1673A when measured with longitudinal varestraint testing. The maximum crack distance for HV1673A was about 0.6 mm while that of heat VX131WXW was about 1.0 mm. The typical diluted weld deposit made with VX131WXW was also resistant to PWSCC due to the chromium content exceeding 24%. These simultaneous results mark progress toward crack-free welds and provide direction for further optimization of Ta-containing filler metals.

30% chromium-nickel alloy↗

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

CGSim: A Simulation Framework for Large Scale Distributed Computing Environment

Large-scale distributed computing infrastructures such as the Worldwide LHC Computing Grid (WLCG) require comprehensive simulation tools for evaluating performance, testing new algorithms, and optimizing resource allocation strategies. However, existing simulators suffer from limited scalability, hardwired algorithms, lack of real-time monitoring, and inability to generate datasets suitable for modern machine learning approaches. We present CGSim, a simulation framework for large-scale distributed computing environments that addresses these limitations. Built upon the validated SimGrid simulation framework, CGSim provides high-level abstractions for modeling heterogeneous grid environments while maintaining accuracy and scalability. Key features include a modular plugin mechanism for testing custom workflow scheduling and data movement policies, interactive real-time visualization dashboards, and automatic generation of event-level datasets suitable for AI-assisted performance modeling. We demonstrate CGSim’s capabilities through a comprehensive evaluation using production ATLAS PanDA workloads, showing significant calibration accuracy improvements across WLCG computing sites. Scalability experiments show near-linear scaling for multi-site simulations, with distributed workloads achieving 6 × better performance compared to single-site execution. The framework enables researchers to simulate WLCG-scale infrastructures with hundreds of sites and thousands of concurrent jobs within practical time budget constraints on commodity hardware.

Vatsavai, Sairam Sri [Brookhaven National Laborato↗

Insights into Molecular Magnetism in Metal–Metal Bonded Systems as Revealed by a Spectroscopic and Computational Analysis of Diiron Complexes

A pair of bimetallic compounds featuring Fe–Fe bonds, [Fe( i PrNPPh 2 ) 3 FeR] (R = PMe 3 , ≡N t Bu), have been investigated using High-Frequency Electron Paramagnetic Resonance (HFEPR) as well as field- and temperature-dependent 57 Fe nuclear γ resonance (Mössbauer) spectroscopy. To gain insight into the local site electronic structure, we have concurrently studied a compound containing a single Fe(II) in a geometry analogous to that of one of the dimer sites. Our spectroscopic studies have allowed for the assessment of the electronic structure via the determination of the zero-field splitting and 57 Fe hyperfine parameters for the entire series. We also report on our efforts to correlate structure with physical properties in metal–metal bonded systems using ligand field theory guided by quantum chemical calculations. Through the insight gained in this study, we discuss strategies for the design of single-molecule magnets based on polymetallic compounds linked via direct metal–metal bonds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Digital Three Level Space Vector Modulator for High Frequency Vector Sequence Generation

This letter proposes a digital high-speed three-level space vector pulse width modulator (3L-SVPWM). A conventional 3L-SVPWM is typically computation-based, involving a sequential execution of sub-tasks on a digital signal processor (DSP) based controller. The resulting high computation time of 5.4 μs limits the implementation of additional control blocks for switching frequencies greater than 100 kHz. This is overcome by transforming sub-tasks into digital blocks with 1-0 decisions and simpler arithmetic operations. The sub-task blocks are executed concurrently on a programmable logic device (PLD). Hence, a fast 3L-SVPWM execution in 140 ns is achieved. The proposed digital 3L-SVPWM enables high switching frequency operation of wide bandgap (WBG) device-based 3 L inverters to generate high fundamental frequency waveforms. A finite state machine is an integral part of the proposed implementation with the ability to generate any vector sequence, maximizing the usage of redundant vector states in 3L-SVPWM. Here, the proposed digital 3L-SVPWM operation is demonstrated with a GaN-based 3 L active neutral point clamped (3L-ANPC) inverter. Experimental results are presented at 250 kHz switching frequency to generate vector sequences for center-aligned SVPWM (CA-SVPWM) and common mode voltage reduced SVPWM (CMVR-SVPWM). The results also showcase a high fundamental frequency generation capability of 10 kHz.

active neutral point clamped inverter↗

2018 LDRD Annual Report (Argonne National Laboratory)

Argonne National Laboratory’s Laboratory Directed Research and Development (LDRD) program encourages the development of novel technical concepts, enhances the Laboratory’s research and development (R&D) capabilities, and enables pursuit of strategic laboratory goals. Argonne’s LDRD projects are proposal based and peer reviewed, supporting ideas that require advanced exploration so they can be sufficiently developed to pursue support through normal programmatic channels. Among the aims of the projects supported by the LDRD program are the establishment of engineering proofs of principle, assessment of design feasibility for prospective facilities, development of instrumentation or computational methods or systems, and discoveries in fundamental science and exploratory development. All LDRD projects have demonstrable ties to one or more of the science, energy, environment, and national security missions of the U.S. Department of Energy (DOE) and its National Nuclear Security Administration (NNSA), and many are also relevant to the missions of other federal agencies that sponsor work at Argonne. A natural consequence of the more “applied” type projects is their concurrent relevance to industry. The LDRD program is managed in overarching portfolios, each containing multiple projects each fiscal year. The LDRD Prime portfolio is further divided into strategic focus areas aligned with Argonne’s strategic plan. The largest component of Argonne’s program is LDRD Prime, which emphasizes R&D explicitly aligned with Laboratory major initiatives in support of Argonne’s strategic plan. The choice of Focus Areas under the LDRD Prime component reflects the major initiatives; the state of development of relevant technical fields; the potential value of advancing those fields to DOE/NNSA and the nation; and the compatibility of the fields with existing facilities, capabilities, and staff expertise at Argonne. Focus Areas with projects that ended in FY18 are: Advanced Computing, Biological and Environmental Science Capability Development, Energy Manufacturing Science and Engineering, Hard X-ray Sciences, Materials and Chemistry, Securing Energy and Critical Resources, and The Universe as Our Laboratory (ULab).

99 GENERAL AND MISCELLANEOUS↗

Parallelizing the Unpacking and Clustering of Detector Data for Reconstruction of Charged Particle Tracks on Multi-core CPUs and Many-core GPUs

We present results from parallelizing the unpacking and clustering steps of the raw data from the silicon strip modules for reconstruction of charged particle tracks. Throughput is further improved by concurrently processing multiple events using nested OpenMP parallelism on CPU or CUDA streams on GPU. The new implementation along with earlier work in developing a parallelized and vectorized implementation of the combinatoric Kalman filter algorithm has enabled efficient global reconstruction of the entire event on modern computer architectures. We demonstrate the performance of the new implementation on Intel Xeon and NVIDIA GPU architectures.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High temporal frequency data from a four turbine, blade-resolved wind farm simulation with ExaWind

The data was generated with ExaWind (https://github.com/Exawind) which couples AMR-Wind (https://github.com/Exawind/amr-wind/), Nalu-Wind (https://github.com/Exawind/nalu-wind), TIOGA (https://github.com/Exawind/tioga), and OpenFAST (https://github.com/OpenFAST/openfast). This is a large-scale simulation of a blade-resolved wind farm using the ExaWind software stack. ExaWind couples together a background flow solver, AMR-Wind, and a near-body solver, Nalu-Wind, through an overset technique from the TIOGA application. Another application, OpenFAST, handles the structural dynamics of the turbine blades and towers, which informs the fluid-structure interaction of the wind turbines with the flow solvers. This particular simulation includes four blade-resolved wind turbines operating in a turbulent atmospheric boundary layer. The AMR-Wind solver uses 500 million cells and is being solved on 256 AMD GPUs of the Oakridge Leadership Computing Facility Frontier supercomputer. Each turbine is assigned its own Nalu-Wind solver with over 13 million elements per turbine and solved using 448 CPU cores, for a total of 1792 CPU cores. For each node, 56 cores contain Nalu-Wind, while 8 cores correspond to AMR-Wind operations on the GPUs. Consequently, ExaWind is entirely utilizing the CPUs and the GPUs of the nodes concurrently. The data used in the visualization is full flow field data output from the simulation. It is lossy-compressed to a specific accuracy using ZFP and written to disk every 16 time-steps to enable real-time flow visualization. The flow fields are sampled at a high temporal frequency to enable real-time, 24fps visualization. The flow fields are sampled every 12 simulation time steps (every 0.04132s).

17 WIND ENERGY↗

Shear strain alters the structure and migration mechanism of self-interstitial atoms in copper

Here, we use atomistic modeling to show that externally applied shear strain causes the lowest energy self-interstitial atom (SIA) structure in copper (Cu) to change from a <100>-type dumbbell to a <110>-type dumbbell. Concurrently, SIA migration switches from the 3-D random walk characteristic of <100>-type dumbbells to a 1-D mechanism analogous to that of crowdion SIAs. Furthermore, the relative energies of these two dumbbell structures as a function of strain are well predicted using elastic dipole tensors computed at zero strain, indicating that examination of these tensors may be used to assess the likelihood of strain-induced SIA structure transitions in other materials. Changes in lowest energy SIA structures and associated migration mechanisms stand to impact predictions of SIA behavior in irradiated solids.

36 MATERIALS SCIENCE↗

Proximity Portability and in Transit , M-to-N Data Partitioning and Movement in SENSEI [Book Chapter]

In high-performance parallel in situ processing, the term in transit processing refers to those configurations where data must move from a producer to a consumer that runs on separate resources. In the context of parallel and distributed computing on an HPC platform one of the central challenges is to determine a mapping of data from producer ranks to consumer ranks. This problem is complicated by the heterogeneity that arises in producer-consumer pairs, such as when producer and consumer codes have different levels of concurrency, different scaling characteristics, or different data models. The resulting mapping and movement of data from M producer to N consumer ranks can have a significant impact on aggregate application performance, particularly when the data consumer requires only a subset of the overall data for its task. This chapter focuses on the design considerations that underlie SENSEI’s implementation to this challenging problem. These design considerations extend the core SENSEI architecture and include ideas like the need to accommodate flexibility in the choice of different partitioning methods, the ability for a data consumer to request and receive only the subset of data needed for its particular operation, and the ability to leverage any of several different data transport tools. The idea of proximity portability, being able to use different data transport methods as part of an in transit workflow, is illustrated through the use of three different transport layers where switching from one transport tool to another is accomplished with only a configuration file change. Here, the chapter also includes a performance analysis summary showing the performance gains that are possible in terms of multiple metrics, such as memory footprint, time to solution, and amount of data moved, when using optimized partitioners in an in transit setting, gains that are made possible by the implementation shaped by specific design considerations.

Bethel, E. Wes↗

The impact of optical measurement techniques on measured aerosol particle size distributions

Ambient aerosol particle size distributions measured by the Ultra-High Sensitivity Aerosol Spectrometer (UHSAS) at various sites around the world exhibit modes at optical diameters near 600 nm and 850 nm. These modes are not present in concurrent measurements with the Grimm 11-D Optical Particle Counter (OPC). Here, in this study, we argue that these modes result from the optical measurement technique itself, and we explain why they appear in measurements by the UHSAS but not in those by the Grimm OPC. We construct computer models of the UHSAS and Grimm (“digital UHSAS” and “digital Grimm”) and use these to investigate the size distribution that would result from measurements of artificial aerosol particle size distributions that do not contain modes. The appearance of modes for the structureless incoming size distributions sampled by the digital UHSAS is explained by the nonlinear behavior of partial scattering cross sections of uniform spherical particles as a function of their diameter. The absence of modes in the digital Grimm is explained by the coarser size resolution of that instrument. Detailed analysis of the relationship between optical and geometric diameters for uniform spherical particles reveals two important results. First, these diameters generally have different numerical values for the same particle, and second, the relationship is nonlinear; thus, widths of size bins in terms of optical diameter differ from those in geometric diameter. These results explain the modes observed in the ambient size distributions and highlight concerns with attempts to create a merged size distribution by combining measurements from different instruments.

54 ENVIRONMENTAL SCIENCES↗

A Performance Model of In-Situ Techniques

The computational capacity of High-Performance Computing (HPC) systems increases continuously with the rapid development of central processing units (CPUs) and graphic processing units (GPUs), while the in-/output (IO) subsystem develops relatively slowly and storage capacity is also limited. Data-intensive applications, which are designed to leverage the high computational capacity of HPC resources, typically generate a considerable amount of data for post-processing visualizations and data analytics. The limited IO speed and storage space could lead to constraints in the actual performance of these applications and, therefore, scientific discovery. In-situ techniques, where data is visualized/analysed while still in memory rather than through disk, can contribute to alleviating these problems as they can reduce or even fully avoid data writing/reading through the IO subsystem to/from storage. However, the overall efficiency of insitu techniques crucially depends on the characteristics of both the in-situ tasks and the applications, and the resource distribution among them. Therefore, choosing the right in-situ approach (synchronous, asynchronous, or hybrid) and resource allocation is essential to minimize overhead and maximize the benefits of concurrent execution. In this paper, we present a performance model of in-situ techniques to find the most beneficial in-situ approach and the preferred resource configuration. We verify the high accuracy of our approach with over 6800 measurements and provide use cases with different applications.

Ju, Yi [Max Planck Computing and Data Facility, Ga↗

Advancing material modeling in hydrocodes using a concurrent finite-element and molecular dynamics multiscale framework

We present a multiscale simulation framework that couples the finite-element method with molecular dynamics. Bypassing traditional equations of state (EOS) by using in-line atomistic simulations, the method offers the advantage of incorporating detailed microscale physics not easily represented with coarse-grained models. Coupling consistency with the continuum code is ensured through the use of lifting and restriction operators, in line with heterogeneous multiscale methods. The concurrent continuum-atomistic framework is validated through comparison with experimental results and conventional EOS models, and demonstrated in a shock-driven hydrodynamic flow simulation under extreme conditions. We further evaluate the framework's usability by comparing it to state-of-the-art EOS models of deuterium. A computational performance study reveals that the atomistic EOS evaluation is a feasible alternative to conventional approaches, and demonstrates a weak scaling of 99% efficiency. These results highlight the framework's potential for large-scale multiscale modeling across a broad range of materials and conditions.

Computer science↗

Static actuator-sharing algorithm for concurrent control of multiple plasma properties

Simultaneous regulation of multiple properties in next-generation tokamaks like ITER and fusion pilot plant may require the integration of different plasma control algorithms. Such integration requires the conversion of individual controller commands into physical actuator requests while accounting for the coupling between different plasma properties. This work proposes a tokamak and scenario-agnostic actuator-sharing algorithm (ASA) to perform the above-mentioned command-request conversion and, hence, integrate multiple plasma controllers. The proposed algorithm implicitly solves a quadratic programming (QP) problem formulated to account for the saturation limits and the relation between the controller commands and physical actuator requests. Since the constraints arising in the QP program are linear, the proposed ASA is highly computationally efficient and can be implemented in the tokamak plasma control system in real time. Furthermore, the proposed algorithm is designed to handle real-time changes in the control objectives and actuators’ availability. Nonlinear simulations carried out using the Control Oriented Transport SIMulator illustrate the effectiveness of the proposed algorithm in achieving multiple control objectives simultaneously.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Selection of a Pair of Experiments to Optimally Reduce Uncertainty in Targeted Nuclear Data

We propose a novel process to select a pair of differential and integral experiments that best reduce uncertainties in targeted 239 ⁢Pu nuclear data while compressing the current nuclear data pipeline from 20 to 3 years. 239⁢ Pu nuclear data are poorly understood for neutrons in the intermediate energy range due to sparsity and uncertainty in historical experiments. New experiments targeting this range will enable better understanding of these nuclear data, but choosing the ideal experiments to conduct is challenging. Beginning with a prior distribution represented by samples of nuclear data generated from theory, generalized least squares adjustments are made to incorporate data from historical experiments. To quantify potential uncertainty reduction obtainable from a pair of candidate experiments, we compute the D-optimality criterion of the posterior covariance of intermediate energy range nuclear data compared to the equivalent covariance after additional adjustment to the pair of candidate experiments. Repeating the process for each of many candidate pairs facilitates the final selection. Results support 63⁢ Cu total cross section measurements for differential experiments and alumina and alumina/graphite configurations for integral experiments. This analysis enables choosing differential and integral experiments to be executed concurrently while shortening decision times relative to the current nuclear data pipeline.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nondestructively Visualizing and Understanding the Mechano-Electro-chemical Origins of “Soft Short” and “Creeping” in All-Solid-State Batteries

All-solid-state Li-metal batteries (ASLMBs) represent a significant breakthrough in the quest to overcome limitations associated with traditional Li-ion batteries, particularly in energy density and safety aspects. However, widespread implementation is stymied due to a lack of profound understanding of the complex mechano-electro-chemical behavior of Li metal in the ASLMBs. Herein, operando neutron imaging and X-ray computed tomography (XCT) are leveraged to nondestructively visualize Li behaviors within ASLMBs. This approach offers real-time observations of Li evolutions, both pre- and post- occurrence of a “soft short”. The coordination of 2D neutron radiography and 3D neutron tomography enables charting of the terrain of Li metal deformation operando. Concurrently, XCT offers a 3D insight into the internal structure of the battery following a “soft short”. Despite the manifestation of a “soft short”, the persistence of Faradaic processes is observed. To study the elusive “soft short”, phase field modeling is coupled with electrochemistry and solid mechanics theory. The research unravels how external pressure curbs dendrite growth, potentially leading to dendrite fractures and thus uncovering the origins of both “soft” and “hard” shorts in ASLMBs. Furthermore, by harnessing finite element modeling, it dive deeper into the mechanical deformation and the fluidity of Li metal.

25 ENERGY STORAGE↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Systematic Study of Solid-State U(VI) Photoreactivity: Long-Lived Radicalization and Electron Transfer in Uranyl Tetrachloride

Reported are the syntheses, structural characterizations and luminescence properties of three novel [UO 2 Cl 4 ] 2- bearing compounds containing substituted 1,1’-dialkyl-4,4’-bipyridinum dications (i.e. viologens). These compounds undergo photoinduced luminescence quenching upon exposure to UV radiation. Further, this reactivity is concurrent with two phenomena: radicalization of the uranyl tetrachloride anion and photoelectron transfer to the viologen which constitutes the formal transfer of one electron from the [UO 2 Cl 4 ] 2- to the viologen species. This behavior is elucidated using electron paramagnetic resonance (EPR) spectroscopy and further probed through a series of characterization and computational techniques including Rehm-Weller analysis, time-dependent density functional theory (TD-DFT), and density of states (DOS). This work provides a systematic study of the photoreactivity of the uranyl unit in the solid state, an under-described aspect of fundamental uranyl chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗