Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Error field detection and correction studies towards ITER operation

In magnetic fusion devices, error field (EF) sources, spurious magnetic field perturbations, need to be identified and corrected for safe and stable (disruption-free) tokamak operation. Within Work Package Tokamak Exploitation RT04, a series of studies have been carried out to test the portability of the novel non-disruptive method, designed and tested in DIII-D (Paz-Soldan et al 2022 Nucl. Fusion62 126007), and to perform an assessment of model-based EF control strategies towards their applicability in ITER. In this paper, the lessons learned, the physical mechanism behind the magnetic island healing, which relies on enhanced viscous torque that acts against the static electro-magnetic torque, and the main control achievements are reported, together with the first design of the asynchronous EF correction current/density controller for ITER.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Magnetic triggering — time-resolved characterisation of silicon strip modules in the presence of switching DC-DC converters

Modules for the ATLAS Inner Tracker (ITk) strip tracker include a DC-DC converter circuit glued directly to the silicon sensor which converts the 11 V supplied to the module to the 1.5 V required for the operation of the readout chips. The DC-DC converter unit, consisting of a copper solenoid and custom ASIC, is located directly above the silicon strip sensor and therefore needs to be shielded to protect the sensor from EMI noise created during the operation of the circuit. Despite dedicated shielding, consisting of an aluminium shield box with continuous solder seams encompassing the surface components and a copper layer in the PCB beneath it, module channels connected to sensor strips located beneath the converter circuit were found to show a noise increase. While the DC-DC converter unit causing the underlying EMI noise operates at a frequency of 2 MHz, module characterisation measurements for ITk strip tracker modules are typically performed asynchronously to the DC-DC switching and are therefore averaged over the full range of time bins with respect to the converter frequency. In order to investigate the time dependence of the noise injection relative to the DC-DC switching frequency, a dedicated setup to understand the time-resolved performance change in modules was developed. By using a magnetic field probe to measure the field leaking through the shield box and triggering on its rising edge, data taking could be synchronised with the DC-DC switching. This paper illustrates the concept and setup of such time-resolved performance measurements using magnetic triggering and presents results for the observed effects on signal and noise for ATLAS ITk strip modules from both laboratory and beam tests.

47 OTHER INSTRUMENTATION↗

Event driven readout architecture with non-priority arbitration for radiation detectors

A novel event driven readout architecture, EDWARD (Event Driven with Access and 8 Reset Decoder) architecture, for highly granular pixel detectors is presented. It incorporates, inter alia, an asynchronous arbitration tree based on Seitz' arbiters, removing the need for an imposed prioritization scheme. It also provides protection against glitches during readout. The system allows not only reading pixel activities, but also retrieving additional data, both analog and digital, from the pixels. A novel in-channel logic allows the entire readout process to be split into consecutive phases for additional flexibility. All operations are controlled by only one edge of the clock signal, seen as an acknowledge token, so there is no dead time between readouts.

47 OTHER INSTRUMENTATION↗

Integration of EDWARD readout architecture in full-field fluorescence imaging detector

Data bandwidth, timing resolution and resource utilization in readouts of radiation detectors are a constant challenge. Event driven solutions are pushing against well-trenched framed solutions. The idea for an asynchronous readout architecture called EDWARD ( E vent- D riven W ith A ccess and R eset D ecoder) was presented at the TWEPP 2021 conference. Here we show the progress of our work which resulted in two chip prototypes. The first one, named 3FI65P1, is a full device with the analog pixel circuitry suited for full-field fluorescence imaging. It is already manufactured, and preliminary results are presented. Finally, the second chip, named EDWARD65P1, contains digital pulse generators with Poisson-exponential distribution in each pixel for extraction of the performance matrix of the EDWARD architecture alone.

47 OTHER INSTRUMENTATION↗

The CMS Phase-2 Fast Beam Condition Monitor prototype test with beam

The Fast Beam Condition Monitor (FBCM) is a standalone luminometer for the High Luminosity LHC (HL-LHC) program of the CMS Experiment at CERN. The detector is under development and features a new, radiation-hard, front-end application-specific integrated circuit (ASIC) designed for beam monitoring applications. The achieved timing resolution of a few nanoseconds enables the measurement of both the luminosity and the beam-induced background. The ASIC, called FBCM23, features six channels with adjustable shaping times, enabling in-field fine-tuning. Each ASIC channel outputs a single binary asynchronous signal encoding time-of-arrival and time-over-threshold information. The FBCM is based on silicon-pad sensors, with two sensor designs presently being considered. This paper presents the results of tests of the FBCM detector prototype using both types of silicon sensors with hadron, muon, and electron beams. Irradiated FBCM23 ASICs and silicon-pad sensors were also tested to simulate the expected conditions near the end of the detector's lifetime in the HL-LHC radiation environment. Based on test results, direct bonding between the sensor and ASIC was chosen, and an optimal bias voltage and ASIC threshold for FBCM operation were proposed. The current design of the front-end test board was validated following the beam test and is now being used for the first front-end module, which is expected to be produced in summer 2025. These results represent a major step forward in validating the FBCM concept, first version of the firmware and establishing a reliable design path for the final detector.

Beam-line instrumentation (beam position and profi↗

Deforestation reshapes land-surface energy-flux partitioning

Land-use and land-cover change significantly modify local land-surface characteristics and water/energy exchanges, which can lead to atmospheric circulation and regional climate changes. In particular, deforestation accounts for a large portion of global land-use changes, which transforms forests into other land cover types, such as croplands and grazing lands. Many previous efforts have focused on observing and modeling land–atmosphere–water/energy fluxes to investigate land–atmosphere coupling induced by deforestation. However, interpreting land–atmosphere–water/energy-flux responses to deforestation is often complicated by the concurrent impacts from shifts in land-surface properties versus background atmospheric forcings. In this study, we used 29 paired FLUXNET sites, to improve understanding of how deforested land surfaces drive changes in surface-energy-flux partitioning. Each paired sites included an intact forested and non-forested site that had similar background climate. We employed transfer entropy, a method based on information theory, to diagnose directional controls between coupling variables, and identify nonlinear cause–effect relationships. Transfer entropy is a powerful tool to detective causal relationships in nonlinear and asynchronous systems. The paired eddy covariance flux measurements showed consistent and strong information flows from vegetation activity (gross primary productivity (GPP)) and physical climate (e.g. shortwave radiation, air temperature) to evaporative fraction (EF) over both non-forested and forested land surfaces. More importantly, the information transfers from radiation, precipitation, and GPP to EF were significantly reduced at non-forested sites, compared to forested sites. We then applied these observationally constrained metrics as benchmarks to evaluate the Energy Exascale Earth System Model (E3SM) land model (ELM). ELM predicted vegetation controls on EF relatively well, but underpredicted climate factors on EF, indicating model deficiencies in describing the relationships between atmospheric state and surface fluxes. Moreover, changes in controls on surface energy flux partitioning due to deforestation were not detected in the model. We highlight the need for benchmarking model simulated surface-energy fluxes and the corresponding causal relationships against those of observations, to improve our understanding of model predictability on how deforestation reshapes land surface energy fluxes.

54 ENVIRONMENTAL SCIENCES↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

The ABACUS cosmological N -body code

We present abacus, a fast and accurate cosmological N-body code based on a new method for calculating the gravitational potential from a static multipole mesh. The method analytically separates the near- and far-field forces, reducing the former to direct 1/r 2 summation and the latter to a discrete convolution over multipoles. The method achieves 70 million particle updates per second per node of the Summit supercomputer, while maintaining a median fractional force error of 10 -5 . We express the simulation time-step as an event-driven ‘pipeline’, incorporating asynchronous events such as completion of co-processor work, input/output, and network communication. abacus has been used to produce the largest suite of N-body simulations to date, the abacussummit suite of 60 trillion particles, incorporating on-the-fly halo finding. abacus enables the production of mock catalogues of the volume and resolution required by the coming generation of cosmological surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Reionization time of the Local Group and Local-Group-like halo pairs

ABSTRACT Patchy cosmic reionization resulted in the ionizing UV background asynchronous rise across the Universe. The latter might have left imprints visible in present-day observations. Several numerical simulation-based studies show correlations between the reionization time and overdensities and object masses today. To remove the mass from the study, as it may not be the sole important parameter, this paper focuses solely on the properties of paired haloes within the same mass range as the Milky Way. For this purpose, it uses CoDaII, a fully coupled radiation hydrodynamics reionization simulation of the local Universe. This simulation holds a halo pair representing the Local Group, in addition to other pairs, sharing similar mass, mass ratio, distance separation, and isolation criteria but in other environments, alongside isolated haloes within the same mass range. Investigations of the paired halo reionization histories reveal a wide diversity although always inside-out, given our reionization model. Within this model, haloes in a close pair tend to be reionized at the same time but being in a pair does not bring to an earlier time their mean reionization. The only significant trend is found between the total energy at z = 0 of the pairs and their mean reionization time: Pairs with the smallest total energy (bound) are reionized up to 50 Myr earlier than others (unbound). Above all, this study reveals the variety of reionization histories undergone by halo pairs similar to the Local Group, that of the Local Group being far from an average one. In our model, its reionization time is ∼625 Myr against 660 ± 4 Myr (z ∼ 8.25 against 7.87 ± 0.02) on average.

79 ASTRONOMY AND ASTROPHYSICS↗

Customized Bayesian optimization for efficient beam tuning at the facility for rare isotope beams

Bayesian optimization (BO) has recently emerged as a powerful approach for on-line beam tuning, and it is rapidly gaining adoption across accelerator facilities due to its flexibility and efficiency in handling complex optimization tasks. At the Facility for Rare Isotope Beams, rapid and reliable tuning is essential to support the delivery of diverse ion species. To improve the practicality of BO in this setting, we implemented several enhancements, including scalarized composite objective construction for multicriteria optimization, asynchronous evaluation for better resource utilization, prior-mean-assisted optimization to accelerate convergence, and a local search strategy for rapid completion of the task. We present the details of these methods, discuss challenges-encountered, and share our experience applying them to specific beam-tuning tasks.

Accelerators & storage rings↗

Entanglement-fidelity limits of photonically networked atomic qubits from recoil and timing

The remote entanglement of two atomic quantum memories through photonic interactions is accompanied by atomic momentum recoil. When the interactions occur at different times, such as from the random emission over the lifetime of the atomic excited state, the difference in recoil timing can expose “which-path” information and ultimately lead to decoherence. Time-bin encoded photonic qubits can be particularly sensitive to asynchronous recoil timing. In this paper we study the limits of entanglement fidelity in atomic systems due to recoil and other timing imbalances and show how these effects can be suppressed or even eliminated through proper experimental design.

Quantum communication↗

Driven responses of periodically patterned superconducting films

Here, we simulate the motion of a commensurate vortex lattice in a periodic lattice of artificial circular pinning sites having different diameters, pinning strengths, and spacings using the time-dependent Ginzburg-Landau formalism. Above some critical DC current density J c , the vortices depin, and the resulting steady-state motion then induces an oscillatory electric field E (t) with a defect "hopping" frequency f 0 , which depends on the applied current density and the pinning landscape characteristics. The frequency generated can be locked to an applied AC current density over some range of frequencies, which depends on the amplitude of the DC as well as the AC current densities. Both synchronous and asynchronous collective hopping behaviors are studied as a function of the supercell size of the simulated system and the (asymptotic) synchronization threshold current densities determined.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Coincident learning for beam-based rf station fault identification using phase information at the SLAC linac coherent light source

Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.

Accelerator Physics (physics.acc-ph)↗

tomoCAM : fast model-based iterative reconstruction via GPU acceleration and non-uniform fast Fourier transforms

X-ray-based computed tomography is a well established technique for determining the three-dimensional structure of an object from its two-dimensional projections. In the past few decades, there have been significant advancements in the brightness and detector technology of tomography instruments at synchrotron sources. These advancements have led to the emergence of new observations and discoveries, with improved capabilities such as faster frame rates, larger fields of view, higher resolution and higher dimensionality. These advancements have enabled the material science community to expand the scope of tomographic measurements towards increasingly in situ and in operando measurements. In these new experiments, samples can be rapidly evolving, have complex geometries and restrictions on the field of view, limiting the number of projections that can be collected. In such cases, standard filtered back-projection often results in poor quality reconstructions. Iterative reconstruction algorithms, such as model-based iterative reconstructions (MBIR), have demonstrated considerable success in producing high-quality reconstructions under such restrictions, but typically require high-performance computing resources with hundreds of compute nodes to solve the problem in a reasonable time. Here, tomoCAM , is introduced, a new GPU-accelerated implementation of model-based iterative reconstruction that leverages non-uniform fast Fourier transforms to efficiently compute Radon and back-projection operators and asynchronous memory transfers to maximize the throughput to the GPU memory. The resulting code is significantly faster than traditional MBIR codes and delivers the reconstructive improvement offered by MBIR with affordable computing time and resources. tomoCAM has a Python front-end, allowing access from Jupyter -based frameworks, providing straightforward integration into existing workflows at synchrotron facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Five-analyzer Johann spectrometer for hard X-ray photon-in/photon-out spectroscopy at the Inner Shell Spectroscopy beamline at NSLS-II: design, alignment and data acquisition

Here, a recently commissioned five-analyzer Johann spectrometer at the Inner Shell Spectroscopy beamline (8-ID) at the National Synchrotron Light Source II (NSLS-II) is presented. Designed for hard X-ray photon-in/photon-out spectroscopy, the spectrometer achieves a resolution in the 0.5–2 eV range, depending on the element and/or emission line, providing detailed insights into the local electronic and geometric structure of materials. It serves a diverse user community, including fields such as physical, chemical, biological, environmental and materials sciences. This article details the mechanical design, alignment procedures and data-acquisition scheme of the spectrometer, with a particular focus on the continuous asynchronous data-acquisition approach that significantly enhances experimental efficiency.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection

Communication overhead in federated learning (FL) poses a significant challenge for network anomaly detection systems, where the myriad of client configurations and network conditions can severely impact system efficiency and detection accuracy. While existing approaches attempt to address this through individual optimization techniques, they often fail to maintain the delicate balance between reduced overhead and detection performance. This paper presents an adaptive FL framework that dynamically combines batch size optimization, client selection, and asynchronous updates to achieve efficient anomaly detection. Through extensive profiling and experimental analysis on two distinct datasets-UNSW-NBIS for general network traffic and ROAD for automotive networks-our framework reduces communication overhead by 97.6%; (from 700.0s to 16.8s) compared to synchronous baseline approaches while maintaining comparable detection accuracy (95.10%; vs. 95.12%;). Statistical validation using Mann-Whitney U test confirms significant improvements (p < 0.05) over existing FL approaches across both datasets, demonstrating the framework's adaptability to different network security contexts. Detailed profiling analysis reveals the efficiency gains through dramatic reductions in GPU operations and memory transfers while maintaining robust detection performance under varying client conditions.

Marfo, William [University of Texas at El Paso]↗

RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows

While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. Here, we experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.

97 MATHEMATICS AND COMPUTING↗

A Task Based Approach for Co-Scheduling Ensemble Workloads on Heterogeneous Nodes

Scientific workflows consist of multiple, connected applications, with data and results flowing from one to another in a pipeline. Traditionally, such workflows are executed in sequential order, storing intermediate data in storage disks. Co-scheduling application workflows concurrently on the same compute nodes would greatly reduce the cost of moving data to/from storage and allow real-time analysis of intermediate results. Nevertheless, most parallel programming runtimes do not allow seamless integration of various applications in a scientific workflow, in part due to the complexity of managing data and resources. The situation is even more complicated for heterogeneous systems. In this work we extend the Minos Computing Library (MCL) runtime to accelerate pipe-lined and parallel workloads where multiple applications are running in the same system. MCL’s asynchronous task library and runtime dynamically manages resources to allow co-scheduling of multiple processes sharing heterogeneous resources. In addition, we design a custom ex- tension of the Open Compute Language (OpenCL) to enable multiple processes to share device memory. We enable MCL to coordinate these shared buffers to allow for easy, fast data sharing between applications. Using malleable micro-benchmarks and two application workflows that combine scientific simulation and AI-based analysis, we show that our method outperforms traditional approaches.

Index Terms—Parallel systems, Scheduling and Task ↗