Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Suppressed electric quadrupole collectivity in 49 Ti

Single-step Coulomb excitation of 46,48,49,50 Ti is presented. A complete set of E2 matrix elements for the quintuplet of states in 49 Ti, centred on the core excitation, was measured for the first time. A total of nine E2 matrix elements are reported, four of which were previously unknown. $^{49}_{22}$Ti 27 shows a 20% quenching in electric quadrupole transition strength as compared to its semi-magic $^{50}_{22}$Ti 28 neighbour. This 20% quenching, while empirically unprecedented, can be explained with a remarkably simple two-state mixing model, which is also consistent with other ground-state properties such as the magnetic dipole moment and electric quadrupole moment. A connection to nucleon transfer data and the quenching of single-particle strength is also demonstrated. The simplicity of the 49 Ti- 50 Ti pair (i.e., approximate single-j 0 7/2 valence space and isolation of yrast states from non-yrast states) provides a unique opportunity to disentangle otherwise competing effects in the ground-state properties of atomic nuclei, the emergence of collectivity, and the role of proton-neutron interactions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

GPU Acceleration of Large-Scale Full-Frequency GW Calculations

Many-body perturbation theory is a powerful method to simulate electronic excitations in molecules and materials starting from the output of density functional theory calculations. By implementing the theory efficiently so as to run at scale on the latest leadership high-performance computing systems it is possible to extend the scope of GW calculations. Here, we present a GPU acceleration study of the full-frequency GW method as implemented in the WEST code. Excellent performance is achieved through the use of (i) optimized GPU libraries, e.g., cuFFT and cuBLAS, (ii) a hierarchical parallelization strategy that minimizes CPU-CPU, CPU-GPU, and GPU-GPU data transfer operations, (iii) nonblocking MPI communications that overlap with GPU computations, and (iv) mixed precision in selected portions of the code. A series of performance benchmarks has been carried out on leadership high-performance computing systems, showing a substantial speedup of the GPU-accelerated version of WEST with respect to its CPU version. Good strong and weak scaling is demonstrated using up to 25 920 GPUs. Finally, we showcase the capability of the GPU version of WEST for large-scale, full-frequency GW calculations of realistic systems, e.g., a nanostructure, an interface, and a defect, comprising up to 10 368 valence electrons.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Diffractive optical computing in free space

Abstract Structured optical materials create new computing paradigms using photons, with transformative impact on various fields, including machine learning, computer vision, imaging, telecommunications, and sensing. This Perspective sheds light on the potential of free-space optical systems based on engineered surfaces for advancing optical computing. Manipulating light in unprecedented ways, emerging structured surfaces enable all-optical implementation of various mathematical functions and machine learning tasks. Diffractive networks, in particular, bring deep-learning principles into the design and operation of free-space optical systems to create new functionalities. Metasurfaces consisting of deeply subwavelength units are achieving exotic optical responses that provide independent control over different properties of light and can bring major advances in computational throughput and data-transfer bandwidth of free-space optical processors. Unlike integrated photonics-based optoelectronic systems that demand preprocessed inputs, free-space optical processors have direct access to all the optical degrees of freedom that carry information about an input scene/object without needing digital recovery or preprocessing of information. To realize the full potential of free-space optical computing architectures, diffractive surfaces and metasurfaces need to advance symbiotically and co-evolve in their designs, 3D fabrication/integration, cascadability, and computing accuracy to serve the needs of next-generation machine vision, computational imaging, mathematical computing, and telecommunication technologies.

36 MATERIALS SCIENCE↗

Freely scalable and reconfigurable optical hardware for deep learning

Abstract As deep neural network (DNN) models grow ever-larger, they can achieve higher accuracy and solve more complex problems. This trend has been enabled by an increase in available compute power; however, efforts to continue to scale electronic processors are impeded by the costs of communication, thermal management, power delivery and clocking. To improve scalability, we propose a digital optical neural network (DONN) with intralayer optical interconnects and reconfigurable input values. The path-length-independence of optical energy consumption enables information locality between a transmitter and a large number of arbitrarily arranged receivers, which allows greater flexibility in architecture design to circumvent scaling limitations. In a proof-of-concept experiment, we demonstrate optical multicast in the classification of 500 MNIST images with a 3-layer, fully-connected network. We also analyze the energy consumption of the DONN and find that digital optical data transfer is beneficial over electronics when the spacing of computational units is on the order of $$>10\,\upmu $$ > 10 μ m.

42 ENGINEERING↗

Coupling a recurrent neural network to SPAD TCSPC systems for real-time fluorescence lifetime imaging

Fluorescence lifetime imaging (FLI) has been receiving increased attention in recent years as a powerful diagnostic technique in biological and medical research. However, existing FLI systems often suffer from a tradeoff between processing speed, accuracy, and robustness. Inspired by the concept of Edge Artificial Intelligence (Edge AI), we propose a robust approach that enables fast FLI with no degradation of accuracy. This approach couples a recurrent neural network (RNN), which is trained to estimate the fluorescence lifetime directly from raw timestamps without building histograms, to SPAD TCSPC systems, thereby drastically reducing transfer data volumes and hardware resource utilization, and enabling real-time FLI acquisition. We train two variants of the RNN on a synthetic dataset and compare the results to those obtained using center-of-mass method (CMM) and least squares fitting (LS fitting). Results demonstrate that two RNN variants, gated recurrent unit (GRU) and long short-term memory (LSTM), are comparable to CMM and LS fitting in terms of accuracy, while outperforming them in the presence of background noise by a large margin. To explore the ultimate limits of the approach, we derive the Cramer-Rao lower bound of the measurement, showing that RNN yields lifetime estimations with near-optimal precision. To demonstrate real-time operation, we build a FLI microscope based on an existing SPAD TCSPC system comprising a 32 x 32 SPAD sensor named Piccolo. Four quantized GRU cores, capable of processing up to 4 million photons per second, are deployed on the Xilinx Kintex-7 FPGA that controls the Piccolo. Powered by the GRU, the FLI setup can retrieve real-time fluorescence lifetime images at up to 10 frames per second. The proposed FLI system is promising and ideally suited for biomedical applications, including biological imaging, biomedical diagnostics, and fluorescence-assisted surgery, etc.

47 OTHER INSTRUMENTATION↗

WLCG Token Usage and Discovery

Since 2017, the Worldwide LHC Computing Grid (WLCG) has been working towards enabling token based authentication and authorisation throughout its entire middleware stack. Following the publication of the WLCG Common JSON Web Token (JWT) Schema v1.0 [1] in 2019, middleware developers have been able to enhance their services to consume and validate the JWT-based [2] OAuth2.0 [3] tokens and process the authorization information they convey. Complex scenarios, involving multiple delegation steps and command line flows, are a key challenge to be addressed in order for the system to be fully operational. This paper expands on the anticipated token based workflows, with a particular focus on local storage of tokens and their discovery by services. The authors include a walk-through of this token flow in the RUCIO managed data-transfer scenario, including delegation to FTS and authorised access to storage elements. Next steps are presented, including the current target of submitting production jobs authorised by Tokens within 2021.

Bockelman, Brian↗

Performance of CUDA Unified Memory in CMS Heterogeneous Pixel Reconstruction

The management of separate memory spaces of CPUs and GPUs brings an additional burden to the development of software for GPUs. To help with this, CUDA unified memory provides a single address space that can be accessed from both CPU and GPU. The automatic data transfer mechanism is based on page faults generated by the memory accesses. This mechanism has a performance cost, that can be with explicit memory prefetch requests. Various hints on the inteded usage of the memory regions can also be given to further improve the performance. The overall effect of unified memory compared to an explicit memory management can depend heavily on the application. In this paper we evaluate the performance impact of CUDA unified memory using the heterogeneous pixel reconstruction code from the CMS experiment as a realistic use case of a GPU-targeting HEP reconstruction software. We also compare the programming model using CUDA unified memory to the explicit management of separate CPU and GPU memory spaces.

Kortelainen, Matti J.↗

Adoption of a token-based authentication model for the CMS Submission Infrastructure

The CMS Submission Infrastructure (SI) is the main computing resource provisioning system for CMS workloads. A number of HTCondor pools are employed to manage this infrastructure, which aggregates geographically distributed resources from the WLCG and other providers. Historically, the model of authentication among the diverse components of this infrastructure has relied on the Grid Security Infrastructure (GSI), based on identities and X509 certificates. In contrast, commonly used modern authentication standards are based on capabilities and tokens. The WLCG has identified this trend and aims at a transparent replacement of GSI for all its workload management, data transfer and storage access operations, to be completed during the current LHC Run 3. As part of this effort, and within the context of CMS computing, the Submission Infrastructure group is in the process of phasing out the GSI part of its authentication layers, in favor of IDTokens and Scitokens. The use of tokens is already well integrated into the HTCondor Software Suite, which has allowed us to fully migrate the authentication between internal components of SI. Additionally, recent versions of the HTCondor-CE support tokens as well, enabling CMS resource requests to Grid sites employing this CE technology to be granted by means of token exchange. After a rollout campaign to sites, successfully completed by the third quarter of 2022, the totality of HTCondor CEs in use by CMS are already receiving Scitoken-based pilot jobs. On the ARC CE side, a parallel campaign was launched to foster the adoption of the REST interface at CMS sites (required to enable token-based job submission via HTCondor-G), which is nearing completion as well. In this contribution, the newly adopted authentication model will be described. We will then report on the migration status and final steps towards complete GSI phase out in the CMS SI.

Pérez-Calero Yzquierdo, Antonio↗

Porting ATLAS Fast Calorimeter Simulation to GPUs with Performance Portable Programming Models

FastCaloSim is a parameterized simulation of the particle energy response and of the energy distribution in the ATLAS calorimeter. It is a relatively small and self-contained package with massive inherent parallelism and captures the essence of GPU offloading via important operations like data transfer, memory initialization, floating point operations, and reduction. It was identified by the High Energy Physics Center for Computational Excellence project as a good testbed for evaluating the performance and ease of portability of programming models. In this paper, we will discuss the results of our evaluation of the porting process to Kokkos, SYCL, Alpaka, OpenMP and std::par (nvc++), and compare performance on NVIDIA, AMD and Intel GPUs, as well as multicore CPUs.

97 MATHEMATICS AND COMPUTING↗

Spatial core-edge coupling of the particle-in-cell gyrokinetic codes GEM and XGC

Two existing particle-in-cell gyrokinetic codes, GEM for the core region and XGC for the edge region, have been successfully coupled with a spatial coupling scheme at the interface in a toroidal geometry. Additionally, a mapping technique is developed for transferring data between GEM's structured and XGC's unstructured meshes. Two examples of coupled simulations are presented to demonstrate the coupling scheme. The optimization of GEM for graphics processing unit is also presented.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Porting fragmentation methods to GPUs using an OpenMP API: Offloading the resolution-of-the-identity second-order Møller–Plesset perturbation method

Here, using an OpenMP Application Programming Interface, the resolution-of-the-identity second-order Møller–Plesset perturbation (RI-MP2) method has been off-loaded onto graphical processing units (GPUs), both as a standalone method in the GAMESS electronic structure program and as an electron correlation energy component in the effective fragment molecular orbital (EFMO) framework. First, a new scheme has been proposed to maximize data digestion on GPUs that subsequently linearizes data transfer from central processing units (CPUs) to GPUs. Second, the GAMESS Fortran code has been interfaced with GPU numerical libraries (e.g., NVIDIA cuBLAS and cuSOLVER) for efficient matrix operations (e.g., matrix multiplication, matrix decomposition, and matrix inversion). The standalone GPU RI-MP2 code shows an increasing speedup of up to 7.5× using one NVIDIA V100 GPU with one IBM 42-core P9 CPU for calculations on fullerenes of increasing size from 40 to 260 carbon atoms using the 6-31G(d)/cc-pVDZ-RI basis sets. A single Summit node with six V100s can compute the RI-MP2 correlation energy of a cluster of 175 water molecules using the correlation consistent basis sets cc-pVDZ/cc-pVDZ-RI containing 4375 atomic orbitals and 14 700 auxiliary basis functions in ~0.85 h. In the EFMO framework, the GPU RI-MP2 component shows near linear scaling for a large number of V100s when computing the energy of an 1800-atom mesoporous silica nanoparticle in a bath of 4000 water molecules. The parallel efficiencies of the GPU RI-MP2 component with 2304 and 4608 V100s are 98.0% and 96.1%, respectively.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High acoustic velocity x -cut lithium niobate sub-terahertz electromechanics

Micromechanical resonators operating above 100 GHz are favorable candidates for quantum physics studies due to their stronger ability to withstand thermal fluctuations, allowing them to remain in the quantum ground state even at kelvin temperatures. Furthermore, electromechanical resonators at sub-terahertz frequencies enable high-speed data transfer in modern communication technologies, making them attractive for communication industries. Recently, sub-terahertz electromechanics has been demonstrated on z-cut thin-film lithium niobate. Yet, the x-cut thin-film lithium niobate is more advantageous for scaling above 100 GHz due to its faster acoustic velocity. Here, we report sub-terahertz electromechanics on x-cut thin-film lithium niobate utilizing the thickness-longitudinal mode. In addition, we study the orientation dependence of these mechanical resonators due to the anisotropy of lithium niobate. We find that devices with a cross section close to the xy plane can be more efficiently excited, in contrast to those near the xz plane. This difference stems from the orientation-dependent nature of the e12 piezoelectric coupling element of the x-cut lithium niobate film. This investigation could assist in optimizing resonator designs by choosing the crystallographic direction that offers the best performance for specific functionalities.

Physics↗

Broadband coherent anti-Stokes Raman scattering (BCARS) microscopy for rapid, label-free biological imaging

Broadband coherent anti-Stokes Raman scattering (BCARS) microscopy is a label-free imaging approach that provides detailed chemical information at high spatial resolution in a sample through nonlinear, coherent excitation of molecular vibrations and detection of Raman spectra. While its utility for biological imaging has been demonstrated, many aspects of this technique must mature before it can be widely adopted. One of the areas of required improvement is imaging speed—most BCARS implementations involve sample rastering, which limits imaging speed. Beam scanning can provide faster BCARS imaging but presents some unique challenges. Here, we describe a beam-scanning BCARS microscopy system that improves spatial resolution twofold and imaging speed by fivefold over a previous beam-scanning implementation. These enhancements were enabled by an improvement in supercontinuum power and the use of a sCMOS camera for its high data transfer rate and low read noise. Implementation of the sCMOS camera required correction for the significant pixel-to-pixel background and photon response nonuniformity. Here, we report on the method that we implemented for calibrating and correcting the pixel-to-pixel differences in sCMOS camera noise.

Dixon, Jessica Z. [Georgia Institute of Technology↗

High-Fidelity Multiphysics Modeling of a Heat Pipe Microreactor Using BlueCrab

Researchers who are actively developing nuclear microreactors are planning to employ innovative designs and features using traditional commercial modeling tools that may be inadequate for their design and licensing activities. The codes developed under the U.S. Department of Energy Office of Nuclear Energy Advanced Modeling and Simulation (NEAMS) program provide flexibility in terms of geometry modeling and multiphysics coupling and are particularly well suited for modeling novel microreactor concepts. To test the maturity of these codes, this paper introduces a conceptual heat pipe microreactor (HP-MR) designed to gather various technologies of interest to microreactor developers such as control drums, heat pipes, and hydride moderators. Here, the objective of this effort is to demonstrate NEAMS tools capability to perform high-fidelity multiphysics simulations, using coupled neutronics (via the Griffin code), heat conduction (via the BISON code), heat pipe modeling (via the Sockeye code), and hydrogen redistribution in hydride metal moderator (via the SWIFT code). Codes are coupled in-memory through the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, which permits flexible multiphysics data transfer schemes. The analysis confirmed two key aspects of the HP-MR concept: (1) its ability to follow the power load requested from the heat pipe and (2) its ability to avoid heat pipe cascading failure unless designed with high power close to operating failure limits of its heat pipes. The developed computational model was distributed publicly on the Virtual Test Bed for training purposes to accelerate adoption by industry and to provide a high-fidelity multiphysics solution for benchmarking against other tools. Additional multiphysics analyses including other transients and coupled physics were identified as necessary future work, together with a focus on validating multiphysics behavior against experiments.

Microreactor↗

Rad-hard readout system for Timepix3 Hybrid Pixel Detectors

The Beam Gas Ionisation (BGI) profile monitor, located in the Proton Synchrotron (PS) and Super Proton Synchrotron (SPS) at CERN, requires a radiation-tolerant readout system to transfer data from the challenging accelerator surroundings to the back-end for processing. The system needs to control and acquire data from four Timepix3 Hybrid Pixel Detectors (HPDs) located directly inside the beam pipe, a highly radioactive environment. It must ensure reliability given limited hardware access and preserve signal integrity for the high-speed data (32 channels at 320 MHz). However, due to the unavailability of a suitable rad-hard Timepix3 readout, the Beam Instrumentation PiXeL (BIPXL) readout system was designed to meet these requirements. This system employs radiation-hardened components such as the GBTx and the FEASTMP, both developed at CERN. It will be compatible with forthcoming hybrid pixel detector initiatives in similarly harsh radiation conditions.

47 OTHER INSTRUMENTATION↗

Bayesian parameter estimation and evaluation of the K -ω shear stress transport model for plane impinging jets

Numerical simulations with semi-empirical turbulence models are commonly used to model impinging jets, often used for cooling solid surfaces. In this work, the constants in the k-ω shear stress transport model in ANSYS FLUENT are calibrated to experimental velocity and heat transfer data for a plane turbulent impinging air jet to determine if Kennedy-O'Hagan calibration (Kennedy and O'Hagan 2001 J. R. Stat. Soc. B 63 425–64) can improve predictions of near-surface velocities and surface Nusselt numbers for similar flows. Impinging jets have been proposed to cool the target plates of the divertor in future magnetic fusion energy reactors, where simulations are used to estimate divertor performance. The flat-plate divertor (Wang et al 2009 Fusion Sci. Technol .56 1023–7) uses a plane jet of helium issuing from a B = 0.5 mm slot to cool a surface with radius of curvature of 44 B at a distance 4 B from the slot. Predictions from the calibrated numerical model are compared with independent experimental data at different flow conditions, as well as surface temperature data for a flat plate divertor test section. The contribution of this work is evaluation of the accuracy of a calibrated turbulence model for modest extrapolations in flow geometry and flow conditions for a plane impinging jet.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

High-resolution crystal structure of the Borreliella burgdorferi PlzA protein in complex with c-di-GMP: new insights into the interaction of c-di-GMP with the novel xPilZ domain

ABSTRACT In the tick-borne pathogens, Borreliella burgdorferi and Borrelia hermsii, c-di-GMP is produced by a single diguanylate cyclase (Rrp1). In these pathogens, the Plz proteins (PlzA, B and C) are the only c-di-GMP receptors identified to date and PlzA is the sole c-di-GMP receptor found in all Borreliella isolates. Bioinformatic analyses suggest that PlzA has a unique PilZN3-PilZ architecture with the relatively uncommon xPilZ domain. Here, we present the crystal structure of PlzA in complex with c-di-GMP (1.6 Å resolution). This is the first structure of a xPilz domain in complex with c-di-GMP to be determined. PlzA has a two-domain structure, where each domain comprises topologically equivalent PilZ domains with minimal sequence identity but remarkable structural similarity. The c-di-GMP binding site is formed by the linker connecting the two domains. While the structure of apo PlzA could not be determined, previous fluorescence resonance energy transfer data suggest that apo and holo forms of the protein are structurally distinct. The information obtained from this study will facilitate ongoing efforts to identify the molecular mechanisms of PlzA-mediated regulation in ticks and mammals.

59 BASIC BIOLOGICAL SCIENCES↗

Skyrmion ratchet in funnel geometries

Here, using a particle-based model, we simulate the behavior of a skyrmion under the influence of asymmetric funnel geometries and ac driving at zero temperature. We specifically investigate possibilities for controlling the skyrmion motion by harnessing a ratchet effect. Our results show that as the amplitude of a unidirectional ac drive is increased, the skyrmion can be set into motion along either the easy or hard direction of the funnel depending on the ac driving direction. When the ac drive is parallel to the funnel axis, the skyrmion flows in the easy direction and its average velocity is quantized. In contrast, when the ac drive is perpendicular to the funnel axis, a Magnus-induced ratchet effect occurs, and the skyrmion moves along the hard direction with a constant average velocity. For biharmonic ac driving of equal amplitude along both the parallel and perpendicular directions, we observe a reentrant pinning phase where the skyrmion ratchet vanishes. For asymmetric biharmonic ac drives, the skyrmion exhibits a combination of effects and can move in either the easy or hard direction depending on the configuration of the ac drives. These results indicate that it is possible to achieve controlled skyrmion motion using funnel geometries, and we discuss ways in which this could be harnessed to perform data transfer operations.

36 MATERIALS SCIENCE↗