Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Field programmable gate array”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Output data formatter for the Electronically Scanned Thinned Array Radiometer (ESTAR) instrument

A prototype Output Data Formatter (ODF) for the ESTAR (Electronically Scanned Thinned Array Radiometer) instrument has been designed and tested. It employs programmable logic devices to format and tag correlator data for transmission to Earth. After accepting 170 bits or correlator and error data in parallel, it appends an identification word and then serially passes the data to the Small Explorer Data System (SEDS) for transmission at a maximum rate of greater than 15 Mb/sec. Implemented with two reprogrammable field programmable gate arrays (FPGA's), each contained in a 132-pin plastic pin grid array (PGA) package, the design is cascadeable, fully testable, and low-power.

Chren, William A., Jr.↗

FPGA Coprocessor Design for an Onboard Multi-Angle Spectro-Polarimetric Imager

A multi-angle spectro-polarimetric imager (MSPI) is an advanced camera system currently under development at JPL for possible future consideration on a satellite-based Aerosol-Cloud-Environ - ment (ACE) interaction study. The light in the optical system is subjected to a complex modulation designed to make the overall system robust against many instrumental artifacts that have plagued such measurements in the past. This scheme involves two photoelastic modulators that are beating in a carefully selected pattern against each other. In order to properly sample this modulation pattern, each of the proposed nine cameras in the system needs to read out its imager array about 1,000 times per second. The onboard processing required to compress this data involves least-squares fits (LSFs) of Bessel functions to data from every pixel in realtime, thus requiring an onboard computing system with advanced data processing capabilities in excess of those commonly available for space flight. As a potential solution to meet the MSPI onboard processing requirements, an LSF algorithm was developed on the Xilinx Virtex-4FX60 field programmable gate array (FPGA). In addition to configurable hardware capability, this FPGA includes Power -PC405 microprocessors, which together enable a combination hardware/ software processing system. A laboratory demonstration was carried out based on a hardware/ software co-designed processing architecture that includes hardware-based data collection and least-squares fitting (computationally), and softwarebased transcendental function computation (algorithmically complex) on the FPGA. Initial results showed that these calculations can be handled using a combination of the Virtex- 4TM Power-PC core and the hardware fabric.

Pingree, Paula J.↗

Controller for the Electronically Scanned Thinned Array Radiometer (ESTAR) instrument

A prototype controller for the ESTAR (electronically scanned thinned array radiometer) instrument has been designed and tested. It manages the operation of the digital data subsystem (DDS) and its communication with the Small Explorer data system (SEDS). Among the data processing tasks that it coordinates are FEM data acquisition, noise removal, phase alignment and correlation. Its control functions include instrument calibration and testing of two critical subsystems, the output data formatter and Walsh function generator. It is implemented in a Xilinx XC3064PC84-100 field programmable gate array (FPGA) and has a maximum clocking frequency of 10 MHz.

Zomberg, Brian G.↗

Digital Interface Board to Control Phase and Amplitude of Four Channels

An increasing number of parts are designed with digital control interfaces, including phase shifters and variable attenuators. When designing an antenna array in which each antenna has independent amplitude and phase control, the number of digital control lines that must be set simultaneously can grow very large. Use of a parallel interface would require separate line drivers, more parts, and thus additional failure points. A convenient form of control where single-phase shifters or attenuators could be set or the whole set could be programmed with an update rate of 100 Hz is needed to solve this problem. A digital interface board with a field-programmable gate array (FPGA) can simultaneously control an essentially arbitrary number of digital control lines with a serial command interface requiring only three wires. A small set of short, high-level commands provides a simple programming interface for an external controller. Parity bits are used to validate the control commands. Output timing is controlled within the FPGA to allow for rapid update rates of the phase shifters and attenuators. This technology has been used to set and monitor eight 5-bit control signals via a serial UART (universal asynchronous receiver/transmitter) interface. The digital interface board controls the phase and amplitude of the signals for each element in the array. A host computer running Agilent VEE sends commands via serial UART connection to a Xilinx VirtexII FPGA. The commands are decoded, and either outputs are set or telemetry data is sent back to the host computer describing the status and the current phase and amplitude settings. This technology is an integral part of a closed-loop system in which the angle of arrival of an X-band uplink signal is detected and the appropriate phase shifts are applied to the Ka-band downlink signal to electronically steer the array back in the direction of the uplink signal. It will also be used in the non-beam-steering case to compensate for phase shift variations through power amplifiers. The digital interface board can be used to set four 5-bit phase shifters and four 5-bit attenuators and monitor their current settings. Additionally, it is useful outside of the closed-loop system for beamsteering alone. When the VEE program is started, it prompts the user to initialize variables (to zero) or skip initialization. After that, the program enters into a continuous loop waiting for the telemetry period to elapse or a button to be pushed. A telemetry request is sent when the telemetry period is elapsed (every five seconds). Pushing one of the set or reset buttons will send the appropriate command. When a command is sent, the interface status is returned, and the user will be notified by a pop-up window if any error has occurred. The program runs until the End Program button is depressed.

Smith, Amy E.↗

Implementing machine learning methods on QICK hardware for qubit readout & control

Quantum readout and control is a fundamental aspect of quantum computing that requires accurate measurement of qubit states. Errors emerge in all stages, from initialization to readout, and identifying errors in post-processing necessitates resource-intensive statistical analysis. In our work, we use a lightweight fully-connected neural network (NN) to classify states of a transmon system with no prior processing. Our NN accelerator yields higher fidelities (92%) than the classical matched filter method (84%). By exploiting the natural parallelism of NNs and their placement near the source of data on field-programmable gate arrays (FPGAs), we can achieve ultra-low latency on the Quantum Instrumentation Control Kit (QICK). Integrating machine learning methods on QICK opens several pathways for efficient real-time processing of quantum circuits.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Distilling particle knowledge for fast reconstruction at high-energy physics experiments

Knowledge distillation is a form of model compression that allows artificial neural networks of different sizes to learn from one another. Its main application is the compactification of large deep neural networks to free up computational resources, in particular on edge devices. In this article, we consider proton-proton collisions at the High-Luminosity Large Hadron Collider (HL-LHC) and demonstrate a successful knowledge transfer from an event-level graph neural network (GNN) to a particle-level small deep neural network (DNN). Our algorithm, DistillNet, is a DNN that is trained to learn about the provenance of particles, as provided by the soft labels that are the GNN outputs, to predict whether or not a particle originates from the primary interaction vertex. The results indicate that for this problem, which is one of the main challenges at the HL-LHC, there is minimal loss during the transfer of knowledge to the small student network, while improving significantly the computational resource needs compared to the teacher. This is demonstrated for the distilled student network on a CPU, as well as for a quantized and pruned student network deployed on a field programmable gate array. Our study proves that knowledge transfer between networks of different complexity can be used for fast artificial intelligence (AI) in high-energy physics that improves the expressiveness of observables over non-AI-based reconstruction algorithms. Such an approach can become essential at the HL-LHC experiments, e.g. to comply with the resource budget of their trigger stages.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Digitally Controlled Slot Coupled Patch Array

A four-element array conformed to a singly curved conducting surface has been demonstrated to provide 2 dB axial ratio of 14 percent, while maintaining VSWR (voltage standing wave ratio) of 2:1 and gain of 13 dBiC. The array is digitally controlled and can be scanned with the LMS Adaptive Algorithm using the power spectrum as the objective, as well as the Direction of Arrival (DoA) of the beam to set the amplitude of the power spectrum. The total height of the array above the conducting surface is 1.5 inches (3.8 cm). A uniquely configured microstrip-coupled aperture over a conducting surface produced supergain characteristics, achieving 12.5 dBiC across the 2-to-2.13- GHz and 2.2-to-2.3-GHz frequency bands. This design is optimized to retain VSWR and axial ratio across the band as well. The four elements are uniquely configured with respect to one another for performance enhancement, and the appropriate phase excitation to each element for scan can be found either by analytical beam synthesis using the genetic algorithm with the measured or simulated far field radiation pattern, or an adaptive algorithm implemented with the digitized signal. The commercially available tuners and field-programmable gate array (FPGA) boards utilized required precise phase coherent configuration control, and with custom code developed by Nokomis, Inc., were shown to be fully functional in a two-channel configuration controlled by FPGA boards. A four-channel tuner configuration and oscilloscope configuration were also demonstrated although algorithm post-processing was required.

D'Arista, Thomas↗

G(sup 4)FET Implementations of Some Logic Circuits

Some logic circuits have been built and demonstrated to work substantially as intended, all as part of a continuing effort to exploit the high degrees of design flexibility and functionality of the electronic devices known as G(sup 4)FETs and described below. These logic circuits are intended to serve as prototypes of more complex advanced programmable-logicdevice-type integrated circuits, including field-programmable gate arrays (FPGAs). In comparison with prior FPGAs, these advanced FPGAs could be much more efficient because the functionality of G(sup 4)FETs is such that fewer discrete components are needed to perform a given logic function in G(sup 4)FET circuitry than are needed perform the same logic function in conventional transistor-based circuitry. The underlying concept of using G(sup 4)FETs as building blocks of programmable logic circuitry was also described, from a different perspective, in G(sup 4)FETs as Universal and Programmable Logic Gates (NPO-41698), NASA Tech Briefs, Vol. 31, No. 7 (July 2007), page 44. A G(sup 4)FET can be characterized as an accumulation-mode silicon-on-insulator (SOI) metal oxide/semiconductor field-effect transistor (MOSFET) featuring two junction field-effect transistor (JFET) gates. The structure of a G(sup 4)FET (see Figure 1) is the same as that of a p-channel inversion-mode SOI MOSFET with two body contacts on each side of the channel. The top gate (G1), the substrate emulating a back gate (G2), and the junction gates (JG1 and JG2) can be biased independently of each other and, hence, each can be used to independently control some aspects of the conduction characteristics of the transistor. The independence of the actions of the four gates is what affords the enhanced functionality and design flexibility of G(sup 4)FETs. The present G(sup 4)FET logic circuits include an adjustable-threshold inverter, a real-time-reconfigurable logic gate, and a dynamic random-access memory (DRAM) cell (see Figure 2). The configuration of the adjustable-threshold inverter is similar to that of an ordinary complementary metal oxide semiconductor (CMOS) inverter except that an NMOSFET (a MOSFET having an n-doped channel and a p-doped Si substrate) is replaced by an n-channel G(sup 4)FET

Mojarradi, Mohammad↗

Evolution of the ATLAS event data model for the HL-LHC

The upcoming high-luminosity run of the CERN Large Hadron Collider (HL-LHC) will yield an unprecedented volume of data. In order to process this data, the ATLAS collaboration is evolving its offline software to be able to use heterogeneous resources such as graphical processing units (GPUs) and field-programmable gate arrays (FPGAs). To reduce conversion overheads, the event data model (EDM) should be compatible with the requirements of these resources. While the ATLAS EDM has long allowed representing data as a structure of arrays, further evolution of the EDM can enable more efficient sharing of data between CPU and GPU resources. Some of this work will be summarized here, including extensions to allow controlling how memory for event data is allocated and the implementation of jagged vectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An Experimental Feasibility Study on Applying Neutron Tomography to Encapsulated Spent Nuclear Fuel - 20050

Visual inspection makes easier to ensure the integrity and safety of spent nuclear fuel (SNF) than any classical techniques. Various classical techniques have been applied but there are no reliable methods to qualitatively and quantitatively verify spent fuel in dry storage. Thus, the present authors have developed the prototype safeguards apparatus for dry storage employing the array of He-4 gas scintillation detectors (S670E, Arktis Radiation Detectors Ltd., Switzerland), newly designed to simultaneously measure thermal and fast neutrons without any moderators. The S670E detector has a cylindrical shape with a diameter of 52 mm and active length of 600 mm (total length: 875 mm). The detector is filled by He-4 gas with an approximate pressure of 180 bar for fast neutron detection, and its inner wall is coated by Li-6 for thermal neutron detection. The scintillation lights generated via Li-6 nuclear reaction and elastic scattering are collected by 24 SiPMs linearly paired at the center of the detector. The detector delivers a TTL (Transistor-Transistor Logic) output for pulse readout and UART (Universal Asynchronous Receiver Transmitter) for device control. In order to assess feasibility of the apparatus, an experimental system has been designed, built, and optimized via computational studies. Cf-252 neutron sources and linearly arrayed detectors, working as a single detector, were occupied for this study due to the difficulties in working with SNF. The laboratory scale cask (diameter: 0.67 m, height: 1.5 m), minimized by a factor of 10 compared to the actual thickness of a commercial TN-32 cask, was also manufactured. The detector array was designed to rotate the lab-scale cask and obtain 36 image profiles at every 10 degrees. All profiles were aligned in single frame image called a sinogram, and the cross-sectional image was then fabricated by the inverse radon transform algorithm. These experiments have been repeated with different configurations and numbers of sources. Some gamma-ray sources were also measured with neutron sources in order to distinguish between neutron and gamma-ray pulses. Basically, a He-4 detector is designed to run on Linux OS so it is difficult to directly apply to Windows-based equipment widely used in S. Korea. Therefore, a new data acquisition board working on Windows OS was designed and built. The board mainly consists of FPGA (Field Programmable Gate Array) and SoC (System on Chip) for TTL pulse readout, sorting measured data, and transferring data to a user interface. In conclusion, the tomographic system at lab-scale has shown considerable potential to detect a partial or gross defect of encapsulated assemblies in dry storage. Next steps of this study will be to 1) repeatedly carry out experiments to demonstrate scientific reliability and validity, and 2) numerically integrate signals with weight factors to enhance the image quality since the suggested system based on passive interrogation method requires longer measurement time. Finally, the system will apply to a commercial dry storage phased out soon in S. Korea. (authors)

07 ISOTOPE AND RADIATION SOURCES↗

Thermal Cycle Reliability of PBGA/CCGA 717 I/Os

Status of thermal cycle test results for a nonfunctional daisy-chained peripheral ceramic column grid array (CCGA) and its plastic ball grid array (PBGA) version. both having 560 I/Os. were presented in last year's conference. Test results included environmental data for three different thermal cycle regimes (-55 C/125 C, -55 C/100 C, and -50 C /75 C). Update information on these - especially failure type for assemblies with high and low solder volumes-were presented. The thermal cycle test procedure followed those recommended IPC-9701 for tin-lead solder joint assemblies. Revision A of this specification covers guideline thermal cycle requirements for Pb-free solder joints. Some background information discussed during release of this specification with its current guideline recommendations were also presented. In a recent reliability investigation a fully populated CCGA with 717 I/Os was also considered for assembly reliability. evaluation. The functional package is a field-programmable gate array that has much higher processing power than its previous version. This new package is smaller in dimension, has no interposer, and has a thinner column wrapped with copper for reliability improvement. This paper will also present thermal cycle test results for this package assembly and its plastic version with 728 I/Os. both of which were exposed to three different cycle regimes. Two cycle profiles were those specified by IPC- 9701A for tin-lead, i.e. -55 to 100 C and -55 to 125 C and one was a cycle profile specified by Mil-Std-883, i.e.. -65 C/150 C which is generally used for ceramic hybrid packages. Per IPC-9701 A, test vehicles were built using daisy chain packages and were continuously monitored. The effects of many process and assembly variables-including corner staking commonly used for improving resistance to mechanical loading such as drop and vibration loads--were also considered as part of the test matrix. Optical photomicrographs were taken at various thermal cycle intervals to document damage progress and behavior. Representative samples of these along with cross-sectional photomicrographs at higher magnification taken by scanning electron microscopy (SEM) to determine crack propagation and failure analyses for packages are also presented.

staking↗

Compact Ku-Band T/R Module for High-Resolution Radar Imaging of Cold Land Processes

Global measurement of terrestrial snow cover is critical to two of the NASA Earth Science focus areas: (1) climate variability and change and (2) water and energy cycle. For radar backscatter measurements, Ku-band frequencies, scattered mainly within the volume of the snowpack, are most suitable for the SWE (snow-water equivalent) measurements. To isolate the complex effects of different snowpack (density and snowgrain size), and underlying soil properties and to distinctly determine SWE, the space-based synthetic aperture radar (SAR) system will require a dual-frequency (13.4 and 17.2 GHz) and dual polarization approach. A transmit/receive (T/R) module was developed operating at Ku-band frequencies to enable the use of active electronic scanning phased-array antenna for wide-swath, high-resolution SAR imaging of terrestrial snow cover. The T/R module has an integrated calibrator, which compensates for all environmental- and time-related changes, and results in very stable power and amplitude characteristics. The module was designed to operate over the full frequency range of 13 to 18 GHz, although only the two frequencies, 13.4 GHz and 17.2 GHz, will be used in this SAR radar application. Each channel of the transmit module produces > 4 W (35 dbm) over the operating bandwidth of 20 MHz. The stability requirements of <0.1 dB receive gain accuracy and <0.1 dB transmit power accuracy over a wide temperature range are achieved using a self-correction scheme, which does real-time amplitude calibration so that the module characteristics are continually corrected. All the calibration circuits are within the T/R module. The timing and calibration sequence is stored in a control FPGA (field-programmable gate array) while an internal 128K 8bit high-speed RAM (random access memory) stores all the calibration values. The module was designed using advanced components and packaging techniques to achieve integration of the electronics in a 2 x6.5x1-in. (5x17x2.5-cm) package. The module size allows 4 T/R modules to feed the 16 16-element subarray on an antenna panel. The T/R module contains four transmit channels and eight receive channels (horizontal and vertical polarizations).

Andricos, Constantine↗

Nanosecond anomaly detection with decision trees and real-time application to exotic Higgs decays

Abstract We present an interpretable implementation of the autoencoding algorithm, used as an anomaly detector, built with a forest of deep decision trees on FPGA, field programmable gate arrays. Scenarios at the Large Hadron Collider at CERN are considered, for which the autoencoder is trained using known physical processes of the Standard Model. The design is then deployed in real-time trigger systems for anomaly detection of unknown physical processes, such as the detection of rare exotic decays of the Higgs boson. The inference is made with a latency value of 30 ns at percent-level resource usage using the Xilinx Virtex UltraScale+ VU9P FPGA. Our method offers anomaly detection at low latency values for edge AI users with resource constraints.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Searching for Clues for a Matter Dominated Universe in Liquid Argon Time Projection Chambers

Liquid Argon Time Projection Chambers (LArTPCs) represent one of the most widely utilized neutrino detection techniques in neutrino experiments, for instance, in the Short Baseline Neutrino (SBN) program and the future large-scale LArTPC: Deep Underground Neutrino Experiment (DUNE). The high-end technique, facilitating excellent spatial and calorimetric reconstruction resolution, also enables testing exotic Beyond Standard Model (BSM) theories, such as baryon number violation (BNV) processes (e.g., proton-decay, neutron-antineutron oscillation). At the same time, Machine Learning (ML) techniques have demonstrated their ubiquitous use in recent decades; ML techniques have also become some of the most powerful tools in high-energy physics (HEP) analyses. Furthermore, the development of algorithms to cater to the needs of problems in HEP (i.e., triggering, reconstruction, improving sensitivity, etc.) has also become an active area of research. By developing a combined approach using Convolutional Neural Network (CNN) and Boosted Decision Tree (BDT) techniques, the sensitivity of neutron-antineutron oscillation in DUNE is evaluated for a projected exposure of 400kton&middot;years. Additionally, to meet the triggering requirement to select such rare events in DUNE, such a search is only supported with highly efficient self-triggering algorithms. An ML-based self-triggering scheme for large-scale LArTPCs, such as DUNE, is also developed with the intention of implementation on field-programmable gate arrays (FPGAs). The ML-based approach for searching for neutron-antineutron oscillation can be demonstrated and validated on the current LArTPC MicroBooNE. The analysis in MicroBooNE represents the first-ever search for neutron-antineutron oscillation in a LArTPC. DUNE's projected 90% C.L. sensitivity to the neutron antineutron oscillation lifetime is 6.45&times;10³² years, assuming 1.327&times;10³⁵ neutron&middot;years, equivalent to 10 years of DUNE far detector exposure (400kton&middot;years). For MicroBooNE, assuming 372 seconds of exposure (equivalent to 3.13&times;10³⁶ neutron&middot;years), the 90% C.L. lifetime sensitivity is found at 3.07&times;10²⁵ yrs, after accounting for Monte-Carlo statistical uncertainty and systematic uncertainty from detector effects.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The Design and Development of the SMEX-Lite Power System

This paper describes the design and development of a 250W orbit average electrical power system electronic Power Node and software for use in Low Earth Orbit missions. The mass of the Power Node is 3.6 Kg (8 lb.). The dimensions of the Power Node are 30cm x 26cm x 7.9cm (11 in. x 10.25 in x 3.1 in.) The design was realized using software, Field Programmable Gate Array (FPGA) digital logic and surface mount technology. The design is generic enough to reduce the non-recurring engineering for different mission configurations. The Power Node charges one to five, low cost, 22-cell 4 AH D-cell battery packs independently. The battery charging algorithms are executed in the power software to reduce the mass and size of the power electronic. The Power Node implements a peak-power tracking algorithm using an innovative hardware/software approach. The power software task is hosted on the spacecraft processor. The power software task generates a MIL-STD-1553 command packet to update the Power Node control settings. The settings for the battery voltage and current limits, as well as minimum solar array voltage used to implement peak power tracking are contained in this packet. Several advanced topologies are used in the Power Node. These include synchronous rectification in the bus regulators, average current control in the battery chargers and quasi-resonant converters for the Field Effect Transistor (FET) transistor drive electronics. Lastly, the main bus regulator uses a feed-forward topology with the PWM implemented in an FPGA.

Rakow, Glenn P.↗

Fast convolutional neural networks on FPGAs with hls4ml

We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on field-programmable gate arrays (FPGAs). By extending the hls4ml library, we demonstrate an inference latency of 5 µs using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

wa-hls4ml: A GNN Surrogate Model for hls4ml

Recent advancements in use of machine learning techniques on field-programmable gate arrays (FPGAs) have allowed for implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area must be strictly bounded. The hls4ml framework is a procedure for converting from trained machine learning model software, to a synthesis result that can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it is possible that the model is unable to be converted into a synthesis result, or that the resource consumption of the model will exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model which uses a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of an arbitrary model when passed through the hls4ml procedure, without the time consumption of actually running the pipeline.

43 PARTICLE ACCELERATORS↗

Bringing heterogeneity to the CMS software framework

The advent of computing resources with co-processors, for example Graphics Processing Units (GPU) or Field-Programmable Gate Arrays (FPGA), for use cases like the CMS High-Level Trigger (HLT) or data processing at leadership-class supercomputers imposes challenges for the current data processing frameworks. These challenges include developing a model for algorithms to offload their computations on the co-processors as well as keeping the traditional CPU busy doing other work. The CMS data processing framework, CMSSW, implements multithreading using the Intel Threading Building Blocks (TBB) library, that utilizes tasks as concurrent units of work. In this paper we will discuss a generic mechanism to interact effectively with non-CPU resources that has been implemented in CMSSW. In addition, configuring such a heterogeneous system is challenging. In CMSSW an application is configured with a configuration file written in the Python language. The algorithm types are part of the configuration. The challenge therefore is to unify the CPU and co-processor settings while allowing their implementations to be separate. We will explain how we solved these challenges while minimizing the necessary changes to the CMSSW framework. We will also discuss on a concrete example how algorithms would offload work to NVIDIA GPUs using directly the CUDA API.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗