Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Field programmable gate array”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

End-to-end codesign of Hessian-aware quantized neural networks for FPGAs

Here, we develop an end-to-end workflow for the training and implementation of co-designed neural networks (NNs) for efficient field-programmable gate array (FPGA) hardware. Our approach leverages Hessian-aware quantization of NNs, the Quantized Open Neural Network Exchange intermediate representation, and the hls4ml tool flow for transpiling NNs into FPGA firmware. This makes efficient NN implementations in hardware accessible to nonexperts in a single open sourced workflow that can be deployed for real-time machine-learning applications in a wide range of scientific and industrial settings. We demonstrate the workflow in a particle physics application involving trigger decisions that must operate at the 40-MHz collision rate of the CERN Large Hadron Collider (LHC). Given the high collision rate, all data processing must be implemented on FPGA hardware within the strict area and latency requirements. Based on these constraints, we implement an optimized mixed-precision NN classifier for high-momentum particle jets in simulated LHC proton-proton collisions.

47 OTHER INSTRUMENTATION↗

The QICK (Quantum Instrumentation Control Kit): Readout and control for qubits and detectors

We introduce a Xilinx RF System-on-Chip (RFSoC)-based qubit controller (called the Quantum Instrumentation Control Kit, or QICK for short), which supports the direct synthesis of control pulses with carrier frequencies of up to 6 GHz. The QICK can control multiple qubits or other quantum devices. The QICK consists of a digital board hosting an RFSoC field-programmable gate array, custom firmware, and software and an optional companion custom-designed analog front-end board. We characterize the analog performance of the system as well as its digital latency, important for quantum error correction and feedback protocols. We benchmark the controller by performing standard characterizations of a transmon qubit. We achieve an average gate fidelity of ℱ avg =99.93%. All of the schematics, firmware, and software are open-source.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

ESnet/JLab FPGA Accelerated Transport

To increase the science rate for high data rates/volumes, Thomas Jefferson National Accelerator Facility (JLab) has partnered with Energy Sciences Network (ESnet) to define an edge to data center traffic shaping / steering transport capability featuring data event aware network shaping and forwarding. The keystone of this ESnet+JLab FPGA Accelerated Transport (EJFAT) is the joint development of an AI/ML directed dynamic compute work Load Balancer (LB) of UDP streamed data. The LB is a suite consisting of a Field Programmable Gate Array (FPGA) executing the dynamically configurable, low fixed latency LB data plane featuring real-time packet redirection and high throughput, and a control plane running on the FPGA host computer that monitors network and compute farm telemetry in order to make dynamic AI/ML guided decisions for destination compute host redirection/load balancing and destination resource provisioning. The LB provides for three-tier horizontal scaling across LB suites, core compute hosts, and CPUs within a host. The LB effectively provides seamless integration of edge/core computing to support direct experimental data processing for immediate use by JLab science programs and others such as the EIC as well as data centers of the future requiring high throughput and low latency for both hot and cooled data for both running experiment data acquisition systems and data center use cases.

97 MATHEMATICS AND COMPUTING↗

End-to-End Workflow for Machine-Learning-Based Qubit Readout With QICK and hls4ml

In this article, we present an end-to-end workflow for superconducting qubit readout that embeds codesigned neural networks into the quantum instrumentation control kit (QICK). Capitalizing on the custom firmware and software of the QICK platform, which is built on Xilinx radiofrequency system-on-chip field-programmable gate arrays (FPGAs), we aim to leverage machine learning (ML) to address critical challenges in qubit readout accuracy and scalability. The workflow utilizes the hls4ml package and employs quantization-aware training to translate ML models into hardware-efficient FPGA implementations via user-friendly Python application programming interfaces. We experimentally demonstrate the design, optimization, and integration of an ML algorithm for single transmon qubit readout, achieving 96% single-shot fidelity with a latency of 32.25 ns and less than 16% FPGA lookup table resource utilization. Our results offer the community an accessible workflow to advance ML-driven readout and adaptive control in quantum information processing applications.

42 ENGINEERING↗

A Re-programmable Platform for Dynamic Burn-in Test of Xilinx Virtexll 3000 FPGA for Military and Aerospace Applications

Field Programmable Gate Arrays (FPGA) have played increasingly important roles in military and aerospace applications. Xilinx SRAM-based FPGAs have been extensively used in commercial applications. They have been used less frequently in space flight applications due to their susceptibility to single-event upsets. Reliability of these devices in space applications is a concern that has not been addressed. The objective of this project is to design a fully programmable hardware/software platform that allows (but is not limited to) comprehensive static/dynamic burn-in test of Virtex-II 3000 FPGAs, at speed test and SEU test. Conventional methods test very few discrete AC parameters (primarily switching) of a given integrated circuit. This approach will test any possible configuration of the FPGA and any associated performance parameters. It allows complete or partial re-programming of the FPGA and verification of the program by using read back followed by dynamic test. Designers have full control over which functional elements of the FPGA to stress. They can completely simulate all possible types of configurations/functions. Another benefit of this platform is that it allows collecting information on elevation of the junction temperature as a function of gate utilization, operating frequency and functionality. A software tool has been implemented to demonstrate the various features of the system. The software consists of three major parts: the parallel interface driver, main system procedure and a graphical user interface (GUI).

Field Programmable Gate Arrays (FPGA)↗

Advanced Data Acquisition Systems

Current and future requirements of the aerospace sensors and transducers field make it necessary for the design and development of new data acquisition devices and instrumentation systems. New designs are sought to incorporate self-health, self-calibrating, self-repair capabilities, allowing greater measurement reliability and extended calibration cycles. With the addition of power management schemes, state-of-the-art data acquisition systems allow data to be processed and presented to the users with increased efficiency and accuracy. The design architecture presented in this paper displays an innovative approach to data acquisition systems. The design incorporates: electronic health self-check, device/system self-calibration, electronics and function self-repair, failure detection and prediction, and power management (reduced power consumption). These requirements are driven by the aerospace industry need to reduce operations and maintenance costs, to accelerate processing time and to provide reliable hardware with minimum costs. The project's design architecture incorporates some commercially available components identified during the market research investigation like: Field Programmable Gate Arrays (FPGA) Programmable Analog Integrated Circuits (PAC IC) and Field Programmable Analog Arrays (FPAA); Digital Signal Processing (DSP) electronic/system control and investigation of specific characteristics found in technologies like: Electronic Component Mean Time Between Failure (MTBF); and Radiation Hardened Component Availability. There are three main sections discussed in the design architecture presented in this document. They are the following: (a) Analog Signal Module Section, (b) Digital Signal/Control Module Section and (c) Power Management Module Section. These sections are discussed in detail in the following pages. This approach to data acquisition systems has resulted in the assignment of patent rights to Kennedy Space Center under U.S. patent # 6,462,684. Furthermore, NASA KSC commercialization office has issued licensing rights to Circuit Avenue Netrepreneurs, LLC , a minority-owned business founded in 1999 located in Camden, NJ.

Perotti, J.↗

Performance of A Real-Time Photon Counting Optical Receiver in the Presence of Emulated Channel Fading

Free-space optical communication links with terrestrial ground stations experience fading due to atmospheric scintillation and beam pointing. Fiber-coupled receiver systems experience additional fading at the interface between the fiber and free-space optics of the telescope. The National Aeronautics and Space Administration (NASA) Glenn Research Center (GRC) has characterized a real-time photon-counting optical ground receiver system with an atmospheric fade emulation system. The receiver system is comprised of a fiber interconnect, an array of superconducting nanowire single photon detectors (SNSPDs), and a field programmable gate array (FPGA) based receive modem. Two fiber interconnect/detector architectures have been studied. One architecture uses a 70-mode photonic lantern coupled to seven single pixel SNSPDs. The other architecture uses a 10-mode few-mode fiber (FMF) coupled to a 15-pixel SNSPD array. The receiver system complies with the Consultative Committee for Space Data Systems (CCSDS) Optical Communications High Photon Efficiency Coding and Synchronization Standard, which uses serially concatenated convolutionally coded pulse-position modulation (SCPPM). The CCSDS standard is designed for use in low photon flux missions, including the Orion Artemis-II Optical (O2O) communications demonstration. The standard utilizes a convolutional symbol interleaver which can be resized to mitigate different fades. The fade emulation system employed in this work emulates scintillation-induced, pointing-induced, and coupling-induced fading. This paper gives an overview of the real-time optical receiver system and the fade emulation system. It presents tests results which show the impact of fading on the performance on the receiver. The test results show that in the presence of channel fading, the 70-mode photonic lantern outperforms the 10-mode FMF under higher (D/r_0=9) turbulence conditions due to high fiber-coupling-induced fading and fiber coupling loss on the 10-mode FMF. When operating in lower turbulence (D/r_0=4), the 10-mode FMF outperforms the 70-mode photonic lantern. The paper also shows a larger convolutional interleaver improves the system performance as long as the receiver does not lose acquisition.

optical communications↗

Performance of a real-time photon counting optical receiver in the presence of emulated channel fading

Free-space optical communication links with terrestrial ground stations experience fading due to atmospheric scintillation and beam pointing. Fiber-coupled receiver systems experience additional fading at the interface between the fiber and free-space optics of the telescope. The National Aeronautics and Space Administration (NASA) Glenn Research Center (GRC) has characterized a real-time photon-counting optical ground receiver system with an atmospheric fade emulation system. The receiver system is comprised of a fiber interconnect, an array of superconducting nanowire single photon detectors (SNSPDs), and a field programmable gate array (FPGA) based receive modem. Two fiber interconnect/detector architectures have been studied. One architecture uses a 70-mode photonic lantern coupled to seven single pixel SNSPDs. The other architecture uses a 10-mode few-mode fiber (FMF) coupled to a 15-pixel SNSPD array. The receiver system complies with the Consultative Committee for Space Data Systems (CCSDS) Optical Communications High Photon Efficiency Coding and Synchronization Standard, which uses serially concatenated convolutionally coded pulse-position modulation (SCPPM). The CCSDS standard is designed for use in low photon flux missions, including the Orion Artemis-II Optical (O2O) communications demonstration. The standard utilizes a convolutional symbol interleaver which can be resized to mitigate different fades. The fade emulation system employed in this work emulates scintillation-induced, pointing-induced, and coupling-induced fading. This paper gives an overview of the real-time optical receiver system and the fade emulation system. It presents tests results which show the impact of fading on the performance on the receiver. The test results show that in the presence of channel fading, the 70-mode photonic lantern outperforms the 10-mode FMF under higher (D/r_0=9) turbulence conditions due to high fiber-coupling-induced fading and fiber coupling loss on the 10-mode FMF. When operating in lower turbulence (D/r_0=4), the 10-mode FMF outperforms the 70-mode photonic lantern. The paper also shows a larger convolutional interleaver improves the system performance as long as the receiver does not lose acquisition.

optical communications↗

Analog Module Architecture for Space-Qualified Field-Programmable Mixed-Signal Arrays

Spacecraft require all manner of both digital and analog circuits. Onboard digital systems are constructed almost exclusively from field-programmable gate array (FPGA) circuits providing numerous advantages over discrete design including high integration density, high reliability, fast turn-around design cycle time, lower mass, volume, and power consumption, and lower parts acquisition and flight qualification costs. Analog and mixed-signal circuits perform tasks ranging from housekeeping to signal conditioning and processing. These circuits are painstakingly designed and built using discrete components due to a lack of options for field-programmability. FPAA (Field-Programmable Analog Array) and FPMA (Field-Programmable Mixed-signal Array) parts exist but not in radiation-tolerant technology and not necessarily in an architecture optimal for the design of analog circuits for spaceflight applications. This paper outlines an architecture proposed for an FPAA fabricated in an existing commercial digital CMOS process used to make radiation-tolerant antifuse-based FPGA devices. The primary concerns are the impact of the technology and the overall array architecture on the flexibility of programming, the bandwidth available for high-speed analog circuits, and the accuracy of the components for high-performance applications.

Edwards, R. Timothy↗

Optimization with the OpenACC-to-FPGA framework on the Arria 10 and Stratix 10 FPGAs

The reconfigurable computing paradigm with field programmable gate arrays (FPGAs) has received renewed interest in the high-performance computing field due to FPGAs’ unique combination of performance and energy efficiency. However, difficulties in programming and optimizing FPGAs have prevented them from being widely accepted as general-purpose computing devices. In accelerator-based heterogeneous computing, portability across diverse heterogeneous devices is also an important issue, but the unique architectural features in FPGAs make this difficult to achieve. To address these issues, a directive-based, high-level FPGA programming and optimization framework was previously developed. In this work, developed optimizations were combined holistically using the directive-based approach to show that each individual benchmark requires a unique set of optimizations to maximize performance. We perform this exploration on Intel Arria 10 and Stratix 10 FPGAs. We also explored the relationships between performance, resource usages, and compilation times, and investigated implications for performance portability. Finally, we present an initial evaluation of a real-world proxy application, LULESH.

97 MATHEMATICS AND COMPUTING↗

Walsh function generator for the Electronically Scanned Thinned Array Radiometer (ESTAR) instrument

A prototype Walsh Function Generator (WFG) for the ESTAR (Electronically Scanned Thinned Array Radiometer) instrument has been designed and tested. Implemented in a single Xilinx XC3020PC68-50 Field Programmable Gate Array (FPGA), it generates a user-programmable set of 32 consecutive Walsh Functions for noise cancellation in the analog circuitry of the Front-End Modules (FEM's). It is implemented in a 68-pin plastic leaded chip carrier (PLCC) package, is fully testable, and can be used for noise cancellation periods as small as 2 msec.

Chren, William A., Jr.↗

Phase aligner for the Electronically Scanned Thinned Array Radiometer (ESTAR) instrument

A prototype Phase Aligner (PA) or the Electronically Scanned Thinned Array Radiometer instrument has been designed and tested. Implemented in a single Xilinx XC3042PC84-125 Field Programmable Gate Array (FPGA), it is a dual-port register file which allows independent storage and phase coherent retrieval of antenna array data by the Central Processing Unit (CPU). It has dimensions of 4 x 20 bits and can be used at clock frequencies as high as 25 MHz. The ESTAR is a passive synthetic-aperture radiometer designed to sense soil moisture and ocean salinity at L-band.

Chren, William A., Jr.↗

Dynamically Reconfigurable Systolic Array Accelerator

A polymorphic systolic array framework has been developed that works in conjunction with an embedded microprocessor on a field-programmable gate array (FPGA), which allows for dynamic and complimentary scaling of acceleration levels of two algorithms active concurrently on the FPGA. Use is made of systolic arrays and a hardware-software co-design to obtain an efficient multi-application acceleration system. The flexible and simple framework allows hosting of a broader range of algorithms, and is extendable to more complex applications in the area of aerospace embedded systems. FPGA chips can be responsive to realtime demands for changing applications needs, but only if the electronic fabric can respond fast enough. This systolic array framework allows for rapid partial and dynamic reconfiguration of the chip in response to the real-time needs of scalability, and adaptability of executables.

Dasu, Aravind↗

Column Grid Array Rework for High Reliability

Due to requirements for reduced size and weight, use of grid array packages in space applications has become common place. To meet the requirement of high reliability and high number of I/Os, ceramic column grid array packages (CCGA) were selected for major electronic components used in next MARS Rover mission (specifically high density Field Programmable Gate Arrays). ABSTRACT The probability of removal and replacement of these devices on the actual flight printed wiring board assemblies is deemed to be very high because of last minute discoveries in final test which will dictate changes in the firmware. The questions and challenges presented to the manufacturing organizations engaged in the production of high reliability electronic assemblies are, Is the reliability of the PWBA adversely affected by rework (removal and replacement) of the CGA package? and How many times can we rework the same board without destroying a pad or degrading the lifetime of the assembly? To answer these questions, the most complex printed wiring board assembly used by the project was chosen to be used as the test vehicle, the PWB was modified to provide a daisy chain pattern, and a number of bare PWB s were acquired to this modified design. Non-functional 624 pin CGA packages with internal daisy chained matching the pattern on the PWB were procured. The combination of the modified PWB and the daisy chained packages enables continuity measurements of every soldered contact during subsequent testing and thermal cycling. Several test vehicles boards were assembled, reworked and then thermal cycled to assess the reliability of the solder joints and board material including pads and traces near the CGA. The details of rework process and results of thermal cycling are presented in this paper.

high reliability↗

RFI Risk Reduction Activities Using New Goddard Digital Radiometry Capabilities

The Goddard Radio-Frequency Explorer (GREX) is the latest fast-sampling radiometer digital back-end processor that will be used for radiometry and radio-frequency interference (RFI) surveying at Goddard Space Flight Center. The system is compact and deployable, with a mass of about 40 kilograms. It is intended to be flown on aircraft. GREX is compatible with almost any aircraft, including P-3, twin otter, C-23, C-130, G3, and G5 types. At a minimum, the system can function as a clone of the Soil Moisture Active Passive (SMAP) ground-based development unit [1], or can be a completely independent system that is interfaced to any radiometer, provided that frequency shifting to GREX's intermediate frequency is performed prior to sampling. If the radiometer RF is less than 200MHz, then the band can be sampled and acquired directly by the system. A key feature of GREX is its ability to simultaneously sample two polarization channels simultaneously at up to 400MSPS, 14-bit resolution each. The sampled signals can be recorded continuously to a 23 TB solid-state RAID storage array. Data captures can be analyzed offline using the supercomputing facilities at Goddard Space Flight Center. In addition, various Field Programmable Gate Array (FPGA) - amenable radiometer signal processing and RFI detection algorithms can be implemented directly on the GREX system because it includes a high-capacity Xilinx Virtex-5 FPGA prototyping system that is user customizable.

Bradley, Damon↗

Searching for metastable particles using graph computing

The reconstruction of charged particle trajectories at the Large Hadron Collider and future colliders relies on energy depositions in sensors placed at distances ranging from a centimeter to a meter from the colliding beams. We propose a method of detecting charged particles that decay invisibly after traversing a short distance of about 25 cm inside the experimental apparatus. One of the decay products may constitute the dark matter known to be 84% of all matter at galactic and cosmological distance scales. Our method uses graph computing to cluster spacepoints recorded by two-dimensional silicon pixel sensors into mathematically-defined patterns. The algorithm may be implemented on silicon-based integrated circuits using field-programmable gate array technology to augment or replace traditional computing platforms.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DRIPS: Dynamic Rebalancing of Pipelined Streaming Applications on CGRAs

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit data streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across application kernels. However, emerging streaming applications at the edge (scientific instruments, sensor networks, network processing) perform much more than digital signal processing and often are data and input dependent. This leads to extremely variable kernel execution times, severely impacting the throughput of the entire pipeline if resources are only statically allocated. Therefore, in this paper, we propose DRIPS — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We present a unified compiler framework to facilitate the mapping of a given streaming application onto the DRIPS CGRA architecture. The experimental results show that DRIPS achieves an average throughput improvement of 1.46$\times$ across a set of representative applications over a statically partitioned solution. The additional area overhead to enable dynamic rebalancing consumes 16.34% of the entire area for a 5x5 CGRA prototype.

Tan, Cheng↗

DynPaC: Coarse-Grained, Dynamic, and Partially Reconfigurable Array for Streaming Applications

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across kernels. However, streaming applications often are data-dependent, leading to variable kernel execution times depending on the input data and impacting the throughput of the entire pipeline if resources are statically allocated. Therefore, in this paper, we discuss the design of DynPaC — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We discuss the required software and hardware components to manage partial dynamic reconfiguration. We demonstrate that by supporting partial dynamic reconfiguration, we can obtain an average speedup of 1.44X for a representative set of applications w.r.t. static partitioning, with a limited area overhead (6.4% of the entire chip).

Tan, Cheng↗