Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Xilinx”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

NEPP Update of Independent Single Event Upset Field Programmable Gate Array Testing

This presentation provides a NASA Electronic Parts and Packaging (NEPP) Program update of independent Single Event Upset (SEU) Field Programmable Gate Array (FPGA) testing including FPGA test guidelines, Microsemi RTG4 heavy-ion results, Xilinx Kintex-UltraScale heavy-ion results, Xilinx UltraScale+ single event effect (SEE) test plans, development of a new methodology for characterizing SEU system response, and NEPP involvement with FPGA security and trust.

Field Programmable Gate Array (FPGA); Triple Modul↗

Virtex-4VQ dynamic and mitigated single event upset characterization summary

This report is the result of funding by the NASA Electronic Parts and Packaging Program (NEPP) and the combined efforts of members within the Xilinx Radiation Test Consortium (XRTC), sometimes known as the Xilinx Single Event Effects (SEE) Test Consortium. The XRTC is a voluntary association of aerospace entities, including leading aerospace companies, universitie and national laboratories, combining resources to characterize reconfigurable, field programmable gate arrays (FPGAs) for aerospace applications. Previous publications of Virtex-4 radiation results are for commercial (non-epitaxial) devices; see, for example, Refs. 1–5. A notable exception is Ref. 6, which presents XRTC upset measurements of storage elements in the PowerPC405s in the XQR4VFX60. This work represents a continuation of the efforts reported in the “Virtex-4QV Static SEU Characterization Summary” [7]. The contents of this report describe various Single Event Functional Interrupt (SEFI) and Single Event Upset (SEU) modes seen while dynamically exercising the clocked resources within Virtex-4 devices and the corresponding mitigation techniques related to the aforementioned observed SEFI modes.

Allen, Gregory↗

NOR Flash Memory Scrubbing Application for Boot File Preservation of NASA’s Descent and Landing Computer (DLC)

Progress on NASA’s Safe and Precise Landing Integrated Capabilities Evolution (SPLICE) project continues, specifically with this development of the Descent and Landing Computer (DLC). One of the DLC’s primary contributions as a SPLICE technology is its implementation of algorithms and operation of sensors to autonomously guide a spacecraft in performing more precise and safer landings on celestial bodies such as the Moon and Mars [1]. The second iteration of the DLC is known as the Engineering Test Unit (ETU) and one of its desired functionalities is the ability to preserve the fidelity of the system’s boot file through the use of memory scrubbing [2]. The ETU has two primary boards, one for housing a Multi-Processor System on a Chip (MPSoC) and the other for housing a Xilinx Kintex Ultrascale FPGA*. To emulate a memory scrubbing function implemented on the ETU’s FPGA board, the design and testing of a software application was performed on a Xilinx KCU105 FPGA evaluation board. The memory scrubbing application had to meet certain key criteria such as (1) properly utilize with the flash memory’s Serial Peripheral Interface (SPI) to read, write, and erase flash memory properly, (2) be able to detect arbitrarily large or small amounts of bit-errors, (3) be able to correct all detected errors, and (4) perform memory scrubbing indefinitely and autonomously. A prototype implementation was constructed and tested, demonstrating successful detection and correction of bit errors in multiple configurations. In the form of burst errors or singular bit flips, and in amounts of errors ranging from one to fifteen (per 256 Bytes), the application was successful in preserving memory fidelity.

Flash memory↗

NOR Flash Memory Scrubbing Application for Boot File Preservation of NASA’s Descent and Landing Computer (DLC)

Progress on NASA’s Safe and Precise Landing Integrated Capabilities Evolution (SPLICE)project continues, specifically with this development of the Descent and Landing Computer(DLC). One of the DLC’s primary contributions as a SPLICE technology is its implementationof algorithms and operation of sensors to autonomously guide a spacecraft in performing moreprecise and safer landings on celestial bodies such as the Moon and Mars. The second iterationof the DLC is known as the Engineering Test Unit (ETU) and one of its desired functionalitiesis the ability to preserve the fidelity of the system’s boot file through the use of memoryscrubbing. The ETU has two primary boards, one for housing a Multi-Processor System ona Chip (MPSoC) and the other for housing a Xilinx Kintex Ultrascale FPGA3. To emulate amemory scrubbing function implemented on the ETU’s FPGA board, the design and testingof a software application was performed on a Xilinx KCU105 FPGA evaluation board. Thememory scrubbing application had to meet certain key criteria such as (1) properly utilize withthe flash memory’s Serial Peripheral Interface (SPI) to read, write, and erase flash memoryproperly, (2) be able to detect arbitrarily large or small amounts of bit-errors, (3) be able tocorrect all detected errors, and (4) perform memory scrubbing indefinitely and autonomously.A prototype implementation was constructed and tested, demonstrating successful detectionand correction of bit errors in multiple configurations. In the form of burst errors or singularbit flips, and in amounts of errors ranging from one to fifteen (per 256 Bytes), the applicationwas successful in preserving memory fidelity.

Radiation tolerant↗

Single-Event Effect (SEE) Survey of Advanced Reconfigurable Field Programmable Gate Arrays: NASA Electronic Parts and Packaging (NEPP) Program Office of Safety and Mission Assurance

The NEPP Reconfigurable Field-Programmable Gate Array (FPGA) task has been charged to evaluate reconfigurable FPGA technologies for use in space. Under this task, the Xilinx single-event-immune, reconfigurable FPGA (SIRF) XQR5VFX130 device was evaluated for SEE. Additionally, the Altera Stratix-IV and SiliconBlue iCE65 were screened for single-event latchup (SEL).

Xilinx Single-Event Effects (SEE) Test Consortium↗

Final Technical Report for CMSC 838L

This paper describes Rahul Vishnoi’s final project supporting in his Graduate School curriculum CMSC 838L, Advanced Topics in Programming Languages and Computer Architecture. This project was selected to intersect with his work as a Pathways Intern supporting Code 583, the Ground Software Systems Branch, at NASA’s Goddard Space Flight Center (GSFC). In this project, Field Programmable Gate Array (FPGA) hardware from Xilinx is used to replace and offload processor and memory-intensive computations from a microcontroller/Processing System (PS) to the FPGA Programmable Logic (PL). An interface between the PL and PS in the form of a C library allows for this bridging of capability.

Microcontroller, FPGA, Embedded Development, Xilin↗

Neuro-Spark: A Submicrosecond Spiking Neural Networks Architecture for In-Sensor Filtering

Neuro-Spark, which is a new neuromorphic architecture with a field-programmable gate array (FPGA) implementation for ultrafast spiking neural network (SNN) inference at the edge, facilitates smart-pixel in-sensor filtering for high-energy physics experiments at the Large Hadron Collider (LHC). Utilizing the evolutionary optimization for neuromorphic systems (EONS) training method, we generate compact SNN models with 91% signal efficiency, akin to convolutional neural networks but with half the parameters. However, deploying near the detector poses a challenge because the SNN must handle a sustained input data rate exceeding 1013 GB/s. To overcome this, we propose a novel hardware architecture that uses high-level synthesis to construct a tuned architecture for the EONS-trained SNN. In addition to the analysis and validation with an AMD Xilinx Artix-A7 FPGA, our solution consumes only ç24% of FPGA LUT and flipflops. We also introduce an innovative quantization method that reduces FPGA resource utilization by ç15% without compromising accuracy. Our FPGA implementation achieves computing latency of ç10 ns for smart-pixel application inference on an edge FPGA.

Miniskar, Narasinga Rao↗

End-to-End Workflow for Machine-Learning-Based Qubit Readout With QICK and hls4ml

In this article, we present an end-to-end workflow for superconducting qubit readout that embeds codesigned neural networks into the quantum instrumentation control kit (QICK). Capitalizing on the custom firmware and software of the QICK platform, which is built on Xilinx radiofrequency system-on-chip field-programmable gate arrays (FPGAs), we aim to leverage machine learning (ML) to address critical challenges in qubit readout accuracy and scalability. The workflow utilizes the hls4ml package and employs quantization-aware training to translate ML models into hardware-efficient FPGA implementations via user-friendly Python application programming interfaces. We experimentally demonstrate the design, optimization, and integration of an ML algorithm for single transmon qubit readout, achieving 96% single-shot fidelity with a latency of 32.25 ns and less than 16% FPGA lookup table resource utilization. Our results offer the community an accessible workflow to advance ML-driven readout and adaptive control in quantum information processing applications.

42 ENGINEERING↗

Compiler-Driven FPGA Virtualization with SYNERGY

FPGAs are increasingly common in modern applications, and cloud providers now support on-demand FPGA acceleration in datacenters. Applications in datacenters run on virtual infrastructure, where consolidation, multi-tenancy, and workload migration enable economies of scale that are fundamental to the provider's business. However, a general strategy for virtualizing FPGAs has yet to emerge. While manufacturers struggle with hardware-based approaches, we propose a compiler/runtime-based solution called Synergy. We show a compiler transformation for Verilog programs that produces code able to yield control to software atsub-clock-tickgranularity according to the semantics of the original program. Synergy uses this property to efficiently support core virtualization primitives: suspend and resume, program migration, and spatial/temporal multiplexing, on hardware which is availabletoday.We use Synergy to virtualize FPGA workloads across a cluster of Intel SoCs and Xilinx FPGAs on Amazon F1. The workloads require no modification, run within 3--4x of unvirtualized performance, and incur a modest increase in FPGA fabric usage.

Computer Science↗

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin [Fermilab] (ORCID:0000000157000288↗

Radio Frequency Field Programable Gate Array Implementation of Reflectometry Cable Monitoring

PNNL has developed an innovative cable testing system that leverages the Xilinx RF System-on-Chip (RFSoC) technology to create a more flexible and capable measurement tool than traditional approaches. The architecture is designed on the ZCU111 development board and takes advantage of the high-speed digital to analog converters (DAC) and analog to digital converters (ADC) to generate and capture digitally synthesized waveforms. The platform allows engineers to easily adjust power, frequency, duration and modulation on the fly rather than being locked into fixed hardware configurations

Tedeschi, Jonathan [Pacific Northwest National Lab↗

RFSoC based digital low level RF control firmware and software suite (mimo_llrf) v1.0

It features a framework of a firmware and software architecture in support of building a digital low-level RF control system for accelerators, where precised digital RF generation and measurement are needed across many RF channels. It primarily supports the Xilinx RFSoC chips (xczu48dr, xczu47dr, xczu29dr) and their evaluation boards (zcu208, zcu216), for a highly integrated solution enabling the need for synchronous low-level RF systems, including: multi-tile synchronization, external reference for sampling clocks, deterministic delay, aligned NCO phase for digital mixers, and built-in EPICS IOC.

Du, Qiang [Lawrence Berkeley National Laboratory (↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Design of digital acquisition for beam current monitor

As a part of the Proton Improvement Plan – II (PIP-II) at Fermilab, instrumentation systems are being modernized to take advantage of the higher speeds and ease of use offered by standardized embedded systems like MicroTCA. A rear-transition module (RTM) is being designed to interface with said embedded systems. In each of the four identical channels on the RTM, the differential signal from an alternating-current current transformer (ACCT) transimpedance amplifier will again be amplified by a differential operation-amplifier, then filtered by a low-pass topology. The conditioned signal is then digitized at a maximum of 10MS/s by an analog to digital converter (ADC) integrated circuit. After digitization, the ADC passes the data to an off the shelf AdvancedMC (AMC) Xilinx FPGA module using low voltage differential signals. This paper will describe the simulation of analog circuitry for signal conditioning, simulation of digital signal integrity based on physical design as well as verification of design characteristics critical to signal integrity. This work aims to create a methodology that can be applied to future RTMs requiring application of high-speed digital design principles.

White, R.Turner [Fermilab]↗

DIRECT RF SAMPLING BASED LLRF CONTROL SYSTEM FOR C-BAND LINEAR ACCELERATOR

Low Level RF (LLRF) control systems of linear accel- erators (LINACs) are typically implemented with hetero- dyne based architectures, which have complex analog RF mixers for up and down conversion. The Gen 3 Radio Fre- quency System-on-Chip (RFSoC) device from AMD Xilinx integrates data converters with maximum RF frequency of 6 GHz. This enables direct RF sampling of C-band LLRF signal typically operated at 5.712 GHz without any analogue mixers, which can significantly simplify the system architec- ture. The data converters sample RF signals in higher order Nyquist zones and then up or down convert digitally by the integrated data path in RFSoC. The closed-loop feedback control firmware implemented in FPGA integrated in RF- SoC can process the base-band signal from the ADC data path and calculate the updated phase and amplitude to be up- mixed by the DAC data path. We have developed a C-band LLRF control RFSoC platform with direct RF sampling, which targets Cool Copper Collider (𝐶3) and other C or S band LINAC research and development projects. In this paper, the architecture of the platform will be described. We have optimized the configuration of the data converter and characterized performance of them with RF pulses. The test results for some of the key performance parameters for the LLRF platform with our custom solid-state amplifier, such as phase and amplitude stability, will be discussed in this paper.

Liu, C↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗