Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Field programmable gate arrays”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

IEEE 1547-2018 Based Interoperable PV Inverter with Advanced Grid-Support Functions

Grid integration of photovoltaic (PV) inverters has been increasing in the past decade. As a result of the uncertainties introduced with high penetrations of PV, better monitoring and control of the PV inverters becomes crucial for improving overall system stability. This paper focuses on the communications capability of the inverter controller and on enabling interoperability. Multiple standards are available to enable interoperability in PV inverters. In this paper, an interoperable controller, enabled by Distributed Network Protocol 3 (DNP3) communications protocols, is developed for a grid-connected, three-phase PV inverter. The DNP3 server for the PV inverter is programmed on the real-time layer of the field-programmable gate array (FPGA)-based inverter controller. Set points for advanced inverter control functions, such as volt/VAr curves, ride-through curves, are sent from a DNP3 client, a simulated distribution management system application, to the PV inverter through DNP3. This communications capability of the inverter controller is validated using a controller-hardware-in-the-loop experimental setup. The code developed to achieve the interoperability is available in the public domain through open-source software licensing. This interoperability will enable smoother grid integration of smart PV inverters with advanced grid-support functions as well as allow better monitoring and control of PV inverters for grid stability.

14 SOLAR ENERGY↗

Real-Time Ethernet Interface for NSTX-U’s Thomson Scattering Diagnostic (2023)

Here, the multipoint Thomson scattering (MPTS) diagnostic system at the National Spherical Torus Experiment Upgrade (NSTX-U) facility is undergoing an upgrade to operate in real-time and interface with the plasma control system (PCS) for NSTX-U. Previous prototyping efforts have shown that spectral analysis and rapid calculations of electron temperature and density are possible on a real-time Linux machine when using up to a 100-Hz laser pulse repetition rate. A remaining challenge was transferring the real-time data to NSTX-U’s PCS, which utilizes the front panel data port (FPDP) protocol. The original proposed method was to convert the real-time data into analog values, but a new solution was developed to keep the output format digital by using an Ethernet controller with a field-programmable gate array (FPGA). This article focuses on a new input module that has been developed to accept incoming user datagram protocol (UDP) packets sent over Ethernet, convert into FPDP format, and integrate into the existing data stream under NSTX-U’s real-time framework.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Toward Evaluating High-Level Synthesis Portability and Performance between Intel and Xilinx FPGAs

Offloading computation from a CPU to a hardware accelerator is becoming a more common solution for improving performance because traditional gains enabled by Moore’s law and Dennard scaling have slowed. GPUs are often used as hardware accelerators, but field-programmable gate arrays (FPGAs) are gaining traction. FPGAs are beneficial because they allow hardware specific to a particular application to be created. However, they are notoriously difficult to program. To this end, two of the main FPGA manufacturers, Intel and Xilinx, have created tools and frameworks that enable the use of higher level languages to design FPGA hardware. Although Xilinx kernels can be designed by using C/C++, both Intel and Xilinx support the use of OpenCL C to architect FPGA hardware. However, not much is known about the portability and performance between these two device families other than the fact that it is theoretically possible to synthesize a kernel meant for Intel to Xilinx and vice versa.In this work, we evaluate the portability and performance of Intel and Xilinx kernels. We use OpenCL C implementations of a subset of the Rodinia benchmarking suite that were designed for an Intel FPGA and make the necessary modifications to create synthesizable OpenCL C kernels for a Xilinx FPGA. We find that the difficulty of porting certain kernel optimizations varies, depending on the construct. Once the minimum amount of modifications is made to create synthesizable hardware for the Xilinx platform, more nontrivial work is needed to improve performance. However, we find that constructs that are known to be performant for an FPGA should improve performance regardless of the platform; the difficulty comes in deciding how to invoke certain kernel optimizations while also abiding by the constraints enforced by a given platform’s hardware compiler.

Cabrera, Anthony↗

A Length Adaptive Algorithm-Hardware Co-design of Transformer on FPGA Through Sparse Attention and Dynamic Pipelining

Transformers are considered one of the most important deep learning models since 2018, in part because it establishes state-of-the-art (SOTA) records and could potentially replace existing Deep Neural Networks (DNNs). Despite the remarkable triumphs, the prolonged turnaround time of Transformer models is a widely recognized roadblock. The variety of sequence lengths imposes additional computing overhead where inputs need to be zero-padded to the maximum sentence length in the batch to accommodate the parallel computing platforms. This paper targets the field-programmable gate array (FPGA) and proposes a coherent sequence length adaptive algorithm–hardware co-design for Transformer acceleration. Particularly, we develop a hardware-friendly sparse attention operator and a length-aware hardware resource scheduling algorithm. The proposed sparse attention operator brings the complexity of attention-based models down to linear complexity and alleviates the off-chip memory traffic. The proposed length-aware resource hardware scheduling algorithm dynamically allocates the hardware resources to fill up the pipeline slots and eliminates bubbles for NLP tasks. Experiments show that our design has very small accuracy loss and has 80.2 × and 2.6 × speedup compared to CPU and GPU implementation, and 4 × higher energy efficiency than state-of-the-art GPU accelerator optimized via CUBLAS GEMM.

Peng, Hongwu↗

SODA Synthesizer: an Open-source, Multi-level, Modular, Extensible Compiler from High-level Frameworks to Silicon

The SODA Synthesizer is an open-source modular, end-to-end hardware compiler framework. The SODA frontend, developed in MLIR, performs system-level design, code partitioning, and high-level optimizations to prepare the specifications for the hardware synthesis. The backend is based on a state-of-the-art high-level synthesis tool, and generates the final hardware design. The backend can interface with logic synthesis tools for field programmable gate arrays or with commercial and open-source logic synthesis tools for application-specific integrated circuits. We discuss the opportunities and challenges in integrating with commercial and open-source tools both at the frontend and backend, and the unique opportunities that an open-source hardware design ecosystem provides.

Bohm Agostini, Nicolas↗

Automatic Qubit Characterization and Gate Optimization with QubiC

As the size and complexity of a quantum computer increases, quantum bit (qubit) characterization and gate optimization become complex and time-consuming tasks. Current calibration techniques require complicated and verbose measurements to tune up qubits and gates, which cannot easily expand to the large-scale quantum systems. We develop a concise and automatic calibration protocol to characterize qubits and optimize gates using QubiC, which is an open source FPGA (field-programmable gate array) based control and measurement system for superconducting quantum information processors. We propose multi-dimensional loss-based optimization of single-qubit gates and full XY-plane measurement method for the two-qubit CNOT gate calibration. We demonstrate the QubiC automatic calibration protocols are capable of delivering high-fidelity gates on the state-of-the-art transmon-type processor operating at the Advanced Quantum Testbed at Lawrence Berkeley National Laboratory. Finally, the single-qubit and two-qubit Clifford gate infidelities measured by randomized benchmarking are of 4.9(1.1) × 10 -4 and 1.4(3) × 10 -2 , respectively.

97 MATHEMATICS AND COMPUTING↗

Towards On-Chip Learning for Low Latency Reasoning with End-to-End Synthesis

The Software Defined Architectures (SODA) Synthesizer is an open-source compiler-based tool able to automatically generate domain-specialized systems targeting Application-Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs) starting from high-level programming. SODA is composed of a frontend, SODA-OPT, which leverages the multilevel intermediate representation (MLIR) framework to interface with productive programming tools (e.g., machine learning frame-works), identify kernels suitable for acceleration, and perform high-level optimizations, and of a state-of-the-art high-level synthesis backend, Bambu from the PandA framework, to generate custom accelerators. One specific application of the SODA Synthesizer is the generation of accelerators to enable ultra-low latency inference and control on autonomous systems for scientific discovery (e.g., electron microscopes, sensors in particle accelerators, etc.). This paper provides an overview of the flow in the context of the generation of accelerators for edge processing to be integrated in transmission electron microscopy (TEM) devices, focusing on use cases from precision material synthesis. We show the tool in action with an example of design space exploration for inference on reconfigurable devices with a conventional deep neural network model (LeNet). Finally, we discuss the research directions and opportunities enabled by SODA in the area of autonomous control for scientific experimental workflows.

Castellana, Vito G.↗

FiberFlex: Real-time FPGA-based Intelligent and Distributed Fiber Sensor System for Pedestrian Recognition

In recent years, security monitoring of public places and critical infrastructure has heavily relied on the widespread use of cameras, raising concerns about personal privacy violations. To balance the need for effective security monitoring with the protection of personal privacy, we explore the potential of optical fiber sensors for this application. This article proposes FiberFlex, an intelligent and distributed fiber sensor system. Ultizing Field Programmable Gate Arrays (FPGA) high-level synthesis (HLS) acceleration, FiberFlex offers real-time pedestrian detection by co-designing the entire pipeline of optical signal acquisition, processing, and recognition networks based on the principles of optical fiber sensing. As a promising alternative to traditional camera-based monitoring systems, FiberFlex achieves pedestrian detection by analyzing the vibration patterns caused by pedestrian footsteps, enabling security monitoring while preserving individual privacy. FiberFlex comprises three modules: First , fiber-optic sensing system: A fiber-optic distributed acoustic sensing (DAS) system is built and used to measure the ground vibration waves generated by people walking. Second , algorithms: We first collect the training data by measuring the ground vibration waves, label the data, and use the data to train the neural network models to perform pedestrian recognition. Third , hardware accelerators: We use HLS tools to design hardware modules on FPGA for data collection and pre-processing and integrate them with the downstream neural network accelerators to perform in-line real-time pedestrian detection. The final detection results are sent back from FPGA to the host CPU. We implement our system FiberFlex with the in-house built DAS system and AMD/Xilinx Kintex7 FPGA KC705 board and verify the whole system using the real-world collected data. We conduct recognition tests on five test subjects of varying ages, heights, and weights in a fixed sensing area. Each subject experienced 20 real-time recognition tests using their daily walking habits, and the subjects were given adequate rest between tests. After 100 tests on five test subjects, the overall real-time recognition accuracy exceeded \(88.0\%\) . The whole system uses 55 W of power, 33 W in the optical DAS system and 22 W in the FPGA. Relying on its end-to-end interdisciplinary design, FiberFlex seamlessly combines fiber-optic sensors with FPGA accelerators to enable low-power real-time security monitoring without compromising privacy, making it a valuable addition to the existing security monitoring network. According to FiberFlex, more valuable research can be conducted in the future, such as fall monitoring for the elderly, migration of identification networks between different application scenarios, and improvement of anti-interference performance in more complex environments. In future perception networks, where the “eyes” are not feasible, let’s use fiber optic touch instead.

Distributed↗

Qubit Control System (QubiC) v1.0

We design a modular FPGA (field-programmable gate array) based system called QubiC to control and measure a superconducting quantum processing unit. The system includes room temperature electronics hardware, FPGA gateware, and engineering software.

Huang, Gang↗

LBNL BPM Firmware (CAL-BPM) v1.0

CAL-BPM is the firmware code base of the Advanced Light Source (ALS) Beam Position Monitor (BPM) Field Programmable Gate Array (FPGA) based electronics.

Norum, William↗

Quantum Instrumentation Control Kit Defect Arbitrary Waveform Generator (QICKDAWG) v.0

SAND2024-08598O The Quantum Instrumentation Control Kit Defect Arbitrary Waveform Generator (QICKDAWG) characterizes nitrogen-vacancy centers in diamond and other defects. It does this by using a radio frequency system-on-a-chip (RFSoC) field programmable gate array (FPGA). QICKDAWG synthesizes microwave pulses from the RFSoC to change the spin state of the defects. The software also allows for laser control using the RFSoC, which optically pumps defects. Ultimately, QICKDAWG supports the implementation of RFSoC FPGAs in defect characterization. This replaces the slow, expensive, traditional hardware, thus lowering the cost and time for defect characterization. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Sentinel

Network intrusion detection systems (NIDS) are commonplace in network security but they frequently employ algorithms that are computational demanding requiring hardware and software with significant power requirements. Two examples of such resource-intensive algorithms used for network security are regular expression matching and broader signature pattern matching which are commonly used in deep packet inspection (DPI). Network security algorithms that have large power requirements may be a challenge for low-power internet-of-things (IoT) environments, which generally lack the power resources to implement complex security measures like computationally expensive DPI at the edge. Furthermore, IoT environments incorporating 5G standalone networks have network latency constraints beyond just power that make DPI at the edge even more difficult. Programmable logic is ideally suited for machine learning inference for DPI because of its deep instruction level parallelism and single-cycle memory access. Machine learning approaches for DPI have been explored before using the programmable logic of field programmable gate arrays (FPGA) as a potential solution for NIDS approaches that would be power-suitable for IoT. However, those previous programmable logic NIDS approaches utilize either a supervised or unsupervised learning model. Sentinel utilizes the ensemble of these two machine learning approaches known as a semi-supervised approach which has shown promise in NIDS implementations. Sentinel provides a programmable logic implementation of a semi-supervised approach for DPI which operates at much lower power and latency than a GPU implementation with negligible loss of accuracy due to quantization through a logistic regressor.

Anderson, MatthewW [Idaho National Laboratory (INL↗

Code Generators for Floating-Point Unit Design in Integrated Circuits (OpenFloat) v1.0

This IP provides a comprehensive set of code generators for various floating-point units (FPUs) essential for integrated circuit design and integration, targeting a broad spectrum of applications, including machine learning and scientific computing. The suite includes FP adders, multipliers, subtractors, dividers, reciprocals, exponentials, square roots, trigonometric functions (sine, cosine, arctangent), and more. It supports customizable hardware design parameters, such as precision (16, 32, 64, and 128 bits) and pipeline depths, offering users enhanced flexibility and productivity. The generated code is in an industry-standard hardware description language, ensuring compatibility with standard design flows, including simulation, verification, synthesis, and implementation on both field-programmable gate arrays (FPGAs) and application-specific integrated circuits (ASICs).

Shalf, JohnM. [Lawrence Berkeley National Laborato↗

Abisko: Deep codesign of an architecture for spiking neural networks using novel neuromorphic materials

The Abisko project aims to develop an energy-efficient spiking neural network (SNN) computing architecture and software system capable of autonomous learning and operation. The SNN architecture explores novel neuromorphic devices that are based on resistive-switching materials, such as memristors and electrochemical RAM. Equally important, Abisko uses a deep codesign approach to pursue this goal by engaging experts from across the entire range of disciplines: materials, devices and circuits, architectures and integration, software, and algorithms. Here, the key objectives of our Abisko project are threefold. First, we are designing an energy-optimized high-performance neuromorphic accelerator based on SNNs. This architecture is being designed as a chiplet that can be deployed in contemporary computer architectures and we are investigating novel neuromorphic materials to improve its design. Second, we are concurrently developing a productive software stack for the neuromorphic accelerator that will also be portable to other architectures, such as field-programmable gate arrays and GPUs. Third, we are creating a new deep codesign methodology and framework for developing clear interfaces, requirements, and metrics between each level of abstraction to enable the system design to be explored and implemented interchangeably with execution, measurement, a model, or simulation. As a motivating application for this codesign effort, we target the use of SNNs for an analog event detector for a high-energy physics sensor.

97 MATHEMATICS AND COMPUTING↗

Optical networking within the Lightwave Energy-Efficient Datacenter project [Invited]

The Lightwave Energy-Efficient Datacenter (LEED) project within the ARPA-e ENLITENED program is developing novel energy-efficient multichannel lightwave networks. These networks are enabled by a new optical “rotor” switch that can reconfigure the network topology in less than 20 µs and a field-programmable-gate-array-based network interface controller called Corundum that can provide precise network-wide synchronization of packets admitted into the lightwave network. Here we review the optical networking research within LEED and discuss future directions.

Mellette, William M.↗

Electron Density Measurements Using USPR (Final Scientific/Technical Report)

UC Davis has fabricated an ultrashort pulse reflectometer (USPR) diagnostic instrument for electron density profile measurements on compact, short duration, magnetically-confined fusion-energy concept devices such as spheromaks and FRCs. The USPR system transmits extremely short duration (~few nsec) chirped waveforms that together span 29 to 75 GHz. These chirped waveforms illuminate and reflect from the target plasma, with each frequency component reflecting from a different density layer (higher frequencies probe deeper into the plasma before reflecting). The reflected waveforms are split into roughly 42 different frequencies; time-of-flight (TOF) measurements made at each frequency with high resolution (~25 psec measurement resolution which corresponds to ~5 mm). These TOF data may then be inverted via software to generate electron density profiles with high time resolution (~10 μsec). At the heart of the system is a field programmable gate array (FPGA) based controller which collects and processes all of the USPR data in addition to generating all of the control signals required for maximum flexibility. The FPGA controller has the software flexibility to be easily reconfigured for different plasma devices, and the entire system sufficiently compact to be easily and quickly transported between devices. A high speed impulse generator was transformed into a set of three ultrashort pulse transmitter chirps using a combination of dispersive waveguide, frequency doublers and high-pass filters. A mm-wave controller was fabricated to sequentially switch between the three chirps, directing the chirps one-by-one to three different mm-wave assemblies spanning 29-75 GHz. Each mm-wave assembly consists of a high power active multiplier chain which converts the transmitter chirp to higher frequencies, and a broadband mixer which downconverts the reflected waveform to the 2-18 GHz range of the UPSR receiver. The 16-channel receiver (shared by all 3 mm-wave assemblies) was fabricated employing custom TOF modules capable of operating at a high 1 MHz sampling rate. Laboratory testing of the full system revealed the presence of unwanted harmonics from the multiplication process, with interference observed in the downconverted reflections at selected frequency channels that could not be completely filtered out. Additional interference effects arising from internal reflections within the mm-wave assemblies were minimized using a high-speed switch which served to “gate out” much of these reflections. The USPR diagnostic was transported and installed onto the HIT-SIU plasma device, becoming operational on 11/08/2022. Although designed to span 3 distinct mm-wave bands, the HIT-SIU plasmas at this time were sufficiently low density such that only the lowest of the three bands was likely to have strong plasma reflections. The system was then set to operate on only the lowest band (assembly #1), with data collected every 1 μsec rather than 3 μsec which would have been the case when cycling through all three bands. Connected to HIT-SIU, time-varying plasma reflections were observed on 9 of 16 possible frequency channels. Close examination of the data collected revealed issues previously unobserved in laboratory testing, associated with (a) reflections from the small aperture horns required for operation within the HIT-SIU device, and (b) a dependence of the recorded TOF with the threshold voltage of a given channel. Plans were made to address each of these issues before undertaking any future campaigns.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Lightweight Embedded Controller in Advanced FPGA SoC for Radar Signal Processing [Poster]

The objective of the project is to demonstrate that critical control functions can be implemented using little resources in modern microelectronics. A finite state machine (Figure 1) is implemented onto a field programmable gate array (FPGA). The functionality of the system is demonstrated by sending binary instructions to the controller. The controller transmits patterns through an LED, controls an electromechanical device, and uses pulse-width modulation (PWM) for radar functions.

42 ENGINEERING↗