Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “application specific integrated circuits”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The CMS Phase-2 Fast Beam Condition Monitor prototype test with beam

The Fast Beam Condition Monitor (FBCM) is a standalone luminometer for the High Luminosity LHC (HL-LHC) program of the CMS Experiment at CERN. The detector is under development and features a new, radiation-hard, front-end application-specific integrated circuit (ASIC) designed for beam monitoring applications. The achieved timing resolution of a few nanoseconds enables the measurement of both the luminosity and the beam-induced background. The ASIC, called FBCM23, features six channels with adjustable shaping times, enabling in-field fine-tuning. Each ASIC channel outputs a single binary asynchronous signal encoding time-of-arrival and time-over-threshold information. The FBCM is based on silicon-pad sensors, with two sensor designs presently being considered. This paper presents the results of tests of the FBCM detector prototype using both types of silicon sensors with hadron, muon, and electron beams. Irradiated FBCM23 ASICs and silicon-pad sensors were also tested to simulate the expected conditions near the end of the detector's lifetime in the HL-LHC radiation environment. Based on test results, direct bonding between the sensor and ASIC was chosen, and an optimal bias voltage and ASIC threshold for FBCM operation were proposed. The current design of the front-end test board was validated following the beam test and is now being used for the first front-end module, which is expected to be produced in summer 2025. These results represent a major step forward in validating the FBCM concept, first version of the firmware and establishing a reliable design path for the final detector.

Beam-line instrumentation (beam position and profi↗

The VMM3a ASIC

The VMM3a is a custom low-noise Application Specific Integrated Circuit. It is intended to be used in the front end readout electronics of both the Micromegas and sTGC detectors of the ATLAS New Small Wheels upgrade project at CERN. It is fabricated in the 130 nm GlobalFoundries 8RF-DM process. Here, the 64 channels with highly configurable parameters, although designed for a specific project, can meet the processing needs of signals from various detector types in several other applications. In this note, the VMM3a, which is the production version of the VMM family, will be presented along with the features incorporated and its performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Timepix4 analog front-end design: Lessons learnt on fundamental limits to noise and time resolution in highly segmented hybrid pixel detectors

This manuscript describes the optimization of the front-end readout electronics for high granularity hybrid pixel detectors. The theoretical study aims at minimizing the noise and jitter. The model presented here is validated with both circuit post layout simulations and measurements on the Timepix4 Application Specific Integrated Circuit (ASIC). The analog front-end circuit and the procedure to optimize the dimensions of the main transistors are described with detail. The Timepix4 is the most recent ASIC designed in the framework of the Medipix4 Collaboration. It was manufactured in 65nm CMOS process, and consists of a four side buttable matrix of 448 × 512 pixels with 55µm pitch. The analog front-end has a gain of ~36mV/ke - when configured in High Gain Mode, and ~20mV/ke - when configured in Low Gain Mode. The Equivalent Noise Charge (ENC) is ~68e - rms and ~80e - rms in High Gain Mode and in Low Gain Mode respectively. In event driven mode the incoming hits can be time stamped within a ~ 200ps time bin and the chip can deal with a maximum flux of ~ 3.6MHzmm –2 s –1 . In photon counting mode, the chip can deal with up to ~ 5GHzmm –2 s –1 . The routine designed to optimize the Timepix4 front-end is then used to analyze the performance limits in terms of jitter and noise for Charge Sensitive Amplifiers in pixel detectors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Modular 512-Channel Neural Signal Acquisition ASIC for High-Density 4096 Channel Electrophysiology

The complexity of information processing in the brain requires the development of technologies that can provide spatial and temporal resolution by means of dense electrode arrays paired with high-channel-count signal acquisition electronics. In this work, we present an ultra-low noise modular 512-channel neural recording circuit that is scalable to up to 4096 simultaneously recording channels. The neural readout application-specific integrated circuit (ASIC) uses a dense 8.2 mm × 6.8 mm 2D layout to enable high-channel count, creating an ultra-light 350 mg flexible module. The module can be deployed on headstages for small animals like rodents and songbirds, and it can be integrated with a variety of electrode arrays. The chip was fabricated in a TSMC 0.18 µm 1.8 V CMOS technology and dissipates a total of 125 mW. Each DC-coupled channel features a gain and bandwidth programmable analog front-end along with 14 b analog-to-digital conversion at speeds up to 30 kS/s. Additionally, each front-end includes programmable electrode plating and electrode impedance measurement capability. We present both standalone and in vivo measurements results, demonstrating the readout of spikes and field potentials that are modulated by a sensory input.

47 OTHER INSTRUMENTATION↗

ASIC Scoping Study Final Report

This report documents the results and findings of a one-year scoping study investigating multichannel readout application specific integrated circuits (ASICs) for interfacing to, and processing data from, silicon photomultiplier (SiPM) arrays. We document ASIC desired and required specifications for four applications supporting national security mission areas: neutron radiography, associated particle imaging, and two versions of kinematic neutron imaging cameras. While each application has a few unique requirements that stress capability, there is generally good agreement among most. Two recently developed ASIC devices were evaluated in a system-like configuration by interfacing these to scintillator crystals exposed to gamma and neutron sources. The 64-channel ORNL device delivered functional capability while meeting most mission requirements for neutron radiography. The Nalu Scientific device, a 32-channel full waveform digitizer, did not demonstrate reliable neutron / gamma separation but it is unclear if this was an ASIC issue or problems with test setup or firmware. A literature survey of other commercial and academic ASICs was undertaken to with the conclusion that existing devices do not meet all requirements.

42 ENGINEERING↗

Investigating resource-efficient neutron/gamma classification ML models targeting eFPGAs

There has been considerable interest and resulting progress in implementing machine learning (ML) models in hardware over the last several years from the particle and nuclear physics communities. A big driver has been the release of the Python package, hls4ml, which has enabled porting models specified and trained using Python ML libraries to register transfer level (RTL) code. So far, the primary end targets have been commercial field-programmable gate arrays (FPGAs) or synthesized custom blocks on application specific integrated circuits (ASICs). However, recent developments in open-source embedded FPGA (eFPGA) frameworks now provide an alternate, more flexible pathway for implementing ML models in hardware. These customized eFPGA fabrics can be integrated as part of an overall chip design. In general, the decision between a fully custom, eFPGA, or commercial FPGA ML implementation will depend on the details of the end-use application. In this work, we explored the parameter space for eFPGA implementations of fully-connected neural network (fcNN) and boosted decision tree (BDT) models using the task of neutron/gamma classification with a specific focus on resource efficiency. We used data collected using an AmBe sealed source incident on Stilbene, which was optically coupled to an OnSemi J-series silicon photomultiplier (SiPM) to generate training and test data for this study. We investigated relevant input features and the effects of bit-resolution and sampling rate as well as trade-offs in hyperparameters for both ML architectures while tracking total resource usage. The performance metric used to track model performance was the calculated neutron efficiency at a gamma leakage of 10 -3 . The results of the study will be used to aid the specification of an eFPGA fabric, which will be integrated as part of a test chip.

47 OTHER INSTRUMENTATION↗

Ultrafast radiographic imaging and tracking: An overview of instruments, methods, data, and applications

Ultrafast radiographic imaging and tracking (U-RadIT) use state-of-the-art ionizing particle and light sources to experimentally study sub-nanosecond transients or dynamic processes in physics, chemistry, biology, geology, materials science and other fields. These processes are fundamental to modern technologies and applications, such as nuclear fusion energy, advanced manufacturing, communication, and green transportation, which often involve one mole or more atoms and elementary particles, and thus are challenging to compute by using the first principles of quantum physics or other forward models. One of the central problems in U-RadIT is to optimize information yield through, e.g. high-luminosity X-ray and particle sources, efficient imaging and tracking detectors, novel methods to collect data, and large-bandwidth online and offline data processing, regulated by the underlying physics, statistics, and computing power. We review and highlight recent progress in: (a.) Detectors such as high-speed complementary metal-oxide semiconductor (CMOS) cameras, hybrid pixelated array detectors integrated with Timepix4 and other application-specific integrated circuits (ASICs), and digital photon detectors; (b.) U-RadIT modalities such as dynamic phase contrast imaging, dynamic diffractive imaging, and four-dimensional (4D) particle tracking; (c.) U-RadIT data and algorithms such as neural networks and machine learning, and (d.) Applications in ultrafast dynamic material science using XFELs, synchrotrons and laser-driven sources. Hardware-centric approaches to U-RadIT optimization are constrained by detector material properties, low signal-to-noise ratio, high cost and long development cycles of critical hardware components such as ASICs. Interpretation of experimental data, including comparisons with forward models, is frequently hindered by sparse measurements, model and measurement uncertainties, and noise. Alternatively, U-RadIT make increasing use of data science and machine learning algorithms, including experimental implementations of compressed sensing. Machine learning and artificial intelligence approaches, refined by physics and materials information, may also contribute significantly to data interpretation, uncertainty quantification and U-RadIT optimization.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

DynPaC: Coarse-Grained, Dynamic, and Partially Reconfigurable Array for Streaming Applications

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across kernels. However, streaming applications often are data-dependent, leading to variable kernel execution times depending on the input data and impacting the throughput of the entire pipeline if resources are statically allocated. Therefore, in this paper, we discuss the design of DynPaC — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We discuss the required software and hardware components to manage partial dynamic reconfiguration. We demonstrate that by supporting partial dynamic reconfiguration, we can obtain an average speedup of 1.44X for a representative set of applications w.r.t. static partitioning, with a limited area overhead (6.4% of the entire chip).

Tan, Cheng↗

DRIPS: Dynamic Rebalancing of Pipelined Streaming Applications on CGRAs

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit data streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across application kernels. However, emerging streaming applications at the edge (scientific instruments, sensor networks, network processing) perform much more than digital signal processing and often are data and input dependent. This leads to extremely variable kernel execution times, severely impacting the throughput of the entire pipeline if resources are only statically allocated. Therefore, in this paper, we propose DRIPS — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We present a unified compiler framework to facilitate the mapping of a given streaming application onto the DRIPS CGRA architecture. The experimental results show that DRIPS achieves an average throughput improvement of 1.46$\times$ across a set of representative applications over a statically partitioned solution. The additional area overhead to enable dynamic rebalancing consumes 16.34% of the entire area for a 5x5 CGRA prototype.

Tan, Cheng↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

FOS: Computer and information sciences↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

Schulte, Jan-Frederik [Purdue U.] (ORCID:000000034↗

Bridging Python to Silicon: The SODA Toolchain

Systems performing scientific computing, data analysis, and machine learning tasks have a growing demand for application-specific accelerators that can provide high computational performance while meeting strict size and power requirements. However, the algorithms and applications that need to be accelerated are evolving at a rate that is incompatible with manual design processes based on hardware description languages. Agile hardware design tools based on compiler techniques can help by quickly producing an application-specific integrated circuit (ASIC) accelerator starting from a high-level algorithmic description. Here, we present the software-defined accelerator (SODA) synthesizer, a modular and open-source hardware compiler that provides automated end-to-end synthesis from high-level software frameworks to ASIC implementation, relying on multilevel representations to progressively lower and optimize the input code. Our approach does not require the application developer to write any register-transfer level code, and it is able to reach up to 364 giga floating point operations per second (GFLOPS)/W efficiency (32-bit precision) on typical convolutional neural network operators.

97 MATHEMATICS AND COMPUTING↗

Towards On-Chip Learning for Low Latency Reasoning with End-to-End Synthesis

The Software Defined Architectures (SODA) Synthesizer is an open-source compiler-based tool able to automatically generate domain-specialized systems targeting Application-Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs) starting from high-level programming. SODA is composed of a frontend, SODA-OPT, which leverages the multilevel intermediate representation (MLIR) framework to interface with productive programming tools (e.g., machine learning frame-works), identify kernels suitable for acceleration, and perform high-level optimizations, and of a state-of-the-art high-level synthesis backend, Bambu from the PandA framework, to generate custom accelerators. One specific application of the SODA Synthesizer is the generation of accelerators to enable ultra-low latency inference and control on autonomous systems for scientific discovery (e.g., electron microscopes, sensors in particle accelerators, etc.). This paper provides an overview of the flow in the context of the generation of accelerators for edge processing to be integrated in transmission electron microscopy (TEM) devices, focusing on use cases from precision material synthesis. We show the tool in action with an example of design space exploration for inference on reconfigurable devices with a conventional deep neural network model (LeNet). Finally, we discuss the research directions and opportunities enabled by SODA in the area of autonomous control for scientific experimental workflows.

Castellana, Vito G.↗

Towards Automated Generation of Chiplet-Based Systems

The Software Defined Architectures (SODA) Synthesizer is an open-source compiler-based tool able to automatically generate domain-specialized systems targeting Application- Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs) starting from high-level programming. SODA is composed of a high-level frontend, SODA-OPT, which leverages the multilevel intermediate representation (MLIR) framework to interface with productive programming tools (e.g., machine learning frameworks), identify kernels suitable for acceleration, and perform high-level optimizations, and of a state-of-the-art high-level synthesis backend, Bambu from the PandA framework, to generate custom accelerators. One specific application of the SODA Synthesizer is the generation of accelerators to enable ultra-low latency inference and control on autonomous systems for scientific discovery (e.g., electron microscopes, sensors in particle accelerators, etc.). This talk will discuss ongoing work on the SODA synthesizer to enable no-human-in-the-loop generation and design space exploration of the chiplets for highly specialized artificial intelligence accelerators. Connecting these highly specialized chiplets to general-purpose cores or programmable accelerators will allow to quickly deploy autonomous systems for scientific discovery.

Limaye, Ankur M.↗

Fail-Safe Logic Design Strategies Within Modern FPGA Architectures

Fail-safe computing refers to computing systems that revert to a non-operational safe state when a fault occurs. In this paper, we investigate a circuit level technique as mitigation for single event upsets (SEUs) and fault injection attacks on field programmable gate arrays (FPGAs), and analyze the effectiveness of the technique as a fail-safe monitor for an encryption algorithm. The propagation of fault effects through FPGA primitives including lookup tables (LUTs) and programmable interconnect points (PIPs) is assessed within an FPGA architecture created using an open source tool, and validated using fault injection experiments on an FPGA. The analysis reveals additional vulnerabilities exist within reconfigurable architectures over those in equivalent fail-safe application specific integrated circuit (ASIC), thus requiring a more elaborate network of redundant circuits and checking logic. The configuration memory bits (CMBs), which configure routing and designate logic functions within the LUTs of the FPGA, add complexity to fail-safe design strategies by introducing additional fault conditions and fault propagation paths. A resource-efficient fail-safe circuit design technique called DEsign for Fail-safe in reCONfigurable systems (DEFCON) is proposed. The benefits and limitations associated with DEFCON are described in the context of fault injection experiments carried out as simulations and in FPGA hardware.

Bhakta, Priya A. [Univ. of New Mexico, Albuquerque↗

Code Generators for Floating-Point Unit Design in Integrated Circuits (OpenFloat) v1.0

This IP provides a comprehensive set of code generators for various floating-point units (FPUs) essential for integrated circuit design and integration, targeting a broad spectrum of applications, including machine learning and scientific computing. The suite includes FP adders, multipliers, subtractors, dividers, reciprocals, exponentials, square roots, trigonometric functions (sine, cosine, arctangent), and more. It supports customizable hardware design parameters, such as precision (16, 32, 64, and 128 bits) and pipeline depths, offering users enhanced flexibility and productivity. The generated code is in an industry-standard hardware description language, ensuring compatibility with standard design flows, including simulation, verification, synthesis, and implementation on both field-programmable gate arrays (FPGAs) and application-specific integrated circuits (ASICs).

Shalf, JohnM. [Lawrence Berkeley National Laborato↗

A MEMS Gyroscope for Reliable Long Duration Measurement While Drilling at 300°C

The 300°C Microelectromechanical system (MEMS) gyroscope project aims to contribute to DOE’s goal of increased geothermal drilling efficiency by 2025 through the development of a 300°C MEMS gyroscope for Measurement While Drilling (MWD). At the conclusion of the 2-year project, the team will develop a 300°C capable MEMS gyroscope containing GE’s patented Multi-Ring Gyroscope Transducer (MRGT) design, custom Silicon-On-Insulator (SOI) based frontend and feedback control electronics, and with demonstrated functionality and lifetime beyond 1000 hours. The project is divided into two budget periods with Go/No-Go decision at the end of the first budget period. The goal for the first budget period is to establish the feasibility of the MRGT and electronics design for meeting the 300°C performance requirements. The goal for the second budget period is to integrate the MRGT with the SOI-based application specific integrated circuit (ASIC) and demonstrate capability to operate at 300°C for 1000 hours. In Budget Period 1 we met the phase 1 goal. We successfully validated the combined MRG, electronics and packaging capability entitlement to achieving 0.5 degrees azimuth uncertainty while enabling operation at significantly higher temperatures than the state-of-the-art. In Budget Period 2, we successfully completed the integration of the MRGT and ASIC with associated high temperature, high reliability packaging to demonstrate the performance and functionality of the integrated gyroscope across the temperature range from room temperature to at 300°C. Furthermore, the team demonstrated operating life of >1,000 hours at 300°C, thus providing a validation of application-relevant lifetime capability.

15 GEOTHERMAL ENERGY↗

Embedded FPGA developments in 130 nm and 28 nm CMOS for machine learning in particle detector readout

Embedded field programmable gate array (eFPGA) technology allows the implementation of reconfigurable logic within the design of an application-specific integrated circuit (ASIC). This approach offers the low power and efficiency of an ASIC along with the ease of FPGA configuration, particularly beneficial for the use case of machine learning in the data pipeline of next-generation collider experiments. An open-source framework called "FABulous" was used to design eFPGAs using 130 nm and 28 nm CMOS technology nodes, which were subsequently fabricated and verified through testing. The capability of an eFPGA to act as a front-end readout chip was assessed using simulation of high energy particles passing through a silicon pixel sensor. A machine learning-based classifier, designed for reduction of sensor data at the source, was synthesized and configured onto the eFPGA. A successful proof-of-concept was demonstrated through reproduction of the expected algorithm result on the eFPGA with perfect accuracy. Finally, further development of the eFPGA technology and its application to collider detector readout is discussed.

47 OTHER INSTRUMENTATION↗