Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “chip scale”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Accelerated Adaptive MGS Phase Retrieval

The Modified Gerchberg-Saxton (MGS) algorithm is an image-based wavefront-sensing method that can turn any science instrument focal plane into a wavefront sensor. MGS characterizes optical systems by estimating the wavefront errors in the exit pupil using only intensity images of a star or other point source of light. This innovative implementation of MGS significantly accelerates the MGS phase retrieval algorithm by using stream-processing hardware on conventional graphics cards. Stream processing is a relatively new, yet powerful, paradigm to allow parallel processing of certain applications that apply single instructions to multiple data (SIMD). These stream processors are designed specifically to support large-scale parallel computing on a single graphics chip. Computationally intensive algorithms, such as the Fast Fourier Transform (FFT), are particularly well suited for this computing environment. This high-speed version of MGS exploits commercially available hardware to accomplish the same objective in a fraction of the original time. The exploit involves performing matrix calculations in nVidia graphic cards. The graphical processor unit (GPU) is hardware that is specialized for computationally intensive, highly parallel computation. From the software perspective, a parallel programming model is used, called CUDA, to transparently scale multicore parallelism in hardware. This technology gives computationally intensive applications access to the processing power of the nVidia GPUs through a C/C++ programming interface. The AAMGS (Accelerated Adaptive MGS) software takes advantage of these advanced technologies, to accelerate the optical phase error characterization. With a single PC that contains four nVidia GTX-280 graphic cards, the new implementation can process four images simultaneously to produce a JWST (James Webb Space Telescope) wavefront measurement 60 times faster than the previous code.

Lam, Raymond K.↗

Tailored Micromagnet Sorting Gate for Simultaneous Multiple Cell Screening in Portable Magnetophoretic Cell‐On‐Chip Platforms

Abstract Conventional magnetophoresis techniques for manipulating biocarriers and cells predominantly rely on large‐scale electromagnetic systems, which is a major obstacle to the development of portable and miniaturized cell‐on‐chip platforms. Herein, a novel magnetic engineering approach by tailoring a nanoscale notch on a disk micromagnet using two‐step optical and thermal lithography is developed. Versatile manipulations are demonstrated, such as separation and trapping, of carriers and cells by mediating changes in the magnetic domain structure and discontinuous movement of magnetic energy wells around the circumferential edge of the micromagnet caused by a locally fabricated nano‐notch in a low magnetic field system. The motion of the magnetic energy well is regulated by the configuration of the nanoscale notch and the strength and frequency of the magnetic field, accompanying the jump motion of the carriers. The proposed concepts demonstrate that multiple carriers and cells can be manipulated and sorted using optimized nanoscale multi‐notch gates for a portable magnetophoretic system. This highlights the potential for developing cost‐effective point‐of‐care testing and lab‐on‐chip systems for various single‐cell‐level diagnoses and analyses.

36 MATERIALS SCIENCE↗

Enabling Microscale Processing: Combined Raman and Absorbance Spectroscopy for Microfluidic On-Line Monitoring

Microfluidics have many potential applications including characterization of chemical processes on a reduced scale, spanning the study of reaction kinetics using on-chip liquid–liquid extractions, sample pretreatment to simplify off-chip analysis, and for portable spectroscopic analyses. The use of in situ characterization of process streams from laboratory-scale and microscale experiments on the same chemical system can provide comprehensive understanding and in-depth analysis of any similarities or differences between process conditions at different scales. A well-characterized extraction of Nd(NO 3 ) 3 from an aqueous phase of varying NO 3– (aq) concentration with tributyl phosphate (TBP) in dodecane was the focus of this microscale study and was compared to an earlier laboratory-scale study utilizing counter current extraction equipment. Here, we verify that this same extraction process can be followed on the microscale using spectroscopic methods adapted for microfluidic measurement. Concentration of Nd (based on UV–vis) and nitrate (based on Raman) was chemometrically measured during the flow experiment, and resulting data were used to determine the distribution ratio for Nd. Extraction distributions measured on the microscale were compared favorably with those determined on the laboratory scale in the earlier study. Both micro-Raman and micro-UV–vis spectroscopy can be used to determine fundamental parameters with significantly reduced sample size as compared to traditional laboratory-scale approaches. This leads naturally to time, cost, and waste reductions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ion Traps and Packaging for Heterogenous Integration - Chimera

Microfabricated surface ion traps and silicon-based photonics are critical technologies for scaling quantum systems. Current ion trap architectures face scalability and integration challenges due to limitations in optical access, fabrication techniques, and material compatibility. State-of-the-art quantum computers and atomic clocks are investigating monolithic integration, which necessitates custom traps for each ion species and has not overcome the integration hurdles presented by merging these technologies. The Chimera (Ion Traps and Packaging for Heterogeneous Integration) project proposes a novel approach utilizing heterogeneous integration (HI) of ion traps and photonic circuits. This separation of components allows for flexibility in ion trap design and reduces fabrication compromises. The Chimera project specifically designed an ion trap to interface vertically with a separately fabricated waveguide chip and demonstrates the first steps to integrating them at the packaging level. The ion trap features a large area of removed silicon, allowing the photonics chip outputs closer to the ion trap, improving alignment and packaging processes. The alignment must be accurate to < 1 µm to ensure that the light from the waveguide can overlap with the trapping region. This fine alignment must also be maintained through an ultra-high vacuum bake, a critical step in preparing an ion trap experiment. By combining separate chips, we demonstrate a new path for scaling trapped ion technology that is less reliant on monolithic integration. We successfully fabricated a trap with a large area of oxide removed, resulting in a region thinned to about 40 µm, a key milestone toward successful integration.

42 ENGINEERING↗

Scaling and Single Event Effects (SEE) Sensitivity

This paper begins by discussing the potential for scaling down transistors and other components to fit more of them on chips in order to increasing computer processing speed. It also addresses technical challenges to further scaling. Components have been scaled down enough to allow single particles to have an effect, known as a Single Event Effect (SEE). This paper explores the relationship between scaling and the following SEEs: Single Event Upsets (SEU) on DRAMs and SRAMs, Latch-up, Snap-back, Single Event Burnout (SEB), Single Event Gate Rupture (SEGR), and Ion-induced soft breakdown (SBD).

Oldham, Timothy R.↗

Validating Extreme Scale Resilience with Veracity (Final Project Report)

As semiconductor technology continues to scale to smaller dimensions and as power-consumption of processor chips continues to increase in importance, there is significant uncertainty about the reliability of future large-scale computers. This uncertainty impacts algorithm design and application development, as well as decisions about system operation policy and provisioning decisions. The Veracity endeavored to answer pertinent questions regarding system reliability and resilience with the goal of enabling reliable science on future leadership machines while keeping overprovisioning low.

97 MATHEMATICS AND COMPUTING↗

Torque and Angular-Momentum Transfer in Merging Rotating Bose-Einstein Condensates

When rotating classical fluid drops merge together, angular momentum can be advected from one to another due to the viscous shear flow at the drop interface. It remains elusive what the corresponding mechanism is in inviscid quantum fluids such as Bose-Einstein condensates (BECs). In this work, we report our theoretical study of an initially static BEC merging with a rotating BECin three-dimensional space along the rotational axis. We show that a solitonlike sheet, resembling a corkscrew, spontaneously emerges at the interface. Furthermore, rapid angular momentum transfer at a constant rate universally proportional to the initial angular-momentum density is observed. Strikingly, this transfer does not necessarily involve fluid advection or drifting of the quantized vortices. We reveal that the corkscrew structure can exert a torque that directly creates angular momentum in the static BEC and annihilates angular momentum in the rotating BEC. Uncovering this intriguing angular momentum transport mechanism may benefit our understanding of various coherent matter-wave systems, spanning from atomtronics on chips to dark matter BECs at cosmic scales.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Software-Hardware Co-design of Heterogeneous SmartNIC System for Recommendation Models Inference and Training

Deep Learning Recommendation Models (DLRMs) are critical applications in various domains and have evolved as one of the single largest machine learning applications. Trillions of DLRM parameters exceed the on-chip memory capacity of GPUs. Large-scale multi-node systems are required for distributed DLRM inference and training, which suffer from the all-to-all communication bottleneck, mainly limiting the scalability of ever-growing DLRMs. In recent years, SmartNICs have evolved with coupled computation and communication capabilities providing opportunities for a powerful heterogeneous device in the system. However, there isn't such a distributed system that fully leverages the abundant smartNIC resources that resolve the scalability issue of DLRMs. In this work, we proposed a software-hardware co-design of a heterogeneous smartNIC system that resolves the communication bottleneck of distributed DLRMs, mitigates the memory bandwidth pressure, and improves computation efficiency. We provide a set of smartNIC designs of cache systems (including local cache and remote cache) and smartNIC computation kernels which reduce data movement, relieve memory lookup intensity, and improve the GPU's computation efficiency. In addition, we propose a graph algorithm that improves the data locality of queries within batches which optimizes the overall system performance with higher data reuse. Our evaluation shows that our system achieves 2.1x latency speedup for inference and 1.6x throughput speedup for training.

Guo, Anqi↗

A systolic architecture for the correlation and accumulation of digital sequences

A fully systolic architecture for the implementation of digital sequence correlator/accumulators is described. These devices consist of a two-dimensional array of processing elements that are conceived for efficient fabrication in Very Large Scale Integrated (VLSI) circuits. A custom VLSI chip that was implemented using these concepts is described. The chip, which contains a four-lag three-level sequence correlator and four bits of accumulation with overflow detection, was designed using the Integrated UNIX-Based Computer Aided Design (CAD) System. Applications of such devices include the synchronization of coded telemetry data, alignment of both real time and non-real time Very Large Baseline Interferometry (VLBI) signals, and the implementation of digital filters and processes of many types.

Deutsch, L. J.↗

On-chip Broadband Load Calibrators for Characterizing Transition-Edge Sensor Bolometers

We describe the design and fabrication of on-chip broadband load calibrators for element-by- element and full circuit characterization of transition-edge sensor (TES) bolometers for cosmic microwave background (CMB) polarimetry. Fabricated in a 4 X 4 grid on a 100 mm silicon wafer, the chips allow systematic testing of individual detector circuit elements without the need for an external blackbody source. This in-situ testing approach provides rapid and efficient characterization of various pixel designs in order to optimize the optical efficiency and control of systematic effects of the detectors. In this work, we demonstrate the on-chip calibrators on the Cosmology Large Angular Scale Surveyor (CLASS) W-band pixels. In particular, we investigate the loss on microstrip transmission lines and measure the performance of various designs of TESs, magic-tees, via-less crossovers, and signal terminations. We also test a filter-bank spectrometer with eight resonators spanning 80 – 105 GHz to characterize the spectral response of circuit elements across the detector bandwidth.

Sumit Dahal↗

Finite element analysis of a micromechanical deformable mirror device

A monolithic spatial light modulator chip was developed consisting of a large number of micrometer-scale mirror cells which can be rotated through an angle by application of an electrostatic field. The field is generated by electronics integral to the chip. The chip has application in photoreceptor based non-impact printing technologies. Chips containing over 16000 cells were fabricated, and were tested to several billions of cycles. Finite Element Analysis (FEA) of the device was used to model both the electrical and mechanical characteristics.

Sheerer, T. J.↗

Generating Weighted Test Patterns for VLSI Chips

Improved built-in self-testing circuitry for very-large-scale integrated (VLSI) digital circuits based on version of weighted-test-pattern-generation concept, in which ones and zeros in pseudorandom test patterns occur with probabilities weighted to enhance detection of certain kinds of faults. Requires fewer test patterns and less computation time and occupies less area on circuit chips. Easy to relate switching activity in outputs with fault-detection activity by use of probabilistic fault-detection techniques.

Siavoshi, Fardad↗

Nanoelectromechanical Control of Spin–Photon Interfaces in a Hybrid Quantum System on Chip

Color centers (CCs) in nanostructured diamond are promising for optically linked quantum technologies. Scaling to useful applications motivates architectures meeting the following criteria: C1 individual optical addressing of spin qubits; C2 frequency tuning of spin-dependent optical transitions; C3 coherent spin control; C4 active photon routing; C5 scalable manufacturability; and C6 low on-chip power dissipation for cryogenic operations. Here, we introduce an architecture that simultaneously achieves C1–C6. We realize piezoelectric strain control of diamond waveguide-coupled tin vacancy centers with ultralow power dissipation necessary. The DC response of our device allows emitter transition tuning by over 20 GHz, combined with low-power AC control. We show acoustic spin resonance of integrated tin vacancy spins and estimate single-phonon coupling rates over 1 kHz in the resolved sideband regime. Combined with high-speed optical routing, our work opens a path to scalable single-qubit control with optically mediated entangling gates.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

The Minos Computing Library: Efficient Parallel Programming for Extremely Heterogeneous Systems

Hardware specialization has become the silver bullet to achieve efficient high performance, from Systems-on-Chip systems, where hardware specialization can be ``extreme'', to large-scale HPC systems. As the complexity of the systems increases, so does the complexity of programming such architectures in a portable way. This work introduces the Minos Computing Library (MCL), as system software, programming model, and programming model runtime that facilitate programming extremely heterogeneous systems. MCL supports the execution of several multi-threaded applications within the same compute node, performs asynchronous execution of application tasks, efficiently balances computation across hardware resources, and provides performance portability. We show that code developed on a personal desktop automatically scales up to fully utilize powerful workstations with 8 GPUs and down to power-efficient embedded systems. MCL provides up to 17.5x speedup over OpenCL on NVIDIA DGX-1 systems and up to 1.88x speedup on single-GPU systems. In multi-application workloads, MCL dynamically resource allocation provides up to 2.43x performance improvement over manual, static allocation of computing resources.

Gioiosa, Roberto↗

Technology achievements and projections for communication satellites of the future

Multibeam systems of the future using monolithic microwave integrated circuits to provide phase control and power gain are contrasted with discrete microwave power amplifiers from 10 to 75 W and their associated waveguide feeds, phase shifters and power splitters. Challenging new enabling technology areas include advanced electrooptical control and signal feeds. Large scale MMIC's will be used incorporating on chip control interfaces, latching, and phase and amplitude control with power levels of a few watts each. Beam forming algorithms for 80 to 90 deg. wide angle scanning and precise beam forming under wide ranging environments will be required. Satelllite systems using these dynamically reconfigured multibeam antenna systems will demand greater degrees of beam interconnectivity. Multiband and multiservice users will be interconnected through the same space platform. Monolithic switching arrays operating over a wide range of RF and IF frequencies are contrasted with current IF switch technology implemented discretely. Size, weight, and performance improvements by an order of magnitude are projected.

Bagwell, J. W.↗

Technology achievements and projections for communication satellites of the future

Multibeam systems of the future using monolithic microwave integrated circuits to provide phase control and power gain are contrasted with discrete microwave power amplifiers from 10 to 75 W and their associated waveguide feeds, phase shifters and power splitters. Challenging new enabling technology areas include advanced electrooptical control and signal feeds. Large scale MMIC's will be used incorporating on chip control interfaces, latching, and phase and amplitude control with power levels of a few watts each. Beam forming algorithms for 80 to 90 deg wide angle scanning and precise beam forming under wide ranging environments will be required. Satellite systems using these dynamically reconfigured multibeam antenna systems will demand greater degrees of beam interconnectivity. Multiband and multiservice users will be interconnected through the same space platform. Monolithic switching arrays operating over a wide range of RF and IF frequencies are contrasted with current IF switch technology implemented discretely. Size, weight, and performance improvements by an order of magnitude are projected.

Bagwell, J. W.↗

Analog VLSI neural network integrated circuits

Two analog very large scale integration (VLSI) vector matrix multiplier integrated circuit chips were designed, fabricated, and partially tested. They can perform both vector-matrix and matrix-matrix multiplication operations at high speeds. The 32 by 32 vector-matrix multiplier chip and the 128 by 64 vector-matrix multiplier chip were designed to perform 300 million and 3 billion multiplications per second, respectively. An additional circuit that has been developed is a continuous-time adaptive learning circuit. The performance achieved thus far for this circuit is an adaptivity of 28 dB at 300 KHz and 11 dB at 15 MHz. This circuit has demonstrated greater than two orders of magnitude higher frequency of operation than any previous adaptive learning circuit.

Kub, F. J.↗

Low-Power CMOS Digital Autocorrelator Spectrometer

Prototype digital autocorrelator spectrometer circuit designed and built as assembly of few very-large-scale integrated (VLSI) complementary metal oxide/semiconductor (CMOS) circuit chips. Spectrometer contains 128 frequency channels and operates at clock rate of as much as 40 MHz. Total dc power needed is only 6 W. Digital autocorrelator spectrometer consists of four 32-channel autocorrelator chips that collectively produce 128-point spectrum of input signal as computed by use of Fourier transform.

Chandra, Kumar M.↗