Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Read Only Memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

394 records · Page 22

Strain-concentration for fast, compact photonic modulation and non-volatile memory

A critical figure of merit (FoM) for electro-optic (EO) modulators is the transmission change per voltage, d T / d V . Conventional approaches in wave-guided modulators maximize d T / d V via a high EO coefficient or longer light-material interaction lengths but are ultimately limited by material losses and nonlinearities. Optical and RF resonances improve d T / d V at the cost of spectral non-uniformity, especially for high- Q optical cavity resonances. Here, we introduce an EO modulator based on piezo-strain-concentration of a photonic crystal cavity to address both trade-offs: (i) it eliminates the trade-off between d T / d V and waveguide loss—i.e., enhancement of the resonance tuning efficiency d v c / d V for the fixed EO coefficient, waveguide length, and cavity Q —and (ii) at high DC strains it exhibits a non-volatile (NV) cavity tuning Δ v c ,NV for passive memory and programming of multiple devices into resonance despite fabrication variations. The device is fabricated on a scalable silicon nitride-on-aluminum nitride platform. We measure d v c / d V =177±1MHz/V, corresponding to Δ v c =40±0.32GHz for a voltage spanning ±120V with an energy consumption of δ U /Δ v c =0.17nW/GHz. The modulation bandwidth is flat up to ω BW,3dB /2 π =3.2±0.07MHz for broadband DC-AC and 142±17MHz for resonant operation near a 2.8 GHz mechanical resonance. Optical extinction up to 25 dB is obtained via Fano-type interference. Strain-induced beam-buckling modes are programmable under a “read-write” protocol with a continuous, repeatable tuning range of 5±0.25GHz, allowing for storage and retrieval, which we quantify with mutual information of 2.4 bits and a maximum non-volatile excursion of 8 GHz. Using a full piezo-optical finite-element-model (FEM) we identify key design principles for optimizing strain-based modulators and chart a path towards achieving performance comparable to lithium niobate-based modulators and the study of high strain physics on-chip.

Wen, Y. Henry (ORCID:0009000685423628)↗

Nonvolatile Memory Solution for Near-Term NASA Missions

Nonvolatile memory (NVM) system that could reliably function in extreme environments is one of the most critical components for many spacecrafts being developed for NASA missions to be launched in next four to seven years. NVM supports the computer system in saving and updating critical state data required for a warm restart after power cycling or in case of a power bus failure. It also provides a power independent mass storage capacity for the scientific data gathered by the instruments. In some cases the window for gathering such data is very small and occurs only once in a given mission. Commercially popular and fully developed Flash NVM technology is inappropriate for many reasons such as the limited read write cycles with slower access speeds, radiation intolerance, higher Single Event Upsets (SEU) rates, etc. It is desirable to have an NVM system based upon a robust cell technology making it immune to the SEUs and with sufficient radiation hardness. Availability of such NVM system seems to be still 5 to 10 years in the future. Meanwhile, it is possible to provide an interim hybrid solution by combining the existing rad-hard technologies. Additional information is contained in the original extended abstract.

Patel, J. U.↗

NASA Tech Briefs, August 2008

Customizable Digital Receivers for Radar Two-Camera Acquisition and Tracking of a Flying Target Visual Data Analysis for Satellites A Data Type for Efficient Representation of Other Data Types Hand-Held Ultrasonic Instrument for Reading Matrix Symbols Broadband Microstrip-to-Coplanar Strip Double-Y Balun A Topographical Lidar System for Terrain-Relative Navigation Programmable Low-Voltage Circuit Breaker and Tester Electronic Switch Arrays for Managing Microbattery Arrays Topics covered include: Lower-Dark-Current, Higher-Blue-Response CMOS Imagers; Fabricating Large-Area Sheets of Single-Layer Graphene by CVD; Support for Diagnosis of Custom Computer Hardware; Providing Goal-Based Autonomy for Commanding a Spacecraft; Dynamic Method for Identifying Collected Sample Mass; Optimal Planning and Problem-Solving; Attitude-Control Algorithm for Minimizing Maneuver Execution Errors; Grants Document-Generation System; Heat-Storage Modules Containing LiNO3 3H2O and Graphite Foam; Precipitation-Strengthened, High-Temperature, High-Force Shape Memory Alloys; Improved Relief Valve Would Be Less Susceptible to Failure; Safety Modification of Cam-and-Groove Hose Coupling; Using Composite Materials in a Cryogenic Pump; Using Electronic Noses to Detect Tumors During Neurosurgery; Producing Newborn Synchronous Mammalian Cells; Smaller, Lower-Power Fast-Neutron Scintillation Detectors; Rotationally Vibrating Electric-Field Mill; Estimating Hardness from the USDC Tool-Bit Temperature Rise; Particle-Charge Spectrometer; Automated Production of Movies on a Cluster of Computers; FIDO-Class Development Rover; and Tone-Based Command of Deep Space Probes Using Ground Antennas.

Source record↗

Energy Exascale Earth System Model v2.0.1

First patch release of v2.0.0 Changes since v2.0.0 [Important change] Fix ocean threading bug seen in debug cases on Chrysalis with Intel 20.0.4. Was introduced around time of v2.0.0 tag. Does not change v2.0.0 answers on Chrysalis because those didn't use threading or debugging. [EAM] Add semi-lagrangian tracer transport for theta-l (F90 and C++), add new algorithm for finding tropopause, add DSCREAM to allow v2 and SCREAM settings in same code such as adjust_ps [EAM-MMF] 60L default, allow C++ back end of RRTMGP (EAM too). [EAMxx] add nu-top functionality, fix forcing functor, add ttype9 and dcmip2012 tests 2.1, 2.2, and 3 HOMME: remove obsolete remap algs, option to specify dynamics alg indep of tracer, new sponge layer, add imex tests [ELM] Add topography-based subgrid (topounits), add FATES-ELM Nitro., Phos. and CH4 coupling, add land-use ts for NARRM, add lulc for SSP3 RCP7, Fix nutrient fertilization exp test and carbon isotope flux, Fix xactive lnd dry deposition, add lake water storage option, fix plant hydraulics 2d params, fix carbon budget calc, fix soil nutrient conc. bug, fix mosart dam bug, add test for new ELM, MOSART features, fix bug in O3 dry dep stomatal resistances, fix plant hydraulics restart BFB error, update mkmapdata. [MOSART] fix bug for reading the latitude from an unstructured input file, fix oversat in bubble test. [MPAS-ocean] Add CFC11, CFC12 tracers, add 2D spherical transport tests, fix del4 tracer mixing, add MARBL ocean tracer mixing, modify harmonic analysis options, add GPU port of vmix routines, fix calc of ML-averaged BV freq. [MPAS-seaice] Change extents of initial polar disks for oRRS18to6v3 grid, fix ice BGC with MARBL, update spherical test cases, fix DON coupling, Remove Cf from sea ice constants. [MPAS-landice] add CRYO1850-4xCO2 compset [CIME] add GCP, ANL GCE, Spock, Perlmutter, deprecate config_compilers.xml, fix and clean-up cmake macros, fix slurm bindings, refactor CIME internal testing, cleanup SCORPIO perf data, allow position independent compset naming, [also] update v2 benchmarking suite, extend e3sm_prod with throughput and memory checks

E3SM Project, DOE↗

Non-volatile, high density, high speed, Micromagnet-Hall effect Random Access Memory (MHRAM)

The micromagnetic Hall effect random access memory (MHRAM) has the potential of replacing ROMs, EPROMs, EEPROMs, and SRAMs because of its ability to achieve non-volatility, radiation hardness, high density, and fast access times, simultaneously. Information is stored magnetically in small magnetic elements (micromagnets), allowing unlimited data retention time, unlimited numbers of rewrite cycles, and inherent radiation hardness and SEU immunity, making the MHRAM suitable for ground based as well as spaceflight applications. The MHRAM device design is not affected by areal property fluctuations in the micromagnet, so high operating margins and high yield can be achieved in large scale integrated circuit (IC) fabrication. The MHRAM has short access times (less than 100 nsec). Write access time is short because on-chip transistors are used to gate current quickly, and magnetization reversal in the micromagnet can occur in a matter of a few nanoseconds. Read access time is short because the high electron mobility sensor (InAs or InSb) produces a large signal voltage in response to the fringing magnetic field from the micromagnet. High storage density is achieved since a unit cell consists only of two transistors and one micromagnet Hall effect element. By comparison, a DRAM unit cell has one transistor and one capacitor, and a SRAM unit cell has six transistors.

Wu, Jiin C.↗

Nonvolatile random access memory

A nonvolatile magnetic random access memory can be achieved by an array of magnet-Hall effect (M-H) elements. The storage function is realized with a rectangular thin-film ferromagnetic material having an in-plane, uniaxial anisotropy and inplane bipolar remanent magnetization states. The thin-film magnetic element is magnetized by a local applied field, whose direction is used to form either a 0 or 1 state. The element remains in the 0 or 1 state until a switching field is applied to change its state. The stored information is detcted by a Hall-effect sensor which senses the fringing field from the magnetic storage element. The circuit design for addressing each cell includes transistor switches for providing a current of selected polarity to store a binary digit through a separate conductor overlying the magnetic element of the cell. To read out a stored binary digit, transistor switches are employed to provide a current through a row of Hall-effect sensors connected in series and enabling a differential voltage amplifier connected to all Hall-effect sensors of a column in series. To avoid read-out voltage errors due to shunt currents through resistive loads of the Hall-effect sensors of other cells in the same column, at least one transistor switch is provided between every pair of adjacent cells in every row which are not turned on except in the row of the selected cell.

Wu, Jiin-Chuan↗

NASA Tech Briefs, November 1995

The contents include: 1) Mission Accomplished; 2) Resource Report: Marshall Space Flight Center; 3) NASA 1995 Software of the Year Award; 4) Microbolometers Based on Epitaxial YBa2Cu3O(sub 7-x) Thin Films; 5) Garnet Random-Access Memory; 6) Fabrication of SNS Weak Links on SOS Substrates; 7) High-Voltage MOSFET Switching Circuit; 8) Asymmetric Switching for a PWM H-Bridge Power Circuit; 9) Better Ohmic Contacts for InP Semiconductor Devices; 10) Low-Bandgap Thermovoltaic Materials and Devices; 11) Digital Frequency-Differencing Circuit; 12) Imaging Magnetometer; 13) Computer-Assisted Monitoring of a Complex System; 14) Buffered Telemetry Demodulator; 15) Compact Multifunction Inspection Head; 16) Optical Detection of Fractures in Ceramic Diaphragms; 17) Eddy-Current Detection of Cracks in Reinforced Carbon/Carbon; 18) Apparent Thermal Conductivity of Multilayer Insulation; 19) Optimizing Misch-Metal Compositions in Metal Hydride Anodes; 20) Device for Sampling Surface Contamination; 21) Probabilistic Failure Assessment for Fatigue; 22) Probabilistic Fatigue and Flaw-Propagation Analysis; 23) Windows Program for Driving the TDU-850 Printer; 24) Subband/Transform MATLAB Functions for Processing Images; 25) Computing Equilibrium Chemical Compositions; 26) Program Processes Thermocouple Readings; 27) ICAN-Second-Generation Integrated Composite Analyzer; 28) Integrated Composite Analyzer with Damping Capabilities; 29) Computing Efficiency of Transfer of Microwave Power; 30) Program Calculates Power Demands of Electronic Designs; 31) Cost-Estimation Program; 32) Program Estimates Areas Required by Electronic Designs; 33) Program to Balance Mapped Turbopump Assemblies; 34) BiblioTech; 35) Controlling Mirror Tilt With a Bimorph Actuator; 36) Burst-Disk Device Simulates Effect of Pyrotechnic Device; 37) Bearing-Mounting Concept Accommodates Thermal Expansion; 38) Parallel-Plate Acoustic Absorbers for Hot Environments; 39) Adjustable-Length Strut Withstands Large Cyclic Loads; 40) Tool Indicates Contact Angles in Bearing Raceways; 41) Gravity Slides With Magnetic Braking; 42) High-Torque, Lightweight, Pneumatically Driven Wrench for Small Spaces; 43) Device for Testing Compatibility of an O-Ring; 44) Magnetic Heat Pump Containing Flow Diverters; 45) Variable-Tilt Helicopter Rotor Mast; 46) "Beach-Ball" Robotic Rovers; 47) Apparatus Would Measure Temperatures of Ball Bearings; 48) Flexible Borescope for Inspecting Ducts; 49) Texturing Copper To Reduce Secondary Emission of Electrons; 50) Automated Laser Cutting in Three Dimensions; 51) Algorithm Helps Monitor Engine Operation; 52) Flexible Revision of Data-Processing Communications; 53) Software for Managing the Use of Land; 54) Thermal Strap Increases Cryocooling Efficiency; 55) Reversible Nut With Engagement Indication; 56) Control Algorithms for Kinematically Redundant Manipulators; 57) Computed Hydrogen-Flow Splits in a Rocket Engine; 58) Pressure and Thermal Modeling of Rocket Launches; 59) Field of View of a Spacecraft Antenna: Analysis and Software; 60) Digital Controller for Laser-Beam-Steering Subsystem; 61) More About Beam-Steering Subsystem for Laser Communication; 62) Digital Controller for Laser-Beam-Steering Subsystem: Part 2; 63) Interface Circuit Board for Space-Shuttle Communications; 64) Automated Planning of Spacecraft Telecommunications; 65) Artifacts of Spectral Analysis of Instrument Readings; 66) Neural-Network Controller for Vibration Suppression; 67) Adaptive Finite-Element Computation in Fracture Mechanics; 68) Attitude Control for the Cassini Spacecraft; 69) Analytical Model for Fluid Dynamics in a Microgravity Environment; 70) Study of Rocket-Engine Joints Bonded by NVCU/NARloy-Z; 71) Improved Silicon Nitride for Advanced Heat Engines; 72) Parameters for Welding Aluminum/Lithium Alloys; 73) Lightweight Composite Intertank Structure; 74) Foil Patches Seal Small Vacuum Leaks; 75) Data Base on Cables and Connectors; 76) Effect of Clock Mode on Radiation Hardnessf an ADC; and 77) Fault-Tolerant Control for a Robotic Inspection System.

Source record↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Data, Photographs, Videos, and Information for the Niwot Ridge Subalpine Forest (US-NR1) AmeriFlux site

This data package contains data and information about the operation of the Niwot Ridge Subalpine Forest AmeriFlux site (US-NR1) between Nov 1998 to the present (2020). This data archive supplements the primary 30-min data storage for the US-NR1 data (i.e., https://doi.org/10.17190/AMF/1246088) by providing the following: (i) five-minute statistics (means, variances, covariances) of all data measured by the data system between Nov 1998 and September 2020 in netCDF format, (ii) CSV data files saved within the memory of the CR23X data loggers (as well as an archive of the data logger programs), (iii) an archive of previous 30-min ASCII data versions of the US-NR1 AmeriFlux data and information related to each data release (a replica of what can be found at http://urquell.colorado.edu/data_ameriflux/), (iv) a web calendar (in HTML format) documenting activity at the site (a replica of http://urquell.colorado.edu/calendar/), (v) photos (over 15,000) and video taken at the site between years 2001 and present day (2020), and (vi) several auxiliary datasets, primary related to trees near the site, soil moisture and soil temperature, and subcanopy radiation data. The data package is setup so that the web calendar, photos, and electronic logbook can be easily accessed on a local computer using a web browser. The provided data files are in either netCDF, CSV, ASCII, or MATLAB format. To obtain a better understanding about the archive, please start by reading the PDF: README_ESS_DIVE_USNR1_readme_first.pdf.

54 ENVIRONMENTAL SCIENCES↗

FPGA Implementation of Stereo Disparity with High Throughput for Mobility Applications

High speed stereo vision can allow unmanned robotic systems to navigate safely in unstructured terrain, but the computational cost can exceed the capacity of typical embedded CPUs. In this paper, we describe an end-to-end stereo computation co-processing system optimized for fast throughput that has been implemented on a single Virtex 4 LX160 FPGA. This system is capable of operating on images from a 1024 x 768 3CCD (true RGB) camera pair at 15 Hz. Data enters the FPGA directly from the cameras via Camera Link and is rectified, pre-filtered and converted into a disparity image all within the FPGA, incurring no CPU load. Once complete, a rectified image and the final disparity image are read out over the PCI bus, for a bandwidth cost of 68 MB/sec. Within the FPGA there are 4 distinct algorithms: Camera Link capture, Bilinear rectification, Bilateral subtraction pre-filtering and the Sum of Absolute Difference (SAD) disparity. Each module will be described in brief along with the data flow and control logic for the system. The system has been successfully fielded upon the Carnegie Mellon University's National Robotics Engineering Center (NREC) Crusher system during extensive field trials in 2007 and 2008 and is being implemented for other surface mobility systems at JPL.

Random access memory↗

RKKY Exchange Bias Mediated Ultrafast All-Optical Switching of a Ferromagnet

The discovery of ultrafast helicity-independent all-optical switching (HI-AOS), as well as picosecond all-electrical switching of a ferrimagnet, has inspired the ultrafast spintronics community to explore ultrafast switching of a ferromagnet to achieve practical ultrafast storage and memory devices. Two explored mechanisms of HI-AOS of a ferromagnet in ferromagnet-ferrimagnet heterostructure are: a) exploiting the indirect exchange coupling with and b) injection of non-local spin current originated from a switching ferrimagnet. Here, in this manuscript, exchange mediated HI-AOS of a Ruderman–Kittel–Kasuya–Yosida (RKKY) exchange coupled “[Co/Pt]-multilayers/Pt spacer/CoGd” heterostructure is demonstrated. The authors have measured layer-resolved static magnetic properties, single-shot HI-AOS, and magnetization dynamics of the ferromagnetic Co/Pt multilayers (MLs), that are ferromagnetically or antiferromagnetically coupled with ferrimagnetic CoGd layers. Time-resolved magnetization dynamics reveal a 3.5 ps switching time of the Co/Pt MLs, which is the fastest switching of a ferromagnet reported to date. Employing an extended microscopic three-temperature model, the temporal dynamics of the exchange coupled ferromagnet–ferrimagnet heterostructure are simulated, qualitatively and quantitatively explaining the experimental switching phenomena. This work experimentally as well as theoretically establishes the mechanism of exchange mediated all-optical switching of ferromagnet-ferrimagnet heterostructures, which can be integrated with a magnetic tunnel junction for efficient reading after ultrafast energy-efficient switching.

36 MATERIALS SCIENCE↗

Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis

Abstract Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units used by hundreds to thousands of users simultaneously. Applications from large numbers of users have diverse characteristics, such as varying computation, communication, memory, and I/O intensity. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, I/O performance is becoming increasingly important as data sizes rapidly increase and large-scale applications, such as simulation and model training, are widely adopted. However, predicting I/O performance is difficult because I/O systems are shared among all users and involve many layers of software and hardware stack, including the application, network interconnect, operating system, file system, and storage devices. Furthermore, updates to these layers and changes in system management policy can significantly alter the I/O behavior of applications and the entire system. To improve the prediction of the I/O performance on HPC systems, we propose integrating information from several different system logs and developing a regression-based approach to predict the I/O performance. Our proposed scheme can dynamically select the most relevant features from the log entries using various feature selection algorithms and scoring functions, and can automatically select the regression algorithm with the best accuracy for the prediction task. The evaluation results show that our proposed scheme can predict the write performance with up to 90% prediction accuracy and the read performance with up to 99% prediction accuracy using the real logs from the Cori supercomputer system at NERSC.

97 MATHEMATICS AND COMPUTING↗

Harmonized Emissions Component (HEMCO) 3.0 as a Versatile Emissions Component for Atmospheric Models: Application in the GEOS-Chem, NASA GEOS, WRF-GC, CESM2, NOAA GEFS-Aerosol, and NOAA UFS Models

Emissions are a central component of atmospheric chemistry models. The Harmonized Emissions Component (HEMCO) is a software component for computing emissions from a user-selected ensemble of emission inventories and algorithms. It allows users to re-grid, combine, overwrite, subset, and scale emissions from different inventories through a configuration file and with no change to the model source code. The configuration file also maps emissions to model species with appropriate units. HEMCO can operate in offline stand-alone mode, but more importantly it provides an online facility for models to compute emissions at runtime. HEMCO complies with the Earth System Modeling Framework (ESMF) for portability across models. We present a new version here, HEMCO 3.0, that features an improved three-layer architecture to facilitate implementation into any atmospheric model and improved capability for calculating emissions at any model resolution including multiscale and unstructured grids. The three-layer architecture of HEMCO 3.0 includes (1) the Data Input Layer that reads the configuration file and accesses the HEMCO library of emission inventories and other environmental data, (2) the HEMCO Core that computes emissions on the user-selected HEMCO grid, and (3) the Model Interface Layer that re-grids (if needed) and serves the data to the atmospheric model and also serves model data to the HEMCO Core for computing emissions dependent on model state (such as from dust or vegetation). The HEMCO Core is common to the implementation in all models, while the Data Input Layer and the Model Interface Layer are adaptable to the model environment. Default versions of the Data Input Layer and Model Interface Layer enable straightforward implementation of HEMCO in any simple model architecture, and options are available to disable features such as re-gridding that may be done by independent couplers in more complex architectures. The HEMCO library of emission inventories and algorithms is continuously enriched through user contributions so that new inventories can be immediately shared across models. HEMCO can also serve as a general data broker for models to process input data not only for emissions but for any gridded environmental datasets. We describe existing implementations of HEMCO 3.0 in (1) the GEOS-Chem “Classic” chemical transport model with shared-memory infrastructure, (2) the high-performance GEOS-Chem (GCHP) model with distributed-memory architecture, (3) the NASA GEOS Earth System Model (GEOS ESM), (4) the Weather Research and Forecasting model with GEOS-Chem (WRF-GC), (5) the Community Earth System Model Version 2 (CESM2), and (6) the NOAA Global Ensemble Forecast System – Aerosols (GEFS-Aerosols), as well as the planned implementation in the NOAA Unified Forecast System (UFS). Implementation of HEMCO in CESM2 contributes to the Multi-Scale Infrastructure for Chemistry and Aerosols (MUSICA) by providing a common emissions infrastructure to support different simulations of atmospheric chemistry across scales.

Haipeng Lin↗

Variation Tolerant and Energy-Efficient Charge Domain Compute-in-Memory Array with Binary and Multi-Level Cell Ferroelectric FET

Here, in this work, we present a variation-tolerant and energy-efficient charge-domain Ferroelectric FET (FeFET) based Compute-in-Memory (CiM) array design that is compatible with both binary and multi-level cell memory sensing. We demonstrate that: 1) by exploiting FeFET as a nonvolatile switch, its high ON/OFF ratio in the subthreshold region can suppress the error introduced by the inaccurate ON state conductance, thus realizing robust CiM operations, unlike the current-domain CiM design where the computation results is highly sensitive to the device conductance variation; 2) by leveraging a dense dynamic random access memory (DRAM)-like 1FeFET1C cell structure, the proposed design benefits from the existing high density DRAM establishment while also significantly relaxing the capacitor retention and transistor leakage requirement; 3) the charge-domain CiM supports both binary FeFET with minimum overhead and MLC FeFET with tolerable latency for MLC state sensing, whose efficacy is validated experimentally on both cell-level and array-level; 4) the proposed CiM shows much better device variation resilience than conventional current-domain CiM, and also improves inference accuracy. Macro-level evaluation results demonstrate significantly higher energy efficiency and area efficiency compared to prior CiM works.

Duan, Jiahui [University of Notre Dame, IN (United↗

PANTHER: A Programmable Architecture for Neural Network Training Harnessing Energy-Efficient ReRAM

The wide adoption of deep neural networks has been accompanied by ever-increasing energy and performance demands due to the expensive nature of training them. Additionally, numerous special-purpose architectures have been proposed to accelerate training: both digital and hybrid digital-analog using resistive RAM (ReRAM) crossbars. ReRAM-based accelerators have demonstrated the effectiveness of ReRAM crossbars at performing matrix-vector multiplication operations that are prevalent in training. However, they still suffer from inefficiency due to the use of serial reads and writes for performing the weight gradient and update step. A few works have demonstrated the possibility of performing outer products in crossbars, which can be used to realize the weight gradient and update step without the use of serial reads and writes. However, these works have been limited to low precision operations which are not sufficient for typical training workloads. Moreover, they have been confined to a limited set of training algorithms for fully-connected layers only. To address these limitations, we propose a bit-slicing technique for enhancing the precision of ReRAM-based outer products, which is substantially different from bit-slicing for matrix-vector multiplication only. We incorporate this technique into a crossbar architecture with three variants catered to different training algorithms. To evaluate our design on different types of layers in neural networks (fully-connected, convolutional, etc.) and training algorithms, we develop PANTHER, an ISA-programmable training accelerator with compiler support. Our design can also be integrated into other accelerators in the literature to enhance their efficiency. Our evaluation shows that PANTHER achieves up to 8.02×, 54.21×, and 103× energy reductions as well as 7.16×, 4.02×, and 16× execution time reductions compared to digital accelerators, ReRAM-based accelerators, and GPUs, respectively.

42 ENGINEERING↗

ASK Magazine; No. 21

THIS ISSUE FEATURES A VISUAL DEPICTION OF THE ACADEMY of Program and Project Leadership (APPL). I imagine a variety of initial reactions to the drawing. One might be, "What is a cartoon doing in a magazine about project management?" Or perhaps, "Wow, nice colors-and fun." Another may be to closely search the image for signs, symbols and meaning. Still another, to read a new level of innovation and creativity into the picture. Undoubtedly, some readers will raise questions about the cost. Of course, any reaction is a sign of engagement. The stronger, the more energized the emotional and cognitive processing, the better. It is a sign of attention and interaction. For I've heard it said, "You only need to worry if they don t care one way or the other." So what is the point of the picture? To stimulate interest, raise questions, promote discussion, and maybe raise a smile.. .That, at least, was my initial reaction when I was introduced to the work of Nancy Hegedus, who helps to create these drawings for Root Learning Inc. At the NASA PM Conference, I was first shown the work Nancy had been doing with the help of Goddard s Knowledge Management Architect, Dr. Ed Rogers. I was immediately drawn into the power of visualization as a tool for more effective learning, communicating, and conveying complex knowledge concepts. We need new tools in today s world, where information and data overwhelms by sheer volume. There are articles, pamphlets, communications, and white papers-all aiming to convince and influence. Reactions to these tend to be either avoidance or mind-numbing, heavy-eyed consent; the message never registers or enters the soul. That s one of the reasons that APPL s Knowledge Sharing Initiative (KSI) has turned to storytelling as a memorable way of transfer- ring knowledge, inspiring imitation of best practices, and spurring reflection. ASK Magazine s recent fourth birthday marks an important milestone in APPL s continuing quest to provide ongoing support to project managers and to promote mission success. And similar to storytelling, the power of visualization is receiving increasing attention in recent years as a way to stimulate engagement. Pictures and visual graphs are viewed as one of the most effective ways for displaying, describing, and generating discussion about quantitative and technically complex information. Prototypes, models, and simulations are considered essential for stimulating innovation through open and engaging discussions. There has also been extensive writing on the use of visual graphics, pictures, and cartoons to facilitate memory, creativity, openness, attention-and even well-being. For many of these reasons, I am excited to have a colorful visual depiction of the APPL world included in ASK. Without the addition of text or slides, the intent is to invite people into the world of the APPL mission-as well as its products, services, customers, and partners- in a fun and engaging manner. As project leaders strive to find ways to encourage engagement, learning, and transmission of knowledge, traditional technologies are proving to be as valuable as modern technologies. (But for those who want more information in the form of texts and slide presentations, we certainly have an abundance of those as well.)

Laufer, Alexander↗