Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Read Only Memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

The ETROC2 as the Final Version for CMS Endcap Timing Layer (ETL) Upgrade

The ETROC (Endcap Timing Readout Chip) is being developed for the LGAD-based CMS Endcap Timing Layer (ETL) at HL-LHC. The ETL on each side of the interaction region will be instrumented with a two-disk system of MIP-sensitive LGAD (Low Gain Avalanche Diodes) silicon devices, read out by ETROCs for precision timing measurement with down to ~30 ps timing resolution per track. The ETROC is designed to handle a 16 x 16 pixel cell matrix, with each pixel being 1.3 mm x 1.3 mm to match the LGAD sensor pixel size. The front-end design for preamplifier and discriminator has been specifically optimized for the reduced LGAD signals, with enough flexibilities to meet the ETL specific needs for time resolution, power budget and radiation profile. The ETROC chip is implemented in a commercial 65nm CMOS process. Each channel consists of a preamplifier, a discriminator, a TDC used for TOA (Time Of Arrival) and TOT (Time Over Threshold) measurements, and a memory for data storage and readout. An in-pixel auto threshold calibration is included, along with a self-testing pattern generator. The TOT is used for time-walk correction of the TOA measurement. The detailed hit information (TOA and TOT) from each cell will be read out from a local circular buffer after each Level-1 Accept (about 1 MHz). In addition, a charge injection circuit is implemented to allow for testing and calibration. For more detailed monitoring of the signal pulses, waveform sampling circuits are included for one pixel. The clock distribution is based on a 16x16 H-tree design with a shielding structure to alleviate potential interference. The global peripheral circuits include a PLL, a phase shifter, an I2C slave controller, a fast control block, a global readout, and a data driver along with an efuse and temperature sensor. The ETROC builds event data frames for each L1A selected event and is also capable of providing L1 trigger information for user-defined delayed hits. The main design challenge is how to extract precision timing information from the small LGAD signals in the presence of high irradiation fluence, while keeping the power consumption and digital activity low. The ETL design goal for the time resolution of 50 ps per hit is required to achieve a 35 ps arrival time measurement for a MIP particle, which has its track registered in two ETL disk layers. The LGAD contribution is known to be about 30 ps, this means that the jitter from the ETROC has to be kept below 40 ps. The ETROC2 is the first full size full functionality prototype design fully compatible with the final chip specifications for CMS ETL and now becomes the final version. The ETROC2 chips have been extensively tested. We will present here new testing results including the bump bonding yield improvement study, the time walk correction (TWC) generality study with one pixel TWC applying to all pixels, the final SEU testing using both heavy ion and proton beam, more beam test studies including different sensors, and readiness for the ETROC2 production for CMS ETL upgrade.

Liu, Tiehui [Fermilab] (ORCID:0009000765225605)↗

Battery Monitoring System

The component that is powered by the battery pack being monitored is a valuable asset and must be in working condition at all times. Battery chemistry and characteristics have a major role in how to evaluate the state of the battery. The battery monitoring system has many parts that lead to an accurate battery reading. The components consist of a coulomb counting device, end of life voltage detection, a consideration of use for a real-time clock (RTC), temperature sensor, and non-volatile random-access memory (NVRAM). The combination of these elements allows the monitoring system to be highly reliant. Moving forward a better implementation of the ideas in this paper and further testing should ensure a high-quality battery monitoring system.

25 ENERGY STORAGE↗

Coherent Control over Nuclear Hyperpolarization Using an Optically Initializable Chromophore-Radical System

Chromophore radicals (CR) are emerging as important components for molecular quantum information science (QIS), especially in the context of quantum sensing. Here, we demonstrate that the optically hyperpolarized electrons in a 1,6,7,12-tetrakis(4-tert-butylphenoxy)-perylene-3,4,9,10-bis(dicarboximide) (tpPDI) covalently linked to a partially deuterated 1,3-bis(diphenylene)-d 16 -2-phenylallyl radical (BDPA-d 16 ) can be coherently manipulated via pulsed dynamic nuclear polarization (DNP) methods to transfer polarization to nuclear spins and back. Under light illumination at 85 K, electron hyperpolarization in BDPA is enhanced 2.1- to 2.4-fold over thermal polarization and lasts for more than 100 ms. By applying nuclear orientation via electron spin-locking (NOVEL) DNP, this optically amplified electron hyperpolarization was successfully transferred to a 1 H nuclear spin within the CR system and efficiently returned to the electron spin for readout via reverse-NOVEL. The NOVEL transfer efficiency of 65% amounts to a 688-fold nuclear spin hyperpolarization of the target nuclear spin, considering the 2.1-fold electron spin hyperpolarization. This reversible coherent manipulation of hyperpolarization transfer highlights the utility of CR systems to initialize and read out nuclear spin states in a disordered matrix at moderate cryogenic temperatures. Coupled with CRs’ environmental compatibility, tunability, and precise state initialization, these results highlight the promising role of nuclear spins in CRs for QIS applications, including quantum sensing and memory.

charge transfer↗

Persistent Memory Object Storage and Indexing for Scientific Computing

This paper presents Mosiqs, a persistent memory object storage framework with metadata indexing and querying for scientific computing. We design Mosiqs based on the key idea that memory objects on shared PM pool can live beyond the application lifetime and can become the sharing currency for applications and scientists. Mosiqs provides an aggregate memory pool atop an array of persistent memory devices to store and access memory objects. Mosiqs uses a lightweight persistent memory key-value store to manage the metadata of memory objects such as persistent pointer mappings, which enables memory object sharing for effective scientific collaborations. Mosiqs is implemented atop PMDK. We evaluate the proposed approach on many-core server with an array of real PM devices. The preliminary evaluation confirms a 100% improvement for write and 30% in read performance against a PM-aware file system approach.

Khan, Awais↗

Estimation of the time for steam generator trip due to cyber intrusions

The time required to trip a pressurized water reactor (PWR) by inserting malicious signals into its steam generator (SG) control system has been studied using the Generic PWR (GPWR) Simulator. A semi-analytical model is developed to approximately reproduce the simulator response and understand the dynamics of the control unit. A series of two proportional-integral controllers determines control action according to preset constants, the readings from the feedwater level sensor, and those from feedwater and steam flowrate transmitters. It is observed that the most important factor that determines whether a trip will occur is how much additional water is added to or withheld from the SG over time compared to normal operating conditions. In order to determine the effects of control action on the SG, changes in mass inventory are considered. This approach models the SG water level as a function of mass inventory and has a backward temporal memory. A Python interface is developed for the GPWR framework to automatically simulate different spoofing scenarios and post-process the related data. We observe that the trip times predominantly depend on flow mismatch and/or level errors. Controller parameters, including the integral time and gain constants, either speed up or slow down the rate of progression to a trip setpoint but do not cause a trip by themselves. The reactor can trip on a high-level signal when the reading crosses above 78%, increased from its reference level of 57%, or a low-level reading when it is below 25%. The present results show roughly how long the operators would have to respond to an attack, given a specific set of spoofing signals within the issue space analyzed. Furthermore, we have generated a simple surface by fitting a combination of exponential functions to the data obtained from the GPWR Simulator. In general, trips on a low level have been observed to occur faster than those on a high level.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Sharing tuples across independent coordination namespace systems

A system and method for federating a tuple storage database across multiple coordinated namespace (CNS) extended memory storage systems allowing the sharing of tuples and tuple data across independent systems. The method provides a federation service for multiple coordination namespace systems. The method retrieves a tuple from connected independent CNS systems wherein a local CNS Controller sends a read request to the local gatekeeper to retrieve a first tuple and creates a local pending remote record. The local gatekeeper at a requesting node sends a broadcast query to a plurality of remote gatekeepers for the tuple and Remote gatekeepers at remote nodes query in its local CNS for the tuple. The Local gatekeeper process at the requesting node receives results from a plurality of remote gatekeepers for the said tuple and selects one remote gatekeeper to receive the requested tuple and broadcasts a read for tuple data with selected gatekeeper.

Jacob, Philip↗

Accelerating Scientific Workflows on HPC Platforms with In Situ Processing

Scientific workflows drive most modern large-scale science breakthroughs by allowing scientists to define their computations as a set of jobs executed in a given order based on their data dependencies. Workflow management systems (WMSs) have become key to automating scientific workflows-executing computational jobs and orchestrating data transfers between those jobs running on complex high-performance computing (HPC) platforms. Traditionally, WMSs use files to communicate between jobs: a job writes out files that are read by other jobs. However, HPC machines face a growing gap between their storage and compute capabilities. To address that concern, the scientific community has adopted a new approach called in situ, which bypasses costly parallel filesystem I/O operations with faster in-memory or in-network communications. When using in situ approaches, communication and computations can be interleaved. In this work, we leverage the Decaf in situ dataflow framework to accelerate task-based scientific workflows managed by the Pegasus WMS, by replacing file communications with faster MPI messaging. We propose a new execution engine that uses Decaf to manage communications within a sub-workflow (i.e., set of jobs) to optimize inter-job communications. We consider two workflows in this study: (i) a synthetic workflow that benchmarks and compares file- and MPI-based communication; and (ii) a realistic bioinformatics workflow that computes mu-tational overlaps in the human genome. Experiments show that in situ communication can improve the bioinformatics workflow execution time by 22% to 30% compared with file communication. Our results motivate further opportunities and challenges for bridging traditional WMSs with in situ frameworks.

Decaf↗

Variation-Resilient FeFET-Based In-Memory Computing Leveraging Probabilistic Deep Learning

Reliability issues stemming from device level nonidealities of nonvolatile emerging technologies like ferroelectric field-effect transistors (FeFETs), especially at scaled dimensions, cause substantial degradation in the accuracy of in-memory crossbar-based AI systems. Here, in this work, we present a variation-aware design technique to characterize the device level variations and to mitigate their impact on hardware accuracy employing a Bayesian neural network (BNN) approach. An effective conductance variation model is derived from the experimental measurements of cycle-to-cycle (C2C) and device-to-device (D2D) variations performed on FeFET devices fabricated using 28 nm high-k metal gate technology. The variations were found to be a function of different conductance states within the given programming range, which sharply contrasts earlier efforts where a fixed variation dispersion was considered for all conductance values. Such variation characteristics formulated for three different device sizes at different read voltages were provided as prior variation information to the BNN to yield a more exact and reliable inference. Near-ideal accuracy for shallow networks (MLP5 and LeNet models) on the MNIST dataset and limited accuracy decline by ~3.8%–16.1% for deeper AlexNet models on CIFAR10 dataset under a wide range of variations corresponding to different device sizes and read voltages, demonstrates the efficacy of our proposed device-algorithm co-design technique.

97 MATHEMATICS AND COMPUTING↗

An Accurate, Error-Tolerant, and Energy-Efficient Neural Network Inference Engine Based on SONOS Analog Memory

In this work, we demonstrate SONOS (silicon-oxide-nitrideoxide- silicon) analog memory arrays that are optimized for neural network inference. The devices are fabricated in a 40nm process and operated in the subthreshold regime for in-memory matrix multiplication. Subthreshold operation enables low conductances to be implemented with low error, which matches the typical weight distribution of neural networks, which is heavily skewed toward near-zero values. This leads to high accuracy in the presence of programming errors and process variations. We simulate the end-to-end neural network inference accuracy, accounting for the measured programming error, read noise, and retention loss in a fabricated SONOS array. Evaluated on the ImageNet dataset using ResNet50, the accuracy using a SONOS system is within 2.16% of floating-point accuracy without any retraining. The unique error properties and high On/Off ratio of the SONOS device allow scaling to large arrays without bit slicing, and enable an inference architecture that achieves 20 TOPS/W on ResNet50, a >10× gain in energy efficiency over state-of-the-art digital and analog inference accelerators.

97 MATHEMATICS AND COMPUTING↗

Differential Power Processing for Ultra-Efficient Data Storage

Here this paper presents the hardware, software, and power codesign of an ultra-efficient data storage server with differential power processing (DPP). DPP can reduce the power conversion stress, improve the efficiency, and enhance the functionality of modular power electronics systems. The power inputs of a large number of hard disk drives (HDDs) were connected in series and supported by a multiport ac-coupled differential power processing (MAC-DPP) converter through a multiwinding transformer. Methods for controlling the multi-input multi-output power flow in the multiwinding transformer while avoiding core saturation were investigated. A ten-port MAC-DPP prototype with 700-W/in 3 power density was built to support a 450-W HDD storage system with ten series-stacked voltage domains. The prototype was tested on a 50-HDD server testbench, and the overall system loss is below 1 W (99.77% system efficiency). The server was able to maintain high-speed reading and writing operation of all 50 HDDs against the worst hot-swapping scenarios. A variety of hardware/software configurations and many cloud storage techniques were tested on the fully functioning server. Experimental results show that the energy efficiency of large-scale information systems (CPU/GPU clusters, memory banks, HDD arrays, etc.) can be greatly improved by software, hardware, and power codesign.

42 ENGINEERING↗

Streaming Data Reorganization at Scale with DeltaFS Indexed Massive Directories

We report complex storage stacks providing data compression, indexing, and analytics help leverage the massive amounts of data generated today to derive insights. It is challenging to perform this computation, however, while fully utilizing the underlying storage media. This is because, while storage servers with large core counts are widely available, single-core performance and memory bandwidth per core grow slower than the core count per die. Computational storage offers a promising solution to this problem by utilizing dedicated compute resources along the storage processing path. We present DeltaFS Indexed Massive Directories (IMDs), a new approach to computational storage. DeltaFS IMDs harvest available (i.e., not dedicated) compute, memory, and network resources on the compute nodes of an application to perform computation on data. We demonstrate the efficiency of DeltaFS IMDs by using them to dynamically reorganize the output of a real-world simulation application across 131,072 CPU cores. DeltaFS IMDs speed up reads by 1,740x while only slightly slowing down the writing of data during simulation I/O for in situ data processing.

97 MATHEMATICS AND COMPUTING↗

Single-shot switching in Tb/Co-multilayer based nanoscale magnetic tunnel junctions

Magnetic tunnel junctions (MTJs) are elementary units of magnetic memory devices. For high-speed and low-power data storage and processing applications, fast reversal of the magnetization by an ultrashort laser pulse is extremely important. Here we demonstrate single-shot switching of Tb/Co-multilayer based nanoscale MTJs by combining the optical writing and the electrical read-out methods. A 90-fs-long laser pulse switches the magnetization of the storage layer (SL). The change in the tunneling magnetoresistance (TMR) between the SL and a reference layer (RL) is probed electrically across the oxide barrier. Single-shot switching is demonstrated by varying the cell diameter from 300 nm to 20 nm. The anisotropy, magnetostatic coupling, and switching probability exhibit cell-size dependence. By suitable association of laser fluence and magnetic field, successive commutation between high-resistance and low-resistance states is achieved. The nature of the magnetization reversal of both SL and RL in a continuous film is probed with a depth-resolved magneto-optical Kerr effect (MOKE) magnetometry. The ultrafast dynamics in the continuous full-MTJ stack is investigated with the time-resolved pump–probe technique. Our experimental findings provide strong support for the growing interest in ultrafast spintronic devices.

36 MATERIALS SCIENCE↗

Educating HPC Users in the use of advanced computing technology

We examine a multi-modal approach to educating and training users of an advanced computing technology testbed at the Institute for Advanced Computational Science at Stony Brook University. Ookami provides researchers worldwide with access to 176 Fujitsu A64FX compute nodes, this being the same processor technology powering the Japanese Fugaku supercomputer, the fastest computer in the world since June 2020. However, achieving high-performance on this Arm-based, leadership computing technology requires that users be familiar with details of computer architecture, performance analysis and modeling, and high-performance programming models that are commonly omitted in introductory programming courses. Indeed, regardless of their seniority, many of the testbed users are surprisingly unfamiliar with basic concepts such as vectorization, pipelining, latency/bandwidth, roofline models, computing energy/power, threads, and non-uniform memory access. These same concepts also pervade mainstream x86 technologies, so this is of widespread concern. Due to the national/global nature of our user community that is also very diverse in both discipline and experience, the inability to offer formal classes, and our experience that most people do not tend to read online documentation or training materials in sufficient depth, we have consciously employed multiple approaches that heavily emphasize (online) personal interactions and transfer of skills. Online documentation has been organized around best-practices and FAQs; twice-weekly hackathons and office hours via Zoom enable deep dives by both the team and the user community with multiple broad benefits; a Slack channel provides both real time and archived answers and discussions; and workshops, training and webinars target community needs as they arise. Furthermore, the perspective that these tools are being used in an educational setting rather than just for project communication makes them more effective and contributes to community success.

A64FX↗

Critical Assessment of Metagenome Interpretation: the second round of challenges

Abstract Evaluating metagenomic software is key for optimizing metagenome interpretation and focus of the Initiative for the Critical Assessment of Metagenome Interpretation (CAMI). The CAMI II challenge engaged the community to assess methods on realistic and complex datasets with long- and short-read sequences, created computationally from around 1,700 new and known genomes, as well as 600 new plasmids and viruses. Here we analyze 5,002 results by 76 program versions. Substantial improvements were seen in assembly, some due to long-read data. Related strains still were challenging for assembly and genome recovery through binning, as was assembly quality for the latter. Profilers markedly matured, with taxon profilers and binners excelling at higher bacterial ranks, but underperforming for viruses and Archaea. Clinical pathogen detection results revealed a need to improve reproducibility. Runtime and memory usage analyses identified efficient programs, including top performers with other metrics. The results identify challenges and guide researchers in selecting methods for analyses.

54 ENVIRONMENTAL SCIENCES↗

Giant Domain Wall Conductivity in Self‐Assembled BiFeO 3 Nanocrystals

Abstract Ever‐increasing demand on electronic devices with ultrahigh‐density non‐volatile data storage has attracted great interest in novel ferroelectric memories based on conductive ferroelectric domain walls. Embedded in an insulating material, ferroelectric domain walls have the capability of being (re)created, displaced, erased, and altered in their spatial configurations and electronic characteristics. However, the domain wall conductivities are in most cases not yet sufficiently high to ensure the current density required to drive read‐out circuits operating at high speeds. In this work, a giant domain wall current (>10 µA) of a single charged domain wall is obtained through conductive atomic force microscopy with a bias field of 4 V. This is achieved in self‐assembled BiFeO 3 nanocrystals grown by sol‐gel method on Nb‐doped SrTiO 3 substrates. Local configurations of domains and domain wall types are studied using vector piezoresponse force microscopy and high‐resolution transmission electronic microscopy. The enhancement of the wall current is shown to be due to the formation of conducting pathways of charged defects accumulated along domain walls and traversing the nanocrystals. The diverse domain walls can be manipulated by electric field in a perpendicular architecture. The perpendicular array structure of BiFeO 3 nanocrystals should have great potentials for developing perpendicular nanoelectronic prototypes.

Liu, Lisha↗

L'Arlesienne de ROOT

Over many years, ROOT users have repeatedly stumbled over—and loudly rediscovered—the infamous 1 GB limit on individual I/O operations, a constraint that somehow survived long past the era when anyone thought a gigabyte was “a lot.” As experiments embraced ever-larger objects and collections, this limit became an increasingly unavoidable rite of passage. This contribution recounts the sustained, multi-year quest by ROOT I/O developers to finally retire this relic, navigating a maze of legacy APIs, memory-management assumptions, and integer boundaries that seemed determined to preserve the status quo. We describe how internal interfaces were carefully modernized to introduce fully 64-bit–capable code paths without breaking the mountains of existing user code that would definitely have noticed. With the limit now lifted, ROOT can finally handle multi-gigabyte objects in a single read or write operation, even when splitting them into an RNTuple is not an option (we’re looking at you, large RooWorkspaces and giant histograms), liberating users from yet another “fun” debugging adventure and clearing the way for the massive analyses of the HL-LHC and beyond.

Canal, Philippe G. [Fermilab] (ORCID:0000000277487↗

L'Arlesienne de ROOT

Over many years, ROOT users have repeatedly stumbled over—and loudly rediscovered—the infamous 1 GB limit on individual I/O operations, a constraint that somehow survived long past the era when anyone thought a gigabyte was “a lot.” As experiments embraced ever-larger objects and collections, this limit became an increasingly unavoidable rite of passage. This contribution recounts the sustained, multi-year quest by ROOT I/O developers to finally retire this relic, navigating a maze of legacy APIs, memory-management assumptions, and integer boundaries that seemed determined to preserve the status quo. We describe how internal interfaces were carefully modernized to introduce fully 64-bit–capable code paths without breaking the mountains of existing user code that would definitely have noticed. With the limit now lifted, ROOT can finally handle multi-gigabyte objects in a single read or write operation, even when splitting them into an RNTuple is not an option (we’re looking at you, large RooWorkspaces and giant histograms), liberating users from yet another “fun” debugging adventure and clearing the way for the massive analyses of the HL-LHC and beyond.

Canal, Philippe G. [Fermilab] (ORCID:0000000277487↗