Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “programmable devices”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Design and Performance of Kokkos Staging Space toward Scalable Resilient Application Couplings

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose Kokkos data staging memory space, an extension of Kokkos' data abstraction (memory space) for heterogeneous computing systems. This new abstraction allows to express data on a virtual shared-space for multiple Kokkos applications, thus extending Kokkos to support inter-application data exchange to build an efficient application workflow. Additionally, we study the effectiveness of asynchronous data layout conversions for applications requiring different memory access patterns for the shared data. Our preliminary evaluation with a synthetic benchmark indicate the effectiveness of this conversion adapted to three different scenarios representing access frequency and use patterns of the shared data.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Hardware and Software Bill of Materials (HBOMs/SBOMs) Extraction Methods

Hardware and software bills of materials (HBOMs and SBOMs) provide important visibility into the components, dependencies, and supply chain relationships within programmable digital devices. This visibility is critical for advanced nuclear reactor applications, where use of common or shared hardware components, software libraries, suppliers, or manufacturing processes may create common cause failure (CCF) vulnerabilities despite apparent diversity. This paper evaluates current approaches for obtaining and analyzing HBOMs and SBOMs in support of CCF, diversity and defense-in-depth (D3) assessments, and begins to explore potential methods for artificial intelligence/machine learning-based analysis. The availability of BOM information from advanced reactor manufacturers and vendors, representative hardware and software categories found in advanced reactor systems continues to limit research [13]. This paper compares commonly used BOM formats, including CycloneDX, SPDX, and SWID. It also surveys publicly available tools for generating BOMs from source code, compiled binaries, and hardware-related information, noting limitations in language coverage, system age, and format interoperability. Finally, this paper evaluates methods for correlating BOM data with vulnerability and exploitability information, including VEX, CVE, and CWE resources. The findings indicate that publicly available nuclear-vendor BOMs are limited, making third-party extraction and research into novel analysis techniques necessary.

Cybersecurity↗

Aquatic organism tracking devices, systems and associated methods

Aquatic organism tracking devices, systems and associated methods are described. According to one aspect, an aquatic organism tracking device includes a housing, a transducer coupled with the housing and configured to transmit a data transmission externally of the housing and an aquatic organism associated with the tracking device, a programmable oscillator coupled with the housing, and wherein the programmable oscillator is configured to generate a clock signal having a selected one of a plurality of different frequencies, processing circuitry coupled with the housing and configured to receive the clock signal from the programmable oscillator and to execute a plurality of executable instructions according to the clock signal, a power source coupled with the housing and configured to store electrical energy, and wherein the processing circuitry is configured to control the provision of the electrical energy from the power source to the transducer to generate the data transmission as a result of the execution of the instructions.

Deng, Z. Daniel↗

LC-MEMENTO: A Memory Model for Accelerated Architectures

With the advent of heterogeneous architectures, in particular, with the ubiquity of multi-GPU systems, it is becoming increasingly important to manage device memory efficiently in order to reap the benefits of the additional core count. To date, such responsibility mainly falls on the programmer where device-to-host data communication (and vice versa), if not done properly, may incur costly memory transfer operations and synchronization. The problem may be compounded by additional requirement to maintain system-wide memory consistency that may involve expensive synchronization overhead. In this paper, we present Location Consistency Memory Model for Enhanced Transfer Operations (LC-MEMENTO). This framework considers incorporating runtime techniques for multi-GPU memory management to support relaxed synchronization semantics and memory transfer operations automatically. Specifically, we implement a relaxed form of a memory consistency model based on the Location Consistency (LC) in an Asynchronous Many-Task Runtime (ARTS) and demonstrate that, this memory model enables additional optimization opportunities for the three representative applications encompassing different computational patterns (scientific computation, graphs, data streaming, etc.).

Memory Models, Accelerators, Adaptive Optimization↗

Two-Dimensional Silk Crystal Films as Matrix Layer for High-Performance Microelectronics

This study explores a bio-inspired approach for memristive devices by combining Keggin-type polyoxometalates (POMs)-[SiW 12 O 40 ] 4 (POM-T) and [PW 12 O 40 ] 3 (POM-P), with silk fibroin (SF) to create 2D SF–POM layers on highly ordered pyrolytic graphite (HOPG) as resistive switching layers for memristors. We propose that the ordered SF layer template 0D POMs facilitate the formation of conductive filaments, thereby enhancing the variability of the manufactured memristors. AFM analysis revealed that both SF and SF–POM layers shared similar morphologies, while SF–POM–T formed larger aggregates, likely due to the stronger acidity of POM-T, which probably caused SF to aggregate and alter its secondary structure. Scanning Kelvin probe microscopy (SKPM) revealed that POMs reduced the contact potential difference of HOPG, resulting in lower work functions. Compared to an SF device, the SF–POM–P device showed improved memristive behavior, with a larger current gap and good repeatability over multiple sweeps; whereas the SF–POM–T device did not exhibit memristor activity, likely due to acidity-induced disruption of the SF template’s order and CF formation. More importantly, SF–POM–P devices also demonstrated programmable memristive states. Finally, combining simulation-driven memristor modeling, we showcase a co-design workflow for advancing bioinspired memristors through new materials design, synthesis, and device modeling and development.

36 MATERIALS SCIENCE↗

Modbus RTU for Embedded Cyber Secure Inverter Controller

The Modbus communication protocol is a widely adopted communication standard in industrial control systems. This communication protocol is known for being reliable and straightforward to implement while being versatile in terms of its operating parameters while supporting multiple formats over various hardware infrastructures and architectures. Many intelligent devices such as Programmable Logic Controllers (PLCs), Human-Machine Interfaces (HMIs), Internet-of-Things (IoT), and various Operational Technologies (OT) utilize Modbus for their communication systems. These types of systems must communicate with each other through a standardized and central communication process. To support the integration of these modular systems, a Field-Programmable Gate Array (FPGA) can act as an embedded central routing fabric for this communication to take place. Embedded systems are versatile enough to interface with various devices and systems to accomplish various goals. Additionally, embedded systems require relatively small physical designs to minimize the required resources to facilitate the intended application by providing low-level system access. This minimization of system resources goes hand in hand with reducing the financial cost of a proposed solution or system. As remotely collaborating researchers often use FPGAs to prototype designs that are required to have a method for data transmission among systems, it is imperative to provide a baseline standard for communications among devices and systems. A typical method of implementing the Modbus RTU communication protocol in an embedded environment is using integrated logic architectures within the FPGA called “Intellectual Property (IP) cores.” IP cores can be designed using integrated logic or circuit designs to function as an embedded processor. These IP cores can then perform the required computational actions to support the Modbus RTU communication protocol by utilizing high-level programming languages such as the C programming language. The hardware description language of Very High-Speed Integrated Circuit Hardware Description Language (VHDL) allows for the control of real hardware at the logic gate and signal level. These logic gates and signals can be designed and controlled to perform desired actions based on the system design. Programming an FPGA using VHDL allows an individual to access the lowest abstraction level of the system during FPGA development. This level of abstraction is referred to as the register-transfer level (RTL), which gives access to manipulating values and variables at the register level. This register-level manipulation provides precision over creating the logical circuit within the FPGA, thus minimizing the required code to perform desired operations. The Modbus RTU communication protocol can be implemented within an FPGA using VHDL programming to establish a standardized and embedded serial communication pathway. This implementation provides a standardized communication protocol to streamline research efforts among researchers, thus increasing the efficiency of research efforts. Additionally, this Modbus RTU implementation requires fewer resources when compared to typical communication protocol implementations that utilize an IP core, reducing the hardware requirement for effective research efforts.

communication↗

Machine learning for arbitrary single-qubit rotations on an embedded device

Here, in this study, we present a technique for using machine learning (ML) for single-qubit gate synthesis on field-programmable logic for a superconducting transmon-based quantum computer based on simulated studies. Our approach is multi-stage. We first “bootstrap” a model based on simulation with access to the full state vector for measuring gate fidelity. We next present an algorithm, named adapted randomized benchmarking (ARB), for fine-tuning the gate on hardware based on measurements of the devices. We also present techniques for deploying the model on programmable devices with care to reduce the required resources. While the techniques here are applied to a transmon-based computer, many of them are portable to other architectures.

97 MATHEMATICS AND COMPUTING↗

Ultrafast jet classification at the HL-LHC

Abstract Three machine learning models are used to perform jet origin classification. These models are optimized for deployment on a field-programmable gate array device. In this context, we demonstrate how latency and resource consumption scale with the input size and choice of algorithm. Moreover, the models proposed here are designed to work on the type of data and under the foreseen conditions at the CERN large hadron collider during its high-luminosity phase. Through quantization-aware training and efficient synthetization for a specific field programmable gate array, we show that O ( 100 ) ns inference of complex architectures such as Deep Sets and Interaction Networks is feasible at a relatively low computational resource cost.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

AEcroscopy: A Software–Hardware Framework Empowering Microscopy Toward Automated and Autonomous Experimentation

Microscopy has been pivotal in improving the understanding of structure-function relationships at the nanoscale and is by now ubiquitous in most characterization labs. However, traditional microscopy operations are still limited largely by a human-centric click-and-go paradigm utilizing vendor-provided software, which limits the scope, utility, efficiency, effectiveness, and at times reproducibility of microscopy experiments. Here, in this work, a coupled software–hardware platform is developed that consists of a software package termed AEcroscopy (short for Automated Experiments in Microscopy), along with a field-programmable-gate-array device with LabView-built customized acquisition scripts, which overcome these limitations and provide the necessary abstractions toward full automation of microscopy platforms. The platform works across multiple vendor devices on scanning probe microscopes and electron microscopes. It enables customized scan trajectories, processing functions that can be triggered locally or remotely on processing servers, user-defined excitation waveforms, standardization of data models, and completely seamless operation through simple Python commands to enable a plethora of microscopy experiments to be performed in a reproducible, automated manner. This platform can be readily coupled with existing machine-learning libraries and simulations, to provide automated decision-making and active theory-experiment optimization to turn microscopes from characterization tools to instruments capable of autonomous model refinement and physics discovery.

47 OTHER INSTRUMENTATION↗

Low-power anisotropic molecular electronic memristors

A molecular electronic memristor, programmable resistive memory device, promises to revolutionize next-generation flexible data storage units, offering fast, dense and ultralow power solutions. Here we report anisotropic resistive switching in molecular κ-(BEDT-TTF) 2 Cu[N(CN) 2 ]Cl memristors, consisting of alternatively segregated bis(ethylenedithio)tetrathiafulvalene (BEDT-TTF) and Cu[N(CN) 2 ]Cl layers. Electron resistance switching behavior controlled by charge tunneling in molecular memristors show a low set voltage of 0.5 V (10 V/cm) with the ON/OFF ratio of 2.3×10 3 along a-axis and a high-level endurance of 1.25×10 4 cycles along all axes. Finally, the findings of such molecular electronic crystals promise for low-power data storage memristors.

36 MATERIALS SCIENCE↗

Short-Depth QAOA circuits and Quantum Annealing on Higher-Order Ising Models (Rev.2)

The Quantum Alternating Operator Ansatz (QAOA) and Quantum Annealing (QA) are quantum algorithms that are both based on the adiabatic theorem and both have the goal of sampling the optimal solution(s) of combinatorial optimization problems. Quantum annealing has been physically instantiated on D-Wave devices using superconducting flux qubits, and QAOA can be programmed on digital gate-model quantum computers such as the programmable superconducting transmon qubits devices of the IBMQ series, for instance ibm washington. QAOA and QA address the same types of problems, but it is unclear how they will scale to large problem sizes and to larger and higher-fidelity quantum computers. In this article, we present a direct comparison between QAOA, one and two rounds, run on all 127 qubits of ibm washington and QA run on D-Wave Advantage system4.1 and Advantage system6.1. The problems which allow for this comparison are random Ising model problems whose connectivity matches the heavy hexagonal lattice topology of ibm washington and the Pegasus graph connectivity of the two D-Wave devices. We create two classes of problem instances for this comparison: one with higher order terms (ZZZ variable interactions), linear terms, and quadratic terms, and a separate problem type with only linear and quadratic terms. Our QAOA circuits are novel and extremely short depth, with a CNOT depth of 6 per round, which allows whole chip usage of ibm washington’s heavy hexagonal lattice and can be applied to future heavy-hex chips. We also test the effectiveness of the error suppression technique digital dynamical decoupling on the QAOA circuits. The QAOA circuits compiled to ibm washington are composed of several thousand circuit instructions, approximately 3, 000 depending on the details of the circuit, making these some the largest quantum circuits ever executed on a digital quantum processor. QAOA and QA are compared against the classical heuristic algorithm of simulated annealing and all problem instances are exactly solved using CPLEX in order to evaluate which samplers, if any, correctly found the ground state solution(s) of the problem instances. We find that (i) QA outperforms QAOA on all problem instances, (ii) QAOA samples the problems better than random sampling, and (iii) QAOA angle computation exhibits clear parameter concentration across the ensemble of Ising models.

127 Qubits↗

Programmable soft valves for digital and analog control

In soft devices, complex actuation sequences and precise force control typically require hard electronic valves and microcontrollers. Existing designs for entirely soft pneumatic control systems are capable of either digital or analog operation, but not both, and are limited by speed of actuation, range of pressure, time required for fabrication, or loss of power through pull-down resistors. Using the nonlinear mechanics intrinsic to structures composed of soft materials—in this case, by leveraging membrane inversion and tube kinking—two modular soft components are developed: a piston actuator and a bistable pneumatic switch. These two components combine to create valves capable of analog pressure regulation, simplified digital logic, controlled oscillation, nonvolatile memory storage, linear actuation, and interfacing with human users in both digital and analog formats. Three demonstrations showcase the capabilities of systems constructed from these valves: 1) a wearable glove capable of analog control of a soft artificial robotic hand based on input from a human user’s fingers, 2) a human-controlled cushion matrix designed for use in medical care, and 3) an untethered robot which travels a distance dynamically programmed at the time of operation to retrieve an object. This work illustrates pathways for complementary digital and analog control of soft robots using a unified valve design.

42 ENGINEERING↗

PLC Integration for the Horn A Magnetic Field Mapping Device

This presentation summarizes the integration of a programmable logic controller (PLC) into the Horn A magnetic field mapping device for the Long-Baseline Neutrino Facility (LBNF). The device is used to position magnetic field probes along the centerline of the Horn A inner conductor to verify that the magnetic field within this region is approximately zero. The presentation covers the PLC hardware and electrical integration, stepper motor and encoder control, ladder logic development, mechanical integration, and system testing. Results include successful bidirectional motion and encoder feedback for the translation and rotation axes, as well as characterization of translational motion for repeatable probe positioning.

Bitakis, Kayla [Unlisted, US, IL]↗

Nonvolatile memory cells from hafnium zirconium oxide ferroelectric tunnel junctions using Nb and NbN electrodes

Ferroelectric tunnel junctions (FTJs) utilizing hafnium zirconium oxide (HZO) have attracted interest as non-volatile memory for microelectronics due to ease of integration into back-end-of-line (BEOL) complementary metal oxide semiconductor fabrication. This work examines asymmetric electrode NbN/HZO/Nb devices with 7 nm thick HZO as FTJs in a memory structure, with an output resistance that can be controlled by read and write voltages. The individual FTJs are measured to have a tunneling electroresistance of 10 during the read state without significant filament conduction formation and reasonable ferroelectric performance. Endurance and remanent polarizations of up to 10 5 cycles and 20 μC/cm 2 , respectively, are measured and are shown to be dependent on the cycling voltage. Electrical measurements demonstrate how magnitude of the write pulse can modulate the high state resistance and the read pulse influences both resistance values as well as separation of resistance states. Then, by using two opposite switching FTJ devices in series, a programmable nonvolatile resistor divider is demonstrated. Measurements of these two FTJ unit memory cells show wide applicability to a BEOL microfabrication process for a re-readable, rewritable, and nonvolatile memory cell.

42 ENGINEERING↗

DRIPS: Dynamic Rebalancing of Pipelined Streaming Applications on CGRAs

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit data streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across application kernels. However, emerging streaming applications at the edge (scientific instruments, sensor networks, network processing) perform much more than digital signal processing and often are data and input dependent. This leads to extremely variable kernel execution times, severely impacting the throughput of the entire pipeline if resources are only statically allocated. Therefore, in this paper, we propose DRIPS — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We present a unified compiler framework to facilitate the mapping of a given streaming application onto the DRIPS CGRA architecture. The experimental results show that DRIPS achieves an average throughput improvement of 1.46$\times$ across a set of representative applications over a statically partitioned solution. The additional area overhead to enable dynamic rebalancing consumes 16.34% of the entire area for a 5x5 CGRA prototype.

Tan, Cheng↗

DynPaC: Coarse-Grained, Dynamic, and Partially Reconfigurable Array for Streaming Applications

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across kernels. However, streaming applications often are data-dependent, leading to variable kernel execution times depending on the input data and impacting the throughput of the entire pipeline if resources are statically allocated. Therefore, in this paper, we discuss the design of DynPaC — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We discuss the required software and hardware components to manage partial dynamic reconfiguration. We demonstrate that by supporting partial dynamic reconfiguration, we can obtain an average speedup of 1.44X for a representative set of applications w.r.t. static partitioning, with a limited area overhead (6.4% of the entire chip).

Tan, Cheng↗

AURORA: Automated Refinement of Coarse-Grained Reconfigurable Accelerators

Coarse-grained reconfigurable arrays (CGRAs), loosely defined as arrays of functional units interconnected through a network-on-chip (NoC), provide higher flexibility than domain-specific ASIC accelerators while offering increased hardware efficiency with respect to fine-grained reconfigurable devices, such as Field Programmable Gate Arrays (FPGAs). Un-fortunately, designing a CGRA for a specific application domain involves enormous software/hardware engineering effort (e.g., designing the CGRA, map operations onto the CGRA, etc) and requires the exploration on a large design space (e.g., applying appropriate loop transformation on each application, specializing the reconfigurable processing elements of the CGRA, refining the network topology, deciding the size of the data memory, etc). Int his paper, we propose AURORA – a software/hardware co-design framework to automatically synthesize optimal CGRA given a set of applications of interest

Tan, Cheng↗

Companion Assisted Software Based Remote Attestation in SCADA Networks

Critical infrastructure such as power generation and water distribution systems have become a priority target in cyber warfare because of their recent computerization and introduction to the internet. As a result, Supervisory Control and Data Acquisition (SCADA) system security has become a hot topic in academic and industrial research. Among these topics, Remote Attestation is a security method intended to detect the presence of fileless malware in remote devices as they continue to operate. This allows for the detection of malware in the absence of long-term storage artifacts before symptoms of compromise begin to appear. In general, a trusted device (the verifier) makes a request for evidence of innocence from the untrusted device (the prover). In software-based schemes, the verifier can then measure the delay between its request and the prover’s response. If this delay is greater than the known computational time of the evidence gathering algorithm performed by the prover, then evidence may have been forged. Multi-hop networks often introduce too much network jitter to allow accurate measurement of prover response time, which limits the effectiveness of software based Remote Attestation in a real-world setting. In this work, we introduce a companion device that the verifier can trust to perform a subset of attestation, thereby removing any network jitter. This device is a Field Programmable Gate Array (FPGA) that is physically connected to the prover. We provide a communication protocol between the verifier, prover, and companion. To evaluate our scheme, we simulate it in a common SCADA network environment under normal and heavy traffic loads. Our simulations are performed in the discrete event network simulator NS-3, and we perform statistical analysis over our results to show that our scheme allows for tight timing constraints to be placed on the prover such that the verifier can more easily determine the validity of the evidence that it receives.

Johnson, William A.↗