Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “program processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Optimization control technology for building energy conservation

A simulation processor generates and stores a simulation model based on conditions associated with a physical structure, such as a building. A neural network processor implements a neural network, having an input layer coupled to receive sensor data from the structure and having an output layer coupled to supply control signals to the at least one electrically operable environmental control device. The neural network is trained using the simulation model. A particle swarm optimization processor programmed to receive the simulation results and perform particle swarm optimization, ascertains optimal parameters for controlling the at least one electrically operable environmental control device and supplies these optimal parameters to the neural network processor. The neural network processor uses the optimal parameters supplied by the particle swarm optimization processor to further train the neural network.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Retargetable Optimizing Compilers for Quantum Accelerators via a Multi-Level Intermediate Representation

In this work, we present a multi-level quantum-classical intermediate representation (IR) that enables an optimizing, retargetable compiler for available quantum languages. Our work builds upon the Multi-level Intermediate Representation (MLIR) framework and leverages its unique progressive lowering capabilities to map quantum languages to the LLVM machine-level IR. We provide both quantum and classical optimizations via the MLIR pattern rewriting sub-system and standard LLVM optimization passes, and demonstrate the programmability, compilation, and execution of our approach via standard benchmarks and test cases. In comparison to other standalone language and compiler efforts available today, our work results in compile times that are 1000x faster than standard Pythonic approaches, and 5-10x faster than comparative standalone quantum language compilers. Our compiler provides quantum resource optimizations via standard programming patterns that result in a 10x reduction in entangling operations, a common source of program noise. We see this work as a vehicle for rapid quantum compiler prototyping.

43 PARTICLE ACCELERATORS↗

Systems, methods, and media for high dynamic range imaging using single-photon and conventional image sensor data

In accordance with some embodiments, systems, methods, and media for high dynamic range imaging using single-photon and conventional image sensor data are provided. In some embodiments, the system comprises: first detectors configured to detect a level of photons proportional to incident photon flux; second detectors configured to detect arrival of individual photons; a processor programmed to: receive, from the first detectors, first values indicative of photon flux from a scene with a first resolution; receive, from the second detectors, second values indicative of photon flux from the scene with a lower resolution; provide a first encoder of a trained machine learning model first flux values based on the first values, provide the second encoder of the model second flux values; receive, as output, values indicative of photon flux from the scene; and generate a high dynamic range image based on the third plurality of values.

Gutierrez Barragan, Felipe↗

Measuring the capabilities of quantum computers

Quantum computers can now run interesting programs, but each processor’s capability—the set of programs that it can run successfully—is limited by hardware errors. These errors can be complicated, making it difficult to accurately predict a processor’s capability. Benchmarks can be used to measure capability directly, but current benchmarks have limited flexibility and scale poorly to many-qubit processors. We show how to construct scalable, efficiently verifiable benchmarks based on any program by using a technique that we call circuit mirroring. With it, we construct two flexible, scalable volumetric benchmarks based on randomized and periodically ordered programs. We use these benchmarks to map out the capabilities of twelve publicly available processors, and to measure the impact of program structure on each one. We find that standard error metrics are poor predictors of whether a program will run successfully on today’s hardware, and that current processors vary widely in their sensitivity to program structure.

97 MATHEMATICS AND COMPUTING↗

Techniques for recovering from errors when executing software applications on parallel processors

In various embodiments, a software program uses hardware features of a parallel processor to checkpoint a context associated with an execution of a software application on the parallel processor. The software program uses a preemption feature of the parallel processor to cause the parallel processor to stop executing instructions in accordance with the context. The software program then causes the parallel processor to collect state data associated with the context. After generating a checkpoint based on the state data, the software program causes the parallel processor to resume executing instructions in accordance with the context.

Hukerikar, Saurabh↗

Characterization of Quantum Frequency Processors

Frequency-bin qubits possess unique synergies with wavelength-multiplexed lightwave communications, suggesting valuable opportunities for quantum networking with the existing fiber-optic infrastructure. Although the coherent manipulation of frequency-bin states requires highly controllable multi-spectral-mode interference, the quantum frequency processor (QFP) provides a scalable path for gate synthesis leveraging standard telecom components. Here, we summarize the state of the art in experimental QFP characterization. Distinguishing between physically motivated “open box” approaches that treat the QFP as a multiport interferometer, and “black box” approaches that view the QFP as a general quantum operation, we highlight the assumptions and results of multiple techniques, including quantum process tomography of a tunable beamsplitter—to our knowledge the first full process tomography of any frequency-bin operation. Our findings should inform future characterization efforts as the QFP increasingly moves beyond proof-of-principle tabletop demonstrations toward integrated devices and deployed quantum networking experiments.

42 ENGINEERING↗

Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep Learning

In this work, we present FastConv, a template-based code auto-generation open source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming convolution layers of convolutional neural networks. ARM CPUs cover a wide range designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. The experiments with layer-wise evaluation on the VGG--16 model confirms a 1.25x performance gains is got by tuning the Winograd library. Integrated comparison results shows 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, Arm NN, and FeatherCNN on the Kunpeng 920 beside few cases. Furthermore, problem size performance portability experiments with various convolution shapes shows that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920 . CPU performance portability evaluation on the VGG--16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x is observed on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively.

97 MATHEMATICS AND COMPUTING↗

Systems and methods for enhanced power system model validation

A system for enhanced power system model validation is provided. The system includes a computing device including at least one processor in communication with at least one memory device. The at least one processor is programmed to store a plurality of models for a plurality of devices and a plurality of input files associated with the plurality of models, receive, from a user, a selection of model of the plurality of models to simulate, retrieve one or more input files of the plurality of input files, perform a model validity check on the selected model, if the selected model passed the model validity check, perform a model calibration on the selected model, and if the selected model passed the model calibration, perform a post evaluation on the selected model.

Wang, Honggang↗

Systems and methods for single-axis tracking via sky imaging and machine leanring comprising a neural network to determine an angular position of a photovoltaic power system

A system and method is disclosed for solar tracking and controlling an angular position of a photovoltaic power system. The solar tracking system includes an imaging device for capturing images of the sky; a solar position data generating module; and a control system comprising a neural network. The neural network has multiple convolutional layers to generate a first output associated with the images, and a solar position data module. A first dense layer module receives the solar position data and generates a second output. A second dense layer module receives the first output and the second output and generates a concatenated data sequence. A processor is programmed to generate a multi-planar irradiance signal (MPIS) in response to the concatenated data sequence, and determine an angular position of the PV power system and adjust the angular position in response to an angle of maximum irradiance.

Stein, Joshua↗

Implementation and Demonstration of P4 Software for Improving ICS Protocol Visibility and Control [Slides]

No prior enabling funded work applicable to this proposal. Programming Protocol-independent Packet Processors (P4) is an open source, domain-specific programming language for network switching devices. P4 complements traditional Software Defined Networking (SDN) which is primarily concerned with the management of packets (e.g. routing/dropping decisions) rather than how each packet is processed. The introduction of P4 provided new capabilities (e.g. firewall, load balancing, enhanced security) but has primarily been deployed in data centers. This effort investigates ways to expand P4 into other niches such as ICS networks.

97 MATHEMATICS AND COMPUTING↗

Equilipy

There is high demand for high-throughput calculations of phase equilibria based on the CALPHAD method (CALculation of PHAse Diagram). High-throughput calculations are possible through HPC, currently by a commercial program called Thermo-Calc. However, the number of nodes/processors are limited to the number of purchased license (16 processors per license). This program provides a toolkit for high-throughput calculations of phase equilibria by the CALPHAD method. The program is prepared for Python environment, so that it can be easily installed and used together with other open-source programs. The program can be run in supercomputer.

Kwon, Sunyong [Oak Ridge National Laboratory (ORNL↗

Method and apparatus for real time, in situ sensing and characterization of roughness, geometrical shapes, geometrical structures, composition, defects, and temperature in three-dimensional manufacturing systems

Methods and apparatuses for manufacturing are disclosed, including (a) providing an apparatus having: a laser; scanner; powder injection system; powder spreading system; dichroic filter; imager-and-processor; and computer; (b) programming the computer with specifications of a sample; (c) using the computer to set initial parameters based on the sample specifications; (d) adjusting a stage to position the sample; (e) focusing and scanning electromagnetic radiation onto the sample while powder is concurrently injected onto the sample in order to deposit a layer; (f) capturing two-dimensional images of the sample and probing the sample to determine whether the deposited layer was manufactured per the specifications; (g) use the computer to adjust the three-dimensional manufacturing parameters based on the determination made in step (f) prior to additively manufacturing a subsequent layer or making repairs; and (h) repeating steps (d), (e), (f), and (g) until the manufacture is complete. Other embodiments are described and claimed.

Liu, Jian↗

HIPLZ: Enabling performance portability for exascale systems

While heterogeneous computing has emerged as a dominant trend in current and future High-Performance Computing (HPC) systems, it is also widely recognized that this shift has led to increased software complexity due to a proliferation of programming systems for different heterogeneous processors. One such example is the Heterogeneous-Compute Interface for Portability from AMD (HIP ), which is composed of a C Runtime API and C++ Kernel Language. Many HPC applications will likely use HIP on future exascale systems (e.g., Frontier and El Capitan), but HIP currently only targets AMD and NVIDIA processors. This limitation creates challenges for users who would also like to run their applications on exascale systems based on other architectures (e.g., Aurora, which is based on Intel hardware) that are currently not targeted by HIP . In this paper, we introduce the design and implementation of HIPLZ , a compiler and runtime system that uses the Intel Level Zero API to support HIP on Intel GPU architectures. We discuss the design of HIPLZ , derived from HIPCL (an implementation of HIP on top of OpenCL ), and portability issues that occur from using the Level Zero runtime as a backend. We evaluate our implementation by running several performance benchmarks and mini-apps written in HIP on Intel architectures using HIPLZ . Our results show that this approach provides competitive performance relative to Intel's OpenCL implementations on Intel Gen9 and UHD Graphics 770 GPUs, while providing good coverage of features needed by HPC applications. Overall, this approach is a promising demonstration of enabling performance portability for exascale systems.

97 MATHEMATICS AND COMPUTING↗

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

Cache management based on access type priority

Systems, apparatuses, and methods for cache management based on access type priority are disclosed. A system includes at least a processor and a cache. During a program execution phase, certain access types are more likely to cause demand hits in the cache than others. Demand hits are load and store hits to the cache. A run-time profiling mechanism is employed to find which access types are more likely to cause demand hits. Based on the profiling results, the cache lines that will likely be accessed in the future are retained based on their most recent access type. The goal is to increase demand hits and thereby improve system performance. An efficient cache replacement policy can potentially reduce redundant data movement, thereby improving system performance and reducing energy consumption.

Yin, Jieming↗

Controlling accesses to a branch prediction unit for sequences of fetch groups

An electronic device is described that handles control transfer instructions (CTIs) when executing instructions in program code. The electronic device has a processor that includes a branch prediction functional block and a sequential fetch logic functional block. The sequential fetch logic functional block determines, based on a record associated with a CTI, that a specified number of fetch groups of instructions that were previously determined to include no CTIs are to be fetched for execution in sequence following the CTI. When each of the specified number of fetch groups is fetched and prepared for execution, the sequential fetch logic prevents corresponding accesses of the branch prediction functional block for acquiring branch prediction information for instructions in that fetch group.

Yalavarti, Adithya↗

EJFAT: Towards Intelligent Compute Destination Load Balancing

To handle increased data flow, Jefferson Lab (JLab) is partnering with ESnet for development of an AI/ML directed compute work Load Balancer (LB) of UDP streamed data. The LB is FPGA based featuring dynamically configurable, low latency and high throughput destination address switching. The LB provides integration of edge and core computing to support JLab experimental programs, the Electron-Ion Collider, as well as data centers of the future. In the ESnet/JLab FPGA Accelerated Transport (EJFAT) initiative, the function of the LB Data Plane (DP) is to redirect data streams to selectable (but unknown to sender) destination hosts based on current worload and within that host to destination ports as a function of sub- stream id. This effects hierarchical scaling, first across compute machines for processing over a series of events and second, across ports so different data source sub-streams may be assigned to different processors for further parallelization. The LB Control Plane (CP) programs the DP using compute farm telemetry to direct and balance workloads across a compute cluster as the operating conditions require. While Proportional/Integrative/Derivative (PID) controllers are often seen in similar applications, here we investigate the feasibility of a Reinforcement Learning (RL) based schedule manager running in the CP to provide dynamic updates to the DP scheduling policy.

Lawrence, David↗