Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “program processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

QuaSiMo: A composable library to program hybrid workflows for quantum simulation

Abstract A composable design scheme is presented for the development of hybrid quantum/classical algorithms and workflows for applications of quantum simulation. The proposed object‐oriented approach is based on constructing an expressive set of common data structures and methods that enables programming of a broad variety of complex hybrid quantum simulation applications. The abstract core of the scheme is distilled from the analysis of the current quantum simulation algorithms. Subsequently, it allows synthesis of new hybrid algorithms and workflows via the extension, specialisation, and dynamic customisation of the abstract core classes defined by the proposed design. The design scheme is implemented using the hardware‐agnostic programming language QCOR into the QuaSiMo library. To validate the implementation, the authors test and show its utility on commercial quantum processors from IBM and Rigetti, running some prototypical quantum simulations.

97 MATHEMATICS AND COMPUTING↗

Equilipy: a python package for calculating phase equilibria

The CALPHAD (CALculation of PHAse Diagram) approach (Nigel Saunders & Miodownik, 1998) provides predictions for thermodynamically stable phases in multicomponent-multiphase materials across a wide range of temperatures. Consequently, the CALPHAD calculations became an essential tool in materials and process design (Luo, 2015). Such design tasks frequently require navigating a high-dimensional space due to multiple components involved in the system. This increasing complexity demands high-throughput CALPHAD calculations, especially in the rapidly evolving field of alloy design. In response to the need, we developed Equilipy an open-source Python package designed for calculating phase equilibria of multicomponent-multiphase systems. Equilipy is specifically tailored for high-throughput CALPHAD calculations, offering parallel computations across multiple processors and nodes with the given NPT input conditions namely elemental compositions (N), pressure (P), and temperature (T). Equilipy utilizes the program structure and Gibbs energy functions from the Fortran-based program, Thermochimica (Piro et al., 2013), with incorporating a new Gibbs energy minimization algorithm. This algorithm, originally developed by Capitani and Brown in 1987 (Capitani & Brown, 1987), has been revised and implemented to enhance the stability and performance of calculations. The Fortran codes are precompiled and interfaced with Python via F2PY, ensuring high computation speed. Benchmark tests shown in Figure 1 demonstrate that Equilipy’s computation speed is comparable to those of established commercial software, TC-Python and PanPython. This result highlights its efficiency and potential applications in various scientific and industrial fields.

97 MATHEMATICS AND COMPUTING↗

Open Radiation Monitoring: Histogram Builder Module Design

The Open Radiation Monitoring Project seeks to develop and demonstrate a modular radiation detection architecture designed specifically for use in arms control treaty verification (ACTV) applications that will facilitate rapid development of trusted systems to meet the needs of potential future treaties. A modular architecture can be used to reduce more complex systems to a series of single purpose building blocks, thereby facilitating equipment inspection and in turn building trust in the equipment by all treaty parties. Furthermore, a modular architecture can be used to control data flow within the measurement system, reducing the risk of "hidden switches" and constraining the amount of sensitive information that could potentially be inadvertently leaked. This report details the first revision of a prototype circuit that will convert analog pulses directly into a histogrammed data set for further processing. The circuit was designed with both spectroscopy and multiplicity analysis in mind but can, in principle, be used to reduce any raw data stream into a histogram. The number of output channels is limited, and the histogram bin ranges are user configurable to allow for non-uniform and discontinuous bins, which makes it possible to restrict the information being passed down stream if desired. Pulse processing relies entirely on analog circuitry and non- programmable logic, which enables operation without the need for a central processor or other programmable control unit. The circuit remains untested under the Open Radiation Monitoring project due to the closure of the sponsoring program. However, further development and testing is scheduled to take place in support of a purpose-built trusted verification system development effort known as COGNIZANT, which demonstrates the potential benefit of developing a suite of modular trusted system components.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Controlling prediction functional blocks used by a branch predictor in a processor

An electronic device includes a processor, a branch predictor in the processor, and a predictor controller in the processor. The branch predictor includes multiple prediction functional blocks, each prediction functional block configured for generating predictions for control transfer instructions (CTIs) in program code based on respective prediction information, the branch predictor configured to select, from among predictions generated by the prediction functional blocks for each CTI, a selected prediction to be used for that CTI. The predictor controller keeps a record of prediction functional blocks from which the branch predictor previously selected predictions for CTIs. The predictor controller uses information from the record for controlling which prediction functional blocks are used by the branch predictor for generating predictions for CTIs.

97 MATHEMATICS AND COMPUTING↗

Modbus RTU for Embedded Cyber Secure Inverter Controller

The Modbus communication protocol is a widely adopted communication standard in industrial control systems. This communication protocol is known for being reliable and straightforward to implement while being versatile in terms of its operating parameters while supporting multiple formats over various hardware infrastructures and architectures. Many intelligent devices such as Programmable Logic Controllers (PLCs), Human-Machine Interfaces (HMIs), Internet-of-Things (IoT), and various Operational Technologies (OT) utilize Modbus for their communication systems. These types of systems must communicate with each other through a standardized and central communication process. To support the integration of these modular systems, a Field-Programmable Gate Array (FPGA) can act as an embedded central routing fabric for this communication to take place. Embedded systems are versatile enough to interface with various devices and systems to accomplish various goals. Additionally, embedded systems require relatively small physical designs to minimize the required resources to facilitate the intended application by providing low-level system access. This minimization of system resources goes hand in hand with reducing the financial cost of a proposed solution or system. As remotely collaborating researchers often use FPGAs to prototype designs that are required to have a method for data transmission among systems, it is imperative to provide a baseline standard for communications among devices and systems. A typical method of implementing the Modbus RTU communication protocol in an embedded environment is using integrated logic architectures within the FPGA called “Intellectual Property (IP) cores.” IP cores can be designed using integrated logic or circuit designs to function as an embedded processor. These IP cores can then perform the required computational actions to support the Modbus RTU communication protocol by utilizing high-level programming languages such as the C programming language. The hardware description language of Very High-Speed Integrated Circuit Hardware Description Language (VHDL) allows for the control of real hardware at the logic gate and signal level. These logic gates and signals can be designed and controlled to perform desired actions based on the system design. Programming an FPGA using VHDL allows an individual to access the lowest abstraction level of the system during FPGA development. This level of abstraction is referred to as the register-transfer level (RTL), which gives access to manipulating values and variables at the register level. This register-level manipulation provides precision over creating the logical circuit within the FPGA, thus minimizing the required code to perform desired operations. The Modbus RTU communication protocol can be implemented within an FPGA using VHDL programming to establish a standardized and embedded serial communication pathway. This implementation provides a standardized communication protocol to streamline research efforts among researchers, thus increasing the efficiency of research efforts. Additionally, this Modbus RTU implementation requires fewer resources when compared to typical communication protocol implementations that utilize an IP core, reducing the hardware requirement for effective research efforts.

communication↗

Risk-informed Graded Approach for Reliability and Performance Assessment of Sensor and Instrumentation Systems within Advanced Condition Monitoring Technologies

Advanced condition monitoring (ACM) technologies, such as digital twins, are innovative strategies designed to provide real-time health insights, including the remaining useful life of components. The primary goal of ACM is to predict and alert operators to potential functional failures before they occur. ACM systems achieve this by integrating predictive models with various sensor instrumentation, analog-to-digital converters, data warehouses, and data pre-processors. These sensor and instrumentation systems (SIS) are essential for forming a comprehensive understanding of component conditions and ensuring the predictive success of ACM programs. Introducing new technologies like ACM involves varying degrees of risk that can impact plant reliability. Therefore, risk mitigation should be commensurate with the performance and reliability of the developed technology, following a risk-informed graded approach (RIGA). Establishing a RIGA process requires a clear understanding of the hazards and reliability of all subsystems, including their interdependencies and potential impacts on the overall system. Given the critical role of SIS in ACM, this work reviews hazard identification and reliability quantification methods for SIS. It also considers these methods' implications when developing a RIGA process for ACM.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Toward coherent quantum computation of scattering amplitudes with a measurement-based photonic quantum processor

In recent years, applications of quantum simulation have been developed to study the properties of strongly interacting theories. This has been driven by two factors: on the one hand, needs from theorists to have access to physical observables that are prohibitively difficult to study using classical computing; on the other hand, quantum hardware becoming increasingly reliable and scalable to larger systems. In this work, we discuss the feasibility of using quantum optical simulation for studying scattering observables that are presently inaccessible via lattice QCD and are at the core of the experimental program at Jefferson Laboratory, the future Electron-Ion Collider, and other accelerator facilities. We show that recent progress in measurement-based photonic quantum computing can be leveraged to provide deterministic generation of required exotic gates and implementation in a single photonic quantum processor. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Design Considerations for GPU-based Mixed Integer Programming on Parallel Computing Platforms

Mixed Integer Programming (MIP) is a powerful abstraction in combinatorial optimization that finds real-life application across many significant sectors. The recent proliferation of graphical processing unit (GPU)-based accelerated computing architectures in large-scale parallel computing or supercomputing presents new opportunities as well as challenges in the advancement of MIP solver technology to effectively use the new accelerated computing platforms and scale to large parallel systems. Here, we recount the conventional processor-based strategies and focus on configurations where the most promising intersection lies between parallel MIP solver approaches and the specific strengths of accelerated parallel platforms. We note that the best potential lies in solving problems whose individual matrix sizes (of the linear program relaxation) fit entirely within one accelerator's memory and whose branch-and-bound (or branch-and-cut) trees cannot be fully contained within a small number of computational nodes. Additionally, we identify ideal features of computational linear algebra support on GPU accelerators that would help advance this direction of scalable parallel solution of MIP problems on GPU-based accelerated computing architectures.

Perumalla, Kalyan↗

Advancements in Multiphysics Microdepletion Analysis of an eVinci TM -like Microreactor Leveraging OpenMC-CRAB Workflow

Nuclear microreactors (MRs) are a class of nuclear reactor technology, characterized by reduced dimensions, modular design, and reduced power output in contrast to conventional Light Water Reactors (LWRs). MRs are proposed for supplying electricity and eventual process heat to remote locations, such as military installations and disaster-affected areas. Current research work sponsored by the US Department of Energy Microreactor Program (MRP) is devoted to the development of novel modeling and simulation tools to better support MR vendors and regulatory bodies. Notably, the NRC is projected to utilize the CRAB multiphysics software driver for executing both design and beyond-design-basis accident analyses. Furthermore, the NRC has been utilizing the MELCOR code to calculate mechanistic source terms during accidents. Since MELCOR relies on isotopic inventory and reactor temperature/power profiles under accident conditions, which theoretically can be derived from CRAB, the goal is to establish a comprehensive CRAB-MELCOR computational framework. Past work was focused on testing and demonstrating CRAB's capability to generate results that can be used to inform mechanistic source term calculations in MELCOR. In particular, a computational workflow leveraging OpenMC-generated microscopic cross sections and CRAB was first applied to perform multiphysics microscopic depletion calculation followed by an accident scenario for a stylized microreactor problem. In fiscal year 2024, the research work has been focused on applying the OpenMC-CRAB workflow, which was first tested in fiscal year 2023, to a realistic 3D heat-pipe cooled MR problem representative of the eVinci TM design. The latter computational problem was developed with inputs from WEC to conserve selected neutronic and thermal characteristics of the eVinci TM design without releasing proprietary data. The results of this simulation, encompassing isotopic inventory, power density distribution, and kinetic parameters, will inform both MELCOR and the WEC-developed FATE code for mechanistic source terms calculations. The results from the two codes will then be compared for code verification purposes. This report contains the design characteristics of the realist heat pipe cooled microreactor developed as a use-case for the verification exercise, and the current results for the multiphysics microscopic depletion performed with the OpenMC-CRAB workflow. The results include eigenvalue as a function of time, power distribution at EOL, in addition to nuclides inventory's time evolution and spatial distribution. Finally, we report improvements to the workflow efficiency achieved through a collaboration with the NEAMS programs. Through this collaborative effort, we were able to strongly decrease the computational time for the multiphysics microdepletion calculation (i.e., from 17.4 hours to 5.7 hours on 280 processors) in addition to simplifying the interface to generate isotopics spatial distribution utilizable by FATE and MELCOR. Future work, including the improvement of the current microscopic cross-sections' library and the simulation of an accident scenario at EOL, is also discussed.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

IceNet for FireBox - A Berkeley Warehouse-Scale Computer

Berkeley’s FireBox is a next-generation warehouse-scale computer (WSC) that utilizes the energy-efficiency and bandwidth density of integrated silicon-photonic interconnects to enable a new high-bandwidth and low-latency network fabric connecting thousands of compute nodes to petabytes of DRAM and Flash storage. The high bandwidth, low latency and high connectivity of FireBox’s WSC network fabric (IceNet) will enable dramatic improvements in the overall system energy efficiency enabling fine-grain power control on system resources (processors, links and memory/storage components). IceNet is a special 3-stage photonic Clos network architected to achieve ultra-low-latency connectivity between processor nodes and memory, drastically cutting down on the energy wasted in resource idling (processors and memory stalled due to pending network requests). This is achieved by integration of the first and last switch stages into processor/memory hub clients and by heavy over-provisioning of the high-radix middle switches (FlareSwitches). A key hardware component developed in this program is an active laser power management photonic integrated circuits called LightSpark. It interacts with the FlareSwitch and provides laser power to a subset of occupied switch ports, increasing the utilization of laser light in the photonic network by an order of magnitude. In addition to guiding the laser power where it is needed, the laser-power management module enables both wavelength and laser redundancy, significantly increasing the robustness of the system. The goal of the IceNet fabric is to enable communication between 1000s of processor nodes and PBs of memory/storage with <100ns latency, <10pJ/b wall-plug energy cost at multiple Pb/s of available connectivity bandwidth. These metrics represent two-orders of magnitude improvement with respect to the status of current data-center technology.

42 ENGINEERING↗

Many-Body Physics in the NISQ Era: Quantum Programming a Discrete Time Crystal

Recent progress in the realm of noisy intermediate-scale quantum (NISQ) devices represents an exciting opportunity for many-body physics by introducing new laboratory platforms with unprecedented control and measurement capabilities. We explore the implications of NISQ platforms for many-body physics in a practical sense: we ask which physical phenomena, in the domain of quantum statistical mechanics, they may realize more readily than traditional experimental platforms. While a universal quantum computer can simulate any system, the eponymous noise inherent to NISQ devices practically favors certain simulation tasks over others in the near term. As a particularly well-suited target, we identify discrete time crystals (DTCs), novel nonequilibrium states of matter that break time translation symmetry. These can only be realized in the intrinsically out-of-equilibrium setting of periodically driven quantum systems stabilized by disorder-induced many-body localization. While promising precursors of the DTC have been observed across a variety of experimental platforms—ranging from trapped ions to nitrogen-vacancy centers to NMR crystals—none have all the necessary ingredients for realizing a fully fledged incarnation of this phase, and for detecting its signature long-range spatiotemporal order. We show that a new generation of quantum simulators can be programmed to realize the DTC phase and to experimentally detect its dynamical properties, a task requiring extensive capabilities for programmability, initialization, and readout. Specifically, the architecture of Google’s Sycamore processor is a remarkably close match for the task at hand. We also discuss the effects of environmental decoherence, and how they can be distinguished from ‘internal’ decoherence coming from closed-system thermalization dynamics. Already with existing technology and noise levels, we find that DTC spatiotemporal order would be observable over hundreds of periods, with parametric improvements to come as the hardware advances.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Field Programmable Gate Array Data Capture for Control Systems

Some Industrial Control Systems (ICS) networks are based on protocols such as Serial and Industrial Ethernet. These protocols currently have no existing cybersecurity monitoring tools, leaving a large gap in the cyber defense of critical infrastructure. In order to analyze such ICS traffic, it is first necessary to implement methods of capturing the ICS data. Whereas traditional methods of analyzing data would use microprocessors, the nature of high-speed analog data can be difficult to implement on such a versatile processor, as they are rather inefficient for doing a single task. Whereas Field Programmable Gate Arrays (FPGAs) provide an adequate tool in analyzing high speed data, as despite the lack of program versatility, Programmable Logic can implement a solution with minimal clock cycles, allowing time for each new packet of data to be captured before a new data sample is taken.

42 ENGINEERING↗

System, method, and computer program for creating an internal conforming structure

A system for creating an internal formation of a tubular structure having an inner surface via additive manufacturing. The system broadly includes a computer modeling system and an additive manufacturing system. The computer modeling system may include a processor for generating a lattice cellular component via computer-aided design software according to inputs received from a user. The processor may also generate an internal formation lattice structure based on the lattice cellular component and modify the lattice structure to follow and/or conform to the curvature of the inner surface of the outer wall of the tubular structure. The additive manufacturing system may be configured to produce the lattice structure and the tubular structure via additive manufacturing material deposited layer by layer according to the lattice structure.

42 ENGINEERING↗

4-Clique network minor embedding for quantum annealers

Quantum annealing is a quantum algorithm for computing solutions to combinatorial optimization problems. This study proposes a method for minor embedding optimization problems onto sparse quantum annealing hardware graphs called 4-clique network minor embedding. This method is in contrast to the standard minor embedding technique of using a path of linearly connected qubits in order to represent a logical variable state. The 4-clique minor embedding is possible on Pegasus graph connectivity, which is the native hardware graph for some of the current D-Wave quantum annealers. The Pegasus hardware graph contains many cliques of size 4, making it possible to form a graph composed entirely of paths of connected 4-cliques on which a problem can be minor-embedded. The 4-clique chains come at the cost of additional qubit usage on the hardware graph, but they allow for stronger coupling within each chain, thereby increasing chain integrity, reducing chain breaks, and allow for greater usage of the available energy scale for programming logical problem coefficients on current quantum annealers. The 4-clique minor embedding technique is compared with the standard linear path minor embedding with experiments on two D-Wave quantum annealing processors with Pegasus hardware graphs. We show proof-of-concept experiments where the 4-clique minor embeddings can use weak chain strengths while successfully carrying out the computation of minimizing random all-to-all spin glass problem instances. Published by the American Physical Society 2024

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A MLIR Dialect for Quantum Assembly Languages

We demonstrate the utility of the Multi-Level Intermediate Representation (MLIR) for quantum computing. Specifically, we extend MLIR with a new quantum dialect that enables the expression and compilation of common quantum assembly languages. The true utility of this dialect is in its ability to be lowered to the LLVM intermediate representation (IR) in a manner that is adherent to the quantum intermediate representation (QIR) specification recently proposed by Microsoft. We leverage a qcor-enabled implementation of the QIR quantum runtime API to enable a retargetable (quantum hardware agnostic) compiler workflow mapping quantum languages to hybrid quantum-classical binary executables and object code. We evaluate and demonstrate this novel compiler workflow with quantum programs written in OpenQASM 2.0. We provide concrete examples detailing the generation of MLIR from OpenQASM source files, the lowering process from MLIR to LLVM IR, and ultimately the generation of executable binaries targeting available quantum processors.

Mccaskey, Alex↗

Clang UPC Compiler (Clang UPC) v3.9.1-1

Clang Unified Parallel C (Clang UPC) provides a compilation and execution environment for programs written in the UPC (Unified Parallel C) language. The Clang UPC compiler extends the capabilities of the Clang LLVM C compiler to comply with the UPC Language Specification version 1.3. It includes support for UPC collectives and a configurable pointer-to-shared representation. The compiler generates programs that run on a wide variety of systems ranging from workstations to leadership-class supercomputers, in conjunction with the Berkeley UPC runtime and GASNet communication system. This compiler generates assembly code / object code directly for Intel processors and IBM PowerPC.

Hargrove, Paul↗

Performance of an Astrophysical Radiation Hydrodynamics Code under Scalable Vector Extension Optimization

We present results of a performance study of an astrophysical radiation hydrodynamics code, V2D, on the Arm-based A64FX processor developed by Fujitsu. The code solves sparse linear systems, a task for which the A64FX architecture should be well suited. Here, we performed the performance analysis study on Ookami, an Apollo 80 platform utilizing the A64FX processor. We explored several compilers and performance anal-ysis packages and found the code did not perform as expected under scalable vector extension optimization, suggesting that a “deeper dive” into analyzing the code is worthwhile. However, a simple driver program that exercised basic sparse linear algebra routines used by V2D did show significant speedup with the use of the scalable vector extension optimization. We present the initial results from the study which used V2D on a relatively simple test problem that emphasized the repeated solution of sparse linear systems.

79 ASTRONOMY AND ASTROPHYSICS↗