Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Hardware generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Experiences readying applications for Exascale

The advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems.

Gottiparthi, Kalyan↗

Generating clock signals for a cycle accurate, cycle reproducible FPGA based hardware accelerator

A method, system and computer program product are disclosed for generating clock signals for a cycle accurate FPGA based hardware accelerator used to simulate operations of a device-under-test (DUT). In one embodiment, the DUT includes multiple device clocks generating multiple device clock signals at multiple frequencies and at a defined frequency ratio; and the FPG hardware accelerator includes multiple accelerator clocks generating multiple accelerator clock signals to operate the FPGA hardware accelerator to simulate the operations of the DUT. In one embodiment, operations of the DUT are mapped to the FPGA hardware accelerator, and the accelerator clock signals are generated at multiple frequencies and at the defined frequency ratio of the frequencies of the multiple device clocks, to maintain cycle accuracy between the DUT and the FPGA hardware accelerator. In an embodiment, the FPGA hardware accelerator may be used to control the frequencies of the multiple device clocks.

Asaad, Sameh W.↗

SODA Synthesizer: an Open-source, Multi-level, Modular, Extensible Compiler from High-level Frameworks to Silicon

The SODA Synthesizer is an open-source modular, end-to-end hardware compiler framework. The SODA frontend, developed in MLIR, performs system-level design, code partitioning, and high-level optimizations to prepare the specifications for the hardware synthesis. The backend is based on a state-of-the-art high-level synthesis tool, and generates the final hardware design. The backend can interface with logic synthesis tools for field programmable gate arrays or with commercial and open-source logic synthesis tools for application-specific integrated circuits. We discuss the opportunities and challenges in integrating with commercial and open-source tools both at the frontend and backend, and the unique opportunities that an open-source hardware design ecosystem provides.

Bohm Agostini, Nicolas↗

Deployment and validation of predictive 6-dimensional beam diagnostics through generative reconstruction with standard accelerator elements

Understanding the 6-dimensional phase space distribution of particle beams is essential for optimizing accelerator performance. Conventional diagnostics such as use of transverse deflecting cavities offer detailed characterization but require dedicated hardware and space. Generative phase space reconstruction (GPSR) methods have shown promise in beam diagnostics, yet prior implementations still rely on such components. Here we present the first experimental implementation and validation of the GPSR methodology, realized by the use of standard accelerator elements including accelerating cavities and dipole magnets, to achieve complete 6-dimensional phase space reconstruction. Through simulations and experiments at the Pohang Accelerator Laboratory X-ray Free Electron Laser facility, we successfully reconstruct complex, nonlinear beam structures. Furthermore, we validate the methodology by predicting independent downstream measurements excluded from training, revealing the reconstruction closely resembling ground truth. This advancement establishes a pathway for predictive diagnostics across beamline segments while reducing hardware requirements and expanding applicability to various accelerator facilities.

Kim, Seongyeol [Pohang Univ. of Science and Techno↗

Towards Precision-Aware Fault Tolerance Approaches for Mixed-Precision Applications

Graphics Processing Units (GPUs), the dominantly adopted accelerators in HPC systems, are susceptible to transient hardware fault. New generation of GPUs feature mixed-precision architectures such as NVIDIA Tensor Cores to accelerate matrix multiplications. While being widely adapted, how would they behave under transient hardware faults remain unclear. In this study, we conduct a large-scale fault injection experiments on GEMM kernels implemented with different floating-point data types on the V100 and A100 Tensor Cores, and show distinct error resilience characteristics for the GEMMS with different formats. In the future, we plan to explore this space by building precision-aware floating-point fault tolerance techniques for applications such as DNNs that exercise low-precision computations.

Fang, Bo↗

S&TR September 2025: Computing Grand Challenge Turns 20

Livermore’s Computing Grand Challenge Program enters its 20th year with more unclassified high-performance computing (HPC) power than ever before. This unique, peer-reviewed competition awards HPC allocations on top supercomputers to multidisciplinary teams with high-impact projects. The Grand Challenge encourages researchers to innovate, pushes scientific discovery to new heights, improves the Laboratory’s HPC capabilities, and extends HPC accessibility to collaborators. Awardees must adapt to successive generations of HPC hardware and learn to run simulations at scale. The feature article spotlights three Grand Challenge teams whose research broke new ground in key scientific pursuits—the essence of dark matter, explosion-generated seismic waves, and protein interactions linked to cancer—while underscoring the importance of academic partnerships and considering the program’s future.

07 ISOTOPE AND RADIATION SOURCES↗

Generation of Tunable Stochastic Sequences Using the Insulator–Metal Transition

Probabilistic computing is a paradigm in which data are not represented by stable bits, but rather by the probability of a metastable bit to be in a particular state. The development of this technology has been hindered by the availability of hardware capable of generating stochastic and tunable sequences of “1s” and “0s”. The options are currently limited to complex CMOS circuitry and, recently, magnetic tunnel junctions. Here, we demonstrate that metal–insulator transitions can also be used for this purpose. We use an electrical pump/probe protocol and take advantage of the stochastic relaxation dynamics in VO 2 to induce random metallization events. A simple latch circuit converts the metallization sequence into a random stream of 1s and 0s. The resetting pulse in between probes decorrelates successive events, providing a true stochastic digital sequence.

97 MATHEMATICS AND COMPUTING↗

Experiences Detecting Defective Hardware in Exascale Supercomputers

In May 2022, the newest supercomputer to top the TOP 500 list was Frontier at Oak Ridge National Laboratory, demonstrating the capability of computing more than 1.1 quintillion (1018) floating-point calculations every second. Driving this ground-breaking rate of computing is Frontier’s more than 37,000 graphics processing units (GPUs) and 9,408 central processing units (CPUs). In total, Frontier contains more than 60 million parts. At this scale, the smallest margin of error may generate hundreds of hardware errors across the system. These errors are capable of directly hindering world-class science performed on Frontier if not found. In this work, we describe and evaluate two strategies for finding hardware-level faults in Frontier’s 9,408 compute nodes. There are two strategies developed: the first uses the Slurm scheduler to scavenge available compute time to run the node screen, the second builds upon the lessons learned in the first strategy and enforces a weekly screen of each node. Using June 2023 as a case study, we find that the first scheduling strategy consumed more than ten times the resources as the second scheduling strategy, but successfully detected five hardware defects in Frontier. We summarize the lessons learned while developing and running a node screen on the world’s first exascale supercomputer.

Hagerty, Nick↗

Quantum Random Number Generator (QRNG)

The Los Alamos Quantum Random Number Generator (QRNG) is a hardware-based, high-performance Random Number Generator capable of generating 200 Mbit/s or more of true random numbers. Like flipping a coin, it is very much random and essential for information security like encrypting data on the internet, checking email, or purchasing something from an online vendor. The device harvests entropy from fluctuations in an optical source that arise from quantum mechanical properties of light. Qrypt, Inc., a company launched in 2017, began making strategic investments and developing partnerships to advance cutting-edge quantum hardware solutions. One of those key investments was licensing QRNG from Los Alamos and subsequently collaborating with advanced quantum materials and technology researcher Dr. Raymond Newell to create high-quality random keys at scale.

97 MATHEMATICS AND COMPUTING↗

All-Electric Nonassociative Learning in Nickel Oxide

Habituation and sensitization represent nonassociative learning mechanisms in both non-neural and neural organisms. They are essential for a range of functions from survival to adaptation in dynamic environments. Design of hardware for neuroinspired computing strives to emulate such features driven by electric bias and can also be incorporated into neural network algorithms. Herein, cellular-like learning in oxygen-deficient NiO x devices is demonstrated. Both habituation learning and sensitization response can be achieved in a single device by simply controlling the magnitude of the electric field. Spontaneous memory relaxations and dynamic redistribution of oxygen vacancies under electric bias enable such learning behavior of NiO x under sequential training. These characteristics in simple device arrays are implemented to learn alphabets as well as demonstrate simulated algorithmic use cases in digit recognition. Transition metal oxides with carefully prepared defect concentrations can be highly sensitive to electronic structure perturbations under moderate electrical stimulus and serve as building blocks for next-generation neuroinspired computing hardware.

36 MATERIALS SCIENCE↗

Electronic Bottleneck Suppression in Next‐Generation Networks with Integrated Photonic Digital‐to‐Analog Converters

Digital‐to‐analog converters (DAC) are indispensable functional units in signal processing instrumentation and wide‐band telecommunication links for both civil and military applications. As photonic systems are capable of high data throughput and low latency, an increasingly found system limitation stems from the required domain crossing such as digital to analog and electronic to optical. A photonic DAC implementation, in contrast, enables a seamless signal conversion with respect to both energy efficiency and short signal delay, often requiring bulky discrete optical components and electric–optic transformation, hence introducing inefficiencies. Herein, a novel coherent parallel photonic DAC concept along with a 4‐bit experimental prototype capable of performing this DAC without optic–electric–optic domain crossing is introduced. This new paradigm guarantees a linear intensity weighting among bits when operating at high sampling rates (50 GHz), featuring an exceptional sampling efficiency (> 100 GS ) and small footprint (≈1 mm 2 ) in an 8‐bit implementation. Importantly, this photonic DAC enables seamless interfaces of next‐generation data processing hardware with high relevance in data centers, task‐specific compute accelerators such as neuromorphic engines, and network edge processing applications.

Meng, Jiawei↗

AutoLabs: cognitive multi-agent systems with self-correction for autonomous chemical experimentation

The automation of chemical research through self-driving laboratories (SDLs) promises to accelerate scientific discovery, yet the reliability and granular performance of the underlying AI agents remain critical, under-examined challenges. In this work, we introduce AutoLabs, a self-correcting, multi-agent architecture designed to autonomously translate natural-language instructions into executable protocols for a high-throughput liquid handler. The system engages users in dialogue, decomposes experimental goals into discrete tasks for specialized agents, performs tool-assisted stoichiometric calculations, and iteratively self-corrects its output before generating a hardware-ready file. We present a comprehensive evaluation framework featuring five benchmark experiments of increasing complexity, from simple sample preparation to multi-plate timed syntheses. Through a systematic ablation study of 20 agent configurations, we assess the impact of reasoning capacity, architectural design (single- vs. multi-agent), tool use, and self-correction mechanisms. Our results demonstrate that agent reasoning capacity is the most critical factor for success, reducing quantitative errors in chemical amounts (nRMSE) by over 85% in complex tasks. When combined with a multi-agent architecture and iterative self-correction, AutoLabs approaches expert-authored reference procedures on the benchmark (F1-score > 0.89) on challenging multi-plate syntheses. These findings establish a clear blueprint for developing robust and trustworthy AI partners for autonomous laboratories, highlighting the synergistic effects of modular design, advanced reasoning, and self-correction to ensure both performance and reliability in high-stakes scientific applications. Code: https://github.com/pnnl/autolabs

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sensitive dependence on initial conditions in a formation of magnetic vortices

The magnetic vortex exhibits promise as a true random number generator for hardware-based encryption and probabilistic computing due to its stochastic formation of energetically equivalent fourfold degenerate states, characterized by two topologies: polarity and chirality. However, a comprehensive understanding of the stochastic formation of magnetic vortices remains elusive. In this work, we show that the magnetization relaxation in asymmetric Permalloy disks evolves along a pitchfork bifurcation, with both bifurcation paths leading to the formation of magnetic vortices with the same chirality. In the bifurcation, one formation path is always chosen under weak in-plane magnetic fields, ultimately determining the final magnetic vortex state. By delaying the in-plane magnetic field, we quantitatively investigate when the final vortex state is determined and find that it is closely associated with the initial conditions rather than the bifurcation point itself. Our findings provide valuable insights into future spintronic-based encryption and probabilistic computing.

Jeong, Suyeong↗

Bridging the Gap Between LLMs and LNS with Dynamic Data Format and Architecture Codesign

Deep neural networks (DNNs) have achieved tremendous success in the past few years. However, their training and inference demand exceptional computational and memory resources. Quantization has been shown as an effective approach to mitigate the cost, with the mainstream data types reduced from FP32 to FP16/BF16 and recently FP8 in the latest NVIDIA H100 GPUs. With increasingly aggressive quantization, however, the conventional floating-point formats suffer from limited precision in representing numbers around zero. Recently, NVIDIA demonstrated the potential of using a Logarithmic Number System (LNS) for the next generation of tensor cores. While LNS mitigates the hurdles in representing small numbers, in this work we observed a mismatch between LNS and the emerging Large Language Models (LLM), where LLM exhibits significant outliers when directly adopting the LNS format. In this paper, we present a data-format/architecture codesign to bright this gap. On the format side, we propose a dynamic LNS format to flexibly represent outliers at a higher precision, by exploiting asymmetry in the LNS representation and identifying outliers through a per-vector basis. On the architecture side, for demonstration, we realize the dynamic LNS format in a systolic array, which can handle the irregularity of the outliers at runtime. We implement our approach on an Alveo U280 FPGA as a prototype. Experimental results show that our design can effectively handle the outliers and resolve the mismatch between LNS and LLM, contributing to an accuracy improvement of 15.4% and 16% over the floating-point and the original LNS baselines, using four state-of-the-art LLM models. Our observation and design lay a solid foundation for the large-scale adoption of the LNS format in the next-generation deep learning hardware.

Haghi, Pouya↗

SRF Thin Films: Not just for Cavities

Recent years have seen renewed interest and activities in developing SRF cavity materials based on thin film technologies. In this framework, considerable progress has been achieved in the development of high quality films and layered structures along with associated deposition techniques such as ECR, HiPIMS, and ALD. Beyond cavity applications, the developments in SRF thin film technologies find a variety of applications in the fields of superconducting metamaterials, electronics, sensors and quantum devices. High quality Nb films with high RRR open the way for enhanced coherence times for quantum qbits. Other SRF thin films such as NbTiN are developed for superconducting backend processes for future generations of computing hardware and radiation-hard sensors for nuclear and high-energy physics. The unifying theme across these technologies is that the same physics and material properties such as extreme low loss, stability, and manufacturability are required. This talk will present an overview of the emerging applications of SRF films and structures beyond accelerator cavities.

Valente-Feliciano, A.-M.↗

Rapid Cryogenic Electrical Characterization of Materials and Devices Using Gifford-McMahon Cryocoolers

Thin-film heterostructures are necessary building blocks for superconducting and phononic quantum computing devices. Many new generations of quantum hardware demand extensive materials research to optimize performances at cryogenic temperatures (below 10 K). Here, we demonstrate compact cryogenic measurement systems capable of reaching sub-10K temperatures in less than three hours with the ability to measure AC/DC resistance and dielectric properties of thin-film materials. Our platform utilizes Gifford-McMahon (GM) cryocoolers as effective tools for providing high throughput cooling-warming cycles. We successfully used the GM-based measurement systems to measure 1) the superconducting transition temperature for Nb thin films (T c ~7.8 K), and 2) the temperature dependence of the dielectric constant in SiO 2 thin films down to 10 K. The fast electrical characterization feedback will be critical in developing robust materials and components for cryogenic computing devices.

36 MATERIALS SCIENCE↗

A Synthesis Methodology for Intelligent Memory Interfaces in Accelerator Systems

Domain-specific systems improve the performance of a specific set of applications compared to general-purpose processing systems by deploying custom hardware accelerators. These hardware accelerators are generated using high-level synthesis (HLS) tools. The HLS tools enable a comprehensive design space exploration to optimize the compute performance of the generated accelerators. However, they often ignore the challenges of implementing the accelerators in a system-on-chip, particularly how the accelerators access memory. Our work introduces a buffering system design that improves accelerators' memory accesses by intelligently employing burst transactions to prefetch useful data from external memory to on-chip local buffers. Our design is dynamic, parametric, and transparent to the accelerators generated by HLS tools. We derive the buffering system parameters using appropriate compiler-based analysis passes and memory channel latency constraints. The proposed buffering system design results in, on average, 8.8x performance improvements while lowering memory channel utilization on average by 53.2% for a set of PolyBench kernels.

Limaye, Ankur M. (ORCID:0000000194062584)↗

High-Fidelity Dataset Generation for Sensor Anomalies in Power Grids using Hardware-in-the-Loop Testbed

Sensor anomalies in power grids can have significant impacts on the operation of the grid due to the increased reliance of the grid operation on data-driven applications. However, there is a lack of datasets that accurately capture these anomalies as many of the anomalies go undetected using the current bad data detectors. High-fidelity labeled datasets are essential for developing robust applications that can detect and mitigate the impacts of anomalies. In this paper, we propose a hardware-in-the-loop testbed model that can emulate the grid behavior with high-fidelity. This testbed is used to inject anomalies at various levels in the grid architecture and generate labeled datasets. These high-fidelity datasets can be used for development and validation of data-driven applications for detection and mitigation of anomalies in grids and other cyber-physical systems.

Hyder, Burhan↗