Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AI hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Use of Legacy Maritime Protocols Increases Exploitability of Virtual Aids to Navigation

With increased reliance on Virtual Aid(s) to Navigation (VAtoN) - also known as electronic Aid(s) to Navigation (eAtoN), or virtual buoys - a cyber event is likely to cause disruption to international maritime shipping. VAtoN has no physical hardware for visual reference and displays only on a vessel’s Electronic Chart Display Information System (ECDIS) and Automatic Radar Plotting Aid (ARPA); therefore, mariners must rely on the accuracy of the information provided. As VAtoN uses the National Maritime Electronics Association (NMEA) 0183 protocol for both Global Navigation Satellite System (GNSS) and Automatic Identification System (AIS), an insecure protocol that has been proven susceptible to spoofing, denial, and manipulation, the likelihood of a cyber-related event increases substantially.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Application of Quantum Machine Learning to High Energy Physics Analysis at LHC Using Quantum Computer Simulators and Quantum Computer Hardware

Machine learning enjoys widespread success in High Energy Physics (HEP) analyses at LHC. However the ambitious HL-LHC program will require much more computing resources in the next two decades. Quantum computing may offer speed-up for HEP physics analyses at HL-LHC, and can be a new computational paradigm for big data analyses in High Energy Physics.We have successfully employed three methods (1) Variational Quantum Classifier (VQC) method, (2) Quantum Support Vector Machine Kernel (QSVM-kernel) method and (3) Quantum Neural Network (QNN) method for two LHC flagship analyses: ttH (Higgs production in association with two top quarks) and H->mumu (Higgs decay to two muons, the second generation fermions). We shall address the progressive improvements in performance from method (1) to method (3).We will present our experiences and results of a study on LHC High Energy Physics data analyses with IBM Quantum Simulator and Quantum Hardware (using IBM Qiskit framework), Google Quantum Simulator (using Google Cirq framework), and Amazon Quantum Simulator (using Amazon Braket cloud service). The work is in the context of a Qubit platform (a gate-model quantum computer). Taking into account the present limitation of hardware access, different quantum machine learning methods are studied on simulators and the results are compared with classical machine learning methods (BDT, classical Support Vector Machine and classical Neural Network). Furthermore, we do apply quantum machine learning on IBM quantum hardware to compare performance between quantum simulator and quantum hardware. The work is performed by an international and interdisciplinary collaboration with the Department of Physics and Department of Computer Sciences of University of Wisconsin, CERN Quantum Technology Initiative, IBM Research Zurich, IBM T.J. Watson Research Center, Fermilab Quantum Institute, BNL Computational Science Initiative, State University of New York at Stony Brook, and Quantum Computing and AI Research of Amazon Web Services. This work pioneers a close collaboration of academic institutions with industrial corporations in the High Energy Physics analyses effort. Though the size of event samples in future HL-LHC physics and the limited number of qubits pose some challenges to the Quantum Machine learning studies for High Energy Physics, more advanced quantum computers with larger number of qubits, reduced noise and improved running time (as envisioned by IBM and Google) may outperform classical machine learning in both classification power and in speed.Although the era of efficient quantum computing may still be years away, we have made promising progress and obtained preliminary results in applying quantum machine learning to High Energy Physics. A PROOF OF PRINCIPLE.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Intelligent experiments through real-time AI: Fast Data Processing and Autonomous Detector Control for sPHENIX and future EIC detectors (Phase-I)

With an ever increasing demand for high precision data from modern detectors for discovery science and precision measurements, all major high energy nuclear and particle experiments, current and future, are facing the challenge on how to deal with the large volume of raw data generated from sophisticated state-of-the-art detectors in high rate collisions. These goals need to be balanced with available hardware and cost limits on DAQ (Data AcQuisition system) bandwidth and offline computing resources to capture, store and process the signal events. Two prototypical examples are the upcoming sPHENIX experiment, the DOE next generation heavy ion physics experiment at the Relativistic Heavy Ion Collider at BNL, and the future EIC experiments that are planned to be online circa 2030.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Digital Twin for Chemical Science (DTCS) v0.01

Directly visualizing the trajectories of chemistry can unravel novel insights into the behavior of catalysts, gas phase reactions, photo-induced dynamics, and building blocks for quantum information processing. The ability of explicitly identifying, tracking, and tagging the exchange of matter, hence the annihilation and creation of new chemical species, can be best realized through a close coupling of theory and experiment. While the synchrotron-based characterization facilities propelled rapidly in its hardware, providing higher brightness, better resolution, and more precision, the software infrastructure is lagging. We developed DTCS (Digital Twin for Chemical Science) v.01, a central platform that faithfully mimics advanced instrumentations in Scientific User Facilities, by solving a variety of technical challenges in data acquisition, analysis, and model-driven interpretation. Rooted in physics and accelerated by AI, we validated this concept by direct comparison with precise experimental X-ray Photoelectron Spectroscopy (XPS) observations using a ubiquitous metal-water interfacial scenario, i.e., Ag/H2O as our main narrative. The DTCS v.01 input mirrors how the bench chemists work, with the output directly linked to the end station computer, thereby providing a user-friendly, knowledge-driven, and accessible user experience with mechanistic insights standardized in a way that are ready to be published, versioned, and transferred flexibly.

Qian, Jin↗

Advancing Dynamic Modeling of Grid-Connected PV Inverter Using Bi-LSTM-Based AI Model

Power electronic converters (PECs) are widely used in modern power systems to facilitate the interconnection between various AC or DC sources and loads. Because of the extensive integration, the power system has grown into a more dynamic system in which the dynamics of the PECs must be adequately modeled. The paper presents a new bidirectional long short-term memory (Bi-LSTM) method for evaluating grid-connected inverter-based resources (IBR) dynamics. The method is tested using real hardware data from a grid -connected commercial inverter in laboratory experiments. Results show the Bi-LSTM model accurately reproduces the detailed IBR model's dynamics, even when the internal structure is unknown and parameters are unknown, preventing the disclosure of the manufacturer's confidential data.

Subedi, Sunil [ORNL] (ORCID:000000034069090X)↗

Advancing Industry 4.0: Multimodal Sensor Fusion for AI-Based Fault Detection in 3D Printing

Additive manufacturing, particularly fused deposition modeling, is transforming modern production by enabling rapid prototyping and complex part fabrication. However, its layer-by-layer process remains vulnerable to faults such as nozzle clogging, filament runout, and layer misalignment, which compromise print quality and reliability. Traditional inspection methods are costly, time-intensive, and often limited to post-process analysis, making them unsuitable for real-time intervention. In this current study, the authors developed a novel, low-cost, and portable faultdetection system that leverages multimodal sensor fusion and artificial intelligence for real-time monitoring in FDM-based 3D printing. The system integrates acoustic, vibration, and thermal sensing into a non-intrusive architecture, capturing complementary data streams that reflect both mechanical and process-related anomalies. Acoustic and thermal sensors operate in a fully contactless manner, while the vibration sensor requires minimal attachment such that it will not interfere with printer hardware, thereby preserving portability and ease of deployment. The multimodal signals are processed into spectrograms and time-frequency features, which are classified using convolutional neural networks for intelligent fault detection. The proposed system advances Industry 4.0 objectives by offering an affordable, scalable, and practical monitoring solution that improves faultdetection accuracy, reduces waste, and supports sustainable, adaptive manufacturing.

42 ENGINEERING↗

Optics Enabled Networks and Architectures for Data Center Cost and Power Efficiency

Bandwidth demand for datacenter networks continues as performance increases and is further fueled by the exploding demand for AI and new HPC workloads. Managing power and costs will require a range of solutions including new networking and workload specialized architectures, composable systems and optical circuit switching. In this study we focus primarily on two topics, examining the benefits of flatter networks (enabled mainly by means of co-packaged-optics-enabled switches) and the utilization improvement potential for composable (disaggregated) systems, while discussing specialized hardware and networks, and optical circuit switching more briefly.

99 GENERAL AND MISCELLANEOUS↗

Low Precision and Efficient Programming Languages for Sustainable AI: Final Report for the Summer Project of 2024

This document contains all relevant material generated during the authors' summer internship at NREL in 2024. This report shows how to improve energy efficiency of a few code samples by using low-precision data types combined with mixed-precision algorithms. The main applications considered here are (i) linear system solvers using mixed precision, and (ii) neural networks using mixed precision. This report also discusses how programming languages affect energy consumption of algorithms, energy metrics for a code and tools, and the available current software and hardware infrastructure.

97 MATHEMATICS AND COMPUTING↗

Machine Learning for Predictive Performance Analysis in Charged Particle Beam Tools

Imaging methods driven by probes, electrons, and ions have played a dominant role in modern science and engineering. Opportunities for machine vision and AI that focus on consumer problems like driving and feature recognition, are now presenting themselves for automating aspects of the scientific processes. This proposal aims to enable and drive discovery in ultra-low energy implantation by taking advantage of faster processing, flexible control and detection methods, and architecture-agnostic workflows that will result in higher efficiency and shorter scientific development cycles. Custom microscope control, collection and analysis hardware will provide a framework for conducting novel in situ experiments revealing unprecedented insight into surface dynamics at the nanoscale. Ion implantation is a key capability for the semiconductor industry. As devices shrink, novel materials enter the manufacturing line, and quantum technologies transition to being more mainstream. Traditional implantation methods fall short in terms of energy, ion species, and positional precision. Here we demonstrate 1 keV focused ion beam Au implantation into Si and validate the results via atom probe tomography. We show the Au implant depth at 1 keV is 0.8 nm and that identical results for low energy ion implants can be achieved by either lowering the column voltage, or decelerating ions using bias – while maintaining a sub-micron beam focus. We compare our experimental results to static calculations using SRIM and dynamic calculations using binary collision approximation codes TRIDYN and IMSIL. A large discrepancy between the static and dynamic simulation is found that is due to lattice enrichment with high stopping power Au and surface sputtering. Additionally, we demonstrate how model details are particularly important to the simulation of these low-energy heavy-ion implantations. Finally, we discuss how our results pave a way to much lower implantation energies, while maintaining high spatial resolution.

47 OTHER INSTRUMENTATION↗

Strategies for Integrating Deep Learning Surrogate Models with HPC Simulation Applications

The emerging trend of the convergence of high performance computing (HPC), machine learning/deep learning (ML/DL), and big data analytics presents a host of challenges for large-scale computing campaigns that seek best practices to interleave traditional scientific simulation-based workloads with ML/DL models. A portfolio of systematic approaches to incorporate deep learning into modeling and simulation serves a vital need when we support AI for science at a computing facility. In this paper, we evaluate several strategies for deploying deep learning surrogate models in a representative physics application on supercomputers at the Oak Ridge Leadership Computing Facility (OLCF). We discuss a set of recommended deployment architectures and implementation approaches. We analyze and evaluate these alternatives and show their performance and scalability up to 1000 GPUs on two mainstream platforms equipped with different deep learning hardware and software stacks.

Yin, Junqi↗

Artificial Intelligence Thermostat to Detect Faults

Residential air conditioners and heat pumps often experience faults due to inadequate maintenance, which can severely reduce efficiency or even cause system failure. Common issues include dirty or clogged air filters and refrigerant leaks. These problems degrade performance and increase energy use and operating costs. This study presents a smart thermostat with embedded artificial intelligence to detect such faults and alert homeowners when maintenance is needed. The thermostat uses low-cost measurements—including return-air temperature, relative humidity, supply-air temperature, outdoor-air temperature, and condenser subcooling—to identify abnormal operations. Because different faults produce distinct response patterns, tailored algorithms are developed to recognize characteristic fault signatures. The investigation is built on a detailed co-simulation platform that couples EnergyPlus with the DOE/ORNL Heat Pump Design Model (HPDM). EnergyPlus represents the building’s dynamic environment, while HPDM is a high-fidelity, hardware-based model that can simulate fault-free performance as well as a wide range of faults, including gradual degradation such as minor refrigerant leakage. This platform provides a virtual training and testing environment that helps distinguish fault-induced behavior from normal operation and supports development of robust diagnostic algorithms. Using this framework, a Dynamic Bayesian Network was developed to identify two common faults—gradual refrigerant charge loss and indoor airflow blockage—and the AI-embedded thermostat was verified through annual building simulations.

Shen, Bo [ORNL] (ORCID:0000000336600393)↗

CGSim: A Simulation Framework for Large Scale Distributed Computing Environment

Large-scale distributed computing infrastructures such as the Worldwide LHC Computing Grid (WLCG) require comprehensive simulation tools for evaluating performance, testing new algorithms, and optimizing resource allocation strategies. However, existing simulators suffer from limited scalability, hardwired algorithms, lack of real-time monitoring, and inability to generate datasets suitable for modern machine learning approaches. We present CGSim, a simulation framework for large-scale distributed computing environments that addresses these limitations. Built upon the validated SimGrid simulation framework, CGSim provides high-level abstractions for modeling heterogeneous grid environments while maintaining accuracy and scalability. Key features include a modular plugin mechanism for testing custom workflow scheduling and data movement policies, interactive real-time visualization dashboards, and automatic generation of event-level datasets suitable for AI-assisted performance modeling. We demonstrate CGSim’s capabilities through a comprehensive evaluation using production ATLAS PanDA workloads, showing significant calibration accuracy improvements across WLCG computing sites. Scalability experiments show near-linear scaling for multi-site simulations, with distributed workloads achieving 6 × better performance compared to single-site execution. The framework enables researchers to simulate WLCG-scale infrastructures with hundreds of sites and thousands of concurrent jobs within practical time budget constraints on commodity hardware.

Vatsavai, Sairam Sri [Brookhaven National Laborato↗

ATHENA: Analytical Tool for Heterogeneous Neuromorphic Architectures

The ASC program seeks to use machine learning to improve efficiencies in its stockpile stewardship mission. Moreover, there is a growing market for technologies dedicated to accelerating AI workloads. Many of these emerging architectures promise to provide savings in energy efficiency, area, and latency when compared to traditional CPUs for these types of applications — neuromorphic analog and digital technologies provide both low-power and configurable acceleration of challenging artificial intelligence (AI) algorithms. If designed into a heterogeneous system with other accelerators and conventional compute nodes, these technologies have the potential to augment the capabilities of traditional High Performance Computing (HPC) platforms [5]. This expanded computation space requires not only a new approach to physics simulation, but the ability to evaluate and analyze next-generation architectures specialized for AI/ML workloads in both traditional HPC and embedded ND applications. Developing this capability will enable ASC to understand how this hardware performs in both HPC and ND environments, improve our ability to port our applications, guide the development of computing hardware, and inform vendor interactions, leading them toward solutions that address ASC’s unique requirements.

97 MATHEMATICS AND COMPUTING↗

A Survey: Handling Irregularities in Neural Network Acceleration with FPGAs

In the last decade, Artificial Intelligence (AI) through Deep Neural Networks (DNNs) has penetrated virtually every aspect of science, technology, and business. Many types of DNNs have been and continue to be developed, including Convolutional Neural Networks (CNNs), Recurrent Neural Net- works (RNNs), and Graph Neural Networks (GNNs). The overall problem for all of these Neural Networks (NNs) is that their target applications generally pose stringent constraints on latency and throughput, while also having strict accuracy requirements. There have been many previous efforts in creating hardware to accelerate NNs. The problem designers face is that optimal NN models typically have significant irregularities, making them hardware-unfriendly. In this paper, we first define the problems in NN acceleration by characterizing common irregularities in NN processing into 4 types; then we summarize the existing works that handle the four types of irregularities efficiently using hardware, especially FPGAs; finally, we provide a new vision of next-generation FPGA-based NN acceleration: that the emerging heterogeneity in the next-generation FPGAs is the key to achieving higher performance.

Geng, Tong↗

Fast, differentiable, and extensible big bang nucleosynthesis package

Here, we introduce light isotope nucleosynthesis with JAX (LINX), a new differentiable public big bang nucleosynthesis code designed for fast parameter estimation. By leveraging JAX, LINX achieves both speed and differentiability, enabling the use of Bayesian inference, including gradient-based methods. We discuss the formalism used in LINX for rapid primordial elemental abundance predictions and give examples of how LINX can be used. When combined with differentiable cosmic microwave background power spectrum emulators, LINX can be used for joint cosmic microwave background and big bang nucleosynthesis analyses without requiring extensive computational resources, including on personal hardware.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Lifetime extension of legacy CEBAF LLRF hardware

A significant portion of the Low-Level Radio Frequency (LLRF) hardware in Jefferson Lab’s CEBAF is from the original construction of the facility using 1980’s CAMAC technology. Of the fifty-three zones in CEBAF, thirty-six of them are legacy hardware. The age of the legacy system has led to difficulties in maintaining the hardware due to parts going obsolete without suitable drop in replacements. Continued operation of the legacy system is required as the installation of LLRF 3.0 systems is costly and cannot be completed in a short period of time with the available resources. The most pressing failure in the legacy system was a failing buffer card, which is responsible for communication between the EPICs network and individual RF control modules. A new buffer card was designed as a transparent, drop in, replacement so that upgrades are simply a matter of swapping the existing legacy hardware. This buffer card upgrades a single point failure component and promises to extend the operable lifetime of CEBAF’s legacy systems.

Accelerator Physics↗

The neurobench framework for benchmarking neuromorphic computing algorithms and systems

Neuromorphic computing shows promise for advancing computing efficiency and capabilities of AI applications using brain-inspired principles. However, the neuromorphic research field currently lacks standardized benchmarks, making it difficult to accurately measure technological advancements, compare performance with conventional methods, and identify promising future research directions. This article presents NeuroBench, a benchmark framework for neuromorphic algorithms and systems, which is collaboratively designed from an open community of researchers across industry and academia. NeuroBench introduces a common set of tools and systematic methodology for inclusive benchmark measurement, delivering an objective reference framework for quantifying neuromorphic approaches in both hardware-independent and hardware-dependent settings. For latest project updates, visit the project website (neurobench.ai).

Yik, Jason [Harvard Univ., Cambridge, MA (United S↗

Neural network methods for radiation detectors and imaging

Recent advances in image data proccesing through deep learning allow for new optimization and performance-enhancement schemes for radiation detectors and imaging hardware. This enables radiation experiments, which includes photon sciences in synchrotron and X-ray free electron lasers as a subclass, through data-endowed artificial intelligence. We give an overview of data generation at photon sources, deep learning-based methods for image processing tasks, and hardware solutions for deep learning acceleration. Most existing deep learning approaches are trained offline, typically using large amounts of computational resources. However, once trained, DNNs can achieve fast inference speeds and can be deployed to edge devices. A new trend is edge computing with less energy consumption (hundreds of watts or less) and real-time analysis potential. While popularly used for edge computing, electronic-based hardware accelerators ranging from general purpose processors such as central processing units (CPUs) to application-specific integrated circuits (ASICs) are constantly reaching performance limits in latency, energy consumption, and other physical constraints. These limits give rise to next-generation analog neuromorhpic hardware platforms, such as optical neural networks (ONNs), for high parallel, low latency, and low energy computing to boost deep learning acceleration (LA-UR-23-32395).

edge computing↗