Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Fast b -tagging at the high-level trigger of the ATLAS experiment in LHC Run 3

The ATLAS experiment relies on real-time hadronic jet reconstruction and b-tagging to record fully hadronic events containing b-jets. These algorithms require track reconstruction, which is computationally expensive and could overwhelm the high-level-trigger farm, even at the reduced event rate that passes the ATLAS first stage hardware-based trigger. In LHC Run 3, ATLAS has mitigated these computational demands by introducing a fast neural-network-based b-tagger, which acts as a low-precision filter using input from hadronic jets and tracks. It runs after a hardware trigger and before the remaining high-level-trigger reconstruction. This design relies on the negligible cost of neural-network inference as compared to track reconstruction, and the cost reduction from limiting tracking to specific regions of the detector. In the case of Standard Model HH → bb̅bb̅, a key signature relying on b-jet triggers, the filter lowers the input rate to the remaining high-level trigger by a factor of five at the small cost of reducing the overall signal efficiency by roughly 2%.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Differentiable Earth mover’s distance for data compression at the high-luminosity LHC

Abstract The Earth mover’s distance (EMD) is a useful metric for image recognition and classification, but its usual implementations are not differentiable or too slow to be used as a loss function for training other algorithms via gradient descent. In this paper, we train a convolutional neural network (CNN) to learn a differentiable, fast approximation of the EMD and demonstrate that it can be used as a substitute for computing-intensive EMD implementations. We apply this differentiable approximation in the training of an autoencoder-inspired neural network (encoder NN) for data compression at the high-luminosity LHC at CERN The goal of this encoder NN is to compress the data while preserving the information related to the distribution of energy deposits in particle detectors. We demonstrate that the performance of our encoder NN trained using the differentiable EMD CNN surpasses that of training with loss functions based on mean squared error.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Semi-Analytical Approach for State-Space Electromagnetic Transient Simulation

Here, this paper proposes a semi-analytical approach for efficient and accurate electromagnetic transient (EMT) simulation of a power grid. The approach first derives a high-order semi-analytical solution (SAS) of the grid’s state-space EMT model using the differential transformation (DT), and then evaluates the solution over enlarged, variable time steps to significantly accelerate the simulations while maintaining its high accuracy on detailed fast EMT dynamics. The approach also addresses switches during large time steps by using a limit violation detection algorithm with a binary search-enhanced quadratic interpolation. Case studies are conducted on EMT models of the IEEE 39-bus system and large-scale systems to demonstrate the merits of the new simulation approach against traditional numerical methods.

electromagnetic transient↗

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

Preliminary Development of Heat Transfer Model-Based Control Algorithms of Liquid Sodium Purification System: Advanced Sensors and Instrumentation Advanced Controls

Monitoring the operation of sodium purification system is essential for efficient operation of sodium fast reactors. In this work, a heat transfer model has been developed for monitoring the plugging meter and cold trap systems at the Mechanisms Engineering Test Loop (METL) liquid sodium facility at Argonne National Laboratory. The model of the purification system was developed by treating the respective aspects of the cold trap purification loop and plugging meter diagnostic loop as two separate control volumes using information from the METL piping and instrumentation diagram (P&ID). A model predictive controller was designed using first order differential equations with the specified boundary conditions. The system behavior was studied with a tuned optimized procedure using the internal cold trap temperature and plugging meter outlet temperature as control variables, and the air blower temperature as an independent variable respectively. Results of computer simulations obtained in this study compared favorably with experimental data showing very good reference tracking response with negligible overshoot as both plugging meter and cold trap physical models approach the setpoint.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Applications and Techniques for Fast Machine Learning in Science

In this community review report, we discuss applications and techniques for fast machine learning (ML) in science—the concept of integrating powerful ML methods into the real-time experimental data processing loop to accelerate scientific discovery. The material for the report builds on two workshops held by the Fast ML for Science community and covers three main areas: applications for fast ML across a number of scientific domains; techniques for training and implementing performant and resource-efficient ML algorithms; and computing architectures, platforms, and technologies for deploying these algorithms. We also present overlapping challenges across the multiple scientific domains where common solutions can be found. This community report is intended to give plenty of examples and inspiration for scientific discovery through integrated and accelerated ML solutions. This is followed by a high-level overview and organization of technical advances, including an abundance of pointers to source material, which can enable these breakthroughs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Stochastic quantum Krylov protocol with double-factorized Hamiltonians

Here we propose a class of randomized quantum Krylov diagonalization (rQKD) algorithms capable of solving the eigenstate estimation problem with modest quantum resource requirements. Compared to previous real-time evolution quantum Krylov subspace methods, our approach expresses the time evolution operator e –i$\widehat{H}$$\tau$ as a linear combination of unitaries and subsequently uses a stochastic sampling procedure to reduce circuit depth requirements. While our methodology applies to any Hamiltonian with fast-forwardable subcomponents, we focus on its application to the explicitly double-factorized electronic-structure Hamiltonian. To demonstrate the potential of the proposed rQKD algorithm on near-term quantum devices, we provide numerical benchmarks for a variety of molecular systems with circuit-based state-vector simulators including the effects of sampling noise, achieving ground-state energy errors of less than 1 kcal mol -1 with circuit depths orders of magnitude shallower than those required for low-rank deterministic Trotter-Suzuki decompositions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

The Blending ToolKit: A simulation framework for evaluation of galaxy detection and deblending

We present an open source Python library for simulating overlapping (i.e., blended) images of galaxies and performing self-consistent comparisons of detection and deblending algorithms based on a suite of metrics. The package, named Blending Toolkit (BTK), serves as a modular, flexible, easy-to-install, and simple-to-use interface for exploring and analyzing systematic effects related to blended galaxies in cosmological surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time (LSST). BTK has three main components: (1) a set of modules that perform fast image simulations of blended galaxies, using the open source image simulation package GalSim; (2) a module that standardizes the inputs and outputs of existing deblending algorithms; (3) a library of deblending metrics commonly defined in the galaxy deblending literature. In combination, these modules allow researchers to explore the impacts of galaxy blending in cosmological surveys. Additionally, BTK provides researchers who are developing a new deblending algorithm a framework to evaluate algorithm performance and make principled comparisons with existing deblenders. BTK includes a suite of tutorials and comprehensive documentation. The source code is publicly available on GitHub at https://github.com/LSSTDESC/BlendingToolKit.

79 ASTRONOMY AND ASTROPHYSICS↗

Real-Time Detection and Localization of Line Trip Event Via Relative Phase Angles

Line trip events widely exist in power systems. They can result in power outages and a huge economic loss if not promptly detected and localized. To provide a fast and precise solution, this paper presents a Complete Coverage of Voltage Measurement (CCVM)-based line trip event detection algorithm and a Relative Phase Angle (RPA)-based line trip event localization algorithm. First, frequency and relative phase angle features during a line trip event are calculated. Then, the CCVM-based algorithm is proposed from both frequency and rate of change of frequency estimation algorithm aspects. Additionally, the RPA-based algorithm is presented, and two cases are studied to demonstrate the uniqueness of the proposed algorithm. Various experiments are conducted, where the simulation results demonstrate that the proposed CCVM-based algorithm can detect a line trip event as short as 2.07 ms. In addition, the RPA-based algorithm has 1.26 times higher localization accuracy compared with the frequency magnitude-based and phase angle-based algorithms. In conclusion, the experiment results on the examples from two interconnected power systems in the U.S. verified the performance of the proposed algorithms in wide-area power systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distributed Inference with Sparse and Quantized Communication

Here, we consider the problem of distributed inference where agents in a network observe a stream of private signals generated by an unknown state, and aim to uniquely identify this state from a finite set of hypotheses. We focus on scenarios where communication between agents is costly, and takes place over channels with finite bandwidth. To reduce the frequency of communication, we develop a novel event-triggered distributed learning rule that is based on the principle of diffusing low beliefs on each false hypothesis. Building on this principle, we design a trigger condition under which an agent broadcasts only those components of its belief vector that have adequate innovation, to only those neighbors that require such information. We prove that our rule guarantees convergence to the true state exponentially fast almost surely despite sparse communication, and that it has the potential to significantly reduce information flow from uninformative agents to informative agents. Next, to deal with finite-precision communication channels, we propose a distributed learning rule that leverages the idea of adaptive quantization. We show that by sequentially refining the range of the quantizers, every agent can learn the truth exponentially fast almost surely, while using just 1 bit to encode its belief on each hypothesis. For both our proposed algorithms, we rigorously characterize the trade-offs between communication-efficiency and the learning rate.

42 ENGINEERING↗

Development of lean, efficient, and fast physics-framed deep-learning-based proxy models for subsurface carbon storage

In this work, we present deep-learning-based surrogate models for CCUS developed with four different algorithms and a physics-framed two-phase flow problem involving displacement of water by CO 2 . The deep-learning models were trained using 3D datasets describing the pressure plume, CO 2 saturation plume, and water extraction rate generated by numerical simulation. The hyperparameters defining the architecture of the neural networks were optimized to determine the slimmest network size and training parameters that give the most efficient performance at the least training cost. To develop a robust model that closely mimics the governing physical laws, the discretized form of the two-phase fluid transport equation was used to formulate the supervised deep-learning task. The algorithms investigated in this study predicted the data to above 95% accuracy, with the multi-layer perceptron model demonstrating the best performance by balancing training speed, prediction time, and prediction accuracy with lean network capacity. Furthermore, the surrogate models simultaneously predict reservoir pressure and CO 2 saturation in every grid block, including the surface well extraction rate and bottomhole pressure, at all simulation times for a given static model realization in just a few seconds on a standard desktop computer. A key outcome of this study is that limits can be placed on network design parameters to avoid over designing neural networks, with associated efficiencies in training and prediction times. This is very useful because large volumes of data may be generated in CCUS projects and over-design of neural network architectures imposes penalties that are antithetical to the goal of near-real time forecasting.

58 GEOSCIENCES↗

Charged particle tracking with quantum annealing optimization

Abstract At the High Luminosity Large Hadron Collider (HL-LHC), traditional track reconstruction techniques that are critical for physics analysis will need to be upgraded to scale with track density. Quantum annealing has shown promise in its ability to solve combinatorial optimization problems amidst an ongoing effort to establish evidence of a quantum speedup. As a step towards exploiting such potential speedup, we investigate a track reconstruction approach by adapting the existing geometric Denby-Peterson (Hopfield) network method to the quantum annealing framework for HL-LHC conditions. We develop additional techniques to embed the problem onto existing and near-term quantum annealing hardware. Results using simulated annealing and quantum annealing with the D-Wave 2X system on the TrackML open dataset are presented, demonstrating the successful application of a quantum annealing algorithm to the track reconstruction challenge. We find that combinatorial optimization problems can effectively reconstruct tracks, suggesting possible applications for fast hardware-specific implementations at the HL-LHC while leaving open the possibility of a quantum speedup for tracking.

97 MATHEMATICS AND COMPUTING↗

Voltage regulation in distribution grids: A survey

Environmental and sustainability concerns have caused a recent surge in the penetration of distributed energy resources into the power grid. This may lead to voltage violations in the distribution systems making voltage regulation more relevant than ever. Owing to this and rapid advancements in sensing, communication, and computation technologies, the literature on voltage control techniques is growing at a rapid pace in distribution networks. In particular, there is a paradigm shift from traditional offline centralized approaches to distributed ones leveraging increased and varied types of actuators, real-time sensing, fast and efficient computations, and an overall distributed situational awareness. This paper reviews state-of-the-art voltage control algorithms, summarizes the underlying methods, and classifies their coordination mechanisms into local, centralized, distributed, and decentralized. The underlying solution methodologies are further classified into two categories, open-loop and feedback-based. Two specific example workflows are provided to illustrate these solutions for voltage regulation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Logical quantum processor based on reconfigurable atom arrays

Suppressing errors is the central challenge for useful quantum computing, requiring quantum error correction (QEC) for large-scale processing. However, the overhead in the realization of error-corrected ‘logical’ qubits, in which information is encoded across many physical qubits for redundancy, poses substantial challenges to large-scale logical quantum computing. Here we report the realization of a programmable quantum processor based on encoded logical qubits operating with up to 280 physical qubits. Using logical-level control and a zoned architecture in reconfigurable neutral-atom arrays, our system combines high two-qubit gate fidelities, arbitrary connectivity, as well as fully programmable single-qubit rotations and mid-circuit readout. Operating this logical processor with various types of encoding, we demonstrate improvement of a two-qubit logic gate by scaling surface-code distance from d = 3 to d = 7, preparation of colour-code qubits with break-even fidelities, fault-tolerant creation of logical Greenberger–Horne–Zeilinger (GHZ) states and feedforward entanglement teleportation, as well as operation of 40 colour-code qubits. Finally, using 3D [[8,3,2]] code blocks, we realize computationally complex sampling circuits with up to 48 logical qubits entangled with hypercube connectivity with 228 logical two-qubit gates and 48 logical CCZ gates. We find that this logical encoding substantially improves algorithmic performance with error detection, outperforming physical-qubit fidelities at both cross-entropy benchmarking and quantum simulations of fast scrambling. These results herald the advent of early error-corrected quantum computation and chart a path towards large-scale logical processors.

74 ATOMIC AND MOLECULAR PHYSICS↗

Real Time implementation of Artificial Intelligence compression algorithm for High-Speed Streaming Readout signals

The new generation of high-energy physics experiments plans to acquire data in streaming mode. With this approach, it is possible to access the information of the whole detector (organized in time slices) for optimal and lossless triggering of data acquisitions. With this approach, data rates, especially in large detectors, are often very high, and the network is likely to be the bottleneck for the entire Streaming Read Out system. The aim of this work is to study the implementation of a lossy compression algorithm based on Artificial Intelligence: an Autoencoder. With Machine Learning it is possible to achieve a high compression ratio and fast inference time with only a small degradation of the signals, almost negligible for the specific application. This work explores different configurations of the Autoencoder and the implementation on different hardware. Different Autoencoder configurations are explored to find the best trade-off between compression ratio and reconstruction loss, both for signals and energy spectrum. Different hardware implementations are also explored to find the best platform to achieve real-time performance for the specific application.

Rossi, Fabio (ORCID:0009000385713885)↗

An Update to the National Renewable Energy Laboratory Baseline Wind Turbine Controller

The National Renewable Energy Laboratory's 5-MW wind turbine model is well established as an industry standard and is often used as a comparison model, or a model on which to build upon. Though effective, the legacy controller for the 5-MW wind turbine uses a simple algorithm that is not up to date with many industry standards. Additionally, as the research community has advanced into fast-paced development cycles, as systems engineering tools such as Wind-Plant Integrated System Design & Engineering Model (WISDEM ®) [1] are employed, and as a greater focus on controls co-design practices is encouraged, demand for a generic wind turbine controller has arisen. This work presents updates for the NREL 5-MW controller to a more modern control architecture, and establishes a generic tuning framework that can be easily adapted to various wind turbines. Based on initial results, the updated generic controller eases the automatic tuning process while maintaining or improving the performance of the legacy NREL 5-MW controller.

17 WIND ENERGY↗

Faster Tensor Network Decoding for Topological Quantum Codes

We present a fast and Bayes-optimal-approximating tensor network decoder for planar quantum LDPC codes based on the tensor renormalization group algorithm, originally proposed by Levin, and Nave. By precomputing the renormalization group flow for the null syndrome, we need only recompute tensor contractions in the causal cone of the measured syndrome at the time of decoding. This allows us to achieve an overall runtime complexity of ($pnχ^6$) where p is the depolarizing noise rate, and χ is the cutoff value used to control singular value decomposition approximations used in the algorithm. We apply our decoder to the surface code in the code capacity noise model and compare its performance to the original matrix product state (MPS) tensor network decoder introduced by Bravyi, Suchara, and Vargo. The MPS decoder has a p-independent runtime complexity of $\mathcal{O}(nχ^3)$ resulting in significantly slower decoding times compared to our algorithm in the low-p regime.

97 MATHEMATICS AND COMPUTING↗

Multiscale Modeling of a Direct Nonoxidative Methane Dehydroaromatization Reactor with a Validated Model for Catalyst Deactivation

Due to the recent boom in shale gas production, aromatics production using direct nonoxidative methane dehydroaromatization (DHA) is being investigated extensively. However, due to rapid coke formation, catalysts in the nonoxidative methane DHA reactors get deactivated, which is one of the critical issues for the commercial success of the methane DHA process. In this paper, a model for catalyst deactivation is developed. Rate models for other DHA reactions are developed by considering the decrease in the catalyst activity with time. Due to the very fast coke formation rate on the fresh catalyst, there is coke formation immediately upon the introduction of the feed. Therefore, an algorithm is developed for estimation of the initial state of the reactor and the kinetic parameters by coupling an iterative direct substitution approach with an optimization approach. Transient experimental data from an in-house reactor are first reconciled and then used for developing the kinetic model including the coke formation model. Using the rate model, a dynamic, heterogeneous, multiscale reactor model with embedded heating is developed. Here, the model couples the catalyst pellet level model with a reactor level model. Impacts of temperature, L/D ratio, and scheduling of reactors on variability in conversion and yield with time are studied.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗