Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “runtime”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Combined Meteorological and Hydrologic Uncertainties Shape Projections of Future Soil Moisture in the Eastern United States

Physical hazards pose risks to many critical systems. Designing adaptive measures to mitigate these risks is challenging due to large uncertainties in modeling future hazards and the associated sectoral responses. Here, we help address this challenge in a hydrologic context by examining the combined role of meteorological forcing and hydrologic parameter uncertainties in shaping projections of future soil moisture. By encoding a simple conceptual water balance model in a differentiable programming framework, we facilitate fast runtimes and an efficient calibration, enabling an improved uncertainty analysis. We characterize uncertainty in model parameters by calibrating against different target data sets and by using several loss functions. We then convolve the resulting parameter ensemble with a set of Earth system model projections to produce a large ensemble (2,340 members) of daily soil moisture simulations. Focusing on the eastern United States, we find that most ensemble members project a drying of soils across the region, although some simulate wetter conditions throughout this century. Our ensemble shows an increase in the frequency and intensity of dry extremes while there is less agreement for wet extremes. We conduct sensitivity analyses on several soil moisture signatures to measure the relative influence of meteorological and hydrologic uncertainties across space and time. Both meteorological and hydrologic factors contribute consistently to uncertainty surrounding long-term trends, while changes to both wet and dry soil extremes are typically more sensitive to hydrologic parameter uncertainty. Our results underscore the need to account for varied sources of uncertainty when developing long-term hydrometeorological projections.

Lafferty, David C. [University of Illinois Urbana‐↗

BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters

Microbial metabolic gene clusters encode the biosynthesis or catabolism of metabolites that facilitate ecological specialization, mediate microbiome interactions and constitute a major source of medicines and crop protection agents. Here, we present BiG-SCAPE and BiG-SLiCE 2.0, next-generation methods that facilitate scalable, accurate and interactive gene cluster analyses. BiG-SCAPE 2.0 updates its classification, alignment methods, and visualizations, enabling more accurate analysis, up to 8x faster runtimes and halved memory requirements. BiG-SLiCE 2.0 updates its distance metric, pHMM database, and classification logic, resulting in increased sensitivity nearing that of BiG-SCAPE. Analysis of 260,630 biosynthetic gene clusters from publicly available genomes reveals that both tools generate concurring estimates of gene cluster diversity, thus providing significantly extended methodological support for recent evidence indicating that the vast majority of natural product diversity remains unexplored. Together, these updates will facilitate global genome mining efforts for natural product discovery and microbiome analyses scalable with current data sizes.

Draisma, Arjan [Wageningen University & Research (↗

Boundaries of quantum supremacy via random circuit sampling

Abstract Google’s quantum supremacy experiment heralded a transition point where quantum computers can evaluate a computational task, random circuit sampling, faster than classical supercomputers. We examine the constraints on the region of quantum advantage for quantum circuits with a larger number of qubits and gates than experimentally implemented. At near-term gate fidelities, we demonstrate that quantum supremacy is limited to circuits with a qubit count and circuit depth of a few hundred. Larger circuits encounter two distinct boundaries: a return of a classical advantage and practically infeasible quantum runtimes. Decreasing error rates cause the region of a quantum advantage to grow rapidly. At error rates required for early implementations of the surface code, the largest circuit size within the quantum supremacy regime coincides approximately with the smallest circuit size needed to implement error correction. Thus, the boundaries of quantum supremacy may fortuitously coincide with the advent of scalable, error-corrected quantum computing.

97 MATHEMATICS AND COMPUTING↗

Solving k –SAT problems with generalized quantum measurement

We generalize the projection–based quantum measurement–driven k –SAT algorithm of Benjamin, Zhao, and Fitzsimons to arbitrary strength quantum measurements, including the limit of continuous monitoring. In doing so, we clarify that this algorithm is a particular case of the measurement–driven quantum control strategy elsewhere referred to as “Zeno dragging”. We argue that the algorithm is most efficient with finite time and measurement resources in the continuum limit, where measurements have an infinitesimal strength and duration. Moreover, for solvable k -SAT problems, the dynamics generated by the algorithm converge deterministically towards target dynamics in the long–time (Zeno) limit, implying that the algorithm can successfully operate autonomously via Lindblad dissipation, without detection. We subsequently study both the conditional and unconditional dynamics of the algorithm implemented via generalized measurements, quantifying the advantages of detection for heralding errors. These strategies are investigated first in a computationally–trivial 2-qubit 2-SAT problem to build intuition, and then we consider the scaling of the algorithm on 3-SAT problems encoded with 4–10 qubits. We numerically investigate the scaling of 3-SAT with respect to algorithmic runtime and find that the optimized time to solution scales with qubit number n as λ n , where λ is slightly larger than $\sqrt{2}$ for unconditional dynamics and less than $\sqrt{2}$ for conditional dynamics. We assess the implications for using this analog measurement–driven approach to quantum computing in practice.

quantum information↗

Critical Assessment of Metagenome Interpretation: the second round of challenges

Abstract Evaluating metagenomic software is key for optimizing metagenome interpretation and focus of the Initiative for the Critical Assessment of Metagenome Interpretation (CAMI). The CAMI II challenge engaged the community to assess methods on realistic and complex datasets with long- and short-read sequences, created computationally from around 1,700 new and known genomes, as well as 600 new plasmids and viruses. Here we analyze 5,002 results by 76 program versions. Substantial improvements were seen in assembly, some due to long-read data. Related strains still were challenging for assembly and genome recovery through binning, as was assembly quality for the latter. Profilers markedly matured, with taxon profilers and binners excelling at higher bacterial ranks, but underperforming for viruses and Archaea. Clinical pathogen detection results revealed a need to improve reproducibility. Runtime and memory usage analyses identified efficient programs, including top performers with other metrics. The results identify challenges and guide researchers in selecting methods for analyses.

54 ENVIRONMENTAL SCIENCES↗

Efficient and assured reinforcement learning-based building HVAC control with heterogeneous expert-guided training

Abstract Building heating, ventilation, and air conditioning (HVAC) systems account for nearly half of building energy consumption and $$20\%$$ of total energy consumption in the US. Their operation is also crucial for ensuring the physical and mental health of building occupants. Compared with traditional model-based HVAC control methods, the recent model-free deep reinforcement learning (DRL) based methods have shown good performance while do not require the development of detailed and costly physical models. However, these model-free DRL approaches often suffer from long training time to reach a good performance, which is a major obstacle for their practical deployment. In this work, we present a systematic approach to accelerate online reinforcement learning for HVAC control by taking full advantage of the knowledge from domain experts in various forms . Specifically, the algorithm stages include learning expert functions from existing abstract physical models and from historical data via offline reinforcement learning, integrating the expert functions with rule-based guidelines, conducting training guided by the integrated expert function and performing policy initialization from distilled expert function. Moreover, to ensure that the learned DRL-based HVAC controller can effectively keep room temperature within the comfortable range for occupants, we design a runtime shielding framework to reduce the temperature violation rate and incorporate the learned controller into it. Experimental results demonstrate up to 8.8 X speedup in DRL training from our approach over previous methods, with low temperature violation rate.

Xu, Shichao↗

Efficient online quantum circuit learning with no upfront training

Optimization is a promising candidate for studying the utility of variational quantum algorithms (VQAs). However, evaluating cost functions using quantum hardware introduces runtime overheads that limit exploration. Surrogate-based methods can reduce calls to a quantum computer, yet existing approaches require hyperparameter pre-training and have been tested only on small problems. Here, we show that surrogate-based methods can enable successful optimization at scale, without pre-training, by using radial basis function interpolation (RBF) to construct an adaptive, hyperparameter-free surrogate. Using the surrogate as an acquisition function drives hardware queries to the vicinity of the true optima. For 16-qubit random 3-regular Max-Cut instances with the Quantum Approximate Optimization Algorithm (QAOA), our method outperforms state-of-the-art approaches, without considering their upfront training costs. Furthermore, we successfully optimize QAOA circuits for 127-qubit random Ising models on an IBM processor using 10 4 −10 5 measurements. Strong empirical performance demonstrates the promise of automated surrogate-based learning for large-scale VQA applications.

97 MATHEMATICS AND COMPUTING↗

Search of strong lens systems in the Dark Energy Survey using convolutional neural networks

We present our search for strong lens, galaxy-scale systems in the first data release of the Dark Energy Survey (DES), based on a color-selected parent sample of 18 745 029 luminous red galaxies (LRGs). We used a convolutional neural network (CNN) to grade this LRG sample with values between 0 (non-lens) and 1 (lens). Our training set of mock lenses is data-driven, that is, it uses lensed sources taken from HST-COSMOS images and lensing galaxies from DES images of our LRG sample. A total of 76 582 cutouts were obtained with a score above 0.9, which were then visually inspected and classified into two catalogs. The first one contains 405 lens candidates, of which 90 present clear lensing features and counterparts, while the other 315 require more evidence, such as higher resolution imaging or spectra, to be conclusive. A total of 186 candidates are newly identified by our search, of which 16 are among the 90 most promising (best) candidates. The second catalog includes 539 ring galaxy candidates. This catalog will be a useful false positive sample for training future CNNs. For the 90 best lens candidates we carry out color-based deblending of the lens and source light without fitting any analytical profile to the data. This method is shown to be very efficient in the deblending, even for very compact objects and for objects with a complex morphology. Finally, from the 90 best lens candidates, we selected 52 systems with one single deflector to test an automated modeling pipeline that has the capacity to successfully model 79% of the sample within an acceptable computing runtime.

79 ASTRONOMY AND ASTROPHYSICS↗

Full Simulation of CMS for Run-3 and Phase-2

In this contribution we report the status of the CMS Geant4 simulation and the prospects for Run-3 and Phase-2. Firstly, we report about our experience during the start of Run-3 with Geant4 10.7.2, the common software package DD4hep for geometry description, and VecGeom runtime geometry library. In addition, FTFP_BERT_EMM Physics List and CMS configuration for tracking in magnetic field have been utilized. For the first time, for the Grid mass production of Monte-Carlo, this combination of components is used. Further simulation improvements are under development targeting Run-3 such as the switch to the new Geant4 11.1 in production, that provides several features important for the optimization of simulation, for example the new transportation process with built-in multiple scattering, neutron general process, custom tracking manager, G4HepEm sub-library, and others. We will present evaluation of various options, validation results, and the final choice of simulation configuration for 2023 production and beyond. The performance of the CMS full simulation for Run-2 and Run-3 will also be discussed. CMS development plan for the Phase-2 Geant4 based simulation is very ambitious, and it includes a new geometry description, physics, and simulation configurations. The progress on new detector descriptions and full simulation will be presented as well as the R&D in progress to reduce compute capacity needs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Boosting RDataFrame performance with transparent bulk event processing

RDataFrame is ROOT’s high-level interface for Python and C++ data analysis. Since it first became available, RDataFrame adoption has grown steadily and it is now poised to be a major component of analysis software pipelines for LHC Run 3 and beyond. Thanks to its design inspired by declarative programming principles, RDataFrame enables the development of highperformance, highly parallel analyses without requiring expert knowledge of multi-threading and I/O: user logic is expressed in terms of self-contained, small computation kernels tied together by a high-level API. This design completely decouples analysis logic from its actual execution, and opens several interesting avenues for workflow optimization. In particular, in this work we explore the benefits of moving internal data processing from an event-by-event to a bulkby-bulk loop. This refactoring dramatically reduces the framework’s runtime overheads; in collaboration with the I/O layer it improves data access patterns; it exposes information that optimizing compilers might use to auto-vectorize the invocation of user-defined computations; finally, while existing user-facing interfaces remain unaffected, it becomes possible to additionally offer interfaces that explicitly expose bulks of events, useful e.g. for the injection of GPU kernels into the analysis workflow. In order to inform similar future R&D, design challenges will be presented, as well as an investigation of the relevant timememory trade-off backed by novel performance benchmarks.

Guiraud, Enrico↗

Experience with the alpaka performance portability library in the CMS software

ion Library for Parallel Kernel Acceleration) is a header-only C++ library that provides performance portability across different back-ends, abstracting the underlying levels of parallelism. It supports serial and parallel execution on CPUs, and extremely parallel execution on NVIDIA, AMD and Intel GPUs.This contribution will show how alpaka is used in the CMS software to develop and maintain a single code base; to use different toolchains to build the code for each supported back-end, and link them into a single application; to seamlessly select the best backend at runtime, and implement portable reconstruction algorithms that run efficiently on CPUs and GPUs from different vendors. It will describe the validation and deployment of the alpaka-based implementation in the CMS High Level Trigger, and highlight how it achieves near-native performance.

Alawieh, Jaafar [CERN]↗

Reconstruction framework advancements to support streaming for the ePIC detector at the EIC

The ePIC collaboration adopted the JANA2 framework to manage its reconstruction algorithms. This framework has since evolved substantially in response to ePIC’s needs. There have been three main design drivers: integrating cleanly with the Podio-based data models and other layers of the key4hep stack, enabling external configuration of existing components, and supporting timeframe splitting for streaming readout. The result is a unified component model featuring a new declarative interface for specifying inputs, outputs, parameters, services, and resources. This interface enables the user to instantiate, configure, and wire components via an external file. One critical new addition to the component model is a hierarchical decomposition of data boundaries into levels such as Run, Timeframe, PhysicsEvent, and Subevent. Two new component abstractions, Folder and Unfolder, are introduced in order to traverse this hierarchy, e.g. by splitting or merging. The pre-existing components can now operate at different event levels, and JANA2 will automatically construct the corresponding parallel processing topology. This means that a user may write an algorithm once, and configure it at runtime to operate on timeframes or on physics events. Overall, these changes mean that the user requires less knowledge about the framework internals, obtains greater flexibility with configuration, and gains the ability to reuse the existing abstractions in new streaming contexts.

Brei, Nathan [Thomas Jefferson National Accelerato↗

TIA: A forward model and analyzer for Talbot interferometry experiments of dense plasmas

Interferometry is one of the most sensitive and successful diagnostic methods for plasmas. However, owing to the design of most common interferometric systems, the wavelengths of operation and, therefore, the range of densities and temperatures that can be probed are severely limited. Talbot–Lau interferometry offers the possibility of extending interferometry measurements to x-ray wavelengths by means of the Talbot effect. While there have been several proof-of-concept experiments showing the efficacy of this method, it is only recently that experiments to probe High Energy Density (HED) plasmas using Talbot–Lau interferometry are starting to take place. To improve these experimental designs, we present here the Talbot-Interferometry Analyzer (TIA) tool, a forward model for generating and postprocessing synthetic x-ray interferometry images from a Talbot–Lau interferometer. Although TIA can work with any two-dimensional hydrodynamic code to study plasma conditions as close to reality as possible, this software has been designed to work by default with output files from the hydrodynamic code FLASH, making the tool user-friendly and accessible to the general plasma physics community. Here, the model has been built into a standalone app, which can be installed by anyone with access to the MATLAB runtime installer and is available upon request to the authors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Direct implicit and explicit energy-conserving particle-in-cell methods for modeling of capacitively coupled plasma devices

Achieving large-scale kinetic modeling is a crucial task for the development and optimization of modern plasma devices. With the trend of decreasing pressure in applications, such as plasma etching, kinetic simulations are necessary to self-consistently capture the particle dynamics. The standard, explicit, electrostatic, momentum-conserving particle-in-cell method suffers from restrictive stability constraints on spatial cell size and temporal time step, requiring resolution of the electron Debye length and electron plasma period, respectively. This results in a very high computational cost, making the technique prohibitive for large volume device modeling. We investigate the direct implicit algorithm and the explicit energy conserving algorithm as alternatives to the standard approach, both of which can reduce computational cost with a minimal (or controllable) impact on results. These algorithms are implemented into the well-tested EDIPIC-2D and LTP-PIC codes, and their performance is evaluated via 2D capacitively coupled plasma discharge simulations. The investigation reveals that both approaches enable the utilization of cell sizes larger than the Debye length, resulting in a reduced runtime, while incurring only minor inaccuracies in plasma parameters. The direct implicit method also allows for time steps larger than the electron plasma period; however, care must be taken to avoid numerical heating or cooling. It is demonstrated that by appropriately adjusting the ratio of cell size to time step, it is possible to mitigate this effect to an acceptable level.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Semantic embedding for quantum algorithms

The study of classical algorithms is supported by an immense understructure, founded in logic, type, and category theory, that allows an algorithmist to reason about the sequential manipulation of data irrespective of a computation’s realizing dynamics. As quantum computing matures, a similar need has developed for an assurance of the correctness of high-level quantum algorithmic reasoning. Parallel to this need, many quantum algorithms have been unified and improved using quantum signal processing (QSP) and quantum singular value transformation (QSVT), which characterize the ability, by alternating circuit ansätze, to transform the singular values of sub-blocks of unitary matrices by polynomial functions. However, while the algebraic manipulation of polynomials is simple (e.g., compositions and products), the QSP/QSVT circuits realizing analogous manipulations of their embedded polynomials are non-obvious. This work constructs and characterizes the runtime and expressivity of QSP/QSVT protocols where circuit manipulation maps naturally to the algebraic manipulation of functional transforms (termed semantic embedding). In this way, QSP/QSVT can be treated and combined modularly, purely in terms of the functional transforms they embed, with key guarantees on the computability and modularity of the realizing circuits. We also identify existing quantum algorithms whose use of semantic embedding is implicit, spanning from distributed search to proofs of soundness in quantum cryptography. The methods used, based in category theory, establish a theory of semantically embeddable quantum algorithms, and provide a new role for QSP/QSVT in reducing sophisticated algorithmic problems to simpler algebraic ones.

Physics↗

Ground and excited state gradients with end-to-end differentiable semiempirical quantum chemistry

Accurate and efficient gradients of molecular energy with respect to nuclear degrees of freedom are essential for geometry optimization and molecular dynamics, including simulations that go beyond the Born–Oppenheimer regime. A common approach involves deriving analytical formulas for new electronic structure methods, which is often conceptually difficult and requires tedious coding. Here, we implement analytical, semi-numerical, and automatic differentiation (AD)-based gradient pathways for semiempirical Hamiltonian models in the PYSEQM software package, leveraging both graphics processing unit (GPU) and central processing unit (CPU) architectures. We further extend these capabilities to excited states calculated using the configuration interaction singles and time-dependent Hartree–Fock ansätze. We benchmark wall time, peak memory usage, and accuracy across three molecular families of varying chemical complexity, including systems of up to a thousand atoms. For ground-state simulations, analytical and AD gradients achieve near-identical GPU runtimes, while semi-numerical gradients are slower on GPU but remain competitive on CPU. For excited states, both analytical and custom AD approaches using implicit differentiation show similar performance and low memory requirements, whereas gradients with full AD are memory-limited. AD gradients match analytical ones in accuracy across all tested systems, aided by a quaternion-based diatomic frame rotation for two-center quantities that ensures smooth energy surfaces. Overall, automatic differentiation emerges as a practical alternative to analytical gradients in semiempirical quantum chemistry, offering high accuracy while allowing seamless integration in AI-driven workflows and popular packages, such as PyTorch and JAX. Our results provide actionable guidance for selecting optimal gradient strategies in large-scale ground- and excited-state molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bootstrap embedding for interacting electrons in phonon coherent-state mean field

Here, we develop a Fermi–Bose bootstrap embedding framework for the ground state of interacting electrons coupled to a phonon mean field. The method combines bootstrap embedding for correlated electrons with a self-consistent coherent-state mean-field treatment for phonons. This method models the interacting electron–phonon problem as a system of correlated electrons traveling in a self-consistently specified potential landscape, allowing for efficient treatment of large lattice systems. Convergence of the methods for fragment size and total system size is demonstrated for the one-dimensional Hubbard–Holstein model for up to 350 sites. Finite-size scaling is performed to extrapolate to the infinite system size. Benchmarking against the density matrix renormalization group for a small 8-site system at half- and quarter-filling shows an orders-of-magnitude runtime advantage. The comparison further reveals that the method performs best in regimes dominated by localization, such as the Mott insulating phase and the strong-coupling tiny polaron regime, where the local embedding ansatz is still valid. However, due to the mean-field treatment for phonons, we find limitations of our methods in the weakly coupled delocalized region and at the Peierls transition, where quantum phonon fluctuations and long-range kinetic correlations become substantial.

Islam, Shariful [North Carolina State University, ↗