Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Model Counting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Computer vision-based rock bolt detection in orthomosaic imagery obtained in the Waste Isolation Pilot Plant underground facility

Assessing structural integrity of large underground tunnel facilities is often a time consuming and human-labor intensive task. Thus, research using various modes of sensing and automated detection of key structural components in mines is posed to aid in safety assessments and establishing overall structural health. We propose an approach utilizing off-the-shelf camera and lidar technology fixed to a custom sensing platform, image stitching techniques, fine-tuned object detection models, and specialized model-inference methods to automatically detect, count, and map roof bolts for assessment of structural safety in man-made underground tunnels. Results show a novel workflow for effective object counting in orthomosaic tunnel ceiling images generated from collections in GPS-denied mining environments. Additionally, we demonstrate effective fine-tuning of EfficientDet object detectors utilizing state-of-the-art image augmentation techniques known as the mosaic and mixup transformations. Our work is demonstrated on sensed data and imagery collected from the Department of Energy (DOE) Waste Isolation Pilot Plant (WIPP) where miles of tunnel ceiling must be assessed for structural integrity.

42 ENGINEERING

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.

Chitty-Venkata, Krishna Teja

On the Prospect of Chemically Transferable Coarse-Grained Electronic Models for Soft Materials

Electronic coarse-graining (ECG) methods predict quantum-mechanical electronic properties directly from coarse-grained (CG) molecular configurations, enabling electronic predictions at mesoscale length scales. Here, we present a diagnostic assessment of the feasibility of chemically transferable ECG models across a broad polymer-relevant chemical space using all-atom, united-atom, and Martini-scale representations. While high-resolution ECG models achieve near-quantitative accuracy, we show that chemically transferable ECG at the Martini resolution fails because the CG force field does not sample the same configurational distribution of local molecular structure as that underlying the DFT-parameterized ECG model. We demonstrate that our proposed Element-Count-Label (ECL) representation, which augments Martini beads with explicit stoichiometric data, significantly improves chemical generalization across diverse polymer chemistries. However, we find that even with improved chemical resolution, the model cannot recover electronic property distributions that are absent from the configurational space sampled by the CG force field. These results demonstrate that chemically transferable ECG requires future Martini-like force fields to explicitly preserve quantum chemistry–compatible local molecular structure in addition to thermodynamic and structural fidelity.

Kidder, Katherine M [Department of Chemistry; Univ

A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing

Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This paper introduces an alternative strategy, called phantom parallelism, to minimize the net energy consumption of traditional tensor (model) parallelism, the most energy-inefficient component of large neural network training. The approach is presented in the context of feed-forward network architectures as a preliminary, but comprehensive, proof-of-principle study of the proposed methodology. We derive new forward and backward propagation operators for phantom parallelism, implement them as custom autograd operations within an end-to-end phantom parallel training pipeline and compare its parallel performance and energy-efficiency against those of conventional tensor parallel training pipelines. Formal analyses that predict lower bandwidth and FLOP counts are presented with supporting empirical results on up to 256 GPUs that corroborate these gains. Experiments are shown to deliver ∼50% reduction in the energy consumed to train FFNs using the proposed phantom parallel approach when compared with conventional tensor parallel methods. Additionally, the proposed approach is shown to train smaller phantom models to the same model loss on smaller GPU counts as larger tensor parallel models on larger GPU counts offering the possibility for even greater energy savings.

Seal, Sudip [ORNL] (ORCID:0000000332330656)

EAGLE-I County Customer Dataset Fall 2025

This dataset provides a combination of modeled and collected county-level electric customer counts derived from 2023 EIA-861 utility customer data, 2021 HIFLD electric retail service territory boundaries, 2021 LandScan population estimates, and 2025 EAGLE-I customer outages. The dataset details county FIPS code, number of customers, and customer type (modeled, collected, mixed). Outage data in included for all 50 U.S. states, Puerto Rico, and the District of Columbia (excluding other U.S. territories).

24 POWER TRANSMISSION AND DISTRIBUTION

Small-scale signatures of primordial non-Gaussianity in k-nearest neighbour cumulative distribution functions

ABSTRACT Searches for primordial non-Gaussianity in cosmological perturbations are a key means of revealing novel primordial physics. However, robustly extracting signatures of primordial non-Gaussianity from non-linear scales of the late-time Universe is an open problem. In this paper, we apply k-Nearest Neighbour cumulative distribution functions, kNN-CDFs, to the quijote-png simulations to explore the sensitivity of kNN-CDFs to primordial non-Gaussianity. An interesting result is that for halo samples with $M_\mathrm{ h}\langle 10^{14}$ M$_\odot$ $h^{-1}$, the kNN-CDFs respond to equilateral PNG in a manner distinct from the other parameters. This persists in the galaxy catalogues in redshift space and can be differentiated from the impact of galaxy modelling, at least within the halo occupation distribution (HOD) framework considered here. kNN-CDFs are related to counts-in-cells and, through mapping a subset of the kNN-CDF measurements into the count-in-cells picture, we show that our results can be modelled analytically. A caveat of the analysis is that we only consider the HOD framework, including assembly bias. It will be interesting to validate these results with other techniques for modelling the galaxy–halo connection, e.g. (hybrid) effective field theory or semi-analytical methods.

Coulton, William R. (ORCID:0000000212973673)

Absolute decay counting of 146 Sm with 4π cryogenic microcalorimetry

We present a methodology for absolute activity counting of long-lived isotopes based on cryogenic Decay Energy Spectroscopy. A 146 Sm source was produced at the TRIUMF Laboratory and then processed and purified at Lawrence Livermore National Laboratory, yielding a pure sample. The source was embedded within a 4π thermal absorber coupled to a magnetic microcalorimeter achieving nearly 100% counting efficiency. Experimental uncertainties were studied and modeled, including thermal coupling of the source to the absorber, pulse pile-up, trigger, and event selection efficiencies. Here, the absolute activity of the pure 146 Sm source was measured to better than 1% uncertainty.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Effect of likelihood misspecification in Gaussian process-driven autonomous experimentation

In recent years, several groups have designed Autonomous Experiment (AE) models with the aim of using them as an alternative method for neutron scattering scanning. In an AE, Gaussian processes (GPs) are most frequently used due to their interpretability, their non-parametric nature, their universal approximation, and their closed-form predictive distribution. GPs have two key components, namely, the model for the likelihood of a neutron count knowing the underlying dynamic structure factor and the acquisition function. In this paper, we investigate the impact, on the quality of an AE, of the likelihood and acquisition function choices, in energy scans and (Q, ω) ones, with respect to the signal-over-noise ratio. While we hypothesized that the quality of GP predictions would decrease when the normal to Poisson likelihood approximation breaks down at low count rates, we found that the use of the correct Poisson likelihood does not improve the quality of the data collected, as well as yields very poor results in (Q, ω) scans at low count rates. In fact, the best results are obtained with a combination of normal likelihood, including the observation noise, and the change in variance acquisition function. In addition, we find that the performance, or quality of the predictive distribution, is a misleading measure of efficiency, that is, of the quality of the data collected.

Perryman, David Elliott [Inst. Laue-Langevin (ILL)

Fault-tolerant resource comparison of qudit and qubit encodings for diagonal quadratic operators

Finite local Hilbert-space truncations arise naturally in quantum simulations of lattice field theories and motivate qudit encodings, but their fault-tolerant advantage over qubit encodings remains unclear. We compare the non-Clifford cost of implementing quadratic diagonal evolutions, exemplified by 𝑈 = 𝑒$^{−𝑖⁢𝑡⁢𝜙^2_𝑥}$ in a uniform field-amplitude discretization of a real scalar field, using either one logical 𝑑-level qudit or 𝑛 𝑏 = ⌈log 2⁡ 𝑑⌉ logical qubits. We analyze two standard settings: product-formula simulation and linear combination of unitaries (LCU) per block encoding, taking the resource metric to be the number of non-Clifford gates after synthesis into a discrete logical gate set. Because tight synthesis bounds for general single-qudit rotations are not known, we express the qudit constructions in terms of embedded two-level SU⁡(2) rotations and derive explicit finite-𝑑 break-even conditions for their synthesis cost; these serve as compiler targets for when qudit encodings can outperform the qubit baseline. Within the constructive models studied here, product-formula implementations would require an exponentially stronger per-primitive synthesis advantage for qudits to win asymptotically, while in the LCU setting the qubit encoding is asymptotically cheaper in 𝑑. Nevertheless, the finite-𝑑 threshold analysis identifies low-dimensional regions in which qudits can yield meaningful constant-factor savings, particularly for LCU-based implementations. As a secondary analysis of the LCU construction, we use an idealized negligible-overhead qubit-qudit code-switching model to give an absolute 𝑇-count comparison and reinterpret the savings as an allowable per-switch overhead budget.

Godwood, Samuel [Univ. of Liverpool (United Kingdo

PhotonIDs: ML-Powered Photon Identification System for Dark Count Elimination

Reliable single photon detection is the foundation for practical quantum communication and networking. However, today's superconducting nanowire single photon detector(SNSPD) inherently fails to distinguish between genuine photon events and dark counts, leading to degraded fidelity in long-distance quantum communication. In this work, we introduce PhotonIDs, a machine learning-powered photon identification system that is the first end-to-end solution for real-time discrimination between photons and dark count based on full SNSPD readout signal waveform analysis. PhotonIDs ~demonstrates: 1) an FPGA-based high-speed data acquisition platform that selectively captures the full waveform of signal only while filtering out the background data in real time; 2) an efficient signal preprocessing pipeline, and a novel pseudo-position metric that is derived from the physical temporal-spatial features of each detected event; 3) a hybrid machine learning model with near 98% accuracy achieved on photon/dark count classification. Additionally, proposed PhotonIDs ~ is evaluated on the dark count elimination performance with two real-world case studies: (1) 20 km quantum link, and (2) Erbium ion-based photon emission system. Our result demonstrates that PhotonIDs ~could improve more than 31.2 times of signal-noise-ratio~(SNR) on dark count elimination. PhotonIDs ~ marks a step forward in noise-resilient quantum communication infrastructure.

Linne, Karl C. [Chicago U.] (ORCID:000900091870358

Rapid neutron and gamma-ray source localization using machine learning

Rapid localization of radiation sources is critical for applications including nuclear emergency response, safeguards, and security. However, conventional imaging systems such as neutron scatter cameras and Compton cameras depend on rare coincidence events, which often result in long acquisition times. In this work, we address the challenge of rapid source localization by developing a machine learning approach to predict the direction of a single radiation source using only count rates from an array of neutron and gamma-ray detectors. The proposed model is a fully connected neural network (FCNN) trained using Monte Carlo simulation data from a 252 Cf source. The model hyperparameters are optimized with a small set of routine 252 Cf measurements. We benchmarked the performance of the trained and optimized machine learning model using additional 252 Cf , 137 Cs , and PuBe measurements under laboratory conditions with varying source-detector configurations. For these measurements, the machine learning model achieved a mean localization error smaller than 30° with 3 x 10 3 system counts, corresponding to 8 s measurement time for the imaging system used in this work. In this low-statistics regime, the method outperformed traditional scatter-based imaging by more than 75% in localization accuracy for the evaluated measurement configurations. These results demonstrate that a machine learning-based approach can significantly reduce the time required for accurate single-source localization, providing a robust and computationally efficient alternative to traditional imaging systems in time-critical nuclear security and emergency response scenarios.

Gamma-ray imaging

FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging

We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.

47 OTHER INSTRUMENTATION

Bayesian chain graph models to characterize microbe-environment dynamics

Microbiome data require statistical models that can simultaneously decode microbes' reaction to the environment and interactions among microbes. While a multiresponse linear regression model seems like a straight-forward solution, we argue that treating it as a graphical model is problematic given that the regression coefficient matrix does not encode the conditional dependence structure between response and predictor nodes. This observation is especially important in biological settings when we have prior knowledge on the edges from specific experimental interventions that can only be properly encoded under a conditional dependence model. Here, we propose a chain graph model with two sets of nodes (predictors and responses) whose solution yields a graph with edges that indeed represent conditional dependence, thus agreeing with the experimenter's intuition on the average behavior of nodes under treatment. The solution to our model is sparse via the Bayesian linear regression (LASSO). In addition, we propose an adaptive extension so that different shrinkages can be applied to different edges to incorporate edge-specific prior knowledge. Our model is computationally inexpensive through an efficient Gibbs sampling algorithm and can account for binary, counting, and compositional responses via an appropriate hierarchical structure. We test the performance of our model in a variety of simulated datasets, thereby showing superior performance to state-of-the-art approaches. We further apply our model to human gut and soil microbial compositional datasets, and we highlight that CG-LASSO can estimate biologically meaningful network structures in the data.

compositional data

From RNNs to Foundation Models: An Empirical Study on Commercial Building Energy Consumption

Accurate short-term energy consumption forecasting for commercial buildings is crucial for smart grid operations. While smart meters and deep learning models enable forecasting using past data from multiple buildings, data heterogeneity from diverse buildings can reduce model performance. The impact of increasing dataset heterogeneity in time series forecasting, while keeping size and model constant, is understudied. We tackle this issue using the ComStock dataset, which provides synthetic energy consumption data for U.S. commercial buildings. Two curated subsets, identical in size and region but differing in building type diversity, are used to assess the performance of various time series forecasting models, including finetuned open-source foundation models (FMs). The results show that dataset heterogeneity and model architecture have a greater impact on post-training forecasting performance than the parameter count. Moreover, despite the higher computational cost, finetuned FMs demonstrate competitive performance compared to base models trained from scratch.

commercial buildings

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling, we study representational differences between LLMs and their smaller counterparts, with the goal of replicating the representational qualities of larger models in smaller models. We observe a geometric phenomenon which we term embedding condensation, where token embeddings collapse into a narrow cone-like subspace in some language models. Through systematic analyses across multiple Transformer families, we show that small models such as GPT2 and Qwen3-0.6B exhibit severe condensation, whereas larger models such as GPT2-x1 and Qwen3-32B are more resistant to this phenomenon. Additional observations show that embedding condensation is not reliably mitigated by knowledge distillation from larger models. To fight against it, we formulate a dispersion loss that explicitly encourages embedding dispersion during training. Experiments demonstrate that it mitigates condensation, recovers dispersion patterns seen in larger models, and yields performance gains across 10 benchmarks. We believe this work offers a principled path toward improving smaller Transformers without additional parameters.

Xiao, Xi [ORNL] (ORCID:0009000009316982)

Beyond real: alternative unitary cluster Jastrow models for molecular electronic structure calculations on near-term quantum computers

Near-term quantum devices require wavefunction ansätze that are expressive while also of shallow circuit depth in order to both accurately and efficiently simulate molecular electronic structure. While the unitary coupled cluster ansatz (e.g., UCCSD) has become a standard, the high gate count associated with the implementation of this limits its feasibility on noisy intermediate-scale quantum (NISQ) hardware. k -Fold unitary cluster Jastrow (uCJ) ansätze mitigate this challenge by providing O( kN 2 ) circuit scaling and favorable linear depth circuit implementation. Previous work has focused on the real orbitalrotation (Re-uCJ) variant of uCJ, which allows an exact (Trotter-free) implementation. Here we extend and generalize the k -fold uCJ framework by introducing two new variants, Im-uCJ and g-uCJ, which incorporate imaginary and fully complex orbital rotation operators, respectively. Similar to Re-uCJ, both of the new variants achieve quadratic gate-count scaling. Our results focus on the simplest k = 1 model, and show that the uCJ models frequently maintain energy errors within chemical accuracy (∼1 kcal mol −1 ). Both g-uCJ and Im-uCJ are more expressive in terms of capturing electron correlation and are also more accurate than the earlier Re-uCJ ansatz. We further show that Im-uCJ and g-uCJ circuits can also be implemented exactly, without any Trotter decomposition. Numerical tests using k = 1 on H 2 , H 3 + , Be 2 , C 2 H 4 , C 2 H 6 and C 6 H 6 in various basis sets confirm the practical feasibility of these shallow Jastrow-based ansätze for applications on near-term quantum hardware.

Tkachenko, Nikolay V. [University of California, B

Impact of dynamics, entanglement and Markovian noise on the fidelity of few-qubit digital quantum simulation

Quantum algorithms have been proposed to accelerate the simulation of the chaotic dynamical systems that are ubiquitous in the physics of plasmas. Quantum computers without error correction might even use noise to their advantage to calculate the Lyapunov exponent by measuring the Loschmidt echo fidelity decay rate. For the first time, digital Hamiltonian simulations of the quantum sawtooth map, performed on the IBM-Q quantum hardware platform, show that the fidelity decay rate of a digital quantum simulation increases during the transition from dynamical localization to chaotic diffusion in the map. The observed error per CNOT gate increases by $1.5{\times }$ as the dynamics varies from localized to diffusive, while only changing the phases of virtual RZ gates and keeping the overall gate count constant. A gate-based Lindblad noise model that captures the effective change in relaxation and dephasing errors during gate operation qualitatively explains the effect of dynamics on fidelity as being due to the localization and entanglement of the states created. Specifically, highly delocalized states that are entangled with random phases show an increased sensitivity to dephasing and, on average, a similar sensitivity to relaxation as localized states. In contrast, delocalized unentangled states show an increased sensitivity to dephasing but a lower sensitivity to relaxation. This gate-based Lindblad model is shown to be a useful benchmarking tool by estimating the effective Lindblad coherence times during CNOT gates and finding a consistent $2\unicode{x2013}3{\times }$ shorter $T_2$ time than reported for idle qubits. Thus, the interplay of the dynamics of a simulation with the noise processes that are active can strongly influence the overall fidelity decay rate.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Seed classification with random forest models

Premise: To improve forest conservation monitoring, we developed a protocol to automatically count and identify the seeds of plant species with minimal resource requirements, making the process more efficient and less dependent on human operators. Methods and Results: Seeds from six North American conifer tree species were separated from leaf litter and imaged on a flatbed scanner. In the most successful species-classification approach, an ImageJ macro automatically extracted measurements for random forest classification in the software R. The method allows for good classification accuracy, and the same process can be used to train the model on other species. Conclusions: This protocol is an adaptable tool for efficient and consistent identification of seed species or potentially other objects. Automated seed classification is efficient and inexpensive, making it a practical solution that enhances the feasibility of large-scale monitoring projects in conservation biology.

59 BASIC BIOLOGICAL SCIENCES