Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Embedding a critical point in a hadron to quark-gluon crossover equation of state

Lattice QCD simulations have shown unequivocally that the transition from hadrons to quarks and gluons is a crossover when the baryon chemical potential is zero or small. Many model calculations predict the existence of a critical point at a value of the chemical potential where current lattice simulations are unreliable. We show how to embed a critical point in a smooth background equation of state so as to yield the critical exponents and critical amplitude ratios expected of a transition in the same universality class as the liquid-gas phase transition and the three-dimensional Ising model. There are only two independent critical exponents; the relations α + 2β + γ = 2 and β(δ-1) = γ arise automatically, as does a relation between the two critical amplitudes. The resulting equation of state has parameters that may be inferred by hydrodynamic modeling of heavy-ion collisions in the Beam Energy Scan II at the BNL Relativistic Heavy Ion Collider or in experiments at other accelerators.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

EnZymClass: Substrate specificity prediction tool of plant acyl-ACP thioesterases based on ensemble learning

Characterizing the functional properties of plant acyl-ACP thioesterases (TEs), a key enzyme class used in the production of renewable oleochemicals in microbial hosts, experimentally, can be an expensive and time consuming process since it requires manual screening of thousands of candidates in a database. Using amino acid sequence to computationally predict an enzyme’s function might accelerate this process; however obtaining the necessary amount of information on previously characterized enzymes and their respective sequences required by standard Machine Learning (ML) based approaches to accurately infer sequence-function relationships can be prohibitive, especially with a low-throughput testing cycle. Experimental noise, unbalanced dataset where high sequence similarity does not always imply identical functional properties will further prevent robust prediction performance. Herein we present a ML method, Ensemble method for enZyme Classification (EnZymClass), that is specifically designed to address these issues. We used EnZymClass to classify TEs into short, long and mixed free fatty acid substrate specificity categories. While general guidelines for inferring substrate specificity have been proposed before, prediction of chain-length preference from primary sequence has remained elusive for plant acyl-ACP TEs. By applying EnZymClass to a subset of TEs in the ThYme database, we identified two medium chain TEs, ClFatB3 and CwFatB2, with previously uncharacterized activity in E. coli fatty acid production hosts.

59 BASIC BIOLOGICAL SCIENCES↗

Virtual Diagnostic Suite for Electron Beam Prediction and Control at FACET-II

We discuss the implementation of a suite of virtual diagnostics at the FACET-II facility currently under commissioning at SLAC National Accelerator Laboratory. The diagnostics will be used for the prediction of the longitudinal phase space along the linac, spectral reconstruction of the bunch profile, and non-destructive inference of transverse beam quality (emittance) while using edge radiation at the injector dogleg and bunch compressor locations. These measurements will be folded into adaptive feedbacks and Machine Learning (ML)-based reinforcement learning controls to improve the stability and optimize the performance of the machine for different experimental configurations. In this paper we describe each of these diagnostics with expected measurement results that are based on simulation data and discuss progress towards implementation in regular operations.

Emma, Claudio (ORCID:0000000247191831)↗

deeprob/ThioesteraseEnzymeSpecificity: EnZymClass-first-release

Characterizing the functional properties of plant acyl-ACP thioesterases (TEs), a key enzyme class used in the production of renewable oleochemicals in microbial hosts, experimentally, can be an expensive and time consuming process since it requires manual screening of thousands of candidates in a database. Using amino acid sequence to computationally predict an enzyme’s function might accelerate this process; however obtaining the necessary amount of information on previously characterized enzymes and their respective sequences required by standard Machine Learning (ML) based approaches to accurately infer sequence-function relationships can be prohibitive, especially with a low-throughput testing cycle. Experimental noise, unbalanced dataset where high sequence similarity does not always imply identical functional properties will further prevent robust prediction performance. Herein we present a ML method, Ensemble method for enZyme Classification (EnZymClass), that is specifically designed to address these issues. We used EnZymClass to classify TEs into short, long and mixed free fatty acid substrate specificity categories. While general guidelines for inferring substrate specificity have been proposed before, prediction of chain-length preference from primary sequence has remained elusive for plant acyl-ACP TEs. By applying EnZymClass to a subset of TEs in the ThYme database, we identified two medium chain TEs, ClFatB3 and CwFatB2, with previously uncharacterized activity in E. coli fatty acid production hosts.

Banerjee, Deepro↗

Nonlinear Optics Measurements in IOTA

Nonlinear integrable optics is a recently proposed accelerator lattice design approach which allows to generate an amplitude dependent tune shift which is needed in high brightness accelerators to mitigate fast coherent instabilities. Whereas usually octupoles are used to achieve this task, this concept allows doing so without exciting any resonances, in turn preventing any particle loss. The concept is based around a special magnet design, together with specific constraints on the optics of the accelerator. To study such a system, the Integrable Optics Test Accelerator (IOTA) was recently constructed and commissioned at Fermilab. For the assessment of the performance of this concept, good knowledge of the optics and the (non-)linear dynamics without the special magnet is of key importance. As such, measurements were conducted in the IOTA ring, using the captured turn-by-turn data by the beam position monitors after excitation to infer quantities such as amplitude detuning and resonance driving terms. In this note, first results of these measurements are presented.

43 PARTICLE ACCELERATORS↗

Adaptive sampling for accelerating neutron diffraction-based strain mapping *

Abstract Neutron diffraction is a useful technique for mapping residual strains in dense metal objects. The technique works by placing an object in the path of a neutron beam, measuring the diffracted signals and inferring the local lattice strain values from the measurement. In order to map the strains across the entire object, the object is stepped one position at a time in the path of the neutron beam, typically in raster order, and at each position a strain value is estimated. Typical dwell times at neutron diffraction instruments result in an overall measurement that can take several hours to map an object that is several tens of centimeters in each dimension at a resolution of a few millimeters, during which the end users do not have an estimate of the global strain features and are at risk of incomplete information in case of instruments outages. In this paper, we propose an object adaptive sampling strategy to measure the significant points first. We start with a small initial uniform set of measurement points across the object to be mapped, compute the strain in those positions and use a machine learning technique to predict the next position to measure in the object. Specifically, we use a Bayesian optimization based on a Gaussian process regression method to infer the underlying strain field from a sparse set of measurements and predict the next most informative positions to measure based on estimates of the mean and variance in the strain fields estimated from the previously measured points. We demonstrate our real-time measure-infer-predict workflow on additively manufactured steel parts—demonstrating that we can get an accurate strain estimate even with 30%–40% of the typical number of measurements—leading the path to faster strain mapping with useful real-time feedback. We emphasize that the proposed method is general and can be used for fast mapping of other material properties such as phase fractions from time-consuming point-wise neutron measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Analytic solutions for Asay foil trajectories with implications for ejecta source models and mass measurements

We consider the trajectory of an Asay foil ejecta diagnostic for scenarios where ejecta are produced at a singly shocked planar surface and fly ballistically through a perfect vacuum to the sensor. We do so by building upon a previously established mathematical framework derived for the analytic study of stationary sensors. First, we derive the momentum conservation equation for the problem, in a form amenable to accelerating sensors, in terms of a generic ejecta source model. The result is an integrodifferential equation of motion for the foil trajectory. This equation yields an easily calculable closed-form implicit solution for the foil trajectory in instant-production scenarios. From there, we derive a boundary condition that particle velocity distributions must satisfy if their associated foil trajectories are to exhibit a smooth initial acceleration, as occurs in some experiments. This condition is identical to one derived previously from a consideration of piezoelectric voltage data obtained in similar experiments. We also compare techniques for inferring accumulated ejecta masses from foil trajectories, first by deriving the exact solution, and then by quantifying the error imposed by a frequently used approximate solution (both subject to the assumption of instantaneous ejecta production). Finally, we examine the common practice of presenting inferred cumulative ejecta masses as a function of implied ejecta velocity, establishing the conditions under which this methodology is most meaningful.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Neutron-Resonance Transmission Analysis with a Compact Deuterium-Tritium Neutron Generator

Neutron Resonance Transmission Analysis (NRTA) is a spectroscopic technique which uses the resonant absorption of neutrons in the epithermal range to infer the isotopic composition of an object. This spectroscopic technique has relevance in many traditional fields of science and nuclear security. NRTA in the past made use of large, expensive accelerator facilities to achieve precise neutron beams, significantly limiting its applicability. Here, we describe a series of NRTA experiments where we use a compact, low-cost deuterium-tritium (DT) neutron generator to produce short neutron beams (2.6 m) along with a 6 Li-glass neutron detector. The time-of-flight spectral data from five elements – silver, cadmium, tungsten, indium, and 238 U – clearly show the corresponding absorption lines in the 1-30 eV range. The experiments show the applicability of NRTA in this simplified configuration, and prove the feasibility of this compact and low-cost approach. This could significantly broaden the applicability of NRTA, and make it practical and applicable in many fields, such as material science, nuclear engineering, and arms control.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

X-ray free electron laser studies of electron and phonon dynamics of graphene adsorbed on copper

Here, we report optical pumping and x-ray absorption spectroscopy experiments at the Pohang Accelerator Laboratory free electron laser that probes the electron dynamics of a graphene monolayer adsorbed on copper in the femtosecond regime. By analyzing the results with ab initio theory we infer that the excitation of graphene is dominated by indirect excitation from hot electron-hole pairs created in the copper by the optical laser pulse. However, once the excitation is created in graphene, its decay follows a similar path as in many previous studies of graphene adsorbed on semiconductors, i.e., rapid excitation of strongly coupled optical phonons and eventual thermalization. It is likely that the lifetime of the hot electron-hole pairs in copper governs the lifetime of the electronic excitation of the graphene.

36 MATERIALS SCIENCE↗

Dynamical Dark Energy Imprints in the Lyman-Alpha Forest

The nature of dark energy (DE) remains elusive, even though it constitutes the dominant energy-density component of the Universe and drives the late-time acceleration of cosmic expansion. By combining measurements of the expansion history from baryon acoustic oscillations, supernova surveys, and cosmic microwave background data, the Dark Energy Spectroscopic Instrument (DESI) Collaboration has inferred that the DE equation of state may evolve over time. The profound implications of a time-variable, ``dynamical" DE (DDE) that departs from a cosmological constant motivate the need for independent observational tests. In this work, we use cosmological hydrodynamical simulations of structure formation to investigate how DDE affects the properties of the Lyman-Alpha ``forest'' of absorption features produced by neutral hydrogen in the cosmic web. We find that DDE models consistent with the DESI constraints induce a spectral tilt in the forest transmitted flux power spectrum, imprinting a scale- and redshift-dependent signature relative to standard Lambda-CDM cosmologies. These models also yield higher intergalactic medium temperatures and reduced Lyman-Alpha opacity compared to Lambda-CDM. We discuss the observational implications of these trends as potential avenues for independent confirmation of DDE.

Garza, Diego [UC, Santa Cruz, Inst. Part. Phys.] (↗

Enhancing the Quality and Reliability of Machine Learning Interatomic Potentials through Better Reporting Practices

Recent developments in machine learning interatomic potentials (MLIPs) have empowered even nonexperts in machine learning to train MLIPs for accelerating materials simulations. However, reproducibility and independent evaluation of presented MLIP results is hindered by a lack of clear standards in current literature. In this Perspective, we aim to provide guidance on best practices for documenting MLIP use while walking the reader through the development and deployment of MLIPs including hardware and software requirements, generating training data, training models, validating predictions, and MLIP inference. We also suggest useful plotting practices and analyses to validate and boost confidence in the deployed models. Finally, we provide a step-by-step checklist for practitioners to use directly before publication to standardize the information to be reported. Altogether, we hope that our work will encourage the reliable and reproducible use of these MLIPs, which will accelerate their ability to make a positive impact in various disciplines including materials science, chemistry, and biology, among others.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Fast Adaptive Neural Control of Resonant Extraction at Fermilab

We present the development of a machine learning (ML) based regulation system for third-order resonant beam extraction in the Mu2e experiment at Fermilab. Classical and ML-based controllers have been optimized using semi-analytic simulations and evaluated in terms of regulation performance and training efficiency. We compare several controller architectures and discuss the integration of neural control into an adaptive framework. We also present progress on surrogate models that predict the controller response given a spill intensity and controller action history. To enable real-time deployment, we report progress on implementing low-latency, edge-based inference suitable for hardware-constrained environments. Our results demonstrate the feasibility and advantages of ML-based control in managing complex, time-varying physical systems, with broader implications for accelerator operations and other domains requiring fast, adaptive regulation.

Berlioz, Jose Rene [Fermilab]↗

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗

Atacama Cosmology Telescope: Constraints on prerecombination early dark energy

The early dark energy (EDE) scenario aims to increase the value of the Hubble constant (H0) inferred from cosmic microwave background (CMB) data over that found in the standard cosmological model (Λ CDM ), via the introduction of a new form of energy density in the early Universe. The EDE component briefly accelerates cosmic expansion just prior to recombination, which reduces the physical size of the sound horizon imprinted in the CMB. Previous work has found that nonzero EDE is not preferred by Planck CMB power spectrum data alone, which yield a 95% confidence level (C.L.) upper limit f EDE < 0.087 on the maximal fractional contribution of the EDE field to the cosmic energy budget. In this paper, we fit the EDE model to CMB data from the Atacama Cosmology Telescope (ACT) data release 4. We find that a combination of ACT, large-scale Planck TT (similar to WMAP), Planck CMB lensing, and BAO data prefers the existence of EDE at >99.7 % C .L . : f EDE = 0.091$^{+0.020}_{-0.036}$, with H 0 = 70.9$^{+1.0}_{-2.0}$ km/s/Mpc (both 68% C.L.). From a model-selection standpoint, we find that EDE is favored over Λ CDM by these data at roughly 3 σ significance. In contrast, a joint analysis of the full Planck and ACT data yields no evidence for EDE, as previously found for Planck alone. We show that the preference for EDE in ACT alone is driven by its TE and EE power spectrum data. The tight constraint on EDE from Planck alone is driven by its high-ℓ TT power spectrum data. Understanding whether these differing constraints are physical in nature, due to systematics, or simply a rare statistical fluctuation is of high priority. Here, the best-fit EDE models to ACT and Planck exhibit coherent differences across a wide range of multipoles in TE and EE, indicating that a powerful test of this scenario is anticipated with near-future data from ACT and other ground-based experiments.

79 ASTRONOMY AND ASTROPHYSICS↗

Impact of new physics on the JUNO-long-baseline synergy in the neutrino mass ordering determination

The determination of the neutrino mass ordering is one of the flagship goals in particle physics. A well-known and powerful synergy emerges when combining high-precision measurements of the effective atmospheric mass-squared splitting from electron antineutrino disappearance in reactor experiments with that from muon (anti)neutrino disappearance in accelerator-based long-baseline experiments. To fully exploit this synergy, percent-level precision in the atmospheric mass splitting is required—a target that JUNO is expected to achieve within a few months of data taking. This motivated the formulation of a mass ordering sum rule for neutrino disappearance channels, which shows that by combining data from T2K and NOvA with JUNO after one year of operation, the neutrino mass ordering can be determined at the 3⁢𝜎 confidence level. Since JUNO has recently started taking data, it is timely to ask whether this sum rule remains robust in the presence of new physics. We identify the necessary conditions for new physics to affect the sum rule and demonstrate that, in some cases, such effects could lead to an incorrect inference of the mass ordering. As concrete examples, we consider scalar nonstandard interactions (SNSI) and neutrinos coupled to an ultralight scalar field. We find that, for SNSI, current constraints render any modification of the sum rule negligible, whereas in the latter case, the inference of the ordering requires caution. Nevertheless, these effects can be disentangled, illustrating how the sum rule can also be used to search for new physics.

Alves, Gustavo F. S. [Fermi National Accelerator L↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Comment on "Table-Top Laser-Based Source of Femtosecond, Collimated, Ultrarelativistic Positron Beams"

Sarri et al. have reported the generation of low divergence (~3 mrad), high-density (10 14 cm –3 ) positron beams using millimeter-scale converter targets and a 50 pC, 200 MeV laser-wakefield accelerated (LWFA) electron source. It was argued that the positron divergence was dominated by the pair-production birth cone angle $θ_{e+} ≈ 1/γ_{e–}$. The small, energy-independent divergence value was used to infer a positron beam density of 2 × 10 14 cm –3 from a 4.2 mm Ta converter target where the divergence and yield measurements agreed with their Monte Carlo (MC) simulations. We have repeated these simulations using experimental conditions and disagree with the reported density, divergence, and yield values by up to factors > 50. In our work, we find that a divergence on the order of milliradian is not physical and can only be achieved if inelastic particle scattering is omitted in the calculations.

47 OTHER INSTRUMENTATION↗