Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “task tuning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Domain-specific text embedding model for accelerator physics

Accelerator physics presents unique challenges for natural language processing (NLP) due to its specialized terminology and complex concepts. A key component in overcoming these challenges is the development of robust text embedding models that transform textual data into dense vector representations, facilitating efficient information retrieval and semantic understanding. In this work, we introduce AccPhysBERT, a sentence embedding model fine-tuned specifically for accelerator physics. Our model demonstrates superior performance across a range of downstream NLP tasks, surpassing existing models in capturing the domain-specific nuances of the field. We further showcase its practical applications, including semantic paper-reviewer matching and integration into retrieval-augmented generation systems, highlighting its potential to enhance information retrieval and knowledge discovery in accelerator physics. Published by the American Physical Society 2025

Hellert, Thorsten (ORCID:0000000227970926)↗

The effect of lighting environment on task performance in buildings – A review

The effects of indoor environmental conditions on human health, satisfaction, and performance have been the focal point of research for decades. This paper reviews and summarizes the impact of lighting environment on task performance, specifically for the built environment audience. Existing studies included a variety of performance tests on cognitive performance and perception, visual acuity and reaction, memory, reasoning, and labor productivity. Illuminance, luminance ratio and correlated color temperature were found to affect performance in different ways, reflecting the impact of experimental techniques, conditions, performance evaluation methods used and data analysis methods. These were reviewed and categorized, with discussion on limitations related to sample size, modeling approach, carryover effects and other factors affecting individual differences in performance, with recommendations for future improvement. Although no universal conclusions can be made, in general, task performance seems to improve with higher illuminances, contrast ratios in the range of 7–11:1 (while always making sure that glare will not occur in the space) and higher correlated color temperature, while spectral tuning in the red or blue wavelengths has also shown positive effects. To obtain more generic evidence, future studies should be more consistent in terms of experimental procedures and overall light conditions, and also consider the effects of vertical illuminance, daylight provision/control, and outside views on task performance. Finally, studying performance with multi-factorial designs in a human-centered optimized manner (such as deploying variable lighting scenarios optimized for various tasks) can lead to deeper understanding of lighting effects on task performance, and ultimately to improved lighting design and operation in buildings overall.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Data Agnostic Feature-Target Analysis & Ranking Machine Learning Pipeline (DAFTAR-ML) v0.1.0

DAFTAR-ML is a specialized machine-learning pipeline that identifies relevant features based on their relationship to a target variable. Many ML pipelines focus solely on prediction, and feature ranking is often absent or lacks robust statistical methods. DAFTAR-ML performs its tasks with this outcome in mind. Model training is robust, using nested cross-validation and hyperparameter tuning. Instead of relying on native feature-importance scores, it employs SHAP (SHapley Additive exPlanations) to quantify feature importance. The pipeline also produces comprehensive results, including publication-quality visualizations.

Melie, Tina [Lawrence Berkeley National Laboratory↗

Can Large Language Models Understand Intermediate Representations?

Intermediate Representations (IRs) are essential in compiler design and program analysis, yet their comprehension by Large Language Models (LLMs) remains underexplored. This paper presents a pioneering empirical study to investigate the capabilities of LLMs, including GPT-4, GPT-3, Gemma 2, LLaMA 3.1, and Code Llama, in understanding IRs. We analyze their performance across four tasks: Control Flow Graph (CFG) reconstruction, decompilation, code summarization, and execution reasoning. Our results indicate that while LLMs demonstrate competence in parsing IR syntax and recognizing high-level structures, they struggle with control flow reasoning, execution semantics, and loop handling. Specifically, they often misinterpret branching instructions, omit critical IR operations, and rely on heuristic-based reasoning, leading to errors in CFG reconstruction, IR decompilation, and execution reasoning. The study underscores the necessity for IR-specific enhancements in LLMs, recommending fine-tuning on structured IR datasets and integration of explicit control flow models to augment their comprehension and handling of IR-related tasks.

Jiang, Hailong↗

Swap Path Network for Robust Person Search Pre-training

This code corresponds to the WACV25 conference paper, "Swap Path Network for Robust Person Search Pre-training". In that paper, we introduce a new model for the person search task called the Swap Path Net (SPNet). The person search task is a problem in computer vision, where we locate and rank matches to an image of a query person in a set of other images where we want to find them. We also introduce a novel pre-training algorithm specific to the Swap Path Net architecture. The code implements pre-training and fine-tuning of the Swap Path Net (SPNet). This includes ingesting image datasets and updating the weights of the SPNet neural network to train it for the person search task. The repository contains code, configs, and instructions to reproduce all results from the paper.

Jaffe, LucasW [Lawrence Livermore National Laborat↗

A Deep State Space Model for Rainfall‐Runoff Simulations

The classical way of studying the rainfall‐runoff processes in the water cycle relies on conceptual or physically‐based hydrologic models. Deep learning (DL) has recently emerged as an alternative and blossomed in the hydrology community for rainfall‐runoff simulations. However, the decades‐old Long Short‐Term Memory (LSTM) network remains the benchmark for this task, outperforming newer architectures like Transformers. In this work, we propose a State Space Model (SSM), specifically the Frequency Tuned Diagonal State Space Sequence (S4D‐FT) model, for rainfall‐runoff simulations. The proposed S4D‐FT is benchmarked against the established LSTM and a physically‐based Sacramento Soil Moisture Accounting model under in‐sample and out‐of‐sample simulation setups across 531 watersheds in the contiguous United States (CONUS). Results show that S4D‐FT is able to outperform the LSTM model across diverse regions under both simulation setups, especially for regions that feature snowmelt‐driven or intermittent flow regimes. In contrast, S4D‐FT tends to underperform in flashier, high‐magnitude flow regimes, likely due to its global state‐space convolution computation that emphasizes slow, storage‐driven dynamics, which makes it less effective at picking up short bursts and noisy spikes in the data. In summary, our pioneering introduction of the S4D‐FT for rainfall‐runoff simulations challenges the dominance of LSTM in the hydrology community and expands the arsenal of DL tools available for hydrological modeling.

Wang, Yihan [Univ. of Oklahoma, Norman, OK (United↗

Design of Hopfield Networks Based on Superconducting Coupled Oscillators

The global energy shortage has driven the development of many energy-efficient computational platforms beyond Moore's law, among which brain-inspired neuromorphic computing is one of the promising solutions. Associative memory and pattern recognition are important computations solved by brain-inspired Hopfield networks. Classical Hopfield networks store memories via fixed point attractors of their dynamics. In oscillatory Hopfield networks, these attractors are replaced by periodic orbits. Here, we design an oscillatory Hopfield network based on coupled superconducting oscillators. We first employ a mathematical phase reduction approach to map networks of coupled superconducting rapid single flux quantum (RSFQ) ring oscillators to coupled Kuramoto phase-oscillator networks. We use this theory to numerically optimize the hardware's mutual inductances in order to directly match the phase-reduced superconducting oscillators to a model of phase-oscillator-based Hopfield networks. The resulting network can store multiple oscillatory phase-locked memory patterns and recover the patterns based on the initial phase conditions. As different pattern recognition tasks, or learning, require tunable connectivity strengths between the oscillatory nodes, we further employ a coupler circuit that enables tuning the coupling strength between two oscillators by applying an external flux. We demonstrate the functionality of our design through numerical simulations of a small example network with oscillators operating at 86 GHz and recognizing patterns within 10 ns. Our approach enables the learning and retrieval of dynamical memory patterns with a wide range of applications where rhythmic dynamic output is beneficial.

Cheng, Ran↗

Enhancing Molecular Isotope Detection with a Chirp Tuned Optical Centrifuge

This paper gives a brief experimental overview on creating a chirp-tunable optical centrifuge with the goal of amplifying the isotope shift of a given molecule. Results for this experiment were recorded with a CO 2 plume as the probed material. Beyond amplification of the isotope shift, this technology could be modified to alter magnetic field characteristics of common molecules as well as simulating high temperature gases found in combustion environments. Our specific design uses double mirror and grating stretcher and will include a unique tuning feature for customization of the beam’s red and blue components' temporal chirp, allowing for quick reactivity to varying task requirements. The design and parameters will be discussed, as well as the methods used to model the optical centrifuge. Although this experiment is ongoing, results have been gathered previously that display the efficacy of the optical centrifuge. For the probed CO 2 sample, Raman Shifts centered around a wavenumber of 60 and 275 were recorded with durations of 1000 and 800 picoseconds, respectively.

47 OTHER INSTRUMENTATION↗

Transferring a Molecular Foundation Model for Polymer Property Predictions

Transformer-based large language models have remarkable potential to accelerate design optimization for applications such as drug development and material discovery. Self-supervised pretraining of transformer models requires large-scale data sets, which are often sparsely populated in topical areas such as polymer science. Further, state-of-the-art approaches for polymers conduct data augmentation to generate additional samples but unavoidably incur extra computational costs. In contrast, large-scale open-source data sets are available for small molecules and provide a potential solution to data scarcity through transfer learning. In this work, we show that using transformers pretrained on small molecules and fine-tuned on polymer properties achieves comparable accuracy to those trained on augmented polymer data sets for a series of benchmark prediction tasks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fermilab Booster loss modelling and rebalancing using Bayesian methods

Fermilab Booster is being upgraded for the PIP-II project to support 20Hz ramp rate at higher intensities. Loss trip limits determine the achievable peak power. To meet PIP-II requirements, losses need to be halved as compared to current levels. Losses primarily occur at injection and transition crossing, with both gradually increasing and threshold-like intensity-dependent behaviors. The existing simulation models are not yet good enough for quantitative loss predictions. In practice, it will be necessary to tune up the Booster using iterative methods and operator intuition. In this paper we present an effort to systematically model Booster losses using active learning (Bayesian exploration) techniques, and subsequently to rebalance them for higher trip limit margins. We first created several sets of spatially and temporally isolated orbit and optics knobs, and trained Gaussian process models for each beam loss monitor as well as beam current. This is a complex task due to safety and timing requirements – we discuss mitigations such as uncertainty constraints and approximate fitting. Once models are stable, we perform large-scale single and multi-objective tuning using scalarized objectives made up of critical beam loss locations. Our results demonstrate significant rebalancing of losses, increasing trip margins, as well as an overall improvement in beam transmission efficiency. We are exploring how to combine existing simulations with experimental data and automate the collection procedure so that more advanced surrogate models can be created over time.

Kuklev, Nikita [Fermilab]↗

Foundation models for atomistic simulation of chemistry and materials

Conventional computational methods for modeling chemical and materials systems are limited by system size and timescale, forcing a trade-off between quantum-mechanical accuracy and the sampling needed for realistic observables. Large language and vision foundation models — pre-trained on massive datasets using transformer architectures — have revolutionized many fields. It is thus interesting to ask whether a foundation model — subject to suitable data, parameter scaling and training — could enable learned simulations of chemistry and materials. Here, in this study, we review the field of machine-learned interatomic potentials (MLIPs) and posit that scaling up large and diverse chemical and materials datasets and highly expressive architectures using advanced training strategies should result in models that are: more efficient, transferable, robust to out-of-distribution scenarios, and easier to fine-tune to a variety of downstream physical observables than models trained from scratch on small datasets corresponding to specific, targeted atomistic simulation tasks. We provide specific criteria for creating such large-scale MLIP foundation models, coordinated strategies for their development, evaluation and deployment, and highlight potential emergent capabilities that could transform predictive simulations in chemistry and materials science and accelerate discovery across multiple technological domains.

Yuan, Eric C.-Y. [University of California, Berkel↗

Fully Convolutional Spatio-Temporal Models for Representation Learning in Plasma Science

We have trained a fully convolutional spatio-temporal model for fast and accurate representation learning in the challenging exemplar application area of fusion energy plasma science. The onset of major disruptions is a critically important fusion energy science issue that must be resolved for advanced tokamak plasmas such as the $25B burning plasma international thermonuclear experimental reactor (ITER) experiment. While a variety of statistical methods have been used to address the problem of tokamak disruption prediction and control, recent approaches based on deep learning have proven particularly compelling. In the present paper, we introduce further improvements to the fusion recurrent neural network (FRNN) software suite, which delivered cross-machine disruption predictions with unprecedented accuracy using a large database of experimental signals from two major tokamaks. Up to now, FRNN was based on the long short-term memory (LSTM) variant of recurrent neural networks to leverage the temporal information in the data. Here, we implement and apply the "temporal convolutional neural network (TCN)" architecture to the time-dependent input signals. Furthermore, this allows highly optimized convolution operations to carry the majority of the computational load of training, thus enabling a reduction in training time, and the effective use of high-performance computing resources for hyperparameter tuning. At the same time, the TCN-based architecture achieves better predictive performance when compared with the LSTM architecture for various tasks for a representative fusion database.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Nonlinear Optics Measurements in IOTA

Nonlinear integrable optics is a recently proposed accelerator lattice design approach which allows to generate an amplitude dependent tune shift which is needed in high brightness accelerators to mitigate fast coherent instabilities. Whereas usually octupoles are used to achieve this task, this concept allows doing so without exciting any resonances, in turn preventing any particle loss. The concept is based around a special magnet design, together with specific constraints on the optics of the accelerator. To study such a system, the Integrable Optics Test Accelerator (IOTA) was recently constructed and commissioned at Fermilab. For the assessment of the performance of this concept, good knowledge of the optics and the (non-)linear dynamics without the special magnet is of key importance. As such, measurements were conducted in the IOTA ring, using the captured turn-by-turn data by the beam position monitors after excitation to infer quantities such as amplitude detuning and resonance driving terms. In this note, first results of these measurements are presented.

43 PARTICLE ACCELERATORS↗

Decoding Golden Eagle Movement Behavior from High-Resolution, Variable-Rate Telemetry Data Through Bayesian Filtering

The recent advances in animal tracking technology have enabled the collection of a vast amount of in situ data regarding the movement of wildlife at high spatiotemporal resolution. These data are usually available at variable time resolutions and contains noise (error) originating from GPS fixes. Decoding movement characteristics, particularly of flying animals, from telemetry data while handling these factors is a challenging yet important task for conservation purposes. Typically, this task is broken into two subtasks: resampling, and model calibration. The resampling subtask converts the variable rate positional data into a constant time interval data, while the model calibration subtask uses the resampled data to tune time-invariant parameters of the proposed models. For telemetry data at high temporal resolutions (order of 1 second), it is very challenging to decouple noise from actual movements using interpolation-based resampling techniques. Any errors introduced during resampling can significantly alter the the calibration and prediction attributes of the movement model. We address this problem through a unified Bayesian state-space framework that can handle both the resampling and calibration tasks in a single step. In addition, we use the speed and heading of the bird from telemetry data to regularize the position information of the bird. We use a Kalman filtering approach to include these nonlinearly related motion parameters within the state space framework. We cross-validated to quantify how this inclusion affects the model performance in estimating true bird movements. The relationship between the true state of the bird and environmental and topographical covariates is then represented parametrically. These parameters are then tuned using stochastic sampling strategies like Markov Chain Monte Carlo (MCMC). We use the telemetry data collected from golden eagles in the western USA to demonstrate the applicability of this approach to build a predictive, probabilistic movement model. Our preliminary results show that this approach provides improved predictive performance in terms of capturing higher-order motion parameters such as angular and horizontal accelerations, which may have simpler and more direct relationships with environmental covariates than corresponding speeds. In this talk, we will demonstrate how this state-space approach benefits the prediction capabilities of a movement model in simulating golden eagle paths through a wind power plant in Wyoming given certain atmospheric conditions. The model outcomes are aimed at informing mitigation strategies that can minimize the potential for collisions of golden eagles with wind turbines.

Bayesian methods↗

Benchmarking materials property prediction methods: the Matbench test set and Automatminer reference algorithm

Abstract We present a benchmark test suite and an automated machine learning procedure for evaluating supervised machine learning (ML) models for predicting properties of inorganic bulk materials. The test suite, Matbench, is a set of 13 ML tasks that range in size from 312 to 132k samples and contain data from 10 density functional theory-derived and experimental sources. Tasks include predicting optical, thermal, electronic, thermodynamic, tensile, and elastic properties given a material’s composition and/or crystal structure. The reference algorithm, Automatminer, is a highly-extensible, fully automated ML pipeline for predicting materials properties from materials primitives (such as composition and crystal structure) without user intervention or hyperparameter tuning. We test Automatminer on the Matbench test suite and compare its predictive power with state-of-the-art crystal graph neural networks and a traditional descriptor-based Random Forest model. We find Automatminer achieves the best performance on 8 of 13 tasks in the benchmark. We also show our test suite is capable of exposing predictive advantages of each algorithm—namely, that crystal graph methods appear to outperform traditional machine learning methods given ~10 4 or greater data points. We encourage evaluating materials ML algorithms on the Matbench benchmark and comparing them against the latest version of Automatminer.

36 MATERIALS SCIENCE↗

HydraGNN_OPF_GFM_2026 - Ensemble of predictive graph foundation models for power grid applications

This dataset supports research on graph foundation models for optimal power flow (OPF) on electric grids using HydraGNN. It contains heterogeneous graph representations of PGLib-OPF cases spanning systems from 14 to 13,659 buses, together with packed HDF5 datasets for pretraining, feasibility classification, and N-1 contingency analysis. The release includes OPF solution data, downstream fine-tuning datasets, pretrained HeteroSAGE and HeteroHEAT model checkpoints, hyperparameter-optimization summaries across multiple heterogeneous GNN architectures, and aggregated fine-tuning results for sample-efficiency studies. The dataset is designed to enable scalable training, evaluation, and transfer-learning studies for OPF surrogate modeling, including node-level AC-OPF solution prediction, graph-level prediction, feasibility classification, operating-condition generalization, and contingency-response tasks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Generative large language models for predictive maintenance planning

Maintenance planning and the generation of necessary components for tasks can prove time-consuming and complex. Automating the creation of recurring or similar tasks by leveraging previous planning packages and data, while uncovering insights to automate planning package generation, presents an opportunity to conserve valuable time and resources. This work aims to harness the textual and probabilistic capabilities of large language models (LLMs) to automate the generation of planning packages. Utilizing diverse data sources ranging from raw data to handwritten text, both singular and collaborative LLMs are trained and tested. Results demonstrate their capability to generate essential planning package components, effectively replicating the statistical patterns in the data. This demonstrates the use of these tools inside a digital asset for automated planning. This work outlines a methodology for constructing datasets, a training suite, and evaluation methods for LLM-based textual and conversational planning tools utilized in an asset digital twin. Results indicate that the fine-tuned models generate estimated planning information within the statistical ranges observed in real maintenance data. The models achieve high accuracy (>90%) in document question-answering and instruction generation tasks. Furthermore, the conversational retrieval-augmented generation (RAG) assistant system achieves 100% document retrieval accuracy, while conversational information capture exceeds 98% across the majority of work-package assistant modules.

97 MATHEMATICS AND COMPUTING↗

Capsule network-based semantic segmentation model for thermal anomaly identification on building envelopes

Thermography technology is widely used to inspect thermal anomalies in building façade systems. Computer vision-based techniques provide opportunities to autonomously detect such heat anomalies to significantly improve the efficiency of decision-making for building envelope retrofitting and maintenance. Here, in this work, we propose a novel Capsule Network-based deep learning model – CapsLab – that detects and identifies thermal anomalies by semantic segmentation. CapsLab is built based on our proposed prediction-tuning capsule (PT-Capsule) layer. Different from a traditional capsule layer, which consists of part-whole transformation and capsule-routing process, the proposed layer is composed of a prediction and tuning process, which helps decreasing the number of model parameters significantly. While the applicability of traditional Capsule Networks (CapsNets) has been limited to simpler tasks and smaller datasets due to their scalability issue, we can leverage the lightweight of the proposed PT-Capsule layer, and apply it to the semantic segmentation task. In this work, we also employ our previously presented performance metric, referred to as the Anomaly Identification Metric (AIM) (Kakillioglua et al. 2021), to evaluate the segmentation outputs. Traditional performance metrics do not accurately reflect the true performance of the segmentation models in thermal anomaly identification due to the high subjectivity in the annotation process and higher overlap ratio sensitivity of the standard metrics. AIM, on the other hand, is robust to these drawbacks. Experimental results show, both qualitatively and quantitatively, that our proposed segmentation method can effectively segment the thermal anomalies. Specifically, our model provides 9.38% and 13.53% improvements over the baseline model – DeepLabV3+ – based on traditional mIoU score and the AIM score, respectively, while requiring less model parameters and less computation at the same time. In addition, the scores that the AIM metric generates better align with the scores provided by building performance experts.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗