Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “foundation models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Biochemistry & Molecular Biology↗

AMVOS: Additive Manufacturing Video Object Segmentation Dataset

This dataset provides labeled video frames from four additive manufacturing (AM) processes for video object segmentation (VOS) tasks. It contains 90 video segments comprising 900 individually annotated frames across five AM datasets: laser hot-wire directed energy deposition (LHW-DED), tungsten inert gas wire arc additive manufacturing (TIG-WAAM), plasma arc welding (PAW), visible-light polymer extrusion (visPolymer), and near-infrared polymer extrusion (irPolymer). Each video segment consists of 10 contiguous frames with corresponding pixel-level object instance annotations. Depending on the process, two of four object classes are labeled per frame: Melt Pool, Feed Wire, Nozzle, or Material. Raw frames are provided as .jpg files and annotations as palettized .png files. The dataset follows the directory structure of established VOS benchmarks (DAVIS, YouTube-VOS, MOSE), enabling direct integration into VOS model training and evaluation pipelines for foundation model fine-tuning, domain adaptation, or zero-shot performance benchmarking. Data was collected at Oak Ridge National Laboratory's Manufacturing Demonstration Facility.

Wetzel, Jon [ORNL]↗

Exocortex Network for AI-Augmented Human-Led Scientific Expedition

AI advances in science can be viewed along two main directions with a fluid boundary: enhancing efficiency through automation and smart tools to accelerate tasks that humans can already perform; and enabling exploration into uncharted territories and potentially toward AGI. These advances manifest in the AI cognitive core through the development and explainability of foundation models; in the physical embodiment of instruments and facilities; and in the integrated agency of AI workflows exemplified by the science exocortex. To address the role of humans in this evolving landscape, in this Perspective, we suggest a third direction: the development of personalized agents that form human-centered networks, supporting both efficiency and exploration while ensuring that AI remains aligned with human vision.

97 MATHEMATICS AND COMPUTING↗

Towards the next generation of Geospatial Artificial Intelligence

Geospatial Artificial Intelligence (GeoAI), as the integration of geospatial studies and AI, has become one of the fastest-developing research directions in spatial data science and geography. This rapid change in the field calls for a deeper understanding of the recent developments and envision where the field is going in the near future. In this work, we provide a quantitative analysis of the GeoAI literature from the spatial, temporal, and semantic aspects. We briefly discuss the history of AI and GeoAI by highlighting some pioneering work. Then we discuss the current landscape of GeoAI by selecting five representative subdomains including remote sensing, urban computing, Earth system science, cartography, and geospatial semantics. Finally, we highlight several unique future research directions of GeoAI which are classified into two groups: GeoAI method development challenges and GeoAI Ethics challenges. Topics include heterogeneity-aware GeoAI, knowledge-guided GeoAI, spatial representation learning, geo-foundation models, fairness-aware GeoAI, privacy-aware GeoAI, as well as interpretable and explainable GeoAI. We hope our review of GeoAI’s past, present, and future is comprehensive and can enlighten the next generation of GeoAI research.

58 GEOSCIENCES↗

Spatiotemporal forecasting of the edge localized modes in tokamak plasmas using neural networks

Artificial intelligence techniques have been increasingly adopted by the plasma and fusion science to address problems like plasma reconstruction, surrogate modeling, and tokamak/stellarator optimization. A key focus in sustained fusion research is the prediction and mitigation of edge-localized-modes (ELMs), instabilities that occur in short, periodic bursts and can cause erosion to the tokamak vessel wall. Recent research has demonstrated the power of neural networks in approximating continuous functions. In this work, we build spatiotemporal forecasting models that can predict the onset of ELMs and their evolution at early stages. We leverage recent advances in generative modeling, sequence-to-sequence modeling, and Fourier neural operators to propose architectures and training strategies that can learn to forecast short to long term dynamics of the noisy signals due to ELMs. We benchmark the developed model against a state-of-the-art foundation model using the beam emission spectroscopy (BES) data that captures the plasma fluctuations due to ELMs over a 8 x 8 spatial grid. Our models demonstrate high accuracy, outperforming the baselines, in predicting the evolution of BES signals during ELM events. Furthermore, the developed models exhibit high accuracy in predicting the rapid rise and relaxation of the signals due to ELMs within 30–80 µs.

edge localized modes↗

Segmentation Model Distillation [Poster]

The process of training object detection (OD) or image segmentation model requires both a substantial amount of data and technical knowledge, which often creates challenges in applying these types of models to their full potential. In order to streamline the process of developing these models, we propose a new pipeline where a foundation model assists in the dataset generation. Then this resulting dataset is used to fine-tune a fast light-weight model to perform the custom segmentation or OD. This resulting model is also fit for real-time image segmentation, such as in a video stream.

97 MATHEMATICS AND COMPUTING↗

High-Fidelity CFD Modeling of Cryogenic Hydrogen Isotope Extrusion for Fusion Reactor Pellet Fueling

This study investigates the extrusion processes of deuterium and protium using ANSYS-Polyflow. The geometries and computational fluid dynamics (CFD) settings closely replicate the experimental setups and data acquired from the extruder experiments at Oak Ridge National Laboratory (ORNL) for validation purposes. We explore the impacts of (1) slip versus non-slip boundary conditions and (2) the use of constant, temperature-, and shear rate–dependent viscosities, concluding that the implementation of non-slip wall boundary conditions combined with shear rate–dependent viscosity produced more accurate predictions. The simulations achieved excellent agreement with the experimental data, with relative differences of only 5% for deuterium, and 3% to 6% for protium. This is the first time that experimental extrusion data at ORNL have been accurately predicted through high-fidelity CFD modeling. In conclusion, the advancements offer valuable insights and a foundational modeling tool for optimizing pellet injectors for ITER and other future reactor-scale devices.

ANSYS-Polyflow↗

Automatic building energy model development and debugging using large language models agentic workflow

Building energy modeling (BEM) is a complex process that demands significant time and expertise, limiting its broader application in building design and operations. While Large Language Models (LLMs) agentic workflow have facilitated complex engineering processes, their application in BEM has not been specifically explored. This paper investigates the feasibility of automating BEM using LLM agentic workflow. Here, we developed a generic LLM-planning-based workflow that takes a building description as input and generates an error-free EnergyPlus building energy model. Our robust workflow includes four core agents: 1) Building Description Pre-Processing, 2) IDF Object Information Extraction, 3) Single IDF Object Generator Suite, and 4) IDF Debugging Agent. These agents divide the complex tasks into manageable sub-steps, enabling LLMs to generate accurate and reliable results at each stage. The case study demonstrates the successful translation of a building description into an error-free EnergyPlus model for the iUnit modular building at the National Renewable Energy Laboratory. The effectiveness of our workflow surpasses: 1) naive prompt engineering, 2) other LLM-based workflows, and 3) manual modeling, in terms of accuracy, reliability, and time efficiency. The paper concludes with a discussion on the interplay between foundational models and LLM agent planning design, advocating for the use of fine-tuned, specialized models to advance this field.

97 MATHEMATICS AND COMPUTING↗

Particle trajectory representation learning with masked point modeling

Liquid argon time projection chambers (LArTPCs) offer millimeter-scale 3D images of particle trajectories, enabling precision studies of neutrino oscillation, detection of supernova and solar neutrinos, searches for exotic dark matter, and proton decay. Current approaches utilize supervised machine learning models, requiring extensive simulations of particle physics and detector response that can introduce bias. Self-supervised learning (SSL), a machine learning approach that learns useful representations of unlabeled data from the data itself, has significantly advanced how large datasets are utilized for representation learning; however, its potential for applications to sensory data in high precision particle physics experiments remains largely unexplored. We introduce the Point-based liquid argon masked autoencoder (PoLAr-MAE), a self-supervised framework that learns physically meaningful representations directly from unlabeled LArTPC images. PoLAr-MAE achieves remarkable data efficiency for a point-level segmentation task, outperforming fully supervised methods in low data regimes. Linear classifiers on model outputs demonstrate robust performance across multiple downstream tasks. Our results position sensor-level SSL as a practical foundation model strategy for LArTPCs.

Young, Samuel [Stanford Univ., CA (United States)]↗

Method to simultaneously facilitate all jet physics tasks

Machine learning has become an essential tool in jet physics. Due to their complex, high-dimensional nature, jets can be explored holistically by neural networks in ways that are not possible manually. However, innovations in all areas of jet physics are proceeding in parallel. We show that specially constructed machine learning models trained for a specific jet classification task can improve the accuracy, precision, or speed of all other jet physics tasks. This is demonstrated by training on a particular multiclass generation and classification task and then using the learned representation for different generation and classification tasks, for datasets with a different (full) detector simulation, for jets from a different collision system ($pp$ versus $ep$), for generative models, for likelihood ratio estimation, and for anomaly detection. We consider our omnilearn approach thus as a jet-physics foundation model. It is made publicly available for use in any area where state-of-the-art precision is required for analyses involving jets and their substructure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions

Instruction finetuning is a popular paradigm to align large language models (LLM) with human intent. Despite its popularity, this idea is less explored in improving the LLMs to align existing foundation models with scientific disciplines, concepts and goals. In this work, we present SciTune as a tuning framework to improve the ability of LLMs to follow scientific multimodal instructions. To test our methodology, we use a human-generated scientific instruction tuning dataset and train a large multimodal model LLaMA-SciTune that connects a vision encoder and LLM for science-focused visual and language understanding. LLaMA-SciTune significantly outperforms the state-of-the-art models in the generated figure types and captions in multiple scientific multimodal benchmarks. In comparison to the models that are fine-tuned with machine generated data only, LLaMA-SciTune surpasses human performance on average and in many sub-categories on the ScienceQA benchmark.

• Artificial intelligence (AI) / machine learning ↗

Location generalizability of image-based air quality models

This paper is to be submitted at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Computer Vision for Earth Observation workshop. The full paper abstract is below: The ability to rapidly quantify atmospheric pollutants is important both for global emissions monitoring and for mitigating the adverse effects that follow a hazardous chemical release. In the aftermath of a chemical release, imagery is often the only available resource to assess local conditions. Recent work has demonstrated initial success in predicting particulate matter pollution from imagery; however, these results are tied to a specific site and do not generalize to new geographic locations. In this work, we seek to understand how easily deep learning models generalize to new locations in the context of image-based air quality assessments, targeting two distinct tasks: (1) broad measures of particulate matter pollution, and (2) the mass of a given chemical released in hazardous plumes. For the latter, we focus on sulfur dioxide, a toxic aerosol and a major component of particulate matter pollution caused by industrial fossil fuel consumption. To develop a model that operates in the widest possible range of environments, we test different training strategies, including the use of new geolocation foundation models. The best performing models achieve >80% accuracy when evaluating unseen imagery at previously seen sites, but we find significant drops in performance when evaluating imagery from unseen sites, at best 65%. Additionally, we present the public release of the National Parks Air Quality Index Dataset, a new medium-sized dataset that pairs imagery with sensor-based air quality measurements at 15 different national parks.

Byler, Eleanor B. [BATTELLE (PACIFIC NW LAB)]↗

Harnessing distributed GPU computing for generalizable graph convolutional networks in power grid reliability assessments

Although machine learning (ML) has emerged as a powerful tool for rapidly assessing grid contingencies, prior studies have largely considered a static grid topology in their analyses. This limits their application, since they need to be re-trained for every new topology. Here, this paper explores the development of generalizable graph convolutional network (GCN) models by pre-training them across a range of grid topologies and contingency types. We found that a GCN model with auto-regressive moving average (ARMA) layers with a line graph representation of the grid offered the best predictive performance in predicting voltage magnitudes (VM) and voltage angles (VA). We introduced the concept of phantom nodes to consider disparate grid topologies with a varying number of nodes and lines. For pre-training the GCN ARMA model across a variety of topologies, distributed graphics processing unit (GPU) computing afforded us significant training scalability. The predictive performance of this model on grid topologies that were part of the training data is substantially better than the direct current (DC) approximation. Although direct application of the pre-trained model to topologies that are not part of the grid is not particularly satisfactory, fine-tuning with small amounts of data from a specific topology of interest significantly improves predictive performance. In general, this paper highlights the feasibility of training large-scale GNN models to assess the reliability of power grids by considering a wide variety of grid topologies and contingency types. With the advent of foundational models in ML and the exponential increase in GPU computing clusters, generalizable ML models will significantly enhance how utilities manage power systems and make decisions in real-time or near-real-time.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Generative AI for Grid Operations [Slides]

In the last few years, the development and use of generative artificial intelligence (AI) and large-language models (LLMs) have changed the landscape of how AI and machine learning (ML) are being used in power systems. LLMs are built on foundational models based on large data sets that can be trained to provide information rapidly and through simple natural language prompts. Generative AI can then perform human-like tasks using ML models to identify and mimic pattens in the data sets. This presentation explores how generative AI can enhance grid operations by improving forecasts, enabling rapid contingency analyses, and offering real-time operational suggestions. By providing grid operators with valuable insights, generative AI will empower them to manage power systems more effectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Agent-based modeling for multimodal transportation of CO 2 for carbon capture, utilization, and storage: CCUS-agent

Here, to understand the system-level interactions between the entities in Carbon Capture, Utilization, and Storage (CCUS), an agent-based foundational modeling tool, CCUS-Agent, is developed for a large-scale study of transportation flows and infrastructure in the United States. Key features of the tool include (i) modular design, (ii) multiple transportation modes, (iii) capabilities for extension, and (iv) testing against various system components and networks of small and large sizes. Five matching algorithms for CO 2 supply agents (e.g., powerplants and industrial facilities) and demand agents (e.g., storage and utilization sites) are explored: Most Profitable First Year (MPFY), Most Profitable All Years (MPAY), Shortest Total Distance First Year (SDFY), Shortest Total Distance All Years (SDAY), and Shortest distance to long-haul transport All Years (ACAY). Before matching, the supply agent, demand agent, and route must be available, and the connection must be profitable. A profitable connection means the supply agent portion of revenue from the 45Q tax credit must cover the supply agent costs and all transportation costs, while the demand agent revenue portion must cover all demand agent costs. A case study employing over 5500 supply and demand agents and multimodal CCUS transportation infrastructure in the contiguous United States is conducted. The results suggest that it is possible to capture over 9 billion tonnes (GT) of CO 2 from 2025 to 2043, which will increase significantly to 22 GT if the capture costs are reduced by 40 %. The MPFY and SDFY algorithms capture more CO 2 earlier in the time horizon, while the MPAY and SDAY algorithms capture more later in the time horizon.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Uncertainty quantification of fireball features extracted from nuclear test films using computer vision

Films from the US’s historic nuclear testing era comprise the only extensive collection of imagery depicting high-yield detonations. These films offer unique insights into the characteristics of flows occurring on scales that are difficult to replicate experimentally, and they are a valuable source of data for the validation of models used to describe nuclear detonations. In recent work, we implemented modern computer vision and machine learning techniques to extract features of the fireball following nuclear detonation. With a training dataset of fireball films, we fine-tuned a You Only Look Once 11 (YOLO11) model to detect and track the fireball. Applied to a video, the outer bounding box produced in each frame by YOLO11 is used as an input prompt to Meta’s Segment Anything Model 2 (SAM2), which is shown to accurately predict the boundary of the fireball over time with high resolution. These state-of-the-art computer vision foundation models exhibit impressive visual accuracy in their results but lack an output of values that robustly quantify uncertainty in scientific applications. In this paper, we develop procedures for uncertainty quantification of extracted fireball features. We outline the application of a parallel attention mechanism to calculate uncertainty ranges that complement and better pose model validation data. This higher quality fireball validation data may serve to improve prognostic models describing nuclear detonations in support of nuclear forensic and emergency response activities.

Khristy, Joel [ORNL] (ORCID:0000000209963060)↗

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES↗

Towards Anomaly Detection at the CMS High-Level Trigger System

Traditional trigger strategies in CMS typically rely on model-dependent selections or rigid kinematic cuts, risking the omission of unexpected exotic signatures. To address this, we propose a novel anomaly detection (AD) algorithm for the High-Level Trigger (HLT), designed to serve as a complementary second layer of filtering to the Level-1 AXOL1TL AD algorithm. We employ a transformer-based foundation model trained on a diverse ensemble of Standard Model processes. By combining a joint contrastive and classification objective, and using particle kinematics as inputs, the model learns to map events to a physics-informed latent space where anomalous events are isolated from dominant backgrounds. Preliminary results show that this strategy enhances the signal-to-background ratio across a range of rare SM and BSM scenarios. Furthermore, this work constitutes foundational R&D for the potential implementation of an analogous AD algorithm in the Level-1 trigger system for Phase-2.

Cruz, Roy [U. Wisconsin, Madison (main)] (ORCID:00↗