Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Machine learning–assisted prediction of heat fluxes through thermally anisotropic building envelopes

Thermally anisotropic building envelope (TABE) is a novel active building envelope that can save energy use to maintain thermal comfort in buildings by redirecting heat and coolness from building envelopes to thermal loops. Finite element models (FEMs) can be used to compute the heat fluxes through TABEs, but the high computational cost of finite element simulations has prevented parametric studies and design optimizations. This paper proposes a domain knowledge–informed, finite element–based machine learning framework to reduce the computation cost for the energy management of buildings installed with TABE that uses a ground thermal loop. First, the training heat flux data set was generated by FEM simulations with different thermal loop schedules. Then, both shallow learning models (i.e., multivariate linear regression and eXtreme Gradient Boost, or XGBoost) and a deep learning model (i.e., deep neural network, or DNN) were trained to predict the heat fluxes. Domain knowledge was used for data preprocessing and feature selection. Finally, the suitability of the selected machine learning model was tested under different thermal loop schedules. Herein, the case study results showed that: (1) XGBoost can be as accurate as DNN (coefficient of determination equal to 0.81) with much less training time; (2) the annual energy cost savings for different thermal loop schedules obtained by the XGBoost-predicted and FEM-calculated heat fluxes are consistent, having a difference of only 4%; and (3) XGBoost can reduce the computation time for the annual energy analysis of the case study building with a given thermal loop schedule from around 12 h by using FEM to less than 1 min.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

Cost function for low-dimensional manifold topology assessment

Abstract In reduced-order modeling, complex systems that exhibit high state-space dimensionality are described and evolved using a small number of parameters. These parameters can be obtained in a data-driven way, where a high-dimensional dataset is projected onto a lower-dimensional basis. A complex system is then restricted to states on a low-dimensional manifold where it can be efficiently modeled. While this approach brings computational benefits, obtaining a good quality of the manifold topology becomes a crucial aspect when models, such as nonlinear regression, are built on top of the manifold. Here, we present a quantitative metric for characterizing manifold topologies. Our metric pays attention to non-uniqueness and spatial gradients in physical quantities of interest, and can be applied to manifolds of arbitrary dimensionality. Using the metric as a cost function in optimization algorithms, we show that optimized low-dimensional projections can be found. We delineate a few applications of the cost function to datasets representing argon plasma, reacting flows and atmospheric pollutant dispersion. We demonstrate how the cost function can assess various dimensionality reduction and manifold learning techniques as well as data preprocessing strategies in their capacity to yield quality low-dimensional projections. We show that improved manifold topologies can facilitate building nonlinear regression models.

42 ENGINEERING↗

Autonomous Tuning and Charge-State Detection of Gate-Defined Quantum Dots

Defining quantum dots in semiconductor-based heterostructures is an essential step in initializing solid-state qubits. With growing device complexity and increasing number of functional devices required for measurements, a manual approach to finding suitable gate voltages to confine electrons electrostatically is impractical. Here, we implement a two-stage device characterization and dot-tuning process, which first determines whether devices are functional and then attempts to tune the functional devices to the single or double quantum-dot regime. We show that automating well-established manual-tuning procedures and replacing the experimenter’s decisions by supervised machine learning is sufficient to tune double quantum dots in multiple devices without premeasured input or manual intervention. The quality of measurement results and charge states are assessed by four binary classifiers trained with experimental data, reflecting real device behavior. We compare and optimize eight models and different data preprocessing techniques for each of the classifiers to achieve reliable autonomous tuning, an essential step towards scalable quantum systems in quantum-dot-based qubit architectures.

Darulova, J↗

Comparative Study of Large Language Model Architectures on Frontier

Large language models (LLMs) have garnered significant attention in both the AI community and beyond. Among these, the Generative Pre-trained Transformer (GPT) has emerged as the dominant architecture, spawning numerous variants. However, these variants have undergone pre-training under diverse conditions, including variations in input data, data preprocessing, and training methodologies, resulting in a lack of controlled comparative studies. Here we meticulously examine two prominent open-sourced GPT architectures, GPT-NeoX and LLaMA, leveraging the computational power of Frontier, the world’s first Exascale supercomputer. Employing the same materials science text corpus and a comprehensive end-to-end pipeline, we conduct a comparative analysis of their training and downstream performance. Our efforts culminate in achieving state-of-the-art performance on a challenging materials science benchmark. Furthermore, we investigate the computation and energy efficiency, and propose a computationally efficient method for architecture design. To our knowledge, these pre-trained models represent the largest available for materials science. Our findings provide practical guidance for building LLMs on HPC platforms.

Yin, Junqi↗

Mass Detection for Heavy-Duty Vehicles using Gaussian Belief Propagation

Predicting vehicle mass is critical to accurately estimate energy use and emissions of commercial trucks. However, data from vehicle telematics is often not at sufficient temporal resolution or accuracy for use in model-based detection methods. In this work, a new statistical mass prediction technique is described for heavy-duty vehicles that incorporates the use Gaussian Belief Propagation (GBP) for probabilistic inference. Similar to Bayesian inference models, the GBP model typically requires less labeled training data than other contemporary machine learning techniques. First, a factor graph is constructed, and a set of Gaussian belief nodes with associated means and variances are fitted to the training data. To better handle noisy input data, the GBP mass prediction model utilizes a k-nearest factors (kNF) algorithm for probabilistic inference on unseen testing data. The proposed method is compared with a classical weighted k-nearest neighbors (kNN) regressor. This statistical kNF-GBP model works even with low-quantity, low-quality initial training data, while being capable of realtime mass estimation. Unlike the kNN regressor, the GBP model produces a measure of uncertainty with its predictions. The proposed method is validated using curve-sampled driving data collected from multiple cloud-connected Class 8 regional haul diesel trucks. Both the kNN regressor and the kNF-GBP mass prediction model were able to predict payload mass with coefficients of determination above 0.97 with minimal data preprocessing.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

gp_blendclass_singleband

The code used for the data preprocessing, image simulation, and model training and analysis reported in the paper "Gaussian Process Classification for Galaxy Blend Identification in LSST" (arXiv:2107.09246).

Buchanan, JamesJ.↗

Detecting Living-off-the-land Attacks Using K-means And Graph Convolutional Networks

The code ingests Zeek logs derived from network packet captures and goes through data preprocessing before it gets passed into a K-Means model that labels each device as either a client or server. Graph Convolutional Network (GCN) model is used to obtain the embeddings to represent the features in lower dimension. Last, K-means cluster analysis is used to cluster the embeddings for each class.

Quach, Anna [Idaho National Laboratory (INL), Idah↗

Modelling Framework For Fire And Smoke Detection In Imagery

This framework was developed to ingest images/video images to train and test artificial neural network structures for image-based detection of fire and smoke. The code includes data preprocessing, model development, and testing. The framework is designed to work with RGB video imagery.

Griffel, LloydM.↗

HydraGNN_GFM_FineTuning4Materials v1.0

This repository enables fine-tuning of the HydraGNN Predictive GFM 2026 — an open-source ensemble of pre-trained graph foundation models for atomistic materials modeling, developed at Oak Ridge National Laboratory. The GFM 2026 is freely available and downloadable via Globus from the OLCF Data Constellation (DOI: 10.13139/OLCF/2562660). Starting from these pre-trained weights, this repository provides a complete transfer learning pipeline for adapting the GFM ensemble to domain-specific molecular and materials property prediction tasks. It includes: 1) Utilities for ensemble fine-tuning with task-specific output heads 2) Example pipelines for eight widely-used materials and molecular datasets 3) Tools for model adaptation and head configuration 4) Data preprocessing utilities for each supported dataset 5) Benchmarking and evaluation scripts

Ungerboeck, Linda↗

Development of Prognostic Models Using Plant Asset Data

The recent growth of machine learning and artificial intelligence technologies provides opportunities for leveraging data-driven algorithms to address the problems of diagnostics and prognostics in the nuclear power industry. The use of machine learning and other statistical methods as prognostic models is of particular interest in the nuclear industry to accurately predict future equipment or plant state given a set of measurements. Such predictive capability will enable predictive assessment of component condition and remaining life and allow for condition-based predictive maintenance. The resulting optimization of maintenance scheduling and reduction in unnecessary maintenance activities will lower overall maintenance costs and improve the economics of nuclear power. This report discusses the various aspects of data processing and model development that are likely to influence the performance of prognostic models. Data from a boiling-water reactor was used to evaluate several prognostic models to identify key considerations for developing such models to predict data-driven plant state and equipment degradation condition. Preliminary results indicate the need for data sets that are relevant to the problem at hand and contain signatures that may be correlated to the prediction problem. Assuming such data exist, development of prognostic models using data-driven methods requires an understanding of the various sources of influence on the prediction accuracy (such as the model architecture, data preprocessing approaches, and potentially external factors influencing the equipment or plant system under assessment). Ongoing research is evaluating these factors in greater detail and examining techniques for calculating prediction uncertainty bounds.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

RockFlow: Fast Generation of Synthetic Source Rock Images Using Generative Flow Models

Image-based evaluation methods are a valuable tool for source rock characterization. The time and resources needed to obtain images has spurred development of machine-learning generative models to create synthetic images of pore structure and rock fabric from limited image data. While generative models have shown success, existing methods for generating 3D volumes from 2D training images are restricted to binary images and grayscale volume generation requires 3D training data. Shale characterization relies on 2D imaging techniques such as scanning electron microscopy (SEM), and grayscale values carry important information about porosity, kerogen content, and mineral composition of the shale. Here, we introduce RockFlow, a method based on generative flow models that creates grayscale volumes from 2D training data. We apply RockFlow to baseline binary micro-CT image volumes and compare performance to a previously proposed model. We also show the extension of our model to 2D grayscale data by generating grayscale image volumes from 2D SEM and dual modality nanoscale shale images. The results show that our method underestimates the porosity and surface area on the binary baseline datasets but is able to generate realistic grayscale image volumes for shales. With improved binary data preprocessing, we believe that our model is capable of generating synthetic porous media volumes for a very broad class of rocks from shale to carbonates to sandstone.

58 GEOSCIENCES↗

Earthquake Nowcasting with Deep Learning

We review previous approaches to nowcasting earthquakes and introduce new approaches based on deep learning using three distinct models based on recurrent neural networks and transformers. We discuss different choices for observables and measures presenting promising initial results for a region of Southern California from 1950–2020. Earthquake activity is predicted as a function of 0.1-degree spatial bins for time periods varying from two weeks to four years. The overall quality is measured by the Nash Sutcliffe efficiency comparing the deviation of nowcast and observation with the variance over time in each spatial region. The software is available as open source together with the preprocessed data from the USGS.

Fox, Geoffrey Charles (ORCID:0000000310171391)↗

Streamlining Ocean Dynamics Modeling with Fourier Neural Operators: A Multiobjective Hyperparameter and Architecture Optimization Approach

Training an effective deep learning model to learn ocean processes involves careful choices of various hyperparameters. We leverage DeepHyper’s advanced search algorithms for multiobjective optimization, streamlining the development of neural networks tailored for ocean modeling. The focus is on optimizing Fourier neural operators (FNOs), a data-driven model capable of simulating complex ocean behaviors. Selecting the correct model and tuning the hyperparameters are challenging tasks, requiring much effort to ensure model accuracy. DeepHyper allows efficient exploration of hyperparameters associated with data preprocessing, FNO architecture-related hyperparameters, and various model training strategies. We aim to obtain an optimal set of hyperparameters leading to the most performant model. Moreover, on top of the commonly used mean squared error for model training, we propose adopting the negative anomaly correlation coefficient as the additional loss term to improve model performance and investigate the potential trade-off between the two terms. The numerical experiments show that the optimal set of hyperparameters enhanced model performance in single timestepping forecasting and greatly exceeded the baseline configuration in the autoregressive rollout for long-horizon forecasting up to 30 days. Utilizing DeepHyper, we demonstrate an approach to enhance the use of FNO in ocean dynamics forecasting, offering a scalable solution with improved precision.

97 MATHEMATICS AND COMPUTING↗

Janus: A Python Package for Agent-Based Modeling of Land Use and Land Cover Change

Janus is an open source Python package for agent-based modeling (ABM) of land use and land cover change (LULCC). Many ABMs of LULCC have been created across platforms, some of which are not ideal for large scale, high resolution scenarios. This model provides a simple object-oriented framework for creating ABMs specific to LULCC. The organizational philosophy of the modeling framework is to create software objects (agents) that are associated with specific and contextual attributes which are isolated from where those agents exist in the spatial setting of the model, yet provide clear linkages between the agent, their environment, and other agents in the simulation. In this way, the framework allows for assembly of LULCC ABMs with low (programmatic) overhead, making the models extensible and providing clear mechanisms for integrating them with process-oriented biophysical models. Provided with Janus is a suite of geospatial data preprocessing tools that can use arbitrary land cover products as an input. Crop choice decisions are based on potential crop prices, these can be created synthetically, or drawn from integrated human-Earth systems models such as the GCAM. Janus is publicly accessible through GitHub and provides an example dataset for testing.

54 ENVIRONMENTAL SCIENCES↗

Layered 'recognition cone' networks that pre-process, classify, and describe.

Discussion of pattern recognition programs for input data preprocessing with simultaneous or subsequent characterization, or characterization into a 'recognition cone,' or description and naming, interrelated descriptions, and conversion. A computer program is described that transforms and characterizes the input through the successive layers of a recognition cone. The program can choose and put forth names of parts of the input scene. It combines pieces of a description into interrelated wholes by using n-tuple characterizers and conducts a simple and stylized conversation about what it has seen. The technique of combining recognition cones with preprocessing transformations and characterizations is expected to contribute to technology in this field.

Uhr, L.↗

Quantitative simultaneous multi-element microprobe analysis using combined wavelength and energy dispersive systems

A combined WDS-EDS system obviates the severe X-ray peak overlap problems encountered with Na, Mg, Al and Si common to pure EDS systems. By application of easily measured empirical correction factors for pulse pile-up and peak overlaps which are normally observed in the analysis of silicate minerals, the accuracy of analysis is comparable with that expected for WDS electron microprobe analyses. The continuum backgrounds are subtracted for the spectra by a spline fitting technique based on integrated intensities between the peaks. The preprocessed data are then reduced to chemical analyses by existing data reduction programs.

Walter, L. S.↗

Satellite laser ranging work at the Goddard Space Flight Center

The paper describes the satellite laser ranging system at the Goddard Space Flight Center, its range and accuracy capabilities, and planned improvements for future systems. Major subsystems are described, including the laser, optical/mechanical, receiver, computer/software, timing, and laser data preprocessing subsystems. Operational considerations are examined, with attention given the mobile station layout, manpower requirements, and transportability. System performance is considered, with emphasis on system accuracy (calibration, stability, clock synchronization, atmospheric propagation correction) and range capability.

Mcgunigal, T. E.↗