Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Maximum likelihood estimation of label imperfections and its use in the identification of mislabeled patterns

The problem of estimating label imperfections and the use of the estimation in identifying mislabeled patterns is presented. Expressions for the maximum likelihood estimates of classification errors and a priori probabilities are derived from the classification of a set of labeled patterns. Expressions also are given for the asymptotic variances of probability of correct classification and proportions. Simple models are developed for imperfections in the labels and for classification errors and are used in the formulation of a maximum likelihood estimation scheme. Schemes are presented for the identification of mislabeled patterns in terms of threshold on the discriminant functions for both two-class and multiclass cases. Expressions are derived for the probability that the imperfect label identification scheme will result in a wrong decision and are used in computing thresholds. The results of practical applications of these techniques in the processing of remotely sensed multispectral data are presented.

Chittineni, C. B.↗

Deep active learning for classifying cancer pathology reports

Abstract Background Automated text classification has many important applications in the clinical setting; however, obtaining labelled data for training machine learning and deep learning models is often difficult and expensive. Active learning techniques may mitigate this challenge by reducing the amount of labelled data required to effectively train a model. In this study, we analyze the effectiveness of 11 active learning algorithms on classifying subsite and histology from cancer pathology reports using a Convolutional Neural Network as the text classification model. Results We compare the performance of each active learning strategy using two differently sized datasets and two different classification tasks. Our results show that on all tasks and dataset sizes, all active learning strategies except diversity-sampling strategies outperformed random sampling, i.e., no active learning. On our large dataset (15K initial labelled samples, adding 15K additional labelled samples each iteration of active learning), there was no clear winner between the different active learning strategies. On our small dataset (1K initial labelled samples, adding 1K additional labelled samples each iteration of active learning), marginal and ratio uncertainty sampling performed better than all other active learning techniques. We found that compared to random sampling, active learning strongly helps performance on rare classes by focusing on underrepresented classes. Conclusions Active learning can save annotation cost by helping human annotators efficiently and intelligently select which samples to label. Our results show that a dataset constructed using effective active learning techniques requires less than half the amount of labelled data to achieve the same performance as a dataset constructed using random sampling.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗

Design of Digital Twin Sensing Strategies Via Predictive Modeling and Interpretable Machine Learning

This work develops a methodology for sensor placement and dynamic sensor scheduling decisions for digital twins. The digital twin data assimilation is posed as a classification problem, and predictive models are used to train optimal classification trees that represent the map from observed data to estimated digital twin states. In addition to providing a rapid digital twin updating capability, the resulting classification trees yield an interpretable mathematical representation that can be queried to inform sensor placement and sensor scheduling decisions. The proposed approach is demonstrated for a structural digital twin of a 12 ft wingspan unmanned aerial vehicle. Offline, training data are generated by simulating scenarios using predictive reduced-order models of the vehicle in a range of structural states. Furthermore, these training data can be further augmented using experimental or other historical data. In operation, the trained classifier is applied to observational data from the physical vehicle, enabling rapid adaptation of the digital twin in response to changes in structural health. Within this context, we study the performance of the optimal tree classifiers and demonstrate how they enable explainable structural assessments from sparse sensor measurements and also inform optimal sensor placement.

47 OTHER INSTRUMENTATION↗

One of These Things IS Like the Other: Pursuing a New Taxonomy of Industry for Improved Energy System Modeling

Industrial processes drive the exchange of materials, energy, and currency throughout the economy. These processes are powered by electricity and direct combustion, with variation in their operation even within the same industry. This heterogeneity makes it difficult for large models, including the National Energy Modeling System (US), to project their energy use while remaining tractable. Decarbonization and ensuing changes to the energy system require changes to industrial processes while offering opportunities for process innovation, but the extent and nature of changes are difficult to model with current classification schemes and corresponding data. The North American Industrial Classification (NAICS) is an economic taxonomy of industries, but its categories are less meaningful from an energy and material flow perspective. For example, a facility that makes steel from iron ore in a blast furnace/basic oxygen furnace is categorized under the same NAICS code as a facility that makes steel from scrap in an electric arc furnace despite the scale, use of recycled scrap versus iron ore, and energy use differences in the two facility types. Exploratory analysis is performed on a large dataset used for plant-level energy assessment in order to detect clusters that can aid in better modeling of industry for energy analysis in an evolving system with breakthrough technologies.

28 EE - Advanced Manufacturing Office (EE-5A)↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

The Earth Model Column Collaboratory (EMC 2 ) v1.1: an open-source ground-based lidar and radar instrument simulator and subcolumn generator for large-scale models

Abstract. Climate models are essential for our comprehensive understanding of Earth's atmosphere and can provide critical insights on future changes decades ahead. Because of these critical roles, today's climate models are continuously being developed and evaluated using constraining observations and measurements obtained by satellites, airborne, and ground-based instruments. Instrument simulators can provide a bridge between the measured or retrieved quantities and their sampling in models and field observations while considering instrument sensitivity limitations. Here we present the Earth Model Column Collaboratory (EMC2), an open-source ground-based lidar and radar instrument simulator and subcolumn generator, specifically designed for large-scale models, in particular climate models, but also applicable to high-resolution model output. EMC2 provides a flexible framework enabling direct comparison of model output with ground-based observations, including generation of subcolumns that may statistically represent finer model spatial resolutions. In addition, EMC2 emulates ground-based (and air- or space-borne) measurements while remaining faithful to large-scale models' physical assumptions implemented in their cloud or radiation schemes. The simulator uses either single particle or bulk particle size distribution lookup tables, depending on the selected scheme approach, to perform the forward calculations. To facilitate model evaluation, EMC2 also includes three hydrometeor classification methods, namely, radar- and sounding-based cloud and precipitation detection and classification, lidar-based phase classification, and a Cloud Feedback Model Intercomparison Project Observational Simulator Package (COSP) lidar simulator emulator. The software is written in Python, is easy to use, and can be straightforwardly customized for different models, radars, and lidars. Following the description of the logic, functionality, features, and software structure of EMC2, we present a case study of highly supercooled mixed-phase cloud based on measurements from the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) West Antarctic Radiation Experiment (AWARE). We compare observations with the application of EMC2 to outputs from four configurations of the NASA Goddard Institute for Space Studies (GISS) climate model (ModelE3) in single-column model (SCM) mode and from a large-eddy simulation (LES) model. We show that two of the four ModelE3 configurations can form and maintain highly supercooled precipitating cloud for several hours, consistent with observations and LES. While our focus is on one of these ModelE3 configurations, which performed slightly better in this case study, both of these configurations and the LES results post-processed with EMC2 generally provide reasonable agreement with observed lidar and radar variables. As briefly demonstrated here, EMC2 can provide a lightweight and flexible framework for comparing the results of both large-scale and high-resolution models directly with observations, with relatively little overhead and multiple options for achieving consistency with model microphysical or radiation scheme physics.

58 GEOSCIENCES↗

The Earth Model Column Collaboratory (EMC2) v1.1: An Open-Source Ground-Based Lidar and Radar Instrument Simulator and Subcolumn Generator for Large-Scale Models

Climate models are essential for our comprehensive understanding of Earth's atmosphere and can provide critical insights on future changes decades ahead. Because of these critical roles, today's climate models are continuously being developed and evaluated using constraining observations and measurements obtained by satellites, airborne, and ground-based instruments. Instrument simulators can provide a bridge between the measured or retrieved quantities and their sampling in models and field observations while considering instrument sensitivity limitations. Here we present the Earth Model Column Collaboratory (EMC2), an open-source ground-based lidar and radar instrument simulator and subcolumn generator, specifically designed for large-scale models, in particular climate models, but also applicable to high-resolution model output. EMC2 provides a flexible framework enabling direct comparison of model output with ground-based observations, including generation of subcolumns that may statistically represent finer model spatial resolutions. In addition, EMC2 emulates ground-based (and air- or space-borne) measurements while remaining faithful to large-scale models' physical assumptions implemented in their cloud or radiation schemes. The simulator uses either single particle or bulk particle size distribution lookup tables, depending on the selected scheme approach, to perform the forward calculations. To facilitate model evaluation, EMC2 also includes three hydrometeor classification methods, namely, radar- and sounding-based cloud and precipitation detection and classification, lidar-based phase classification, and a Cloud Feedback Model Intercomparison Project Observational Simulator Package (COSP) lidar simulator emulator. The software is written in Python, is easy to use, and can be straightforwardly customized for different models, radars, and lidars. Following the description of the logic, functionality, features, and software structure of EMC2, we present a case study of highly supercooled mixed-phase cloud based on measurements from the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) West Antarctic Radiation Experiment (AWARE). We compare observations with the application of EMC2 to outputs from four configurations of the NASA Goddard Institute for Space Studies (GISS) climate model (ModelE3) in single-column model (SCM) mode and from a large-eddy simulation (LES) model. We show that two of the four ModelE3 configurations can form and maintain highly supercooled precipitating cloud for several hours, consistent with observations and LES. While our focus is on one of these ModelE3 configurations, which performed slightly better in this case study, both of these configurations and the LES results post-processed with EMC2 generally provide reasonable agreement with observed lidar and radar variables. As briefly demonstrated here, EMC2 can provide a lightweight and flexible framework for comparing the results of both large-scale and high-resolution models directly with observations, with relatively little overhead and multiple options for achieving consistency with model microphysical or radiation scheme physics.

Earth Model Column Collaboratory↗

Onboard Hyperspectral Image Classification via Transfer Learning for Communication-Limited Spacecraft

Employing deep-learning and artificial-intelligence (AI) techniques onboard spacecraft can dramatically improve priority data selection to ensure more effective use of the available downlink. However, deployment of effective deep-learning models requires significant training on the ground, which may not be feasible, due to limited data available in an unexplored environment. Therefore, this research explores building robust classification models for onboard data processing where training data is highly limited using transfer-learning techniques. In this paper, we focus on the use case of hyperspectral imaging for remote sensing, a domain where the high dimensionality of the data from the sensor can rapidly saturate the downlink bandwidth. With this bottleneck, there is an impending need to autonomously and robustly classify data onboard to optimize downlink of high-impact measurements, thus maximizing the scientific utility per bit transmitted to the ground. This paper examines the use of deep neural networks onboard for hyperspectral image classification in a communication-limited scenario to analyze how the models perform with limited training data. The use of transfer learning can ameliorate the issue of poor generalization by transferring features learned from training on a large source dataset for one classification task to the target classification task with limited training data. For two deep-learning models from literature, we compare the accuracy of the models trained using transfer learning to models trained from scratch using a random weight initialization with varying amounts of training data. We demonstrate the feasibility and performance of running inference of the deep-learning models on representative flight-like hardware.

Advanced Avionics, Machine Learning, Data Processi↗

On the discovery of stars, quasars, and galaxies in the Southern Hemisphere with S-PLUS DR2

ABSTRACT This paper provides a catalogue of stars, quasars, and galaxies for the Southern Photometric Local Universe Survey Data Release 2 (S-PLUS DR2) in the Stripe 82 region. We show that a 12-band filter system (5 Sloan-like and 7 narrow bands) allows better performance for object classification than the usual analysis based solely on broad bands (regardless of infrared information). Moreover, we show that our classification is robust against missing values. Using spectroscopically confirmed sources retrieved from the Sloan Digital Sky Survey DR16 and DR14Q, we train a random forest classifier with the 12 S-PLUS magnitudes + 4 morphological features. A second random forest classifier is trained with the addition of the W1 (3.4 $\mu\mathrm{m} $) and W2 (4.6 $\mu\mathrm{m} $) magnitudes from the Wide-field Infrared Survey Explorer (WISE). Forty-four per cent of our catalogue have WISE counterparts and are provided with classification from both models. We achieve 95.76 per cent (52.47 per cent) of quasar purity, 95.88 per cent (92.24 per cent) of quasar completeness, 99.44 per cent (98.17 per cent) of star purity, 98.22 per cent (78.56 per cent) of star completeness, 98.04 per cent (81.39 per cent) of galaxy purity, and 98.8 per cent (85.37 per cent) of galaxy completeness for the first (second) classifier, for which the metrics were calculated on objects with (without) WISE counterpart. A total of 2926 787 objects that are not in our spectroscopic sample were labelled, obtaining 335 956 quasars, 1347 340 stars, and 1243 391 galaxies. From those, 7.4 per cent, 76.0 per cent, and 58.4 per cent were classified with probabilities above 80 per cent. The catalogue with classification and probabilities for Stripe 82 S-PLUS DR2 is available for download.

79 ASTRONOMY AND ASTROPHYSICS↗

Bhutan Agriculture: Developing a Crop Mask for Rice and Creating a Data Collection Protocol Utilizing Remotely Sensed Data in Bhutan

Rice cultivation in Bhutan has been increasingly threatened by deteriorating soil health and outbreaks of diseases and pests associated with the global change in climate patterns. Field surveys, which the national government of Bhutan has relied on to monitor remote agricultural lands, are becoming increasingly overwhelmed by growing threats to agricultural health. To address these concerns, NASA DEVELOP partnered with the Department of Agriculture of Bhutan, the Bhutan Foundation, and the Ugyen Wangchuck Institute of Conservation and Environmental Research (UWICER) and worked to increase the government of Bhutan’s agricultural monitoring capacity. Utilizing Earth observations including Landsat 8 Operational Land Imager (OLI), Sentinel-1 C-band Synthetic Aperture Radar (C-SAR), Shuttle Radar Topography Mission (SRTM), and Planet imagery, the DEVELOP team worked with NASA SERVIR and created a sampling protocol to identify rice plantations and supplement field surveys for more efficient agriculture monitoring. The analysis focused on districts Paro, Punakha, Samtse, Sarpang, Trongsa, Zhemgang, Wangdue Phodrang, and Samdrup Jongkhar in the year 2020 during the period of transplantation (June) to harvesting of rice (November). The team provided the partners with a sampling protocol for integrating NASA Earth observations into their crop monitoring methods, as well as a crop mask for rice identification and to aid crop management. The crop mask for rice was developed using the Random Forest (RF) classifier for the eight districts of Bhutan. Visually, the random forest model has proved to be more accurate and precise than the classification and Regression Tree model. Statistically, the Random Forest model was 91.8% accurate in identifying rice in Bhutan.

Yeshey Seldon↗

Atomic resolution convergent beam electron diffraction analysis using convolutional neural networks

Two types of convolutional neural network (CNN) models, a discrete classification network and a continuous regression network, were trained to determine local sample thickness from convergent beam diffraction (CBED) patterns of SrTiO 3 collected in a scanning transmission electron microscope (STEM) at atomic column resolution. Acquisition of atomic resolution CBED patterns for this purpose requires careful balancing of CBED feature size in pixels, acquisition speed, and detector dynamic range. The training datasets were derived from multislice simulations, which must be convolved with incoherent source broadening. Sample thicknesses were also determined using quantitative high-angle annular dark-field (HAADF) STEM images acquired simultaneously. The regression CNN performed well on sample thinner than 35 nm, with 70% of the CNN results within 1 nm of HAADF thickness, and 1.0 nm overall root mean square error between the two measurements. The classification CNN was trained for a thicknesses up to 100 nm and yielded 66% of CNN results within one classification increment of 2 nm of HAADF thickness. Our approach depends on methods from computer vision including transfer learning and image augmentation.

36 MATERIALS SCIENCE↗

Image processing pipeline for AI-driven nanoparticle megalibrary characterization

Recent innovations have made it possible to produce megalibraries, millions of structurally and compositionally distinct nanoparticles on a chip. These megalibraries yield vast volumes of data that are impossible to analyze manually, necessitating the development of automated tools. In previous work, we created a binary classification machine learning model to select quality nanoparticle images for downstream analysis. In this work, we show that adding a custom image processing step before training can produce significantly higher-performing models in a fraction of the time and make them more robust to different image noise levels and microscope acquisition settings. The image processing pipeline proposed here effectively cleans raw nanoparticle images, enhances key features, and allows us to use much lower resolution images and simpler neural network model architectures. These features result in higher performance and significant cost savings. Experiments demonstrate superior performance relative to baseline, including an 18.2% improvement in recall and a 13.1% increase in accuracy. Given the high cost of downstream analysis, it is critical to minimize false positives, and our best-performing model reaches a precision of 95.9% and a weighted F-score of 95.1% on an unseen test set. Additionally, model training time is reduced from hours to less than a minute. We also show that, using this custom image processing pipeline, model performance is significantly improved at lower pixel resolutions compared to downsizing alone. We expect that adopting this pipeline for AI-driven automated nanoparticle characterization will allow researchers to rapidly and accurately analyze much greater volumes of data, thereby accelerating materials discovery.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Automatic Detection and Classification of Radio Galaxy Images by Deep Learning

Abstract Surveys conducted by radio astronomy observatories, such as SKA, MeerKAT, Very Large Array, and ASKAP, have generated massive astronomical images containing radio galaxies (RGs). This generation of massive RG images has imposed strict requirements on the detection and classification of RGs and makes manual classification and detection increasingly difficult, even impossible. Rapid classification and detection of images of different types of RGs help astronomers make full use of the observed astronomical image data for further processing and analysis. The classification of FRI and FRII is relatively easy, and there are more studies and literature on them at present, but FR0 and FRI are similar, so it is difficult to distinguish them. It poses a greater challenge to image processing. At present, deep learning has made breakthrough progress in the field of image analysis and processing and has preliminary applications in astronomical data processing. Compared with classification algorithms that can only classify galaxies, object detection algorithms that can locate and classify RGs simultaneously are preferred. In target detection algorithms, YOLOv5 has outstanding advantages in the classification and positioning of small targets. Therefore, we propose a deep-learning method based on an improved YOLOv5 object detection model that makes full use of multisource data, combining FIRST radio with SDSS optical image data, and realizes the automatic detection of FR0, FRI, and FRII RGs. The innovation of our work is that on the basis of the original YOLOv5 object detection model, we introduce the SE Net attention mechanism, increase the number of preset anchors, adjust the network structure of the feature pyramid, and modify the network structure, thereby allowing our model to demonstrate galaxy classification and position detection effects. Our improved model produces satisfactory results, as evidenced by experiments. Overall, the mean average precision (mAP@0.5) of our improved model on the test set reaches 89.4%, which can determine the position (R.A. and decl.) and automatically detect and classify FR0s, FRIs, and FRIIs. Our work contributes to astronomy because it allows astronomers to locate FR0, FRI, and FRII galaxies in a relatively short time and can be further combined with other astronomically generated data to study the properties of these galaxies. The target detection model can also help astronomers find FR0s, FRIs, and FRIIs in future surveys and build a large-scale star RG catalog. Moreover, our work is also useful for the detection of other types of galaxies.

Astronomy & Astrophysics↗

Revisiting Power Systems Time-Domain Simulation Methods and Models

The changing nature of power systems dynamics is challenging present practices related to modeling and study of system-level dynamic behavior. While developing new techniques and models to handle the new modeling requirements, it is also critical to review some of the terminology used to describe existing simulation approaches and the embedded assumptions. This article provides a first-principles review of the simplifications and transformations commonly used in the formulation of time-domain simulation models. It introduces a taxonomy and classification of time-domain simulation models depending on their frequency bandwidth, network representation, and software availability. Furthermore, it focuses on the fundamental aspects of averaging techniques, and model reduction approaches that result in modeling choices, and discusses the associated challenges and opportunities of applying these methods in systems with large shares of Inverter Based Resources (IBRs). The article concludes with an illustrative simulation that compares the trajectories of an IBR-dominated system.

behavioral sciences↗

Panel-Segmentation [SWR-21-18]

Panel-Segmentation contains the scripts for automated metadata extraction of solar PV installations, using satellite imagery coupled with computer vision techniques. In this package, the user can perform the following actions: *Automatically generate a satellite image using a set of lat-long coordinates, and a Google Maps API key. Users would need to set up a Google Cloud account and get a Maps Static API key. Please refer to Setting Up Google Maps Static API Key section for this process. *Perform image segmentation on the satellite image, to locate the solar array(s) in the image on a pixel-by-pixel basis, using an image segmentation model (panel_detection_model.pth). Get classification of the installation (rooftop, ground mounted fixed-tilt or tracking, carport, etc). *Perform azimuth estimation on each solar array cluster in the masked image. *Detect solar panels and get its latitude, longitude, and address within a geographic bounding box through the SOL-Searcher Pipeline. *Detect and calculate hurricane damage on solar installations given pre-hurricane and post-hurricane satellite imagery through the Hurricane Detection Pipeline. *Detect and calculate hail damage on solar installations given satellite imagery through the Hail Detection pipeline. *Convert NOAA MESH (Maximum Estimated Size of Hail) grib2 files into kml or geojson files. *Estimate tilt and azimuth of a solar array by processing USGS LiDAR data for the array’s location.

Edun, Ayobami↗

Gauntlet

Gauntlet (Geographic Augmentation of Extracted Building Features Tool) generates 65 measures of a building’s morphology. These morphology features can be used for various classification tasks and modeling the built environment.

Hauser, Taylor [Oak Ridge National Laboratory (ORN↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally corrected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a surrogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [Illinois U., Chicago]↗