Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Decoding Ethiopian Abodes: Towards Classifying Buildings by Occupancy Type Using Footprint Morphology

Building occupancy classification plays a crucial role in urban planning, disaster management, and population modeling. Traditional methods often require extensive field surveys or detailed datasets, which can be time-consuming, expensive, and may yield incomplete or erroneous data. In this paper, we present a novel approach for classifying buildings as residential or non-residential using only building footprint data. By extracting geometric shape derivatives that characterize building morphology, we developed a high-accuracy classification model employing a combination of unsupervised and supervised learning methods. We utilized open-source data from Open Street Map, aggregating it to create binary labels for buildings based on their respective human use type. Our approach demonstrates the potential for scalability without the need for additional data sources other than building footprints and labels, offering a more efficient solution for building occupancy classification.

Adams, Daniel↗

Nonlinear dynamics and quantum chaos of a family of kicked p -spin models

Herein we introduce kicked p-spin models describing a family of transverse Ising-like models for an ensemble of spin-1/2 particles with all-to-all p-body interaction terms occurring periodically in time as delta-kicks. This is the natural generalization of the well-studied quantum kicked top (p = 2) [Haake, Kus', and Scharf, Z. Phys. B 65, 381 (1987)]. We fully characterize the classical nonlinear dynamics of these models, including the transition to global Hamiltonian chaos. The classical analysis allows us to build a classification for this family of models, distinguishing between p = 2 and p > 2, and between models with odd and even p's. Quantum chaos in these models is characterized in both kinematic and dynamic signatures. For the latter, we show numerically that the growth rate of the out-of-time-order correlator is dictated by the classical Lyapunov exponent. Finally, we argue that the classification of these models constructed in the classical system applies to the quantum system as well.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

High-throughput validation of phase formability and simulation accuracy of Cantor alloys

High-throughput methods enable accelerated discovery of novel materials in complex systems such as high-entropy alloys, which exhibit intricate phase stability across vast compositional spaces. Computational approaches, including Density Functional Theory (DFT) and calculation of phase diagrams (CALPHAD), facilitate screening of phase formability as a function of composition and temperature. However, the integration of computational predictions with experimental validation remains challenging in high-throughput studies. In this work, we introduce a quantitative confidence metric to assess the agreement between predictions and experimental observations, providing a quantitative measure of the confidence of machine learning models trained on either DFT or CALPHAD input in accounting for experimental evidence. The experimental dataset was generated via high-throughput in-situ synchrotron X-ray diffraction on compositionally varied FeNiMnCr alloy libraries, heated from room temperature to ~1000 °C. Agreement between the observed and predicted phases was evaluated using either temperature-independent phase classification or a model that incorporates a temperature-dependent probability of phase formation. This integrated approach demonstrates where strong overall agreement between computation and experiment exists, while also identifying key discrepancies, particularly in FCC/BCC predictions at Mn-rich regions to inform future model refinement.

36 - MATERIALS SCIENCE↗

Deep active learning for classifying cancer pathology reports

Abstract Background Automated text classification has many important applications in the clinical setting; however, obtaining labelled data for training machine learning and deep learning models is often difficult and expensive. Active learning techniques may mitigate this challenge by reducing the amount of labelled data required to effectively train a model. In this study, we analyze the effectiveness of 11 active learning algorithms on classifying subsite and histology from cancer pathology reports using a Convolutional Neural Network as the text classification model. Results We compare the performance of each active learning strategy using two differently sized datasets and two different classification tasks. Our results show that on all tasks and dataset sizes, all active learning strategies except diversity-sampling strategies outperformed random sampling, i.e., no active learning. On our large dataset (15K initial labelled samples, adding 15K additional labelled samples each iteration of active learning), there was no clear winner between the different active learning strategies. On our small dataset (1K initial labelled samples, adding 1K additional labelled samples each iteration of active learning), marginal and ratio uncertainty sampling performed better than all other active learning techniques. We found that compared to random sampling, active learning strongly helps performance on rare classes by focusing on underrepresented classes. Conclusions Active learning can save annotation cost by helping human annotators efficiently and intelligently select which samples to label. Our results show that a dataset constructed using effective active learning techniques requires less than half the amount of labelled data to achieve the same performance as a dataset constructed using random sampling.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗

Design of Digital Twin Sensing Strategies Via Predictive Modeling and Interpretable Machine Learning

This work develops a methodology for sensor placement and dynamic sensor scheduling decisions for digital twins. The digital twin data assimilation is posed as a classification problem, and predictive models are used to train optimal classification trees that represent the map from observed data to estimated digital twin states. In addition to providing a rapid digital twin updating capability, the resulting classification trees yield an interpretable mathematical representation that can be queried to inform sensor placement and sensor scheduling decisions. The proposed approach is demonstrated for a structural digital twin of a 12 ft wingspan unmanned aerial vehicle. Offline, training data are generated by simulating scenarios using predictive reduced-order models of the vehicle in a range of structural states. Furthermore, these training data can be further augmented using experimental or other historical data. In operation, the trained classifier is applied to observational data from the physical vehicle, enabling rapid adaptation of the digital twin in response to changes in structural health. Within this context, we study the performance of the optimal tree classifiers and demonstrate how they enable explainable structural assessments from sparse sensor measurements and also inform optimal sensor placement.

47 OTHER INSTRUMENTATION↗

One of These Things IS Like the Other: Pursuing a New Taxonomy of Industry for Improved Energy System Modeling

Industrial processes drive the exchange of materials, energy, and currency throughout the economy. These processes are powered by electricity and direct combustion, with variation in their operation even within the same industry. This heterogeneity makes it difficult for large models, including the National Energy Modeling System (US), to project their energy use while remaining tractable. Decarbonization and ensuing changes to the energy system require changes to industrial processes while offering opportunities for process innovation, but the extent and nature of changes are difficult to model with current classification schemes and corresponding data. The North American Industrial Classification (NAICS) is an economic taxonomy of industries, but its categories are less meaningful from an energy and material flow perspective. For example, a facility that makes steel from iron ore in a blast furnace/basic oxygen furnace is categorized under the same NAICS code as a facility that makes steel from scrap in an electric arc furnace despite the scale, use of recycled scrap versus iron ore, and energy use differences in the two facility types. Exploratory analysis is performed on a large dataset used for plant-level energy assessment in order to detect clusters that can aid in better modeling of industry for energy analysis in an evolving system with breakthrough technologies.

28 EE - Advanced Manufacturing Office (EE-5A)↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

The Earth Model Column Collaboratory (EMC 2 ) v1.1: an open-source ground-based lidar and radar instrument simulator and subcolumn generator for large-scale models

Abstract. Climate models are essential for our comprehensive understanding of Earth's atmosphere and can provide critical insights on future changes decades ahead. Because of these critical roles, today's climate models are continuously being developed and evaluated using constraining observations and measurements obtained by satellites, airborne, and ground-based instruments. Instrument simulators can provide a bridge between the measured or retrieved quantities and their sampling in models and field observations while considering instrument sensitivity limitations. Here we present the Earth Model Column Collaboratory (EMC2), an open-source ground-based lidar and radar instrument simulator and subcolumn generator, specifically designed for large-scale models, in particular climate models, but also applicable to high-resolution model output. EMC2 provides a flexible framework enabling direct comparison of model output with ground-based observations, including generation of subcolumns that may statistically represent finer model spatial resolutions. In addition, EMC2 emulates ground-based (and air- or space-borne) measurements while remaining faithful to large-scale models' physical assumptions implemented in their cloud or radiation schemes. The simulator uses either single particle or bulk particle size distribution lookup tables, depending on the selected scheme approach, to perform the forward calculations. To facilitate model evaluation, EMC2 also includes three hydrometeor classification methods, namely, radar- and sounding-based cloud and precipitation detection and classification, lidar-based phase classification, and a Cloud Feedback Model Intercomparison Project Observational Simulator Package (COSP) lidar simulator emulator. The software is written in Python, is easy to use, and can be straightforwardly customized for different models, radars, and lidars. Following the description of the logic, functionality, features, and software structure of EMC2, we present a case study of highly supercooled mixed-phase cloud based on measurements from the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) West Antarctic Radiation Experiment (AWARE). We compare observations with the application of EMC2 to outputs from four configurations of the NASA Goddard Institute for Space Studies (GISS) climate model (ModelE3) in single-column model (SCM) mode and from a large-eddy simulation (LES) model. We show that two of the four ModelE3 configurations can form and maintain highly supercooled precipitating cloud for several hours, consistent with observations and LES. While our focus is on one of these ModelE3 configurations, which performed slightly better in this case study, both of these configurations and the LES results post-processed with EMC2 generally provide reasonable agreement with observed lidar and radar variables. As briefly demonstrated here, EMC2 can provide a lightweight and flexible framework for comparing the results of both large-scale and high-resolution models directly with observations, with relatively little overhead and multiple options for achieving consistency with model microphysical or radiation scheme physics.

58 GEOSCIENCES↗

On the discovery of stars, quasars, and galaxies in the Southern Hemisphere with S-PLUS DR2

ABSTRACT This paper provides a catalogue of stars, quasars, and galaxies for the Southern Photometric Local Universe Survey Data Release 2 (S-PLUS DR2) in the Stripe 82 region. We show that a 12-band filter system (5 Sloan-like and 7 narrow bands) allows better performance for object classification than the usual analysis based solely on broad bands (regardless of infrared information). Moreover, we show that our classification is robust against missing values. Using spectroscopically confirmed sources retrieved from the Sloan Digital Sky Survey DR16 and DR14Q, we train a random forest classifier with the 12 S-PLUS magnitudes + 4 morphological features. A second random forest classifier is trained with the addition of the W1 (3.4 $\mu\mathrm{m} $) and W2 (4.6 $\mu\mathrm{m} $) magnitudes from the Wide-field Infrared Survey Explorer (WISE). Forty-four per cent of our catalogue have WISE counterparts and are provided with classification from both models. We achieve 95.76 per cent (52.47 per cent) of quasar purity, 95.88 per cent (92.24 per cent) of quasar completeness, 99.44 per cent (98.17 per cent) of star purity, 98.22 per cent (78.56 per cent) of star completeness, 98.04 per cent (81.39 per cent) of galaxy purity, and 98.8 per cent (85.37 per cent) of galaxy completeness for the first (second) classifier, for which the metrics were calculated on objects with (without) WISE counterpart. A total of 2926 787 objects that are not in our spectroscopic sample were labelled, obtaining 335 956 quasars, 1347 340 stars, and 1243 391 galaxies. From those, 7.4 per cent, 76.0 per cent, and 58.4 per cent were classified with probabilities above 80 per cent. The catalogue with classification and probabilities for Stripe 82 S-PLUS DR2 is available for download.

79 ASTRONOMY AND ASTROPHYSICS↗

Atomic resolution convergent beam electron diffraction analysis using convolutional neural networks

Two types of convolutional neural network (CNN) models, a discrete classification network and a continuous regression network, were trained to determine local sample thickness from convergent beam diffraction (CBED) patterns of SrTiO 3 collected in a scanning transmission electron microscope (STEM) at atomic column resolution. Acquisition of atomic resolution CBED patterns for this purpose requires careful balancing of CBED feature size in pixels, acquisition speed, and detector dynamic range. The training datasets were derived from multislice simulations, which must be convolved with incoherent source broadening. Sample thicknesses were also determined using quantitative high-angle annular dark-field (HAADF) STEM images acquired simultaneously. The regression CNN performed well on sample thinner than 35 nm, with 70% of the CNN results within 1 nm of HAADF thickness, and 1.0 nm overall root mean square error between the two measurements. The classification CNN was trained for a thicknesses up to 100 nm and yielded 66% of CNN results within one classification increment of 2 nm of HAADF thickness. Our approach depends on methods from computer vision including transfer learning and image augmentation.

36 MATERIALS SCIENCE↗

Image processing pipeline for AI-driven nanoparticle megalibrary characterization

Recent innovations have made it possible to produce megalibraries, millions of structurally and compositionally distinct nanoparticles on a chip. These megalibraries yield vast volumes of data that are impossible to analyze manually, necessitating the development of automated tools. In previous work, we created a binary classification machine learning model to select quality nanoparticle images for downstream analysis. In this work, we show that adding a custom image processing step before training can produce significantly higher-performing models in a fraction of the time and make them more robust to different image noise levels and microscope acquisition settings. The image processing pipeline proposed here effectively cleans raw nanoparticle images, enhances key features, and allows us to use much lower resolution images and simpler neural network model architectures. These features result in higher performance and significant cost savings. Experiments demonstrate superior performance relative to baseline, including an 18.2% improvement in recall and a 13.1% increase in accuracy. Given the high cost of downstream analysis, it is critical to minimize false positives, and our best-performing model reaches a precision of 95.9% and a weighted F-score of 95.1% on an unseen test set. Additionally, model training time is reduced from hours to less than a minute. We also show that, using this custom image processing pipeline, model performance is significantly improved at lower pixel resolutions compared to downsizing alone. We expect that adopting this pipeline for AI-driven automated nanoparticle characterization will allow researchers to rapidly and accurately analyze much greater volumes of data, thereby accelerating materials discovery.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Automatic Detection and Classification of Radio Galaxy Images by Deep Learning

Abstract Surveys conducted by radio astronomy observatories, such as SKA, MeerKAT, Very Large Array, and ASKAP, have generated massive astronomical images containing radio galaxies (RGs). This generation of massive RG images has imposed strict requirements on the detection and classification of RGs and makes manual classification and detection increasingly difficult, even impossible. Rapid classification and detection of images of different types of RGs help astronomers make full use of the observed astronomical image data for further processing and analysis. The classification of FRI and FRII is relatively easy, and there are more studies and literature on them at present, but FR0 and FRI are similar, so it is difficult to distinguish them. It poses a greater challenge to image processing. At present, deep learning has made breakthrough progress in the field of image analysis and processing and has preliminary applications in astronomical data processing. Compared with classification algorithms that can only classify galaxies, object detection algorithms that can locate and classify RGs simultaneously are preferred. In target detection algorithms, YOLOv5 has outstanding advantages in the classification and positioning of small targets. Therefore, we propose a deep-learning method based on an improved YOLOv5 object detection model that makes full use of multisource data, combining FIRST radio with SDSS optical image data, and realizes the automatic detection of FR0, FRI, and FRII RGs. The innovation of our work is that on the basis of the original YOLOv5 object detection model, we introduce the SE Net attention mechanism, increase the number of preset anchors, adjust the network structure of the feature pyramid, and modify the network structure, thereby allowing our model to demonstrate galaxy classification and position detection effects. Our improved model produces satisfactory results, as evidenced by experiments. Overall, the mean average precision (mAP@0.5) of our improved model on the test set reaches 89.4%, which can determine the position (R.A. and decl.) and automatically detect and classify FR0s, FRIs, and FRIIs. Our work contributes to astronomy because it allows astronomers to locate FR0, FRI, and FRII galaxies in a relatively short time and can be further combined with other astronomically generated data to study the properties of these galaxies. The target detection model can also help astronomers find FR0s, FRIs, and FRIIs in future surveys and build a large-scale star RG catalog. Moreover, our work is also useful for the detection of other types of galaxies.

Astronomy & Astrophysics↗

Revisiting Power Systems Time-Domain Simulation Methods and Models

The changing nature of power systems dynamics is challenging present practices related to modeling and study of system-level dynamic behavior. While developing new techniques and models to handle the new modeling requirements, it is also critical to review some of the terminology used to describe existing simulation approaches and the embedded assumptions. This article provides a first-principles review of the simplifications and transformations commonly used in the formulation of time-domain simulation models. It introduces a taxonomy and classification of time-domain simulation models depending on their frequency bandwidth, network representation, and software availability. Furthermore, it focuses on the fundamental aspects of averaging techniques, and model reduction approaches that result in modeling choices, and discusses the associated challenges and opportunities of applying these methods in systems with large shares of Inverter Based Resources (IBRs). The article concludes with an illustrative simulation that compares the trajectories of an IBR-dominated system.

behavioral sciences↗

Panel-Segmentation [SWR-21-18]

Panel-Segmentation contains the scripts for automated metadata extraction of solar PV installations, using satellite imagery coupled with computer vision techniques. In this package, the user can perform the following actions: *Automatically generate a satellite image using a set of lat-long coordinates, and a Google Maps API key. Users would need to set up a Google Cloud account and get a Maps Static API key. Please refer to Setting Up Google Maps Static API Key section for this process. *Perform image segmentation on the satellite image, to locate the solar array(s) in the image on a pixel-by-pixel basis, using an image segmentation model (panel_detection_model.pth). Get classification of the installation (rooftop, ground mounted fixed-tilt or tracking, carport, etc). *Perform azimuth estimation on each solar array cluster in the masked image. *Detect solar panels and get its latitude, longitude, and address within a geographic bounding box through the SOL-Searcher Pipeline. *Detect and calculate hurricane damage on solar installations given pre-hurricane and post-hurricane satellite imagery through the Hurricane Detection Pipeline. *Detect and calculate hail damage on solar installations given satellite imagery through the Hail Detection pipeline. *Convert NOAA MESH (Maximum Estimated Size of Hail) grib2 files into kml or geojson files. *Estimate tilt and azimuth of a solar array by processing USGS LiDAR data for the array’s location.

Edun, Ayobami↗

Gauntlet

Gauntlet (Geographic Augmentation of Extracted Building Features Tool) generates 65 measures of a building’s morphology. These morphology features can be used for various classification tasks and modeling the built environment.

Hauser, Taylor [Oak Ridge National Laboratory (ORN↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally corrected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a surrogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [Illinois U., Chicago]↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗