Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Adversarial Attacks on Deep Neural Network-based Power System Event Classification Models

Online event classification is essential to strengthening the reliability of the power transmission system. Recently, deep learning based methods have achieved great success in numerous domains such as computer vision and natural language processing. Researchers began to adopt deep learning based methods to solve the power system event identification problem and achieved effective results. However, these previous works do not consider that deep learning models are vulnerable to adversarial attacks, potentially influencing real-world applications' reliability. In this paper, we adopt several adversarial attack mechanisms by adding tailored noise signal to the input Phasor Measurement Units (PMU) time series and make the deep learning model misclassify the power system event. This numerical study discloses that current state-of-the-art deep learning based power system event classifiers are extremely vulnerable to adversarial attacks, which may jeopardize the reliability of the power transmission system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

Reference Shapefiles and Pre-trained Random Forest Classification Models for Detecting Aufeis on the North Slope of Alaska in Landsat Imagery

This dataset provides shapefiles and trained machine learning models used for aufeis detection at four sites on the North Slope of Alaska. It includes reference data for evaluating Landsat-based detection methods, supporting research on remote sensing approaches for identifying aufeis. The ReferenceData folder contains ArcGIS shapefiles of semi-automated land cover classifications for 217 Landsat Collection 2 images, categorizing pixels into six classes: aufeis, snow, ground, none, water, and cloud. The SiteBuffers.zip file includes 10-kilometer buffer shapefiles defining regions of interest around four aufeis fields (Canning21, FH1, Firth, and Kuparuk), used to test three detection techniques. Additionally, the TrainedRFModels folder contains six pre-trained Scikit-Learn Random Forest classifiers (100 trees, max depth = 30) designed to predict aufeis presence in Landsat Collection 2 Surface Reflectance images using Red, Blue, SWIR2, NDVI, and NDWI bands. This dataset supports the development and validation of remote sensing methods for mapping aufeis in Arctic environments.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

A Novel Method to Train Classification Models for Structure Detection in In Situ Spacecraft Data

We present a method for creating spacecraft-like data which can be used to train Machine Learning (ML) models to detect and classify structures in in situ spacecraft data. First, we use the Grad-Shafranov equation to numerically solve for several magnetohydrostatic equilibria which are variations on a known analytic equilibrium. These equilibria are then used as the initial conditions for Particle-In-Cell simulations in which the structures of interest are observed and labeled. We then take one-dimensional slices through the simulations to replicate what a spacecraft collecting data from the simulation would observe. This sliced data then can be used as training data for the initial training of ML models intended for use on spacecraft data. We demonstrate the method applied to the problem of detecting small-scale plasmoids in the magnetotail, which is important for understanding complex magnetotail reconnection dynamics. The simple 1D classifier we train is able to detect more than 70% of the plasmoid points in the data set but also produces a large number of false positives. Our further work on this example problem is detailed, and further potential uses of the method are discussed.

79 ASTRONOMY AND ASTROPHYSICS↗

Land use classification for hydrologic models using interactive machine classification of LANDSAT data

Models designed to simulate the hydrology of urban areas require input parameters describing the land use and degree of imperviousness of the watershed. Unfortunately, the magnitude and spatial distribution of these parameters are rather difficult to estimate when a large watershed is involved. Trade-offs between accuracy of the model parameters and the time or money available for their determination must be made. Because of the necessity of such trade-offs, a study was developed to investigate the use of computer aided analysis of LANDSAT multispectral data in estimating percent of imperviousness and associated land uses needed in urban hydrologic modeling. An interactive computer was used to delineate seven land use classifications in the 342 sq. km. Maryland portion of the Anacostia River Basin from LANDSAT data. These results compared favorably with those of an earlier study which obtained the same information through analysis of aerial photographs having a scale of 1:4800. Approximately 94 man days were required to complete the land use analysis using the aerial photographs while less than three man days were required to accomplish similar tasks using the LANDSAT data.

Thomas J. Jackson↗

Active microwave responses - An aid in improved crop classification

A study determined the feasibility of using visible, infrared, and active microwave data to classify agricultural crops such as corn, sorghum, alfalfa, wheat stubble, millet, shortgrass pasture and bare soil. Visible through microwave data were collected by instruments on board the NASA C-130 aircraft over 40 agricultural fields near Guymon, OK in 1978 and Dalhart, TX in 1980. Results from stepwise and discriminant analysis techniques indicated 4.75 GHz, 1.6 GHz, and 0.4 GHz cross-polarized microwave frequencies were the microwave frequencies most sensitive to crop type differences. Inclusion of microwave data in visible and infrared classification models improved classification accuracy from 73 percent to 92 percent. Despite the results, further studies are needed during different growth stages to validate the visible, infrared, and active microwave responses to vegetation.

Rosenthal, W. D.↗

Mapping the Extent of Mangrove Ecosystem Degradation by Integrating an Ecological Conceptual Model with Satellite Data

Anthropogenic and natural disturbances can cause degradation of ecosystems, reducing their capacity to sustain biodiversity and provide ecosystem services. Understanding the extent of ecosystem degradation is critical for estimating risks to ecosystems, yet there are few existing methods to map degradation at the ecosystem scale and none using freely available satellite data for mangrove ecosystems. In this study, we developed a quantitative classification model of mangrove ecosystem degradation using freely available earth observation data. Crucially, a conceptual model of mangrove ecosystem degradation was established to identify suitable remote sensing variables that support the quantitative classification model, bridging the gap between satellite-derived variables and ecosystem degradation with explicit ecological links. We applied our degradation model to two case-studies, the mangroves of Rakhine State, Myanmar, which are severely threatened by anthropogenic disturbances, and Shark River within the Everglades National Park, USA, which is periodically disturbed by severe tropical storms. Our model suggested that 40% (597 km2) of the extent of mangroves in Rakhine showed evidence of degradation. In the Everglades, the model suggested that the extent of degraded mangrove forest increased from 5.1% to 97.4% following the Category 4 Hurricane Irma in 2017. Quantitative accuracy assessments indicated the model achieved overall accuracies of 77.6% and 79.1% for the Rakhine and the Everglades, respectively. We highlight that using an ecological conceptual model as the basis for building quantitative classification models to estimate the extent of ecosystem degradation ensures the ecological relevance of the classification models. Our developed method enables researchers to move beyond only mapping ecosystem distribution to condition and degradation as well. These results can help support ecosystem risk assessments, natural capital accounting, and restoration planning and provide quantitative estimates of ecosystem degradation for new global biodiversity targets.

Calvin K. F. Lee↗

Computer discrimination procedures applicable to aerial and ERTS multispectral data

Two statistical models are compared in the classification of crops recorded on color aerial photographs. A theory of error ellipses is applied to the pattern recognition problem. An elliptical boundary condition classification model (EBC), useful for recognition of candidate patterns, evolves out of error ellipse theory. The EBC model is compared with the minimum distance to the mean (MDM) classification model in terms of pattern recognition ability. The pattern recognition results of both models are interpreted graphically using scatter diagrams to represent measurement space. Measurement space, for this report, is determined by optical density measurements collected from Kodak Ektachrome Infrared Aero Film 8443 (EIR). The EBC model is shown to be a significant improvement over the MDM model.

Richardson, A. J.↗

Computer identification of ground pattern from aerial photographs.

Two statistical models are comapared in the classification of crops recorded on color aerial photographs. A theory of error ellipses is applied to the pattern recognition problem. An elliptical boundary condition classification model (EBC), useful for recognition of candidate patterns, evolves out of error ellipse theory. The EBC model is compared with the minimum distance to the mean (MDM) classification model in terms of pattern recognition ability. The pattern recognition results of both models are interpreted graphically using scatter diagrams to represent measurement space. Measurement space, for this report, is determined by optical density measurements collected from Kodak Ektachrome Infrared Aero Film 8443 (EIR). The EBC model is shown to be a significant improvement over the MDM model.

Richardson, A. J.↗

Performance Evaluation of Vertical Federated Machine Learning Against Adversarial Threats on Wide-Area Control System: Preprint

Federated machine learning (FL) is gaining significant popularity to develop cybersecurity solutions in power grids because of its advanced capability to support decentralized data handing at local devices, its privacy preservation, and its low-bandwidth requirement. However, the evolving adversarial machine learning (AML) threats raise significant concerns for the cybersecurity of FL architectures. The FL-based split neural network (SplitNN) achieves high performance through the decentralized training of local neural network models while preserving data privacy across multiple entities. In this paper, we propose a methodology for evaluating the performance of a vertical FLbased anomaly detector against different types of AML attacks, including denial-of-service attacks, adversarial data injection attacks, and replay attacks on the trained local models deployed in the grid network. For a case study, we consider the modified IEEE 13-bus system, and we develop SplitNN-based binary and multiclass classification models to detect, locate, and identify different types of data integrity attacks on the volt-watt control with two pooling layers: maximum pooling and AvgPool. Our experimental results, computed through performance metrics, reveal that the severity of these AML attacks varies with the integrated pooling mechanism, the type of classification model, and the nature of the cyberattack. Further, the AML attacks negatively impacted the prediction time per sample for the pretrained SplitNN during the online testing.

adversarial threats↗

Investigating resource-efficient neutron/gamma classification ML models targeting eFPGAs

There has been considerable interest and resulting progress in implementing machine learning (ML) models in hardware over the last several years from the particle and nuclear physics communities. A big driver has been the release of the Python package, hls4ml, which has enabled porting models specified and trained using Python ML libraries to register transfer level (RTL) code. So far, the primary end targets have been commercial field-programmable gate arrays (FPGAs) or synthesized custom blocks on application specific integrated circuits (ASICs). However, recent developments in open-source embedded FPGA (eFPGA) frameworks now provide an alternate, more flexible pathway for implementing ML models in hardware. These customized eFPGA fabrics can be integrated as part of an overall chip design. In general, the decision between a fully custom, eFPGA, or commercial FPGA ML implementation will depend on the details of the end-use application. In this work, we explored the parameter space for eFPGA implementations of fully-connected neural network (fcNN) and boosted decision tree (BDT) models using the task of neutron/gamma classification with a specific focus on resource efficiency. We used data collected using an AmBe sealed source incident on Stilbene, which was optically coupled to an OnSemi J-series silicon photomultiplier (SiPM) to generate training and test data for this study. We investigated relevant input features and the effects of bit-resolution and sampling rate as well as trade-offs in hyperparameters for both ML architectures while tracking total resource usage. The performance metric used to track model performance was the calculated neutron efficiency at a gamma leakage of 10 -3 . The results of the study will be used to aid the specification of an eFPGA fabric, which will be integrated as part of a test chip.

47 OTHER INSTRUMENTATION↗

Understanding the Effect of Land Cover Classification on Model Estimates of Regional Carbon Cycling in the Boreal Forest Biome

The original objectives of this proposed 3-year project were to: 1) quantify the respective contributions of land cover and disturbance (i.e., wild fire) to uncertainty associated with regional carbon source/sink estimates produced by a variety of boreal ecosystem models; 2) identify the model processes responsible for differences in simulated carbon source/sink patterns for the boreal forest; 3) validate model outputs using tower and field- based estimates of NEP and NPP; and 4) recommend/prioritize improvements to boreal ecosystem carbon models, which will better constrain regional source/sink estimates for atmospheric C02. These original objectives were subsequently distilled to fit within the constraints of a 1 -year study. This revised study involved a regional model intercomparison over the BOREAS study region involving Biome-BGC, and TEM (A.D. McGuire, UAF) ecosystem models. The major focus of these revised activities involved quantifying the sensitivity of regional model predictions associated with land cover classification uncertainties. We also evaluated the individual and combined effects of historical fire activity, historical atmospheric CO2 concentrations, and climate change on carbon and water flux simulations within the BOREAS study region.

Kimball, John↗

Line Faults Classification Using Machine Learning on Three Phase Voltages Extracted from Large Dataset of PMU Measurements

An end-to-end supervised learning method is developed to classify transmission line faults in a twoyear field-recorded dataset that includes synchronized measurements of three-phase voltages recorded by 38 Phasor Measurement Units (PMU) sparsely located in in the US Western Grid interconnection. Statistical analysis is performed to extract features from this large dataset to train Support Vector Machine (SVM), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) classifiers initially. The training further leverages a simulated dataset from a synthetic grid with 12 PMUs to increase the number of faults of types infrequently seen in the field-recorded dataset. Training the classification models with the combined dataset resulted in a classification accuracy of 97.7%. This is a significant improvement over 89.7% to 92.5% accuracy obtained by relying on the field-recorded dataset alone.

47 OTHER INSTRUMENTATION↗

Data resolution versus forestry classification and modeling

This paper examines the effects on timber stand computer classification accuracies caused by changes in the resolution of remotely sensed multispectral data. This investigation is valuable, especially for determining optimal sensor and platform designs. Theoretical justification and experimental verification support the finding that classification accuracies for low resolution data could be better than the accuracies for data with higher resolution. The increase in accuracy is constructed as due to the reduction of scene inhomogeneity at lower resolution. The computer classification scheme was a maximum likelihood classifier.

Kan, E. P.↗

Deformable phrase level attention: A flexible approach for improving AI based medical coding

Objective: Improving the AI-driven automated medical encoding of clinical text plays a vital role in gathering information on the occurrence of diseases to improve population-level health. This work presents a novel attention mechanism designed to enhance text classification models and ensure appropriate classification of medical concepts in unstructured electronic health records. Materials and Methods: We developed a deformable, phrase-level attention mechanism to identify important lexical word-level and contextual phrase-level information from clinical text documents. We evaluated conventional and transformer-based deep learning models that we extended with our attention mechanism on the extraction of critical cancer information (e.g., site, subsite, laterality, histology, behavior) from 629,908 electronic pathology reports and on the automated medical encoding of 52,722 hospital discharge summaries. Results: Transformer-based models with the deformable, phrase-level attention mechanism achieved the best performance on the extraction of critical cancer information from pathology reports. Conventional- and transformer-based models show similar or better performance than their baseline counterparts on the automated medical encoding of clinical documents. Discussion: The addition of phrase-level information allowed models extended with our proposed method to outperform standard word-level attention. Our method showed favorable properties for the real-world application in terms of model robustness and phenotyping. These results indicate that our method is promising for automated data harmonization for common data models. Conclusion: This work proposes a novel deformable, phrase-level attention mechanism that enhances text classification models in the extraction of medical concepts from clinical text documents. We demonstrate strong performances on two clinical text datasets and showcase real-world deployability of our method.

Automated medical encoding↗

Segmentation, modeling and classification of the compact objects in a pile

The problem of interpreting dense range images obtained from the scene of a heap of man-made objects is discussed. A range image interpretation system consisting of segmentation, modeling, verification, and classification procedures is described. First, the range image is segmented into regions and reasoning is done about the physical support of these regions. Second, for each region several possible three-dimensional interpretations are made based on various scenarios of the objects physical support. Finally each interpretation is tested against the data for its consistency. The superquadric model is selected as the three-dimensional shape descriptor, plus tapering deformations along the major axis. Experimental results obtained from some complex range images of mail pieces are reported to demonstrate the soundness and the robustness of our approach.

Gupta, Alok↗