Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Self-Supervised Cloud Classification

Abstract Low-level marine clouds play a pivotal role in Earth’s weather and climate through their interactions with radiation, heat and moisture transport, and the hydrological cycle. These interactions depend on a range of dynamical and microphysical processes that result in a broad diversity of cloud types and spatial structures, and a comprehensive understanding of cloud morphology is critical for continued improvement of our atmospheric modeling and prediction capabilities moving forward. Deep learning has recently accelerated our ability to study clouds using satellite remote sensing, and machine learning classifiers have enabled detailed studies of cloud morphology. A major limitation of deep learning approaches to this problem, however, is the large number of hand-labeled samples that are required for training. This work applies a recently developed self-supervised learning scheme to train a deep convolutional neural network (CNN) to map marine cloud imagery to vector embeddings that capture information about mesoscale cloud morphology and can be used for satellite image classification. The model is evaluated against existing cloud classification datasets and several use cases are demonstrated, including training cloud classifiers with very few labeled samples, interrogation of the CNN’s learned internal feature representations, cross-instrument application, and resilience against sensor calibration drift and changing scene brightness. The self-supervised approach learns meaningful internal representations of cloud structures and achieves comparable classification accuracy to supervised deep learning methods without the expense of creating large hand-annotated training datasets. Significance Statement Marine clouds heavily influence Earth’s weather and climate, and improved understanding of marine clouds is required to improve our atmospheric modeling capabilities and physical understanding of the atmosphere. Recently, deep learning has emerged as a powerful research tool that can be used to identify and study specific marine cloud types in the vast number of images collected by Earth-observing satellites. While powerful, these approaches require hand-labeling of training data, which is prohibitively time intensive. This study evaluates a recently developed self-supervised deep learning method that does not require human-labeled training data for processing images of clouds. We show that the trained algorithm performs competitively with algorithms trained on hand-labeled data for image classification tasks. We also discuss potential downstream uses and demonstrate some exciting features of the approach including application to multiple satellite instruments, resilience against changing image brightness, and its learned internal representations of cloud types. The self-supervised technique removes one of the major hurdles for applying deep learning to very large atmospheric datasets.

54 ENVIRONMENTAL SCIENCES↗

Learning to identify semi-visible jets

We train a network to identify jets with fractional dark decay (semi-visible jets) using the pattern of their low-level jet constituents, and explore the nature of the information used by the network by mapping it to a space of jet substructure observables. Semi-visible jets arise from dark matter particles which decay into a mixture of dark sector (invisible) and Standard Model (visible) particles. Such objects are challenging to identify due to the complex nature of jets and the alignment of the momentum imbalance from the dark particles with the jet axis, but such jets do not yet benefit from the construction of dedicated theoretically-motivated jet substructure observables. A deep network operating on jet constituents is used as a probe of the available information and indicates that classification power not captured by current high-level observables arises primarily from low-p T jet constituents.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗

Quantum computing based hybrid deep learning for fault diagnosis in electrical power systems

Quantum computing (QC) and deep learning have shown promise of supporting transformative advances and have recently gained popularity in a wide range of areas. Here, this paper proposes a hybrid QC-based deep learning framework for fault diagnosis of electrical power systems that combine the feature extraction capabilities of conditional restricted Boltzmann machine with an efficient classification of deep networks. Computational challenges stemming from the complexities of such deep learning models are overcome by QC-based training methodologies that effectively leverage the complementary strengths of quantum assisted learning and classical training techniques. The proposed hybrid QC-based deep learning framework is tested on a simulated electrical power system with 30 buses and wide variations of substation and transmission line faults, to demonstrate the framework’s applicability, efficiency, and generalization capabilities. High computational efficiency is enjoyed by the proposed hybrid approach in terms of computational effort required and quality of diagnosis performance over classical training methods. In addition, superior and reliable fault diagnosis performance with faster response time is achieved over state-of-the-art pattern recognition methods based on artificial neural networks (ANN) and decision trees (DT).

42 ENGINEERING↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Global Geo-processed Data of Aquifer Properties by 0.5° Grid, Country and Water Basins

This repository of global hydrogeologic datasets contains aquifer properties on 0.5° scale, including depth to groundwater (Fan et al., 2013), aquifer thickness (de Graaf et al., 2015), WHYMap aquifer classes (Richts et al., 2011), recharge (Döll and Fiedler, 2008; Gleeson et al., 2016), lakes (Messager et al., 2016), porosity and permeability (Gleeson et al., 2014), digitized and geo-processed from their respective sources. Globally gridded aquifer properties could be used independently to estimate global groundwater availability or used as critical inputs to the superwell model to simulate groundwater extraction and provide estimates of pumped volumes and unit costs under user-specific scenarios. Key resources related to this data are: Niazi, H., Ferencz, S. B., Graham, N. T., Yoon, J., Wild, T. B., Hejazi, M., Watson, D. J., & Vernon, C. R. (2025). Long-term hydro-economic analysis tool for evaluating global groundwater cost and supply: Superwell v1.1. Geoscientific Model Development, 18(5), 1737-1767. https://doi.org/10.5194/gmd-18-1737-2025 superwell model repository which uses this data to simulate groundwater extraction and provides estimates of the global extractable volumes and unit-costs ($/km3) of accessible groundwater production under user-specified extraction scenarios. Repository Overview Main output: aquifer_properties_rec.csv contains all processed outputs, including aquifer properties like porosity, permeability, recharge, lake areas, aquifer thickness, and depth to groundwater. shapefiles.zip: contains all digitized GIS databases and shapefile for all aquifer properties prep_inputs.R and prep_inputs_recharge_lakes.R: R scripts that process the shapefiles to produce the aquifer_properties_rec.csv file plot_inputs.R: R script for plotting the maps and conducting preliminary analysis on the available groundwater volume basin_to_country_mapping.csv, basin_country_region_mapping.csv and continent_county_mapping.csv provide the mapping between continents, 32 energy-economic macro regions, countries, and water basins for post-processing aquifer_properties_rec.csv Maps: Each map visualizes the spatial distribution of one of the aquifer properties across the globe map_in_Porosity.png map_in_Permeability.png map_in_Aquifer_thickness.png map_in_Depth_to_water.png map_in_Recharge.png map_in_Grid_area_km.png map_in_Lake_area_km.png map_in_WHYClass.png Sample inputs sample_inputs.py: this script samples inputs from the aquifer_properties_rec dataset, ensuring the sampled and original inputs maintain the same distributions sampled_data_100.csv contains 100 sampled data points and sampled_data_100.png compares their distributions Dataset Overview The main outputs are consolidated in a comprehensive aquifer_properties_rec.csv file and include the following fields: GridCellID: Unique identifier for each (roughly 0.5°) grid cell Continent: Continent name Country: Country name GCAM_basin_ID: Identifier for GCAM hydrologic basin Basin_long_name: Full name of the basin WHYClass: Hydrogeologic classification based on WHYMap aquifer classes (Richts et al., 2011) Porosity: Soil porosity (%) (Gleeson et al., 2014) Permeability: Soil permeability (in square meters; Gleeson et al., 2014) Aquifer_thickness: Thickness of the aquifer (in meters; de Graaf et al., 2015) Depth_to_water: Depth to groundwater (in meters; Fan et al., 2013) Recharge: long-term annual averaged recharge rates (in m/yr; Döll and Fiedler, 2008; Gleeson et al., 2016) Grid_area: Area of the grid cell (in square meters) Lakes_area: Area of inland lakes (in square meters; Messager et al., 2016) Key References The datasets are digitized versions of global hydrogeologic properties from the following key literature sources: Depth to Groundwater: Fan, Y., Li, H., & Miguez-Macho, G. (2013). Global Patterns of Groundwater Table Depth. Science, 339(6122), 940-943. https://doi.org/10.1126/science.1229881 Aquifer Thickness: de Graaf, I. E. M., Sutanudjaja, E. H., van Beek, L. P. H., & Bierkens, M. F. P. (2015). A high-resolution global-scale groundwater model. Hydrol. Earth Syst. Sci., 19(2), 823-837. https://doi.org/10.5194/hess-19-823-2015 Porosity and Permeability: Gleeson, T., Moosdorf, N., Hartmann, J., & van Beek, L. P. H. (2014). A glimpse beneath earth's surface: GLobal HYdrogeology MaPS (GLHYMPS) of permeability and porosity. Geophysical Research Letters, 41(11), 3891-3898. https://doi.org/10.1002/2014GL059856 Aquifer classes: Richts, A., Struckmeier, W. F., & Zaepke, M. (2011). WHYMAP and the Groundwater Resources Map of the World 1:25,000,000. In J. A. A. Jones (Ed.), Sustaining Groundwater Resources: A Critical Element in the Global Water Crisis (pp. 159-173). Springer Netherlands. https://doi.org/10.1007/978-90-481-3426-7_10 Recharge: Döll, P., & Fiedler, K. (2008). Global-scale modeling of groundwater recharge. Hydrol. Earth Syst. Sci., 12(3), 863-885. https://doi.org/10.5194/hess-12-863-2008; Gleeson, T., Befus, K. M., Jasechko, S., Luijendijk, E., & Cardenas, M. B. (2016). The global volume and distribution of modern groundwater. Nature Geoscience, 9(2), 161-167. https://doi.org/10.1038/ngeo2590 Inland Lakes: Messager, M. L., Lehner, B., Grill, G., Nedeva, I., & Schmitt, O. (2016). Estimating the volume and age of water stored in global lakes using a geo-statistical approach. Nature Communications, 7(1), 13603. https://doi.org/10.1038/ncomms13603 Cite as Niazi, H., Watson, D., Hejazi, M., Yonkofski, C., Ferencz, S., Vernon, C., Graham, N., Wild, T., & Yoon, J. (2024). Global Geo-processed Data of Aquifer Properties by 0.5° Grid, Country and Water Basins. MultiSector Dynamics-Living, Intuitive, Value-adding, Environment. https://doi.org/10.57931/2484226 Contact Reach out to Hassan Niazi or open an issue in the superwell repository for questions or suggestions.

aquifer thickness↗

KG-Hub—building and exchanging biological knowledge graphs

Knowledge graphs (KGs) are a powerful approach for integrating heterogeneous data and making inferences in biology and many other domains, but a coherent solution for constructing, exchanging, and facilitating the downstream use of KGs is lacking. Here we present KG-Hub, a platform that enables standardized construction, exchange, and reuse of KGs. Features include a simple, modular extract–transform–load pattern for producing graphs compliant with Biolink Model (a high-level data model for standardizing biological data), easy integration of any OBO (Open Biological and Biomedical Ontologies) ontology, cached downloads of upstream data sources, versioned and automatically updated builds with stable URLs, web-browsable storage of KG artifacts on cloud infrastructure, and easy reuse of transformed subgraphs across projects. Current KG-Hub projects span use cases including COVID-19 research, drug repurposing, microbial–environmental interactions, and rare disease research. KG-Hub is equipped with tooling to easily analyze and manipulate KGs. KG-Hub is also tightly integrated with graph machine learning (ML) tools which allow automated graph ML, including node embeddings and training of models for link prediction and node classification.

59 BASIC BIOLOGICAL SCIENCES↗

The miniJPAS survey quasar selection – I. Mock catalogues for classification

In this series of papers, we employ several machine learning (ML) methods to classify the point-like sources from the miniJPAS catalogue, and identify quasar candidates. Since no representative sample of spectroscopically confirmed sources exists at present to train these ML algorithms, we rely on mock catalogues. In this first paper, we develop a pipeline to compute synthetic photometry of quasars, galaxies, and stars using spectra of objects targeted as quasars in the Sloan Digital Sky Survey . To match the same depths and signal-to-noise ratio distributions in all bands expected for miniJPAS point sources in the range 17.5 ≤ r < 24, we augment our sample of available spectra by shifting the original r-band magnitude distributions towards the faint end, ensure that the relative incidence rates of the different objects are distributed according to their respective luminosity functions, and perform a thorough modelling of the noise distribution in each filter, by sampling the flux variance either from Gaussian realizations with given widths, or from combinations of Gaussian functions. Finally, we also add in the mocks the patterns of non-detections which are present in all real observations. Although the mock catalogues presented in this work are a first step towards simulated data sets that match the properties of the miniJPAS observations, these mocks can be adapted to serve the purposes of other photometric surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Preserving isohydricity: vertical environmental variability explains Amazon forest water-use strategies

Abstract Increases in hydrological extremes, including drought, are expected for Amazon forests. A fundamental challenge for predicting forest responses lies in identifying ecological strategies which underlie such responses. Characterization of species-specific hydraulic strategies for regulating water-use, thought to be arrayed along an ‘isohydric–anisohydric’ spectrum, is a widely used approach. However, recent studies have questioned the usefulness of this classification scheme, because its metrics are strongly influenced by environments, and hence can lead to divergent classifications even within the same species. Here, we propose an alternative approach positing that individual hydraulic regulation strategies emerge from the interaction of environments with traits. Specifically, we hypothesize that the vertical forest profile represents a key gradient in drought-related environments (atmospheric vapor pressure deficit, soil water availability) that drives divergent tree water-use strategies for coordinated regulation of stomatal conductance (gs) and leaf water potentials (ΨL) with tree rooting depth, a proxy for water availability. Testing this hypothesis in a seasonal eastern Amazon forest in Brazil, we found that hydraulic strategies indeed depend on height-associated environments. Upper canopy trees, experiencing high vapor pressure deficit (VPD), but stable soil water access through deep rooting, exhibited isohydric strategies, defined by little seasonal change in the diurnal pattern of gs and steady seasonal minimum ΨL. In contrast, understory trees, exposed to less variable VPD but highly variable soil water availability, exhibited anisohydric strategies, with fluctuations in diurnal gs that increased in the dry season along with increasing variation in ΨL. Our finding that canopy height structures the coordination between drought-related environmental stressors and hydraulic traits provides a basis for preserving the applicability of the isohydric-to-anisohydric spectrum, which we show here may consistently emerge from environmental context. Our work highlights the importance of understanding how environmental heterogeneity structures forest responses to climate change, providing a mechanistic basis for improving models of tropical ecosystems.

Forestry↗

Metabolic Response in Patients With Post-treatment Lyme Disease Symptoms/Syndrome

Abstract Background Post-treatment Lyme disease symptoms/syndrome (PTLDS) occurs in approximately 10% of patients with Lyme disease following antibiotic treatment. Biomarkers or specific clinical symptoms to identify patients with PTLDS do not currently exist and the PTLDS classification is based on the report of persistent, subjective symptoms for ≥6 months following antibiotic treatment for Lyme disease. Methods Untargeted liquid chromatography–mass spectrometry metabolomics was used to determine longitudinal metabolic responses and biosignatures in PTLDS and clinically cured non-PTLDS Lyme patients. Evaluation of biosignatures included (1) defining altered classes of metabolites, (2) elastic net regularization to define metabolites that most strongly defined PTLDS and non-PTLDS patients at different time points, (3) changes in the longitudinal abundance of metabolites, and (4) linear discriminant analysis to evaluate robustness in a second patient cohort. Results This study determined that observable metabolic differences exist between PTLDS and non-PTLDS patients at multiple time points. The metabolites with differential abundance included those from glycerophospholipid, bile acid, and acylcarnitine metabolism. Distinct longitudinal patterns of metabolite abundance indicated a greater metabolic variability in PTLDS versus non-PTLDS patients. Small numbers of metabolites (6 to 40) could be used to define PTLDS versus non-PTLDS patients at defined time points, and the findings were validated in a second cohort of PTLDS and non-PTLDS patients. Conclusions These data provide evidence that an objective metabolite-based measurement can distinguish patients with PTLDS and help understand the underlying biochemistry of PTLDS.

Immunology↗

Particle hit clustering and identification using point set transformers in liquid argon time projection chambers

Liquid argon time projection chambers are often used in neutrino physics and dark-matter searches because of their high spatial resolution. The images generated by these detectors are extremely sparse, as the energy values detected by most of the detector are equal to 0, meaning that despite their high resolution, most of the detector is unused in a particular interaction. Instead of representing all of the empty detections, the interaction is usually stored as a sparse matrix, a list of detection locations paired with their energy values. Traditional machine learning methods that have been applied to particle reconstruction such as convolutional neural networks (CNNs), however, cannot operate over data stored in this way and therefore must have the matrix fully instantiated as a dense matrix. Operating on dense matrices requires a lot of memory and computation time, in contrast to directly operating on the sparse matrix. We propose a machine learning model using a point set neural network that operates over a sparse matrix, greatly improving both processing speed and accuracy over methods that instantiate the dense matrix, as well as over other methods that operate over sparse matrices. Compared to competing state-of-the-art methods, our method improves classification performance by 14%, segmentation performance by more than 22%, while taking 80% less time and using 66% less memory. Compared to state-of-the-art CNN methods, our method improves classification performance by more than 86%, segmentation performance by more than 71%, while reducing runtime by 91% and reducing memory usage by 61%.

calibration and fitting methods↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Data from "A Bayesian Record Linkage Approach to Applications in Tree Demography Using Overlapping LiDAR Scans"

Processed LiDAR data and environmental covariates from 2015 and 2019 LiDAR scans in the Vicinity of Snodgrass Mountain (Western Colorado, USA), in a geographic subset used in primary analysis for the research paper.This package contains LiDAR-derived canopy height maps for 2015 and 2019, crown polygons derived from the height maps using a segmentation algorithm, and environmental covariates supporting the model of forest growth. Source datasets include August 2015 and August 2019 discrete-return LiDAR point clouds collected by Quantum Geospatial for terrain mapping purposes on behalf of the Colorado Hazard Mapping Program and the Colorado Water Conservation Board. Both datasets adhere to the USGS QL2 quality standard. The point cloud data were processed using the R package lidR to generate a canopy height model representing maximum vegetation height above the ground surface, using a pit-free algorithm.This dataset was compiled to assess how spatial patterns of tree growth in montane and subalpine forests are influenced by water and energy availability. Understanding these growth patterns can provide insight into forest dynamics in the Southern Rocky Mountains under changing climatic conditions.This dataset contains .tif, .csv, and .txt files. This dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Array-Based Machine Learning for Functional Group Detection in Electron Ionization Mass Spectrometry

Mass spectrometry is a ubiquitous technique capable of complex chemical analysis. The fragmentation patterns that appear in mass spectrometry are an excellent target for artificial intelligence methods to automate and expedite the analysis of data to identify targets such as functional groups. To develop this approach, we trained models on electron ionization (a reproducible hard fragmentation technique) mass spectra so that not only the final model accuracies but also the reasoning behind model assignments could be evaluated. The convolutional neural network (CNN) models were trained on 2D images of the spectra using transfer learning of Inception V3, and the logistic regression models were trained using array-based data and Scikit Learn implementation in Python. Our training dataset consisted of 21,166 mass spectra from the United States’ National Institute of Standards and Technology (NIST) Webbook. The data was used to train models to identify functional groups, both specific (e.g., amines, esters) and generalized classifications (aromatics, oxygen-containing functional groups, and nitrogen-containing functional groups). We found that the highest final accuracies on identifying new data were observed using logistic regression rather than transfer learning on CNN models. It was also determined that the mass range most beneficial for functional group analysis is 0–100 m/z. We also found success in correctly identifying functional groups of example molecules selected from both the NIST database and experimental data. Beyond functional group analysis, we also have developed a methodology to identify impactful fragments for the accurate detection of the models’ targets. The results demonstrate a potential pathway for analyzing and screening substantial amounts of mass spectral data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The System for Classification of Low-Pressure Systems (SyCLoPS): An All-In-One Objective Framework for Large-Scale Data Sets

We propose the first unified objective framework (SyCLoPS) for detecting and classifying all types of low-pressure systems (LPSs) in a given data set. We use the state-of-the-art automated feature tracking software TempestExtremes (TE) to detect and track LPS features globally in ERA5 and compute 16 parameters from commonly found atmospheric variables for classification. A Python classifier is implemented to classify all LPSs at once. The framework assigns 16 different labels (classes) to each LPS data point and designates four different types of high-impact LPS tracks, including tracks of tropical cyclone (TC), monsoonal system, subtropical storm and polar low. The classification process involves disentangling high-altitude and drier LPSs, differentiating tropical and non-tropical LPSs using novel criteria, and optimizing for the detection of the four types of high-impact LPS. A comparison of our labels with those in the International Best Track Archive for Climate Stewardship (IBTrACS) revealed an overall accuracy of 95% in distinguishing between tropical systems, extratropical cyclones, and disturbances. SyCLoPS produces a better TC detection skill compared to the previous algorithms, highlighted by an approximately 6% reduction in the false alarm rate compared to the previous TE algorithm. The vertical cross section composite of the four types of high-impact LPS we detect each shows distinct structural characteristics. Finally, we demonstrate that SyCLoPS is valuable for investigating various aspects of LPSs in climate data, such as the evolution of a single LPS track, patterns of LPS frequencies, and precipitation or wind influence associated with a particular LPS class.

54 ENVIRONMENTAL SCIENCES↗

Northern Hemisphere Winter Air Temperature Patterns and Their Associated Atmospheric and Ocean Conditions

The Northern Hemisphere (NH) has experienced winter Arctic warming and continental cooling in recent decades, but the dominant patterns in winter surface air temperature (SAT) are not well understood. Here, a self-organizing map (SOM) analysis is performed to identify the leading patterns in winter daily SAT fields from 1979 to 2018, and their associated atmospheric and ocean conditions are also examined. Three distinct winter SAT patterns with two phases of nearly opposite signs and a time scale of 7–12 days are found: one pattern exhibits concurrent SAT anomalies of the same sign over North America (NA) and northern Eurasia, while the other two patterns show SAT anomalies of opposite signs between, respectively, NA and the Bering Sea, and the Kara Sea and East Asia (EA). Winter SAT variations may arise from changes in the SOM frequencies. Specifically, the observed increasing trends of winter cold extremes over NA, central Eurasia, and EA during 1998–2013 can be understood as a result of the increasing occurrences of some specific SAT patterns. These SOMs are closely related to poleward advection of midlatitude warm air and equatorward movements of polar cold airmass. These meridional displacements of cold and warm airmasses cause concurrent anomalies over different regions not only in SAT but also in water vapor and surface downward longwave radiation. Anomalous sea surface temperatures in the tropical Pacific, midlatitude North Pacific, and North Atlantic and anomalous Arctic sea ice concentrations also concur to support and maintain the anomalous atmospheric circulation that causes the SAT anomalies.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

Integrating very-high-resolution UAS data and airborne imaging spectroscopy to map the fractional composition of Arctic plant functional types in Western Alaska

Widespread changes in vegetation cover and composition are driving strong impacts on Arctic ecosystem functioning and global climate feedbacks. An accurate characterization of tundra vegetation composition is required to understand how the Arctic will respond to future climate change. However, quantifying tundra vegetation composition over large areas is challenging as commonly-used satellite observations are too coarse, spatially and spectrally, to differentiate low-lying tundra vegetation types. Recent airborne and spaceborne imaging spectroscopy platforms provide better data to characterize vegetation composition. Yet, our ability to characterize vegetation composition with imaging spectroscopy remains largely unexplored in the Arctic, particularly due to a lack of ground observations needed to train and test classification models. To address this problem, we collected very-high-resolution (VHR, ~5 cm) unoccupied aerial system (UAS) imagery at three low-Arctic tundra sites located on the Seward Peninsula, western Alaska. In this paper, we examine the feasibility of integrating imagery from the UAS and the hyperspectral Airborne Visible/Infrared Imaging Spectrometer, Next Generation (AVIRIS-NG) airborne instrument to map the fractional composition of 12 key Arctic plant functional types (PFTs). To this end, we first mapped the 12 PFTs from our VHR UAS imagery using random forest classification. We then used these UAS-derived PFT maps as ground truth to develop partial least squares regression (PLSR) models to predict the fractional cover (FCover) of each PFT from AVIRIS-NG imagery. Further, we evaluated the performance of our PLSR models using reserved UAS samples, as well as by mapping PFT FCover and dominant PFT for large tundra landscapes. Our results show that 1) Arctic PFTs can be effectively mapped using VHR UAS imagery, with overall accuracy between 86% and 92%, 2) when the UAS mapped PFTs were used to inform PLSR scaling models, the FCover of the 12 PFTs could be effectively estimated from AVIRIS-NG imagery with a mean absolute error (MAE) <0.13, and 3) our PLSR models outperformed traditional, fully constrained least-squares (FCLS) linear mixture analysis and produced high-quality, spatially contiguous PFT FCover and PFT maps that captured vegetation spatial patterns with similar accuracy to those developed from UAS imagery. The developed PLSR models have the potential to be broadly applied for quantifying vegetation composition with AVIRIS-NG images to help monitor tundra vegetation dynamics and improve process-based modeling of tundra ecosystems.

54 ENVIRONMENTAL SCIENCES↗