Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Interpretation of Pennsylvania agricultural land use from ERTS-1 data

The author has identified the following significant results. To study the complex agricultural patterns in Pennsylvania, a portion of an ERTS scene was selected for detailed analysis. Various photographic products were made and were found to be only of limited value. This necessitated the digital processing of the ERTS data. Using an unsupervised classification procedure, it was possible to delineate the following categories: (1) forest land with a northern aspect, (2) forest land with a southern aspect, (3) valley trees, (4) wheat, (5) corn, (6) alfalfa, grass, pasture, (7) disturbed land, (8) builtup land, (9) strip mines, and (10) water. These land use categories were delineated at a scale of approximately 1:20,000 on the line printer output. Land use delineations were also made using the General Electric IMAGE 100 interactive analysis system.

Mcmurtry, G. J.↗

Tone Noise and Nearfield Pressure Produced by Jet-Cavity Interaction

Cavity flow resonance can cause numerous problems in aerospace applications. While our long-term goal is to understand cavity flows well enough to devise effective cavity resonance suppression techniques, this paper describes a fundamental study of resonant tones produced by jet-cavity interaction at subsonic and supersonic speeds. Our specific jet-cavity configuration can also be used as a test bed for evaluating active and passive flow resonance control concepts. Two significant findings emerge from this study. 1) Originally, we expected that tones produced by jet-cavity interaction would resemble cavity tones or jet tones or would involve some simple combinations of each. The experimental data do not support these expectations: instead, the jet cavity interaction produce a unique set of tones. We propose simple yet and physically insightful correlations for these tones. Although the pressure patterns on the cavity floor display very complex variations with the Mach number for a length/depth = 8 cavity, the tones correspond to the acoustic modes of the cavity-independent of flow. For a length/ depth = 3 cavity, however, a surprise emerges: the pressure patterns on the cavity floor are not so complex but the tones depend significantly on the flow. Additionally, we examine the role of external feedback unique to jet-cavity interaction. 2) Previous research led us to expect that traditional classifications (open, transitional, or closed) for cavities in an infinite flight stream would be insensitive to small changes in Mach number and would depend primarily on cavity length/depth ratios. Use of the novel high resolution photoluminescent pressure sensitive paint shows that the classifications are actually quite sensitive to jet Mach number for a length/depth = 8 cavity. However, these classifications provide no guidance whatsoever for tone amplitude or frequency. Detailed experimental data and insights presented here will assist researchers who are performing numerical simulations of jet-cavity flows as a first step toward devising resonance suppression methods.

Raman, Ganesh↗

A Deterministic Self-Organizing Map Approach and its Application on Satellite Data based Cloud Type Classification

A self-organizing map (SOM) is a type of competitive artificial neural network, which projects the high dimensional input space of the training samples into a low dimensional space with the topology relations preserved. This makes SOMs supportive of organizing and visualizing complex data sets and have been pervasively used among numerous disciplines with different applications. Notwithstanding its wide applications, the self-organizing map is perplexed by its inherent randomness, which produces dissimilar SOM patterns even when being trained on identical training samples with the same parameters every time, and thus causes usability concerns for other domain practitioners and precludes more potential users from exploring SOM based applications in a broader spectrum. Motivated by this practical concern, we propose a deterministic approach as a supplement to the standard self-organizing map. In accordance with the theoretical design, the experimental results with satellite cloud data demonstrate the effective and efficient organization as well as simplification capabilities of the proposed approach.

Initialization method↗

Automated Classification of Vehicle Movements at Signalized Intersections Using Vehicle Trajectories

Accurate vehicle movement classification through signalized intersections is of paramount importance to the analysis of intersection performance and the optimization of traffic control strategies. Conventional techniques for tracking vehicle turning movements depend on infrastructure-based strategies like human counts, loop detectors, and video analytics, all of which are costly, prone to errors, and spatially constrained. High-frequency trajectory data can be utilized to determine vehicle movement patterns in a scalable and infrastructure-independent method due to the adoption of connected vehicles (CVs). In recent years, several studies have utilized connected vehicle data to generate performance measures. Most of the trajectory-based performance measures approaches, however, require map matching-i.e., extracting geospatial references from maps to identify the movements that individual vehicles make at a signalized intersection. These approaches are often time-consuming and hinder scalability since geographic features need to be provided for an analysis to be conducted. Map matching methods are prone to errors as different map versions change these geographic features. This research presents a novel automatic classification pipeline that uses CV trajectory data to classify vehicle movements at signalized crossings, specifically pass-through left-turn and right-turn maneuvers. The process starts by filtering trips that cross a spatial bounding box that has been defined at the target intersection. Approach and departure headings for each trajectory crossing the boundary are computed and are clustered together to identify dominant movements. The proposed algorithm is used to classify the movement of vehicles at 10 intersections in the state of California, and the results indicate that the algorithm can classify movements at these intersections with varying traffic volumes and road network configurations, all in a map-less framework with no need for conflation of vehicle trajectories to a digital base map.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decoding substance use disorder severity from clinical notes using a large language model

Substance use disorder (SUD) poses a major concern due to its detrimental effects on health and society. SUD identification and treatment depend on a variety of factors such as severity, co-determinants (e.g., withdrawal symptoms), and social determinants of health. Existing diagnostic coding systems used by insurance providers, like the International Classification of Diseases (ICD-10), lack granularity for certain diagnoses, but American clinicians will add this granularity (as that found within the Diagnostic and Statistical Manual of Mental Disorders classification or DSM-5) as supplemental unstructured text in clinical notes. Traditional natural language processing (NLP) methods face limitations in accurately parsing such diverse clinical language. Large language models (LLMs) offer promise in overcoming these challenges by adapting to diverse language patterns. This study investigates the application of LLMs for extracting severity-related information for various SUD diagnoses from clinical notes. We propose a workflow employing zero-shot learning of LLMs with carefully crafted prompts and post-processing techniques. Through experimentation with Flan-T5, an open-source LLM, we demonstrate its superior recall compared to the rule-based approach. Focusing on 11 categories of SUD diagnoses, we show the effectiveness of LLMs in extracting severity information, contributing to improved risk assessment and treatment planning for SUD patients.

60 APPLIED LIFE SCIENCES↗

Change in land use in the Phoenix (1:250,000) Quadrangle, Arizona between 1970 and 1973: ERTS as an aid in a nationwide program for mapping general land use

Changes in land use between 1970 and 1973 in the Phoenix (1:250,000 scale) Quadrangle in Arizona have been mapped using only the images from ERTS-1, tending to verify the utility of a standard land use classification system proposed for use with ERTS images. Types of changes detected have been: (1) new residential development of former cropland and rangeland; (2) new cropland from the desert; and (3) new reservoir fill-up. The seasonal changing of vegetation patterns in ERTS has complemented air photos in delimiting the boundaries of some land use types. ERTS images, in combination with other sources of information, can assist in mapping the generalized land use of the fifty states by the standard 1:250,000 quadrangles. Several states are already working cooperatively in this type of mapping.

Place, J. L.↗

Improved assessment of mangrove forests in Sundarbans East Wildlife Sanctuary using WorldView 2 and TanDEM-X high resolution imagery

Recent developments of remote sensing techniques which can capture both the structure and function of the ecosystem provide a more representative view of the landscape. These unique Earth observations were used to help improve traditional forestry surveys by providing species-specific land cover classes for mangrove forests in the Sundarbans East Wildlife Sanctuary. By combining optical data from WorldView2 (WV2; 2 m pixel) and a canopy height model derived using radar data from TanDEM-X (TDX; 12 m pixel), we identified nine mangrove and five non-mangrove classes by following an Iterative Self-Organizing Data Analysis Algorithm. Three dominant mangrove species accounted for nearly 50% of the sanctuary. Heritieria fomes disproportionately covered the largest area at 43%, overturning previous field-based estimates of Excoecaria agallocha dominance. E. agallocha and Sonneratia apetala, covered 3% and 1.47% of the sanctuary, respectively. Four mixed species classes were also identified with clear vegetation zonation patterns that trended toward species homogeneity with increasing distance from shore. The overall land cover accuracy (WV2: 89.33%; WV2-TDX: 89.89%), the Kappa Coefficient (WV2:0.88; WV2-TDX: 0.89) and change statistics between WV2 and WV2-TDX landcover classifications indicate that the WV2 imagery can separate mangrove community types without structural data. The combination of the land cover classifications and the canopy height model indicated that H. fomes were not only the most dominant forest but also, on average, the tallest (12.3 m) among the other eight mangrove types. Our large-scale mapping with high resolution optical and radar platforms can capture subtle changes in mangrove vegetation and canopy structural gradients more accurately and be used to monitor biodiversity changes and Aichi Biodiversity Targets and Indicators, which would contribute to biodiversity policy updating.

Md Mizanur Rahman↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

Heterogeneous Graph Neural Network for identifying hadronically decayed tau leptons at the High Luminosity LHC

Here, we present a new algorithm that identifies reconstructed jets originating from hadronic decays of tau leptons against those from quarks or gluons. No tau lepton reconstruction algorithm is used. Instead, the algorithm represents jets as heterogeneous graphs with tracks and energy clusters as nodes and trains a Graph Neural Network to identify tau jets from other jets. Different attributed graph representations and different GNN architectures are explored. We propose to use differential track and energy cluster information as node features and a heterogeneous sequentially-biased encoding for the inputs to final graph-level classification.

47 OTHER INSTRUMENTATION↗

Behavioral Segmentation and Clustering of Geospatial Trajectories

The rapid growth of global positioning system (GPS) devices has led to a corresponding increase in the size of GPS datasets. While these large GPS datasets contain a wealth of information about the behaviors of the moving objects in them, manual classification and anomaly detection are prohibitively time consuming. We utilize unsupervised machine learning techniques to first identify the behaviors for individual moving objects and then cluster those objects by their behavioral sequences. In this way, trajectories behaving unusually as well as common patterns of behavior are both detectable in large datasets without requiring an a priori definition of "unusual" or "common."

97 MATHEMATICS AND COMPUTING↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Use of robust estimators in parametric classifiers

The parametric approach to density estimation and classifier design is a well studied subject. The parametric approach is desirable because basically it reduces the problem of classifier design to that of estimating a few parameters for each of the pattern classes. The class parameters are usually estimated using maximum-likelihood (ML) estimators. ML estimators are, however, very sensitive to the presence of outliers. Several robust estimators of mean and covariance matrix and their effect on the probability of error in classification are examined. Comments are made about alpha-ranked (alpha-trimmed) estimators.

Safavian, S. Rasoul↗

Requirements for Soldering Fluxes Research using the B-53 Test Board

IPC J-STD-004B standard prescribes general requirements for the classification and testing of soldering flux for high qualify interconnections. This standard defines the classification of soldering materials through specifications of test methods and inspection criteria. The materials include liquid flux, paste flux, solderpaste flux, solder preform flux, and flux-cored solder. This research will use the proposed IPC-53 Surface Insulation Resistance (SIR) test patterns by means of an open comb (2D) and closed comb (3D to simulate a component over the comb pattern). The 2D open comb has uniformity of conductor spacing, sheet resistance, and flux outgassing. The 3D-closed comb simulates the effect of leadless or bottom-terminated components, which have non-uniform sheet resistance and flux outgassing. The response variables will include SIR testing and visual imaging. The objective is to investigate IPC test method improvements for characterizing soldering fluxes when using leadless components with narrow pad-to-pad spacing.

Diamond, Louis↗

A study of the utilization of ERTS-1 data from the Wabash River Basin

The author has identified the following significant results. Preliminary acreage estimates for corn, soybeans, and other cover types made from classification of ERTS-1 data compared well with those made by the U.S. Department of Agriculture. Registration of multiple frames of ERTS-1 CCT data over Lynn County, Texas and DeKalb County, Illinois was achieved to a high degree of accuracy. Spectral/temporal computer pattern recognition analysis was carried out for the first time using satellite data.

Landgrebe, D. A.↗

Separability of agricultural cover types in spectral channels and wavelength regions

This study was a continuation of a more complete evaluation of the spectral channels as well as wavelength regions - visible, near infrared, middle infrared, and thermal infrared - with respect to their estimated probability of correct classification Pc in discriminating agricultural cover types reported previously by Kumar and Silva (1977). Multispectral scanner data in twelve spectral channels in the wavelength range of 0.4-11.7 microns acquired in the middle of July for three flightlines were analyzed by applying automatic pattern recognition techniques. The same analysis was performed for the data acquired in the middle of August 1971, over the same three flightlines, to investigate the effect of time on the results. The effect of deletion of each spectral channel as well as each wavelength region on Pc is given. Values of Pc for all possible combinations of wavelength regions in the subsets of one to twelve spectral channels are also given. The overall values of Pc were found to be greater for the data of the middle of August than the data of the middle of July.

Kumar, R.↗

Characterizing Signatures of Geothermal Exploration Data with Machine Learning Techniques: An Application to the Nevada Play Fairway Analysis

We are introducing machine learning methods to the play fairway analysis to generate geothermal potential maps to support the evaluation of geothermal resource potential and the exploration for undiscovered blind geothermal systems in the Nevada Great Basin region. Our project aims to identify new ways to combine the play fairway data and empirically organize relationships between feature weights and labels in an improved workflow. As a means of doing this, we introduce machine learning methods to evaluate the influence of certain geological and geophysical features/feature sets in predicting geothermal favorability. This report highlights promising approaches based on supervised and unsupervised learning methods. First, we demonstrate a filter method applied to supervised classification modeling. The supervised filter method is based on permutation analysis to evaluate every possible feature combination/drop out scenario and rank feature influence based on the performance variance of supervised classification models. Additionally, we present an unsupervised factor analysis based on principal component analysis coupled with a semi-supervised kmeans clustering algorithm. This analysis allows us to identify the optimal number of groups/clusters for training sites and structural settings to identify feature patterns including correlation, variance, and latent and dominant feature relationships. The results from these methods offer a promising avenue for identifying favorable sources of predictive information to identify the locations of blind geothermal systems and furthering our understanding of complex geothermal feature and label relationships in the Great Basin region and beyond.

15 GEOTHERMAL ENERGY↗

Software and Algorithms for Biomedical Image Data Processing and Visualization

A new software equipped with novel image processing algorithms and graphical-user-interface (GUI) tools has been designed for automated analysis and processing of large amounts of biomedical image data. The software, called PlaqTrak, has been specifically used for analysis of plaque on teeth of patients. New algorithms have been developed and implemented to segment teeth of interest from surrounding gum, and a real-time image-based morphing procedure is used to automatically overlay a grid onto each segmented tooth. Pattern recognition methods are used to classify plaque from surrounding gum and enamel, while ignoring glare effects due to the reflection of camera light and ambient light from enamel regions. The PlaqTrak system integrates these components into a single software suite with an easy-to-use GUI (see Figure 1) that allows users to do an end-to-end run of a patient s record, including tooth segmentation of all teeth, grid morphing of each segmented tooth, and plaque classification of each tooth image. The automated and accurate processing of the captured images to segment each tooth [see Figure 2(a)] and then detect plaque on a tooth-by-tooth basis is a critical component of the PlaqTrak system to do clinical trials and analysis with minimal human intervention. These features offer distinct advantages over other competing systems that analyze groups of teeth or synthetic teeth. PlaqTrak divides each segmented tooth into eight regions using an advanced graphics morphing procedure [see results on a chipped tooth in Figure 2(b)], and a pattern recognition classifier is then used to locate plaque [red regions in Figure 2(d)] and enamel regions. The morphing allows analysis within regions of teeth, thereby facilitating detailed statistical analysis such as the amount of plaque present on the biting surfaces on teeth. This software system is applicable to a host of biomedical applications, such as cell analysis and life detection, or robotic applications, such as product inspection or assembly of parts in space and industry.

Talukder, Ashit↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗