Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Knowledge Graph Entity Linking using Graph Embeddings

Details the use of a custom embedding model on knowledge graphs to aid in downstream natural language processing (NLP) models for Derivative Classification Assist. Motivations, algorithms, and results were discussed.

Mahesh, Aarav [Sandia National Laboratories (SNL-N↗

A Hierarchical Feature-Based Methodology to Perform Cervical Cancer Classification

Prevention of cervical cancer could be performed using Pap smear image analysis. This test screens pre-neoplastic changes in the cervical epithelial cells; accurate screening can reduce deaths caused by the disease. Pap smear test analysis is exhaustive and repetitive work performed visually by a cytopathologist. This article proposes a workload-reducing algorithm for cervical cancer detection based on analysis of cell nuclei features within Pap smear images. We investigate eight traditional machine learning methods to perform a hierarchical classification. We propose a hierarchical classification methodology for computer-aided screening of cell lesions, which can recommend fields of view from the microscopy image based on the nuclei detection of cervical cells. We evaluate the performance of several algorithms against the Herlev and CRIC databases, using a varying number of classes during image classification. Results indicate that the hierarchical classification performed best when using Random Forest as the key classifier, particularly when compared with decision trees, k-NN, and the Ridge methods.

60 APPLIED LIFE SCIENCES↗

Effect of transient liquid flow on retention characteristics of screen acquisition systems

A design analysis, is developed based on experimental data, to predict the effects of transient flow and pressure surges (caused either by valve or pump operation, or by boiling of liquids in warm lines) on the retention performance of screen acquisition systems. A survey of screen liquid acquisition system applications was performed to determine appropriate system environment and classification. A screen model was developed which assumed that the screen device was a uniformly distributed composite orthotropic structure, and which accounted for liquid inflow/outflow, gas ingestion quality, screen stress, and liquid spill. A series of 177 tests using 13 specimens (5 screen meshes, 4 screen device construction/backup methods, and 2 orientations) with three test fluids (isopropyl alcohol, Freon 114, and LH2) provided data which verified important features of the screen model and resulted in a design tool which could accurately predict the transient startup performance acquisition devices.

Cady, E. C.↗

Digital techniques for processing Landsat imagery

An overview of the basic techniques used to process Landsat images with a digital computer, and the VICAR image processing software developed at JPL and available to users through the NASA sponsored COSMIC computer program distribution center is presented. Examples of subjective processing performed to improve the information display for the human observer, such as contrast enhancement, pseudocolor display and band rationing, and of quantitative processing using mathematical models, such as classification based on multispectral signatures of different areas within a given scene and geometric transformation of imagery into standard mapping projections are given. Examples are illustrated by Landsat scenes of the Andes mountains and Altyn-Tagh fault zone in China before and after contrast enhancement and classification of land use in Portland, Oregon. The VICAR image processing software system which consists of a language translator that simplifies execution of image processing programs and provides a general purpose format so that imagery from a variety of sources can be processed by the same basic set of general applications programs is described.

Green, W. B.↗

An investigation of vegetation and other Earth resource/feature parameters using LANDSAT and other remote sensing data. 1: LANDSAT. 2: Remote sensing of volcanic emissions

A fanning technique based on a simplistic physical model provided a classification algorithm for mixture landscapes. Results of applications to LANDSAT inventory of 1.5 million acres of forest land in Northern Maine are presented. Signatures for potential deer year habitat in New Hampshire were developed. Volcanic activity was monitored in Nicaragua, El Salvador, and Guatemala along with the Mt. St. Helens eruption. Emphasis in the monitoring was placed on the remote sensing of SO2 concentrations in the plumes of the volcanoes.

Birnie, R. W.↗

High Performance Input/Output for Parallel Computer Systems

The goal of our project is to study the I/O characteristics of parallel applications used in Earth Science data processing systems such as Regional Data Centers (RDCs) or EOSDIS. Our approach is to study the runtime behavior of typical programs and the effect of key parameters of the I/O subsystem both under simulation and with direct experimentation on parallel systems. Our three year activity has focused on two items: developing a test bed that facilitates experimentation with parallel I/O, and studying representative programs from the Earth science data processing application domain. The Parallel Virtual File System (PVFS) has been developed for use on a number of platforms including the Tiger Parallel Architecture Workbench (TPAW) simulator, The Intel Paragon, a cluster of DEC Alpha workstations, and the Beowulf system (at CESDIS). PVFS provides considerable flexibility in configuring I/O in a UNIX- like environment. Access to key performance parameters facilitates experimentation. We have studied several key applications fiom levels 1,2 and 3 of the typical RDC processing scenario including instrument calibration and navigation, image classification, and numerical modeling codes. We have also considered large-scale scientific database codes used to organize image data.

Ligon, W. B.↗

The Urban Environmental Monitoring/100 Cities Project: Legacy of the First Phase and Next Steps

The Urban Environmental Monitoring (UEM) project, now known as the 100 Cities Project, at Arizona State University (ASU) is a baseline effort to collect and analyze remotely sensed data for 100 urban centers worldwide. Our overarching goal is to use remote sensing technology to better understand the consequences of rapid urbanization through advanced biophysical measurements, classification methods, and modeling, which can then be used to inform public policy and planning. Urbanization represents one of the most significant alterations that humankind has made to the surface of the earth. In the early 20th century, there were less than 20 cities in the world with populations exceeding 1 million; today, there are more than 400. The consequences of urbanization include the transformation of land surfaces from undisturbed natural environments to land that supports different forms of human activity, including agriculture, residential, commercial, industrial, and infrastructure such as roads and other types of transportation. Each of these land transformations has impacted, to varying degrees, the local climatology, hydrology, geology, and biota that predate human settlement. It is essential that we document, to the best of our ability, the nature of land transformations and the consequences to the existing environment. The focus in the UEM project since its inception has been on rapid urbanization. Rapid urbanization is occurring in hundreds of cities worldwide as population increases and people migrate from rural communities to urban centers in search of employment and a better quality of life. The unintended consequences of rapid urbanization have the potential to cause serious harm to the environment, to human life, and to the resulting built environment because rapid development constrains and rushes decision making. Such rapid decision making can result in poor planning, ineffective policies, and decisions that harm the environment and the quality of human life. Slower, more thought-out, decision making could result in more favorable outcomes. The harm to the environment includes poor air quality, soil erosion, polluted rivers and aquifers, and loss of wildlife habitat. Human life is then threatened because of increased potential for disease spreading, human conflict, environmental hazards, and diminished quality of life. The built environment is potentially threatened when cities are built in areas that can be impacted by events such as hurricanes, tsunamis, earthquakes, fires, and landslides. Our goals include assessing the threat of such events on cities and the people living there.

Stefanov, William L.↗

Estimating Soil Moisture in a Boreal Old Jack Pine Forest

Polarimetric L- and P-band AIRSAR data, corresponding model simulations, and classification algorithms have shown that in a boreal old jack pine (OJP) stand, the principal scattering mechanism responsible for radar backscatter is the double-bounce mechanism between the tree trunks and the ground.

Soil↗

Predictive models enhance feedstock quality of corn stover via air classification

Feedstock heterogeneity is a fundamental obstacle to cost-competitive biobased products. Agricultural products like corn stover have anatomical components that vary in their chemical composition, mechanical properties, structure, and response to chemical and biological treatments. A technique that can enrich streams in select anatomical fractions would allow a tailored deconstruction approach to increase overall process efficiency. Air classification can be leveraged for such refining, however, fundamental characterization and understanding of the particle properties that underly the physics of air classification are only modestly documented. Here, we determine fundamental particle properties including mass-to-area ratio, drag coefficient, and partition velocity that describe how anatomical tissues of corn stover behave during air classification. In this work, mass-to-area ratios of anatomical tissues vary by nearly two orders of magnitude from 2.3 mg/mm 2 for cob to 0.04 mg/mm 2 for leaf. Drag coefficients of longer, fibrous materials (i.e., rind, husk, and sheath) are shown to correlate with particle area (p-value < 0.001) whereas granular tissues (i.e., cob, pith, and leaf) correlate better with mass-to-area ratio (p-values < 0.001). When compared to experimental observations, a simulated two-stage air classification and size reduction scenario predicts the overall partitioning of anatomical tissues within 15% for pith, husk, rind, and cob tissues. The model predicts an air-classified fraction preferentially enriched in cob (purity = 20%), rind (purity = 74%), and pith (purity = 4.5%) with a mass yield of 47%. Empirical relations for these properties can be used to predict the partitioning of corn stover during air classification based on anatomical type and size.

09 BIOMASS FUELS↗

Detecting Arsenic Contamination Using Satellite Imagery and Machine Learning

Arsenic, a potent carcinogen and neurotoxin, affects over 200 million people globally. Current detection methods are laborious, expensive, and unscalable, being difficult to implement in developing regions and during crises such as COVID-19. This study attempts to determine if a relationship exists between soil’s hyperspectral data and arsenic concentration using NASA’s Hyperion satellite. It is the first arsenic study to use satellite-based hyperspectral data and apply a classification approach. Four regression machine learning models are tested to determine this correlation in soil with bare land cover. Raw data are converted to reflectance, problematic atmospheric influences are removed, characteristic wavelengths are selected, and four noise reduction algorithms are tested. The combination of data augmentation, Genetic Algorithm, Second Derivative Transformation, and Random Forest regression (R 2 =0.840 and normalized root mean squared error (re-scaled to [0,1]) = 0.122) shows strong correlation, performing better than past models despite using noisier satellite data (versus lab-processed samples). Three binary classification machine learning models are then applied to identify high-risk shrub-covered regions in ten U.S. states, achieving strong accuracy (=0.693) and F1-score (=0.728). Overall, these results suggest that such a methodology is practical and can provide a sustainable alternative to arsenic contamination detection.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Evaluation of artificial neural network performance for classification of potato plants infected with potato virus Y using spectral data on multiple varieties and genotypes

Potato virus Y (Potyviridae, PVY) is a plant virus that poses a significant threat to potato producers on a global basis. The pathogen has disrupted seed potato supplies and negatively impacted yield and quality of commercial potato crops. The potato industry currently manages PVY infection levels via insecticide applications, regional seed certification programs that rely on field scouting to visually assess individual plants for infection status, and destructive and costly tissue sampling coupled with laboratory assays. Despite these efforts, PVY continues to confound potato industry stakeholders resulting in economic harm. Remote sensing and machine learning provide for the development of new tools to more accurately detect and spatially quantify PVY-infected plants versus the current state of the art. However, there is a need to understand how the occurrence of many different potato varieties impact the dynamics of developing models to detect potato plants impacted with PVY and their potential effectiveness. This study evaluates classification modelling outcomes using spectral datasets collected in different temporal and spatial environments (greenhouse and a production field) on multiple potato varieties consisting of labelled instances of plants infected with PVY and those not infected with the virus. A modelling framework was developed to support iterative modelling runs using artificial neural network (ANN) architectures configured as binary classifiers to develop sample populations to support statistical analysis on model performance using specific spectral subsets. When using spectral data to detect PVY-infected plants, ANN models achieved the highest mean accuracy of 0.894 on a single variety. Conversely, the same ANN model architecture only achieved a mean accuracy of 0.575 on a spectral data set representing 29 potato breeding lines. Additionally, statistical analysis indicates spectral regions including the red edge, near infrared and shortwave infrared contain more important spectral features for the ANN classifier introduced in this research.

60 APPLIED LIFE SCIENCES↗

Extracting scene feature vectors through modeling, volume 3

The remote estimation of the leaf area index of winter wheat at Finney County, Kansas was studied. The procedure developed consists of three activities: (1) field measurements; (2) model simulations; and (3) response classifications. The first activity is designed to identify model input parameters and develop a model evaluation data set. A stochastic plant canopy reflectance model is employed to simulate reflectance in the LANDSAT bands as a function of leaf area index for two phenological stages. An atmospheric model is used to translate these surface reflectances into simulated satellite radiance. A divergence classifier determines the relative similarity between model derived spectral responses and those of areas with unknown leaf area index. The unknown areas are assigned the index associated with the closest model response. This research demonstrated that the SRVC canopy reflectance model is appropriate for wheat scenes and that broad categories of leaf area index can be inferred from the procedure developed.

Berry, J. K.↗

Evaluating Limits of Machine Learning-Assisted Raman Spectroscopy in Classification of Biological Samples

Machine learning (ML)-assisted Raman spectroscopy has become a powerful analytical tool for the classification and identification of analytes; however, technical challenges impacting its detection accuracy have not been thoroughly investigated. This study explores experimental factors affecting classification performance. Among the evaluated ML models, ML algorithms show minimal impact on classification accuracy. Instead, experimental factors, including spectral similarity between tested samples and data quality, dominate detection performance. Increases in spectral noise and spectral similarity significantly reduce classification accuracy. In well-controlled samples with low experimental noise, ML-assisted Raman spectroscopy can discriminate lipid mixtures with a composition difference of 1.85 mol %. To assess the effect of biological heterogeneity, we analyzed single-cell Raman spectra from Saccharomyces cerevisiae strains carrying single, double, or triple gene mutations. Intrinsic cell-to-cell variability introduced substantial spectral differences, severely reducing the accuracy of multiclass classification of these genetically similar strains at the single-cell level. Averaging Raman spectra across multiple cells improved classification accuracy by reducing this spectral variability. We also assess the effectiveness of transfer learning across different Raman spectrometers, specifically by applying an ML model trained on one instrument to another Raman spectrometer. Transfer learning can be improved with proper instrument calibration, highlighting the importance of instrument standardization. Overall, our results demonstrate that data quality and spectral similarity are the primary bottlenecks in ML-assisted Raman spectroscopy. Careful attention to sample preparation, data acquisition, measurement conditions, and instrument calibration is critical to achieving robust and reliable classification performance.

Fungi↗

Jet classification using high-level features from anatomy of top jets

Recent advancements in deep learning models have significantly enhanced jet classification performance by analyzing low-level features (LLFs). However, this approach often leads to less interpretable models, emphasizing the need to understand the decision-making process and to identify the high-level features (HLFs) crucial for explaining jet classification. To address this, we consider the top jet tagging problems and introduce an analysis model (AM) that analyzes selected HLFs designed to capture important features of top jets. Our AM mainly consists of the following three modules: a relation network analyzing two-point energy correlations, mathematical morphology and Minkowski functionals for generalizing jet constituent multiplicities, and a recursive neural network analyzing subjet constituent multiplicity to enhance sensitivity to subjet color charges. We demonstrate that our AM achieves performance comparable to the Particle Transformer (ParT) while requiring fewer computational resources in a comparison of top jet tagging using jets simulated at the hadronic calorimeter angular resolution scale. Furthermore, as a more constrained architecture than ParT, the AM exhibits smaller training uncertainties because of the bias-variance tradeoff. We also compare the information content of AM and ParT by decorrelating the features already learned by AM. Lastly, we briefly comment on the results of AM with finer angular resolution inputs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗