Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

The use of multidate multichannel radiance data in urban feature analysis

Two images were obtained from thematic mappers on Landsats 4 and 5 over the Washington, DC area during November 1982 and March 1984. Selected training areas containing different types of urban land use were examined,one area consisting entirely of forest. Mean digital radiance values for each bandpass in each image were examined, and variances, standard deviations, and covariances between bandpasses were calculated. It has been found that two bandpasses caused forested areas to stand out from other land use types, especially for the November 1982 image. In order to evaluate quantitatively the possible utility of the principal components analysis in selected feature extraction, the eigenvectors were evaluated for principal axes rotations which rendered each selected land use type most separable from all other land use types. The evaluated eigenvectors were plotted as a function of land use type, whose order was decided by considering anticipated shadow component and by examining the relative loadings indicative of vegetation for each of the principal components for the different features considered. The analysis was performed for each seven-band image separately and for the two combined images. It was found that by combining the two images, more dramatic land use type separation could be obtained.

Duggin, M. J.↗

Testing of Haar-Like Feature in Region of Interest Detection for Automated Target Recognition (ATR) System

The objectives of this project were to develop a ROI (Region of Interest) detector using Haar-like feature similar to the face detection in Intel's OpenCV library, implement it in Matlab code, and test the performance of the new ROI detector against the existing ROI detector that uses Optimal Trade-off Maximum Average Correlation Height filter (OTMACH). The ROI detector included 3 parts: 1, Automated Haar-like feature selection in finding a small set of the most relevant Haar-like features for detecting ROIs that contained a target. 2, Having the small set of Haar-like features from the last step, a neural network needed to be trained to recognize ROIs with targets by taking the Haar-like features as inputs. 3, using the trained neural network from the last step, a filtering method needed to be developed to process the neural network responses into a small set of regions of interests. This needed to be coded in Matlab. All the 3 parts needed to be coded in Matlab. The parameters in the detector needed to be trained by machine learning and tested with specific datasets. Since OpenCV library and Haar-like feature were not available in Matlab, the Haar-like feature calculation needed to be implemented in Matlab. The codes for Adaptive Boosting and max/min filters in Matlab could to be found from the Internet but needed to be integrated to serve the purpose of this project. The performance of the new detector was tested by comparing the accuracy and the speed of the new detector against the existing OTMACH detector. The speed was referred as the average speed to find the regions of interests in an image. The accuracy was measured by the number of false positives (false alarms) at the same detection rate between the two detectors.

neural network↗

A minicomputer based software system for the selection of optimal subsets of Thematic Mapper channels

A software system has been developed and implemented on a minicomputer for feature selection based on two inter-dependent methods. The first is an enhancement of the traditional approach based on optimizing interclass average separabilities. The second is based on a Monte Carlo simulation of multispectral data and machine classification with subsequent estimation of classification accuracy as a function of channel subset. The two methods are mutually supportive - the first allows rapid screening whereas the second is based on the more solid theoretical foundation of maximizing classification accuracy.

Card, D. H.↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗

Automated Data Accountability for Missions in Mars Rover Data

As the Mars Curiosity Rover transmits data to the JPL Ground Data System (GDS), it frequently observes data loss and corruption, requiring re-transmits from the rover and Ground Data System Analysts (GDSA) to monitor the downlink process. As new missions are launched, the GDSA team redistributes analysts to these new missions, causing shortages in previous missions. The GDSA team can significantly benefit from the automation and optimization of the downlink process of telemetry data. In fact, there is a need for a better understanding of why the data is corrupted, so that the GDSA team can best determine the root cause of the issues in the GDS. This paper presents machine learning and deep learning based approaches to automate and optimize the detection of data loss. We first created a pipeline to automatically accumulate data from the telemetry databases (MAROS, Telemetry Data Storage, and GDS Elastic Search Database) in the downlink process. With our newly created datasets, we perform feature selection to supplement the GDSA understanding of the downlink process and provide supplemental analysis on the importance of different features. We implement various machine learning and deep learning based models, including support vector machines, ensemble methods, and deep neural networks and evaluate their accuracies in identifying whether a downlink process is complete or incomplete. We utilize fast hyperparameter optimization methods that allow our models to quickly be re-trained, allowing them to quickly be tuned and optimized on daily incoming data in real time. This hyperparameter optimization also allows our methods to be quickly integrated into other JPL missions. Our results show that our best-performing machine learning and deep learning based models outperform the existing GDSA detection software by 6 accuracy points and can aid analysts by providing insights into the data accountability problem. Since these various machine learning and deep learning approaches vary significantly in interpretability, we provide a discussion on the tradeoffs between their performance and trustworthiness in helping detect issues in data transmission.

Divsalar, Dariush↗

Global 3D Data Visualization and Analysis Platform With Advanced Machine Learning Capabilities in Support of Lunar Exploration

Introduction: The science goals for NASA’s Artemis program include: a) Understanding the character and origin of lunar polar volatiles, b) Conducting experimental science in the lunar environment and c) Investigating and mitigating exploration risks [1]. The permanently shadowed regions (PSRs) on the Lunar south pole are expected to host large quantities of water-ice and volatiles that are important for sustainable Lunar exploration [2]. There are several missions such as onboard Korea Pathfinder Lunar Orbiter (KPLO: Korean name Danuri) with onboard ShadowCam camera [3], Astrobotic Peregrine Mission One [4], and other efforts underway to obtain high resolution topographic, minerals, volatiles and other information on the moon. We envision a need in immediate future for platforms to integrate these data sets, provide rendering and visualization capabilities in the context of a 3D Lunar globe for easier information access and analysis. NASA's Celestial Mapping System (CMS) [5] is developed to address the need for 3D tools for planetary science investigations, mission planning, in-situ operations, in a 3D-first design constructed around a unified view of a planetary globe. At present CMS provides many critical functionalities that include: 1) equipment planning and optimized placement on Lunar surface 2) line-of-sight (LOS) analysis 3) powerful measurement tools based on 3D terrain with realistic 3D models to represent rovers, astronauts and equipment 4) visualization of derived mapping products (e.g. resource maps), and 5) a data engine for hosting new observations that are not available in other contemporary lunar data tools [5, 7]. Planetary Data Ingestion: CMS can consume and analyze data from locally hosted and external third party sources. It is compatible with Open Geospatial Consortium (OGC) data and file standards and currently integrates datasets from the Astrogeology Science Center of USGS. This includes global and local data acquired from NASA (LRO, Clementine, Lunar Orbiter) and JAXA (SELENE/Kaguya), with capability of integrating more datasets. In addition, users can specify other WMS-hosted data endpoints, which CMS can then query and stream data from automatically. To set-up an automated process for ingestion and accurate rendering, visualization and analysis of external 3rd party planetary datasets within CMS, we initiated the process of ingesting unique dataset of super-enhanced images of the permanently shadowed regions (PSRs) at the lunar poles which were produced by the Hyper-effective nOise Removal U-net Software (HORUS) tool [8]. This tool was developed to enhance the extremely low-light images of the interior of PSRs and provide the ability to see within these regions at and discern surface features (i.e. boulders and craters) down to 3 meters in size. We focused on the Nobile region on the Lunar south pole, selected site for VIPER mission and stitched several images to create a high-resolution map within one of the PSR of Nobile crater. Figure 1 shows the dark PSR zone form the original NAC layer of LRO as the base layer (left image) and the illuminated areas within that crater (center) which was created by ingesting and merging several of HORUS generated images. At present we employ a semi-automated process to ensure spatial accuracy and merger of several overlapping zones. However, we are in the process of completely automating this process by employing AI based techniques that would rank, sort, and stack the images based on their information density. The georectification of the images would employ selected features. Analysis on Ingested Planetary Datasets: Once an external planetary data-set is successfully ingested, georectified and merged seamlessly as a data-layer; CMS’ numerous analysis tools can be used on this data. A Line of Sight (LOS) tool has been developed for CMS which analyzes terrain profiles and obstructions to determine visibility for remote observers [5,6]. Figure 1 (right image) shows the viewshed analysis on the same PSR in the Nobile region. The yellow pin shows the observer location outside the PSR. The yellow area shows the visible part of PSR. The obstructed area with no visibility for the observer is shown in red. The Measurements tool allows the user to take area and distance measurements of features on the terrain using various shapes. Measurement type can be specified in a number of ways: Line, Path, Polygon, Circle, Ellipse, Square, Rectangle or Freehand. Once the shape is specified, elevation information can then be extracted along each of these shapes. Figure 2 (left) shows the measurements performed on a crater n illuminated PSR in Nobile region. The equipment placement tool allows the user to place a 3D equipment model at a desired location and analyze its coverage area. The equipment placement tool is coupled with LOS to determine the coverage. Figure 2 (right) shows an equipment placed on the Lunar terrain and it’s coverage area. The red rays are blocked sight lines and the green rays are non-obstructed sight lines with the cyan lines showing the point of intersection with the terrain. More details are provided in the video demonstrations in Reference 5. Overcoming Polar Distortions: 3D geospatial applications exhibit significant distortions in polar imagery due to several reasons: 1) distortions in the source imagery, 2) incompatible tessellation algorithms at the poles, and 3) map projections. We are leveraging new tessellation algorithms and reprojecting data using projections that are better suited for Lunar poles. The goal is to seamlessly switch to polar projections while maintaining 3D view and navigation.

Maps↗

Global 3D Data Visualization and Analysis Platform with Advanced Machine Learning Capabilities in Support of Lunar Exploration

Introduction: The science goals for NASA’s Artemis program include: a) Understanding the character and origin of lunar polar volatiles, b) Conducting experimental science in the lunar environment and c) Investigating and mitigating exploration risks. The permanently shadowed regions (PSRs) on the Lunar south pole are expected to host large quantities of water-ice and volatiles that are important for sustainable Lunar exploration. There are several missions such as onboard Korea Pathfinder Lunar Orbiter (KPLO: Korean name Danuri) with onboard ShadowCam camera, Astrobotic Peregrine Mission One [4], and other efforts underway to obtain high resolution topographic, minerals, volatiles and other information on the moon. We envision a need in immediate future for platforms to integrate these data sets, provide rendering and visualization capabilities in the context of a 3D Lunar globe for easier information access and analysis. NASA's Celestial Mapping System (CMS) is developed to address the need for 3D tools for planetary science investigations, mission planning, in-situ operations, in a 3D-first design constructed around a unified view of a planetary globe. At present CMS provides many critical functionalities that include: 1) equipment planning and optimized placement on Lunar surface 2) line-of-sight (LOS) analysis 3) powerful measurement tools based on 3D terrain with realistic 3D models to represent rovers, astronauts and equipment 4) visualization of derived mapping products (e.g. resource maps), and 5) a data engine for hosting new observations that are not available in other contemporary lunar data tools. Planetary Data Ingestion: CMS can consume and analyze data from locally hosted and external third party sources. It is compatible with Open Geospatial Consortium (OGC) data and file standards and currently integrates datasets from the Astrogeology Science Center of USGS. This includes global and local data acquired from NASA (LRO, Clementine, Lunar Orbiter) and JAXA (SELENE/Kaguya), with capability of integrating more datasets. In addition, users can specify other WMS-hosted data endpoints, which CMS can then query and stream data from automatically. To set-up an automated process for ingestion and accurate rendering, visualization and analysis of external 3rd party planetary datasets within CMS, we initiated the process of ingesting unique dataset of super-enhanced images of the permanently shadowed regions (PSRs) at the lunar poles which were produced by the Hyper-effective nOise Removal U-net Software (HORUS) tool. This tool was developed to enhance the extremely low-light images of the interior of PSRs and provide the ability to see within these regions at and discern surface features (i.e. boulders and craters) down to 3 meters in size. We focused on the Nobile region on the Lunar south pole, selected site for VIPER mission and stitched several images to create a high-resolution map within one of the PSR of Nobile crater. Figure 1 shows the dark PSR zone form the original NAC layer of LRO as the base layer (left image) and the illuminated areas within that crater (center) which was created by ingesting and merging several of HORUS generated images. At present we employ a semi-automated process to ensure spatial accuracy and merger of several overlapping zones. However, we are in the process of completely automating this process by employing AI based techniques that would rank, sort, and stack the images based on their information density. The georectification of the images would employ selected features. Analysis on Ingested Planetary Datasets: Once an external planetary data-set is successfully ingested, georectified and merged seamlessly as a data-layer; CMS’ numerous analysis tools can be used on this data. A Line of Sight (LOS) tool has been developed for CMS which analyzes terrain profiles and obstructions to determine visibility for remote observers. Figure 1 (right image) shows the viewshed analysis on the same PSR in the Nobile region. The yellow pin shows the observer location outside the PSR. The yellow area shows the visible part of PSR. The obstructed area with no visibility for the observer is shown in red. The Measurements tool allows the user to take area and distance measurements of features on the terrain using various shapes. Measurement type can be specified in a number of ways: Line, Path, Polygon, Circle, Ellipse, Square, Rectangle or Freehand. Once the shape is specified, elevation information can then be extracted along each of these shapes. Figure 2 (left) shows the measurements performed on a crater n illuminated PSR in Nobile region. The equipment placement tool allows the user to place a 3D equipment model at a desired location and analyze its coverage area. The equipment placement tool is coupled with LOS to determine the coverage. Figure 2 (right) shows an equipment placed on the Lunar terrain and it’s coverage area. The red rays are blocked sight lines and the green rays are non-obstructed sight lines with the cyan lines showing the point of intersection with the terrain. More details are provided in the video demonstrations in Reference 5. Overcoming Polar Distortions: 3D geospatial applications exhibit significant distortions in polar imagery due to several reasons: 1) distortions in the source imagery, 2) incompatible tessellation algorithms at the poles, and 3) map projections. We are leveraging new tessellation algorithms and reprojecting data using projections that are better suited for Lunar poles. The goal is to seamlessly switch to polar projections while maintaining 3D view and navigation.

Maps↗

cTULIP: application of a human-based RNA-seq primary tumor classification tool for cross-species primary tumor classification in canine

The domestic dog, Canis familiaris, is quickly gaining traction as an advantageous model for use in the study of cancer, one of the leading causes of death worldwide. Naturally occurring canine cancers share clinical, histological, and molecular characteristics with the corresponding human diseases. In this study, we take a deep-learning approach to test how similar the gene expression profile of canine glioma and bladder cancer (BLCA) tumors are to the corresponding human tumors. We likewise develop a tool for identifying misclassified or outlier samples in large canine oncological datasets, analogous to that which was developed for human datasets. We test a number of machine learning algorithms and found that a convolutional neural network outperformed logistic regression and random forest approaches. We use a recently developed RNA-seq-based convolutional neural network, TULIP, to test the robustness of a human-data-trained primary tumor classification tool on cross-species primary tumor prediction. Our study ultimately highlights the molecular similarities between canine and human BLCA and glioma tumors, showing that protein-coding one-to-one homologs shared between humans and canines, are sufficient to distinguish between BLCA and gliomas. The results of this study indicate that using protein-coding one-to-one homologs as the features in the input layer of TULIP performs good primary tumor prediction in both humans and canines. Furthermore, our analysis shows that our selected features also contain the majority of features with known clinical relevance in BLCA and gliomas. Our success in using a human-data-trained model for cross-species primary tumor prediction also sheds light on the conservation of oncological pathways in humans and canines, further underscoring the importance of the canine model system in the study of human disease.

60 APPLIED LIFE SCIENCES↗

Analysis of pseudocolor transformations of ERTS-1 images of Southern California area

The author has identified the following significant results. Representative faults and lineaments, natural features on the Mojave Desert, and cultural features of the southern California area were studied on ERTS-1 images. The relative appearances of the features were compared on a band 4 and 5 subtraction image, its pseudocolor transformation, and pseudocolor images of bands 4, 5, and 7. Selected features were also evaluated in a test given students at the University of California, Los Angeles. Observations and the test revealed no significant improvement in the ability to detect and locate faults and lineaments on the pseudocolor transformations. With the exception of dry lake surfaces, no enhancement of the features studied was observed on the bands 4 and 5 subtraction images. Geologic and geographic features characterized by minor tonal differences on relatively flat surfaces were enhanced on some of the pseudocolor images.

Merifield, P. M.↗

A new μ -high energy resolution fluorescence detection microprobe imaging spectrometer at the Stanford Synchrotron Radiation Lightsource beamline 6-2

In this report we describe a new synchrotron X-ray Fluorescence (XRF) imaging instrument with an integrated High Energy Fluorescence Detection X-ray Absorption Spectroscopy (HERFD-XAS) spectrometer at the Stanford Synchrotron Radiation Lightsource at beamline 6-2. The X-ray beam size on the sample can be defined via a range of pinhole apertures or focusing optics. XRF imaging is performed using a continuous rapid scan system with sample stages covering a travel range of 250 × 200 mm 2 , allowing for multiple samples and/or large samples to be mounted. The HERFD spectrometer is a Johann-type with seven spherically bent 100 mm diameter crystals arranged on intersecting Rowland circles of 1 m diameter with a total solid angle of about 0.44% of 4π sr. A wide range of emission lines can be studied with the available Bragg angle range of ~64.5°–82.6°. With this instrument, elements in a sample can be rapidly mapped via XRF and then selected features targeted for HERFD-XAS analysis. Furthermore, utilizing the higher spectral resolution of HERFD for XRF imaging provides better separation of interfering emission lines, and it can be used to select a much narrower emission bandwidth, resulting in increased image contrast for imaging specific element species, i.e., sparse excitation energy XAS imaging. This combination of features and characteristics provides a highly adaptable and valuable tool in the study of a wide range of materials.

47 OTHER INSTRUMENTATION↗

Evaluation of the potential format and content of a cockpit display of traffic information

The types and formats of information most suitable to be displayed in a cockpit display of traffic information (CDTI) are investigated. Twenty three airline pilots and 13 instrumentated general aviation pilots were asked to select from sets of symbols of various complexities incorporating various levels of information that would contain all information necessary for monitoring the traffic situation, detecting errors, maintaining separation and merging. Display features selected by a significant number of pilots were then evaluated for their capabilities in helping pilots to assess the lateral or vertical separation between their own and another aircraft in a dynamic simulation. It is found that while some of the features initially chosen by the pilots, such as flightpath predictors, aided the pilots in perceiving the traffic situation correctly, others, such as ground speed and climb/descend arrows and relative altitude encoding of symbols for other aircraft, did not contribute to improved performance speed or accuracy.

Hart, S. G.↗

Prediction of Self-Diffusion in Binary Fluid Mixtures Using Artificial Neural Networks

Artificial neural networks (ANNs) were developed to accurately predict the self-diffusion constants for individual components in binary fluid mixtures. The ANNs were tested on an experimental database of 4328 self-diffusion constants from 131 mixtures containing 75 unique compounds. The presence of strong hydrogen bonding molecules may lead to clustering or dimerization resulting in non-linear diffusive behavior. To address this, self- and binary association energies were calculated for each molecule and mixture to provide information on intermolecular interaction strength and were used as input features to the ANN. An accurate, generalized ANN model was developed with an overall average absolute deviation of 4.1%. Forward input feature selection reveals the importance of critical properties and self-association energies along with other fluid properties. Additional ANNs were developed with subsets of the full input feature set to further investigate the impact of various properties on model performance. The results from two specific mixtures are discussed in additional detail: one providing an example of strong hydrogen bonding and the other an example of extreme pressure changes, with the ANN models predicting self-diffusion well in both cases.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

MER-DIMES : a planetary landing application of computer vision

During the Mars Exploration Rovers (MER) landings, the Descent Image Motion Estimation System (DIMES) was used for horizontal velocity estimation. The DIMES algorithm combines measurements from a descent camera, a radar altimeter and an inertial measurement unit. To deal with large changes in scale and orientation between descent images, the algorithm uses altitude and attitude measurements to rectify image data to level ground plane. Feature selection and tracking is employed in the rectified data to compute the horizontal motion between images. Differences of motion estimates are then compared to inertial measurements to verify correct feature tracking. DIMES combines sensor data from multiple sources in a novel way to create a low-cost, robust and computationally efficient velocity estimation solution, and DIMES is the first use of computer vision to control a spacecraft during planetary landing. In this paper, the detailed implementation of the DIMES algorithm and the results from the two landings on Mars are presented.

landing systems↗

Combining artificial intelligence and physics-based modeling to directly assess atomic site stabilities: from sub-nanometer clusters to extended surfaces

The performance of functional materials is dictated by chemical and structural properties of individual atomic sites. In catalysts, for instance, the thermodynamic stability of constituting atomic sites is a key descriptor from which more complex properties, such as molecular adsorption energies and reaction rates, can be derived. In this study, we present a widely applicable machine learning (ML) approach to instantaneously compute the stability of individual atomic sites in structurally and electronically complex nano-materials. Conventionally, we determine such site stabilities using computationally intensive first-principles calculations. With our approach, we predict the stability of atomic sites in sub-nanometer metal clusters of 3–55 atoms with mean absolute errors in the range of 0.11–0.14 eV. To extract physical insights from the ML model, we introduce a genetic algorithm (GA) for feature selection. This algorithm distills the key structural and chemical properties governing the stability of atomic sites in size-selected nanoparticles, allowing for physical interpretability of the models and revealing structure–property relationships. The results of the GA are generally model and materials specific. In the limit of large nanoparticles, the GA identifies features consistent with physics-based models for metal–metal interactions. By combining the ML model with the physics-based model, we predict atomic site stabilities in real time for structures ranging from sub-nanometer metal clusters (3–55 atom) to larger nanoparticles (147 to 309 atoms) to extended surfaces using a physically interpretable framework. Finally, we present a proof of principle showcasing how our approach can determine stable and active nanocatalysts across a generic materials space of structure and composition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage Systems

Monitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue.We address such issues by building an end-to-end machine learn- ing (ML) workflow that can generate I/O traces for large-scale HPC applications. We leverage ML based feature selection and gener- ative models for I/O trace generation. The generative models are trained on I/O traces collected by the darshan I/O characterization tool over a period of one year. We present a two-step generation process consisting of two deep-learning models, called the feature generator and the trace generator. The combination of two-step generative models provides robustness by reducing the bias of the model and accounting for the stochastic nature of the I/O traces across different runs of an application. We evaluate the performance of the generative models and show that the two-step model can generate time-series I/O traces with less than 20% root mean square error.

Paul, Arnab↗

Documentation of procedures for textural/spatial pattern recognition techniques

A C-130 aircraft was flown over the Sam Houston National Forest on March 21, 1973 at 10,000 feet altitude to collect multispectral scanner (MSS) data. Existing textural and spatial automatic processing techniques were used to classify the MSS imagery into specified timber categories. Several classification experiments were performed on this data using features selected from the spectral bands and a textural transform band. The results indicate that (1) spatial post-processing a classified image can cut the classification error to 1/2 or 1/3 of its initial value, (2) spatial post-processing the classified image using combined spectral and textural features produces a resulting image with less error than post-processing a classified image using only spectral features and (3) classification without spatial post processing using the combined spectral textural features tends to produce about the same error rate as a classification without spatial post processing using only spectral features.

Haralick, R. M.↗

Support Vector Machines for Classification of Direct Energy Deposition Standoff Distance for Improved Process Control

A critical factor in the implementation of direct energy deposition is the ability to maintain the standoff distance between the nozzle and the build surface, as this influences powder capture efficiency and overall part quality. Due to process-related variations, layer height may vary, causing unintended variation in standoff distance and poor build quality. While prior work has utilized contact probing to qualify standoff distance during processing, in situ methods for qualification of standoff distance are of major interest. The present work seeks to understand efficacy of image-based methods for classifying standoff distance variation in real-time using support vector machines (SVMs). It was hypothesized that the size of the melt pool and the amount of spatter will have significant correlations with deviations in the standoff distance; thus, SVMs were used on a dataset that is comprised of morphological features of melt pool size and image entropy. The SVM model was used to classify melt pool images into categories according to standoff distance variation from nominal. K-folds cross validation was used to find the optimal hyperparameters for the SVM model. To understand the impact of the selected features on the classification performance and inference speed, multiple models were trained with differing numbers of included features. Results for classification score, inference time, and image preprocessing/feature extraction from these data are reported. The present results show that the SVM model was able to predict the standoff distance classification with an accuracy of 97 percent and a speed of 0.122 s per image, making it a viable solution for real-time control of standoff distance.

Klesmith, Zoe↗

Image-analysis techniques for determination of morphology and kinematics in arctic sea ice

SAR data have been used to study sea ice with respect to its motion and formation/deformation. With the prospect of the Alaska SAR Facility development in the near future, there is a great need for robust and efficient sea-ice analysis techniques. This paper presents a sea-ice motion analysis technique that can be used for: (1) local motion analysis of a selected ice patch, and (2) a global ice motion over the entire image area. In order to meet the operational speed requirement (over fifty images per day) a sea-ice motion analysis technique has been developed which requires very little human interaction. The proposed technique uses a subset of easily distinguishable features to predict global motion characteristics. The developed technique is applied to two pairs of SEASAT SAR images, one pair with a minor motion of 'ice pack' and another with a larger and discontinuous motion of 'fast ice'. The new approach enables the development of a set of computer-aided tools for feature selection and registration and the implementation of an optimal search strategy for automatic template matching via a motion prediction model.

Lee, Meemong↗