Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Lowering and Runtime Support for Fortran’s Multi-Image Parallel Features using LLVM Flang, PRIF, and Caffeine

This paper provides an overview of the multi-image parallel features in Fortran 2023 and their implementation in the LLVM flang compiler and the Caffeine parallel runtime library. The features of interest support a Single-Program, Multiple-Data (SPMD) programming model based on executing multiple “images”, each of which is a program instance. The features also support a Partitioned Global Address Space (PGAS) in the form of “coarray” distributed data structures. The paper discusses the lowering of multi-image features to the Parallel Runtime Interface for Fortran (PRIF) and the implementation of PRIF in the Caffeine parallel runtime library. This paper also provides an early view into the design of a new multi-image dialect of the LLVM Multi-Level Intermediate Representation (MLIR). We describe validation and testing of the resulting software stack, and demonstrate that performance compares favorably to another open-source compiler and runtime library: GNU Compiler Collection (GCC) gfortran and OpenCoarrays, respectively.

Bonachea, Dan↗

Glass Refraction Distortion Object Detection via Abstract Features

Glass reflection and refraction lead to missing and distorted object feature data, affecting the accuracy of object detection. In order to solve the above problems, this paper proposed a glass refraction distortion object detection via abstract features. The number of parameters of the algorithm is reduced by introducing skip connections and expansion modules with different expansion rates. The abstract feature information of the object is extracted by binary cross-entropy loss. Meanwhile, the abstract feature distance between the object domain and source domain is reduced by a loss function, which improves the accuracy of object detection under glass interference. To verify the effectiveness of the algorithm in this paper, the GRI dataset is produced and made public on GitHub. The algorithm of this paper is compared with the current state-of-the-art Deep Face, VGG Face, TBE-CNN, DA-GAN, PEN-3D, LMZMPM, and the average detection accuracy of our algorithm is 92.57% at the highest, and the number of parameters is only 5.13 M.

Cai, Lei↗

Data Agnostic Feature-Target Analysis & Ranking Machine Learning Pipeline (DAFTAR-ML) v0.1.0

DAFTAR-ML is a specialized machine-learning pipeline that identifies relevant features based on their relationship to a target variable. Many ML pipelines focus solely on prediction, and feature ranking is often absent or lacks robust statistical methods. DAFTAR-ML performs its tasks with this outcome in mind. Model training is robust, using nested cross-validation and hyperparameter tuning. Instead of relying on native feature-importance scores, it employs SHAP (SHapley Additive exPlanations) to quantify feature importance. The pipeline also produces comprehensive results, including publication-quality visualizations.

Melie, Tina [Lawrence Berkeley National Laboratory↗

Climatology of Severe Local Storm Environments and Synoptic-Scale Features over North America in ERA5 Reanalysis and CAM6 Simulation

Severe local storm (SLS) activity is known to occur within specific thermodynamic and kinematic environments. These environments are commonly associated with key synoptic-scale features—including southerly Great Plains low-level jets, drylines, elevated mixed layers, and extratropical cyclones—that link the large-scale climate to SLS environments. This work analyzes spatiotemporal distributions of both extreme values of SLS environmental parameters and synoptic-scale features in the ERA5 reanalysis and in the Community Atmosphere Model, version 6 (CAM6), historical simulation during 1980–2014 over North America. Compared to radiosondes, ERA5 successfully reproduces SLS environments, with strong spatiotemporal correlations and low biases, especially over the Great Plains. Both ERA5 and CAM6 reproduce the climatology of SLS environments over the central United States as well as its strong seasonal and diurnal cycles. ERA5 and CAM6 also reproduce the climatological occurrence of the synoptic-scale features, with the distribution pattern similar to that of SLS environments. Compared to ERA5, CAM6 exhibits a high bias in convective available potential energy over the eastern United States primarily due to a high bias in surface moisture and, to a lesser extent, storm-relative helicity due to enhanced low-level winds. Additionally, composite analysis indicates consistent synoptic anomaly patterns favorable for significant SLS environments over much of the eastern half of the United States in both ERA5 and CAM6, though the pattern differs for the southeastern United States. Overall, our results indicate that both ERA5 and CAM6 are capable of reproducing SLS environments as well as the synoptic-scale features and transient events that generate them.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Inferring the Focal Depths of Small Earthquakes in Southern California Using Physics-Based Waveform Features

Determining the depths of small crustal earthquakes is challenging in many regions of the world, because most seismic networks are too sparse to resolve trade-offs between depth and origin time with conventional arrival-time methods. Precise and accurate depth estimation is important, because it can help seismologists discriminate between earthquakes and explosions, which is relevant to monitoring nuclear test ban treaties and producing earthquake catalogs that are uncontaminated by mining blasts. Here, we examine the depth sensitivity of several physics-based waveform features for ~8000 earthquakes in southern California that have well-resolved depths from arrival-time inversion. We focus on small earthquakes (2 < M L < 4) recorded at local distances (<150 km), for which depth estimation is especially challenging. We find that differential magnitudes (M w /M L –M c ) are positively correlated with focal depth, implying that coda wave excitation decreases with focal depth. We analyze a simple proxy for relative frequency content, Φ≡log 10 (M 0 )+3log 10 (f c ), and find that source spectra are preferentially enriched in high frequencies, or “blue-shifted,” as focal depth increases. Here, we also find that two spectral amplitude ratios Rg 0.5–2 Hz/Sg 0.5–8 Hz and Pg/Sg at 3–8 Hz decrease as focal depth increases. Using multilinear regression with these features as predictor variables, we develop models that can explain 11%–59% of the variance in depths within 10 subregions and 25% of the depth variance across southern California as a whole. We suggest that incorporating these features into a machine learning workflow could help resolve focal depths in regions that are poorly instrumented and lack large databases of well-located events. Some of the waveform features we evaluate in this study have previously been used as source discriminants, and our results imply that their effectiveness in discrimination is partially because explosions generally occur at shallower depths than earthquakes.

58 GEOSCIENCES↗

Adversarial Perturbations Are Not So Weird: Entanglement of Robust and Non-Robust Features in Neural Network Classifiers

Neural networks trained on visual data are well-known to be vulnerable to often imperceptible adversarial perturbations. The reasons for this vulnerability are still being debated in the literature. Recently Ilyas et al. (2019) showed that this vulnerability arises, in part, because neural network classifiers rely on highly predictive but brittle “non-robust” features. In this paper we extend the work of Ilyas et al. by investigating the nature of the input patterns that give rise to these features. In particular, we hypothesize that in a neural network trained in a standard way, non-robust features respond to small, “non-semantic” patterns that are typically entangled with larger, robust patterns, known to be more human-interpretable, as opposed to solely responding to statistical artifacts in a dataset. Thus, adversarial examples can be formed via minimal perturbations to these small, entangled patterns. In addition, we demonstrate a corollary of our hypothesis: robust classifiers are more effective than standard (non-robust) ones as a source for generating transferable adversarial examples in both the untargeted and targeted settings. The results we present in this paper provide new insight into the nature of the non-robust features responsible for adversarial vulnerability of neural network classifiers.

97 MATHEMATICS AND COMPUTING↗

Examination of 3013 Containers Baseline Surface Features

The Surveillance and Monitoring Program at Los Alamos National Laboratory (LANL) was tasked with evaluating the baseline features of 3013 containers. This baseline is to be used as a basis for comparison for 3013 containers that had been packaged with corrosive plutonium materials. The LANL team evaluated an unwelded container, an unused welded container, and a welded container that had held plutonium metal without any corrosive impurities. These three containers had features with depths no larger than 6 µm and had similar depth distributions. The features observed in the baseline containers were all shallower than those seen in containers packaged with corrosive plutonium materials. This work establishes workflows for feature identification and measurement, as well as establishment of baseline data for future comparison with corroded containers.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Taxonomy and Feature set for Server-Side Identification of Proxies

Malicious actors frequently use proxies and VPNs to evade detection and hide their origin. Current challenges to information security include the use of residential proxies to blend in with normal traffic and Man-in-the-Middle phishing proxies that are used to compromise accounts protected with mult-factor authentication. We advance a taxonomy and feature set for the identification of proxied traffic based on the network layer where proxying occurs. We describe how these features apply to common proxy types and how to use these features in the classification of the proxied traffic. Collection of these additional features is feasible using existing network sensors and web servers, while only adding about 30% volume to commonly deployed network sensor logs.

97 MATHEMATICS AND COMPUTING↗

Detection of Critical Surface Features in PTLs and GDLs for Improved Device Performance and Manufacturing Reliability

High points, or features that protrude above the surface of the material, on porous transport layers (PTLs) and gas diffusion layers (GDLs) can be critical features that may affect the manufacturing process and the performance of the device containing the feature. High points on PTLs and GDLs may stress the membrane of a polymer electrolyte membrane (PEM) during lamination and cell operation of a PEM electrolyzer or fuel cell. Additionally, high points on GDLs may impact the reliability of the manufacturing process. Thus, understanding these critical features and developing procedures to detect them are a key part of developing quality control techniques for PTLs and GDLs. This work evaluates the effectiveness of the Keyence VR6200 benchtop-scale structured light optical profilometer for detection of surface protrusions on PTLs and GDLs. Standard testing procedures for detecting and measuring high points were created for use on both material types. These procedures were evaluated using Gage Repeatability and Reproducibility (Gage R&R), where the repeatability, reproducibility, and effectiveness of the system to detect and measure high points were quantified. We have shown with high statistical power that the system is very effective in detection and measurement of high points, with Gage R&R contributions measured to be 2.2% and 3.6% for PTLs and GDLs, respectively.

36 MATERIALS SCIENCE↗

Massive Early-type Galaxies in the HSC-SSP: Flux Fraction of Tidal Features and Merger Rates

Here we present a statistical study on tidal features around massive early-type galaxies (ETGs). Utilizing the imaging data of the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP), we measure the flux fraction of tidal features (f tidal ) in 2649 ETGs with stellar mass M * > 10 11 M ⊙ and redshift 0.05 < z < 0.15 using automated techniques. The Wide layer of HSC-SSP reaches a depth of ~28.5 mag arcsec –2 in the i band. Under this surface brightness limit, we find that about 28% of these galaxies harbor prominent tidal features with f tidal > 1%, among which the number of ETGs decreases exponentially with f tidal , with a logarithmic slope of ~100. Within the stellar mass range we probe, we note that f tidal increases by a factor of 2 from M * ≈ 10 11 to 10 12 M ⊙ . We also perform a pair count to estimate the merger rate of these massive ETGs. Combining the merger rates with f tidal , we estimate that the typical lifetime of tidal features is ~3 Gyr, consistent with previous studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Real-World Driving Features for Identifying Intelligent Driver Model Parameters

Driver behavior models play a significant role in representing different driving styles and the associated relationships with traffic patterns and vehicle energy consumption in simulation studies. The models often serve as a proxy for baseline human driving when assessing energy-saving strategies that alter vehicle velocity. Such models are especially important in connectivity-enabled energy-saving strategy research because they can easily adapt to changing driving conditions like posted speed limits or change in traffic light state. While numerous driver models exist, parametric driver models provide the flexibility required to represent variability in real-world driving through different combinations of model parameters. These model parameters must be informed by a representative set of parameter values for the driver model to adequately represent a real-world driver. It stands to reason that determining the parameter values from real-world driving data would serve the purpose of representing a real-world driver. Although the published literature is replete with techniques and consequences of using real-world driving data to determine parameter values for parametric driver models, none have explored them in the context of using shorter driving features where the parameter values may change over the course of a single trip for the same driver. In this study we consider the “intelligent driver model” (IDM) as our driver behavior model and use real-world driving data from the Transportation Secure Data Center (TSDC) maintained at the National Renewable Energy Laboratory (NREL). The TSDC includes real-world travel data from across the United States, from which NREL has created a wide range of driving routes consisting of road features such as speed limits, stop locations, and turn locations. The real-world driving data are categorized into different driving regimes and extracted into driving features. The driving features are then used to calibrate the parameter values for the IDM. The distribution of the parameters and the relationships among them are reported. The insights obtained from this study enable judicious usage of IDM or similar parametric driver models to represent baseline human driver behavior in simulations.

27 ARPA - Advanced Research Projects Agency-Energy↗

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)↗

Decreasing wind speed extrapolation error via domain-specific feature extraction and selection

Abstract. Model uncertainty is a significant challenge in the wind energy industry and can lead to mischaracterization of millions of dollars' worth of wind resources. Machine learning methods, notably deep artificial neural networks (ANNs), are capable of modeling turbulent and chaotic systems and offer a promising tool to produce high-accuracy wind speed forecasts and extrapolations. This paper uses data collected by profiling Doppler lidars over three field campaigns to investigate the efficacy of using ANNs for wind speed vertical extrapolation in a variety of terrains, and it quantifies the role of domain knowledge in ANN extrapolation accuracy. A series of 11 meteorological parameters (features) are used as ANN inputs, and the resulting output accuracy is compared with that of both standard log-law and power-law extrapolations. It is found that extracted nondimensional inputs, namely turbulence intensity, current wind speed, and previous wind speed, are the features that most reliably improve the ANN's accuracy, providing up to a 65 % and 52 % increase in extrapolation accuracy over log-law and power-law predictions, respectively. The volume of input data is also deemed important for achieving robust results. One test case is analyzed in depth using dimensional and nondimensional features, showing that the feature nondimensionalization drastically improves network accuracy and robustness for sparsely sampled atmospheric cases.

17 WIND ENERGY↗

DMTN-118: Review of Timeseries Features

Rubin Observatory will compute timeseries variability features on lightcurves to aid users in identifying objects of interest, both during realtime Alert Production as well as in the annual Data Releases. The Data Products Definition Document (DPDD; LSE-163) allocates space for pre-computed timeseries features, and a sample set is baselined in LDM-151. However, in the subsequent decade the scientfic community has made a great deal of further progress in this area. This technote reviews the relevant literature, grouping related features where possible; discusses potential concerns and open questions; and proposes a new baseline feature set.

79 ASTRONOMY AND ASTROPHYSICS↗

Real-World Driving Features for Identifying Intelligent Driver Model Parameters: Preprint

Driver behavior models play a significant role in representing different driving styles and the associated relationships with traffic patterns and vehicle energy consumption in simulation studies. The models often serve as a proxy for baseline human driving when assessing energy-saving strategies that alter vehicle velocity. Such models are especially important in connectivity-enabled energy-saving strategy research because they can easily adapt to changing driving conditions like posted speed limits or change in traffic light state. While numerous driver models exist, parametric driver models provide the flexibility required to represent variability in real-world driving through different combinations of model parameters. These model parameters must be informed by a representative set of parameter values for the driver model to adequately represent a real-world driver. It stands to reason that determining the parameter values from real-world driving data would serve the purpose of representing a real-world driver. While the published literature is replete with techniques and consequences of using real-world driving data to determine parameter values for parametric driver models, none have explored them in the context of using shorter driving features where the parameter values may change over the course of a single trip for the same driver. In this study we consider the “Intelligent Driver Model” (IDM) as the driver behavior model to explore and real-world driving data from the Transportation Secure Data Center (TSDC) maintained at the National Renewable Energy Laboratory (NREL). The TSDC includes real-world travel data from across the United States, from which NREL has created a wide range of driving routes consisting of road features such as speed limits, stop locations and turn locations. The real-world driving data are categorized into different driving regimes and extracted into driving features. The driving features are then used to calibrate the parameter values for the IDM. The distribution of the parameters and the relationships among them are reported. The insights obtained from this study enable judicious usage of IDM or similar parametric driver models to represent baseline human driver behavior in simulations.

27 ARPA - Advanced Research Projects Agency-Energy↗

A Single-Feature Machine Learning Method for Detecting Multiple Types of Events from PMU Data

This paper describes simple and efficient machine learning (ML) methods for efficiently detecting multiple types of power system events captured by PMUs scarcely placed in a large power grid. It uses a single feature from each PMU based on a rectangle area enclosing the event in a given data window. This single feature is sufficient to enable commonly used ML models to detect different types of events quickly and accurately. The feature is used by five ML models on four different data-window sizes. The results indicated a tradeoff between the execution speed and detection accuracy in variety of data-window size choices. The proposed method is insensitive to most data quality issues typical for data from field PMUs, and thus it does not require major data cleansing efforts prior to feature extraction.

Dokic, Tatjana↗

Exploiting Multi-Domain Features for Detection of Unclassified Electromagnetic Signals

Deep Learning based classification techniques have shown excellent performance in static environments, where the training and testing samples are drawn from the same distribution. However, real world scenarios often present samples that do not belong to the known set of classes chosen during training. This is quite common for electromagnetic signals, where it is impractical to assume that all possible waveforms are known a-priori, specially in scenarios like warfare. To address this problem, we propose a deep learning based adversarial model where the generator learns to generate waveform features that can deceive the discriminator model as true samples. We introduce domain knowledge of wireless signals by decomposing the signal into a lower dimensional unique feature set, which is used for classifying known versus unknown signals. We further introduce multiple domain representations of the signal to extract features and combine them together to accurately classify new waveforms as an unknown class. Our results show that combined features from multiple domains outperform any single domain representation, especially at low SNR regimes with fewer number of samples to classify.

99 - GENERAL AND MISCELLANEOUS↗

Exploiting Multi-Domain Features for Detection of Unclassified Electromagnetic Signals (Presentation)

Deep Learning based classification techniques have shown excellent performance in static environments, where the training and testing samples are drawn from the same distribution. However, real world scenarios often present samples that do not belong to the known set of classes chosen during training. This is quite common for electromagnetic signals, where it is impractical to assume that all possible waveforms are known a-priori, specially in scenarios like warfare. To address this problem, we propose a deep learning based adversarial model where the generator learns to generate waveform features that can deceive the discriminator model as true samples. We introduce domain knowledge of wireless signals by decomposing the signal into a lower dimensional unique feature set, which is used for classifying known versus unknown signals. We further introduce multiple domain representations of the signal to extract features and combine them together to accurately classify new waveforms as an unknown class. Our results show that combined features from multiple domains outperform any single domain representation, especially at low SNR regimes with fewer number of samples to classify.

99 - GENERAL AND MISCELLANEOUS↗