Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Insight into the deformation features and capacity loss mechanisms of lithium-ion pouch cells under spherical indentation conditions

Mechanical deformation under extreme conditions is one of the important reasons for the failure of lithium-ion batteries in automotive application. However, the deformation features and component failure of lithium-ion cells to external loading has never been a design consideration. Here, in this study, we conduct spherical indentation tests on a dozen of lithium-ion cells with different capacities under different control mode conditions to investigate their deformation features and capacity loss mechanisms. The experimental results show that, under mechanical deformation conditions, internal faults of cells occur in stages, and energy accumulation and sudden release are two key processes of cell's mechanical failure. The cells' state of charge is the main factor affecting their thermal runaway behaviors. In addition, a finite element model is developed to simulate the deformation features and the failure mechanism of key components of lithium-ion pouch cells; the 3D x-ray computed tomography is employed to demonstrate its internal configuration. With this model, the force-strain response, the deformation features as well as the size of the failure area of lithium-ion cells under spherical indentation conditions are accurately predicted. In 3D x-ray computed tomography images, unique mud cracks in cooper current collector are observed, and the influence mechanisms of the isolated fragments on the cell capacities are revealed. These results may provide useful information for the mechanical structure design of the components of lithium-ion pouch cells.

25 ENERGY STORAGE↗

Empowering Machine Learning Forecasting of Labquake Using Event‐Based Features and Clustering Characteristics

Abstract Following recent advances of machine learning (ML), we present a novel approach to extract spatiotemporal seismo‐mechanical features from Acoustic Emission (AE) catalogs to empower ML‐based forecasting. The AE data were recorded during laboratory stick‐slip experiments on granite samples cut by rough faults. Based on the features computed for a past time window, a random forest (RF) classifier is used to forecast the occurrence of a large magnitude event ( M AE > 3.5) in the next time window. Event‐based features allow us to associate informative time‐space characteristics to each feature and nearest‐neighbor clustering analysis enables us to separate background and clustered seismicity and train individual models. The results show that the separation of AEs enhances the forecasting accuracy from 73.2% for the entire catalog up to 82.1% and 89.0% if background and clustered events are used separately. The presented new approach may be upscaled for applications to forecast tectonic earthquakes.

Karimpouli, Sadegh↗

Feature engineering descriptors, transforms, and machine learning for grain boundaries and variable-sized atom clusters

Abstract Obtaining microscopic structure-property relationships for grain boundaries is challenging due to their complex atomic structures. Recent efforts use machine learning to derive these relationships, but the way the atomic grain boundary structure is represented can have a significant impact on the predictions. Key steps for property prediction common to grain boundaries and other variable-sized atom clustered structures include: (1) describing the atomic structure as a feature matrix, (2) transforming the variable-sized feature matrix to a fixed length common to all structures, and (3) applying a machine learning algorithm to predict properties from the transformed matrices. We examine how these steps and different combinations of engineered features impact the accuracy of grain boundary energy predictions using a database of over 7000 grain boundaries. Additionally, we assess how different engineered features support interpretability, offering insights into the physics of the structure-property relationships.

36 MATERIALS SCIENCE↗

Explainable machine learning for incipient anomaly detection in compact molten salt heat exchanger with overlapping feature distributions

High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.

Prantikos, Konstantinos [Argonne National Laborato↗

Deep optical imaging of star-forming blue early-type galaxies: Color map structures and faint features indicative of recent mergers

Blue early-type galaxies with galaxy-scale ongoing star formation are interesting targets in order to understand the stellar mass buildup in elliptical and S0 galaxies in the local Universe. We study the star-forming population of blue early-type galaxies to understand the origin of star formation in these otherwise red and dead stellar systems. The legacy survey imaging data taken with the dark energy camera in the g, r, and z bands for 55 star-forming blue early-type galaxies were examined, and g – r color maps were created. We identified low surface brightness features near 37 galaxies, faint-level interaction signatures near 15 galaxies, and structures indicative of recent merger activity in the optical color maps of all 55 galaxies. These features are not visible in the shallow Sloan Digital Sky Survey imaging data in which these galaxies were originally identified. Low surface brightness features found around galaxies could be remnants of recent merger events. The star-forming population of blue early-type galaxies could be post-merger systems that are expected to be the pathway for the formation of elliptical galaxies. We hypothesize that the star-forming population of blue early-type galaxies is a stage in the evolution of early-type galaxies. The merger features will eventually disappear, fuel for star formation will cease, and the galaxy will move to the passive population of normal early-type galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

TopFusion: Using Topological Feature Space for Fusion and Imputation in Multi-Modal Data

We present a novel multi-modal data fusion technique using topological features. The method, TopFusion, leverages the flexibility of topological data analysis tools (namely persistent homology and persistence images) to map multi-modal datasets into a common feature space by forming a new multi-channel persistence image. Each channel in the image is representative of a view of the data from a modality-dependent filtration. We demonstrate that the topological perspective we take allows for more effective data reconstruction, i.e. imputation. In particular, by performing imputation in topological feature space we are able to outperform the same imputation techniques applied to raw data or alternatively derived features. We show that TopFusion representations can be used as input to downstream deep learning-based computer vision models and doing so achieves comparable performance to other fusion methods for classification on two multi-modal datasets.

Myers, Audun D.↗

Versatile feature learning with graph convolutions and graph structures

Graphs represent real world relationships, and graph embedding projects nodes in a graph to a latent space that can help simplify downstream tasks. Recent development of graph convolutions in deep learning significantly improves the performance of many learning tasks on graphs. Unfortunately, prior embedding methods either do not embed graphs with node features, or fail to produce high-quality embeddings for downstream learning tasks that result in large performance gap in comparison to direct learning on graphs.We present a versatile and effective embedding method, Conv2Vec, for embedding graphs with or without node features. It is based on graph convolutions with objective functions motivated by concepts and structures from classical graph algorithms. Conv2Vec produce high-quality embedding for both plain graphs and graphs with node features for downstream tasks.We evaluate the embeddings generated by Conv2Vec with a transductive node classification task. With the generated embeddings and very simple machine learning approaches, we are able to achieve accuracies similar to those achieved by direct learning with graph convolutions. Interestingly, if we strip the node features from the graph and thus learning an embedding has to rely entirely on the graph topology, node classification with our embedding significantly outperforms direct learning with various graph convolutions. This suggests that structures from classical graph algorithms may play an important role in learning on graphs.

Cong, Guojing↗

Dual Context: Leveraging Structured Application Context for Code Generation and Runtime Feature Activation via Chat Interfaces

Integrating artificial intelligence (AI) capabilities into software applications typically involves two common paths. For developers, AI assists in generating and documenting source code and other related software engineering efforts. For users, AI assists them through question-and-answer exchanges via chatbots. Both approaches have their value, but neither effectively leverages the modularity of component-based architectures that modern web application frameworks offer. We implement a proof of concept within a centralized suite of applications used for the Atmospheric Radiation Measurement (ARM) Data Center Operational Tools, where we introduce a third integration path through the ARM Context Engine (ACE). ACE is a context driven system that uses structured contextual specifications to enable Large Language Models (LLMs) to render interactive and feature-rich user interface (UI) components directly within chat responses, alongside or in place of conventional text outputs. These specifications serve two important purposes across what we call code context and UI context. Code context provides AI-assisted development tools with structured application knowledge beyond raw code, including component relationships, architectural patterns and schematic information, enabling the generation of consistent, well-structured code. UI context defines the rules for enabling and rendering component features at runtime based on the user's natural language input, allowing end users to activate capabilities such as data export, filtering, and pagination within chat responses, without requiring code changes or redeployment. We demonstrate, through a comparative evaluation against general-purpose AI chatbots, that context-driven component rendering provides interactive capabilities that text-based responses cannot replicate, including deterministic component behavior, application-consistent design language, and on-demand feature activation. A development effort comparison further shows that features that traditionally require multi-step development cycles can be activated with a single naturallanguage request. In this ongoing work, we present ACE as an emerging approach to AI integration that positions modular, well-documented software architecture as the foundation for AI-ready applications. ACE treats context as a shared resource across both development and user-facing AI, bringing cohesion to conventionally disconnected efforts, bridging developer tooling and end-user capabilities within a single framework.

Tadimeti, Vijay [ORNL]↗

Machine Learning Using a Simple Feature for Detecting Multiple Types of Events From PMU Data

This paper describes simple and efficient machine learning (ML) methods for efficiently detecting multiple types of power system events captured by PMUs scarcely placed in a large power grid. It uses a single feature from each PMU based on a rectangle area enclosing the event in a given data window. This single feature is sufficient to enable commonly used ML models to detect different types of events quickly and accurately. The feature is used by five ML models on four different data-window sizes. The results indicated a tradeoff between the execution speed and detection accuracy in variety of data-window size choices. Here, the proposed method is insensitive to most data quality issues typical for data from field PMUs, and thus it does not require major data cleansing efforts prior to feature extraction.

Big data↗

Neural Architecture and Feature Search for Predicting the Ridership of Public Transportation Routes

Accurately predicting the ridership of public-transit routes provides substantial benefits to both transit agencies, who can dispatch additional vehicles proactively before the vehicles that serve a route become crowded, and to passengers, who can avoid crowded vehicles based on publicly available predictions. The spread of the coronavirus disease has further elevated the importance of ridership prediction as crowded vehicles now present not only an inconvenience but also a public-health risk. At the same time, accurately predicting ridership has become more challenging due to evolving ridership patterns, which may make all data except for the most recent records stale. One promising approach for improving prediction accuracy is to fine-tune the hyper-parameters of machine-learning models for each transit route based on the characteristics of the particular route, such as the number of records. However, manually designing a machine-learning model for each route is a labor-intensive process, which may require experts to spend a significant amount of their valuable time. To help experts with designing machine-learning models, we propose a neural-architecture and feature search approach, which optimizes the architecture and features of a deep neural network for predicting the ridership of a public-transit route. Our approach is based on a randomized local hyper-parameter search, which minimizes both prediction error as well as the complexity of the model. We evaluate our approach on real-world ridership data provided by the public transit agency of Chattanooga, TN, and we demonstrate that training neural networks whose architectures and features are optimized for each route provides significantly better performance than training neural networks whose architectures and features are generic.

Ayman, Afiya↗

Invariant Features for Accurate Predictions of Quantum Chemical UV-vis Spectra of Organic Molecules

Including invariance of global properties of a phys-ical system as an intrinsic feature in graph neural networks (GNNs) enhances the model's robustness and generalizability and reduces the amount of training data required to obtain a desired accuracy for predictions of these properties. Existing open source GNN libraries construct invariant features only for specific GNN architectures. This precludes the generalization of invariant features to arbitrary message passing neural network (MPNN) layers which, in turn, precludes the use of these libraries for new, user-specified predictive tasks. To address this limitation, we implement invariant MPNNs into the flexible and scalable HydraGNN architecture. HydraGNN enables a seamless switch between various MPNNs in a unified layer sequence and allows for a fair comparison between the predictive performance of different MPNNs. We trained this enhanced HydraGNN archi-tecture on the ultraviolet-visible (UV-vis) spectrum of GDB-9 molecules, a feature that describes the molecule's electronic exci-tation modes, computed with time-dependent density functional tight binding (TD-DFTB) and available open source through the GDB-9-Ex dataset. We assess the robustness (i.e., accuracy and generalizability) of the predictions obtained using different invariant MPNNs with respect to different values of the full width at half maximum (FWHM) for the Gaussian smearing of the theoretical peaks. Our numerical results show that incorporating invariance in the HydraGNN architecture significantly enhances both accuracy and generalizability in predicting UV-vis spectra of organic molecules.

Baker, Justin↗

In-Situ Detection and Prediction of WAAM Cross Feature Geometry

Abstract Wire arc additive manufacturing (WAAM) is increasingly used by manufacturers due to its relatively low cost and high deposition rate compared to other metal AM methods, but the parts produced by WAAM can be subject to localized variations in part quality. One such variation is the cross-feature defect, whereby a localized part height increase occurs due to the crossing of deposition toolpaths. Mitigation of this defect is typically achieved using manual path planning strategies, but closed-loop control is underutilized. Since the nature of this defect and of the WAAM process is such that the previous layer’s geometry influences that of the subsequent layer’s, the cross-feature defect geometry changes throughout the deposition. Therefore, any closed-loop control strategy will need to incorporate the dynamic trait of this defect. The present work seeks to implement an in-situ process modeling approach where a regression model can be continuously updated to predict the defect geometry of the subsequent deposition layer based on the historical process data. Several multi-layer cross-feature geometries are deposited and current, voltage, and optical camera data is taken for each layer. The resulting cross-feature geometries are characterized using 3D scanning and the performance and accuracy of the in-situ modeling approach is evaluated.

Thien, Austen↗

Foundations of automatic feature extraction at LHC–point clouds and graphs

Abstract Deep learning algorithms will play a key role in the upcoming runs of the Large Hadron Collider (LHC), helping bolster various fronts ranging from fast and accurate detector simulations to physics analysis probing possible deviations from the Standard Model. The game-changing feature of these new algorithms is the ability to extract relevant information from high-dimensional input spaces, often regarded as “replacing the expert” in designing physics-intuitive variables. While this may seem true at first glance, it is far from reality. Existing research shows that physics-inspired feature extractors have many advantages beyond improving the qualitative understanding of the extracted features. In this review, we systematically explore automatic feature extraction from a phenomenological viewpoint and the motivation for physics-inspired architectures. We also discuss how prior knowledge from physics results in the naturalness of the point cloud representation and discuss graph-based applications to LHC phenomenology.

Bhardwaj, Akanksha↗

Lowering and Runtime Support for Fortran’s Multi-Image Parallel Features using LLVM Flang, PRIF, and Caffeine

This paper provides an overview of the multi-image parallel features in Fortran 2023 and their implementation in the LLVM flang compiler and the Caffeine parallel runtime library. The features of interest support a Single-Program, Multiple-Data (SPMD) programming model based on executing multiple “images”, each of which is a program instance. The features also support a Partitioned Global Address Space (PGAS) in the form of “coarray” distributed data structures. The paper discusses the lowering of multi-image features to the Parallel Runtime Interface for Fortran (PRIF) and the implementation of PRIF in the Caffeine parallel runtime library. This paper also provides an early view into the design of a new multi-image dialect of the LLVM Multi-Level Intermediate Representation (MLIR). We describe validation and testing of the resulting software stack, and demonstrate that performance compares favorably to another open-source compiler and runtime library: GNU Compiler Collection (GCC) gfortran and OpenCoarrays, respectively.

Bonachea, Dan↗

Glass Refraction Distortion Object Detection via Abstract Features

Glass reflection and refraction lead to missing and distorted object feature data, affecting the accuracy of object detection. In order to solve the above problems, this paper proposed a glass refraction distortion object detection via abstract features. The number of parameters of the algorithm is reduced by introducing skip connections and expansion modules with different expansion rates. The abstract feature information of the object is extracted by binary cross-entropy loss. Meanwhile, the abstract feature distance between the object domain and source domain is reduced by a loss function, which improves the accuracy of object detection under glass interference. To verify the effectiveness of the algorithm in this paper, the GRI dataset is produced and made public on GitHub. The algorithm of this paper is compared with the current state-of-the-art Deep Face, VGG Face, TBE-CNN, DA-GAN, PEN-3D, LMZMPM, and the average detection accuracy of our algorithm is 92.57% at the highest, and the number of parameters is only 5.13 M.

Cai, Lei↗

Data Agnostic Feature-Target Analysis & Ranking Machine Learning Pipeline (DAFTAR-ML) v0.1.0

DAFTAR-ML is a specialized machine-learning pipeline that identifies relevant features based on their relationship to a target variable. Many ML pipelines focus solely on prediction, and feature ranking is often absent or lacks robust statistical methods. DAFTAR-ML performs its tasks with this outcome in mind. Model training is robust, using nested cross-validation and hyperparameter tuning. Instead of relying on native feature-importance scores, it employs SHAP (SHapley Additive exPlanations) to quantify feature importance. The pipeline also produces comprehensive results, including publication-quality visualizations.

Melie, Tina [Lawrence Berkeley National Laboratory↗

Climatology of Severe Local Storm Environments and Synoptic-Scale Features over North America in ERA5 Reanalysis and CAM6 Simulation

Severe local storm (SLS) activity is known to occur within specific thermodynamic and kinematic environments. These environments are commonly associated with key synoptic-scale features—including southerly Great Plains low-level jets, drylines, elevated mixed layers, and extratropical cyclones—that link the large-scale climate to SLS environments. This work analyzes spatiotemporal distributions of both extreme values of SLS environmental parameters and synoptic-scale features in the ERA5 reanalysis and in the Community Atmosphere Model, version 6 (CAM6), historical simulation during 1980–2014 over North America. Compared to radiosondes, ERA5 successfully reproduces SLS environments, with strong spatiotemporal correlations and low biases, especially over the Great Plains. Both ERA5 and CAM6 reproduce the climatology of SLS environments over the central United States as well as its strong seasonal and diurnal cycles. ERA5 and CAM6 also reproduce the climatological occurrence of the synoptic-scale features, with the distribution pattern similar to that of SLS environments. Compared to ERA5, CAM6 exhibits a high bias in convective available potential energy over the eastern United States primarily due to a high bias in surface moisture and, to a lesser extent, storm-relative helicity due to enhanced low-level winds. Additionally, composite analysis indicates consistent synoptic anomaly patterns favorable for significant SLS environments over much of the eastern half of the United States in both ERA5 and CAM6, though the pattern differs for the southeastern United States. Overall, our results indicate that both ERA5 and CAM6 are capable of reproducing SLS environments as well as the synoptic-scale features and transient events that generate them.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Inferring the Focal Depths of Small Earthquakes in Southern California Using Physics-Based Waveform Features

Determining the depths of small crustal earthquakes is challenging in many regions of the world, because most seismic networks are too sparse to resolve trade-offs between depth and origin time with conventional arrival-time methods. Precise and accurate depth estimation is important, because it can help seismologists discriminate between earthquakes and explosions, which is relevant to monitoring nuclear test ban treaties and producing earthquake catalogs that are uncontaminated by mining blasts. Here, we examine the depth sensitivity of several physics-based waveform features for ~8000 earthquakes in southern California that have well-resolved depths from arrival-time inversion. We focus on small earthquakes (2 < M L < 4) recorded at local distances (<150 km), for which depth estimation is especially challenging. We find that differential magnitudes (M w /M L –M c ) are positively correlated with focal depth, implying that coda wave excitation decreases with focal depth. We analyze a simple proxy for relative frequency content, Φ≡log 10 (M 0 )+3log 10 (f c ), and find that source spectra are preferentially enriched in high frequencies, or “blue-shifted,” as focal depth increases. Here, we also find that two spectral amplitude ratios Rg 0.5–2 Hz/Sg 0.5–8 Hz and Pg/Sg at 3–8 Hz decrease as focal depth increases. Using multilinear regression with these features as predictor variables, we develop models that can explain 11%–59% of the variance in depths within 10 subregions and 25% of the depth variance across southern California as a whole. We suggest that incorporating these features into a machine learning workflow could help resolve focal depths in regions that are poorly instrumented and lack large databases of well-located events. Some of the waveform features we evaluate in this study have previously been used as source discriminants, and our results imply that their effectiveness in discrimination is partially because explosions generally occur at shallower depths than earthquakes.

58 GEOSCIENCES↗