Engineering PapersSearch

SEARCH · Engineering Papers

Results for “feature importance analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Design of diverse, functional mitochondrial targeting sequences across eukaryotic organisms using variational autoencoder

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

59 BASIC BIOLOGICAL SCIENCES

Scalable Computation of Topological Abstractions for Scalar Data

Topological data analysis has become an important tool for large scale scalar data analysis and visualization, efficiently extracting the inherent structure and features of interest of the data. However, with growing dataset sizes and complexity, it is increasingly becoming infeasible to compute topological abstractions of interest in serial and on single machines. This paper presents the state of the art in the scalable computation of topological abstractions on scalar data, in shared memory parallel on single machines, and in distributed memory parallel on multiple machines. We highlight results for set‐based, graph‐based and complex‐based abstractions and organize the state of the art based on this taxonomy. The paper identifies parallelization and distribution techniques common in topological algorithms and highlights further areas of interest with underdeveloped efforts.

97 MATHEMATICS AND COMPUTING

Unlocking Scholarly Insights: Leveraging Machine Learning Approaches for Citation Analysis and Intent Classification

Publicly funded organizations, notably institutions like the Los Alamos National Laboratory (LANL), are deeply vested in acquiring robust productivity metrics to gauge the entirety of their research output. Motivated by the imperative to enhance institutional productivity assessment, this study investigates the utilization of Large Language Models (LLM) such as BERT-based models, as well as local LLaMa-30b-instruct and Mixtral-8x7b-instruct architectures for classifying type of URL referenced resources in academic papers such as software, dataset, as well as authorship intent. Challenges in discerning resource types from context are highlighted, along with the potential of BERT and LLMs to address these challenges. Through comprehensive analysis, this research unveils a notable surge in documents featuring URL citations, indicative of the escalating importance of digital resources in scholarly publications. Moreover, citations to datasets and software demonstrate consistent growth over time, underscoring their increasing significance. Our findings also reveal that LANL authors contribute substantially to accessible science, comprising about 10% of dataset and software mentions in LANL

Large Language Models, BERT, citation classificati

Data for "Design of Diverse, Functional Mitochondrial Targeting Sequences Across Eukaryotic Organisms Using Variational Autoencoder"

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

AI/ML

Fundamental Insights into Cathode Stability: Linking Compositional Tuning and Local Coordination in Complex Metal Oxides under Aqueous Transformations

Compositional tuning of complex metal oxides in Li-ion battery materials influences their performance as well as their end-of-life behavior, in particular, the tendency to release toxic metal cations in aqueous solution. We modeled ternary variants of a parent LiCoO 2 delafossite structure by varying the metal identity and relative amounts. This yielded ten model formulations of Li(A 4/6 B 1/6 C 1/6 )O 2 , where the material is enriched with the A metal and doped with B and C, with Ni, Mn, Co, Fe, Al, V, and Ti as constituent metals. To assess their stability in aqueous conditions, metal release energetics were calculated using a combination of Density Functional Theory calculations and thermodynamics. Metal release in ternary oxides is dictated by subtle variations in the coordination environment of the leaving group. To identify governing chemical features across diverse compositions with varying local coordination environments, we leverage random forest regression and descriptor importance analysis. A key result is that metal–oxygen orbital hybridization, quantified using a projected density-of-states-derived descriptor, H d/p , provides a physically grounded measure of interaction strength that governs metal release energetics. This refined perspective goes beyond conventional oxidation state considerations and offers more robust insights for materials science. Finally, we model defect surface-bound O 2 dimer formation as a proxy for reactive oxygen species (ROS) generation. The results show that Ni-rich compositions more readily stabilize spin-polarized O 2 dimers, corroborating experimental reports of an increased ROS-driven biological response. In conclusion, our results establish a compositional and electronic basis for metal release and surface oxygen reactivity that form a rationale for complex metal oxide design principles.

36 MATERIALS SCIENCE

Raman spectroscopic investigation of ianthinite [U$_2^{4+}$(UO$_2$)$_4$O$_6$(OH)$_4$(H$_2$O)$_4$]·$5$H$_2$O, a rare mixed-valence uranium oxide hydrate

Ianthinite ([[U$_2^{4+}$(UO$_2$)$_4$O$_6$(OH)$_4$(H$_2$O)$_4$]·$5$H$_2$O) is an exotic mineral that possesses U in both tetravalent and hexavalent oxidation states and is structurally related to the U 3 O 8 polymorphs, which are commonly encountered technogenic materials in the nuclear fuel cycle. Despite the similarities between U 3 O 8 and ianthinite, and the importance of ianthinite in U paragenesis, no Raman spectra have been reported for this mineral. Here, to gain a more complete understanding of how structural attributes of ianthinite give rise to observable spectroscopic features and how these may relate to important materials in the nuclear fuel cycle, we provide, for the first time, Raman spectra of ianthinite. Ianthinite readily oxidizes at ambient conditions, complicating analysis of phase-pure material. Several analytical methods are employed herein to decouple the Raman features of ianthinite from its alteration product(s). First, a simple difference spectrum is presented, then results of Raman spectroscopic mapping are employed, and finally, we use a novel processing and analysis method. Each analysis method provides different insight into structural features that are unique to ianthinite, in particular, features that are attributable to U(IV) in distorted octahedral coordination in both ianthinite and U 3 O 8 phases.

Spano, Tyler L. [Oak Ridge National Laboratory (OR

Machine Learning in the Context of Laser-Induced Breakdown Spectroscopy

The integration of machine learning (ML) with Laser-Induced Breakdown Spectroscopy (LIBS) has revolutionized the analytical capabilities of LIBS. The combi-nation of both methods enables more accurate and efficient data analysis. While LIBS itself is a powerful technique for elemental analysis, the vast amount of spectral data it generates can be hard to interpret. Machine learning addresses these challenges by leveraging algorithms that can learn from data, identify patterns, and make predictions without explicit programming for the interpretation of each specific task. In LIBS application, ML techniques are used to enhance various analytical processes. For example, ML algorithms can classify materials based on their spectral fingerprints, predict the concentration of elements in a sample, and identify underlying patterns within complex datasets. Here, this application improves the precision of LIBS analyses while significantly reducing the time required for data processing and interpretation. In this chapter, the fundamental concepts of ML will be discussed first. Following this, the process of data splitting and the importance of feature selection will be examined. Several machine learning methods will then be closely examined, exploring how each can benefit LIBS analysis and highlighting their respective advantages and shortcomings. This structured approach will provide a comprehensive understanding of the integration of ML in the context of LIBS analysis.

47 OTHER INSTRUMENTATION

ToF-SIMS spectral data analysis of Paenibacillus sp. 300A biofilms and planktonic cells

Analysis of bacterial biofilms is particularly challenging and important with diverse applications from systems biology to biotechnology. Among the variety of techniques that have been applied, time-of-flight secondary ion mass spectrometry (ToF-SIMS) has many promising features in studying the surface characteristics of biofilms. ToF-SIMS offers high spatial resolution and high mass accuracy, which permit surface sensitive analysis of biofilm components. Thus, ToF-SIMS provides a powerful solution to addressing the challenge of bacterial biofilm analysis. This dataset covers ToF-SIMS analysis of Paenibacillus sp. 300A (300A) isolated from the Hanford site in Richland, WA. The strain is known to have metal and sulfur reducing properties and can be used for bioremediation, wastewater treatment, bioengineering and technology development. There is a current need to identify small molecules and fragments produced from bacterial biofilms. Static ToF-SIMS spectra of 300A were obtained using an IONTOF TOF-SIMS V instrument equipped with a 25 keV Bi 3 + metal ion gun. Identified molecules and molecular fragments are compared against known biological databases and the reported peaks have at least 65 ppm mass accuracy. These molecules range from lipids and fatty acids to flavonoids, quinolones, and other naturally occurring organic compounds. It is anticipated that the spectral identification of key peaks will assist detection of metabolites, extracellular polymeric substance molecules like polysaccharides, and biologically relevant small molecules using ToF-SIMS in future surface and interface research of bacterial biofilms.

Biofilms

ToF-SIMS spectral analysis of Shewanella oneidensis MR-1 biofilms

Analysis of bacterial biofilms is particularly challenging and important with diverse applications from systems biology to biotechnology. Among the variety of techniques that have been applied, time-of-flight secondary ion mass spectrometry (ToF-SIMS) has many powerful features in studying the surface characteristics of biofilms. ToF-SIMS offers high spatial resolution, mass resolution, and mass accuracy, which permit surface sensitive analysis of biofilm components. Thus, ToF-SIMS provides a powerful solution to addressing the challenge of bacterial biofilm analysis. This dataset covers ToF-SIMS analysis of Shewanella oneidensis MR-1 isolated from freshwater lake sediment in New York state. The MR-1 strain is known to have metal and sulfur reducing properties and it can be used for bioremediation and wastewater treatment. There is a current need to identify small molecules and fragments produced from bacterial biofilms, especially those from extracellular polymeric substance (EPS). Static ToF-SIMS spectra of MR-1 were obtained using an IONTOF TOF.SIMS V instrument equipped with a 25 keV Bi$^+_3$ metal ion gun. Identified molecules and molecular fragments are compared against known biological databases and the reported peaks have at least 65 ppm mass accuracy. These molecules range from lipids, fatty acids, flavonoids, and quinolones to other naturally occurring organic compounds. It is anticipated that the mass spectral identification of key peaks will assist detection of metabolites, EPS molecules like polysaccharides, and biologically relevant small organic molecules using ToF-SIMS in future surface and interface research.

59 BASIC BIOLOGICAL SCIENCES

Advancing AI-Driven Analysis in X-ray Absorption Spectroscopy: Spectral Domain Mapping and Universal Models

In recent years, rapid progress has been made in developing artificial intelligence (AI) and machine learning (ML) methods for X-ray absorption spectroscopy (XAS) analysis. Compared to traditional XAS analysis methods, AI/ML approaches offer dramatic improvements in efficiency and help eliminate human bias. To advance this field, we advocate an AI-driven XAS analysis pipeline that features several interconnected key building blocks: benchmarks, workflows, databases, and AI/ML models. Specifically, we present two case studies for XAS ML. In the first study, we demonstrate the importance of reconciling the discrepancies between simulation and experiment using spectral domain mapping (SDM). Our ML model, which is trained solely on simulated spectra, predicts an incorrect oxidation state trend for Ti atoms in a combinatorial zinc titanate film. After transforming the experimental spectra into a simulation-like representation using SDM, the same model successfully recovers the correct oxidation state trend. In the second study, we explore the development of universal XAS ML models that are trained on the entire periodic table, which enables them to leverage common trends across elements. Looking ahead, we envision that an AI-driven pipeline can unlock the potential of real-time XAS analysis to accelerate scientific discovery.

36 MATERIALS SCIENCE

Phase Picking Beyond Local Distances: Where Waveform Filtering Still Matters for Deep Learning Models

Waveform filtering is a standard step in traditional seismic phase picking but often receives little attention in deep learning workflows, where models are typically trained on raw or minimally processed waveforms. Although this strategy performs well for local events, we show that performance can degrade substantially at regional distances. To address this limitation, we introduce two ways to incorporate multiband-filtered waveforms into deep learning phase pickers. The stacking approach concatenates filtered inputs along the channel dimension, while the branching approach processes each frequency band through a dedicated network branch before feature fusion. Both approaches can substantially improve performance across epicentral distances of 0° to 20°, but their effectiveness depends strongly on the selected frequency bands. Tests with multiple filter banks show that filter-bank design should be treated as part of model optimization rather than as a fixed preprocessing choice. Grad-CAM analysis of the branching model indicates that band importance varies among waveform samples and across training realizations, with only a weak overall preference for the 0.25 to 0.5 Hz band. These results show that no single filter band is consistently optimal and demonstrate that explicit feature engineering remains valuable for robust deep learning-based seismic phase picking.

58 GEOSCIENCES

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES

Explainable AI classification for parton density theory

Quantitatively connecting properties of parton distribution functions (PDFs, or parton densities) to the theoretical assumptions made within the QCD analyses which produce them has been a longstanding problem in HEP phenomenology. To confront this challenge, we introduce an ML-based explainability framework, XAI4PDF, to classify PDFs by parton flavor or underlying theoretical model using ResNet-like neural networks (NNs). By leveraging the differentiable nature of ResNet models, this approach deploys guided backpropagation to dissect relevant features of fitted PDFs, identifying x-dependent signatures of PDFs important to the ML model classifications. By applying our framework, we are able to sort PDFs according to the analysis which produced them while constructing quantitative, human-readable maps locating the x regions most affected by the internal theory assumptions going into each analysis. This technique expands the toolkit available to PDF analysis and adjacent particle phenomenology while pointing to promising generalizations.

Artificial Intelligence

Ascribe XR v0.1.0

Ascribe XR is an immersive visualization software designed for scientists and engineers working with 3D data sets. Its key features include interactive exploration, multi-user collaboration, and flexible data import capabilities, supporting various formats such as meshes, volumes, and terrain maps. The software utilizes Godot, OpenXR and PC-VR technology to provide an immersive experience. Ascribe XR is used for data analysis, visualization, and collaboration in various fields, enabling users to gain deeper insights into complex data sets. Its advantages over similar technologies include its flexibility, customizability, and ease of use. Ascribe XR's interactive and immersive environment facilitates collaboration and accelerates the discovery process. Compared to traditional 2D visualization tools, Ascribe XR offers a more engaging and intuitive experience, allowing users to explore complex data sets in a more natural and interactive way. Its ability to support multi-user collaboration and flexible data import capabilities make it a versatile tool for various applications. Overall, Ascribe XR provides a unique combination of features, usability, and performance, making it an attractive solution for scientists and engineers working with 3D data sets.

Pandolfi, Ronald [Lawrence Berkeley National Labor

Pooled Rideshare in the U.S.: An Exploratory Study of User Preferences

Pooled ridesharing offers on-demand, one-way, cost-effective transportation for passengers traveling in similar directions via a shared vehicle ride with others they do not know. Despite its potential benefits, the adoption of pooled rideshare remains low in the United States. This exploratory study aims to evaluate potential service improvements and features that may increase users’ willingness to adopt the service. The study analyzed transportation behaviors, rideshare preferences, and willingness to adopt pooled rideshare services among 8296 U.S. participants in 2025, building on findings from a 2021 nationwide survey of 5385 U.S. participants. The study incorporated 77 actionable items developed from the results of the 2021 survey to assess whether addressing specific user-generated topics such as safety, reliability, convenience, and privacy can improve pooled rideshare use. A side-by-side comparison of the 2021 and 2025 data revealed shifts in transportation behavior, with personal rideshare usage increasing from 22% to 28%, public transportation from 21% to 27%, and pooled rideshare from 6% to 8%, while personal vehicle (79%) use remained dominant. Participants rated features such as driver verification (94%), vehicle information (93%), peak time reliability (93%), and saving time and money (92–93%) as most important for improving rideshare services. A pre-to-post analysis of willingness to use pooled rideshare utilizing the actionable items as per respondents’ preferences showed improvement: “definitely will” increased from 15.9% to 20.1% and “probably will” rose from 35.6% to 47.7%. These results suggest that well-targeted service improvements may meaningfully enhance pooled rideshare acceptance. This study offers practical guidance for Transportation Network Companies (TNCs) and policymakers aiming to improve pooled rideshare as well as potential future research opportunities.

Transportation Network Companies (TNCs)

The importance of electron scattering in the analysis of actinide X-ray spectroscopy

Abstract Manifestations of electron scattering in X-ray spectroscopy have been evident for decades. Here, it will be shown that the proper interpretation of variants of X-ray Absorption Spectroscopy (XAS) of actinide materials must include an accurate treatment of features caused by electron scattering, i.e., EXAFS or Extended X-ray Absorption Fine Structure. These EXAFS features can be of such low energy that they are within ten to twenty electron volts of the Unoccupied Density of States (UDOS), immediately above the Fermi Energy or Band Gap. The adaption of simple models using the FEFF simulation program will be presented, including the demonstration of the robust nature of the results from different models. Graphical abstract

Tobin, J. G. (ORCID:0000000322943301)

Simulation and analysis of small angle scattering (SAS) patterns of Ni-based superalloy microstructures generated by a phase-field model

This paper investigates the relationship between microstructural features and small-angle scattering (SAS) patterns in Ni-based superalloys using a combined phase-field and SAS simulation approach coupled with microstructure analyses. The simulated SAS patterns accurately capture key experimental observations previously reported in the literature, including the time-dependent transition from circular to square-shaped precipitates and the development of anisotropic SAS patterns. Importantly, our analysis reveals the correlations between characteristic length scales extracted from SAS profiles and microstructural descriptors, such as precipitate size and inter-precipitate distance. These findings provide a comprehensive understanding of the link between SAS profiles and microstructure evolution in Ni-based superalloys, offering valuable insights for materials characterization and design.

Microstructure