Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Street context of various demographic groups in their daily mobility

Abstract We present an urban science framework to characterize phone users’ exposure to different street context types based on network science, geographical information systems (GIS), daily individual trajectories, and street imagery. We consider street context as the inferred usage of the street, based on its buildings and construction, categorized in nine possible labels. The labels define whether the street is residential, commercial or downtown, throughway or not, and other special categories. We apply the analysis to the City of Boston, considering daily trajectories synthetically generated with a model based on call detail records (CDR) and images from Google Street View. Images are categorized both manually and using artificial intelligence (AI). We focus on the city’s four main racial/ethnic demographic groups (White, Black, Hispanic and Asian), aiming to characterize the differences in what these groups of people see during their daily activities. Based on daily trajectories, we reconstruct most common paths over the street network. We use street demand (number of times a street is included in a trajectory) to detect each group’s most relevant streets and regions. Based on their street demand, we measure the street context distribution for each group. The inclusion of images allows us to quantitatively measure the prevalence of each context and points to qualitative differences on where that context takes place. Other AI methodologies can further exploit these differences. This approach presents the building blocks to further studies that relate mobile devices’ dynamic records with the differences in urban exposure by demographic groups. The addition of AI-based image analysis to street demand can power up the capabilities of urban planning methodologies, compare multiple cities under a unified framework, and reduce the crudeness of GIS-only mobility analysis. Shortening the gap between big data-driven analysis and traditional human classification analysis can help build smarter and more equal cities while reducing the efforts necessary to study a city’s characteristics.

Salgado, Ariel (ORCID:0000000177015372)↗

Design Principles in Engineering of Multigrain Nanocatalysts via Multiscale Electronic Structure Characterization

Engineering grain boundary (GB) strain provides a promising pathway to tune the catalytic properties of nanocrystals. However, structural heterogeneity from random grain orientation and geometry has limited clear structure–property correlations. Here, we utilize a multigrain Co3O4/Mn3O4 core/shell nanocrystal platform as a model system to systematically investigate how geometric misfit strain at GBs serves as catalytically active sites for the oxygen reduction reaction. Through precise subnanometer-level control over grain morphology and by integrating multiscale electronic structure characterization, we identify the electronic structural signature of GB defects and establish a direct correlation between localized strain fields and modified electronic states. Strain modulation at GBs alters the eg orbital energy levels, with elongation along the z-axis combined with shear strain stabilizing the eg states, in contrast to the destabilization observed under pure shear strain. This stabilization mechanism enhances the electrocatalytic activity and selectivity of strained GBs compared with strain-relaxed grain surfaces. Furthermore, we reveal that GBs exhibit a radial strain gradient, producing a spatial energy shift that further modulates local electronic structures, as resolved through the classification of electron energy loss spectroscopy data. Together, these findings demonstrate that geometric misfit strain enables precise tuning of grain geometry and the resulting electronic structures, offering a robust strategy for engineering next-generation nanocatalysts.

Cho, Min Gee↗

Real-time tracking of structural evolution in 2D MXenes using theory-enhanced machine learning

In situ Electron Energy Loss Spectroscopy (EELS) combined with Transmission Electron Microscopy (TEM) has traditionally been pivotal for understanding how material processing choices affect local structure and composition. However, the ability to monitor and respond to ultrafast transient changes, now achievable with EELS and TEM, necessitates innovative analytical frameworks. Here, we introduce a machine learning (ML) framework tailored for the real-time assessment and characterization of in operando EELS Spectrum Images (EELS-SI). We focus on 2D MXenes as the sample material system, specifically targeting the understanding and control of their atomic-scale structural transformations that critically influence their electronic and optical properties. This approach requires fewer labeled training data points than typical deep learning classification methods. By integrating computationally generated structures of MXenes and experimental datasets into a unified latent space using Variational Autoencoders (VAE) in a unique training method, our framework accurately predicts structural evolutions at latencies pertinent to closed-loop processing within the TEM. This study presents a critical advancement in enabling automated, on-the-fly synthesis and characterization, significantly enhancing capabilities for materials discovery and the precision engineering of functional materials at the atomic scale.

47 OTHER INSTRUMENTATION↗

Optimizing training trajectories in variational autoencoders via latent Bayesian optimization approach *

Unsupervised and semi-supervised ML methods such as variational autoencoders (VAE) have become widely adopted across multiple areas of physics, chemistry, and materials sciences due to their capability in disentangling representations and ability to find latent manifolds for classification and/or regression of complex experimental data. Like other ML problems, VAEs require hyperparameter tuning, e.g. balancing the Kullback–Leibler and reconstruction terms. However, the training process and resulting manifold topology and connectivity depend not only on hyperparameters, but also their evolution during training. Because of the inefficiency of exhaustive search in a high-dimensional hyperparameter space for the expensive-to-train models, here we have explored a latent Bayesian optimization (zBO) approach for the hyperparameter trajectory optimization for the unsupervised and semi-supervised ML and demonstrated for joint-VAE with rotational invariances. We have demonstrated an application of this method for finding joint discrete and continuous rotationally invariant representations for modified national institute of standards and technology database (MNIST) and experimental data of a plasmonic nanoparticles material system. The performance of the proposed approach has been discussed extensively, where it allows for any high dimensional hyperparameter trajectory optimization of other ML models.

42 ENGINEERING↗

ZTF_image_classification

This repo demonstrates how to use the LLNL developed Gaussian process hyperparameter estimation method, MuyGPS, for image classification using Zwicky Transient Facility (ZTF) data.

Pruett, Kerianne↗

Magnetic Gears for a Marine Hydrokinetic Generator Component Model

The goal of this project is to design, fabricate, and test a hermetically sealed 50 kilowatt (kW) multistage magnetically geared generator (MGG). The Component Content Model provides data submitters with an easy and consistent means of uploading data and associated meta data about a component that is currently under development. The data fields include generic information about the component, technology classifications, current costs and performance, proposed target goals, and the environment that the component is operated in. These data are important to DOE and will be used to develop data products that provide quantitative information to guide and support programmatic decisions. Data will also be used by DOE in general assessments of MHK component readiness, performance, costs, and proposed plans. The ultimate goal is to use these data to perform research and tailor programs to best benefit the industry.

16 TIDAL AND WAVE POWER↗

Constituent Data Replacement Tool

The purpose of this tool is to estimate key parameters that may be missing in public wastewater composition datasets. The tool can be applied to develop complete treatment and critical mineral extraction profiles for leachate, produced water and other aqueous waste streams. The tool applies machine learning algorithms to replace missing data in a user’s water data set that are adjusted based on user preferences for options including algorithm type, number of features, and classification variables. The tool can use the user’s data alone or combine user data with the NEWTS USGS Produced Water Database for more robust training. This research was funded by the U.S. Department of Energy’s Office Fossil Energy and Carbon Management (FECM) through National Energy Technology Laboratory’s ongoing research under the Water Management for Power System Field Work Proposal, DE-FECM 1022428 and Critical Minerals Field Work Proposal, DE-FECM 1022420.

Aqueous Chemistry↗

Advanced Distributed Optical Fiber Sensor Systems for Pipeline Integrity Monitoring

Distributed fiber optic sensors allow the measurement of structural parameters such as static/dynamic strain, temperature, pressure, and vibrations at thousands of locations along a single fiber cable. Deep neural network (DNN) algorithms were developed for rapid data processing speed and vibration event classification.

Lalam, Nageswara↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

Using Machine Learning Algorithm to Detect Blowing Snow and Fog in Antarctica Based on Ceilometer and Surface Meteorology Systems

Blowing snow is a common weather phenomenon in Antarctica and plays an important role in the water vapor cycle and ice sheet mass balance. Although it has a significant impact on the climate of Antarctica, people do not know much about this process. Fog events are difficult to distinguish from blowing snow events using existing detection algorithms by a ceilometer. In this study, based on ceilometer, the meteorological parameters observed by surface meteorology systems are further combined to detect blowing snow and fog using the AdaBoost algorithm. The weather phenomena recorded by human observers are ‘true’. The dataset is collected from 1 January 2016 to 31 December 2016 at the AWARE site. Among them, three-quarters of the data are used as the training set and the rest of the data as the testing set. The classification accuracy of the proposed algorithm for the testing set is about 94%. Compared with the Loeb method, the proposed algorithm can detect 89.12% of blowing snow events and 76.10% of fog events, while the Loeb method can only identify 64.29% of blowing snow events and 31.87% of fog events.

54 ENVIRONMENTAL SCIENCES↗

The MOST Hosts Survey: Spectroscopic Observation of the Host Galaxies of ∼40,000 Transients Using DESI

We present the Multi-Object Spectroscopy of Transient (MOST) Hosts survey. The survey is planned to run throughout the 5 yr of operation of the Dark Energy Spectroscopic Instrument (DESI) and will generate a spectroscopic catalog of the hosts of most transients observed to date, in particular all the supernovae observed by most public, untargeted, wide-field, optical surveys (Palomar Transient Factory, PTF/intermediate PTF, Sloan Digital Sky Survey II, Zwicky Transient Facility, DECAT, DESIRT). Science cases for the MOST Hosts survey include Type Ia supernova cosmology, fundamental plane and peculiar velocity measurements, and the understanding of the correlations between transients and their host-galaxy properties. Here we present the first release of the MOST Hosts survey: 21,931 hosts of 20,235 transients. These numbers represent 36% of the final MOST Hosts sample, consisting of 60,212 potential host galaxies of 38,603 transients (a transient can be assigned multiple potential hosts). Of all the transients in the MOST Hosts list, only 26.7% have existing classifications, and so the survey will provide redshifts (and luminosities) for nearly 30,000 transients. A preliminary Hubble diagram and a transient luminosity–duration diagram are shown as examples of future potential uses of the MOST Hosts survey. The survey will also provide a training sample of spectroscopically observed transients for classifiers relying only on photometry, as we enter an era when most newly observed transients will lack spectroscopic classification. The MOST Hosts DESI survey data will be released on a rolling cadence and updated to match the DESI releases.

79 ASTRONOMY AND ASTROPHYSICS↗

Nuclear Forensics Scanning

Explore the source record for details and available documents.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Catalog of Candidate Double and Lensed Quasars from Gaia and WISE Data

Making use of strong correlations between closely separated multiple or double sources and photometric and astrometric metadata in Gaia Early Data Release 3 (EDR3), we generate a catalog of candidate double- and multiply imaged lensed quasars and active galactic nuclei (AGNs), comprising 3140 systems. It includes two partially overlapping parts: a sample of distant (redshifts mostly greater than 1) sources with perturbed data; and systems that have been resolved into separate components by Gaia at separations less than 2''. For the first part, which is roughly one-third of the published catalog, we synthesized 0.617 million redshifts using multiple machine-learning prediction and classification methods, using independent photometric and astrometric data from Gaia EDR3 and the Wide-field Infrared Survey Explorer, with accurate spectroscopic redshifts from the Sloan Digital Sky Survey (SDSS) as a training set. Using these synthetic redshifts, we estimate a 4.9% rate of interlopers with spectroscopic redshifts below 1 in this part of the catalog. Unresolved candidate double and dual AGNs and quasars are selected as sources with a marginally high BP/RP excess factor (phot_bp_rp_excess_factor), which is sensitive to source extent, limiting our search to high-redshift quasars. For the second part of the catalog, additional filters on measured parallax and near-neighbor statistics are applied to diminish the propagation of the remaining stellar contaminants. The estimated rate of the positives (double or multiple sources) is 98%, and the estimated rate of dual (physically related) quasars is greater than 54%. A few dozen serendipitously found objects of interest are discussed in more detail, including known and new lensed images, planetary nebulae, young IR stars of peculiar morphology, and quasars with catastrophic redshift errors in SDSS.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep Cellular Recurrent Network for Efficient Analysis of Time-Series Data With Spatial Information

Efficient processing of large-scale time series data is an intricate problem in machine learning. Conventional sensor signal processing pipelines with hand engineered feature extraction often involve huge computational cost with high dimensional data. Deep recurrent neural networks have shown promise in automated feature learning for improved time-series processing. However, generic deep recurrent models grow in scale and depth with increased complexity of the data. This is particularly challenging in presence of high dimensional data with temporal and spatial characteristics. Consequently, this work proposes a novel deep cellular recurrent neural network (DCRNN) architecture to efficiently process complex multi-dimensional time series data with spatial information. Here, the cellular recurrent architecture in the proposed model allows for location-aware synchronous processing of time series data from spatially distributed sensor signal sources. Extensive trainable parameter sharing due to cellularity in the proposed architecture ensures efficiency in the use of recurrent processing units with high-dimensional inputs. This study also investigates the versatility of the proposed DCRNN model for classification of multi-class time series data from different application domains. Consequently, the proposed DCRNN architecture is evaluated using two time-series datasets: a multichannel scalp EEG dataset for seizure detection, and a machine fault detection dataset obtained in-house. The results suggest that the proposed architecture achieves state-of-the-art performance while utilizing substantially less trainable parameters when compared to comparable methods in the literature.

60 APPLIED LIFE SCIENCES↗

Atmospheric condition identification in multivariate data through a metric for total variation

Identification of atmospheric conditions within a multivariable atmospheric data set is a necessary step in the validation of emerging and existing high-fidelity models used to simulate wind plant flows and operation.Atmospheric conditions relevant for wind energy research include stationary conditions, given the need for well-converged statistics for model validation, as well as conditions observed less frequently, such as extreme atmospheric events, which are used in wind turbine and wind plant design.Aggregation of observations without regard to covariance between time series discounts the dynamical nature of the atmosphere and is not sufficiently representative of atmospheric conditions.Identification and characterization of continuous time periods with atmospheric conditions that have a high value for analysis or simulation set the stage for more advanced model validation and the development of real-time control and operational strategies.The current work explores a single metric for variation in a multivariate data sample that quantifies variability within each channel as well as covariance between channels.The total variation is used to identify conditions of interest that conform to desired objective functions, such as stationary conditions, ramps or waves of wind speed, and changes in wind direction.Total variation is somewhat sensitive to the presence of outliers in the input data, and the method is best complemented by quality-control procedures to ensure reliable results.The direct detection and classification of events or conditions of interest within atmospheric data sets is vital to developing our understanding of wind plant response and to the formulation of forecasting and control models.

17 WIND ENERGY↗

Quantifying Epistemic Uncertainty in Binary Classification via Accuracy Gain

ABSTRACT Recently, a surge of interest has been given to quantifying epistemic uncertainty (EU), the reducible portion of uncertainty due to lack of data. We propose a novel EU estimator in the binary classification setting, as the posterior expected value of the empirical gain in accuracy between the current prediction and the optimal prediction. In order to validate the performance of our EU estimator, we introduce an experimental procedure where we take an existing dataset, remove a set of points, and compare the estimated EU with the observed change in accuracy. Through real and simulated data experiments, we demonstrate the effectiveness of our proposed EU estimator.

97 MATHEMATICS AND COMPUTING↗

A Machine Learning Approach for the Prediction of Formability and Thermodynamic Stability of Single and Double Perovskite Oxides

Perovskite oxides continue to attract huge interest due to their fascinating and wide-ranging properties for diverse applications. The tunability of these properties may be further enhanced by increasing their compositional complexity via double perovskite-ordered configurations containing multiple cations. In this work, we focus on an exhaustive chemical space of single and double oxide perovskites and optimally explore this space to identify novel compositions that are likely to form stable compounds. Critically, we examine the relationship between formability, the practical ability to synthesize a compound, and stability, the thermodynamic preference to form the structure. Our formability and stability training data sets were enumerated from the available experimental literature and in-house density functional theory computations and contained 1505 and 3469 examples, respectively, representing state-of-the-art in the current open literature in perovskite and double perovskite compounds. Subsequently, cross-validated and highly accurate machine learning classification models are built using these training data sets and employed to screen for novel stable oxide perovskites. The study identifies (1) atomic features relevant to prediction of formability and stability in perovskite and double perovskite compounds, (2) the importance of including energy contributions due to local structural relaxations going beyond the high symmetry perovskite phase, and (3) 437,828 double perovskite compounds that are likely to be stable and 891,188 compounds that are likely to be formable. From the intersection of this large chemical space of formable and stable oxide perovskites, 414 compositions are identified as the most promising candidates for future experimental synthesis of novel oxide perovskites. The developed models may be generalized and have implications beyond perovskite discovery if applied to other families of compounds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Semi-supervised Learning of Dynamical Systems with Neural Ordinary Differential Equations: A Teacher-Student Model Approach

Modeling dynamical systems is crucial for a wide range of tasks, but it remains challenging due to complex nonlinear dynamics, limited observations, or lack of prior knowledge. Recently, data-driven approaches such as Neural Ordinary Differential Equations (NODE) have shown promising results by leveraging the expressive power of neural networks to model unknown dynamics. However, these approaches often suffer from limited labeled training data, leading to poor generalization and suboptimal predictions. On the other hand, semi-supervised algorithms can utilize abundant unlabeled data and have demonstrated good performance in classification and regression tasks. We propose TS-NODE, the first semi-supervised approach to modeling dynamical systems with NODE. TS-NODE explores cheaply generated synthetic pseudo rollouts to broaden exploration in the state space and to tackle the challenges brought by lack of ground-truth system data under a teacher-student model. TS-NODE employs an unified optimization framework that corrects the teacher model based on the student's feedback while mitigating the potential false system dynamics present in pseudo rollouts. TS-NODE demonstrates significant performance improvements over a baseline Neural ODE model on multiple dynamical system modeling tasks.

Wang, Yu↗