Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Time series segmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Real-time correlation of crystallization and segmental order in conjugated polymers

Modulating the segmental order in the morphology of conjugated polymers is widely recognized as a crucial factor for achieving optimal electronic properties and mechanical deformability. However, it is worth noting that the segmental order is typically associated with the crystallization process, which can result in rigid and brittle long-range ordered crystalline domains. To precisely control the morphology, a comprehensive understanding of how highly anisotropic conjugated polymers form segmentally ordered structures with ongoing crystallization is essential, yet currently elusive. To fill this knowledge gap, we developed a novel approach with a combination of stage-type fast scanning calorimetry and micro-Raman spectroscopy to capture the series of specimens with a continuum in the polymer percent crystallinity and detect the segmental order in real-time. Through the investigation of conjugated polymers with different backbones and side-chain structures, we observed a generally existing phenomenon that the degree of segmental order saturates before the maximum crystallinity is achieved. This disparity allows the conjugated polymers to achieve good charge carrier mobility while retaining good segmental dynamic mobility through the tailored treatment. Moreover, the crystallization temperature to obtain optimal segmental order can be predicted based on T g and T m of conjugated polymers. In conclusion, this in-depth characterization study provides fundamental insights into the evolution of segmental order during crystallization, which can aid in designing and controlling the optoelectronic and mechanical properties of conjugated polymers.

36 MATERIALS SCIENCE↗

Localized keyhole pore prediction during laser powder bed fusion via multimodal process monitoring and X-ray radiography

Systematic fault detection and control during laser powder bed fusion (L-PBF) has been a long-standing objective for system manufacturers and researchers in the additive manufacturing (AM) industry. This manuscript investigates a data fusion approach for detection of keyhole porosity formation during laser irradiation of Ti-6Al-4V substrates by concurrent recording of thermally induced optical emission measured using both off-axis and coaxial photodiode sensors, and acoustic emission. Subsurface defect formation was monitored via high-speed synchrotron X-ray imaging at 20,000 frames per second, enabling temporal registration of keyhole pore formation events to the monitoring signals at a resolution of 50 µs. We developed data fusion machine learning (ML) models for localized prediction of keyhole pore formation at various time scales ranging from 0.5 ms to 2 ms. The signal segments were featurized using two independent approaches: (1) power spectral density (PSD) and (2) highly comparative time series analysis (HCTSA) framework. The extracted features from different sensor modalities were fused together to construct a multimodal feature space and sequential feature selection was used to determine the most informative features for training the ML models. The predictive performance was evaluated for three classifying algorithms: Support Vector Machine (SVM), K-Nearest Neighbor (KNN), and Gaussian Naive Bayes (GNB). As a result, pore formation events were predicted with up to 0.95 F1-score, 1.0 recall and 0.94 accuracy. The most heavily weighted features indicate that model performance is chiefly governed by the acoustic monitoring signal, with a secondary contribution from the optical emission sensors.

36 MATERIALS SCIENCE↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University ↗

TICC Clustering Library v.1.0

SAND2024-01234O TICC is a clustering algorithm that labels a sequence of data points according to numerical properties. This library is a Python implementation of the algorithm described in "Toeplitz Inverse Covariance-Based Clustering of Multivariate Time Series Data" (Hallac et al. 2017). It includes documentation, performance improvements, examples, and test coverage. This library allows users to automatically segment a series of multivariate data points according to their covariance—that is, the way the values at each data point are changing in relation to one another. This is useful for identifying periods in which a system is behaving. For example, if a sensor is measuring a car's velocity, steering wheel angle, braking and acceleration, TICC can determine when the car was stopped, beginning/exiting a turn, slowing or accelerating at an intersection, or driving on straight or curved roads. TICC can be applied to measure multiple quantities at known times. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525

Dalbey, Keith↗

Distribution System Segmented Model Simplification with Independent Dynamically Changing End-Use Loads

With more emerging technologies in power distribution systems, to support quasi-static time-series (QSTS) simulation over short time frames, accurately modeling the dynamic behavior of end-use load under various conditions is im- portant. However, time and cost can be a burden for the QSTS simulation of thousands of houses representing dynamic load behaviors on large distribution feeders. There are two main contributions of this paper in order to reduce the QSTS simulation complexity. First, the paper presents a method to stochasti- cally model the dynamic behavior of end-use load for QSTS simulations. Second, using the stochastically modeled load behavior as an input to a distribution sys- tem, and building on our previous QSTS simulation simplification/acceleration method that relies on uniform load behavior, this paper presents a segment substitution approach for the simplification of distribution system models with independent dynamically changing end-use loads. Our numerical results show that compared to GridLAB-D, a software tool that simulates end-use load be- havior through physical models, the proposed stochastic modeling approach achieved an 89.3% simulation time reduction for end-use loads, while accurately capturing the temporally changing total load behavior. Moreover, the QSTS simulation of a large test distribution system through OpenDSS showed that the proposed simplification method decreased the feeder simulation time by a factor of three at less than 0.3% voltage error. Compared to our previous method with independent dynamic loads, the voltage error is reduced by a factor of four.

Power distribution system modeling, stochastic pro↗

Deep learning-enabled probing of irradiation-induced defects in time-series micrographs

Modeling time-series data with convolutional neural networks (CNNs) requires building a model to learn in batches as opposed to training sequentially. Coupling CNNs with in situ or operando techniques opens the possibility of accurately segmenting dynamic reactions and mass transport phenomena to understand how materials behave under the conditions in which they are used. In this article, in situ ion irradiation transmission electron microscopy (TEM) images are used as inputs into the CNN to assess the defect generation rate, defect cluster density, and saturation of defects. We then use the output segmentation maps to correlate with conventional TEM micrographs to assess the model’s ability to detail nanoscale interactions. Next, we discuss the implications of preprocessing and hyperparameters on model variability, accuracy when expanded to other datasets, and the role of regularization when controlling model variance. Ultimately, we eliminate human bias when extrapolating physical metrics, speed up analysis time, decouple reactions that happen at 100 ms intervals, and deploy models that are both accurate and transferable to similar experiments.

Burns, Kory (ORCID:0000000158019909)↗

Wildfires identification: Semantic segmentation using support vector machine classifier

This paper deals with wildfire identification in the Alaska regions as a semantic segmentation task using support vector machine classifiers. Instead of colour information represented by means of BGR channels, we proceed with a normalized reflectance over 152 days so that such time series is assigned to each pixel. We compare models associated with $\mathcal{l}1$-loss and $\mathcal{l}2$-loss functions and stopping criteria based on a projected gradient and duality gap in the presented benchmarks.

Pecha, Marek↗

Hydrologic Regionalization under Data Scarcity: Implications for Streamflow Prediction

Continuous streamflow prediction is crucial in many applications of water resources planning and management. However, streamflow prediction is challenging, particularly in data-scarce regions. Here, we demonstrate an approach to regionalize the flow duration curve for predicting daily streamflow in the data-scare region of the central Himalayas. We developed a regression-based model to estimate streamflow at various segments of a flow duration curve by incorporating basin characteristics and climate variables. This study analyzes the sensitivities of proximity and characteristics between the donor (gauged) and receptor (ungauged) basins for time-series streamflow prediction. Our results show that regionalization techniques perform better in low to medium flows over high flows. Our findings are significant in the central Himalayan regional context to inform operational and management decisions in water sector projects like hydropower plants, which generally rely on low-to-medium streamflow information. Although the quantitative results are region-specific, the approach and insights are generalizable to the Himalayan region.

54 ENVIRONMENTAL SCIENCES↗

Vehicle Residual Value Analysis by Powertrain Type and Impacts on Total Cost of Ownership

Vehicle depreciation is a key factor in determining the total cost of vehicle ownership and consumer purchase behavior. This report examines how light-duty-vehicle residual values have evolved over time for conventional and advanced vehicle technologies, accounting for important factors such as market segment, size class, and country of assembly. Advancements in electric vehicle technology have led to plug-in vehicles exhibiting depreciation curves similar to those of conventional vehicles. This report compares two methods for determining depreciation trends (snapshot method and time-series method) in order to identify potential impacts on calculating vehicle total cost of ownership given differing data availability.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

The Large-Scale Vorticity Balance of the Antarctic Continental Margin in a Fine-Resolution Global Simulation

The depth-integrated vorticity budget of a global, eddy-permitting ocean/sea ice simulation over the Antarctic continental margin (ACM) is diagnosed to understand the physical mechanisms implicated in meridional transport. The leading-order balance is between the torques due to lateral friction, nonlinear effects, and bottom vortex stretching, although details vary regionally. Maps of the time-averaged depth-integrated vorticity budget terms and time series of the spatially averaged, depth-integrated vorticity budget terms reveal that the flow in the Amundsen, Bellingshausen, and Weddell Seas and, to a lesser extent, in the western portion of East Antarctica, is closer to an approximate topographic Sverdrup balance (TSB) compared to other segments of the ACM. Correlation and coherence analyses further support these findings, and also show that inclusion of the vorticity tendency term in the response (the planetary vorticity advection and the bottom vortex stretching term) increases the correlation with the forcing (the vertical net stress curl), and also increases the coherence between forcing and response at high frequencies across the ACM, except for the West Antarctic Peninsula. Finally, these findings suggest that the surface stress curl, imparted by the wind and the sea ice, has the potential to contribute to the meridional, approximately cross-slope, transport to a greater extent in the Amundsen, Bellingshausen, Weddell, and part of the East Antarctic continental margin than elsewhere in the ACM.

58 GEOSCIENCES↗

Improved Data Interpretation through Identification of Time Series Periodicity Changes

Analysis and interpretation of time series data is easiest when the data values occur at uniform intervals in time, but actual data may have differing data sampling frequencies, such as monthly and daily readings. Applying data analysis techniques, such as smoothing, to such a data set may not give a representative result between time segments. The ability to automatically distinguish time segments of differing data frequency would provide a means for applying data analysis independently to each segment, though a suitable blending at segment boundaries would be required. A method for detecting frequency changes was developed and applied to Gaussian and median smoothing of hydraulic head data from groundwater wells at the U.S. Department of Energy Hanford Site in southeastern Washington state. The process identifies time segments of high-frequency (daily) or low-frequency (greater than daily) data using adjusted-bandwidth Gaussian kernel density estimation and a threshold value, which are further refined to address small blocks of low-frequency data within larger blocks of high-frequency data. User-selectable levels of smoothing are then applied independently to the time segments prior to combining the segment results for a single smoothed data set. This time segment identification approach provides effective low- and high-frequency data separation, which provides a method to apply data analysis independently to each time segment.

97 MATHEMATICS AND COMPUTING↗

Vehicle Residual Value Analysis by Powertrain Type and Impacts on Total Cost of Ownership

Vehicle depreciation is a key factor in determining the total cost of vehicle ownership and consumer purchase behavior. This report examines how light-duty-vehicle residual values have evolved over time for conventional and advanced vehicle technologies, accounting for important factors such as market segment, size class, and country of assembly. Although plug-in vehicles have historically depreciated faster than conventional vehicles, advancements in electric vehicle technology have led to plug-in vehicles exhibiting depreciation curves similar to those of conventional vehicles. This report compares two methods for determining depreciation trends (snapshot method and time-series method) in order to identify potential impacts on calculating vehicle total cost of ownership given differing data availability.

33 ADVANCED PROPULSION SYSTEMS↗

Timeseries Photos of a Variably Inundated Stream: Umtanum Creek, Washington, United States

This dataset is associated with a broader study using game camera timeseries photos collected to evaluate stream variable inundation via changes in width (i.e. wet fraction). Four game cameras were deployed along Umtanum Creek (Washington, United States) to track changes in stream inundation over time. Drone imagery was collected at the same location on October 18, 2024 which was used to construct a digital elevation model (DEM) of the streambed topography. The associated paper and data can be found at https://doi.org/10.1016/j.envsoft.2025.106715 (Bao et al., 2025a)) and https://doi.org/10.15485/2589885 (Bao et al., 2025b), respectively. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) files that describes each file and a data dictionaries (DD) that describe all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) field protocol; and (5) folders containing game camera photos. Game camera photos are organized into folders for each camera (CDL, CUL, CDR, CUR; see readme for information on camera naming) by the month photos were collected. All files are .csv, .jpg, or .pdf.

AI image segmentation↗

Evaluating Image Classification Deep Convolutional Neural Network Architectures for Remaining Useful Life Estimation of Turbofan Engines

Accurate estimation of the remaining useful life (RUL) is a key component of condition-based maintenance (CBM) and prognosis and health management (PHM). Data-based models for the estimation of RUL are of particular interest because expert knowledge of systems is not always available, and physical modeling is often not feasible. Additionally, using data-based models, which make decisions based on raw sensor data, allow features to be learned instead of manually determined. In this work, deep convolutional neural network (CNN) architectures are investigated for their ability to estimate the RUL of turbofan engines. To improve the accuracy of the models, CNN architectures, which have proven successful in image classification, are implemented and tested. Specifically, the blocks used in the Visual Geometry Group (VGG) architecture, inception modules used in the GoogLeNet architecture, and residual blocks used in the ResNet architecture are incorporated. To account for varying flight lengths, the input to the models is a window of time series data collected from the engine under test. Window locations at the climb, cruise, and descent stages are considered. To further improve the RUL estimations, multiple overlapping windows at each location are used. This increases the amount of training data available and is found to increase the accuracy of the resulting RUL estimations by averaging the estimates from all overlapping segments. The model is trained and tested using the new Commercial Modular Aero-Propulsion System Simulation (N-CMAPSS) data set, and high prognosis accuracy was achieved. Furthermore, this work expands on the model developed and used in the 2021 PHM Society Data Challenge, which received second place.

convolutional neural networks↗

Subject-specific modeling framework for particle deposition using computational fluid dynamics

Quantifying particle deposition and dose in the respiratory tract requires a physiologically realistic representation and reproducible computational workflows. However, existing modeling frameworks, such as the International Commission on Radiological Protection (ICRP) compartmental models and the Multiple Path Particle Dosimetry (MPPD) tool, lack detailed deposition profiles and subject-specific capabilities. The combination of advances in computer vision algorithms applied to the respiratory tract and Computational Fluid and Particle Dynamics (CFPD) allows high-fidelity simulations of particle behavior in anatomically accurate geometries derived from individual CT scans. The segmentation, preprocessing, and file preparation task for a CFPD simulation was often time-consuming, and no prior studies to-date have yet presented a fully automated framework. This work presents a fully automated workflow to obtain individualized particle deposition profiles in the human respiratory tract. The pipeline starts with segmenting upper and lower airway geometries using morphological and deep learning-based methods, generating three-dimensional (3D) models from CT imaging data. Next, a series of algorithms are presented to quality check and prepare the 3D geometry for a CFD or CFPD simulation. The preprocessing step includes correcting geometric artifacts, enforcing a physically consistent mesh, and automatically identifying and capping multiple outlets, which is required for CFD/CFPD simulations. These processed models are then input into open-source (OpenFOAM) or commercial (StarCCM+) CFD solvers, where flow and transient particle transport equations — including turbulence and particle–wall interactions are solved under realistic breathing conditions. Finally, the resulting particle deposition profiles can be integrated with Monte Carlo radiation transport codes and state-of-the-art computational phantoms to assess organ-specific absorbed doses in scenarios of radioactive aerosol inhalation. The presented work streamlines respiratory tract segmentation, preprocessing for CFD/CFPD simulations, and integration with dose assessment workflows, reducing manual intervention and improving access to high-fidelity, subject-specific modeling. The high precision in predicted particle deposition and dose distributions can improve personalized treatment strategies in respiratory medicine and refine dose estimates for radiation protection.

AI↗

Dataset_for_Molecular_Motion_Below_the_Glass_Transition_A_Solid-State_NMR_Study_of_Siloxane_Polymer_Dynamics Study

This dataset contains solid-state 1H and 13C NMR relaxometry data, differential scanning calorimetry (DSC) data, and size exclusion chromatography (SEC/GPC) data supporting the study of sub-glass-transition (sub-Tg) molecular dynamics in a composition- and sequence-controlled series of diphenyl-substituted polysiloxanes (PDMS, 14Ph, 33Ph, 50Ph, 67Ph, and 100Ph; 0–100% diphenylsiloxane content by mole).All solid-state NMR data were acquired on a 200 MHz Bruker Avance III HD spectrometer using a static 7 mm HX probe or a 4 mm HX probe under 4 kHz magic-angle spinning. Raw Bruker TopSpin experiment folders are included for: (1) variable-temperature 1H lineshape measurements used to determine linewidth (FWHM) as a function of temperature across the glass transition; (2) 1H T1 (saturation recovery with solid-echo detection), probing nanosecond-scale dynamics near the 1H Larmor frequency; (3) 1H T1rho (direct spin-lock, 62.5 kHz), probing microsecond-scale segmental dynamics; (4) 13C-detected Lee–Goldburg cross-polarization 1H T1rho (LGCPH T1rho) for 33Ph and 50Ph, resolving aromatic and aliphatic proton environments; and (5) 13C T1 relaxation for 33Ph and 50Ph. Differential scanning calorimetry data (TA Instruments DSC 25, −150 to +120 °C, up to +300 °C for 100Ph, 10 °C/min) are included for all six compositions and support the glass-transition temperatures in Table 1 and Figure 1. Size exclusion chromatography data (Agilent 1200 Series, PL-Gel 300 mixed-C column, THF mobile phase, polystyrene calibration standards) are included for the three synthesized copolymers (33Ph, 50Ph, 67Ph) and support the number-average molecular weights in Table 1. Processed data include per-composition relaxation-time summaries (Excel), curve-fitting and Bloembergen-Purcell-Pound (BPP) model analysis notebooks (Jupyter/Python), and Igor Pro (.pxp) master files used to generate the manuscript's figures.

Bloembergen-Purcell-Pound theory↗

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as time series data shifts. In this research, a changepoint detection (CPD) algorithm that automatically detects data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window across the labeled data sets. Pending approval, we plan to release the labeled data sets for this research on NREL's DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package. By supplying the training sets and algorithm, we hope to encourage further development in this research space.

data shift↗