Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Feature Based Qualification of 17-4PH Stainless Steel to Evaluate Location-Specific Variability in Wire Arc Additive Manufacturing

Qualifying large-scale metal additive manufacturing (M-AM) technologies such as wire arc additive manufacturing (WAAM) can be challenging. This is especially significant in precipitation hardened martensitic stainless steels like SS 17-4PH, where thermal histories induce location-specific microstructural variability and property anisotropy. The Department of Defense (DOD) and the United States Army Combat Capabilities Development Command Ground Vehicle Systems Center (GVSC) Ground Vehicle Materials Engineering (GVME) aim to build robust and qualified large-scale M-AM workflows that could reduce the time and cost through quick and informed evaluation, testing, and development of feedstock, processes, and parts. The report presents the findings from the collaborative efforts between Oak Ridge National Laboratory (ORNL) and the U.S. Army GVSC GVME. The aim of this project was to develop a geometric feature-based qualification framework for WAAM of SS 17-4PH components. This report outlines selection methodology of representative build geometries, optimization of WAAM process parameters, in-situ monitoring, microstructure-property evaluation, thermal simulations, as well as data visualization techniques incorporated in this project. The results from this project demonstrate a clear understanding of thermal history dependent phase evolution and consequent location-specific property variations in WAAM of SS 17-4PH. These results in conjunction with the data-driven methodologies used in this project are expected to reduce qualification timelines, improve predictability, and accelerate the development of reliable feature-based qualification strategies for part production via large-scale M-AM technologies.

36 MATERIALS SCIENCE↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Towards a Robust Adaptive Digital Twin for Fusion Applications

The development of a digital twin system for fusion applications is essential for enhancing the prediction, analysis, and optimization of complex plasma processes. Machine learning (ML), particularly deep learning has demonstrated strong capabilities in modeling such highly nonlinear and intricate systems. However, two critical challenges limit the deployment of deep learning-based digital twins: Uncertainty Quantification (UQ) and data drift. UQ is vital for ensuring trustworthy predictions, especially in decision-support scenarios. Additionally, data-driven models are often sensitive to changes in the underlying data distribution, such as shot-to-shot variations in fusion experiments, which can lead to performance degradation over time. To address these challenges, we are developing an uncertainty-aware, adaptive digital twin framework. Our approach incorporates deep learning models enhanced with Gaussian Process approximations for predictive uncertainty estimation, coupled with an online learning mechanism that enables continuous model adaptation to new experimental data. This adaptive capability allows the data driven models to respond effectively to evolving plasma behaviors and equipment conditions. Specifically, to mitigate the effects of shot-to-shot drift, our system updates itself incrementally as new data becomes available, improving both robustness and fidelity. Our vision is to evolve this data driven model into a self-sustaining digital twin system that leverages UQ based feedback to continuously refine itself and potentially support real-time decision making. This presentation will cover a brief background on uncertainty quantification for ML, our ongoing effort on development of UQ capabilities for ML, our data science pipeline from data collection to model development and analysis and online learning framework for modeling coil deflection at DIII-D. I will also briefly touch upon opportunities and challenges in development of digital twin framework.

Sammuli, Brian [General Atomics]↗

Comparing Outdoor to Indoor Performance for Bifacial Modules Affected by Polarization-Type Potential-Induced Degradation

Bifacial photovoltaic (PV) modules have the advantage of using light reflected off of the ground to contribute to power production. Predicting the energy gain is challenging and requires complex models to do so accurately. Often, module degradation over time is neglected in models for the sake of simplicity or is underestimated. Comparing outdoor and indoor current–voltage (I–V) performance for bifacial modules is more challenging than for monofacial modules, as there are additional variables to consider such as rear albedo non-uniformity, cell mismatch, and their effects on temperature. This challenge is compounded when heterogeneous degradation modes occur, such as polarization-type potential-induced degradation (PID-p). To examine the effects of PID-p on I–V predictions using an empirical data-driven approach, 16 bifacial PERC modules are installed outdoors on racks with different albedo conditions. A subset is exposed to high-voltage biases of −1500 V or +1500 V. Outdoor data are traced at irradiance ranges of 150–250 W/m 2 , 500–600 W/m 2 , and 900–1000 W/m 2 . These curves are corrected using control module temperature, wire resistivity, and module resistance measured indoors. We examine several methods to transform indoor I–V curves to accurately, and more simply than existing methods, approximate outdoor performance for bifacial modules without and with varying levels of PID-p degradation. This way, bifacial performance modeling can be more accessible and informed by fielded, degraded modules. Distributions of percent errors between indoor and outdoor performance parameters and Mean Absolute Percent Errors (MAPEs) are used to assess method quality. Results including low-irradiance data (150–250 W/m 2 ) are discussed but are filtered for quantifying method quality as these data introduce substantial errors. The method with the most optimal tradeoff between low MAPE and analysis simplicity involves measuring the front side of a module indoors at an irradiance equal to plane-of-array irradiance plus the product of module bifaciality and albedo irradiance. This method gives MAPE values of 1–6.5% for non-degraded and 1.6–5.9% for PID-p degraded module performance.

14 SOLAR ENERGY↗

Equation-Free Coarse Control of Distributed Parameter Systems via Local Neural Operators

The control of high-dimensional distributed parameter systems (DPS) remains a challenge when explicit coarse-grained equations are unavailable. Classical equation-free (EF) approaches rely on fine-scale simulators treated as black-box timesteppers. However, repeated simulations for steady-state computation, linearization, and control design are often computationally prohibitive, or the microscopic timestepper may not even be available, leaving us with data as the only resource. We propose a data-driven alternative that uses local neural operators, trained on spatiotemporal microscopic/mesoscopic data, to obtain efficient short-time solution operators. These surrogates are employed within Krylov subspace methods to compute coarse steady and unsteady-states, while also providing Jacobian information in a matrix-free manner. Krylov-Arnoldi iterations then approximate the dominant eigenspectrum, yielding reduced models that capture the open-loop slow dynamics without explicit Jacobian assembly. Both discrete-time Linear Quadratic Regulator (dLQR) and pole-placement (PP) controllers are based on this reduced system and lifted back to the full nonlinear dynamics, thereby closing the feedback loop.

93B52, 93C20, 47N70, 65J15, 65M32, 68T07, 68T20, 6↗

Estimation and Bias Correction of Aerosol Abundance using Data-driven Machine Learning and Remote Sensing

Air quality information is increasingly becoming a public health concern, since some of the aerosol particles pose harmful effects to peoples health. One widely available metric of aerosol abundance is the aerosol optical depth (AOD). The AOD is the integrated light extinction coefficient over a vertical atmospheric column of unit cross section, which represents the extent to which the aerosols in that vertical profile prevent the transmission of light by absorption or scattering. The comparison between the AOD measured from the ground-based Aerosol Robotic Network (AERONET) system and the satellite MODIS instruments at 550 nm shows that there is a bias between the two data products. We performed a comprehensive analysis exploring possible factors which may be contributing to the inter-instrumental bias between MODIS and AERONET. The analysis used several measured variables, including the MODIS AOD, as input in order to train a neural network in regression mode to predict the AERONET AOD values. This not only allowed us to obtain an estimate, but also allowed us to infer the optimal sets of variables that played an important role in the prediction. In addition, we applied machine learning to infer the global abundance of ground level PM2.5 from the AOD data and other ancillary satellite and meteorology products. This research is part of our goal to provide air quality information, which can also be useful for global epidemiology studies.

Malakar, Nabin K.↗

Human Performance Contributions to Safety in Commercial Aviation

Every day in aviation, pilots, air traffic controllers, and other front-line personnel perform countless correct judgments and actions in a variety of operational environments. These judgments and actions are often the difference between an accident and a non-event. Ironically, data on these behaviors are rarely collected or analyzed. Data-driven decisions about safety management and design of safety-critical systems are limited by the available data, which influence how decision makers characterize problems and identify solutions. Large volumes of data are collected on the failures and errors that result in infrequent incidents and accidents, but in the absence of data on behaviors that result in routine successful outcomes, safety management and system design decisions are based on a small sample of nonrepresentative safety data. This assessment aimed to find and document “safety successes” made possible by human operators. With many Aeronautics Research Mission Directorate (ARMD) Programs and Projects focusing on increased automation and autonomy and decreased human involvement, failure to fully consider the human contributions to successful system performance in civil aviation represents a significant risk — a risk that has not been recognized to date. Without understanding how humans contribute to safety, any estimate of predicted safety of autonomous capabilities is incomplete and inherently suspect. Furthermore, understanding the ways in which humans contribute to safety can promote strategic interactions among safety technologies, functions, procedures and the people using them. Without this understanding, the full benefits of an integrated, optimized human/technology or autonomous system will not be realized. Historically, safety has been consistently defined in terms of the occurrence of accidents or recognized risks (i.e., in terms of things that go wrong). These adverse outcomes are explained by identifying their causes, and safety is restored by eliminating or mitigating these causes. An alternative to this approach is to focus on what goes right and identify how to replicate that process. Focusing on the rare cases of failures attributed to “human error” provides little information about why human performance routinely prevents adverse events. Hollnagel has proposed that things go right because people continuously adjust their work to match their operating conditions. These adjustments become increasingly important as systems continue to grow in complexity. Thus, the definition of safety should reflect not only “avoiding things that go wrong” but “ensuring that things go right.” The basis for safety management requires developing an understanding of everyday activities. However, few mechanisms to monitor everyday work exist in the aviation domain, which limits opportunities to learn how designs function in reality. This concept of safety thinking and safety management is reflected in the emerging field of resilience engineering. According to Hollnagel, a system is resilient if it can sustain required operations under expected and unexpected conditions by adjusting its functioning prior to, during, or following changes, disturbances, and opportunities. To explore “positive” behaviors that contribute to resilient performance in commercial aviation, the assessment team examined a range of existing sources of data about pilot and air traffic control (ATC) tower controller performance, including subjective interviews with domain experts and objective aircraft flight data records. These data were used to identify strategies that support resilient performance, methods for exploring and refining those strategies in existing data, and proposed methods for capturing and analyzing new data.

Null, Cynthia H.↗

Prediction of Pushback Times and Ramp Taxi Times for Departures at Charlotte Airport

When optimizing the takeoff sequence and schedule for departures at busy airports, it is important to accurately predict the taxi times from gate to runway because those are used to calculate the earliest possible takeoff times. Several airports like Charlotte Douglas International Airport show relatively long taxi times inside the ramp area with large variations, with respect to the travel times in the airport movement area. Also, the pushback process times have not been accurately modeled so far mainly due to the lack of accurate data. The recent deployment of the integrated arrival, departure, and surface traffic management system at Charlotte airport by NASA enables more accurate flight data in the airport surface operations to be obtained. Taking advantage of this system, actual pushback times and ramp taxi times from historical flight data at this airport are analyzed. Based on the analysis, a simple, data-driven prediction model is introduced for estimating pushback times and ramp transit times of individual departure flights. To evaluate the performance of this prediction model, several machine learning techniques are also applied to the same dataset. The prediction results show that the data-driven prediction model is as good as the machine learning algorithms when comparing various prediction performance metrics.

Lee, Hanbong↗

Prediction of Pushback Times and Ramp Taxi Times for Departures at Charlotte Airport

When optimizing the takeoff sequence and schedule for departures at busy airports, it is important to accurately predict the taxi times from gate to runway because those are used to calculate the earliest possible takeoff times. Several airports like Charlotte Douglas International Airport show relatively long taxi times inside the ramp area with large variations, with respect to the travel times in the airport movement area. Also, the pushback process times have not been accurately modeled so far mainly due to the lack of accurate data. The recent deployment of the integrated arrival, departure, and surface traffic management system at Charlotte airport by NASA enables more accurate flight data in the airport surface operations to be obtained. Taking advantage of this system, actual pushback times and ramp taxi times from historical flight data at this airport are analyzed. Based on the analysis, a simple, data-driven prediction model is introduced for estimating pushback times and ramp transit times of individual departure flights. To evaluate the performance of this prediction model, several machine learning techniques are also applied to the same dataset. The prediction results show that the data-driven prediction model is as good as the machine learning algorithms when comparing various prediction performance metrics.

airport surface operations↗

Object-Based Comparison of Data-Driven and Physics-Driven Satellite Estimates of Extreme Rainfall

The Global Precipitation Measurement (GPM) constellation of spaceborne sensors provides a variety of direct and indirect measurements of precipitation processes. Such observations can be employed to derive spatially and temporally consistent gridded precipitation estimates either via data-driven retrieval algorithms or by assimilation into physically based numerical weather models. We compare the data-driven Integrated Multisatellite Retrievals for GPM (IMERG) and the assimilation-enabled NASA-Unified Weather Research and Forecasting (NU-WRF) model against Stage IV reference precipitation for four major extreme rainfall events in the southeastern United States using an object-based analysis framework that decomposes gridded precipitation fields into storm objects. As an alternative to conventional ‘‘grid-by-grid analysis,’’ the object-based approach provides a promising way to diagnose spatial properties of storms, trace them through space and time, and connect their accuracy to storm types and input data sources. The evolution of two tropical cyclones are generally captured by IMERG and NU-WRF, while the less organized spatial patterns of two mesoscale convective systems pose challenges for both. NU-WRF rain rates are generally more accurate, while IMERG better captures storm location and shape. Both show higher skill in detecting large, intense storms compared to smaller, weaker storms. IMERG’s accuracy depends on the input microwave and infrared data sources; NU-WRF does not appear to exhibit this dependence. Findings highlight that an object-oriented view can provide deeper insights into satellite precipitation performance and that the satellite precipitation community should further explore the potential for ‘‘hybrid’’ data-driven and physics-driven estimates in order to make optimal usage of satellite observations.

extreme events↗

Natural Language Processing Analysis of Notices to Airmen for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized.

Natural Language Processing↗

Natural Language Processing (NLP) Analysis of NOTAMs for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.

Natural Language Processing↗

NASA’s Digital Information Platform to Accelerate the Transformation of the National Airspace System

In order to accelerate the digital transformation of airspace operations, a foundational framework and infrastructure for providing sustainable, data-driven, and cohesive decision-making digital services for both traditional and emergent air vehicles is being developed. The reference implementation of Digital Information Platform builds an ecosystem for the aviation community by providing access to a secure and trusted source of aviation data and services. Several key features and services have been implemented to enable secure data sharing, communication, and service registration on the Platform. The technical approach used to implement these features is presented here. NASA-developed integrated aviation data and machine learning based prediction services to optimize airspace operations are available on the Platform. These services are being evaluated in an operational environment by flight operators and the real-world benefits are being captured. The Platform fosters collaboration among industry and researchers to develop complex aviation services and the aim is to make it publicly accessible for consumption by the aviation community.

Digital Transformation↗

NASA’s Digital Information Platform to Accelerate the Transformation of the National Airspace System

In order to accelerate the digital transformation of airspace operations, a foundational framework and infrastructure for providing sustainable, data-driven, and cohesive decision-making digital services for both traditional and emergent air vehicles is being developed. The reference implementation of Digital Information Platform builds an ecosystem for the aviation community by providing access to a secure and trusted source of aviation data and services. Several key features and services have been implemented to enable secure data sharing, communication, and service registration on the Platform. The technical approach used to implement these features is presented here. NASA-developed integrated aviation data and machine learning based prediction services to optimize airspace operations are available on the Platform. These services are being evaluated in an operational environment by flight operators and the real-world benefits are being captured. The Platform fosters collaboration among industry and researchers to develop complex aviation services and the aim is to make it publicly accessible for consumption by the aviation community.

Digital Transformation↗

Improving the Fidelity of Capability & Resource Weighting in A Probalistic Risk Assessment Model for Spaceflight

INTRODUCTION NASA’s Informing Mission Planning via Analysis of Complex Tradespaces (IMPACT) tool uses Probabilistic Risk Assessment (PRA) to provide an evidence-based, data-driven estimate of how medical system capabilities affect mission outcomes. IMPACT maps condition incidence to available resources thereby facilitating the calculation of outcome metrics that allow the estimation of mission medical risk. Conditions can be nominally categorized as either treated or untreated depending on the availability of necessary diagnostic and therapeutic capabilities. This categorization enables IMPACT to estimate the effect of various medical system configurations on mission outcomes such as crew mortality, disability, crew member down-time, return to duty/recovery, and need for evacuation. IMPACT currently employs an equal weighting, “partial credit” approach to define treatment in which each of the capabilities associated with a given condition contributes an equal amount to management of the condition. This feature enables IMPACT to report values in between “fully untreated” and “fully treated” based on the proportion of capabilities available within the model. However, as is normal in medical/clinical practice, not all individual capabilities contribute equally to medical care. For example, the ability to provide intramuscular epinephrine during an anaphylactic episode contributes more likelihood of overall management success than does the administration of oral diphenhydramine. We hypothesize that weighting the relative contribution of each capability to each specific condition will improve outcome prediction and therefore will better provide mission planners with more nuanced and accurate options when designing space medical systems. METHODS Using a five-point Fibonacci scaling sequence (1, 2, 3, 5, 8) subject matter experts from NASA’s Exploration Medical Capabilities (ExMC) element assigned relative contribution weighting values to each identified capability within IMPACT. Since the relative importance of each capability varies depending on the specific condition, the resulting “partial” weighting was completed for more than 1,600 individual weighting assignments for 666 capabilities across 121 conditions. Each assignment required three-physician concurrence based on the overall importance of the capability to the diagnosis and management of the condition being considered and the difficulty with which it could be improvised by the crew. Once complete, 100,000 IMPACT simulations were run for a 6-month Lunar mission with a 30-day surface stay to evaluate the effect of this modification of the model. RESULTS Partial weighting significantly decreased predicted task time loss (TTL), evacuation, and loss of crew life without causing significant changes to the recommended medical system design. CONCLUSIONS The paucity of real-world referent data to support long-duration space missions of this type limits the ability to judge one predictive analytics method as superior to another. However, since the proposed method significantly reduces and optimizes outcome risks—without changing the medical system design—incorporating a partial weighting methodology is likely to provide a more accurate and operationally-relevant representation of medical risk without compromising IMPACTs ability to inform overarching medical system requirements.

Steller JG↗

Planning Bias: Planning as a Source of Sampling Bias

Many data-driven planning methods are trained on data generated by planners. It is well known that many statistical learning methods are sensitive to sampling bias, and yet there has been little or no attention to planning as a sampling method and its role in introducing sampling bias into planner-generated training data. Recently, it has been demonstrated that A**,* in the presence of problems with variable heuristic error, prefers some solutions over other equally cost-optimal solutions. But, as we discuss in this paper, mitigation may not be as simple as resolving arbitrary tie-breaking by sampling from ties uniformly at random. In this paper, we formalize an intuition of planning bias. We focus on problems which output a single solution. Diverse planning only complicates the problem by generalizing it to bias in the set of sets; we show how it is subject to bias in the single solution. We make some useful observations about deterministic algorithms in contrast to non-deterministic algorithms. We explain how information entropy may be a good way to measure planning bias, and discuss some issues in evaluating practical approaches to measurement. We address the intuition that uniform random tiebreaking should mitigate bias; and sketch a novel approach to constructing an appropriate random distribution for duplicate detection during forward search for unbiased A*. Finally, we suggest directions for future work.

Planning Scheduling Algorithms↗

EVs@Scale Next-Gen Profiles - Fleet Utilization 2024

As part of the U.S. Department of Energy’s EVs@Scale initiative, the Next-Gen Profiles (NGP) project provides a comprehensive, data-driven analysis of electric vehicle (EV) and electric vehicle supply equipment (EVSE) operations across real-world fleet deployments. This paper presents findings from the NGP’s Fleet Utilization study, which investigates operational behavior and asset usage across seventeen EV fleets and two EVSE fleets, encompassing a wide range of vehicle types and use cases. Data collected from diverse sources—varying in format and temporal resolution—are first reformatted into a unified structure. From this harmonized dataset, a suite of rigorously defined performance metrics is calculated at an hourly cadence, enabling consistent cross-comparison of charging, routing, and other key operational behaviors. Amid rapidly increasing EV adoption and growing demands for energy-efficient fleet operations, the analysis reveals clear utilization trends—including diurnal and weekly activity cycles, differences in short versus long charging session dependencies, and route-specific energy usage patterns. These findings highlight the need for tailored infrastructure strategies and the deployment of advanced energy management systems, such as Distributed Energy Resource Management Systems (DERMS) and Site Energy Management Systems (SEMS), which can optimize charging schedules and mitigate peak loads. By leveraging anonymized, harmonized datasets and standardized metrics, this study offers critical insights into fleet behavior and performance, providing a foundation to improve operational efficiency, reduce costs, and enable the scalable deployment of electrified transportation.

Wells, Landon↗

Individual Data Sparsity in Smart Thermostat Big Data: Impacts on Modeling Thermostat Use Behavior Dynamics

This study explores the impacts of the sparsity of individual thermostat interaction data on modeling thermostat use behavior dynamics using a dataset of over 100,000 smart thermostats. In developing a data-driven model of Thermal Frustration Theory (TFT), we investigate the challenges and trade-offs in clustering occupant data to enhance predictive accuracy. Our findings reveal that a single, aggregated model fails to capture the diversity of occupant behaviors, resulting in extremely poor prediction performance. Conversely, excessive clustering exacerbates data sparsity, undermining model reliability. By identifying an optimal clustering strategy, we achieve a balance that significantly improves the prediction of manual setpoint changes during demand response (DR) events, enhancing energy management and occupant comfort

Fannon, David↗