Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

ESnet Data and AI Workshop Report

In February 2025, the DOE user facility Energy Sciences Network (ESnet) held a three-day Data and AI Workshop in Berkeley, California. The objective of the workshop was to identify challenges within ESnet that could be addressed through data-driven methods, to help define ESnet’s data-analysis requirements, and to shape its AI strategy, guiding data-stewardship efforts and the direction of AI research and AIOps exploration for ESnet7, the next iteration of ESnet’s network. This report summarizes the multi-faceted discussions and findings and presents a set of recommendations for next steps.

97 MATHEMATICS AND COMPUTING↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Advancing Grid Resilience through Smart Charge Management: Findings from Maryland’s Pilot

This report presents research findings from a four-year Smart Charge Management (SCM) pilot program conducted by Maryland’s largest electric utilities—Baltimore Gas and Electric (BGE), Potomac Electric Power Company (Pepco), and Delmarva Power & Light (DPL)—to evaluate strategies for optimizing electric vehicle (EV) charging loads and enhancing grid stability. Supported by the U.S. Department of Energy (DOE), Argonne National Laboratory collaborated with all project partners and examined the effectiveness of Time-of-Use (TOU) and Load Balancing (LB) strategies in managing peak demand, deferring costly infrastructure upgrades, and reducing grid constraints at the feeder level. Using charging data from over 4,600 EV drivers, the study analyzed SCM’s impact on the distribution systems of BGE and Pepco, which consists of over 2000 feeders. Unlike prior research that focused on system-wide trends or synthetic feeders, this analysis offers granular, feeder-level insights based on real-world operational data. It highlights how transformer density, load profiles, and infrastructure constraints influence smart charging performance. Results show feeder-level conditions play a crucial role in SCM effectiveness, with most feeders benefiting more from LB, while TOU-based SCM may be sufficient for others. By 2035, LB reduced peak charging loads by 27% on average, compared to 23% under TOU-based SCM, though some feeders saw reductions exceeding 35%, while others experienced minimal impact. Feeders with higher transformer utilization and limited capacity benefited more from LB, which more effectively distributed charging demand during off-peak hours. Beyond reducing grid constraints, SCM offers long-term operational and financial benefits. By shifting EV charging demand strategically, utilities can optimize asset utilization, delay infrastructure investments, and enhance grid performance. In terms of infrastructure upgrade deferrals, at the feeder level, LB consistently reduced peak charging loads and resulting infrastructure upgrade costs, particularly in high EV enrollment areas, decreasing the number of overloaded transformers by up to 35%, while TOU-based SCM achieved 20-30% reductions depending on feeder characteristics. At the system level, LB has the potential to defer total upgrade costs by $\$$186 million for BGE, compared to $\$$159 million under TOU-based SCM. For Pepco, TOU-based SCM performed slightly better, deferring upgrade costs by $\$$30 million, compared to $\$$29 million under LB. Section 4.5 reviews some of the system differences between BGE and Pepco. However, as EV adoption scales, TOU-based SCM will introduce secondary peak charging loads, reinforcing the need for more advanced, adaptive SCM approaches to prevent new grid challenges. As EV adoption continues to grow, feeder-level managed charging strategies will be essential for mitigating grid stress, improving infrastructure efficiency, and maintaining energy affordability for consumers. This report provides critical insights for utilities, Public Utility Commissions (PUCs), and state agencies on the role of feeder-specific smart charging in infrastructure planning, policy development, and grid modernization. The findings underscore the importance of tailored, data-driven SCM solutions that align with local grid conditions, ensuring a resilient, cost-effective transition to increasing EV adoption while safeguarding distribution system performance.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Data-Driven Atomic Physics: Harnessing Machine Learning and High-Repetition-Rate Experiments for Laser-driven HED

High-energy-density plasma experiments are central to progress in atomic physics, fusion energy, and national security science, but they have traditionally been constrained by slow data collection and manual, time-intensive analysis. This project targeted that bottleneck by enabling high-repetition-rate experiments to produce and interpret much larger volumes of data quickly enough to guide experiments while they run.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Towards a Robust Adaptive Digital Twin for Fusion Applications

The development of a digital twin system for fusion applications is essential for enhancing the prediction, analysis, and optimization of complex plasma processes. Machine learning (ML), particularly deep learning has demonstrated strong capabilities in modeling such highly nonlinear and intricate systems. However, two critical challenges limit the deployment of deep learning-based digital twins: Uncertainty Quantification (UQ) and data drift. UQ is vital for ensuring trustworthy predictions, especially in decision-support scenarios. Additionally, data-driven models are often sensitive to changes in the underlying data distribution, such as shot-to-shot variations in fusion experiments, which can lead to performance degradation over time. To address these challenges, we are developing an uncertainty-aware, adaptive digital twin framework. Our approach incorporates deep learning models enhanced with Gaussian Process approximations for predictive uncertainty estimation, coupled with an online learning mechanism that enables continuous model adaptation to new experimental data. This adaptive capability allows the data driven models to respond effectively to evolving plasma behaviors and equipment conditions. Specifically, to mitigate the effects of shot-to-shot drift, our system updates itself incrementally as new data becomes available, improving both robustness and fidelity. Our vision is to evolve this data driven model into a self-sustaining digital twin system that leverages UQ based feedback to continuously refine itself and potentially support real-time decision making. This presentation will cover a brief background on uncertainty quantification for ML, our ongoing effort on development of UQ capabilities for ML, our data science pipeline from data collection to model development and analysis and online learning framework for modeling coil deflection at DIII-D. I will also briefly touch upon opportunities and challenges in development of digital twin framework.

Sammuli, Brian [General Atomics]↗

Meeting Global Health Needs via Infectious Disease Forecasting: Development of a Reliable Data-Driven Framework

Infectious diseases (IDs) have a significant detrimental impact on global health. Timely and accurate ID forecasting can result in more informed implementation of control measures and prevention policies. To meet the operational decision-making needs of real-world circumstances, we aimed to build a standardized, reliable, and trustworthy ID forecasting pipeline and visualization dashboard that is generalizable across a wide range of modeling techniques, IDs, and global locations. We forecasted 6 diverse, zoonotic diseases (brucellosis, campylobacteriosis, Middle East respiratory syndrome, Q fever, tick-borne encephalitis, and tularemia) across 4 continents and 8 countries. We included a wide range of statistical, machine learning, and deep learning models (n=9) and trained them on a multitude of features (average n=2326) within the One Health landscape, including demography, landscape, climate, and socioeconomic factors. The pipeline and dashboard were created in consideration of crucial operational metrics—prediction accuracy, computational efficiency, spatiotemporal generalizability, uncertainty quantification, and interpretability—which are essential to strategic data-driven decisions. While no single best model was suitable for all disease, region, and country combinations, our ensemble technique selects the best-performing model for each given scenario to achieve the closest prediction. For new or emerging diseases in a region, the ensemble model can predict how the disease may behave in the new region using a pretrained model from a similar region with a history of that disease. The data visualization dashboard provides a clean interface of important analytical metrics, such as ID temporal patterns, forecasts, prediction uncertainties, and model feature importance across all geographic locations and disease combinations. As the need for real-time, operational ID forecasting capabilities increases, this standardized and automated platform for data collection, analysis, and reporting is a major step forward in enabling evidence-based public health decisions and policies for the prevention and mitigation of future ID outbreaks.

60 APPLIED LIFE SCIENCES↗

Comparing Outdoor to Indoor Performance for Bifacial Modules Affected by Polarization-Type Potential-Induced Degradation

Bifacial photovoltaic (PV) modules have the advantage of using light reflected off of the ground to contribute to power production. Predicting the energy gain is challenging and requires complex models to do so accurately. Often, module degradation over time is neglected in models for the sake of simplicity or is underestimated. Comparing outdoor and indoor current–voltage (I–V) performance for bifacial modules is more challenging than for monofacial modules, as there are additional variables to consider such as rear albedo non-uniformity, cell mismatch, and their effects on temperature. This challenge is compounded when heterogeneous degradation modes occur, such as polarization-type potential-induced degradation (PID-p). To examine the effects of PID-p on I–V predictions using an empirical data-driven approach, 16 bifacial PERC modules are installed outdoors on racks with different albedo conditions. A subset is exposed to high-voltage biases of −1500 V or +1500 V. Outdoor data are traced at irradiance ranges of 150–250 W/m 2 , 500–600 W/m 2 , and 900–1000 W/m 2 . These curves are corrected using control module temperature, wire resistivity, and module resistance measured indoors. We examine several methods to transform indoor I–V curves to accurately, and more simply than existing methods, approximate outdoor performance for bifacial modules without and with varying levels of PID-p degradation. This way, bifacial performance modeling can be more accessible and informed by fielded, degraded modules. Distributions of percent errors between indoor and outdoor performance parameters and Mean Absolute Percent Errors (MAPEs) are used to assess method quality. Results including low-irradiance data (150–250 W/m 2 ) are discussed but are filtered for quantifying method quality as these data introduce substantial errors. The method with the most optimal tradeoff between low MAPE and analysis simplicity involves measuring the front side of a module indoors at an irradiance equal to plane-of-array irradiance plus the product of module bifaciality and albedo irradiance. This method gives MAPE values of 1–6.5% for non-degraded and 1.6–5.9% for PID-p degraded module performance.

14 SOLAR ENERGY↗

Creating Accurate Methane Emission Inventories through Data-Driven Airborne Survey Strategies

Because natural gas emits less carbon than other fossil fuels, it holds promise as a green energy transition fuel. However, the overall carbon footprint of natural gas is significantly elevated by methane emissions that occur during its production and transmission (Cusworth et al. 2022). Methane “super-emitters,” while comprising only about 1% of sites, are responsible for the majority of oil- and gas-sourced methane emissions, making their detection and mitigation critical in reducing the climate impact of natural gas and in meeting national and global sustainability goals (Sherwin et al. 2024). Yet, despite advancements in detection, significant uncertainties remain regarding the size, frequency, and duration distributions of methane emissions (e.g., Frankenberg et al. 2016, Cusworth et al. 2022, Chen, Sherwin et al. 2022, Conrad et al. 2023, Johnson et al. 2023, Sherwin et al. 2024) underscoring the need for comprehensive emissions inventories segmented by basin across the US. Airborne surveys are well-suited for collecting data to build these comprehensive, basin-level inventories because they allow for extensive spatial coverage, and have the spatial resolution, and the sensitivity to pinpoint individual methane sources. As remote sensing technologies enable rapid basin-scale surveys, it is imperative to establish scientifically and statistically robust standards to generate reliable and actionable emissions inventories. Recent work has shown that differences in airborne sampling strategies, detection technologies, and analysis can lead to large differences between survey conclusions if not correctly accounted for (Chen et al. 2024). This elevates the importance of incorporating proper sampling and analysis techniques when designing a methane emissions monitoring campaign to produce accurate results and facilitate cross-study comparisons. In this paper, we describe a survey strategy designed using the latest conclusions from the literature to align results from different aerial surveys. We identify several sampling and analysis principles, including large sample sizes, balanced sampling across oil and gas production, careful survey area definition, and a unified protocol for analysis, to be vital to producing an unbiased estimate of basin-scale emissions. We present results from a Department of Energy-funded project that deployed this survey strategy in two understudied oil and gas- producing regions in the United States: the Haynesville Basin in Texas and Louisiana, and the Woodford Shale in the Anadarko Basin in Oklahoma.

03 NATURAL GAS↗

Equation-Free Coarse Control of Distributed Parameter Systems via Local Neural Operators

The control of high-dimensional distributed parameter systems (DPS) remains a challenge when explicit coarse-grained equations are unavailable. Classical equation-free (EF) approaches rely on fine-scale simulators treated as black-box timesteppers. However, repeated simulations for steady-state computation, linearization, and control design are often computationally prohibitive, or the microscopic timestepper may not even be available, leaving us with data as the only resource. We propose a data-driven alternative that uses local neural operators, trained on spatiotemporal microscopic/mesoscopic data, to obtain efficient short-time solution operators. These surrogates are employed within Krylov subspace methods to compute coarse steady and unsteady-states, while also providing Jacobian information in a matrix-free manner. Krylov-Arnoldi iterations then approximate the dominant eigenspectrum, yielding reduced models that capture the open-loop slow dynamics without explicit Jacobian assembly. Both discrete-time Linear Quadratic Regulator (dLQR) and pole-placement (PP) controllers are based on this reduced system and lifted back to the full nonlinear dynamics, thereby closing the feedback loop.

93B52, 93C20, 47N70, 65J15, 65M32, 68T07, 68T20, 6↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs↗

Verification, Validation, and Calibration Through a Causal Lens

While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

New U.S. Data Tools are Playing a Crucial Role in Decarbonizing Buildings at Speed, Scale, and Low Cost

Preparing buildings for retrofits traditionally requires expensive on-site audits or timeintensive simulation models. As a result, the majority of buildings fail to pursue cost-saving retrofits. To address these barriers, the U.S. Department of Energy (DOE) has introduced the Building Efficiency Targeting Tool for Energy Retrofits (BETTER)-a new, free, on-line tool that utilizes a data-driven analytical engine and user-friendly web interface to automatically analyze a building's monthly energy usage in response to weather conditions. The tool benchmarks a building's electric and fossil energy usage against peers; estimates energy, cost, and emissions reductions at the building and portfolio levels; recommends energy efficiency measures; and prioritizes buildings for net-zero energy retrofits. Thanks to interoperability with the DOE's Standard Energy Efficiency Data (SEED) platform, BETTER is supporting U.S. jurisdictions to prepare buildings for retrofit at speed, scale, and low cost to comply with energy policies. This paper discusses the use of BETTER and SEED by one of the branches of the California state government to streamline a retrofit program across 455 public non-residential buildings to align with state goals to reduce greenhouse gas emissions. It describes the organization's challenge to reduce energy consumption across a geographically diverse, aging portfolio; explores how BETTER and SEED improved workflow efficiency; presents preliminary results, including avoiding audit costs of $3.28 million and developing the groundwork for retrofit projects estimated to prevent emission of 2,271 t CO2e annually; and provides guidance for other jurisdictions seeking similar results.

BETTER↗

Invertible Temper Modeling using Normalizing Flows and the Effects of Structure Preserving Loss

Advanced manufacturing research and development is typically small-scale, owing to costly experiments associated with these novel processes. Deep learning techniques could help accelerate this development cycle but frequently struggle in small-data regimes like the advanced manufacturing space. While prior work has applied deep learning to modeling visually plausible advanced manufacturing microstructures, little work has been done on data-driven modeling of how microstructures are affected by heat treatment, or assessing the degree to which synthetic microstructures are able to support existing workflows. We propose to address this gap by using invertible neural networks (normalizing flows) to model the effects of heat treatment, e.g., tempering. The model is developed using scanning electron microscope imagery from samples produced using shear-assisted processing and extrusion (ShAPE) manufacturing. This approach not only produces visually and topologically plausible samples, but also captures information related to a sample’s material properties or experimental process parameters. We also demonstrate that topological data analysis, used in prior work to characterize microstructures, can also be used to stabilize model training, preserve structure, and improve downstream results. We assess directions for future work and identify our approach as an important step towards end-to-end deep learning system for accelerating advanced manufacturing research and development.

Howland, Sylvia↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Data-driven modeling of dynamic occupant thermostat override behavior for demand response applications

Buildings consume nearly 40% of global energy and produce similar emissions. Whiletechnological advances address efficiency, occupant behavior causes energy use variations up to 300% between identical buildings. This gap between predicted and actual building performance impacts building design, operations, and grid demand management programs. Through analyses of smart thermostat data from 1,400 single-occupant homes, the researchdemonstrates that occupants respond to 8°F thermostat setpoint changes within a median of 15 minutes, while 2°F changes trigger responses within a median of 30 minutes. This highlights an understudied temporal relationship between thermostat setbacks and response time of occupant behaviors. Models of such behavior dynamics are required to incorporate occupant impacts into building performance simulation. A key contribution of this dissertation is the Thermal Frustration Theory (TFT), which positsthat thermal discomfort driven behaviors are caused by the time-accumulation of discomfort, not simply a temperature deviation threshold or a delay from an initiating event. Using a dataset of 634 thermostats, each with 25+ manual setpoint changes, a comparative analysis of TFT and comfort zone and a delayed response theories demonstrated that personalized TFT models better predict when manual setpoint change occur. This was measured by the area under the curve statistical measure (AUC); all three models perform similarly by a Matthews Correlation Coefficient measure. Higher AUC performance is especially important for modeling occupant behavior in demand response programs where false negatives of rare occupant interactions could adversely affect grid stability. EnergyPlus based simulations were conducted with TFT-derived occupant models, demonstrating the ability to identify parameters of known TFT models from only data observable with smart thermostats, even under the presence of noise from routine overrides. Overall, the dissertation highlights that thermostat interactions are neither static,instantaneous, nor driven solely by the environment. Instead, temporal accumulation of discomfort and routine-based behavior play important roles. The methodology and results offer a pathway towards more accurate modeling of human-building interactions for policy assessment, building design, and demand response programs.

Sharma, Kunind [Northeastern University] (ORCID:00↗

Excited-state uncertainties in lattice-QCD calculations of multi-hadron systems

Excited-state effects lead to hard-to-quantify systematic uncertainties in lattice quantum chromodynamics (LQCD) spectroscopy calculations when computationally accessible imaginary times are smaller than inverse excitation gaps, as often arises for multi-hadron systems with signal-to-noise problems. Lanczos residual bounds address this by providing two-sided constraints on energies that do not require assumptions beyond Hermiticity, but often give very conservative systematic uncertainty estimates. Here, a more-constraining set of gap bounds is introduced for hadron spectroscopy. These bounds provide tighter constraints whose validity requires an explicit assumption about an energy gap. Exactly solvable lattice field theory correlators are used to test the utility of residual and gap bounds at finite and infinite statistics. Two-sided bounds and other analysis methods are then applied to a high-statistics LQCD calculation of nucleon-nucleon scattering at $m_π\sim 800$ MeV. Generalized eigenvalue problem (GEVP) and Lanczos energy estimators are compatible when applied to the same correlator data, but analyses including different interpolating operators show statistically significant inconsistencies. However, two-sided bounds from all operators are consistent. Under the assumption that the number of energy levels below $NΔ$ and $ΔΔ$ thresholds is the same as for non-interacting nucleons, gap bounds are sufficient to constrain nucleon-nucleon scattering amplitudes at phenomenologically relevant precision. Lanczos methods further reveal that energy-eigenstate estimates from previously studied asymmetric correlators have not converged over accessible imaginary times. Nevertheless, data-driven examples demonstrate why assumptions are required to draw conclusions about the natures of two-nucleon ground states at these masses.

Detmold, William [MIT, Cambridge, CTP]↗