Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-driven method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Particle Filtering for Model-Based Anomaly Detection in Sensor Networks

A novel technique has been developed for anomaly detection of rocket engine test stand (RETS) data. The objective was to develop a system that postprocesses a csv file containing the sensor readings and activities (time-series) from a rocket engine test, and detects any anomalies that might have occurred during the test. The output consists of the names of the sensors that show anomalous behavior, and the start and end time of each anomaly. In order to reduce the involvement of domain experts significantly, several data-driven approaches have been proposed where models are automatically acquired from the data, thus bypassing the cost and effort of building system models. Many supervised learning methods can efficiently learn operational and fault models, given large amounts of both nominal and fault data. However, for domains such as RETS data, the amount of anomalous data that is actually available is relatively small, making most supervised learning methods rather ineffective, and in general met with limited success in anomaly detection. The fundamental problem with existing approaches is that they assume that the data are iid, i.e., independent and identically distributed, which is violated in typical RETS data. None of these techniques naturally exploit the temporal information inherent in time series data from the sensor networks. There are correlations among the sensor readings, not only at the same time, but also across time. However, these approaches have not explicitly identified and exploited such correlations. Given these limitations of model-free methods, there has been renewed interest in model-based methods, specifically graphical methods that explicitly reason temporally. The Gaussian Mixture Model (GMM) in a Linear Dynamic System approach assumes that the multi-dimensional test data is a mixture of multi-variate Gaussians, and fits a given number of Gaussian clusters with the help of the wellknown Expectation Maximization (EM) algorithm. The parameters thus learned are used for calculating the joint distribution of the observations. However, this GMM assumption is essentially an approximation and signals the potential viability of non-parametric density estimators. This is the key idea underlying the new approach.

Solano, Wanda↗

Towards a State Based Control Architecture for Large Telescopes: Laying a Foundation at the VLT

Large telescopes are characterized by a high level of distribution of control-related tasks and will feature diverse data flow patterns and large ranges of sampling frequencies; there will often be no single, fixed server-client relationship between the control tasks. the architecture is also challenged by the task of integrating heterogeneous subsystems which will be delivered by multiple different contractors. Due to the high number of distributed components, the control system needs to effectively detect errors and faults, impede their propagation, and accurately mitigate them in the shortest time possible, enabling the service to be restored. The presented Data-Driven Architecture is based on a decentralized approach with an end-to-end integration of disparate, independently developed software components. These components employ a high-performance standards-based communication middle-ware infrastructure, based on the Data Distribution Service. A set of rules and principles, based on JPL's State Analysis method and architecture, are use to constrain component-to component interactions, where the Control System and System Under Control are clearly separated. State Analysis provide a model-based process for capturing system and software requirements and design, greatly reducing the gap between the requirements on software specified by systems engineers and the implementation by software engineers. The method and architecture has been field tested at the Very Large Telescope, where it has been integrated into an operational system.

European Extremely Large Telescope (E-ELT)↗

A Satellite Data-Driven, Client-Server Decision Support Application for Agricultural Water Resources Management

Water cycle extremes such as droughts and floods present a challenge for water managers and for policy makers responsible for the administration of water supplies in agricultural regions. In addition to the inherent uncertainties associated with forecasting extreme weather events, water planners need to anticipate water demands and water user behavior in a typical circumstances. This requires the use decision support systems capable of simulating agricultural water demand with the latest available data. Unfortunately, managers from local and regional agencies often use different datasets of variable quality, which complicates coordinated action. In previous work we have demonstrated novel methodologies to use satellite-based observational technologies, in conjunction with hydro-economic models and state of the art data assimilation methods, to enable robust regional assessment and prediction of drought impacts on agricultural production, water resources, and land allocation. These methods create an opportunity for new, cost-effective analysis tools to support policy and decision-making over large spatial extents. The methods can be driven with information from existing satellite-derived operational products, such as the Satellite Irrigation Management Support system (SIMS) operational over California, the Cropland Data Layer (CDL), and using a modified light-use efficiency algorithm to retrieve crop yield from the synergistic use of MODIS and Landsat imagery. Here we present an integration of this modeling framework in a client-server architecture based on the Hydra platform. Assimilation and processing of resource intensive remote sensing data, as well as hydrologic and other ancillary information occur on the server side. This information is processed and summarized as attributes in water demand nodes that are part of a vector description of the water distribution network. With this architecture, our decision support system becomes a light weight 'app' that connects to the server to retrieve the latest information regarding water demands, land use, yields and hydrologic information required to run different management scenarios. Furthermore, this architecture ensures all agencies and teams involved in water management use the same, up-to-date information in their simulations.

Agricultural↗

Human Performance Contributions to Safety in Commercial Aviation

Every day in aviation, pilots, air traffic controllers, and other front-line personnel perform countless correct judgments and actions in a variety of operational environments. These judgments and actions are often the difference between an accident and a non-event. Ironically, data on these behaviors are rarely collected or analyzed. Data-driven decisions about safety management and design of safety-critical systems are limited by the available data, which influence how decision makers characterize problems and identify solutions. Large volumes of data are collected on the failures and errors that result in infrequent incidents and accidents, but in the absence of data on behaviors that result in routine successful outcomes, safety management and system design decisions are based on a small sample of nonrepresentative safety data. This assessment aimed to find and document “safety successes” made possible by human operators. With many Aeronautics Research Mission Directorate (ARMD) Programs and Projects focusing on increased automation and autonomy and decreased human involvement, failure to fully consider the human contributions to successful system performance in civil aviation represents a significant risk — a risk that has not been recognized to date. Without understanding how humans contribute to safety, any estimate of predicted safety of autonomous capabilities is incomplete and inherently suspect. Furthermore, understanding the ways in which humans contribute to safety can promote strategic interactions among safety technologies, functions, procedures and the people using them. Without this understanding, the full benefits of an integrated, optimized human/technology or autonomous system will not be realized. Historically, safety has been consistently defined in terms of the occurrence of accidents or recognized risks (i.e., in terms of things that go wrong). These adverse outcomes are explained by identifying their causes, and safety is restored by eliminating or mitigating these causes. An alternative to this approach is to focus on what goes right and identify how to replicate that process. Focusing on the rare cases of failures attributed to “human error” provides little information about why human performance routinely prevents adverse events. Hollnagel has proposed that things go right because people continuously adjust their work to match their operating conditions. These adjustments become increasingly important as systems continue to grow in complexity. Thus, the definition of safety should reflect not only “avoiding things that go wrong” but “ensuring that things go right.” The basis for safety management requires developing an understanding of everyday activities. However, few mechanisms to monitor everyday work exist in the aviation domain, which limits opportunities to learn how designs function in reality. This concept of safety thinking and safety management is reflected in the emerging field of resilience engineering. According to Hollnagel, a system is resilient if it can sustain required operations under expected and unexpected conditions by adjusting its functioning prior to, during, or following changes, disturbances, and opportunities. To explore “positive” behaviors that contribute to resilient performance in commercial aviation, the assessment team examined a range of existing sources of data about pilot and air traffic control (ATC) tower controller performance, including subjective interviews with domain experts and objective aircraft flight data records. These data were used to identify strategies that support resilient performance, methods for exploring and refining those strategies in existing data, and proposed methods for capturing and analyzing new data.

Null, Cynthia H.↗

Certification Gap Analysis

This report describes a generic method for addressing any new technology to its associated set of regulations and certification criteria. The result is a framework under which a detailed assessment can be conducted. Using just such a framework, the report maps the detailed updated regulations and evolving ASTM standards to the particular technology planning and tests. As a result, a roadmap of NASA technology is documented that shows clear transfer of technology data to industry (standards developers, as well as technology developers) and the FAA regulatory policy and certification staff upon whom certification and policy will be data-driven. A clear description of benefits and gaps are identified, as well.

Herbert Schlickenmaier↗

Natural Language Processing Methods for Air Traffic Management Text and Speech Data

This presentation discusses two efforts of the NARI AI/ML Intern team during the Fall 2021 OSTEM Internship term. For Letters of Agreement (LoA), we have studied how LoAs are structured and explored the question ‘What is an LoA constraint?’ To do this, our approach is data-driven, iterative, and assisted by machine learning when available. In this presentation, we will walk through our tasks of manually scanning through documents, performing a preliminary entity labelling task, and our unsupervised analysis on LoA procedures sections. After this research phase, we define the smallest constraint unit in an LoA, and start to perform entity extraction. Looking towards constraint extraction, we are also exploring the use of a one-class support vector machine (OneClassSVM) model to identify patterns within the data. The second effort of our team this term is focused on Air Traffic Control System Command Center (ATCSCC) advisory meetings, and the subsequent advisory documents that get published from their content. These advisory documents are important to give readily accessible summaries of daily operations, so that data centers, airline officials, and other stakeholders can easily understand the context of these meetings in real time. In applying machine learning to this scenario, two natural language processing tasks are used. First is developing machine learning models to convert the meeting speech data into text. With this text, use of extractive and abstractive text summarization models are used to automatically generate preliminary versions of the advisory documents.

Natural Language Processing↗

Capturing Complex Multivariate Time Series Interactions to Detect High-Risk Adverse Events During Flight

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Capturing Multivariate Time Series Interactions to Detect High‑Risk Instability During Approach

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Leveraging Human Performance Data to Change the Narrative that People are the Safety Problem

The study of errors and failure has a long and productive history in the behavioral sciences. By studying how systems fail, we rule out various mechanisms for how those systems might work, thereby refining our theories of how they actually work. Human performance, however, includes more than errors; human performance comprises both failures and successes. A systematic bias to collect and analyze data only on error affects the decisions we make as a community by promoting the narrative that “people are the safety problem.” This narrative manifests in both obvious and subtle ways in the design of systems intended for human use. When the only safety data that are available are about human failure, then “data-driven” designs can only consider that humans fail. Changing this narrative will depend on new data and new ways to examine data – specifically, data on the processes by which human create and contribute to safety. An alternate narrative is that people represent a primary source of safety, through their capability to anticipate, monitor for, respond to, and learn from expected and unexpected change. This presentation will describe research efforts to expand the range of safety-relevant events to include not just rare safety failures but frequent safety successes. These efforts include use of data from both operations and simulations to develop methods and metrics for learning from structured observation, self-report, and system data.

Jon Holbrook↗

An Advanced Open-Source Platform for Air Quality Analysis, Visualization, and Prediction

Ambient air pollution is the largest environmental health risk factor, leading to several million premature deaths globally per year. The challenge of combating poor air quality is exacerbated by growing urban populations, changing emissions, and a warming climate. While there have been many advances monitoring and modeling of atmospheric composition, reflected in the dramatic increase in archived Earth Observations, there is no single measurement or method that alone can provide an accurate depiction of the entire atmosphere. The rapidly growing collections of observational and modeling data require us to be smarter about what data to include, and how such data is used. In recent years, NASA has invested significantly in advancing the concepts for Analytics Collaborative Framework (ACF) [5] and New Observing Strategies (NOS) [4] to tackle our software infrastructure need for harmonized data management and dynamic acquisition of diverse measurements for on-demand, interactive, multivariate analysis, and access [3]. It is not enough to have a big data, standalone analytics solution; it is critical that we start integrating data from remote sensing, modeling, and in-situ networks in a harmonized manner that enables timely and data-driven decision-making for air quality management. This work presents the design and development of an Air Quality Analytics Collaborative Framework (AQ ACF), as part of NASA’s Advanced Information Systems Technology (AIST) effort, to establish a data, machine-learning, and numerically driven platform for air quality analysis, visualization, and prediction.

Liu, Qian↗

Impact Real World System Validation

Introduction NASA has developed a new evidence-based data-driven probabilistic risk assessment and tradespace analysis tool as a successor to the Integrated Medical Model. This updated decision support tool is known as IMPACT (Informing Mission Planning via Analysis of Complex Tradespaces). IMPACT estimates the frequency and consequences of medical conditions that might arise during exploration missions. A validation analysis of IMPACT was performed with respect to a set of International Space Station (ISS) and Shuttle Transportation System (STS) real world system (RWS) referent data due to the limited referent data available from exploration missions. Methods Observed mission and crew characteristics from STS and ISS missions were used as model inputs within MEDPRAT (Medical Extensible Dynamic Probabilistic Risk Assessment Tool). For each mission, two hundred thousand simulations were generated. For each mission, model outputs included occurrence counts for each condition, total medical events (TME), and the probability of loss of crew life (LOCL). These simulated model outputs were compared to the RWS referent data. Results The predicted number of total medical events exceeded the total RWS medical events for ISS missions and combined ISS and STS missions and fell within the 90% confidence interval for STS missions. For the 32 ISS missions simulated by IMPACT, the number of total medical events was overpredicted for 19 missions and fell within the 90% confidence interval for 13 missions. For the 21 STS missions, the total number of medical events was overpredicted for 3 missions, fell within the 90% confidence interval for 16 missions, and was underpredicted for 2 missions. Combined, 29 missions were in range, 22 were overpredicted, and 2 were underpredicted. The predicted LOCL probability for the 32 ISS missions, the 21 STS missions, and the combined ISS and STS missions was consistent with the zero LOCL events observed in the RWS referent data. The validation analysis included a comparison of the number of medical events predicted by IMPACT and the number of medical events observed in the RWS data on a condition-by-condition basis. For ISS missions, 50 conditions were in range, 52 conditions were statistically underpowered (not enough observed sample to draw any conclusions on precision), 8 conditions were overpredicted, and 9 conditions were underpredicted. Overall, only 14% (17/119) of conditions were out of range for STS missions, 40 conditions were in range, 59 conditions were statistically underpowered, 10 conditions were overpredicted, and 10 conditions were underpredicted. Overall, only 17% (20/119) of conditions were out of range. For combined ISS and STS missions, 11 conditions were overpredicted, and 11 conditions were underpredicted. Overall, only 18% (22/119) of conditions were out of range. For combined ISS and STS missions, 49 conditions were in range, 46 conditions were statistically underpowered, 18 conditions were overpredicted, and 8 conditions were underpredicted. Overall, 21% (26/121) of conditions were out of range. Conclusion The results of this validation analysis should not be interpreted as a pass/fail test of the validity of IMPACT. Instead, this validation analysis should be used to assess some of the IMPACT outcomes in terms of consistencies and inconsistencies with the ISS and STS RWS referent data.

L. Boley↗

Development of Level of Detail System and First-Person Camera for the GCAS Visualization Suite

The use of data-driven simulations has become standard practice as part of planning for future space missions. These simulations allow visualizing the data interactively to show what the data represents, as well as the importance of the data in the context of the mission. Using this visualized data can enhance users’ understanding of it and accelerate analysis efforts related to missions planned around it. Three-dimensional (3D) visualization software was developed to allow creating 3D representations of various communication systems, as well as the physical terrain of the Moon, for upcoming missions. The goal of this software development effort was to create interactive visualization capabilities in the Glenn Research Center Communication Analysis Suite (GCAS) using data exported from MATLAB® (MathWorks, Inc.) scripts. This software had the functionality to visualize the line of sight and dynamic link margins of the communication satellites orbiting the Earth and the Moon. One important addition to this was the visualization of the terrain data located within the GeoTIFF files, which were produced in an effort to understand the Moon’s terrain. Proper displacement values of this data have to be visualized to showcase where craters are located and how the shadow casting works with said craters at different points of the day, as well as analysis of possible landing sites for future lunar expeditions. The graphics library coded in JavaScript, three.js, had been previously selected for developing this visualization software. The software was revised to conform to modern standards, then further developed to convert the MATLAB® data into JavaScript 3D objects and Blender GL Transmission Format Binary file (GLB) objects, which were to be imported into the scene. In the process, a variety of other testing projects were created to be combined with this project at a later point; these included the first-person camera movements around spherical objects to portray human movement around the Moon, GeoTIFF loading methods, data transfer methods for incorporating the elevation data into the scene, and level of detail (LOD) capabilities to decrease memory usage and rendering time.

Visualization↗

The Unintended Consequences of Focusing on Human Error (And How You Can Help)

The literature on human performance is rich with findings of cognitive failures and methods to identify, label, and measure them. In many real-world contexts, however, outcomes are driven far more by successful than failed cognition. Designers of systems intended for human use, in an effort to be “data driven,” rely upon findings from the cognitive performance literature to inform their system designs. When most available data are about human error, data-driven designs focus on the human primarily as a source of failure. Designs intended to support or replace humans often fail to acknowledge or understand the capabilities that humans routinely contribute to successful performance. Consequently, designs intended to “protect” the system from “error-prone” humans can design-out the capability for the human to effectively intervene or adapt. The development of paradigms to study successful human performance represents a significant and largely untapped opportunity for research in cognition.

Jon Holbrook↗

Gaussian Process for Flight Delay Prediction: Learning a Stochastic Process

This paper presents a machine-learning approach to predict flight delays. Whereas neural networks are extensively studied for predictive capabilities, they involve non-intuitive design and extensive analysis, particularly in training and optimization processes. Instead, the proposed framework employs Gaussian Processes as a supervised learning technique for flight delay prediction. This data-driven approach trains the model using prior information, specifically the mean and covariance tied to existing data. The proposed Gaussian Process Regression (GPR) model employs the day of flight as a pivotal feature for delay forecasting. We analyze flights from various routes and gauge the accuracy of the presented learning technique by comparing the predicted delays with the actual ones. Given the inherent challenges in precisely forecasting delays, we predict the delays with a 95 % confidence interval. Also, an error propagation analysis in the prediction horizon is carried out to determine the optimal time frame for prediction. The proposed method for flight delay prediction is important as airlines can strategize flight operations and issue timely advisories.

stochastic↗

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids: Preprint

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid↗

Data-Informed Evaluation Framework for Integrated Energy Systems: Insights from Power, Process Heat, and Hydrogen Production Applications

The multi-criteria decision analysis (MCDA) framework provides a systematic evaluation of the diverse preferences and performance metrics associated with alternative solutions. This approach is advantageous over a single-criterion methodology, which are only valid under conditions that assume ceteris paribus or an "apples-to-apples" comparison. However, selecting suitable technologies for integrated energy systems (IES) can be likened to an "apples-to-oranges" comparison, given the heterogeneous factors at stake. These factors include economics and performance parameters, geological compatibility, and environmental impacts. Consequently, past research has often employed a mixture of qualitative and quantitative criteria tailored to the specific interests of each study. While the method proves effective in handling the intricate interplay of criteria, the resulting rankings and scores can vary from study to study. This inconsistency is introduced from the use of subjectively defined thresholds and weights. As a result, decision-makers frequently find it challenging to establish clear connections between specific criteria and the resulting scores, as the transformation of criteria into ordinal scores results in a substantial loss of information. To address this challenge, we introduce a data-informed IES evaluation framework that offers comprehensive, interpretable, and traceable evaluations backed by quantifiable rationale. First, we identified key IES evaluation criteria from a decade of literature, focusing on relevant IES applications in power, process heat, and hydrogen production. We leveraged state-of-the-art cost estimates from the Idaho National Laboratory (INL) and technical data from 78 reactor designs from the International Atomic Energy Agency (IAEA) and the Organization for Economic Co-operation and Development - Nuclear Energy Agency (OECD-NEA). Lastly, we established thresholds by analyzing the mean, variance, root mean square, and slope of values across alternatives, categorizing the preferences of decision-makers into distinct utility functions, such as linear, saturating, exponential, and stepwise. Our approach yielded two main outcomes: (1) it provided consistent assessments across different stakeholder groups and (2) it visualized uncertainties in the decision-making context via comprehensive sensitivity analysis. To demonstrate the impact of our framework, we conducted case studies on 6 reactor designs (AP1000, NuScale, BWRX-300, Xe-100, eVinci, iMSR) for the three applications. Our data-driven framework proved to be highly effective in addressing heterogenous uncertainties faced by varied decision-makers? preferences and IES applications, as well as cost and technical estimates of advanced reactors.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Ultra-filtration(UF) units are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square (RMSE) metric. Accurate prediction of initial TMP is critical for optimizing CCRO operations, as it enables the development of robust modelling frameworks that enhance process efficiency and reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗