Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Advancements and opportunities to improve bottom–up estimates of global wetland methane emissions

Wetlands are the single largest natural source of atmospheric methane (CH 4 ), contributing approximately 30% of total surface CH 4 emissions, and they have been identified as the largest source of uncertainty in the global CH 4 budget based on the most recent Global Carbon Project CH 4 report. High uncertainties in the bottom–up estimates of wetland CH 4 emissions pose significant challenges for accurately understanding their spatiotemporal variations, and for the scientific community to monitor wetland CH 4 emissions from space. In fact, there are large disagreements between bottom–up estimates versus top–down estimates inferred from inversion of atmospheric CH 4 concentrations. To address these critical gaps, we review recent development, validation, and applications of bottom–up estimates of global wetland CH 4 emissions, as well as how they are used in top–down inversions. These bottom–up estimates, using (1) empirical biogeochemical modeling (e.g. WetCHARTs: 125–208 TgCH 4 yr -1 ); (2) process-based biogeochemical modeling (e.g. WETCHIMP: 190 ± 39 TgCH 4 yr -1 ); and (3) data-driven machine learning approach (e.g. UpCH4: 146 ± 43 TgCH 4 yr -1 ). Bottom–up estimates are subject to significant uncertainties (~80 Tg CH 4 yr -1 ), and the ranges of different estimates do not overlap, further amplifying the overall uncertainty when combining multiple data products. These substantial uncertainties highlight gaps in our understanding of wetland CH 4 biogeochemistry and wetland inundation dynamics. Major tropical and arctic wetland complexes are regional hotspots of CH 4 emissions. However, the scarcity of satellite data over the tropics and northern high latitudes offer limited information for top–down inversions to improve bottom–up estimates. Recent advances in surface measurements of CH 4 fluxes (e.g. FLUXNET-CH 4 ) across a wide range of ecosystems including bogs, fens, marshes, and forest swamps provide an unprecedented opportunity to improve existing bottom–up estimates of wetland CH 4 estimates. We suggest that continuous long-term surface measurements at representative wetlands, high fidelity wetland mapping, combined with an appropriate modeling framework, will be needed to significantly improve global estimates of wetland CH 4 emissions. There is also a pressing unmet need for fine-resolution and high-precision satellite CH 4 observations directed at wetlands.

54 ENVIRONMENTAL SCIENCES↗

The Poisson tensor completion non-parametric differential entropy estimator

We introduce the Poisson tensor completion (PTC) estimator, a non-parametric differential entropy estimator. The PTC estimator leverages inter-sample relationships to compute a low-rank Poisson tensor decomposition of the frequency histogram. Our crucial observation is that the histogram bins are an instance of a space partitioning of counts and thus can be identified with a spatial Poisson process. The Poisson tensor decomposition leads to a completion of the intensity measure over all bins—including those containing few to no samples—and leads to our proposed PTC differential entropy estimator. A Poisson tensor decomposition models the underlying distribution of the count data and guarantees non-negative estimated values and so can be safely used directly in entropy estimation. Our estimator is the first tensor-based estimator that exploits the underlying spatial Poisson process related to the histogram explicitly when estimating the probability density with low-rank tensor decompositions for the purpose of tensor completion. Furthermore, we demonstrate that our PTC estimator is a substantial improvement over standard histogram-based estimators for sub-Gaussian probability distributions because of the concentration of norm phenomenon.

42 ENGINEERING↗

Extreme metrics from large ensembles: investigating the effects of ensemble size on their estimates

Abstract. We consider the problem of estimating the ensemble sizes required to characterize the forced component and the internal variability of a number of extreme metrics. While we exploit existing large ensembles, our perspective is that of a modeling center wanting to estimate a priori such sizes on the basis of an existing small ensemble (we assume the availability of only five members here). We therefore ask if such a small-size ensemble is sufficient to estimate accurately the population variance (i.e., the ensemble internal variability) and then apply a well-established formula that quantifies the expected error in the estimation of the population mean (i.e., the forced component) as a function of the sample size n, here taken to mean the ensemble size. We find that indeed we can anticipate errors in the estimation of the forced component for temperature and precipitation extremes as a function of n by plugging into the formula an estimate of the population variance derived on the basis of five members. For a range of spatial and temporal scales, forcing levels (we use simulations under Representative Concentration Pathway 8.5) and two models considered here as our proof of concept, it appears that an ensemble size of 20 or 25 members can provide estimates of the forced component for the extreme metrics considered that remain within small absolute and percentage errors. Additional members beyond 20 or 25 add only marginal precision to the estimate, and this remains true when statistical inference through extreme value analysis is used. We then ask about the ensemble size required to estimate the ensemble variance (a measure of internal variability) along the length of the simulation and – importantly – about the ensemble size required to detect significant changes in such variance along the simulation with increased external forcings. Using the F test, we find that estimates on the basis of only 5 or 10 ensemble members accurately represent the full ensemble variance even when the analysis is conducted at the grid-point scale. The detection of changes in the variance when comparing different times along the simulation, especially for the precipitation-based metrics, requires larger sizes but not larger than 15 or 20 members. While we recognize that there will always exist applications and metric definitions requiring larger statistical power and therefore ensemble sizes, our results suggest that for a wide range of analysis targets and scales an effective estimate of both forced component and internal variability can be achieved with sizes below 30 members. This invites consideration of the possibility of exploring additional sources of uncertainty, such as physics parameter settings, when designing ensemble simulations.

54 ENVIRONMENTAL SCIENCES↗

Solution Irregularity Remediation for Spatial Discretization Error Estimation for S N Transport Solutions

The discrete ordinates linear Boltzmann transport equation is typically solved in its spatially discretized form, incurring spatial discretization error. Quantification of this error for purposes such as adaptive mesh refinement or error analysis requires an a posteriori estimator, which utilizes the numerical solution to the spatially discretized equation to compute an estimate. Because the quality of the numerical solution informs the error estimate, irregularities, present in the true solution for any realistic problem configuration, tend to cause the largest deviation in the error estimate vis-a-vis the true error. In this paper, an analytical partial singular characteristic tracking (pSCT) procedure for reducing the estimator’s error is implemented within our novel residual source estimator for a zeroth-order discontinuous Galerkin scheme, at the additional cost of a single inner iteration. Here, a metric-based evaluation of the pSCT scheme versus the standard residual source estimator is performed over the parameter range of a Method of Manufactured Solutions test suite. The pSCT scheme generates near-ideal accuracy in the estimate in problems where the dominant source of the estimator’s error is the solution irregularity, namely, problems where the true solution is discontinuous and problems where the true solution’s first derivative is discontinuous and the scattering ratio is low. In problems where the scattering ratio is high and the true solution is discontinuous in the first derivative, the error in the scattering source, which is not converged by the pSCT scheme, is greater than the error incurred due to the irregularity. Ultimately, a pSCT scheme is judged to be useful for error estimation in problems where the computational cost of the scheme is justified. In the presence of many irregularities, such a scheme may be intractable for general use, but in benchmarks, as an analytical tool, or in problems that have nondissipative discontinuities, the scheme may prove invaluable.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A New Estimator to Correct for Bias from Tag Rate Expansion on Natural-Origin Fish Attributes in Mixed-Stock Analysis Using Parentage-Based Tagging

Abstract In fisheries where hatchery- and natural-origin conspecifics occur as mixed stocks, it is often important to estimate both the natural-origin proportion of the mixture and the composition of attributes within the natural-origin portion (e.g., genetic stock, sex, and age-class). These estimates are facilitated by parentage-based tagging, which allows large numbers of hatchery fish to be efficiently tagged and later identified. When tag rates are less than 100%, the untagged fish in the mixture include both untagged hatchery fish and natural-origin fish. Unbiased estimation of the abundance and attribute composition of the natural-origin portion requires an estimator that accounts for tag rates. Two estimators are described: one “accounting-style” estimator similar to previously described approaches and one maximum likelihood method. These estimators were evaluated and compared using simulations mimicking estimation of the composition of fish migrating past a dam. The two estimators performed similarly at reducing bias when tag rates were high, but the maximum likelihood method had smaller mean square error when tag rates were low. We provide an R package to allow usage of this estimator in a wide variety of fisheries applications.

Delomas, Thomas A. (ORCID:000000015154759X)↗

Comparison of Global Aboveground Biomass Estimates From Satellite Observations and Dynamic Global Vegetation Models

The global forest carbon stocks represent the amount of carbon stored in woody vegetation and are important for quantifying the ability of the global forests to sequester atmospheric CO 2 and to provide ecosystem services (e.g., timber) under climate change. The forest ecosystem carbon pool estimates are highly variable and poorly quantified in areas lacking forest inventory estimates. Here, we compare and analyze aboveground biomass (AGB) estimates from five satellite-based global data sets and nine dynamic global vegetation models (DVGMs). We find that across the data sets, mean AGB exhibits the largest variability around the tropical area. In addition, AGB shows a similar latitudinal trend but large variability among the data sets. Satellite-based AGB estimates are lower than those simulated by DVGMs. The divergence among the satellite-based AGB estimates can be driven by the methodology, input satellite products, and the forested areas used to estimate AGB. The modeled NPP, autotrophic respiration, and carbon allocation mostly drive the variability of AGB simulated by DGVMs. The future availability of a high-quality global forest area map is anticipated to improve AGB estimate accuracy and to reduce the discrepancies among different satellite- and model-based AGB estimates. Furthermore, we suggest the carbon-modeling community reexamine the methodology used to estimate AGB and forested areas for a more robust global forest carbon stock estimation.

54 ENVIRONMENTAL SCIENCES↗

Using spatio-temporal graph neural networks to estimate fleet-wide photovoltaic performance degradation patterns

Accurate estimation of photovoltaic (PV) system performance is crucial for determining its feasibility as a power generation technology and financial asset. PV-based energy solutions offer a viable alternative to traditional energy resources due to their superior Levelized Cost of Energy (LCOE). A significant challenge in assessing the LCOE of PV systems lies in understanding the Performance Loss Rate (PLR) for large fleets of PV systems. Estimating the PLR of PV systems becomes increasingly important in the rapidly growing PV industry. Precise PLR estimation benefits PV users by providing real-time monitoring of PV module performance, while explainable PLR estimation assists PV manufacturers in studying and enhancing the performance of their products. However, traditional PLR estimation methods based on statistical models have notable drawbacks. Firstly, they require user knowledge and decision-making. Secondly, they fail to leverage spatial coherence for fleet-level analysis. Additionally, these methods inherently assume the linearity of degradation, which is not representative of real world degradation. To overcome these challenges, we propose a novel graph deep learning-based decomposition method called the Spatio-Temporal Graph Neural Network for fleet-level PLR estimation (PV-stGNN-PLR). PV-stGNN-PLR decomposes the power timeseries data into aging and fluctuation components, utilizing the aging component to estimate PLR. PV-stGNN-PLR exploits spatial and temporal coherence to derive PLR estimation for all systems in a fleet and imposes flatness and smoothness regularization in loss function to ensure the successful disentanglement between aging and fluctuation. We have evaluated PV-stGNN-PLR on three simulated PV datasets consisting of 100 inverters from 5 sites. Experimental results show that PV-stGNN-PLR obtains a reduction of 33.9% and 35.1% on average in Mean Absolute Percent Error (MAPE) and Euclidean Distance (ED) in PLR degradation pattern estimation compared to the state-of-the-art PLR estimation methods.

14 SOLAR ENERGY↗

Estimating Field-Level Perennial Bioenergy Grass Biomass Yields Using the Normalized Difference Red-Edge Index and Linear Regression Analysis for Central Virginia, USA

We investigated the indicative power of the normalized difference red-edge index (NDRE) for estimating field-level perennial bioenergy grass biomass yields utilizing Sentinel-2 imagery and a linear regression model as a rapid, cost-effective method for biomass yield estimations for bioenergy. We used 2019 data from three study sites containing mature perennial bioenergy grass stands in central Virginia, USA. Of the simulated daily NDRE values based on the temporally weighted averaging of two temporal neighbors, we found the strongest index–yield correlation on 11 August (R = 0.85). We estimated the perennial bioenergy grass biomass yields for (1) all sites using the data pooled from the three sites (all-site estimation) and (2) each site using the data pooled from the other two sites (cross-site estimation). The estimated field-level perennial bioenergy grass biomass yields strongly correlated with the recorded yields (average R2 = 0.76), with a root mean square error (RMSE) of 1.5 Mg/ha and a mean absolute error (MAE) of 1.2 Mg/ha for the all-site estimation. For the cross-site estimation, the site with diverse perennial grass types had the weakest correlation (R2 = 0.44) of the sites, indicating a difficulty in accounting for heterogeneous index–yield relationships in a single model. In addition to identifying a strong indicative power of the NDRE for estimating the overall perennial bioenergy grass biomass yields at a field level, the findings from this study call for an analysis across multiple perennial grasses and a comparison using multiple sites to understand (1) if the indicative power of the index shifts from the biomass of the specific perennial bioenergy grass type to the overall biomass during the growing season and (2) the level of perennial bioenergy grass heterogeneity that may hinder the remotely sensed biomass yield estimation using a single model.

09 BIOMASS FUELS↗

Regularized Differentiation for Bioburden Density Estimation in Planetary Protection

In this paper, we propose and investigate the performance of two novel shrinkage estimators for bioburden density estimation in planetary protection. The estimators are based on the regularized differentiation of a cumulative count of colony forming units collected throughout the data collecting session or the life cycle of the entire mission. The regularized differentiation recasts the problem of bioburden density estimation as a linear least squares problem. The least squares problem is then solved through regularization techniques, such as truncated singular value decomposition and penalized least squares. The regularization is necessary to avoid noise amplification during the differentiation of noisy data. The two regularization estimators are compared with four other commonly used estimators to simultaneously evaluate the means of multivariable independent Poisson distributions: the maximum likelihood, noninformative Bayes estimator with Jeffreys prior, Empirical Bayes using conjugate gamma-Poisson model with gamma parameters selected by method of moments, and the Clevenson-Zidek estimator. It is shown through computer-simulated data that the regularized differentiation based on ridge regression has the smallest mean-squared error among all estimators. The analysis of shrinkage mechanism implemented by regularized differentiation is performed, and it is shown that the regularized differentiation amounts to performing a weighted averaging of all the samples. The weights are determined by the regularization parameter automatically selected by the L-curve technique. Since the method of least squares makes no distributional assumptions about the data, it presents an attractive technique for bioburden density estimation when there are concerns about the misspecification of the distributional model. The paper concludes with the analysis of the bioburden data collected during InSight mission and directions for future work.

97 - MATHEMATICS AND COMPUTING↗

A Multi-Fidelity Gaussian Process Regression Method for Probabilistic Wind Farm Power Curve Estimation

Accurate estimation of the power curve for wind turbines or wind farms is crucial to ensure their efficient operation and management. However, conventional methods for power curve estimation rely either on expensive and infrequent measurements or on low-quality numerical simulations. Moreover, the majority of previous studies on power curve estimation for wind turbines or wind farms focused on deterministic estimation, which provides a point estimate of the relationship between wind speed and power generation. Nevertheless, the deterministic approach fails to consider the inherent uncertainty associated with wind energy production resulting from varying turbine characteristics. This can lead to inaccurate power generation estimation and suboptimal decisions regarding energy management. In this paper, a kernel density estimation (KDE) based Multi-Fidelity Gaussian Process Regression (MFGPR) model is proposed to fuse theoretical power curve data and the ground true measurements to create a mapping of wind speed and wind power. By conducting a case study on an actual wind farm in China, the efficacy of the proposed MFGPR model was demonstrated in characterizing the variability of wind power. The probabilistic MFGPR model was also able to generate confidence intervals that encompassed the measured power, thereby improving the accuracy and confidence in wind power estimation or wind resource assessment. Overall, the proposed MFGPR model offers a reliable approach to integrate high-fidelity ground measurements and theoretical power curve data, resulting in precise wind resource assessment and power estimation.

Gaussian process regression↗

Two-Stage Optimization Framework for Detecting and Correcting Parameter Cyber-Attacks in Power System State Estimation

One major tool of Energy Management Systems for monitoring the status of the power grid is State Estimation. Since the results of state estimation are used within the energy management system, the security of the state estimation process is most important. The focus research in this area is on detecting False Data Injection attacks on measurements. While this is important, State Estimation also rely on database that are used to describe the relationship between measurements and systems' states. This paper presents a two-stage programming framework to detect and correct attacks in the parameters of the measurement model used by the state estimation process in the Energy Management System. In the first stage, an estimate of the line parameters ratios are obtained. In the second stage, the estimated ratios from stage I are used in a Bi-Level model for obtaining a final estimate of the measurements' model parameters. Hence, the presented framework does not only unify the detection and correction in a single optimization run, but also provide a monitoring scheme for the SE database that is typically considered static. In addition, in the two stages, linear programming framework is preserved. For validation, the IEEE 118 bus system is used for implementation. The results of this paper illustrate the effectiveness of the proposed model for detecting attacks in the database used in the state estimation process.

state estimation, two-stage optimization, cyber-ph↗

Improved Gas Plume Identification Using Nearest Neighbor Methods for Background Estimation

Longwave infrared (LWIR) hyperspectral imaging (HSI) can be used for many tasks in remote sensing, including detecting and identifying effluent gases by LWIR sensors on airborne platforms. Identification is used after detection to increase confidence in weakly detected plumes, reduce false positives from detection, and distinguish between similar and confounding material signatures. Background estimation is an important step used to reveal the unique spectral characteristics of the detected gas, allowing the identification model to determine what the gas is specifically. The importance of proper background estimation increases when dealing with weak signals, large libraries of gases of interest, and uncommon or heterogeneous backgrounds. In this article, we propose two methods for background estimation: a novel k-nearest segments (KNS) algorithm and the standard k-nearest neighbors (KNN) algorithm. We test our methods and three existing background estimation methods for comparison against global background estimation to determine which performs best at estimating the true background radiance under a plume and for increasing identification confidence using a neural network classification model. We compare the different methods using 640 simulated weak plumes in an urban environment. For identification, our KNS algorithm improves median neural network identification confidence by 53.2%. For background radiance estimation, the KNN algorithm provides a median of 49 times less RMSE than global background estimation. Furthermore, KNN is the easiest method to tune for different plumes, making it an excellent “out of the box” background estimator.

47 OTHER INSTRUMENTATION↗

A life cycle and product type based estimator for quantifying the carbon stored in wood products

Background: Timber harvesting and industrial wood processing laterally transfer the carbon stored in forest sectors to wood products creating a wood products carbon pool. The carbon stored in wood products is allocated to end-use wood products (e.g., paper, furniture), landfill, and charcoal. Wood products can store substantial amounts of carbon and contribute to the mitigation of greenhouse effects. Therefore, accurate accounts for the size of wood products carbon pools for different regions are essential to estimating the land-atmosphere carbon exchange by using the bottom-up approach of carbon stock change. Results: To quantify the carbon stored in wood products, we developed a state-of-the-art estimator (Wood Products Carbon Storage Estimator, WPsCS Estimator) that includes the wood products disposal, recycling, and waste wood decomposition processes. The wood products carbon pool in this estimator has three subpools: (1) end-use wood products, (2) landfill, and (3) charcoal carbon. In addition, it has a user-friendly interface, which can be used to easily parameterize and calibrate an estimation. To evaluate its performance, we applied this estimator to account for the carbon stored in wood products made from the timber harvested in Maine, USA, and the carbon storage of wood products consumed in the United States. Conclusion: The WPsCS Estimator can efficiently and easily quantify the carbon stored in harvested wood products for a given region over a specific period, which was demonstrated with two illustrative examples. In addition, WPsCS Estimator has a user-friendly interface, and all parameters can be easily modified.

54 ENVIRONMENTAL SCIENCES↗

Prediction of Solar Irradiance and Photovoltaic Solar Energy Product Based on Cloud Coverage Estimation Using Machine Learning Methods

Cloud cover estimation from images taken by sky-facing cameras can be an important input for analyzing current weather conditions and estimating photovoltaic power generation. The constant change in position, shape, and density of clouds, however, makes the development of a robust computational method for cloud cover estimation challenging. Accurately determining the edge of clouds and hence the separation between clouds and clear sky is difficult and often impossible. Toward determining cloud cover for estimating photovoltaic output, we propose using machine learning methods for cloud segmentation. We compare several methods including a classical regression model, deep learning methods, and boosting methods that combine results from the other machine learning models. To train each of the machine learning models with various sky conditions, we supplemented the existing Singapore whole sky imaging segmentation database with hazy and overcast images collected by a camera-equipped Waggle sensor node. We found that the U-Net architecture, one of the deep neural networks we utilized, segmented cloud pixels most accurately. However, the accuracy of segmenting cloud pixels did not guarantee high accuracy of estimating solar irradiance. We confirmed that the cloud cover ratio is directly related to solar irradiance. Additionally, we confirmed that solar irradiance and solar power output are closely related; hence, by predicting solar irradiance, we can estimate solar power output. This study demonstrates that sky-facing cameras with machine learning methods can be used to estimate solar power output. This ground-based approach provides an inexpensive way to understand solar irradiance and estimate production from photovoltaic solar facilities.

14 SOLAR ENERGY↗

Improving Abundance Estimates of Spring–Summer Snake River Chinook Salmon for Fisheries Management

Abstract The Columbia River basin is home to a run of spring–summer Chinook Salmon Oncorhynchus tshawytscha that returns to the Snake River drainage of Idaho, Oregon, and Washington in the Pacific Northwest. Historically, the run was one of the more productive throughout the Columbia River basin. However, Snake River spring–summer Chinook Salmon have experienced declines in abundance due to overfishing, habitat degradation, and dams. Several stocks are listed as threatened under the U.S. Endangered Species Act and are supported by mitigation hatcheries funded by Idaho Power Company, the Lower Snake River Compensation Plan, and the Bonneville Power Administration. To maximize tribal and state harvest of returning hatchery adults, minimize impacts on wild fish, and ensure that enough hatchery fish return to meet broodstock needs, careful fisheries management is required. Since 2008, managers have used hatchery adults, PIT-tagged as juveniles and detected at Lower Granite Dam, to generate adult abundance estimates. In season, these estimates inform state and tribal harvest shares and ensure that broodstock needs are met. Postseason, they provide smolt-to-adult survival and return rates. Since 2012, parentage-based tagging (PBT) has provided an alternative method to estimate stock- and age-specific returns at Lower Granite Dam, since returning hatchery adults sampled at Lower Granite Dam can be assigned to their parents. We compared stock-specific abundance estimates between PIT- and PBT-derived methodologies for return years 2016–2019. Across all years, PIT tag estimates accounted for 65% of the PBT-based estimates at Lower Granite Dam across all age-groups and release sites combined. This underrepresentation across all groups equated to 49,833 fish that were not accounted for in PIT tag abundance estimates. It is clear that PBT-based estimates should aide in-season harvest management and postseason run reconstruction to avoid the known bias of estimates from PIT tags, especially during years of low returns when increased accuracy is critical.

Coykendall, D. Katharine (ORCID:0000000211482397)↗

Calibration approach and range of observed sap flow influences transpiration estimates from thermal dissipation sensors

Calibrating thermal dissipation (TD) sap flow sensors has become increasingly important to accurately estimate whole-tree transpiration, but it is unclear how the calibration approach itself influences the resulting coefficients and estimates. Here, we compare the two most common calibration approaches, gravimetric and potometric, using TD sensors inserted into Eucalyptus benthamii tree stems. The gravimetric approach uses an excised stem segment devoid of branches and leaves and pushes water through the stem using gravity, a positive force. The potometric approach uses a severed stem containing an intact canopy placed upright in a reservoir where water is pulled through the stem via transpiration, a negative force. We hypothesized that the positive pressure associated with gravimetric calibration would overestimate conductive sapwood area relative to that estimated from potometric calibration and that coefficients from these different approaches would result in different estimates of transpiration when applied to intact trees. We also predicted that calibrations could improve transpiration estimates by targeting the range of observed sap flow rates (i.e., K values) in intact trees. Conductive sapwood area was higher under gravimetric calibrations and resulting estimates of transpiration were lower compared to potometric calibrations. Segmented calibration curves, which fit two separate curves for the relationship between sap flux density (Fd) and sap flux index (K) based on the range of sap flow rates observed in intact trees, increased transpiration estimates from both gravimetric and potometric coefficients and diminished the magnitude of difference in transpiration estimates between approaches. Researchers should be aware that calibration approach and range of observed sap flow profoundly influences transpiration estimates from TD sensors and this likely applies to calibrations of other heat-based sap flow sensors.

54 ENVIRONMENTAL SCIENCES↗

Validating and Comparing Energy Estimation Methods at Water Resource Recovery Facilities

Water resource recovery facilities play a crucial role in the water-energy nexus, consuming a substantial amount of energy in the United States. Growing treatment volumes and more stringent water quality standards are expected to increase the amount of energy needed to treat wastewater, but accurately estimating energy consumption and potential remains challenging due to variability in scale, treatment methods, and effluent treatment standards. In this study, we used publicly available data to evaluate the accuracy of methods for estimating energy consumption and generation, then quantified uncertainty based on key factors like flow rate, treatment level, and geographic location. To validate methods, we estimated energy consumption and generation at the facility-level, then compared estimates to self-reported data from utilities in major U.S. cities. We found that process models of treatment trains under best practice configurations were accurate relative to other methods for estimating electricity use, total energy use, and electricity generation from biogas utilization, and less complex methods based on effluent treatment level and prime movers also performed well for estimating electricity consumption and generation, respectively. Applying the evaluated methods to a national inventory of treatment facilities, we estimate that annual energy consumption ranged from 56.3 x 10^3 to 82.5 x 10^3 TJ in 2012 and 83.6 x 10^3 to 127 x 10^3 TJ in 2042. Our results indicate that not all estimation methods are suited for every use case, so we recommend that researchers and practitioners select an estimation method based on data availability and desired computational intensity.

Hodson, Abigayle↗

Do-calculus enables estimation of causal effects in partially observed biomolecular pathways

Abstract Motivation Estimating causal queries, such as changes in protein abundance in response to a perturbation, is a fundamental task in the analysis of biomolecular pathways. The estimation requires experimental measurements on the pathway components. However, in practice many pathway components are left unobserved (latent) because they are either unknown, or difficult to measure. Latent variable models (LVMs) are well-suited for such estimation. Unfortunately, LVM-based estimation of causal queries can be inaccurate when parameters of the latent variables are not uniquely identified, or when the number of latent variables is misspecified. This has limited the use of LVMs for causal inference in biomolecular pathways. Results In this article, we propose a general and practical approach for LVM-based estimation of causal queries. We prove that, despite the challenges above, LVM-based estimators of causal queries are accurate if the queries are identifiable according to Pearl’s do-calculus and describe an algorithm for its estimation. We illustrate the breadth and the practical utility of this approach for estimating causal queries in four synthetic and two experimental case studies, where structures of biomolecular pathways challenge the existing methods for causal query estimation. Availability and implementation The code and the data documenting all the case studies are available at https://github.com/srtaheri/LVMwithDoCalculus. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗