Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Generalized additive models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Tree-ring evidence marks year 2022 as the driest spring season in nearly four centuries in the Western Himalayas

The Hindukush-Karakoram-Himalayan region is a crucial freshwater source for billions across South Asia, yet its climate remains poorly understood due to limited long-term records. Winter and spring precipitation govern snow accumulation and downstream water availability in dry months particularly across the Western Himalayas (WH), but recent decades show intensifying droughts with unclear long-term context. Here, we have reconstructed a nearly four-century-long spring i.e. February to May (FMAM) precipitation for the Lahaul region of the (WH), an area dominated by the Western Disturbances. This record was developed using moisture-sensitive Cedrus deodara (Deodar) tree-rings from three high-elevation sites. A regional composite tree-ring-width chronology, developed through a Nested Principal Component Analysis and modeled with a nonlinear Generalized Additive Model (GAM) that explains 71 % of the variance during the calibration period. We identified the last two decades as the most precipitation deficit phase and the year 2022 showing the driest FMAM on record. The observed rise in the FMAM dry episodes post 1999 CE in our reconstruction, corresponds to the meteorological records. This recent drying is linked to a northward shift of the subtropical westerly jet and reduced moisture transport, both associated with unusual sea surface temperature patterns in the tropical Indian Ocean and the Western Pacific Ocean. Our results provide compelling evidence of long-term hydroclimatic instability in the WH and emphasize the value of tree-ring records in extending precipitation histories beyond the instrumental observations. Such reconstructions can be benchmarks to validate high-resolution climate models and formulate adaptation policies to mitigate future risks.

Cedrus deodara

Data and scripts associated with “Moisture content modulates DOM thermodynamic regulation of oxygen consumption in drying streambed sediments”

This data package is associated with the publication “Moisture content modulates DOM thermodynamic regulation of oxygen consumption in drying streambed sediments” published in Scientific Reports (Garayburu-Caruso et al., 2026). The package contains processed data products and scripts used to quantify how drying and re-inundation of riverbed sediments influence dissolved organic matter (DOM) thermodynamic properties and their relationship with sediment oxygen (O₂) consumption across 33 stream sites in the contiguous United States. The data package contains DOM thermodynamic metrics (e.g., Gibbs free energy of carbon oxidation and thermodynamic efficiency), and O₂ consumption along with watershed-scale climate and land-cover metrics used as explanatory variables in the analyses. Underlying unprocessed and processed ultrahigh-resolution mass spectrometry data, oxygen consumption rates from laboratory moisture-manipulation experiments, within-sample environmental properties, sediment moisture content and contextual field measurements are archived separately at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2428003 (Laan et al., 2024) and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689 (Forbes et al.,2023). A preliminary version of this data package was published in February 2026 at the time of manuscript submission. It was updated in June 2026, at the time of manuscript acceptance, to include the finalized data and additional metadata (readme, data dictionary, and file level metadata). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. At the top level, the data package is organized into five main folders: (1) Data, (2)Figures, (3) Map, (4) GAM_Reulsts, and (5) src. The Data folder contains analysis-ready tabular files with oxygen consumption rates, DOM thermodynamic properties by site and treatment, site-level environmental variables, watershed-scale metrics, and other derived variables referenced in the manuscript. The Figures folder contains static image files associated with the main text and supplemental figures, while the Map folder includes spatial data and map-layer files used to create the sampling-location map. The GAM results folder contains the results for each of the general additive model (GAM).The src folder contains R scripts used to perform data processing, statistical analyses (including clustering, generalized additive models, and threshold analysis), and figure generation. This data package is associated with a GitHub repository found at https://github.com/WHONDRS-Hub/ECA_DOM_Thermodynamics.

Dissolved organic matter

Exploring the Whole Set of Accurate Sparse Interpretable Models

In data science applications, there are often many models that fit the data well. This phenomenon was called the Rashomon Effect by Leo Breiman. The set of good models is called the Rashomon Set, and the goal of this project is to locate, store, and study the Rashomon sets for classes of interpretable models, including decision trees and generalized additive models.

97 MATHEMATICS AND COMPUTING

Hyperspectral Mapping of the Invasive Species Pepperweed and the Development of a Habitat Suitability Model

Mapping and predicting the spatial distribution of invasive plant species is central to habitat management, however difficult to implement at landscape and regional scales. Remote sensing techniques can reduce the impact field campaigns have on these ecologically sensitive areas and can provide a regional and multi-temporal view of invasive species spread. Invasive perennial pepperweed (Lepidium latifolium) is now widespread in fragmented estuaries of the South San Francisco Bay, and is shown to degrade native vegetation in estuaries and adjacent habitats, thereby reducing forage and shelter for wildlife. The purpose of this study is to map the present distribution of pepperweed in estuarine areas of the South San Francisco Bay Salt Pond Restoration Project (Alviso, CA), and create a habitat suitability model to predict future spread. Pepperweed reflectance data were collected in-situ with a GER 1500 spectroradiometer along with 88 corresponding pepperweed presence and absence points used for building the statistical models. The spectral angle mapper (SAM) classification algorithm was used to distinguish the reflectance spectrum of pepperweed and map its distribution using an image from EO-1 Hyperion. To map pepperweed, we performed a supervised classification on an ASTER image with a resulting classification accuracy of 71.8%. We generated a weighted overlay analysis model within a geographic information system (GIS) framework to predict areas in the study site most susceptible to pepperweed colonization. Variables for the model included propensity for disturbance, status of pond restoration, proximity to water channels, and terrain curvature. A Generalized Additive Model (GAM) was also used to generate a probability map and investigate the statistical probability that each variable contributed to predict pepperweed spread. Results from the GAM revealed distance to channels, distance to ponds and curvature were statistically significant (p < 0.01) in determining the locations of suitable pepperweed habitats.

Pepperweed

TPSAS-NF1676L-13204-DND

Invasive perennial pepperweed (Lepidium latifolium) has spread rapidly throughout the western United States in the past fifteen years. Pepperweed outcompetes many native species for water and nutrients, disturbing sensitive ecosystems. The purpose of this study is to map the contemporary distribution of pepperweed throughout the wetland ecosystem of the restored South San Francisco Bay Salt Ponds and create a habitat suitability model to predict future spread. Pepperweed reflectance data were collected in-situ with the GER 1500 spectroradiometer along with presence and absence pepperweed GPS data points. A Spectral Angle Mapper (SAM) classification algorithm (Method A) was used to distinguish pepperweed spectra and map its distribution on an EO-1 Hyperion image of the study area. Similarly, a supervised classification was run on an ASTER image of the study area. A GIS multivariate habitat suitability model was created to predict areas most susceptible to pepperweed colonization. Variables incorporated into the model were tidal extent, propensity for disturbance, pond salinity, proximity to levees and channels, and terrain curvature. A Generalized Additive Model (GAM) was also used to generate a suitability map (Method B) and investigate the statistical probability that each variable contributed to predict pepperweed spread. Results from the GAM revealed distance to channels, distance to disturbance factors, extreme pond salinity and curvature as statistically significant in determining the locations of suitable pepperweed habitats

Andrew Nguyen

Chrono-Validation of Near-Real-Time Landslide Susceptibility Models via Plugin Statistical Simulations

The idea behind any validation scheme in landslide susceptibility studies is to test whether a model calibrated on a certain data can predict an unknown dataset of the same nature (landslide presences/absences and covariates). Almost the entirety of landslide susceptibility studies are validated by subsetting a single dataset into a training and test sets. This dataset usually corresponds either to event-specific or to historical inventories. Very rarely, a multi-temporal inventory is available and, in the few cases where this condition is met, the validation practices involve training a model on a specific landslide inventory, deriving a single predictive equation and validating it on a subsequent landslide inventory. This commonly leads landslide predictive studies, even those with a strong statistical rigor, to neglect the uncertainty estimation in their modeling scheme. In statistics, validation can also be performed via statistical simulations. This means that after fitting a given model, one can generate any number of predictive functions and test their predictive skills on any type and number of unknown datasets. In this work, we take a similar direction and we apply it to model and validate three separate co-seismic inventories, including an uncertainty estimation phase. We mapped these inventories within the same area in Indonesia, for three earthquakes occurred in 2012, 2017 and 2018. Specifically, we build three event-specific Bayesian Generalize Additive Models of the binomial family. From each model we then simulate 1000 predictive realizations over the remaining two inventories, by using a plug-in scheme where all the morphometric covariates are kept fixed and only the ground motion is replaced according to the prediction target. By doing so, we introduce a new analytical tool for near-real-time landslide predictive purposes, which is able to produce a probabilistic model which stands in between the definitions of susceptibility and hazard. In fact, our model is able to accurately estimate “where” and “when” - although not “how frequently” - landslide have occurred by featuring the multitemporal information of the trigger. In our findings, the simulations are quite similar to the fitted models; and the nine combinations we analyse produce excellent performance. This result confirms the assumption that “the past is the key to the future”, as we show that the relative contribution of each variable and their interactions in each probabilistic model remains practically the same across temporal replicates. This information is not trivial because it supports the routines implemented in global near-real-time applications.

Temporal validation

Polarization Decomposition and Temperature Bias Resolution for SMAP Passive Soil Moisture Retrieval Using Time Series Brightness Temperature Observations

In passive microwave remote sensing of soil moisture, the tau-omega (τ-ω) model has often been used to provide soil moisture estimates at a spatial scale representative of the satellite footprint dimensions. For modeling simplicity, model parameters such as the single scattering albedo (ω) and vegetation opacity (τ) that go into the geophysical inversion process are often assumed to be independent of polarizations. Although this absence of polarization dependence can often be justified in special cases as in low-frequency remote sensing or under dense vegetation conditions, it is not a robust assumption in general. Additional model parameterization errors arising from this assumption are possible, leading to degradation in soil moisture estimation accuracy. In this paper, we propose a time series approach to try to resolve the polarization dependence of several τ-ω model parameters as well as the temperature bias arising from the ancillary temperature data. The Version 4 of the Soil Moisture Active Passive (SMAP) Level 1B brightness temperature time series observations were used to illustrate the mechanics of this approach, with an emphasis on a comparison between resulting satellite soil moisture retrievals and in situ data collected at several core validation sites. It was found that this time series approach resulted in significant reduction of the dry bias exhibited in the current SMAP passive soil moisture data products, while retaining the same performance in other metrics of the current baseline passive soil moisture retrieval algorithm.

time series

Observational benchmarks inform representation of soil organic carbon dynamics in land surface models

Abstract. Representing soil organic carbon (SOC) dynamics in Earth system models (ESMs) is a key source of uncertainty in predicting carbon–climate feedbacks. Machine learning models can help identify dominant environmental controllers and establish their functional relationships with SOC stocks. The resulting knowledge can be integrated into ESMs to reduce uncertainty and improve predictions of SOC dynamics over space and time. In this study, we used a large number of SOC field observations (n=54 000), geospatial datasets of environmental factors (n=46), and two machine learning approaches (namely random forest, RF, and generalized additive modeling, GAM) to (1) identify dominant environmental controllers of global and biome-specific SOC stocks, (2) derive functional relationships between environmental controllers and SOC stocks, and (3) compare the identified environmental controllers and predictive relationships with those in models used in Phase 6 of the Coupled Model Intercomparison Project (CMIP6). Our results showed that the diurnal temperature, drought index, cation exchange capacity, and precipitation were important observed environmental predictors of global SOC stocks. While the RF model identified 14 environmental factors that describe climatic, vegetation, and edaphic conditions as important predictors of global SOC stocks (R2=0.61, RMSE = 0.46 kg m−2), current ESMs oversimplify the relationships between environmental factors and SOC, with precipitation, temperature, and net primary productivity explaining > 96 % of the variability in ESM-modeled SOC stocks. Further, our study revealed notable disparities among the functional relationships between environmental factors and SOC stocks simulated by ESMs compared with observed relationships. To improve SOC representations in ESMs, it is imperative to incorporate additional environmental controls, such as the cation exchange capacity, and refine the functional relationships to align more closely with observations.

54 ENVIRONMENTAL SCIENCES

Assessing Heterogeneity of Surface Water Temperature Following Stream Restoration and a High-Intensity Fire from Thermal Imagery

Thermal heterogeneity of rivers is essential to support freshwater biodiversity. Salmon behaviorally thermoregulate by moving from patches of warm water to cold water. When implementing river restoration projects, it is essential to monitor changes in temperature and thermal heterogeneity through time to assess the impacts to a river’s thermal regime. Lightweight sensors that record both thermal infrared (TIR) and multispectral data carried via unoccupied aircraft systems (UASs) present an opportunity to monitor temperature variations at high spatial (<0.5 m) and temporal resolution, facilitating the detection of the small patches of varying temperatures salmon require. Here, we present methods to classify and filter visible wetted area, including a novel procedure to measure canopy cover, and extract and correct radiant surface water temperature to evaluate changes in the variability of stream temperature pre- and post-restoration followed by a high-intensity fire in a section of the river corridor of the South Fork McKenzie River, Oregon. We used a simple linear model to correct the TIR data by imaging a water bath where the temperature increased from 9.5 to 33.4 °C. The resulting model reduced the mean absolute error from 1.62 to 0.35 °C. We applied this correction to TIR-measured temperatures of wetted cells classified using NDWI imagery acquired in the field. We found warmer conditions (+2.6 °C) after restoration (p < 0.001) and median absolute deviation for pre-restoration (0.30) to be less than both that of post-restoration (0.85) and post-fire (0.79) orthomosaics. In addition, there was statistically significant evidence to support the hypothesis of shifts in temperature distributions pre- and post-restoration (KS test 2009 vs. 2019, p < 0.001, D = 0.99; KS test 2019 vs. 2021, p < 0.001, D = 0.10). Moreover, we used a Generalized Additive Model (GAM) that included spatial and environmental predictors (i.e., canopy cover calculated from multispectral NDVI and photogrammetrically derived digital elevation model) to model TIR temperature from a transect along the main river channel. This model explained 89% of the deviance, and the predictor variables showed statistical significance. Collectively, our study underscored the potential of a multispectral/TIR sensor to assess thermal heterogeneity in large and complex river systems.

Barker, Matthew I. (ORCID:0000000252864930)

Weather and Climate Change Impacts on Human Mortality in Bangladesh

Weather and climate profoundly affect human health. Several studies have demonstrated a U-, V-, or J-shaped temperature-mortality relationship with increasing death rates at the lower and particularly upper end of the temperature distribution. The objectives of this study were (1) to analyze the relationship between temperature and mortality in Bangladesh for different subpopulations and (2) to project future heat-related mortality under climate change scenarios. We used (non-)parametric Generalized Additive Models adjusted for trend, season and day of the month to analyze the effect of temperature on daily mortality. We found a decrease in mortality with increasing temperature over a wide range of values; between the 90th and 95th percentile an abrupt increase in mortality was observed which was particularly pronounced for the elderly above the age of 65 years, for males, as well as in urban areas and in areas with a high socio-economic status. Daily historical and future temperature values were obtained from the NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset. This dataset is comprised of downscaled climate scenarios for the globe that are derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5). The derived dose-response functions were used to estimate the number of heat-related deaths occurring during the 1990s (1980-2005), the 2020s (2011-2040) and the 2050s (2041-2070). We estimated that excess deaths due to heat will triple from the 1990s to the 2050s, with an annual number of 0.5 million excess deaths in 1990 to and expected number of 1.5 millions in 2050.

mortality

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator

Barge Site - Avian Radar System / Derived Data

This is a combined data set of 67,410 bird/bat tracks from an avian radar system deployed on a research barge (MERLIN True3D, DeTect, Panama City, Florida, USA) and concurrent wind measurements from two scanning lidars (WindCube v2.1, Vaisala, Vantaa, Finland, and Halo XR+, Halo Photonics, Lannion, France). The research barge (16.5 m x 61 m) was deployed as part of the Wind Forecast Improvement Project (WFIP-3) off the northeast coast of the United States south of Massachusetts (40.9 deg N, 70.79 deg W). This data set comprises 5 weeks of data between August 27th 2024 and September 27th 2024. Radar data were provided by DeTect and Lidar data were accessed through the Wind Data Hub (wfip3/barg.WINDPROF.z01.a0) The data have been filtered and sorted into two size groups ("big" and "small") based on a clustering approach. See Snortland, A., Clerc, J., Hein, C., & Cotter, E. (2025). Wind as Driver of Bird and Bat Abundance, Flight Direction, Altitude, and Speed on the North Atlantic Shelf. arXiv preprint arXiv:2511.14983 for complete details. Data are provided in 2 files: "Birds" and "Birds_hourly" Birds: This file contains information about each of the 67,410 flying animal tracks detected by the radar during the data collection period, including parameters measured by the radar and wind information interpolated from the lidar wind measurements. We note that the raw radar dataset contained 301,618 tracks; tracks in this processed dataset were filtered based on the requirements described in Snortland et al. (2025). Birds_hourly: This file contains timeseries of the number of tracks detected per hour over the course of the data collection period, including wind conditions and sun position for each hour. These data were used for generalized additive modeling in Snortland et al. (2025).

17 WIND ENERGY

A Method for Incorporating Changing Structural Characteristics Due to Propellant Mass Usage in a Launch Vehicle Ascent Simulation

Launch vehicles consume large quantities of propellant quickly, causing the mass properties and structural dynamics of the vehicle to change dramatically. Currently, structural load assessments account for this change with a large collection of structural models representing various propellant fill levels. This creates a large database of models complicating the delivery of reduced models and requiring extensive work for model changes. Presented here is a method to account for these mass changes in a more efficient manner. The method allows for the subtraction of propellant mass as the propellant is used in the simulation. This subtraction is done in the modal domain of the vehicle generalized model. Additional computation required is primarily for constructing the used propellant mass matrix from an initial propellant model and further matrix multiplications and subtractions. An additional eigenvalue solution is required to uncouple the new equations of motion; however, this is a much simplier calculation starting from a system that is already substantially uncoupled. The method was successfully tested in a simulation of Saturn V loads. Results from the method are compared to results from separate structural models for several propellant levels, showing excellent agreement. Further development to encompass more complicated propellant models, including slosh dynamics, is possible.

McGhee, D. S.

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns

Projected Climate Impacts to South African Maize and Wheat Production in 2055: A Comparison of Empirical and Mechanistic Modeling Approaches

Crop model-specific biases are a key uncertainty affecting our understanding of climate change impacts to agriculture. There is increasing research focus on intermodel variation, but comparisons between mechanistic (MMs) and empirical models (EMs) are rare despite both being used widely in this field. We combined MMs and EMs to project future (2055) changes in the potential distribution (suitability) and productivity of maize and spring wheat in South Africa under 18 downscaled climate scenarios (9 models run under 2 emissions scenarios). EMs projected larger yield losses or smaller gains than MMs. The EMs' median-projected maize and wheat yield changes were 3.6% and 6.2%, respectively, compared to 6.5% and 15.2% for the MM. The EM projected a 10% reduction in the potential maize growing area, where the MM projected a 9% gain. Both models showed increases in the potential spring wheat production region (EM = 48%, MM = 20%), but these results were more equivocal because both models (particularly the EM) substantially overestimated the extent of current suitability. The substantial water-use efficiency gains simulated by the MMs under elevated CO2 accounted for much of the EMMM difference, but EMs may have more accurately represented crop temperature sensitivities. Our results align with earlier studies showing that EMs may show larger climate change losses than MMs. Crop forecasting efforts should expand to include EMMM comparisons to provide a fuller picture of crop-climate response uncertainties.

DSSAT

A New Hybrid Spatio-temporal Model for Estimating Daily Multi-year PM2.5 Concentrations Across Northeastern USA Using High Resolution Aerosol Optical Depth Data

The use of satellite-based aerosol optical depth (AOD) to estimate fine particulate matter PM(sub 2.5) for epidemiology studies has increased substantially over the past few years. These recent studies often report moderate predictive power, which can generate downward bias in effect estimates. In addition, AOD measurements have only moderate spatial resolution, and have substantial missing data. We make use of recent advances in MODIS satellite data processing algorithms (Multi-Angle Implementation of Atmospheric Correction (MAIAC), which allow us to use 1 km (versus currently available 10 km) resolution AOD data.We developed and cross validated models to predict daily PM(sub 2.5) at a 1X 1 km resolution across the northeastern USA (New England, New York and New Jersey) for the years 2003-2011, allowing us to better differentiate daily and long term exposure between urban, suburban, and rural areas. Additionally, we developed an approach that allows us to generate daily high-resolution 200 m localized predictions representing deviations from the area 1 X 1 km grid predictions. We used mixed models regressing PM(sub 2.5) measurements against day-specific random intercepts, and fixed and random AOD and temperature slopes. We then use generalized additive mixed models with spatial smoothing to generate grid cell predictions when AOD was missing. Finally, to get 200 m localized predictions, we regressed the residuals from the final model for each monitor against the local spatial and temporal variables at each monitoring site. Our model performance was excellent (mean out-of-sample R(sup 2) = 0.88). The spatial and temporal components of the out-of-sample results also presented very good fits to the withheld data (R(sup 2) = 0.87, R(sup)2 = 0.87). In addition, our results revealed very little bias in the predicted concentrations (Slope of predictions versus withheld observations = 0.99). Our daily model results show high predictive accuracy at high spatial resolutions and will be useful in reconstructing exposure histories for epidemiological studies across this region.

Air pollution

Construction and Use of Resting 12-Lead High Fidelity ECG "SuperScores" in Screening for Heart Disease

We investigated the accuracy of several conventional and advanced resting ECG parameters for identifying obstructive coronary artery disease (CAD) and cardiomyopathy (CM). Advanced high-fidelity 12-lead ECG tests (approx. 5-min supine) were first performed on a "training set" of 99 individuals: 33 with ischemic or dilated CM and low ejection fraction (EF less than 40%); 33 with catheterization-proven obstructive CAD but normal EF; and 33 age-/gender-matched healthy controls. Multiple conventional and advanced ECG parameters were studied for their individual and combined retrospective accuracies in detecting underlying disease, the advanced parameters falling within the following categories: 1) Signal averaged ECG, including 12-lead high frequency QRS (150-250 Hz) plus multiple filtered and unfiltered parameters from the derived Frank leads; 2) 12-lead P, QRS and T-wave morphology via singular value decomposition (SVD) plus signal averaging; 3) Multichannel (12-lead, derived Frank lead, SVD lead) beat-to-beat QT interval variability; 4) Spatial ventricular gradient (and gradient component) variability; and 5) Heart rate variability. Several multiparameter ECG SuperScores were derivable, using stepwise and then generalized additive logistic modeling, that each had 100% retrospective accuracy in detecting underlying CM or CAD. The performance of these same SuperScores was then prospectively evaluated using a test set of another 120 individuals (40 new individuals in each of the CM, CAD and control groups, respectively). All 12-lead ECG SuperScores retrospectively generated for CM continued to perform well in prospectively identifying CM (i.e., areas under the ROC curve greater than 0.95), with one such score (containing just 4 components) maintaining 100% prospective accuracy. SuperScores retrospectively generated for CAD performed somewhat less accurately, with prospective areas under the ROC curve typically in the 0.90-0.95 range. We conclude that resting 12-lead high-fidelity ECG employing and combining the results of several advanced ECG software techniques shows great promise as a rapid and inexpensive tool for screening of heart disease.

Schlegel, T. T.